Large Language Models
Chatbots like ChatGPT (30 Nov 2022) write, translate and program in natural language: now an everyday tool.
Open in the interactive tree →Language models predict the next word, learned from huge amounts of text; with growing size (GPT-3, 2020: 175 billion parameters) abilities such as reasoning and programming appeared. Fine-tuning with human feedback made them usable as chatbots. ChatGPT reached roughly 100 million users within two months.
As of October 2026
Frontier models in autumn 2026 (for example Anthropic’s Claude Opus 5.5 from 22 September and OpenAI’s GPT-6 models in ChatGPT from October) solve expert-level exam and coding tasks and work as agents for hours. Per the Stanford AI Index 2026, US and Chinese models have traded the top spot several times since early 2025; in March 2026 the leading model was ahead by just 2.7%. Generative AI reached 53% population adoption within three years, and four in five US high-school and college students use AI for schoolwork, yet the models still hallucinate and have a “jagged” capability profile (the best model reads analog clocks correctly only about half the time).
Open steps
- Fewer confident wrong answers High AI leverageModels still state falsehoods fluently; training and benchmarks reward guessing over saying 'I do not know', so calibrated uncertainty is missing.
- Even skills across tasks High AI leverageModels solve expert exams yet fail simple tasks such as reading clocks; finding and closing such gaps needs broad tests that are not memorised.
- Learning beyond the text supply Medium AI leveragePublic human text may be used up between 2026 and 2032 at current scaling; data-efficient training, verified synthetic data and multimodal sources are needed.
- Models that keep learning safely Medium AI leverageDeployed models do not learn from use; adding knowledge after training without forgetting or being poisoned by bad data is unsolved.
Where AI could help
Medium AI leverage. Agents already do part of AI research engineering, but compute, data quality and new ideas remain the limits.
- Run, debug and analyze many training experiments in parallel
- Generate synthetic data, verifiers and training environments
- Write and tune training and inference kernels
- Red-team and evaluate models at scale
Shown so far
- In September 2026 OpenAI said (company claim, preliminary) its researchers use about 3.1 agent-workdays per human workday, and over half of successful 4-8 hour tasks needed human intervention. source
- In November 2024 METR's RE-Bench found AI agents scored 4x higher than human experts at a 2-hour budget on ML engineering tasks, but humans did better at 8 hours and scored 2x at 32 hours. source
- In May 2025 Google said (own claim) that a scheduling heuristic found by AlphaEvolve recovers about 0.7% of its worldwide compute, and a 23% faster kernel cut Gemini training time by about 1%. source
Prerequisites
- The Internet1969
- Unicode1991Language-model tokenisers work on Unicode text
- Graphics Processor (GPU)1999
- Wikipedia2001Wikipedia is a core part of the text used to train language models
- Learning from Human Feedback2017Human-feedback training turned language models into usable assistants
- Transformer Architecture2017
- Pretrained Language Models2018Pretraining on raw text is the recipe of every large language model
- Neural Scaling Laws2020Scaling laws guided the size of models and training data
Unlocks
- AI in medical diagnosis2018–2026
- Open-Weight Language Models2023-2026
- Animal-sound foundation models2024-2026
- EU AI Act2024-2028
- Reasoning Models2024
- AI & the Labor Market2025-2026
- AI Wins Math Olympiad Gold2025
- AI as Research Partner2025-2026
- Humanoid Robots2026
- AI Alignment & Interpretabilityopen
- Artificial General Intelligenceopen
- Digital Identity & Trustopen
- Disinformation & Trustopen
- Reliable, Honest AIopen
- Vetting the AI Paper Floodopen
- Personal Tutor for Everyone2030s?