Artificial General Intelligence
An AI system that learns and performs practically any mental task as well as a human: whether and when is disputed.
Open in the interactive tree →There is no generally accepted definition of AGI: some mean parity with humans on most economically relevant tasks, others independent learning in any new situation. Today’s systems are superhuman in parts but fail at tasks children can do. The question touches the economy, security and politics.
As of October 2026
OpenAI CEO Sam Altman remarked in December 2025 that “we built AGIs” (an offhand comment, not an independently tested claim) and suggested shifting the debate to superintelligence. The largest survey, AI Impacts (October 2023, 2,778 AI researchers), gave a 10% chance of machines beating humans at every task by 2027 and 50% by 2047 (about 2060 in the 2022 round); older 2012-13 polls gave a median near 2040-2050. Measurements show a jagged profile: METR estimates agents manage tasks over 16 hours long at 50% success but only about 3 hours at 80% success, and robots solve only 12% of household tasks per the AI Index. The interactive test ARC-AGI-3 (ARC Prize 2026) asks whether agents can beat novel game worlds as efficiently as humans; the ARC Prize holds that AGI does not exist while that gap remains.
What is missing
- Continual learning during operation instead of frozen models
- Reliability over long tasks (80% and 99% horizons instead of 50%)
- Understanding the physical world (world models, acting in space)
- Data efficiency: learning from few examples like a human
- A measurable, widely accepted definition and test method
Becomes possible once solved
- Automated research and development
- Broad automation of cognitive work
- AI as a personal expert in medicine, law and education
- Economic growth driven by computation instead of labor
Open steps
- Learning during operation Medium AI leverageLet a deployed model absorb new knowledge and skills over time without forgetting old ones or retraining from scratch.
- 80% and 99% task horizons High AI leverageRaise reliability on multi-hour tasks from 50% success to 80% or 99%, not only lengthen the tasks.
- Physical-world models High AI leverageLearn models of space, objects and physics from video and action so agents can plan and act in the real world.
- Learning from few examples Medium AI leverageReach human-like sample efficiency: learn a new task from a handful of examples instead of millions.
- Hard-to-game general-ability tests Medium AI leverageDesign interactive tests that measure efficient learning of new tasks and resist memorization and benchmark overfitting.
Where AI could help
Medium AI leverage. AI now does part of AI research engineering, but evidence is limited to well-specified tasks; missing ideas and an agreed definition remain.
- Run, debug and analyze many ML experiments in parallel, speeding up routine research engineering
- Write and tune kernels, training code and data pipelines
- Generate and test hypotheses on architectures, long-horizon memory and continual learning
- Build harder evaluations for long-task reliability and generalization
Shown so far
- In November 2024 METR's RE-Bench found AI agents scored 4x higher than human experts at a 2-hour budget on ML engineering tasks, but humans did better at 8 hours and scored 2x at 32 hours. source
- In September 2026 OpenAI said (company claim, preliminary) its researchers use about 3.1 agent-workdays per human workday, and over half of successful 4-8 hour tasks needed human intervention. source
Prerequisites
- Large Language Models2022
- Reasoning Models2024
- AI Agents2025
Unlocks
- AI Research Mathematician2030s?
- Autonomous AI Research2030s?