Human Tech Tree
Current research2025 · Present (2015 – Oct 2026)

Information / AI & Robotics

AI Agents

AI systems that use tools on their own, write code and work through hours-long tasks: in 2026 the biggest lever of AI.

Open in the interactive tree →

An agent plans multiple steps, calls programs, searches and files, and checks its own results. Coding agents in particular entered developers’ daily work in 2025. What matters is how long a task can be completed reliably in one go.

As of October 2026

METR measures the task length agents complete with 50% success: about 2 hours for o3 (April 2025) versus at least 16 hours for an early Claude Mythos Preview (spring 2026), though METR warns that values above 16 hours are unreliable and the 80%-success horizon was only about 3 hours. Since 2024 the doubling time has been about 3 months (89 days; 131 days since 2023). Per the Stanford AI Index 2026, agent success on OSWorld computer tasks rose from about 12% to 66.3%, within 6 points of human performance. According to Anthropic, Claude Mythos Preview was withheld in April 2026 because the model found thousands of zero-day flaws in major operating systems and browsers (company claim); only selected partners (Project Glasswing) got access.

Open steps

  • 80-99% success on long tasks High AI leverageAgents solve hour-long tasks about half the time; deployment needs error detection, recovery and consistent results across repeated runs.
  • Agents that resist hidden instructions Medium AI leverageWeb pages, mails and files can carry text that hijacks an agent; defences must hold when the agent has real permissions and tools.
  • Measuring tasks beyond a day High AI leverageTime-horizon tests become unreliable above about 16 hours; new long, cheaply scored tasks are needed to tell real progress from benchmark gaming.
  • Oversight of work humans cannot recheck Medium AI leverageReviewing an agent's day of work takes a day; reviewers and checkers must catch errors and sabotage at far lower cost than redoing the task.

Where AI could help

Medium AI leverage. Agents help build agent tasks, tools and tests, but reliability, security and cost matter more than speed of research.

  • Build tasks, tools and test environments for agent training automatically
  • Analyze agent failures in logs and fix scaffolds
  • Write more of the agent scaffolding and evaluation code
  • Red-team prompt-injection defenses at scale

Shown so far

  • In March 2025 METR found that the task length frontier models complete at 50% success had doubled about every seven months since 2019 (preprint). source
  • In September 2026 OpenAI said (company claim, preliminary) its researchers use about 3.1 agent-workdays per human workday, and over half of successful 4-8 hour tasks needed human intervention. source

Prerequisites

Unlocks

Sources

More in AI & Robotics · Present

All 39 points in AI & Robotics →

Open in the interactive tree →