Human Tech Tree
Unsolvedopen · Research Frontier · Today (unsolved as of Oct 2026)

Information / AI & Robotics

Reliable, Honest AI

Language models still invent facts and sources while sounding convincing: real reliability is missing.

Open in the interactive tree →

Models produce plausible-sounding text, not checked truth; where knowledge is missing they often guess instead of saying “I don’t know”. Benchmarks often reward guessing more than admitting ignorance. For medicine, law and engineering this is an obstacle.

As of October 2026

Error rates depend strongly on the task: when summarizing supplied text, the best-ranked models on Vectara’s leaderboard (22 September 2026) invent content in about 2-3% of summaries. On a test of factual knowledge (AA-Omniscience), the Stanford AI Index 2026 found hallucination rates across 26 top models ranging from 22% to 94%. Grounding a model in supplied text lowers the rate sharply, but does not eliminate it.

What is missing

  • Calibration: models that reliably know and report their uncertainty
  • Training and benchmarks that reward admitting ignorance
  • Checkable evidence: automatic verification of claims against sources or formal proofs
  • Understanding of the internal mechanisms that lead to hallucinations

Becomes possible once solved

  • AI decisions in medicine and law without constant checking
  • Fully automatic reports and research logs with verifiable sources
  • Reliable AI agents in government and business

Open steps

  • Calibrated uncertainty Medium AI leverageMake models say when they do not know, with confidence that matches real accuracy across topics and phrasings.
  • Rewarding 'I don't know' Medium AI leverageChange training and benchmark scoring so that admitting ignorance beats confident guessing.
  • Checking claims against sources High AI leverageVerify each claim in an answer automatically against documents or databases, with citations a reader can trust.
  • Machine-checkable proofs High AI leverageTurn claims in math, code and specifications into formal proofs that a proof checker verifies, beyond competition problems.
  • Why models hallucinate, mechanistically Medium AI leverageIdentify the internal circuits that make a model answer when it should refuse, and fix them without hurting helpfulness.

Where AI could help

Medium AI leverage. AI can flag likely confabulations and produce machine-checkable proofs, but calibration and training incentives are still open research problems.

  • Estimate answer uncertainty by sampling many answers and comparing their meanings
  • Check claims against sources automatically with retrieval agents and verifier models
  • Generate formal proofs for code and math so correctness is machine-checkable
  • Create benchmarks and training data that reward admitting ignorance over confident guessing

Shown so far

  • In June 2024 Oxford researchers showed in Nature that measuring uncertainty over meanings (semantic entropy) detects confabulated answers without task-specific training. source
  • In July 2024 DeepMind's AlphaProof and AlphaGeometry 2 solved 4 of 6 IMO problems with machine-checkable formal proofs (according to the company; method published in Nature, November 2025). source

Prerequisites

Unlocks

Sources

More in AI & Robotics · Research Frontier · Today

All 39 points in AI & Robotics →

Open in the interactive tree →