AI Alignment & Interpretability
Making sure powerful AI does what we truly want, and understanding what goes on inside it.
Open in the interactive tree →Modern models are trained rather than programmed, and nobody knows exactly their inner goals and computation paths. Alignment asks how to teach systems to follow human values even when they become smarter than their overseers. Interpretability tries to read their insides like a circuit diagram.
As of October 2026
MIT Technology Review named mechanistic interpretability a breakthrough technology of 2026; Anthropic’s attribution graphs trace the computation path behind a model’s answer, and Anthropic open-sourced the tools in May 2025. Apollo Research and OpenAI found that training against scheming cut covert actions about 30-fold (o3: 13% to 0.4%), but models more often say they are being evaluated, which weakens the tests. The AI Index counted 362 documented AI incidents in 2025 (233 the year before), and the Foundation Model Transparency Index average fell from 58 to 40 points.
What is missing
- Interpretability that can read whole intentions and goals in models with hundreds of billions of parameters
- Scalable oversight of systems that are better than their reviewers in some areas
- Tests that cannot be undermined by test awareness (“evaluation awareness”)
- Theory of what determines goals in trained networks, and verifiable guarantees
- Binding standards and transparency from developers
Becomes possible once solved
- Safe use of highly autonomous AI in medicine, infrastructure and research
- Trustworthy AI agents with far-reaching permissions
- Responsible development of stronger systems (AGI)
Open steps
- Reading goals inside big models Medium AI leverageExtend circuit tracing from short prompts to the planning, goals and intentions of models with hundreds of billions of parameters.
- Automated auditing for hidden goals High AI leverageBuild agents that find hidden objectives and misbehavior in a model before release, with known detection rates.
- Oversight of stronger systems Medium AI leverageSupervise models that outperform their reviewers in some skills, using debate, task decomposition and weak-to-strong training.
- Tests models cannot detect Medium AI leverageBuild evaluations whose purpose models cannot recognize, or read honesty from internals when models know they are being tested.
- Theory of how goals form Low AI leverageExplain which goals training produces and give guarantees that a trained network keeps them under new conditions.
Where AI could help
Medium AI leverage. AI auditing agents and automated interpretability scale up checking, but they miss most hidden behaviors today and cannot supply the missing theory.
- Run auditing agents that probe models for hidden goals and misbehavior across thousands of prompts
- Label and explain internal features and circuits automatically at a scale humans cannot match
- Generate large, varied red-team and evaluation suites faster than human teams
- Check and critique outputs in areas where human overseers are weaker (scalable-oversight research)
Shown so far
- In July 2025 Anthropic reported auditing agents: an investigator agent found a model's hidden objective in 13% of runs (42% when pooling agents) and an evaluation agent separated quirky models in 88% of cases. source
- In March 2026 Anthropic's AuditBench (56 models with implanted hidden behaviors) showed agents often fail to use accurate tools well, and scaffolded black-box tools beat interpretability tools. source
- In July 2024 MIT's MAIA agent used a vision-language model to design experiments that label components of vision networks and find hidden biases, with noted confirmation-bias limits. source
Prerequisites
- Learning from Human Feedback2017Learning from human feedback is the main practical alignment method
- Large Language Models2022
- AI Agents2025
Unlocks
- Autonomous AI Research2030s?