Replication Crisis
Many published findings fail when others try to repeat them; the problem has been known since 2005 and is not solved.
Open in the interactive tree →John Ioannidis's 2005 essay 'Why most published research findings are false' and the Open Science Collaboration's 2015 psychology replications (only about 36-47% of results confirmed, depending on criterion) started the debate. In the cancer biology replication project (50 experiments from 23 papers), 46% of replications succeeded on more criteria than they failed, and effect sizes were on average 85% smaller. Causes include publish-or-perish incentives, small samples, flexible analyses and missing raw data.
As of October 2026
Paper mills and compromised peer review drive record retraction numbers: Nature counted more than 10,000 retractions in 2023, and Hindawi (Wiley) alone retracted over 8,000 articles that year. In the cancer replication project only 2% of the original experiments had openly available data and no protocol was completely described, which hampered replication.
What is missing
- Incentives that reward replications and null results, not just novelty
- Mandatory preregistration or registered reports plus open raw data and code
- Standing funding lines and journals for replication studies
- Better training in statistics, sample size and analytic flexibility
- Large-scale automatic detection of fraud and paper-mill output
Becomes possible once solved
- Reliable evidence base for medicine and policy decisions
- Automated research without compounding errors
- Less wasted research funding
Open steps
- Large-scale error screening High AI leverageAutomatically find statistical, numerical and citation errors across the literature, with few false alarms.
- Re-running published analyses Medium AI leverageRe-run published code on shared data to test computational reproducibility, and find what stops it.
- Predicting which claims will fail Medium AI leveragePredict which published claims are fragile so that replication effort goes where it matters.
- Paper-mill and image fraud Medium AI leverageDetect paper-mill output and reused or manipulated figures across publishers at scale.
- Testing fixes for replicability Low AI leverageMeasure which policies, such as registered reports, data-sharing rules or replication funding, actually raise replication rates.
Where AI could help
Medium AI leverage. AI can screen papers for errors and fraud at scale and re-run analyses; journals and funders still have to reward replication.
- Scanning papers for statistical, numerical and citation errors at scale
- Re-running published code and analyses on shared data to test computational reproducibility
- Predicting which claims are fragile so replication effort goes where it matters
- Detecting paper-mill text and reused or manipulated figures
Shown so far
- In the DARPA-funded SCORE program (2019-2023) human experts predicted replication with 76 to 78% accuracy; early AI methods did poorly, and a later open competition showed AI getting close to human accuracy, according to project lead Brian Nosek. source
- In March 2025 Nature-reported LLM tools YesNoError and the Black Spatula Project checked tens of thousands of papers for errors, but the latter had about 10% false alarms and every flag needs expert checking. source
Prerequisites
- Scientific Method~1600-1620
- Statistics & Least Squares1809
- Falsifiability (Popper)1934Replication is the test culture that falsifiability demands
- Peer Review & Research Bodies1945-1975
Unlocks
- Self-Running Research Labs2030s?
- Verified Science2030s?