Human Tech Tree
Unsolvedopen · Research Frontier · Today (unsolved as of Oct 2026)

Civilization / Foundations of Civilization

Replication Crisis

Many published findings fail when others try to repeat them; the problem has been known since 2005 and is not solved.

Open in the interactive tree →

John Ioannidis's 2005 essay 'Why most published research findings are false' and the Open Science Collaboration's 2015 psychology replications (only about 36-47% of results confirmed, depending on criterion) started the debate. In the cancer biology replication project (50 experiments from 23 papers), 46% of replications succeeded on more criteria than they failed, and effect sizes were on average 85% smaller. Causes include publish-or-perish incentives, small samples, flexible analyses and missing raw data.

As of October 2026

Paper mills and compromised peer review drive record retraction numbers: Nature counted more than 10,000 retractions in 2023, and Hindawi (Wiley) alone retracted over 8,000 articles that year. In the cancer replication project only 2% of the original experiments had openly available data and no protocol was completely described, which hampered replication.

What is missing

  • Incentives that reward replications and null results, not just novelty
  • Mandatory preregistration or registered reports plus open raw data and code
  • Standing funding lines and journals for replication studies
  • Better training in statistics, sample size and analytic flexibility
  • Large-scale automatic detection of fraud and paper-mill output

Becomes possible once solved

  • Reliable evidence base for medicine and policy decisions
  • Automated research without compounding errors
  • Less wasted research funding

Open steps

  • Large-scale error screening High AI leverageAutomatically find statistical, numerical and citation errors across the literature, with few false alarms.
  • Re-running published analyses Medium AI leverageRe-run published code on shared data to test computational reproducibility, and find what stops it.
  • Predicting which claims will fail Medium AI leveragePredict which published claims are fragile so that replication effort goes where it matters.
  • Paper-mill and image fraud Medium AI leverageDetect paper-mill output and reused or manipulated figures across publishers at scale.
  • Testing fixes for replicability Low AI leverageMeasure which policies, such as registered reports, data-sharing rules or replication funding, actually raise replication rates.

Where AI could help

Medium AI leverage. AI can screen papers for errors and fraud at scale and re-run analyses; journals and funders still have to reward replication.

  • Scanning papers for statistical, numerical and citation errors at scale
  • Re-running published code and analyses on shared data to test computational reproducibility
  • Predicting which claims are fragile so replication effort goes where it matters
  • Detecting paper-mill text and reused or manipulated figures

Shown so far

  • In the DARPA-funded SCORE program (2019-2023) human experts predicted replication with 76 to 78% accuracy; early AI methods did poorly, and a later open competition showed AI getting close to human accuracy, according to project lead Brian Nosek. source
  • In March 2025 Nature-reported LLM tools YesNoError and the Black Spatula Project checked tens of thousands of papers for errors, but the latter had about 10% false alarms and every flag needs expert checking. source

Prerequisites

Unlocks

Sources

More in Foundations of Civilization · Research Frontier · Today

All 47 points in Foundations of Civilization →

Open in the interactive tree →