Vetting the AI Paper Flood
AI-written papers and reviews flood journals and conferences; human review cannot keep up.
Open in the interactive tree →Peer review rests on unpaid expert time and trust. Language models raise the volume of submissions, and reviews themselves are increasingly machine-written. Detection of AI text is error-prone, and publishers' policies differ.
As of October 2026
A Pangram Labs analysis with its own AI-text detector (18 November 2025) classed 15,899 reviews (21%) at the ICLR 2026 AI conference as fully AI-generated; 9% of roughly 19,000 submissions contained over 50% AI text. Such detector-based figures are themselves estimates, and no reliable large-scale check of AI use in submissions and reviews exists.
What is missing
- Verifiable labeling of AI contributions and provenance
- Paid or credited reviewing work instead of free volunteering
- Open peer-review and data-checking models
- Automated tools that check statistics, images and citations
- Common rules for AI use in submission and review
Becomes possible once solved
- Trustworthy AI-assisted research
- Scalable quality assurance
- Less fraud in science
Open steps
- Detecting AI-written text Low AI leverageReliably detect AI-written reviews and papers at scale without wrongly accusing human authors.
- Automated checks before review High AI leverageAutomated checks of statistics, images, references and citations before a human reviewer opens a paper.
- AI help for human reviewers High AI leverageTools that help reviewers write clearer, more accurate reviews instead of replacing them.
- Provenance of AI contributions Low AI leverageVerifiable labels showing which text, code or analysis an AI produced, so that editors can see it.
- Testing new review models Low AI leverageTest which review models work against the flood: open review, credited or paid review, or lotteries for borderline papers.
Where AI could help
Medium AI leverage. AI tools for images, paper mills and reviewer feedback help, but detection is an arms race and reviewer labor and incentives are the real limit.
- Automated checks of statistics, images, references and citations before review
- Paper-mill and duplicate-submission screening shared across publishers
- LLM feedback that helps reviewers write clearer, more specific reviews
- Provenance and disclosure checks for AI-written text, used as a flag for human editors
Shown so far
- In April 2025 a randomized study of over 20,000 ICLR 2025 reviews found 27% of reviewers who got LLM feedback updated their reviews, which became longer and more informative. source
- In a 2023 pilot the American Society for Microbiology screened 2,627 accepted manuscripts with the AI image tool Imagetwin, which flagged image-duplication concerns in 3.9% (mBio study). source
- In March 2025 LLM-based YesNoError reportedly checked over 37,000 papers for errors in two months, while the Black Spatula Project saw about 10% false alarms and needs expert verification. source
Prerequisites
- Peer Review & Research Bodies1945-1975
- Open Access & Open Science2018-2026
- Large Language Models2022
Unlocks
- Self-Running Research Labs2030s?
- Verified Science2030s?