Avoiding a Digital Dark Age
Digital records rely on formats, software and hardware that keep changing: 38% of 2013 web pages were gone by 2023, and no medium has a centuries-long record.
Open in the interactive tree →Digital data depends on working hardware, software that understands the format and an institution that pays to keep copies. The 1986 BBC Domesday discs needed later migration and emulation, and NASA's 1976 Viking tapes sat unprocessed for over a decade and were then in an unknown format. Migration can introduce errors, emulation needs the old software, and proprietary formats are hardest for outsiders to rebuild.
As of October 2026
Pew found in May 2024 that 38% of 2013 webpages were gone by October 2023 and 54% of English Wikipedia pages had a dead reference link. The Internet Archive passed one trillion archived pages in October 2025 and adds about 498 million a day. Microsoft reported in Nature on 18 February 2026 a 2 mm borosilicate glass piece holding 4.8 TB, with a 10,000-year lifetime estimated by accelerated ageing (company claim, not observed). Software Heritage held over 22 billion unique source files from more than 340 million projects (2024 annual report; check live counter).
What is missing
- Storage whose lifetime is observed rather than extrapolated
- Custodians with stable funding for centuries
- Open formats and preserved software that can always be rebuilt
- Legal and technical means to copy and emulate protected software
- Capture of web content before it disappears
Becomes possible once solved
- Today's science, law and literature available to future historians
- Government and court records that stay verifiable
- Fewer losses from format change and dead links
Open steps
- Keep formats and software readable Medium AI leverageFiles need software that understands them; migration can introduce errors and emulation needs the original software and the right to run it.
- Media with proven lifetimes Low AI leverageGlass written by lasers holds 4.8 TB in a 2 mm piece and is estimated to last 10,000 years by accelerated ageing (Microsoft, Nature, 18 February 2026); lifetimes are not observed.
- Capture the web before it vanishes Medium AI leverage38% of 2013 webpages were gone by 2023 (Pew) and the Internet Archive has passed one trillion pages; capture depends on crawlers reaching a page first.
- Funding and custody for centuries Low AI leverageArchives need institutions that keep copies and pay for migration for centuries; Software Heritage started at Inria in 2015 and has UNESCO backing.
- Open formats over locked ones Low AI leverageOpen standards such as PDF/A and OpenDocument can be rebuilt by outsiders; locked or encrypted formats add a step that must survive as long as the data.
Where AI could help
Low AI leverage. Funding, custody and law decide this; AI helps at the margins with format identification, migration checks and web capture.
- Identify unknown file formats and check migrated files against originals
- Find at-risk web pages and prioritise their capture
- Extract text and metadata from old or damaged media
- Describe archived items automatically so they can be found
Prerequisites
- Hard Disk Drive1956
- World Wide Web1991
- Cloud Computing2006