Human Tech Tree
Current research2022-2026 · Present (2015 – Oct 2026)

Civilization / Humanities & Culture

AI Reads Archives & Inscriptions

Handwriting recognition makes archive pages searchable, and DeepMind's Ithaca (2022) and Aeneas (2025) restore and date damaged Greek and Latin inscriptions.

Open in the interactive tree →

Handwriting models learn an old hand from a few thousand transcribed pages: the Dutch National Archives reached 7% character error with 6,000 training pages for about 3 million pages of East India Company and notarial records (2020-21). Ithaca restored damaged Greek text with 62% accuracy (25% for historians alone, 72% together), and Aeneas, trained on over 176,000 Latin inscriptions, dates texts to within 13 years of historians' ranges.

As of October 2026

Aeneas appeared in Nature on 23 July 2025 with open code and data; in a study with 23 historians the best results came from historians using its suggestions. It restores gaps of unknown length only 58% of the time (top 20), so its output is a lead, not a reading. Harvard released 983,004 public-domain volumes (242 billion tokens) from its Google Books scans in June 2025. Transkribus, run since July 2019 by the non-profit READ-COOP, was used by the Dutch National Archives, which aim to digitise a tenth of their holdings over 15 years.

Open steps

  • Rare scripts and hands Medium AI leverageExtend handwriting recognition to rare scripts and hands with few transcribed pages; the Dutch project needed 6,000 training pages for 7% character error.
  • Treat restorations as hypotheses Medium AI leverageAeneas restores gaps of known length with 73% top-20 accuracy but 58% when the length is unknown; tools must show uncertainty and parallels.
  • Most holdings are not scanned Low AI leverageScanning and cataloguing set the pace: the Dutch National Archives aim to digitise a tenth of their holdings over 15 years.
  • Searchable names, places, dates Medium AI leverageTurn transcripts into searchable records; the Dutch archives' platform adds named-entity recognition, and Harvard released 983,004 public-domain volumes as text in 2025.

Where AI could help

Medium AI leverage. Reading and restoring are solved well enough to help scholars, but scanning, access and expert checks decide how much becomes usable.

  • Transcribe old handwriting after training on a few thousand pages
  • Propose restorations and dates for damaged inscriptions with parallels
  • Tag people, places and dates so collections can be searched
  • Find parallel passages across large digitised corpora

Shown so far

  • The Dutch National Archives got 7% character error from 6,000 training pages with Transkribus, against a 20% target, on a first batch of about 3 million pages. source
  • Google DeepMind's Aeneas restores gaps of up to ten characters with 73% top-20 accuracy, attributes inscriptions to 62 provinces with 72% accuracy and dates within 13 years (company announcement, 23 July 2025). source
  • DeepMind's Ithaca restored damaged Greek inscriptions with 62% accuracy; historians scored 25% alone and 72% with Ithaca (company announcement, 9 March 2022). source

Prerequisites

Sources

More in Humanities & Culture · Present

All 41 points in Humanities & Culture →

Open in the interactive tree →