Researched
Unicode
Unicode 1.0 (1991) gives every character of every writing system a number, so computers can store all human scripts.
Open in the interactive tree →Before it, hundreds of incompatible code pages (ASCII covered only English) garbled non-English text. Unicode, with UTF-8 since 1993, now covers over 150,000 characters including emoji and is the base of the web, phones and language-model tokenisers.
Prerequisites
- Alphabet~1800 BC
- Typewriter & Keyboard1868
- Personal Computer1981
Unlocks
- Smartphone2007Phones show every language and emoji thanks to Unicode
- AI for Low-Resource Languages2022-2026
- Large Language Models2022Language-model tokenisers work on Unicode text
- Saving Endangered Languagesopen
Sources
More in Humanities & Culture · Digital Age
- Wikipedia2001