Researched
Transformer Architecture
Google’s “Attention Is All You Need” (2017) replaces sequential reading with attention: trainable in parallel and scalable almost without limit.
Open in the interactive tree →The transformer architecture by Vaswani and colleagues lets every word attend to all others and parallelizes very well on GPUs. BERT (2018) and GPT (2018) showed that pretraining on huge amounts of text is universally useful. Since then the same architecture has carried language, image, audio and robotics models.
Prerequisites
- LSTM memory networks1997Attention was first added to LSTM translators, which transformers replaced
- Deep Learning2012
- Word Embeddings (word2vec)2013Word embeddings are the input layer of transformers
- Attention Mechanism2014The transformer is built on attention, first used for translation in 2014
Unlocks
- Learning from Human Feedback2017
- Pretrained Language Models2018
- Neural Scaling Laws2020
- AI for Low-Resource Languages2022-2026
- Generative Image & Video AI2022
- Large Language Models2022
- AI protein design2023Protein language and generative design models are built on transformer networks
- Animal-sound foundation models2024-2026
- Project CETI: sperm whale codas2024-2026