Researched
Attention Mechanism
Bahdanau, Cho and Bengio (2014) let a translation network look back at the most relevant input words while writing each output word.
Open in the interactive tree →Attention solved the bottleneck of squeezing a whole sentence into one vector. Its generalisation, self-attention, is the core of the transformer.
Prerequisites
Unlocks
- Transformer Architecture2017The transformer is built on attention, first used for translation in 2014