The transformer, introduced in the 2017 paper Attention Is All You Need, is the architecture behind today's language models and increasingly behind vision and audio models too.
Self-Attention
For each token in a sequence, self-attention decides how much to draw on every other token when building that token's representation. In "The animal didn't cross the road because it was tired", attention helps the model link "it" to "animal".
Mechanically, each token is projected into a query, a key and a value. The similarity between one token's query and every token's key gives attention weights, which are used to combine the values.
Multi-Head Attention
Several attention "heads" run in parallel, each free to focus on different relationships — syntax, reference, position — and their outputs are combined.
Positional Information
Attention itself ignores order, so transformers add positional encodings that tell the model where each token sits in the sequence.
Building Blocks
A transformer layer combines attention with a feed-forward network, residual connections and normalisation. Models stack dozens of such layers.
Encoders and Decoders
- Encoder-only models (such as BERT) build rich representations for understanding tasks.
- Decoder-only models (most chat LLMs) generate text one token at a time.
- Encoder–decoder models (such as T5 and Whisper) map an input sequence to an output sequence.
Why It Won
Transformers process sequences in parallel, scale well with data and compute, and capture long-range relationships — which made large pretrained models possible.