011 min
The problem self-attention solves
By 2014, the standard way to give a decoder access to a source sentence was the mechanism Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio described that year: instead of compressing an entire sentence into one fixed-length vector, the decoder learned to weight the encoder's hidden states by relevance before producing each word. That fixed a real problem β translation quality had been dropping sharply on long sentences because a single vector could not hold everything a sentence needed.
But the encoder producing those hidden states was still a recurrent network. It read the source sentence one token at a time, and the hidden state at position 50 could only exist once the hidden state at position 49 had been computed. Attention let the decoder reach freely into that stack of states; nothing let the encoder build the stack any faster, and nothing let two distant source positions compare with each other directly.
The specific gap self-attention closes is that second one: relating positions within a single sequence β encoder to encoder, or decoder to decoder β without a chain of sequential steps in between. Once positions in the same sequence can score their relevance to each other directly, in one step, the recurrence that used to move information between distant words becomes unnecessary β which is exactly what let the Transformer drop it in 2017.
023 min
How self-attention works
Self-attention takes a sequence of vectors β one per token β and returns a new sequence of the same length, where every output vector is a weighted mix of all the input vectors. The steps are the same whatever the tokens represent; the example below follows one sentence the whole way through.
Three vectors per token
Take the sentence The trophy would not fit in the suitcase because it was too big. Deciding what it refers to β the trophy or the suitcase β is not obvious from the word it alone; the answer has to come from the rest of the sentence.
Self-attention gives every token three roles, each produced by multiplying its embedding by a separate learned matrix. The query is what a token is asking about. The key is what a token offers when another token asks. The value is what actually gets passed along once a match is found. For it, the query stands for something close to "which noun does this word point back to?" The keys for trophy and suitcase stand for what kind of question each word can answer.
Scoring, then scaling
To score how much it should attend to trophy, the model takes the dot product of it's query vector and trophy's key vector β one number, large when the two vectors point in a similar direction and small when they do not. It computes this same number against every token's key, including it's own. Vaswani et al. (2017) divide every one of these scores by the square root of the key vectors' dimension before going further, because unscaled dot products grow large as the vector size grows, which pushes the next step into a range where it barely responds to further change.
From scores to weights to a new vector
A softmax function turns the full set of scaled scores for it into weights that are positive and add up to one. The new vector for it is the weighted sum of every token's value vector, using those weights. When trophy's weight dominates, the new vector for it ends up carrying mostly whatever trophy's value vector encodes β which is the sense in which self-attention has resolved the reference, with no hand-written rule involved.
Working through the arithmetic
Real weights are learned and have hundreds of numbers each. The table below runs the same steps on small, hand-picked two-number vectors, to show how a score turns into a weight β not to claim these are numbers a trained model actually produced.
| Token | Toy key vector | Dot product with it's query | Scaled (Γ·β2) | Softmax weight |
|---|---|---|---|---|
| trophy | (1, β1) | 2.00 | 1.41 | 0.57 |
| big | (0.5, β0.5) | 1.00 | 0.71 | 0.28 |
| suitcase | (0.2, 0.1) | 0.10 | 0.07 | 0.15 |
With it's query set to (1, β1), trophy's key lines up with it exactly, big partly agrees β a big object is the reason nothing fits β and suitcase barely registers. Summing the value vectors by those weights pulls the result strongly toward whatever trophy's value vector holds.
Several of these at once
One query-key-value comparison can track only one kind of relationship. The 2017 paper runs eight of these in parallel, each working on a smaller slice of the vector, because β in the authors' own words β this "allows the model to jointly attend to information from different representation subspaces at different positions" (Vaswani et al., 2017). One head can end up tracking pronoun reference, another the subject of a verb, without either being assigned that job. This entire computation β query, key, value, score, weight, sum β is what a large language model repeats, layer after layer, for every token in its input at once.
031 min
Where self-attention came from
Self-attention itself was not new in 2017. Jianpeng Cheng, Li Dong, and Mirella Lapata used it a year earlier inside a recurrent reading model, and Zhouhan Lin and colleagues used it in 2017 to build fixed-length sentence embeddings β in both cases as one component added to a network that still processed tokens in sequence.
What Ashish Vaswani and his co-authors did differently was remove everything else. Their own paper states it plainly: to the best of their knowledge, the Transformer was "the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution." The operation already existed; making it the whole model, with nothing sequential left, was the new part.
042 min
What self-attention delivered, measured
Vaswani and colleagues did not just claim self-attention worked; they measured it against the two things that mattered for translation at the time: quality and training cost.
On the WMT 2014 English-to-German task, the Transformer β built almost entirely from self-attention layers β reached 28.4 BLEU, more than 2 BLEU above the best previous result, ensembles included. On WMT 2014 English-to-French, it reached a single-model state of the art of 41.8 BLEU, after training for 3.5 days on eight GPUs, which the paper describes as a small fraction of the training cost of the best models published before it.
Both numbers trace back to the same mechanical property covered above: because every position's query can compare against every other position's key in one step, none of the model's layers had to wait for an earlier position to finish computing before starting on a later one. A recurrent encoder-decoder, including the kind Bahdanau, Cho, and Bengio described in 2014, has no such option β position 50 depends on position 49 having been computed first, in every layer, on every training step. Self-attention's constant-length path between any two positions, not a bigger model or more training data, is what the paper credits for reaching a better score in less training time.
BLEU itself is a score from 0 to 100 that compares machine output against human reference translations; a gain of more than 2 points at that level was a large jump for the field in 2017.
052 min
What self-attention means for your work
Self-attention's cost does not scale gently, and that shows up on more than one desk before it shows up as an incident.
For an engineer: every position compares against every other position, so the work inside a self-attention layer grows with the square of how many tokens go in. Doubling the input from 2,000 to 4,000 tokens roughly quadruples that layer's arithmetic, not doubles it. This is the concrete reason a retrieval step that hands the model only the passages it needs β see retrieval-augmented generation β can be cheaper than simply widening the context window and sending everything.
For a product manager: whether an answer correctly connects two facts far apart in a long document is a direct test of self-attention working as intended, not a vague "the model got confused" failure. If a support bot misses an instruction stated at the top of a long thread and repeated differently at the bottom, every position technically could have attended to every other position β the training data or the model's capacity just did not make that particular connection reliable. Reproducing the failure with a shorter input, and checking whether it goes away, tells you whether the problem is length or the connection itself.
For a designer: self-attention is what makes a chatbot's reply sensitive to something said many turns earlier in the same conversation, as long as that earlier turn is still inside the window of text being sent to the model. Once it scrolls out of that window, the mechanism has nothing left to attend to β it was never a memory system, only a comparison over whatever is currently in view.
061 min
What self-attention costs
The quadratic term is not a footnote. Vaswani et al. (2017) lay it out in their own comparison table: a self-attention layer's work per layer grows with the square of the sequence length, while a recurrent layer's work grows only with the length itself (though with the square of the vector size instead). In exchange, self-attention needs a constant number of sequential steps to connect any two positions, against a number that grows with sequence length for a recurrent layer β the trade the whole mechanism is built around.
That trade is why processing very long documents in one pass gets expensive fast, and why engineering effort since 2017 has mostly gone into making this same computation cheaper rather than replacing it β reordering how data moves through hardware memory rather than approximating the attention scores themselves. A Mixture of Experts layer addresses a different cost inside the same block β the feed-forward step, not attention β which is why the two are often combined rather than treated as alternatives. As of September 2026, the quadratic relationship between length and cost still holds inside the self-attention layer itself; what has changed since 2017 is how efficiently that arithmetic gets executed.
072 min
What self-attention does not solve
Self-attention decides which positions to combine; it does not remember anything once a conversation or document scrolls past the window sent to the model, and it adds no sense of order on its own β a Transformer only knows token 3 came before token 7 because a separate positional signal was added to the vectors before attention ever ran.
One specific overclaim is worth naming and correcting. Attention weights are often shown to users, or cited in papers, as if they explain why a model produced an answer β the weight on a word treated as how much the model "looked at" it. Jain and Wallace (2019) tested this directly and reported that learned attention weights are frequently uncorrelated with other measures of how much an input actually mattered to the output, and that very different weight patterns can produce the same prediction. That finding is contested rather than settled: Wiegreffe and Pinter (2019) argue the original tests depended on assumptions about what counts as an explanation, and that attention can pass more careful ones. Reading attention weights as a debugging signal, rather than as proof of the model's reasoning, is the safer habit either way.
The other real limit is the one covered above: past a few thousand tokens, the quadratic cost makes applying self-attention to the whole input the wrong tool, which is why retrieval and chunking exist as workarounds rather than luxuries.
081 min
Self-attention vs. nearby concepts
| Compared with | The one fact that decides |
|---|---|
| Cross-attention (encoder-decoder attention) | Which sequence the keys and values come from. In self-attention, the queries, keys and values all come from the same sequence. In cross-attention, the queries come from one sequence (the decoder) and the keys and values come from another (the encoder's output). |
| Attention, the general term (Bahdanau, Cho and Bengio, 2014) | Whether recurrence is still doing the sequential work. Bahdanau, Cho and Bengio's 2014 paper describes the problem as a fixed-length vector being a bottleneck, and proposes a decoder that can automatically search the source sentence for relevant parts β but its encoder was still a recurrent network processing tokens in order. |
| Multi-head attention | Not a different mechanism. Multi-head attention is several self-attention computations run in parallel on different slices of the same vectors, then combined into one result. |
| Transformer | Self-attention is one operation. The Transformer is the full architecture built almost entirely out of that operation, plus positional encoding and feed-forward layers, stacked several times. |
The question that actually needs an answer in practice is the first row's: whether the sequence being attended to is the same one doing the attending, or a different one. Everything else in this table follows from that.
?5 questions
Questions people ask
Is self-attention the same thing as attention?
Is self-attention the same as the Transformer?
Why does self-attention scale scores by the square root of the key dimension?
Does self-attention need positional information to work?
Why is self-attention expensive on long inputs?
Β§7 sources
Sources
Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv:1409.0473
Cheng, J., Dong, L., and Lapata, M. (2016). Long Short-Term Memory-Networks for Machine Reading. arXiv:1601.06733
Lin, Z. et al. (2017). A Structured Self-Attentive Sentence Embedding. arXiv:1703.03130
Show all 7 sourcesShow fewer sources
Jain, S., and Wallace, B. C. (2019). Attention is not Explanation. arXiv:1902.10186
Wiegreffe, S., and Pinter, Y. (2019). Attention is not not Explanation. arXiv:1908.04626
Alammar, J. The Illustrated Transformer





