Transformer

A neural network architecture that processes a whole sequence at once, letting every token draw information directly from every other token through attention.

13 min read

Β· Also in

By Ravi SuranaUpdated 9 sources

Quick answer

~20 sec

A Transformer is a neural network architecture that processes a whole sequence at once rather than one item at a time. Its core step is attention: every token compares itself with every other token, then takes information from the ones that matter to it. Introduced in 2017 for translation, it now underlies most large language models.

012 min

The problem the Transformer solves

Before 2017, the standard architecture for translating a sentence was a recurrent neural network, usually a long short-term memory network with an attention layer bolted on. A recurrent network reads one token at a time and keeps a running state. To compute the state at position 50, it must first compute positions 1 to 49.

That design fails in two specific, measurable ways, and Vaswani and colleagues put both in a table in the 2017 paper.

The first is training speed. Because each position depends on the one before it, the work cannot be split across a sequence. The paper counts this as the number of sequential operations a layer needs: a recurrent layer needs a number that grows with the sequence length, while a self-attention layer needs a constant number.

The second is distance. For information at position 1 to reach position 50, a recurrent network passes it through 49 intermediate states. The paper calls this the maximum path length. Every step is a chance for the signal to be diluted, which is why recurrent translation degraded on long sentences.

Layer typeWork per layerSequential stepsLongest path between two positions
Self-attentiongrows with the square of sequence lengthconstantconstant
Recurrentgrows linearly with sequence lengthgrows with sequence lengthgrows with sequence length

Source: Vaswani et al. (2017), Table 1. n is sequence length; the recurrent row also scales with the square of the representation size.

The Transformer takes the second problem to zero. Any two positions are one step apart. It pays for that with the first column: attention work grows with the square of the sequence length, which is the cost the rest of this entry keeps returning to.

023 min

How a Transformer works

Follow one sentence the whole way through: The server refused the request because it was overloaded. A model that handles this sentence correctly has to decide whether it refers to the server or to the request. Nothing in the word it carries the answer.

From text to vectors

The sentence is first split into tokens, which are common chunks of text rather than whole words. The 2017 English-to-German model used a shared vocabulary of about 37,000 byte-pair tokens. Each token is then looked up in a table and becomes a list of numbers, called a vector. In the paper's base model each vector has 512 numbers.

Attention has no built-in sense of order, so position has to be added explicitly. The original model added a fixed pattern of sine and cosine values to each token's vector, one wavelength per dimension.

The attention step

Every token now produces three vectors from its own vector. The query stands for what this token is looking for. The key stands for what this token offers to others. The value is what the token passes on when another token selects it.

The model takes the query for it and compares it against the key of every token in the sentence, using a dot product, which is a single number measuring how well two vectors line up. Those numbers are divided by the square root of the key size to keep them in a workable range, then turned into weights that add up to one. The output for it is the sum of every token's value vector, weighted by those numbers.

Attention(Q, K, V) = softmax(QKα΅€ / √d_k) V

Where the interesting thing happens

For it, the query lines up more strongly with the key of server than with the key of request, because the surrounding words push it that way during training. So the value vector of server dominates the sum, and the vector representing it now carries information about the server. The word overloaded helps: servers are overloaded, requests are not.

Two things about that step matter more than the arithmetic. It happened for every token at the same time, not one after another. And the distance between it and server cost nothing, because they were compared directly.

Heads and layers

One set of queries and keys can only track one kind of relation. The paper runs eight sets in parallel, called heads, each working on 64 numbers instead of 512, so the total cost stays similar to a single large head. Different heads end up tracking different relations, such as which noun a pronoun refers to and which verb a subject belongs to.

Each token's result then goes through a small feed-forward network on its own, and the whole block repeats. The base model stacks six of these blocks, so the answer for it is refined six times before it leaves the encoder.

031 min

What the 2017 results actually were

The paper that introduced the Transformer reported two models on the WMT 2014 translation benchmarks.

Base modelBig model
Blocks66
Vector size5121024
Attention heads816
Parameters65 million213 million
Training steps100,000300,000
Training time12 hours3.5 days

Both were trained on one machine with eight NVIDIA P100 GPUs. The big model scored 28.4 BLEU on English-to-German, which the authors state improved on the previous best results, ensembles included, by more than 2 BLEU. On English-to-French it scored 41.8 BLEU, a single-model record at the time, and the paper describes the 3.5-day training run as a small fraction of the training costs of the best models then published.

BLEU is a score from 0 to 100 that compares machine output against human reference translations. A 2-point gain at that level was large for the field.

The number worth carrying away is 213 million parameters. That is roughly one thousandth the size of the language models this architecture is used for today. The architecture that won on translation in 2017 was not changed much to become the architecture behind text generation; it was mostly made bigger and trained on far more data.

042 min

What the Transformer means for your work

Four consequences follow from the mechanism, and each one lands on a different desk.

For an engineer: attention work grows with the square of the input length. Doubling a prompt from 4,000 to 8,000 tokens roughly quadruples the attention arithmetic, even though the number of tokens only doubled. This is the reason chunking a long document into retrieved passages can be cheaper than sending the document, and the reason a prompt that grew slowly over six months eventually becomes a cost line somebody asks about.

For a product manager: the model keeps nothing between calls. The architecture has no state that survives a request. Everything the model appears to remember about a conversation is text that your application re-sent. A chat that feels like it has a memory is a chat that is paying to resend its own history on every turn.

For a designer: input is read all at once, output is produced one token at a time. Generating token 200 requires token 199 to exist first. That asymmetry is why streaming interfaces exist, and why time-to-first-token and tokens-per-second are two separate numbers that move independently. A change that shortens the prompt improves the first; only a faster model or a shorter answer improves the second.

For a founder: the architecture is not the advantage. The 2017 paper is public, and a working implementation is a few hundred lines. What is not public is the training data, the trained weights, the evaluation set, and the serving infrastructure. A product plan that treats "we use Transformers" as a differentiator is describing an ingredient every competitor also has.

051 min

What the Transformer costs

The quadratic term is not a rounding error. Dao and colleagues open the 2022 FlashAttention paper by stating plainly that the time and memory of self-attention are both quadratic in sequence length, and that this is what makes Transformers slow on long inputs.

Their fix was not a new approximation. FlashAttention computes exactly the same attention, but reorders the work so that far less data moves between the GPU's main memory and its much faster on-chip memory. The reported gains were 15 percent faster training on BERT-large at 512 tokens, three times faster on GPT-2 at 1,000 tokens, and 2.4 times faster on long-range tasks between 1,000 and 4,000 tokens.

Two things are worth reading off that list. The speedup grows with sequence length, which tells you the bottleneck was memory movement rather than arithmetic. And the largest published gains came from engineering under the architecture rather than from changing it, which is where much of the last few years of work has gone.

061 min

Transformer vs. nearby concepts

Compared withThe one fact that decides
Not in the library yetRecurrent neural networkWhether positions are processed in order. A recurrent network must finish position 49 before position 50; a Transformer computes all positions at once.
PsychologyAttentionAttention is one operation. The Transformer is an architecture built almost entirely out of that operation, plus positional encoding, feed-forward layers and normalisation. Attention existed before 2017 and was used inside recurrent models; the paper's claim was that it was enough on its own.
Not in the library yetLarge language modelA Transformer is the shape of the network. A large language model is a specific trained set of weights, almost always in that shape. Comparing them is like comparing a blueprint with a building.
AIMixture of expertsThese are not alternatives. Mixture of experts replaces the feed-forward part of a Transformer block with several feed-forward networks and a router that picks a few per token. It is a Transformer variant, not a rival architecture.

The question "should I use a Transformer or an LLM" has no answer because the two are not at the same level. The question "should I use a Transformer or a state space model" is real, and the section on contested evidence below is where it is decided.

072 min

A second case: the same architecture on images

In 2020, Dosovitskiy and colleagues at Google published An Image is Worth 16x16 Words. They cut an image into fixed 16-by-16 pixel squares, flattened each square into a vector, and fed the resulting sequence into a standard Transformer. The paper reports that a pure Transformer applied directly to sequences of image patches performs very well on image classification, without the convolutional layers that had defined computer vision for a decade.

One variable differs between this case and the 2017 translation model: what counts as a token. In 2017 a token was a chunk of text. In 2020 it was a patch of pixels. Everything downstream β€” queries, keys, values, heads, stacked blocks β€” was left alone.

What the contrast teaches is not available from either case alone. The translation result could be read as a good method for language. The vision result shows that the architecture assumes almost nothing about language specifically. It assumes only that the input can be cut into a sequence of pieces, and that which pieces matter depends on the other pieces present. Audio, video frames and protein sequences were all brought under the same architecture on that basis.

This also explains a practical pattern. When a new modality gets a capable model quickly, it is usually because somebody found a sensible way to turn it into tokens, not because a new architecture was designed for it.

082 min

Where the evidence is contested

The strongest current objection comes from Albert Gu and Tri Dao, whose 2023 Mamba paper attacks the quadratic cost directly. Their argument is that attention's cost on long sequences is not an implementation problem to be optimised away but a property of the operation, and that a selective state space model can do the same job with cost that grows linearly rather than quadratically.

Stated at its strongest, their case is empirical rather than theoretical. Earlier sub-quadratic architectures existed and lost on language, and the paper says so. Mamba's claim is that it closes that gap: five times higher throughput than Transformers at inference, linear scaling in sequence length, and a three-billion-parameter model that the authors report matches Transformers twice its size on language modelling.

The counter-case is that Transformer results at frontier scale are far more numerous, the tooling is far more mature, and the largest deployed systems as of September 2026 remain attention-based, often with attention and state space layers mixed rather than one replacing the other.

For a practitioner the disagreement resolves more cleanly than the literature does. If your sequences are short enough that attention cost is not your largest line item, this debate does not reach your decision. If you are processing very long sequences and have measured attention as the bottleneck, it is worth evaluating, and worth evaluating on your own data rather than on the benchmark tables.

092 min

How the Transformer changed since

The 2017 architecture had an encoder that read the input and a decoder that wrote the output. Almost none of that survives unchanged in current practice.

2018 β€” the encoder alone. BERT kept the encoder, threw away the decoder, and trained by hiding words and predicting them from both directions at once. For several years this was the default for classification and search ranking.

2020 β€” the decoder alone. GPT-3 kept the decoder, threw away the encoder, and trained purely on predicting the next token at 175 billion parameters. This is the branch that current chat models descend from. The encoder-decoder pair from the original paper is now the minority design.

2021 β€” position handled differently. Su and colleagues introduced rotary position embedding, which rotates the query and key vectors by an angle that depends on position rather than adding a fixed pattern to the input. It replaced the sinusoidal scheme in most new models.

2023 β€” cheaper inference. Ainslie and colleagues introduced grouped-query attention, which lets several query heads share one set of keys and values. The paper reports quality close to full multi-head attention at close to the speed of the single-key-value version, and it can be added to an existing model using about 5 percent of the original pre-training compute.

So the name is stable and the details are not. A model described today as a Transformer is usually decoder-only, with rotary positions and grouped-query attention, none of which are in the paper that named it.

?8 questions

Questions people ask

Is a Transformer the same thing as an LLM?

No. A Transformer is an architecture, meaning the shape of the network. A large language model is a specific set of trained weights, almost always built in that shape. Transformers are also used for images, audio and protein sequences.

Why is it called a Transformer?

The 2017 paper named it without explaining the choice. The architecture transforms one sequence into another, which is what machine translation needs. The name has no relation to electrical transformers or to the toy franchise.

What does attention actually do in a Transformer?

It lets each token pull information from other tokens in the same sequence. Each token issues a query, every token offers a key, and the tokens whose keys match the query contribute most to the result.

Why do Transformers struggle with long inputs?

Every token attends to every other token, so the work grows with the square of the sequence length. FlashAttention (2022) reduced the constant cost substantially by moving less data around, but the quadratic relationship itself remains.

Do I need to understand Transformers to build with language models?

Not the arithmetic. Three consequences are worth knowing: long prompts cost more than proportionally, the model keeps no state between calls, and output is generated one token at a time while input is read at once.

What is the difference between a Transformer and a recurrent neural network?

Order of processing. A recurrent network computes positions one after another and passes information along a chain. A Transformer computes all positions simultaneously and connects any two positions directly, whatever the distance between them.

Is anything replacing the Transformer?

State space models such as Mamba (2023) are the serious challenger, claiming linear rather than quadratic scaling and five times the inference throughput. As of September 2026 the largest deployed systems remain attention-based, sometimes mixing both layer types.

Does a Transformer need a GPU?

To train at any scale, yes. The original 2017 models used eight NVIDIA P100 GPUs for 12 hours and 3.5 days respectively. Small trained Transformers run acceptably on ordinary processors for inference.

Β§9 sources

Sources

  1. Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762

  2. Devlin, J. et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805

  3. Brown, T. et al. (2020). Language Models are Few-Shot Learners (GPT-3). arXiv:2005.14165

  4. Dosovitskiy, A. et al. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv:2010.11929

Show all 9 sources
  1. Dao, T. et al. (2022). FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. arXiv:2205.14135

  2. Su, J. et al. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864

  3. Ainslie, J. et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245

  4. Gu, A. and Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752

  5. Alammar, J. The Illustrated Transformer

Keep reading

More from AI

All of AI
All of AI