Transformer Architecture
The neural network design, built around 'attention,' that made today's large language models possible.
Quick Answer
A transformer is a neural network architecture that processes an entire sequence of text at once and uses self-attention to weigh how relevant every word is to every other word, regardless of distance between them. Introduced in 2017, it replaced older sequence models and now underlies essentially all modern large language models, including GPT and Claude.
The Full Picture
Before transformers, models that processed language read text one word at a time in order, which made it hard to connect a word to something said much earlier in a long passage — the connection had to survive being passed along step by step. The transformer architecture, introduced in the 2017 paper 'Attention Is All You Need,' solved this by letting the model look at an entire sequence simultaneously and directly weigh the relevance of every word to every other word.
That mechanism is called self-attention. For each word, the model computes how much it should 'attend to' every other word in the sequence when building its understanding of that word in context — so in a sentence about a beam and a column, the model can directly link a pronoun to whichever noun it actually refers to, no matter how far apart they sit in the text. Stacking many attention layers lets the model build increasingly rich representations of meaning and context.
This design has two practical payoffs that drove the current AI boom: transformers can be trained in parallel across huge amounts of text (rather than one word at a time), which made training on internet-scale data feasible, and they handle long-range context far better than earlier architectures. That combination of scalability and context-handling is why transformers underlie the large language models used for chat, summarization, code generation, and document understanding.
For anyone evaluating AI tools, the practical implication is context window and coherence: a transformer-based model can reference something stated pages earlier in a document because attention connects distant text directly, which is part of why modern AI can summarize a long spec section or contract coherently rather than losing track of earlier clauses.
Real Examples
Common Misconceptions
People assume: 'Transformer' and 'large language model' mean the same thing.
Actually: A transformer is the underlying architecture — the mathematical design for processing sequences with attention. An LLM is a specific model built using that architecture, trained on large text data. Transformers are also used outside language, for images and other data types.
People assume: Transformers understand meaning the way humans do.
Actually: Self-attention lets a transformer track statistical relationships between tokens with remarkable effectiveness, but that's pattern-based context-weighing, not comprehension in the human sense — which is part of why transformer-based models can still produce fluent but incorrect output (hallucination).
Frequently Asked Questions
What is self-attention in a transformer?
A mechanism that lets the model weigh how relevant every element in a sequence is to every other element when building context, rather than only looking at nearby words. It's what lets transformers connect distant but related pieces of text directly.
Why did transformers replace earlier language models?
Earlier recurrent models processed text sequentially and struggled to retain context over long passages, plus they were slow to train because each step depended on the last. Transformers process sequences in parallel and use attention to handle long-range context, making them faster to train at scale and better at coherence.
Is a transformer the same as an LLM?
No. A transformer is the architecture; a large language model is a specific model built on that architecture and trained on large amounts of text. Nearly all current LLMs are transformer-based, but transformers are also used for non-language tasks like image and protein modeling.
Are transformers used outside of text and chatbots?
Yes. The same attention-based architecture has been adapted for images (vision transformers), audio, and multimodal systems that combine text with images — relevant to AI that reads both drawings and written specs together.
What is a context window and how does it relate to transformers?
The context window is the amount of text a transformer-based model can consider at once when generating a response. Because attention operates across the whole window, a larger context window lets the model reference more of a long document — a spec section or contract — in a single pass.