Attention and transformers
Attention networks provide an alternative to RNNs:
- they remove the need for a state memory
- they enable parallel processing
When multiple attention layers are combined, they form transformers, powerful architectures used for tasks like translation.
Transformers also serve as the foundation for advanced models (LLMs).
Embedding
Word (or token) embedding represents words as vectors (lists of numbers) instead of a single number.
This allows models to manipulate word meanings and their relationships mathematically.
- animals plotted on a 2D chart (using e.g., speed horizontally and weight vertically)
- each animal's position can be treated like a vector from the origin
- the axes aren't labeled but that's ok: embeddings place items in a space where relationships matter more than the meaning of each dimension (real vectors might have thousands of dimensions)
- here even without knowing the axis meanings, the layout still captures useful relationships
Basic vector arithmetic:
- addition
A + B:- place the tail of
Bat the head ofA - the resulting tail
A + Bruns from the tail ofAto the head ofB
- place the tail of
- subtraction
A - B:- flip
B180° to make–B - add
Aand–Busing the same tip-to-tail method - the resulting tail
A – Bgoes from the tail ofAto the head of–B
- flip
By adding and subtracting vectors, we can create new combinations of traits:
This illustrates two core ideas of embeddings:
- items are arranged in a space where similar properties place them near each other
- even without knowing what each dimension represents, we can perform vector arithmetic to generate new relationships or meanings
Embedding words
Word embeddings represent words as vectors in a high-dimensional space.
An algorithm called embedder:
- determines what each axis in this space should mean
- places similar words close to each other
- gives every word a list of numbers describing its coordinates in this space
Because the space is high-dimensional, words can belong to multiple overlapping clusters.
Word embeddings support vector arithmetic:
king - man + woman ≈ queen- this works because the embedding captures relationships such as gender or royalty
Embeddings are not perfect but capture many subtle relationships well.
In this similarity graph:

- each word is most similar to itself
- related words cluster together as small dark blocks
- sometimes strange similarities appear (like fish matching well with chocolate, or blue with caramel) likely due to quirks in the training data
Embeddings improve predictions: because similar words are close together, a model predicting the next word might output a vector near "roar", producing related words like "blast" instead of something unrelated like "daffodil".
Pretrained embeddings models like GloVe, word2vec, and fastText are freely available.
Word embeddings can also be extended to sentence embeddings, allowing entire sentences to be compared semantically.
ELMo
ELMo (Embedding from Language Models) improves on traditional word embeddings by addressing word meaning ambiguity (for example, "train" as a noun vs. a verb).
It generates contextualized word embeddings: the same word gets different vectors depending on both preceding and following context.
ELMo's architecture:
- two two-layer-deep RNNs are grouped by direction
- input: a sequence of words
- each word is transformed into two tensors:
- forward tensor: computed by the forward RNNs (F1 and F2), capturing the context from preceding words
- backward tensor: computed by the backward RNNs (B1 and B2), capturing the context from following words
- outputs from both groups are combined to create contextualized word embeddings
- ELMo diagrams use a red background referencing the Sesame Street character Elmo
As a result, ELMo can distinguish between different meanings of the same word based on context.
Pretrained ELMo models are widely available. They are trained on large general datasets and can be fine-tuned for specific domains like medicine or law.
Algorithms like ELMo are typically used as an embedding layer at the very beginning of NLP systems, where they convert input words into embeddings before any further processing happens.
Icon for an embedding layer, representing the transformation of words into context-aware, semantically rich embeddings:
Attention
When translating a sentence, not all words are equally relevant for translating a specific word.
Attention (or self-attention) identifies which words influence each other, allowing the model to focus on the important ones.
Attention is a flexible concept that can be applied beyond NLP, such as in CNNs, to emphasize relevant parts of the input.
Modern attention mechanisms often use the query-key-value (QKV) framework.
Self-attention
Self-attention lets each word in a sentence determine how much to focus on every other word, including itself.
- we want to translate the word
dog - this example is limited to comparing
dogtodinner, but in practicedogis compared to every word in the sentence - each tensor is transformed into a new tensor using three small neural networks:
- red network:
dog→Q(query) - blue network:
dinner→K(key) - green network:
dinner→V(value)
- red network:
- a scoring function compares
Q(dog)withK(dinner)and produces a similarity score - that score is used to scale
V(dinner): higher similarity → larger score →dinneris more relevant to the translation ofdog
Using attention to determine the contribution of all the words to the word dog:
- attention is applied to all words simultaneously to compute the final representation for
dog - shared networks: only three neural networks are used in total (Q, K, and V), applied to every word once
- softmax: normalizes scores, keeps numbers stable, emphasizes strongest matches
- output for
dog= sum of all word values, each scaled by its attention score (this includes the value ofdogitself) - why does a word usually scores most highly with itself?
- its query and key come from the same word, making them very similar
- this is not a problem:
- often the word itself carries the most important information
- in other cases, context matters more: when word order changes, when a word has no direct translation and must rely on other words, or pronoun resolution
The same attention process is applied independently to every word: each word serves as a query and compares to all words in the sentence.
Self-attention has no sequential dependencies, so it can be computed in parallel in four steps, regardless of sentence length (with sufficient compute/memory):
- transform inputs to queries, keys, and values
- score queries vs. keys
- scale values using scores
- sum scaled values to produce outputs
Self-attention is implemented as a dedicated attention layer in deep neural networks.
It relies on word embeddings. This is essential because it ensures that query–key similarity scores reflect true linguistic similarity.
Q/KV attention
In self-attention, QKV come from the same input data, which led to the name self-attention.
Q/KV attention is a variation where the slash indicates that queries come from one source and keys/values from another.
It is used to add attention to an encoder–decoder model (like seq2seq), where queries come from one part and keys/values from the other. Because of this, it is also called an encoder–decoder attention layer.
Multi-head attention
Words can be similar in many ways. For example, in songwriting, words may be considered similar based on how they sound (e.g., rhyme or rhythm).
Multi-head attention allows a model to compare words using multiple criteria simultaneously.
It works by running several independent attention networks, called heads, in parallel:
- each head is initialized independently and learns to focus on different patterns or aspects of the input
- this enables the model to capture diverse relationships at the same time
The outputs of all heads are combined and passed through a fully connected layer, keeping the output the same shape as the input and allowing layers to be stacked easily.
Multi-head attention can be applied to both self-attention and Q/KV networks.
Icons for attention layers
Transformers
Transformers were designed to replace RNNs:
- they use attention mechanisms instead of recurrence to learn relationships between words (e.g., for translation)
- they were introduced in the paper "Attention Is All You Need":
- authors named the model a transformer
- but the name is ambiguous and doesn't clearly reflect the model's reliance on attention
Transformers rely on three additional concepts:
- skip connections
- norm-add
- positional encoding
Skip connections
Skip connections (residual connections) allow a layer to focus only on the changes it needs to make, rather than processing the entire input.
The layer computes small changes, which are then added to the original input:
output = input + small changes
The direct path that carries the original input to the output is called a skip connection or residual connection (due to its mathematical interpretation):
Skip connections make layers smaller, faster, and easier to train by improving gradient flow during backpropagation.
In transformers, skip connections are used not only for efficiency but also to preserve positional information in the input.
Norm-add
norm-add is a shorthand for combining:
- layer normalization (a regularization technique)
- skip connection
Layer normalization can be applied before the layer or either before or after the skip-connection addition, and in practice, all placements perform similarly.
Positional encoding
Transformers process all words in parallel, so they don't know word order like RNNs which naturally process words sequentially.
Positional encoding solves this by injecting each word's position into its representation.
Positional embedding:
- use a mathematical function (often sine/cosine functions) to convert a word's position into a vector
- add this vector to the word's embedding vector (in contrast to appending, which requires more bits)
- the model learns to distinguish meaning and position
The transformer architecture preserves positional information through skip connections, which re-add each layer's input (including positional information) so it remains available to subsequent layers.
Assembling a transformer
Generic version of a transformer model using word-level translation as an example:
- a transformer consists of encoder and decoder blocks
- dashed lines indicate multiple identical layers
A single encoder block:
- starts with a multi-head self-attention layer:
- QKV all come from the same input
- wrapped in a norm-add node
- this layer captures relationships between all words in the input sentence
- followed by two layers referred to collectively as a pointwise feed-forward layer:
- two consecutive 1×1 convolution layers (described as a pair of modified fully connected layers by the original paper)
- they refine the attention output, removing redundancy and focusing on the most important information
- the first uses ReLU, the second has no activation
- wrapped in a norm-add node
A single decoder block:
- starts with a multi-head self-attention layer:
- uses masking
- input: the words generated so far (
[START]token if at the beginning) - purpose: to determine which words are most strongly related to each other
- wrapped in a norm-add node
- followed by a multi-head Q/KV attention layer:
- Q come from the previous layer, KV come from the concatenated outputs of all the encoder blocks
- wrapped in a norm-add node
- followed by a pointwise feed-forward layer
A key detail in the decoder is masking:
- if a transformer sees future words, it could "cheat" instead of learning to predict correctly
- example:
- sentence: "My dog loves taking long walks"
- to predict "long", the model should only see "My dog loves taking", not "long" and "walks"
- to predict "walks", the model should only see all words up to "long"
- masking ensures the model learns proper sequential dependencies without "cheating" while allowing parallel computation
Masking is applied in the first self-attention layer of each decoder block, making it a masked (multi-head self-attention) layer.
Transformers in action
Transformers outperform RNNs because they can be trained in parallel and don't rely on recurrent internal states.
However, their attention mechanism can require significant memory, though adjustments exist to mitigate this.
BERT and GPT-2
These two architectures have used transformer blocks for a wide range of tasks extending far beyond translation.
BERT
BERT (Bidirectional Encoder Representations from Transformers) uses a series of encoder blocks:
- operations are run in parallel (even if only a single line is used)
- the dashed line indicates additional stacked encoder blocks
- input: sentence split into tokens
- structure:
- word embedder
- position embedder
- followed by multiple transformer encoder blocks: each one uses self-attention, allowing BERT to consider how every word relates to every other word in the sentence
- BERT's output: a contextualized representation for each word in the sentence
- the output can be used for downstream tasks (classification, question answering, etc.)
- after the final encoder block, a simple downstream classifier is attached (often a fully connected layer)
- this classifier takes the sentence embedding produced by BERT and outputs a task-specific result
- e.g., determining if an input sentence is grammatical or not, which is basically a classification problem producing a yes/no answer
- BERT diagrams use a yellow background referencing the Sesame Street character Bert
Originally trained on two tasks:
- next sentence prediction (NSP): decide if the second sentence logically follows the first
- cloze task (masked language modeling): some words are removed in sentences, and BERT learns to fill in the blanks using context
The original large BERT has 340 million parameters and was trained on Wikipedia and 10,000+ books. Many pre-trained versions are freely available.
Unlike RNNs that are either unidirectional (left-to-right) or shallowly bidirectional (like ELMo, which has just two layers in each direction), BERT uses attention layers to consider the influence of every word on every other word across multiple stacked encoder blocks.
This is why BERT is sometimes called deeply bidirectional:
- "bidirectional": looks at all words on both sides of a token simultaneously
- "deep": repeats this across many stacked encoder layers
- some call it deeply dense, since every word is considered simultaneously
GPT-2
GPT-2 (Generative Pretrained Transformer 2) uses a series of decoder blocks:
- operations are run in parallel even if only a single line is used
- the dashed line indicates additional stacked decoder blocks
- input: sentence split into tokens
- structure:
- word embedder
- position embedder
- followed by multiple transformer decoder blocks: each one uses a masked self-attention layer and two 1×1 convolutions
- output: a probability distribution over the next token (word)
The original GPT-2 model had 1.5 billion weights, 48 decoder blocks, and 12 attention heads, processing up to 512 tokens at a time. It was trained on the WebText dataset (40 GB, ~8 million documents).
Text generation can use different prompting setups:
- zero-shot:
- the model is given only a prompt
- example: "Describe today's outfit"
- GPT-2 generates text based purely on what it learned during training
- one-shot:
- the model is given one example to guide its output
- example:
- "Yesterday I wore a blue shirt and black pants"
- prompt: "Describe today's outfit"
- GPT-2 uses the example to produce output in a similar style
- few-shot:
- the model is given a few examples (usually 2–5) to guide its output
- gives the model more context and helps it generate more consistent text
- more shots guide the model better, but the goal is to use as few as possible while still getting good results
GPT-2 can produce repetitive loops or boring output because the model chooses the highest-probability next word, and some phrases can form high-probability cycles in the probability distribution. It doesn't know when repetition happens unless guided otherwise.
Techniques to improve output quality:
- n-gram penalties → reduce repetition
- penalize sequences of
nwords (n-grams) if they repeat
- penalize sequences of
- beam search → consider multiple continuations
- choose several likely next words
- extend each one with additional words to form several possible sequences (paths)
- score each sequence based on its probability
- keep the most probable one and discard the others
- temperature adjustment → add randomness
- instead of always picking the most probable word, choose among several likely words
- temperature = 0 → always pick the highest-probability word
- higher temperature → more varied and creative output
After fine-tuning, GPT-2 can produce grammatical, context-aware narratives.
GPT-2 can also perform tasks such as question answering, summarization, and translation.
Generators discussion
Bigger models can produce better text (up to a point):
- GPT-2 showed that large neural networks can generate convincing text
- instead of inventing a new design, researchers asked: what if we just make everything bigger?
- GPT-3 followed this idea:
- much longer input sequences
- many more layers
- many more attention heads
- 175 billion parameters vs. GPT-2's 1.5 billion
- result: much more fluent and flexible text generation
Enormous cost and resources:
- training GPT-3 required:
- hundreds of GPU-years
- millions of dollars
- it was trained on a massive dataset from the web and books, cleaned down to hundreds of billions of words
- this is a game only the big and rich can play
What GPT-3 can do:
- write computer code
- draft legal-style text
- simplify complex documents
- generate creative writing (fiction, poetry)
- power role-playing and interactive storytelling
Scaling might keep working… but maybe not forever:
- researchers speculate that even larger models, trained on trillions of words, might capture most useful language patterns
- but this is uncertain, and the costs grow rapidly
- fine-tuning extremely large models for specific tasks also becomes harder
Core limitation:
- studies show that GPT-3 still performs poorly on tasks requiring deep understanding, reasoning, morality, or common sense
- these systems do not understand meaning
- they generate words based on statistical patterns
- they reproduce biases present in their training data, raising concerns about accuracy, fairness, and social impact
Data poisoning
Data poisoning is an attack on NLP systems where malicious actors manipulate the training data so the model learns false or misleading information leading to incorrect or unreliable outputs.
Poisoning is hard to detect:
- NLP models are trained on massive datasets
- poisoned data can blend in with normal-looking text
- some attacks use indirect or hidden language (concealed data poisoning)
Data poisoning is dangerous because attackers do not need access to the system itself. They can publish manipulative content online that later becomes part of training data.
Real-world risks and consequences:
- incorrect decisions can seriously impact people's lives
- especially when NLP systems are used in sensitive areas (education, medicine, law, fraud detection)
Currently, there are no robust methods to detect, prevent, or certify NLP systems as free from poisoning or bias.