Attention and transformers

Attention networks provide an alternative to RNNs:

When multiple attention layers are combined, they form transformers, powerful architectures used for tasks like translation.

Transformers also serve as the foundation for advanced models (LLMs).

Embedding

Word (or token) embedding represents words as vectors (lists of numbers) instead of a single number.

This allows models to manipulate word meanings and their relationships mathematically.

Embedding analogy

Basic vector arithmetic:

Basic vector arithmetic

By adding and subtracting vectors, we can create new combinations of traits:

New combinations of traits through vector arithmetic

This illustrates two core ideas of embeddings:

Embedding words

Word embeddings represent words as vectors in a high-dimensional space.

An algorithm called embedder:

Because the space is high-dimensional, words can belong to multiple overlapping clusters.

Word embeddings support vector arithmetic:

Embeddings are not perfect but capture many subtle relationships well.

In this similarity graph:

Similarity graph

Embeddings improve predictions: because similar words are close together, a model predicting the next word might output a vector near "roar", producing related words like "blast" instead of something unrelated like "daffodil".

Pretrained embeddings models like GloVe, word2vec, and fastText are freely available.

Word embeddings can also be extended to sentence embeddings, allowing entire sentences to be compared semantically.

ELMo

ELMo (Embedding from Language Models) improves on traditional word embeddings by addressing word meaning ambiguity (for example, "train" as a noun vs. a verb).

It generates contextualized word embeddings: the same word gets different vectors depending on both preceding and following context.

ELMo's architecture:

ELMo

As a result, ELMo can distinguish between different meanings of the same word based on context.

Pretrained ELMo models are widely available. They are trained on large general datasets and can be fine-tuned for specific domains like medicine or law.

Algorithms like ELMo are typically used as an embedding layer at the very beginning of NLP systems, where they convert input words into embeddings before any further processing happens.

Icon for an embedding layer, representing the transformation of words into context-aware, semantically rich embeddings:

Icon for an embedding layer

Attention

When translating a sentence, not all words are equally relevant for translating a specific word.

Attention (or self-attention) identifies which words influence each other, allowing the model to focus on the important ones.

Attention is a flexible concept that can be applied beyond NLP, such as in CNNs, to emphasize relevant parts of the input.

Modern attention mechanisms often use the query-key-value (QKV) framework.

Self-attention

Self-attention lets each word in a sentence determine how much to focus on every other word, including itself.

Self attention core mechanism

Using attention to determine the contribution of all the words to the word dog:

Self attention

The same attention process is applied independently to every word: each word serves as a query and compares to all words in the sentence.

Self-attention has no sequential dependencies, so it can be computed in parallel in four steps, regardless of sentence length (with sufficient compute/memory):

  1. transform inputs to queries, keys, and values
  2. score queries vs. keys
  3. scale values using scores
  4. sum scaled values to produce outputs

Self-attention is implemented as a dedicated attention layer in deep neural networks.

It relies on word embeddings. This is essential because it ensures that query–key similarity scores reflect true linguistic similarity.

Q/KV attention

In self-attention, QKV come from the same input data, which led to the name self-attention.

Q/KV attention is a variation where the slash indicates that queries come from one source and keys/values from another.

It is used to add attention to an encoder–decoder model (like seq2seq), where queries come from one part and keys/values from the other. Because of this, it is also called an encoder–decoder attention layer.

Multi-head attention

Words can be similar in many ways. For example, in songwriting, words may be considered similar based on how they sound (e.g., rhyme or rhythm).

Multi-head attention allows a model to compare words using multiple criteria simultaneously.

It works by running several independent attention networks, called heads, in parallel:

The outputs of all heads are combined and passed through a fully connected layer, keeping the output the same shape as the input and allowing layers to be stacked easily.

Multi-head attention can be applied to both self-attention and Q/KV networks.

Icons for attention layers

Icons for the different types of attention layers

Transformers

Transformers were designed to replace RNNs:

Transformers rely on three additional concepts:

  1. skip connections
  2. norm-add
  3. positional encoding

Skip connections

Skip connections (residual connections) allow a layer to focus only on the changes it needs to make, rather than processing the entire input.

The layer computes small changes, which are then added to the original input:

output = input + small changes

The direct path that carries the original input to the output is called a skip connection or residual connection (due to its mathematical interpretation):

Skip connection

Skip connections make layers smaller, faster, and easier to train by improving gradient flow during backpropagation.

In transformers, skip connections are used not only for efficiency but also to preserve positional information in the input.

Norm-add

norm-add is a shorthand for combining:

  1. layer normalization (a regularization technique)
  2. skip connection

Norm-add

Layer normalization can be applied before the layer or either before or after the skip-connection addition, and in practice, all placements perform similarly.

Positional encoding

Transformers process all words in parallel, so they don't know word order like RNNs which naturally process words sequentially.

Positional encoding solves this by injecting each word's position into its representation.

Positional embedding:

Positional embedding

The transformer architecture preserves positional information through skip connections, which re-add each layer's input (including positional information) so it remains available to subsequent layers.

Assembling a transformer

Generic version of a transformer model using word-level translation as an example:

Transformer

A single encoder block:

A single encoder block

A single decoder block:

A single decoder block

A key detail in the decoder is masking:

Masking is applied in the first self-attention layer of each decoder block, making it a masked (multi-head self-attention) layer.

Transformers in action

Transformers outperform RNNs because they can be trained in parallel and don't rely on recurrent internal states.

However, their attention mechanism can require significant memory, though adjustments exist to mitigate this.

BERT and GPT-2

These two architectures have used transformer blocks for a wide range of tasks extending far beyond translation.

BERT

BERT (Bidirectional Encoder Representations from Transformers) uses a series of encoder blocks:

BERT architecture

Originally trained on two tasks:

  1. next sentence prediction (NSP): decide if the second sentence logically follows the first
  2. cloze task (masked language modeling): some words are removed in sentences, and BERT learns to fill in the blanks using context

The original large BERT has 340 million parameters and was trained on Wikipedia and 10,000+ books. Many pre-trained versions are freely available.

Unlike RNNs that are either unidirectional (left-to-right) or shallowly bidirectional (like ELMo, which has just two layers in each direction), BERT uses attention layers to consider the influence of every word on every other word across multiple stacked encoder blocks.

This is why BERT is sometimes called deeply bidirectional:

GPT-2

GPT-2 (Generative Pretrained Transformer 2) uses a series of decoder blocks:

GPT-2 architecture

The original GPT-2 model had 1.5 billion weights, 48 decoder blocks, and 12 attention heads, processing up to 512 tokens at a time. It was trained on the WebText dataset (40 GB, ~8 million documents).

Text generation can use different prompting setups:

GPT-2 can produce repetitive loops or boring output because the model chooses the highest-probability next word, and some phrases can form high-probability cycles in the probability distribution. It doesn't know when repetition happens unless guided otherwise.

Techniques to improve output quality:

After fine-tuning, GPT-2 can produce grammatical, context-aware narratives.

GPT-2 can also perform tasks such as question answering, summarization, and translation.

Generators discussion

Bigger models can produce better text (up to a point):

Enormous cost and resources:

What GPT-3 can do:

Scaling might keep working… but maybe not forever:

Core limitation:

Data poisoning

Data poisoning is an attack on NLP systems where malicious actors manipulate the training data so the model learns false or misleading information leading to incorrect or unreliable outputs.

Poisoning is hard to detect:

Data poisoning is dangerous because attackers do not need access to the system itself. They can publish manipulative content online that later becomes part of training data.

Real-world risks and consequences:

Currently, there are no robust methods to detect, prevent, or certify NLP systems as free from poisoning or bias.

Previous Recurrent neural networks All ⏎ Next Reinforcement learning

A Kemar Joint