Recurrent neural networks

Recurrent neural networks (RNNs) are sequential models that remember information from earlier elements in a sequence.

A sequence is a set of samples where the order matters, like the words in a sentence.

In sequence processing, each element in the input is called a token (often a word).

Algorithms that can understand sequences can also generate new ones, including stories, music, or lyrics.

Working with language

Common natural language processing tasks

NLP (natural language processing) is the processing of natural language information by a computer.

Common NLP tasks include:

These tasks rely on language models, which don't understand language or the meaning of words. Instead, they create outputs that seem correct by relying purely on statistical methods.

Transforming text into numbers

Written text must first be converted into numbers.

Two main ways to convert text into numbers:

  1. character-based encoding: each character gets a unique number
  2. word-based encoding: each word gets a unique number

Then numbers are fed into an autoregressive neural network, which predicts the next number in the sequence and then turns those numbers back into words to see the generated text.

Fine-tuning and downstream networks

Models can start as general-purpose systems and be adapted for specialized tasks.

This specialization leverages the knowledge the original model has already learned.

Ways to specialize a system:

In practice, many systems combine fine-tuning and downstream networks.

Neural networks fail at language prediction

A tiny fully connected neural network is used to predict sequences:

A tiny fully connected neural network

The network is trained to make predictions using the sliding window approach:

The sliding window approach

After training it, this tiny fully connected neural network fails to predict words accurately. Such a tiny network cannot handle complex data like language.

Even a larger fully connected neural network can't handle the complexities of natural language because:

  1. it can't capture the structure of the text (its semantics)
    • predicting the next word requires understanding context from many words earlier
    • without this, too many possible continuations remain plausible
  2. tiny errors in predictions lead to incomprehensible text
    • words are stored as single numbers
    • tiny differences in predictions can give unrelated words
    • example: words "keep" and "flint" are assigned consecutive numbers: 1003 and 1004
      • the network predicts the next word as 1003.49
      • it is rounded to the nearest integer to convert to a word → 1003 → "keep"
      • if it predicts 1003.51 instead, rounding gives 1004 → "flint"
  3. it doesn't track the order of words in the input
    • this makes it impossible to resolve references (like pronouns) that depend on sequence order
    • e.g., "Bob told John that he was hungry" → who does "he" refer to?

We need something smarter than fully connected layers and representing words as single numbers.

Recurrent neural networks

Recurrent neural networks (RNNs) are designed to manage language as an ordered sequence.

They involve new concepts.

State

State is the condition of a system at a given moment which includes:

  1. the system's current situation
  2. any information it needs to remember from earlier inputs (possibly in a compressed or transformed form)

State evolves over time. With each new input, the system:

The order of inputs matters: changing the sequence leads to different states and different outputs.

Each input in a sequence is called a time step.

Recurrent cells

Sequential data is processed using recurrent cells, self-contained modules that both compute outputs and manage state over time:

A recurrent neural cell

The system contains one or multiple neural networks that learn how to manage state and produce outputs during training.

The cell's internal state:

A recurrent cell on its own layer forms a recurrent layer, and networks built mainly from these are RNNs. Depending on context, "RNN" may refer to the network, the layer, or the cell itself.

Recurrent cell and recurrent layer diagrams

Diagrams of long sequences can be unrolled (showing each step separately) or rolled-up (a compact version).

Recurrent cells in action

An unrolled RNN diagram illustrates how a recurrent cell predicts the next word in a five-word sequence:

Recurrent cells in action

As the sequence progresses, the hidden state accumulates a compact representation of the words seen:

Early in training, predictions may be inaccurate, but after enough training on real text, the RNN learns to represent the sequence effectively in its hidden state and assign high probability to the correct next word.

Backpropagation through time

Even though the above diagram shows several "cells", it is really the same recurrent cell reused at each time step with shared weights.

To train that recurrent cell:

To account for this, the gradient of the final error must be propagated backward through all previous steps to update the shared weights.

The solution is backpropagation through time (BPTT):

However, backpropagating gradients through many time steps can cause fundamental challenges in training RNNs:

Long short-term memory and gated recurrent networks

A long short-term memory (LSTM) network is a type of recurrent cell designed to prevent vanishing and exploding gradients.

It maintains an internal state (memory) that can be selectively updated over time:

An LSTM uses three internal networks that implement gates, collectively known as the gating mechanism:

An LSTM uses three internal networks

"Forgetting" means pushing values stored in the memory state toward zero, while "remembering" means adding new values to the memory state.

The core difference with basic RNN is the cell state:

LSTMs are so common that "RNN" often means an "LSTM network".

A common variation of the LSTM is the gated recurrent unit (GRU), and both LSTM and GRU are often tested to see which works best.

Different architectures

Recurrent cells can be either stacked or combined with other networks to handle complex sequence tasks.

CNN-LSTM Networks:

Deep RNNs:

Bidirectional RNNs (bi-RNNs):

Bidirectional RNNs

Seq2Seq

Seq2Seq ("sequence to sequence") is an algorithm for translating entire sentences between languages, rather than word by word.

Translation challenges:

Seq2Seq is conceptually similar to autoencoders, but the "latent vector" is called the context vector:

Seq2Seq

It uses two RNNs:

  1. encoder:
    • reads the input sentence word by word
    • updates its hidden state at each step
    • ignores outputs at each step; only the final hidden state matters
    • the final hidden state is the context vector, summarizing the input sentence
  2. decoder:
    • starts with the context vector as its initial hidden state
    • starts the output sequence with a [START]token
    • generates a sequence autoregressively: each generated word becomes the next input
    • stops when it produces an [END] token

This approach allows Seq2Seq to generate variable-length output sequences from fixed-length input sequences.

Limitations of Seq2Seq:

Seq2Seq is simple, widely used, and easy to implement. It works well for short sentences but struggles with long or complex ones.

Previous Autoencoders All ⏎ Next Attention and transformers

A Kemar Joint