Backpropagation

Backpropagation is the core algorithm used to train neural networks.

It works by adjusting the network's weights (initially random) to reduce error and produce the correct outputs.

To reduce error, we need to know whether each weight should be increased or decreased.

To figure that out, each neuron is given a delta value.

These delta values are calculated starting from the final output layer and moving backward to the first.

Because this process passes information backward through the layers, it's called backpropagation.

Backpropagation doesn't "propagate the error" but instead "propagates the gradient of the error".

Gradient descent

The gradient (or derivative) is the slope of the error curve at a particular weight.

Error curves and gradients

The sign of the gradient tells us the direction to move the weight to reduce error:

The process of moving each weight in the direction that reduces error is named gradient descent because it uses the gradient to "descend" (reduce) the error.

By calculating gradients for all weights, we can update them simultaneously.

However, there's a challenge:

To manage this: small adjustments are made to each weight in the hopes that any such mistakes won't drown out improvements.

Backpropagation

For clarity, we ignore activation functions: they are essential in real networks, but here they only add extra detail.

Without activation functions, neurons only perform multiplication and addition. This leads to a key insight:

When a neuron's output changes, the final output error changes by a proportional amount.

A tiny neural network

Consider a tiny network (4 neurons, 8 weights) whose task is to classify 2D points into two categories (class 1 and class 2).

A single output neuron could handle two classes, but a larger network is used to demonstrate backpropagation principles:

Tiny neural network

Conventions Explanation
Ao, Bo, etc. Output of neuron A, B, etc.
Aδ, Bδ, etc. Delta value for neuron A, B, etc.
P1, P2 Predictions made by the network. Note: P1 = Co and P2 = Do.
Am A change in the output of neuron A due to a change by an amount m.
E The network's total error.
Em Change in the error.

Output neurons

Backpropagation starts at the end of the network.

The output layer is a special case because there is no next layer.

So gradients come directly from the network error.

Calculating the network error

To obtain a single number representing the network's error:

The difference between prediction and label is measured using cross-entropy:

Network Error = CrossEntropy( (P1, P2), TrueLabel )

Deep learning libraries provide this function, which returns both the error and its gradient:

From gradient to delta

We can't see the full error curve for a neuron's output, but we can calculate its slope at any point:

Network's error changes as P1 changes

The delta (δ or Δ) is the gradient evaluated at a specific neuron's output:

Finding deltas for P1 and P2

For P1:

As soon as Co changes, the error curve changes too, so the derivative and Cδ must be recalculated.

The same process applies to find Dδ for P2.

The goal is to nudge all weights together until the error is as small as possible. Getting to exactly 0 error isn't always desirable, as continuing to train beyond a certain point can lead to overfitting.

Using deltas to change weights

Once we know the delta values for each neuron in the output layer, we can adjust the weights going into that layer.

Consider the weight AC:

To reduce the error, we adjust the weight in the opposite direction:

This gives the weight update rule: AC = AC − (Ao × Cδ).

This rule applies to all weights feeding into the output layer.

Other neurons' deltas

Next, we compute deltas for neurons in the previous hidden layer.

If the output of a neuron (like neuron A) changes, that change affects the network error through every neuron it connects to in the next layer (neurons C and D in our example).

A change in A's output affects the error via C:

Similarly, a change in A's output affects the error via D:

This means that a change in the output of neuron A influences the network's error through multiple paths at the same time:

Because neurons C and D do not affect each other, their contributions to the error are independent.

The total change to the error is just the sum of both contributions to the error:

Total change to the error

So the true delta for A must include both effects added together: Aδ = (AC x Cδ) + (AD x Dδ).

This is how to get the value of delta for any hidden neuron in any network:

Feed-forward and feed-backward are mirror processes

Deltas flow backward in the network, while outputs flow forward.

Example for any arbitrary neuron named H (we still ignore the activation function):

Forward and backward flows

Backpropagation on a larger neural network

Step Explanation Image

Setup

Example network:

  • 2 input neurons and 2 output neurons
  • four layers
  • three hidden layers (of 2, 4 and 3 neurons from left to right)

Start by evaluating a sample:

  • provide the values of its X and Y features to the input layer
  • the network produces predictions: P1 and P2

Compute deltas for the output layer

Compute the error for the output layer. Assume it's greater than zero (i.e., prediction isn't perfect).

Compute the delta for the upper neuron (arbitrarily chosen) using a loss function (e.g., cross-entropy) and save it with the neuron.

Repeat this process for all the other neurons in the output layer.

At this point, every neuron in the output layer has a delta.

Compute deltas for the layer 3

Move one layer backward.

For each neuron, compute its delta:

  • find the deltas of all next-layer neurons that use its output
  • multiply each by the connecting weight
  • sum the results

Compute deltas for the layer 2

Move one layer backward.

For each neuron, compute its delta:

  • find the deltas of all next-layer neurons that use its output
  • multiply each by the connecting weight
  • sum the results

Compute deltas for the layer 1

Finally, move to the first hidden layer.

For each neuron, compute its delta:

  • find the deltas of all next-layer neurons that use its output
  • multiply each by the connecting weight
  • sum the results

Once all deltas are known, we can update the weights.

The final step is deciding how much to adjust each weight: a question answered by choosing the right learning rate.

The learning rate

Delta values indicate the direction in which weights should change to reduce error, but they are only accurate for tiny changes.

If weights change too much, the network may overshoot the minimum and become unstable; if too little, learning becomes slow.

The amount of changes to the weights is controlled by the learning rate (η, eta), a hyperparameter that scales weight changes:

How the learning rate affects backpropagation?

Consider a simple binary classifier to separate the "two moons" dataset where about 1500 training data points are divided into two classes:

The two moons dataset

Classifier architecture:

Network architecture for the two moons dataset

The goal is to adjust those 37 weights using backpropagation so the output matches the correct label.

Running backprop successfully relies on making small changes to the weights, but what counts as "small" depends on the network and the data.

We have to experiment with different learning rates to find out:

Learning rate Explanation Image

0.5 (large)

Decision boundaries are terrible:

  • no clear decision boundaries
  • everything is assigned to a single class (light background)

Accuracy and loss:

  • accuracy stays around 0.5, meaning half the data points are misclassified
  • loss (or error) starts high and doesn't improve, even after many epochs

Weights during training:

  • one weight dominates the graph (it keeps overshooting its target)
  • other weights are changing too little to make a visible difference

These poor results aren't surprising because 0.5 is a very high learning rate and causes unstable updates.

0.05 (reduce by a factor of 10)

Major improvement: clear and accurate class boundaries appear.

Accuracy and loss: the network reaches 100% accuracy on both training and test sets after 10–16 epochs.

Weights during training:

  • many weights are now changing meaningfully
  • some weights become quite large, which can slow or inhibit learning (ideally weights stay within a small range like [-1,1], this can be controlled with regularization)

This setup works well: the network learns quickly and classifies the data perfectly.

Reducing the learning rate further to 0.01 has no real benefit here because the classification was already perfect at 0.05. For this network and this data, a smaller learning rate just makes training slower.

The learning rate must be tuned for each specific network and dataset (not too big, or too small, but just right).

If a network isn't learning, lowering the learning rate is often a good first step.

Algorithms can automatically adjust the learning rate in sophisticated ways.

Previous Neural networks All ⏎ Next Optimizers

A Kemar Joint