Overfitting and underfitting

Learning from limited examples involves balancing generalization and specificity:

Overfitting is the more common issue and is managed using regularization techniques.

Detecting and addressing overfitting

Overfitting happens when a model learns the training data too well (including its noise and specific details) so it struggles to make accurate predictions on new data.

Overfitting can be detected by monitoring the training error and the validation error (an estimate of how well the model is likely to perform on new data):

Overfitting and underfitting

Example:

Background music curves

Good models strike a balance:

The main goal is to perform well on new data, not just the training data.

Early stopping (prevent overfitting)

When training a machine learning model:

Early stopping (or last-minute stopping):

In practice:

Regularization

While early stopping halts training once validation error starts rising, regularization aims to delay that rise so we can train longer and improve both training and validation performance.

Regularization is achieved by:

Implementation:

For multi-layer architectures, specialized regularization methods like dropout, batchnorm, layer norm, and weight regularization provide additional ways to control overfitting.

Bias and variance

Imagine we've been measuring daily wind speed on a mountain over several months.

The observed data shows:

Since the exact amount of noise is unknown, the goal is to approximate the underlying curve by trying to fit a curve that captures the general trend:

Data for daily wind speed on a mountain over several months

Bias and variance can explain why different curve-fitting choices succeed or fail at recovering the "ideal" underlying curve.

We generate smaller datasets by randomly sampling 30 points each and then fit simple and complex curves to them:

Bias and variance

Simple curves have:

Complex curves follow the data much more closely and have:

Bias and variance:

Bias and variance help explain underfitting and overfitting:

Ideally, we'd like curves with a low bias and low variance, but in practice, when one goes down, the other often goes up.

The goal is to find the best bias-variance tradeoff for each specific situation.

In some applications:

Fitting a line with Bayes' rule

Bias and variance are concepts commonly used in frequentist analysis.

The Bayesian approach treats all possible curves as having probabilities, updating those probabilities as more data is collected, but never settling on a single "final true" answer.

We'll use Bayes' Rule to approximate the noisy atmospheric dataset with a straight line for simplicity.

Some tricks to make lines easier to visualize:

Title Explanation Image

Slope–intercept form

A line can be represented as a single point using two numbers:

  1. slope (horizontal axis)
    • measures how much the line is tilted from horizontal
    • 0 → horizontal line
    • ±1 → diagonal line
    • infinity → vertical line
    • for simplicity, we restrict slopes to the range [-2, 2] (the light-green region)
  2. intercept (vertical axis)
    • the value of the line as it crosses (intercepts) the Y axis
Slope–intercept form

Dual representation

XY and SI spaces enable a dual mapping between them:

  • XY space (Cartesian space): where lines are drawn
  • SI space (slope-intercept space): where each line is a point

This allows us to visualize many possible lines compactly.

Dual representation

Elegant point about the dual representation

If points are arranged along a line in SI space, the equation linking slope and intercept ensures that all the corresponding XY lines intersect at the same point.

Dual representation intersection

Bayesian curve fitting assigns a probability to every possible line based on how well it fits the data.

Example: likelihood of all possible straight lines for a single data point of the noisy dataset:

Fitting lines for just one data point of the noisy dataset

Full example: find the likelihood of all possible straight lines for the full noisy data set:

Step Image Explanation
1

Left:

  • prior: our initial belief before seeing any data is a Gaussian distribution showing that any line is possible
  • this prior assumes a horizontal line is most likely but allows others

Right:

  • 20 lines with high probability picked at random from this guess
2

Left:

  • choose one random data point (shown in red)
  • which lines would pass through this point?

Right:

  • this gives us the likelihood: a new probability distribution over lines
  • lines close to the point are more likely

Apply Bayes' Rule: posterior = prior × likelihood.

3

Left:

  • this gives us the posterior: an updated probability distribution over lines

Right:

  • lines drawn from this posterior show the system has learned that some areas (like the top) don't match the data

The posterior becomes the new prior.

4

Left:

  • choose another random data point (shown in red)
  • which lines would pass through both points reasonably well?

Right:

  • a new likelihood is computed for all lines that go through or near both points

Apply Bayes' Rule: posterior = new prior (step 3) × likelihood (step 4).

5

Left:

  • the posterior shrinks, meaning fewer lines now fit both points

Right:

  • lines drawn from this posterior cluster in the same direction as the two points we just learned from
6

Repeat this process again and again.

Each step refines the posterior further.

Over time, the system converges toward a small set of lines that fit the data well.

7

Bayes' Rule is useful in training a learning system:

In Bayesian terms:

Previous Training and testing All ⏎ Next Data preparation

A Kemar Joint