Curves and surfaces

Derivatives and gradients are important properties of curves and surfaces.

We can think of the error as a surface. To reduce error, we move toward a minimum on its surface. The steepest downhill direction is given by the negative gradient.

These ideas are at the heart of backpropagation.

The nature of functions

Functions:

Functions are deterministic if no randomness is involved.

Curves and surfaces are graphical versions of mathematical functions.

We assume all our curves and surfaces satisfy certain conditions since these are the kinds most useful in machine learning and optimization:

  1. continuous
    • the curve can be drawn without lifting the pen
    • it has no jumps or breaks (no discontinuities)
  2. smooth
    • the curve has no sharp corners or cusps
    • it changes direction gradually
  3. single-valued
    • for each input (horizontal position), there is only one output (vertical value)
    • if you follow the curve from left to right, it never loops back or reverses direction

The derivative

We minimize the system's error by imagining it as a curve and finding its smallest value.

Maximums and minimums

Global maximum and global minimum are the absolute highest and lowest points along the entire curve:

Global maximum and minimum

But curves can go on forever or repeat, making it hard to identify absolute highs and lows.

To handle this, we look for local maximums and minimums (high or low points near a given location):

Local maximum and minimum

There is only one global maximum and one global minimum, but there can be many local maximums and minimums.

Tangent lines

A tangent line touches a curve at a single point and matches the curve's slope at that point:

Tangent lines

To find a tangent line:

  1. pick two points at equal distances along the curve around a target point
  2. draw a line through them
  3. move them closer together - as they merge, the line becomes the tangent, the best straight-line approximation of the curve at that spot

Finding a tangent line

The derivative is simply the slope of this tangent line.

The rules that said our curves need to be continuous, smooth, and single-valued guarantee that we can always find a tangent line, and thus a derivative, for every point on the curve.

Thinking of a function as a curve, the derivative shows how y changes as x changes:

The derivative

Finding minimums and maximums with derivatives

The derivative at a point can be used to find local maximums and minimums on a curve.

It tells us which direction to move along the curve.

Step Description Image

Find a local maximum

  • find the derivative at an initial point (rightmost point)
  • take a small step along the x axis in the direction of the sign of the derivative
    • e.g., at the starting point the derivative is negative and relatively large
    • so a large step to the left is taken
  • repeat until the derivative = 0
Find a local maximum

Find a local minimum

Do the same, but move opposite the sign of the derivative.

Find a local minimum

Note:

The gradient

The gradient generalizes the derivative into three or more dimensions, allowing us to find maximums and minimums on complex surfaces.

Water, gravity, and the gradient

Imagine a smooth, continuous sheet of fabric.

The surface of this fabric satisfies the rules previously stated:

The surface

If we pour water on it:

Thus:

Finding maximums and minimums with gradients

In a landscape:

Finding maximums and minimums with gradients

At certain special points:

When the gradient is zero, we say it has vanished — no further ascent or descent is possible.

Saddle points

In three dimensions, we encounter saddle points — the surface in the local neighborhood of a point looks like a saddle that horse riders use:

Finding maximums and minimums with gradients

At the center of a saddle:

In deep learning, the gradient represents how to adjust a model to reduce error:

Monitoring the error over time helps detect when progress stalls and guides adjustments to escape zero-gradient regions.

Previous Bayes' rule All ⏎ Next Information theory

A Kemar Joint