Curves and surfaces
Derivatives and gradients are important properties of curves and surfaces.
We can think of the error as a surface. To reduce error, we move toward a minimum on its surface. The steepest downhill direction is given by the negative gradient.
These ideas are at the heart of backpropagation.
The nature of functions
Functions:
- can accept any number of input values, called arguments
- can provide one or more output values, called returned values or returns
Functions are deterministic if no randomness is involved.
Curves and surfaces are graphical versions of mathematical functions.
We assume all our curves and surfaces satisfy certain conditions since these are the kinds most useful in machine learning and optimization:
- continuous
- the curve can be drawn without lifting the pen
- it has no jumps or breaks (no discontinuities)
- smooth
- the curve has no sharp corners or cusps
- it changes direction gradually
- single-valued
- for each input (horizontal position), there is only one output (vertical value)
- if you follow the curve from left to right, it never loops back or reverses direction
The derivative
We minimize the system's error by imagining it as a curve and finding its smallest value.
Maximums and minimums
Global maximum and global minimum are the absolute highest and lowest points along the entire curve:
But curves can go on forever or repeat, making it hard to identify absolute highs and lows.
To handle this, we look for local maximums and minimums (high or low points near a given location):
There is only one global maximum and one global minimum, but there can be many local maximums and minimums.
Tangent lines
A tangent line touches a curve at a single point and matches the curve's slope at that point:
- at local maxima and minima, tangent lines are horizontal (slope = 0)
- tangents help us find these key points
To find a tangent line:
- pick two points at equal distances along the curve around a target point
- draw a line through them
- move them closer together - as they merge, the line becomes the tangent, the best straight-line approximation of the curve at that spot
The derivative is simply the slope of this tangent line.
The rules that said our curves need to be continuous, smooth, and single-valued guarantee that we can always find a tangent line, and thus a derivative, for every point on the curve.
Thinking of a function as a curve, the derivative shows how y changes as x changes:
- if
yincreases asxincreases, the derivative is positive - if
ydecreases asxincreases, the derivative is negative
Finding minimums and maximums with derivatives
The derivative at a point can be used to find local maximums and minimums on a curve.
It tells us which direction to move along the curve.
| Step | Description | Image |
|---|---|---|
|
Find a local maximum |
|
|
|
Find a local minimum |
Do the same, but move opposite the sign of the derivative. |
|
Note:
- the derivative is
0on flat regions, which can fool the algorithm - techniques like momentum help avoid getting stuck at these flat spots
The gradient
The gradient generalizes the derivative into three or more dimensions, allowing us to find maximums and minimums on complex surfaces.
Water, gravity, and the gradient
Imagine a smooth, continuous sheet of fabric.
The surface of this fabric satisfies the rules previously stated:
- it's both smooth and continuous (because it's a single piece of fabric)
- it's single-valued (because the fabric never curls over on itself like a crashing wave)

If we pour water on it:
- the water flows downhill along the steepest path — the direction of maximum descent
- the opposite direction, where we climb most steeply, is the direction of maximum ascent, or the gradient
- the magnitude of the gradient shows how steep the slope is
Thus:
- a hiker seeking the highest peak follows the gradient
- a drop of water seeking the lowest point follows the negative gradient
Finding maximums and minimums with gradients
In a landscape:
- to reach a local maximum, move repeatedly in the direction of the gradient
- to reach a local minimum, move in the opposite direction — the negative gradient

At certain special points:
- very top of the hill (maximum) → no direction to go up → gradient = 0
- valley bottom (minimum) → all directions go up → gradient = 0
- flat plateau or plain → no slope in any direction → gradient = 0
When the gradient is zero, we say it has vanished — no further ascent or descent is possible.
Saddle points
In three dimensions, we encounter saddle points — the surface in the local neighborhood of a point looks like a saddle that horse riders use:

At the center of a saddle:
- the gradient also vanishes, it looks like a plateau, so there's no gradient
- but the point is neither a true maximum nor a minimum
- moving slightly in a new direction reveals curvature, and the gradient reappears, pointing toward the maximum ascent for this point
In deep learning, the gradient represents how to adjust a model to reduce error:
- ideally, training moves toward the true minimum
- but algorithms can get stuck on plateaus, peaks, or saddles where the gradient vanishes
- optimization methods help "unstick" learning, but sometimes the only fix is a major change (e.g., more data or altered parameters)
Monitoring the error over time helps detect when progress stalls and guides adjustments to escape zero-gradient regions.