Deep learning overview

Deep learning refers to machine learning algorithms that use a series of steps (or layers) of computation.

The neuron

The fundamental unit of a neural network is the neuron (or node):

A neuron or node

The neuron operates like this:

  1. multiply every input value, \( x_0 \), \( x_1 \) and \( x_2 \), by its associated weight, \( w_0 \), \( w_1 \) and \( w_2 \)
  2. add all the products from step 1 together along with the bias value, \( b \)
  3. give the resulting single number to \( h \) to produce the output (also a single number)

Virtually all the fantastic accomplishments of modern AI are due to this primitive construct.

Neural networks (multi-node models)

Neural networks are collections of individual neurons arranged in layers:

A simple deep network

The term deep comes from stacking layers:

We need one weight for each arrow (except the output arrow) and one bias value for each node:

Training a network

General training algorithm:

  1. select the model's architecture
    • number of hidden layers
    • nodes per layer
    • activation function
  2. randomly (but intelligently) initialize all the weights and biases associated with the selected architecture
  3. run the training data through the model
    • and calculate the average error (also called the loss)
    • this is the forward pass
  4. determine how much each weight and bias contributes to the average error
    • with the backpropagation algorithm
  5. update the weights and biases
    • with the gradient descent algorithm
    • this and the previous step make up the backward pass
  6. repeat from step 3 until the network is considered "good enough" (as close to zero as possible)

Neural networks are randomly initialized

Training a neural network assigns values to the weights and biases by iteratively adjusting them.

How should we pick the initial set of values?

At first "at random" meant picking a "small random number" (0.001 or -0.0056) but that didn't work consistently.

Researchers revisited the "small random value" idea:

Good enough and overfitting

Sometimes the network is "good enough" because it is overfitting: it's learned all the details of the training data without learning the general trends of the data.

Overfitting is addressed in several ways:

Training a network is a bit of a fluke

Training neural networks is hard because they have very complex error surfaces.

In theory, it shouldn't work this well, but Fortuna smiled on humanity:

  1. she's given us the ability to train complex models with first-order gradient descent:
    • it's a simple method that works best on smooth, simple problems
    • yet it works to train complex models because, although there are many local minima (places where training could get "stuck"), they tend to be similarly good in practice
    • so getting stuck isn't a big problem, most solutions are "good enough"
  2. she's arranged things so that the "wrong" gradient direction found by stochastic gradient descent is often what we need:
    • using random mini-batches makes training faster at the cost of less accurate gradient estimates
    • this means the algorithm often moves in a slightly "wrong" direction
    • that sounds bad, but it's actually helpful
    • it helps avoid getting stuck too early in the training process and allows better exploration of the error surface

Previous AI overview All ⏎ Next AI history

A Kemar Joint