Deep learning overview
Deep learning refers to machine learning algorithms that use a series of steps (or layers) of computation.
The neuron
The fundamental unit of a neural network is the neuron (or node):
- \( x_0 \), \( x_1 \) and \( x_2 \) are the inputs to the neuron (numbers)
- \( h \) is the node or the neuron:
- it's an activation function (most often the ReLU)
- \( w_0 \), \( w_1 \) and \( w_2 \) are the weights (numbers)
- \( b \) is the bias (a number)
The neuron operates like this:
- multiply every input value, \( x_0 \), \( x_1 \) and \( x_2 \), by its associated weight, \( w_0 \), \( w_1 \) and \( w_2 \)
- add all the products from step 1 together along with the bias value, \( b \)
- give the resulting single number to \( h \) to produce the output (also a single number)
Virtually all the fantastic accomplishments of modern AI are due to this primitive construct.
Neural networks (multi-node models)
Neural networks are collections of individual neurons arranged in layers:
The term deep comes from stacking layers:
- the middle layers are known as hidden layers
- nodes of the hidden layers usually use ReLU activation functions
- the output of the current layer is the input to the following layer
- the topmost node usually uses a sigmoid activation function
We need one weight for each arrow (except the output arrow) and one bias value for each node:
- as the number of nodes in a layer increases, the number of weights increases even faster:
- the 3-node model needs 12 weights and 4 bias values
- the 8-node model needs 32 weights and 9 bias values
- this fact alone restrained neural networks for years (too big for a single computer's memory)
- OpenAI's GPT-3 has over 175 billion weights, rumor puts GPT-4 weights at 1.7 trillion weights
Training a network
General training algorithm:
- select the model's architecture
- number of hidden layers
- nodes per layer
- activation function
- randomly (but intelligently) initialize all the weights and biases associated with the selected architecture
- run the training data through the model
- and calculate the average error (also called the loss)
- this is the forward pass
- determine how much each weight and bias contributes to the average error
- with the backpropagation algorithm
- update the weights and biases
- with the gradient descent algorithm
- this and the previous step make up the backward pass
- repeat from step 3 until the network is considered "good enough" (as close to zero as possible)
Neural networks are randomly initialized
Training a neural network assigns values to the weights and biases by iteratively adjusting them.
How should we pick the initial set of values?
- the answer for our current level of understanding is "at random"
- we roll dice to get the initial value for each weight and bias
- the iterative process then refines these values to arrive at the final set
- however, the iterative process doesn't always end in the same place
- so networks should be trained multiple times (if feasible):
- to gather data on their effectiveness
- or to understand that a bad set of initial values was used purely by chance
At first "at random" meant picking a "small random number" (0.001 or -0.0056) but that didn't work consistently.
Researchers revisited the "small random value" idea:
- three factors need to be considered:
- the form of the activation function
- the number of connections coming from the layer below (fan-in)
- the number of outputs to the layer above (fan-out)
- formulas were devised to use all three factors to select the initial weights for each layer
- bias values are usually initialized to zero
Good enough and overfitting
Sometimes the network is "good enough" because it is overfitting: it's learned all the details of the training data without learning the general trends of the data.
Overfitting is addressed in several ways:
- acquiring more training data
- tweaking the training algorithm to introduce things that keep the network from focusing on irrelevant details
- data augmentation to invent some data by slightly modifying the data you already have
Training a network is a bit of a fluke
Training neural networks is hard because they have very complex error surfaces.
In theory, it shouldn't work this well, but Fortuna smiled on humanity:
- she's given us the ability to train complex models with first-order gradient descent:
- it's a simple method that works best on smooth, simple problems
- yet it works to train complex models because, although there are many local minima (places where training could get "stuck"), they tend to be similarly good in practice
- so getting stuck isn't a big problem, most solutions are "good enough"
- she's arranged things so that the "wrong" gradient direction found by stochastic gradient descent is often what we need:
- using random mini-batches makes training faster at the cost of less accurate gradient estimates
- this means the algorithm often moves in a slightly "wrong" direction
- that sounds bad, but it's actually helpful
- it helps avoid getting stuck too early in the training process and allows better exploration of the error surface