Neural networks

Deep learning algorithms are built from networks of connected computational units called artificial neurons, which are inspired by the neurons in the human brain.

Neurons can be arranged into layers to form deep learning networks.

Biological neurons

A biological neuron:

A simplified biological neuron (shown in red below) sends its output to another neuron (in blue):

A simplified biological neuron

Neuronal connections:

Connectomes:

Scientists have tried to emulate the brain by creating large networks of simplified neurons in hardware or software.

So far, this has not produced human-like intelligence, but simplified neuron models can still solve many real-world problems effectively.

Artificial neurons

Artificial neurons in machine learning are basic versions of biological neurons.

They are improved versions of the original perceptron (1957).

An artificial neuron

Drawing the neurons

Diagrams are often simplified to stay clear and easy to read:

A simplified artificial neuron

Even if not shown: biases and weights are always part of the neuron and are the only parameters that change during training.

Naming weights:

Named weights

Feed-forward networks

In deep learning, neurons are arranged in layers:

The most common way to organize layers is the feed-forward network:

Neural network graphs

Neural networks are often shown as directed graphs:

Directed graph

How data flows:

No loops or cycles:

Training and reverse flow (feed-backward):

Initializing the weights

Training a neural network means adjusting its weights to improve performance.

How we initialize the weights helps the network learn faster and avoid issues like vanishing or exploding gradients during training.

Common initialization methods derived from theoretical research:

Deep networks

The phrase deep learning comes from stacking many layers vertically:

The depth of a network is the number of layers that contain neurons (excluding the support layer).

Diagrams may show the data flow vertically or left-to-right.

A simple 3-layer network:

A simple 3-layer network

Fully connected layers

A fully connected layer (also called an FC, linear, or dense layer) connects every neuron to all neurons in the previous layer.

Fully connected layers

Other layer types include convolutional and pooling layers, which organize neurons differently.

Tensors

Neural network layer outputs are often stored as arrays of numbers called tensors.

A tensor is a structured, multidimensional block of numbers with no gaps or extra elements.

Its shape is defined by the number of dimensions and the size of each.

Example of 1D, 2D, and 3D tensors represented as list, grid, and volume respectively:

1D, 2D, and 3D tensors represented as list, grid, and volume

Preventing network collapse

Activation functions are essential in neural networks.

Without activation functions:

This phenomenon is called a network collapse: neurons can collapse (or combine) into a single neuron.

Activation functions prevent this by introducing nonlinearity, allowing the network to model more complex relationships.

Activation functions

An activation function takes a floating-point number as input and returns a new floating-point number as output.

It can be visualized as a graph:

Each neuron could use a different function, but in practice the same one is used for all neurons within a layer.

Straight-line functions

Straight-line functions are linear functions because they involve only addition and multiplication.

They can be represented as straight lines.

Linear functions can cause network collapse but are useful in two cases:

  1. on output neurons: safe from collapse because no layers follow
  2. as a processing step between the summation step in a neuron and its nonlinear activation function to pass values through without change
Name Description Image

Identity function

The output always equals the input (y = x).

Identity function

Linear function

Example 1.

Linear function 1

Linear function

Example 2 tilted to a different slope.

Linear function 2

Step functions

Step functions are a variation on linear functions.

Name Description Image

Stair-step function

  • The curve remains single-valued (each input x has only one output y).
  • The output stays constant over intervals and then jumps to a new value at certain points.
  • Here, a filled circle indicates that the y value is included.
Stair-step function

Step function

One constant value on one side of the threshold and another value on the other side.

Step function

Unit step function

0 to the left of the threshold, 1 to the right.

Unit step function

Heaviside step function

  • A unit step where the threshold is 0.
  • The "sign function" is a variation of this one where values to the left of 0 are -1 (rather than 0) and values to the right are 1.
  • The original perceptron used a heaviside step function as its activation function.
Heaviside step function

Piecewise linear functions (ReLU)

Piecewise linear functions are made up of multiple straight-line segments.

Name Description Image

ReLU

  • ReLU (e is lowercase) = Rectified Linear Unit.
  • Named after electronic rectifiers that block negative voltage.
  • A ReLU does the same for numbers:
    • outputs 0 for negative inputs
    • outputs the input itself for non-negative values
  • Although it consists of two straight lines, the "bend" at 0 makes it nonlinear.

Drawback: negative inputs always output 0, which can halt learning in a network.

To address this, several ReLU variations have been developed.

ReLU function

Leaky ReLU

Outputs a small scaled-down value for negative inputs, instead of 0.

Leaky ReLU

Shifted ReLU

Moves the bend of the ReLU function down and left.

Shifted ReLU

Other variations:

Smooth functions (Sigmoid)

Training a neural network requires computing derivatives of each neuron's output, which necessarily involves their activation functions.

Earlier activation functions (like ReLU) are made of straight-line segments and have sharp corners where the derivative doesn't exist.

Even so, ReLU works because we can mathematically approximate the derivative at those corners, which is why ReLU remains popular.

An alternative approach is to use smooth activation functions:

Example of smooth functions:

Name Description Image

Softplus

  • Smooth approximation of ReLU.
  • Avoids the sharp bend at zero.
Softplus

Exponential ReLU or ELU

  • Smooth version of Shifted ReLU.
  • Allows negative output values.
Exponential ReLU

Swish

  • Another way to smooth out the ReLU.
  • Adds a small bump just left of 0.
  • Transitions more gradually.
Swish

Sigmoid (or logistic function/curve)

  • Smooth version of the Heaviside Step.
  • Output range [0, 1] (0 for very negative inputs, 1 for very positive).
  • The name "sigmoid" comes from the resemblance of the curve to an "S" shape.
Sigmoid

Hyperbolic tangent (tanh)

  • Similar to sigmoid but output range [-1, 1].
  • More centered around 0.
Hyperbolic tangent and sigmoid

Sine wave

  • Output range: [-1, 1] (like tanh).
  • Unlike tanh/sigmoid: it doesn't saturate (or stop changing) for inputs that are far from 0.
Sine wave

Comparing activation functions

There's no universal rule for choosing the best activation function.

It's based on experience and experimentation.

A few rules of thumb:

Softmax

Softmax is a function applied to the output layer of classification neural networks when there are two or more output neurons.

It converts raw output scores from the network into probabilities:

Softmax

Softmax in action

Previous Ensembles All ⏎ Next Backpropagation

A Kemar Joint