Convolutional neural networks

Convolutional neural networks (CNNs) or convnets are a deep learning technique based on convolution.

CNNs learn to "see" the world of their inputs.

Introducing convolution

Detecting yellow

A convolutional filter (or kernel) is a set of weights used in a CNN to detect a specific feature from an input.

Example - detecting yellow in a color image:

Step Explanation Image

Input

RGB image.

Each pixel has 3 values: R, G, B ∈ [0,1].

Yellow detector (neuron)

Compute a yellowness score per pixel:

  • yellow ≈ high R, high G, low B (1, 1, 0)
  • detection can be done using a neuron with 3 weights: (+1, +1, −1) for (R, G, B)
  • these 3 weights form a kernel (+1, +1, −1)

Kernel size = 1×1×3:

  • size = 1×1 (because it looks at 1 pixel at a time)
  • depth = 3 (because each pixel has 3 channels)

Apply to the whole image

The neuron is applied by moving pixel by pixel, row by row across the image.

At each pixel: output = (+1 × Red) + (+1 × Green) + (−1 × Blue)

This produces a new image (or tensor):

  • same width and height
  • but only one channel

Output

If a pixel is very yellow, the output is large.

In a grayscale image: larger numbers → brighter pixels → closer to white.

So whiter pixels = more yellow in the input.

The sliding (or scanning, or sweeping) operation is called convolution (the mathematical operation).

In CNNs, convolution refers to sliding a learned kernel across an input and computing weighted sums at each position to produce feature maps. This enables detection of specific features (such as yellow).

Weight sharing

Weight sharing means all neurons use the same filter (set of weights) stored in shared memory:

Especially useful in CNNs, since convolutions apply the same filter across many positions in an input.

Larger filters

The area of pixels a filter "sees" is called its local receptive field or footprint:

A filter with a footprint larger than one pixel is called a spatial filter.

When convolving across the image, the center of the filter is called the anchor or reference point.

Filters can be any size, but small odd-numbered squares (1×1 to 9×9) are most common to keep the anchor centered.

Filters and features

Neuroscience suggests that our visual system doesn't just send all visual data to the brain to process.

Instead, some early pattern recognition happens right in the eyes:

In CNNs, filters (or feature detectors) mimic this process by scanning images for features (or patterns) like edges, lines, or textures.

The output of a filter is a feature map, showing how well each part of the image matches the filter.

Example - detecting vertical stripes:

Step Explanation Image

Input

Binary 20×20 image with values 0 (black) and 1 (white).

Filter (Kernel)

3×3 filter designed to detect vertical white stripes, with values -1 (black) and 1 (white).

Apply the kernel

Slide the filter one pixel at a time over the image.

At each position, multiply the filter values by the pixels beneath it and sum the results to produce one single output value.

Result = feature map

Sliding a 3×3 filter across a 20×20 image one pixel at a time produces 18 positions where the filter fully fits (in both width and height), so the resulting feature map has size = 18×18.

The feature map shows match strength with values from -6 to 3 (scaled to [0,1] for display).

Higher (whiter) values = better matches.

Pixels with perfect matches +3 are highlighted with a red border.

Areas of the input where those matches occurred highlighted with a red border.

CNNs learn filter weights automatically during training:

The fact that this process works and achieves high accuracy across many tasks is one of the great discoveries in deep learning.

Padding

When a convolution filter reaches the edge of an image, part of its receptive field can fall off the side. There are no pixels there, but we still need an output value.

Simply avoiding edges reduces the output size and loses information.

Padding solves this:

Padding

Multidimensional convolution

A convolution filter must have the same number of channels as the input tensor it is applied to.

This ensures that every input value is paired with a filter value:

Multidimensional convolution

The process generalizes to tensors with any number of channels.

Convolution layers

A convolution layer consists of multiple filters (often hundreds) applied independently to an input tensor.

Each filter scans the same input independently to detect different features (like vertical lines, horizontal lines, or dots).

Each filter produces a single-channel feature map, and these outputs are stacked together to form the final output tensor.

Multiple filters

Main rules:

Special types of convolution layer

1D convolution

In a 1D convolution layer, multiple filters move in only one direction (height or width).

It's often used in text processing, where data can be represented as a grid of letters or words.

Each filter covers the entire width (or height) of the input and slides along the other dimension.

1D convolution

1x1 convolutions (reducing the number of channels)

1×1 convolutions are used to reduce the number of channels in a tensor while preserving its spatial dimensions (width and height).

It works with multiple 1×1 filters:

The network automatically learns the best way to combine the input channels during training.

Reducing channels from 300 → 175:

1x1 convolution

Benefits:

Changing output size

In addition to the number of channels, width and height of a tensor can also be reduced to save computing resources.

Pooling (downsampling)

Reducing tensor size is a useful side effect of pooling, though its main role is robust feature detection if the input is slightly shifted.

A normal convolution filter can fail when a feature is slightly misaligned

The mathematical way to make the filter "see more" is to make the input image a bit blurry:

The mathematical way to make the filter see more is to make the input image a bit blurry

Pooling is a method for blurring (or downsampling) a tensor by reducing its size.

Two types of pooling:

  1. average pooling:
    • replaces each block with its average value
  2. max pooling:
    • replaces each block with its largest value
    • networks using max pooling tend to learn faster since this preserves the strongest feature in each block
    • "pooling" usually means max pooling

Two types of pooling

Pooling becomes useful when applying multiple convolution layers in succession:

When a tensor has multiple channels, pooling is applied separately to each channel:

Pooling with multiple channels

Pooling makes convolutions shift-invariant because it allows features to be detected even when their positions shift slightly.

Striding (downsampling)

Striding is the process of moving a convolutional filter by more than one pixel at a time (e.g., stride = 2), both horizontally and vertically, which reduces the output size.

Strided convolution is faster than convolution followed by a separate pooling layer because:

  1. fewer filter evaluations
    • the filter is applied fewer times compared to moving one pixel at a time and then pooling
    • this reduces the total number of computations
  2. no separate pooling step
    • strided convolution reduces output size directly, combining convolution and downsampling in one step

Striding

Depending on the stride, some input pixels may be reused, or each pixel may be used only once.

Strided convolution can replace pooling for downsampling, often producing similar results faster, because it avoids extra pooling computations.

However, the learned filters differ, so they can't be swapped one for the other without retraining.

Transposed convolution (upsampling)

Spatial dimensions can also be increased while keeping the same number of channels.

Standard upsampling simply repeats the input values as many times as we request.

Transposed convolution instead works by inserting zeros between and around input elements, then applying a convolution filter:

Transposed convolution

There is a limit to how much we can increase the input size. In practice, models double resolution using a 3×3 transposed convolution.

The output of transposed convolution is not identical to the result of upsampling followed by convolution because the operations behave differently mathematically.

Transposed convolution upsamples feature maps while preserving learned patterns, making it useful for e.g., image generation, segmentation, and decoding in autoencoders.

Hierarchies of filters

Convolutional networks, like biological visual systems, work in hierarchies of processing layers.

Each layer looks for patterns that are a bit more complex than the layer before.

To demonstrate this, we'll build a CNN to recognize simple binary faces with some simplifications:

  1. images and filters are binary (only 0s and 1s)
  2. filters look for exact matches
  3. filters are hand-designed instead of learned
  4. no padding is used (to keep diagrams simple)
  5. use very small images (12×12) (so diagrams fit on the page)

Finding face masks

The recognition task is to determine whether a new image (the candidate) matches a reference mask even if small parts are slightly shifted.

A direct pixel-by-pixel comparison fails when the candidate differs slightly, even though the overall structure of the face is clearly the same:

Pixel-by-pixel comparison

Instead, we can use convolution filters to check for structural similarity by detecting features and their arrangements.

Because it's easier to work backward, we start by designing the final filter:

This face detection filter (labeled F) can support different arrangements. For example, we create another filter (labeled P) that detects a face in profile:

Face detection filters

Finding eyes, noses, and mouths

Because the final filter detects eyes, a nose, and a mouth (3×3×3), the previous layer must produce those features arranged in a 3×3 grid.

For 12×12 input images, one approach is to use large 4×4 filters (E4, N4, M4) to detect eyes, noses, and mouths:

Face detection 4×4 filters

However, this is computationally expensive and inflexible for detecting new features.

A more efficient and flexible approach is to use an additional convolution layer to build the 4×4 filters from smaller 2×2 blocks:

Build the 4×4 filters from smaller 2×2 blocks

Stacking filters

Stacking these filters produces a series of convolution layers where each layer uses the output of the layer below it:

This hierarchical design is efficient (smaller filters reduce computation) and flexible (new features can be added without rebuilding large filters).

Applying filters

The overall result is an all-convolutional network (because all layers use convolution plus pooling):

All-convolutional network

A candidate image is evaluated using three hierarchical layers.

Pooling plays a key role by allowing robustness to small positional changes and reducing computation and memory use.

Step Explanation Image

Layer 1 — detecting basic patterns

Apply 2×2 filters T, Q, L, R to the 12×12 input image → T-map, Q-map, L-map, R-map

  • no padding → convolving 1 pixel at a time produces 11×11 feature maps
  • matches = 1 (green), non-matches = 0 (pink)
Example: T filter → T-map

Apply 2×2 max pooling to each feature map → T-pool, Q-pool, L-pool, R-pool

  • each 2×2 block becomes 1 if there is any green inside
  • output size becomes 6×6
Example: T-map → T-pool

Stack the four pooled maps → 6×6×4 output tensor

Layer 2 — detecting higher-level patterns

Apply filters E, N, M to the 6×6×4 tensor → E-map, N-map, M-map

Apply 2×2 max pooling to each feature map → E-pool, N-pool, M-pool.

Pooling shrinks them to 3×3.

Example: applying 2×2 max pooling to the E-map

Stack the 3 pooled maps → 3×3×3 output tensor.

Layer 3 — detecting whole-face patterns

Apply filters F and P to the 3×3×3 tensor:

  • filters are the same size as the input (3×3×3), so they don't slide
  • direct application produces a 1×1×2 tensor

Interpret the 1×1×2 tensor output

1st channel → result of the F filter (F-out)

  • green → F matches → accept candidate
  • beige → reject candidate

2nd channel → result of the P filter (P-out)

  • green → P matches → accept candidate
  • beige → reject candidate

By stacking convolution layers hierarchically, networks detect increasingly complex patterns:

This also enables flexible pattern recognition: to recognize more face types, we can simply add more filters to the final layer.

There are many CNN architectures

Properly implementing a CNN supporting a multitude of layer types is not trivial.

A few go-to CNN architectures (some with over 100 layers):

Previous Optimizers All ⏎ Next Convnets in practice

A Kemar Joint