Autoencoders

Autoencoders are neural networks that learn to compress input data into a smaller set of latent variables through a bottleneck, and then reconstruct it.

This bottleneck forces lossy compression, since some information is discarded.

Using only the decoder part, the model can generate new data.

Different architectures, including fully connected and convolutional layers, affect the structure and quality of latent representations.

Common uses:

Introduction to encoding

Encoding and decoding:

Lossless vs. lossy encoding:

Blending representations

Blending is a way of combining numerical representations of multiple inputs to create new outputs that contain aspects of each input.

Two main approaches to blending data:

Content blending

Parametric blending

  • mixes raw data directly rather than underlying structure
  • example: mixing images of a cow and a zebra
  • result: a simple superposition of both images, not a meaningful "zebracow" hybrid
  • mixes the parameters that describe the underlying structure of objects
  • example: blending the center, radius, and color of two circles
  • result: an "in-between" circle with blended properties, something meaningful
Content blending Parametric blending

However, parametric blending does not work well with arbitrary compressed representations, because compression can destroy the internal structure needed for meaningful interpolation.

Autoencoders solve this problem:

The simplest autoencoder

An autoencoder is a type of neural network that compresses data and then reconstructs it:

The simplest autoencoder

This network is trained so that the output matches the input.

Since the input itself is the target, this is called semi-supervised learning (no labels, just the input data).

If trained on a single image (like a tiger):

Training on a larger set of images reveals the bottleneck's limitations:

A better autoencoder

To work well, an autoencoder needs:

  1. enough latent dimensions to represent the data
  2. enough computational capacity (neurons and layers) to encode and decode data effectively

Below, we compare two autoencoder architectures using the MNIST dataset (60,000 labeled 28×28 grayscale images of handwritten digits).

Both models are trained for 50 epochs (= we run through all 60,000 samples 50 times).

Architecture Explanation Image

Shallow autoencoder

  • input: 28×28 image flattened into 784 values
  • encoder: one fully connected "bottleneck" layer reducing 784 → 20 latent variables
  • decoder: one fully connected layer expanding 20 → 784
  • (flattening layer at the start and unflattening layer at the end are omitted in the diagrams)
  • result: reconstructed images are blurry but still recognizable
Shallow autoencoder

Deep autoencoder

  • input: 28×28 image flattened into 784 values
  • encoder (3 fully connected layers):
    • 784 → 512 latent variables
    • 512 → 256 latent variables
    • 256 → 20 latent variables
  • decoder (3 fully connected layers):
    • 20 → 256 numbers
    • 256 → 512 numbers
    • 512 → 784 numbers
  • (these sizes commonly halve and then double, though this is a guideline, not a rule)
  • result: produces much clearer reconstructions than the shallow 20-latent version
Deep autoencoder

Exploring the autoencoder

A closer look at the latent variables

Latent variables are compressed representations of input data that encode the essential information needed to reconstruct inputs.

The network learns its own private encoding that makes little sense to humans.

For example, using the deep autoencoder seen above:

This lack of interpretability is not a problem, because the latent variables serve as an internal code optimized for accurate reconstruction.

The parameter space

Latent variables are high-dimensional and abstract.

To understand what they represent, we can reduce the encoder to compress each input into just two numbers that capture its key features, allowing us to plot and visualize them as (x, y) points.

Plotting 10,000 MNIST images onto a 2D plane shows the latent variable space (or latent space):

Latent space

Decoding these 2D points back into images reveals what each region represents:

Blending latent variables

Blending latent variables means mixing the compressed representations of two inputs so the decoder can create smooth intermediate outputs.

Intermediate outputs reveal relationships between points in the latent space.

Having more latent variables makes blended results more detailed and realistic.

Example using the 6-layer deep autoencoder:

Results:

Blending latent variables

Limitations:

Predicting from novel input

Using the deep autoencoder trained on MNIST digits, we compress and decompress a tiger image, resized to 28×28 pixels to fit the network:

Predicting from novel input

This shows that autoencoders only work well on data similar to what they were trained on, because their latent variables are designed to represent that specific type of data.

Convolutional autoencoders

Since convolutional layers are well-suited for image data, they provide a natural way to build an autoencoder for MNIST.

Convolutional autoencoder

Training this model for 50 epochs produces reconstructions that are very close to the original digits.

Blending the latent variables of two images, however, gives mixed results. The intermediate images are often messy because the network has never seen "between" digits during training.

The same limitation appears with out-of-distribution inputs. When given the low-resolution tiger image, the model cannot map it meaningfully into the digit-based latent space, so the output is not tiger-like.

Like simpler autoencoders, convolutional autoencoders cannot generalize to new inputs or latent-variable combinations that lie outside the data they were trained on.

Denoising

A popular use of autoencoders is to remove noise from images by learning to reconstruct clean data.

In this example, noise is added to images from the MNIST dataset:

The autoencoder is trained using noisy images as input and clean images as targets.

The model learns a latent representation that captures the underlying structure of the digits while ignoring random noise.

Two architectures are tested.

Model 1 - convolutional autoencoder with explicit downsampling and upsampling layers:

Model 2 - simplified architecture using only convolution layers:

Both architectures work well for denoising, though determining which is better would require more detailed evaluation.

Variational autoencoders

Traditional autoencoders suffer from an unstructured latent space, where regions may be mixed or empty, leading to poor-quality outputs.

A Variational Autoencoder (VAE) addresses this by organizing the latent space more effectively:

VAEs are nondeterministic, introducing randomness during encoding, so they don't always produce the same output for the same input.

Training

A Variational Autoencoder (VAE) is designed to generate realistic outputs from random latent variables by organizing the latent space during training.

To do this, it enforces three key properties:

  1. all latent variables occupy a defined region (so we know where to sample when generating new data)
  2. similar inputs produce latent variables that cluster together
  3. empty regions in latent space are minimized

This structure is achieved using a specialized loss function that penalizes violations of these rules.

Clustering the latent variables

The specialized error term of VAEs encourages the values of each latent variable to follow a unit Gaussian distribution:

In practice, perfect Gaussians are rarely achieved:

Clumping digits together

A VAE learns to group similar inputs in latent space.

For example:

A high-dimensional latent space enables the model to group together all variations of these features.

The VAE discovers these features and groups the images automatically (no manual labeling needed).

In latent space, distance between latent vectors (points in multidimensional space) reflects how much two images resemble each other.

Training must balance two competing goals:

To balance these goals, VAEs introduce randomness, letting the model "usually" cluster similar inputs while still keeping the Gaussian structure.

Introducing randomness

To generate varied outputs from similar inputs, a VAE introduces randomness into the latent variables produced by the encoder.

A simple idea:

VAEs solve this using the reparameterization trick:

Implementation requires splitting since the encoder must output two parameters (μ and σ) for each latent variable:

The reparameterization trick

This is why trained VAEs are nondeterministic and give slightly different results every time: the encoder is deterministic up to the mean and standard deviation, but the sampling step introduces variation.

Exploring the VAE

A fully connected VAE is similar to a standard autoencoder with two key enhancements:

A fully connected variational autoencoder

Working with the MNIST samples

Using the above VAE to reconstruct MNIST digits demonstrates how the model encodes and generates data:

Input and output of the VAE

The VAE produces good matches of input digits, but because it samples from a latent distribution, the output varies for the same input:

The VAE produces different outputs for the same input

The latent space is stable even when we add noise to test its structure:

Blending the latent variables of two images is used to test the continuity of latent space:

Out-of-distribution inputs don't produce meaningful results:

Working with two latent variables

We trained a VAE with only 2 latent variables (instead of 20) to better visualize how it organizes the latent space.

Plotted latent variables for 10,000 MNIST images:

Latent variables for 10,000 MNIST images

Decoding the latent space:

Decoding the latent space

Even with just two latent variables, the digits are well-clumped and structured, much better than earlier visualizations. Some images are slightly fuzzy, but most still look digit-like.

Producing new input

A VAE can generate new images by:

The number of latent variables affects the output:

Once trained, the decoder can be used on its own as a generator to create unlimited new images.

This approach works for any type of data the VAE has been trained on.

Previous Convnets in practice All ⏎ Next Recurrent neural networks

A Kemar Joint