Generative adversarial networks

A Generative Adversarial Network (GAN) consists of two neural networks trained together:

  1. generator: creates new data samples
  2. discriminator: tries to tell real data apart from generated data

Through competition, both networks improve, leading the generator to produce highly realistic data.

Generating new data is an important goal in machine learning. GANs support this goal by offering an alternative to autoencoders and can be used to generate many types of data, including images, music, and speech.

GANs principles

Forging money with GANs:

Learning process:

There are four possibilities when the discriminator classifies a bill:

  1. a real bill is correctly labeled real (true positive):
    • no learning occurs
  2. a real bill is incorrectly labeled fake (false negative):
    • feedback loop: the discriminator needs to better recognize real bills
  3. a fake bill is incorrectly labeled real (false positive):
    • feedback loop: the discriminator needs to better detect flaws in forgeries
  4. a fake bill is correctly labeled fake (true negative):
    • feedback loop: the generator needs to improve its output

Overall:

Learning from experience

Training starts with both networks untrained, producing mostly random outputs and classifications.

Learning happens by trial and error:

Over time:

Why adversarial?

In GANs, the generator and discriminator are viewed as opponents, not teammates.

The term adversarial comes from game theory:

When neither network can improve its performance without the other network changing its strategy, a state of Nash equilibrium is reached.

Implementing GANs

GANs consist of a generator and a discriminator trained together, but after training, only the generator is used.

Depending on context, the term "GAN" can refer to:

Discriminator and generator

The discriminator

The generator

The discriminator is a neural network that evaluates whether a given input is real (1) or fake (0).

Its architecture can vary freely:

  • can be shallow or deep
  • can use any type of layers: fully connected, convolutional, recurrent, transformers, etc.

The generator is a neural network that creates synthetic data (like images) from random numbers.

Its architecture can vary freely:

  • can be shallow or deep
  • can use any type of layers: fully connected, convolutional, recurrent, transformers, etc.

Sometimes it doesn't have its own loss function:

  • this is because it doesn't learn on its own
  • its goal is defined entirely in relation to the discriminator
The discriminator The generator

Training

A GAN is trained by alternating updates to the discriminator and the generator during each training round.

Each training round has four steps:

Step Explanation Image

Step 1

Real data are given to the discriminator:

  • the generator is not used
  • if the discriminator labels real data as fake (false negative), an error signal punishes it
  • the error drives a backpropagation step through the discriminator, updating its weight

Step 2

Fake data are given to the discriminator:

  • if the discriminator correctly identifies the fake (true negative), the error is used to update only the generator
  • gradients are computed through both networks, but the discriminator is frozen: its weights are not changed
  • only the generator learns

Step 3

Fake data are given to the discriminator:

  • if it labels them as real (false positive), it is penalized
  • only the discriminator's weights are updated

Step 4

Fake data are given to the discriminator (repeat step 2):

  • if the discriminator catches a fake bill (true negative), the generator learns again
  • step 2 is repeated because the discriminator learns from two types of errors, while the generator learns from only one
  • so repeating this step keeps both networks improving at roughly the same pace

Training alternates between the two networks, ensuring that only one network's weights are updated at any given step.

GANs in action

Imagine the training data as points:

Simple training set

The generator takes random numbers as input and tries to produce points that look like they were sampled from the Gaussian distribution.

The discriminator tries to tell real points (from the Gaussian) apart from fake points made by the generator.

Determining whether a single point is real or fake is difficult because it could plausibly come from many distributions:

Therefore, the discriminator uses mini-batches, since examining a group of points reveals patterns that make it easier to determine whether they follow the real distribution.

For example, these sets of points produced by the generator clearly do not match the true Gaussian distribution:

Building a discriminator and generator

GANs are hard to train:

Because our example dataset is a simple 2D Gaussian distribution, both the generator and discriminator can use simple networks.

A minimal GAN design:

A minimal GAN design

Training

The generator and discriminator are trained one at a time, in alternating steps.

When training the generator:

Training proceeds in mini-batches, alternating between updating the discriminator and the generator.

Testing

We trained and tested the simple GAN using:

Results over training (epochs 1–13):

Results over training

The generator starts with rough, misaligned points. Its output gradually improves. It produces data closer to the real data's center and shape around epoch 13.

The loss curves suggest the discriminator becomes nearly unable to distinguish real from generated data:

The loss curves

DCGANs

To work with images, convolutional layers are better than dense layers.

GANs built with convolutional layers are called DCGANs (Deep Convolutional Generative Adversarial Networks).

Training a DCGAN for MNIST using Gildenblat's model:

DCGAN for MNIST

Training setup:

Results:

DCGANs:

Challenges

GANs are tricky to train because they are highly sensitive to their structure and hyperparameters:

Finding the right settings is hard, though some general rules can help.

There is currently no proof that GANs will converge. Successful training is based on empirical methods rather than proven outcomes.

Using big samples

Training GANs to generate large images (e.g., 1,000×1,000 pixels) is challenging because it requires massive computing power.

Progressive GANs (ProGANs) make it easier:

Modal collapse

Sometimes, GANs find a "shortcut" to succeed in training, but in a way that is useless.

For example, if we want a GAN to create cat images, it might just make one cat image over and over. The discriminator thinks it's real, so the generator stops learning.

This problem is called modal collapse:

This happens because neural networks try to satisfy the training goal without creating diverse results.

Ways to fix it:

Training with generated data

Using generated data to train other neural networks might seem like a good idea, but it's risky and discouraged.

GANs are imperfect:

Training on such flawed synthetic data:

The key principle is "bias in, bias out".

Previous Reinforcement learning All ⏎ Next Large language models

A Kemar Joint