Generative adversarial networks
A Generative Adversarial Network (GAN) consists of two neural networks trained together:
- generator: creates new data samples
- discriminator: tries to tell real data apart from generated data
Through competition, both networks improve, leading the generator to produce highly realistic data.
Generating new data is an important goal in machine learning. GANs support this goal by offering an alternative to autoencoders and can be used to generate many types of data, including images, music, and speech.
GANs principles
Forging money with GANs:
- a generator creates counterfeit bills
- a discriminator determines whether bills are real or fake
- their goal: improve over time, pushing each other to get better
Learning process:
- the generator creates forgeries based on its current knowledge
- early attempts may be random scribbles or drawings
- the discriminator collects real bills and forgeries:
- labels them "real" or "fake"
- shuffles the bills together
- tries to classify them without looking at the labels
There are four possibilities when the discriminator classifies a bill:
- a real bill is correctly labeled real (true positive):
- no learning occurs
- a real bill is incorrectly labeled fake (false negative):
- feedback loop: the discriminator needs to better recognize real bills
- a fake bill is incorrectly labeled real (false positive):
- feedback loop: the discriminator needs to better detect flaws in forgeries
- a fake bill is correctly labeled fake (true negative):
- feedback loop: the generator needs to improve its output
Overall:
- FN and FP cause the discriminator to improve
- TN causes the generator to avoid repeating mistakes
- TP has no effect on either network
Learning from experience
Training starts with both networks untrained, producing mostly random outputs and classifications.
Learning happens by trial and error:
- because the generator never sees the data it is trying to imitate, it can only learn through trial and error
- the system improves gradually:
- the discriminator learns from correctly labeled data and becomes better at spotting fakes
- the generator experiments with variations to fool the discriminator
- each improvement in one network pushes the other to improve, creating a feedback loop
Over time:
- the discriminator becomes highly accurate
- the generator produces outputs that are nearly indistinguishable from real data
Why adversarial?
In GANs, the generator and discriminator are viewed as opponents, not teammates.
The term adversarial comes from game theory:
- a branch of mathematics that studies competition
- the generator and discriminator are like players in a game: one tries to create convincing fakes, while the other tries to catch them
When neither network can improve its performance without the other network changing its strategy, a state of Nash equilibrium is reached.
Implementing GANs
GANs consist of a generator and a discriminator trained together, but after training, only the generator is used.
Depending on context, the term "GAN" can refer to:
- the training method
- the combined network
- the trained generator
Discriminator and generator
|
The discriminator |
The generator |
|---|---|
|
The discriminator is a neural network that evaluates whether a given input is real ( Its architecture can vary freely:
|
The generator is a neural network that creates synthetic data (like images) from random numbers. Its architecture can vary freely:
Sometimes it doesn't have its own loss function:
|
|
|
|
Training
A GAN is trained by alternating updates to the discriminator and the generator during each training round.
Each training round has four steps:
| Step | Explanation | Image |
|---|---|---|
|
Step 1 |
Real data are given to the discriminator:
|
|
|
Step 2 |
Fake data are given to the discriminator:
|
|
|
Step 3 |
Fake data are given to the discriminator:
|
|
|
Step 4 |
Fake data are given to the discriminator (repeat step 2):
|
|
Training alternates between the two networks, ensuring that only one network's weights are updated at any given step.
GANs in action
Imagine the training data as points:

- the 3D view shows the distribution as points belonging to a Gaussian distribution:
- the x and y axes represent the 2D coordinates of the data points
- the height (z-axis) represents the probability density
- the Gaussian distribution is centered at (5, 5), meaning points near (5,5) are the most likely to occur
- the 2D view shows the same distribution but from above:
- the circle represents one standard deviation from the center (5,5), along with some representative points randomly drawn from the distribution
- in a Gaussian distribution, most samples fall within one standard deviation of the center
The generator takes random numbers as input and tries to produce points that look like they were sampled from the Gaussian distribution.
The discriminator tries to tell real points (from the Gaussian) apart from fake points made by the generator.
Determining whether a single point is real or fake is difficult because it could plausibly come from many distributions:

Therefore, the discriminator uses mini-batches, since examining a group of points reveals patterns that make it easier to determine whether they follow the real distribution.
For example, these sets of points produced by the generator clearly do not match the true Gaussian distribution:

Building a discriminator and generator
GANs are hard to train:
- two networks must be tuned together
- GANs are notoriously sensitive to architectural choices and hyperparameters
- even small changes can lead to large performance differences
- effective GAN design usually involves many small experiments on subsets of the data
Because our example dataset is a simple 2D Gaussian distribution, both the generator and discriminator can use simple networks.
A minimal GAN design:
- the generator's output feeds directly into the discriminator
- generator:
- despite its simplicity, the generator must learn to map random noise to points resembling a Gaussian distribution without explicit knowledge of that goal
- discriminator output = a probability indicating how likely the point is real (
1) or fake (0)
Training
The generator and discriminator are trained one at a time, in alternating steps.
When training the generator:
- the discriminator remains in the network, so gradients can flow through it
- the discriminator's weights are frozen, so only the generator learns
- this prevents both models from being trained at once and helps balance their learning
Training proceeds in mini-batches, alternating between updating the discriminator and the generator.
Testing
We trained and tested the simple GAN using:
- training set: 10,000 points drawn from the Gaussian distribution
- mini-batches: 32 points per batch
- epoch: one pass through all 10,000 points
Results over training (epochs 1–13):

- blue points = original dataset
- orange points = points produced by the generator
The generator starts with rough, misaligned points. Its output gradually improves. It produces data closer to the real data's center and shape around epoch 13.
The loss curves suggest the discriminator becomes nearly unable to distinguish real from generated data:
DCGANs
To work with images, convolutional layers are better than dense layers.
GANs built with convolutional layers are called DCGANs (Deep Convolutional Generative Adversarial Networks).
Training a DCGAN for MNIST using Gildenblat's model:
- generator:
- the second dense layer has 6,272 neurons to match the discriminator's final convolution output of shape 7 × 7 × 128
- batch normalization prevents overfitting
- reshape layer converts 1D input into a 3D tensor for the following layers
- uses explicit upsampling layers
- discriminator:
- uses explicit downsampling (pooling) layers
Training setup:
- loss function: binary cross-entropy
- optimizer: Nesterov SGD with learning rate 0.0005 and momentum 0.9
Results:
- after 1 epoch: outputs are mostly unintelligible
- after 100 epochs: generator produces recognizable MNIST digits
DCGANs:
- use convolutional layers for image generation
- carefully match tensor sizes between generator and discriminator
- can learn to create realistic images from random input
- but success depends heavily on design choices and training settings
Challenges
GANs are tricky to train because they are highly sensitive to their structure and hyperparameters:
- the generator and discriminator need to be balanced
- if one improves too quickly, the other can't catch up, and training fails
Finding the right settings is hard, though some general rules can help.
There is currently no proof that GANs will converge. Successful training is based on empirical methods rather than proven outcomes.
Using big samples
Training GANs to generate large images (e.g., 1,000×1,000 pixels) is challenging because it requires massive computing power.
Progressive GANs (ProGANs) make it easier:
- start by training the GAN on very small images (4×4 pixels), then gradually use larger sizes (8×8, 16×16, etc.)
- add more layers to the network as the image size increases
- by scaling up step by step, the GAN learns faster and can generate high-resolution images more effectively
Modal collapse
Sometimes, GANs find a "shortcut" to succeed in training, but in a way that is useless.
For example, if we want a GAN to create cat images, it might just make one cat image over and over. The discriminator thinks it's real, so the generator stops learning.
This problem is called modal collapse:
- full modal collapse: the generator produces only one repeated output
- partial modal collapse: the generator produces a few repeated outputs or minor variations
This happens because neural networks try to satisfy the training goal without creating diverse results.
Ways to fix it:
- use mini-batches of data
- modify the discriminator's loss to penalize lack of diversity
Training with generated data
Using generated data to train other neural networks might seem like a good idea, but it's risky and discouraged.
GANs are imperfect:
- discriminators may fail to catch subtle flaws
- generators can produce incomplete or biased outputs, like focusing on only part of the data
Training on such flawed synthetic data:
- propagates errors and biases into new models
- causes unfair, unsafe, or unreliable results
- is especially dangerous in areas like medicine, finance, or hiring
The key principle is "bias in, bias out".

