Training and testing

Training is the process of starting with a system that has default or random settings and gradually improving it.

Testing is used afterward to estimate how well the trained system performs on new, unseen data.

Training

In supervised learning:

Training process:

Training process

As this predict–compare–update loop runs, the classifier's internal variables are nudged toward values that improve its prediction accuracy.

Completing one full pass through the entire training set is called an epoch, and training usually involves many epochs.

Training continues as long as the system is still learning and improving its performance on the training data.

After training, the next step is to evaluate the classifier's accuracy.

Testing the performance

Training

Systems are trained on labeled training data, but strong performance on this data does not guarantee good real-world performance due to overfitting.

There is no formula that can guarantee how well a trained model will perform in the real world.

To estimate real-world performance, we must test systems through experiments on data that goes beyond the training set.

Test data

To estimate real-world performance, we must use unseen data called the test set (or test data).

A model is trained using training data, then evaluated once on the test data to estimate real-world performance.

Guidelines:

How to avoid leakage:

Test data is created by splitting the original dataset, commonly around 75% for training and 25% for testing.

Validation data

Up to this point:

That strategy is slow.

We want a rough estimate of the system's performance as we go along.

To make this estimate, we split the input data into three sets:

  1. 60 % training set
  2. 20 % test set
  3. 20 % validation set (chunk of data that's meant to be a good proxy of the real world data)

Updated workflow with validation:

  1. train the system on the training set for one epoch
  2. evaluate its performance on the validation set after each epoch
  3. use this feedback to decide to stop training or to adjust hyperparameters (learning rate, model complexity, etc.)
  4. evaluate the system once on the test set

Always reserve the test set for final evaluation to prevent overestimating the model due to subtle data contamination.

Cross-validation

Cross-validation (or rotation validation) is a technique used when datasets are small and rare (e.g., Pluto photos), so every sample is precious.

Instead of permanently splitting the dataset, the model is trained and tested multiple times on different temporary splits of the data.

The core idea is to run a loop:

After all iterations, all the recorded scores are averaged to get an overall estimate of the model's performance.

Estimates are less reliable than those from a dedicated test set, but worth it when data is scarce.

This algorithm avoids data leakage because each iteration trains a fresh model on one subset of data and evaluates it on a separate, unseen subset.

k-fold cross-validation

K-fold cross-validation is a variant of cross-validation where the data is split into k equal-sized groups, called folds.

One smaller group is allowed if the data can't be split evenly.

Each sample belongs to exactly one fold.

For example, in 5-fold cross-validation:

k-fold cross-validation

The process can be repeated or randomized for more robust results.

Previous Classification All ⏎ Next Overfitting and underfitting

A Kemar Joint