Essential statistics

Averages

Common types of averages of a list of numbers:

  1. mean
  2. mode
    • the value that occurs the most often in the list
    • if no value occurs more often than any other, the list has no mode
  3. median
    • the number in the middle of the sorted list

While averages summarize data, they don't reveal how values are distributed.

Random variables and probability distributions

A probability distribution tells us how likely different outcomes are in a random process.

If we assign a category to 950 cars (sedan, pickup, minivan, SUV, wagon):

Available cars as probability distribution

A random variable is a function associated with a probability distribution:

Selecting a value from a distribution is called drawing a value from the random variable.

Probability distributions are classified into two main categories:

  1. continuous probability distribution
    • used when the outcomes can take any real value in a range (e.g., the amount of oil in a car)
    • also called a probability density function (pdf)
  2. discrete probability distribution
    • used when the outcomes are countable and can only take a fixed number of possible values (e.g., vehicle types)
    • also called a probability mass function (pmf)

In practice, random values are generated by selecting a probability distribution and drawing a random variable from it.

Continuous distributions

The uniform distribution

Every value within a specific range is equally likely to be selected.

The most basic version is the uniform distribution from 0 to 1 (by convention here, a closed circle = the point is included):

Uniform distribution

Two key features:

  1. fixed range:
    • we get 0 everywhere except between values 0 and 1
    • the distribution is called finite because it has a specific minimum and maximum
  2. uniform likelihood:
    • every value in the range 0 to 1 is equally probable
    • this is why the graph is described as uniform, flat, or constant in the range 0 to 1

The normal distribution

Also called the Gaussian distribution or bell curve.

Normal distribution

A normal distribution is defined by two numbers: the mean and the standard deviation.

Term Explanation Image

Mean (average)

Determines the center location of the peak.

It's also the median and the mode (a special property of the normal distribution).

Standard deviation σ (écart type)

σ (lowercase sigma) is a unit of distance (like a step size) that measures how spread out the data are around the mean.

To find it:

  • start at the center of the peak
  • move outward symmetrically until you enclose about 68 % of the area under the curve
  • the distance from the center to either end of this interval is one standard deviation

From −σ to +σ = two steps total distance, but still called within 1 standard deviation because we're only going −1σ away from the center in each direction.

Extending further:

  • by another standard deviation encloses 95 % of the area under the curve
  • by another standard deviation encloses 99.7 % of the area under the curve
  • this is known as the three-sigma-rule or 68-95-99.7 rule

For 1,000 samples drawn from a normal distribution:

  • about 680 are within ±1σ (one standard deviation)
  • about 950 are within ±2σ (two standard deviations)
  • about 997 are within ±3σ (three standard deviations)

Sometimes, instead of standard deviation, the variance is used for convenience:

The normal distribution appears frequently because it naturally describes many real-world observations.

Discrete distributions

The Bernoulli distribution

Simple discrete distribution that produces only two possible outcomes: 0 and 1.

A common example is a coin flip:

Bernoulli distribution

The Multinoulli distribution (or categorical distribution)

Extends the Bernoulli distribution for more than two outcomes, where an event will have one of K possible outcomes (e.g., a 20-sided die).

This kind of distribution is important for training classifiers that sort inputs into many classes.

The output is a list where all entries are 0 except for the chosen category, which is 1.

For example:

Multinoulli distribution

Expectation, dependence, and independence

Expected value

The average (or mean) of a list of numbers is called the expected value.

This is useful in many situations:

Dependence

Independent variables don't depend on each other.

Dependent variables do depend on each other:

Independent and identically distributed variables

Values called i.i.d. (independent and identically distributed) are:

Sampling and replacement

It's useful to build new datasets from existing ones by randomly selecting some of the elements of the existing set.

Selection with replacement (SWR)

We make a copy of each item we select, so the original stays in the dataset:

Selection with replacement

This is called selection with replacement because we can think of it as:

  1. removing the object
  2. making a copy
  3. replacing the original

Selection without replacement (SWOR)

We remove the item we pick from the original dataset:

Selection without replacement

Bootstrapping

Bootstrapping (or bagging, or resampling) is a statistical technique used to estimate properties of a large population by working with smaller, manageable samples.

Bootstrapping involves two steps:

  1. create a sample set from the population using sampling without replacement (SWOR)
  2. resample that sample set many times with replacement (SWR) to create multiple new samples called bootstraps

Example: estimating the average height of babies

Histogram of the 1,000 bootstrap means:

Histogram of bootstrap means

Bootstrapping is useful:

Covariance and correlation

Variables can have relationships with each other:

Identifying and measuring these relationships is useful:

To measure these relationships, we use covariance and correlation.

Covariance

Two variables covary when one variable's value increases or decreases and the other does the same by a fixed multiple of that amount.

A number called the covariance measures:

The sign of the covariance tells us the direction of the relationship:

Covariance

Correlation

Because of the way covariance is defined mathematically, it doesn't take into account relationships between the units of the two variables.

The correlation coefficient (or correlation) solves this problem to let us make these comparisons:

Correlation

When two variables have a perfect positive or negative correlation (+1 or -1), the variables are linearly correlated.

Correlation is more useful because it isn't affected by the scale of the numbers.

Statistics don't tell us everything

The Anscombe's quartet is a famous example of how statistics can be misleading:

We shouldn't assume that statistics tell us everything about any data set.

Whenever we work with new data, it's important to:

High-dimensional spaces

A single sample (or piece of data) often contains many numbers called features. For example, a piece of fruit might be described by its weight, color, and size.

Each sample can be thought of as a point in a space, where each feature corresponds to one dimension:

Many datasets have far more features. For instance, a 1,000 x 1,000 grayscale image has 1,000,000 pixel values → a million dimensions.

We cannot visualize or imagine such high-dimensional spaces, but the mathematics and algorithms can handle such high-dimensional spaces.

While intuition from lower dimensions helps, high-dimensional spaces can behave in unexpected ways.

Previous AI history All ⏎ Next Measuring performance

A Kemar Joint