Averages and measures of spread
A data set needs more than one number to describe it. A measure of location says where the data are centred. A measure of spread says how variable they are.
Two classes can have the same mean mark but very different consistency. For example,
both have mean , but the second set is much more spread out. A useful comparison therefore pairs a measure of location with a measure of spread.
Prerequisites
Section titled “Prerequisites”You should be able to:
- order numerical data and use sigma notation ;
- calculate with squares and square roots;
- interpret frequency tables and class intervals;
- use the statistics functions on your calculator.
Review presenting and interpreting data if frequency tables or grouped data are unfamiliar.
Measures of location
Section titled “Measures of location”For observations :
| Measure | Definition | Particularly useful when |
|---|---|---|
| Mean | all values should influence the centre | |
| Median | middle value after ordering | the distribution is skewed or contains extremes |
| Mode | most frequent value | the most common outcome matters |
The mean uses every value, so it is sensitive to extreme observations. The median depends on position, not distance, so it is resistant to extremes. There can be no mode, one mode, or several modes.
Worked example 1: raw data
Section titled “Worked example 1: raw data”The journey times, in minutes, for seven journeys are
First order the data:
The fourth value is the median, so
The value occurs most often, so the mode is . The mean is
The unusually long journey raises the mean above the median. For a typical journey time, the median is more representative here.
Finding the median position
Section titled “Finding the median position”For ordered observations, the median is at position
If this position is an integer, use that observation. If it is halfway between two integers, average the two corresponding observations. For , for example, the median is halfway between the fourth and fifth values.
Measures of spread
Section titled “Measures of spread”Common measures of spread are:
The range uses only the two extreme values. The interquartile range, or IQR, measures the spread of the middle of the ordered data. Both variance and standard deviation use every observation.
Variance has squared units. If is measured in centimetres, the variance is in . Standard deviation has the same units as the data, so it is usually easier to interpret in context.
Why deviations are squared
Section titled “Why deviations are squared”The deviations from the mean always sum to zero:
Simply averaging signed deviations would therefore give zero for every data set. Squaring makes all contributions non-negative and gives greater weight to observations far from the mean.
The efficient variance formula
Section titled “The efficient variance formula”Expanding gives
Hence
This is sometimes called the population standard deviation formula. In A level data questions, use division by unless the question or calculator convention specifies otherwise.
Worked example 2: variance and standard deviation
Section titled “Worked example 2: variance and standard deviation”For the data ,
Therefore
and
The standard deviation is
Do not round the variance before taking its square root. Retaining full calculator accuracy prevents avoidable error.
Self-check 1
Section titled “Self-check 1”Find the mean, variance and standard deviation of .
Answer
Here , and . Thus
and .
Frequency tables
Section titled “Frequency tables”If value occurs with frequency , there are
observations. Repeated values can be handled without writing them all out:
Worked example 3: discrete frequency data
Section titled “Worked example 3: discrete frequency data”The number of goals scored by a team in matches is summarised below.
| Goals, | |||||
|---|---|---|---|---|---|
| Frequency, |
Create the required totals:
| Total |
Then
and
Therefore
For the median, use cumulative frequency. The sixth observation is and the seventh is , so
The median need not itself be an observed value.
Grouped continuous data
Section titled “Grouped continuous data”When only class intervals and frequencies are known, the original values have been lost. Use each class midpoint as a representative value:
These answers are estimates, because the observations are unlikely all to equal their class midpoints.
Worked example 4: estimating from grouped data
Section titled “Worked example 4: estimating from grouped data”The masses of parcels are grouped as follows.
| Mass in kg | Frequency | Midpoint | ||
|---|---|---|---|---|
| Total |
Notice that the last midpoint is , not : the interval has width .
The estimated mean is
The estimated variance is
so the estimated standard deviation is
Self-check 2
Section titled “Self-check 2”Ten values lie in the classes below.
| Class | |||
|---|---|---|---|
| Frequency |
Estimate the mean.
Answer
The midpoints are , and . Therefore
It is an estimate because the exact values within each class are unknown.
Quartiles and the interquartile range
Section titled “Quartiles and the interquartile range”After ordering the data, marks approximately the position, the median marks the position, and marks the position. Then
Different accepted conventions can give slightly different quartiles for a small raw data set. In exam questions, follow the convention indicated by the data presentation or calculator guidance. For cumulative frequency data, quartiles are read at cumulative frequencies , and , usually by interpolation from a graph.
Worked example 5: comparing distributions
Section titled “Worked example 5: comparing distributions”Two machines fill packets. Their masses, in grams, have these summaries.
| Median | IQR | Mean | Standard deviation | |
|---|---|---|---|---|
| Machine A | ||||
| Machine B |
Machine A has the larger typical packet mass because both its median and mean are larger. It is also more consistent because both its IQR and standard deviation are smaller.
A complete contextual comparison names the statistic, gives the direction, and interprets it:
Machine A has the smaller standard deviation, , so its packet masses are less variable.
Do not say that Machine A is “better” without defining what better means. A larger centre could represent better performance in a test, but worse waiting times in a hospital.
Choosing suitable measures
Section titled “Choosing suitable measures”Pair measures that respond similarly to extremes:
| Shape or issue | Suitable pair | Reason |
|---|---|---|
| roughly symmetric data without extremes | mean and standard deviation | both use every observation |
| skewed data or possible outliers | median and IQR | both resist extreme observations |
| categorical or most common outcome | mode | mean and median may be meaningless |
For the ordered salaries
in thousands of pounds, the mean is approximately , which is not representative of most employees. The median is . The large salary is genuine data, but it makes the median and IQR more informative summaries of a typical salary and its spread.
Read outliers and cleaning data before deciding that an unusual observation should be removed.
Linear transformations and coding
Section titled “Linear transformations and coding”Suppose every observation is transformed by
Then
and
Adding shifts every value equally, so it changes location but not spread. Multiplying by scales all distances by .
Worked example 6: changing units
Section titled “Worked example 6: changing units”Temperatures measured in degrees Celsius have mean and standard deviation . Fahrenheit temperature is
Therefore
and
The changes the mean but has no effect on the standard deviation.
Worked example 7: reversing a code
Section titled “Worked example 7: reversing a code”Large values are coded using
The coded data have mean and standard deviation . Since ,
and
Self-check 3
Section titled “Self-check 3”The data have mean and variance . Let . Find the mean, variance and standard deviation of .
Answer
The multiplier is , so
and
A standard deviation cannot be negative.
Missing values from summary statistics
Section titled “Missing values from summary statistics”Summary statistics can be reversed to recover totals:
and
Worked example 8: combining two groups
Section titled “Worked example 8: combining two groups”Group A contains observations with mean . Group B contains observations with mean . The combined mean is not the simple average of and , because the groups have different sizes.
The group totals are
Hence
\bar{x}_{\text{combined}}=rac{90+88}{12+8}=8.9.This is a weighted mean: the larger group contributes more.
Worked example 9: correcting an error
Section titled “Worked example 9: correcting an error”The mean of recorded values was calculated as . One value had been entered as instead of .
The recorded total was
Correcting the error gives
Therefore the corrected mean is
For a corrected variance, both and must be corrected. Replacing by changes the square total by .
Common misconceptions
Section titled “Common misconceptions”- A larger mean does not imply a larger spread.
- Standard deviation is not the average value. It measures typical distance from the mean.
- is the standard deviation. Do not forget the square root.
- Adding a constant does not change standard deviation or IQR.
- Grouped-data calculations are estimates and should be described as such.
- A summary statistic should be interpreted in context, with units where appropriate.
- An outlier is not automatically an error and should not automatically be deleted.
Final self-check
Section titled “Final self-check”A data set has , and .
- Find the mean and standard deviation.
- One further value, , is added. Find the new mean.
- Predict, without calculating the new standard deviation, whether the spread is likely to increase or decrease. Explain.
Answers
-
The mean is
The variance is
so the standard deviation is .
-
The new total is and the new number of observations is . Thus
-
The spread is likely to increase. The added value lies far above both the old mean and the new mean , so it contributes a large squared deviation. A full calculation confirms the increase.
Next steps
Section titled “Next steps”Use these summaries alongside presenting and interpreting data, then study histograms and outliers and cleaning data. Mean and standard deviation later become parameters of the normal distribution.