Statistics foundations: averages, spread and data diagrams
Statistics turns data into evidence. A calculation alone is not a conclusion: you must know what was measured, how the data were collected, what a diagram represents and what uncertainty remains.
This lesson secures the Higher-tier GCSE statistics needed at A-level. The central habits are to match the method to the data, distinguish exact values from estimates, and compare distributions using both their centre and their spread.
Prerequisites
Section titled “Prerequisites”You should be able to:
- calculate with fractions, decimals and percentages;
- substitute into a formula;
- find the midpoint of an interval;
- read scales and coordinates;
- calculate areas of rectangles.
Review exact arithmetic, ratio and proportion or coordinate geometry if needed.
Statistical language
Section titled “Statistical language”The population is the complete set of people or objects of interest. A sample is a subset from which data are collected. A variable is the characteristic recorded.
| Type | Meaning | Examples |
|---|---|---|
| qualitative | labels or categories | eye colour, transport used |
| quantitative | numerical values | height, number of siblings |
| discrete | takes separate, countable values | goals scored, shoe size |
| continuous | can take any value in an interval | time, mass, temperature |
A continuous measurement is recorded to finite accuracy. A mass written as kg to the nearest kilogram represents
The upper endpoint is excluded because kg would round to kg.
Worked example 1: classify variables
Section titled “Worked example 1: classify variables”For each variable, identify its type.
- A student’s blood group is qualitative.
- The number of emails received in a day is quantitative and discrete.
- The time taken to run m is quantitative and continuous.
Decimals do not decide whether data are continuous. Shoe sizes such as belong to a fixed set, so shoe size is discrete.
Populations, samples and bias
Section titled “Populations, samples and bias”A census collects data from every member of the population. It avoids sampling variation, but may be expensive, slow or impossible. A well-designed sample is usually faster and may permit more careful measurement.
A sample should represent the population relevant to the question. Common methods include:
- simple random sampling, where every population member has an equal chance of selection;
- systematic sampling, where every th member is selected after a random starting point;
- stratified sampling, where each relevant group is represented in population proportion.
For a stratum of size in a population of size , a sample of total size should contain
members from that stratum, rounded sensibly while keeping the final sample size equal to .
Worked example 2: stratified sampling
Section titled “Worked example 2: stratified sampling”A college has students, of whom are in Year 12. A stratified sample of students is required.
The Year 12 allocation is
Therefore select Year 12 students, randomly within that group.
Bias and questionnaire design
Section titled “Bias and questionnaire design”A large sample can still be biased. Asking only members of a school athletics club about weekly exercise under-represents less active students. An online voluntary poll may attract people with unusually strong opinions.
Questions should be neutral, precise and use non-overlapping response intervals. Compare:
Do you agree that the excellent new timetable should continue?
with:
Should the new timetable continue? Yes / No / Unsure
Intervals such as and overlap at . Use boundaries such as and .
Self-check 1
Section titled “Self-check 1”- State the population when bulbs from a day’s factory output are tested for lifetime.
- A school has pupils in Key Stage 3 and in Key Stage 4. Find the Key Stage 4 allocation in a stratified sample of .
- Explain one problem with asking shoppers leaving a sports shop whether the town needs more sports facilities.
Answers
- All bulbs produced by the factory that day.
- pupils.
- The location creates selection bias because sports-shop customers are likely to be more interested in sport than the town population. Other precise, relevant explanations are possible.
Frequency tables and averages
Section titled “Frequency tables and averages”For values , the mean is
In a frequency table, a value occurring times contributes to the total:
The median is the middle value after ordering the data. The mode is the most frequent value or category. The range is
Worked example 3: an ungrouped frequency table
Section titled “Worked example 3: an ungrouped frequency table”The numbers of books read by students are summarised below.
| Books, | |||||
|---|---|---|---|---|---|
| Frequency, | |||||
Hence
For values, the median lies halfway between the th and th ordered values. Cumulative frequencies are , so both positions contain . Thus the median is . The greatest frequency is , so the mode is also .
Choosing an average
Section titled “Choosing an average”- The mean uses every value, but is affected by extreme values.
- The median is resistant to extreme values and is often better for skewed data.
- The mode is the only one of these suitable for qualitative data.
Suppose five salaries, in thousands of pounds, are
Their mean is , but their median is . The mean is correct, yet is not typical of four of the five employees. Context decides which summary is useful.
Grouped data and estimates
Section titled “Grouped data and estimates”When values are grouped into intervals, their exact values are lost. To estimate the mean, represent every value in a class by its midpoint.
Worked example 4: estimate a grouped mean
Section titled “Worked example 4: estimate a grouped mean”| Time in minutes | Frequency, | Midpoint, | |
|---|---|---|---|
Therefore
The symbol matters. We do not know where the observations lie within each interval. The estimate assumes that each class can be represented by its midpoint.
The class containing the largest frequency is the modal class. Here it is , with frequency . This does not prove that is the mode.
Misconception: comparing raw frequencies
Section titled “Misconception: comparing raw frequencies”The final two classes both have width , but the first two have width . When class widths differ, frequency alone does not describe concentration. Histograms use frequency density.
Quartiles and spread
Section titled “Quartiles and spread”Quartiles split ordered data into four parts:
- lower quartile ;
- median ;
- upper quartile .
The interquartile range is
It measures the spread of the middle and is less affected by extremes than the range.
Worked example 5: median and quartiles
Section titled “Worked example 5: median and quartiles”Consider the ordered data
There are values, so
The lower half is , whose median is , so . The upper half is , whose median is , so . Therefore
Different calculator and examination conventions can locate quartiles slightly differently. If a question provides a convention or a cumulative frequency graph, use that convention consistently.
Outliers
Section titled “Outliers”A common rule identifies a value as an outlier if it is below
or above
For the data above, the fences are
Thus is not an outlier by this rule, although it is relatively large. An outlier rule flags observations for investigation. It does not automatically justify deleting them.
Cumulative frequency
Section titled “Cumulative frequency”Cumulative frequency is a running total. Plot cumulative frequency against the upper class boundary, because the total answers a question of the form “how many observations are below this boundary?”
For the grouped times in Worked example 4:
| Upper boundary | ||||
|---|---|---|---|---|
| Cumulative frequency |
The graph should also begin at here. Join the points with a smooth increasing curve.
Worked example 6: estimate quartiles from cumulative frequency
Section titled “Worked example 6: estimate quartiles from cumulative frequency”There are observations, so read from the graph at cumulative frequencies
Using straight-line interpolation within the relevant classes:
- lies in . It is of the observations into that class, so
- the median lies in . It is of the observations into that class, so
- is of the observations into the same class, so
Hence
These are estimates because grouped data do not reveal the positions within each class.
Box plots
Section titled “Box plots”A box plot displays the five-number summary:
The box runs from to , a line marks the median, and whiskers extend towards the extremes. Modified box plots may show outliers separately, so check the stated convention.
Box plots are especially useful for comparing distributions on the same scale. A complete comparison discusses:
- location, usually the medians;
- spread, usually the IQRs or ranges;
- the context and units.
Worked example 7: compare distributions
Section titled “Worked example 7: compare distributions”Two delivery services have these summaries, in minutes.
| Service | Median | IQR | Range |
|---|---|---|---|
| A | |||
| B |
Service B is typically faster because its median delivery time is lower. Service A is more consistent through the middle half because its IQR is smaller. However, Service A has the larger overall range, perhaps because of an extreme delay.
Do not write only “A is better”. That depends on whether a customer values a lower typical time or greater consistency.
Histograms and frequency density
Section titled “Histograms and frequency density”A histogram represents continuous grouped data. The area of each bar is proportional to frequency. Therefore
and consequently
Unlike a bar chart, a histogram has a continuous numerical horizontal scale and adjacent class bars touch.
Worked example 8: construct histogram heights
Section titled “Worked example 8: construct histogram heights”For the time data from Worked example 4:
| Interval | Frequency | Width | Frequency density |
|---|---|---|---|
Although the third class has the greatest frequency, the second has the tallest bar because its data are more concentrated per minute.
Worked example 9: recover a frequency
Section titled “Worked example 9: recover a frequency”A histogram bar covers and has frequency density . Its frequency is
If a vertical axis is marked only with an arbitrary scale, use a known bar to determine the area-to-frequency scale before finding other frequencies.
Scatter diagrams and correlation
Section titled “Scatter diagrams and correlation”A scatter diagram plots paired observations to investigate association between two quantitative variables.
- positive correlation: larger tends to accompany larger ;
- negative correlation: larger tends to accompany smaller ;
- no correlation: no clear linear trend.
A line of best fit follows the centre of a roughly linear cloud, with points balanced above and below. It can estimate from .
Interpolation estimates within the observed data range. Extrapolation estimates beyond it and is less reliable because the relationship may change.
Misconception: correlation proves causation
Section titled “Misconception: correlation proves causation”Correlation does not by itself show that changing one variable causes the other to change. Ice-cream sales and sunburn cases may rise together because both are influenced by warm weather. A third variable is sometimes called a lurking or confounding variable.
Self-check 2
Section titled “Self-check 2”- Find the mean and range of .
- A grouped class has frequency . Find its midpoint and frequency density.
- For and , find the IQR and the upper outlier fence.
- Dataset P has median and IQR . Dataset Q has median and IQR . Compare them.
- Explain why predicting a person’s mass at age from a best-fit line based only on children aged to is unsafe.
Answers
- and the range is .
- Midpoint . Class width , so frequency density .
- . The upper fence is .
- Q has the higher median, so its values are typically larger. P has the smaller IQR, so its middle half is less spread out. A contextual conclusion would require units and knowledge of what is being measured.
- This is extreme extrapolation beyond the observed age range. Growth is unlikely to continue according to the same linear relationship into adulthood.
Exam habits and common errors
Section titled “Exam habits and common errors”- Order raw data before finding the median or quartiles.
- Divide by , not by the number of rows, when using a frequency table.
- Use class midpoints only for estimating a grouped mean.
- Plot cumulative frequency at upper class boundaries.
- Use frequency density, not frequency, for histogram height when widths differ.
- Compare both centre and spread, using values and units.
- Write for estimates from grouped data or graphs.
- Treat an outlier as a point to investigate, not an automatic error.
- Distinguish association from causation and interpolation from extrapolation.
What you should now be able to do
Section titled “What you should now be able to do”You should be able to:
- classify data and identify populations, samples and possible bias;
- calculate and interpret the mean, median, mode, range and IQR;
- estimate statistics from grouped data;
- construct and interpret cumulative frequency graphs, box plots and histograms;
- compare distributions with evidence;
- interpret scatter diagrams without claiming unjustified causation.
Next, deepen these ideas in populations and sampling, presenting and interpreting data, measures of location and spread, histograms and correlation and regression. Probability is the other major foundation for A-level statistics, so continue with probability foundations.