Statistics
Statistics turns data into reasoned conclusions. It asks three connected questions:
- How were the data obtained? A biased sample cannot be repaired by elegant calculation.
- What do the data show? Tables, diagrams and summary statistics reveal location, spread and association.
- What can reasonably be inferred? Probability models quantify uncertainty, but conclusions must retain their context and limitations.
The central habit is to distinguish the observed sample from the wider population. A sample statistic such as is known once the data are collected. A population parameter such as is generally unknown and is the subject of inference.
Prerequisites
Section titled “Prerequisites”You should be able to:
- calculate with fractions, percentages, powers and square roots;
- rearrange formulae and use sigma notation ;
- interpret inequalities, coordinates, gradients and areas;
- use combinations and the statistical functions on your calculator;
- communicate conclusions in complete sentences.
Review GCSE statistics and probability foundations if frequency diagrams, averages or basic probability rules are insecure. Calculator output is useful only when you know which quantity you requested and what it means.
The statistical investigation cycle
Section titled “The statistical investigation cycle”A statistical investigation is an iterative process:
The conclusion often exposes a weakness or suggests a new question, so the process may begin again.
1. Define the question and population
Section titled “1. Define the question and population”State what is being measured, in whom or what, and when. “Do pupils sleep enough?” is vague. “What is the mean number of hours slept on school nights by Year 12 pupils at this school this term?” identifies a variable, population and time period.
2. Choose a sample
Section titled “2. Choose a sample”A census observes every member of the population. A sample observes only some members. Sampling is usually quicker and cheaper, but introduces sampling variation and possible bias.
A simple random sample of size gives every possible sample of size an equal chance of selection. Opportunity and voluntary response samples are convenient, but can systematically overrepresent some groups.
Learn to identify sampling units, frames, strata and sources of bias in populations, samples and sampling.
3. Clean and present the data
Section titled “3. Clean and present the data”Check impossible values, inconsistent units, duplicates and missing observations before calculating. An unusual value is not automatically an error. Removing a genuine extreme observation merely because it is inconvenient distorts the evidence.
Choose a display that preserves the important structure:
| Data or purpose | Useful display |
|---|---|
| categorical frequencies | bar chart |
| discrete numerical data | vertical line chart |
| continuous grouped data | histogram |
| cumulative frequencies | cumulative frequency graph |
| compare distributions | box plots |
| relationship between two variables | scatter diagram |
Study data presentation, histograms and outliers and cleaning data.
4. Analyse and interpret
Section titled “4. Analyse and interpret”Use measures of location and spread together. The mean alone does not describe variability, and the standard deviation alone does not locate the data.
For observations ,
The median and interquartile range are resistant to extreme values. The mean and standard deviation use every value, so they are more sensitive to skew and outliers. See averages and measures of spread.
Worked example 1: compare two distributions
Section titled “Worked example 1: compare two distributions”Two groups complete the same task. Their times, in minutes, are summarised below.
| Group | Mean | Median | Standard deviation | Interquartile range |
|---|---|---|---|---|
| A | ||||
| B |
Group A has the lower mean and median, so it was generally faster. Group B has the smaller standard deviation and interquartile range, so its times were more consistent.
For A, , which is evidence consistent with positive skew: a few long times may have pulled the mean upwards. This does not prove the exact shape without seeing the raw data or a diagram.
A strong comparison names the statistics, compares both location and spread, and interprets them in context.
Probability as a model of uncertainty
Section titled “Probability as a model of uncertainty”Probability assigns numbers from to to events. For an event ,
If and are mutually exclusive,
If they are independent,
These statements describe different ideas. Mutually exclusive events cannot occur together. Independent events do not alter one another’s probabilities. Two events with positive probability cannot be both mutually exclusive and independent.
Conditional probability makes the available information explicit:
Build this strand through probability, conditional probability and probability modelling.
Worked example 2: conditional probability and independence
Section titled “Worked example 2: conditional probability and independence”In a year group, pupils study physics, study chemistry, and study both. One of the pupils is selected at random. Let and denote studying physics and chemistry.
The probability that a chemistry pupil also studies physics is
To test independence, compare with :
but
Since , the events are not independent. Notice that the denominator in is the number studying chemistry because the condition restricts the relevant population to that group.
Self check 1
Section titled “Self check 1”Suppose , and . Find and . Are and independent?
Answer
Also,
The events are not independent because
Random variables and distributions
Section titled “Random variables and distributions”A random variable assigns a numerical value to each outcome of a random process. Its probability distribution lists possible values and their probabilities. For a discrete random variable ,
The expected value is the long run mean over many repetitions. It need not be a possible individual outcome. Learn the underlying ideas in discrete random variables.
Two central A level models are:
for a count of successes in independent trials with constant success probability , and
for a continuous, symmetric bell shaped model with mean and variance .
The model must fit the mechanism, not merely the appearance of the numbers. Use the binomial distribution, the normal distribution and choosing a distribution.
Worked example 3: choose and use a binomial model
Section titled “Worked example 3: choose and use a binomial model”A component is defective with probability , independently of other components. Let be the number of defective components among the next .
There is a fixed number of trials, two classifications per trial, independence and constant probability. Therefore
The probability of at least one defective component is most efficiently found by a complement:
The expected number is
This does not mean exactly defects will occur. It is the mean count over many groups of components.
Correlation and regression
Section titled “Correlation and regression”A scatter diagram can show association between paired variables. The product moment correlation coefficient, , measures the strength and direction of a linear relationship:
A value near indicates strong positive linear correlation, a value near strong negative linear correlation, and a value near little linear correlation. It does not rule out a strong curved relationship.
A regression line predicts one variable from another. Use the correct regression direction, interpolate within the observed range where possible, and avoid treating association as causation. A lurking variable may influence both measured variables.
Study correlation and regression before correlation hypothesis tests.
Hypothesis testing
Section titled “Hypothesis testing”A hypothesis test asks whether sample evidence is unusually extreme under a stated null model. The general structure is:
- State hypotheses about a population parameter.
- Assume and identify the sampling distribution.
- Calculate a tail probability or compare with a critical region.
- Reject if the evidence is significant at the stated level.
- Give a contextual conclusion with the correct degree of uncertainty.
The key logical rule is
Worked example 4: interpret a p value
Section titled “Worked example 4: interpret a p value”A test uses
The calculated p value is .
At the significance level,
so reject . There is sufficient evidence at the level that the population proportion exceeds .
At the level,
so do not reject . There is insufficient evidence at the level that the proportion exceeds .
The p value is not the probability that is true. It is a probability about results, calculated under the assumption that is true.
Begin with hypothesis testing language, then study binomial hypothesis tests and normal hypothesis tests.
Self check 2
Section titled “Self check 2”A two tailed test gives p value . State the conclusion at the and significance levels.
Answer
At , , so reject . There is sufficient evidence for the alternative hypothesis at the level.
At , , so do not reject . There is insufficient evidence for the alternative hypothesis at the level.
The contextual wording of the alternative hypothesis should replace the generic phrase “for the alternative hypothesis” in a full solution.
Common misconceptions
Section titled “Common misconceptions”- A large sample removes bias. A large biased sample can estimate the wrong quantity very precisely.
- Correlation proves causation. Association alone cannot establish a causal mechanism.
- Mutually exclusive means independent. For events with positive probability, mutual exclusivity makes one event impossible when the other occurs.
- Expected value is the most likely outcome. It is a probability weighted long run mean and may not be attainable.
- A normal distribution has standard deviation . In , the second parameter is the variance; the standard deviation is .
- A significant result proves the alternative. It provides evidence against at a stated error threshold.
- Calculator output is an interpretation. A decimal needs notation, method, context and an appropriately cautious conclusion.
A route through the topic
Section titled “A route through the topic”Follow this order if you are learning statistics for the first time:
- Populations, samples and sampling
- Data presentation, histograms and averages and measures of spread
- Outliers and cleaning data and the large data set
- Probability, conditional probability and probability modelling
- Discrete random variables, the binomial distribution and the normal distribution
- Correlation and regression
- Hypothesis testing language, followed by the binomial, normal and correlation tests
When solving a mixed question, return to the investigation cycle. Identify the population and variables, inspect how the data were obtained, choose a justified model or summary, calculate accurately, and finish by answering the question in context.