Presenting and interpreting data
Good statistical diagrams do more than display numbers. They reveal the shape, centre, spread and unusual features of a distribution. The correct display depends on what the variable means and how the data were recorded.
Prerequisites
Section titled “Prerequisites”You should be able to:
- order positive and negative numbers;
- read scales and plot coordinates;
- calculate percentages;
- find a midpoint and interpret inequalities such as .
This lesson introduces the representations needed before averages and measures of spread and histograms.
Classifying data
Section titled “Classifying data”A variable is a characteristic recorded for each individual or item.
| Type | Meaning | Examples |
|---|---|---|
| Qualitative | labels or categories | eye colour, method of travel |
| Quantitative | numerical values with meaningful arithmetic | height, number of calls |
| Discrete | separate, usually countable values | number of siblings, shoe size |
| Continuous | any value in an interval is possible | mass, time, temperature |
Quantitative data may be discrete or continuous. Qualitative data are not continuous merely because their categories have been assigned numbers. For example, coding bus, car and walk as does not make the mean code meaningful.
Data can also be described by their source:
- primary data are collected first hand for the investigation;
- secondary data were collected previously by somebody else.
And by how observations are paired:
- univariate data record one variable per individual;
- bivariate data record two variables on the same individual, such as height and arm span.
The pairing in bivariate data matters. A scatter diagram must use matched pairs , not two independently sorted lists. See correlation and regression for bivariate displays.
Self-check 1
Section titled “Self-check 1”Classify each variable as qualitative, discrete quantitative or continuous quantitative.
- The time taken to complete a race.
- The number of emails received in a day.
- A customer’s satisfaction category: poor, fair, good or excellent.
Answer
- Continuous quantitative. Time can, in principle, take any value in an interval, even if the stopwatch rounds it.
- Discrete quantitative. Emails are counted in whole numbers.
- Qualitative. The categories are ordered, but the gaps between them have no numerical size.
Choosing a presentation
Section titled “Choosing a presentation”The display must preserve the feature you want to study.
| Situation | Suitable presentation | Main strength |
|---|---|---|
| Categories | bar chart | compares category frequencies |
| Small numerical data set | ordered list or stem and leaf diagram | retains every value |
| Discrete values with repetitions | frequency table or vertical line chart | shows exact frequencies |
| Grouped continuous data | histogram | compares frequency density |
| Ordered numerical data | cumulative frequency graph | estimates median, quartiles and percentiles |
| Comparing distributions | box plots on a common scale | compares centre, spread and skewness |
In a bar chart, bars are separated because categories are distinct. In a histogram, class intervals form a continuous scale, so bars touch and area, not height alone, represents frequency. That distinction is developed in histograms.
Frequency tables and class intervals
Section titled “Frequency tables and class intervals”A frequency table compresses repeated observations. Its frequencies must satisfy
where is the total number of observations. The relative frequency of a class is
and its percentage frequency is .
Worked example 1: constructing classes
Section titled “Worked example 1: constructing classes”The journey times, in minutes, are
Using classes of width gives:
| Time in minutes | Frequency | Cumulative frequency |
|---|---|---|
The convention includes but excludes . Every possible value belongs to exactly one class, so there are no gaps or overlaps.
The cumulative frequency means that journeys took less than minutes. It does not mean that journeys lie in the class .
Grouping loses detail. From the table we can tell that one journey took between and minutes, but not that its time was minutes.
Stem and leaf diagrams
Section titled “Stem and leaf diagrams”A stem and leaf diagram is a compact ordered list. It shows the distribution while retaining each observation. It needs:
- leaves in ascending order;
- every observation, including repeated values;
- a key that fixes the place value.
Worked example 2: constructing a stem and leaf diagram
Section titled “Worked example 2: constructing a stem and leaf diagram”For the data
use the tens digits as stems and sort the units digits within each row:
1 | 82 | 1 2 2 6 73 | 0 5 94 | 1Key: represents .
There are leaves, confirming that all observations are present. The repeated leaf on stem is essential because occurs twice.
The ordered values can now be read directly. The fifth and sixth are and , so the median is
Back to back stem and leaf diagrams
Section titled “Back to back stem and leaf diagrams”Two small data sets can share stems. Leaves for one set appear on the left and leaves for the other on the right. Left hand leaves are written so that values increase away from the stem to the left, while right hand leaves increase away from the stem to the right.
Group A Group B 8 4 1 | 1 | 2 5 99 7 3 0 | 2 | 1 1 6 8 5 2 | 3 | 0 4Key: represents in Group A and in Group B.
Always make the key unambiguous. A key such as differs by a factor of from .
Self-check 2
Section titled “Self-check 2”The leaves on stem are , and the key says represents . Write the four values and state their mode.
Answer
The values are . The mode is .
Cumulative frequency graphs
Section titled “Cumulative frequency graphs”For grouped data, cumulative frequency is plotted against the upper class boundary. The points represent statements of the form
Include the lower boundary of the first class with cumulative frequency , then join the points with a smooth increasing curve or straight segments, according to the expected convention.
Worked example 3: plotting cumulative frequency
Section titled “Worked example 3: plotting cumulative frequency”For the table in Worked example 1, plot
Notice that uses the upper boundary and cumulative frequency . Plotting the class midpoint would change the meaning and give incorrect quartile estimates.
For observations, common A level graph estimates use cumulative frequency levels
Draw horizontally from each frequency level to the curve, then vertically to the data axis. Because grouped data do not reveal the positions within a class, graphically read quartiles are estimates.
Worked example 4: interpolating a median
Section titled “Worked example 4: interpolating a median”Suppose values are grouped and the cumulative frequencies at and are and . The median is at cumulative frequency
Assuming values are evenly distributed through this class, the required fraction across it is
Therefore the estimated median is
This is linear interpolation. It is an estimate based on an even spread within the class, not a claim that an observed value was exactly .
Percentiles and proportions
Section titled “Percentiles and proportions”The th percentile is the value below which approximately of observations lie. For , the th percentile is read at cumulative frequency
To estimate the number above a value , read the cumulative frequency and calculate
Self-check 3
Section titled “Self-check 3”A cumulative frequency curve represents measurements. At , the cumulative frequency is approximately .
- Estimate how many measurements are below .
- Estimate how many are at least .
- At what cumulative frequency should be read?
Answer
- Approximately .
- Approximately .
- At .
Box plots
Section titled “Box plots”A box plot displays the five number summary:
The box runs from to , with a line at the median. Whiskers extend to the extremes, unless a question uses a stated outlier convention. The box length is the interquartile range
Worked example 5: drawing and interpreting a box plot
Section titled “Worked example 5: drawing and interpreting a box plot”A data set has
Draw all five values on one linear scale. The box extends from to , its median line is at , and the whiskers end at and .
The IQR is
Each interval from the minimum to , from to the median, from the median to , and from to the maximum contains approximately of the data. Their unequal physical lengths show differing concentration. The interval to is relatively long, so that quarter of the observations is more spread out.
The longer upper half of the box and longer upper whisker suggest positive skew. This is evidence about shape, not a proof of an exact model.
Outliers on box plots
Section titled “Outliers on box plots”A common rule labels values as outliers if
If this rule is specified, whiskers usually end at the most extreme non-outlying observations and outliers are plotted separately. Do not assume a convention that the question has not given. Learn how to investigate suspicious observations in outliers and cleaning data.
For the values above, the fences are
Neither nor is an outlier under this rule.
Comparing distributions
Section titled “Comparing distributions”A complete comparison uses context and normally addresses both:
- location, usually the median;
- spread, usually the IQR or range.
Shape and outliers may add useful evidence. Use comparative language such as greater, smaller, more consistent or more positively skewed.
Worked example 6: writing a comparison
Section titled “Worked example 6: writing a comparison”Two delivery services have these summaries, in minutes:
| Service | Median | IQR | Range |
|---|---|---|---|
| A | |||
| B |
A strong interpretation is:
Service A is typically faster because its median delivery time is minutes lower. Service B is more consistent because both its IQR and range are smaller.
Do not write only “A is better”. Faster and more consistent are different properties, and the data do not say which matters more to a customer.
If two box plots are being compared, they must use the same scale. Otherwise physical box lengths are not comparable.
Common misconceptions
Section titled “Common misconceptions”- A cumulative frequency point goes at the midpoint. It goes at the upper class boundary because it counts all observations below that boundary.
- The four quarters of a box plot have equal lengths. They contain roughly equal numbers of observations, but their lengths show spread.
- A large range proves most data are variable. The range may be caused by one extreme value. The IQR describes the middle .
- Joining bar chart bars makes a histogram. A histogram uses a continuous scale and frequency density, particularly when class widths differ.
- A precise calculator answer makes grouped data exact. Grouped representations have already lost information, so interpolated statistics remain estimates.
Final self-check
Section titled “Final self-check”The five number summaries for test scores are:
- Find the IQR for each class.
- Which class has the higher typical score?
- Which class is more consistent by IQR?
- What proportion of Class P scored between and ?
Answer
- Class P: . Class Q: .
- Class P, because its median is , compared with for Class Q.
- Class Q, because its IQR is smaller.
- Approximately , because and .
A concise comparison is: Class P achieved a slightly higher typical score, but Class Q’s scores were more consistent.
Next steps
Section titled “Next steps”- Use summaries from data displays in averages and measures of spread.
- Learn why unequal class widths require histograms.
- Decide how data should be gathered in sampling.
- Study paired variables in correlation and regression.