Notes/Mathematics/Paper 5/Representation of Data
CAIEA Level9709§5.1

Representation of Data

Choosing a diagram; drawing and reading stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs; medians, quartiles and fair comparisons; and the mean and standard deviation from raw data, grouped data, coded totals and two combined data sets.

300 min read 13 sub-topics
128
question parts
2021–2025 · 37 papers
9 marks
per paper
≈ 19% of the paper
2.0/3
avg difficulty
moderate
#5
most examined
of 5 topics by marks

Statistics starts where the data arrives: a list of numbers that means nothing until it is summarised. This topic teaches the two kinds of summary that every later statistics topic relies on. One is a picture: a stem-and-leaf diagram, a box-and-whisker plot, a histogram or a cumulative frequency graph. The other is a pair of numbers, a centre and a spread: the median with the interquartile range, or the mean with the standard deviation.

Across 2021–2025 this topic carried 343 marks over 128 tagged parts, about 9.3 of the 50 marks on every Paper 5. Every one of the 37 papers in that window had a question on it, and 10 papers had two. It ranks fifth of the five Paper 5 topics by marks, but Paper 5 is unusually flat, so fifth still means roughly a fifth of the paper:

topicmarks/paper
Discrete Random Variables13.2
The Normal Distribution11.6
Probability11.2
Permutations and Combinations9.8
Representation of Data9.3

It is also the easiest-rated topic on the paper (mean difficulty 2.052.05 on the bank's 11–44 scale), and it usually comes early: 36 of its 47 questions in the window were Question 1, 2, 3 or 4. The marks go to things that can be checked: a key on a stem-and-leaf diagram, bars at the right class boundaries, a midpoint instead of a class limit, a comparison written in context.

Inside the topic, the bank splits the marks five ways (a part can carry more than one tag, so the rows overlap):

sub-topicpartsmarkssections
Stem-and-leaf diagrams, box-and-whisker plots, histograms, cumulative frequency graphs4615502, 04, 06, 08
Measures of central tendency and variation5210803, 05, 07
Mean and standard deviation from raw, grouped or coded data3410710–13
Using cumulative frequency graphs (percentiles and proportions)193409
Selecting and evaluating statistical diagrams51201

The question shapes repeat almost exactly from year to year. A typical question gives two teams' times, asks for a back-to-back stem-and-leaf diagram, the median and interquartile range of one team, a pair of box-and-whisker plots, and then two comparisons. Another gives a grouped table and asks for a histogram or a cumulative frequency graph, a reading off that graph, and estimates of the mean and standard deviation. A third gives summary totals such as ∑(x−k)\sum(x-k) and ∑(x−k)2\sum(x-k)^2, or the totals of two separate groups, and asks for a mean, a standard deviation or an unknown constant. Each shape has its own sections below.

Before you start you should be able to
  • Putting a list of numbers in order, and finding the middle value of an ordered list

  • Working out a mean: add the values and divide by how many there are

  • Reading a value off a graph with a linear scale, and drawing axes with a sensible linear scale

  • Substituting numbers into a formula, and rearranging a simple formula to make a different letter the subject

  • Nothing from the other Paper 5 topics: this is the first topic, and it starts from scratch

By the end of this page you can
  • Choose a suitable diagram for a data set and state, in one sentence, an advantage or disadvantage that follows from what the diagram keeps or loses

  • Draw a stem-and-leaf diagram, single or back-to-back, with ordered and aligned leaves and a key that names both data sets and the units

  • Find the median, quartiles and interquartile range from an ordered list or from a printed stem-and-leaf diagram, reading the left-hand side of a back-to-back diagram correctly

  • Draw a box-and-whisker plot, or a pair of them on one scale, from five key values, and apply an outlier rule when a question defines one

  • Write comparisons of two data sets in context, one about the centre and one about the spread, and say whether the mean or the median is more suitable, and why

  • Find class boundaries for rounded, truncated and inequality-defined classes; draw a histogram using frequency density; and read frequencies back off a histogram

  • Find which class contains the median or a quartile, and the greatest or least possible interquartile range, from a grouped frequency table

  • Draw a cumulative frequency graph from a frequency table or a cumulative frequency table, starting at the correct point

  • Use a cumulative frequency graph to estimate medians, quartiles, percentiles and the number or percentage of values above, below or between given values, and turn it into a box-and-whisker plot

  • Calculate the mean and standard deviation from raw data or from ∑x\sum x and ∑x2\sum x^2, find a missing value from a mean, and recover ∑x2\sum x^2 from a mean and a standard deviation

  • Calculate estimates of the mean and standard deviation of grouped data using class midpoints, including when the table is a cumulative frequency table

  • Use coded totals ∑(x−a)\sum(x-a) and ∑(x−a)2\sum(x-a)^2 to find a mean, a variance, the constant aa or the number of values nn

  • Combine two data sets by adding their totals, and work backwards from a combined mean or standard deviation to a missing total

01

Choosing a representation

Syllabus requirement · §5.1

“

select a suitable way of presenting raw statistical data, and discuss advantages and/or disadvantages that particular representations may have.

”

This topic uses four diagrams. Each one is taught in full later in the note (§02, §04, §06 and §08); this section is about choosing between them, which the exam tests in one- and two-mark parts.

Every diagram throws some information away. That is its purpose: 150 raw numbers are unreadable, while a picture of their shape is not. So the question "which diagram?" is really "which information can I afford to lose?", and the answer follows from what each diagram keeps:

  • A stem-and-leaf diagram lists every value, sorted into rows. Nothing is lost: every original value can be read back.
  • A box-and-whisker plot is drawn from just five numbers: the smallest value, the lower quartile, the median, the upper quartile and the largest value. All other values are lost.
  • A histogram is drawn from a grouped frequency table, which sorts the values into classes (intervals such as 20⩽t<3020 \leqslant t < 30) and records only how many values, the frequency, fall in each class. The histogram shows these frequencies as areas. The individual values inside each class are lost.
  • A cumulative frequency graph is also drawn from a grouped table. It shows running totals (the number of values up to each point), which makes it the best diagram for estimating medians, quartiles and percentiles (§09). Individual values are lost.

Diagram

Keeps

Loses

Good for

Stem-and-leaf

every original value

nothing

small data sets (up to about 30 values); exact median and quartiles; comparing two small sets back-to-back

Box-and-whisker

the five key values

every other value

comparing the centre and spread of two or more sets at a glance

Histogram

class frequencies, as areas

the values inside each class

large data sets; showing the shape of the distribution, even with unequal classes

Cumulative frequency graph

running totals at class boundaries

the values inside each class

large data sets; estimating the median, quartiles, percentiles, and how many values lie above or below a value

The stem-and-leaf diagram is the only one that keeps the raw data, which is why it is the standard answer to “state an advantage over a box-and-whisker plot”. Its disadvantage is the mirror image: with hundreds of values it becomes impractical to draw or read.

How these one-mark answers are marked

The answer must say what the diagram does with the data, not that it is "clearer" or "easier". Mark schemes accept answers like these:

  • Advantage of a stem-and-leaf diagram over a box-and-whisker plot: "it shows all the original data values", or "further statistics such as the mean or the mode can be found from it".
  • Disadvantage of a box-and-whisker plot compared with a stem-and-leaf diagram: "it does not show the individual data values".
  • Advantage of a pair of box-and-whisker plots over two cumulative frequency graphs: "you can see at a glance which group is generally larger, and which is more spread out".
  • Suitable diagram for a large data set: "a histogram, because there are too many values for a stem-and-leaf diagram; it shows the shape of the distribution".

One trap: saying a stem-and-leaf diagram lets you find the median, the interquartile range or the range does not score as an advantage over a box-and-whisker plot. A box plot shows those too. The advantage must be something only the full data gives.

How mark schemes label their marks

This note often says which mark a step earns, using the codes printed in Cambridge mark schemes:

  • B1: a mark for a correct value or statement on its own;
  • M1: a method mark, for a correct method even if the arithmetic then slips;
  • A1: an accuracy mark for a correct answer, which needs the M1 before it;
  • DM1: a method mark that depends on an earlier method mark;
  • SC: a special-case mark, a partial mark for a particular incomplete answer;
  • CAO: correct answer only; "condone": accepted although not ideal; FT ("follow through"): marked using your own earlier answer.

Choosing a diagram for three situations

For each situation, name a suitable diagram and give a reason.

(a) A swimming coach has the times of the 1414 swimmers in her squad and wants to display them so that each individual time can still be seen.

(b) A council has the ages of 24002400 residents, grouped into classes of different widths, and wants to show the shape of the age distribution.

(c) A teacher wants to compare the spread of the marks of two classes, each of 3030 students, at a glance.

Show full working
  1. 1

    (a) Count the values and note what is wanted. There are only 1414 values, and the coach wants every individual time to stay visible.

    The size of the data set and the purpose are the two facts that decide the diagram. Write them down first.

  2. 2

    (a) Choose the diagram that keeps every value. A stem-and-leaf diagram lists all 1414 times, so each individual time can still be read, and 1414 values is few enough to draw easily.

    The reason names what the diagram keeps (every value), which is what the mark is for.

  3. 3

    (b) Count the values and note what is wanted. There are 24002400 values, already grouped into classes of different widths, and the aim is the shape of the distribution.

    2400 leaves could never be drawn, so any diagram that lists individual values is ruled out straight away.

  4. 4

    (b) Choose the diagram built for grouped data with unequal classes. A histogram: each bar's area represents the frequency of its class, so the shape is shown correctly even though the class widths differ.

    A bar chart of frequencies would exaggerate the wide classes. The histogram's use of area is exactly what unequal classes need (§06).

  5. 5

    (c) Note what is wanted. The teacher wants to compare spread, for two groups, at a glance.

    Comparison of two groups is the key word. The diagram should put both groups side by side on one scale.

  6. 6

    (c) Choose the diagram built for comparisons. A pair of box-and-whisker plots drawn on the same scale: the widths of the two boxes (the interquartile ranges) and the lengths of the whiskers show at once which class's marks are more spread out.

    A back-to-back stem-and-leaf diagram would also compare two groups, and would be acceptable. The box plots win on 'at a glance' because spread is shown directly as a length.

Answer

(a) A stem-and-leaf diagram: it shows every individual time, and 14 values is a small data set. (b) A histogram: it shows the shape of a large, grouped data set, and its areas handle the unequal class widths. (c) A pair of box-and-whisker plots on the same scale: they show the spread of each class (box width and whiskers) side by side.

Name the diagram, then give a reason that says what the diagram keeps or shows. A reason about neatness or preference scores nothing.

Advantage of a stem-and-leaf diagram over a box plot

9709/52 M/J 2021 Q7(a)1 mark

The heights, in cm, of the 1111 basketball players in each of two clubs, the Amazons and the Giants, are shown below.

Amazons205198181182190215201178202196184
Giants175182184187189192193195195195204

State an advantage of using a stem-and-leaf diagram compared to a box-and-whisker plot to illustrate this information.

Show full working
  1. 1

    Say what each diagram keeps. A stem-and-leaf diagram of these heights would list all 2222 heights. A box-and-whisker plot would keep only five values for each club: the shortest, the lower quartile, the median, the upper quartile and the tallest.

    The advantage has to come from this difference. Everything else about the two diagrams is the same.

  2. 2

    Rule out the answers that do not score. The median, the quartiles, the interquartile range and the range can all be read from a box-and-whisker plot as well, so "you can find the median" is not an advantage.

    The mark scheme says explicitly that median, IQR, range or spread, 'which can be found from both', score nothing.

  3. 3

    State the advantage in one sentence. The stem-and-leaf diagram shows all the original data values, so every individual height can be seen, and further statistics such as the mean or the mode can be worked out from it.

    Either half of this sentence earns the mark: 'includes all the raw data', or 'further statistics, such as the mean, mode or standard deviation, can be found'.

Answer

A stem-and-leaf diagram shows all the original data (every individual height), so, for example, the mean or the mode could also be found; a box-and-whisker plot shows only five values.

If your answer would also be true of a box-and-whisker plot, it is not an advantage over one.

Your turn

Each answer is a sentence or two. Before revealing a solution, check that your reason says what the diagram keeps or loses.

  1. 1

    A factory records the masses of 500500 bags of flour. Give one disadvantage of representing these masses with a stem-and-leaf diagram.

    Stuck? Show hint

    Think about how many leaves the diagram would need.

    Show solution
    1. 1

      Count what the diagram would need. A stem-and-leaf diagram shows every value as a leaf, so this one would need 500500 leaves.

      The feature that is an advantage for a small data set (it keeps every value) becomes the problem for a large one.

    2. 2

      State the disadvantage. With 500500 values the diagram would be impractical: very long rows, slow to draw and hard to read.

      A grouped diagram (histogram or cumulative frequency graph) would be the sensible alternative.

    Answer

    With 500 values a stem-and-leaf diagram would be impractical: far too many leaves to draw or read clearly.

  2. 29709/61 O/N 2013 Q4(iii)1 mark

    The following are the house prices in thousands of dollars, arranged in ascending order, for 5151 houses from a certain area.

    253 270 310 354 386 428 433 468 472 477 485 520 520 524 526 531 535
    536 538 541 543 546 548 549 551 554 572 583 590 605 614 638 649 652
    666 670 682 684 690 710 725 726 731 734 745 760 800 854 863 957 986

    Give one disadvantage of using a box-and-whisker plot rather than a stem-and-leaf diagram to represent this set of data.

    Stuck? Show hint

    A box-and-whisker plot is drawn from only five of these 51 prices.

    Show solution
    1. 1

      Say what the box plot keeps. A box-and-whisker plot of these prices is drawn from five values only: the lowest price, the lower quartile, the median, the upper quartile and the highest price.

      The disadvantage is whatever the box plot throws away.

    2. 2

      Name what is lost. The other 4646 individual prices cannot be seen on it, whereas a stem-and-leaf diagram would show every one of the 5151 prices.

      The mark scheme asks to see 'individual items' (or equivalent) in the answer.

    Answer

    A box-and-whisker plot does not show the individual house prices; a stem-and-leaf diagram shows every one of them.

  3. 39709/63 M/J 2012 Q1(i)(ii)4 marks

    Ashfaq and Kuljit have done a school statistics project on the prices of a particular model of headphones for MP3 players. Ashfaq collected prices from 2121 shops. Kuljit used the internet to collect prices from 163163 websites.

    (i) Name a suitable statistical diagram for Ashfaq to represent his data, together with a reason for choosing this particular diagram.

    (ii) Name a suitable statistical diagram for Kuljit to represent her data, together with a reason for choosing this particular diagram.

    Stuck? Show hint

    21 values is a small data set; 163 is too many to list one by one.

    Show solution
    1. 1

      (i) Size of Ashfaq's data set. 2121 prices is a small data set, small enough to write out every value.

      Small data sets suit the diagram that keeps every value.

    2. 2

      (i) Diagram and reason. A stem-and-leaf diagram, because it shows all 2121 prices while still showing the shape and spread of the data.

      The mark scheme also accepted a box-and-whisker plot with a reason such as 'shows the spread' or 'the median can be read'. The reason mark depends on naming an acceptable diagram first.

    3. 3

      (ii) Size of Kuljit's data set. 163163 prices is a large data set: a stem-and-leaf diagram with 163163 leaves would be impractical.

      Too many values to list is itself the reason for grouping them.

    4. 4

      (ii) Diagram and reason. A histogram, because it groups the 163163 prices into classes and shows the shape and spread of the distribution (and the modal class, the class with the highest frequency density).

      The mark scheme also accepted a cumulative frequency graph (it easily gives the median) or a box-and-whisker plot (it shows the spread), each with a matching reason.

    Answer

    (i) A stem-and-leaf diagram: the data set is small, and it shows every price as well as the shape (a box-and-whisker plot showing the spread was also accepted). (ii) A histogram: there are too many prices to list, and it shows the shape and spread of the grouped data (a cumulative frequency graph or box-and-whisker plot with a matching reason was also accepted).

  4. 49709/63 O/N 2016 Q5(iii)1 mark

    The tables summarise the heights, hh cm, of 6060 girls and 6060 boys.

    Height of girls (cm)140<h≤150140 < h \le 150150<h≤160150 < h \le 160160<h≤170160 < h \le 170170<h≤180170 < h \le 180180<h≤190180 < h \le 190
    Frequency122117100
    Height of boys (cm)140<h≤150140 < h \le 150150<h≤160150 < h \le 160160<h≤170160 < h \le 170170<h≤180170 < h \le 180180<h≤190180 < h \le 190
    Frequency02023125

    (In an earlier part of the question, cumulative frequency graphs of both sets of heights were drawn on the same axes.)

    The students are asked to compare the heights of the girls and the boys. State one advantage of using a pair of box-and-whisker plots instead of the cumulative frequency graphs to do this.

    Stuck? Show hint

    What can you see directly on two box plots drawn on one scale that you would have to work out from two curves?

    Show solution
    1. 1

      Say what a pair of box plots shows directly. Drawn on one scale, the two boxes show each group's median and interquartile range as positions and lengths, so the comparison can be seen without reading anything off a curve.

      On cumulative frequency graphs the medians and quartiles have to be found by reading across and down first.

    2. 2

      Write the advantage in context. You can see at a glance which group is taller on the whole (whose box sits further right), and which group's heights are more spread out (whose box and whiskers are longer).

      The mark scheme accepted any sensible comment in context, such as 'can see which is taller' or 'can see which of boys or girls is more spread out'.

    Answer

    The box plots show the medians and spreads directly side by side, so you can see at a glance which group is generally taller and whose heights are more spread out.

Practise choosing and evaluating statistical diagramsReal past-paper questions · Selecting and evaluating statistical diagrams
02

Stem-and-leaf diagrams

Syllabus requirement · §5.1

“

draw and interpret stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs (Including back-to-back stem-and-leaf diagrams.)

”

A stem-and-leaf diagram lists every value in a data set, sorted, in a way that also shows its shape. Each value is split into two parts:

  • the stem: the leading digit or digits, written once as a row label down the middle or left;
  • the leaf: the single final digit, written in that row.

For example, 3434 splits into stem 33 and leaf 44. The row "3∣3 4 83 \mid 3\ 4\ 8" therefore means the three values 3333, 3434 and 3838. Because each leaf takes up the same width, a longer row means more values, so the diagram doubles as a sideways bar chart of the data.

Choosing the stem. The stem is everything except the last digit that you want to show:

dataexample valuestemleafkey
two-digit whole numbers343433443∣43 \mid 4 means 3434
three-digit whole numbers178178 cm17178817∣817 \mid 8 means 178178 cm
one decimal place7.97.9 s77997∣97 \mid 9 means 7.97.9 s
large values, recorded to the nearest hundred$31 20031312231∣231 \mid 2 means $31 200

The last row is why the key matters: the digits "31∣231 \mid 2" mean nothing until the key says they stand for $31 200.

What the four marks are for

A "draw a back-to-back stem-and-leaf diagram" part is almost always worth 4 marks, and the mark schemes award them for four separate things:

  1. The stem — correct, in order with the smallest at the top, each stem written once (not split, not upside down).
  2. The left-hand data set — labelled, leaves in order increasing from right to left, lined up in neat columns, no commas.
  3. The right-hand data set — labelled, leaves in order increasing from left to right, lined up, no commas.
  4. One key that names both data sets and states the units, e.g. "8∣2∣58 \mid 2 \mid 5 means 2828 minutes for Smarts and 2525 minutes for Teasers".

A single missing leaf, a comma between leaves or a key without units each costs a whole mark, so check each of the four before moving on.

Drawing a stem-and-leaf diagram

The scores, out of 5050, of 1212 students in a class test are

34, 12, 28, 45, 19, 33, 27, 41, 15, 22, 38, 934,\ 12,\ 28,\ 45,\ 19,\ 33,\ 27,\ 41,\ 15,\ 22,\ 38,\ 9

Draw a stem-and-leaf diagram to represent these scores.

Show full working
  1. 1

    Sort the scores into ascending order. 9, 12, 15, 19, 22, 27, 28, 33, 34, 38, 41, 459,\ 12,\ 15,\ 19,\ 22,\ 27,\ 28,\ 33,\ 34,\ 38,\ 41,\ 45

    Sorting first means the leaves land in order automatically in each row. It is also the list you will need later for the median and quartiles.

  2. 2

    Choose the stems. The scores are two-digit numbers (think of 99 as 0909), so the stem is the tens digit. The scores run from 99 to 4545, so the stems are 0, 1, 2, 3, 40,\ 1,\ 2,\ 3,\ 4.

    Write every stem from the smallest to the largest, even one that turns out to have no leaves.

  3. 3

    Write the leaves of the stem-00 and stem-11 rows. From the sorted list: stem 00 gets 99; stem 11 gets the units digits of 12,15,1912, 15, 19, which are 2,5,92, 5, 9. 0912 5 9\begin{array}{c|l} 0 & 9 \\ 1 & 2\ 5\ 9 \end{array}

    Each leaf is only the last digit. The '1' of 12 is already carried by the stem.

  4. 4

    Write the remaining rows the same way. Stem 22: 22,27,28→2,7,822, 27, 28 \to 2, 7, 8. Stem 33: 33,34,38→3,4,833, 34, 38 \to 3, 4, 8. Stem 44: 41,45→1,541, 45 \to 1, 5. 0912 5 922 7 833 4 841 5\begin{array}{c|l} 0 & 9 \\ 1 & 2\ 5\ 9 \\ 2 & 2\ 7\ 8 \\ 3 & 3\ 4\ 8 \\ 4 & 1\ 5 \end{array}

    Keep the leaves in equally spaced columns, with spaces and no commas, so that longer rows really look longer.

  5. 5

    Count the leaves. 1+3+3+3+2=121 + 3 + 3 + 3 + 2 = 12, which matches the 1212 students.

    A dropped value is invisible once the diagram is drawn. The count catches it.

  6. 6

    Add a key saying what the digits mean. Key: 1∣2 means a score of 12\text{Key: } 1 \mid 2 \text{ means a score of } 12

    Without the key a reader cannot tell 12 from 1.2 or 120. When the data has units (cm, s, kg) they go in the key too.

Answer

0912 5 922 7 833 4 841 5\begin{array}{c|l} 0 & 9 \\ 1 & 2\ 5\ 9 \\ 2 & 2\ 7\ 8 \\ 3 & 3\ 4\ 8 \\ 4 & 1\ 5 \end{array} with key 1∣21 \mid 2 means a score of 1212.

Back-to-back stem-and-leaf diagrams

To compare two data sets, both are put on one diagram with a shared stem down the middle. The right-hand set is written exactly as above. The left-hand set is its mirror image: its leaves also increase moving away from the stem, and on the left "away from the stem" means leftwards. So the left-hand leaves are written in reverse order: the smallest leaf sits right next to the stem and the largest is furthest out.

Reading works the same way: on either side, start at the stem and read outwards.

Group PGroup Q1892976325835103684014leaves increase leftwardsleaves increase rightwardsread outwards:22, 23, 26, 27, 29Key: 3 | 2 | 5 means 23 seconds for P and 25 seconds for Q

A back-to-back diagram of two groups of 9 times. On both sides the leaves increase moving away from the shared stem, so the left-hand row “9 7 6 3 2 | 2” is read from the stem outwards: 22, 23, 26, 27, 29. The single key covers both groups and states the units.

Drawing a back-to-back stem-and-leaf diagram

The times, in seconds, taken by 99 pupils in each of two groups, PP and QQ, to solve a puzzle are:

P: 23, 31, 18, 27, 35, 22, 29, 40, 26Q: 30, 19, 25, 33, 38, 41, 28, 36, 44P:\ 23,\ 31,\ 18,\ 27,\ 35,\ 22,\ 29,\ 40,\ 26 \qquad Q:\ 30,\ 19,\ 25,\ 33,\ 38,\ 41,\ 28,\ 36,\ 44

Draw a back-to-back stem-and-leaf diagram with group PP on the left.

Show full working
  1. 1

    Sort each group separately. P: 18, 22, 23, 26, 27, 29, 31, 35, 40P:\ 18,\ 22,\ 23,\ 26,\ 27,\ 29,\ 31,\ 35,\ 40 Q: 19, 25, 28, 30, 33, 36, 38, 41, 44Q:\ 19,\ 25,\ 28,\ 30,\ 33,\ 36,\ 38,\ 41,\ 44

    The two groups are never mixed. Each gets its own sorted list.

  2. 2

    Choose the shared stems. Both groups run from the teens to the forties, so the stems are 1, 2, 3, 41,\ 2,\ 3,\ 4, written once down the middle.

    One stem column serves both sides, so it must cover the smallest and largest values of either group.

  3. 3

    Write QQ's leaves on the right, increasing left to right. 1925 830 3 6 841 4\begin{array}{c|l} 1 & 9 \\ 2 & 5\ 8 \\ 3 & 0\ 3\ 6\ 8 \\ 4 & 1\ 4 \end{array}

    The right-hand side is an ordinary stem-and-leaf diagram.

  4. 4

    Write PP's stem-22 row on the left, in reverse. PP's values in the twenties are 22,23,26,27,2922, 23, 26, 27, 29, with leaves 2,3,6,7,92, 3, 6, 7, 9. Written so that they increase away from the stem (right to left), the row reads 9 7 6 3 2  ∣  29\ 7\ 6\ 3\ 2 \;\big|\; 2

    This is the step most often done backwards. Check: the leaf touching the stem is the smallest, and the leaf furthest left is the largest.

  5. 5

    Write PP's other rows the same way. Stem 11: 18→818 \to 8. Stem 33: 31,35→31, 35 \to written 5 15\ 1. Stem 44: 40→040 \to 0. PQ8199 7 6 3 225 85 130 3 6 8041 4\begin{array}{r|c|l} P & & Q \\ \hline 8 & 1 & 9 \\ 9\ 7\ 6\ 3\ 2 & 2 & 5\ 8 \\ 5\ 1 & 3 & 0\ 3\ 6\ 8 \\ 0 & 4 & 1\ 4 \end{array}

    Label the two sides with the group names. That labelling is part of each side's mark.

  6. 6

    Count each side. PP: 1+5+2+1=91 + 5 + 2 + 1 = 9. QQ: 1+2+4+2=91 + 2 + 4 + 2 = 9. Both match the 99 pupils.

    Count each side separately. A leaf written on the wrong side keeps the total the same but breaks both counts.

  7. 7

    Write one key that covers both sides and gives the units. Pick any row, for example stem 22 with PP's leaf 33 and QQ's leaf 55: Key: 3∣2∣5 means 23 seconds for P and 25 seconds for Q\text{Key: } 3 \mid 2 \mid 5 \text{ means } 23 \text{ seconds for } P \text{ and } 25 \text{ seconds for } Q

    The key reads 'left leaf | stem | right leaf'. It must name both groups and state the units, or the key mark is lost.

Answer

PQ8199 7 6 3 225 85 130 3 6 8041 4\begin{array}{r|c|l} P & & Q \\ \hline 8 & 1 & 9 \\ 9\ 7\ 6\ 3\ 2 & 2 & 5\ 8 \\ 5\ 1 & 3 & 0\ 3\ 6\ 8 \\ 0 & 4 & 1\ 4 \end{array} with key 3∣2∣53 \mid 2 \mid 5 means 2323 s for PP and 2525 s for QQ.

Before drawing, sort both lists. After drawing, count both sides. Those two habits catch almost every lost mark.

A back-to-back diagram from a past paper

9709/55 M/J 2025 Q5(a)4 marks

The Smarts and the Teasers are two quiz teams that each contain 1111 members. Both complete a puzzle and the following table gives the times taken, in minutes, by the members of each team.

Smarts383013291822281811941
Teasers3937183625253221151239

Represent this information in a back-to-back stem-and-leaf diagram with Smarts on the left-hand side.

Show full working
  1. 1

    Sort the Smarts' times. 9, 11, 13, 18, 18, 22, 28, 29, 30, 38, 419,\ 11,\ 13,\ 18,\ 18,\ 22,\ 28,\ 29,\ 30,\ 38,\ 41

    Neither row of the table is in order, so sorting comes first.

  2. 2

    Sort the Teasers' times. 12, 15, 18, 21, 25, 25, 32, 36, 37, 39, 3912,\ 15,\ 18,\ 21,\ 25,\ 25,\ 32,\ 36,\ 37,\ 39,\ 39

    Repeated values (18, 18 and 25, 25 and 39, 39) each get their own leaf.

  3. 3

    Choose the shared stems. The smallest time is 99 (a Smart) and the largest is 4141 (also a Smart), so the stems are 0, 1, 2, 3, 40,\ 1,\ 2,\ 3,\ 4.

    The stem 0 row is needed for the value 9, and stem 4 for 41, even though the Teasers have nothing on either row.

  4. 4

    Write the Teasers' leaves on the right, increasing outwards. Stem 00: none. Stem 11: 12,15,18→2 5 812, 15, 18 \to 2\ 5\ 8. Stem 22: 21,25,25→1 5 521, 25, 25 \to 1\ 5\ 5. Stem 33: 32,36,37,39,39→2 6 7 9 932, 36, 37, 39, 39 \to 2\ 6\ 7\ 9\ 9. Stem 44: none.

    The ordinary side first. Count: 3 + 3 + 5 = 11 ✓.

  5. 5

    Write the Smarts' leaves on the left, increasing outwards (right to left). Stem 00: 99. Stem 11: 11,13,18,18→11, 13, 18, 18 \to written 8 8 3 18\ 8\ 3\ 1. Stem 22: 22,28,29→22, 28, 29 \to written 9 8 29\ 8\ 2. Stem 33: 30,38→30, 38 \to written 8 08\ 0. Stem 44: 41→141 \to 1.

    The leaf next to the stem is the smallest in each row. Count: 1 + 4 + 3 + 2 + 1 = 11 ✓.

  6. 6

    Assemble the diagram with labels. SmartsTeasers908 8 3 112 5 89 8 221 5 58 032 6 7 9 914\begin{array}{r|c|l} \text{Smarts} & & \text{Teasers} \\ \hline 9 & 0 & \\ 8\ 8\ 3\ 1 & 1 & 2\ 5\ 8 \\ 9\ 8\ 2 & 2 & 1\ 5\ 5 \\ 8\ 0 & 3 & 2\ 6\ 7\ 9\ 9 \\ 1 & 4 & \end{array}

    The 9 belongs on the Smarts' side of stem 0: it is a Smarts time. The Teasers have nothing on this row.

  7. 7

    Write the key, naming both teams and the units. Key: 8∣2∣5 means 28 minutes for Smarts and 25 minutes for Teasers\text{Key: } 8 \mid 2 \mid 5 \text{ means } 28 \text{ minutes for Smarts and } 25 \text{ minutes for Teasers}

    This is the mark scheme's own key. Any correct row would do, as long as both teams and 'minutes' appear.

Answer

SmartsTeasers908 8 3 112 5 89 8 221 5 58 032 6 7 9 914\begin{array}{r|c|l} \text{Smarts} & & \text{Teasers} \\ \hline 9 & 0 & \\ 8\ 8\ 3\ 1 & 1 & 2\ 5\ 8 \\ 9\ 8\ 2 & 2 & 1\ 5\ 5 \\ 8\ 0 & 3 & 2\ 6\ 7\ 9\ 9 \\ 1 & 4 & \end{array} Key: 8∣2∣58 \mid 2 \mid 5 means 2828 minutes for Smarts and 2525 minutes for Teasers.

The same data comes back in §04 (box-and-whisker plots) and §05 (comparisons): a stem-and-leaf part is usually the first step of a longer question, so an error here costs marks later too.

Common mistakes
  • Left-hand leaves written smallest to largest, left to right

    Left-hand leaves increase away from the stem, i.e. from right to left

    Both sides grow outwards from the shared stem. The leaf touching the stem is the smallest on either side.

  • A key such as "1∣21 \mid 2 means 1212" on a back-to-back diagram

    "3∣2∣53 \mid 2 \mid 5 means 2323 s for PP and 2525 s for QQ": both groups named, units stated

    Mark schemes require both data sets to be identified and the units to appear in the key, the headings or a title.

  • Leaves separated by commas, or bunched unevenly

    Leaves in neat, equally spaced columns with no punctuation

    The alignment is what lets the rows act as bars. Commas and misalignment each cost a mark.

  • A stem left out because no value falls in it

    Write every stem from the smallest to the largest, even if its row is empty on one side

    Missing stems distort the shape, and the stem mark needs the complete, ordered stem.

In the exam
11 of the 47 questions on this topic in 2021–2025 asked for a back-to-back stem-and-leaf diagram, every time for 4 marks

It is the most common single part in the whole topic. The data is always two teams or groups of 11 or 15 values each, and the later parts (median and IQR, box plots, comparisons) are built on the diagram, so the four marks here are worth securing with a sort-draw-count-key routine.

Your turn

For every diagram: sort first, count the leaves on each side at the end, and write a key with both names and units.

  1. 1

    Draw a stem-and-leaf diagram for the following masses, in kg, of 1010 parcels: 21, 34, 18, 27, 33, 40, 25, 19, 31, 2221,\ 34,\ 18,\ 27,\ 33,\ 40,\ 25,\ 19,\ 31,\ 22

    Stuck? Show hint

    Sort the list first, then use the tens digit as the stem.

    Show solution
    1. 1

      Sort the masses. 18, 19, 21, 22, 25, 27, 31, 33, 34, 4018,\ 19,\ 21,\ 22,\ 25,\ 27,\ 31,\ 33,\ 34,\ 40

      Sorting puts the leaves in order automatically.

    2. 2

      Choose the stems: 1, 2, 3, 41,\ 2,\ 3,\ 4 (tens digits from 1818 up to 4040).

      The stem is the tens digit.

    3. 3

      Write each row. 18 921 2 5 731 3 440\begin{array}{c|l} 1 & 8\ 9 \\ 2 & 1\ 2\ 5\ 7 \\ 3 & 1\ 3\ 4 \\ 4 & 0 \end{array}

      Each leaf is the units digit.

    4. 4

      Count the leaves. 2+4+3+1=102 + 4 + 3 + 1 = 10 ✓

      The count matches the 10 parcels.

    5. 5

      Add the key. Key: 1∣81 \mid 8 means 1818 kg.

      The key states what the digits mean, with the units.

    Answer

    18 921 2 5 731 3 440\begin{array}{c|l} 1 & 8\ 9 \\ 2 & 1\ 2\ 5\ 7 \\ 3 & 1\ 3\ 4 \\ 4 & 0 \end{array} with key 1∣81 \mid 8 means 1818 kg.

  2. 29709/51 M/J 2025 Q3(a)4 marks

    Last Sunday, teams of runners took part in a charity event. The time taken, in seconds, to run 50 m50\text{ m} was recorded, correct to 1 decimal place, for each runner. The times recorded for 1111 runners from each of the Gulls and the Herons are shown in the table.

    Gulls7.98.28.38.68.68.89.29.79.810.010.4
    Herons9.59.98.58.19.210.88.39.79.39.98.7

    Draw a back-to-back stem-and-leaf diagram to represent this information, with Gulls on the left-hand side.

    Stuck? Show hint

    With one decimal place, the stem is the whole-number part and the leaf is the tenths digit: 7.97.9 is 7∣97 \mid 9. The stems run from 77 to 1010.

    Show solution
    1. 1

      Gulls are already in order. 7.9, 8.2, 8.3, 8.6, 8.6, 8.8, 9.2, 9.7, 9.8, 10.0, 10.47.9,\ 8.2,\ 8.3,\ 8.6,\ 8.6,\ 8.8,\ 9.2,\ 9.7,\ 9.8,\ 10.0,\ 10.4

      Always check: this row happens to be sorted, the Herons' row is not.

    2. 2

      Sort the Herons. 8.1, 8.3, 8.5, 8.7, 9.2, 9.3, 9.5, 9.7, 9.9, 9.9, 10.88.1,\ 8.3,\ 8.5,\ 8.7,\ 9.2,\ 9.3,\ 9.5,\ 9.7,\ 9.9,\ 9.9,\ 10.8

      Eleven values, including the repeated 9.9.

    3. 3

      Choose the stems: the whole-number parts 7, 8, 9, 107,\ 8,\ 9,\ 10. The leaf is the tenths digit, so 10.410.4 is stem 1010, leaf 44.

      The stem 10 is written as a single stem, not split into '1' and '0'.

    4. 4

      Write the Gulls on the left, increasing outwards. Stem 77: 99. Stem 88: 2,3,6,6,8→8 6 6 3 22, 3, 6, 6, 8 \to 8\ 6\ 6\ 3\ 2. Stem 99: 2,7,8→8 7 22, 7, 8 \to 8\ 7\ 2. Stem 1010: 0,4→4 00, 4 \to 4\ 0.

      Reverse order on the left. Count: 1 + 5 + 3 + 2 = 11 ✓.

    5. 5

      Write the Herons on the right. Stem 77: none. Stem 88: 1 3 5 71\ 3\ 5\ 7. Stem 99: 2 3 5 7 9 92\ 3\ 5\ 7\ 9\ 9. Stem 1010: 88.

      Count: 4 + 6 + 1 = 11 ✓.

    6. 6

      Assemble the diagram. GullsHerons978 6 6 3 281 3 5 78 7 292 3 5 7 9 94 0108\begin{array}{r|c|l} \text{Gulls} & & \text{Herons} \\ \hline 9 & 7 & \\ 8\ 6\ 6\ 3\ 2 & 8 & 1\ 3\ 5\ 7 \\ 8\ 7\ 2 & 9 & 2\ 3\ 5\ 7\ 9\ 9 \\ 4\ 0 & 10 & 8 \end{array}

      This matches the published mark scheme.

    7. 7

      Key. 7∣9∣57 \mid 9 \mid 5 means 9.79.7 seconds for Gulls and 9.59.5 seconds for Herons.

      The key must name both teams and give the units (s); writing 9.7 and 9.5 also shows where the decimal point goes.

    Answer

    GullsHerons978 6 6 3 281 3 5 78 7 292 3 5 7 9 94 0108\begin{array}{r|c|l} \text{Gulls} & & \text{Herons} \\ \hline 9 & 7 & \\ 8\ 6\ 6\ 3\ 2 & 8 & 1\ 3\ 5\ 7 \\ 8\ 7\ 2 & 9 & 2\ 3\ 5\ 7\ 9\ 9 \\ 4\ 0 & 10 & 8 \end{array} Key: 7∣9∣57 \mid 9 \mid 5 means 9.79.7 s for Gulls and 9.59.5 s for Herons.

  3. 39709/55 O/N 2025 Q4(a)4 marks

    The heights, in cm, of 1515 players from each of two sports teams, Pelicans and Swans, are given in the table.

    Pelicans156160164165167170171173178182182184185186187
    Swans170180183165174158170181162178174163191182174

    Draw a back-to-back stem-and-leaf diagram to represent the heights of the players from Pelicans and Swans, with Pelicans on the left-hand side.

    Stuck? Show hint

    Three-digit values: the stem is the first two digits (1515 to 1919) and the leaf is the units digit.

    Show solution
    1. 1

      Pelicans are already in order. 156, 160, 164, 165, 167, 170, 171, 173, 178, 182, 182, 184, 185, 186, 187156,\ 160,\ 164,\ 165,\ 167,\ 170,\ 171,\ 173,\ 178,\ 182,\ 182,\ 184,\ 185,\ 186,\ 187

      15 values, already sorted.

    2. 2

      Sort the Swans. 158, 162, 163, 165, 170, 170, 174, 174, 174, 178, 180, 181, 182, 183, 191158,\ 162,\ 163,\ 165,\ 170,\ 170,\ 174,\ 174,\ 174,\ 178,\ 180,\ 181,\ 182,\ 183,\ 191

      Check the count: 15 values.

    3. 3

      Stems 15, 16, 17, 18, 1915,\ 16,\ 17,\ 18,\ 19; each leaf is the units digit.

      178 is stem 17, leaf 8.

    4. 4

      Pelicans on the left, increasing outwards. Stem 1515: 66. Stem 1616: 7 5 4 07\ 5\ 4\ 0. Stem 1717: 8 3 1 08\ 3\ 1\ 0. Stem 1818: 7 6 5 4 2 27\ 6\ 5\ 4\ 2\ 2. Stem 1919: none.

      Count: 1 + 4 + 4 + 6 = 15 ✓.

    5. 5

      Swans on the right. Stem 1515: 88. Stem 1616: 2 3 52\ 3\ 5. Stem 1717: 0 0 4 4 4 80\ 0\ 4\ 4\ 4\ 8. Stem 1818: 0 1 2 30\ 1\ 2\ 3. Stem 1919: 11.

      Count: 1 + 3 + 6 + 4 + 1 = 15 ✓. The published mark scheme prints the stem-17 row as 0 0 4 4 4 and leaves out the 8, although its own key uses 178 cm for a Swan. The data has a Swan of 178 cm, so the 8 belongs in this row.

    6. 6

      Assemble the diagram. PelicansSwans61587 5 4 0162 3 58 3 1 0170 0 4 4 4 87 6 5 4 2 2180 1 2 3191\begin{array}{r|c|l} \text{Pelicans} & & \text{Swans} \\ \hline 6 & 15 & 8 \\ 7\ 5\ 4\ 0 & 16 & 2\ 3\ 5 \\ 8\ 3\ 1\ 0 & 17 & 0\ 0\ 4\ 4\ 4\ 8 \\ 7\ 6\ 5\ 4\ 2\ 2 & 18 & 0\ 1\ 2\ 3 \\ & 19 & 1 \end{array}

      The stem 19 row is needed for the Swan of 191 cm.

    7. 7

      Key. 1∣17∣81 \mid 17 \mid 8 represents 171171 cm for Pelicans and 178178 cm for Swans.

      Both teams named, units (cm) stated.

    Answer

    PelicansSwans61587 5 4 0162 3 58 3 1 0170 0 4 4 4 87 6 5 4 2 2180 1 2 3191\begin{array}{r|c|l} \text{Pelicans} & & \text{Swans} \\ \hline 6 & 15 & 8 \\ 7\ 5\ 4\ 0 & 16 & 2\ 3\ 5 \\ 8\ 3\ 1\ 0 & 17 & 0\ 0\ 4\ 4\ 4\ 8 \\ 7\ 6\ 5\ 4\ 2\ 2 & 18 & 0\ 1\ 2\ 3 \\ & 19 & 1 \end{array} Key: 1∣17∣81 \mid 17 \mid 8 represents 171171 cm for Pelicans and 178178 cm for Swans.

  4. 49709/52 O/N 2024 Q6(a)4 marks

    Teams of 1515 runners took part in a charity run last Saturday. The times taken, in minutes, to complete the course by the runners from the Falcons and the runners from the Kites are shown in the table.

    Falcons383942444648505152565859646976
    Kites324040454748525458595960616365

    Draw a back-to-back stem-and-leaf diagram to represent this information, with the Falcons on the left-hand side.

    Stuck? Show hint

    Both rows are already sorted. The stems run from 33 to 77; only the Falcons have a value on stem 77.

    Show solution
    1. 1

      Stems 3, 4, 5, 6, 73,\ 4,\ 5,\ 6,\ 7 (from 3232 up to 7676).

      Stem 7 is needed for the Falcon who took 76 minutes.

    2. 2

      Falcons on the left, increasing outwards. Stem 33: 38,39→9 838, 39 \to 9\ 8. Stem 44: 42,44,46,48→8 6 4 242, 44, 46, 48 \to 8\ 6\ 4\ 2. Stem 55: 50,51,52,56,58,59→9 8 6 2 1 050, 51, 52, 56, 58, 59 \to 9\ 8\ 6\ 2\ 1\ 0. Stem 66: 64,69→9 464, 69 \to 9\ 4. Stem 77: 76→676 \to 6.

      Count: 2 + 4 + 6 + 2 + 1 = 15 ✓.

    3. 3

      Kites on the right. Stem 33: 22. Stem 44: 0 0 5 7 80\ 0\ 5\ 7\ 8. Stem 55: 2 4 8 9 92\ 4\ 8\ 9\ 9. Stem 66: 0 1 3 50\ 1\ 3\ 5. Stem 77: none.

      Count: 1 + 5 + 5 + 4 = 15 ✓.

    4. 4

      Assemble the diagram. FalconsKites9 8328 6 4 240 0 5 7 89 8 6 2 1 052 4 8 9 99 460 1 3 567\begin{array}{r|c|l} \text{Falcons} & & \text{Kites} \\ \hline 9\ 8 & 3 & 2 \\ 8\ 6\ 4\ 2 & 4 & 0\ 0\ 5\ 7\ 8 \\ 9\ 8\ 6\ 2\ 1\ 0 & 5 & 2\ 4\ 8\ 9\ 9 \\ 9\ 4 & 6 & 0\ 1\ 3\ 5 \\ 6 & 7 & \end{array}

      This matches the published mark scheme, including the Falcons' 6 on stem 7.

    5. 5

      Key. 1∣5∣41 \mid 5 \mid 4 means 5151 minutes for Falcons and 5454 minutes for Kites.

      Both teams named, units (minutes) stated.

    Answer

    FalconsKites9 8328 6 4 240 0 5 7 89 8 6 2 1 052 4 8 9 99 460 1 3 567\begin{array}{r|c|l} \text{Falcons} & & \text{Kites} \\ \hline 9\ 8 & 3 & 2 \\ 8\ 6\ 4\ 2 & 4 & 0\ 0\ 5\ 7\ 8 \\ 9\ 8\ 6\ 2\ 1\ 0 & 5 & 2\ 4\ 8\ 9\ 9 \\ 9\ 4 & 6 & 0\ 1\ 3\ 5 \\ 6 & 7 & \end{array} Key: 1∣5∣41 \mid 5 \mid 4 means 5151 minutes for Falcons and 5454 minutes for Kites.

Practise stem-and-leaf diagrams and the other statistical diagramsReal past-paper questions · Stem-and-leaf diagrams, box-and-whisker plots, histograms, cumulative frequency graphs
03

Median, quartiles and interquartile range

Syllabus requirement · §5.1

“

understand and use different measures of central tendency (mean, median, mode) and variation (range, interquartile range, standard deviation)

”

Once data is in order, a few landmark values describe it well.

  • The median is the middle value. Half the data lies below it and half above.
  • The lower quartile Q1Q_1 is the middle of the lower half of the data, and the upper quartile Q3Q_3 is the middle of the upper half. So Q1Q_1, the median and Q3Q_3 cut the ordered data into four quarters.
  • The range is largest value −- smallest value.
  • The interquartile range is IQR=Q3−Q1,\text{IQR} = Q_3 - Q_1, the spread of the middle half of the data. Unlike the range, it ignores the most extreme quarter at each end, so one unusually large or small value cannot distort it.
  • The mode is the most common value. It is rarely asked for on this paper.

A typical 3-mark part asks for "the median and the interquartile range": B1 for the median, M1 for Q3−Q1Q_3 - Q_1 with values in the right place, A1 for the answer.

Median and quartiles of a list of n values
  1. 1

    Sort the values into ascending order and number their positions 1,2,…,n1, 2, \ldots, n.

    Every later step counts positions in this sorted list.

  2. 2

    Median at position 12(n+1)\tfrac12(n+1). If nn is odd this is a whole number, and the median is that value. If nn is even it ends in .5.5, and the median is the mean of the two middle values.

    For n = 11, ½(12) = 6, the 6th value. For n = 12, ½(13) = 6.5, the mean of the 6th and 7th.

  3. 3

    Split the data into a lower half and an upper half. If nn is odd, leave the median itself out of both halves. If nn is even, the halves are simply the first n2\tfrac n2 values and the last n2\tfrac n2.

    Each half then has the same number of values.

  4. 4

    Q1Q_1 = the median of the lower half; Q3Q_3 = the median of the upper half. Find each one exactly as you found the median.

    Mark schemes describe this as finding the quartiles 'from the two halves'. For odd n it puts Q₁ at position ¼(n+1) and Q₃ at position ¾(n+1): for n = 11 that is the 3rd and 9th values, for n = 19 the 5th and 15th, for n = 27 the 7th and 21st.

  5. 5

    IQR =Q3−Q1= Q_3 - Q_1. Show the subtraction.

    The M1 is for the subtraction with Q₃ and Q₁ in acceptable ranges, so write 'IQR = 65 − 54' before the answer.

11 values (odd): leave the median out of both halves467911131518202225Q₁ = 7median = 13Q₃ = 20lower half (5 values)upper half (5 values)12 values (even): the halves are the first 6 and the last 635681011131416192124Q₁ = 7median = 12Q₃ = 17.5lower half (6 values)upper half (6 values)

Top: 11 values (odd). The median is the 6th value; it is left out, and each half of 5 values has its own middle value: Q₁ is the 3rd value, Q₃ the 9th. Bottom: 12 values (even). The median is halfway between the 6th and 7th values; each half of 6 values has an even count, so each quartile is the mean of two values.

Median, quartiles and IQR when n is odd

Find the median, the quartiles, the interquartile range and the range of these 1111 values: 4, 15, 7, 22, 11, 18, 9, 25, 13, 20, 64,\ 15,\ 7,\ 22,\ 11,\ 18,\ 9,\ 25,\ 13,\ 20,\ 6

Show full working
  1. 1

    Sort and number the values. 41, 62, 73, 94, 115, 136, 157, 188, 209, 2210, 2511\underset{1}{4},\ \underset{2}{6},\ \underset{3}{7},\ \underset{4}{9},\ \underset{5}{11},\ \underset{6}{13},\ \underset{7}{15},\ \underset{8}{18},\ \underset{9}{20},\ \underset{10}{22},\ \underset{11}{25}

    Numbering the positions makes each later step a matter of counting.

  2. 2

    Find the median's position. n=11n = 11, so 12(n+1)=12×12=6\tfrac12(n+1) = \tfrac12 \times 12 = 6

    A whole number, so the median is a single value.

  3. 3

    Read the median. The 6th value: median=13\text{median} = 13

    Five values lie below 13 and five above.

  4. 4

    Find Q1Q_1 from the lower half. Leaving out the median, the lower half is 4,6,7,9,114, 6, 7, 9, 11. Its middle (3rd) value is Q1=7Q_1 = 7

    The lower half has 5 values, and the middle of 5 is the 3rd. That is position 3 = ¼(11 + 1) in the full list.

  5. 5

    Find Q3Q_3 from the upper half. The upper half is 15,18,20,22,2515, 18, 20, 22, 25. Its middle value is Q3=20Q_3 = 20

    This is position 9 = ¾(11 + 1) in the full list.

  6. 6

    Subtract to get the IQR. IQR=Q3−Q1=20−7=13\text{IQR} = Q_3 - Q_1 = 20 - 7 = 13

    Upper quartile minus lower quartile, never the other way round.

  7. 7

    Subtract to get the range. range=25−4=21\text{range} = 25 - 4 = 21

    The range uses the smallest and largest values (4 and 25); the IQR uses only the middle half, so it is smaller.

Answer

Median 1313, Q1=7Q_1 = 7, Q3=20Q_3 = 20, IQR =13= 13, range =21= 21.

Median, quartiles and IQR when n is even

Find the median and the interquartile range of these 1212 values: 13, 5, 21, 8, 16, 3, 24, 10, 14, 6, 19, 1113,\ 5,\ 21,\ 8,\ 16,\ 3,\ 24,\ 10,\ 14,\ 6,\ 19,\ 11

Show full working
  1. 1

    Sort and number the values. 31, 52, 63, 84, 105, 116, 137, 148, 169, 1910, 2111, 2412\underset{1}{3},\ \underset{2}{5},\ \underset{3}{6},\ \underset{4}{8},\ \underset{5}{10},\ \underset{6}{11},\ \underset{7}{13},\ \underset{8}{14},\ \underset{9}{16},\ \underset{10}{19},\ \underset{11}{21},\ \underset{12}{24}

    Twelve values, so n is even.

  2. 2

    Find the median's position. 12(n+1)=12×13=6.5\tfrac12(n+1) = \tfrac12 \times 13 = 6.5

    Position 6.5 means halfway between the 6th and 7th values.

  3. 3

    Average the 6th and 7th values. median=11+132=12\text{median} = \frac{11 + 13}{2} = 12

    Never round 6.5 to 6 or 7. The median here is not one of the data values at all.

  4. 4

    Split into halves of 66. Lower half: 3,5,6,8,10,113, 5, 6, 8, 10, 11. Upper half: 13,14,16,19,21,2413, 14, 16, 19, 21, 24.

    With n even, no value is left out: the halves are the first 6 and the last 6.

  5. 5

    Q1Q_1 is the median of the lower half, halfway between its 3rd and 4th values: Q1=6+82=7Q_1 = \frac{6 + 8}{2} = 7

    Six values have no single middle, so average the two middle ones.

  6. 6

    Q3Q_3 is the median of the upper half: Q3=16+192=17.5Q_3 = \frac{16 + 19}{2} = 17.5

    Same rule for the upper half: its 3rd and 4th values are 16 and 19.

  7. 7

    Subtract to get the IQR. IQR=17.5−7=10.5\text{IQR} = 17.5 - 7 = 10.5

    Quartiles that are not data values are perfectly normal when n is even.

Answer

Median 1212, Q1=7Q_1 = 7, Q3=17.5Q_3 = 17.5, IQR =10.5= 10.5.

Recent Paper 5 questions almost always use an odd number of values (11, 15, 19 or 27), so the quartiles are data values. The mark schemes also accept quartiles within a small range around the correct position, but the halves method gives the values they print.

Median and IQR from a sorted table

9709/55 O/N 2025 Q4(b)3 marks

The heights, in cm, of 1515 players from each of two sports teams, Pelicans and Swans, are given in the table.

Pelicans156160164165167170171173178182182184185186187
Swans170180183165174158170181162178174163191182174

Find the median and the interquartile range of the heights of the Pelicans.

Show full working
  1. 1

    Check the order and count. The Pelicans' row is already in ascending order, and n=15n = 15.

    Always check. The Swans' row, for comparison, is not sorted.

  2. 2

    Median position. 12(15+1)=8\tfrac12(15 + 1) = 8

    n = 15 is odd, so the median is the 8th value.

  3. 3

    Read the median. The 8th Pelican is 173173, so median=173 cm\text{median} = 173 \text{ cm}

    The B1 needs the median clearly identified: write 'median =', not just the number.

  4. 4

    Lower half. Leaving out the 8th value, the lower half is the first 77 values: 156,160,164,165,167,170,171156, 160, 164, 165, 167, 170, 171. Its middle (4th) value is Q1=165Q_1 = 165

    Middle of 7 is the 4th. In the full list that is position ¼(15 + 1) = 4.

  5. 5

    Upper half. The last 77 values are 178,182,182,184,185,186,187178, 182, 182, 184, 185, 186, 187. Its middle value is Q3=184Q_3 = 184

    Position ¾(15 + 1) = 12 in the full list.

  6. 6

    IQR. IQR=184−165=19 cm\text{IQR} = 184 - 165 = 19 \text{ cm}

    The M1 accepts 182 ≤ UQ ≤ 185 and 164 ≤ LQ ≤ 167, but it is earned by showing the subtraction.

Answer

Median =173= 173 cm; IQR =184−165=19= 184 - 165 = 19 cm.

Reading a printed stem-and-leaf diagram

When the data comes as a printed diagram, the method is the same. Two extra things need care:

  1. Converting leaves back to values with the key. In a key like "6∣32∣76 \mid 32 \mid 7 means $32 600 for Browns and $32 700 for Greens", the stem is thousands of dollars and the leaf is hundreds.
  2. Reading the left-hand side outwards from the stem. A left-hand row printed as "9 8 4∣309\ 8\ 4 \mid 30" holds the salaries $30 400, $30 800 and $30 900, smallest first: the digit next to the stem is the smallest.

To find, say, the 14th value, count down the rows, keeping a running total of leaves, until the total passes 14. Some papers print each row's count in brackets, which saves the counting.

Median and IQR from a printed back-to-back diagram

9709/51 O/N 2025 Q3(a)3 marks

The back-to-back stem-and-leaf diagram shows the annual salaries, in dollars, of 2727 employees at each of two companies, Browns and Greens.

Find the median and interquartile range for the annual salaries of employees at Browns.

The printed diagram. Browns is on the left, so each Browns row is read from the stem outwards. The bracketed numbers are the number of leaves in each row.

The printed diagram. Browns is on the left, so each Browns row is read from the stem outwards. The bracketed numbers are the number of leaves in each row.

Show full working
  1. 1

    Read the key. 6∣32∣76 \mid 32 \mid 7 means $32 600 for Browns, so a stem of 3232 is $32 000 and each leaf counts hundreds of dollars.

    Get the units right before reading any value, or every answer is out by a factor of 100.

  2. 2

    Median position. n=27n = 27, so n+1=28n + 1 = 28 and 12×28=14th\tfrac12 \times 28 = 14\text{th}

    n is odd, so the median is a single value.

  3. 3

    Quartile positions. Q1:14×28=7th,Q3:34×28=21stQ_1: \tfrac14 \times 28 = 7\text{th}, \qquad Q_3: \tfrac34 \times 28 = 21\text{st}

    These are exactly the halves method: the lower half is the first 13 values and its middle is the 7th.

  4. 4

    Running totals of the Browns' row counts. The brackets give 3,7,7,6,3,13, 7, 7, 6, 3, 1 for stems 3030 to 3535, so the running totals are 3, 10, 17, 23, 26, 273,\ 10,\ 17,\ 23,\ 26,\ 27

    Row 30 holds values 1–3, row 31 holds 4–10, row 32 holds 11–17, row 33 holds 18–23, and so on.

  5. 5

    Locate the 14th value. It is in row 3232 (values 11–17), and it is the 14−10=414 - 10 = 4th leaf of that row, reading outwards from the stem. Row 3232 read outwards is 0,2,2,4,6,7,90, 2, 2, 4, 6, 7, 9, so the 4th leaf is 44: median=$32 400\text{median} = \$32\,400

    Reading outwards on the left: the printed row '9 7 6 4 2 2 0 | 32' is 0, 2, 2, 4, 6, 7, 9 from the stem outwards.

  6. 6

    Locate the 7th value (Q1Q_1). It is in row 3131 (values 4–10), the 7−3=47 - 3 = 4th leaf. Row 3131 read outwards is 0,1,3,3,5,8,80, 1, 3, 3, 5, 8, 8, so the 4th leaf is 33: Q1=$31 300Q_1 = \$31\,300

    Same counting, one row higher.

  7. 7

    Locate the 21st value (Q3Q_3). It is in row 3333 (values 18–23), the 21−17=421 - 17 = 4th leaf. Row 3333 read outwards is 1,3,5,5,7,81, 3, 5, 5, 7, 8, so the 4th leaf is 55: Q3=$33 500Q_3 = \$33\,500

    Check the row count: 6 leaves, matching the bracketed (6).

  8. 8

    IQR. IQR=33 500−31 300=$2200\text{IQR} = 33\,500 - 31\,300 = \$2200

    The mark scheme accepts 335[00] ≤ UQ ≤ 337[00] and 313[00] ≤ LQ ≤ 315[00] for the M1; the answer 2200 is marked CAO.

Answer

Median = $32 400; IQR = 33 500 − 31 300 = $2200.

Use running totals of the row counts to find which row a position falls in, then count leaves outwards from the stem within that row.

Common mistakes
  • Reading a left-hand row of a back-to-back diagram from left to right

    Read every row from the stem outwards

    On the left the smallest leaf is next to the stem. Reading the printed order gives the row backwards and moves the quartiles.

  • Rounding a position such as 6.56.5 to 66 or 77

    Average the two neighbouring values

    Position 6.5 is halfway between the 6th and 7th values, so neither one alone is the median.

  • Ignoring the key, e.g. giving a median of 324324 instead of $32 400

    Convert every value you quote using the key

    Mark schemes give only a special-case mark when the key is ignored consistently.

  • Giving the IQR without showing Q3−Q1Q_3 - Q_1

    Write "IQR =65−54=11= 65 - 54 = 11"

    The method mark is for the subtraction with both quartiles in range; a bare wrong answer gets nothing.

Your turn

Work out the positions from n first, then count. For printed diagrams, convert with the key and read each row outwards from the stem.

  1. 19709/53 O/N 2025 Q6(b)2 marks

    Last Saturday, a cycling competition for teams of 1111 cyclists took place. For each cyclist, the time taken to complete the course was recorded to the nearest minute. The times taken by the cyclists from two teams, the Linnets and the Puffins, are shown in the following table.

    Linnets4851545759606464656870
    Puffins4549515555585962646474

    Find the interquartile range of the times taken by the Linnets.

    Stuck? Show hint

    n=11n = 11: the quartiles are the 3rd and 9th values.

    Show solution
    1. 1

      Positions. n=11n = 11 and the Linnets are sorted. The median is the 6th value (6060), so the lower half is 48,51,54,57,5948, 51, 54, 57, 59 and the upper half is 64,64,65,68,7064, 64, 65, 68, 70.

      Leave the median out of both halves when n is odd.

    2. 2

      Lower quartile. The middle of the lower half: Q1=54Q_1 = 54

      The 3rd value of the full list.

    3. 3

      Upper quartile. The middle of the upper half: Q3=65Q_3 = 65

      The 9th value of the full list.

    4. 4

      IQR. IQR=65−54=11 minutes\text{IQR} = 65 - 54 = 11 \text{ minutes}

      Show the subtraction for the M1.

    Answer

    IQR =65−54=11= 65 - 54 = 11 minutes.

  2. 29709/53 M/J 2023 Q4(a)3 marks

    The times taken, in minutes, to complete a cycle race by 1919 cyclists from each of two clubs, the Cheetahs and the Panthers, are represented in the following back-to-back stem-and-leaf diagram.

    Key: 7∣9∣17 \mid 9 \mid 1 means 9797 minutes for Cheetahs and 9191 minutes for Panthers

    Find the median and the interquartile range of the times of the Cheetahs.

    The printed diagram. Cheetahs are on the left: read each row from the stem outwards.

    The printed diagram. Cheetahs are on the left: read each row from the stem outwards.

    Stuck? Show hint

    n=19n = 19: the median is the 10th value, Q1Q_1 the 5th and Q3Q_3 the 15th. Count the Cheetahs' leaves row by row.

    Show solution
    1. 1

      Median position. n=19n = 19, so the median is at 12(20)=10\tfrac12(20) = 10th.

      n is odd, so the median is a single value.

    2. 2

      Quartile positions. Q1Q_1 at 14(20)=5\tfrac14(20) = 5th, Q3Q_3 at 34(20)=15\tfrac34(20) = 15th.

      These come straight from the halves method: the middle of the first 9 and of the last 9.

    3. 3

      Cheetahs' row counts and running totals. Rows 77 to 1212 have 2,5,3,5,3,12, 5, 3, 5, 3, 1 leaves, so the running totals are 2, 7, 10, 15, 18, 192,\ 7,\ 10,\ 15,\ 18,\ 19

      The totals show which row each position falls in.

    4. 4

      Median (10th). The running total reaches 1010 at the end of row 99, so the 10th value is the last leaf of that row. Row 99 read outwards is 7,8,97, 8, 9: median=99 minutes\text{median} = 99 \text{ minutes}

      Row 9 holds values 8 to 10, and '9 8 7 | 9' read outwards is 97, 98, 99.

    5. 5

      Q1Q_1 (5th). Row 88 holds values 3 to 7; read outwards it is 0,2,3,7,80, 2, 3, 7, 8, and the 5th value is its 3rd leaf: Q1=83Q_1 = 83

      5 − 2 = 3, so the 3rd leaf of row 8.

    6. 6

      Q3Q_3 (15th). Row 1010 holds values 11 to 15; read outwards it is 1,3,3,5,61, 3, 3, 5, 6, and the 15th value is its last leaf: Q3=106Q_3 = 106

      15 − 10 = 5, the 5th leaf of row 10.

    7. 7

      IQR. IQR=106−83=23 minutes\text{IQR} = 106 - 83 = 23 \text{ minutes}

      Subtraction shown for the method mark.

    Answer

    Median =99= 99 minutes; IQR =106−83=23= 106 - 83 = 23 minutes.

  3. 39709/52 M/J 2024 Q4(a)3 marks

    The back-to-back stem-and-leaf diagram shows the annual salaries of 1919 employees at each of two companies, Petral and Ravon.

    Key: 2∣31∣52 \mid 31 \mid 5 means $31 200 for a Petral employee and $31 500 for a Ravon employee.

    Find the median and the interquartile range of the salaries of the Petral employees.

    The printed diagram. Petral is on the left: read each row from the stem outwards.

    The printed diagram. Petral is on the left: read each row from the stem outwards.

    Stuck? Show hint

    Stems are thousands of dollars and leaves are hundreds. With n=19n = 19 you need the 5th, 10th and 15th Petral salaries.

    Show solution
    1. 1

      Petral's rows, read outwards, with running totals. 3030: 0,0,30, 0, 3 (total 33). 3131: 1,2,2,8,9,91, 2, 2, 8, 9, 9 (total 99). 3232: 0,4,5,50, 4, 5, 5 (total 1313). 3333: 3,5,73, 5, 7 (total 1616). 3434: 0,10, 1 (total 1818). 3535: none. 3636: 88 (total 1919).

      The final total, 19, confirms nothing was missed.

    2. 2

      Median (10th). Row 3232 holds values 10 to 13, so the 10th is its first leaf, 00: median=$32 000\text{median} = \$32\,000

      Position ½(19 + 1) = 10.

    3. 3

      Q1Q_1 (5th). Row 3131 holds values 4 to 9; the 5th is its 2nd leaf, 22: Q1=$31 200Q_1 = \$31\,200

      Position ¼(20) = 5.

    4. 4

      Q3Q_3 (15th). Row 3333 holds values 14 to 16; the 15th is its 2nd leaf, 55: Q3=$33 500Q_3 = \$33\,500

      Position ¾(20) = 15.

    5. 5

      IQR. IQR=33 500−31 200=$2300\text{IQR} = 33\,500 - 31\,200 = \$2300

      The mark scheme accepts 33 300 ≤ UQ ≤ 33 700 and 31 100 ≤ LQ ≤ 31 200.

    Answer

    Median = $32 000; IQR = 33 500 − 31 200 = $2300.

  4. 49709/52 M/J 2022 Q3(a)3 marks

    The back-to-back stem-and-leaf diagram shows the diameters, in cm, of 1919 cylindrical pipes produced by each of two companies, AA and BB.

    Key: 1∣35∣31 \mid 35 \mid 3 means the pipe diameter from company AA is 0.351 cm0.351\text{ cm} and from company BB is 0.353 cm0.353\text{ cm}.

    Find the median and interquartile range of the pipes produced by company AA.

    The printed diagram. Company A is on the left: read each row from the stem outwards.

    The printed diagram. Company A is on the left: read each row from the stem outwards.

    Stuck? Show hint

    The stem 3535 with leaf 11 means 0.3510.351 cm. With n=19n = 19 you need the 5th, 10th and 15th values.

    Show solution
    1. 1

      Company AA's rows, read outwards, with running totals. 3333: 44 (total 11). 3434: 0,2,3,8,90, 2, 3, 8, 9 (total 66). 3535: 1,1,4,5,7,81, 1, 4, 5, 7, 8 (total 1212). 3636: 2,5,6,92, 5, 6, 9 (total 1616). 3737: 1,3,41, 3, 4 (total 1919).

      The key turns stem 35, leaf 1 into 0.351 cm.

    2. 2

      Median (10th). Row 3535 holds values 7 to 12; the 10th is its 4th leaf, 55: median=0.355 cm\text{median} = 0.355 \text{ cm}

      10 − 6 = 4.

    3. 3

      Q1Q_1 (5th). Row 3434 holds values 2 to 6; the 5th is its 4th leaf, 88: Q1=0.348 cmQ_1 = 0.348 \text{ cm}

      5 − 1 = 4.

    4. 4

      Q3Q_3 (15th). Row 3636 holds values 13 to 16; the 15th is its 3rd leaf, 66: Q3=0.366 cmQ_3 = 0.366 \text{ cm}

      15 − 12 = 3.

    5. 5

      IQR. IQR=0.366−0.348=0.018 cm\text{IQR} = 0.366 - 0.348 = 0.018 \text{ cm}

      Giving 355 and 18 (ignoring the key) earns only a special-case mark.

    Answer

    Median =0.355= 0.355 cm; IQR =0.366−0.348=0.018= 0.366 - 0.348 = 0.018 cm.

Practise medians, quartiles and measures of spreadReal past-paper questions · Measures of central tendency (mean, median, mode) and variation (range, IQR, standard deviation)
04

Box-and-whisker plots

Syllabus requirement · §5.1

“

draw and interpret stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs

”

A box-and-whisker plot (box plot for short) draws the five key values of §03 on a scale:

  • a box from Q1Q_1 to Q3Q_3, so the length of the box is the interquartile range;
  • a vertical line inside the box at the median;
  • whiskers: lines from the ends of the box out to the smallest and largest values.

So a box plot shows the centre (the median line), the spread of the middle half (the box) and the full extent of the data (the whiskers) in one small picture. Its purpose is comparison, which is why exam questions nearly always ask for two box plots on a single diagram.

614192736minQ₁medianQ₃maxinterquartile range = box widthrange = max − min

The five key values drawn as a box-and-whisker plot. The box spans the interquartile range, the line inside it marks the median, and the whiskers reach the smallest and largest values.

What the marks are for

Box plots are drawn on a printed grid, and the marks are awarded for accuracy and presentation:

  • one linear scale for both plots, with at least three values marked at equal spacing, and a label with units (e.g. "Time (minutes)");
  • all five key values of each plot drawn accurately on that scale;
  • each plot labelled with its group's name;
  • whiskers drawn from the middle of each end of the box, not from its corners, and not through the box.

Each of these can cost a mark on its own. If the median and quartiles are not given, finding them (§03) is usually worth a separate mark.

Drawing a pair of box plots

In §02, groups PP and QQ had these sorted puzzle times, in seconds:

P: 18, 22, 23, 26, 27, 29, 31, 35, 40Q: 19, 25, 28, 30, 33, 36, 38, 41, 44P:\ 18,\ 22,\ 23,\ 26,\ 27,\ 29,\ 31,\ 35,\ 40 \qquad Q:\ 19,\ 25,\ 28,\ 30,\ 33,\ 36,\ 38,\ 41,\ 44

Draw box-and-whisker plots for both groups on a single diagram.

Show full working
PQ15202530354045Time (seconds)

The finished diagram: one linear time scale, two labelled plots, whiskers from the middle of each end of the box.

  1. 1

    Positions for n=9n = 9. Median at 12(10)=5\tfrac12(10) = 5th. The lower half is the first 44 values and the upper half the last 44, so each quartile is the mean of the 2nd and 3rd values of its half.

    n = 9 is odd, so the median is left out of both halves; halves of 4 have no single middle value.

  2. 2

    PP's median. The 5th value: 2727.

    Four values below it and four above.

  3. 3

    PP's lower quartile. Lower half 18,22,23,2618, 22, 23, 26: Q1=22+232=22.5Q_1 = \frac{22 + 23}{2} = 22.5

    The mean of the 2nd and 3rd values of the half.

  4. 4

    PP's upper quartile. Upper half 29,31,35,4029, 31, 35, 40: Q3=31+352=33Q_3 = \frac{31 + 35}{2} = 33

    Same rule for the upper half.

  5. 5

    PP's extremes. Smallest 1818, largest 4040.

    Write all five values down before drawing anything.

  6. 6

    QQ's median. The 5th value: 3333.

    Same method for the second group.

  7. 7

    QQ's lower quartile. Lower half 19,25,28,3019, 25, 28, 30: Q1=25+282=26.5Q_1 = \frac{25 + 28}{2} = 26.5

    Mean of the 2nd and 3rd values of the half.

  8. 8

    QQ's upper quartile. Upper half 36,38,41,4436, 38, 41, 44: Q3=38+412=39.5Q_3 = \frac{38 + 41}{2} = 39.5

    Mean of the 2nd and 3rd values of the half.

  9. 9

    QQ's extremes. Smallest 1919, largest 4444.

    Now all ten key values are known.

  10. 10

    Draw one linear scale covering both groups: the smallest value overall is 1818 and the largest 4444, so a scale from 1515 to 4545 marked every 55 seconds works. Label it "Time (seconds)".

    A shared scale is what makes the comparison possible. The label with units is part of a mark.

  11. 11

    Draw PP's plot. Box from 22.522.5 to 3333, median line at 2727, whiskers from the middle of each end of the box out to 1818 and 4040. Label it "PP".

    Whiskers stop exactly at the smallest and largest values, and must not be drawn through the box.

  12. 12

    Draw QQ's plot below it on the same scale. Box from 26.526.5 to 39.539.5, median line at 3333, whiskers to 1919 and 4444. Label it "QQ".

    Stack the plots vertically so equal values line up; that is what lets the reader compare them by eye.

Answer

PP: 18, 22.5, 27, 33, 4018,\ 22.5,\ 27,\ 33,\ 40. QQ: 19, 26.5, 33, 39.5, 4419,\ 26.5,\ 33,\ 39.5,\ 44. Two labelled plots on one linear scale from 1515 to 4545 seconds.

List the five key values of each group in a small table first (min, Q₁, median, Q₃, max). The drawing is then just plotting ten numbers.

A pair of box plots from a past paper

9709/55 M/J 2025 Q5(b)4 marks

The Smarts and the Teasers are two quiz teams that each contain 1111 members. Both complete a puzzle and the following table gives the times taken, in minutes, by the members of each team.

Smarts383013291822281811941
Teasers3937183625253221151239

For the Teasers, the values of the lower quartile, median and upper quartile are 1818, 2525 and 3737 minutes respectively.

On a single diagram draw box-and-whisker plots for the two teams.

The grid printed with this part. You choose and label the scale.

The grid printed with this part. You choose and label the scale.

Show full working
The mark scheme's diagram: both plots on one labelled time scale.

The mark scheme's diagram: both plots on one labelled time scale.

  1. 1

    Sort the Smarts' times (as in §02): 9, 11, 13, 18, 18, 22, 28, 29, 30, 38, 419,\ 11,\ 13,\ 18,\ 18,\ 22,\ 28,\ 29,\ 30,\ 38,\ 41

    The Smarts' quartiles are not given, so they must be found. That is the first of the four marks.

  2. 2

    Smarts' median. n=11n = 11, so the median is the 6th value: 2222.

    Position ½(11 + 1) = 6.

  3. 3

    Smarts' lower quartile. Lower half 9,11,13,18,189, 11, 13, 18, 18, so Q1=13Q_1 = 13.

    The 3rd value of the full list.

  4. 4

    Smarts' upper quartile. Upper half 28,29,30,38,4128, 29, 30, 38, 41, so Q3=30Q_3 = 30.

    The 9th value. The mark scheme's first B1 is for LQ 13, M 22, UQ 30.

  5. 5

    Collect all ten key values. minQ1medianQ3maxSmarts913223041Teasers1218253739\begin{array}{l|ccccc} & \text{min} & Q_1 & \text{median} & Q_3 & \text{max} \\ \hline \text{Smarts} & 9 & 13 & 22 & 30 & 41 \\ \text{Teasers} & 12 & 18 & 25 & 37 & 39 \end{array}

    The Teasers' smallest (12) and largest (39) come from the table; their quartiles and median are given.

  6. 6

    Choose one scale. Values run from 99 to 4141, so a linear scale from 00 to 4545 (or 55 to 4545) marked every 55 minutes, labelled "Time (minutes)".

    One scale, at least three values marked, labelled with 'time' and 'minutes'.

  7. 7

    Draw the Smarts' plot and label it. Box 1313 to 3030, median line 2222, whiskers to 99 and 4141.

    B1 (follow-through) for the Smarts' five values plotted accurately.

  8. 8

    Draw the Teasers' plot and label it. Box 1818 to 3737, median line 2525, whiskers to 1212 and 3939.

    The last B1 needs two plots, whiskers not through the boxes nor from the corners, and the single labelled linear scale.

Answer

Smarts: 9, 13, 22, 30, 419,\ 13,\ 22,\ 30,\ 41; Teasers: 12, 18, 25, 37, 3912,\ 18,\ 25,\ 37,\ 39, drawn as two labelled box plots on one linear scale labelled "Time (minutes)".

When a question defines an "outlier"

An outlier is a value that lies an unusually long way from the rest of the data. Older papers (up to 2019) sometimes gave this definition in the question:

An outlier is any data value which is more than 1.51.5 times the interquartile range above the upper quartile, or more than 1.51.5 times the interquartile range below the lower quartile.

No 2021–2025 question has used it, and the current syllabus does not name it, but the method is quick if it ever appears. The definition sets two boundaries, sometimes called fences:

upper fence=Q3+1.5×IQR,lower fence=Q1−1.5×IQR\text{upper fence} = Q_3 + 1.5 \times \text{IQR}, \qquad \text{lower fence} = Q_1 - 1.5 \times \text{IQR}

A value above the upper fence or below the lower fence is an outlier. Only the most extreme values can possibly cross a fence, so check the largest and smallest values first. Recent papers ask the related question in words instead: "the mean is unduly affected by the extreme value $36 800" (§05).

Testing for outliers

The times, in minutes, taken by 1111 runners to complete a fun run are 12, 14, 15, 17, 18, 19, 21, 22, 23, 24, 4512,\ 14,\ 15,\ 17,\ 18,\ 19,\ 21,\ 22,\ 23,\ 24,\ 45 Using the definition above, determine whether there are any outliers.

Show full working
1015202530354045lower fence 3upper fence 3545 — outlierQ₁ = 15, Q₃ = 23, IQR = 8, so 1.5 × IQR = 12 → fences at 15 − 12 = 3 and 23 + 12 = 35

The eleven times on a number line with the two fences marked. Only 45 lies beyond a fence.

  1. 1

    Lower quartile. n=11n = 11 and the list is sorted, so Q1Q_1 is the 3rd value: Q1=15Q_1 = 15

    Every outlier test starts from correct quartiles (§03).

  2. 2

    Upper quartile. Q3Q_3 is the 9th value: Q3=23Q_3 = 23

    Position ¾(11 + 1) = 9.

  3. 3

    IQR. IQR=23−15=8\text{IQR} = 23 - 15 = 8

    The spread of the middle half.

  4. 4

    One and a half IQRs. 1.5×8=121.5 \times 8 = 12

    Work this out on its own: it is added to Q₃ and subtracted from Q₁.

  5. 5

    Upper fence. Q3+12=23+12=35Q_3 + 12 = 23 + 12 = 35

    Measured outwards from the upper quartile.

  6. 6

    Lower fence. Q1−12=15−12=3Q_1 - 12 = 15 - 12 = 3

    Measured outwards from the lower quartile.

  7. 7

    Compare the extremes with the fences. The smallest time is 12>312 > 3, so it is not an outlier. The largest is 45>3545 > 35, so it is an outlier.

    Every other value lies between these two, so nothing else can cross a fence.

Answer

4545 is an outlier (it is above the upper fence of 3535); there are no other outliers.

Common mistakes
  • Two box plots drawn on two different scales

    One linear scale for both plots, labelled with the quantity and units

    Plots on different scales cannot be compared by eye, and the scale mark is lost.

  • Whiskers drawn from the top and bottom corners of the box, or through the box

    Whiskers start at the middle of each end of the box and go outwards only

    Mark schemes penalise both explicitly.

  • Whiskers stopped at a rounded scale value instead of the actual smallest or largest value

    Whiskers end exactly at the minimum and maximum

    The end points are two of the five key values being tested.

  • Adding 1.5×IQR1.5 \times \text{IQR} to Q1Q_1 (or subtracting it from Q3Q_3)

    Upper fence =Q3+1.5 IQR= Q_3 + 1.5\,\text{IQR}; lower fence =Q1−1.5 IQR= Q_1 - 1.5\,\text{IQR}

    The fences lie outside the box. Moving towards the median puts them inside the data.

Your turn

For each plot: list the five key values first, then check your drawing against the four presentation points in the callout above.

  1. 19709/52 M/J 2024 Q4(b)3 marks

    The back-to-back stem-and-leaf diagram shows the annual salaries of 1919 employees at each of two companies, Petral and Ravon.

    Key: 2∣31∣52 \mid 31 \mid 5 means $31 200 for a Petral employee and $31 500 for a Ravon employee.

    The median salary of the Ravon employees is $33 800, the lower quartile is $32 000 and the upper quartile is $34 400.

    Represent the data shown in the back-to-back stem-and-leaf diagram by a pair of box-and-whisker plots in a single diagram.

    The stem-and-leaf diagram (the same one as in the §03 exercise).

    The stem-and-leaf diagram (the same one as in the §03 exercise).

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    Petral's median and quartiles were found in the §03 exercise. The smallest and largest salaries of both companies come from the first and last rows of the diagram.

    Show solution
    The mark scheme's pair of box plots.

    The mark scheme's pair of box plots.

    1. 1

      Petral's five values. From §03: Q1=31 200Q_1 = 31\,200, median =32 000= 32\,000, Q3=33 500Q_3 = 33\,500. From the diagram: smallest 30 00030\,000 (row 3030, leaf 00), largest 36 80036\,800 (row 3636, leaf 88).

      The extremes are the leaves nearest the stem in the first row and furthest from it in the last row.

    2. 2

      Ravon's five values. Given: Q1=32 000Q_1 = 32\,000, median =33 800= 33\,800, Q3=34 400Q_3 = 34\,400. From the diagram: smallest 30 20030\,200 (row 3030, first leaf 22), largest 36 90036\,900 (row 3636, last leaf 99).

      Ravon is on the right, read left to right as usual.

    3. 3

      Scale. Salaries run from $30 000 to $36 900: use a linear scale from $30 000 to $37 000, marked every $1000, labelled "Salary ($)".

      The mark scheme wants a scale no smaller than 1 cm to $1000, labelled "salaries" and $.

    4. 4

      Draw Petral's plot. Box 31 20031\,200 to 33 50033\,500, median line 32 00032\,000, whiskers to 30 00030\,000 and 36 80036\,800; label it P.

      B1 (follow-through from part (a)).

    5. 5

      Draw Ravon's plot on the same scale. Box 32 00032\,000 to 34 40034\,400, median line 33 80033\,800, whiskers to 30 20030\,200 and 36 90036\,900; label it R.

      B1; the third B1 is for whiskers not through the boxes and the single labelled scale.

    Answer

    Petral: 30 000, 31 200, 32 000, 33 500, 36 80030\,000,\ 31\,200,\ 32\,000,\ 33\,500,\ 36\,800; Ravon: 30 200, 32 000, 33 800, 34 400, 36 90030\,200,\ 32\,000,\ 33\,800,\ 34\,400,\ 36\,900, as two labelled box plots on one linear salary scale.

  2. 29709/53 M/J 2024 Q4(b)3 marks

    The times taken, in seconds, by 1515 members of each of two swimming clubs, the Penguins and the Dolphins, to swim 5050 metres are shown in the following table.

    Penguins353942444545485056585961666872
    Dolphins364143484949505154565660616471

    The diagram shows a box-and-whisker plot representing the times for the Penguins.

    On the same diagram, draw a box-and-whisker plot to represent the times for the Dolphins.

    The Penguins' box plot, already drawn on the scale you must use.

    The Penguins' box plot, already drawn on the scale you must use.

    Stuck? Show hint

    n=15n = 15: the median is the 8th time and the quartiles are the 4th and 12th.

    Show solution
    The mark scheme's diagram with the Dolphins' plot added.

    The mark scheme's diagram with the Dolphins' plot added.

    1. 1

      Dolphins' median. The Dolphins' times are sorted; the 8th is median=51\text{median} = 51

      Position ½(15 + 1) = 8.

    2. 2

      Dolphins' lower quartile. Lower half (first 7): 36,41,43,48,49,49,5036, 41, 43, 48, 49, 49, 50, so Q1=48Q_1 = 48.

      The 4th value of the half.

    3. 3

      Dolphins' upper quartile. Upper half (last 7): 54,56,56,60,61,64,7154, 56, 56, 60, 61, 64, 71, so Q3=60Q_3 = 60.

      The 4th value of the half.

    4. 4

      Extremes. Smallest 3636, largest 7171.

      The whisker end points, one of the three marks.

    5. 5

      Draw on the printed scale. Box from 4848 to 6060, median line at 5151, whiskers to 3636 and 7171, labelled "Dolphins".

      Use the scale that the Penguins' plot is already drawn on.

    Answer

    Dolphins: min 3636, Q1=48Q_1 = 48, median 5151, Q3=60Q_3 = 60, max 7171, drawn on the given scale and labelled.

  3. 39709/53 O/N 2025 Q6(c)3 marks

    Last Saturday, a cycling competition for teams of 1111 cyclists took place. For each cyclist, the time taken to complete the course was recorded to the nearest minute. The times taken by the cyclists from two teams, the Linnets and the Puffins, are shown in the following table.

    Linnets4851545759606464656870
    Puffins4549515555585962646474

    On the grid below, draw a box-and-whisker plot to represent the information for the Linnets and the Puffins.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    You found the Linnets' quartiles in the §03 exercise. For both teams, n=11n = 11: the median is the 6th value and the quartiles the 3rd and 9th.

    Show solution
    The mark scheme's pair of box plots.

    The mark scheme's pair of box plots.

    1. 1

      Linnets' five values. Smallest 4848, Q1=54Q_1 = 54, median 6060 (6th), Q3=65Q_3 = 65, largest 7070.

      The quartiles are the ones found in §03.

    2. 2

      Puffins' five values. Sorted already: smallest 4545, Q1=51Q_1 = 51 (3rd), median 5858 (6th), Q3=64Q_3 = 64 (9th), largest 7474.

      Same positions, since n = 11 again.

    3. 3

      Scale. A single linear scale from 4545 to 7575 minutes, marked every 55 minutes, labelled "Time (minutes)".

      The mark scheme wanted 2 cm to 10 or 20 minutes, at least three equally spaced values and a label with time and minutes.

    4. 4

      Draw the Linnets' plot. Box 5454 to 6565, median 6060, whiskers to 4848 and 7070; label it.

      Each plot must be labelled; whiskers not through the box.

    5. 5

      Draw the Puffins' plot on the same scale. Box 5151 to 6464, median 5858, whiskers to 4545 and 7474; label it.

      Same scale as the Linnets, or the comparison is lost.

    Answer

    Linnets: 48, 54, 60, 65, 7048,\ 54,\ 60,\ 65,\ 70; Puffins: 45, 51, 58, 64, 7445,\ 51,\ 58,\ 64,\ 74, as two labelled box plots on one linear time scale.

  4. 49709/61 O/N 2013 Q4(i)(ii)6 marks

    The following are the house prices in thousands of dollars, arranged in ascending order, for 5151 houses from a certain area.

    253 270 310 354 386 428 433 468 472 477 485 520 520 524 526 531 535
    536 538 541 543 546 548 549 551 554 572 583 590 605 614 638 649 652
    666 670 682 684 690 710 725 726 731 734 745 760 800 854 863 957 986

    (i) Draw a box-and-whisker plot to represent the data.

    (ii) An expensive house is defined as a house which has a price that is more than 1.51.5 times the interquartile range above the upper quartile.

    For the above data, give the prices of the expensive houses.

    Stuck? Show hint

    n=51n = 51: the median is the 26th price and the quartiles are the 13th and 39th. Each printed line holds 17 prices.

    Show solution
    The mark scheme's sketch for part (i). It is schematic and not drawn accurately to its own scale; the correct values are 253, 520, 554, 690 and 986.

    The mark scheme's sketch for part (i). It is schematic and not drawn accurately to its own scale; the correct values are 253, 520, 554, 690 and 986.

    1. 1

      (i) Median. n=51n = 51, so the median is at 12(52)=26\tfrac12(52) = 26th. Each printed line holds 1717 prices, so the 26th is the 9th on line 2: 554554.

      Counting in blocks of 17 is quicker than counting all the way from the start.

    2. 2

      (i) Lower quartile. 14(52)=13\tfrac14(52) = 13th, which is on line 1: Q1=520Q_1 = 520.

      n is odd, so the halves method gives a whole-number position.

    3. 3

      (i) Upper quartile. 34(52)=39\tfrac34(52) = 39th, the 5th on line 3: Q3=690Q_3 = 690.

      34 prices fill the first two lines, so the 39th is the 5th on line 3.

    4. 4

      (i) Draw the plot. Extremes 253253 and 986986. On a linear scale labelled "Price (thousands of dollars)": box from 520520 to 690690, median line at 554554, whiskers to 253253 and 986986.

      The units mark needs 'thousands of dollars' in the label or heading.

    5. 5

      (ii) IQR. IQR=690−520=170\text{IQR} = 690 - 520 = 170

      From the quartiles in part (i).

    6. 6

      (ii) One and a half IQRs. 1.5×170=2551.5 \times 170 = 255

      The M1 is for multiplying the IQR by 1.5.

    7. 7

      (ii) The boundary. Q3+255=690+255=945Q_3 + 255 = 690 + 255 = 945

      'More than 1.5 times the IQR above the upper quartile' means above 945.

    8. 8

      (ii) List the prices above the boundary. Only 957957 and 986986 exceed 945945.

      The next price down, 863, is below 945.

    Answer

    (i) Box plot with minimum 253253, Q1=520Q_1 = 520, median 554554, Q3=690Q_3 = 690, maximum 986986 (thousands of dollars). (ii) Prices above 945945: 957957 and 986986 thousand dollars.

Practise box-and-whisker plots and the other statistical diagramsReal past-paper questions · Stem-and-leaf diagrams, box-and-whisker plots, histograms, cumulative frequency graphs
05

Comparing data sets and choosing an average

Syllabus requirement · §5.1

“

understand and use different measures of central tendency (mean, median, mode) and variation (range, interquartile range, standard deviation) (e.g. in comparing and contrasting sets of data.)

”

Once two data sets have been summarised, the paper asks you to compare them, usually for 1 or 2 marks. The mark schemes are strict about what counts, and three rules cover almost every case.

Rule 1: one comment about centre, one about spread. If two comparisons are asked for, make one about the central tendency (which group is generally larger, taller, slower…) and one about the spread (which group is more consistent or more varied). Two comments about the median count as one idea.

Rule 2: write it in context. The comment must be about the actual quantity, in words. "The median of AA is higher" scores nothing; "the pipes from company AA generally have larger diameters" scores. Mark schemes say it directly: a simple numerical comparison of statistics is not enough. Quoting the numbers as support is fine, as long as the sentence says what they mean.

Rule 3: use the right everyday word. For times, a smaller median means faster or quicker. For heights, "taller". For spread, "more consistent" (smaller IQR or range) or "more varied / more spread out" (larger).

To decide which group is more spread out, use the IQR (or the standard deviation of §10, if that is what you have). The range depends on just two extreme values, so it can disagree with the IQR.

Scores

Does not score

“The Smarts were generally quicker than the Teasers.”

“The median for the Smarts is lower.” (no context)

“The Smarts' times were more consistent (IQR 17 min against 19 min).”

“Smarts 17, Teasers 19.” (numbers only)

“The range of weights of the Rebels is greater.”

“The range of the Rebels is greater.” (no quantity named)

“The Cheetahs' times were more spread out than the Panthers'.”

“The Cheetahs have a bigger IQR.” (a statistic, not the times)

Examples modelled on the mark schemes. Each comment in the left column names the quantity and says what the comparison means.

Writing two comparisons

In §04, the puzzle times of groups PP and QQ gave these key values, in seconds:

P: 18, 22.5, 27, 33, 40Q: 19, 26.5, 33, 39.5, 44P:\ 18,\ 22.5,\ 27,\ 33,\ 40 \qquad Q:\ 19,\ 26.5,\ 33,\ 39.5,\ 44

Make two comparisons between the times of the two groups.

Show full working
  1. 1

    Compare the centres. Median of P=27P = 27 s, median of Q=33Q = 33 s, so group PP's times are generally lower.

    Compare medians for the centre. Box plots and stem-and-leaf diagrams give the median directly.

  2. 2

    Translate into context. For times, lower means quicker: “Group P generally solved the puzzle more quickly than group Q.”\text{“Group } P \text{ generally solved the puzzle more quickly than group } Q.\text{”}

    This sentence is about the pupils and the puzzle, not about 'the median'. That is what earns the mark.

  3. 3

    PP's IQR. 33−22.5=10.5 s33 - 22.5 = 10.5 \text{ s}

    Work out both IQRs; don't judge them by eye.

  4. 4

    QQ's IQR. 39.5−26.5=13 s39.5 - 26.5 = 13 \text{ s}

    10.5 < 13, so P's middle half is less spread out.

  5. 5

    Translate into context. “Group P’s times were more consistent than group Q’s.”\text{“Group } P\text{'s times were more consistent than group } Q\text{'s.”}

    A second, different idea: consistency (spread), not speed (centre).

Answer

Group PP generally solved the puzzle more quickly (median 2727 s against 3333 s), and group PP's times were more consistent (IQR 10.510.5 s against 1313 s).

Two comparisons from a past paper

9709/55 M/J 2025 Q5(c)2 marks

The Smarts and the Teasers are two quiz teams that each contain 1111 members. Both complete a puzzle and the following table gives the times taken, in minutes, by the members of each team.

Smarts383013291822281811941
Teasers3937183625253221151239

(From part (b):) For the Teasers, the values of the lower quartile, median and upper quartile are 1818, 2525 and 3737 minutes respectively.

Make two comparisons between the times for the two teams.

Show full working
  1. 1

    Collect the key values (found in §04): Smarts Q1=13Q_1 = 13, median 2222, Q3=30Q_3 = 30; Teasers Q1=18Q_1 = 18, median 2525, Q3=37Q_3 = 37.

    Comparisons are made from statistics you already have.

  2. 2

    Centre. The Smarts' median, 2222 minutes, is lower than the Teasers' 2525 minutes, so “The Smarts were generally quicker at the puzzle.”\text{“The Smarts were generally quicker at the puzzle.”}

    Lower times means quicker. The mark scheme's answer is simply 'Smarts are quicker'.

  3. 3

    Smarts' IQR. 30−13=17 minutes30 - 13 = 17 \text{ minutes}

    Use the IQR. The ranges (41 − 9 = 32 and 39 − 12 = 27) point the other way, because the Smarts have the two most extreme times.

  4. 4

    Teasers' IQR. 37−18=19 minutes37 - 18 = 19 \text{ minutes}

    From the given quartiles.

  5. 5

    Spread, in context. 17<1917 < 19, so “The Smarts’ times were more consistent.”\text{“The Smarts' times were more consistent.”}

    The mark scheme's second comment, word for word.

Answer

The Smarts were generally quicker (median 2222 against 2525 minutes), and the Smarts' times were more consistent (IQR 1717 against 1919 minutes).

Decide spread with the IQR. It is the measure that goes with the median, and it is not thrown by one extreme time.

Mean or median?

The mean uses the size of every value. That is its strength, but it also means one extreme value drags it towards itself. The median only depends on which value is in the middle, so an extreme value barely moves it.

So when the data contains an extreme value, or is skewed (bunched at one end with a long tail at the other), the median is the more representative average, and it goes with the IQR as its measure of spread. When the data is roughly symmetrical with no extreme values, the mean (with the standard deviation) is fine and uses all the information.

The exam asks this in two ways:

  • "Is the median or the mean more suitable here?" Answer: the median, because the mean is affected by the extreme value, and name it, e.g. "the extreme value of 19.419.4 m".
  • "State what feature of the distribution accounts for the mean and median being different." Answer: the distribution is not symmetrical (it is skewed). If it were symmetrical, the mean and median would be close together.
$300$500$700$900$1100extreme value $1200median $330mean $453.57pulled towards $1200

Seven weekly wages. Six are between $310 and $350; one is $1200. The median ($330) stays in the middle of the cluster, but the mean ($453.57) is dragged towards the extreme value and sits above every wage except that one.

Mean or median with an extreme value

The weekly wages, in dollars, of 77 employees are 310, 320, 325, 330, 340, 350, 1200310,\ 320,\ 325,\ 330,\ 340,\ 350,\ 1200 Find the mean and the median, and say which is the better measure of a typical wage.

Show full working
  1. 1

    Add the wages. 310+320+325+330+340+350+1200=3175310 + 320 + 325 + 330 + 340 + 350 + 1200 = 3175

    The mean needs the total first.

  2. 2

    Divide by 77. xˉ=31757=453.57 (2 d.p.)\bar{x} = \frac{3175}{7} = 453.57 \text{ (2 d.p.)}

    Mean = total ÷ number of values. The symbol x̄ ("x-bar") is the usual name for the mean; §10 uses it throughout.

  3. 3

    Median. The wages are in order and n=7n = 7, so the median is the 4th: median=330\text{median} = 330

    Position ½(7 + 1) = 4.

  4. 4

    Compare each with the data. Six of the seven wages are between $310 and $350. The median, $330, is in the middle of them. The mean, $453.57, is higher than all six of them.

    Measure each average against the bulk of the data. That comparison shows which average is representative.

  5. 5

    Conclude, naming the extreme value. The median is the better measure: the mean is pulled up by the one extreme value of $1200, so it does not represent a typical wage.

    The mark needs the reason tied to this data: the specific extreme value, not a general statement.

Answer

Mean = $453.57, median = $330. The median is better, because the mean is distorted by the extreme value $1200.

positive skewmean > medianboth heresymmetricalmean = mediannegative skewmean < medianmedianmean

Shape and the two averages. In a symmetrical distribution the mean and median coincide. A long tail to the right (positive skew) pulls the mean above the median; a long tail to the left (negative skew) pulls it below. A mean and median that differ noticeably means the distribution is not symmetrical.

Mean, median and a reason

9709/53 M/J 2022 Q2(a)(b)(c)3 marks

Twenty children were asked to estimate the height of a particular tree. Their estimates, in metres, were as follows.

4.14.24.44.54.64.85.05.25.35.45.55.86.06.26.36.46.66.86.919.4\begin{matrix} 4.1 & 4.2 & 4.4 & 4.5 & 4.6 & 4.8 & 5.0 & 5.2 & 5.3 & 5.4 \\ 5.5 & 5.8 & 6.0 & 6.2 & 6.3 & 6.4 & 6.6 & 6.8 & 6.9 & 19.4 \end{matrix}

(a) Find the mean of the estimated heights.

(b) Find the median of the estimated heights.

(c) Give a reason why the median is likely to be more suitable than the mean as a measure of the central tendency for this information.

Show full working
  1. 1

    (a) Total of the first row. 4.1+4.2+4.4+4.5+4.6+4.8+5.0+5.2+5.3+5.4=47.54.1 + 4.2 + 4.4 + 4.5 + 4.6 + 4.8 + 5.0 + 5.2 + 5.3 + 5.4 = 47.5

    Adding row by row keeps the arithmetic checkable.

  2. 2

    (a) Total of the second row. 5.5+5.8+6.0+6.2+6.3+6.4+6.6+6.8+6.9+19.4=75.95.5 + 5.8 + 6.0 + 6.2 + 6.3 + 6.4 + 6.6 + 6.8 + 6.9 + 19.4 = 75.9

    The second row includes the extreme value 19.4.

  3. 3

    (a) Total of all 20. 47.5+75.9=123.447.5 + 75.9 = 123.4

    Add the two row totals.

  4. 4

    (a) Mean. xˉ=123.420=6.17 m\bar{x} = \frac{123.4}{20} = 6.17 \text{ m}

    The mark scheme shows 123.4 ÷ 20 = 6.17.

  5. 5

    (b) Median position. n=20n = 20, so 12(21)=10.5\tfrac12(21) = 10.5: halfway between the 10th and 11th values.

    n is even, so there is no single middle value.

  6. 6

    (b) Median. The data is printed in order; the 10th value is 5.45.4 (end of the first row) and the 11th is 5.55.5: median=5.4+5.52=5.45 m\text{median} = \frac{5.4 + 5.5}{2} = 5.45 \text{ m}

    Average the two middle values.

  7. 7

    (c) Compare with the bulk of the data. Nineteen of the estimates lie between 4.14.1 and 6.96.9 m; the mean, 6.176.17, has been pulled up towards one estimate of 19.419.4 m.

    This is the evidence: the mean sits high because of one value.

  8. 8

    (c) State the reason in context. The mean is unduly influenced by the extreme value 19.419.4 m, whereas the median is not.

    This is the mark scheme's answer, and it names the extreme value.

Answer

(a) 6.176.17 m. (b) 5.455.45 m. (c) The mean is unduly influenced by the extreme value 19.419.4 m; the median is not.

Common mistakes
  • "The median of the Cheetahs is lower."

    "The Cheetahs were generally faster than the Panthers."

    A comparison of statistics without context scores nothing.

  • Two comments about the average, when two comparisons are asked for

    One about the centre and one about the spread

    Two comments on the same idea earn one mark.

  • Saying a group with a bigger median time was "better" or "faster"

    For times, smaller is quicker; say which group was quicker

    Think about what the numbers measure before choosing the word.

  • "The median is better because it is more accurate."

    "The median is better because the mean is affected by the extreme value $1200."

    The reason must point to a feature of this data: the extreme value or the skew.

Your turn

Each comparison sentence needs context. Each reason needs a feature of this data.

  1. 19709/53 M/J 2023 Q4(b)2 marks

    The times taken, in minutes, to complete a cycle race by 1919 cyclists from each of two clubs, the Cheetahs and the Panthers, are represented in the following back-to-back stem-and-leaf diagram.

    Key: 7∣9∣17 \mid 9 \mid 1 means 9797 minutes for Cheetahs and 9191 minutes for Panthers

    The median and interquartile range for the Panthers are 103103 minutes and 1414 minutes respectively.

    Make two comparisons between the times taken by the Cheetahs and the times taken by the Panthers.

    The printed diagram.

    The printed diagram.

    Stuck? Show hint

    You found the Cheetahs' median (9999) and IQR (2323) in the §03 exercise.

    Show solution
    1. 1

      Centre. Cheetahs' median 9999 min, Panthers' 103103 min. Lower times are faster: "The Cheetahs were generally faster than the Panthers."

      A comparison of central tendency, in context.

    2. 2

      Spread. Cheetahs' IQR 2323 min, Panthers' 1414 min: "The Cheetahs' times were more spread out (less consistent) than the Panthers'."

      A comparison of spread, in context.

    Answer

    The Cheetahs were generally faster (median 9999 against 103103 min); the Cheetahs' times were more spread out (IQR 2323 against 1414 min).

  2. 29709/51 O/N 2022 Q3(b)(c)5 marks

    The Lions and the Tigers are two basketball clubs. The heights, in cm, of the 1111 players in each of their first team squads are given in the table.

    Lions178186181187179190189190180169196
    Tigers194179187190183201184180195191197

    (b) Find the median and the interquartile range of the heights of the Lions first team squad.

    It is given that for the Tigers, the lower quartile is 183183 cm, the median is 190190 cm and the upper quartile is 195195 cm.

    (c) Make two comparisons between the heights of the players in the Lions first team squad and the heights of the players in the Tigers first team squad.

    Stuck? Show hint

    Sort the Lions first; n=11n = 11, so use the 3rd, 6th and 9th values.

    Show solution
    1. 1

      (b) Sort the Lions. 169, 178, 179, 180, 181, 186, 187, 189, 190, 190, 196169,\ 178,\ 179,\ 180,\ 181,\ 186,\ 187,\ 189,\ 190,\ 190,\ 196

      The table row is not in order.

    2. 2

      (b) Median. The 6th value: 186186 cm.

      Position ½(11 + 1) = 6.

    3. 3

      (b) Lower quartile. The 3rd value: Q1=179Q_1 = 179.

      Middle of the lower half 169, 178, 179, 180, 181.

    4. 4

      (b) Upper quartile. The 9th value: Q3=190Q_3 = 190.

      Middle of the upper half 187, 189, 190, 190, 196.

    5. 5

      (b) IQR. 190−179=11 cm190 - 179 = 11 \text{ cm}

      Show the subtraction.

    6. 6

      (c) Centre. Tigers' median 190190 cm against the Lions' 186186 cm: "The Tigers are generally taller."

      Central tendency, in context.

    7. 7

      (c) Tigers' IQR. 195−183=12 cm195 - 183 = 12 \text{ cm}

      From the given quartiles.

    8. 8

      (c) Spread, in context. 1212 cm against the Lions' 1111 cm: "The Tigers' heights are slightly less consistent than the Lions'."

      The mark scheme also condoned 'similar spread', since 12 and 11 are close.

    Answer

    (b) Median 186186 cm; IQR =190−179=11= 190 - 179 = 11 cm. (c) The Tigers are generally taller; the Tigers' heights are slightly less consistent (IQR 1212 against 1111 cm).

  3. 39709/52 O/N 2023 Q4(c)1 mark

    The heights, in cm, of the 1111 players in each of two teams, the Aces and the Jets, are shown in the following table.

    Aces180174169182181166173182168171164
    Jets175174188168166174181181170188190

    Give one comment comparing the spread of the heights of the Aces with the spread of the heights of the Jets.

    Stuck? Show hint

    Sort both teams and find each IQR (or range). Then write one sentence about spread, in context.

    Show solution
    1. 1

      Sort both teams. Aces: 164,166,168,169,171,173,174,180,181,182,182164, 166, 168, 169, 171, 173, 174, 180, 181, 182, 182. Jets: 166,168,170,174,174,175,181,181,188,188,190166, 168, 170, 174, 174, 175, 181, 181, 188, 188, 190.

      n = 11 for both.

    2. 2

      Aces' IQR. Q3−Q1=181−168=13Q_3 - Q_1 = 181 - 168 = 13 cm.

      The 9th and 3rd values.

    3. 3

      Jets' IQR. 188−170=18188 - 170 = 18 cm.

      The ranges agree: 18 cm for the Aces, 24 cm for the Jets.

    4. 4

      Comment in context. "The Jets have a greater spread of heights than the Aces."

      Only one comment about spread is wanted. A comment about the average would score nothing here.

    Answer

    The Jets' heights are more spread out than the Aces' (IQR 1818 against 1313 cm).

  4. 49709/52 M/J 2024 Q4(c)1 mark

    The back-to-back stem-and-leaf diagram shows the annual salaries of 1919 employees at each of two companies, Petral and Ravon.

    Key: 2∣31∣52 \mid 31 \mid 5 means $31 200 for a Petral employee and $31 500 for a Ravon employee.

    Comment on whether the mean or the median would be a better representation of the data for the employees at Petral.

    The printed diagram.

    The printed diagram.

    Stuck? Show hint

    Look at the largest Petral salary compared with the rest.

    Show solution
    1. 1

      Look for an extreme value. Every Petral salary is between $30 000 and $34 100 except one: $36 800, alone on stem 3636.

      Row 35 is empty on Petral's side, so $36 800 stands apart from the rest.

    2. 2

      Conclude. The median is better, because the mean would be affected by the extreme value of $36 800.

      The mark scheme needs 'median', and a reference to the extreme value (or the skew) in context.

    Answer

    The median, because there is an extreme value ($36 800) that would distort the mean.

  5. 59709/52 M/J 2023 Q3(c)1 mark

    The following back-to-back stem-and-leaf diagram represents the monthly salaries, in dollars, of 2727 employees at each of two companies, AA and BB.

    Comment on whether the mean would be a more appropriate measure than the median for comparing the given information for the two companies.

    The printed diagram, with its key.

    The printed diagram, with its key.

    Stuck? Show hint

    Look at the largest salary on each side.

    Show solution
    1. 1

      Look for extreme values. Company BB's salaries run from $2540 to $2820, and then one salary of $3090 on stem 3030, with nothing in between on BB's side of row 2929.

      Company A's largest values continue smoothly, from 2950 to 3010, so it has no isolated value.

    2. 2

      Conclude. No: the mean would be less appropriate than the median, because company BB has an extreme value ($3090) that would distort its mean.

      The mark scheme needs a reference to company B (or $3090) and a statement that the mean is not appropriate, with no contradictory comment.

    Answer

    No. The median is more appropriate, because of the extreme value $3090 in company BB.

Practise comparing data sets and choosing measuresReal past-paper questions · Measures of central tendency (mean, median, mode) and variation (range, IQR, standard deviation)
06

Histograms

Syllabus requirement · §5.1

“

draw and interpret stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs

”

A histogram shows a grouped frequency table for a continuous quantity (a time, a length, a mass). It looks like a bar chart, but it differs in one important way: the area of each bar, not its height, represents the frequency.

Why? Grouped tables often have classes of different widths. Suppose the class 00–1010 minutes holds 2020 people and the class 4040–7070 minutes holds 3030. If bar height were frequency, the 4040–7070 bar would be taller and three times wider, so it would look far more than one-and-a-half times as important. Using area fixes this. The height is chosen so that height ×\times width == frequency, which means

frequency density=frequencyclass width\text{frequency density} = \frac{\text{frequency}}{\text{class width}}

and the vertical axis is labelled frequency density. Because the quantity is continuous, the bars touch: each bar runs exactly from its class's lower boundary to its upper boundary.

Reading a histogram backwards uses the same rule turned round: frequency=frequency density×class width.\text{frequency} = \text{frequency density} \times \text{class width}.

Class boundaries

Class widths must be measured between the true class boundaries, which are not always the numbers printed in the table. How to find them depends on how the data was recorded:

The table says

Example class

True boundaries

Width

inequalities

20⩽t<3020 \leqslant t < 30

2020 to 3030

1010

“correct to the nearest minute” (whole numbers)

121121–130130

120.5120.5 to 130.5130.5

1010

“to the nearest hundred”

900900–12001200

850850 to 12501250

400400

age “in completed years”

1010–1919

1010 to 2020

1010

Rounded data: move each printed limit out by half a unit of rounding (half a minute, 50 for “nearest hundred”). Truncated data such as completed years: the lower limit stays and the upper limit moves up a whole unit, because someone who is 19 years and 11 months is still recorded as 19.

Inequalities: 20 ≤ t < 30, 30 ≤ t < 40203040To the nearest minute: 121–130, 131–140120.5130.5140.5printed limits 121, 130 | 131, 140 move out by 0.5Age in completed years: 10–19, 20–29102030

The same printed classes, three different recording rules. For times to the nearest minute, 121–130 really covers 120.5 up to 130.5, so neighbouring bars meet at 130.5 with no gap.

What the four marks are for

"Draw a histogram" is almost always a 4-mark part:

  • M1 at least four frequency densities worked out (a table of widths and densities is the safest way to show them);
  • A1 every bar the correct height;
  • B1 bar ends at the correct class boundaries, on a linear horizontal scale;
  • B1 axes labelled: "frequency density" on the vertical axis, and the quantity with its units on the horizontal axis (e.g. "time (minutes)"), each with a linear scale.

A histogram with frequency on the vertical axis, gaps between the bars, or bars drawn from the printed class limits instead of the boundaries loses marks even if every number is right.

10204070246fd 2f = 20fd 5f = 50fd 3f = 60fd 1f = 30frequency density time (minutes)The trapThe 20–40 class has thelargest frequency (60),but it is not the tallestbar — it is just wider.frequency density= frequency ÷ widthArea = frequency20 × 3 = 60 ✓Bars touch: the scale iscontinuous, unlike abar chart.

The histogram of the first worked example below. The 20–40 class has the largest frequency (60), but its bar is not the tallest: it is twice as wide as the 10–20 class, so its height is only 60 ÷ 20 = 3.

Drawing a histogram with unequal classes

The times, tt minutes, that 160160 people spent in a museum are summarised below.

Time (tt minutes)0⩽t<100 \leqslant t < 1010⩽t<2010 \leqslant t < 2020⩽t<4020 \leqslant t < 4040⩽t<7040 \leqslant t < 70
Frequency2020505060603030

Draw a histogram to represent this information.

Show full working
  1. 1

    Class boundaries. The classes are written as inequalities, so the boundaries are exactly the printed numbers: 0,10,20,40,700, 10, 20, 40, 70.

    No adjustment is needed when the classes are given as inequalities.

  2. 2

    Class widths. 10−0=10,20−10=10,40−20=20,70−40=3010 - 0 = 10, \quad 20 - 10 = 10, \quad 40 - 20 = 20, \quad 70 - 40 = 30

    Work out the widths as their own step. A wrong width is otherwise invisible.

  3. 3

    Frequency density of the first two classes. 2010=2,5010=5\frac{20}{10} = 2, \qquad \frac{50}{10} = 5

    Frequency ÷ width, class by class.

  4. 4

    Frequency density of the last two classes. 6020=3,3030=1\frac{60}{20} = 3, \qquad \frac{30}{30} = 1

    The 20–40 class has the most people, 60, but a smaller density than the 10–20 class, because its 60 people are spread over twice the width.

  5. 5

    Set up the axes. Horizontal axis: "time (minutes)", linear from 00 to 7070. Vertical axis: "frequency density", linear from 00 to at least 55.

    Both labels, with units on the horizontal axis, are one of the four marks.

  6. 6

    Draw four touching bars. 00 to 1010 at height 22; 1010 to 2020 at height 55; 2020 to 4040 at height 33; 4040 to 7070 at height 11.

    Each bar runs exactly between its two boundaries, so neighbouring bars share an edge.

  7. 7

    Check with areas. 2×10=202 \times 10 = 20, 5×10=505 \times 10 = 50, 3×20=603 \times 20 = 60, 1×30=301 \times 30 = 30: every area gives back its frequency.

    Area = frequency is the defining property; this check catches a wrong width at once.

Answer

Frequency densities 2, 5, 3, 12,\ 5,\ 3,\ 1 on the classes 00–1010, 1010–2020, 2020–4040, 4040–7070; four touching bars on axes labelled "frequency density" and "time (minutes)".

Rounded data: finding the boundaries first

The heights of 7676 plants, measured correct to the nearest cm, are summarised below.

Height (cm)150150–159159160160–164164165165–169169170170–184184
Frequency1212202026261818

Find the class widths and frequency densities needed to draw a histogram.

Show full working
  1. 1

    Boundaries. "Correct to the nearest cm" means a height recorded as 159159 could really be anything from 158.5158.5 up to 159.5159.5. So each printed limit moves out by 0.50.5: 149.5, 159.5, 164.5, 169.5, 184.5149.5,\ 159.5,\ 164.5,\ 169.5,\ 184.5

    The upper boundary of one class is the lower boundary of the next, so the bars touch.

  2. 2

    Widths. 159.5−149.5=10,164.5−159.5=5,169.5−164.5=5,184.5−169.5=15159.5 - 149.5 = 10, \quad 164.5 - 159.5 = 5, \quad 169.5 - 164.5 = 5, \quad 184.5 - 169.5 = 15

    Using the printed limits (159 − 150 = 9) would make every width one too small.

  3. 3

    Frequency densities. 1210=1.2,205=4,265=5.2,1815=1.2\frac{12}{10} = 1.2, \quad \frac{20}{5} = 4, \quad \frac{26}{5} = 5.2, \quad \frac{18}{15} = 1.2

    The two narrow classes get the tallest bars.

Answer

Boundaries 149.5,159.5,164.5,169.5,184.5149.5, 159.5, 164.5, 169.5, 184.5; widths 10,5,5,1510, 5, 5, 15; frequency densities 1.2, 4, 5.2, 1.21.2,\ 4,\ 5.2,\ 1.2.

Lay the working out as a table with rows for boundaries, width, frequency and frequency density. Mark schemes are laid out the same way.

A histogram from data rounded to the nearest hundred

9709/51 M/J 2023 Q5(a)4 marks

The populations of 150150 villages in the UK, to the nearest hundred, are summarised in the table.

Population100 – 800900 – 12001300 – 20002100 – 32003300 – 4800
Number of villages812504832

Draw a histogram to represent this information.

The grid printed with this part.

The grid printed with this part.

Show full working
The mark scheme's histogram.

The mark scheme's histogram.

  1. 1

    Boundaries. "To the nearest hundred" means a population recorded as 800800 could be anything from 750750 up to 850850. So each printed limit moves out by 5050: 50, 850, 1250, 2050, 3250, 485050,\ 850,\ 1250,\ 2050,\ 3250,\ 4850

    Half the rounding unit: half of 100 is 50. This is the step the question is testing.

  2. 2

    Widths. 850−50=800,1250−850=400,2050−1250=800,3250−2050=1200,4850−3250=1600850 - 50 = 800, \quad 1250 - 850 = 400, \quad 2050 - 1250 = 800, \quad 3250 - 2050 = 1200, \quad 4850 - 3250 = 1600

    The mark scheme's widths are exactly these.

  3. 3

    Frequency densities of the first three classes. 8800=0.01,12400=0.03,50800=0.0625\frac{8}{800} = 0.01, \quad \frac{12}{400} = 0.03, \quad \frac{50}{800} = 0.0625

    Small numbers, because the widths are large; that is fine. The M1 needs at least four of them.

  4. 4

    Frequency densities of the last two classes. 481200=0.04,321600=0.02\frac{48}{1200} = 0.04, \quad \frac{32}{1600} = 0.02

    All five heights must be right for the A1.

  5. 5

    Axes. Horizontal: "population", linear, covering 5050 to 48504850. Vertical: "frequency density", linear from 00 to at least 0.06250.0625.

    The B1 for axes needs both labels and a vertical scale that reaches the tallest bar.

  6. 6

    Draw five touching bars with ends at 50,850,1250,2050,3250,485050, 850, 1250, 2050, 3250, 4850 and heights 0.01,0.03,0.0625,0.04,0.020.01, 0.03, 0.0625, 0.04, 0.02.

    The B1 for bar ends is read at the axis: the first bar starts at 50, not at 100.

Answer

Widths 800,400,800,1200,1600800, 400, 800, 1200, 1600; frequency densities 0.01,0.03,0.0625,0.04,0.020.01, 0.03, 0.0625, 0.04, 0.02; bars with ends at 50,850,1250,2050,3250,485050, 850, 1250, 2050, 3250, 4850.

Reading a frequency off a histogram

9709/62 O/N 2013 Q4(i)1 mark

The following histogram summarises the times, in minutes, taken by 190190 people to complete a race.

Show that 7575 people took between 200200 and 250250 minutes to complete the race.

The printed histogram.

The printed histogram.

Show full working
  1. 1

    Read the bar. The bar from 200200 to 250250 minutes has frequency density 1.51.5.

    Read the height off the vertical scale: 1.5 lies between the 1.4 and 1.6 grid lines.

  2. 2

    Width of the class. 250−200=50 minutes250 - 200 = 50 \text{ minutes}

    The width comes from the horizontal scale.

  3. 3

    Frequency = density × width. 1.5×50=751.5 \times 50 = 75

    The area of the bar is the frequency. The mark scheme requires 1.5 × 50 to be seen, since the answer 75 is given.

Answer

Frequency =1.5×50=75= 1.5 \times 50 = 75 people.

In a 'show that' part the working is the answer: write the density, the width and their product.

Common mistakes
  • Bar height == frequency

    Bar height == frequency density == frequency ÷\div class width

    Only frequency density makes the area represent the frequency when widths differ.

  • Width of the class 121121–130130 taken as 99

    Boundaries 120.5120.5 and 130.5130.5, width 1010

    Rounded data needs its boundaries before any width is measured.

  • Gaps between the bars, or bars starting at the printed class limits

    Bars touch, and run between the true boundaries

    The quantity is continuous. Gaps belong to bar charts.

  • Vertical axis labelled "frequency", or no units on the horizontal axis

    "Frequency density" and, e.g., "time (minutes)"

    The axes labels are a whole mark.

In the exam
10 of the 47 questions on this topic in 2021–2025 asked for a histogram, every time for 4 marks

Every one of those tables had classes of unequal width, and four of the ten gave rounded data ("to the nearest minute", "to the nearest hundred"), where the boundaries are not the printed numbers. The histogram part is usually followed by an estimate of the mean (§11) or a "which class contains…" part (§07) on the same table.

Your turn

Boundaries first, then widths, then densities, in a table. Check each bar with area = frequency.

  1. 1

    A histogram of the lengths of some fish has a bar from 2020 cm to 3535 cm of height 2.42.4, and a bar from 3535 cm to 4040 cm whose height is unknown. The 3535–4040 class contains 1919 fish.

    (a) How many fish are in the 2020–3535 class?

    (b) Find the height of the 3535–4040 bar.

    Stuck? Show hint

    Frequency = density × width, and density = frequency ÷ width.

    Show solution
    1. 1

      (a) Width. 35−20=1535 - 20 = 15 cm.

      Measured between the boundaries.

    2. 2

      (a) Frequency. 2.4×15=36 fish2.4 \times 15 = 36 \text{ fish}

      Area of the bar.

    3. 3

      (b) Width. 40−35=540 - 35 = 5 cm.

      A narrower class.

    4. 4

      (b) Frequency density. 195=3.8\frac{19}{5} = 3.8

      Height = frequency ÷ width.

    Answer

    (a) 3636 fish. (b) Height 3.83.8.

  2. 29709/52 F/M 2022 Q3(a)4 marks

    At a summer camp an arithmetic test is taken by 250250 children. The times taken, to the nearest minute, to complete the test were recorded. The results are summarised in the table.

    Time taken, in minutes1 – 3031 – 4546 – 6566 – 7576 – 100
    Frequency2130688645

    Draw a histogram to represent this information.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    To the nearest minute: the first class runs from 0.50.5 to 30.530.5.

    Show solution
    1. 1

      Boundaries. 0.5, 30.5, 45.5, 65.5, 75.5, 100.50.5,\ 30.5,\ 45.5,\ 65.5,\ 75.5,\ 100.5.

      Each printed limit moves out by half a minute.

    2. 2

      Widths. 30, 15, 20, 10, 2530,\ 15,\ 20,\ 10,\ 25.

      Differences of consecutive boundaries.

    3. 3

      Frequency densities. 2130=0.7,3015=2,6820=3.4,8610=8.6,4525=1.8\frac{21}{30} = 0.7,\quad \frac{30}{15} = 2,\quad \frac{68}{20} = 3.4,\quad \frac{86}{10} = 8.6,\quad \frac{45}{25} = 1.8

      These match the mark scheme.

    4. 4

      Draw. Five touching bars with ends at 0.5,30.5,45.5,65.5,75.5,100.50.5, 30.5, 45.5, 65.5, 75.5, 100.5 and heights 0.7,2,3.4,8.6,1.80.7, 2, 3.4, 8.6, 1.8, on axes labelled "frequency density" and "time (minutes)", both with linear scales.

      The mark scheme condones the first bar starting at 0.

    Answer

    Widths 30,15,20,10,2530, 15, 20, 10, 25; frequency densities 0.7,2,3.4,8.6,1.80.7, 2, 3.4, 8.6, 1.8; bars from 0.50.5 to 100.5100.5.

  3. 39709/51 M/J 2024 Q3(a)4 marks

    The heights, in cm, of 200200 adults in Barimba are summarised in the following table.

    Height (hh cm)130≤h<150130 \le h < 150150≤h<160150 \le h < 160160≤h<170160 \le h < 170170≤h<175170 \le h < 175175≤h<195175 \le h < 195
    Frequency1632766412

    Draw a histogram to represent this information.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    Inequality classes: the boundaries are the printed numbers.

    Show solution
    The mark scheme's histogram.

    The mark scheme's histogram.

    1. 1

      Widths. 20, 10, 10, 5, 2020,\ 10,\ 10,\ 5,\ 20.

      Straight from the inequalities.

    2. 2

      Frequency densities. 1620=0.8,3210=3.2,7610=7.6,645=12.8,1220=0.6\frac{16}{20} = 0.8,\quad \frac{32}{10} = 3.2,\quad \frac{76}{10} = 7.6,\quad \frac{64}{5} = 12.8,\quad \frac{12}{20} = 0.6

      The narrow 170–175 class gives the tallest bar.

    3. 3

      Draw. Bars with ends at 130,150,160,170,175,195130, 150, 160, 170, 175, 195 and heights 0.8,3.2,7.6,12.8,0.60.8, 3.2, 7.6, 12.8, 0.6; axes "frequency density" and "height (cm)".

      The mark scheme wants a horizontal scale no smaller than 1 cm to 10 cm.

    Answer

    Frequency densities 0.8,3.2,7.6,12.8,0.60.8, 3.2, 7.6, 12.8, 0.6 on bars 130130–150150–160160–170170–175175–195195.

  4. 49709/52 O/N 2025 Q5(a)4 marks

    The times of 240240 competitors taking part in an event are recorded correct to the nearest minute. The results are summarised in the table.

    Time (minutes)1 – 1011 – 2021 – 2526 – 3031 – 50
    Frequency1238687646

    Draw a histogram to represent this information.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    Boundaries 0.5,10.5,20.5,25.5,30.5,50.50.5, 10.5, 20.5, 25.5, 30.5, 50.5.

    Show solution
    The mark scheme's histogram.

    The mark scheme's histogram.

    1. 1

      Boundaries. 0.5, 10.5, 20.5, 25.5, 30.5, 50.50.5,\ 10.5,\ 20.5,\ 25.5,\ 30.5,\ 50.5.

      Half a minute out from each printed limit.

    2. 2

      Widths. 10, 10, 5, 5, 2010,\ 10,\ 5,\ 5,\ 20.

      Differences of consecutive boundaries.

    3. 3

      Frequency densities. 1210=1.2,3810=3.8,685=13.6,765=15.2,4620=2.3\frac{12}{10} = 1.2,\quad \frac{38}{10} = 3.8,\quad \frac{68}{5} = 13.6,\quad \frac{76}{5} = 15.2,\quad \frac{46}{20} = 2.3

      These match the mark scheme.

    4. 4

      Draw. Touching bars from 0.50.5 to 50.550.5 at these heights, on axes labelled "frequency density" and "time (minutes)".

      The mark scheme wants each axis to use at least half the grid.

    Answer

    Widths 10,10,5,5,2010, 10, 5, 5, 20; frequency densities 1.2,3.8,13.6,15.2,2.31.2, 3.8, 13.6, 15.2, 2.3.

  5. 59709/63 O/N 2013 Q12 marks

    The distance of a student's home from college, correct to the nearest kilometre, was recorded for each of 5555 students. The distances are summarised in the following table.

    Distance from college (km)1 – 34 – 56 – 89 – 1112 – 16
    Number of students18138124

    Dominic is asked to draw a histogram to illustrate the data. Dominic's diagram is shown below.

    Give two reasons why this is not a correct histogram.

    Dominic's diagram.

    Dominic's diagram.

    Stuck? Show hint

    Check the vertical axis label, the gaps, and where each bar starts and stops.

    Show solution
    1. 1

      Reason 1: the bar heights. The vertical axis is "number of students", so the heights are frequencies. The classes have different widths (3,2,3,3,53, 2, 3, 3, 5 km), so the heights should be frequency densities; as drawn, the areas do not represent the frequencies.

      For example, the 12–16 bar is made to look larger than its 4 students justify, because it is the widest.

    2. 2

      Reason 2: the gaps. There are gaps between the bars. The distances are continuous and rounded to the nearest km, so the class 11–33 runs from 0.50.5 to 3.53.5 and the next begins at 3.53.5: the bars should touch.

      The bars are drawn between the printed limits (1 to 3, 4 to 5, …) instead of the boundaries. Gaps and wrong class limits count as the same mark-scheme point, so one of your two reasons must be about frequency density.

    Answer

    One reason of each kind: (1) the heights are frequencies, not frequency densities, so the areas do not represent the frequencies (the vertical axis should be frequency density); (2) there are gaps between the bars, which should touch, running between the class boundaries 0.50.5, 3.53.5, 5.55.5, 8.58.5, 11.511.5, 16.516.5.

Practise histograms and the other statistical diagramsReal past-paper questions · Stem-and-leaf diagrams, box-and-whisker plots, histograms, cumulative frequency graphs
07

Grouped data: the median class and the greatest possible IQR

Syllabus requirement · §5.1

“

understand and use different measures of central tendency (mean, median, mode) and variation (range, interquartile range, standard deviation)

”

With a grouped frequency table the individual values are gone, so the median and quartiles cannot be found exactly. Two kinds of question are still possible, and both appear regularly after a histogram part:

  1. "Which class interval contains the median (or the lower or upper quartile)?" (1 mark)
  2. "Find the greatest possible value of the interquartile range." (2 marks)

Both rest on running totals: add up the frequencies class by class. Each running total is the number of values up to the end of that class. (This is the cumulative frequency of §08.)

For grouped data, the median is the 12n\tfrac12 nth value, the lower quartile the 14n\tfrac14 nth and the upper quartile the 34n\tfrac34 nth. (No "+1+1" here: with grouped data we can only estimate, treating the values as spread evenly through each class, and then the halfway point is simply 12n\tfrac12 n. The 12(n+1)\tfrac12(n+1) rule of §03 is for short lists of individual values.) The class containing one of these is the first class whose running total reaches or passes that position.

The greatest and least possible interquartile range

Once you know which class holds Q1Q_1 and which holds Q3Q_3, you know each quartile only somewhere inside its class. So:

  • the IQR is largest when Q3Q_3 is at the top of its class and Q1Q_1 is at the bottom of its class: greatest IQR=upper boundary of the Q3 class−lower boundary of the Q1 class\text{greatest IQR} = \text{upper boundary of the } Q_3 \text{ class} - \text{lower boundary of the } Q_1 \text{ class}
  • the IQR is smallest when Q3Q_3 is at the bottom of its class and Q1Q_1 at the top of its class: least IQR=lower boundary of the Q3 class−upper boundary of the Q1 class\text{least IQR} = \text{lower boundary of the } Q_3 \text{ class} - \text{upper boundary of the } Q_1 \text{ class} (If both quartiles are in the same class, the least possible IQR is 00.)

Use the true class boundaries (§06), not the printed limits. For a count such as a population, which must be a whole number, the top of a class is one less than its upper boundary (e.g. 32493249 rather than 32503250 for a class rounded to 21002100–32003200); mark schemes accept either.

greatest possible IQR = 10 − 2 = 8least = 6 − 4 = 2f = 6f = 18Q₁ classf = 21f = 22Q₃ classf = 1302461015kgrunning total624456780Q₁ is the 20th parcel (6 < 20 ≤ 24); Q₃ is the 60th (45 < 60 ≤ 67)

The worked example below. Q₁ lies somewhere in the 2–4 class and Q₃ somewhere in the 6–10 class. Pushing Q₁ to the bottom of its class and Q₃ to the top gives the greatest possible IQR, 10 − 2 = 8; pushing them towards each other gives the least, 6 − 4 = 2.

Median class, quartile classes and the IQR limits

The masses, mm kg, of 8080 parcels are summarised below.

Mass (mm kg)0⩽m<20 \leqslant m < 22⩽m<42 \leqslant m < 44⩽m<64 \leqslant m < 66⩽m<106 \leqslant m < 1010⩽m<1510 \leqslant m < 15
Frequency661818212122221313

(a) Which class contains the median?

(b) Find the greatest and least possible values of the interquartile range.

Show full working
  1. 1

    Running totals. 6,6+18=24,24+21=45,45+22=67,67+13=806, \quad 6 + 18 = 24, \quad 24 + 21 = 45, \quad 45 + 22 = 67, \quad 67 + 13 = 80

    Each total is the number of parcels up to the end of that class. The last one must equal n = 80.

  2. 2

    (a) Median position. 12n=12×80=40\tfrac12 n = \tfrac12 \times 80 = 40

    For grouped data use ½n, ¼n and ¾n.

  3. 3

    (a) Locate it. The running total is 2424 after the 22–44 class and 4545 after the 44–66 class. The 40th parcel lies after the 24th and by the 45th, so it is in the 4⩽m<64 \leqslant m < 6 class.

    The median class is the first one whose running total reaches 40.

  4. 4

    (b) Quartile positions. 14×80=20,34×80=60\tfrac14 \times 80 = 20, \qquad \tfrac34 \times 80 = 60

    The same idea, for the quartiles.

  5. 5

    (b) Locate Q1Q_1. The 20th parcel: the running total is 66 after the first class and 2424 after the second, so Q1Q_1 is in the 2⩽m<42 \leqslant m < 4 class.

    6 < 20 ≤ 24.

  6. 6

    (b) Locate Q3Q_3. The 60th parcel: the running total is 4545 after the third class and 6767 after the fourth, so Q3Q_3 is in the 6⩽m<106 \leqslant m < 10 class.

    45 < 60 ≤ 67.

  7. 7

    (b) Greatest possible IQR. Q3Q_3 at the top of its class, Q1Q_1 at the bottom of its class: 10−2=8 kg10 - 2 = 8 \text{ kg}

    Largest possible Q₃ minus smallest possible Q₁.

  8. 8

    (b) Least possible IQR. Q3Q_3 at the bottom of its class, Q1Q_1 at the top of its class: 6−4=2 kg6 - 4 = 2 \text{ kg}

    Smallest possible Q₃ minus largest possible Q₁.

Answer

(a) 4⩽m<64 \leqslant m < 6. (b) Greatest possible IQR =8= 8 kg; least possible IQR =2= 2 kg.

The greatest possible IQR from a past paper

9709/51 M/J 2021 Q5(c)2 marks

The times taken by 200200 players to solve a computer puzzle are summarised in the following table.

Time (tt seconds)0≤t<100 \le t < 1010≤t<2010 \le t < 2020≤t<4020 \le t < 4040≤t<6040 \le t < 6060≤t<10060 \le t < 100
Number of players1654783220

Find the greatest possible value of the interquartile range of these times.

Show full working
  1. 1

    Running totals. 16,70,148,180,20016, \quad 70, \quad 148, \quad 180, \quad 200

    16 + 54 = 70, 70 + 78 = 148, 148 + 32 = 180, 180 + 20 = 200.

  2. 2

    Quartile positions. 14×200=50,34×200=150\tfrac14 \times 200 = 50, \qquad \tfrac34 \times 200 = 150

    ¼n and ¾n for grouped data.

  3. 3

    Class of Q1Q_1. 16<50⩽7016 < 50 \leqslant 70, so Q1Q_1 is in the 10⩽t<2010 \leqslant t < 20 class.

    The running total first reaches 50 in the second class.

  4. 4

    Class of Q3Q_3. 148<150⩽180148 < 150 \leqslant 180, so Q3Q_3 is in the 40⩽t<6040 \leqslant t < 60 class.

    The running total is only 148 at the end of the 20–40 class, so the 150th time is just into the next class.

  5. 5

    Greatest possible IQR. 60−10=50 seconds60 - 10 = 50 \text{ seconds}

    Top of the Q₃ class minus bottom of the Q₁ class. The mark scheme also condones 49.9 recurring, since t < 60.

Answer

60−10=5060 - 10 = 50 seconds.

Watch the borderline: 148 is just short of 150, so Q₃ is in the 40–60 class, not the 20–40 class. Always compare the position with the running totals explicitly.

Common mistakes
  • Giving the class with the largest frequency as the median class

    The median class is where the running total first reaches 12n\tfrac12 n

    The class with the most values tells you where values bunch (for equal widths it is the modal class), not where the middle value is.

  • Greatest IQR == top of the Q3Q_3 class −- top of the Q1Q_1 class

    Top of the Q3Q_3 class −- bottom of the Q1Q_1 class

    To make a difference as large as possible, make the larger number as large as possible and the smaller number as small as possible.

  • Using printed limits (e.g. 13001300 and 32003200) for rounded data

    Use the true boundaries (e.g. 12501250 and 32503250, or 32493249 for whole-number counts)

    The quartile could be any value that rounds into the class.

Your turn

Write the running totals out every time, then compare each position with them.

  1. 19709/52 F/M 2022 Q3(b)(c)2 marks

    At a summer camp an arithmetic test is taken by 250250 children. The times taken, to the nearest minute, to complete the test were recorded. The results are summarised in the table.

    Time taken, in minutes1 – 3031 – 4546 – 6566 – 7576 – 100
    Frequency2130688645

    (b) State which class interval contains the median.

    (c) Given that an estimate of the mean time is 61.0561.05 minutes, state what feature of the distribution accounts for the median and the mean being different.

    Stuck? Show hint

    Median position 12×250=125\tfrac12 \times 250 = 125. For (c), think back to §05.

    Show solution
    1. 1

      (b) Running totals. 21, 51, 119, 205, 25021,\ 51,\ 119,\ 205,\ 250.

      Up to the end of each class.

    2. 2

      (b) Locate the median. Position 125125: 119<125⩽205119 < 125 \leqslant 205, so the median is in the 6666–7575 class.

      The mark scheme also condones 65.5–75.5.

    3. 3

      (c) Compare. The median is somewhere between 65.565.5 and 75.575.5, but the mean is 61.0561.05, noticeably lower.

      A mean below the median suggests a long tail of smaller values.

    4. 4

      (c) Name the feature. The distribution is not symmetrical (it is skewed).

      The mark scheme ignores which way the skew goes; 'not symmetrical' or 'skewed' is enough.

    Answer

    (b) 6666–7575. (c) The distribution is not symmetrical (it is skewed).

  2. 29709/51 M/J 2023 Q5(b)(c)3 marks

    The populations of 150150 villages in the UK, to the nearest hundred, are summarised in the table.

    Population100 – 800900 – 12001300 – 20002100 – 32003300 – 4800
    Number of villages812504832

    (b) Write down the class interval which contains the median for this information.

    (c) Find the greatest possible value of the interquartile range for the populations of the 150150 villages.

    Stuck? Show hint

    The boundaries were found in §06: 50,850,1250,2050,3250,485050, 850, 1250, 2050, 3250, 4850.

    Show solution
    1. 1

      Running totals. 8, 20, 70, 118, 1508,\ 20,\ 70,\ 118,\ 150.

      8 + 12 = 20, 20 + 50 = 70, 70 + 48 = 118.

    2. 2

      (b) Median. Position 7575: 70<75⩽11870 < 75 \leqslant 118, so the median is in the 21002100–32003200 class.

      The mark scheme also accepts 2050–3250.

    3. 3

      (c) Class of Q1Q_1. Position 14×150=37.5\tfrac14 \times 150 = 37.5: 20<37.5⩽7020 < 37.5 \leqslant 70, so the 13001300–20002000 class.

      The running total first passes 37.5 in the third class.

    4. 4

      (c) Class of Q3Q_3. Position 34×150=112.5\tfrac34 \times 150 = 112.5: 70<112.5⩽11870 < 112.5 \leqslant 118, so the 21002100–32003200 class.

      The two quartiles are in neighbouring classes.

    5. 5

      (c) Smallest possible Q1Q_1. 12501250, the smallest population that rounds to 13001300.

      The bottom of the Q₁ class.

    6. 6

      (c) Largest possible Q3Q_3. 32493249, the largest whole-number population that rounds to 32003200 (32503250 would round up to 33003300).

      Populations are whole numbers, so the top of the class is 3249 rather than 3250 (see the note on counts above).

    7. 7

      (c) Greatest possible IQR. 3249−1250=19993249 - 1250 = 1999

      The mark scheme's answer; it also condones 3250 − 1250 = 2000.

    Answer

    (b) 21002100–32003200. (c) 3249−1250=19993249 - 1250 = 1999 (the boundary value 20002000 is condoned).

  3. 39709/51 M/J 2024 Q3(b)2 marks

    The heights, in cm, of 200200 adults in Barimba are summarised in the following table.

    Height (hh cm)130≤h<150130 \le h < 150150≤h<160150 \le h < 160160≤h<170160 \le h < 170170≤h<175170 \le h < 175175≤h<195175 \le h < 195
    Frequency1632766412

    The interquartile range is RR cm. Show that RR is not greater than 1515.

    Stuck? Show hint

    Find the classes containing the 50th and the 150th heights.

    Show solution
    1. 1

      Running totals. 16, 48, 124, 188, 20016,\ 48,\ 124,\ 188,\ 200.

      Up to the end of each class.

    2. 2

      Class of Q1Q_1. The 50th height: 48<50⩽12448 < 50 \leqslant 124, so 160⩽h<170160 \leqslant h < 170.

      ¼ × 200 = 50.

    3. 3

      Class of Q3Q_3. The 150th height: 124<150⩽188124 < 150 \leqslant 188, so 170⩽h<175170 \leqslant h < 175.

      ¾ × 200 = 150. The M1 is for identifying both classes.

    4. 4

      Greatest possible IQR. 175−160=15175 - 160 = 15 so RR cannot be greater than 1515.

      The largest Q₃ can be is 175 and the smallest Q₁ can be is 160.

    Answer

    Q1Q_1 is in 160⩽h<170160 \leqslant h < 170 and Q3Q_3 in 170⩽h<175170 \leqslant h < 175, so R⩽175−160=15R \leqslant 175 - 160 = 15.

  4. 49709/52 F/M 2024 Q3(c)1 mark

    The times taken, in minutes, by 150150 students to complete a puzzle are summarised in the table.

    Time taken (tt minutes)0≤t<200 \le t < 2020≤t<3020 \le t < 3030≤t<3530 \le t < 3535≤t<4035 \le t < 4040≤t<5040 \le t < 5050≤t<7050 \le t < 70
    Frequency82335522012

    In which class interval does the lower quartile of the times lie?

    Stuck? Show hint

    14×150=37.5\tfrac14 \times 150 = 37.5.

    Show solution
    1. 1

      Running totals. 8, 31, 66, 118, 138, 1508,\ 31,\ 66,\ 118,\ 138,\ 150.

      Only the first few are needed.

    2. 2

      Locate Q1Q_1. Position 37.537.5: 31<37.5⩽6631 < 37.5 \leqslant 66, so Q1Q_1 is in the 30⩽t<3530 \leqslant t < 35 class.

      The mark scheme condones '3rd interval' or '30–35'.

    Answer

    30⩽t<3530 \leqslant t < 35.

Practise measures of centre and spread from grouped dataReal past-paper questions · Measures of central tendency (mean, median, mode) and variation (range, IQR, standard deviation)
08

Drawing cumulative frequency graphs

Syllabus requirement · §5.1

“

draw and interpret stem-and-leaf diagrams, box-and-whisker plots, histograms and cumulative frequency graphs

”

The cumulative frequency at a value xx is the number of observations less than or equal to xx: the running total of §07. A cumulative frequency graph plots these running totals, and from it you can read how many values lie below any value you like (§09).

Three rules decide where the points go:

  1. Plot each running total at the upper class boundary of its class. The running total for the class 20⩽t<3020 \leqslant t < 30 counts everything up to 3030, so it belongs at t=30t = 30, not at the midpoint 2525.
  2. Start the curve at zero, at the lower boundary of the first class. Nothing has been counted before the first class begins, so the first point is (lower boundary of the first class, 00). It is (0,0)(0, 0) only when the first class starts at 00. One exception: if that boundary would be negative for a quantity that cannot be negative (a time, a distance), start at 00 instead; for 00–44 km to the nearest km the curve starts at (0,0)(0, 0), not (−0.5,0)(-0.5, 0). And if a cumulative table has no column with frequency 00, start at the smallest possible value, usually (0,0)(0, 0).
  3. Join the points with a smooth, increasing curve. Mark schemes usually withhold the final accuracy mark for straight line segments, and the curve must never go above the total nn.

Sometimes the table is already cumulative, with headings like "t⩽20t \leqslant 20". Then the numbers in it are the running totals, and each heading gives the plotting position directly.

What the marks are for

  • The cumulative frequencies (cf) correct (often a B1 of their own, when you have to work them out).
  • Axes labelled "cumulative frequency" and the quantity with units, each with a linear scale with at least three values marked, covering the whole range of the data (and using at least half the grid).
  • Points plotted at the upper boundaries.
  • A smooth curve through all the points, joined to the correct starting point and staying at or below nn.

From a frequency table to a cumulative frequency graph

The lengths, correct to the nearest cm, of 120120 leaves are summarised below.

Length (cm)1010–14141515–19192020–24242525–29293030–3939
Frequency882222404032321818

Draw a cumulative frequency graph to illustrate the data.

Show full working
1015202530354020406080100120(14.5, 8)(19.5, 30)(24.5, 70)(29.5, 102)(39.5, 120)start(9.5, 0)length (cm)cumulative frequency

The finished graph: six points at the upper class boundaries, starting from (9.5, 0), joined by a smooth increasing curve that ends at (39.5, 120).

  1. 1

    Upper class boundaries. The lengths are rounded to the nearest cm, so the class 1010–1414 really runs from 9.59.5 to 14.514.5. The upper boundaries are 14.5, 19.5, 24.5, 29.5, 39.514.5,\ 19.5,\ 24.5,\ 29.5,\ 39.5

    Same boundary rule as for histograms (§06): half a unit out from each printed limit.

  2. 2

    Running totals. 8,8+22=30,30+40=70,70+32=102,102+18=1208, \quad 8 + 22 = 30, \quad 30 + 40 = 70, \quad 70 + 32 = 102, \quad 102 + 18 = 120

    Add each new frequency to the previous total. The last total must equal n = 120.

  3. 3

    Starting point. The first class begins at its lower boundary, 9.59.5, so the curve starts at (9.5,0)(9.5, 0).

    No leaf is shorter than 9.5 cm, so the cumulative frequency is 0 there. Starting at (0, 0) would be wrong.

  4. 4

    Points to plot. (9.5,0), (14.5,8), (19.5,30), (24.5,70), (29.5,102), (39.5,120)(9.5, 0),\ (14.5, 8),\ (19.5, 30),\ (24.5, 70),\ (29.5, 102),\ (39.5, 120)

    Each running total paired with its upper boundary.

  5. 5

    Axes. Horizontal: "length (cm)", linear, covering 9.59.5 to 39.539.5. Vertical: "cumulative frequency", linear from 00 to 120120.

    The scales must be linear. Don't bunch the uneven boundaries (the last class is twice as wide) into equal spacing.

  6. 6

    Draw the curve. Plot the six points and join them with a smooth curve that always rises (or stays level), ending at (39.5,120)(39.5, 120).

    Steepest where the class frequencies are largest (the 20–24 class, with 40 leaves).

Answer

Points (9.5,0),(14.5,8),(19.5,30),(24.5,70),(29.5,102),(39.5,120)(9.5, 0), (14.5, 8), (19.5, 30), (24.5, 70), (29.5, 102), (39.5, 120), joined by a smooth increasing curve, on labelled linear axes.

Plot the starting point on purpose. It is the point candidates most often put in the wrong place, at (0, 0) or at (10, 0).

A cumulative frequency graph from rounded data

9709/52 F/M 2025 Q3(a)4 marks

The lengths of 250250 leaves of a certain type of plant are measured, correct to the nearest centimetre. The results are summarised in the table below.

Length (cm)5 – 910 – 1415 – 1920 – 2425 – 2930 – 39
Frequency182860724824

On the grid below, draw a cumulative frequency graph to illustrate this information.

The grid printed with this part. You choose and label both scales.

The grid printed with this part. You choose and label both scales.

Show full working
The mark scheme's graph, starting at (4.5, 0). The horizontal line at 155 is the reading for part (b), used in the §09 exercises.

The mark scheme's graph, starting at (4.5, 0). The horizontal line at 155 is the reading for part (b), used in the §09 exercises.

  1. 1

    Upper boundaries. Nearest cm, so half a unit out: 9.5, 14.5, 19.5, 24.5, 29.5, 39.59.5,\ 14.5,\ 19.5,\ 24.5,\ 29.5,\ 39.5

    The M1 is for plotting at these upper end points.

  2. 2

    Running totals. 18,46,106,178,226,25018, \quad 46, \quad 106, \quad 178, \quad 226, \quad 250

    18 + 28 = 46, 46 + 60 = 106, 106 + 72 = 178, 178 + 48 = 226, 226 + 24 = 250. The first B1 is for these.

  3. 3

    Starting point. The first class, 55–99, begins at 4.54.5, so the curve starts at (4.5,0)(4.5, 0).

    The mark scheme's final A1 requires the curve to be joined to (4.5, 0).

  4. 4

    Axes. Vertical "cumulative frequency", linear from 00 to 250250; horizontal "length (cm)", linear, covering at least 55 to 39.539.5.

    The second B1: both labels, linear scales with at least three values marked, using more than half the grid.

  5. 5

    Plot the points. (4.5,0),(9.5,18),(14.5,46),(19.5,106),(24.5,178),(29.5,226),(39.5,250)(4.5, 0), (9.5, 18), (14.5, 46), (19.5, 106), (24.5, 178), (29.5, 226), (39.5, 250).

    The M1 needs at least four of these at the upper end points.

  6. 6

    Draw the curve. Join the points with a smooth increasing curve that does not go above 250250.

    The A1 is lost for straight line segments.

Answer

Points (4.5,0),(9.5,18),(14.5,46),(19.5,106),(24.5,178),(29.5,226),(39.5,250)(4.5, 0), (9.5, 18), (14.5, 46), (19.5, 106), (24.5, 178), (29.5, 226), (39.5, 250), joined by a smooth increasing curve on labelled linear axes.

A graph from a cumulative frequency table

9709/53 O/N 2023 Q4(a)2 marks

The weights, x kgx\text{ kg}, of 120120 students in a sports college are recorded. The results are summarised in the following table.

Weight (x kgx\text{ kg})x≤40x \le 40x≤60x \le 60x≤65x \le 65x≤70x \le 70x≤85x \le 85x≤100x \le 100
Cumulative frequency0143860106120

Draw a cumulative frequency graph to represent this information.

The grid printed with this part.

The grid printed with this part.

Show full working
The mark scheme's two acceptable versions (horizontal scale from 0 or from 40); both curves start at (40, 0).

The mark scheme's two acceptable versions (horizontal scale from 0 or from 40); both curves start at (40, 0).

  1. 1

    Read the points straight from the table. Each heading "x⩽x \leqslant value" gives the plotting position, and the numbers are already cumulative: (40,0), (60,14), (65,38), (70,60), (85,106), (100,120)(40, 0),\ (60, 14),\ (65, 38),\ (70, 60),\ (85, 106),\ (100, 120)

    No adding up is needed. Don't add the cumulative frequencies again.

  2. 2

    Starting point. The table itself says no student weighs 4040 kg or less, so the curve starts at (40,0)(40, 0), not at the origin.

    The mark scheme's A1 requires the curve joined to (40, 0).

  3. 3

    Axes. Horizontal "weight (kg)", linear, covering 4040 to 100100; vertical "cumulative frequency", linear from 00 to 120120.

    The M1 needs at least three points plotted accurately on linear scales with at least three values marked on each.

  4. 4

    Draw. A smooth increasing curve through all six points.

    Label both axes. That is part of the A1.

Answer

Points (40,0),(60,14),(65,38),(70,60),(85,106),(100,120)(40, 0), (60, 14), (65, 38), (70, 60), (85, 106), (100, 120), joined by a smooth increasing curve on labelled linear axes.

Common mistakes
  • Plotting running totals at class midpoints or lower boundaries

    Plot each running total at the upper boundary of its class

    The running total only includes the whole class once you reach its top end.

  • Starting every curve at (0,0)(0, 0)

    Start at (lower boundary of the first class, 00), e.g. (4.5,0)(4.5, 0) or (40,0)(40, 0)

    Only a first class that begins at 0 starts at the origin.

  • Adding up a table that is already cumulative

    If the headings read "x⩽…x \leqslant \ldots", the numbers are already running totals

    Adding them again gives totals far bigger than n.

  • Joining the points with straight line segments

    A smooth increasing curve through the points

    Mark schemes routinely give A0 for ruled segments.

In the exam
10 of the 47 questions on this topic in 2021–2025 asked for a cumulative frequency graph (2 to 4 marks)

Six of the ten gave the data as a cumulative frequency table ("t⩽20t \leqslant 20" headings), where the points can be read straight off; the other four gave an ordinary frequency table that had to be totalled first. Every one was followed by a reading from the graph (§09), and most by a grouped mean or standard deviation (§11).

Your turn

For each graph, write the list of points (including the starting point) before drawing anything.

  1. 19709/52 F/M 2021 Q5(a)4 marks

    A driver records the distance travelled in each of 150150 journeys. These distances, correct to the nearest km, are summarised in the following table.

    Distance (km)0 – 45 – 1011 – 2021 – 3031 – 4041 – 60
    Frequency12163266204

    Draw a cumulative frequency graph to illustrate the data.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    Upper boundaries 4.5,10.5,20.5,30.5,40.5,60.54.5, 10.5, 20.5, 30.5, 40.5, 60.5. A distance cannot be negative, so the curve starts at the origin.

    Show solution
    The mark scheme's graph (a low-resolution scan; the extra line is the reading for a later part).

    The mark scheme's graph (a low-resolution scan; the extra line is the reading for a later part).

    1. 1

      Upper boundaries. 4.5, 10.5, 20.5, 30.5, 40.5, 60.54.5,\ 10.5,\ 20.5,\ 30.5,\ 40.5,\ 60.5.

      Half a km above each printed upper limit.

    2. 2

      Running totals. 12, 28, 60, 126, 146, 15012,\ 28,\ 60,\ 126,\ 146,\ 150.

      12 + 16 = 28, + 32 = 60, + 66 = 126, + 20 = 146, + 4 = 150. The first B1 is for these.

    3. 3

      Starting point. Distances start at 00 km (none can be negative), so the curve starts at (0,0)(0, 0).

      The mark scheme's A1 requires the curve joined to (0, 0).

    4. 4

      Points. (0,0),(4.5,12),(10.5,28),(20.5,60),(30.5,126),(40.5,146),(60.5,150)(0, 0), (4.5, 12), (10.5, 28), (20.5, 60), (30.5, 126), (40.5, 146), (60.5, 150).

      Each running total at its upper boundary.

    5. 5

      Draw. Axes "distance (km)" from 00 to 6060 and "cumulative frequency" from 00 to 150150; a smooth increasing curve through the points.

      Linear scales on both axes (the second B1); no straight segments.

    Answer

    Points (0,0),(4.5,12),(10.5,28),(20.5,60),(30.5,126),(40.5,146),(60.5,150)(0, 0), (4.5, 12), (10.5, 28), (20.5, 60), (30.5, 126), (40.5, 146), (60.5, 150), joined smoothly on labelled linear axes.

  2. 29709/52 F/M 2023 Q1(a)3 marks

    Each year the total number of hours, xx, of sunshine in Kintoo is recorded during the month of June. The results for the last 6060 years are summarised in the table.

    xx30≤x<6030 \le x < 6060≤x<9060 \le x < 9090≤x<11090 \le x < 110110≤x<140110 \le x < 140140≤x<180140 \le x < 180180≤x≤240180 \le x \le 240
    Number of years48142572

    Draw a cumulative frequency graph to illustrate the data.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    The first class starts at 3030, so the curve starts at (30,0)(30, 0).

    Show solution
    The mark scheme's graph, starting at (30, 0) (a low-resolution scan; the extra line is the reading for a later part).

    The mark scheme's graph, starting at (30, 0) (a low-resolution scan; the extra line is the reading for a later part).

    1. 1

      Running totals. 4, 12, 26, 51, 58, 604,\ 12,\ 26,\ 51,\ 58,\ 60.

      The B1 is for all of these.

    2. 2

      Points. (30,0),(60,4),(90,12),(110,26),(140,51),(180,58),(240,60)(30, 0), (60, 4), (90, 12), (110, 26), (140, 51), (180, 58), (240, 60).

      Inequality classes: the upper boundaries are the printed numbers.

    3. 3

      Draw. A smooth increasing curve through the points; axes "hours of sunshine" (linear, 3030 to 240240) and "cumulative frequency" (linear, 00 to 6060).

      The A1 needs the curve joined to (30, 0) with no ruled segments.

    Answer

    Points (30,0),(60,4),(90,12),(110,26),(140,51),(180,58),(240,60)(30, 0), (60, 4), (90, 12), (110, 26), (140, 51), (180, 58), (240, 60) joined smoothly.

  3. 39709/53 O/N 2024 Q4(a)4 marks

    On a certain day, the heights of 150150 sunflower plants grown by children at a local school are measured, correct to the nearest cm. These heights are summarised in the following table.

    Height (cm)10–1920–2930–3940–4445–4950–5455–59
    Frequency1018324228146

    Draw a cumulative frequency graph to illustrate the data.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    Upper boundaries 19.5,29.5,…,59.519.5, 29.5, \ldots, 59.5; the curve starts at (9.5,0)(9.5, 0).

    Show solution
    The mark scheme's graph.

    The mark scheme's graph.

    1. 1

      Upper boundaries. 19.5, 29.5, 39.5, 44.5, 49.5, 54.5, 59.519.5,\ 29.5,\ 39.5,\ 44.5,\ 49.5,\ 54.5,\ 59.5.

      Nearest cm: half a unit out.

    2. 2

      Running totals. 10, 28, 60, 102, 130, 144, 15010,\ 28,\ 60,\ 102,\ 130,\ 144,\ 150.

      The first B1 is for these.

    3. 3

      Points. (9.5,0),(19.5,10),(29.5,28),(39.5,60),(44.5,102),(49.5,130),(54.5,144),(59.5,150)(9.5, 0), (19.5, 10), (29.5, 28), (39.5, 60), (44.5, 102), (49.5, 130), (54.5, 144), (59.5, 150).

      Start at the lower boundary of the first class.

    4. 4

      Draw. Axes "height (cm)" from 9.59.5 to 59.559.5 and "cumulative frequency" from 00 to 150150, both linear; a smooth increasing curve through the points.

      A0 for straight segments or a curve that rises above 150.

    Answer

    Points (9.5,0),(19.5,10),(29.5,28),(39.5,60),(44.5,102),(49.5,130),(54.5,144),(59.5,150)(9.5, 0), (19.5, 10), (29.5, 28), (39.5, 60), (44.5, 102), (49.5, 130), (54.5, 144), (59.5, 150), joined smoothly.

  4. 49709/52 M/J 2025 Q5(a)2 marks

    The times taken, tt minutes, by 300300 students to travel to Hollowton College are recorded. The results are summarised in the table below.

    Time (tt minutes)t≤10t \leq 10t≤20t \leq 20t≤30t \leq 30t≤40t \leq 40t≤60t \leq 60t≤90t \leq 90
    Cumulative frequency3486142208265300

    On the grid, draw a cumulative frequency graph to illustrate this information.

    The grid printed with this part.

    The grid printed with this part.

    Stuck? Show hint

    The table is already cumulative. Times cannot be negative, so start at the origin.

    Show solution
    1. 1

      Points. (0,0),(10,34),(20,86),(30,142),(40,208),(60,265),(90,300)(0, 0), (10, 34), (20, 86), (30, 142), (40, 208), (60, 265), (90, 300).

      Read straight from the headings; start at (0, 0).

    2. 2

      Draw. Linear axes "time (minutes)" 00 to 9090 and "cumulative frequency" 00 to 300300, using over half the grid; a smooth curve through the points, staying below 300300 until t=90t = 90.

      Bars score B0 on the first mark; the second B1 needs a curve with no line segments, joined to (0, 0).

    Answer

    Points (0,0),(10,34),(20,86),(30,142),(40,208),(60,265),(90,300)(0, 0), (10, 34), (20, 86), (30, 142), (40, 208), (60, 265), (90, 300), joined smoothly.

Practise cumulative frequency graphs and the other statistical diagramsReal past-paper questions · Stem-and-leaf diagrams, box-and-whisker plots, histograms, cumulative frequency graphs
09

Reading a cumulative frequency graph

Syllabus requirement · §5.1

“

use a cumulative frequency graph (e.g. to estimate medians, quartiles, percentiles, the proportion of a distribution above (or below) a given value, or between two values.)

”

A cumulative frequency graph turns a count into a value, and a value into a count:

  • count → value: find the count on the vertical axis, go across to the curve, then down to the horizontal axis;
  • value → count: find the value on the horizontal axis, go up to the curve, then across to the vertical axis.

The ppth percentile is the value with p%p\% of the data at or below it: the median is the 5050th percentile, Q1Q_1 the 2525th and Q3Q_3 the 7575th. Below, "cf" is short for cumulative frequency.

The only real skill is turning the words of the question into the right count. Everything is measured from the bottom of the data, because a cumulative frequency counts the values below a point:

The question says

Read the graph at cumulative frequency…

the median

12n\tfrac12 n

the lower / upper quartile

14n\tfrac14 n / 34n\tfrac34 n

the ppth percentile

p100 n\tfrac{p}{100}\, n

x%x\% of the values are kk or more (or longer than kk)

100−x100 n\tfrac{100 - x}{100}\, n, then read off kk

NN of the values are more than kk

n−Nn - N, then read off kk

Wording that counts from the top (“or more”, “longer than”, “more than”) must be turned round, because the graph counts from the bottom.

The question asks for

Read

the number of values below kk

the cumulative frequency at kk

the number of values above kk

n−(cumulative frequency at k)n - (\text{cumulative frequency at } k)

the number of values between aa and bb

(cf at b)−(cf at a)(\text{cf at } b) - (\text{cf at } a)

a percentage or proportion

the number found, ÷ n\div\, n (then ×100\times 100 for a percentage)

Value → count questions. Read the cumulative frequency at each value off the curve, then combine.

Show the reading on the graph

Mark schemes give the method mark for the right count seen or shown on the graph (e.g. "0.65×120=780.65 \times 120 = 78"), and the accuracy mark only if use of the graph is seen. Draw the across-and-down lines on the grid, and write the count you used. A correct number with no visible reading can drop to a single special-case mark.

On a cumulative frequency graph the quartiles are at 14n\tfrac14 n and 34n\tfrac34 n, with no "+1". The 12(n+1)\tfrac12(n+1) positions of §03 are for small lists of individual values.

1015202530354020406080100120length (cm)cumulative frequencyQ₁: cf 30 → 19.5 cmmedian: cf 60 → 23.2 cmQ₃: cf 90 → 27.3 cm30% are k or more:cf 0.7 × 120 = 84 → 26.4 cm

The leaves graph from §08 (n = 120) with four readings: Q₁ at cumulative frequency 30, the median at 60 and Q₃ at 90, plus the reading at 84 used for “30% of the leaves are k cm or longer”.

Median, quartiles and a percentile from a graph

Use the cumulative frequency graph of the lengths of the 120120 leaves in §08 to estimate (a) the median, (b) the interquartile range, (c) the 9090th percentile.

Show full working
  1. 1

    (a) Count for the median. 12×120=60\tfrac12 \times 120 = 60

    ½n, with no '+1', on a cumulative frequency graph.

  2. 2

    (a) Read across and down. Across from 6060 on the vertical axis to the curve, then down to the length axis: median≈23.2 cm\text{median} \approx 23.2 \text{ cm}

    The curve passes through (19.5, 30) and (24.5, 70), so a reading between 19.5 and 24.5 is expected.

  3. 3

    (b) Count for Q1Q_1. 14×120=30\tfrac14 \times 120 = 30

    ¼n on a graph.

  4. 4

    (b) Read Q1Q_1. At cumulative frequency 3030 the curve is exactly at a plotted point: Q1=19.5Q_1 = 19.5 cm.

    Each quartile gets its own across-and-down line.

  5. 5

    (b) Count for Q3Q_3. 34×120=90\tfrac34 \times 120 = 90

    ¾n on a graph.

  6. 6

    (b) Read Q3Q_3. Across from 9090 and down: Q3≈27.3Q_3 \approx 27.3 cm.

    Between (24.5, 70) and (29.5, 102).

  7. 7

    (b) Subtract. IQR≈27.3−19.5=7.8 cm\text{IQR} \approx 27.3 - 19.5 = 7.8 \text{ cm}

    Show the two readings and the subtraction.

  8. 8

    (c) Count for the 9090th percentile. 90100×120=108\tfrac{90}{100} \times 120 = 108

    The pth percentile has p% of the data at or below it.

  9. 9

    (c) Read across and down. 90th percentile≈31.3 cm90\text{th percentile} \approx 31.3 \text{ cm}

    Between 29.5 (cf 102) and 39.5 (cf 120), close to 29.5 because the curve is still fairly steep there.

Answer

(a) Median ≈23.2\approx 23.2 cm. (b) IQR ≈27.3−19.5=7.8\approx 27.3 - 19.5 = 7.8 cm. (c) 9090th percentile ≈31.3\approx 31.3 cm.

Readings from a hand-drawn curve are estimates. Mark schemes accept a range around the true value, but only if the reading lines are visible.

“Or more”, “more than” and “between”

Using the same graph of the 120120 leaves:

(a) 30%30\% of the leaves have length kk cm or more. Estimate kk.

(b) Estimate how many leaves are longer than 3232 cm.

(c) Estimate how many leaves have lengths between 2222 cm and 3232 cm.

Show full working
  1. 1

    (a) Turn the wording round. If 30%30\% are kk or more, then 70%70\% are below kk. The count to read at is 0.7×120=840.7 \times 120 = 84

    The graph counts from the bottom, so 'the top 30%' means reading at 70%. Reading at 0.3 × 120 = 36 is the classic error.

  2. 2

    (a) Read across and down. k≈26.4k \approx 26.4

    In an exam, write the 84 and draw the reading line: past mark schemes give one mark for the count and one for a value read from the graph.

  3. 3

    (b) Read up and across at 3232 cm. The cumulative frequency at 3232 is about 110110: about 110110 leaves are 3232 cm or shorter.

    Value → count: up from 32 to the curve, then across.

  4. 4

    (b) Subtract from the total. 120−110=10 leaves120 - 110 = 10 \text{ leaves}

    'Longer than' counts from the top, so subtract from n.

  5. 5

    (c) Read the cumulative frequency at 2222 cm. About 4949.

    About 49 leaves are 22 cm or shorter.

  6. 6

    (c) Subtract the two readings (110110 from part (b)). 110−49=61 leaves110 - 49 = 61 \text{ leaves}

    Those up to 32 cm, minus those up to 22 cm, leaves those in between.

Answer

(a) k≈26.4k \approx 26.4. (b) About 1010 leaves. (c) About 6161 leaves.

Interquartile range and a “longer than” reading

9709/51 O/N 2023 Q1(a)(b)4 marks

The times taken by 120120 children to complete a particular puzzle are represented in the cumulative frequency graph.

(a) Use the graph to estimate the interquartile range of the data.

(b) 35%35\% of the children took longer than TT seconds to complete the puzzle. Use the graph to estimate the value of TT.

The printed graph. Draw your reading lines on it.

The printed graph. Draw your reading lines on it.

Show full working
  1. 1

    (a) Count for Q1Q_1. 14×120=30\tfrac14 \times 120 = 30

    n = 120 is given in the stem.

  2. 2

    (a) Read the lower quartile. Across from 3030 to the curve and down: Q1≈23.7Q_1 \approx 23.7 s.

    The curve is very steep here, so read carefully. The mark scheme accepts 23.25 < LQ ≤ 24.

  3. 3

    (a) Count for Q3Q_3. 34×120=90\tfrac34 \times 120 = 90

    ¾n.

  4. 4

    (a) Read the upper quartile. Across from 9090 to the curve and down: Q3≈31Q_3 \approx 31 s.

    Accepted range 30.5 < UQ < 31.25.

  5. 5

    (a) Subtract. IQR≈31−23.7=7.3 s\text{IQR} \approx 31 - 23.7 = 7.3 \text{ s}

    Answers from 7.0 to 7.5 were accepted, with graph use seen.

  6. 6

    (b) Turn the wording round. 35%35\% took longer than TT, so 65%65\% took TT or less: 0.65×120=780.65 \times 120 = 78

    The first B1 is for 78, seen or shown on the graph.

  7. 7

    (b) Read across and down from 7878. T≈28.5 secondsT \approx 28.5 \text{ seconds}

    Accepted range 28 < T < 29.

Answer

(a) IQR ≈31−23.7=7.3\approx 31 - 23.7 = 7.3 s (accept 7.07.0 to 7.57.5). (b) T≈28.5T \approx 28.5 (accept between 2828 and 2929).

A box-and-whisker plot from a cumulative frequency graph

9709/53 M/J 2025 Q4(a)(b)5 marks

8484 people attempt a particular puzzle. The times taken, in minutes, to complete the puzzle are recorded. These times are represented in the cumulative frequency graph below.

(a) Use the graph to estimate how many people took between 44 and 7.57.5 minutes to complete the puzzle.

(b) On the grid below, draw a box-and-whisker plot to summarise the information in the cumulative frequency graph.

Fig. 4.1: the cumulative frequency graph.

Fig. 4.1: the cumulative frequency graph.

Fig. 4.2: the grid for the box plot.

Fig. 4.2: the grid for the box plot.

Show full working
The mark scheme's box-and-whisker plot.

The mark scheme's box-and-whisker plot.

  1. 1

    (a) Cumulative frequency at 7.57.5 minutes. Up from 7.57.5 to the curve and across: about 6969.

    About 69 people took 7.5 minutes or less.

  2. 2

    (a) Cumulative frequency at 44 minutes. Up from 44 and across: about 25.525.5.

    About 25 or 26 people took 4 minutes or less.

  3. 3

    (a) Subtract. 69−25.5≈43.5, so about 43 or 44 people69 - 25.5 \approx 43.5, \text{ so about } 43 \text{ or } 44 \text{ people}

    The mark scheme accepts 43 or 44.

  4. 4

    (b) Count for the median. 12×84=42\tfrac12 \times 84 = 42

    ½n on a graph.

  5. 5

    (b) Read the median. Across from 4242 and down: median≈4.8 minutes\text{median} \approx 4.8 \text{ minutes}

    A B1 on its own, for the median plotted on the box.

  6. 6

    (b) Count for Q1Q_1. 14×84=21\tfrac14 \times 84 = 21

    ¼n on a graph.

  7. 7

    (b) Read Q1Q_1. Across from 2121 and down: Q1≈3.7Q_1 \approx 3.7 (or 3.83.8).

    The mark scheme accepts either.

  8. 8

    (b) Count for Q3Q_3. 34×84=63\tfrac34 \times 84 = 63

    ¾n on a graph.

  9. 9

    (b) Read Q3Q_3. Across from 6363 and down: Q3≈6.5Q_3 \approx 6.5 (or 6.66.6).

    A second B1, for both quartiles plotted on the box.

  10. 10

    (b) Smallest and largest times. The curve starts at 22 minutes (cumulative frequency 00) and reaches 8484 at 1212 minutes, so the whiskers end at 22 and 1212.

    These are the lowest and highest times the graph allows. The true shortest and longest times are not known, so the whisker ends are estimates, like the quartiles.

  11. 11

    (b) Draw the plot. On a linear scale from 22 to 1212, labelled "time (minutes)": box from 3.73.7 to 6.56.5, median line at 4.84.8, whiskers to 22 and 1212.

    The last B1 is for the linear, labelled scale (time and minutes, at least three equally spaced values).

Answer

(a) About 4343 or 4444 people. (b) Box plot with minimum 22, Q1≈3.7Q_1 \approx 3.7, median ≈4.8\approx 4.8, Q3≈6.5Q_3 \approx 6.5, maximum 1212, on a labelled linear time scale.

Common mistakes
  • Reading at 0.3n0.3n for "30%30\% are kk or more"

    Read at 0.7n0.7n: 70%70\% are below kk

    The graph counts from the bottom.

  • Giving the cumulative frequency at kk as the number above kk

    Number above kk =n−(cf at k)= n - (\text{cf at } k)

    The cumulative frequency is the number at or below k.

  • Using 14(n+1)\tfrac14(n+1) for a quartile on a cumulative frequency graph

    Use 14n\tfrac14 n, 12n\tfrac12 n, 34n\tfrac34 n

    The '+1' belongs to small ordered lists.

  • A correct value with no reading lines on the graph

    Draw the across-and-down lines and state the count used

    Use of the graph must be seen for the accuracy mark.

Your turn

For each part: write the count you will read at, draw the lines, then read. For graphs you draw yourself, the answers are the mark-scheme values; your own reading should be close.

  1. 19709/53 M/J 2021 Q1(a)(b)(c)5 marks

    The heights in cm of 160160 sunflower plants were measured. The results are summarised on the following cumulative frequency curve.

    (a) Use the graph to estimate the number of plants with heights less than 100100 cm.

    (b) Use the graph to estimate the 6565th percentile of the distribution.

    (c) Use the graph to estimate the interquartile range of the heights of these plants.

    The printed curve.

    The printed curve.

    Stuck? Show hint

    (a) is a value → count reading. (b) needs 0.65×1600.65 \times 160. (c) needs readings at 4040 and 120120.

    Show solution
    1. 1

      (a) Read up from 100100 cm and across. About 6060 plants.

      Value → count. The mark scheme accepts 60 or 61.

    2. 2

      (b) Count. 0.65×160=1040.65 \times 160 = 104.

      The M1 is for 104, seen or used on the graph.

    3. 3

      (b) Read. Across from 104104 and down: 6565th percentile ≈136\approx 136 cm.

      Use of the graph must be seen.

    4. 4

      (c) Count for Q1Q_1. 14×160=40\tfrac14 \times 160 = 40.

      No '+1' on a graph.

    5. 5

      (c) Count for Q3Q_3. 34×160=120\tfrac34 \times 160 = 120.

      ¾n.

    6. 6

      (c) Read Q3Q_3. Across from 120120 and down: Q3≈150Q_3 \approx 150 cm.

      Accepted range 148 ≤ UQ ≤ 152.

    7. 7

      (c) Read Q1Q_1. Across from 4040 and down: Q1≈76Q_1 \approx 76 cm.

      Accepted range 74 ≤ LQ ≤ 78.

    8. 8

      (c) Subtract. IQR≈150−76=74 cm\text{IQR} \approx 150 - 76 = 74 \text{ cm}

      The A1 must come from 150 − 76.

    Answer

    (a) About 6060. (b) About 136136 cm. (c) IQR ≈150−76=74\approx 150 - 76 = 74 cm.

  2. 29709/53 O/N 2023 Q4(b)2 marks

    The weights, x kgx\text{ kg}, of 120120 students in a sports college are recorded. The results are summarised in the following table.

    Weight (x kgx\text{ kg})x≤40x \le 40x≤60x \le 60x≤65x \le 65x≤70x \le 70x≤85x \le 85x≤100x \le 100
    Cumulative frequency0143860106120

    It is found that 35%35\% of the students weigh more than W kgW\text{ kg}.

    Use your graph to estimate the value of WW.

    Stuck? Show hint

    This is the graph drawn in the third worked example of §08 ("A graph from a cumulative frequency table"). "More than WW" counts from the top.

    Show solution
    1. 1

      Turn the wording round. 35%35\% weigh more than WW, so 65%65\% weigh WW or less: 0.65×120=780.65 \times 120 = 78

      The M1 is for 78.

    2. 2

      Read. Across from 7878 to the curve and down: W≈76W \approx 76 kg.

      Accepted range 75 < W < 79, with the reading shown on the graph.

    Answer

    W≈76W \approx 76 (between 7575 and 7979).

  3. 39709/52 F/M 2025 Q3(b)2 marks

    The lengths of 250250 leaves of a certain type of plant are measured, correct to the nearest centimetre. The results are summarised in the table below.

    Length (cm)5 – 910 – 1415 – 1920 – 2425 – 2930 – 39
    Frequency182860724824

    38%38\% of these leaves are of length kk cm or more.

    Use your graph to find an estimate for kk.

    Stuck? Show hint

    Use the graph from the second §08 worked example ("A cumulative frequency graph from rounded data"). Read at 62%62\% of 250250.

    Show solution
    1. 1

      Turn the wording round. 38%38\% are kk or more, so 62%62\% are below kk: 0.62×250=1550.62 \times 250 = 155

      The M1 requires a clear reading at 155 on the graph.

    2. 2

      Read. Across from 155155 and down: k≈23k \approx 23.

      The curve passes (19.5, 106) and (24.5, 178), so 155 falls between them. Accepted range 22.5 ≤ k ≤ 23.5.

    Answer

    k≈23k \approx 23 (between 22.522.5 and 23.523.5).

  4. 49709/52 M/J 2025 Q5(b)2 marks

    The times taken, tt minutes, by 300300 students to travel to Hollowton College are recorded. The results are summarised in the table below.

    Time (tt minutes)t≤10t \leq 10t≤20t \leq 20t≤30t \leq 30t≤40t \leq 40t≤60t \leq 60t≤90t \leq 90
    Cumulative frequency3486142208265300

    120120 students take more than kk minutes to travel to college. Use your graph to estimate the value of kk.

    Stuck? Show hint

    Use the graph from the last §08 exercise. "More than kk" counts from the top: read at 300−120300 - 120.

    Show solution
    1. 1

      Turn the count round. 120120 take more than kk, so 300−120=180300 - 120 = 180 take kk minutes or less.

      The M1 is for the line drawn from 180.

    2. 2

      Read. Across from 180180 to the curve and down: k≈35k \approx 35 minutes.

      The curve passes (30, 142) and (40, 208), so 180 falls between them.

    Answer

    k≈35k \approx 35 minutes.

  5. 59709/51 O/N 2024 Q3(a)(b)5 marks

    The time taken, in minutes, to walk to school was recorded for 200200 pupils at a certain school. These times are summarised in the following table.

    Time taken (tt minutes)t≤15t \leq 15t≤25t \leq 25t≤30t \leq 30t≤40t \leq 40t≤50t \leq 50t≤70t \leq 70
    Cumulative frequency184688140176200

    (a) Draw a cumulative frequency graph to illustrate the data.

    (b) Use your graph to estimate the median and the interquartile range of the data.

    The grid printed with part (a).

    The grid printed with part (a).

    Stuck? Show hint

    (a) The table is cumulative; start at (0,0)(0, 0). (b) Read at 100100, 5050 and 150150.

    Show solution
    The mark scheme's graph for part (a).

    The mark scheme's graph for part (a).

    1. 1

      (a) Points. (0,0),(15,18),(25,46),(30,88),(40,140),(50,176),(70,200)(0, 0), (15, 18), (25, 46), (30, 88), (40, 140), (50, 176), (70, 200).

      Read straight from the cumulative table; times cannot be negative, so start at (0, 0).

    2. 2

      (a) Draw. A smooth increasing curve through the points on linear axes "time (minutes)" and "cumulative frequency".

      The A1 needs the curve joined to (0, 0) and both axes labelled.

    3. 3

      (b) Median. Across from 12×200=100\tfrac12 \times 200 = 100 and down: median ≈33\approx 33 minutes.

      Between (30, 88) and (40, 140). 33 is the mark scheme's value; the B1 follows through from your own curve (within half a square), so a reading of about 32 from a well-drawn curve also scores.

    4. 4

      (b) Lower quartile. Across from 5050 and down: Q1≈26Q_1 \approx 26.

      Accepted range 25 < LQ ≤ 27.

    5. 5

      (b) Upper quartile. Across from 150150 and down: Q3≈42Q_3 \approx 42.

      Accepted range 41 ≤ UQ ≤ 43.

    6. 6

      (b) IQR. 42−26=16 minutes42 - 26 = 16 \text{ minutes}

      Show the subtraction.

    Answer

    (a) Points (0,0),(15,18),(25,46),(30,88),(40,140),(50,176),(70,200)(0, 0), (15, 18), (25, 46), (30, 88), (40, 140), (50, 176), (70, 200) joined smoothly. (b) Median ≈33\approx 33 minutes; IQR ≈42−26=16\approx 42 - 26 = 16 minutes.

Practise estimating percentiles and proportions from cumulative frequency graphsReal past-paper questions · Using cumulative frequency graphs to estimate percentiles and proportions
10

Mean and standard deviation from data and from totals

Syllabus requirement · §5.1

“

calculate and use the mean and standard deviation of a set of data (including grouped data) either from the data itself or from given totals Σx\Sigma x and Σx2\Sigma x^2

”

The mean of nn values xx is their total divided by nn: xˉ=∑xn.\bar{x} = \frac{\sum x}{n}. Here ∑x\sum x ("sigma xx") means "add up all the values of xx".

The standard deviation σ\sigma measures how far the values typically are from the mean. It is built in three stages:

  1. find each value's deviation from the mean, x−xˉx - \bar{x};
  2. square each deviation (so values below the mean do not cancel values above it) and take the mean of the squares: this is the variance, σ2=∑(x−xˉ)2n;\sigma^2 = \frac{\sum (x - \bar{x})^2}{n};
  3. take the square root, to get back to the original units: σ=variance\sigma = \sqrt{\text{variance}}.

A bigger standard deviation means the values are more spread out around the mean. If every value is the same, every deviation is 00 and σ=0\sigma = 0.

The working formula

Working out every deviation is slow. The formula booklet gives a second version of the variance that needs only two totals, ∑x\sum x and ∑x2\sum x^2:

σ2=∑x2n−xˉ2"the mean of the squares minus the square of the mean"\sigma^2 = \frac{\sum x^2}{n} - \bar{x}^2 \qquad \text{"the mean of the squares minus the square of the mean"}

It comes from the definition by expanding the bracket. Remember that xˉ\bar{x} is a fixed number, the same for every term:

∑(x−xˉ)2=∑(x2−2xˉ x+xˉ2)\sum (x - \bar{x})^2 = \sum \left(x^2 - 2\bar{x}\,x + \bar{x}^2\right)

Sum each of the three terms separately. The middle term has the constant 2xˉ2\bar{x} outside the sum, and the last term is the constant xˉ2\bar{x}^2 added nn times:

=∑x2−2xˉ∑x+nxˉ2= \sum x^2 - 2\bar{x}\sum x + n\bar{x}^2

Divide by nn, and use ∑xn=xˉ\dfrac{\sum x}{n} = \bar{x} in the middle term:

∑(x−xˉ)2n=∑x2n−2xˉ⋅xˉ+xˉ2=∑x2n−xˉ2\frac{\sum (x - \bar{x})^2}{n} = \frac{\sum x^2}{n} - 2\bar{x}\cdot\bar{x} + \bar{x}^2 = \frac{\sum x^2}{n} - \bar{x}^2

This version is also the one that lets you work backwards: rearranged, it gives ∑x2\sum x^2 from a mean and a standard deviation, ∑x2=n(σ2+xˉ2),\sum x^2 = n\left(\sigma^2 + \bar{x}^2\right), which is needed whenever a value is added to or removed from a data set whose raw values you don't have.

Mean and standard deviation (raw data)
xˉ=∑xn\bar{x} = \frac{\sum x}{n}

Mean

σ2=∑x2n−xˉ2\sigma^2 = \frac{\sum x^2}{n} - \bar{x}^2

Variance: mean of the squares minus the square of the mean (in the formula booklet)

σ=σ2\sigma = \sqrt{\sigma^2}

Standard deviation

∑x=nxˉ\sum x = n\bar{x}

A total from a mean

∑x2=n(σ2+xˉ2)\sum x^2 = n\left(\sigma^2 + \bar{x}^2\right)

A sum of squares from a mean and a standard deviation

Mean and standard deviation of a list

Find the mean and standard deviation of these 66 values: 4, 7, 5, 9, 6, 54,\ 7,\ 5,\ 9,\ 6,\ 5

Show full working
  1. 1

    Total. ∑x=4+7+5+9+6+5=36\sum x = 4 + 7 + 5 + 9 + 6 + 5 = 36

    The first total the formulas need.

  2. 2

    Mean. xˉ=366=6\bar{x} = \frac{36}{6} = 6

    n = 6 values.

  3. 3

    Square each value. 42=16, 72=49, 52=25, 92=81, 62=36, 52=254^2 = 16,\ 7^2 = 49,\ 5^2 = 25,\ 9^2 = 81,\ 6^2 = 36,\ 5^2 = 25

    Σx² means square each value first, then add. It is not (Σx)².

  4. 4

    Sum of squares. ∑x2=16+49+25+81+36+25=232\sum x^2 = 16 + 49 + 25 + 81 + 36 + 25 = 232

    The second total the formulas need.

  5. 5

    Mean of the squares. ∑x2n=2326=38.666…\frac{\sum x^2}{n} = \frac{232}{6} = 38.666\ldots

    Keep all the digits for now.

  6. 6

    Subtract the square of the mean. σ2=38.666…−62=38.666…−36=2.666…\sigma^2 = 38.666\ldots - 6^2 = 38.666\ldots - 36 = 2.666\ldots

    Mean of squares minus square of mean, in that order. The other order gives a negative 'variance', which is impossible.

  7. 7

    Square root. σ=2.666…=1.63 (3 s.f.)\sigma = \sqrt{2.666\ldots} = 1.63 \text{ (3 s.f.)}

    Round only at the end.

Answer

xˉ=6\bar{x} = 6, σ=1.63\sigma = 1.63 (3 s.f.).

Adding a value when only the mean and standard deviation are known

Ten values have mean 2020 and standard deviation 33. An eleventh value, 3131, is added. Find the mean and standard deviation of all 1111 values.

Show full working
  1. 1

    Total of the ten values. ∑x=nxˉ=10×20=200\sum x = n\bar{x} = 10 \times 20 = 200

    The raw values are unknown, so rebuild the totals from the summary.

  2. 2

    Rearrange the variance formula for ∑x2\sum x^2. From σ2=∑x2n−xˉ2\sigma^2 = \dfrac{\sum x^2}{n} - \bar{x}^2: ∑x2=n(σ2+xˉ2)\sum x^2 = n\left(\sigma^2 + \bar{x}^2\right)

    Add x̄² to both sides, then multiply by n.

  3. 3

    Sum of squares of the ten values. ∑x2=10(32+202)=10×409=4090\sum x^2 = 10\left(3^2 + 20^2\right) = 10 \times 409 = 4090

    Substitute n = 10, σ = 3 and x̄ = 20.

  4. 4

    New total. ∑x=200+31=231\sum x = 200 + 31 = 231

    Totals can simply be added to. That is why everything is converted to totals first.

  5. 5

    New sum of squares. ∑x2=4090+312=4090+961=5051\sum x^2 = 4090 + 31^2 = 4090 + 961 = 5051

    Add the square of the new value, not the value itself.

  6. 6

    New mean. xˉ=23111=21\bar{x} = \frac{231}{11} = 21

    Now n = 11.

  7. 7

    Mean of the squares. 505111=459.18…\frac{5051}{11} = 459.18\ldots

    The same formula, with the new totals and the new n.

  8. 8

    Subtract the square of the mean. σ2=459.18…−212=459.18…−441=18.18…\sigma^2 = 459.18\ldots - 21^2 = 459.18\ldots - 441 = 18.18\ldots

    This is the new variance.

  9. 9

    New standard deviation. σ=18.18…=4.26 (3 s.f.)\sigma = \sqrt{18.18\ldots} = 4.26 \text{ (3 s.f.)}

    The spread has grown, because 31 is a long way from the old mean of 20.

Answer

New mean =21= 21; new standard deviation =4.26= 4.26 (3 s.f.).

Adding or removing a value: convert to totals (Σx and Σx²), change the totals, then convert back. Never try to adjust a mean or a standard deviation directly.

A missing value from a mean

9709/53 M/J 2023 Q4(c)3 marks

The times taken, in minutes, to complete a cycle race by 1919 cyclists from each of two clubs, the Cheetahs and the Panthers, are represented in the following back-to-back stem-and-leaf diagram.

Key: 7∣9∣17 \mid 9 \mid 1 means 9797 minutes for Cheetahs and 9191 minutes for Panthers

Another cyclist, Kenny, from the Cheetahs also took part in the race. The mean time taken by the 2020 cyclists from the Cheetahs was 9999 minutes.

Find the time taken by Kenny to complete the race.

The printed diagram. The Cheetahs are on the left.

The printed diagram. The Cheetahs are on the left.

Show full working
  1. 1

    List the 1919 Cheetahs' times, reading each left-hand row outwards from the stem: 78,79,80,82,83,87,88,97,98,99,101,103,103,105,106,112,118,119,12478, 79, 80, 82, 83, 87, 88, 97, 98, 99, 101, 103, 103, 105, 106, 112, 118, 119, 124

    The same list as in the §03 exercise.

  2. 2

    Row totals. 78+79=157,80+82+83+87+88=420,97+98+99=294,78 + 79 = 157,\quad 80 + 82 + 83 + 87 + 88 = 420,\quad 97 + 98 + 99 = 294, 101+103+103+105+106=518,112+118+119=349,124101 + 103 + 103 + 105 + 106 = 518,\quad 112 + 118 + 119 = 349,\quad 124

    Adding row by row keeps the arithmetic checkable.

  3. 3

    Total of the 1919 known times. 157+420+294+518+349+124=1862157 + 420 + 294 + 518 + 349 + 124 = 1862

    This is Σx for the 19 cyclists.

  4. 4

    Total for all 2020 Cheetahs, from the mean. nxˉ=20×99=1980n\bar{x} = 20 \times 99 = 1980

    The B1 is for 1980. The mean of 99 is for all 20 cyclists, Kenny included.

  5. 5

    Kenny's time is the difference. 1980−1862=118 minutes1980 - 1862 = 118 \text{ minutes}

    The total for 20 is the total for 19 plus Kenny's time.

Answer

Kenny took 118118 minutes.

Common mistakes
  • σ2=xˉ2−∑x2n\sigma^2 = \bar{x}^2 - \dfrac{\sum x^2}{n}

    σ2=∑x2n−xˉ2\sigma^2 = \dfrac{\sum x^2}{n} - \bar{x}^2

    The mean of the squares comes first. A negative variance is the tell-tale sign of the wrong order.

  • ∑x2=(∑x)2\sum x^2 = \left(\sum x\right)^2

    Square each value, then add

    For 2 and 3: 2² + 3² = 13, but (2 + 3)² = 25.

  • Rounding the mean before squaring it

    Keep the mean exact (or to several more figures) until the end

    Squaring a rounded mean can move the answer outside the accepted range.

  • Giving the variance when the standard deviation was asked for

    Check the question and take the square root if needed

    The two differ only by the final square root, so it is an easy slip.

Your turn

Build the totals Σx and Σx² first, whatever form the data arrives in.

  1. 1

    Find the mean and standard deviation of 2, 3, 3, 5, 72,\ 3,\ 3,\ 5,\ 7.

    Stuck? Show hint

    Find ∑x\sum x and ∑x2\sum x^2 separately.

    Show solution
    1. 1

      Total. ∑x=2+3+3+5+7=20\sum x = 2 + 3 + 3 + 5 + 7 = 20.

      The first total.

    2. 2

      Mean. xˉ=205=4\bar{x} = \dfrac{20}{5} = 4.

      Five values.

    3. 3

      Sum of squares. ∑x2=4+9+9+25+49=96\sum x^2 = 4 + 9 + 9 + 25 + 49 = 96.

      Square, then add.

    4. 4

      Mean of the squares. 965=19.2\dfrac{96}{5} = 19.2.

      Divide by n.

    5. 5

      Variance. σ2=19.2−42=19.2−16=3.2\sigma^2 = 19.2 - 4^2 = 19.2 - 16 = 3.2.

      Mean of squares minus square of mean.

    6. 6

      Standard deviation. σ=3.2=1.79\sigma = \sqrt{3.2} = 1.79 (3 s.f.).

      Square root at the end.

    Answer

    xˉ=4\bar{x} = 4, σ=1.79\sigma = 1.79.

  2. 29709/52 M/J 2021 Q7(d)3 marks

    The heights, in cm, of the 1111 basketball players in each of two clubs, the Amazons and the Giants, are shown below.

    Amazons205198181182190215201178202196184
    Giants175182184187189192193195195195204

    Four new players join the Amazons. The mean height of the 1515 players in the Amazons is now 191.2 cm191.2\text{ cm}. The heights of three of the new players are 180 cm180\text{ cm}, 185 cm185\text{ cm} and 190 cm190\text{ cm}.

    Find the height of the fourth new player.

    Stuck? Show hint

    Total of the original 1111; total of all 1515 from the new mean; the difference is the four new players.

    Show solution
    1. 1

      Total of the original 1111 Amazons. 205+198+181+182+190+215+201+178+202+196+184=2132205 + 198 + 181 + 182 + 190 + 215 + 201 + 178 + 202 + 196 + 184 = 2132

      The total of the 11 original heights, from the table.

    2. 2

      Total of all 1515. 191.2×15=2868191.2 \times 15 = 2868

      The B1 is for both totals.

    3. 3

      Equation for the fourth height hh. 2868=2132+180+185+190+h2868 = 2132 + 180 + 185 + 190 + h

      The M1 is for forming this equation.

    4. 4

      Add the three known new heights. 180+185+190=555, so 2868=2132+555+h=2687+h180 + 185 + 190 = 555, \text{ so } 2868 = 2132 + 555 + h = 2687 + h

      Collect all the known numbers on the right.

    5. 5

      Subtract 26872687. h=2868−2687=181 cmh = 2868 - 2687 = 181 \text{ cm}

      A1.

    Answer

    181181 cm.

  3. 39709/63 M/J 2014 Q4(i)(ii)6 marks

    The heights, xx cm, of a group of 2828 people were measured. The mean height was found to be 172.6172.6 cm and the standard deviation was found to be 4.584.58 cm. A person whose height was 161.8161.8 cm left the group.

    (i) Find the mean height of the remaining group of 2727 people.

    (ii) Find ∑x2\sum x^2 for the original group of 2828 people. Hence find the standard deviation of the heights of the remaining group of 2727 people.

    Stuck? Show hint

    Rebuild ∑x\sum x and ∑x2\sum x^2 for the 2828 people, take away the person who left, then use the formulas with n=27n = 27.

    Show solution
    1. 1

      (i) Original total. ∑x=28×172.6=4832.8\sum x = 28 \times 172.6 = 4832.8

      Total from a mean.

    2. 2

      (i) Remove the person who left. 4832.8−161.8=46714832.8 - 161.8 = 4671

      The total of the remaining 27.

    3. 3

      (i) New mean. 467127=173 cm\frac{4671}{27} = 173 \text{ cm}

      Divide by the new n.

    4. 4

      (ii) Inside the bracket of ∑x2=n(σ2+xˉ2)\sum x^2 = n(\sigma^2 + \bar{x}^2). 4.582+172.62=20.9764+29 790.76=29 811.73644.58^2 + 172.6^2 = 20.9764 + 29\,790.76 = 29\,811.7364

      σ² + x̄² for the original group.

    5. 5

      (ii) Multiply by 2828. ∑x2=28×29 811.7364=834 728.6\sum x^2 = 28 \times 29\,811.7364 = 834\,728.6

      Keep the unrounded value.

    6. 6

      (ii) Remove the person's square. 834 728.6−161.82=834 728.6−26 179.24=808 549.4834\,728.6 - 161.8^2 = 834\,728.6 - 26\,179.24 = 808\,549.4

      Take 161.8², not 161.8, from the sum of squares.

    7. 7

      (ii) Mean of the squares. 808 549.427=29 946.27\frac{808\,549.4}{27} = 29\,946.27

      Divide by the new n, 27.

    8. 8

      (ii) New variance. 29 946.27−1732=29 946.27−29 929=17.2729\,946.27 - 173^2 = 29\,946.27 - 29\,929 = 17.27

      The new mean from part (i) is exactly 173.

    9. 9

      (ii) New standard deviation. 17.27=4.16 cm\sqrt{17.27} = 4.16 \text{ cm}

      The mark scheme's answer (it also accepted 4.15).

    Answer

    (i) 173173 cm. (ii) ∑x2=834 728.6\sum x^2 = 834\,728.6 (about 835 000835\,000); new standard deviation =4.16= 4.16 cm.

Practise calculating the mean and standard deviationReal past-paper questions · Calculating mean and standard deviation from raw, grouped or coded data
11

Estimating the mean and standard deviation of grouped data

Syllabus requirement · §5.1

“

calculate and use the mean and standard deviation of a set of data (including grouped data)

”

In a grouped frequency table the individual values are lost: we only know how many values fall in each class. To estimate the mean and standard deviation, every value in a class is treated as if it were the class midpoint xx, and the class frequency ff says how many times that midpoint counts:

xˉ=∑fx∑f,σ2=∑fx2∑f−xˉ2\bar{x} = \frac{\sum fx}{\sum f}, \qquad \sigma^2 = \frac{\sum fx^2}{\sum f} - \bar{x}^2

These are the §10 formulas with each value counted ff times; ∑f\sum f is just nn, the total frequency. The results are estimates, and questions say "calculate an estimate of the mean", because the midpoint is only a stand-in for the real values in the class.

Midpoints come from the true class boundaries (§06), not from the printed limits:

  • 20⩽t<3020 \leqslant t < 30 has midpoint 20+302=25\dfrac{20 + 30}{2} = 25;
  • 121121–130130 to the nearest minute has boundaries 120.5120.5 and 130.5130.5, so midpoint 125.5125.5;
  • 00–44 km to the nearest km has boundaries 00 and 4.54.5 (a distance cannot be negative), so midpoint 2.252.25.

What the marks are for

A "mean and standard deviation" part is typically 5 or 6 marks: B1 for the midpoints (at least four or five correct), sometimes B1 for the frequencies, M1 for a correct mean expression, A1 for the mean, M1 for a correct variance expression that includes "−xˉ2-\bar{x}^2", A1 for the standard deviation.

The method marks are only given if the xx values are midpoints, not upper or lower boundaries, class widths, frequency densities, frequencies or cumulative frequencies. Setting the working out as a table with columns xx, ff, fxfx, fx2fx^2 shows the examiner everything.

Mean and standard deviation of a grouped table

The ages, in completed years, of 4040 members of a gym are grouped as follows.

Age (years)1010–19192020–29293030–39394040–5959
Frequency881616101066

Calculate estimates of the mean and standard deviation of the ages.

Show full working
  1. 1

    Class boundaries. Ages in completed years are truncated (§06), so 1010–1919 runs from 1010 to 2020, 2020–2929 from 2020 to 3030, 3030–3939 from 3030 to 4040, and 4040–5959 from 4040 to 6060.

    A member recorded as 19 could be 19 years and 11 months, so the class really ends at 20.

  2. 2

    Midpoints. 10+202=15,20+302=25,30+402=35,40+602=50\frac{10 + 20}{2} = 15, \quad \frac{20 + 30}{2} = 25, \quad \frac{30 + 40}{2} = 35, \quad \frac{40 + 60}{2} = 50

    Average of each class's two boundaries.

  3. 3

    fxfx for each class. 8×15=120,16×25=400,10×35=350,6×50=3008 \times 15 = 120, \quad 16 \times 25 = 400, \quad 10 \times 35 = 350, \quad 6 \times 50 = 300

    Each midpoint counted as many times as the class frequency.

  4. 4

    Total frequency. ∑f=8+16+10+6=40\sum f = 8 + 16 + 10 + 6 = 40

    Σf must equal the stated n = 40.

  5. 5

    Total of fxfx. ∑fx=120+400+350+300=1170\sum fx = 120 + 400 + 350 + 300 = 1170

    The first column total.

  6. 6

    Estimated mean. xˉ=117040=29.25 years\bar{x} = \frac{1170}{40} = 29.25 \text{ years}

    Σfx ÷ Σf.

  7. 7

    Square each midpoint. 152=225,252=625,352=1225,502=250015^2 = 225, \quad 25^2 = 625, \quad 35^2 = 1225, \quad 50^2 = 2500

    Square first, before multiplying by f.

  8. 8

    fx2fx^2 for each class. 8×225=1800,16×625=10 000,10×1225=12 250,6×2500=15 0008 \times 225 = 1800, \quad 16 \times 625 = 10\,000, \quad 10 \times 1225 = 12\,250, \quad 6 \times 2500 = 15\,000

    Squared midpoint times f. Not (fx)².

  9. 9

    Sum. ∑fx2=1800+10 000+12 250+15 000=39 050\sum fx^2 = 1800 + 10\,000 + 12\,250 + 15\,000 = 39\,050

    The second column total.

  10. 10

    Mean of the squares. 39 05040=976.25\frac{39\,050}{40} = 976.25

    Σfx² ÷ Σf.

  11. 11

    Subtract the square of the mean. σ2=976.25−29.252=976.25−855.5625=120.6875\sigma^2 = 976.25 - 29.25^2 = 976.25 - 855.5625 = 120.6875

    Use the unrounded mean.

  12. 12

    Estimated standard deviation. σ=120.6875=11.0 years (3 s.f.)\sigma = \sqrt{120.6875} = 11.0 \text{ years (3 s.f.)}

    Square root last.

Answer

Estimated mean =29.25= 29.25 years; estimated standard deviation =11.0= 11.0 years (3 s.f.).

Lay out columns x, f, fx, fx² and total the last three. The mark scheme follows the same layout.

Mean and standard deviation with rounded data

9709/51 O/N 2023 Q4(b)5 marks

The times, to the nearest minute, of 150150 athletes taking part in a charity run are recorded. The results are summarised in the table.

Time in minutes101 − 120121 − 130131 − 135136 − 145146 − 160
Frequency1848343218

Calculate estimates for the mean and standard deviation of the times taken by the athletes.

Show full working
  1. 1

    Boundaries. To the nearest minute: 100.5, 120.5, 130.5, 135.5, 145.5, 160.5100.5,\ 120.5,\ 130.5,\ 135.5,\ 145.5,\ 160.5.

    Half a minute out from each printed limit.

  2. 2

    Midpoints. 110.5, 125.5, 133, 140.5, 153110.5,\ 125.5,\ 133,\ 140.5,\ 153

    For example (130.5 + 135.5) ÷ 2 = 133. The B1 needs at least four correct.

  3. 3

    fxfx values. 18×110.5=1989,48×125.5=6024,34×133=4522,32×140.5=4496,18×153=275418 \times 110.5 = 1989, \quad 48 \times 125.5 = 6024, \quad 34 \times 133 = 4522, \quad 32 \times 140.5 = 4496, \quad 18 \times 153 = 2754

    The mark scheme shows exactly these products.

  4. 4

    ∑fx\sum fx. 1989+6024+4522+4496+2754=19 7851989 + 6024 + 4522 + 4496 + 2754 = 19\,785

    Σf = 150, as stated.

  5. 5

    Estimated mean. xˉ=19 785150=131.9 minutes\bar{x} = \frac{19\,785}{150} = 131.9 \text{ minutes}

    M1 A1.

  6. 6

    Square each midpoint. 110.52=12 210.25,125.52=15 750.25,1332=17 689,140.52=19 740.25,1532=23 409110.5^2 = 12\,210.25, \quad 125.5^2 = 15\,750.25, \quad 133^2 = 17\,689, \quad 140.5^2 = 19\,740.25, \quad 153^2 = 23\,409

    Square first, then multiply by f.

  7. 7

    fx2fx^2 values. 18×12 210.25=219 784.5,48×15 750.25=756 012,34×17 689=601 426,18 \times 12\,210.25 = 219\,784.5, \quad 48 \times 15\,750.25 = 756\,012, \quad 34 \times 17\,689 = 601\,426, 32×19 740.25=631 688,18×23 409=421 36232 \times 19\,740.25 = 631\,688, \quad 18 \times 23\,409 = 421\,362

    Squared midpoint times frequency.

  8. 8

    ∑fx2\sum fx^2. 219 784.5+756 012+601 426+631 688+421 362=2 630 272.5219\,784.5 + 756\,012 + 601\,426 + 631\,688 + 421\,362 = 2\,630\,272.5

    The column total, as the mark scheme gives it.

  9. 9

    Mean of the squares. 2 630 272.5150=17 535.15\frac{2\,630\,272.5}{150} = 17\,535.15

    Σfx² ÷ Σf.

  10. 10

    Subtract the square of the mean. σ2=17 535.15−131.92=17 535.15−17 397.61=137.54\sigma^2 = 17\,535.15 - 131.9^2 = 17\,535.15 - 17\,397.61 = 137.54

    The M1 needs the '− mean²' in the expression.

  11. 11

    Estimated standard deviation. σ=137.54=11.7 minutes (3 s.f.)\sigma = \sqrt{137.54} = 11.7 \text{ minutes (3 s.f.)}

    The mark scheme accepts 11.6 ≤ σ < 11.95.

Answer

Estimated mean =131.9= 131.9 minutes; estimated standard deviation =11.7= 11.7 minutes.

When the table is cumulative

If the table gives cumulative frequencies ("t⩽10t \leqslant 10: 3434, t⩽20t \leqslant 20: 8686, …"), the class frequencies must be recovered first by subtracting consecutive running totals: the class 10<t⩽2010 < t \leqslant 20 holds 86−34=5286 - 34 = 52 values. Using cumulative frequencies as ff is one of the errors the mark schemes name explicitly.

Mean and standard deviation from a cumulative frequency table

9709/52 M/J 2025 Q5(c)6 marks

The times taken, tt minutes, by 300300 students to travel to Hollowton College are recorded. The results are summarised in the table below.

Time (tt minutes)t≤10t \leq 10t≤20t \leq 20t≤30t \leq 30t≤40t \leq 40t≤60t \leq 60t≤90t \leq 90
Cumulative frequency3486142208265300

Calculate estimates of the mean and standard deviation of the times taken to travel to college by the 300300 students.

Show full working
  1. 1

    Recover the class frequencies. Subtract each running total from the next: 34,86−34=52,142−86=56,208−142=66,265−208=57,300−265=3534, \quad 86 - 34 = 52, \quad 142 - 86 = 56, \quad 208 - 142 = 66, \quad 265 - 208 = 57, \quad 300 - 265 = 35

    The first B1 is for these frequencies. They must add to 300.

  2. 2

    Classes and midpoints. The classes are 00–1010, 1010–2020, 2020–3030, 3030–4040, 4040–6060, 6060–9090, with midpoints 5, 15, 25, 35, 50, 755,\ 15,\ 25,\ 35,\ 50,\ 75

    The second B1. The first class starts at 0, since a time cannot be negative.

  3. 3

    fxfx values. 34×5=170,52×15=780,56×25=1400,66×35=2310,57×50=2850,35×75=262534 \times 5 = 170, \quad 52 \times 15 = 780, \quad 56 \times 25 = 1400, \quad 66 \times 35 = 2310, \quad 57 \times 50 = 2850, \quad 35 \times 75 = 2625

    Frequency times midpoint, class by class.

  4. 4

    ∑fx\sum fx. 170+780+1400+2310+2850+2625=10 135170 + 780 + 1400 + 2310 + 2850 + 2625 = 10\,135

    The numerator of the mean.

  5. 5

    Estimated mean. xˉ=10 135300=33.78…=33.8 minutes (3 s.f.)\bar{x} = \frac{10\,135}{300} = 33.78\ldots = 33.8 \text{ minutes (3 s.f.)}

    The mark scheme gives A0 for leaving the answer as 10 135/300: evaluate it.

  6. 6

    fx2fx^2 values. 34×25=850,52×225=11 700,56×625=35 000,34 \times 25 = 850, \quad 52 \times 225 = 11\,700, \quad 56 \times 625 = 35\,000, 66×1225=80 850,57×2500=142 500,35×5625=196 87566 \times 1225 = 80\,850, \quad 57 \times 2500 = 142\,500, \quad 35 \times 5625 = 196\,875

    Each midpoint squared first (5² = 25, 15² = 225, …), then multiplied by f.

  7. 7

    ∑fx2\sum fx^2. 850+11 700+35 000+80 850+142 500+196 875=467 775850 + 11\,700 + 35\,000 + 80\,850 + 142\,500 + 196\,875 = 467\,775

    The numerator of the mean of the squares.

  8. 8

    Mean of the squares. 467 775300=1559.25\frac{467\,775}{300} = 1559.25

    Σfx² ÷ Σf.

  9. 9

    Subtract the square of the mean. σ2=1559.25−(10 135300)2=1559.25−1141.31…=417.94\sigma^2 = 1559.25 - \left(\frac{10\,135}{300}\right)^2 = 1559.25 - 1141.31\ldots = 417.94

    Use the exact mean 10 135/300 here, not 33.8.

  10. 10

    Estimated standard deviation. σ=417.94=20.4 minutes (3 s.f.)\sigma = \sqrt{417.94} = 20.4 \text{ minutes (3 s.f.)}

    The mark scheme accepts 20.4 ≤ σ < 20.45.

Answer

Estimated mean =33.8= 33.8 minutes; estimated standard deviation =20.4= 20.4 minutes.

Cumulative table → difference first. Then it is an ordinary grouped-data calculation.

Common mistakes
  • Using upper boundaries, class widths or frequency densities as xx

    Use class midpoints

    The method marks require midpoints that lie inside their classes.

  • Using cumulative frequencies as ff

    Subtract consecutive cumulative frequencies first

    Cumulative totals count the early classes over and over.

  • ∑fx2\sum fx^2 worked out as ∑(fx)2\sum (fx)^2

    Square the midpoint, then multiply by ff

    (fx)² squares the frequency too.

  • Midpoint of 121121–130130 (nearest minute) taken as 125125

    Boundaries 120.5120.5 and 130.5130.5, midpoint 125.5125.5

    Midpoints come from the true boundaries.

Your turn

Lay out x, f, fx and fx² in columns. Keep the mean unrounded for the variance.

  1. 19709/52 F/M 2021 Q5(c)3 marks

    A driver records the distance travelled in each of 150150 journeys. These distances, correct to the nearest km, are summarised in the following table.

    Distance (km)0 – 45 – 1011 – 2021 – 3031 – 4041 – 60
    Frequency12163266204

    Calculate an estimate of the mean distance travelled for the 150150 journeys.

    Stuck? Show hint

    The first class runs from 00 to 4.54.5 km (a distance cannot be negative), so its midpoint is 2.252.25.

    Show solution
    1. 1

      Midpoints. 2.25, 7.5, 15.5, 25.5, 35.5, 50.52.25,\ 7.5,\ 15.5,\ 25.5,\ 35.5,\ 50.5.

      Boundaries 0, 4.5, 10.5, 20.5, 30.5, 40.5, 60.5. The B1 needs at least five correct.

    2. 2

      fxfx values. 12×2.25=27, 16×7.5=120, 32×15.5=496, 66×25.5=1683, 20×35.5=710, 4×50.5=20212 \times 2.25 = 27,\ 16 \times 7.5 = 120,\ 32 \times 15.5 = 496,\ 66 \times 25.5 = 1683,\ 20 \times 35.5 = 710,\ 4 \times 50.5 = 202

      The mark scheme shows these products.

    3. 3

      ∑fx\sum fx. 27+120+496+1683+710+202=323827 + 120 + 496 + 1683 + 710 + 202 = 3238.

      Σf = 150.

    4. 4

      Estimated mean. 3238150=21.6 km (3 s.f.)\frac{3238}{150} = 21.6 \text{ km (3 s.f.)}

      Accept 21.5866… too.

    Answer

    21.621.6 km.

  2. 29709/52 O/N 2022 Q4(b)3 marks

    The times taken, in minutes, to complete a word processing task by 250250 employees at a particular company are summarised in the table.

    Time taken (tt minutes)0≤t<200 \le t < 2020≤t<4020 \le t < 4040≤t<5040 \le t < 5050≤t<6050 \le t < 6060≤t<10060 \le t < 100
    Frequency3246965224

    From the data, the estimate of the mean time taken by these 250250 employees is 43.243.2 minutes.

    Calculate an estimate for the standard deviation of these times.

    Stuck? Show hint

    The mean is given, so only ∑fx2\sum fx^2 is needed.

    Show solution
    1. 1

      Midpoints. 10, 30, 45, 55, 8010,\ 30,\ 45,\ 55,\ 80.

      Inequality classes.

    2. 2

      fx2fx^2 values. 32×100=3200, 46×900=41 400, 96×2025=194 400, 52×3025=157 300, 24×6400=153 60032 \times 100 = 3200,\ 46 \times 900 = 41\,400,\ 96 \times 2025 = 194\,400,\ 52 \times 3025 = 157\,300,\ 24 \times 6400 = 153\,600

      Midpoint squared times frequency.

    3. 3

      ∑fx2\sum fx^2. 3200+41 400+194 400+157 300+153 600=549 9003200 + 41\,400 + 194\,400 + 157\,300 + 153\,600 = 549\,900

      The column total.

    4. 4

      Mean of the squares. 549 900250=2199.6\frac{549\,900}{250} = 2199.6

      Divide by Σf = 250.

    5. 5

      Variance. 2199.6−43.22=2199.6−1866.24=333.362199.6 - 43.2^2 = 2199.6 - 1866.24 = 333.36

      Use the given mean.

    6. 6

      Standard deviation. 333.36=18.3 minutes (3 s.f.)\sqrt{333.36} = 18.3 \text{ minutes (3 s.f.)}

      18.258… to at least 3 s.f.

    Answer

    18.318.3 minutes.

  3. 39709/51 M/J 2022 Q3(b)(d)4 marks

    The times taken to travel to college by 25002500 students are summarised in the table.

    Time taken (tt minutes)0≤t<200 \le t < 2020≤t<3020 \le t < 3030≤t<4030 \le t < 4040≤t<6040 \le t < 6060≤t<9060 \le t < 90
    Frequency440720920300120

    (b) From the data, the estimate of the mean value of tt is 31.4431.44.

    Calculate an estimate of the standard deviation of the times taken to travel to college.

    (d) It was later discovered that the times taken to travel to college by two students were incorrectly recorded. One student's time was recorded as 1515 instead of 55 and the other's time was recorded as 6565 instead of 7575.

    Without doing any further calculations, state with a reason whether the estimate of the standard deviation in part (b) would be increased, decreased or stay the same.

    Stuck? Show hint

    (d) Which classes were the wrong and the correct times in?

    Show solution
    1. 1

      (b) Midpoints. 10, 25, 35, 50, 7510,\ 25,\ 35,\ 50,\ 75.

      B1 for at least four.

    2. 2

      (b) fx2fx^2 values. 440×100=44 000, 720×625=450 000, 920×1225=1 127 000, 300×2500=750 000, 120×5625=675 000440 \times 100 = 44\,000,\ 720 \times 625 = 450\,000,\ 920 \times 1225 = 1\,127\,000,\ 300 \times 2500 = 750\,000,\ 120 \times 5625 = 675\,000

      Midpoints squared (100, 625, 1225, 2500, 5625), times frequencies.

    3. 3

      (b) ∑fx2\sum fx^2. 44 000+450 000+1 127 000+750 000+675 000=3 046 00044\,000 + 450\,000 + 1\,127\,000 + 750\,000 + 675\,000 = 3\,046\,000

      The column total.

    4. 4

      (b) Mean of the squares. 3 046 0002500=1218.4\frac{3\,046\,000}{2500} = 1218.4

      Divide by Σf = 2500.

    5. 5

      (b) Variance. 1218.4−31.442=1218.4−988.4736=229.92641218.4 - 31.44^2 = 1218.4 - 988.4736 = 229.9264

      Given mean.

    6. 6

      (b) Standard deviation. 229.9264=15.2 minutes (3 s.f.)\sqrt{229.9264} = 15.2 \text{ minutes (3 s.f.)}

      Accept 15.16…

    7. 7

      (d) Check the classes. 1515 and 55 are both in 0⩽t<200 \leqslant t < 20; 6565 and 7575 are both in 60⩽t<9060 \leqslant t < 90.

      The estimate depends only on the class frequencies and midpoints.

    8. 8

      (d) Conclude. The class frequencies do not change, so the estimate stays the same.

      The mark scheme: 'stays the same, data still in same intervals'.

    Answer

    (b) 15.215.2 minutes. (d) It stays the same: each corrected time is in the same class as the wrong one, so the frequencies are unchanged.

  4. 49709/53 O/N 2022 Q3(c)4 marks

    The times, tt minutes, taken to complete a walking challenge by 250250 members of a club are summarised in the table.

    Time taken (tt minutes)t≤20t \le 20t≤30t \le 30t≤35t \le 35t≤40t \le 40t≤50t \le 50t≤60t \le 60
    Cumulative frequency3266112178228250

    It is given that an estimate for the mean time taken to complete the challenge by these 250250 members is 34.434.4 minutes.

    Calculate an estimate for the standard deviation of the times taken to complete the challenge by these 250250 members.

    Stuck? Show hint

    Difference the cumulative frequencies first. The first class is 00 to 2020.

    Show solution
    1. 1

      Frequencies. 32, 34, 46, 66, 50, 2232,\ 34,\ 46,\ 66,\ 50,\ 22.

      66 − 32 = 34, 112 − 66 = 46, 178 − 112 = 66, 228 − 178 = 50, 250 − 228 = 22.

    2. 2

      Midpoints. 10, 25, 32.5, 37.5, 45, 5510,\ 25,\ 32.5,\ 37.5,\ 45,\ 55.

      Classes 0–20, 20–30, 30–35, 35–40, 40–50, 50–60.

    3. 3

      fx2fx^2 values. 32×100=3200, 34×625=21 250, 46×1056.25=48 587.5,32 \times 100 = 3200,\ 34 \times 625 = 21\,250,\ 46 \times 1056.25 = 48\,587.5, 66×1406.25=92 812.5, 50×2025=101 250, 22×3025=66 55066 \times 1406.25 = 92\,812.5,\ 50 \times 2025 = 101\,250,\ 22 \times 3025 = 66\,550

      Midpoints squared, times frequencies.

    4. 4

      ∑fx2\sum fx^2. 3200+21 250+48 587.5+92 812.5+101 250+66 550=333 6503200 + 21\,250 + 48\,587.5 + 92\,812.5 + 101\,250 + 66\,550 = 333\,650

      The mark scheme's total.

    5. 5

      Mean of the squares. 333 650250=1334.6\frac{333\,650}{250} = 1334.6

      Divide by Σf = 250.

    6. 6

      Variance. 1334.6−34.42=1334.6−1183.36=151.241334.6 - 34.4^2 = 1334.6 - 1183.36 = 151.24

      Given mean.

    7. 7

      Standard deviation. 151.24=12.3 minutes (3 s.f.)\sqrt{151.24} = 12.3 \text{ minutes (3 s.f.)}

      Rounded to 3 s.f.

    Answer

    12.312.3 minutes.

  5. 59709/62 O/N 2013 Q4(ii)(iii)7 marks

    The following histogram summarises the times, in minutes, taken by 190190 people to complete a race.

    (ii) Calculate estimates of the mean and standard deviation of the times of the 190190 people.

    (iii) Explain why your answers to part (ii) are estimates.

    The printed histogram (the same one as in §06).

    The printed histogram (the same one as in §06).

    Stuck? Show hint

    First turn each bar into a frequency: density × width. The bars run 100–150, 150–175, 175–200, 200–250 and 250–350.

    Show solution
    1. 1

      Frequencies from the bars. 0.2×50=10,1.0×25=25,2.0×25=50,1.5×50=75,0.3×100=300.2 \times 50 = 10,\quad 1.0 \times 25 = 25,\quad 2.0 \times 25 = 50,\quad 1.5 \times 50 = 75,\quad 0.3 \times 100 = 30

      These add to 190 ✓. The mark scheme gives M1 A1 for the frequencies.

    2. 2

      Midpoints. 125, 162.5, 187.5, 225, 300125,\ 162.5,\ 187.5,\ 225,\ 300.

      The centre of each bar.

    3. 3

      fxfx values. 10(125)=1250, 25(162.5)=4062.5, 50(187.5)=9375, 75(225)=16 875, 30(300)=900010(125) = 1250,\ 25(162.5) = 4062.5,\ 50(187.5) = 9375,\ 75(225) = 16\,875,\ 30(300) = 9000

      Frequency times midpoint for each bar.

    4. 4

      ∑fx\sum fx. 1250+4062.5+9375+16 875+9000=40 562.51250 + 4062.5 + 9375 + 16\,875 + 9000 = 40\,562.5

      The column total.

    5. 5

      Estimated mean. xˉ=40 562.5190=213.48…=213 minutes (3 s.f.)\bar{x} = \frac{40\,562.5}{190} = 213.48\ldots = 213 \text{ minutes (3 s.f.)}

      Σf = 190. Keep 213.48… for the variance.

    6. 6

      fx2fx^2 values. 10(1252)=156 250, 25(162.52)=660 156.25, 50(187.52)=1 757 812.5,10(125^2) = 156\,250,\ 25(162.5^2) = 660\,156.25,\ 50(187.5^2) = 1\,757\,812.5, 75(2252)=3 796 875, 30(3002)=2 700 00075(225^2) = 3\,796\,875,\ 30(300^2) = 2\,700\,000

      Midpoint squared, times frequency.

    7. 7

      ∑fx2\sum fx^2. 156 250+660 156.25+1 757 812.5+3 796 875+2 700 000=9 071 093.75156\,250 + 660\,156.25 + 1\,757\,812.5 + 3\,796\,875 + 2\,700\,000 = 9\,071\,093.75

      The column total.

    8. 8

      Mean of the squares. 9 071 093.75190=47 742.6\frac{9\,071\,093.75}{190} = 47\,742.6

      Divide by Σf = 190.

    9. 9

      Variance. σ2=47 742.6−213.48…2=47 742.6−45 576.6=2166.0\sigma^2 = 47\,742.6 - 213.48\ldots^2 = 47\,742.6 - 45\,576.6 = 2166.0

      Minus the square of the unrounded mean.

    10. 10

      Estimated standard deviation. σ=2166.0=46.5 minutes (3 s.f.)\sigma = \sqrt{2166.0} = 46.5 \text{ minutes (3 s.f.)}

      The mark scheme accepts 46.5 or 46.6.

    11. 11

      (iii) Why estimates. The midpoint of each class has been used instead of the actual (raw) times.

      The mark scheme's reason: mid-points used, not the raw data.

    Answer

    (ii) Mean ≈213\approx 213 minutes; standard deviation ≈46.5\approx 46.5 minutes. (iii) Class midpoints were used in place of the actual data values.

Practise the mean and standard deviation of grouped dataReal past-paper questions · Calculating mean and standard deviation from raw, grouped or coded data
12

Coded totals

Syllabus requirement · §5.1

“

calculate and use the mean and standard deviation of a set of data … or coded totals Σ(x−a)\Sigma(x - a) and Σ(x−a)2\Sigma(x - a)^2

”

Some questions don't give ∑x\sum x and ∑x2\sum x^2. Instead they give totals of x−ax - a for some constant aa: ∑(x−a)and∑(x−a)2.\sum (x - a) \quad \text{and} \quad \sum (x - a)^2. Subtracting a constant from every value is called coding. It is used to keep numbers small (heights near 170170 cm coded as x−170x - 170, for instance), and the questions then test whether you know what coding does to the mean and to the spread.

What coding does to the mean. Sliding every value down by aa slides the mean down by aa. Writing x−a‾\overline{x - a} for the mean of the coded values: x−a‾=xˉ−a,i.e.xˉ=∑(x−a)n+a.\overline{x - a} = \bar{x} - a, \qquad \text{i.e.} \qquad \bar{x} = \frac{\sum (x - a)}{n} + a.

What coding does to the spread. Nothing. Every value moves by the same amount, so every gap between values, and every distance from the mean, stays the same. The variance of x−ax - a is the variance of xx: σ2=∑(x−a)2n−(∑(x−a)n)2\sigma^2 = \frac{\sum (x - a)^2}{n} - \left(\frac{\sum (x - a)}{n}\right)^2 This is the §10 formula with x−ax - a in place of xx, and nothing is added back at the end.

Turning a coded total into an ordinary one. ∑(x−a)\sum (x - a) means (x1−a)+(x2−a)+⋯+(xn−a)(x_1 - a) + (x_2 - a) + \cdots + (x_n - a). Collecting the xx terms and the nn copies of aa: ∑(x−a)=∑x−na.\sum (x - a) = \sum x - na. This one line is what most "find kk" and "find nn" questions need.

-20246810−5x̄x̄range 9 − 4 = 5original: 4, 5, 5, 6, 7, 9 — mean 6coded (subtract 5): −1, 0, 0, 1, 2, 4 — mean 6 − 5 = 1, range still 5each point movesthe same distance,so every gap keepsits size: σ² unchanged

The six values from §10 before and after subtracting 5, on one scale. Every point slides 5 to the left, so the mean moves by 5 but every gap (and so the standard deviation) is unchanged.

Coded totals
∑(x−a)=∑x−na\sum (x - a) = \sum x - na

Removing the brackets: n copies of a

xˉ=∑(x−a)n+a\bar{x} = \frac{\sum (x - a)}{n} + a

The mean shifts by a

σ2=∑(x−a)2n−(∑(x−a)n)2\sigma^2 = \frac{\sum (x - a)^2}{n} - \left(\frac{\sum (x - a)}{n}\right)^2

The variance is unchanged by coding

Seeing that coding leaves the spread alone

The values 4, 7, 5, 9, 6, 54,\ 7,\ 5,\ 9,\ 6,\ 5 have mean 66 and variance 2.666…2.666\ldots (§10). Code them by subtracting 55, and use the coded values to find the mean and variance of the original data.

Show full working
  1. 1

    Coded values. Subtract 55 from each: −1, 2, 0, 4, 1, 0-1,\ 2,\ 0,\ 4,\ 1,\ 0

    Each value has become smaller and simpler to work with.

  2. 2

    Coded total. ∑(x−5)=−1+2+0+4+1+0=6\sum (x - 5) = -1 + 2 + 0 + 4 + 1 + 0 = 6

    Check with Σ(x − a) = Σx − na: 36 − 6 × 5 = 6 ✓.

  3. 3

    Coded mean. 66=1\frac{6}{6} = 1

    The mean of the coded values.

  4. 4

    Original mean. Add back the 55: xˉ=1+5=6\bar{x} = 1 + 5 = 6

    Matches the mean found directly in §10.

  5. 5

    Coded sum of squares. ∑(x−5)2=1+4+0+16+1+0=22\sum (x - 5)^2 = 1 + 4 + 0 + 16 + 1 + 0 = 22

    Square each coded value, then add.

  6. 6

    Mean of the coded squares. 226=3.666…\frac{22}{6} = 3.666\ldots

    Divide by n = 6.

  7. 7

    Variance. σ2=3.666…−12=2.666…\sigma^2 = 3.666\ldots - 1^2 = 2.666\ldots

    Exactly the variance found in §10. Nothing is added back for the variance.

Answer

Mean =1+5=6= 1 + 5 = 6; variance =2.666…= 2.666\ldots, the same as for the original data.

Coding shifts the mean and leaves the variance and standard deviation alone. That single fact covers most coded questions.

Finding the constant, then the variance

9709/51 O/N 2021 Q2(a)(b)4 marks

A summary of 4040 values of xx gives the following information:

∑(x−k)=520,∑(x−k)2=9640,\sum (x - k) = 520, \quad \sum (x - k)^2 = 9640,

where kk is a constant.

(a) Given that the mean of these 4040 values of xx is 3434, find the value of kk.

(b) Find the variance of these 4040 values of xx.

Show full working
  1. 1

    (a) Coded mean. ∑(x−k)40=52040=13\frac{\sum (x - k)}{40} = \frac{520}{40} = 13

    The mean of the values x − k.

  2. 2

    (a) Link it to the real mean. The coded mean is the real mean minus kk: 34−k=1334 - k = 13

    The M1 is for an equation linking Σx (or the mean), Σ(x − k) and k.

  3. 3

    (a) Solve. k=34−13=21k = 34 - 13 = 21

    A1.

  4. 4

    (b) Mean of the coded squares. 964040=241\frac{9640}{40} = 241

    Divide by n = 40.

  5. 5

    (b) Variance. σ2=241−132=241−169=72\sigma^2 = 241 - 13^2 = 241 - 169 = 72

    The coded variance formula (coded mean 13 from part (a)) gives the variance of x itself: no adjustment for k.

Answer

(a) k=21k = 21. (b) Variance =72= 72.

Part (b) never uses k. If you find yourself adding k to a variance, stop.

Mixed totals: a coded sum with an ordinary sum of squares

9709/53 O/N 2022 Q13 marks

5050 values of the variable xx are summarised by

Σ(x−20)=35andΣx2=25 036.\Sigma(x - 20) = 35 \quad \text{and} \quad \Sigma x^2 = 25\,036.

Find the variance of these 5050 values.

Show full working
  1. 1

    Notice the mismatch. One total is coded, ∑(x−20)\sum (x - 20), but the sum of squares is not: it is ∑x2\sum x^2. They cannot go into one formula together until both are about xx.

    This is the trap. Using 25 036 with the coded mean 35/50 gives a wrong answer.

  2. 2

    Convert the coded total. ∑x−50×20=35\sum x - 50 \times 20 = 35

    Σ(x − a) = Σx − na, with n = 50 and a = 20.

  3. 3

    Solve for ∑x\sum x. ∑x=35+1000=1035\sum x = 35 + 1000 = 1035

    The B1 is for Σx = 1035 (or the mean 20.7).

  4. 4

    Mean of xx. xˉ=103550=20.7\bar{x} = \frac{1035}{50} = 20.7

    Now the mean and Σx² both describe x itself.

  5. 5

    Mean of the squares. 25 03650=500.72\frac{25\,036}{50} = 500.72

    Σx² is uncoded, so it goes with the uncoded mean.

  6. 6

    Variance. σ2=500.72−20.72=500.72−428.49=72.23\sigma^2 = 500.72 - 20.7^2 = 500.72 - 428.49 = 72.23

    The ordinary formula. The mark scheme wants the exact answer 72.23.

Answer

Variance =72.23= 72.23.

Common mistakes
  • Adding aa (or a2a^2) to the variance of coded data

    The variance of x−ax - a is the variance of xx

    Shifting every value by the same amount doesn't change the spread.

  • ∑(x−a)=∑x−a\sum (x - a) = \sum x - a

    ∑(x−a)=∑x−na\sum (x - a) = \sum x - na

    a is subtracted from each of the n values.

  • Using a coded mean with an uncoded ∑x2\sum x^2 (or the other way round)

    Convert so that both totals describe the same variable

    The formula needs the mean and the sum of squares of the same quantity.

  • Forgetting to add aa back to the coded mean

    xˉ=∑(x−a)n+a\bar{x} = \dfrac{\sum (x - a)}{n} + a

    The coded mean is the mean of x − a, not of x.

Your turn

Decide first which totals are coded and which are not. Then use Σ(x − a) = Σx − na to link them.

  1. 19709/52 M/J 2022 Q13 marks

    For nn values of the variable xx, it is given that

    Σ(x−200)=446andΣx=6846.\Sigma(x - 200) = 446 \quad \text{and} \quad \Sigma x = 6846.

    Find the value of nn.

    Stuck? Show hint

    ∑(x−200)=∑x−200n\sum (x - 200) = \sum x - 200n.

    Show solution
    1. 1

      Remove the brackets. ∑x−200n=446\sum x - 200n = 446

      B1 for this three-term equation; B1 for Σ200 = 200n.

    2. 2

      Substitute ∑x\sum x. 6846−200n=4466846 - 200n = 446

      Σx is given.

    3. 3

      Add 200n200n to both sides. 6846=446+200n6846 = 446 + 200n

      Get the n term on its own side.

    4. 4

      Subtract 446446. 200n=6846−446=6400200n = 6846 - 446 = 6400

      Now only the n term is on the right.

    5. 5

      Divide by 200200. n=32n = 32

      Final B1.

    Answer

    n=32n = 32.

  2. 29709/51 M/J 2023 Q1(a)(b)4 marks

    A summary of 5050 values of xx gives

    ∑(x−q)=700,∑(x−q)2=14 235,\sum(x - q) = 700, \quad \sum(x - q)^2 = 14\,235,

    where qq is a constant.

    (a) Find the standard deviation of these values of xx.

    (b) Given that ∑x=2865\sum x = 2865, find the value of qq.

    Stuck? Show hint

    (a) needs only the coded totals. (b) Use ∑(x−q)=∑x−50q\sum (x - q) = \sum x - 50q.

    Show solution
    1. 1

      (a) Coded mean. 70050=14\frac{700}{50} = 14

      The mean of x − q.

    2. 2

      (a) Mean of the coded squares. 14 23550=284.7\frac{14\,235}{50} = 284.7

      Divide by n = 50.

    3. 3

      (a) Variance. 284.7−142=284.7−196=88.7284.7 - 14^2 = 284.7 - 196 = 88.7

      The coded variance is the variance of x.

    4. 4

      (a) Standard deviation. 88.7=9.42 (3 s.f.)\sqrt{88.7} = 9.42 \text{ (3 s.f.)}

      9.4180… to at least 3 s.f.

    5. 5

      (b) Remove the brackets. ∑x−50q=700\sum x - 50q = 700

      The M1 is for forming this equation.

    6. 6

      (b) Substitute ∑x\sum x. 2865−50q=7002865 - 50q = 700

      Σx is given.

    7. 7

      (b) Add 50q50q to both sides. 2865=700+50q2865 = 700 + 50q

      Get the q term on its own side.

    8. 8

      (b) Subtract 700700. 50q=216550q = 2165

      2865 − 700 = 2165.

    9. 9

      (b) Divide by 5050. q=43.3q = 43.3

      A1.

    Answer

    (a) 9.429.42. (b) q=43.3q = 43.3.

  3. 39709/53 M/J 2025 Q1(a)(b)4 marks

    For a set of 4040 values of xx, it is found that

    ∑(x−k)=836.0,∑(x−k)2=25410.8,\sum (x - k) = 836.0, \quad \sum (x - k)^2 = 25410.8,

    where kk is a constant.

    (a) Given that the mean of these 4040 values is 124.0124.0, find the value of kk.

    (b) Find the standard deviation of these 4040 values of xx.

    Stuck? Show hint

    (a) ∑x=40×124\sum x = 40 \times 124. (b) Coded variance formula.

    Show solution
    1. 1

      (a) ∑x\sum x from the mean. ∑x=40×124=4960\sum x = 40 \times 124 = 4960

      Total from a mean.

    2. 2

      (a) Equation. 4960−40k=8364960 - 40k = 836

      Σ(x − k) = Σx − 40k. B1.

    3. 3

      (a) Add 40k40k to both sides. 4960=836+40k4960 = 836 + 40k

      Get the k term on its own side.

    4. 4

      (a) Subtract 836836. 40k=412440k = 4124

      4960 − 836 = 4124.

    5. 5

      (a) Divide by 4040. k=103.1k = 103.1

      B1.

    6. 6

      (b) Coded mean. 83640=20.9\frac{836}{40} = 20.9

      The mean of x − k.

    7. 7

      (b) Mean of the coded squares. 25 410.840=635.27\frac{25\,410.8}{40} = 635.27

      Divide by n = 40.

    8. 8

      (b) Variance. 635.27−20.92=635.27−436.81=198.46635.27 - 20.9^2 = 635.27 - 436.81 = 198.46

      Coded totals give the variance of x directly.

    9. 9

      (b) Standard deviation. 198.46=14.1 (3 s.f.)\sqrt{198.46} = 14.1 \text{ (3 s.f.)}

      14.087… to 3 s.f.

    Answer

    (a) k=103.1k = 103.1. (b) 14.114.1.

  4. 49709/62 M/J 2011 Q3(i)(ii)7 marks

    A sample of 3636 data values, xx, gave ∑(x−45)=−148\sum(x - 45) = -148 and ∑(x−45)2=3089\sum(x - 45)^2 = 3089.

    (i) Find the mean and standard deviation of the 3636 values.

    (ii) One extra data value of 2929 was added to the sample. Find the standard deviation of all 3737 values.

    Stuck? Show hint

    (ii) Code the new value too: 29−45=−1629 - 45 = -16. Add it to the coded total, and its square to the coded sum of squares.

    Show solution
    1. 1

      (i) Coded mean. −14836=−4.11…\frac{-148}{36} = -4.11\ldots

      The mean of x − 45.

    2. 2

      (i) Add back 4545. xˉ=−4.11…+45=40.9 (3 s.f.)\bar{x} = -4.11\ldots + 45 = 40.9 \text{ (3 s.f.)}

      The mean shifts back by a.

    3. 3

      (i) Mean of the coded squares. 308936=85.81\frac{3089}{36} = 85.81

      Divide by n = 36.

    4. 4

      (i) Variance. 85.81−(−4.11…)2=85.81−16.90=68.9085.81 - (-4.11\ldots)^2 = 85.81 - 16.90 = 68.90

      Coded variance formula; nothing added back.

    5. 5

      (i) Standard deviation. 68.90=8.30\sqrt{68.90} = 8.30

      3 s.f.

    6. 6

      (ii) Code the new value. 29−45=−1629 - 45 = -16.

      Keep everything in coded form.

    7. 7

      (ii) New coded total. ∑(x−45)=−148+(−16)=−164\sum (x - 45) = -148 + (-16) = -164

      Add the coded new value. M1.

    8. 8

      (ii) New coded sum of squares. ∑(x−45)2=3089+(−16)2=3089+256=3345\sum (x - 45)^2 = 3089 + (-16)^2 = 3089 + 256 = 3345

      Add the square of the coded new value. M1.

    9. 9

      (ii) New mean of the coded squares. 334537=90.41\frac{3345}{37} = 90.41

      Now n = 37.

    10. 10

      (ii) New variance. 90.41−(−16437)2=90.41−19.65=70.7690.41 - \left(\frac{-164}{37}\right)^2 = 90.41 - 19.65 = 70.76

      The new coded mean is −164/37 = −4.43…

    11. 11

      (ii) New standard deviation. 70.76=8.41\sqrt{70.76} = 8.41

      The spread grows slightly: 29 is about 12 below the mean of 40.9, more than one standard deviation (8.30) away.

    Answer

    (i) Mean 40.940.9, standard deviation 8.308.30. (ii) 8.418.41.

  5. 59709/61 M/J 2018 Q13 marks

    In a statistics lesson 1212 people were asked to think of a number, xx, between 11 and 2020 inclusive. From the results Tom found that ∑x=186\sum x = 186 and that the standard deviation of xx is 4.54.5. Assuming that Tom's calculations are correct, find the values of ∑(x−10)\sum (x - 10) and ∑(x−10)2\sum (x - 10)^2.

    Stuck? Show hint

    First ∑(x−10)=∑x−12×10\sum (x - 10) = \sum x - 12 \times 10. Then put everything into the coded variance formula and solve for ∑(x−10)2\sum (x - 10)^2.

    Show solution
    1. 1

      Coded total. ∑(x−10)=186−12×10=66\sum (x - 10) = 186 - 12 \times 10 = 66

      B1.

    2. 2

      Coded variance formula with σ=4.5\sigma = 4.5. ∑(x−10)212−(6612)2=4.52\frac{\sum (x - 10)^2}{12} - \left(\frac{66}{12}\right)^2 = 4.5^2

      The variance of x − 10 equals the variance of x. The M1 is for this substitution.

    3. 3

      Evaluate the known terms. ∑(x−10)212−30.25=20.25\frac{\sum (x - 10)^2}{12} - 30.25 = 20.25

      (66/12)² = 5.5² = 30.25 and 4.5² = 20.25.

    4. 4

      Add 30.2530.25 to both sides. ∑(x−10)212=50.5\frac{\sum (x - 10)^2}{12} = 50.5

      Isolate the fraction.

    5. 5

      Multiply by 1212. ∑(x−10)2=606\sum (x - 10)^2 = 606

      B1.

    Answer

    ∑(x−10)=66\sum (x - 10) = 66, ∑(x−10)2=606\sum (x - 10)^2 = 606.

Practise coded totalsReal past-paper questions · Calculating mean and standard deviation from raw, grouped or coded data
13

Combining two data sets

Syllabus requirement · §5.1

“

calculate and use the mean and standard deviation of a set of data … and use such totals in solving problems which may involve up to two data sets.

”

When two groups are put together (two teams, two classes, two companies) the combined mean and standard deviation come from adding the totals, never from averaging the means or the standard deviations.

If the first group has nxn_x values xx and the second has nyn_y values yy, the combined group has nx+nyn_x + n_y values, and

combined mean=∑x+∑ynx+ny,combined variance=∑x2+∑y2nx+ny−(combined mean)2.\text{combined mean} = \frac{\sum x + \sum y}{n_x + n_y}, \qquad \text{combined variance} = \frac{\sum x^2 + \sum y^2}{n_x + n_y} - (\text{combined mean})^2.

These are just the §10 formulas applied to the combined totals. The only work is getting all four totals, ∑x\sum x, ∑y\sum y, ∑x2\sum x^2, ∑y2\sum y^2:

  • given directly: use them;
  • given a mean instead of a total: ∑y=ny yˉ\sum y = n_y\,\bar{y};
  • given a standard deviation instead of a sum of squares: ∑y2=ny(σy2+yˉ2)\sum y^2 = n_y\left(\sigma_y^2 + \bar{y}^2\right) (§10);
  • given coded totals with the same code for both groups: add the coded totals directly, then add aa back to the combined coded mean (§12).

A question can also run backwards: given the combined standard deviation, find one of the missing totals. Write the combined variance formula with the unknown in it, then solve.

Why the means cannot simply be averaged

Group AA has 22 values with mean 1010. Group BB has 88 values with mean 2020. Find the mean of all 1010 values, and compare it with the average of 1010 and 2020.

Show full working
  1. 1

    Total of group AA. ∑x=2×10=20\sum x = 2 \times 10 = 20

    A total from a mean: Σ = n × mean.

  2. 2

    Total of group BB. ∑y=8×20=160\sum y = 8 \times 20 = 160

    Same for the second group.

  3. 3

    Combined total. 20+160=18020 + 160 = 180

    Totals can be added; means cannot.

  4. 4

    Combined mean. 1802+8=18010=18\frac{180}{2 + 8} = \frac{180}{10} = 18

    Divide by the combined count.

  5. 5

    Compare with averaging the two means. 10+202=15≠18\frac{10 + 20}{2} = 15 \ne 18

    Eight of the ten values come from group B, so the combined mean must be much nearer 20 than 10. Averaging the means treats a group of 2 as equal to a group of 8.

Answer

Combined mean =18= 18, not 1515.

Combined mean and standard deviation

Class AA has 1515 students with mean mark 6262 and ∑x2=58 320\sum x^2 = 58\,320. Class BB has 1010 students with mean mark 7070 and ∑y2=49 300\sum y^2 = 49\,300. Find the mean and standard deviation of the marks of all 2525 students.

Show full working
  1. 1

    Total of class AA. ∑x=15×62=930\sum x = 15 \times 62 = 930

    The sums of squares are given, but the plain totals must be rebuilt from the means.

  2. 2

    Total of class BB. ∑y=10×70=700\sum y = 10 \times 70 = 700

    Same for the second class.

  3. 3

    Combined total. 930+700=1630930 + 700 = 1630

    Add the two totals.

  4. 4

    Combined mean. 163025=65.2\frac{1630}{25} = 65.2

    Divide by the combined count, 15 + 10 = 25.

  5. 5

    Combined sum of squares. 58 320+49 300=107 62058\,320 + 49\,300 = 107\,620

    Sums of squares add in exactly the same way.

  6. 6

    Mean of the squares. 107 62025=4304.8\frac{107\,620}{25} = 4304.8

    Divide by the combined count.

  7. 7

    Combined variance. 4304.8−65.22=4304.8−4251.04=53.764304.8 - 65.2^2 = 4304.8 - 4251.04 = 53.76

    Minus the square of the combined mean.

  8. 8

    Combined standard deviation. 53.76=7.33 (3 s.f.)\sqrt{53.76} = 7.33 \text{ (3 s.f.)}

    Square root last.

Answer

Combined mean =65.2= 65.2; combined standard deviation =7.33= 7.33.

Combining two teams from their totals

9709/53 M/J 2021 Q3(a)(b)5 marks

A sports club has a volleyball team and a hockey team. The heights of the 66 members of the volleyball team are summarised by ∑x=1050\sum x = 1050 and ∑x2=193 700\sum x^2 = 193\,700, where xx is the height of a member in cm. The heights of the 1111 members of the hockey team are summarised by ∑y=1991\sum y = 1991 and ∑y2=366 400\sum y^2 = 366\,400, where yy is the height of a member in cm.

(a) Find the mean height of all 1717 members of the club.

(b) Find the standard deviation of the heights of all 1717 members of the club.

Show full working
  1. 1

    (a) Combined total. ∑x+∑y=1050+1991=3041\sum x + \sum y = 1050 + 1991 = 3041

    Add the two totals.

  2. 2

    (a) Combined mean. 30416+11=304117=178.88…=178.9 cm (1 d.p.)\frac{3041}{6 + 11} = \frac{3041}{17} = 178.88\ldots = 178.9 \text{ cm (1 d.p.)}

    The mark scheme accepts 178.9, 178.88 or 179. Keep 3041/17 for part (b).

  3. 3

    (b) Combined sum of squares. 193 700+366 400=560 100193\,700 + 366\,400 = 560\,100

    Add the sums of squares. First M1.

  4. 4

    (b) Mean of the squares. 560 10017=32 947.06\frac{560\,100}{17} = 32\,947.06

    Divide by 17.

  5. 5

    (b) Subtract the square of the mean. σ2=32 947.06−(304117)2=32 947.06−31 998.90=948.16\sigma^2 = 32\,947.06 - \left(\frac{3041}{17}\right)^2 = 32\,947.06 - 31\,998.90 = 948.16

    Second M1: the variance formula with the mean squared. Using the exact 3041/17 avoids rounding trouble.

  6. 6

    (b) Standard deviation. σ=948.16=30.8 cm (3 s.f.)\sigma = \sqrt{948.16} = 30.8 \text{ cm (3 s.f.)}

    The mark scheme also accepts 30.7 (from a rounded mean).

Answer

(a) 178.9178.9 cm. (b) 30.830.8 cm.

Working backwards from a combined standard deviation

9709/51 M/J 2025 Q3(c)(d)5 marks

Last Sunday, teams of runners took part in a charity event. The time taken, in seconds, to run 50 m50\text{ m} was recorded, correct to 1 decimal place, for each runner. (Part (a), on the Gulls and Herons, is in the §02 exercises.)

Two other teams of runners, the Eagles and the Swifts, also took part in the event. The recorded times in seconds for 2020 runners from the Eagles and 3030 runners from the Swifts are denoted by xx and yy respectively.

It is given that ∑x=175.0\sum x = 175.0 and that the mean of yy is 8.48.4.

(c) Find the mean of the times taken by all 5050 runners.

It is given that ∑x2=1823.0\sum x^2 = 1823.0.

It is also known that the standard deviation of the times taken by all 5050 runners is 1.381.38 seconds.

(d) Find the value of ∑y2\sum y^2, correct to 11 decimal place.

Show full working
  1. 1

    (c) The Swifts' total from their mean. ∑y=30×8.4=252.0\sum y = 30 \times 8.4 = 252.0

    Only the mean of y is given, so rebuild the total.

  2. 2

    (c) Combined total. 175.0+252.0=427.0175.0 + 252.0 = 427.0

    Add the two totals.

  3. 3

    (c) Combined mean. 427.050=8.54 s\frac{427.0}{50} = 8.54 \text{ s}

    M1 A1.

  4. 4

    (d) Write the combined variance with the unknown. 1.382=1823.0+∑y250−8.5421.38^2 = \frac{1823.0 + \sum y^2}{50} - 8.54^2

    The combined standard deviation is 1.38, so the combined variance is 1.38². This is the M1.

  5. 5

    (d) Evaluate the numbers. 1.9044=1823.0+∑y250−72.93161.9044 = \frac{1823.0 + \sum y^2}{50} - 72.9316

    1.38² = 1.9044 and 8.54² = 72.9316.

  6. 6

    (d) Add 72.931672.9316 to both sides. 74.836=1823.0+∑y25074.836 = \frac{1823.0 + \sum y^2}{50}

    Start isolating the unknown. The DM1 is for rearranging to find Σy².

  7. 7

    (d) Multiply by 5050. 3741.8=1823.0+∑y23741.8 = 1823.0 + \sum y^2

    Clear the fraction.

  8. 8

    (d) Subtract 1823.01823.0. ∑y2=1918.8\sum y^2 = 1918.8

    To 1 decimal place, as asked.

Answer

(c) 8.548.54 s. (d) ∑y2=1918.8\sum y^2 = 1918.8.

Working backwards: write the combined formula with the unknown total in it, substitute every number you know, then undo the operations one at a time.

Common mistakes
  • Combined mean =xˉ+yˉ2= \dfrac{\bar{x} + \bar{y}}{2}

    Combined mean =∑x+∑ynx+ny= \dfrac{\sum x + \sum y}{n_x + n_y}

    Averaging means is only right when the groups are the same size.

  • Combined standard deviation found by averaging the two standard deviations

    Add the sums of squares and use the variance formula

    Standard deviations don't add, and the spread between the two group means matters too.

  • Using ∑x2\sum x^2 and ∑y2\sum y^2 when only means and standard deviations were given

    First rebuild ∑x2=nx(σx2+xˉ2)\sum x^2 = n_x(\sigma_x^2 + \bar{x}^2) for each group

    The formulas need totals, so convert everything to totals first.

  • Answers left in coded units or in thousands

    Convert back: add aa to a coded mean; multiply by 10001000 if the data was in thousands

    The final answer must be in the units of the original quantity.

Your turn

Get all four totals first, converting where necessary. Then add them and use the formulas.

  1. 19709/52 O/N 2024 Q6(c)3 marks

    Teams of 1515 runners took part in a charity run last Saturday. Let xx and yy denote the times, in minutes, of a runner from the Falcons and a runner from the Kites respectively.

    It is given that

    ∑x=792,∑x2=43 504,∑y=783,∑y2=42 223.\sum x = 792, \quad \sum x^2 = 43\,504, \quad \sum y = 783, \quad \sum y^2 = 42\,223.

    Find the mean and the standard deviation of the times taken by all 3030 runners from the two teams.

    Stuck? Show hint

    All four totals are given.

    Show solution
    1. 1

      Combined total. 792+783=1575792 + 783 = 1575

      Add the two totals.

    2. 2

      Combined mean. 157530=52.5 minutes\frac{1575}{30} = 52.5 \text{ minutes}

      B1 needs the value 52.5, not just 1575/30.

    3. 3

      Combined sum of squares. 43 504+42 223=85 72743\,504 + 42\,223 = 85\,727

      Add the sums of squares.

    4. 4

      Mean of the squares. 85 72730=2857.57\frac{85\,727}{30} = 2857.57

      Divide by the combined count.

    5. 5

      Variance. 2857.57−52.52=2857.57−2756.25=101.322857.57 - 52.5^2 = 2857.57 - 2756.25 = 101.32

      M1.

    6. 6

      Standard deviation. 101.32=10.1 minutes (3 s.f.)\sqrt{101.32} = 10.1 \text{ minutes (3 s.f.)}

      A1; it must be identified as the standard deviation.

    Answer

    Mean 52.552.5 minutes; standard deviation 10.110.1 minutes.

  2. 29709/51 M/J 2024 Q1(a)(b)4 marks

    A summary of 2020 values of xx gives

    ∑(x−30)=439,∑(x−30)2=12 405.\sum(x - 30) = 439, \quad \sum(x - 30)^2 = 12\,405.

    A summary of another 2525 values of xx gives

    ∑(x−30)=470,∑(x−30)2=11 346.\sum(x - 30) = 470, \quad \sum(x - 30)^2 = 11\,346.

    (a) Find the mean of all 4545 values of xx.

    (b) Find the standard deviation of all 4545 values of xx.

    Stuck? Show hint

    Both groups use the same code, so the coded totals can be added directly.

    Show solution
    1. 1

      (a) Combined coded total. 439+470=909439 + 470 = 909

      Same code (x − 30) for both groups.

    2. 2

      (a) Combined coded mean. 90945=20.2\frac{909}{45} = 20.2

      The mean of x − 30.

    3. 3

      (a) Mean of xx. 20.2+30=50.220.2 + 30 = 50.2

      Add the 30 back.

    4. 4

      (b) Combined coded sum of squares. 12 405+11 346=23 75112\,405 + 11\,346 = 23\,751

      Add them the same way.

    5. 5

      (b) Mean of the coded squares. 23 75145=527.8\frac{23\,751}{45} = 527.8

      Divide by 45.

    6. 6

      (b) Variance. 527.8−20.22=527.8−408.04=119.76527.8 - 20.2^2 = 527.8 - 408.04 = 119.76

      Coded mean squared; nothing is added back for the variance.

    7. 7

      (b) Standard deviation. 119.76=10.9 (3 s.f.)\sqrt{119.76} = 10.9 \text{ (3 s.f.)}

      10.94…

    Answer

    (a) 50.250.2. (b) 10.910.9.

  3. 39709/51 O/N 2025 Q3(b)5 marks

    The back-to-back stem-and-leaf diagram shows the annual salaries, in dollars, of 2727 employees at each of two companies, Browns and Greens.

    The annual salary of an employee at Browns is denoted by xx thousand dollars and the annual salary of an employee at Greens is denoted by yy thousand dollars. It is given that, for the 2727 employees at each of the companies,

    ∑x=878.3,∑x2=28 616.09,∑y=896.5,∑y2=29 815.63.\sum x = 878.3,\quad \sum x^2 = 28\,616.09,\quad \sum y = 896.5,\quad \sum y^2 = 29\,815.63.

    Find the mean and standard deviation of the annual salaries of these 5454 employees.

    The printed diagram (the same one as in the §03 worked example).

    The printed diagram (the same one as in the §03 worked example).

    Stuck? Show hint

    Work in thousands of dollars, then convert both answers to dollars at the end.

    Show solution
    1. 1

      Combined total, in thousands. 878.3+896.5=1774.8878.3 + 896.5 = 1774.8

      Add the two totals.

    2. 2

      Combined mean, in thousands. 1774.854=32.866…\frac{1774.8}{54} = 32.866\ldots

      B1 for this expression.

    3. 3

      Convert to dollars. 32.866…×1000=$32 900 (3 s.f.)32.866\ldots \times 1000 = \$32\,900 \text{ (3 s.f.)}

      Second B1: the answer in dollars.

    4. 4

      Combined sum of squares. 28 616.09+29 815.63=58 431.7228\,616.09 + 29\,815.63 = 58\,431.72

      Add the sums of squares.

    5. 5

      Mean of the squares. 58 431.7254=1082.069\frac{58\,431.72}{54} = 1082.069

      Divide by 54.

    6. 6

      Variance, in (thousands of dollars)². 1082.069−(1774.854)2=1082.069−1080.218=1.85111082.069 - \left(\frac{1774.8}{54}\right)^2 = 1082.069 - 1080.218 = 1.8511

      M1 A1. Keep plenty of figures: the two terms are very close, so rounding early wrecks the answer.

    7. 7

      Standard deviation. 1.8511=1.36 thousand dollars=$1360\sqrt{1.8511} = 1.36 \text{ thousand dollars} = \$1360

      Final A1, converted to dollars.

    Answer

    Mean ≈ $32 900; standard deviation ≈ $1360.

  4. 49709/63 M/J 2018 Q4(i)(ii)7 marks

    Farfield Travel and Lacket Travel are two travel companies which arrange tours abroad. The numbers of holidays arranged in a certain week are recorded in the table below, together with the means and standard deviations of the prices.

    Number of holidaysMean price ($)Standard deviation ($)
    Farfield Travel301500230
    Lacket Travel212400160

    (i) Calculate the mean price of all 5151 holidays.

    (ii) The prices of individual holidays with Farfield Travel are denoted by $xFx_F and the prices of individual holidays with Lacket Travel are denoted by $xLx_L. By first finding ΣxF2\Sigma x_F^2 and ΣxL2\Sigma x_L^2, find the standard deviation of the prices of all 5151 holidays.

    Stuck? Show hint

    ∑x2=n(σ2+xˉ2)\sum x^2 = n\left(\sigma^2 + \bar{x}^2\right) for each company.

    Show solution
    1. 1

      (i) Farfield's total. 30×1500=45 00030 \times 1500 = 45\,000

      Total from a mean.

    2. 2

      (i) Lacket's total. 21×2400=50 40021 \times 2400 = 50\,400

      Total from a mean.

    3. 3

      (i) Combined total. 45 000+50 400=95 40045\,000 + 50\,400 = 95\,400

      Add the totals.

    4. 4

      (i) Combined mean. 95 40051=1870.59…=$1870 (3 s.f.)\frac{95\,400}{51} = 1870.59\ldots = \$1870 \text{ (3 s.f.)}

      Keep 1870.59 for part (ii).

    5. 5

      (ii) Farfield's sum of squares. ∑xF2=30(2302+15002)=30(52 900+2 250 000)=69 087 000\sum x_F^2 = 30\left(230^2 + 1500^2\right) = 30(52\,900 + 2\,250\,000) = 69\,087\,000

      Rearranged variance formula. M1 A1.

    6. 6

      (ii) Lacket's sum of squares. ∑xL2=21(1602+24002)=21(25 600+5 760 000)=121 497 600\sum x_L^2 = 21\left(160^2 + 2400^2\right) = 21(25\,600 + 5\,760\,000) = 121\,497\,600

      A1.

    7. 7

      (ii) Combined sum of squares. 69 087 000+121 497 600=190 584 60069\,087\,000 + 121\,497\,600 = 190\,584\,600

      Add the sums of squares.

    8. 8

      (ii) Mean of the squares. 190 584 60051=3 736 953\frac{190\,584\,600}{51} = 3\,736\,953

      Divide by 51.

    9. 9

      (ii) Combined variance. 3 736 953−1870.592=3 736 953−3 499 100=237 8533\,736\,953 - 1870.59^2 = 3\,736\,953 - 3\,499\,100 = 237\,853

      M1: subtract the combined mean squared.

    10. 10

      (ii) Standard deviation. 237 853=$488 (3 s.f.)\sqrt{237\,853} = \$488 \text{ (3 s.f.)}

      The mark scheme accepts anything from 486 to 490.

    Answer

    (i) $1870. (ii) $488.

Practise combining two data setsReal past-paper questions · Calculating mean and standard deviation from raw, grouped or coded data

Everything on one page

median at position 12(n+1);Q1,Q3=medians of the lower and upper halves\text{median at position } \tfrac12(n+1); \quad Q_1, Q_3 = \text{medians of the lower and upper halves}

A list of n values (for odd n: Q₁ at ¼(n+1), Q₃ at ¾(n+1))

IQR=Q3−Q1,range=max−min\text{IQR} = Q_3 - Q_1, \qquad \text{range} = \text{max} - \text{min}

Measures of spread

frequency density=frequencyclass width,frequency=frequency density×class width\text{frequency density} = \frac{\text{frequency}}{\text{class width}}, \qquad \text{frequency} = \text{frequency density} \times \text{class width}

Histogram height; area = frequency

median at 12n,Q1 at 14n,Q3 at 34n\text{median at } \tfrac12 n, \quad Q_1 \text{ at } \tfrac14 n, \quad Q_3 \text{ at } \tfrac34 n

Grouped data and cumulative frequency graphs (no +1)

greatest IQR=top of Q3 class−bottom of Q1 class\text{greatest IQR} = \text{top of } Q_3 \text{ class} - \text{bottom of } Q_1 \text{ class}

From a grouped frequency table

least IQR=bottom of Q3 class−top of Q1 class\text{least IQR} = \text{bottom of } Q_3 \text{ class} - \text{top of } Q_1 \text{ class}

From a grouped frequency table (0 if both quartiles share a class)

upper fence=Q3+1.5 IQR,lower fence=Q1−1.5 IQR\text{upper fence} = Q_3 + 1.5\,\text{IQR}, \quad \text{lower fence} = Q_1 - 1.5\,\text{IQR}

Outliers, when a question defines them this way

xˉ=∑xn,σ2=∑x2n−xˉ2\bar{x} = \frac{\sum x}{n}, \qquad \sigma^2 = \frac{\sum x^2}{n} - \bar{x}^2

Mean and variance (in the formula booklet)

∑x=nxˉ,∑x2=n(σ2+xˉ2)\sum x = n\bar{x}, \qquad \sum x^2 = n\left(\sigma^2 + \bar{x}^2\right)

Totals back from a mean and a standard deviation

xˉ=∑fx∑f,σ2=∑fx2∑f−xˉ2\bar{x} = \frac{\sum fx}{\sum f}, \qquad \sigma^2 = \frac{\sum fx^2}{\sum f} - \bar{x}^2

Grouped data, x = class midpoint (estimates)

∑(x−a)=∑x−na,xˉ=∑(x−a)n+a\sum (x - a) = \sum x - na, \qquad \bar{x} = \frac{\sum (x - a)}{n} + a

Coded totals: the mean shifts by a

σ2=∑(x−a)2n−(∑(x−a)n)2\sigma^2 = \frac{\sum (x - a)^2}{n} - \left(\frac{\sum (x - a)}{n}\right)^2

Coded totals: the variance is unchanged

combined mean=∑x+∑ynx+ny,combined variance=∑x2+∑y2nx+ny−(combined mean)2\text{combined mean} = \frac{\sum x + \sum y}{n_x + n_y}, \quad \text{combined variance} = \frac{\sum x^2 + \sum y^2}{n_x + n_y} - (\text{combined mean})^2

Two data sets: add the totals

Can you do all of these?

  • State an advantage or disadvantage of a diagram by saying what it keeps or loses (not median, IQR or range for a stem-and-leaf against a box plot)

  • Draw a back-to-back stem-and-leaf diagram: complete stem, left leaves increasing right to left, aligned, no commas, one key naming both sets and the units

  • Find the median and quartiles from a list or a printed stem-and-leaf diagram, reading left-hand rows outwards from the stem and converting with the key

  • Draw a pair of box plots on one labelled linear scale, whiskers from the middle of each end of the box to the exact extremes, each plot labelled

  • Apply an outlier definition given in a question: Q₃ + 1.5 IQR and Q₁ − 1.5 IQR

  • Write two comparisons in context, one about centre and one about spread (for times: smaller = quicker)

  • Say why the median beats the mean by naming the extreme value; say “not symmetrical” when the mean and median differ

  • Find class boundaries for rounded, truncated and inequality classes before any width or midpoint

  • Draw a histogram with frequency density, touching bars at the boundaries, and both axes labelled; read a frequency back as density × width

  • Find the class containing the median or a quartile from running totals, and the greatest and least possible IQR

  • Plot cumulative frequency at upper boundaries, starting at (lower boundary of first class, 0), with a smooth curve

  • Start a cumulative frequency curve at 0, not a negative boundary, for a quantity that cannot be negative

  • Read a cumulative frequency graph at ¼n, ½n, ¾n or p% of n, turning “or more” and “more than” round, and show the reading lines

  • Find the mean and standard deviation from data or from Σx and Σx², and rebuild Σx² = n(σ² + x̄²) when a value is added or removed

  • Estimate a grouped mean and standard deviation with midpoints, differencing a cumulative table first

  • Use Σ(x − a) = Σx − na; shift the mean by a; leave the variance alone

  • Combine two data sets by adding totals, and work backwards from a combined standard deviation

Now do the questions
128 real Paper 5 parts from 2021–2025, sorted by difficulty, with mark schemes