Notes/Mathematics/Paper 5/The Normal Distribution
CAIEAS Level9709§5.5

The Normal Distribution

The normal curve and its tables, standardising to find a probability, expected numbers and intervals about the mean, working backwards to an unknown boundary, mean or standard deviation (one or two at a time), and the normal approximation to the binomial with a continuity correction.

300 min read 9 sub-topics
112
question parts
2021–2025 · 37 papers
12 marks
per paper
≈ 23% of the paper
2.7/3
avg difficulty
demanding
#2
most examined
of 5 topics by marks

Every distribution in the Discrete Random Variables topic (syllabus 5.4) was discrete: a list of separate values (0, 1, 2, …), each with its own probability. This topic is about a continuous variable, one that can take any value in a range, such as a height, a mass or a time. For a continuous variable a single exact value has probability 00. Only an interval of values has a probability, and that probability is an area under a curve. The curve used for this whole topic is the bell-shaped normal curve.

It is one of the two heaviest topics on Paper 5. Across 2021–2025 it carried 429 marks over 112 tagged parts, about 11.6 of the 50 marks on every paper, and all 37 papers in that window had a normal distribution question:

topicmarks/paper
Discrete Random Variables13.2
The Normal Distribution11.6
Probability11.2
Permutations and Combinations9.8
Representation of Data9.3

Inside the topic the marks split three ways (a part can carry more than one tag, so the rows overlap):

sub-topicpartsmarkssections
Solving problems using standardisation (ZZ-scores)8731502–08
Normal approximation to the binomial, with continuity correction3014209
The normal distribution as a model for a continuous variable258901

Its parts are also harder than average: the mean difficulty tag is 2.722.72 on the bank's 11–44 scale, the highest of the five Paper 5 topics.

The questions follow a small number of patterns, and this note gives each one its own section. Forwards (§03–§05): you are given μ\mu, σ\sigma and a boundary, and you find a probability, often turned into an expected number of items. Backwards (§06–§08): you are given a probability and must find a boundary, the mean, the standard deviation, or both of them together. The approximation (§09) is always a 5-mark part, and it turned up on 28 of the 37 papers.

One habit runs through every section: sketch the curve and shade the region before you write any algebra. A ten-second sketch tells you whether the answer must be above or below 0.50.5 and whether a zz-value must be positive or negative. Using Φ(z)\Phi(z) where 1−Φ(z)1 - \Phi(z) was needed, or giving zz the wrong sign, loses more marks in this topic than any other slip. (Φ\Phi and zz are explained in §02 and §03.)

Reading the notes in this topic.

  • "§03" means section 03 of this note. Other syllabus topics are named in words, e.g. "the Probability topic (5.3)".
  • A source such as "9709/52 M/J 2024 Q3(b)" means Paper 5, variant 2, the May/June 2024 series, question 3(b). F/M is February/March and O/N is October/November. Papers up to 2019 were called Statistics 1 and numbered 61–63.
  • The "why" notes on worked past papers say which mark each line earns, using Cambridge's codes. M1 is a method mark, for a correct method even if the arithmetic later slips. A1 is an accuracy mark, for a correct answer, and needs the method mark before it. B1 is a mark for a correct value on its own, such as a zz-value. FT (follow through) means the mark can still be earned using your own wrong earlier answer correctly. CAO means correct answer only. AWRT means any answer that rounds to the stated value is accepted.
Before you start you should be able to
  • Discrete Random Variables (5.4): a binomial count X∼B(n,p)X \sim \mathrm{B}(n, p) has mean npnp and variance np(1−p)np(1-p), and "at least", "more than", "fewer than", "at most" pick out different whole numbers

  • Probability (5.3): P(not A)=1−P(A)\mathrm{P}(\text{not } A) = 1 - \mathrm{P}(A); for independent events, multiply the probabilities; for mutually exclusive events, add them

  • Representation of Data (5.1): the standard deviation is the square root of the variance

  • Rearranging a linear equation such as a−μσ=z\dfrac{a - \mu}{\sigma} = z to make any one letter the subject, and solving two simultaneous linear equations

By the end of this page you can
  • Describe the normal distribution as a model for a continuous variable, read X∼N(μ,σ2)X \sim \mathrm{N}(\mu, \sigma^2) correctly (the second number is the variance), sketch normal curves, and say when a normal model is suitable

  • Read Φ(z)\Phi(z) from the MF19 table, including a third decimal place of zz from the ADD columns, and use symmetry to find left tails, right tails, negative-zz areas and areas between two values

  • Read the table backwards to find zz from a probability, using the critical-values table or the body of the table

  • Standardise with Z=X−μσZ = \dfrac{X - \mu}{\sigma} to find P(X<a)\mathrm{P}(X < a), P(X>a)\mathrm{P}(X > a) and P(a<X<b)\mathrm{P}(a < X < b), showing the working the mark scheme needs

  • Turn a probability into an expected number of items, split a population into categories, and use a normal probability for several independent items

  • Handle intervals centred on the mean ("within kk of the mean", "differs from the mean by less than") and boundaries given in standard deviations

  • Find an unknown boundary, including the edge of a category and the unknown end of an interval, from a given probability

  • Find an unknown mean or standard deviation from one probability statement, including one given as a sample frequency

  • Find both μ\mu and σ\sigma from two probability statements, or from one statement and a given relationship between μ\mu and σ\sigma

  • State and check the conditions np>5np > 5 and nq>5nq > 5, approximate B(n,p)\mathrm{B}(n, p) by N(np,npq)\mathrm{N}(np, npq), and apply the correct continuity correction for every wording

01

The normal distribution as a model for continuous data

Syllabus requirement · §5.5

“

understand the use of a normal distribution to model a continuous random variable, and use normal distribution tables

”

Discrete and continuous variables

A discrete random variable takes separate values you can list: the number of heads in 10 tosses can be 0,1,2,…,100, 1, 2, \dots, 10 and nothing in between. Each value has its own probability, and the probabilities add up to 11 (the Discrete Random Variables topic, 5.4).

A continuous random variable can take any value in a range. The mass of a bag of rice could be 1.481.48 kg, 1.4831.483 kg, 1.48311.4831 kg, … There are infinitely many possible values packed into any interval, so the probability of hitting one exact value is 00:

P(X=1.48)=0for a continuous X\mathrm{P}(X = 1.48) = 0 \quad \text{for a continuous } X

Probabilities only make sense for intervals, such as P(1.4<X<1.5)\mathrm{P}(1.4 < X < 1.5). A continuous distribution is described by a curve, and

P(a<X<b)=the area under the curve between a and b\mathrm{P}(a < X < b) = \text{the area under the curve between } a \text{ and } b

The total area under the curve is 11, because XX is certain to take some value.

One useful consequence: since a single value has probability 00, including or excluding an endpoint makes no difference. P(X<5)\mathrm{P}(X < 5) and P(X⩽5)\mathrm{P}(X \leqslant 5) are the same number for a continuous variable. (This is not true for a discrete variable, and §09 is all about that difference.)

The normal curve

Many measured quantities (heights, masses, lengths, times) cluster around a central value, with values further from the centre becoming steadily rarer, equally on both sides. The normal distribution is the model for this. Its curve is the symmetric bell shape below. We write

X∼N(μ,σ2)X \sim \mathrm{N}(\mu, \sigma^2)

read "XX is normally distributed with mean μ\mu and variance σ2\sigma^2". Here:

  • μ\mu (the Greek letter "mu") is the mean, the centre of the bell;
  • σ\sigma (the Greek letter "sigma") is the standard deviation, which measures how spread out the values are;
  • σ2\sigma^2 is the variance. The second number inside N( , )\mathrm{N}(\ ,\ ) is always the variance, not the standard deviation.

Properties you can read off the picture:

  • the curve is symmetric about μ\mu, so exactly half the area is on each side: P(X<μ)=P(X>μ)=0.5\mathrm{P}(X < \mu) = \mathrm{P}(X > \mu) = 0.5;
  • the mean, median and mode are all equal to μ\mu;
  • the curve never quite touches the axis, but almost all of the area lies within 3σ3\sigma of the mean (between μ−3σ\mu - 3\sigma and μ+3σ\mu + 3\sigma).
X ~ N(μ, σ²): bell-shaped and symmetric about μmean = median = mode = μμ−3σμ−2σμ−σμ+σμ+2σμ+3σμtotalarea = 1almost all of the area lies between μ − 3σ and μ + 3σ,so a sketch only needs to run about 3σ either side of the mean

The normal curve: symmetric about μ, with mean = median = mode, total area 1, and almost all of the area within 3σ of the mean.

Changing μ\mu slides the whole curve left or right without changing its shape. Changing σ\sigma changes the spread: a larger σ\sigma gives a wider, flatter curve, and a smaller σ\sigma a narrower, taller one. The curve has to get taller as it gets narrower because the total area must stay equal to 11.

μ slides the curve sideways; σ stretches itμ = 0μ = 2.2same σ, different μlarge σsmall σsame μ, different σeach curve still encloses an area of 1, so a narrower curve must be taller

Left: same σ, different μ (the curve slides). Right: same μ, different σ (the curve stretches; the narrow one is taller because each area is 1).

Reading the notation

For each random variable, state the mean and the standard deviation.

(a) X∼N(20,9)X \sim \mathrm{N}(20, 9) (b) Y∼N(5.2,1.52)Y \sim \mathrm{N}(5.2, 1.5^2) (c) W∼N(31.4,3.6)W \sim \mathrm{N}(31.4, 3.6)

(d) For the variable in (a), write down P(X=20)\mathrm{P}(X = 20) and P(X⩽20)\mathrm{P}(X \leqslant 20).

Show full working
  1. 1

    (a) Read the two numbers. In N(20,9)\mathrm{N}(20, 9) the first number is the mean and the second is the variance: μ=20,σ2=9\mu = 20, \qquad \sigma^2 = 9

    Always say out loud which number is which: mean first, variance second.

  2. 2

    Square-root the variance to get the standard deviation. σ=9=3\sigma = \sqrt{9} = 3

    Every later calculation divides by σ, so this is the number you will actually use. Dividing by 9 instead of 3 is a common way to lose a whole question.

  3. 3

    (b) The variance is written as a square. N(5.2,1.52)\mathrm{N}(5.2, 1.5^2) has variance 1.521.5^2, so μ=5.2,σ=1.5\mu = 5.2, \qquad \sigma = 1.5

    Writing the variance as 1.5² is Cambridge's way of handing you σ directly. Don't square 1.5 and then forget to undo it.

  4. 4

    (c) A variance that is not a perfect square. N(31.4,3.6)\mathrm{N}(31.4, 3.6) has σ2=3.6\sigma^2 = 3.6, so μ=31.4,σ=3.6=1.897…\mu = 31.4, \qquad \sigma = \sqrt{3.6} = 1.897\ldots

    Keep σ = √3.6 on the calculator rather than typing a rounded 1.9, so later answers stay accurate.

  5. 5

    (d) A single exact value. XX is continuous, so P(X=20)=0\mathrm{P}(X = 20) = 0

    There is no area above a single point on the axis.

  6. 6

    Including the endpoint changes nothing. P(X⩽20)=P(X<20)\mathrm{P}(X \leqslant 20) = \mathrm{P}(X < 20)

    The single value 20 has probability 0, so adding it to the region adds no area.

  7. 7

    Use symmetry about the mean. 2020 is the mean, so the area to its left is half the total: P(X⩽20)=0.5\mathrm{P}(X \leqslant 20) = 0.5

    Exactly half of the area lies on each side of the mean.

Answer

(a) μ=20\mu = 20, σ=3\sigma = 3. (b) μ=5.2\mu = 5.2, σ=1.5\sigma = 1.5. (c) μ=31.4\mu = 31.4, σ=3.6≈1.90\sigma = \sqrt{3.6} \approx 1.90. (d) P(X=20)=0\mathrm{P}(X = 20) = 0 and P(X⩽20)=0.5\mathrm{P}(X \leqslant 20) = 0.5.

Sketching normal curves

A sketch question gives you one or more distributions and asks for their curves on one diagram. The mark scheme looks for three things, so build each curve from them:

  1. the peak sits directly above the mean μ\mu;
  2. the curve runs down to the axis about 3σ3\sigma either side, so it spans roughly μ−3σ\mu - 3\sigma to μ+3σ\mu + 3\sigma;
  3. relative heights: a curve with a smaller σ\sigma is narrower and taller; two curves with the same σ\sigma have the same shape and height.

Sketching two curves on one diagram

A∼N(10,4)A \sim \mathrm{N}(10, 4) and B∼N(16,1)B \sim \mathrm{N}(16, 1). Describe how you would sketch both curves on one diagram with the horizontal axis from 00 to 2020.

Show full working
  1. 1

    Find each standard deviation. AA has variance 44, so σA=4=2\sigma_A = \sqrt{4} = 2. BB has variance 11, so σB=1\sigma_B = 1.

    The spread of a sketch depends on σ, not on the variance.

  2. 2

    Where does curve AA start and finish? μA−3σA=10−6=4,μA+3σA=10+6=16\mu_A - 3\sigma_A = 10 - 6 = 4, \qquad \mu_A + 3\sigma_A = 10 + 6 = 16 So AA is a bell centred at 1010, meeting the axis at about 44 and 1616.

    Three standard deviations either side covers almost all of the area, so this is where the curve should visibly reach the axis.

  3. 3

    Where does curve BB start and finish? μB−3σB=16−3=13,μB+3σB=16+3=19\mu_B - 3\sigma_B = 16 - 3 = 13, \qquad \mu_B + 3\sigma_B = 16 + 3 = 19 So BB is centred at 1616 and runs from about 1313 to 1919.

    Same rule, different numbers.

  4. 4

    Compare the heights. BB has half the standard deviation of AA, so BB is half as wide and must be about twice as tall to enclose the same area of 11.

    This is the mark examiners most often withhold: a narrower curve drawn at the same height as a wider one.

Answer

AA: a bell with its peak above 1010, reaching the axis near 44 and 1616. BB: a narrower bell, about twice as tall, with its peak above 1616, reaching the axis near 1313 and 1919.

A sketch question from a real paper

9709/61 O/N 2013 Q13 marks

It is given that X∼N(30,49)X \sim \mathrm{N}(30, 49), Y∼N(30,16)Y \sim \mathrm{N}(30, 16) and Z∼N(50,16)Z \sim \mathrm{N}(50, 16). On a single diagram, with the horizontal axis going from 0 to 70, sketch three curves to represent the distributions of XX, YY and ZZ.

Show full working
The mark scheme's own sketch: X and Y share a centre at 30, Y is narrower and taller, and Z is the same shape as Y moved to 50.

The mark scheme's own sketch: X and Y share a centre at 30, Y is narrower and taller, and Z is the same shape as Y moved to 50.

  1. 1

    Standard deviations. σX=49=7,σY=16=4,σZ=16=4\sigma_X = \sqrt{49} = 7, \qquad \sigma_Y = \sqrt{16} = 4, \qquad \sigma_Z = \sqrt{16} = 4

    The second number in each bracket is a variance, so take square roots first.

  2. 2

    Curve XX. Peak above 3030. It spans about 30−3(7)=930 - 3(7) = 9 to 30+3(7)=5130 + 3(7) = 51.

    The mark scheme accepts X running roughly from 10 to 50 (or 15 to 45).

  3. 3

    Curve YY. Same centre, 3030, but σY=4\sigma_Y = 4 is smaller than σX=7\sigma_X = 7, so YY spans only about 1818 to 4242 and is taller than XX.

    One mark is for 'same mean as X but higher and thinner'. Both words matter.

  4. 4

    Curve ZZ. Same variance as YY, so it is an identical shape (same width, same height), with its peak moved to 5050. It spans about 3838 to 6262.

    Z and Y differ only in μ, and changing μ only slides a curve.

Answer

Three bells: XX centred at 3030, wide and low; YY centred at 3030, narrower and taller than XX; ZZ exactly the shape of YY but centred at 5050.

In a sketch, label each curve and its mean on the axis. Relative width and height are what earn the marks, not exact heights.

When is a normal model suitable?

A normal distribution is a sensible model for data that is:

  • continuous (measurements such as mass or time, not counts);
  • symmetric, with a single peak in the middle;
  • tailing off evenly on both sides, with few values far from the centre.

It is not suitable for data that is clearly skewed (for example waiting times with a long right tail), has two peaks, or is a small whole-number count. If a question shows you a diagram of the data and asks whether a normal model fits, give two features: "symmetric" plus one more ("peaks in the middle" or "tails off on both sides").

Judging a model from a frequency table

The masses, mm grams, of 100 tomatoes are summarised below. Is a normal distribution a suitable model? Give reasons.

mm60–7070–8080–9090–100100–110110–120
frequency3163130173
Show full working
  1. 1

    Is the variable continuous? Mass is a measurement, so yes.

    A count such as 'number of seeds' would already rule out a normal model.

  2. 2

    Where is the peak? The biggest frequencies, 3131 and 3030, are in the two middle classes, so there is a single peak in the middle.

    A normal curve has one peak, at the mean.

  3. 3

    Is it symmetric? Reading outwards from the middle: 3131 and 3030, then 1616 and 1717, then 33 and 33. The two sides almost match.

    Pair up classes the same distance from the centre and compare them.

  4. 4

    Do the tails thin out? Yes: only 33 tomatoes in each end class.

    Few values far from the centre, on both sides.

  5. 5

    Conclude, with two reasons. A normal model is suitable: the masses are symmetrical about the middle and peak in the centre, tailing off at both ends.

    Symmetry plus one more feature is what earns both marks.

Answer

Yes: the data is continuous, roughly symmetrical, with a single central peak and thin tails on both sides.

Naming a model from a stem-and-leaf diagram

9709/62 F/M 2017 Q4(ii)2 marks

The weights in kilograms of packets of cereal were noted correct to 4 significant figures. The following stem-and-leaf diagram shows the data.

Key: 748∣5748 \mid 5 represents 0.7485 kg0.7485\text{ kg}.

Name a distribution that might be a suitable model for the weights of this type of cereal packet. Justify your answer.

The printed stem-and-leaf diagram: 59 weights, with the row frequencies in brackets.

The printed stem-and-leaf diagram: 59 weights, with the row frequencies in brackets.

Show full working
  1. 1

    Look at the shape. Reading the row frequencies down the page: 1,6,12,15,13,11,11, 6, 12, 15, 13, 11, 1. They rise to a single peak in the middle (the 750750 row) and fall away on both sides.

    A stem-and-leaf diagram turned on its side is a bar chart, so the shape of the distribution can be read directly.

  2. 2

    Check symmetry and tails. The two extremes are both 11, and the frequencies fall away on both sides of the peak, so the shape is roughly symmetrical.

    Symmetry plus a peak in the middle is exactly the picture of a normal curve.

  3. 3

    Name the model and give two reasons. A normal distribution: the data is (roughly) symmetrical and peaks in the middle, tailing off at both ends.

    The mark scheme requires 'symmetrical' plus another reason for the second mark.

Answer

A normal distribution, because the weights are roughly symmetrical and peak in the middle (tailing off quickly at both ends).

Common mistakes
  • Reading N(20,9)\mathrm{N}(20, 9) as "mean 20, standard deviation 9"

    The second number is the variance: σ=9=3\sigma = \sqrt{9} = 3

    Every mark scheme in this topic says 'not σ², not √σ' next to the standardising mark.

  • Drawing a narrow and a wide normal curve at the same height

    The narrower curve (smaller σ\sigma) must be taller, since both areas are 11

    Sketch questions award a separate mark for 'higher and thinner'.

  • Justifying a normal model with "symmetrical" alone

    Give two features: symmetrical and peaks in the middle (or tails off at both ends)

    The mark scheme needs symmetry plus another reason.

Your turn

  1. 1

    State the mean and standard deviation of (a) N(12,0.25)\mathrm{N}(12, 0.25), (b) N(−3,22)\mathrm{N}(-3, 2^2), (c) N(100,50)\mathrm{N}(100, 50).

    Stuck? Show hint

    Mean first, variance second. Square-root the variance.

    Show solution
    1. 1

      (a) μ=12\mu = 12 and σ2=0.25\sigma^2 = 0.25, so σ=0.25=0.5\sigma = \sqrt{0.25} = 0.5

      The square root of a number less than 1 is bigger than the number.

    2. 2

      (b) μ=−3\mu = -3 and the variance is written as 222^2, so σ=2\sigma = 2

      A negative mean is fine. Only σ has to be positive.

    3. 3

      (c) μ=100\mu = 100 and σ2=50\sigma^2 = 50, so σ=50=7.07 (3 s.f.)\sigma = \sqrt{50} = 7.07 \ (\text{3 s.f.})

      Not a perfect square, so keep the exact value on your calculator.

    Answer

    (a) 1212 and 0.50.5; (b) −3-3 and 22; (c) 100100 and 50≈7.07\sqrt{50} \approx 7.07.

  2. 2

    For each variable, say whether a normal distribution is likely to be a suitable model, with a reason. (a) The heights of adult women in a large city. (b) The number of heads when a coin is tossed 3 times. (c) The time customers wait in a queue, where most wait under 2 minutes but a few wait over 20 minutes.

    Stuck? Show hint

    Is it continuous? Is it symmetric with one central peak?

    Show solution
    1. 1

      (a) Suitable: height is continuous, and heights cluster symmetrically around an average with few very short or very tall people.

      This is the textbook normal situation.

    2. 2

      (b) Not suitable: the number of heads is a discrete count with only four possible values (00, 11, 22, 33). It is binomial, B(3,0.5)\mathrm{B}(3, 0.5).

      A normal model needs a continuous variable (or, as in §09, a binomial with a large n).

    3. 3

      (c) Not suitable: the waiting times are heavily skewed, bunched near 00 with a long tail to the right, so they are not symmetric.

      A normal curve has equal tails on both sides.

    Answer

    (a) Yes: continuous and symmetric about a central value. (b) No: a small discrete count. (c) No: strongly skewed.

Practise the normal distribution as a modelReal past-paper questions · Normal distribution as a model for continuous random variables

The rest of this note

Checking your access…

Can you do all of these?

  • Sketch the curve and shade the region before any algebra; decide whether the answer is above or below 0.5

  • N(μ, σ²): the second number is the variance; divide by σ, never by σ² or √σ

  • Write the standardising line with the numbers substituted: the first method mark is for it

  • Φ is the area to the left: a right tail is 1 − Φ, and Φ(−z) = 1 − Φ(z)

  • Use the ADD columns for a third decimal place of z, and quote table values to 4 d.p.

  • For 0.75, 0.9, 0.95, 0.975, 0.99 use the critical values 0.674, 1.282, 1.645, 1.960, 2.326 exactly

  • Expected number = n × p with p to at least 4 s.f.; give a single whole number, no '≈'

  • 'Within k of the mean' is 2Φ(k) − 1; 'more than k from the mean' is both tails; 'above' is one tail

  • A distance given in standard deviations is already the z-value

  • Working backwards, the sign of z comes from where the boundary is: above μ positive, below μ negative

  • Equate the standardised expression to a z-value, never to a probability

  • Signs must be consistent: x₁ − μ and z have the same sign, and σ must come out positive

  • For an interval with one unknown end, add the known tail to the slice to get the area to the left

  • Two unknowns: two z-values, each with its own sign; clear fractions, subtract to remove μ, solve for σ, substitute back

  • A given relationship between μ and σ is the second equation: substitute it at once; with a boundary of 0 it cancels

  • Binomial approximation: evaluate np and nq and compare both with 5

  • Mean np, variance npq, standard deviation √(npq)

  • Continuity correction: list the whole numbers wanted and cut half-way to the first one not wanted

Now do the questions
112 real Paper 5 parts from 2021–2025, sorted by difficulty, with mark schemes