The Poisson distribution as a model for random events
“
understand the relevance of the Poisson distribution to the distribution of random events, and use the Poisson distribution as a model.
Think about the number of emails arriving in an inbox in an hour, or the number of misprints on a printed page. There's no fixed number of "trials": an email can arrive at any moment, and most moments nothing arrives at all. That's a different situation from the binomial, so it needs a different model.
A Poisson distribution is the right model when three conditions hold:
- events occur singly: two events never happen at precisely the same moment;
- events occur independently and at random: one event doesn't make the next one more or less likely, or change when it happens;
- events occur at a constant average rate: the mean number of events is proportional to the length of the interval (twice the time, twice the expected count).
Here the interval is whatever stretch you're counting over (an hour, a metre of cloth, a page), and the rate is the average number of events per unit of it (3 calls per hour, 1.5 flaws per square metre).
Ask yourself: is there a fixed number of trials? A binomial counts successes out of attempts that you could list and tick off one by one: patients, free throws. A Poisson counts events that can happen at any moment in a stretch of time or space, with no list of attempts, just an average rate. If you can say what is, it's binomial. If you can only say "about so many per hour", it's Poisson. (A binomial with a huge and a tiny can still be treated as Poisson; that's §04.)
Stating the conditions in an exam
When a question asks you to justify a Poisson model, state the conditions in the context of the question. "Computers are donated singly, independently and at a constant average rate" scores. "The events are independent" on its own usually doesn't. There's more on this, with past-paper wording, later in this section.
When all three conditions hold, the random variable counting the number of events in a fixed interval is written
read " has a Poisson distribution with parameter ", where (lambda) is the mean number of events in that interval. Unlike the binomial's , there is no upper limit on how large can be: in principle any non-negative integer is possible, however unlikely the larger values become.
The probability formula itself is:
The Poisson probability formula
λ is the mean number of events in the interval; r is the count you want. The next two blocks explain the symbols and where the formula comes from.
What each symbol means
- (the Greek letter lambda) is the mean number of events in the interval you're asking about. If calls arrive at 3 per hour and the question is about one hour, . The interval can be a length of time, a length of cloth, an area of lawn or a volume of liquid; whatever it is, must match it.
- is the particular count you want the probability of:
- is a fixed number, , the same you met with in Pure Maths. You'll never work out by hand; use the key on your calculator.
- (" factorial") means , so . By definition , and as for any power, so .
The formula booklet has no Poisson tables, so every Poisson probability comes from this formula and your calculator.
Where the formula comes from
You don't need to reproduce this in the exam, but seeing it once makes the rest of the topic make sense.
Take calls arriving at an average of per hour. Chop the hour into very short slots, say one-second slots. Each slot is so short that it holds either one call or none, and the chance of a call in any one slot is . Now the number of calls is binomial: .
Make the slots shorter and shorter, so grows and shrinks while stays fixed. Two things happen to the binomial formula:
The first fraction tends to because each factor on top is almost . The second limit is the standard one that defines . Put them together and you get .
So a Poisson distribution is what a binomial turns into when there are a huge number of chances for the event and each chance is tiny. Two results follow straight away. The mean is . The variance is also very nearly , because is almost . That is why the mean and variance are equal (§03), and why a binomial with large and small can be swapped for a Poisson (§04).
Po(λ) for λ = 1, 4 and 10. Small λ gives a right-skewed shape; as λ grows, the distribution moves right, spreads out and becomes more symmetric.
Substituting into the Poisson formula, piece by piece
Telephone calls arrive at a small business independently, singly and at a constant average rate of per hour. Find the probability that exactly calls arrive in a randomly chosen hour.
Show full working
- 1
Step 1 — check the conditions and state the model. Calls arrive independently, singly, and at a constant average rate, so a Poisson model applies. The mean number of calls in one hour is , so , where is the number of calls in one hour.
Writing X ~ Po(3) and saying what X counts usually earns the first mark, and it keeps λ and r from getting mixed up later.
- 2
Step 2 — identify each piece the formula needs, separately. Here and .
Naming λ and r out loud before touching the formula stops the classic swap: putting the count 2 where the mean 3 belongs.
- 3
Step 3 — find .
grows fast, so this is where an exponent slip does the most damage. Work it out on its own.
- 4
Step 4 — find .
The factorial sits in the denominator. Forget it and the answer comes out too big.
- 5
Step 5 — find , using a calculator.
Keep at least 5 s.f. of . Rounding it to 0.05 here already moves the third significant figure of the answer.
- 6
Step 6 — assemble the formula, substituting all three pieces from Steps 3–5.
The formula written with numbers in place, before any evaluating, is the line that earns the method mark.
- 7
Step 7 — evaluate.
Only now round to 3 s.f., at the very end.
(3 s.f.).
Work out , and one at a time, then multiply. Doing it all in one go on the calculator is how a wrong power or a missing factorial gets through.
Treating as automatically zero, since "nothing happens" sounds like it shouldn't count
, since and . It's a real probability, and often quite a large one
You'll need P(X = 0) all the time, especially for 'at least one' questions. It's the easiest Poisson probability of all, as long as you remember 0! = 1.
Assuming a described situation is binomial because it involves "counting how many times something happens"
Look for a fixed number of trials first. If events can happen at any moment, with no upper limit, it's Poisson
Binomial needs a fixed number of trials. Poisson has no such limit.
Your turn
- 1
Flaws occur in a certain fabric independently, singly and at a constant average rate of per square metre. Find the probability that a randomly chosen square metre of the fabric contains exactly flaw.
Stuck? Show hint
X ~ Po(1.5). Find , and separately before multiplying.
Show solution
- 1
Model. = number of flaws in , so with and .
The rate and the area already match (per square metre), so no scaling is needed.
- 2
Pieces. , ,
Three separate numbers first, the same routine as the worked example.
- 3
Substitute.
Writing the formula with numbers in place is the method line examiners look for.
- 4
Evaluate.
Round only now: 0.335.
Answer(3 s.f.).
- 1
- 2
A radioactive source emits particles independently, singly and at a constant average rate of per minute. Find the probability that no particles are emitted in a randomly chosen minute.
Stuck? Show hint
directly, since and .
Show solution
- 1
Model. = number of particles in one minute, so , and here .
The rate is per minute and the question asks about one minute, so λ = 4 as given.
- 2
Pieces. , .
Both equal 1, which is why is always just .
- 3
Substitute.
Showing the formula with numbers in place, even when it simplifies to , earns the method mark.
- 4
Evaluate.
Three significant figures of a small number means 0.0183, not 0.018.
Answer(3 s.f.).
- 1
Stating the conditions, and spotting when the model breaks
Short questions worth 1 or 2 marks ask you either to state a condition for a Poisson model or to explain why a Poisson model does not fit. Both are easy marks once you know what the examiners look for.
Stating conditions. Mark schemes accept any of these four (the three conditions from the start of this section, with "at random" and "independently" counted as separate points), as long as each is written about the thing being counted:
| Marking point | Written in context (orders arriving at a shop) |
|---|---|
| random | orders arrive at random |
| independent | orders arrive independently of each other |
| singly | orders arrive one at a time (singly) |
| constant mean rate | orders arrive at a constant mean rate |
Two details cost students marks every year. First, the word constant on its own is not enough: it has to be a constant mean or a constant rate. Second, a condition written without context ("events are independent", "it is random") scores nothing, or only a special-case mark when two are asked for.
Spotting when the model breaks. A Poisson model is ruled out by any of these:
- the rate changes over the interval: busier at lunchtime, busier in the day than at night, or a count that has been rising year on year;
- events are not independent or do not happen singly: customers who arrive in groups, or one accident that tends to cause another;
- the variable can take impossible values: a Poisson variable is a count, so it can only be . A variable that can be negative, or can be a non-integer such as , is not Poisson;
- the mean and variance are not equal, for example has mean but variance (see §03).
Judging a model for customers at a café
A café owner models the number of customers who walk in during a randomly chosen -minute period by a Poisson distribution. She notices that many customers arrive in groups of friends, and that the café is much busier between and pm than in the middle of the afternoon.
(a) State two conditions needed for the Poisson model to be valid, in context.
(b) Use her observations to comment on whether the model is suitable.
Show full working
- 1
Step 1 — pick the first condition and write it about customers. "Customers arrive independently of each other."
Mentioning customers puts the condition in context, and that's where the mark comes from.
- 2
Step 2 — pick a second condition, again in context. "Customers arrive at a constant mean rate."
Write 'mean rate' in full. 'Customers arrive at a constant' is the half-answer that loses the mark.
- 3
Step 3 — test the first condition against her observations. Groups of friends walk in together, so one customer arriving makes others arrive at the same moment. Customers are not arriving independently or singly.
Take each condition in turn and ask whether the evidence in the question breaks it.
- 4
Step 4 — test the second condition. The café is much busier at lunchtime, so the mean number of arrivals per minutes is not the same all afternoon. The mean rate is not constant.
A changing rate is the reason examiners use more than any other to reject a Poisson model.
- 5
Step 5 — conclude. Both conditions fail, so a single Poisson model for every -minute period is not suitable. (Counting groups instead of people, within lunchtime only, might be closer to Poisson.)
A clear conclusion that refers back to the evidence finishes the answer.
(a) e.g. customers arrive independently; customers arrive at a constant mean rate. (b) Groups arrive together, so arrivals are not independent or single; lunchtime is busier, so the rate is not constant. The model is not suitable.
For any 'comment on the model' question, go through the conditions one at a time and point to the sentence in the question that breaks each one.
Two assumptions, stated in context
The number of orders arriving at a shop during an -hour working day is modelled by the random variable with distribution . State two assumptions that are required for the Poisson model to be valid in this context.
Show full working
- 1
Step 1 — identify what is being counted. The events are orders arriving at the shop, so every assumption must mention orders.
The mark scheme says 'must be in context'. Two correct assumptions without context score only a special-case B1.
- 2
Step 2 — write the first assumption. "Orders arrive at random."
One B1 for any one correct assumption.
- 3
Step 3 — write the second assumption. "Orders arrive at a constant mean rate."
The mark scheme adds 'must say mean or rate'. 'Independently' or 'singly' would also score.
Any two of: orders arrive at random; orders arrive independently; orders arrive singly; orders arrive at a constant mean rate.
Start every condition with the thing being counted ('Orders arrive…'). It guarantees context without any extra effort.
When the rate changes between day and night
A website owner finds that, on average, his website receives hits per minute. He believes that the number of hits per minute follows a Poisson distribution.
A friend agrees that the website receives, on average, hits per minute. However, she notices that the number of hits during the day-time ( am to pm) is usually about twice the number of hits during the night-time ( pm to am).
(i) Explain why this fact contradicts the owner's belief that the number of hits per minute follows a Poisson distribution.
(ii) Specify separate Poisson distributions that might be suitable models for the number of hits during the day-time and during the night-time.
Show full working
- 1
Step 1 — (i) find the condition the friend's fact breaks. Day-time minutes get about twice as many hits as night-time minutes, so the mean number of hits per minute changes during the day.
Ask which of the four conditions the new information is about. 'Twice as many in the day' is about the rate.
- 2
Step 2 — (i) write the answer. The mean number of hits per minute is not constant, so a single Poisson model does not fit.
The mark scheme accepts 'mean not constant' or 'not a constant rate'.
- 3
Step 3 — (ii) give the night rate a letter. Let the night-time rate be hits per minute. Then the day-time rate is .
One unknown is enough, because the day rate is given as twice the night rate.
- 4
Step 4 — (ii) use the overall average. Day and night are each hours, so the overall rate is the plain average of the two:
Equal lengths of time means an ordinary average works. This equation (written as 2p + p = 2 × 0.3) is the mark scheme's M1.
- 5
Step 5 — (ii) solve.
Multiply both sides by 2, then divide by 3.
- 6
Step 6 — (ii) state both models. Day-time rate per minute and night-time rate per minute, so the number of hits per minute is in the day and at night.
Give each model with its time unit. Po(24) and Po(12) per hour are also accepted.
(i) The mean number of hits per minute is not constant. (ii) Day-time: per minute; night-time: per minute.
When a question says the rate is different at different times, the fix is not to abandon Poisson but to use a separate Poisson model for each time with its own constant rate.
Your turn: conditions and critique
- 19709/62 O/N 2025 Q1(a)1 mark
The number, , of used computers donated to a charity has a constant average rate of computers per week. State a necessary condition for to have a Poisson distribution.
Stuck? Show hint
The constant rate is already given. Choose one of the other conditions and write it about computers or donations.
Show solution
- 1
Identify what is counted. Computers being donated.
The mark scheme requires 'computers' or 'donations' to appear.
- 2
State a condition not already given. "Computers are donated independently of each other." ("singly" or "at random" also score.)
'Events are independent' or 'It is independent' scores B0, because there is no context.
Answere.g. Computers are donated independently (or singly, or at random).
- 1
- 29709/63 M/J 2021 Q3(b)1 mark
The local council claims that the average number of accidents per year on a particular road is . Jane claims that the true average is greater than . She looks at the records for a random sample of recent years and finds that the total number of accidents during those years was . Assume that the number of accidents per year follows a Poisson distribution, and a test is carried out (you do not need to carry it out here).
Jane finds that the number of accidents per year has been gradually increasing over recent years. State how this might affect the validity of the test.
Stuck? Show hint
Which Poisson condition does a gradually increasing number of accidents break?
Show solution
- 1
Link the fact to a condition. If the number of accidents per year is increasing, the mean number per year is not constant.
A trend over time always points to the constant-rate condition.
- 2
State the effect. The Poisson model is not valid, so the test based on it may not be valid.
The mark scheme wants both parts: mean not constant, so the Poisson model is not valid.
AnswerThe mean is not constant, so a Poisson model is not valid and the test may not be valid.
- 1
- 39709/71 M/J 2014 Q8(i)2 marks
The following tables show the probability distributions for the random variables and .
For each of the variables and state how you can tell from its probability distribution that it does NOT have a Poisson distribution.
Stuck? Show hint
Ignore the probabilities. Look at the values each variable can take, and compare with 0, 1, 2, …
Show solution
- 1
Recall the possible values of a Poisson variable. It is a count, so it can only take
The probabilities here are the same as those of Po(1), so they are a distraction. The values are what give it away.
- 2
Check . can take the value . A count cannot be negative, so is not Poisson.
One B1 for 'cannot have a negative value'.
- 3
Check . can take the value . A count must be a whole number, so is not Poisson.
One B1 for 'cannot have a non-integer value'.
Answertakes a negative value (); takes a non-integer value (). A Poisson variable takes only the values
- 1
Calculating probabilities: ranges, equations and the most likely value
“
use formulae to calculate probabilities for the distribution Po(λ).
As with the binomial, most questions ask about a range of values: "at least ", "fewer than ", "between and ". You turn the words into -values in the same way as in §5.4. The big difference is that a Poisson variable has no upper limit. "At most " is a short list () that you can add up. "At least " would mean adding forever. For anything with no upper limit, work out the complement and take it away from .
Phrase | Means | How to find it |
|---|---|---|
at most / no more than / fewer than | add directly | |
fewer than | add directly | |
between and inclusive | add the finite list | |
at least / not fewer than | use , as there's no top end | |
more than | use , as there's no top end |
Anything with a top end (at most, fewer than, between) can be added up directly. Anything without one (at least, more than) needs the complement.
A quicker way to get the next term
Each Poisson probability is the one before it multiplied by :
So for : , then , then . It saves retyping the whole formula each time. Keep 5 or 6 significant figures on each term and round only at the end.
- 1
Check the conditions and identify , the mean number of events in the interval the question asks about.
- 2
Turn the wording into a list of -values, reading strict and non-strict inequalities carefully.
- 3
Decide: is there a top end? A finite list (at most , fewer than , between and ) is added up directly. Anything with no top end (at least, more than) needs or .
- 4
Work out each term from the formula, keeping , and as separate pieces until the final multiplication.
- 5
Add or subtract as needed, keeping several significant figures until the last line.
'At least': use the complement
Telephone calls arrive at a small business independently, singly and at a constant average rate of per hour. Find the probability that at least calls arrive in a randomly chosen hour.
Show full working
The bars for 'at least 2' go on forever, so there's no last one to stop a direct sum at.
- 1
Step 1 — state the model. As before, , the number of calls in one hour.
This is the same situation as the §01 example, so λ = 3 carries over. Say so rather than assume it.
- 2
Step 2 — translate "at least 2". This means , which has no top end. You can't add it up directly, so use the complement:
'At least 2' includes 2, so the complement stops at 1: it is 1 − P(X ≤ 1), not 1 − P(X ≤ 2).
- 3
Step 3 — find .
and , so this term is just .
- 4
Step 4 — find .
Keep the unrounded and multiply it by 3. Don't restart from a rounded 0.0498.
- 5
Step 5 — add the two excluded terms.
Add the excluded terms before subtracting. Writing 1 − P(0) + P(1) instead of 1 − [P(0) + P(1)] is a common sign slip.
- 6
Step 6 — subtract from 1.
Round only at the end: 0.801.
(3 s.f.).
'At least' or 'more than' means there's no top end, so go straight to 1 minus the complement.
A range between two values
The random variable has the distribution . Find .
Show full working
P(2 < X < 5) on Po(3): strict inequalities drop both end values, so only the bars for x = 3 and x = 4 are summed.
- 1
Step 1 — translate the phrase. "" is a strict inequality on both sides, so it means or : a finite, bounded list, short enough to sum directly.
Strict '<' on both sides excludes 2 and 5. The mark scheme allows one end error for the method mark, but not for the answer.
- 2
Step 2 — find , using and .
and are worked out as separate pieces before multiplies in.
- 3
Step 3 — find , using .
Reuse the same . Only the power and the factorial change from one term to the next.
- 4
Step 4 — add the two terms.
X = 3 and X = 4 cannot both happen, so their probabilities add.
(3 s.f.).
Read the inequality signs carefully. '2 < X < 5' leaves out 2 and 5, so only 3 and 4 are included.
Solving equations in λ
Some questions give you a relationship between Poisson probabilities and ask you to find (or the value of ). The method never changes. Write every probability out in full from the formula. Cancel , which appears in every term. Then divide by the smallest power of you can see.
What's left depends on how far apart the -values are. If they're one apart, such as and , you get a linear equation. If they're further apart, a is left over; that case comes after the next two examples.
Equating two consecutive probabilities to find λ
The random variable has the distribution . It is given that . Find the value of .
Show full working
- 1
Step 1 — write both probabilities out in full from the formula.
Write both sides out in full before cancelling, so you can see that each factor you cancel really is on both sides.
- 2
Step 2 — set them equal and cancel the common factor .
is never zero, so dividing both sides by it is always safe.
- 3
Step 3 — divide both sides by , allowed because a Poisson mean is strictly positive ().
Dividing by λ is safe only because λ ≠ 0, so say so. Otherwise you lose the root λ = 0 without explaining why it doesn't count.
- 4
Step 4 — multiply both sides by 2.
Quick check: and . They are equal.
.
When the two r-values are one apart, cancelling and the smaller power of λ leaves a linear equation.
Finding r when two probabilities are equal
The random variable has the distribution . It is given that . Write down an equation in , and hence find the value of .
Show full working
- 1
Step 1 — write both probabilities using the formula, with .
λ is known (15) but r is the unknown n, so the factorials stay symbolic.
- 2
Step 2 — set the two expressions equal, since is given.
Write the equation out in full first. Examiners want to see it, and it makes the cancelling easy to check.
- 3
Step 3 — cancel the common factor from both sides.
is a non-zero common factor of both sides.
- 4
Step 4 — divide both sides by , using .
Splitting into lets the powers cancel. The mark scheme's M1 needs the powers reduced to a single 15.
- 5
Step 5 — use to cancel the remaining factorial.
(n+1)! = (n+1) × n! is the factorial identity. The same M1 needs the factorials reduced to n + 1.
- 6
Step 6 — solve the resulting equation for .
Quick check: Po(15) has its two equal tallest bars at 14 and 15, which fits n = 14.
.
Same pattern as before: cancels straight away, then (n+1)! = (n+1) × n! lets the factorials cancel.
When you get a quadratic
Both examples above ended in a linear equation because the two -values were one apart. If they're two apart, such as and , you're left with equal to a number, which you solve by square-rooting. If the equation has three terms, such as , and , both a term and a term survive, and you factorise or use the quadratic formula. The method doesn't change: write every term out in full, cancel , then divide by the lowest power of . A Poisson mean is always positive, so throw away the negative root.
Three probabilities that cancel down to a quadratic
The random variable has the distribution . It is given that . Find the value of .
Show full working
- 1
Step 1 — write every probability in the equation out in full from the formula.
With three terms, write each one out before substituting. It's easy to drop a term if you try to do it in your head.
- 2
Step 2 — substitute these into the given equation.
Keep the coefficient 3 attached to its own term only.
- 3
Step 3 — cancel the common factor from all three terms.
multiplies every term, so it cancels from all three at once.
- 4
Step 4 — clear the fractions by multiplying every term by 6.
6 is the lowest common multiple of the denominators 1, 2 and 6.
- 5
Step 5 — divide every term by , which is allowed since .
After dividing by λ, a λ² term is still there, so this time the equation is quadratic.
- 6
Step 6 — rearrange into the standard quadratic form .
Everything on one side, equal to zero, before factorising.
- 7
Step 7 — factorise (or use the quadratic formula).
Check the factorisation: 6 × (−3) = −18 and 6 + (−3) = 3, matching the −18 and the −3λ.
- 8
Step 8 — reject the negative root. A Poisson mean can never be negative, so is rejected, leaving .
λ is a mean number of events, so it can't be negative. Only λ = 6 makes sense.
.
Three terms in the equation is the signal that a full quadratic (not a linear equation) is coming. Expect to factorise or use the quadratic formula, and always discard the negative root.
A quadratic in λ from three probabilities
The random variable has the distribution , where . It is given that . Find the value of .
Show full working
- 1
Step 1 — write every probability out in full from the formula.
The mark scheme's B1 is for this full equation, with or without the factors.
- 2
Step 2 — substitute into the given equation.
Keep the 5/2 attached only to the P(Y=3) term.
- 3
Step 3 — cancel the common factor from every term.
Replace 3!, 4! and 5! by 6, 24 and 120 now, so the next step's common multiple is visible.
- 4
Step 4 — clear every fraction by multiplying all three terms by 120.
120 = 5! is the lowest common multiple of 6, 24 and 120.
- 5
Step 5 — divide every term by , allowed since .
is the lowest power present, and λ > 0 is given, so dividing loses nothing.
- 6
Step 6 — rearrange into standard quadratic form.
Everything on one side, equal to zero. This is the line the mark scheme's M1 is looking for.
- 7
Step 7 — factorise.
Check: 10 × (−5) = −50 and 10 + (−5) = 5. That confirms the factorisation before you trust it.
- 8
Step 8 — reject the negative root, since is given (and a Poisson mean can never be negative anyway).
The mark scheme awards the final A1 for λ = 10 alone, so don't leave −5 standing as a second answer.
.
Same shape as the last example. The r-values run from 3 to 5, so dividing by λ³ leaves a quadratic with one positive and one negative root.
Attempting to sum "at least " or "more than " directly, term by term
These have no top end, so rewrite them as or first
A Poisson variable has no largest value, so a direct sum would never end.
Cancelling or a power of from only one side of an equation
Only cancel something that multiplies both sides. Write both sides out in full first
If you skip straight to a simplified equation, it's easy to cancel something that isn't actually on both sides.
Assuming every equated-probability question ends in a linear equation, because the first few practised did
Check what's left after dividing by the lowest power of present. Two terms two apart leave a pure , solved by square-rooting; three terms leave a full quadratic, needing factorising or the formula
Three terms (or a gap of two between terms) is the signal a quadratic is coming. Treating it as linear from habit leads to a lost or mishandled root.
Your turn
- 1
. Find .
Stuck? Show hint
No top end, so use 1 − P(X ⩽ 2), adding P(X=0), P(X=1) and P(X=2).
Show solution
- 1
has no top end, so
'At least 3' keeps 3, so the complement stops at 2.
- 2
, so this is just .
- 3
.
- 4
.
- 5
Three excluded terms. Count them: 0, 1 and 2.
- 6
Round only at the end: 0.762.
Answer(3 s.f.).
- 1
- 2
The random variable has the distribution . Given that , find the value of .
Stuck? Show hint
Write both sides out from the formula, cancel , then divide both sides by λ.
Show solution
- 1
Writing both probabilities out in full: and .
Full expressions first, so you can see each cancellation is allowed.
- 2
Set them equal as given, and cancel the common factor :
is non-zero, so it cancels from both sides.
- 3
Divide both sides by (allowed since ):
λ = 0 is not a valid Poisson mean, so no solution is lost.
- 4
Multiply both sides by 2:
Check: and .
Answer.
- 1
- 39709/72 M/J 2014 Q4(iii)3 marks
The random variable has the distribution , where . Given that , find .
Stuck? Show hint
Write both probabilities out in full, cancel and one power of μ, then solve the resulting equation in μ².
Show solution
- 1
Writing both out in full: and .
Write both sides in full first. The mark scheme's B1 is for this equation.
- 2
Substituting into the given equation and cancelling :
is non-zero, so it cancels from both sides.
- 3
Dividing both sides by (allowed since ):
μ ≠ 0 is given precisely so that you can divide by μ.
- 4
Multiplying both sides by 6:
Clear the fraction before square-rooting.
- 5
Square-rooting, and rejecting the negative root:
√144 gives ±12, but only the positive root can be a mean.
Answer.
- 1
Finding λ from one given probability
Sometimes a single probability is given as a number and is unknown. The easiest case, and the one examiners use, is , because it is just . Setting equal to the number and taking natural logs gives in one line:
Since , is negative, so comes out positive, as a mean must. A question phrased as "the probability of at least one event is " comes to the same thing: , so .
Finding λ when P(X=0) is known
The number of misprints on a page of a newspaper has the distribution . The probability that a randomly chosen page has no misprints is . Find .
Show full working
- 1
Step 1 — write from the formula.
, so is just . This is why P(X=0) questions always lead to a logarithm.
- 2
Step 2 — set it equal to the given value.
One equation with one unknown.
- 3
Step 3 — take natural logs of both sides.
ln undoes e, which brings λ down out of the power.
- 4
Step 4 — multiply both sides by .
A positive answer is a quick check: a Poisson mean can never be negative.
(3 s.f.).
is the only Poisson probability you can undo with a single logarithm. If you see P(X=0) given as a number, reach for ln straight away.
Finding λ to 3 significant figures
The random variable has the distribution . Given that , find the value of correct to significant figures.
Show full working
- 1
Step 1 — write in terms of .
The same first line as the previous example.
- 2
Step 2 — form the equation.
This equation earns the first B1.
- 3
Step 3 — take natural logs.
ln of a number below 1 is negative, and that makes λ positive in the next step.
- 4
Step 4 — solve and round.
The question asks for 3 significant figures, and that is the second B1.
(3 s.f.).
Keep the full calculator value of ln until the very last line, then round to the accuracy the question asks for.
Your turn: λ from a probability
- 19709/62 F/M 2022 Q7(b)4 marks
The random variables and have the distributions and respectively. It is given that
- ,
- , where is a non-zero constant.
Find the value of .
Stuck? Show hint
The first fact is about P(·=0), so it gives , i.e. λ = 2μ. Put that into the second fact.
Show solution
- 1
Write the first fact using the formula. and , so
Squaring doubles the power: .
- 2
Compare the powers.
If then , because is one-to-one.
- 3
Write the second fact using the formula.
Write each probability in full before substituting anything.
- 4
Substitute on the left.
(2μ)² = 4μ², then divide by 2! = 2.
- 5
Set the two sides equal and cancel. Dividing both sides by (which is not zero) gives .
is never zero, and μ ≠ 0 because it is a Poisson mean, so the division is allowed.
Answer.
Where the probabilities go up and down: the most likely value
A Poisson bar chart rises, reaches a peak, then falls away (look back at the diagram in §01). Some questions ask which values of are on the rising side, or which single value is most likely. Both come from comparing neighbouring probabilities.
Divide by , writing each out from the formula:
The cancels, , and , so the factorials leave . (This is the multiplier from the "quicker way to get the next term" tip at the start of this section, with moved up by one.)
So only when this ratio is bigger than , that is when
While the bars are still climbing. The last value reached by climbing is the most likely value (the mode). If is not a whole number, the mode is the whole number just below . (If is a whole number, and tie for most likely.)
Which values are on the rising side, and which is most likely?
The random variable has the distribution .
(a) Find the set of values of for which .
(b) Hence find the most likely value of .
Show full working
- 1
Step 1 — write the inequality out in full.
Both sides in full first. The mark scheme gives the M1 for seeing both expressions.
- 2
Step 2 — divide both sides by .
is positive, so dividing by it does not flip the inequality.
- 3
Step 3 — divide both sides by .
. Again a positive divisor, so the sign stays.
- 4
Step 4 — multiply both sides by , using .
(r+1)!/r! = r + 1. This is the ratio result above, reached step by step.
- 5
Step 5 — solve for . Since is a whole number , the set is .
r counts events, so only 0, 1, 2, … are allowed. List every one that fits.
- 6
Step 6 — (b) read off the most likely value. is the last 'rising' comparison, and for the inequality fails, so the bars fall after . The most likely value is .
The rising stops at r + 1 = 4. Check with numbers: P(X=3) = 0.163, P(X=4) = 0.188, P(X=5) = 0.173.
(a) . (b) The most likely value is .
Every 'P(X=r) < P(X=r+1)' question collapses to r + 1 < λ. The most likely value is the largest r + 1 that satisfies it.
The rising side and the mode of Po(2.4)
A random variable has the distribution . It is given that .
(i) Find the set of possible values of .
(ii) Hence find the value of for which is greatest.
Show full working
- 1
Step 1 — (i) write both probabilities in full.
This line is the M1.
- 2
Step 2 — cancel and .
Both are positive, so the inequality keeps its direction.
- 3
Step 3 — use .
This is the A1 line (or r < 1.4).
- 4
Step 4 — (i) list the values. , so or .
Only whole numbers from 0 upwards are possible values of r.
- 5
Step 5 — (ii) find the peak. The probabilities rise from to and from to , then stop rising. So is greatest at .
The last rising step is from r = 1 to r = 2, so the peak is at 2, not 1. Check: P(1) = 0.218, P(2) = 0.261, P(3) = 0.209.
(i) . (ii) .
The answer to 'greatest' is one more than the largest r on the rising list, because each r in the list means the next value is bigger.
Your turn: the most likely value
- 19709/73 O/N 2019 Q5(iii)3 marks
The random variable has the distribution and it is given that .
(a) Write down an inequality in .
(b) Hence or otherwise find the largest possible value of .
Stuck? Show hint
Write both probabilities in full, cancel and , and use (n+1)! = (n+1) × n!.
Show solution
- 1
(a) Write the inequality in full.
The mark scheme accepts this, or the version with already cancelled.
- 2
Cancel and .
Both are positive, so the sign stays.
- 3
Clear the factorials.
(n+1)!/n! = n + 1.
- 4
(b) Solve. , so the largest whole number is .
Check: P(Z=4) = 0.168 < P(Z=5) = 0.175, but P(Z=5) > P(Z=6) = 0.151.
Answer(a) (or ). (b) .
- 1
Mean equals variance, scaling λ, and combining probabilities
“
use the fact that if X ~ Po(λ) then the mean and variance of X are each equal to λ. Proofs are not required.
For the mean and the variance are both equal to :
You aren't asked to prove this, but §01 showed why it's true. A Poisson count is a binomial with a huge number of tiny time-slices, each with a tiny chance of an event, where . The binomial mean is . The binomial variance is , and when is tiny, is almost , so the variance is almost too. You can see it in the §01 diagram: as gets bigger, the bars move right and spread out.
This gives you a quick check on real data. If the mean and variance of some counts come out close to each other, a Poisson model is reasonable. If they're far apart, it probably isn't.
Two more consequences come up in exam questions:
- the standard deviation is , because the standard deviation is always the square root of the variance;
- mean = variance is a test a variable must pass. If a variable's mean and variance are different, it cannot be Poisson. For example, if and , then but (using the rules and , which the next topic, Linear Combinations of Random Variables, covers in full). These are not equal, so is not Poisson. also takes only even values , which a count of random events would not do.
Standard deviation, and a doubled variable that is not Poisson
The number of letters delivered to an office in a day has the distribution . The office is charged units of postage for each letter, so the total charge is units.
(a) Find the standard deviation of .
(b) Find and , and give a reason why does not have a Poisson distribution.
Show full working
- 1
Step 1 — (a) state the variance. For a Poisson variable the variance equals the mean:
Start from the property, not from memory of a 'standard deviation formula'.
- 2
Step 2 — (a) square-root it.
Writing 6.25 as the standard deviation is the usual slip. The s.d. is always the square root of the variance.
- 3
Step 3 — (b) find .
Multiplying a variable by 2 multiplies its mean by 2.
- 4
Step 4 — (b) find .
The multiplier is squared for the variance, because variance is measured in squared units.
- 5
Step 5 — (b) give the reason. but . A Poisson variable has mean equal to variance, so is not Poisson. (Also, can only be even.)
Quote the two numbers and say which property fails. Either reason scores.
(a) . (b) , ; mean variance, so is not Poisson.
Whenever a question asks 'why is this not Poisson?', check two things: are the mean and variance equal, and can the variable take every value 0, 1, 2, …?
Your turn: mean equals variance
- 19709/62 M/J 2023 Q2(a)1 mark
The random variable has a Poisson distribution. State the relationship between and .
Stuck? Show hint
One symbol answers it.
Show solution
- 1
State the property. .
The mark scheme insists on '=', not '≈'. Writing E(W) = λ and Var(W) = λ is also condoned.
Answer.
- 1
- 29709/71 O/N 2011 Q14 marks
The random variable has the distribution . The random variable is defined by .
(i) Find the mean and variance of .
(ii) Give a reason why the variable does not have a Poisson distribution.
Stuck? Show hint
E(2X) = 2E(X) and Var(2X) = 4Var(X), with E(X) = Var(X) = 1.3.
Show solution
- 1
(i) Mean.
B1 for 2.6.
- 2
(i) Variance.
Var(X) = 1.3 because X is Poisson; the 2 is squared for the variance.
- 3
(ii) Reason. , so is not Poisson. (Or: takes only even values, so it cannot take every whole number.)
Either reason scores the B1.
Answer(i) Mean , variance . (ii) The variance is not equal to the mean (or cannot take odd values).
- 1
Scaling λ to a different interval
Condition 3 from §01, a constant average rate, means the mean count grows in proportion to the length of the interval. If events happen at a certain rate per unit (per hour, per metre, per page), then over an interval units long,
This is where more marks are lost than anywhere else in the topic. The question gives a rate for one interval (per hour, say) and then asks about another (a -minute period, weeks). Always rescale to the interval in the question before you use the formula.
The rate stays fixed at 3 per hour and only the window length changes: 3 hours gives λ = 9, and 20 minutes (1/3 hour) gives λ = 1.
- 1
Find the rate, and note the interval it's quoted over (per hour, per metre, per week).
- 2
Find the interval the question asks about, in the same units. Convert minutes to hours, or months to years, if you need to.
- 3
Multiply: . Write down the new model, such as , with a new letter if it helps.
- 4
Only now use the Poisson formula.
Scaling the mean before substituting into the formula
Telephone calls arrive at a small business independently, singly and at a constant average rate of per hour. Find the probability that no calls arrive in a randomly chosen -hour period.
Show full working
- 1
Step 1 — identify the rate and the new interval length, separately. The rate is per hour; the interval asked about is hours.
Write both down so you can see they don't match: 1 hour against 2 hours.
- 2
Step 2 — scale the mean to the new interval, multiplying the rate by the interval length.
Scale first, as a separate line. Using λ = 3 here would answer a question about one hour, not two.
- 3
Step 3 — state the new model. Let be the number of calls in a -hour period: .
A new letter for the new interval stops you mixing up the 1-hour and 2-hour variables.
- 4
Step 4 — substitute into the formula for .
, so .
- 5
Step 5 — evaluate.
Three significant figures of a small number: 0.00248.
(3 s.f.).
Write λ for the new interval on its own line before any probability.
Scaling down to a smaller interval, then using a complement
At a certain shop, customers arrive independently and randomly at a constant average rate of per hour. Find the probability that, in a randomly chosen -minute period, at least customers arrive.
Show full working
- 1
Step 1 — convert the rate to the interval actually asked about. The rate is per hour ( minutes), so per minute: (the mark scheme's B1 needs seen)
Scaling down works the same way as scaling up: multiply the rate by the interval length, which is now a fraction of an hour.
- 2
Step 2 — state the model for a 1-minute period. Let be the number of arrivals in minute: .
Stating X ~ Po(0.39) with the interval named shows the scaling was intended.
- 3
Step 3 — translate "at least 2". Unbounded above, so use the complement:
'At least 2' includes 2, so the complement is 1 − P(X ≤ 1).
- 4
Step 4 — find .
.
- 5
Step 5 — find .
Reuse the unrounded .
- 6
Step 6 — add the two excluded terms.
Keep the sum unrounded, because it is about to be subtracted from 1.
- 7
Step 7 — subtract from 1.
When the answer is small, subtracting from 1 wipes out leading digits. That is why the terms needed 6 decimal places.
(3 s.f.).
Work in the units the question uses (here, minutes) before writing down λ.
Scaling to a longer interval, then a bounded range
The number, , of used computers donated to a charity has a constant average rate of computers per week. Assume that has a Poisson distribution. Calculate the probability that the number of computers donated during a -week period is more than and less than .
Show full working
P(6 < Y < 9) on Po(9.6): only the bars for y = 7 and y = 8 are shaded.
- 1
Step 1 — scale the mean to the 4-week period.
The rate is per week and the question asks about 4 weeks, so multiply by 4. The mark scheme gives B1 for 9.6.
- 2
Step 2 — state the model. Let be the number of computers donated in weeks: .
A new letter for the 4-week count keeps it separate from the weekly X.
- 3
Step 3 — translate "more than 6 and less than 9". This means , i.e. or : a finite, bounded list, short enough to sum directly.
'More than 6' excludes 6 and 'less than 9' excludes 9.
- 4
Step 4 — find .
With λ = 9.6, both and are huge. Work out their ratio (1490.97) as a single piece.
- 5
Step 5 — find .
Shortcut check: P(Y=8) = P(Y=7) × 9.6/8 = 0.100981 × 1.2.
- 6
Step 6 — add the two terms.
Different values of Y are mutually exclusive, so their probabilities add.
(3 s.f.).
Scaling and choosing sum-or-complement are two separate decisions. Do the scaling first.
Working backwards: when the interval length is unknown
So far the interval has always been given. Some questions give you a target probability instead and ask how long an interval you need. Call the unknown length , write , and solve the inequality.
Finding how long to wait for at least one event
Emails arrive at a desk independently, singly and at a constant average rate of per hour. Find the smallest whole number of minutes you should wait to be at least certain of receiving at least one email.
Show full working
- 1
Step 1 — express as a function of the unknown waiting time minutes. The rate is per minutes, so over minutes:
Same scaling as before, except that the length of time is now the unknown.
- 2
Step 2 — translate "at least one email" into a complement, since it has no top end (§02).
'At least one' is everything except zero, so only is needed.
- 3
Step 3 — impose the condition as an inequality.
'At least 99% certain' means ≥ 0.99.
- 4
Step 4 — isolate the exponential term. Subtract 1 from both sides, , then multiply by , which flips the inequality:
Multiplying by a negative number flips the inequality. It is the flip students most often forget.
- 5
Step 5 — take of both sides.
Because ln is an increasing function, the inequality sign stays the same.
- 6
Step 6 — multiply both sides by , flipping the inequality again since it's negative.
A second negative multiplier means a second flip.
- 7
Step 7 — round UP to a whole number of minutes. Waiting only minutes gives , just under ; whole minutes is the first that meets it.
This isn't normal rounding. 69 minutes gives slightly less than 99%, so it doesn't meet the condition.
At least minutes.
For 'at least' conditions on time, always round up, never to the nearest.
Finding the minimum waiting time from a target probability
The number of customers arriving at service desk during a -minute period has the distribution . An inspector waits at desk . She wants to wait long enough to be certain of seeing at least one customer arrive at the desk. Find the minimum time for which she should wait, giving your answer correct to the nearest minute.
Show full working
- 1
Step 1 — express as a function of the unknown waiting time minutes, scaling the rate as in the examples above.
Scale the rate from per 10 minutes to per minute: 2.1 ÷ 10 = 0.21.
- 2
Step 2 — write "at least one arrival" as a complement, since it has no top end (§02).
'At least one' is everything except zero.
- 3
Step 3 — set up the inequality the question describes: this probability must be at least .
The mark scheme's first M1 is for this inequality (it condones '=').
- 4
Step 4 — rearrange to isolate the exponential term. Subtract 1 from both sides, , then multiply by , which flips the inequality:
Multiplying by −1 flips the inequality sign.
- 5
Step 5 — take of both sides.
ln undoes e and brings x out of the power. Because ln is increasing, the inequality sign stays the same.
- 6
Step 6 — divide by , flipping the inequality since it's negative.
Dividing by a negative number is the second flip. The mark scheme also accepts working in 10-minute units: 2.302585/2.1 = 1.096 units, which is 10.96 minutes.
- 7
Step 7 — round to a whole number of minutes, rounding UP. Waiting only minutes gives , below ; only whole minutes meets the target.
Don't round to the nearest whole number here. 10 minutes gives slightly less than 90%, so it fails the condition.
She should wait at least minutes.
Whenever a question asks for the smallest time that makes you 'at least X% certain', round up to the next whole unit, even if the answer is much closer to the number below.
The other direction: the largest period, rounded down
The two examples above wanted at least one event, so a longer wait helped and the answer was a minimum, rounded up. Turn the question round and ask for no events: now a longer period makes "nothing happens" less likely, so the condition gives a maximum length, and a maximum rounds down.
| Question asks for | Condition | Solving gives | Round |
|---|---|---|---|
| minimum time to see at least one event | a number | up | |
| largest period with no events | a number | down |
The algebra is the same as before: write in terms of the unknown length, take logs, and watch the inequality sign every time you multiply or divide by a negative number.
The largest number of days with no breakdowns
Breakdowns of a lift occur at random at a constant mean rate of per day. Find the largest whole number of days, , for which the probability of no breakdowns in days is greater than .
Show full working
- 1
Step 1 — write in terms of .
The usual scaling: rate × length, with the length left as a letter.
- 2
Step 2 — write the condition.
'No breakdowns' is X = 0, and .
- 3
Step 3 — take natural logs.
ln is increasing, so the direction stays the same.
- 4
Step 4 — divide by , flipping the inequality.
Dividing by a negative number flips > into <. This flip is why the answer is a maximum.
- 5
Step 5 — round DOWN. must be less than , so the largest whole number is . Check: gives , but gives .
Rounding 4.46 to the nearest whole number also gives 4 here, but only by luck. Always round down for 'largest', and check both neighbours.
days.
After the last flip, look at the sign. 'n < number' means round down; 'n ≥ number' means round up.
Largest number of days with no accidents
The number of accidents on a certain road has a Poisson distribution with mean per -day period. The probability that there will be no accidents during a period of days is greater than . Find the largest possible value of .
Show full working
- 1
Step 1 — write the condition in terms of first.
The mark scheme's first M1 is for this. It allows '=' throughout.
- 2
Step 2 — take natural logs.
This is the second M1: 'attempt ln both sides'. ln is increasing, so the inequality sign stays the same.
- 3
Step 3 — multiply by , flipping the sign.
Multiplying by a negative number flips > into <, which is why λ, and so n, has a maximum.
- 4
Step 4 — link to . The rate is per days, so over days
Scale the rate to the unknown length, as in the rest of this section.
- 5
Step 5 — solve for .
0.008 is positive, so dividing by it keeps the sign.
- 6
Step 6 — round down. The largest whole number of days is . Check: gives ; gives .
The mark scheme accepts n = 6 or n ≤ 6, but not n < 6 or n ≥ 6.
The largest possible value of is .
Solving for λ first and then converting to n keeps each step short. Either order is fine, as long as the sign is tracked.
Substituting the rate given in the question directly into the Poisson formula, without checking it matches the interval being asked about
Always compare the interval the rate is quoted over against the interval named in the question, and scale if they differ
A weekly rate used unscaled in a '4-week period' question gives an answer that looks fine and is wrong, with no arithmetic slip to warn you.
Scaling using addition (e.g. '4 weeks means add 4') instead of multiplication
Scaling always multiplies:
The Poisson mean is proportional to the length of the interval. Doubling the interval doubles λ.
Your turn
- 1
A weaving machine produces flaws in cloth independently, singly and at a constant average rate of flaws per metre. Find the probability that a randomly chosen -metre length of cloth contains at least flaw.
Stuck? Show hint
Scale λ to the 5-metre length first (0.8 × 5), then use the complement for "at least 1".
Show solution
- 1
Scale. , so for a m length.
The rate is per metre and the question asks about 5 m, so scale first.
- 2
Complement.
'At least 1' has no top end; its complement is just the single value 0.
- 3
Find .
.
- 4
Subtract.
Round at the end: 0.982.
Answer(3 s.f.).
- 1
- 2
Accidents occur at a certain road junction independently and at random, at a constant average rate of per year. Find the probability that fewer than accidents occur in a randomly chosen -month period.
Stuck? Show hint
Scale the yearly rate down to 1 month (÷12) first, then sum P(X=0) and P(X=1) directly.
Show solution
- 1
Scale. per month, so .
A year is 12 months, so divide the yearly rate by 12.
- 2
Translate. "Fewer than " means or : bounded, so sum directly.
Strict '<' excludes 2 itself.
- 3
Find .
.
- 4
Find .
Reuse the unrounded .
- 5
Add.
Different values are mutually exclusive, so add.
Answer(3 s.f.).
- 1
- 34 marks
Meteorites are recorded striking a certain region independently, singly and at a constant average rate of per year. Find the minimum number of whole years of observation needed to be at least certain of recording at least one meteorite strike.
Stuck? Show hint
Write λ(t) = 0.6t, set up , solve for t using logarithms, then round up.
Show solution
- 1
Scale with unknown. Over years, .
The same scaling as always, with the length left as a letter.
- 2
Set up the condition. , and we need
'At least one' is everything except zero; 'at least 95%' means ≥.
- 3
Subtract 1, then multiply both sides by , flipping the inequality:
Multiplying by a negative number flips the inequality.
- 4
Take of both sides:
ln is increasing, so the inequality sign stays the same.
- 5
Divide by the negative number , flipping the inequality again:
Dividing by a negative is the second flip.
- 6
Round up. , so . Check: gives , while gives .
A minimum under an 'at least' condition always rounds up. Substituting both neighbours is a quick proof that 5 works and 4 doesn't.
AnswerAt least years.
- 1
- 49709/71 O/N 2013 Q4(ii)5 marks
The number of radioactive particles emitted per -minute period by some material has a Poisson distribution with mean . Find, in minutes, the longest time period for which the probability that no particles are emitted is at least .
Stuck? Show hint
Solve for λ first, then convert λ to minutes using 0.7 per 150 minutes. The answer is a length of time, so give it to 3 s.f.
Show solution
- 1
Condition. "No particles" is , so we need
M1 for this (the mark scheme allows '=').
- 2
Take logs.
The sign doesn't change when you take logs.
- 3
Multiply by , flipping.
A larger λ would make 'no particles' less likely, so λ has a maximum.
- 4
Convert to minutes. Over minutes, , so
Rate 0.7 per 150 minutes means 0.7/150 per minute.
- 5
Round. The longest period is minutes (3 s.f.).
Time is continuous here, so 3 s.f. is fine; 2.15 is below 2.1536, so it still satisfies the condition. Rounding up to 2.16 would not.
Answerminutes (3 s.f.).
- 1
Combining Poisson probabilities
Many questions build a bigger event out of simple Poisson probabilities. You already have every tool; the skill is choosing the right way to join them. Work out each simple probability first, on its own line, then combine:
| The question says | How to combine |
|---|---|
| " and ", for two independent counts | multiply: |
| one thing in one period and a different thing in the other, either order | multiply, then × 2 for the two orders |
| the same thing in each of periods | raise to the power: |
| exactly of periods | binomial: with |
| given |
Different periods of time that do not overlap are independent, because Poisson events occur independently. That is what allows the multiplying.
For a conditional probability on a single variable, look at what actually is. For example " given ": the values with and are just , so the top of the fraction is .
(Adding two independent Poisson counts together, such as "the total number of cars and trucks", uses the fact that is itself Poisson. That belongs to the next topic, Linear Combinations of Random Variables.)
Emails and texts: 'and', either order, and each of several hours
Emails arrive at an office at random at a constant mean rate of per hour, and, independently, texts arrive at random at a constant mean rate of per hour.
(a) Find the probability that, in a randomly chosen hour, at least email and exactly text arrive.
(b) Two separate hours are chosen at random. Find the probability that no emails arrive in one of these hours and at least email arrives in the other.
(c) Find the probability that at least email arrives in each of separate hours.
Show full working
- 1
Step 1 — state the models. Per hour, emails and texts , independent.
Name each count and its λ before combining anything.
- 2
Step 2 — (a) find .
'At least 1' is everything except 0.
- 3
Step 3 — (a) find .
Straight from the formula with λ = 1 and r = 1.
- 4
Step 4 — (a) multiply, since the counts are independent.
'And' for independent events means multiply.
- 5
Step 5 — (b) find the two single-hour probabilities.
Different hours don't overlap, so they are independent.
- 6
Step 6 — (b) multiply for one particular order. First hour none, second hour at least one:
This covers only one of the two ways it can happen.
- 7
Step 7 — (b) double for the two orders.
'In one of these hours… in the other' does not say which hour is which, so the reverse order counts too.
- 8
Step 8 — (c) raise to the power 5.
'Each of 5 hours' means the same event five times over, independently: multiply 0.8647 by itself 5 times.
(a) (b) (c) (all 3 s.f.).
Find each simple probability on its own line first. The combining step then becomes a one-line multiplication you can check easily.
Exactly 2 of 5 hours, and a conditional probability
Emails arrive at random at a constant mean rate of per hour, so the number in an hour is .
(a) Five separate hours are chosen. Find the probability that no emails arrive in exactly of the hours.
(b) Find the probability that exactly email arrives in an hour, given that at most emails arrive in that hour.
Show full working
- 1
Step 1 — (a) find the probability for one hour.
This single-hour probability becomes the 'success' probability of a binomial.
- 2
Step 2 — (a) recognise a binomial. Each hour either has no emails (probability ) or not (), independently, over hours. The number of empty hours is .
Fixed number of hours, two outcomes each, same p, independent: the four binomial conditions from §5.4.
- 3
Step 3 — (a) substitute into the binomial formula.
⁵C₂ = 10 counts which 2 of the 5 hours are the empty ones.
- 4
Step 4 — (a) evaluate.
Keep p unrounded right through.
- 5
Step 5 — (b) write the conditional formula.
Always start a conditional probability from the definition.
- 6
Step 6 — (b) simplify the top. If then is automatically true, so the top is just :
Ask which values satisfy both conditions. Only E = 1 does.
- 7
Step 7 — (b) find the bottom.
P(0) + P(1) + P(2), with taken out as a common factor.
- 8
Step 8 — (b) divide. (as a fraction, ).
The cancels, which is a nice check on the arithmetic.
(a) (3 s.f.) (b) .
'Exactly j of m periods' is a binomial whose p is a Poisson probability. 'Given' is a fraction whose top is the overlap of the two events.
At least 4 cars and at least 2 trucks
The numbers of cars and trucks arriving per minute at a fuel station are modelled by independent variables with distributions and respectively. Find the probability that at least cars and at least trucks arrive at the fuel station during a randomly chosen -minute period.
Show full working
- 1
Step 1 — scale both means to minutes. Cars: . Trucks: .
Both rates are per minute and the question asks about 5 minutes.
- 2
Step 2 — cars: set up the complement. With ,
'At least 4' has no top end, so use the complement, which stops at 3.
- 3
Step 3 — cars: list the four terms.
One term per value, each from the formula. The mark scheme wants the full expression or terms seen.
- 4
Step 4 — cars: add the four terms.
Add the excluded terms first, then subtract once. That avoids the 1 − P(0) + P(1) sign slip.
- 5
Step 5 — cars: subtract from 1.
Keep 5 or 6 significant figures, because this is about to be multiplied.
- 6
Step 6 — trucks: set up the complement. With ,
'At least 2' stops the complement at 1.
- 7
Step 7 — trucks: find the two terms.
and , so both terms are multiples of .
- 8
Step 8 — trucks: subtract.
This is the second M1 of the mark scheme.
- 9
Step 9 — multiply, because the counts are independent.
'At least 4 cars AND at least 2 trucks' with independent counts: multiply. Adding the λ's would answer a different question (about the total).
(3 s.f.).
Two different conditions on two independent counts means two separate Poisson calculations, then multiply. Only combine the λ's when the question asks about the total.
Exactly 1 order in one hour and at least 2 in the other
The number of orders arriving at a shop during an -hour working day is modelled by the random variable with distribution . Find the probability that, in two randomly chosen -hour periods, exactly order will arrive in one of the -hour periods, and at least orders will arrive in the other -hour period.
Show full working
- 1
Step 1 — scale to one hour.
The distribution is given per 8 hours; the question is about 1-hour periods.
- 2
Step 2 — find .
Straight from the formula.
- 3
Step 3 — write as a complement.
The complement of 'at least 2' is 0 or 1, and is a common factor of both terms.
- 4
Step 4 — evaluate it.
; keep it unrounded for the product.
- 5
Step 5 — multiply for one order.
The two hours are separate, so they are independent. This is the M1 for a product of two Poisson probabilities.
- 6
Step 6 — double for the two orders.
'One of the periods… the other' means either hour could be the one with exactly 1. The mark scheme has a separate M1 for this × 2.
(3 s.f.).
When two different outcomes are shared between two periods without saying which is which, multiply by 2.
Your turn: combining probabilities
- 19709/62 O/N 2023 Q7(b)3 marks
A random variable has the distribution . Two independent values of are chosen. Find the probability that both of these values are greater than .
Stuck? Show hint
Find P(X > 1) = 1 − P(X = 0) − P(X = 1) for one value, then square it.
Show solution
- 1
One value.
'Greater than 1' excludes 0 and 1.
- 2
Evaluate. , so
Keep this unrounded; it is about to be squared.
- 3
Both values.
Two independent values, same event each time: square. Using Po(4.8) instead would be the total of the two, which is a different question.
Answer(3 s.f.).
- 1
- 29709/62 O/N 2020 Q5(d)3 marks
Customers arrive at a shop at a constant average rate of per minute, and the number arriving per minute has the distribution . Five -minute periods are chosen at random. Find the probability that no customers arrive during exactly of these periods.
Stuck? Show hint
p = P(no customers in a minute) = . Then use the binomial formula with n = 5, r = 2.
Show solution
- 1
Single period.
M1 for P(none arrive) clearly identified.
- 2
Binomial. The number of empty periods out of is , so
5 independent periods, each empty or not, with the same p.
- 3
Substitute.
⁵C₂ = 10.
- 4
Evaluate.
The mark scheme accepts 0.0732 or 0.0733.
Answer(3 s.f.).
- 1
- 39709/71 M/J 2012 Q5(i)5 marks
A random variable has the distribution .
(a) Find .
(b) Find the probability that given that .
Stuck? Show hint
For (b), X = 3 and X ≥ 3 together is just X = 3, so divide P(X = 3) by your answer to (a).
Show solution
- 1
(a) Complement.
'At least 3' stops the complement at 2. 1 + 3.2 + 5.12 = 9.32.
- 2
(a) Evaluate.
Keep it unrounded for part (b).
- 3
(b) Top of the fraction. " and " is just :
Which values satisfy both conditions? Only 3.
- 4
(b) Divide.
Conditional probability = overlap ÷ condition.
Answer(a) (b) (3 s.f.).
- 1
The Poisson approximation to the binomial distribution
“
use the Poisson distribution as an approximation to the binomial distribution where appropriate. The conditions that n is large and p is small should be known; n > 50 and np < 5, approximately.
Some situations really are binomial: a fixed number of trials , each with the same probability . But sometimes is very large and very small, like the number of faulty items in a batch of , or the number of typing errors in characters. Working out is hopeless, and you don't need to. §01 showed that a binomial with large and small is almost the same as a Poisson, so you can use
with , and work with the Poisson formula from there.
How large, how small?
The syllabus says " large, small", with the working guideline and , approximately. When a question asks you to justify the approximation, mark schemes want the values, not the words: " large, small" on its own scores nothing. Write, for example, " and " (or "" in place of the check).
Why these conditions? The closer a real binomial is to the "huge , tiny " picture from §01, the better the match. You can also see it through mean and variance. A Poisson has mean = variance. A binomial has mean and variance , and these are only close when is close to , that is, when is small. The half is the exam's dividing line: when is bigger, you're expected to use the normal approximation (§5.5) instead. The table at the end of §05 puts all the choices side by side.
B(200, 0.02) against Po(4): with n large and p small the exact binomial bars and the Poisson approximation are almost identical (at x = 3, 0.1963 against 0.1954).
- 1
Identify the exact model: state from the situation described, as in §5.4.
- 2
Check both conditions, with values: and (or ). Writing " large, small" without the numbers does not score.
- 3
Find , and state the approximating model, .
- 4
Forget and from this point on. Every remaining step (§01–§03) uses alone.
Comparing the exact binomial value against its Poisson approximation
Each component made by a machine is faulty with probability , independently of the others. A batch of components is checked. Use a Poisson approximation to find the probability that exactly are faulty, and compare your answer with the exact binomial probability.
Show full working
- 1
Step 1 — identify the exact model, and check the approximation conditions. The number of faulty components is . Here is large and is small, so a Poisson approximation should be appropriate (Step 2 confirms ).
Name the exact binomial first, so it is clear what is being approximated. Quote the numbers against the guideline; the words alone don't score.
- 2
Step 2 — find for the approximating distribution.
λ = np is its own line; it is the only thing carried forward.
- 3
Step 3 — state the approximating model. .
Write '≈' rather than '~', because this is an approximation.
- 4
Step 4 — substitute into the Poisson formula for .
and , worked out as separate pieces.
- 5
Step 5 — evaluate the Poisson approximation.
Keep unrounded until this multiplication.
- 6
Step 6 — for comparison, the exact binomial value is The two answers differ by less than , so the approximation is good here.
You won't usually be asked for this comparison. It's here so you can see how close the two answers are.
Poisson approximation: (3 s.f.); exact binomial value: (3 s.f.).
Once you have λ = np, you're finished with n and p. Everything after that uses λ.
Recognising the approximation is needed, then using a cumulative range
The data produced by a certain data entry firm always include a small number of incorrect characters that occur at random. The proportion of incorrect characters is denoted by , and experience has shown that . A particular data set from the firm contains characters, of which characters are incorrect. Use a suitable approximating distribution to find .
Show full working
'Less than 4' on Po(1.45) stops at x = 3, so the four bars x = 0, 1, 2, 3 are summed directly.
- 1
Step 1 — identify the exact model, and check the approximation conditions. . Here is large and is tiny, so a Poisson approximation is appropriate (the next step confirms ).
The question says 'suitable approximating distribution'. Name the exact model first, so it is clear what is being approximated.
- 2
Step 2 — find .
The mark scheme's B1 needs λ = 1.45 together with an indication that you are using a Poisson.
- 3
Step 3 — state the approximating model. .
From here on, only λ is used; n and p are finished with.
- 4
Step 4 — translate "". This means : a finite, bounded list, so sum directly.
Strict '<' excludes 4 itself.
- 5
Step 5 — find .
.
- 6
Step 6 — find .
Reuse the unrounded .
- 7
Step 7 — find .
1.45² = 2.1025, divided by 2! = 2.
- 8
Step 8 — find .
1.45³ = 3.048625, divided by 3! = 6. This is the term most often dropped, but 'X < 4' ends at 3.
- 9
Step 9 — add all four terms.
Different values of X are mutually exclusive, so add.
(3 s.f.).
With four terms to add, keep 5 or 6 significant figures on each one. Rounding every term early can change the final answer.
Finding an unknown p from a target P(X=0)
Each of light bulbs independently fails in its first week with probability , where is small. Using a Poisson approximation, find the value of for which the probability that no bulbs fail in the first week is .
Show full working
- 1
Step 1 — write in terms of . , approximated by with
n is known and p is not, so the Poisson mean is left in terms of p.
- 2
Step 2 — write .
, so . The unknown now sits in the exponent.
- 3
Step 3 — set it equal to the target.
One equation, one unknown.
- 4
Step 4 — take of both sides.
ln undoes e, which brings p out of the exponent.
- 5
Step 5 — divide both sides by .
Check the approximation still holds: np = 0.693 < 5 and p is tiny.
(3 s.f.).
An unknown p inside a Poisson approximation always sits in λ = np. If the question involves P(X=0), expect and one logarithm.
Working backwards through the approximation to find an unknown p
Continuing the data-entry example above (, characters, ): the firm's management wishes to decrease the value of by giving their employees some training. Their aim is that, for a data set containing characters, the value of for the new value of should be double the value of when . Use a suitable approximating distribution to find the new value of .
Show full working
- 1
Step 1 — write the new mean in terms of . With unchanged,
Only p changes, so the approximating mean is np with p left as a letter.
- 2
Step 2 — write for this mean.
, so the unknown sits in the exponent and logs will be needed.
- 3
Step 3 — find the target value: double the original . From part (a) of the data-entry example, when , so the target is
Keep it as if you can. The mark scheme accepts either form.
- 4
Step 4 — set the new equal to this target, forming an equation in .
This equation is the mark scheme's M1: an equation in p, ready to solve.
- 5
Step 5 — take of both sides to bring out of the exponent.
ln undoes e. Using avoids any rounding.
- 6
Step 6 — divide both sides by to isolate .
Dividing by the negative coefficient makes both sides positive. A negative p here would mean the 2 went on the wrong side of the equation.
(3 s.f.).
When a target is described as a multiple of an earlier probability ('double', 'half'), write that earlier value down first. Then the equation has a definite number on the right.
Justifying with values, then finding the sample size n
Most plants of a certain type have three leaves. However, it is known that, on average, in of these plants have four leaves, and plants with four leaves are called 'lucky'. The number of lucky plants in a random sample of plants is denoted by .
(a) State, with a justification, an approximating distribution for , giving the values of any parameters.
(d) The number of lucky plants in a random sample of plants, where is large, is denoted by . Given that , correct to significant figures, use a suitable approximating distribution to find the value of .
Show full working
- 1
Step 1 — (a) identify the exact model. Each plant is lucky with probability , independently, so .
Name n and p now, because you'll quote them in the justification.
- 2
Step 2 — (a) find .
This number is needed both for the distribution and for the justification.
- 3
Step 3 — (a) state the approximation. .
B1: 'Poisson with mean 2.5'. Writing only np = 2.5 is not enough.
- 4
Step 4 — (a) justify with the numbers. and (or ).
The mark scheme says 'must see 2.5 (or 0.0001) and 25000'. Quoting only n > 50 and np < 5 without the values does not score.
- 5
Step 5 — (d) write for the new sample. With plants, where
The sample size is now the unknown, so λ is left in terms of n.
- 6
Step 6 — (d) write as a complement.
'At least 1' leaves only the single term .
- 7
Step 7 — (d) set it equal to the given value.
This is the first M1.
- 8
Step 8 — (d) take natural logs.
The second M1 is for correct use of ln.
- 9
Step 9 — (d) convert to .
λ = np, so n = λ ÷ p.
- 10
Step 10 — (d) round. Because was only given to s.f., is only known to s.f.: .
The mark scheme accepts any whole number from 32 950 to 33 050.
(a) ; and . (d) (3 s.f.).
When n is the unknown, keep λ = np as a letter expression, solve for λ with a logarithm, then divide by p.
Checking only that is large, without also checking that is small, or writing " large, small" with no numbers
Check both with values: and (or ). If is larger than about , CAIE expects a normal approximation to the binomial (§5.5) instead
Poisson needs mean ≈ variance, and np(1−p) is only close to np when p is small. A large n alone doesn't make p small.
Carrying the original and forward into later working after switching to the Poisson approximation
Once you have , work only with
Mixing binomial and Poisson pieces in one calculation gives an answer that belongs to neither model.
Your turn
- 19709/63 M/J 2023 Q6(b)2 marks
It is known that in people in a certain country have a particular blood condition. A random sample of people is chosen, and the number having the condition is exactly. Find and , and explain briefly why your answers suggest that a Poisson approximation may be reasonable here.
Stuck? Show hint
Compare np and np(1−p) for the exact binomial. If they're close, that supports a Poisson approximation.
Show solution
- 1
Mean.
Use the binomial mean formula on the exact distribution.
- 2
Variance.
The mark scheme says writing just 2.5 for the variance is not sufficient. Show 2.4995 (or 4999/2000).
- 3
Compare. and are almost equal, and a Poisson distribution has its mean equal to its variance, so is likely to be a good approximation.
The explanation mark needs the comparison ('almost equal') tied to the Poisson property.
Answerand . These are almost equal, which supports a Poisson approximation .
- 1
- 24 marks
A large company finds that of invoices contain an error, independently of each other. In a random sample of invoices, use a Poisson approximation to find the probability that more than invoices contain an error.
Stuck? Show hint
λ = np = 300 × 0.005. "More than 2" has no top end, so use the complement, 1 − P(X ⩽ 2).
Show solution
- 1
Conditions and . and is small; , so .
Check both conditions with numbers, then compute λ = np on its own.
- 2
Complement.
'More than 2' excludes 2, so the complement includes it.
- 3
.
- 4
Reuse the unrounded .
- 5
1.5² = 2.25, divided by 2! = 2.
- 6
, so
Round at the end: 0.191.
Answer(3 s.f.).
- 1
- 39709/62 O/N 2025 Q4(b)1 mark
An inspector believes that of cups made at a certain factory contain flaws. The factory owner claims that the true percentage is less than . The inspector examines a random sample of cups and finds that of them contain flaws, and a test is carried out using the binomial distribution. Explain why it would not be appropriate to use the Poisson approximation to the binomial distribution to carry out the test.
Stuck? Show hint
Check n > 50 and np < 5 with the actual numbers. One failed condition, quoted with its value, is enough.
Show solution
- 1
Identify and . and .
The exact model is B(40, 0.18).
- 2
Check each condition with its value. , which is more than . (Also is not more than , and is not small.)
The mark scheme needs context: the number 7.2, 40 or 0.18 must appear. One correct reason scores.
Answer(or is not , or is not small), so the Poisson approximation is not appropriate.
- 1
- 49709/62 M/J 2023 Q2(b)1 mark
The random variable has the distribution . Jyothi wishes to use a Poisson distribution as an approximate distribution for . Use the formulae for and to explain why it is necessary for to be close to for this to be a reasonable approximation.
Stuck? Show hint
A Poisson distribution has mean = variance. Compare np with np(1 − p).
Show solution
- 1
Write both binomial formulae. and .
The mark scheme requires the formulae to be seen.
- 2
Use the Poisson property. A Poisson distribution has mean equal to variance, so we need .
This is the link between the two distributions.
- 3
Conclude. Dividing by gives , so must be close to .
The conclusion about 1 − p (or q) is what earns the B1.
Answerand ; for a Poisson these must be (almost) equal, so must be close to , i.e. close to .
- 1
- 59709/61 M/J 2021 Q5(b)4 marks
On average, in adults has a certain genetic disorder. In a random sample of people, where is large, the probability that no-one has the genetic disorder is more than . Find the largest possible value of .
Stuck? Show hint
λ = n/75 000. Solve with logs, watch the sign, then round DOWN (see §03).
Show solution
- 1
Model. .
B1 for the mean in terms of n.
- 2
Condition.
'No-one' is X = 0, and .
- 3
Take logs.
ln is increasing, so the sign stays as >.
- 4
Multiply by , flipping.
A negative multiplier flips > into <, so n has a maximum.
- 5
Round down. The largest possible value is .
'Largest' with n < 7902.04 means round down. The mark scheme needs an integer.
Answer.
- 1
- 69709/72 M/J 2015 Q7(iii)4 marks
In a certain lottery, tickets have been sold altogether and each ticket has a probability of of winning a prize. The random variable denotes the number of prize-winning tickets that have been sold. Use a Poisson approximating distribution to find the conditional probability that , given that .
Stuck? Show hint
λ = 10 500 × 0.0002 = 2.1. The overlap of 'X < 4' and 'X ≥ 1' is X = 1, 2 or 3.
Show solution
- 1
Approximation. with and , so .
λ = np = 2.1.
- 2
Bottom of the fraction.
M1 for P(X ≥ 1).
- 3
Overlap. " and " means :
2.1 + 2.205 + 1.5435 = 5.8485. Leave out 0, since the condition is X ≥ 1.
- 4
Divide.
Overlap ÷ condition, as in §03.
Answer(3 s.f.).
- 1
The normal approximation to the Poisson, and choosing an approximation
“
use the normal distribution, with continuity correction, as an approximation to the Poisson distribution where appropriate. The condition that λ is large should be known; λ > 15, approximately.
Look at the picture in §01. It already looks like a fairly symmetric hump. As gets larger the shape gets closer to a normal curve, and adding up dozens of Poisson terms by hand is out of the question anyway. So for large we approximate the Poisson with a normal distribution.
A Poisson variable has and (§03), so the matching normal distribution is
with mean and variance . The standard deviation is therefore , not .
How large is 'large'?
The working guide is . For small the Poisson is clearly skewed to the right (look at in §01), and a symmetric normal curve would even give some probability to negative counts. By about the shape is close enough to symmetric for the normal curve to fit well, and the bigger gets, the better the fit. When you justify the approximation, quote the value: "".
A quick recap from Paper 5
Standardising. If , then has the standard normal distribution , and is read from the tables. For a negative , use .
Working backwards. If you're given a probability and need , read the table the other way. For the common values use the critical-values table at the bottom: for example, gives .
Why a continuity correction? A Poisson variable only takes whole numbers, but a normal variable is continuous, and for a continuous variable . To make the two match, think of each whole number as a bar of width running from to . The probability of is then the area under the normal curve over that bar. So "more than " starts at , and "at least " starts at . This works just like the normal approximation to the binomial in §5.5; revise it there if it feels shaky.
Po(25) with its matching N(25, 25): the bar for 31 starts at 30.5, so 'more than 30' becomes the area to the right of 30.5, and the spread is σ = √λ = 5, not λ.
Phrase | Continuity-corrected boundary |
|---|---|
use | |
use | |
use | |
use |
The boundary always moves half a unit towards the values being included.
- 1
Find for the interval named in the question, scaling first if needed (§03).
- 2
Check is large, quoting the value (), and state .
- 3
Apply the continuity correction to the boundary, using the table above. Never skip it when a discrete variable is approximated by a continuous one.
- 4
Standardise, using (never itself) in the denominator.
- 5
Express as an area under , using symmetry to rewrite a negative if needed, and evaluate.
Standardising a large-λ Poisson variable, with a continuity correction
The number of accidents at a certain factory has a Poisson distribution with mean per year. Use a suitable normal approximation to find the probability that more than accidents occur in a randomly chosen year.
Show full working
- 1
Step 1 — check the approximation is appropriate. is large, so a normal approximation applies.
Quote the guideline number, not just the word 'large'.
- 2
Step 2 — state the approximating normal distribution, using as both the mean and the variance.
Both parameters are λ: mean 25 and variance 25.
- 3
Step 3 — apply the continuity correction. "More than " means ; from the table above, this needs the boundary .
'More than 30' starts at 31, whose bar begins at 30.5. So use 30.5, not 29.5.
- 4
Step 4 — identify and . and .
The second parameter in N(25, 25) is the variance. The standard deviation is its square root.
- 5
Step 5 — substitute and evaluate .
Use the corrected boundary 30.5 in the numerator, not 30.
- 6
Step 6 — express the required probability using .
Tables give Φ(z) = P(Z < z), the area to the left. 'More than' is the area to the right, so use 1 − Φ.
- 7
Step 7 — read from the normal tables and subtract.
Φ(1.100) = 0.8643 from the table. For comparison, the exact Poisson answer is 0.1367, so the approximation is out by only 0.001.
(3 s.f.).
The standard deviation is √λ, not λ. Dividing by λ is the easiest mistake to make here.
Scaling the mean first, then a normal approximation with continuity correction
The number, , of used computers donated to a charity has a constant average rate of computers per week, and has a Poisson distribution. Use a suitable approximating distribution to calculate the probability that more than computers are donated during a -week period.
Show full working
- 1
Step 1 — scale the mean to the 20-week period (§03).
Scale first. The unscaled λ = 2.4 would not even count as large.
- 2
Step 2 — check the normal approximation is appropriate. is large, so approximate by a normal distribution.
Say '48 > 15', not just 'large'.
- 3
Step 3 — state the approximating distribution.
The mark scheme's B1 is for N(48, 48): mean and variance both λ.
- 4
Step 4 — apply the continuity correction to "more than 50". needs the boundary .
'More than 50' starts at 51, whose bar begins at 50.5.
- 5
Step 5 — identify and . and
σ = √48, not 48.
- 6
Step 6 — substitute and evaluate .
Keep z to at least 3 decimal places for the table.
- 7
Step 7 — express as an area.
This is the right-hand tail, so use 1 − Φ.
- 8
Step 8 — read the table and evaluate. , so
Use z to 3 d.p. (0.361). Rounding z to 0.36 gives 0.3594, which happens to round to 0.359 here, but that habit costs marks elsewhere.
(3 s.f.).
The normal approximation uses whatever λ you get after scaling, so scale first.
A 'fewer than' boundary and a negative z
The number of emails a company receives in a day has the distribution . Use a suitable approximating distribution to find the probability that fewer than emails arrive on a randomly chosen day.
Show full working
- 1
Step 1 — check and state the approximation. , so .
λ is large; the mean and the variance are both λ.
- 2
Step 2 — apply the continuity correction. "Fewer than " means , whose bar ends at , so use .
The boundary moves half a unit towards the included values, which is downwards for 'fewer than'.
- 3
Step 3 — identify and . and .
σ = √λ, never λ.
- 4
Step 4 — standardise.
A boundary below the mean always gives a negative z. Use that as a sign check.
- 5
Step 5 — use symmetry.
Tables give only positive z. The left tail beyond −z equals the right tail beyond +z.
- 6
Step 6 — read the table and subtract. , so
Use the table's 'ADD' columns for the third decimal place of z.
(3 s.f.).
'Fewer than k' uses k − 0.5, lands left of the mean and gives a negative z. Then use 1 − Φ(|z|).
Scaling up, then a 'less than' boundary
Sales of cell phones at a certain shop occur singly, randomly and independently, at a constant average rate of per hour. Use a suitable approximating distribution to find the probability that the number of sales during a randomly chosen -month period ( hours) will be less than .
Show full working
For 'less than 150' the boundary moves left to 149.5: the bar for 150 is excluded and z is negative.
- 1
Step 1 — scale the mean to the 140-hour period.
The rate is per hour and the period is 140 hours, so multiply.
- 2
Step 2 — check the approximation. , so a normal approximation is appropriate.
Quote the guideline number.
- 3
Step 3 — state the approximating distribution.
The mark scheme's B1 is for N(168, 168), stated or implied.
- 4
Step 4 — apply the continuity correction to "less than 150". needs the boundary .
'Less than 150' ends at 149, whose bar ends at 149.5.
- 5
Step 5 — identify and . and
σ = √168, not 168.
- 6
Step 6 — substitute and evaluate .
The boundary is below the mean, so z is negative, as expected.
- 7
Step 7 — rewrite the negative-z probability using symmetry.
Tables list only positive z. By symmetry, the area to the left of −1.427 equals the area to the right of +1.427.
- 8
Step 8 — read from the tables and subtract from 1.
Use z to 3 d.p. (1.427) in the table, not a z already rounded to 1.43. That premature rounding gives 1 − 0.9236 = 0.0764, which is outside the mark scheme's accepted 0.0767 or 0.0768.
(3 s.f.; from a calculator's exact is also accepted).
A 'less than' boundary below the mean gives a negative z. Rewrite it as 1 − Φ(|z|) before using the tables, which only list positive z.
Ranges with two ends, stating the distribution, and justifying it
Two-sided ranges. "Between and " needs a continuity correction at both ends, each moved half a unit towards the values being kept. Read the words carefully, because "inclusive" and strict inequalities move the ends in different directions:
| Range | Lower boundary | Upper boundary |
|---|---|---|
| ("between and inclusive") | ||
Then standardise each boundary separately, and find the area between them: .
Stating the approximating distribution. When a question says "state a suitable approximating distribution, giving the values of any parameters", write with both numbers. Mark schemes often give one mark for the mean and a separate mark for the variance, so alone loses a mark.
Justifying it. Quote the actual value against the guideline: "". The words " is large", or "" without the value, do not score.
A two-sided range: correcting both ends
The number of parcels delivered to a depot in a day has the distribution .
(a) State a suitable approximating distribution, giving its parameters, and justify its use.
(b) Use it to find the probability that between and parcels inclusive are delivered on a randomly chosen day.
Show full working
- 1
Step 1 — (a) state and justify. , so .
Both parameters written, and the justification quotes the value 50.
- 2
Step 2 — (b) correct the lower end. "Inclusive" keeps , whose bar starts at . Lower boundary: .
Move outwards, towards the values being kept.
- 3
Step 3 — (b) correct the upper end. "Inclusive" keeps , whose bar ends at . Upper boundary: .
Again outwards. Both ends move away from the middle for an inclusive range.
- 4
Step 4 — (b) identify and . and
The standard deviation is √λ, not λ.
- 5
Step 5 — (b) standardise the lower boundary.
Below the mean, so negative.
- 6
Step 6 — (b) standardise the upper boundary.
The range is symmetric about 50 here, so z₂ = −z₁. That is not always the case.
- 7
Step 7 — (b) write the area between them.
Φ(−z) = 1 − Φ(z), from the symmetry of the normal curve.
- 8
Step 8 — (b) read the table and evaluate. , so
Use z to 3 d.p. in the table.
(a) , since . (b) (3 s.f.).
Draw a quick sketch of the two boundaries on a bell curve before writing Φ's. It shows at a glance whether you need Φ(z₂) − Φ(z₁) or something else.
Stating N(λ, λ), then a strict two-sided range
At a certain shop, customers arrive independently and randomly at a constant average rate of per hour. The random variable denotes the number of customers who arrive in a randomly chosen -hour period.
(i) State a suitable approximating distribution for , giving the value(s) of any parameter(s).
(ii) Use your approximating distribution to find .
Show full working
- 1
Step 1 — (i) find . The rate is per hour and counts one hour, so .
No scaling needed, but check it.
- 2
Step 2 — (i) state the approximation. , so .
B1 for the mean 23.4 and a separate B1 for the variance 23.4. The mark scheme notes that marks lost here cannot be recovered in (ii).
- 3
Step 3 — (ii) correct the lower end. "" starts at , whose bar begins at .
Strict inequality: 20 itself is excluded, so the boundary moves up, not down.
- 4
Step 4 — (ii) correct the upper end. "" ends at , whose bar ends at .
30 is excluded, so the boundary moves down.
- 5
Step 5 — (ii) identify and . and
The mark scheme needs square roots in both standardisations.
- 6
Step 6 — (ii) standardise the lower boundary.
Below the mean, so negative.
- 7
Step 7 — (ii) standardise the upper boundary.
Standardise each end separately.
- 8
Step 8 — (ii) look up each value. and .
Negative z: use symmetry. −0.5995… is −0.600 to 3 d.p.
- 9
Step 9 — (ii) subtract.
Area between = upper Φ minus lower Φ.
(i) . (ii) (3 s.f.).
For a strict range a < X < b, both boundaries move inwards (a + 0.5 and b − 0.5). For an inclusive range they move outwards.
Working backwards: λ itself is unknown
So far you've been given and asked for a probability. Some questions turn this round: they give you a probability and ask for (often called ). A probability given to 4 decimal places, like , is a clue. It has come straight out of the normal table, so you can read the -value back from it.
The awkward part is that now appears twice in the standardising equation, once on its own and once inside . The way round this is to write , so . The equation becomes a quadratic in , which you solve in the usual way. Take the positive root only, because a square root can't be negative, and then square it to get .
Solving a quadratic in √μ to find an unknown mean
The random variable has the distribution , where is large. Using a suitable approximating distribution, it is found that . Find .
Show full working
- 1
Step 1 — state the approximating model, with still unknown. Since is large, .
μ is large, so use N(μ, μ), with the unknown in both parameters.
- 2
Step 2 — apply the continuity correction to "". From the continuity-correction table above, this needs the boundary .
'Y < 41' ends at 40, whose bar ends at 40.5.
- 3
Step 3 — find the -value from the given probability. is less than , so must be negative: since , , so .
The given 0.1587 is a probability, but the equation you need links a boundary to μ through z. Converting the probability to z first (reading the table backwards) gives you something to put in the standardising formula. Sign check: a probability below 0.5 always means a negative z.
- 4
Step 4 — write the standardising equation, with as the unknown.
The corrected boundary goes in the numerator and √μ goes in the denominator.
- 5
Step 5 — substitute (so and ), and clear the fraction by multiplying both sides by .
u² replaces μ and u replaces √μ, so the equation has no square roots left.
- 6
Step 6 — rearrange into standard quadratic form in .
Everything on one side, equal to zero.
- 7
Step 7 — solve with the quadratic formula.
Here a = 1, b = −1, c = −40.5.
- 8
Step 8 — reject the negative root, since cannot be negative.
u was defined as a square root, so it cannot be negative.
- 9
Step 9 — square to recover .
Square the unrounded u. Squaring 6.88 instead gives 47.3, which is wrong at 3 s.f.
(3 s.f.).
Put u = √μ before you clear the fraction. It turns the equation into an ordinary quadratic in u.
Finding μ from a probability given to 4 decimal places
The random variable has the distribution where is large. Using a suitable approximating distribution, it is found that , correct to 4 decimal places. Find .
Show full working
- 1
Step 1 — state the approximating model. Since is large, .
The mark scheme gives an M1 for N(μ, μ): mean and variance both μ.
- 2
Step 2 — apply the continuity correction to "". The boundary needed is .
'Y < 46' ends at 45, whose bar ends at 45.5.
- 3
Step 3 — find from the given probability. is less than , so is negative: since , , so .
This is the mark scheme's first M1: . A probability below 0.5 means a negative z.
- 4
Step 4 — write the standardising equation, with unknown.
The corrected boundary goes in the numerator and √μ in the denominator.
- 5
Step 5 — substitute and clear the fraction.
u² replaces μ and u replaces √μ, so no square roots remain.
- 6
Step 6 — rearrange into standard quadratic form.
Get everything on one side, equal to zero, ready to solve.
- 7
Step 7 — solve with the quadratic formula.
Here a = 1, b = −1.5, c = −45.5.
- 8
Step 8 — reject the negative root.
u = √μ cannot be negative.
- 9
Step 9 — square to recover .
Square the unrounded u.
(3 s.f.).
Same steps as the last example. The hard part is spotting that a probability given to 4 d.p. means 'read z from the table, then solve for μ'.
Standardising with instead of
, so the standard deviation used in standardising is
Everything else can be right, but dividing by λ instead of √λ makes z, and so the answer, wrong.
Omitting the continuity correction, or applying it in the wrong direction
Always move the boundary half a unit towards the values being included. Check against the table above rather than guessing
Mark schemes give a separate mark for the continuity correction, and you lose it if it's missing or goes the wrong way.
When λ itself is unknown, trying to rearrange the standardising equation directly for λ without substituting u = √λ first
Put as its own step whenever appears both on its own and inside a square root. The equation then becomes a quadratic in
Rearranging for λ directly goes round in circles: move the λ term and the √λ term is still there, and vice versa.
Keeping the negative root of the quadratic in u, or squaring both roots and reporting two values of λ
Reject the negative root straight away, since is positive. Only the positive root is squared to get
Squaring the negative root gives a positive number that looks like a possible λ, but it doesn't satisfy the original standardising equation, and the mark scheme wants one answer only.
Your turn
- 19709/73 O/N 2015 Q25 marks
The number of calls received per 5-minute period at a large call centre has a Poisson distribution with mean , where . If more than calls are received in a 5-minute period, the call centre is overloaded. It has been found that the probability of being overloaded during a randomly chosen 5-minute period is . Use the normal approximation to the Poisson distribution to obtain a quadratic equation in and hence find the value of .
Stuck? Show hint
P(Z > z) = 0.01 gives z = Φ⁻¹(0.99) = 2.326. Apply the continuity correction to "more than 55", then substitute u = √λ.
Show solution
- 1
Model. Let be the number of calls in a 5-minute period. Then .
λ > 30 is large, so use the normal approximation with mean and variance both λ.
- 2
from the probability. , so the area to the left of is : .
The table's critical value for 0.99 is 2.326. Use it rather than reading the main table backwards.
- 3
Continuity correction. "More than " means , whose bar begins at .
Without the correction you get λ = 40.2, and the mark scheme withholds the final A1.
- 4
Standardise. ( is positive because : the overload region is the right tail.)
The boundary is above the mean, so z is +2.326, not −2.326.
- 5
Substituting and clearing the fraction: , i.e.
u² replaces λ and u replaces √λ.
- 6
Here , , :
Identify a, b and c before substituting into the formula.
- 7
Rejecting the negative root:
The other root, −8.703…, is negative, and u = √λ cannot be.
- 8
Squaring:
Use the full calculator value of u when you square it.
Answer(3 s.f.).
- 1
- 2
A shop sells an average of items of a certain product per day, and the number sold has a Poisson distribution. Use a normal approximation to find the probability that fewer than items are sold on a randomly chosen day.
Stuck? Show hint
N(40, 40). "Fewer than 35" needs the boundary 35 − 0.5 = 34.5.
Show solution
- 1
Check and state. , so .
The approximation needs a large λ.
- 2
Continuity correction. "Fewer than " means , so use .
Move towards the included values, which is downwards here.
- 3
Identify. and
σ = √λ.
- 4
Standardise.
The boundary is below the mean, so z is negative.
- 5
Symmetry.
Tables give only positive z.
- 6
Evaluate. , so
Round at the end: 0.192.
Answer(3 s.f.).
- 1
- 35 marks
Minor faults occur in a production process at a constant average rate of per hour. Use a Poisson model, scaled to a suitable interval, together with a normal approximation, to find the probability that more than faults occur during a -hour period.
Stuck? Show hint
Scale λ = 3 × 24 first, then use N(λ, λ) with a continuity correction on "more than 80".
Show solution
- 1
Scale. faults in hours.
The rate is per hour and the question asks about 24 hours.
- 2
Check and state. , so .
λ is large; the mean and variance are both 72.
- 3
Continuity correction. "More than " means , so use .
Move towards the included values, which is upwards here.
- 4
Identify. and
σ = √72, not 72.
- 5
Standardise.
Keep z to 4 d.p.
- 6
Area.
Φ gives the area to the left, so the right-hand tail is 1 − Φ.
Answer(3 s.f.).
- 1
- 49709/61 M/J 2022 Q5(b)4 marks
Cars arrive at a fuel station at random and at a constant average rate of per hour. Use an approximating distribution to find the probability that the number of cars that arrive during a -hour period is between and inclusive.
Stuck? Show hint
λ = 13.5 × 12 = 162. Inclusive range: boundaries 149.5 and 160.5. Both are below the mean.
Show solution
- 1
Scale and state. , so .
B1 for 162 and the normal model.
- 2
Continuity corrections. Inclusive at both ends: and .
Move outwards to keep 150 and 160.
- 3
Identify. and
σ = √λ, not λ.
- 4
Standardise both.
Both boundaries are below the mean, so both z's are negative.
- 5
Area between.
The 1's cancel. This is the form the mark scheme shows.
- 6
Evaluate.
Φ(0.982) = 0.8370 and Φ(0.118) = 0.5470 from the tables.
Answer(3 s.f.).
- 1
- 59709/62 M/J 2024 Q15 marks
A random variable has the distribution .
(a) Use a suitable approximating distribution to calculate .
(b) Justify the use of your approximating distribution in this case.
Stuck? Show hint
N(145, 145). 'At most 150' keeps 150, so use 150.5. For (b), quote the number 145.
Show solution
- 1
(a) State the approximation. .
B1, stated or implied.
- 2
(a) Continuity correction. keeps , whose bar ends at .
Move towards the values being kept: upwards.
- 3
(a) Standardise.
σ = √145.
- 4
(a) Area.
'At most' is the left-hand area, which is what Φ gives directly.
- 5
(b) Justify. .
The mark scheme says 'λ > 15' scores B0 if 145 is not stated. The value must be there.
Answer(a) (3 s.f.). (b) .
- 1
Choosing the right approximation
You've now met three approximations, two in this topic and one from Paper 5. Questions often just say "use a suitable approximating distribution", so you have to pick the right one and say why, with numbers. Start from what the exact distribution is, then check the conditions:
- Binomial, and : use . Both are discrete, so there's no continuity correction.
- Binomial, and : use from §5.5, with a continuity correction.
- Poisson, : use , with a continuity correction.
- None of these: use the exact distribution.
If is close to (say ), count the failures instead. The number of failures is , and that can be approximated by a Poisson.
You have | Conditions (quote the values) | Use | Continuity correction? |
|---|---|---|---|
small enough to work out directly | the binomial formula exactly | No | |
and | No (both discrete) | ||
and | (§5.5) | Yes | |
Yes |
Discrete to discrete needs no continuity correction; discrete to continuous always does.
Which model? A binomial with n large and p small becomes Poisson; a Poisson with λ large becomes normal, always with a continuity correction and σ = √λ.
Your turn: choosing an approximation
- 18 marks
For each random variable, state a suitable approximating distribution, giving the values of its parameters, and justify your choice with values.
(a) (b) (c) (d)
Stuck? Show hint
For each binomial, work out np first. Below 5 (with n > 50) points to Poisson; above 5 (with nq > 5 too) points to normal. For the Poisson, compare λ with 15.
Show solution
- 1
(a) . Since and , use .
Large n and a small expected count point to Poisson.
- 2
(b) and . Both are more than , so use , since .
np = 32 is far too big for Poisson. The normal variance is npq, not np.
- 3
(c) , so use .
A Poisson with a large mean goes to the normal, with mean and variance both λ.
- 4
(d) . Since and , use .
Quote both numbers, 500 and 1, in the justification.
Answer(a) (b) (c) (d) , each justified with the values shown.
- 1
Everything on one page
Poisson Po(λ)
λ from the probability of no events
Neighbouring probabilities (rising while r + 1 < λ)
Mean, variance and standard deviation of the Poisson
Scaling λ to a different interval
Poisson approximation to the binomial (guideline n > 50, np < 5)
Normal approximation to the Poisson (guideline λ > 15)
Getting each term from the one before
Continuity correction: move half a unit towards the values included
Can you do all of these?
State Poisson conditions in context (random, independently, singly, constant mean rate), always naming what is counted; 'constant' alone must say mean or rate
A Poisson model fails if the rate changes (day/night, a trend), events come in groups or affect each other, the variable can be negative or non-integer, or its mean and variance differ
Standard deviation of Po(λ) is √λ; Y = 2X is not Poisson because Var(Y) = 4λ ≠ 2λ = E(Y)
Scale λ to the exact interval named in the question before substituting into any formula
'At least' and 'more than' have no top end, so always use the complement, never a direct sum
0! = 1 and , so . You'll need it all the time
When equating two Poisson probabilities algebraically, write both out in full before cancelling anything
If the leftover terms include λ², the equation is quadratic, not linear. A pure λ² = k just needs square-rooting; otherwise factorise or use the quadratic formula, and always reject the negative root
Given P(X=0) = c, solve with ln: λ = −ln c
P(X=r) < P(X=r+1) simplifies to r + 1 < λ; the most likely value is the whole number just below λ (λ − 1 and λ tie if λ is whole)
'At least one' in a minimum time → round UP; 'no events' in a largest period → round DOWN. Check both neighbouring values
Combining: 'and' → multiply; one outcome in each of two periods either way → × 2; 'each of k' → power k; 'exactly j of m' → binomial with p = Poisson probability; 'given' → overlap ÷ condition
Poisson approximation to the binomial: justify with values (n > 50 and np < 5, or p < 0.1), since 'n large, p small' alone scores nothing; then use λ = np
Normal approximation to the Poisson needs λ large; state N(λ, λ) with both parameters, justify with the value (e.g. 145 > 15), and standardise with σ = √λ, never σ = λ
Two-sided ranges: correct both ends (inclusive moves outwards, strict moves inwards), standardise each, then Φ(z₂) − Φ(z₁)
Any normal approximation to a discrete distribution needs a continuity correction: move the boundary half a unit towards the included values
If λ itself is unknown, put u = √λ before clearing the fraction, so the standardising equation becomes a quadratic in u
Choosing an approximation: B(n, p) with n > 50, np < 5 → Po(np); np > 5 and nq > 5 → N(np, npq); Po(λ) with λ > 15 → N(λ, λ). Always quote the values