Mathematics 9709/65 — October/November 2025
Cambridge A-Level · Paper 6 Probability & Statistics 2 · worked solutions for every part, with the mark scheme
Topics Hypothesis Tests · The Poisson Distribution · Sampling and Estimation · Linear Combinations of Random Variables · Continuous Random Variables
A student notes the length, minutes, of certain lectures. The results for a random sample of 80 lectures are summarised as follows.
Approach
Since is large, the Central Limit Theorem ensures the sample mean is approximately normally distributed. We estimate the population mean and variance from the sample, then form a two-sided 96% confidence interval using the critical value .
Working
The sample mean is
The unbiased estimate of the population variance is
The standard error of the sample mean is
For a 96% confidence interval, the total tail probability is 4%, so each tail is 2%, giving .
The confidence interval is
Answer
The 96% confidence interval for the mean lecture length is to minutes (3 s.f.).
79.0 to 81.8 minutes (3 s.f.)
Walkthrough
First, calculate the sample mean ; this is the point estimate of the population mean.
Next, calculate the unbiased estimate of the population variance. Because the sample variance with denominator is biased, multiply by , i.e. use
This gives for the variance estimate.
Then identify the critical value for a 96% confidence interval. Since the total tail probability is 4%, and there are two tails, each tail is 2%, so the required -value is .
Finally, substitute into the confidence interval formula
using , giving the interval to minutes. Each of these steps corresponds to a mark: the estimate of the mean (B1), the unbiased variance calculation (M1 A1), the correct (B1), the interval formula (M1), and the final interval (A1).
Key Takeaways
This question tests the large-sample confidence interval for a population mean. The key ideas are the Central Limit Theorem, the use of an unbiased variance estimate , and the correct critical value for the confidence level.
Common Mistakes
- Using the biased variance instead of multiplying by . The mark scheme awards M0 A0 if a biased variance is used.
- Using for a 96% interval; the correct value is about .
- Forgetting to divide the variance by before taking the square root.
- Quoting a point instead of an interval for the final answer.
Things to Be Careful About
- Keep enough decimal places in intermediate steps; rounding too early can change the final interval.
- The mark scheme accepts or .
- Give the final interval to 3 significant figures: to .
The method used in part (a) is valid because the sample size was large. Explain why the method would not be valid if the sample size were small.
Approach
The method in part (a) uses the Central Limit Theorem to justify that the sample mean is approximately normal.
Working
For a small sample size, the Central Limit Theorem does not apply. Without knowing the population distribution, the distribution of the sample mean is not necessarily normal, so the -based confidence interval cannot be justified.
Answer
It would not be valid because the Central Limit Theorem only applies to large samples; for a small sample the sample mean may not be approximately normally distributed.
Because the Central Limit Theorem does not apply to small samples, so the sample mean may not be approximately normal.
Walkthrough
The confidence interval in part (a) is based on the fact that, for a large sample, the sample mean is approximately normally distributed by the Central Limit Theorem. If the sample is small, this theorem cannot be used, so we cannot assume normality of unless the population itself is normal. Since the question does not state that the lecture lengths are normally distributed, the interval would not be valid.
Key Takeaways
The Central Limit Theorem is the justification for using normal-distribution confidence intervals when the population distribution is unknown but the sample size is large.
Common Mistakes
- Stating only that the sample mean is not normal without mentioning the Central Limit Theorem.
- Saying the method is invalid because the variance is unknown; the issue is the normality of the sampling distribution, not the variance.
Things to Be Careful About
- A small sample could still be valid if the population is normal, but the data here give no such information, so the method cannot be assumed valid.
A researcher is investigating whether the proportion of families who do not own a car in his town is different from the proportion of the population in the whole country, which is 10.1%. He takes a large random sample of families in his town and finds the proportion of families that do not own a car.
Approach
A two-tailed test is used when the alternative hypothesis does not specify a direction. Here the researcher is only investigating whether the town proportion is different from the national proportion, not whether it is greater or less.
Working
The claim under investigation is that the proportion in the town is different from the national proportion. This could mean the town proportion is greater than 10.1% or less than 10.1%. Therefore the rejection region must be split between the two tails of the distribution.
Answer
A two-tailed test is appropriate because the researcher is looking for evidence that the proportion is different from 10.1%, not specifically greater than or less than 10.1%.
A two-tailed test is appropriate because the researcher is looking for evidence that the proportion is different from 10.1%, not specifically greater or less.
Walkthrough
The key word is "different from". It has no direction. If the researcher suspected "greater than", a one-tailed test would be appropriate; if the researcher suspected "less than", a one-tailed test in the other direction would be appropriate. Since the researcher is only asking whether there is a difference, both directions must be considered, so the test is two-tailed.
Key Takeaways
A two-tailed test is used when the alternative hypothesis is of the form "not equal to". The significance level is split between the two tails of the distribution.
Common Mistakes
- Using a one-tailed test when the question says "different from".
- Thinking that "different" implies greater, or less, rather than either.
Things to Be Careful About
The explanation should make it clear that the choice is based on the absence of a direction in the research question, not on the size of the sample or the value of the test statistic.
Approach
Let be the proportion of families in the town who do not own a car. The null hypothesis states that this is equal to the national proportion. The alternative hypothesis states that it is different.
Working
where is the proportion of families in the town who do not own a car.
Answer
; , where is the proportion of families in the town who do not own a car.
H0: p = 0.101; H1: p is not equal to 0.101, where p is the proportion of families in the town who do not own a car.
Walkthrough
We need to state the hypotheses in terms of the population proportion, not the sample proportion. Let represent the proportion of families in the town who do not own a car. The null hypothesis is the claim that there is no difference from the national proportion, so . The alternative hypothesis is the claim that the proportion is different, so .
Key Takeaways
Hypotheses should always be written in terms of the population parameter, using a decimal rather than a percentage. The null hypothesis always contains equality, and a two-sided alternative uses .
Common Mistakes
- Defining as the sample proportion instead of the population proportion.
- Writing 10.1 instead of 0.101.
- Using a one-sided alternative such as or .
Things to Be Careful About
The parameter should be defined clearly. If is used in the hypotheses, the conclusion must also use the same definition.
The researcher calculates the value of the test statistic and finds that . He carries out the test at the 5% significance level.
State the conclusion of the test, explaining your answer.
Approach
For a two-tailed test at the 5% significance level, the critical value is . Compare the observed value with this critical value, then state the conclusion in context.
Working
The test statistic is not in the critical region. Equivalently, the one-tail probability is .
Answer
Do not reject . There is insufficient evidence that the proportion of families in the town who do not own a car is different from 10.1%.
Do not reject H0; insufficient evidence that the town proportion differs from 10.1%.
Walkthrough
At the 5% significance level, a two-tailed test has critical values . The test statistic is . Since , the test statistic lies inside the acceptance region, so the result is not significant.
Equivalently, the probability in one tail beyond is approximately . For a two-tailed 5% test, each tail has area . Since , the result is not significant.
Therefore we do not reject . This means there is not enough evidence to say that the town proportion is different from the national proportion of 10.1%.
Key Takeaways
- For a two-tailed 5% test, the critical value is .
- Compare the test statistic with the critical value to decide whether it lies in the critical region.
- The conclusion must be stated in context and should say "insufficient evidence", not "the proportion is equal to 10.1%" as a proven fact.
Common Mistakes
- Using 1.645, which is the critical value for a one-tailed 5% test.
- Comparing 0.0344 with 0.05 instead of 0.025 in a two-tailed test.
- Saying "accept " instead of "do not reject ".
- Giving a conclusion without context.
Things to Be Careful About
At the 5% two-tailed level, the critical value is 1.96, not 1.645. The conclusion should not overstate the result: there is insufficient evidence of a difference, not proof that the proportion is exactly 10.1%.
A certain website receives an average of hits per hour. In the past the value of was 14.4. After making some improvements, the owner of the website wishes to test whether the value of has increased. He chooses a 10-minute period at random and finds that there were 6 hits during this period. You may assume that the number of hits the website receives in any given time period follows a Poisson distribution.
Approach
Set up the null and alternative hypotheses for the Poisson mean. Convert the hourly mean to the 10-minute test period, compute the p-value using the Poisson distribution, compare it with the 2.5% significance level, and draw a conclusion in context.
Working
The mean rate is 14.4 hits per hour. Since the test period is 10 minutes, which is of an hour, the mean number of hits in the test period is:
State the hypotheses:
This is a one-tailed test at the 2.5% significance level. The observed value is 6 hits, so the p-value is:
Using the Poisson distribution with :
Therefore:
Compare with the significance level:
Since the p-value is greater than the significance level, we do not reject .
Answer
There is insufficient evidence that the mean number of hits per hour has increased.
Do not reject the null hypothesis. There is insufficient evidence that the mean number of hits per hour has increased.
Walkthrough
The first step is to recognise that the mean is given per hour (14.4 hits per hour) but the observation is over a 10-minute period. Since 10 minutes is one-sixth of an hour, the expected number of hits in the test period is . This conversion is essential — using 14.4 directly would give a completely wrong p-value.
Next, set up the hypotheses. The null hypothesis states that the mean has not changed: (equivalently per hour). The alternative hypothesis states that the mean has increased: . This is a one-tailed test because the owner only wants to detect an increase, not a decrease.
The observed value is 6 hits. Since the alternative is , the p-value is the probability of observing 6 or more hits: . We compute this as because it is easier to sum the probabilities from 0 to 5.
The Poisson probability formula gives . Summing from to 5 and subtracting from 1 gives the p-value .
Finally, compare the p-value with the significance level 0.025. Since , the result is not statistically significant, so we do not reject . The conclusion must be stated in context: there is insufficient evidence that the mean number of hits per hour has increased.
Key Takeaways
- When the mean is given for one time period but the observation is over a different period, convert the mean proportionally.
- For a one-tailed test with , the p-value is .
- The p-value is compared with the significance level: reject if the p-value is less than the significance level.
- The conclusion must be phrased in the context of the problem.
Common Mistakes
- Using instead of for the 10-minute period.
- Computing instead of — the wrong tail.
- Not showing the full Poisson expression; the mark scheme requires the expression to be seen for the method mark. An unjustified answer of 0.0357 scores only B1M0B1.
- Stating the conclusion without context (e.g., just "do not reject " without mentioning the mean/hits).
Things to Be Careful About
- The significance level is 2.5%, which is 0.025.
- The comparison must be valid: , so we do not reject .
- The final answer should be given to 3 significant figures (0.0357).
Explain whether it is possible that a Type I error or a Type II error or both may have been made in carrying out the test.
Approach
Since was not rejected in part (a), determine which of the two types of error could possibly have been made.
Working
In part (a) we did not reject .
A Type I error is rejecting when it is actually true. Since we did not reject , a Type I error cannot have been made.
A Type II error is failing to reject when it is actually false. Since we did not reject , and it is possible that is false (the mean may have increased without us detecting it), a Type II error may have been made.
Since a Type I error requires rejecting and a Type II error requires not rejecting , both cannot have been made in this single test.
Answer
A Type II error may have been made, but a Type I error cannot have been made. Both cannot have been made.
A Type II error may have been made; a Type I error cannot have been made. Both cannot have been made.
Walkthrough
In part (a) we did not reject . Now we need to think about which errors are possible given this decision.
A Type I error is defined as rejecting when it is actually true. Since we did not reject , a Type I error cannot have been made — we never made the decision that would constitute a Type I error.
A Type II error is defined as failing to reject when it is actually false. Since we did not reject , and it is entirely possible that is false (the mean may have genuinely increased, but the evidence was not strong enough to detect it), a Type II error may have been made.
Because a Type I error requires the decision "reject " and a Type II error requires the decision "do not reject ", the two errors are mutually exclusive in a single test — both cannot have been made.
Key Takeaways
- Type I error: rejecting a true null hypothesis.
- Type II error: failing to reject a false null hypothesis.
- The two errors are mutually exclusive in a single test outcome.
Common Mistakes
- Confusing the two error types.
- Claiming both errors may have been made — impossible since they require opposite decisions.
- Not linking the reasoning to the fact that was not rejected.
Things to Be Careful About
- The mark scheme requires stating that was not rejected (at least once) to justify the reasoning.
- The answer must be in context — mention the mean/hits where appropriate.
A sports fan produces a magazine each month.
On average 1 in 540 characters in the magazine is incorrect.
Use an appropriate approximating distribution to find the probability that, in a magazine containing 2430 characters, there are at least 4 incorrect characters.
Approach
The number of incorrect characters follows a binomial distribution . Since is large and is small, we use the Poisson approximation with .
Working
So . We require :
Answer
0.658
Walkthrough
The number of incorrect characters is a binomial random variable, since each of the 2430 characters is independently either correct or incorrect, with probability of being incorrect. Computing directly from the binomial formula would be extremely tedious, so we use the Poisson approximation, which is valid because is large and is small. The parameter is , so we model .
To find , we avoid summing infinitely many terms by using the complement: . Each term is computed for and summed. Finally we subtract this sum from 1, giving to 3 significant figures.
Key Takeaways
This question tests the Poisson approximation to the binomial: when is large and is small. It also tests computing cumulative Poisson probabilities and using the complement rule for "at least" probabilities.
Common Mistakes
- Computing but forgetting to subtract from to obtain .
- Misplacing the factorials: and .
- Trying to sum infinitely many Poisson terms instead of using the complement.
- The mark scheme warns that an unjustified answer of scores only B1M0B1 — the full probability expression must be shown.
Things to Be Careful About
- Use exactly , obtained from .
- Show the full expression to earn the M1 mark.
- Using directly with the same answer scores B2, but the Poisson route is the intended one.
- Round to 3 significant figures.
Approach
State the conditions that make the binomial-to-Poisson approximation valid, giving the specific values from the question.
Working
For approximated by , we need:
and
Equivalently, . Both conditions hold, so the Poisson approximation is valid.
Answer
and .
n = 2430 > 50 and np = 4.5 < 5
Walkthrough
The Poisson approximation of a binomial distribution is valid only when certain conditions are met. We must state them explicitly with the values from the problem: the number of trials is greater than 50, and is less than 5 (equivalently ). Both conditions must be quoted for the mark.
Key Takeaways
The formal conditions for the Poisson approximation to the binomial: and (or ).
Common Mistakes
- Saying only " is large and is small" — the mark scheme states this is insufficient; specific numerical conditions must be given.
- Quoting only one of the two conditions.
Things to Be Careful About
- Both and (or ) must be stated explicitly for the full mark.
On average the number of copies, , of the magazine sold per month is 123.4.
Approach
Recall a condition under which a random variable is well modelled by a Poisson distribution, stated in the context of the problem.
Working
The number of copies sold per month may follow a Poisson distribution if, for example, the mean number of copies sold per month is constant (at 123.4), or if sales occur randomly and independently of one another.
Answer
The mean number of copies sold per month is constant (or: sales occur independently of each other).
The mean number of copies sold per month is constant.
Walkthrough
A Poisson distribution models the number of events in a fixed interval when events occur at a constant average rate, independently of one another, and singly. For copies sold per month, we need to state one such condition in context, e.g. the mean number of copies sold per month is constant, or each sale is independent of the others.
Key Takeaways
The defining assumptions of the Poisson model: constant mean rate, events occurring independently, and events occurring singly.
Common Mistakes
- Not phrasing the condition in context (e.g., just writing " is constant" without mentioning copies or months).
- Giving a non-numerical or irrelevant condition.
Things to Be Careful About
- The mark scheme awards B1 for any of the following stated in context: mean is constant; sales occur singly, randomly, or independently.
You are now given that has a Poisson distribution.
Use an appropriate approximating distribution to find the probability that in a randomly chosen month, more than 130 copies of the magazine are sold.
Approach
Since has , we approximate by a normal distribution with . We then apply a continuity correction for the strict inequality and standardise.
Working
"More than 130 copies" means , so with the continuity correction we standardise :
From normal tables, , so
Answer
0.261
Walkthrough
We are told is Poisson with mean 123.4. When is large (here ), the Poisson distribution is closely approximated by a normal distribution having the same mean and variance, so . Because is discrete, the event "more than 130 copies" means ; the continuity correction moves the boundary by 0.5 to 130.5 to bridge the discrete and continuous scales. We then standardise: . Since we need the upper tail , we read from tables and subtract from 1, giving .
Key Takeaways
The normal approximation to the Poisson, valid when is large: since mean = variance = . The continuity correction for a discrete-to-continuous approximation. Standardising: and reading upper-tail probabilities from normal tables.
Common Mistakes
- Forgetting the continuity correction and using 130 instead of 130.5. The mark scheme condones wrong or missing cc for M1 only, but you then lose the A1.
- Using the wrong variance: for a Poisson variable the variance equals .
- Reporting directly instead of for the "more than" tail.
Things to Be Careful About
- State explicitly — both B1 marks in the mark scheme depend on it (one for the mean, one for the variance).
- Use 130.5, not 130, for .
- Make sure your standardising uses the same values you quoted: .
- Round the final answer to 3 significant figures: 0.261.
The masses, in kg, of large and small bags of tomatoes have the distributions and respectively.
Approach
Define the difference . Since and are independent normal variables, is also normal. Combine the means by subtraction and the variances by addition, then standardise and use the symmetry of the normal distribution.
Working
Given
Let . Then
The required probability is
Standardise:
Using symmetry,
From normal tables,
Answer
0.578
Walkthrough
We are asked for the probability that a randomly chosen large bag is more than 0.5 kg heavier than a randomly chosen small bag. Because both masses are normally distributed and the bags are chosen independently, the difference is also normal.
The mean of a difference is the difference of the means:
A common trap is to subtract variances, but for independent variables the variance of a difference is the sum of the variances:
The event is the same as . We standardise by subtracting its mean and dividing by its standard deviation, giving a negative -score. Because the normal distribution is symmetric, , which we read directly from the normal table.
Key Takeaways
- A linear combination of independent normal variables is normal.
- For a difference of independent variables, means subtract but variances add.
- Standardising converts any normal probability into a standard normal probability.
- Symmetry lets us handle negative -scores.
Common Mistakes
- Subtracting variances instead of adding them.
- Using instead of .
- Forgetting to divide by the standard deviation when standardising.
- Looking up directly and giving the wrong tail.
Things to Be Careful About
- The inequality must be rearranged to before standardising.
- Use the unrounded mean and variance when calculating .
- Give the final probability to 3 significant figures as required.
The price of tomatoes is $4.30 per kg.
A large bag of tomatoes and a small bag of tomatoes are chosen at random. Find the probability that the total price of the tomatoes in the two bags is less than $16.
Approach
Let be the total mass. Since and are independent normal variables, is normal. The total price is , which is also normal. Find the mean and variance of , then standardise to find .
Working
Given
The total mass is
The total price is , so
and
Therefore
Standardise:
So
Answer
The probability that the total price is less than $16 is
0.596
Walkthrough
The total price is found by multiplying the total mass by $4.30 per kg. First combine the two masses into a total mass . Since and are independent normal variables, is normal with
Multiplying a normal variable by a constant keeps it normal. The mean is multiplied by the constant, but the variance is multiplied by the square of the constant:
So . We need . Standardise:
Then .
An equivalent method is to convert the price limit into a mass limit: the total price is less than $16 exactly when . Standardising this with gives the same .
Key Takeaways
- The sum of independent normal variables is normal.
- Scaling a normal variable by a constant preserves normality.
- , not .
- Probability questions can often be solved in original units or scaled units.
Common Mistakes
- Using instead of for the variance.
- Using instead of .
- Forgetting that the total price includes both bags.
- Using the wrong tail: the question asks for less than $16.
- Rounding the mean or variance before standardising.
Things to Be Careful About
- The price per kg is $4.30, so the total price is times the total mass in kg.
- If using the alternative method, divide $16 by $4.30 to get the critical mass, not multiply.
- Strict inequality does not matter for a continuous normal distribution.
- Give the final answer to 3 significant figures.
The diagram shows the graph of the probability density function, , of a random variable . The graph is a straight line from to where and are constants. Elsewhere .
Approach
The probability density function forms a right-angled triangle with the coordinate axes. Since is a valid PDF, the total area under the curve must equal 1. Use the formula for the area of a triangle to find in terms of .
Working
The graph of is a straight line from to , forming a right-angled triangle with vertices at , , and .
The area of this triangle is:
Since is a probability density function, the total area must equal 1:
Solving for :
Answer
b = 2/a
Walkthrough
A probability density function must satisfy the property that the total area under the curve equals 1. The graph shows as a straight line from on the y-axis to on the x-axis, with elsewhere. This forms a right-angled triangle with the axes, having base and height . The area of this triangle is . Setting this equal to 1 gives , which rearranges to .
Key Takeaways
- The total area under any valid PDF equals 1.
- For simple geometric shapes (triangles, rectangles), the area can be found using basic geometry rather than integration.
- This property is used to find unknown constants in a PDF definition.
Common Mistakes
- Forgetting the factor in the triangle area formula.
- Using the wrong base or height for the triangle.
- Not recognising that outside , so only the triangular region contributes to the area.
Things to Be Careful About
- Ensure the area calculation covers only the region where , which is .
- The constants and must both be positive for a valid PDF.
Approach
First, find the equation of the straight line forming the PDF. Substitute from part (a) to express purely in terms of . Then use the definition to set up an equation in , and solve it using the given value .
Working
The straight line passes through and . Its equation is:
Substituting :
The expected value is:
Integrating:
Evaluating at the upper limit :
Simplifying:
Given :
Cross-multiplying:
Answer
a = 3/2
Walkthrough
To find , we first need the explicit form of . The line from to has gradient , so the equation is . From part (a), , so . Thus for .
The expected value formula is . Substituting and the upper limit :
At : and . The difference is .
Setting and solving gives .
Key Takeaways
- The expected value of a continuous random variable is .
- When the PDF contains unknown parameters, you may need to use multiple properties (area = 1, given mean) to solve for them.
- Substitution of results from earlier parts is essential.
Common Mistakes
- Forgetting to multiply by when computing — integrating instead of gives , not the mean.
- Algebra errors when substituting into .
- Incorrect integration of terms.
Things to Be Careful About
- The upper limit of integration is , not a fixed number.
- Ensure all terms are simplified correctly when combining fractions with in the denominator.
Approach
With found in part (b), determine the full PDF and compute . Set up the equation , solve the resulting quadratic for , and reject any value outside the domain .
Working
From part (b), and .
The PDF is:
We need :
Integrating:
Setting equal to :
Multiplying through by 16:
Rearranging:
Dividing by 3:
Factoring:
So or .
Since must lie within the domain of the PDF, . The value exceeds and is therefore rejected.
Thus .
Answer
k = 2/3
Walkthrough
With and , the PDF is fully determined: for .
The probability is the area under the PDF from 0 to . Setting up the integral:
Setting this equal to and multiplying by 16 to clear denominators:
Factoring gives , so or .
Since the PDF is only defined for , we must have . The value is outside this range, so we reject it. The answer is .
Key Takeaways
- for a PDF defined on .
- When solving for a probability boundary, always check that the solution lies within the valid domain of the PDF.
- Quadratic equations from probability problems often have two roots, but only one may be physically meaningful.
Common Mistakes
- Forgetting to check whether the solutions for lie within the domain .
- Accepting without verification, even though for .
- Algebra errors when clearing denominators or factoring the quadratic.
Things to Be Careful About
- The domain restriction is crucial. Without it, both roots would appear valid.
- Ensure the integral limits are correct: from 0 to , not from to .
- The value is indeed less than , confirming it is valid.
