Mathematics 9709/63 — October/November 2025
Cambridge A-Level · Probability & Statistics 2 · worked solutions for every part, with the mark scheme
Topics The Poisson Distribution · Linear Combinations of Random Variables · Sampling and Estimation · Hypothesis Tests · Continuous Random Variables
The random variables and have independent distributions and respectively.
Approach
Since is a Poisson random variable, it takes integer values only. The inequality therefore includes and . Add the two Poisson probabilities.
Working
Answer
0.392
Walkthrough
We are told that follows a Poisson distribution with mean . The question asks for the probability that is strictly between 2 and 5. Since is discrete, the only whole-number values satisfying are and .
For a Poisson distribution,
Substitute , and , then add the two probabilities.
Key Takeaways
- Poisson probabilities are calculated using .
- Strict inequalities must be interpreted using integer values because the Poisson distribution is discrete.
Common Mistakes
- Including or : the inequality is strict, so only 3 and 4 are included.
- Using the wrong value of in the formula.
Things to Be Careful About
- Give the final answer to 3 significant figures.
- Do not round intermediate values too early; keep enough decimal places until the end.
Approach
Since and are independent Poisson variables, their sum is also Poisson with mean . Use the complement rule: .
Working
Answer
0.875
Walkthrough
Because and are independent Poisson random variables, the sum is also Poisson. Its mean is the sum of the means: .
The event is the complement of . It is easier to calculate the probability of 0, 1, or 2 and subtract from 1.
For ,
So add the probabilities for , and , then subtract the total from 1.
Key Takeaways
- The sum of independent Poisson random variables is Poisson with mean equal to the sum of the means.
- Using the complement rule simplifies probabilities of the form .
Common Mistakes
- Forgetting that is Poisson with mean 5, and instead using the original means separately.
- Calculating as instead of subtracting all values up to 2.
Things to Be Careful About
- The inequality is strict, so must be included in the subtracted probability.
- Keep sufficient decimal places before giving the final answer to 3 significant figures.
The total of 100 random values of and 150 random values of is denoted by .
Use a suitable approximating distribution to find .
Approach
The sum of 100 values of is , and the sum of 150 values of is also . Therefore . Since is large, approximate by a normal distribution with mean 600 and variance 600. Use a continuity correction for .
Working
For a discrete variable, . With the continuity correction, use 559.5:
Answer
(Using tables, 0.0491 is also accepted.)
0.0492 (or 0.0491)
Walkthrough
First find the distribution of . Each of the 100 values of is , so their sum is . Each of the 150 values of is , so their sum is also . Adding these gives .
Because is large, the Poisson distribution can be approximated by a normal distribution with the same mean and variance: .
Since is discrete, the event means can be at most 559. The continuity correction uses the boundary 559.5. Standardise:
Then use the symmetry of the normal distribution to write as and read the probability from the normal table.
Key Takeaways
- The sum of independent Poisson variables is Poisson, with mean equal to the sum of the means.
- A Poisson distribution with a large mean can be approximated by a normal distribution with the same mean and variance.
- A continuity correction is needed when approximating a discrete distribution by a continuous one.
Common Mistakes
- Forgetting the continuity correction, or using 560 instead of 559.5.
- Using the standard deviation instead of the variance when writing .
- Forgetting to subtract the table probability from 1 when the -value is negative.
Things to Be Careful About
- The variance of is 600, so the standard deviation is , not 600.
- The continuity correction must move the boundary in the correct direction: becomes , so use 559.5.
- The final answer may be 0.0492 or 0.0491 depending on the normal table used; both are accepted.
The mean mass of packets of Trueleaf tea is supposed to be 500 grams. An inspector wishes to test whether this value is correct. He weighs 60 randomly chosen packets and notes the mass, grams, of each packet. The results are summarised as follows.
Test, at the 5% significance level, whether the population mean mass is 500 grams.
Approach
We want to test whether the true mean mass is 500 grams. Because the sample is large and the population variance is unknown, we first estimate the population variance from the sample, then use the Central Limit Theorem to standardise the sample mean as an approximately normal statistic.
Working
The sample mean is
The unbiased estimate of the population variance is
Substituting:
The hypotheses are
Under , the standardised sample mean is
For a two-tailed test at the 5% significance level, the critical values are . Since
we do not reject . Equivalently, the tail probability beyond is approximately , which is greater than , so the result is not significant.
Answer
There is insufficient evidence at the 5% significance level to conclude that the mean mass of packets of Trueleaf tea is not 500 grams.
There is insufficient evidence at the 5% significance level that the mean mass of packets of Trueleaf tea is not 500 grams.
Walkthrough
This is a hypothesis test about a population mean, using a large sample. The first step is to find the sample mean from the given sums:
This is the point estimate of the population mean. Next we need an estimate of the variability of the sample mean. The unbiased estimate of the population variance is
Using rather than is essential here because we are estimating a population variance from sample data.
The hypotheses are
The alternative is two-tailed because the inspector is testing whether the mean is correct, not whether it is specifically too high or too low.
Because is large, the Central Limit Theorem tells us that the sample mean is approximately normally distributed. So, assuming is true,
For a 5% two-tailed test, the critical region is or . Since lies between the critical values, we do not reject the null hypothesis. Equivalently, the tail probability is larger than .
Key Takeaways
- For a large sample from a distribution with unknown variance, the sample mean is approximately normal by the Central Limit Theorem.
- An unbiased estimate of the population variance uses denominator .
- A two-tailed test at the 5% level uses tail area in each tail and critical value .
- A non-significant result means there is insufficient evidence against , not that is proven true.
- Summary statistics and are enough to compute both the sample mean and the estimated variance.
Common Mistakes
- Dividing by rather than when estimating the variance. The mark scheme specifically does not allow the biased variance ; it should be .
- Using a one-tailed critical value such as when the alternative hypothesis is .
- Forgetting to divide by when constructing the standard error.
- Writing a conclusion as though the mean is 500 grams; instead, say there is insufficient evidence that it is not 500 grams.
- Not giving the hypotheses with a symbol such as for the population mean.
Things to Be Careful About
- The test statistic is negative because , but the comparison uses distance, so use or check both tails.
- The 5% significance level is split between two tails, so each tail is 2.5%, giving .
- If using the p-value approach, use for one tail, or double it and compare with .
- The final conclusion must be in context: packets of Trueleaf tea and mass in grams.
- Avoid contradictory statements such as "reject" followed by "there is no evidence"; the conclusion must be consistent with the comparison.
The data produced by a certain data entry firm always include a small number of incorrect characters that occur at random. The proportion of incorrect characters is denoted by , and experience has shown that . A particular data set from the firm contains 14500 characters, of which characters are incorrect.
Approach
The random variable counts incorrect characters in 14500 independent trials, so exactly . Because is large and is small, use the Poisson approximation with
Then
Working
The Poisson distribution is , with
So
Answer
0.940
Walkthrough
This is a binomial situation because each character in the data set is an independent trial with probability of being incorrect. With , the exact distribution is , but the question asks for a suitable approximating distribution. Since is large and is very small, the Poisson approximation applies, with mean .
In a Poisson distribution, . The event means or . Substituting and adding the four probabilities gives to 3 significant figures.
Key Takeaways
- A binomial distribution with large and small can be approximated by a Poisson distribution with .
- For cumulative probabilities, sum the individual Poisson probabilities over the required values.
- The smaller the value of and the larger the value of , the better the Poisson approximation becomes.
Common Mistakes
- Not stating that a Poisson approximation is being used; the mark scheme requires an indication of Poisson.
- Using as instead of stopping at .
- Forgetting the factorial in the denominator when writing .
- Trying to use a normal approximation with a continuity correction here; this is a Poisson situation.
Things to Be Careful About
- is strict, so exclude .
- The value must come from .
- Rounding: the final answer may be given as or .
The firm’s management wishes to decrease the value of by giving their employees some training. Their aim is that, for a data set containing 14500 characters, the value of for the new value of should be double the value of when .
Use a suitable approximating distribution to find the new value of .
Approach
For a new proportion , the Poisson approximation has mean , so the probability of no incorrect characters is . The original mean is , so the original probability of no incorrect characters is . Setting the new probability equal to double this gives an equation in ; solve it by taking natural logarithms.
Working
Original:
New:
The requirement is:
Taking natural logarithms of both sides:
Therefore
Using ,
Answer
0.0000522
Walkthrough
From part (a), the original Poisson parameter is , so the original probability of no incorrect characters is .
If the new proportion of incorrect characters is , then the same Poisson approximation gives a new mean of , so the new probability of no incorrect characters is
The aim is that this new probability should be double the original probability, giving
Taking natural logarithms uses and , so
Rearranging gives
Since , this gives .
Key Takeaways
- The Poisson parameter for is always , so when changes, the mean changes proportionally.
- The probability of zero events in a Poisson distribution is simply .
- Exponential equations of the form can be solved by taking natural logarithms of both sides.
Common Mistakes
- Using the old parameter for the new situation instead of .
- Writing incorrectly as ; the correct value is .
- Setting the factor of 2 on the wrong side of the equation, which gives an incorrect value of .
- Forgetting to divide by 14500 after taking logarithms.
Things to Be Careful About
- When rearranging, note that gives . Be careful with signs.
- The final answer must be given to 3 significant figures: .
- The numerical size of should be a small proportion, so a result near is sensible.
The masses of a certain species of animal are known to be normally distributed with standard deviation kg. A researcher obtains the masses of a random sample of animals of this species and uses these masses to find two confidence intervals (% and 90%) for the population mean. The width of the % confidence interval is the width of the 90% confidence interval.
Approach
For a normal-distribution confidence interval, the width is twice the critical value multiplied by the standard error. Equate the width of the interval to times the width of the 90% interval to find the critical value, then convert this critical value to a confidence percentage.
Working
The width of a confidence interval for a population mean, using a known standard deviation and sample size , is
where is the two-tailed standard normal critical value for the confidence level.
For a 90% confidence interval, . Let be the critical value for the confidence interval. Since the interval has width times the 90% interval:
Cancelling the common factor :
Now convert this critical value to a two-tailed confidence level:
Since :
Answer
98
Walkthrough
The confidence interval here has the form
so its total width is . We are told the interval is times as wide as the 90% interval. Because the sample size and standard deviation are the same for both intervals, the factor appears in both widths and cancels. This leaves a simple relation between the two critical values: the critical value is times the 90% critical value.
Using the 90% critical value , the critical value is . A critical value of means that the central area of the standard normal distribution between and is , so the confidence level is 98%.
Key Takeaways
- The width of a confidence interval scales with the critical value, so common factors such as , , and cancel when comparing intervals from the same sample.
- A two-tailed confidence level is the total central probability, found from .
- The standard normal critical value can be converted into a percentage confidence level.
Common Mistakes
- Forgetting the factor of 2 in the width. It cancels here, but it must be included in the formula.
- Using a one-tailed probability for the critical value instead of the two-tailed confidence level.
- Not recognising that both intervals use the same standard error, and unnecessarily trying to solve for or .
- Approximating incorrectly instead of using .
Things to Be Careful About
The confidence level is expressed as a percentage, so the decimal probability corresponds to . The mark scheme accepts either 98 or 98%. Also, the critical value for 90% confidence is , not (which would be for an 80% confidence interval); although the method mark could be awarded with 1.282, the final relation is based on 90%.
Find the probability that the 90% confidence interval contains the population mean given that the % confidence interval contains the population mean.
Approach
Since both intervals have the same centre, the 90% interval is contained inside the wider 98% interval. The event that the 90% interval contains the mean is therefore a subset of the event that the 98% interval contains the mean, so use the conditional probability formula for nested events.
Working
Let . The 90% interval contains when
The 98% interval contains when
Since , the event is a subset of . Therefore
Substituting the confidence levels as probabilities:
Answer
45/49 or 0.918
Walkthrough
Both confidence intervals are centred on the same sample mean ; only their widths differ. A 90% interval uses a smaller critical value () than the 98% interval (), so whenever the 90% interval contains the population mean, the 98% interval must also contain it. This makes the 90% event a subset of the 98% event.
For conditional probability, when one event is a subset of another, . Here is the event that the 90% interval contains , whose probability is 0.90, and is the event that the 98% interval contains , whose probability is 0.98. So the required conditional probability is .
Key Takeaways
- A lower-confidence interval is narrower and is contained within a higher-confidence interval centred at the same sample mean.
- If event is contained in event , then .
- Confidence level probabilities can be used directly in probability calculations.
Common Mistakes
- Assuming the probability is 1 because the wider interval already contains the mean.
- Dividing in the wrong order, i.e. instead of .
- Forgetting to use the confidence probabilities (0.90 and 0.98) as decimals.
Things to Be Careful About
The mark scheme allows follow-through from the value of found in part (a), but if were less than 90, the subset relationship would not hold and the conditional-probability ratio would not apply. Here , so the calculation is valid. Express the answer as a fraction or as to 3 significant figures.
It is known that 20% of households in a certain country contain more than 4 people. Laxmi believes that, in her town, the percentage is lower than 20%. She chooses a random sample of 40 households in her town and notes the number which contain more than 4 people. She then carries out a test at the 2.5% significance level using a binomial distribution.
Approach
Let be the number of households in the sample with more than 4 people. Under , . Since Laxmi believes the percentage is lower, this is a one-tailed lower-tail test:
The rejection region is the largest set for which . The probability of a Type I error is .
Working
Calculate the binomial probabilities:
So
This is less than . Check the next value:
Thus the critical region is , and
Answer
0.00794 (3 sf)
Walkthrough
We start by defining the random variable as the number of households in the sample of 40 that contain more than 4 people. The null hypothesis is , because the national percentage is 20%. Laxmi believes the percentage in her town is lower, so the alternative hypothesis is . This is a one-tailed lower-tail test.
A Type I error is made when the test rejects even though is true. Its probability is therefore the probability that the test statistic falls in the rejection region when . We need to find the largest number such that , because the test is at the 2.5% significance level.
We calculate , and using the binomial formula. Their sum is , which is less than 0.025. We must also check the next value: adding gives , which is greater than 0.025. This shows that the rejection region cannot include 3, so the critical region is . The probability of a Type I error is the actual probability of being in this critical region, namely 0.00794.
Key Takeaways
This question tests how to find the critical region for a one-tailed binomial test and how to convert that region into the probability of a Type I error. The probability of a Type I error is not automatically the nominal significance level; it is the actual probability of the rejection region under . You must compare cumulative binomial probabilities with the significance level to decide where the critical region ends.
Common Mistakes
- Giving only without showing the binomial expressions. The mark scheme requires the expressions and the comparison with for full marks.
- Forgetting to check . Without this comparison, you have not justified that the critical region is .
- Using the wrong tail. Since Laxmi believes the percentage is lower, the test must be lower-tailed.
- Using instead of . The probability of a household containing more than 4 people is 0.2 under .
Things to Be Careful About
The significance level is 2.5%, so the cumulative probability of the rejection region must be at most 0.025. Because is well below 0.025 while is above, the boundary is at . Round only at the end: use the full values when adding the terms. The final probability should be given as 0.00794 (3 sf).
Approach
The rejection region is the set of values of for which is rejected. For a lower-tail binomial test, it has the form , where is the largest value such that when .
Working
From part (a), and . Therefore the largest value in the rejection region is 2.
Answer
Rejection region: .
X ≤ 2
Walkthrough
The rejection region is the set of values of for which Laxmi will reject . In a lower-tail binomial test, this region has the form . From part (a), , but . Therefore the largest value that can be included is 2, and the rejection region is .
Key Takeaways
The rejection region is determined by the cumulative probability being no greater than the significance level. It must be stated in terms of the test statistic , not just as a probability.
Common Mistakes
- Writing instead of .
- Writing the probability 0.00794 as the rejection region. The rejection region is the set of observed values, not the probability.
- Including 3 in the rejection region because is small; the correct condition is on the cumulative probability , which exceeds 0.025.
Things to Be Careful About
The rejection region for a lower-tail test is always of the form . Make sure the boundary value is included when its cumulative probability is below the significance level.
Laxmi finds that exactly 2 households in her sample contain more than 4 people.
Explain why it is impossible for Laxmi to make a Type II error.
Approach
A Type II error occurs when is false but the test does not reject . If the observed value lies in the rejection region, is rejected, so a Type II error cannot occur.
Working
Laxmi observed . The rejection region from part (b) is , so the observed value lies in the rejection region. Therefore she rejects . A Type II error would require failing to reject , which is not the case.
Answer
Because is in the rejection region, is rejected; a Type II error is only possible when is not rejected.
Because X = 2 lies in the rejection region, H0 is rejected, so a Type II error cannot occur.
Walkthrough
A Type II error occurs when is false but the test fails to reject . Laxmi observed . From part (b), the rejection region is , so the observed value lies in the rejection region. This means she rejects . Since she rejects , she cannot fail to reject a false , so a Type II error is impossible in this outcome.
Key Takeaways
Type II error is conditional on being false and on the test not rejecting it. If the observed value falls in the rejection region, the test rejects , so a Type II error cannot occur.
Common Mistakes
- Saying a Type II error is impossible because the sample proportion is low. The reason is specifically that the observed value lies in the rejection region.
- Confusing Type I and Type II errors. A Type I error is rejecting a true ; a Type II error is failing to reject a false .
Things to Be Careful About
The observed value is exactly the boundary of the rejection region, so it is included. Because it is included, is rejected and no Type II error can be made.
The masses, in kilograms, of large and small bags of potatoes have the independent distributions and respectively.
Find the probability that the total mass of a randomly chosen large bag of potatoes and a randomly chosen small bag of potatoes is more than 3.55kg.
Approach
Let be the mass of a large bag and the mass of a small bag. The total mass is . Since and are independent normal variables, is also normal with mean equal to the sum of the means and variance equal to the sum of the variances. Standardise and use the standard normal table to find the required probability.
Working
Given:
The total mass is . Since and are independent:
Standardise:
Using the normal distribution table:
Answer
0.172 (3 s.f.)
Walkthrough
We are given the masses of two independent normal distributions. The key idea is that when we add two independent normal random variables, the result is also normal. The mean of the sum is the sum of the means, and the variance of the sum is the sum of the variances (this is because the variables are independent, so the covariance is zero).
Once we know the distribution of the total mass , we need to find . To use the standard normal table, we standardise by subtracting the mean and dividing by the standard deviation. The standard deviation is , not — this is a common error.
The standardised value is . Since we want the probability that is greater than 3.55, we want the area to the right of under the standard normal curve. The table gives the area to the left, so we subtract from 1.
Key Takeaways
The sum of independent normal variables is normal. For independent variables, . To find probabilities for a normal distribution, always standardise first. The normal table gives , the area to the left.
Common Mistakes
Using the variance (0.07) instead of the standard deviation () when standardising. Forgetting to subtract from 1 when the question asks for the probability of being "more than" a value. Adding variances incorrectly or forgetting that independence is required.
Things to Be Careful About
The variance of the sum is the sum of the variances only because and are independent. The final answer must be given to 3 significant figures as requested. Make sure to standardise using the standard deviation, not the variance.
Find the probability that the mass of a randomly chosen large bag of potatoes is less than 3 times the mass of a randomly chosen small bag of potatoes.
Approach
We need . Rearrange to and define . Since and are independent normal variables, is normal. Find its mean and variance, then standardise and use the normal table.
Working
Define . Since and are independent:
So .
We need:
Standardise:
By symmetry of the normal distribution:
Answer
0.417 (3 s.f.)
Walkthrough
The inequality can be rearranged to . This is a linear combination of the two independent normal variables. The key difference from part (a) is that is multiplied by 3, so when we find the variance, the coefficient 3 must be squared: .
The mean of is . The variance is . Note that the variance of is the same as the variance of because variance is unaffected by sign.
We then standardise: . Since we want , we want the area to the left of . By symmetry, this equals .
Key Takeaways
An inequality like can be converted to a linear combination . When a variable is multiplied by a constant, its variance is multiplied by the square of that constant. The variance of for independent is . The standard normal distribution is symmetric about 0, so .
Common Mistakes
Forgetting to square the coefficient 3 when finding the variance of . Using or instead of adding the squared-coefficient terms. Confusing the direction of the inequality when standardising.
Things to Be Careful About
The variance of a linear combination is always a sum of squared-coefficient times variance terms (for independent variables) — the minus sign on does not make the variance negative. The standardised value is negative here; use the symmetry property of the normal distribution correctly. The answer may be given as 0.417 or 0.418 depending on rounding; the mark scheme allows both.
The time, in minutes, taken by students to complete a test is modelled by the random variable with probability density function
Find the probability that a randomly chosen student takes longer than 4.5 minutes to complete the test.
Approach
We want . Since the pdf is zero outside , this is the integral of from to .
Working
Expand the integrand:
So
Evaluate at the limits:
Thus
Answer
5/32 or 0.156
Walkthrough
We need the probability that exceeds . For a continuous random variable, probabilities are found by integrating the pdf over the required interval. Because the pdf is defined only for , the event corresponds to the interval . First expand to make the integral easier. Then integrate term by term: , , and . Substitute and , subtract, and multiply by . The negative sign in the pdf is important; the bracket difference is negative, so the final probability comes out positive.
Key Takeaways
A continuous probability is an area under the pdf. When the pdf is a quadratic, expand it before integrating. Always use the correct limits for the event described.
Common Mistakes
Forgetting the factor when integrating. Substituting only one limit. Confusing with . Not expanding the quadratic correctly.
Things to Be Careful About
The pdf is only non-zero on , so the upper limit must be , not infinity. Keep the negative factor throughout; it is essential for the final positive probability.
Approach
The pdf is symmetric about . For a symmetric continuous distribution, the median is the centre of symmetry, so the median is .
Working
For ,
and
so the distribution is symmetric about . Therefore the areas on either side of are equal.
Answer
4
Walkthrough
The median is the value such that half the probability lies below and half lies above it. Here the pdf is symmetric about : replacing by and gives the same value. Hence the area under the curve is split equally at , so the median is simply .
Key Takeaways
Recognising symmetry in a pdf can give the median immediately without integration. A symmetric continuous distribution has its median at the axis of symmetry.
Common Mistakes
Attempting to integrate to find the median unnecessarily. Forgetting that the median is a value of , not a probability.
Things to Be Careful About
The symmetry must be checked around the proposed centre. Here confirms it.
Approach
Use the symmetry of the pdf about . From part (a), . By symmetry, as well. The required interval is the remaining probability, so subtract both tails from .
Working
Answer
11/16 or 0.6875
Walkthrough
The interval is centred at . The distribution is symmetric about , so the probability below equals the probability above . Part (a) gave the upper tail as , so the lower tail is also . Since the total probability is , subtract both tails from : . This avoids any new integration.
Key Takeaways
Symmetry plus complement can find a central probability from a tail probability. Always check that the two tails are equal before using this method.
Common Mistakes
Forgetting to subtract both tails, giving . Integrating from to despite the instruction not to integrate. Using the wrong tail probability.
Things to Be Careful About
The event uses strict inequalities, but for a continuous distribution the endpoints have probability zero, so and give the same probability. The mark scheme requires working: show explicitly.