Mathematics 9709/61 — May/June 2025
Cambridge A-Level · Probability & Statistics 2 · worked solutions for every part, with the mark scheme
Topics Sampling and Estimation · The Poisson Distribution · Hypothesis Tests · Linear Combinations of Random Variables · Continuous Random Variables
It is known that of houses in a certain area have a wind turbine. A random sample of houses in this area is chosen for a survey on domestic heating. The number of houses in the sample that have a wind turbine is denoted by .
Use a suitable approximating distribution to find .
Approach
The number of houses with a wind turbine follows a binomial distribution . Since is large and is small, we approximate this binomial distribution by a Poisson distribution with
Then we calculate using the Poisson probabilities.
Working
Let . Then
Using :
Answer
0.433
Walkthrough
The random variable counts the number of houses in a sample of that have a wind turbine. Each house independently has a wind turbine with probability , so the exact distribution is binomial:
Because is large and is small, the Poisson approximation is appropriate. The Poisson parameter is the mean of the binomial distribution:
We now need . Since the Poisson distribution has no upper limit, we add the probabilities for :
Factor out and evaluate each term. This gives to 3 significant figures.
Key Takeaways
- A binomial distribution with large and small can be approximated by a Poisson distribution with .
- To find for a Poisson variable, sum the individual probabilities from up to .
- The Poisson formula is .
Common Mistakes
- Forgetting to include the term when calculating .
- Using a normal approximation without a continuity correction; the Poisson approximation is the intended method here.
- Quoting the final answer without showing the sum of the four Poisson probabilities; the mark scheme requires method marks.
- Confusing with the probability .
Things to Be Careful About
- The parameter is , not .
- The term for is , which is often forgotten.
- Give the final answer to 3 significant figures: .
- If the normal approximation is used, the mark scheme only awards a mark for identifying the mean , not for the full calculation.
The random variable has the distribution . A random sample of values of is chosen, and the sample mean, , is found.
Approach
For , find the mean and variance of the binomial distribution, then apply the Central Limit Theorem to the sample mean , standardise it, and evaluate the normal tail probability.
Working
For with and :
For the sample mean of 100 values:
By the Central Limit Theorem, is approximately normal:
Standardise:
Therefore:
From normal tables, :
Answer
0.0512 or 0.0513
Walkthrough
We begin with the binomial random variable . The first step is to compute its mean and variance using the standard binomial formulas: and . With and , we get and .
Next, we take the sample mean of 100 independent values of . The mean of the sample mean equals the population mean: . The variance of the sample mean is the population variance divided by the sample size: .
Because is binomial (not normal), the Central Limit Theorem is required to justify that is approximately normal for a large sample. With , the CLT applies, so we treat .
To find , we standardise: . Then . Using normal tables, , so the probability is .
Key Takeaways
- The mean and variance of a binomial distribution are and .
- For a sample mean, and .
- The Central Limit Theorem allows the sample mean to be treated as normal for large even when the population is not normal.
- Standardising a normal variable to enables probability tables from standard normal tables to be used.
Common Mistakes
- Forgetting to divide the variance by the sample size when computing .
- Using the population standard deviation instead of when standardising.
- Forgetting to subtract from 1 when the question asks for a "greater than" probability.
- Applying a continuity correction when the question explicitly says not to use one.
Things to Be Careful About
- The standard deviation of is , not .
- The mean of is 6, not .
- The standardisation must include in the denominator.
- The final answer should be given to 3 significant figures: or .
Approach
The Central Limit Theorem is needed because the underlying variable is not normally distributed.
Working
The random variable has a binomial distribution , which is not a normal distribution. The Central Limit Theorem states that for a sufficiently large sample, the sample mean is approximately normally distributed regardless of the underlying population distribution. Since is large, the CLT justifies treating as normal in part (a).
Answer
The population () is not normally distributed; it is binomial.
The population (X) is not normally distributed; it is binomial.
Walkthrough
The Central Limit Theorem is needed because the underlying variable follows a binomial distribution, not a normal distribution. The CLT guarantees that the sample mean is approximately normal for a large sample regardless of the population distribution. Without it, we could not assume was normal in part (a).
Key Takeaways
- The CLT is the justification for using the normal distribution for a sample mean when the population is not normal.
- A sample size of 100 is considered large enough for the CLT to apply.
Common Mistakes
- Saying the CLT is needed because the sample is large — this is not the reason; the reason is that the population is not normal.
- Confusing the CLT with the normal approximation to the binomial distribution.
Things to Be Careful About
- The answer should mention that the underlying distribution is not normal — it is binomial.
- The mark scheme accepts "Population () is not normally distributed" or "Population () is Binomial" or equivalent.
The time, minutes, for a certain daily bus journey is normally distributed. The bus company claims that the mean of is . A passenger believes that the mean of is actually greater than . She notes the times taken for this journey on a random sample of days. The results are summarised below.
Approach
Use the sample mean as an unbiased estimate of the population mean, and use the sample variance with denominator as an unbiased estimate of the population variance.
Working
The unbiased estimate of the population mean is the sample mean:
The unbiased estimate of the population variance is
Substituting the given values:
Equivalently, using the form in the mark scheme:
Now compute:
Answer
Unbiased estimates:
Mean estimate = 45.8, variance estimate = 16.2
Walkthrough
We are given only the summary statistics from a sample of 60 days, so the best point estimate of the population mean is the sample mean. We therefore divide the sum of the times by the sample size.
For the variance, we need an unbiased estimate. The formula uses in the denominator:
This corrects for the fact that using the sample mean in place of the true population mean tends to underestimate the variance. We substitute , and , then simplify to obtain , which is to 3 significant figures.
Key Takeaways
- The sample mean is an unbiased estimator of the population mean.
- The unbiased estimate of the population variance uses denominator , not .
- Summary statistics and can be substituted directly into the estimator formulas.
Common Mistakes
- Using denominator instead of gives a biased estimate, which the mark scheme explicitly disallows.
- Confusing with .
- Forgetting to square the sample mean inside the variance formula.
Things to Be Careful About
- The variance formula can be written in two equivalent forms; either is acceptable.
- Keep enough precision in the variance for part (b) to avoid rounding errors.
- The mean estimate is to 3 significant figures, and the variance estimate is to 3 significant figures.
Approach
This is a one-tailed test for the population mean. Since the population is normal and the sample size is , the sample mean is normally distributed. Use the unbiased variance estimate from part (a) to standardise the observed sample mean, then compare the test statistic with the critical value for a one-tailed test.
Working
Let be the mean journey time. State the hypotheses:
The sample mean is
From part (a), . The standard error of the sample mean is
Standardise:
For a one-tailed test at the significance level, the critical value is
Compare:
The test statistic is not in the critical region, so we do not reject .
Answer
There is insufficient evidence at the significance level to reject the company's claim that the mean time is minutes, or to support the passenger's belief that the mean time is greater than minutes.
Insufficient evidence at the 5% significance level to reject H0; the passenger's belief is not supported.
Walkthrough
We want to test whether the mean time is greater than minutes. The null hypothesis represents the company's claim, so . The passenger's belief is the alternative hypothesis, , making this a one-tailed test.
Because the population is normal and the sample size is , the sample mean is normally distributed. Under , the mean of the sample mean is and its variance is . Since is unknown, we use the unbiased estimate from part (a).
The standard error is
We standardise the observed sample mean:
For a one-tailed test, the critical value is . Since , the observed result is not unusual enough to reject . Therefore we conclude that there is insufficient evidence to support the passenger's belief.
Key Takeaways
- State and clearly before performing the test.
- For a normal population with an estimated variance, use the standardised test statistic .
- Compare the test statistic with the one-tailed critical value.
- The conclusion must be in context and should not claim a definite effect when the result is not significant.
Common Mistakes
- Using the two-tailed critical value instead of the one-tailed value .
- Forgetting the in the denominator when standardising.
- Using a biased variance estimate from part (a) when the question requires the unbiased estimate.
- Concluding "the mean is not greater than 45" rather than "there is insufficient evidence that it is greater than 45".
- Stating the hypotheses without defining the parameter or without using the correct inequality direction.
Things to Be Careful About
- Use the unbiased variance estimate from part (a), not a prematurely rounded value.
- The test statistic is approximately ; values between and are acceptable depending on rounding.
- The critical value for a one-tailed test is .
- A critical-value method is also acceptable: the critical sample mean is , and since , the conclusion is the same.
- The final conclusion must mention the company's claim or the passenger's belief in context.
At an entertainment centre, the cost for using a particular video game is $0.40 per minute. The number of minutes for which people use the video game has mean and variance .
Approach
Let be the number of minutes a person uses the video game. The amount paid is , so this is a linear scaling of .
Working
Answer
Mean , variance .
Mean = 6, variance = 1.44
Walkthrough
Let be the number of minutes. The amount paid is , so we need the mean and variance of .
For the mean, scaling a variable by a constant scales its mean by the same constant:
For the variance, the constant is squared when it is taken out:
So the mean amount paid is 6 and the variance is 1.44.
Key Takeaways
This question tests the linear transformation of a random variable:
A common pattern is to multiply the mean by the constant but square the constant for the variance.
Common Mistakes
- Multiplying the variance by instead of .
- Forgetting that variance is in squared units.
- Using the standard deviation instead of the variance in the calculation.
Things to Be Careful About
The variance is , not the standard deviation . If units are used, the variance is in squared monetary units, but units are not required here.
Each day, people independently use the video game.
Find the mean and variance of the total amount paid by people.
Approach
Let be the amount paid by the -th person. The total paid by 35 independent people is . Use the mean and variance rules for sums of independent random variables: the mean of the sum is the sum of the means, and the variance of the sum is the sum of the variances.
Working
From part (a), each has mean and variance .
Answer
Mean , variance .
Mean = 210, variance = 50.4
Walkthrough
From part (a), each person's payment has mean and variance . Let be the payment of the -th person and let be the total.
Because the 35 people are independent, the mean of the total is the sum of the individual means:
The variance of the total is the sum of the individual variances:
The key point is that when adding independent variables, variances add, not standard deviations, and the number of variables is multiplied once, not squared.
Key Takeaways
For independent variables , the sum has
This is why the total variance is times the individual variance, not times.
Common Mistakes
- Multiplying the variance by instead of .
- Multiplying standard deviations instead of variances.
- Forgetting that the variance addition formula requires independence.
Things to Be Careful About
- Use the answers from part (a) for the mean and variance of one person's payment.
- The variance is in squared monetary units, but units are not required here.
- The independence of the 35 people is given in the question, so no covariance terms are needed.
The random variables and have the independent distributions and respectively.
Approach
Since and are independent Poisson variables, their sum is also Poisson with parameter equal to the sum of the parameters. Then calculate the three individual probabilities and add them.
Working
.
Answer
0.537
Walkthrough
Since and are independent Poisson variables, their sum is also Poisson. The parameter of is the sum of the parameters, so .
The probability includes the three values , and . For a Poisson distribution, . So we calculate each of the three probabilities and add them:
Adding these gives , which rounds to .
Key Takeaways
- The sum of independent Poisson random variables is Poisson, with parameter equal to the sum of the parameters.
- To find a range probability for a discrete distribution, sum the individual point probabilities.
- The Poisson probability formula is .
Common Mistakes
- Forgetting to add the means, and using or instead of .
- Omitting one of the endpoints or .
- Giving only the final answer without showing the expression; the mark scheme gives only SCB1 for an unsupported answer of .
Things to Be Careful About
- The condition is , so both and must be included.
- Use the combined mean , not the individual means.
- Show the expression with and the three factorial terms to earn the method mark.
The random variable is the sum of independent values of and independent values of .
Use a suitable approximation to find .
Approach
is the sum of independent Poisson variables, so is Poisson with mean . Since the mean is large, approximate by a normal distribution with the same mean and variance. Use a continuity correction for , standardise, and find the upper-tail probability.
Working
.
Approximate:
Continuity correction:
Answer
0.197
Walkthrough
is the sum of independent values of and independent values of . Since and are independent Poisson variables, their sum is Poisson. Therefore
Because the mean is large, the Poisson distribution can be approximated by a normal distribution with the same mean and variance:
We need . Since is discrete and the normal distribution is continuous, we apply a continuity correction. The event means , so we use as the boundary:
Finally,
Key Takeaways
- The sum of independent Poisson variables is Poisson with mean equal to the sum of the means.
- A Poisson distribution with a large mean can be approximated by a normal distribution with the same mean and variance.
- A continuity correction is needed when a discrete distribution is approximated by a continuous one.
Common Mistakes
- Using instead of for the continuity correction. The mark scheme allows an omitted or incorrect continuity correction for the method mark, but using is the correct approach.
- Forgetting that the variance of a Poisson distribution equals its mean, so both mean and variance are .
- Forgetting to subtract from when finding the upper tail .
- Using the square root of the mean as the mean, or using incorrectly.
Things to Be Careful About
- Check that the normal approximation is appropriate: here the mean is , which is large.
- The continuity correction for is , not .
- Give the final probability to 3 significant figures: .
The random variable has the distribution , where .
It is given that .
Find the value of .
Approach
Write each Poisson probability using the formula , cancel the common factor , simplify to a quadratic in , and choose the positive root.
Working
Cancel :
Multiply by :
Since , divide by :
Since , .
Answer
λ = 10
Walkthrough
Start with the given equation:
For , . Substitute this into the equation:
Since is never zero, it can be cancelled from every term:
Multiply through by :
Since , divide by :
Factorise:
So or . Since , the only valid solution is .
Key Takeaways
- The Poisson probability formula is .
- When an equation contains the same exponential factor in every term, it can be cancelled.
- Always check the domain of the parameter when solving; here removes the negative root.
Common Mistakes
- Forgetting to cancel and making the algebra more complicated.
- Making factorial errors when simplifying , and .
- Dividing by without noting that , so the division is valid.
- Giving as an answer without rejecting it because .
Things to Be Careful About
- The mark scheme accepts the equation with or without the factors, but the probabilities must be written correctly.
- After multiplying by , rearrange carefully to get the quadratic in the standard form.
- The final answer must be the positive root only.
A manufacturer of cell phones claims that of students own a Pumpkin phone. Jeyeraj thinks that the proportion of students at his large college who own a Pumpkin phone is less than . He plans to test the manufacturer's claim. He chooses a random sample of students at his college. If the number of students who own a Pumpkin phone is less than , Jeyeraj will reject the manufacturer's claim.
Approach
State the null hypothesis as the manufacturer's claim and the alternative hypothesis as Jeyeraj's belief that the proportion is lower. Since the test is one-tailed in the 'less than' direction, use .
Working
Let be the proportion of students at the college who own a Pumpkin phone.
Answer
H0: p = 0.25, H1: p < 0.25
Walkthrough
The manufacturer's claim is that the population proportion is exactly 25%, so this is the null hypothesis. Jeyeraj thinks the proportion is less, so the alternative hypothesis must be one-sided in the downward direction. Therefore and .
Key Takeaways
Hypotheses are statements about the population proportion, not about the sample. The null hypothesis always contains the equality, and the alternative expresses the direction of suspicion.
Common Mistakes
- Writing when the test is specifically one-sided.
- Stating hypotheses using the sample proportion instead of the population proportion.
- Omitting the inequality direction in .
Things to Be Careful About
Use for the population proportion. Keep as and make sure is because Jeyeraj believes the proportion is less.
Given that the true proportion of students at the college who own a Pumpkin phone is , use a binomial distribution to find the probability of a Type II error.
Approach
A Type II error is accepting when is false. Here is rejected when the number of owners , so a Type II error occurs when . Under the true proportion , , so we calculate .
Working
Let be the number of sampled students who own a Pumpkin phone. The rejection rule is , so the test does not reject when .
Since the true proportion is ,
Therefore
Now
Evaluating the terms:
Thus
Answer
(to 3 sf)
0.175
Walkthrough
The rejection rule is to reject the manufacturer's claim if fewer than 5 students own a Pumpkin phone. A Type II error is failing to reject the null hypothesis when the alternative is true. Since the true proportion is 10%, the number of owners in 30 students follows . The test fails to reject when , so we need . It is easier to calculate using the binomial probability formula for . Substituting the probabilities gives , which rounds to .
Key Takeaways
- A Type II error is the probability of not rejecting when is false.
- The true proportion under the alternative is used to set up the binomial distribution.
- Tail probabilities are often easier computed as minus the lower tail.
Common Mistakes
- Using instead of the true proportion .
- Finding instead of .
- Forgetting to subtract from 1 after summing the lower-tail probabilities.
- Giving an unsupported answer; the mark scheme requires the binomial expression and terms to be shown.
Things to Be Careful About
The rejection region is , so the non-rejection region is . Use the true proportion in the binomial model. Show the full binomial sum and the individual terms to earn the method marks. Round the final answer to 3 significant figures.
At Florence's college, in a random sample of students, it was found that own a Pumpkin phone.
Calculate an approximate confidence interval for the proportion of students at Florence's college who own a Pumpkin phone.
Approach
Use the sample proportion as the estimate of the population proportion and construct a 95% confidence interval using the normal approximation:
with for a 95% confidence interval.
Working
The sample proportion is
The standard error is
For a 95% confidence interval, , so
Therefore the interval is
Answer
0.0225 to 0.227
Walkthrough
We have a random sample of 40 students, of whom 5 own a Pumpkin phone. The sample proportion is . For a 95% confidence interval for a proportion, we use , where . The standard error is approximately . Multiplying by gives the margin of error . Adding and subtracting this margin from gives the interval from to , which rounds to to at 3 significant figures.
Key Takeaways
The confidence interval for a proportion is centered at the sample proportion and uses a normal approximation. The critical value for 95% confidence is . The interval is written as a pair of lower and upper bounds.
Common Mistakes
- Using , which is for a one-sided 95% interval, instead of .
- Using in the denominator of the standard error.
- Forgetting to give both endpoints of the interval.
- Rounding intermediate values too early, which can change the final 3 significant figures.
Things to Be Careful About
The mark scheme requires a -value; any used must be a standard normal critical value. Only calculating one side of the interval can lose marks, so always show both the lower and upper bounds. Use , not , and give the final interval to 3 significant figures.
X is a random variable with probability density function given by
Approach
The probability is the integral of the probability density function over the interval . On this interval , so
Working
Integrate:
Substitute the limits:
Answer
1/2 + 1/π
Walkthrough
We need the probability that is less than . For a continuous random variable, this is the integral of the probability density function over the interval . The PDF is zero outside , and lies inside that interval, so we integrate from to .
The integral splits into two pieces: and . So the antiderivative is .
We evaluate this antiderivative at the upper limit and the lower limit , subtracting the lower value from the upper. At , , so the first term is . At , , so the lower value is . The result is .
Key Takeaways
- For a continuous random variable, the probability of falling in an interval is the integral of its PDF over that interval.
- Integrating gives ; the factor in the denominator is essential.
Common Mistakes
- Forgetting the factor when integrating .
- Using the wrong limits (e.g., starting at ); the PDF is zero outside , so the written working should start at .
- Not showing the intermediate antiderivative . Since the answer is given, a bare final number is not sufficient.
Things to Be Careful About
- This is an "answer given" (AG) question: the working must be convincing and free of errors. The mark scheme awards the marks only if the intermediate step is shown and no errors are seen.
- Evaluate and correctly; both are standard values.
Approach
The expected value is . Since outside , this becomes
The term requires integration by parts.
Working
Set up the integral:
The first part integrates to
For , let and , so and . Integration by parts gives
So the full integral becomes
Now , giving
Substitute the limits:
Since , , :
Answer
1/2 - 2/π²
Walkthrough
The expected value of a continuous random variable is . Since for and , the integral reduces to .
The first term integrates to . The second term is a product of a polynomial and a trigonometric function, so we use integration by parts. Set and . Then and . This gives . The remaining integral is .
Combining, the antiderivative is . Substitute and and subtract. Using , , , we get .
Key Takeaways
- The expectation of a continuous random variable is over the whole real line; the PDF being zero outside its support simplifies the limits.
- Integration by parts is the standard tool for ; choosing removes the polynomial.
- The integrals and are needed.
Common Mistakes
- Forgetting to multiply by inside the integral — integrating alone gives the total area , not the expectation.
- Errors in the integration by parts, especially the sign of the term. The minus sign in front of the integral must be handled correctly.
- Forgetting the or factors when integrating the trigonometric functions.
- Not showing the intermediate expression with the integral sign, which the mark scheme requires for the M1 mark.
Things to Be Careful About
- The mark scheme requires that the integration by parts reach an expression with at least two terms correct (and more than two terms). Write out the full antiderivative before substituting limits.
- This is an AG question: the final answer is given, so the working must be legitimate and contain no errors.
- When substituting the lower limit , , , and .