Mathematics 9709/62 — February/March 2024
Cambridge A-Level · Probability & Statistics 2 · worked solutions for every part, with the mark scheme
Topics Sampling and Estimation · The Poisson Distribution · Linear Combinations of Random Variables · Hypothesis Tests · Continuous Random Variables
The lengths, , of a sample of 100 insects of a certain type were summarised as follows.
Approach
The sample mean is an unbiased estimate of the population mean . The unbiased estimate of the population variance is obtained using the divisor :
Working
The unbiased estimate of the population mean is the sample mean:
So
For the variance, use the unbiased formula:
First calculate :
Then:
Answer
Unbiased estimate of the population mean: .
Unbiased estimate of the population variance: (3 s.f.).
Mean = 0.368, variance = 0.0384 (3 s.f.)
Walkthrough
We are given summary statistics for a sample of 100 insects: , and . Because these are sample data, we want unbiased estimates of the population mean and variance.
The sample mean is an unbiased estimator of the population mean, so we divide the total of the observations by the sample size:
For the population variance, we cannot simply divide by . Since we have estimated the mean from the same data, one degree of freedom is lost, so we use as the divisor. This gives the unbiased estimate:
Substituting the given values gives to 3 significant figures.
Key Takeaways
- The sample mean is an unbiased estimate of the population mean.
- An unbiased estimate of the population variance uses the divisor , not .
- Summary statistics such as and are sufficient to calculate these estimates.
Common Mistakes
- Using instead of in the variance formula, which gives the biased sample variance.
- Forgetting to subtract when using .
- Using instead of .
- Rounding intermediate values too early, which can change the final answer.
Things to Be Careful About
- The unbiased variance estimate is , not ; the latter would come from dividing by instead of .
- Give the variance to 3 significant figures as required by the mark scheme.
- The mean has units cm, while the variance has units .
Approach
The reliability of the estimates depends on how the sample was chosen, not just on the sample size. The necessary condition is that the sample must be random.
Working
For the estimates to be reliable, the sample must be a random sample. This means that every insect in the population has an equal chance of being selected, so that the sample is representative of the population.
Answer
The sample must be a random sample.
The sample must be a random sample.
Walkthrough
Part (b) asks for a necessary condition for the estimates in part (a) to be reliable. The unbiased formulas used in part (a) assume that the data come from a random sample. If the sample is not random, the estimates may be biased even though the formulas are mathematically correct. Therefore the necessary condition is that the sample must be a random sample.
Key Takeaways
- Random sampling is essential for estimates to be reliable.
- A random sample gives every member of the population an equal chance of being selected.
- The reliability of an estimate depends on how the sample is collected, not only on its size.
Common Mistakes
- Saying that the sample must be large. A large sample is not the necessary condition asked for here.
- Saying that the population must be normally distributed. This is not required for unbiased estimates of the mean and variance.
- Giving only "the sample should be representative" without mentioning randomness.
Things to Be Careful About
- The mark scheme accepts several equivalent answers, such as "randomly selected", "representative of the population", or "all values should have equal chance of being selected".
- Do not confuse reliability with precision; a random sample ensures representativeness, not necessarily a small variance.
A random sample of 250 people living in Barapet was chosen. It was found that 78 of these people owned a BETEC phone.
Calculate an approximate 98% confidence interval for the proportion of people living in Barapet who own a BETEC phone.
Approach
The confidence interval for a population proportion is based on the sample proportion and the standard error of . For a 98% confidence interval, the critical value from the standard normal distribution is . Substitute into
and evaluate the two endpoints.
Working
The sample proportion is
For a 98% confidence interval,
The standard error is
So the confidence interval is
Evaluating the margin of error:
Therefore,
and
Answer
0.244 to 0.380
Walkthrough
We are estimating the unknown proportion of all Barapet residents who own a BETEC phone. From a sample of people, 78 own one, so the sample proportion is . Because the sample is large, is approximately normally distributed with standard error . For a 98% confidence interval, we want the central 98% of the normal distribution, so the upper tail is 1% and the critical value is . We then form standard error. This gives the lower and upper endpoints. The interval is , meaning we are 98% confident that the true proportion lies in this range.
Key Takeaways
- The sample proportion is the point estimate for the population proportion.
- The standard error of is .
- For a 98% CI, use , not 1.96.
- The interval is standard error.
Common Mistakes
- Using for a 95% interval or for a 99% interval instead of .
- Forgetting to divide by inside the square root.
- Using a value such as instead of the sample proportion in the standard error.
- Giving only the margin of error rather than the full interval.
Things to Be Careful About
- The formula uses in the numerator, not .
- The interval must be written as two values, not just one endpoint.
- The answer should be rounded to 3 significant figures: to .
- This is an approximate interval, valid because the sample size is large.
Manjit claims that more than 40% of the people living in Barapet own a BETEC phone.
Use your answer to part (a) to comment on this claim.
Approach
Compare the claimed proportion with the confidence interval from part (a). If lies outside the interval, the claim is not supported by the sample.
Working
From part (a),
The claimed proportion is
Since is greater than the upper endpoint , it is not contained in the confidence interval.
Answer
Unlikely to be true, because is not in the confidence interval.
Unlikely to be true because 0.4 is not in the confidence interval.
Walkthrough
Manjit claims that more than 40% of residents own a BETEC phone. The 98% confidence interval from part (a) estimates the plausible range for the true proportion. Since the whole interval lies below 0.4, the claimed value 0.4 is not inside the interval. Therefore the sample evidence does not support the claim; it is unlikely to be true.
Key Takeaways
- A confidence interval gives a range of plausible values for the unknown proportion.
- If a claimed value is outside the interval, the claim is not supported by the data.
- The word "unlikely" should be used when commenting, rather than a definite statement.
Common Mistakes
- Saying only "the interval goes up to 0.38" without explicitly stating that 0.4 is not in the interval.
- Saying "it lies outside the confidence interval" without making clear that the claim is unlikely to be true.
- Making a definite conclusion such as "impossible" rather than "unlikely".
Things to Be Careful About
- The mark scheme requires both the reason (0.4 is not in the interval) and the conclusion "unlikely".
- Use the interval from your part (a); if your interval is different, compare 0.4 with your endpoints.
- 40% must be written as 0.4 when comparing with the interval.
In a certain lottery, on average 1 in every 10 000 tickets is a prize-winning ticket. An agent sells 6000 tickets.
Use a suitable approximating distribution to find the probability that at least 3 of the tickets sold by the agent are prize-winning tickets.
Approach
Let be the number of prize-winning tickets sold by the agent. The exact distribution is binomial, but with large and small we approximate it by a Poisson distribution with the same mean.
Working
There are tickets and the probability that one ticket wins is . Therefore
so approximately. We need . Using the complement rule,
For a Poisson distribution . Hence
Numerically,
So
Answer
0.0231
Walkthrough
Let be the number of prize-winning tickets among the 6000 sold. Since each ticket either wins or does not win, has a binomial distribution . Computing exactly with would involve very large numbers, so we look for an approximation. When is large and is small, the Poisson distribution is a standard approximation to the binomial. The mean of the binomial is , so we use .
To find , the complement is easier: . We calculate by adding the Poisson probabilities for , and . The formula is . Substituting gives the three terms shown. Their sum is approximately , so the required probability is .
Key Takeaways
This question tests the Poisson approximation to the binomial: using when is large and is small. It also tests routine Poisson probability calculations and the complement rule for an inequality such as at least 3.
Common Mistakes
- Using the exact binomial distribution and attempting a lengthy calculation instead of the approximation.
- Forgetting to include when computing .
- Computing incorrectly; the complement requires .
- Omitting the expression in part (a). The mark scheme requires the method expression to be seen for the method mark; a bare answer without working can lose marks.
- Confusing with a probability; is the mean, not the per-ticket probability.
Things to Be Careful About
The final numerical value is to 3 significant figures. Use , not . When using the complement, subtract from 1 after summing exactly . The mark scheme accepts the algebraic expression ; showing that expression is a clear way to secure the method mark.
Approach
The underlying distribution of the number of prize-winning tickets is binomial. It is appropriate to approximate this binomial distribution by a Poisson distribution when is large and either or is small. We check both conditions explicitly.
Working
There are tickets, so
The probability of a prize-winning ticket is
Equivalently,
Both conditions are satisfied, so the Poisson approximation with is justified.
Answer
Poisson approximation is justified because and (equivalently ).
n = 6000 > 50 and np = 0.6 < 5 (equivalently p = 0.0001 < 0.1)
Walkthrough
The exact distribution is binomial: there are tickets and each has success probability . The standard rule for a Poisson approximation is that should be large and should be small, with moderate. The mark scheme wants the values to be stated. Here , which is greater than 50, so is large. Also , or equivalently . Since these conditions hold, the Poisson model with mean is appropriate.
Key Takeaways
The justification for the Poisson approximation is not just large , small . The examiner requires explicit numerical checks: and either or . These conditions make the approximation reliable and show that you have understood why the Poisson model is suitable.
Common Mistakes
- Writing only since is large and is small is insufficient for the mark.
- Stating but not also checking , or vice versa.
- Using but writing it incorrectly.
- Claiming the distribution becomes Poisson exactly; it is only an approximation to the binomial distribution.
Things to Be Careful About
The two acceptable justifications are and , or and . State both values clearly. Note that is also the parameter used in the approximating Poisson distribution.
Each year a transport firm uses litres of gasoline and litres of diesel fuel, where and have the independent distributions and .
Find the probability that in a randomly chosen year the firm uses more gasoline than diesel fuel.
Approach
Let . Since and are independent normal variables, is also normal. The firm uses more gasoline than diesel when , i.e. . Compute the mean and variance of , standardise, and use the normal distribution tables.
Working
Since and are independent,
and
Thus .
The required probability is
Standardise:
Therefore
Answer
0.0396
Walkthrough
We are asked for the probability that gasoline use exceeds diesel use, i.e. . It is easier to work with one variable: define . Because and are independent and both normal, is normal, so we only need its mean and variance.
For the mean, expectation is linear: . For independent variables, the variance of a difference is the sum of the variances: . This is a common point to remember: subtracting does not subtract variances.
Then has mean and variance . We need . Convert to the standard normal by subtracting the mean and dividing by the standard deviation. Here is standard deviations above the mean, so . Reading the normal tables gives , so the probability is about .
Key Takeaways
- A linear combination of independent normal variables is normal.
- .
- For independent and , , so .
- To find , rewrite it as and standardise.
Common Mistakes
- Writing ; variances always add for independent variables.
- Finding but forgetting to use the correct sign of the mean; the -score uses .
- Using instead of the upper tail .
- Not showing the standardisation step; the mark scheme requires a method mark for standardising with the correct mean and variance.
Things to Be Careful About
- The question asks for "more gasoline than diesel", so the event is , not .
- If you instead use , the event becomes ; the final probability is the same.
- Use the variance, not the standard deviation, in the formula before taking the square root.
- Give the final answer to 3 significant figures as requested.
The costs per litre of gasoline and diesel fuel are $0.80 and $0.85 respectively.
Find the probability that the total cost of gasoline and diesel fuel in a randomly chosen year is between $20,000 and $22,000.
Approach
Let the total cost be . Since and are independent normal variables, is normal. Find and , standardise both boundary values $20,000 and $22,000, and subtract the cumulative probabilities.
Working
Since and are independent,
So .
Standardise the upper limit:
Standardise the lower limit:
Hence
Answer
0.430
Walkthrough
The total cost is . Since and are independent normal variables, is also normal. We need .
First find the mean: . Then find the variance. The coefficients are squared when they multiply variances: .
Now standardise the two boundaries. The upper boundary gives . The lower boundary gives . Therefore the probability is .
Key Takeaways
- For a weighted sum , the mean uses the coefficients directly, but the variance uses their squares.
- A linear combination of independent normal variables is normal, so standard normal tables can be used.
- For a probability between two values, standardise both and subtract cumulative probabilities: .
Common Mistakes
- Forgetting to square the cost coefficients when computing variance, e.g. writing .
- Using or mixing up addition and multiplication.
- Subtracting the cumulative probabilities in the wrong order, giving a negative probability.
- Not using the variance in the denominator; the denominator must be .
Things to Be Careful About
- The total cost is measured in dollars; the boundaries are $20,000 and $22,000.
- Keep enough precision in the -scores; small rounding can change the final 3 significant figures.
- The mark scheme allows the final answer ; do not give an unsupported answer.
- When standardising, the mean is $19,950, which is below both boundaries, so both -scores are positive.
A teacher models the numbers of girls and boys who arrive late for her class on any day by the independent random variables and respectively.
Approach
Over two independent days, the number of girls arriving late is the sum of two independent Poisson variables, so it is Poisson with mean .
Working
Answer
The probability is (3 s.f.).
0.819 (3 s.f.)
Walkthrough
The number of girls late on each day is modelled by . Over a 2-day period, the total number of girls late is the sum of two independent Poisson variables, so it is also Poisson with mean . We need the probability that this total is 0. For a Poisson distribution, , so here . Rounded to 3 significant figures this is 0.819.
Key Takeaways
When independent Poisson counts are added, their means add. The probability of zero events in a Poisson distribution is always .
Common Mistakes
A common error is to use the one-day mean 0.10 instead of the two-day mean 0.20. Another is to forget that the question asks for no girls late, so only the term is needed. An unsupported decimal answer is acceptable here because the mark scheme allows as the final answer, but showing the rate used is safer.
Things to Be Careful About
The period is 2 days, so the rate must be doubled. The answer may be given as or as 0.819 to 3 significant figures. Do not confuse the girls' rate with the boys' rate, since this part concerns girls only.
Find the probability that during a randomly chosen 5-day period the total number of students who arrive late is less than 3.
Approach
Combine the girls and boys rates per day, then scale to 5 days. The total number of late students is Poisson, so sum the probabilities for 0, 1 and 2 late students.
Working
Per day, the total rate is
Over 5 days,
Answer
The probability is (3 s.f.).
0.868 (3 s.f.)
Walkthrough
On any one day, the total number of students late is . Because and are independent Poisson variables, is Poisson with mean . Over 5 independent days, the total is Poisson with mean . The event 'total less than 3' means , or . Using , we sum these three probabilities:
This equals , so the required probability is 0.868 to 3 significant figures.
Key Takeaways
Independent Poisson counts can be combined by adding their means, both across different sources and across repeated periods. Cumulative Poisson probabilities are found by summing individual probabilities.
Common Mistakes
A frequent error is using 0.25 as the mean for the 5-day period instead of 1.25. Another is omitting or when finding . The mark scheme allows one end error for the method mark, but the final answer must be correct for full marks.
Things to Be Careful About
'Less than 3' includes 0, 1 and 2, not 3. Use for the 5-day total. Give the final probability to 3 significant figures.
It is given that the values of and for are very small and can be ignored.
Find the probability that on a randomly chosen day more girls arrive late than boys.
Approach
We need . Since values with can be ignored, only the pairs , and contribute. Use independence to multiply Poisson probabilities, then add the mutually exclusive cases.
Working
Answer
The probability is (3 s.f.).
0.0824 (3 s.f.)
Walkthrough
We need . Since can be ignored, the only possible values of and are 0, 1 and 2. For , the favourable pairs are , and . Because and are independent, the probability of each pair is the product of the individual Poisson probabilities. We compute:
Then add the three products:
Key Takeaways
Discrete probability questions often require enumerating all favourable outcomes. Independence means multiplying probabilities, and mutually exclusive cases are added. The instruction to ignore very small tail probabilities reduces the number of cases.
Common Mistakes
Missing the case is common. Another common error is using instead of . Students sometimes add the two Poisson means or probabilities instead of multiplying independent probabilities. The mark scheme gives a method mark for one correct product and a second method mark for adding all three terms.
Things to Be Careful About
'More girls than boys' is a strict inequality. Only three cases contribute because is negligible. Use the exact Poisson probabilities and keep enough decimal places before rounding the final answer to 3 significant figures.
Following a timetable change the teacher claims that on average more students arrive late than before the change. During a randomly chosen 5-day period a total of 4 students are late.
Test the teacher's claim at the 5% significance level.
Approach
Set up a one-tailed hypothesis test for the Poisson mean over a 5-day period. Calculate the probability of observing 4 or more late students under the null hypothesis, compare it with the 5% significance level, and conclude in context.
Working
Before the change, the mean number of late students per day is
so over 5 days
Let be the total number of late students in 5 days. The hypotheses are
Under , . The probability of 4 or more is
Since , reject at the 5% significance level.
Answer
There is sufficient evidence to support the teacher's claim that more students arrive late on average after the timetable change.
Reject H0 at the 5% level; there is sufficient evidence to support the teacher's claim.
Walkthrough
Before the change, the mean number of late students per day is , so over 5 days the mean is . The teacher claims the mean has increased, so this is a one-tailed test:
Under , the total number late in 5 days is . The observed value is 4, so we calculate the probability of 4 or more:
Since , we reject at the 5% significance level. There is sufficient evidence to support the teacher's claim that more students arrive late on average after the timetable change.
Key Takeaways
A hypothesis test for a Poisson mean uses the Poisson distribution to find the probability of the observed result or more extreme. The comparison of this probability with the significance level determines whether to reject the null hypothesis. The conclusion must be in context and should not claim certainty.
Common Mistakes
Using the per-day mean 0.25 instead of the 5-day mean 1.25 is a common error. Another is testing two-tailed when the claim is clearly one-sided ('more'). Some students calculate instead of . The mark scheme notes that 0.0383 with no working scores only B1, so the full expression must be shown.
Things to Be Careful About
State both hypotheses clearly. Use the 5-day mean . Compare the tail probability with 0.05 and reject because . The final conclusion should be in context, for example 'there is sufficient evidence to suggest that the teacher's claim is true', not 'the claim is correct' or 'more students are late' as a definite statement.
The graph of the probability density function of a random variable is symmetrical about the line . It is given that .
Approach
The PDF is symmetric about , so the probability of being within 3 units to the right of 2 equals the probability of being within 3 units to the left of 2. Use this to find the left tail , then use the complement rule.
Working
Since the graph is symmetric about ,
The total probability is 1, and by symmetry the two outer tails are equal:
Therefore,
So
Hence
Answer
P(X > -1) = 245/256
Walkthrough
The key idea is symmetry. The PDF is symmetric about , so the probability of being between and (three units to the left of ) equals the probability of being between and (three units to the right of ). We are told the latter is . The total probability under any PDF is 1. By symmetry, the two outer tails and are equal. So the total probability is
which becomes
Solving gives . Then is the complement of , so
Key Takeaways
Symmetry of a PDF about a line means equal probabilities for equal distances on either side. The complement rule is useful for finding the probability of the opposite event. For a continuous random variable, single-point probabilities are zero, so strict and non-strict inequalities give the same probability.
Common Mistakes
- Forgetting that there are two outer tails; subtracting from 1 without accounting for both tails gives the wrong result.
- Not using the symmetry of the outer tails .
- For an answer given (AG) question, not showing enough numerical working to reach .
Things to Be Careful About
- The interval mirrors because both are three units from .
- Since is continuous, , so .
- The question says to use only the symmetry information; do not introduce a formula for .
It is now given that, for in a suitable domain,
where is a constant.
Find the value of .
Approach
Since , integrate the given PDF over the interval to , set the result equal to , and solve for .
Working
For a continuous random variable, probability is the integral of the PDF:
Integrate term by term:
Evaluate from to :
So
Thus
Answer
k = 3/256
Walkthrough
We know . For a continuous random variable, this probability is the integral of the PDF over that interval. So integrate from to and set it equal to . First integrate term by term: . Evaluate at and and subtract. The result is . Then , so . This uses the given probability on the interval rather than the total area over the whole domain.
Key Takeaways
A PDF integrates to 1 over its whole domain, and the probability of an interval is the integral of the PDF over that interval. A known probability can be used to determine the normalising constant .
Common Mistakes
- Forgetting to include the factor when integrating.
- Making arithmetic errors when evaluating the definite integral, especially with fractions.
- Equating the integral to 1 instead of to .
Things to Be Careful About
- Use the correct limits and .
- Simplify the fraction carefully: .
- The mark scheme also accepts alternative limits if the total area is used, but with the information given, integrating over to is the direct method.
A different random variable has probability density function . The domain of is all values of for which .
Find .
Approach
First find the domain by solving . Then use symmetry to find the mean, integrate to find , and use .
Working
Solve :
so or . Since between these roots, the domain is .
The function is symmetric about , so the mean is
Now
Integrate:
At :
At :
Therefore
Then
Answer
Var(X) = 9/20 = 0.45
Walkthrough
First find where . Since is a quadratic with a negative leading coefficient, it is nonnegative between its roots. Solve to get and , so the domain is .
Next, note that is symmetric about , and the domain is also symmetric about . Therefore the mean is ; no integration is needed for the mean.
To find , integrate over the domain:
Integrate term by term and evaluate at and . At the antiderivative value is ; at it is . The difference is , and multiplying by gives .
Finally use the variance formula:
Key Takeaways
The domain of a PDF is where the density is nonnegative. Symmetry of a PDF and its domain can give the mean without integration. The variance is , not alone.
Common Mistakes
- Choosing the wrong domain, such as or instead of .
- Forgetting the factor when integrating .
- Using as the variance without subtracting the square of the mean.
- Sign errors when evaluating at , particularly for .
Things to Be Careful About
- At , , so .
- The mean is , not .
- The mark scheme requires both and the mean squared to be numerical before subtracting.
- The final answer may be written as or .
The heights, in centimetres, of adult females in Litania have mean and standard deviation . It is known that in 2004 the values of and were 163.21 and 6.95 respectively. The government claims that the value of this year is greater than it was in 2004. In order to test this claim a researcher plans to carry out a hypothesis test at the 1% significance level. He records the heights of a random sample of 300 adult females in Litania this year and finds the value of the sample mean.
Approach
A Type I error is rejecting the null hypothesis when it is in fact true. In a hypothesis test, the probability of making this error is exactly the significance level.
Working
The test is carried out at the 1% significance level.
Answer
This is the same as 1%.
0.01 or 1%
Walkthrough
The significance level of a hypothesis test is defined as the probability of rejecting the null hypothesis when it is actually true. That is precisely a Type I error. Since the researcher is using a 1% significance level, the probability of a Type I error is simply 0.01.
Key Takeaways
- A Type I error occurs when a true null hypothesis is rejected.
- The probability of a Type I error is fixed by the chosen significance level.
- Here the significance level is 1%, so the answer is 0.01.
Common Mistakes
- Writing an inequality such as instead of the exact value is not accepted here.
- Confusing Type I error with Type II error, which concerns failing to reject a false null hypothesis.
Things to Be Careful About
- Do not involve the sample size or the standard deviation in part (a); the Type I error probability is determined solely by the significance level.
- Give the answer as 0.01, not 0.05 or 0.025.
You should assume that the value of after 2004 remains at 6.95 .
Given that the value of this year is actually 164.91, find the probability of a Type II error.
Approach
Let be the sample mean height. Since is known and the sample is large,
Under the null hypothesis and the alternative . A Type II error is failing to reject when in fact , so it is the probability that falls below the critical value of the rejection region.
Working
Find the critical value for the rejection region . At the 1% significance level,
Therefore
The rejection region is .
If the true mean is , the probability of a Type II error is
Standardising using :
Thus
Answer
0.0275
Walkthrough
To test the claim that the population mean is greater than in 2004, set
Since is known, the sample mean is normally distributed with mean and standard deviation .
A Type II error occurs when the test fails to reject a false null hypothesis. Here we are told that the true mean is , so we need the probability that the sample mean falls in the acceptance region of the test.
First find the boundary of the rejection region. For a one-tailed test at the 1% significance level, the critical -value is . Equating this to the standardised critical sample mean gives
and hence
The rejection region is therefore , and the acceptance region is .
Now, if the true mean is , the Type II error probability is
Standardise this boundary using the true mean :
Thus
Key Takeaways
- A Type I error has probability exactly equal to the significance level.
- The critical region is calculated using the null hypothesis mean.
- A Type II error is calculated under the true/alternative mean, not the null mean.
- For a sample mean test with known variance, the standard error is .
Common Mistakes
- Using the two-tailed critical value, such as or , instead of the one-tailed value .
- Finding the critical value using instead of .
- Finding the Type II probability using ; it must be found using the given true value .
- Forgetting to divide by when standardising.
Things to Be Careful About
- The rejection region is ; the Type II probability is the probability of the acceptance region , because the test fails to reject there.
- One-tailed 1% significance corresponds to , not .
- Since the boundary was rounded to , the standardised value is approximately ; the final probability may be quoted as to depending on rounding.