Mathematics 9709/65 — May/June 2025
Cambridge A-Level · Probability & Statistics 2 · worked solutions for every part, with the mark scheme
Topics Linear Combinations of Random Variables · Sampling and Estimation · The Poisson Distribution · Hypothesis Tests · Continuous Random Variables
The random variable has the distribution . The sum of three independent values of is denoted by .
Approach
Since is the sum of three independent values of , and , we have . To find , add the probabilities for .
Working
Answer
0.342 (3 sf)
Walkthrough
First, identify the distribution of . Since and is the sum of three independent values of , . This is because the sum of independent Poisson variables is Poisson with mean equal to the sum of the means.
To find , we need . Using the Poisson formula , each term contributes:
Evaluate the fractions: and . Then multiply each term by and add to get to 3 significant figures.
Key Takeaways
The sum of independent Poisson random variables is Poisson, with parameter equal to the sum of the individual parameters. A cumulative Poisson probability is found by summing individual probabilities from up to the required value.
Common Mistakes
Using instead of for . Missing the term in the cumulative sum. Rounding intermediate values too early. In the mark scheme, an unsupported answer of only scores one mark, so the expression must be shown.
Things to Be Careful About
The mean of is , not . Include all terms . Give the final answer to 3 significant figures. Show the Poisson expression to earn full marks.
Approach
Write down the cumulative probabilities up to 1 and up to 2 for , then divide the two expressions exactly.
Working
Divide:
Answer
125/44
Walkthrough
We need the exact ratio . First write both cumulative probabilities using .
For :
For :
The factor cancels when we divide. Since , dividing by is the same as multiplying by :
This is the exact value required.
Key Takeaways
When taking ratios of Poisson probabilities, the common exponential factor cancels, so exact fractions can be used. The sum of independent Poisson variables is Poisson, so .
Common Mistakes
Using the wrong parameter for . Omitting the term in either cumulative probability. Using decimals too early and failing to obtain the exact fraction . Not showing both cumulative expressions and the division, which are required by the mark scheme.
Things to Be Careful About
This is a "show that" question, so the final answer must be exact and convincingly obtained. The mark scheme requires both and expressions, with no end error and , before the division is attempted. Avoid decimal approximations in the working.
A random sample of 200 values of a random variable gives the following results.
Approach
Let . The coded sums are and . Estimate the mean and unbiased variance of , then form a 95% confidence interval using .
Working
Estimate the mean:
Estimate the unbiased variance of (and hence of ):
The 95% confidence interval is:
Compute the margin:
So the interval is:
Answer
2.29 to 2.31 (3 s.f.)
Walkthrough
The data are given in coded form: instead of listing the -values, we are told and . This is a common way to simplify calculations. Define . Then , so the estimate of the mean of is .
Next, estimate the variance. The sample variance formula uses . Since we need an unbiased estimate, divide by . This gives
Adding a constant to a variable does not change its variance, so this is also the unbiased estimate of the variance of .
For a 95% confidence interval, use with . The sample size is large, so the sample mean is approximately normal by the Central Limit Theorem. Substituting gives
so the interval is , which rounds to .
Key Takeaways
- Coded data and can be used to find the sample mean and variance.
- The unbiased variance estimate uses divisor .
- A confidence interval for the mean has the form .
- For large samples, the CLT justifies using the normal -value even when the population distribution is unknown.
Common Mistakes
- Using the biased variance (dividing by instead of ); the mark scheme gives M0 for a biased estimate.
- Forgetting to add 2 back when converting to .
- Using a -value or the wrong -value instead of for 95% confidence.
- Giving only one endpoint of the interval instead of the full interval.
Things to Be Careful About
- The variance estimate must be unbiased: use .
- The expression is ; do not forget the division by .
- The final answer must be an interval, not just a point estimate.
- Rounding to 3 significant figures gives to ; keep more decimal places in intermediate working.
- If the biased variance is used, the mark scheme allows at most 4 of the 6 marks.
Approach
Identify the step in part (a) where the normal distribution is used.
Working
In part (a), the confidence interval is constructed as
The value is the critical value from the standard normal distribution. The Central Limit Theorem justifies using this normal value for the sample mean because the sample size is large.
Answer
The use of in part (a).
Use of z = 1.96 in part (a).
Walkthrough
Part (b) asks you to state where the Central Limit Theorem was used in part (a). In part (a), the confidence interval was built using , the critical value from the standard normal distribution. This is the step that relies on the CLT: because is large, the distribution of the sample mean is approximately normal, regardless of the distribution of . Therefore the standard normal value can be used.
Key Takeaways
- The CLT is what allows us to use normal-theory confidence intervals for large samples even when the population is not normal.
- The phrase 'use of ' is the expected answer, because the -value comes from the normal distribution.
Common Mistakes
- Saying 'X is not normally distributed' — this is not the correct identification.
- Saying 'Assume is normally distributed' — the CLT does not require assuming; it justifies the approximation for large samples.
- Not mentioning or at all.
Things to Be Careful About
- The mark scheme requires a reference to the -value, not a general statement about normality.
- The CLT applies to the sample mean, not to itself.
Maroulla’s calculator can generate random numbers between 0.000 and 0.999 inclusive, correct to 3 significant figures. She plans to use her calculator to choose a sample of members from the 851 members in her health club. She numbers the members from 1 to 851. Then she uses her calculator to generate some random numbers. She multiplies each random number by 851 and rounds up to the next whole number to give the number of a member in the sample. This is called a ‘member number’.
Maroulla’s first random number is 0.401.
Find the member number that is produced by this random number.
Approach
Multiply the random number by 851 to scale it across the 851 members, then round the result up to the next whole number.
Working
The scaled value is:
Rounding up to the next whole number gives:
Answer
342
Walkthrough
The first random number represents a point in the interval to . Maroulla multiplies by , the number of members, to turn this into a member number. Here . The instruction is to round up, so the next whole number above is . Therefore the member selected is member 342.
Key Takeaways
This part tests the mechanical step of turning a random decimal into a member number by scaling and rounding up. It also establishes the context for checking whether each member is equally likely to be chosen.
Common Mistakes
- Rounding to the nearest whole number instead of rounding up: would be by normal rounding, but the question explicitly says round up, so the correct answer is .
- Miscalculating ; be careful with the decimal multiplication.
Things to Be Careful About
The phrase “round up to the next whole number” is strict: any non-integer result must go to the next integer, even if the decimal part is less than .
Find all possible random numbers, correct to 3 decimal places, that would produce the following member numbers.
Approach
Work backwards from the member number. A member number of 680 is obtained when the scaled value is greater than 679 and not more than 680, because rounding up sends every value in to 680. Solve this inequality for , then list the possible values that can be written correct to 3 decimal places.
Working
For member number 680:
Dividing by 851:
So:
The random numbers correct to 3 decimal places in this interval are:
Answer
0.798, 0.799
Walkthrough
We need the values of the random number such that rounding up gives 680. If , then rounding up gives 680; if , rounding up also gives 680. However, if exactly, rounding up gives 679. So the condition is .
Dividing each part by 851 gives the interval . The calculator can produce numbers correct to 3 decimal places, so we list the three-decimal values inside this interval. is too small because is below 679 and would round up to 679. and both lie in the interval. is too large because it would round up to 681. Thus only those two values work.
Key Takeaways
This part shows how to invert a ‘round up’ mapping: for a member number , use the strict interval for the scaled value.
Common Mistakes
- Using instead of ; this reverses which endpoint is included and gives incorrect values.
- Listing because it is close to the lower bound, without checking that .
- Including ; it is greater than the upper bound and would round up to 681.
Things to Be Careful About
Because the calculator gives numbers to 3 decimal places, list only three-decimal numbers. Also include the upper endpoint, since a scaled value of exactly 680 is allowed. The mark scheme accepts only and .
Approach
Use the same reverse mapping as in part (b)(i). Member number 850 corresponds to , because rounding up sends every value in to 850.
Working
So:
The only random number correct to 3 decimal places in this interval is:
Answer
0.998
Walkthrough
For member number 850, rounding up produces 850 only when the scaled value is above 849 and not more than 850. Hence . Dividing by 851 gives . The three-decimal values in this interval are checked: is too small, since and rounding up gives 849. lies in the interval. is too large, since and rounding up gives 851. So only is possible.
Key Takeaways
The same interval argument applies for any member number: a member comes from the scaled interval . Here the interval is narrow enough to contain only one three-decimal random number.
Common Mistakes
- Including , forgetting that is already greater than 850.
- Missing that is below the lower bound even though it looks close to 0.998.
- Using instead of at the upper end, which would incorrectly exclude .
Things to Be Careful About
The upper endpoint is included: a scaled value of exactly 850 is allowed because rounding 850 up gives 850. The lower endpoint is excluded: a scaled value of exactly 849 would round up to 849, not 850.
Explain briefly how your answers to part (b) show that Maroulla’s method does not produce a random sample.
Approach
A random sample requires every member to be equally likely to be selected. Compare the number of random numbers that produce member 680 with the number that produce member 850.
Working
From part (b)(i), member 680 can be selected by:
From part (b)(ii), member 850 can be selected by only:
Since every random number is equally likely, member 680 is twice as likely to be selected as member 850.
Answer
Member 680 and member 850 are not equally likely to be chosen, so the selection method does not produce a random sample.
680 and 850 are not equally likely to be chosen.
Walkthrough
A random sample should give every member of the population the same chance of being chosen. In part (b), the two member numbers have different numbers of random numbers that produce them: 680 can be produced by 0.798 and 0.799, while 850 can be produced only by 0.998. If each random number is equally likely, then member 680 is twice as likely to be selected as member 850. Therefore the members are not equally likely to be chosen, so the method is not a random sample.
Key Takeaways
The crucial idea is that random inputs do not automatically give a random sample; the mapping from inputs to members must also give equal chances to every member. A method is flawed if some members correspond to more input values than others.
Common Mistakes
- Saying the calculator is not truly random; this is not the reason and would not gain the mark.
- Giving a vague answer such as “rounding makes it unfair” without referring to part (b). The mark scheme requires the answer to relate to part (b), for example by noting the different numbers of possible random numbers.
Things to Be Careful About
The mark scheme allows follow-through if the numbers of solutions in part (b)(i) and part (b)(ii) are different. The explanation must compare those counts and conclude that the member numbers are not equally likely to be chosen.
The numbers of cars and trucks arriving per minute at a fuel station are modelled by independent variables with distributions and respectively.
Find the probability that at least 4 cars and at least 2 trucks arrive at the fuel station during a randomly chosen 5-minute period.
Approach
For a 5-minute period, the numbers of cars and trucks are Poisson with means and . Compute each at least probability by taking the complement of the cumulative Poisson, then multiply by the variables are independent.
Working
Let be the number of cars in 5 minutes, so .
Let be the number of trucks in 5 minutes, so .
Since the arrivals are independent,
Answer
0.404
Walkthrough
For the 5-minute period, multiply each per-minute rate by 5: the cars have mean , and the trucks have mean . To find , use the complement because Poisson tables give cumulative probabilities from 0 up to a value. Write the cumulative sum as for . Similarly, with . Since the two arrival processes are independent, the joint probability is the product of the two separate probabilities.
Key Takeaways
The Poisson mean scales with the time interval. At least probabilities are easiest as complements. For independent events, .
Common Mistakes
Using the per-minute rate instead of the 5-minute rate. Forgetting to include all terms in the cumulative sum. Calculating only instead of . Adding probabilities instead of multiplying them. Omitting the full cumulative expression, which is needed for the method mark.
Things to Be Careful About
Be sure to include all terms up to for cars and up to for trucks. Keep enough decimal places in intermediate steps so the final answer is correct to 3 significant figures. At least 4 means , so the complement is .
Use a suitable approximating distribution to find the probability that a total of fewer than 145 cars and trucks arrive at the fuel station during a randomly chosen 2-hour period.
Approach
The total number of arrivals per minute is the sum of two independent Poisson variables, so it is Poisson with rate per minute. Over a 2-hour period (120 minutes), the mean is . Because is large, approximate the Poisson distribution by a normal distribution with mean 156 and variance 156. Use a continuity correction because the Poisson variable is discrete, then standardise and use the normal distribution table.
Working
Total arrivals per minute: .
Over 120 minutes, .
Approximation: .
We need . Since is discrete, fewer than 145 means . Applying the continuity correction:
Then
From standard normal tables, , so
Answer
(Alternatively, depending on interpolation.)
0.179 (or 0.178)
Walkthrough
First, combine the two rates: the total number of arrivals per minute is Poisson with parameter . Over 120 minutes, the mean becomes . Because is large, the normal approximation is suitable. The Poisson variable is discrete, so fewer than 145 means ; the continuity correction moves the upper boundary to 144.5. Standardise using . The normal table gives left-tail probabilities, so , and the result is approximately 0.179.
Key Takeaways
The sum of independent Poisson variables is Poisson with parameter equal to the sum of the parameters. For large , a Poisson distribution can be approximated by a normal distribution with the same mean and variance. A continuity correction is required when approximating a discrete distribution by a continuous one.
Common Mistakes
Forgetting to scale the per-minute rate by 120, leaving . Using 145 instead of 144.5 in the continuity correction. Using the wrong standard deviation, for example instead of . Looking up and forgetting to subtract it from 1.
Things to Be Careful About
Fewer than 145 means , so the continuity-corrected boundary is 144.5, not 145.5. The normal approximation is valid because is large. Remember that . The final answer is acceptable as 0.179 or 0.178 depending on interpolation.
Candidates for a certain diploma take two tests. Their marks for the first test and the second test are modelled by the independent variables with distributions and respectively. The final mark, , for each candidate is found by doubling the mark in the first test and adding the result to the mark in the second test.
Find the probability that the mean, , of the final marks of a random sample of 25 candidates is greater than 143.
Approach
Let be the first-test mark and the second-test mark. Since and are independent normal variables and , is also normal. Find and , then use the sampling distribution of the sample mean for a sample of 25. Standardise and find the upper-tail probability.
Working
So .
For a random sample of 25 candidates, the sample mean has
Standardise the value 143:
Therefore
Answer
0.0754
0.0754
Walkthrough
We are told the two test marks are independent normal variables. The final mark is , where is the first-test mark and is the second-test mark. Because a linear combination of independent normal variables is normal, itself is normal. We only need its mean and variance.
The mean is found by taking the linear combination of the means: .
For the variance, the coefficient 2 must be squared: . The independence of and is what allows us to add the variances with no covariance term.
Now we consider the mean of a random sample of 25 final marks. The sample mean has the same mean as an individual final mark, , but its variance is divided by the sample size: . Since the population distribution of is normal, is exactly normal.
To find , standardise 143 by subtracting the mean and dividing by the standard deviation of :
The required probability is the area to the right of this -value:
No continuity correction is needed because is a continuous normal variable.
Key Takeaways
This question combines two ideas: linear combinations of independent normal variables and the sampling distribution of the sample mean. When a variable is a linear combination of independent normals, it remains normal, and its mean and variance are obtained by linear combination rules. For a sample mean, the mean is unchanged but the variance is divided by . Finally, probabilities about the sample mean are found by standardising and using the standard normal table.
Common Mistakes
A common mistake is to forget to square the coefficient 2 when finding the variance, writing instead of . Another is to forget to divide the variance by 25 when moving from an individual final mark to the sample mean. Some candidates standardise incorrectly by using the individual variance instead of . Also, ensure the correct tail is used: the question asks for greater than 143, so the probability is , not .
Things to Be Careful About
The distribution of is normal because it is a linear combination of independent normal variables. The sample mean is also exactly normal because the population is normal, so no Central Limit Theorem approximation is required. The mark scheme requires in the standardisation. A continuity correction is not required for a continuous normal variable, although the mark scheme allows an attempted one. Give the final probability to 3 significant figures, so 0.0754.
It is claimed that 28% of voters in a certain town support the Forward Now political party. A researcher suspects that the true figure is less than 28%. She interviews a random sample of 30 voters from the town and she finds that 4 voters in the sample say that they support the Forward Now party. She plans to carry out a hypothesis test at the 10% significance level.
Approach
Set up the hypotheses for a one-tailed lower-tail test. Let be the number of supporters in the sample. Under , . Compute and compare it with the 10% significance level.
Working
Let be the true proportion of voters supporting Forward Now.
Let be the number of voters in the sample of 30 who support Forward Now. Under , .
The observed value is . For a lower-tail test:
Since , the result is significant at the 10% level. We reject .
Answer
Reject . There is sufficient evidence at the 10% significance level to suggest that the percentage of voters supporting Forward Now is less than 28%.
Reject H₀. There is sufficient evidence at the 10% significance level to suggest that the percentage of voters supporting Forward Now is less than 28%.
Walkthrough
This is a hypothesis test about a population proportion, using the binomial distribution.
Step 1: State the hypotheses. The claim is that 28% of voters support Forward Now, so the null hypothesis is . The researcher suspects the true figure is less, so the alternative hypothesis is . This is a one-tailed test because the alternative is directional.
Step 2: Identify the test statistic and its distribution. Let be the number of voters in the sample who support Forward Now. Under , each voter independently has probability 0.28 of supporting the party, so .
Step 3: Compute the tail probability. Since is , we look at the lower tail. The observed value is , so we compute . This is the sum of through , each found using the binomial probability formula .
Step 4: Compare with the significance level. The tail probability is 0.0495, which is less than 0.10. This means the observed result is unlikely (less than 10% likely) if is true, so the result is significant.
Step 5: Draw a conclusion. Since the result is significant, we reject and conclude that there is sufficient evidence to suggest the true percentage is less than 28%. The conclusion must be stated in context.
Key Takeaways
- Setting up hypotheses correctly: is the claim being tested, is the suspicion.
- Using the binomial distribution for a hypothesis test about a proportion.
- Computing a tail probability as a sum of individual binomial probabilities.
- Comparing the tail probability with the significance level to decide whether to reject .
- Stating the conclusion in context, without being definitive.
Common Mistakes
- Using a two-tailed test when the alternative is clearly one-tailed ().
- Computing instead of .
- Forgetting to compare the probability with 0.10.
- Stating the conclusion without context (e.g., just "reject H₀" without mentioning the percentage of voters).
- Saying the percentage "is" less than 28% rather than "there is sufficient evidence to suggest" it is less.
Things to Be Careful About
- The significance level is 10%, so the comparison is with 0.10.
- The probability must be a tail probability (), not a single point probability.
- The conclusion should not be over-stated: we have evidence, not proof.
- The mark scheme requires the expression or terms to be seen for the method mark — don't just write the final answer.
State, with a reason, whether it is possible that a Type I error was made in carrying out the test.
Approach
Recall the definition of a Type I error and apply it to the conclusion from part (a).
Working
A Type I error occurs when is true but is rejected. In part (a), was rejected. Therefore, if in reality the true proportion is 28%, a Type I error would have been made.
Answer
Yes, it is possible that a Type I error was made, because was rejected.
Yes, because H₀ was rejected.
Walkthrough
A Type I error is defined as rejecting when is actually true. In part (a), we rejected . So if the true proportion of voters supporting Forward Now really is 28%, then we have made a Type I error. Since we don't know the true proportion, we can't say for certain whether a Type I error was made — but it is possible.
Key Takeaways
- A Type I error can only occur when is rejected.
- The probability of a Type I error is controlled by the significance level.
Common Mistakes
- Saying "no, because the result was significant" — this is wrong; significance doesn't rule out Type I error.
- Confusing Type I error (rejecting a true H₀) with Type II error (failing to reject a false H₀).
Things to Be Careful About
- The answer depends on the conclusion from part (a). If part (a) had not rejected , the answer would be "no" (a Type I error would not be possible because H₀ was not rejected).
Later the researcher carries out a similar test at the 10% significance level, using a new random sample of 30 voters from the town.
Find the probability of a Type I error.
Approach
Determine the critical region for the test at the 10% significance level, then find the probability of falling in the critical region under .
Working
The critical region is the set of values of for which .
So the critical region is .
Answer
0.0495
Walkthrough
The probability of a Type I error is the probability of rejecting when is true. To find this, we need to know the critical region — the set of values of the test statistic that lead to rejection.
For a lower-tail test at the 10% significance level, the critical region is , where is the largest value such that .
From part (a), we know , which is less than 0.10. But we must also check whether 5 belongs in the critical region. Computing , which is greater than 0.10, tells us that 5 is not in the critical region.
Therefore the critical region is , and the probability of a Type I error is .
Key Takeaways
- The actual probability of a Type I error (the size of the test) may be less than the nominal significance level.
- To find the critical region, you must check the cumulative probabilities until they exceed the significance level.
Common Mistakes
- Just stating 0.0495 without showing that — the mark scheme requires this for full credit.
- Confusing the significance level (0.10) with the actual probability of a Type I error (0.0495).
Things to Be Careful About
- The probability of a Type I error is the probability under of falling in the critical region.
- You must show that to justify that the critical region is .
A firm makes a certain type of battery-powered toy. The battery life is denoted by hours and the population mean of is supposed to be 12. The Quality Control department wished to test whether the population mean of is actually less than 12. They tested a random sample of 50 of these toys and found that the sample mean, , was 11.4.
Approach
State the null hypothesis as the current assumed value of the population mean, and the alternative hypothesis as the claim being tested: that the population mean is less than 12.
Working
Let be the population mean battery life.
Answer
, .
H0: μ = 12, H1: μ < 12
Walkthrough
We are testing whether the population mean battery life is less than 12 hours. The null hypothesis always represents the status quo or the value currently assumed, so . The alternative hypothesis represents the claim we want to provide evidence for, so . This is a one-tailed test because the alternative is directional (less than).
Key Takeaways
A null hypothesis is a statement about a population parameter, not a sample statistic. The alternative hypothesis is determined by the wording of the question: here 'actually less than 12' means a lower-tail one-tailed test.
Common Mistakes
Writing or . Hypotheses must be about the population mean , not the sample mean , and must use strict inequality in the alternative.
Things to Be Careful About
The mark scheme requires 'population mean', not just 'mean'. Use and with strict inequality.
You may assume that the standard deviation of the battery life is 2.3 hours.
Show that the value leads to rejection of the null hypothesis at the 5% significance level.
Approach
Since the sample size is large (), the sample mean is approximately normally distributed. Standardise the observed sample mean using the assumed population mean and the known standard deviation, then compare the test statistic with the lower 5% critical value.
Working
With , , and :
For a one-tailed test at the 5% significance level, the critical value is:
Since , the observed value lies in the critical region. Equivalently, the -value is .
Answer
leads to rejection of at the 5% significance level.
Reject H0 at the 5% significance level
Walkthrough
We are testing against . The sample mean is approximately normal with mean and standard deviation , because the sample size is large. We convert the observed to a -score: . For a lower-tail test at 5%, the critical value is . Since is further from 0 than , it falls in the rejection region, so we reject .
Key Takeaways
For a test about a population mean with known standard deviation and a large sample, use . Compare the test statistic with the critical value from the standard normal distribution, or compare the -value with the significance level.
Common Mistakes
Forgetting to divide by when standardising. Using as the standard deviation of instead of . Using the wrong tail: the alternative requires the lower tail.
Things to Be Careful About
The mark scheme requires the in the denominator. The comparison must be correct and consistent: (or , or ). Do not compare absolute values inconsistently.
It is given that the value leads to rejection of the null hypothesis at the % significance level.
Find the set of possible values of .
Approach
Find the probability of obtaining a sample mean as low as 11.4 under , i.e. the -value. Rejection at the significance level occurs when this -value is smaller than .
Working
Using the test statistic from part (b), :
As a percentage:
Rejection at the level requires:
Answer
(or ).
α > 3.25 (or α ≥ 3.25)
Walkthrough
The -value is the probability, assuming is true, of observing a sample mean as extreme as 11.4 or more in the direction of the alternative. Since the test is lower-tailed, . Using the standard normal table, , so . We reject whenever the significance level is larger than the -value. In percentage terms, this means . If the significance level is exactly 3.25%, the result is borderline; the mark scheme accepts as well.
Key Takeaways
Rejection occurs when the -value is less than the significance level. The significance level can be written as a decimal or a percentage, but the comparison must be consistent. The threshold significance level is exactly the -value.
Common Mistakes
Confusing the -value with the significance level. Writing instead of . Forgetting to convert to before comparing with .
Things to Be Careful About
The mark scheme accepts or . The -value should be given to 3 significant figures as or . If you only write without converting, make sure the inequality is still correct.
The diagram shows the graph of the probability density function of a random variable . Between and the graph consists of a straight line through with gradient , where and are positive constants. Elsewhere .
It is given that the median of is .
Approach
The median of a continuous random variable is the value such that the probability . This corresponds to the area under the probability density function (PDF) from the lower bound to being equal to . Here, the lower bound is and the median is given as . We can use the area of the triangle formed by the graph or integrate the function from to .
Working
The function is for . At , the value of the function is .
Using the area of the triangle under the graph from to :
Since the median is , this area must be equal to :
Alternatively, using integration:
Setting this equal to gives .
Answer
k = 1/2
Walkthrough
The median of a continuous random variable is the value that splits the total probability (area under the PDF) into two equal halves. Since the PDF is defined for , the area from to the median must be exactly .
The graph from to is a right-angled triangle with base and height . The area of this triangle is . Setting this area equal to immediately gives .
Alternatively, you can integrate the function from to :
Equating this to yields the same result: .
Key Takeaways
- The median of a continuous random variable satisfies .
- The area under a PDF can be calculated using geometric formulas (like the area of a triangle) or definite integration.
Common Mistakes
- Assuming the total area is calculated up to before finding . The median condition only involves the area up to .
- Forgetting that the area under the entire PDF must be , which is needed in part (b).
Things to Be Careful About
- Ensure the limits of integration match the region for the median (from to , not to ).
- The gradient must be positive, and the area must be positive.
Approach
First, we find the value of using the property that the total area under the PDF is . Once is known, we can calculate the expected value using the formula .
Working
The total area under the PDF from to is . The shape is a triangle with base and height (using from part (a)).
Using the area of the triangle:
Since is positive, .
Alternatively, using integration:
Now, calculate the expected value :
Evaluate the integral:
Answer
E(X) = 4/3
Walkthrough
The total probability for any valid PDF must sum to . This means the total area under the graph from to is .
The graph is a triangle with base and height . Since and we found , the height is .
Area = .
Setting this equal to : (since ).
Now we find the expected value . The formula for the mean of a continuous random variable is:
Since outside , the limits are to :
Integrating:
Key Takeaways
- The total area under a PDF is always . This is used to find unknown constants like the upper bound .
- The expected value (mean) is calculated as .
- Always multiply by the PDF inside the integral for expected value.
Common Mistakes
- Forgetting to multiply by when calculating . The integral gives probability (which is ), not the mean.
- Using the wrong limits of integration. The PDF is only non-zero between and .
- Arithmetic errors in integrating (remember , not ).
Things to Be Careful About
- Ensure you use the value of found in part (a) correctly.
- The variable must be positive.
- Check that the final answer for is within the possible range of (i.e., ). Here , which is valid.
