Mathematics 9709/51 — October/November 2024
Cambridge A-Level · Probability & Statistics 1 · worked solutions for every part, with the mark scheme
Topics Discrete Random Variables · Probability · The Normal Distribution · Representation of Data · Permutations and Combinations
Nicola throws an ordinary fair six-sided dice. The random variable is the number of throws that she takes to obtain a 6.
Approach
Since is the number of throws until the first 6, has a geometric distribution with success probability . The event means the first 6 appears on one of the first 7 throws. It is easier to use the complement: means no 6 appears in the first 7 throws.
Working
No 6 in any of the first 7 throws has probability:
Therefore:
Answer
0.721
Walkthrough
Nicola keeps throwing until she gets a 6, so each throw is independent with probability of success and of failure. The random variable is geometric. For , we need the first 6 to happen on throw 1, 2, 3, 4, 5, 6, or 7. Instead of adding seven geometric probabilities, use the complement: means she fails to get a 6 on all of the first 7 throws. That probability is . Subtracting from 1 gives the required probability. The exact value is , which rounds to 0.721.
Key Takeaways
- A waiting time for the first success is modelled by a geometric distribution.
- The complement is often quicker than summing many terms.
- includes throws 1 through 7, not 1 through 8.
Common Mistakes
- Using instead of , which would calculate .
- Adding only six terms instead of seven.
- Forgetting to subtract from 1 and giving .
- Writing only the decimal without showing the complement method may lose method marks; the mark scheme gives only special-case credit for an unsupported correct value.
Things to Be Careful About
- The inequality is strict: excludes .
- The complement event is no 6 in the first 7 throws.
- Give the answer to 3 significant figures as 0.721, or the exact fraction .
Approach
For the second 6 to occur on the 8th throw, the 8th throw must be a 6, and exactly one of the first 7 throws must also be a 6. Choose which of the first 7 throws is the first 6, then multiply the probabilities of the 8 independent throws.
Working
Choose the position of the first 6 among the first 7 throws:
There is one 6 and six non-6s in the first 7 throws, and the 8th throw is a 6:
Answer
0.0651
Walkthrough
We want the second 6 to occur exactly on the 8th throw. This means the 8th throw must be a 6, and among the first 7 throws there must be exactly one 6. The first 6 could be in any of the 7 positions, so there are choices. For a fixed choice, the probability of one 6 and six non-6s in the first 7 throws is , and the 8th throw being a 6 contributes another . Multiplying gives , which is approximately 0.0651.
Key Takeaways
- For the second success to occur on a given throw, fix the last throw as a success and place the first success among the earlier throws.
- Use a binomial coefficient to count the positions of the earlier success.
- This is a negative-binomial/geometric-type counting problem.
Common Mistakes
- Using instead of , which would allow the second 6 to occur before the 8th throw.
- Forgetting to multiply by 7, giving only the probability of one particular arrangement.
- Writing for the first success on the 8th throw instead of the second success.
- Adding probabilities instead of multiplying independent throw probabilities; the mark scheme allows no inappropriate addition.
Things to Be Careful About
- The 8th throw must definitely be a 6, so it is not part of the choice.
- Exactly one of the first 7 throws is a 6; the other six are not 6s.
- Round the final probability to 0.0651, or give the exact fraction .
The random variable takes the values . It is given that , where is a positive constant.
Draw up the probability distribution table for , giving the probabilities as numerical fractions.
Approach
For a probability distribution, the sum of all the probabilities is . Substitute each value of into , sum the resulting expressions, and solve for . Then present the probabilities as fractions.
Working
For the values :
Since the probabilities sum to :
The probability distribution table is:
Equivalently, the probabilities are .
Answer
k = 1/28; probabilities are 6/28, 3/28, 2/28, 6/28, 11/28
Walkthrough
The stem says that takes the values and that . To find a valid probability distribution, substitute each value of into the formula. This gives the probabilities . Since the total probability must be , add these five probabilities and set the sum equal to . This gives , so . Substituting this value back into each probability gives the required numerical fractions. The mark scheme awards the first mark for identifying , a method mark for a table with the correct outcomes and at least two correct probabilities, and the final mark for a fully correct table.
Key Takeaways
A probability distribution table must list all possible outcomes and their probabilities, and the probabilities must sum to . This question uses the total-sum condition to find an unknown constant in a probability distribution. Once is found, each probability can be written as a fraction.
Common Mistakes
- Forgetting to include all five possible values of in the sum, giving an incorrect total such as .
- Leaving the probabilities in terms of instead of substituting .
- Including extra values of that are not in the given set. The mark scheme does not allow extra values unless their probability is .
- Rounding probabilities to decimals too early; the question asks for numerical fractions.
Things to Be Careful About
- The total probability must equal exactly , so is found from .
- The mark scheme accepts a table with probabilities written as , , , , or etc. as part of the method, but the final table should contain numerical fractions.
- Decimal answers are accepted by the mark scheme only if given to at least 3 significant figures, but the question specifically asks for fractions.
Approach
Use the probability distribution from part (a). First calculate , then calculate , and use .
Working
Then
Thus .
Answer
825/196 = 4 41/196 ≈ 4.21
Walkthrough
Using the distribution from part (a), first find the expected value by summing each value of multiplied by its probability. This gives . Next find by summing each value of multiplied by its probability. This gives . The variance is then . The mark scheme awards a method mark for the calculation of , a second method mark for using the correct variance formula, and the final mark for the simplified answer.
Key Takeaways
- The expectation of a discrete random variable is .
- The variance is .
- When probabilities are fractions, keeping all terms over a common denominator avoids rounding errors.
- The variance may be written as an improper fraction, a mixed number, or a decimal.
Common Mistakes
- Using instead of .
- Forgetting to square the values of when computing .
- Forgetting to multiply each by its probability.
- Making arithmetic errors when subtracting the fractions.
Things to Be Careful About
- The method marks allow follow-through from the table in part (a), provided the probabilities are between and and sum to .
- The mark scheme accepts the variance as , , or .
- If the final answer is correct with no working, the mark scheme gives a special case mark of B1, but full method marks require showing both and the variance formula.
The time taken, in minutes, to walk to school was recorded for 200 pupils at a certain school. These times are summarised in the following table.
| Time taken ( minutes) | ||||||
|---|---|---|---|---|---|---|
| Cumulative frequency | 18 | 46 | 88 | 140 | 176 | 200 |
Approach
To draw the cumulative frequency graph, we plot the cumulative frequency against the upper bound of each time interval. We start at because at time , the cumulative frequency is 0. The points to plot are , , , , , and . A smooth S-shaped curve (ogive) is drawn through these points, starting from .
Working
The coordinates to plot are:
Plot these points on the grid with time on the horizontal axis (0 to 70) and cumulative frequency on the vertical axis (0 to 200). Join them with a smooth curve.
Answer
Cumulative frequency graph drawn with points and a smooth curve.
Graph plotted with points (0,0), (15,18), (25,46), (30,88), (40,140), (50,176), (70,200) and smooth curve
Walkthrough
A cumulative frequency graph (or ogive) is used to represent cumulative data. The horizontal axis represents the variable (time in minutes), and the vertical axis represents the cumulative frequency (number of pupils).
- Identify points: The table gives cumulative frequencies at upper bounds. We add as the starting point. The points are , , , , , .
- Plot and draw: Plot these points on the grid. Join them with a smooth curve that starts at and ends at . The curve should be S-shaped, generally increasing.
Key Takeaways
- Cumulative frequency graphs plot cumulative frequency against upper class boundaries.
- Always include the origin if the data starts from 0.
- The curve must be smooth, not a series of straight line segments (unless specified as a polygon).
Common Mistakes
- Forgetting to plot the point .
- Joining points with straight lines instead of a smooth curve.
- Incorrect axis scales (e.g., not using enough of the grid).
Things to Be Careful About
- Ensure the axes are labelled correctly: 'cumulative frequency' and 'time taken / minutes'.
- The scale must be linear and large enough to use at least half the grid.
Approach
The total number of pupils is .
- The median is the value at cumulative frequency .
- The lower quartile (LQ) is the value at cumulative frequency .
- The upper quartile (UQ) is the value at cumulative frequency .
- The interquartile range (IQR) is .
Working
From the cumulative frequency graph (or using the estimated values from the mark scheme):
-
Median: Find when . Looking at the graph, .
-
Quartiles: Find when and .
- For , . So .
- For , . So .
-
IQR:
Answer
Median minutes, IQR minutes.
Median = 33, IQR = 16
Walkthrough
To estimate the median and interquartile range from a cumulative frequency graph, we use the total frequency .
-
Median: The median corresponds to the -th value, which is the 100th value. Draw a horizontal line from to the curve, then drop a vertical line to the time axis. The value is approximately 33 minutes.
-
Lower Quartile (LQ): Corresponds to . Draw a line from to the curve and down to the axis. Value is approximately 26 minutes.
-
Upper Quartile (UQ): Corresponds to . Draw a line from to the curve and down to the axis. Value is approximately 42 minutes.
-
IQR: The difference between the upper and lower quartiles: .
Note: Values read from graphs have some tolerance (typically square on the axis). The mark scheme accepts median around 33, LQ between 25-27, UQ between 41-43.
Key Takeaways
- Median is at cumulative frequency.
- Quartiles are at and cumulative frequencies.
- IQR is UQ - LQ.
- Reading from graphs requires care with scales and interpolation.
Common Mistakes
- Using instead of for the median (i.e., looking for cf=200).
- Calculating IQR as UQ + LQ or UQ * LQ.
- Reading the wrong axis (time vs frequency).
Things to Be Careful About
- Ensure you are reading the time value from the horizontal axis, not the cumulative frequency.
- Graph reading estimates will vary slightly between students; mark schemes allow tolerance.
Calculate an estimate for the mean value of the times taken by the 200 pupils to walk to school.
Approach
To estimate the mean from grouped data, we need the frequency () and midpoint () for each class interval. The frequencies are found by subtracting consecutive cumulative frequencies. The midpoints are the average of the lower and upper bounds of each interval.
Working
First, determine the class intervals, frequencies, and midpoints:
| Time interval | Cumulative Frequency | Frequency () | Midpoint () | |
|---|---|---|---|---|
| 18 | 18 | |||
| 46 | ||||
| 88 | ||||
| 140 | ||||
| 176 | ||||
| 200 |
Total frequency .
Total .
Answer
33.65
Walkthrough
The data is given in cumulative frequency form. To calculate the mean, we must convert this to a frequency distribution table.
-
Frequencies: Subtract each cumulative frequency from the next to get the frequency for that interval. For the first interval, the frequency is just the first cumulative frequency (assuming it starts from 0).
- :
- :
- :
- :
- :
- :
-
Midpoints: Calculate the midpoint for each interval as .
- .
-
Mean Calculation: Multiply each frequency by its midpoint (), sum these products, and divide by the total frequency ().
- .
- Mean .
Key Takeaways
- Cumulative frequency can be converted to grouped frequency data by subtraction.
- The mean of grouped data is estimated using class midpoints.
- Formula: .
Common Mistakes
- Using the upper bound of the interval as the midpoint instead of the average of bounds.
- Using cumulative frequencies in the mean formula instead of actual frequencies.
- Arithmetic errors in calculating .
Things to Be Careful About
- Ensure the first interval starts at 0 (implied by having cf 18, usually meaning ).
- The midpoint of is 60, not 50 or 70.
- Acceptable answers include 33.7 or .
Rahul has two bags, and . Bag contains 4 red marbles and 2 blue marbles. Bag contains 3 red marbles and 4 blue marbles. Rahul also has a coin which is biased so that the probability of obtaining a head when it is thrown is .
Rahul throws the coin.
- If he obtains a head, he chooses at random a marble from bag . He notes the colour and replaces the marble in bag . He then chooses at random a second marble from bag .
- If he obtains a tail, he chooses at random a marble from bag . He notes the colour and discards the marble. He then chooses at random a second marble from bag .
Approach
Let and denote the outcomes of the biased coin. If occurs, two draws are made from bag with replacement; if occurs, two draws are made from bag without replacement. The two marbles are the same colour in four mutually exclusive ways: , , , . Multiply along each branch and add the results.
Working
The coin probabilities are
For bag (with replacement):
For bag (without replacement):
After a red marble is taken from bag , red and blue remain, so
After a blue marble is taken from bag , red and blue remain, so
Now calculate the four same-colour probabilities:
These four events are mutually exclusive, so add them:
Using a common denominator of :
Answer
29/63 or 0.460
Walkthrough
Start with the coin. It is biased, so the probability of heads is and the probability of tails is . If heads occurs, Rahul uses bag and replaces the marble after the first draw, so the second draw has exactly the same probabilities as the first: red and blue. If tails occurs, Rahul uses bag and discards the first marble, so the second draw depends on the colour of the first marble.
The four ways to get two marbles of the same colour are: head then red-red, tail then red-red, head then blue-blue, and tail then blue-blue. For each way, multiply the probabilities along the route. Since these four routes are mutually exclusive, add them together.
Key Takeaways
This question tests the multiplication law for probabilities, the addition law for mutually exclusive events, and the important difference between sampling with replacement and sampling without replacement.
Common Mistakes
- Forgetting to include the coin probability in each branch.
- Treating bag as if the marble were replaced; after the first draw from , the second draw has denominator , not .
- Adding probabilities that are not mutually exclusive, or missing one of the four same-colour routes.
Things to Be Careful About
- In bag , with replacement, the probabilities are red and blue for both draws.
- In bag , without replacement, after a red is drawn there are red and blue left; after a blue is drawn there are red and blue left.
- Show the unsimplified products, for example , because the mark scheme awards a method mark for two clearly identified unsimplified probabilities.
- The final answer can be given as or to 3 significant figures.
Find the probability that the two marbles that Rahul chooses are both from bag given that both marbles are blue.
Approach
Let denote the event that both marbles are blue. We need , the probability that the marbles came from bag given that they are both blue. Using
The numerator is from part (a), and the denominator is .
Working
From part (a),
Also from part (a),
So the total probability that both marbles are blue is
Therefore
Answer
54/61 or 0.885
Walkthrough
This part asks for a conditional probability: the probability that the marbles came from bag , given that both are blue. Write the conditioning event as . The formula is
The numerator is the probability that the coin is a tail and both marbles are blue. This is exactly from part (a), which is . The denominator is the total probability that both marbles are blue, and it includes both the head route and the tail route: . Dividing gives the required conditional probability.
Key Takeaways
Conditional probability is a renormalisation: restrict attention to the cases where the conditioning event occurs, then compare the probability of the intersection with that total. Here the conditioning event is 'both marbles are blue', so we divide by the total probability of both blue.
Common Mistakes
- Using as the numerator instead of .
- Forgetting to include in the denominator.
- Confusing the numerator and denominator when writing the fraction.
Things to Be Careful About
- The numerator and denominator must both refer to the event 'both marbles are blue'.
- The denominator is , not just .
- The final answer can be given as or to 3 significant figures.
- If a previous part was answered incorrectly, follow-through marks allow using the candidate's from part (a).
The weights of the green apples sold by a shop are normally distributed with mean 90 grams and standard deviation 8 grams.
Find the probability that a randomly chosen green apple weighs between 83 grams and 95 grams.
Approach
Let grams be the weight of a green apple. Since is normally distributed, standardise using and use the standard normal distribution table.
Working
Given , the required probability is
Using symmetry:
From the standard normal table:
Answer
0.543
Walkthrough
The question gives a normal distribution for weights, so the variable is continuous. The probability is the area under the normal curve between 83 and 95 grams. We standardise each boundary with .
With and , the lower boundary becomes and the upper boundary becomes .
Because one z-value is negative, tables for positive values are used with symmetry: the area to the left of is . Therefore the required interval area is
Substituting the tabulated values gives , rounded to .
Key Takeaways
This question tests standardisation of a normal random variable, converting an interval into a standard normal interval, and using symmetry to handle negative z-values. It also checks that you know which value is the mean and which is the standard deviation.
Common Mistakes
- Dividing by the variance instead of the standard deviation in the standardisation formula.
- Treating incorrectly; it must be converted using symmetry before subtracting.
- Adding or subtracting the two table probabilities in the wrong order.
- Applying a continuity correction to the normal distribution itself; it is only needed when approximating a discrete distribution.
- Giving unsupported or insufficiently rounded final probabilities.
Things to Be Careful About
Use and , not .
The two z-values are and . The final probability should be greater than , which is a useful check because the interval is close to the mean and covers a large central region.
The mark scheme expects the standardisation formula to be shown with 90 and 8; if the method is not shown, only special-credit marks may be available.
The shop also sells red apples. 60% of the red apples sold by the shop weigh more than 80 grams. 160 red apples are chosen at random from the shop.
Use a suitable approximation to find the probability that fewer than 105 of the chosen red apples weigh more than 80 grams.
Approach
Let be the number of red apples, out of 160, that weigh more than 80 grams. Each apple independently does this with probability , so . Since and are both greater than 5, the normal approximation is appropriate. Use a continuity correction because is discrete.
Working
For the binomial distribution:
Approximately, .
"Fewer than 105" means , so with continuity correction:
Using the standard normal table:
Answer
0.915
Walkthrough
There are 160 independent red apples, each with the same probability of weighing more than 80 grams, so the count has a binomial distribution . A normal approximation is suitable because and are both large enough.
For the approximation we need the binomial mean and variance:
Since is discrete, the event is the same as . The continuity correction halfway between 104 and 105 is 104.5, so we standardise 104.5 using mean 96 and standard deviation .
The required probability is therefore . It is greater than 0.5 because 104.5 is above the mean 96 but not far into the upper tail.
Key Takeaways
This question combines the binomial distribution with the normal approximation. You must know the binomial mean and variance, recognise when a normal approximation is valid, and apply a continuity correction correctly when converting a discrete event to a continuous one.
Common Mistakes
- Omitting the continuity correction entirely.
- Using 105.5 instead of 104.5 for a "fewer than 105" probability; 104.5 is correct because the event is equivalent to .
- Writing the variance as and then using it as if it were still the variance; the standard deviation is .
- Confusing the success category: here success is "weighs more than 80 grams", so .
- Not checking that and are large enough for the normal approximation.
Things to Be Careful About
Use and . In the standardisation formula the denominator is , not .
The boundary for "fewer than 105" is 104.5. Keep enough decimal places when computing so the final table reading is accurate. The final probability should be about 0.915 and should be greater than 0.5, which is a good sanity check.
The mark scheme allows unsimplified forms such as and , but if the variance is clearly identified as the standard deviation the mean/variance mark may be lost.
The heights of the female students at Breven college are normally distributed:
- 90% of the female students have heights less than 182.7 cm.
- 40% of the female students have heights less than 162.5 cm.
Find the mean and the standard deviation of the heights of the female students at Breven college.
Approach
Let the heights be . Convert each given percentage into a standard normal -value, then write the two standardisation equations and solve them for and .
Working
90% of the students have height less than 182.7 cm, so
and
40% of the students have height less than 162.5 cm, so
and
Rewrite both equations:
Subtract the second equation from the first:
Substitute back to find :
Answer
The mean is cm and the standard deviation is cm.
mean = 165.8 cm, standard deviation = 13.2 cm
Walkthrough
We are given two cumulative probabilities from a normal distribution. The key idea is to convert each height into a standard normal score using the known inverse -values. For 90% below, the -value is ; for 40% below, the -value is . Writing each height as gives two linear equations in and . Subtracting one equation from the other eliminates , allowing us to solve for . Substituting back gives .
Key Takeaways
This question tests the relationship between a normal distribution and the standard normal distribution. It shows how to recover the mean and standard deviation from two cumulative probability statements. The key skill is converting a percentage into a -value and then solving simultaneous equations.
Common Mistakes
- Using the probabilities 0.9 and 0.4 directly instead of the corresponding -values.
- Using or in the standardisation formula.
- Making a sign error: a percentage below the mean gives a negative -value.
- Rounding too early, which can lead to final answers that are not accurate enough.
Things to Be Careful About
The mark scheme requires the critical values and to be seen. Ensure the two equations are written with and on the same side before solving. Give the final answers to at least one decimal place because the context is real measurements.
Ten female students are chosen at random from those at Breven college.
Find the probability that fewer than 8 of these 10 students have heights more than 162.5 cm.
Approach
From part (a), , so the probability that one student has height more than 162.5 cm is . Let be the number of the 10 students with height more than 162.5 cm. Then . We need , which is easier to find using the complement .
Working
Using :
Therefore
Answer
0.833
Walkthrough
From part (a), 40% of the students have height less than 162.5 cm, so 60% have height more than 162.5 cm. Therefore the probability of success for each randomly chosen student is . With 10 students, the number of successes follows a binomial distribution. Let be the number of students with height more than 162.5 cm. “Fewer than 8” means or . It is easier to calculate the complement: the probability that 8, 9 or 10 students have height more than 162.5 cm. Each binomial term is computed using , and the three terms are summed and subtracted from 1.
Key Takeaways
This part combines a normal distribution result with a binomial distribution. The important step is identifying from the 40% probability of being below 162.5 cm. The complement rule is useful when “fewer than” would require summing many terms.
Common Mistakes
- Using instead of .
- Misreading “fewer than 8” as “8 or fewer”; it means to .
- Forgetting the combinatorial coefficients , , and .
- Omitting one of the three complement terms, or subtracting the complement incorrectly.
Things to Be Careful About
Make sure the binomial probability uses the correct exponents for and . Since , the term for exactly 10 successes is simply . Give the final answer to three decimal places, matching the mark scheme value .
How many different arrangements are there of the 9 letters in the word INTELLECT in which the two Ts are together?
Approach
Treat the two T's as a single block, since they must be together. This leaves 8 items to arrange, but the two E's and the two L's are identical, so divide by for each repeated pair.
Working
The letters of INTELLECT are
With the two T's joined as one block, arrange:
Number of arrangements:
Answer
10080
Walkthrough
The phrase "the two T's are together" means the two T's must appear as the single block TT somewhere in the word. Once you join them, the problem becomes arranging eight items: TT, I, N, E, E, L, L, C.
If all eight of those items were different, the answer would be . They are not all different: there are two E's and two L's. For any arrangement, swapping the two E's gives the same word, and swapping the two L's gives the same word. These two swaps are independent, so each distinct word has been counted times. Dividing by gives 10080.
Why don't we also divide by for the T's? Because the T's are treated as one block; they are not being arranged separately inside the count. The only order inside the block is TT, so there is no extra factor.
Key Takeaways
- A constraint that two letters are adjacent can be handled by forming a block and arranging one fewer object.
- Repeated letters reduce the number of distinct arrangements: divide by the factorial of each repeated letter frequency.
Common Mistakes
- Counting directly without considering repeated letters.
- Writing an extra division by for the T's after blocking them; the T's should contribute no separate factorial.
- Forgetting to divide by for both E and L.
Things to Be Careful About
The standard calculation is . The mark scheme also accepts equivalent forms such as with or and or , but the simplest is to treat TT as one block and divide by the repeated E and L factorials.
How many different arrangements are there of the 9 letters in the word INTELLECT in which there is a T at each end and the two Es are not next to each other?
Approach
Fix the two T's at the ends, one at each end, so only the 7 middle letters need to be arranged. Start with all middle arrangements and subtract those in which the two E's are adjacent.
Working
Middle letters:
All middle arrangements:
If the two E's are together, treat them as one block. The middle items become:
So the number with the E's together is
Required number:
Answer
900
Walkthrough
With a T fixed at each end, the two end positions are already occupied. Since the two T's are identical, there is no extra factor for choosing which T goes to which end.
The middle seven letters are I, N, E, E, L, L, C. Count their arrangements without any restriction on the E's:
Now subtract the forbidden arrangements in which the two E's are next to each other. Treat EE as one block, so the middle items become EE, I, N, L, L, C. There are six items, with L repeated twice, giving
So the number satisfying the condition is .
An alternative gap method is also valid: arrange the non-E middle letters, choose two gaps for the E's, and multiply; it gives the same result.
Key Takeaways
- When identical letters must be placed at the two ends, fix them first and do not multiply by an extra factor.
- Counting arrangements that avoid adjacency is often easiest by subtracting the arrangements where the two letters ARE adjacent.
Common Mistakes
- Forgetting that the two T's are identical, so there is no factor of 2 for swapping the end T's.
- Subtracting the E-together case incorrectly, e.g. using without dividing by for the repeated L.
- Counting the middle as 8 positions instead of 7.
Things to Be Careful About
The mark scheme accepts either the subtraction method or the gap method. In the subtraction method, the E-together block is just one item among the middle six, and the repeated L still requires a division by .
Four letters are selected at random from the 9 letters in the word INTELLECT.
Find the percentage of the possible selections which contain at least one E and exactly one T.
Approach
Count selections of 4 letters that contain exactly one T and at least one E. The word contains 2 T's, 2 E's, and 5 other letters. Split into selections with exactly one E and selections with two E's. Divide the favourable count by the total number of 4-letter selections, then convert to a percentage.
Working
Case 1: exactly one T and exactly one E
Case 2: exactly one T and two E's
Therefore the number of favourable selections is
The total number of selections of 4 letters is
Required percentage:
Answer
39.7%
Walkthrough
We are choosing 4 letters, not arranging them, so combinations are appropriate. The denominator is because there are 9 physical letter positions and any 4 can be chosen.
For the favourable count, exactly one T means choose 1 of the two T's. The condition "at least one E" can happen in two mutually exclusive ways.
Exactly one E: choose 1 T, 1 E, and then 2 letters from the 5 letters that are neither T nor E. These 5 letters are L, L, I, N, C. This gives .
Exactly two E's: choose 1 T, both E's, and then 1 of the remaining 5 letters. This gives .
Adding the two disjoint cases gives 50 favourable selections. The total possible selections are . Therefore the percentage is
Key Takeaways
- "Selections" means combinations, not arrangements.
- Repeated letters are still counted as separate physical letter positions when using .
- "At least one E" is handled by considering the disjoint cases: exactly one E and exactly two E's.
Common Mistakes
- Using arrangements rather than combinations.
- Counting only one E copy or only one T copy, forgetting that the word has two of each.
- Double-counting the case with two E's by also including it in the one-E case.
- Forgetting to convert the probability fraction into a percentage.
Things to Be Careful About
The mark scheme requires the denominator to be written as . The five non-T, non-E letters are L, L, I, N and C, so there are 5 choices for the remaining spots. The final percentage should be in the range , so is correct.
