Notes/Mathematics/Paper 6/Linear Combinations of Random Variables
CAIEA Level9709§6.2

Linear Combinations of Random Variables

How the mean and variance of a random variable behave under scaling, shifting and combining — and the two special cases, normal and Poisson, where combining variables produces another variable of the same family.

220 min read 4 sub-topics
95
question parts
2021–2025 · 37 papers
10 marks
per paper
≈ 21% of the paper
2.8/3
avg difficulty
demanding
#4
most examined
of 5 topics by marks

Every distribution studied so far — binomial, Poisson, normal — describes a single random variable on its own. But real questions rarely stop there: a taxi fare is a fixed charge plus a rate per kilometre; a factory's total income is the sum of two different products' takings; a sample mean is built from several independent measurements added together. This topic gives the rules for finding the mean and variance of any such combination, without ever having to rebuild a distribution table from scratch.

Across 2021–2025 this topic carried 383 marks over 95 tagged parts, about 10.4 of the 50 marks on every Paper 6. Its average difficulty, 2.762.76, is the highest of the five S2 topics — not because any single rule is hard, but because a question typically chains several rules together (find a variance, combine two variables, recognise normality, standardise) with no algebra to fall back on if a rule is misremembered.

The marks split across four jobs:

  • E(aX+b)E(aX+b) and Var(aX+b)Var(aX+b) — scaling and shifting one random variable;
  • E(aX+bY)E(aX+bY) and Var(aX+bY)Var(aX+bY) for independent XX and YY — combining two different variables, including sums of several independent copies of the same variable;
  • linear combinations of normal variables are themselves normal — the single fact that turns a combined mean and variance into an actual probability;
  • sums of independent Poisson variables are themselves Poisson — the analogous fact for counts.
Before you start you should be able to
  • E(X)E(X) and Var(X)Var(X) from a probability distribution, and Var(X)=E(X2)−[E(X)]2Var(X)=E(X^2)-[E(X)]^2 (§5.4)

  • The binomial distribution's mean and variance, npnp and np(1−p)np(1-p) (§5.4)

  • Standardising a normal variable and reading Φ(z)\Phi(z) from tables (§5.5)

  • The Poisson distribution Po(λ)Po(\lambda), its formula, and E(X)=Var(X)=λE(X)=Var(X)=\lambda (§6.1)

  • Approximating Po(λ)Po(\lambda) by N(λ,λ)N(\lambda,\lambda) when λ\lambda is large, with a continuity correction (§6.1)

  • Reading the normal table backwards — finding zz from a given probability (§5.5)

  • Conditional probability, P(A∣B)=P(A∩B)/P(B)P(A\mid B)=P(A\cap B)/P(B) (§5.3)

By the end of this page you can
  • Use E(aX+b)=aE(X)+bE(aX+b)=aE(X)+b and Var(aX+b)=a2Var(X)Var(aX+b)=a^2Var(X)

  • Use the fact that if XX is normal then so is aX+baX+b, and find probabilities from it

  • Use E(aX+bY)=aE(X)+bE(Y)E(aX+bY)=aE(X)+bE(Y) and, for independent X,YX,Y, Var(aX+bY)=a2Var(X)+b2Var(Y)Var(aX+bY)=a^2Var(X)+b^2Var(Y)

  • Distinguish the sum of nn independent copies of XX from a single copy scaled by nn, and find the mean and variance of each

  • Use the fact that a linear combination of independent normal variables is itself normal, and use this to find probabilities

  • Use the fact that the sum of independent Poisson variables is itself Poisson, with parameter equal to the sum of the individual parameters

  • Put Poisson rates on a common interval before adding them, and approximate a large Poisson total by a normal distribution

  • Handle "differ by more than kk" as two tails of a single difference variable, added

  • Find an unknown sample size nn from a stated probability, standardising with σn\sigma\sqrt{n} and solving the resulting quadratic in n\sqrt{n}

01

E(aX + b) and Var(aX + b) — scaling and shifting one variable

Syllabus requirement · §6.2

“

use, in the course of solving problems, the results that E(aX + b) = aE(X) + b and Var(aX + b) = a²Var(X).

”

Plenty of quantities are just a fixed rule applied to a random variable you already understand: a total cost built from a fixed fee plus a rate per unit, a temperature converted from one scale to another, a price converted into another currency. If XX is a random variable and Y=aX+bY=aX+b for constants aa and bb, the mean and variance of YY follow directly from those of XX, with no new distribution to build:

E(aX+b)=aE(X)+b,Var(aX+b)=a2 Var(X)E(aX+b) = aE(X)+b, \qquad Var(aX+b) = a^2\,Var(X)

The syllabus does not ask you to prove these, but the reason they take this shape is short — and knowing it means you can rebuild either formula if you forget it. Adding the constant bb shifts every possible value of XX up by bb — so the average shifts up by bb too, but the gaps between values are unchanged, so the spread (variance) is completely unaffected by bb. Scaling by aa stretches every value, and every gap between values, by a factor of aa — so the mean scales by aa, but since variance is built from squared deviations, the spread scales by a2a^2, not aa.

The constant b vanishes from the variance

This is the single most common slip in the whole topic: bb shifts the mean but never appears in the variance at all — not as +b+b, not as +b2+b^2. Only the multiplier aa affects the spread, and it does so as a2a^2.

If X is normal, so is aX + b

The two rules above work for any random variable XX — they say nothing about its shape. But when XX happens to be normal, there is a bonus: aX+baX+b is normal too, with exactly the mean and variance those rules give.

X∼N(μ,σ2)⟹aX+b∼N(aμ+b, a2σ2)X\sim N(\mu,\sigma^2) \quad\Longrightarrow\quad aX+b \sim N\big(a\mu+b,\ a^2\sigma^2\big)

That matters because a mean and a variance on their own cannot produce a probability — you need the shape as well. This is the one-variable version of §03's result, and the reasoning there is the same, applied to two variables at once.

Substituting into both formulas, piece by piece

A taxi charges a fixed $2.50 plus $3 per kilometre travelled. The distance travelled on a randomly chosen journey, XX km, has mean 1010 and variance 44. Find the mean and variance of the total fare, Y=3X+2.5Y = 3X + 2.5.

Show full working
  1. 1

    Step 1 — identify aa and bb in Y=aX+bY=aX+b. Here a=3a=3 (the rate per km) and b=2.5b=2.5 (the fixed charge).

    Getting a and b the right way round is the whole question. The rate per kilometre multiplies the random distance, so it is a; the fixed charge is paid whatever happens, so it is b.

  2. 2

    Step 2 — find E(Y)E(Y), substituting E(X)=10E(X)=10 into the mean formula. E(Y)=aE(X)+b=3×10+2.5E(Y) = aE(X)+b = 3\times10+2.5

  3. 3

    Step 3 — evaluate. E(Y)=30+2.5=32.5E(Y) = 30+2.5 = 32.5

  4. 4

    Step 4 — find Var(Y)Var(Y), substituting Var(X)=4Var(X)=4 into the variance formula. Var(Y)=a2 Var(X)=32×4Var(Y) = a^2\,Var(X) = 3^2\times4

    The fixed charge b = 2.5 plays no part in this step at all — only a is squared.

  5. 5

    Step 5 — evaluate. Var(Y)=9×4=36Var(Y) = 9\times4 = 36

Answer

E(Y)=32.5E(Y) = 32.5, Var(Y)=36Var(Y) = 36.

Find E(Y) and Var(Y) as two entirely separate calculations — the constant b enters only the mean formula, and only the multiplier a (squared) enters the variance formula.

A linear transformation of a normal variable

9709/61 O/N 2023 Q4(a)2 marks

The mass, in kilograms, of chemical AA produced per day by a factory is modelled by the random variable X∼N(10.3,5.76)X \sim N(10.3, 5.76). The income generated by chemical AA is $2.50 per kilogram. Find the mean and variance of the daily income generated by chemical AA.

Show full working
  1. 1

    Step 1 — write the income as a linear transformation of XX. Income =2.5X= 2.5X, so a=2.5a=2.5 and b=0b=0 (no fixed fee here).

    There is no fixed fee here, so b = 0. Write it down anyway — it is what tells you nothing gets added on at the end.

  2. 2

    Step 2 — find the mean income. E(2.5X)=2.5×E(X)=2.5×10.3E(2.5X) = 2.5\times E(X) = 2.5\times10.3

  3. 3

    Step 3 — evaluate. E(2.5X)=25.75E(2.5X) = 25.75

  4. 4

    Step 4 — find the variance of the income, squaring the multiplier. Var(2.5X)=2.52×Var(X)=6.25×5.76Var(2.5X) = 2.5^2\times Var(X) = 6.25\times5.76

  5. 5

    Step 5 — evaluate. Var(2.5X)=36Var(2.5X) = 36

  6. 6

    Step 6 — name the distribution of the income. XX is normal, so 2.5X2.5X is normal too, carrying the mean and variance just found: 2.5X∼N(25.75, 36)2.5X \sim N(25.75,\ 36)

    Worth writing even though this part asks only for a mean and a variance — it is what makes a follow-on part such as 'find the probability the income exceeds 30 dollars' answerable at all.

Answer

Mean income == $25.75, variance =36=36 (in $2^2); the income is distributed N(25.75,36)N(25.75, 36).

Combining with a distribution's own mean/variance formula first

9709/62 M/J 2021 Q2(a)3 marks

The random variable XX has the distribution B(400,0.01)B(400, 0.01). Find Var(4X+2)Var(4X+2).

Show full working
  1. 1

    Step 1 — find Var(X)Var(X) first, using the binomial variance formula (§5.4), since it isn't given directly. Var(X)=np(1−p)=400×0.01×0.99Var(X) = np(1-p) = 400\times0.01\times0.99

    The question gives you a distribution, not a variance. Var(X) has to be produced from B(400, 0.01) first; the linear rule cannot start until it exists.

  2. 2

    Step 2 — evaluate. Var(X)=3.96Var(X) = 3.96

  3. 3

    Step 3 — identify aa in 4X+24X+2. Here a=4a=4; the constant 22 will not appear in the variance at all.

  4. 4

    Step 4 — substitute into Var(aX+b)=a2 Var(X)Var(aX+b) = a^2\,Var(X). Var(4X+2)=42×3.96Var(4X+2) = 4^2\times3.96

  5. 5

    Step 5 — evaluate. Var(4X+2)=16×3.96=63.36Var(4X+2) = 16\times3.96 = 63.36

Answer

Var(4X+2)=63.36Var(4X+2) = 63.36.

When Var(X) isn't given directly, find it first using whatever distribution X actually has — the linear-transformation rule only ever needs Var(X) as an input, however that number was obtained.

Common mistakes
  • Writing Var(aX+b)=a2 Var(X)+b2Var(aX+b) = a^2\,Var(X) + b^2, or +b+b

    Var(aX+b)=a2 Var(X)Var(aX+b) = a^2\,Var(X) exactly — the additive constant never appears in the variance

    A constant shift moves every value by the same amount, so it can't change how spread out those values are relative to each other.

  • Using Var(aX+b)=a Var(X)Var(aX+b) = a\,Var(X), forgetting to square the multiplier

    Square aa before multiplying: Var(aX+b)=a2 Var(X)Var(aX+b)=a^2\,Var(X)

    Variance is built from squared deviations, so any linear scaling of X carries through as the square of that scale factor.

The same idea, drawn as a graph

Once a random variable's probability density function has a graph (§6.3), these two rules have a picture: adding bb slides the whole graph sideways without changing its shape or height; scaling by aa stretches it horizontally by a factor of aa and, to keep the total area under the curve equal to 11, shrinks its height by the same factor. Cambridge sets exactly this as a sketching question — "sketch the density of aX+baX+b" — once §6.3 has introduced what that density graph actually is. Keep the E/Var rules from this section in mind; §6.3 comes back to them for the picture.

Your turn

  1. 1

    A random variable XX has mean 88 and variance 55. Find E(2X−3)E(2X-3) and Var(2X−3)Var(2X-3).

    Stuck? Show hint

    E(aX+b) = aE(X)+b uses both a and b; Var(aX+b) = a²Var(X) uses only a, squared.

    Show solution
    1. 1

      Step 1 — identify aa and bb. Comparing 2X−32X-3 with aX+baX+b gives a=2a=2 and b=−3b=-3.

      The minus sign belongs to b. Writing b = −3, rather than thinking 'subtract 3', is what keeps the sign right when it is substituted.

    2. 2

      Step 2 — substitute into the mean formula. E(2X−3)=aE(X)+b=2×8−3E(2X-3) = aE(X)+b = 2\times8-3

    3. 3

      Step 3 — evaluate. E(2X−3)=16−3=13E(2X-3) = 16-3 = 13

    4. 4

      Step 4 — substitute into the variance formula, which uses only aa. Var(2X−3)=a2 Var(X)=22×5Var(2X-3) = a^2\,Var(X) = 2^2\times5

      The −3 is not written anywhere in this line. That is the rule, not an oversight.

    5. 5

      Step 5 — evaluate. Var(2X−3)=4×5=20Var(2X-3) = 4\times5 = 20

    Answer

    E(2X−3)=13E(2X-3)=13, Var(2X−3)=20Var(2X-3)=20.

  2. 23 marks

    A company converts a measured temperature XX (in an internal sensor unit, mean 5050, variance 1616) to degrees Celsius using C=0.5X−10C = 0.5X - 10. Find the mean and standard deviation of CC.

    Stuck? Show hint

    Find E(C) and Var(C) separately, then take a square root for the standard deviation.

    Show solution
    1. 1

      Step 1 — identify aa and bb in C=0.5X−10C = 0.5X - 10. Here a=0.5a=0.5 and b=−10b=-10.

    2. 2

      Step 2 — find the mean. E(C)=0.5×50−10=25−10=15E(C) = 0.5\times50-10 = 25-10 = 15

    3. 3

      Step 3 — find the variance, squaring aa. Var(C)=0.52×16=0.25×16=4Var(C) = 0.5^2\times16 = 0.25\times16 = 4

      0.5 squared is 0.25, not 0.5. A multiplier below 1 shrinks the spread by more than it shrinks the mean.

    4. 4

      Step 4 — the question asks for the standard deviation, so square-root the variance. s.d.(C)=4=2\text{s.d.}(C) = \sqrt{4} = 2

      Read the last line of the question again before answering. Quoting the variance where a standard deviation was asked for throws away the final mark.

    Answer

    E(C)=15E(C)=15, standard deviation of C=2C = 2.

Practise E(aX + b) and Var(aX + b) from real Paper 6 papersReal past-paper questions · E(aX + b) and Var(aX + b)

The rest of this note

Checking your access…

Can you do all of these?

  • The additive constant b never appears in Var(aX+b) — only the multiplier a, squared

  • Var(aX+bY) = a²Var(X) + b²Var(Y) needs independence — and adds variances even for a difference

  • Distinguish 'n independent copies summed' (variance × n) from 'one copy scaled by n' (variance × n²) — the means agree, the variances never do

  • Convert a standard deviation to a variance before combining — mark schemes reject any SD/variance mix

  • Rearrange any inequality between two variables (e.g. L < 3S) onto one side before reading off a and b

  • A linear combination is normal only when every variable in it is independent AND individually normal

  • The sum of independent Poisson variables is Poisson, with parameter equal to the sum of the individual parameters

  • State the new distribution N(mean, variance) or Po(λ) as its own explicit step before standardising or substituting further

  • When a question says 'stating a necessary assumption', write the independence assumption out in words — it carries its own mark

  • Put every Poisson rate on the interval the question asks about BEFORE adding the parameters — never scale the combined λ afterwards

  • 'Differ by more than k' is P(D > k) + P(D < −k) added — not twice one tail, unless E(D) = 0

  • The total of n independent copies has standard deviation σ√n, never σn

Now do the questions
95 real Paper 6 parts from 2021–2025, sorted by difficulty, with mark schemes