Notes/Mathematics/Paper 6/Hypothesis Tests
CAIEA Level9709§6.5

Hypothesis Tests

Turning sample evidence into a formal decision about a claimed population parameter — setting up hypotheses, testing binomial, Poisson and mean parameters, and the two ways a test can go wrong.

300 min read 5 sub-topics
155
question parts
2021–2025 · 37 papers
14 marks
per paper
≈ 27% of the paper
2.4/3
avg difficulty
moderate
#1
most examined
of 5 topics by marks

§6.4 turned sample data into an estimate, and a range of plausible values, for an unknown parameter. A hypothesis test asks a sharper question: is a specific claimed value of that parameter still believable, given what the sample actually showed? A factory claims only 2%2\% of its output is flawed; an inspector's sample suggests more — is that difference real, or just the kind of variation any random sample would show even if the claim were true? This topic gives a formal, repeatable procedure for answering exactly that.

Across 2021–2025 this topic carried 502 marks over 155 tagged parts, about 13.6 of the 50 marks on every Paper 6 — the single heaviest of the five S2 topics, and often the final, longest question on the paper. Its average difficulty, 2.372.37, sits mid-range for S2 — harder than sampling and estimation (2.072.07), easier per step than linear combinations (2.762.76). No single step is exotic. What makes the topic expensive is length: a full test chains five or six steps together (hypotheses, distribution, tail probability, comparison, conclusion in context), so losing the thread halfway costs far more than it would on a shorter question.

The marks split five ways, and this note takes them in that order:

  • the mechanics of a test — hypotheses, one-tailed vs two-tailed, significance level, critical region;
  • exact tests for a binomial or Poisson parameter;
  • a test for a population mean, using a normal population or a large sample (§6.4's Central Limit Theorem);
  • Type I and Type II errors — the two distinct ways a test's conclusion can be wrong;
  • calculating the probability of each type of error.
Before you start you should be able to
  • The binomial and Poisson probability formulas, and cumulative probabilities (§5.4/§6.1)

  • The distribution of the sample mean Xˉ\bar X, and the Central Limit Theorem (§6.4)

  • Standardising a normal variable and using Φ\Phi, including working backwards from a probability (§5.5)

By the end of this page you can
  • State null and alternative hypotheses in terms of a population parameter, and identify a one-tailed or two-tailed test from the wording of a claim

  • Explain significance level and critical region, and reach a conclusion using the critical-region method or the tail-probability (p-value) method

  • Carry out an exact hypothesis test for a binomial or Poisson parameter

  • Use a normal approximation, with a continuity correction, to test a binomial or Poisson parameter when nn (or λ\lambda) is too large for exact evaluation

  • Carry out a hypothesis test for a population mean, using a normal population or a large sample

  • Define a Type I error and a Type II error, and state the condition under which each can occur

  • Calculate the probability of a Type I error, and — given a specific true alternative parameter value — the probability of a Type II error

01

Setting up a test: hypotheses, tails, significance level, critical region

Syllabus requirement · §6.5

“

understand the nature of a hypothesis test, the difference between one-tailed and two-tailed tests, and the terms null hypothesis, alternative hypothesis, significance level, rejection region (or critical region), acceptance region and test statistic. Outcomes of hypothesis tests are expected to be interpreted in terms of the contexts in which questions are set.

”

Every hypothesis test compares two competing statements about a population parameter — a proportion pp, a Poisson mean λ\lambda, a population mean μ\mu:

  • the null hypothesis, H0H_0, is the assumed or claimed value, always written as an equality: H0:θ=θ0H_0: \theta=\theta_0;
  • the alternative hypothesis, H1H_1, is what the investigator suspects instead — and its direction comes entirely from the wording of the claim being investigated.

The sample then provides a test statistic (an observed count, or a sample mean) which is compared against what H0H_0 would predict. If the observed value is implausible enough under H0H_0, H0H_0 is rejected in favour of H1H_1; otherwise there simply isn't enough evidence to reject it.

Hypotheses are about the population, never the sample

H0H_0 and H1H_1 are always statements about the unknown population parameter (pp, λ\lambda, μ\mu) — never about the sample proportion, sample count or sample mean actually observed. Writing H0:xˉ=500H_0: \bar x = 500 instead of H0:μ=500H_0: \mu=500 is a common way to lose the hypotheses mark even when everything that follows is correct.

One-tailed or two-tailed — read it from the claim

If the suspicion has a direction ("less than", "has decreased", "is higher than"), H1H_1 uses << or >> and the test is one-tailed. If the suspicion is only that the value has changed, with no stated direction ("is different from", "is no longer"), H1H_1 uses ≠\neq and the test is two-tailed.

Wording of the claim

H₁

Tail(s)

"...is less than...", "has decreased", "fewer than claimed"

θ<θ0\theta<\theta_0

one-tailed, lower

"...is more than...", "has increased", "greater than claimed"

θ>θ0\theta>\theta_0

one-tailed, upper

"...is different from...", "has changed", "is no longer"

θ≠θ0\theta\neq\theta_0

two-tailed, both

The direction of the suspicion in the question is what fixes H₁ — get this from the words, not from which way the sample happened to differ.

Where the significance level α = 5% sits depends on H₁'s direction−1.645α = 5%one-tailed: H₁: θ < θ₀all of α in the lower tail−1.960+1.960α/2α/2two-tailed: H₁: θ ≠ θ₀2½% in each tail — so the cut-off moves out

A one-tailed test puts the whole significance level α in a single tail; a two-tailed test splits it evenly between both.

Significance level and critical region

The significance level, α\alpha (commonly 5%5\%, 1%1\% or 10%10\%), is chosen before the test — the maximum probability of rejecting H0H_0 that the investigator is willing to risk, given that H0H_0 is actually true. The critical region (or rejection region) is the set of test-statistic values extreme enough, under H0H_0, to trigger rejection — built to be as large as possible without its probability under H0H_0 exceeding α\alpha.

Everything outside the critical region is the acceptance region. That is the syllabus's name for it, but it is worth reading as the non-rejection region: landing there means the evidence was not strong enough to reject H0H_0, not that H0H_0 has been shown to be true. §04 explains why that distinction matters.

There are two equivalent routes to a conclusion, and either is acceptable:

  1. Critical-region method — find the critical region once, then simply check whether the observed value falls inside it.
  2. Tail-probability (p-value) method — find the probability, under H0H_0, of a result at least as extreme as the one observed, and compare that probability directly with α\alpha.

For every test in this topic the two routes agree, because both are built from exactly the same tail probability.

What a hypothesis test is actually asking

Every test asks one question: "if H0H_0 were true, how surprising would this sample be?" A result that would be unusually rare under H0H_0 (probability below α\alpha) is treated as evidence against H0H_0. A test never proves H0H_0 false — it only says the observed data would have been an uncomfortably rare coincidence if H0H_0 were true.

Identifying hypotheses and the tail from a described claim

A vending machine is designed to dispense a can with probability 0.90.9 of it being correctly filled. A technician suspects the true probability is now lower than this. Write down suitable null and alternative hypotheses, and state whether the test is one-tailed or two-tailed.

Show full working
  1. 1

    Step 1 — identify the population parameter being tested. The proportion of cans correctly filled, pp.

  2. 2

    Step 2 — write H0H_0 as the claimed/assumed value, using equality. H0:p=0.9H_0: p=0.9

  3. 3

    Step 3 — read the direction of the suspicion from the wording. "Suspects the true probability is now lower" — a one-directional claim.

  4. 4

    Step 4 — write H1H_1 using the matching inequality. H1:p<0.9H_1: p<0.9

  5. 5

    Step 5 — state the tail. Since H1H_1 uses <<, this is a one-tailed (lower-tailed) test.

Answer

H0:p=0.9H_0: p=0.9, H1:p<0.9H_1: p<0.9; one-tailed (lower).

Underline the direction word in the question (lower, higher, different) before writing anything — it's the single piece of wording that fixes H₁ and the tail, and misreading it here derails everything that follows.

Stating the tail and the reason — the two-part answer

9709/61 O/N 2024 Q5(a)1 mark

The lengths, in centimetres, of worms of a certain kind are normally distributed with mean μ\mu and standard deviation 2.32.3. An article in a magazine states that the value of μ\mu is 12.712.7. A scientist wishes to test whether this value is correct. He measures the lengths, xx cm, of a random sample of 5050 worms of this kind and finds that ∑x=597.1\sum x = 597.1. He plans to carry out a test, at the 1%1\% significance level, of whether the true value of μ\mu is different from 12.712.7.

State, with a reason, whether he should use a one-tailed or a two-tailed test.

Show full working
  1. 1

    Step 1 — read what the command word is asking for. "State, with a reason" wants two separate things in the answer: which test, and why.

    One mark, two halves — and the mark is only paid when both are there. Writing just 'two-tailed' gives no reason, and giving a reason without naming the tail never makes the statement. This is the usual way a correct idea scores zero here.

  2. 2

    Step 2 — find the sentence that says what is being tested. It is the last one: "...a test ... of whether the true value of μ\mu is different from 12.712.7."

    Not the opening sentence, and not the sample data. The tail comes from the sentence describing the purpose of the test; ∑x=597.1\sum x=597.1, n=50n=50 and the 1%1\% level all belong to the calculation in part (b).

  3. 3

    Step 3 — check that sentence for a direction word. "Different from" is not one. There is no "less than", no "greater than", no "has increased" or "has decreased" anywhere in it.

  4. 4

    Step 4 — map that to H1H_1. With no direction claimed, H1H_1 takes ≠\neq: H1:μ≠12.7H_1: \mu\neq12.7 A value of μ\mu below 12.712.7 and a value above it are equally good evidence against H0H_0, so both tails are needed.

  5. 5

    Step 5 — write the answer in its two halves: the tail, then the reason. Two-tailed, because he is looking for a difference — he is testing whether μ\mu has changed, not whether it is larger or smaller.

    The mark scheme's own wording is 'Two-tailed because looking for difference'. Echoing the question's own word back at it is the safest reason to give; on a similar question (9709/65 O/N 2025 Q2(a)) the examiners also accepted 'the researcher is not looking for less than or more than'.

Answer

Two-tailed, because he is looking for a difference — the test is of whether μ\mu is different from 12.712.7, with no direction claimed.

Answer every one of these in one fixed shape: [one-tailed / two-tailed] because [the direction word the question used]. Both halves, every time. Notice too that neither half depends on any of the sample data, so this part can be answered in full before a calculator is touched.

Your turn — one-tailed or two-tailed?

  1. 19709/62 F/M 2023 Q6(a)1 mark

    Last year, the mean time taken by students at a school to complete a certain test was 2525 minutes. Akash believes that the mean time taken by this year's students was less than 2525 minutes. In order to test this belief, he takes a large random sample of this year's students and he notes the time taken by each student. He carries out a test, at the 2.5%2.5\% significance level, for the population mean time, μ\mu minutes. Akash uses the null hypothesis H0:μ=25H_0: \mu = 25.

    Give a reason why Akash should use a one-tailed test.

    Stuck? Show hint

    This one asks for the reason only — the question has already told you the test is one-tailed, so the first half of the template is done for you. Find the sentence saying what Akash believes.

    Show solution
    1. 1

      Find the sentence stating what is being tested. "Akash believes that the mean time taken by this year's students was less than 2525 minutes."

    2. 2

      Read the direction out of it. "Less than" is a direction: he expects μ\mu to have gone down, not simply to have changed. Only departures below 2525 would count as evidence for him, so only one tail is needed.

    Answer

    Because he is expecting a decrease in μ\mu — he believes the mean time is less than 2525 minutes, so only one direction of departure from H0H_0 is being tested.

Common mistakes
  • Writing H0H_0 with an inequality, e.g. H0:p⩽0.9H_0: p \leqslant 0.9

    H0H_0 is always an equality, H0:p=0.9H_0: p=0.9 — the inequality belongs to H1H_1 only

    The whole test is built on modelling the sample assuming H0's exact claimed value is true; there's no single distribution to use if H0 itself is a range.

  • Treating the tail as worth one mark, so getting it wrong is a small slip — a two-tailed test on a question whose wording gives a direction, or a one-tailed test on one that only says "different"

    Fix the tail from the wording before writing anything else: the wrong tail caps the marks on the whole question, not just the hypotheses line

    Mark schemes cap these outright. On 9709/62 O/N 2025 Q6(a) (6 marks) a two-tailed attempt "scores max B1B0M1A1M1 ... A0", i.e. 44 out of 66; on 9709/62 M/J 2024 Q6(a) (5 marks) "Two tail test scores maximum B0 M1 A1 M1 A0", 33 out of 55; on 9709/62 O/N 2024 Q7(b) (7 marks) "max 5/7". It runs the other way too — 9709/63 M/J 2023 Q5(b) is a two-tailed question, and a one-tailed method there is capped at "max 3/5". The standardising and comparison marks survive, because that arithmetic is still done correctly. What always goes is the hypotheses mark at the start and the conclusion mark at the end, since both are stated about a direction the question never asked about.

  • Choosing the tail based on which way the sample statistic happened to differ from the claimed value, rather than the direction stated in the question

    The tail comes from the suspicion being tested (the wording), decided before looking at how the data came out

    A test's direction has to be fixed in advance — choosing it after seeing which way the data leans is a form of bias that invalidates the significance level.

Your turn

  1. 1

    A seed packet claims that 75%75\% of seeds will germinate. A gardener suspects the true germination rate is different from this. Write down suitable hypotheses and state whether the test is one-tailed or two-tailed.

    Stuck? Show hint

    "Different from" gives no direction — this is the two-tailed case.

    Show solution
    1. 1

      H0:p=0.75H_0: p=0.75, H1:p≠0.75H_1: p\neq0.75 — two-tailed, since no direction is claimed.

    Answer

    H0:p=0.75H_0: p=0.75, H1:p≠0.75H_1: p\neq0.75; two-tailed.

  2. 2

    A café claims the mean waiting time for a coffee is 33 minutes. A regular customer believes it now takes longer. Write down suitable hypotheses for a test of this claim.

    Stuck? Show hint

    "Longer" is a direction — one-tailed, upper.

    Show solution
    1. 1

      H0:μ=3H_0: \mu=3, H1:μ>3H_1: \mu>3 — one-tailed (upper), since "longer" gives a clear direction.

    Answer

    H0:μ=3H_0: \mu=3, H1:μ>3H_1: \mu>3; one-tailed (upper).

Practise setting up hypotheses and critical regions from real Paper 6 papersReal past-paper questions · Null and alternative hypotheses, significance level, critical region

The rest of this note

Checking your access…

Can you do all of these?

  • H₀ and H₁ are always statements about the population parameter, and H₀ always uses equality

  • Read the claim's direction carefully to fix H₁ and the tail: 'more/increased' → upper; 'less/decreased' → lower; 'different/changed' → two-tailed

  • For binomial/Poisson tests, use the exact distribution under H₀ — find P(X ⩽ observed) or P(X ⩾ observed), never a single point probability

  • For a mean test, state the necessary assumption (population normal, or large n for the CLT) and always divide by σ/√n, never σ alone

  • Match the critical value to the tails identified — a two-tailed test's critical value is more extreme than the same α's one-tailed value

  • Getting the tail wrong caps the marks on the WHOLE question, not just the hypotheses line — fix it from the wording before writing anything

  • Write H₀ and H₁ about the population parameter using its symbol: 'μ = 510' scores, 'mean = 510' does not, and any symbol you invent must be defined

  • Conclude in context, in the language of the original claim, without asserting certainty — 'insufficient/sufficient evidence', never 'H₀ is true'

  • Conclusions must be in context, not definite, and free of contradictions — 'insufficient evidence that the mean has decreased', never 'the mean has decreased' and never 'the mean has not decreased'

  • P(Type I error) uses the null distribution: for a DISCRETE test it is the critical region's actual probability (usually a little under α), but for a CONTINUOUS mean test it equals α exactly

  • P(Type II error) needs a specific true alternative value, and uses the SAME fixed critical region/value, re-standardised under that true value

Now do the questions
155 real Paper 6 parts from 2021–2025, sorted by difficulty, with mark schemes