Setting up a test: hypotheses, tails, significance level, critical region
“
understand the nature of a hypothesis test, the difference between one-tailed and two-tailed tests, and the terms null hypothesis, alternative hypothesis, significance level, rejection region (or critical region), acceptance region and test statistic. Outcomes of hypothesis tests are expected to be interpreted in terms of the contexts in which questions are set.
Every hypothesis test compares two competing statements about a population parameter — a proportion , a Poisson mean , a population mean :
- the null hypothesis, , is the assumed or claimed value, always written as an equality: ;
- the alternative hypothesis, , is what the investigator suspects instead — and its direction comes entirely from the wording of the claim being investigated.
The sample then provides a test statistic (an observed count, or a sample mean) which is compared against what would predict. If the observed value is implausible enough under , is rejected in favour of ; otherwise there simply isn't enough evidence to reject it.
Hypotheses are about the population, never the sample
and are always statements about the unknown population parameter (, , ) — never about the sample proportion, sample count or sample mean actually observed. Writing instead of is a common way to lose the hypotheses mark even when everything that follows is correct.
One-tailed or two-tailed — read it from the claim
If the suspicion has a direction ("less than", "has decreased", "is higher than"), uses or and the test is one-tailed. If the suspicion is only that the value has changed, with no stated direction ("is different from", "is no longer"), uses and the test is two-tailed.
Wording of the claim | H₁ | Tail(s) |
|---|---|---|
"...is less than...", "has decreased", "fewer than claimed" | one-tailed, lower | |
"...is more than...", "has increased", "greater than claimed" | one-tailed, upper | |
"...is different from...", "has changed", "is no longer" | two-tailed, both |
The direction of the suspicion in the question is what fixes H₁ — get this from the words, not from which way the sample happened to differ.
A one-tailed test puts the whole significance level α in a single tail; a two-tailed test splits it evenly between both.
Significance level and critical region
The significance level, (commonly , or ), is chosen before the test — the maximum probability of rejecting that the investigator is willing to risk, given that is actually true. The critical region (or rejection region) is the set of test-statistic values extreme enough, under , to trigger rejection — built to be as large as possible without its probability under exceeding .
Everything outside the critical region is the acceptance region. That is the syllabus's name for it, but it is worth reading as the non-rejection region: landing there means the evidence was not strong enough to reject , not that has been shown to be true. §04 explains why that distinction matters.
There are two equivalent routes to a conclusion, and either is acceptable:
- Critical-region method — find the critical region once, then simply check whether the observed value falls inside it.
- Tail-probability (p-value) method — find the probability, under , of a result at least as extreme as the one observed, and compare that probability directly with .
For every test in this topic the two routes agree, because both are built from exactly the same tail probability.
Every test asks one question: "if were true, how surprising would this sample be?" A result that would be unusually rare under (probability below ) is treated as evidence against . A test never proves false — it only says the observed data would have been an uncomfortably rare coincidence if were true.
Identifying hypotheses and the tail from a described claim
A vending machine is designed to dispense a can with probability of it being correctly filled. A technician suspects the true probability is now lower than this. Write down suitable null and alternative hypotheses, and state whether the test is one-tailed or two-tailed.
Show full working
- 1
Step 1 — identify the population parameter being tested. The proportion of cans correctly filled, .
- 2
Step 2 — write as the claimed/assumed value, using equality.
- 3
Step 3 — read the direction of the suspicion from the wording. "Suspects the true probability is now lower" — a one-directional claim.
- 4
Step 4 — write using the matching inequality.
- 5
Step 5 — state the tail. Since uses , this is a one-tailed (lower-tailed) test.
, ; one-tailed (lower).
Underline the direction word in the question (lower, higher, different) before writing anything — it's the single piece of wording that fixes H₁ and the tail, and misreading it here derails everything that follows.
Stating the tail and the reason — the two-part answer
The lengths, in centimetres, of worms of a certain kind are normally distributed with mean and standard deviation . An article in a magazine states that the value of is . A scientist wishes to test whether this value is correct. He measures the lengths, cm, of a random sample of worms of this kind and finds that . He plans to carry out a test, at the significance level, of whether the true value of is different from .
State, with a reason, whether he should use a one-tailed or a two-tailed test.
Show full working
- 1
Step 1 — read what the command word is asking for. "State, with a reason" wants two separate things in the answer: which test, and why.
One mark, two halves — and the mark is only paid when both are there. Writing just 'two-tailed' gives no reason, and giving a reason without naming the tail never makes the statement. This is the usual way a correct idea scores zero here.
- 2
Step 2 — find the sentence that says what is being tested. It is the last one: "...a test ... of whether the true value of is different from ."
Not the opening sentence, and not the sample data. The tail comes from the sentence describing the purpose of the test; , and the level all belong to the calculation in part (b).
- 3
Step 3 — check that sentence for a direction word. "Different from" is not one. There is no "less than", no "greater than", no "has increased" or "has decreased" anywhere in it.
- 4
Step 4 — map that to . With no direction claimed, takes : A value of below and a value above it are equally good evidence against , so both tails are needed.
- 5
Step 5 — write the answer in its two halves: the tail, then the reason. Two-tailed, because he is looking for a difference — he is testing whether has changed, not whether it is larger or smaller.
The mark scheme's own wording is 'Two-tailed because looking for difference'. Echoing the question's own word back at it is the safest reason to give; on a similar question (9709/65 O/N 2025 Q2(a)) the examiners also accepted 'the researcher is not looking for less than or more than'.
Two-tailed, because he is looking for a difference — the test is of whether is different from , with no direction claimed.
Answer every one of these in one fixed shape: [one-tailed / two-tailed] because [the direction word the question used]. Both halves, every time. Notice too that neither half depends on any of the sample data, so this part can be answered in full before a calculator is touched.
Your turn — one-tailed or two-tailed?
- 19709/62 F/M 2023 Q6(a)1 mark
Last year, the mean time taken by students at a school to complete a certain test was minutes. Akash believes that the mean time taken by this year's students was less than minutes. In order to test this belief, he takes a large random sample of this year's students and he notes the time taken by each student. He carries out a test, at the significance level, for the population mean time, minutes. Akash uses the null hypothesis .
Give a reason why Akash should use a one-tailed test.
Stuck? Show hint
This one asks for the reason only — the question has already told you the test is one-tailed, so the first half of the template is done for you. Find the sentence saying what Akash believes.
Show solution
- 1
Find the sentence stating what is being tested. "Akash believes that the mean time taken by this year's students was less than minutes."
- 2
Read the direction out of it. "Less than" is a direction: he expects to have gone down, not simply to have changed. Only departures below would count as evidence for him, so only one tail is needed.
AnswerBecause he is expecting a decrease in — he believes the mean time is less than minutes, so only one direction of departure from is being tested.
- 1
Writing with an inequality, e.g.
is always an equality, — the inequality belongs to only
The whole test is built on modelling the sample assuming H0's exact claimed value is true; there's no single distribution to use if H0 itself is a range.
Treating the tail as worth one mark, so getting it wrong is a small slip — a two-tailed test on a question whose wording gives a direction, or a one-tailed test on one that only says "different"
Fix the tail from the wording before writing anything else: the wrong tail caps the marks on the whole question, not just the hypotheses line
Mark schemes cap these outright. On 9709/62 O/N 2025 Q6(a) (6 marks) a two-tailed attempt "scores max B1B0M1A1M1 ... A0", i.e. out of ; on 9709/62 M/J 2024 Q6(a) (5 marks) "Two tail test scores maximum B0 M1 A1 M1 A0", out of ; on 9709/62 O/N 2024 Q7(b) (7 marks) "max 5/7". It runs the other way too — 9709/63 M/J 2023 Q5(b) is a two-tailed question, and a one-tailed method there is capped at "max 3/5". The standardising and comparison marks survive, because that arithmetic is still done correctly. What always goes is the hypotheses mark at the start and the conclusion mark at the end, since both are stated about a direction the question never asked about.
Choosing the tail based on which way the sample statistic happened to differ from the claimed value, rather than the direction stated in the question
The tail comes from the suspicion being tested (the wording), decided before looking at how the data came out
A test's direction has to be fixed in advance — choosing it after seeing which way the data leans is a form of bias that invalidates the significance level.
Your turn
- 1
A seed packet claims that of seeds will germinate. A gardener suspects the true germination rate is different from this. Write down suitable hypotheses and state whether the test is one-tailed or two-tailed.
Stuck? Show hint
"Different from" gives no direction — this is the two-tailed case.
Show solution
- 1
, — two-tailed, since no direction is claimed.
Answer, ; two-tailed.
- 1
- 2
A café claims the mean waiting time for a coffee is minutes. A regular customer believes it now takes longer. Write down suitable hypotheses for a test of this claim.
Stuck? Show hint
"Longer" is a direction — one-tailed, upper.
Show solution
- 1
, — one-tailed (upper), since "longer" gives a clear direction.
Answer, ; one-tailed (upper).
- 1
The rest of this note
Can you do all of these?
H₀ and H₁ are always statements about the population parameter, and H₀ always uses equality
Read the claim's direction carefully to fix H₁ and the tail: 'more/increased' → upper; 'less/decreased' → lower; 'different/changed' → two-tailed
For binomial/Poisson tests, use the exact distribution under H₀ — find P(X ⩽ observed) or P(X ⩾ observed), never a single point probability
For a mean test, state the necessary assumption (population normal, or large n for the CLT) and always divide by σ/√n, never σ alone
Match the critical value to the tails identified — a two-tailed test's critical value is more extreme than the same α's one-tailed value
Getting the tail wrong caps the marks on the WHOLE question, not just the hypotheses line — fix it from the wording before writing anything
Write H₀ and H₁ about the population parameter using its symbol: 'μ = 510' scores, 'mean = 510' does not, and any symbol you invent must be defined
Conclude in context, in the language of the original claim, without asserting certainty — 'insufficient/sufficient evidence', never 'H₀ is true'
Conclusions must be in context, not definite, and free of contradictions — 'insufficient evidence that the mean has decreased', never 'the mean has decreased' and never 'the mean has not decreased'
P(Type I error) uses the null distribution: for a DISCRETE test it is the critical region's actual probability (usually a little under α), but for a CONTINUOUS mean test it equals α exactly
P(Type II error) needs a specific true alternative value, and uses the SAME fixed critical region/value, re-standardised under that true value