Most people never read a methods section. They read a league table: a list of school names and percentages, sorted. This part applies the series to that table. Every idea here has already appeared, and the point is to see them in the place where they are most often ignored.
A league-table figure is a rate
"School X: 72.4 per cent." If School X has 500 pupils in the year, that means 362 passed and 138 failed, and Part 2 said what to do with a count of successes: put a Wilson interval on it, which is the range of true pass rates that a count like 362 of 500 cannot rule out, built by a formula that behaves itself even when the rate is near 0 or 100 per cent. 72.4 per cent of 500 pupils is 68.3 to 76.1 per cent. School Y, one row down at 70.4 per cent, is 66.3 to 74.2. The intervals overlap by most of their width.
| Pass rate | Wilson 95 per cent interval | Width | |
|---|---|---|---|
| School X | 72.4% of 500 pupils | 68.3% to 76.1% | 7.8 points |
| School Y | 70.4% of 500 pupils | 66.3% to 74.2% | 7.9 points |
| School X, if it had 100 pupils | 72.0% | 62.5% to 79.9% | 17.3 points |
The gap between the two schools is 2 points. The uncertainty on each is 4 points either way. A league table sorted to one decimal place is presenting an ordering that the data do not contain, and for a school with 100 pupils in the year, which is an ordinary size, the interval is wider than the gap between the top five. Some league tables now print the interval, or a grouping that says which schools cannot be told apart. When one does not, you can work it out from the percentage and the number of pupils in your head: the half-width, as a fraction, is about 1 divided by the square root of the number of pupils, so with 500 pupils it is about 4 points and with 100 about 9.
Best of k sittings
Schools often publish a pass rate that counts each pupil's best result over several sittings: a pupil who failed in June and passed the resit in November counts as a pass. Call that figure best of k sittings: the chance that a pupil allowed k sittings passes at least once. Best of 1 is the ordinary first-sitting pass rate. Best of 10 answers "if a pupil could sit the paper ten times and the best one counted, how often would they get there", which is the relevant number for a school that lets pupils resit until they pass, and not the number a parent has in mind when they read "pass rate".
The subtle part is how it is estimated. Suppose a school's records hold n = 10 sittings of the paper by pupils at a given standard, and c = 3 of them were passes. The tempting formula is to take the pass rate 0.3 and compute 1 − (1 − 0.3)⁵ = 0.83 as the best-of-5 rate. That is biased, meaning it lands on the wrong side of the truth on average, and noisy, because it plugs a noisy estimate into a curved formula. The estimator that gets it right, an estimator being nothing more than a recipe for turning the records into a figure, counts directly: deal yourself k of the n recorded sittings at random and ask how often none of them is a pass. The chance of that is C(n − c, k) / C(n, k), the number of ways to choose k failures over the number of ways to choose k sittings, and the best-of-k rate is one minus it. For 3 of 10 it gives 0.92 at k = 5, against the plug-in 0.83. The difference is the bias, and it runs in the pessimistic direction here and grows as n shrinks.
| k | best of k sittings, 3 passes in 10 recorded |
|---|---|
| 1 | 0.30 |
| 2 | 0.53 |
| 5 | 0.92 |
| 10 | 1.00 |
Two things to check when you see a best-of-k figure. How many sittings n were recorded, since best of k at k close to n is nearly meaningless (best of 10 from 10 recorded sittings is 1.0 whenever a single sitting was a pass). And whether the first-sitting rate printed beside it was worked out from those same records, which it should be, or taken from somewhere else, in which case the two figures are not about the same pupils and cannot be read against each other.
Why one sitting is not the pupil
People assume that a pupil's mark is the pupil's: sit the same paper on another day and the same mark comes out. It does not, reliably. A bad night's sleep, a cold, which questions happened to come up, a guess that went the other way, all mean that the same pupil on a different day gets a different mark, and a long paper, where a stumble early on unsettles everything after it, drifts further. Even a multiple-choice paper marked by machine moves, because the guesses change. Part 1 showed six pupils with the same preparation spread over 16 marks. Most of that is that they were six different people. Some of it is that each sat the paper on one particular day, and the same six on another day would have spread differently. The second kind of spread is smaller than the first and it is not gone.
The consequence for a league table is that a figure built from one sitting per pupil is a single draw from a distribution: the same school with the same pupils on another day would have printed a different percentage. The teacher's trial dealt with the same problem by never letting one observation stand for a cell, with at least six pupils in every school-and-condition cell by its third year. A league table rests on one sitting per pupil, and so the interval above is an underestimate, because it counts only the pupil-to-pupil variation and not the pupil-day variation on top.
Comparing two papers on the same pupils
Here is the paired design of Part 11 in its most common form. Say School X's 500 pupils also sat the new exam paper that the board is bringing in next year, as a trial. 362 passed the old paper, which is the 72.4 per cent at the top of this part, and 352 passed the new. Is one paper harder than the other?
The unpaired reading compares 72.4 per cent with 70.4 per cent using two independent intervals, as if different pupils had sat the two papers, and concludes nothing, because the intervals overlap. But the pupils are shared, so the right question is about the pupils on whom the two papers disagree. Split the 500 four ways:
| Passed the new | Failed the new | |
|---|---|---|
| Passed the old | 322 | 40 |
| Failed the old | 30 | 108 |
The 322 pupils who passed both papers and the 108 who failed both say nothing about which paper is harder. All the information is in the 70 pupils the papers disagree on: 40 passed only the old, 30 passed only the new. McNemar's test asks whether a 40 to 30 split could have arisen if the two papers were equally hard, which is a coin-toss question, and the answer is p = 0.28: a split that lopsided or more would turn up about 28 times in 100 by chance. Not distinguishable. The paired bootstrap, which redraws the 500 pupils with replacement, keeping each pupil's two results together, and recomputes the gap each time, gives it as +2.0 points with an interval from −1.0 to +5.2, and tells the same story with a size attached.
Notice what the pairing did here compared with Part 11: there, pairing by topic shrank the interval fivefold because every topic moved the same way. Here it did not help much, because the papers disagree in both directions about equally, which is what "they are about the same" looks like from inside the data. Pairing does not manufacture a difference. It removes the noise that is not about the difference, and then reports what is left.
Twenty league tables
A school's prospectus reports a dozen or more subject tables, and the summary is usually a count: "comes top in 12 of 20". Set it against the school down the road. Two schools of identical quality would split twenty subject tables around ten each, and one of them would come out ahead in 14 or more about 6 per cent of the time by chance alone. Twelve of twenty is a coin coming up heads twelve times in twenty, which nobody would call evidence of anything. Part 7 applies in full: twenty comparisons are twenty lottery tickets, and the honest presentation is every subject with its interval, or a correction across them, not a count of the ones that came out ahead.
Three more things a league table cannot tell you, each a threat to validity from Part 9. Whether the pupils were taught the exact questions, which is a construct problem, meaning the number is measuring something other than what it claims: a pupil who has been drilled on the paper is not being measured on the skill. Whether the questions resemble what pupils will need later, which is the external one, meaning whether the result carries beyond the exam hall. And whether the paper is multiple choice marked by machine or long answers marked by a reader, which decides what the percentage even means. None of that appears in the percentage, and the percentage is what gets quoted.
Try it: best of k sittings and the interval
Set the number of sittings recorded and how many of them were passes, and read the best-of-k rate for every k. Then set the number of pupils behind a school's figure and see how wide the interval on it is.
Set sittings recorded to 10 and passes to 1: best of 1 is 0.10 and best of 10 is 1.00, which is the "already 1.0" trap. Then set the school to 100 pupils and watch the interval on a 72 per cent pass rate stretch to nearly 18 points.
In Python
The interval, the estimator, McNemar and the paired bootstrap, on the numbers above.
import numpy as np
from math import comb
from scipy import stats
from statsmodels.stats.contingency_tables import mcnemar
# 1. A league-table figure is a rate, so it has an interval
n_pupils, rate_x, rate_y = 500, 0.724, 0.704
for name, rate in (("School X", rate_x), ("School Y", rate_y)):
ci = stats.binomtest(round(rate * n_pupils), n_pupils).proportion_ci(method="wilson")
print(f"{name}: {100*rate:.1f}% of {n_pupils} pupils passed, Wilson 95% {100*ci.low:.1f}% to {100*ci.high:.1f}%")
# 2. Best of k sittings from n recorded sittings with c passes: the unbiased estimator
def best_of_k(n, c, k):
return 1.0 if n - c < k else 1.0 - comb(n - c, k) / comb(n, k)
n, c = 10, 3
print("best of k sittings, 3 passes in 10 recorded: " + ", ".join(f"k={k}: {best_of_k(n, c, k):.2f}" for k in (1, 2, 5, 10)))
print(f"the plug-in 1-(1-p)^k with p=0.3 at k=5: {1 - 0.7**5:.2f} (biased, here too low, when p is estimated from few sittings)")
# 3. Two papers sat by the same 500 pupils: only the pupils the papers disagree on carry information
both, only_old, only_new, neither = 322, 40, 30, 108 # passed both, old paper only, new paper only, failed both
table = [[both, only_old], [only_new, neither]]
res = mcnemar(table, exact=True)
print(f"passed the old paper {both + only_old}, the new {both + only_new}, disagreements {only_old} vs {only_new}, McNemar exact p = {res.pvalue:.3f}")
# 4. The paired bootstrap for the gap in pass rate, resampling pupils
rng = np.random.default_rng(20260703)
old = np.array([1] * both + [1] * only_old + [0] * only_new + [0] * neither)
new = np.array([1] * both + [0] * only_old + [1] * only_new + [0] * neither)
idx = np.arange(500)
gaps = [(old[s] - new[s]).mean() for s in (rng.choice(idx, 500) for _ in range(5000))]
print(f"gap {100*(old.mean() - new.mean()):+.1f} points, paired bootstrap 95% {100*np.percentile(gaps, 2.5):+.1f} to {100*np.percentile(gaps, 97.5):+.1f}")
# 5. Twenty league tables: how often would two equal schools split them 14 to 6 or worse by chance?
print(f"two schools of equal quality across 20 subject tables: chance one of them comes top in 14 or more = {1 - stats.binom.cdf(13, 20, 0.5):.1%}")
which prints:
School X: 72.4% of 500 pupils passed, Wilson 95% 68.3% to 76.1%
School Y: 70.4% of 500 pupils passed, Wilson 95% 66.3% to 74.2%
best of k sittings, 3 passes in 10 recorded: k=1: 0.30, k=2: 0.53, k=5: 0.92, k=10: 1.00
the plug-in 1-(1-p)^k with p=0.3 at k=5: 0.83 (biased, here too low, when p is estimated from few sittings)
passed the old paper 362, the new 352, disagreements 40 vs 30, McNemar exact p = 0.282
gap +2.0 points, paired bootstrap 95% -1.0 to +5.2
two schools of equal quality across 20 subject tables: chance one of them comes top in 14 or more = 5.8%
Where this leaves us
A league table is a column of rates without their intervals, and the intervals are usually wider than the gaps. A best-of-k pass rate is a well-defined quantity with a correct estimator and a common wrong one. One sitting per pupil hides the pupil-day noise that sits underneath the pupil-to-pupil spread Part 1 was about. Two papers sat by the same pupils should be compared on the pupils where they disagree, with McNemar's test or the paired bootstrap. And a count of subjects in which a school comes top is a count of coin tosses until it comes with intervals or a correction. Everything else in the series was about earning the right to a number. This part was about reading one that someone else printed. The last part steps outside the whole framework and asks the question that p-values were never able to answer.
Next: The Bayesian Alternative: the probability that the hypothesis is true, which nothing so far has provided. Priors, posteriors and credible intervals, why a credible interval for a rate lands almost where Wilson did, how a sceptical prior shrinks a gap, the Bayes factor, and the region of practical equivalence, which is Part 8's bound in Bayesian clothes.