Most people who read about AI models never read a methods section. They read a leaderboard: a table of model names and percentages, sorted. This part applies the series to that table. Every idea here has already appeared, and the point is to see them in the place where they are most often ignored.
A benchmark score is a rate
"Model A: 72.4 per cent." On a benchmark of 500 tasks that means 362 tasks passed and 138 failed, and Part 2 said what to do with a count of successes: put a Wilson interval on it. 72.4 per cent on 500 tasks is 68.3 to 76.1 per cent. Model B, one row down at 70.4 per cent, is 66.3 to 74.2. The intervals overlap by most of their width.
| Score | Wilson 95 per cent interval | Width | |
|---|---|---|---|
| Model A | 72.4% on 500 tasks | 68.3% to 76.1% | 7.8 points |
| Model B | 70.4% on 500 tasks | 66.3% to 74.2% | 7.9 points |
| Model A, if the benchmark had 100 tasks | 72.0% | 62.5% to 79.9% | 17.3 points |
The gap between the two models is 2 points. The uncertainty on each is 4 points either way. A leaderboard sorted to one decimal place is presenting an ordering that the data do not contain, and on a 100-task benchmark, which is common for agent evaluations, the interval is wider than the gap between the top five. Some leaderboards now print the interval, or a "confidence" grouping that says which models are indistinguishable. When one does not, you can compute it from the score and the task count in your head: the half-width, as a fraction, is about 1 divided by the square root of the number of tasks, so on 500 tasks it is about 4 points and on 100 about 9.
pass@k
Coding benchmarks often report pass@k: the probability that at least one of k attempts at a task passes its tests. pass@1 is the ordinary success rate. pass@10 answers "if I let the model try ten times and keep the best, how often does it get there", which is the relevant number for a workflow that does exactly that.
The subtle part is how it is estimated. Suppose you sample n = 10 attempts at a task and c = 3 of them pass. The tempting formula is to take the pass rate 0.3 and compute 1 − (1 − 0.3)⁵ = 0.83 as pass@5. That is biased, and noisy, because it plugs a noisy estimate into a curved formula. The estimator that Chen and colleagues introduced with Codex in 2021, and that every serious benchmark now uses, counts directly: the chance that none of k attempts drawn from the n passes is C(n − c, k) / C(n, k), the number of ways to choose k failures over the number of ways to choose k attempts, and pass@k is one minus that. For 3 of 10 it gives 0.92 at k = 5, against the naive 0.83. The difference is the bias, and it runs in the pessimistic direction here and grows as n shrinks.
| k | pass@k, 3 of 10 samples correct |
|---|---|
| 1 | 0.30 |
| 2 | 0.53 |
| 5 | 0.92 |
| 10 | 1.00 |
Two things to check when you see pass@k. How many samples n were drawn per task, since pass@k at k close to n is nearly meaningless (pass@10 from 10 samples is 1.0 whenever a single sample passed). And whether pass@1 was estimated from those same n samples, which it should be, or from one, which is the noisy version.
Seeds, and why temperature zero is not enough
A language model sampled at temperature zero picks the most likely token every time, and people assume that makes a run reproducible. It does not, reliably. Batched inference, floating-point arithmetic that depends on what else was running, and provider-side changes all mean that the same prompt at temperature zero can produce different outputs on different days, and agent runs, which chain hundreds of such calls, diverge from one another almost immediately. Part 1 showed six runs of the same configuration spread over 16 points, and that was with sampling on. It is smaller at temperature zero and it is not gone.
The consequence for a leaderboard is that a single run per task is a single draw from a distribution, and the score should be reported as a mean over several seeds with its spread, exactly as the verification study ran every cell four times. When a benchmark entry says "single run", the interval above is an underestimate, because it counts only the task-to-task variation and not the run-to-run variation on top.
Comparing two models on the same tasks
Here is the paired design of Part 11 in its most common form. Two models, the same 500 tasks. Model A passed 362, model B passed 352. Are they different?
The unpaired reading compares 72.4 per cent with 70.4 per cent using two independent intervals and concludes nothing, because the intervals overlap. But the tasks are shared, so the right question is about the tasks on which the models disagree. Split the 500 four ways:
| B passed | B failed | |
|---|---|---|
| A passed | 322 | 40 |
| A failed | 30 | 108 |
The 322 tasks both passed and the 108 both failed say nothing about which model is better. All the information is in the 70 discordant tasks: A won 40 of them, B won 30. McNemar's test asks whether a 40 to 30 split could have arisen if the two models were equally good, which is a coin-toss question, and the answer is p = 0.28. Not distinguishable. The paired bootstrap, resampling tasks, gives the gap as +2.0 points with an interval from −1.0 to +5.2, and tells the same story with a size attached.
Notice what the pairing did here compared with Part 11: there, pairing shrank the interval fivefold because every task moved the same way. Here it did not help much, because the models disagree in both directions about equally, which is what "they are about the same" looks like from inside the data. Pairing does not manufacture a difference. It removes the noise that is not about the difference, and then reports what is left.
Twenty benchmarks
Model cards report a dozen or more benchmarks, and the summary is usually a count: "wins on 12 of 20". Two models of identical ability would split twenty benchmarks around ten each, and one of them would win 14 or more about 6 per cent of the time by chance alone. Twelve of twenty is a coin coming up heads twelve times in twenty, which nobody would call evidence of anything. Part 7 applies in full: twenty comparisons are twenty tickets, and the honest presentation is every benchmark with its interval, or a correction across them, not a count of the ones that came out ahead.
Three more things a benchmark table cannot tell you, each a threat to validity from Part 9. Whether the tasks were in the training data, which is a construct problem: a model that has seen the answers is not being measured on the skill. Whether the tasks resemble the work you care about, which is the external one. And whether the grading is automatic and exact or a rubric and a judge, which decides what the score even means. None of that appears in the percentage, and the percentage is what gets quoted.
Try it: pass@k and the interval
Set the number of samples per task and how many passed, and read pass@k for every k. Then set the size of the benchmark and see how wide the interval on a score is.
Set samples to 10 and passes to 1: pass@1 is 0.10 and pass@10 is 1.00, which is the "already 1.0" trap. Then set the benchmark to 100 tasks and watch the interval on a 72 per cent score stretch to nearly 18 points.
In Python
The interval, the estimator, McNemar and the paired bootstrap, on the numbers above.
import numpy as np
from math import comb
from scipy import stats
from statsmodels.stats.contingency_tables import mcnemar
# 1. A benchmark score is a rate, so it has an interval
n_tasks, acc_a, acc_b = 500, 0.724, 0.704
for name, acc in (("model A", acc_a), ("model B", acc_b)):
ci = stats.binomtest(round(acc * n_tasks), n_tasks).proportion_ci(method="wilson")
print(f"{name}: {100*acc:.1f}% on {n_tasks} tasks, Wilson 95% {100*ci.low:.1f}% to {100*ci.high:.1f}%")
# 2. pass@k from n samples per task with c correct: the unbiased estimator
def pass_at_k(n, c, k):
return 1.0 if n - c < k else 1.0 - comb(n - c, k) / comb(n, k)
n, c = 10, 3
print("pass@k for a task with 3 of 10 samples correct: " + ", ".join(f"k={k}: {pass_at_k(n, c, k):.2f}" for k in (1, 2, 5, 10)))
print(f"the naive 1-(1-p)^k with p=0.3 at k=5: {1 - 0.7**5:.2f} (biased, here too low, when p is estimated from few samples)")
# 3. Two models on the same 500 tasks: only the tasks they disagree on carry information
both, only_a, only_b, neither = 322, 40, 30, 108 # A right & B right, A only, B only, both wrong
table = [[both, only_a], [only_b, neither]]
res = mcnemar(table, exact=True)
print(f"A right on {both + only_a}, B right on {both + only_b}; discordant {only_a} vs {only_b}; McNemar exact p = {res.pvalue:.3f}")
# 4. The paired bootstrap for the accuracy gap, resampling tasks
rng = np.random.default_rng(20260703)
a = np.array([1] * both + [1] * only_a + [0] * only_b + [0] * neither)
b = np.array([1] * both + [0] * only_a + [1] * only_b + [0] * neither)
idx = np.arange(500)
gaps = [(a[s] - b[s]).mean() for s in (rng.choice(idx, 500) for _ in range(5000))]
print(f"gap {100*(a.mean() - b.mean()):+.1f} points, paired bootstrap 95% {100*np.percentile(gaps, 2.5):+.1f} to {100*np.percentile(gaps, 97.5):+.1f}")
# 5. Twenty benchmarks: how many "wins" would two equal models split by chance?
print(f"two identical models on 20 benchmarks: chance one of them wins 14 or more = {1 - stats.binom.cdf(13, 20, 0.5):.1%}")
which prints:
model A: 72.4% on 500 tasks, Wilson 95% 68.3% to 76.1%
model B: 70.4% on 500 tasks, Wilson 95% 66.3% to 74.2%
pass@k for a task with 3 of 10 samples correct: k=1: 0.30, k=2: 0.53, k=5: 0.92, k=10: 1.00
the naive 1-(1-p)^k with p=0.3 at k=5: 0.83 (biased, here too low, when p is estimated from few samples)
A right on 362, B right on 352; discordant 40 vs 30; McNemar exact p = 0.282
gap +2.0 points, paired bootstrap 95% -1.0 to +5.2
two identical models on 20 benchmarks: chance one of them wins 14 or more = 5.8%
Where this leaves us
A leaderboard is a column of rates without their intervals, and the intervals are usually wider than the gaps. pass@k is a well-defined quantity with a correct estimator and a common wrong one. A single run per task hides the run-to-run noise that Part 1 was about. Two models on the same tasks should be compared on the tasks where they disagree, with McNemar's test or the paired bootstrap. And a count of wins across many benchmarks is a count of coin tosses until it comes with intervals or a correction. Everything else in the series was about earning the right to a number. This part was about reading one that someone else printed. The last part steps outside the whole framework and asks the question that p-values were never able to answer.
Next: The Bayesian Alternative: the probability that the hypothesis is true, which nothing so far has provided. Priors, posteriors and credible intervals, why a credible interval for a rate lands almost where Wilson did, how a sceptical prior shrinks a gap, the Bayes factor, and the region of practical equivalence, which is Part 8's bound in Bayesian clothes.