Part 5 set the 0.05 line: call a result surprising if a world with no effect would produce it less than once in twenty. That is a fine rule for one question. The verification study asked six. This part is about what asking six does to the rule, and about the correction the study applied, which is why three of its p-values read 0.0006 and why two results that would have passed on their own did not.

Six chances

Suppose none of the six tools does anything at all. Every hypothesis is false, and every p-value is a draw from pure noise. Each test, on its own, has a 5 per cent chance of landing below 0.05 by bad luck, which is what the line means. The chance that at least one of six does is not 5 per cent. It is 1 minus the chance that all six stay above the line, which is 1 − 0.95⁶ = 26 per cent.

Hypotheses tested Chance of at least one false positive at the 0.05 line
1 5%
2 10%
3 14%
6 26%
10 40%
20 64%
50 92%

One in four. A study that tests six things and reports the one that came out significant is, a quarter of the time, reporting noise, and it will look exactly like a real finding. With twenty hypotheses it is more likely than not. This is the multiple comparisons problem, and the chance of at least one false alarm across a set of tests is called the family-wise error rate.

The problem is not that the tests are wrong. Each one is right on its own terms. The problem is the word "one of them". Ask enough questions of noise and noise will answer one of them.

Bonferroni: raise the bar

The oldest fix is the bluntest. If you want the family-wise rate back down to 5 per cent across six tests, demand more of each: use 0.05 / 6 = 0.0083 as the line for every test. This is the Bonferroni correction, and the equivalent way of stating it is to multiply each p-value by six and compare the result to 0.05, which is how papers usually print it.

It works. It is also wasteful, because it treats every test as if it were the one you were most worried about, and in doing so it throws away real effects. With six tests a p-value of 0.01 becomes 0.06 and fails. The waste is called a loss of power, the ability to detect an effect that is really there, and Part 8 has more to say about power.

Holm: raise the bar, then lower it step by step

Sture Holm's 1979 method gives the same protection as Bonferroni with less waste, and it is the one the verification study used, so it is worth doing by hand. The rule: sort the p-values from smallest to largest. Compare the smallest to 0.05 / 6. If it passes, compare the next to 0.05 / 5. Then 0.05 / 4, and so on. Stop at the first one that fails, and everything from there on fails with it.

Here it is on the study's six. The paper prints only the corrected values, but the raw ones can be recovered from them, and the table below reproduces the paper's corrected column exactly.

Rank Hypothesis Raw p-value Bar for this rank Passes? Corrected p-value
1 P1 shell, functional score 0.0001 0.05 / 6 = 0.0083 yes 0.0001 × 6 = 0.0006
2 P2 boot probe, survival 0.0001 0.05 / 5 = 0.0100 yes 0.0001 × 5 = 0.0005, raised to 0.0006
3 P3 shell, tokens 0.0001 0.05 / 4 = 0.0125 yes 0.0001 × 4 = 0.0004, raised to 0.0006
4 P6 shell, performance task 0.0225 0.05 / 3 = 0.0167 no, stop 0.0225 × 3 = 0.0675
5 P4 screenshots, interface score 0.0413 0.05 / 2 = 0.0250 no 0.0413 × 2 = 0.0826
6 P5 modification vs fresh build 0.9989 0.05 / 1 = 0.0500 no 0.9989 × 1 = 0.9989

Three details that puzzle people on first reading.

Why 0.0001 for the first three? Because 10,000 shuffles cannot produce a p-value smaller than one in ten thousand. Zero shuffles beat the observed gap, the procedure reports the floor, and the floor is 0.0001. Multiply by six and you get 0.0006, the number in the paper. It is not a measurement. It is the smallest value the method can print.

Why is the corrected p-value for rank 2 the same as rank 1? Because a corrected p-value is never allowed to be smaller than the one before it in the ranking. Otherwise a hypothesis could fail at rank 1 and its neighbour pass at rank 2 with a raw p-value that was larger, which would make no sense. So each corrected value is raised to at least the previous one, which is why the first three are all 0.0006.

Why did P4 fail at 0.0413 when P6 failed at 0.0225 and both would have passed alone? Because Holm stops at the first failure. Once P6 missed its bar of 0.0167, the procedure declared everything below it not significant, whatever its raw value, and the corrected values simply record that. Both P4 and P6 are results that would have been called significant as single tests and were not as members of a family of six. That is not the correction being unkind. It is the correction doing exactly what the table at the top says it should: refusing to hand chance six chances.

Holm and Bonferroni both hold the family-wise error rate at 5 per cent. Holm rejects everything Bonferroni does and sometimes more, and never less, so there is no reason to prefer Bonferroni except that it is easier to explain in a sentence.

What counts as a family

The correction needs a list of tests to correct across, and the list is a judgement. The verification study drew the line at its six primary hypotheses: those were the headline claims, decided in advance, so they formed one family, and chance was given one bar to clear across all of them.

The study also ran secondary contrasts, questions asked afterwards to understand the primary results: does the boot probe alone capture most of the shell's benefit (a gap of 2.4 points with an interval that includes zero, p = 0.14 uncorrected), and did the models that took more screenshots gain more from them (they did not). Those were reported with raw p-values and the word exploratory attached, which is the honest way to handle a question you did not pre-specify: report it, correct nothing, and label it so that no reader mistakes it for a claim that cleared the bar. What a study must not do is run twenty secondary comparisons, report the two that came out at p = 0.03, and leave the other eighteen in a drawer. The reader cannot correct for tests they cannot see, and Part 10 is about the design choices that prevent it.

One consequence is worth stating plainly. The intervals in the study's table were built for each hypothesis on its own and were not widened for the family, which is standard practice but means the intervals and the corrected p-values are answering slightly different questions. P4's interval of +0.8 to +13.4 excludes zero. P4's corrected p-value is 0.0826. Both are correct, and the paper says which is which.

When there are hundreds

Family-wise control is the right promise for a handful of headline claims: you want to be able to say that, with 95 per cent confidence, none of the supported hypotheses is a false alarm. It is the wrong promise for a screen of a thousand comparisons, where demanding that none be false would mean detecting almost nothing. There the usual method is the Benjamini-Hochberg procedure, which controls the false discovery rate: the expected fraction of the results you call significant that are false. Set it at 10 per cent and you accept that one in ten of your discoveries is noise, in exchange for finding the other nine. You will meet it in any paper with a heat map. For a six-hypothesis agent study, Holm is the right tool, and the paper chose it.

Try it: six dice

The first demo runs a set of tests on pure noise, over and over, and counts how often at least one comes up significant. The second is a Holm calculator: type any list of p-values and watch the step-down happen.

rankraw pbar 0.05 / (m − rank + 1)passesHolm-corrected pBonferroni p

Two things to try in the first demo. Leave the correction at none and slide the hypotheses up to 20: more than half the studies go red, every one of them a study of nothing. Then switch to Holm and slide to 50: the red stays around one in twenty. In the calculator, change the fourth value from 0.0225 to 0.0150 and watch P6 pass its bar and P4 fail its own, so that four survive instead of three: the whole table moves when one number does, because the ranks and the bars move together.

In Python

multipletests does the step-down in one call and reproduces the paper's column. Bonferroni is shown beside it so that you can see what Holm saves: the same three survive, but Bonferroni pushes P4 to 0.25 where Holm leaves it at 0.08.

import numpy as np
from statsmodels.stats.multitest import multipletests

# The six raw p-values, recovered from the paper's Holm-corrected column
names = ["P1 shell, functional", "P2 boot probe, survival", "P3 shell, tokens",
         "P4 screenshots, interface", "P5 modification vs fresh", "P6 shell, performance"]
raw = np.array([0.0001, 0.0001, 0.0001, 0.0413, 0.9989, 0.0225])

reject, holm, _, _ = multipletests(raw, alpha=0.05, method="holm")
_, bonf, _, _ = multipletests(raw, alpha=0.05, method="bonferroni")
for name, p, h, b, r in zip(names, raw, holm, bonf, reject):
    print(f"{name:<28} raw {p:.4f}  Holm {h:.4f}  Bonferroni {b:.4f}  {'supported' if r else 'not supported'}")

# How often noise wins at least once, without a correction
for m in (1, 6, 20):
    print(f"{m:>2} tests at 0.05: chance of at least one false positive = {1 - 0.95**m:.0%}")

which prints:

P1 shell, functional         raw 0.0001  Holm 0.0006  Bonferroni 0.0006  supported
P2 boot probe, survival      raw 0.0001  Holm 0.0006  Bonferroni 0.0006  supported
P3 shell, tokens             raw 0.0001  Holm 0.0006  Bonferroni 0.0006  supported
P4 screenshots, interface    raw 0.0413  Holm 0.0826  Bonferroni 0.2478  not supported
P5 modification vs fresh     raw 0.9989  Holm 0.9989  Bonferroni 1.0000  not supported
P6 shell, performance        raw 0.0225  Holm 0.0675  Bonferroni 0.1350  not supported
 1 tests at 0.05: chance of at least one false positive = 5%
 6 tests at 0.05: chance of at least one false positive = 26%
20 tests at 0.05: chance of at least one false positive = 64%

Where this leaves us

Every hypothesis you test is a ticket in a lottery that noise plays. Six tickets at the 0.05 line give noise a one-in-four chance of a win, so a study that tests six things must raise the bar, and Holm's step-down raises it just enough: each p-value is compared to 0.05 divided by the number of hypotheses still in play, and the first failure ends the run. Applied to the verification study's six, it turns three raw values of 0.0001 into 0.0006, and two results that would have passed alone into 0.0675 and 0.0826, which the paper reports without complaint. The family is the set of headline claims made in advance. Anything asked afterwards is exploratory, and is labelled so. What remains is the hardest kind of result to report, the one where the honest answer is "nothing happened", and the next part is about how to say it and be believed.


Next: Defending a Null: why "not significant" cannot mean "no effect", the equivalence test that can, choosing the smallest effect worth caring about before the data arrive, statistical power, and how many runs it takes to bound a null within five points.