Part 4 produced two numbers by counting shuffles: 0.093 and 0.0014. This part is about what such a number means, and it is short on new machinery and long on wording, because the p-value is the most misread number in science and the misreadings are all small slips of language.

The definition, in one sentence

The p-value is the probability of seeing a result at least as extreme as the one you saw, if the null hypothesis were true.

Every word is doing work. "At least as extreme" is why Part 4 counted shuffles with a gap of 11.9 or more, not exactly 11.9. "If the null hypothesis were true" is the assumption under which the whole calculation runs: we dealt the cards as if the tool did nothing, so every shuffled gap was generated by a world in which the tool does nothing. The p-value describes that world. It says how often that world would produce what we saw.

So p = 0.0014 means: in a world where the tool does nothing, and where labels are shuffled within each model, a gap as big as ours shows up about once in 740 tries. Either we are looking at that one try in 740, or the world is not the one where the tool does nothing. That is the entire argument, and it is a good one, but it is an argument about surprise, not about truth.

Four misreadings

What people say Why it is wrong What is true
"p = 0.0014, so there is a 0.14 per cent chance the tool does nothing" The p-value was computed assuming the tool does nothing. It cannot then tell you how likely that assumption is. That would need to know how plausible the claim was before the study, which the p-value never sees The data would be rare if the tool did nothing. How rare the null was to begin with is a separate question
"p = 0.0014, so there is a 99.86 per cent chance the tool helps" Same error, reversed Nothing in a p-value is a probability about the claim
"p = 0.093, so the result was probably due to chance" The p-value is the chance of the result given chance alone, not the chance of chance alone given the result. The two are as different as "most cats have four legs" and "most four-legged things are cats" A gap this big is not unusual under chance. That is all
"p = 0.0006 is a bigger effect than p = 0.04" The p-value mixes the size of the effect with the size of the study. A trivial effect with enormous n gets a tiny p-value Look at the effect size and its interval (Part 6) for how big. The p-value only says how surprising

The last row matters most in practice, and deserves its own section.

Where 0.05 came from

In 1925 Ronald Fisher, working out how to analyse crop experiments, wrote that one in twenty was a convenient line to draw: results rarer than that under the null were worth treating as real, results commoner were not. He did not derive the number and did not claim it was special. It stuck because everyone needed the same line, and a shared arbitrary line is more useful than a hundred private reasonable ones.

That is what the convention is for. Set the line before the study, and nobody can move it afterwards to let a result through. It is a standard of proof, the "beyond reasonable doubt" of Part 3's courtroom, and it is set at 95 per cent confidence because that is what the field agreed, not because 94 is wrong.

The cost is cliff-edge thinking. A result at p = 0.049 gets called statistically significant and one at p = 0.051 does not, and the two are nearly identical evidence. The best practice, which the verification study follows, is to report the p-value itself rather than just the verdict, and to report the effect and its interval beside it, so that a reader can see the 0.051 for what it is: a result that just missed a line, not a result that vanished.

Significant is not the same as important

The word "significant" has a technical meaning here that is narrower than its everyday one. It means "would be surprising under the null". It does not mean big, useful, or worth acting on. Two examples from the same table.

Suppose a tool raised the functional score by half a point, from 62.0 to 62.5, and you ran 20,000 runs on each side. The noise between runs is about 15 points, so the average of 20,000 runs is pinned down to within about 0.15 points, and a gap of 0.5 is more than three times that. The p-value would be around 0.001. Statistically significant, and of no practical interest to anyone: half a point on a 100-point rubric is less than the difference between two human graders.

Now P6 from the verification study: on the one performance-critical task, agents with a shell scored 15.3 points higher on the human interface score, the biggest effect in the paper, with about 28 runs a side. Its corrected p-value came out at 0.0675, above the line, and the hypothesis was reported as not supported. Fifteen points is a lot. Twenty-eight runs is not many. The p-value blends the two and says "not quite surprising enough".

Effect Runs per side p-value Significant? Important?
Imaginary half-point tool +0.5 20,000 about 0.001 yes no
P6, the shell on the performance task +15.3 about 28 0.0675 no probably, if it holds up

The p-value answers "how sure", not "how much". Reading it as "how much" is the error that makes big studies look impressive and small studies look worthless, when the opposite is often closer to the truth.

Not significant is not no effect

The mirror image, and the one that does the most damage in write-ups. P4 in the verification study: the screenshot tool raised the interface score by 6.9 points on the two visual tasks, and the interval around that gap ran from 0.8 to 13.4 points, which is to say the data were consistent with a modest benefit and not with none. The corrected p-value was 0.0826. The hypothesis was reported as not supported.

What the study did not say, and what a careless reader would take from it, is that screenshots do not help. The honest sentence is the one Part 3 gave for the courtroom: the evidence was not enough. The effect may well be real. A study with more visual tasks would probably settle it. "Not supported" is a statement about this study's power to detect, not about the world, and Part 8 is about what it would take to make the stronger claim that an effect is absent.

Fisher's exact test: the same idea, for counts

The first study's headline was a pair of counts, not a pair of averages: at the higher reasoning setting 16 of 18 runs were perfect on the first try, at the lower setting 5 of 18. The permutation idea works for counts exactly as it does for scores, and the version for a two-by-two table has a name, Fisher's exact test, after the same Fisher.

Perfect first try Not perfect Total
Higher reasoning 16 2 18
Lower reasoning 5 13 18
Total 21 15 36

The shadow says reasoning effort does not matter. If so, the 21 perfect runs were going to be perfect regardless, and which 18 of the 36 runs carried the "higher" label is just dealing. So: deal 18 of the 36 runs to the "higher" pile at random, and ask how often 16 or more of the 21 perfect ones land there, or the mirror image, 5 or fewer. Every possible deal can be counted, which is why the test is called exact, and the answer is p = 0.0005. Fewer than one deal in two thousand produces a split as lopsided as the one observed. The first study reported it as p < 0.001, and it was the only formal test in the paper, used for the one comparison that had enough runs to bear it.

Fisher's exact test is the tool for any "k of n versus j of m" claim with small numbers, which describes most pass-rate comparisons in agent research. Its interval-shaped companion is Newcombe's method from Part 2, which for the same table gave a gap of 61 points with a range of 29 to 78.

The six p-values, read aloud

Here are the verification study's six results again, this time with the p-values. These are corrected for the fact that six hypotheses were tested at once, a step explained in Part 7, and the correction is why three of them are exactly 0.0006 and why one is 0.9989.

Effect Corrected p-value How to read it
P1 shell, functional score +12.3 points 0.0006 None of the 10,000 shuffles beat the observed gap. The smallest p-value the procedure can produce is one in ten thousand, and 0.0006 is that, multiplied by six for the correction. Read it as "fewer than one shuffle in ten thousand"
P2 boot probe, survival +13 points 0.0006 Same. The floor of the method, not a measurement of exactly 0.0006
P3 shell, tokens ×2.2 0.0006 Same. The cost is as certain as the benefit
P4 screenshots, interface score +6.9 points 0.0826 Just above the line. A modest gap on two tasks, suggestive, not confirmed. The uncorrected p-value was about 0.04, and the correction for six tests pushed it over
P5 modification vs fresh build −0.7 points 0.9989 Shuffled gaps beat the observed one almost every time, because the observed gap is nearly zero. This is what "no detectable difference" looks like, and Part 8 explains why it is not the same as "no difference"
P6 shell, performance task +15.3 points 0.0675 The largest effect and a small sample. Above the line. The uncorrected p-value was about 0.02

Notice that P4 and P6 would both have been called significant on their own. They were not, because they were two of six, and the next part is about why that is the right call and not an excess of caution.

Try it: the p-dial

The demo draws a fresh experiment each time you press it, from two conditions with the true gap and noise you set, and computes the p-value by shuffling, exactly as Part 4 did. Press it twenty times and watch the p-value dance.

Three things to try. Set the true gap to 0 and press Run it 20 times: you will usually see one green dot, a false alarm, which is the one-in-twenty the convention accepts. Set the gap back to 12 with 6 runs a side and run 20 times: the dots spread from 0.001 to 0.9, all from the same experiment, which is why a single p-value from a small study should be held loosely. Then set 60 runs a side and watch them all crowd into the left.

In Python

Fisher's exact test is one call. SciPy's permutation_test does Part 4's shuffling for you, and the Welch t-test, the classical alternative, agrees with it here because the scores are roughly bell-shaped.

from scipy import stats

# Fisher's exact test on the first-try perfect runs: 16 of 18 vs 5 of 18
table = [[16, 2],    # higher reasoning: perfect, not perfect
         [5, 13]]    # lower reasoning
result = stats.fisher_exact(table, alternative="two-sided")
print(f"Fisher exact p = {result.pvalue:.4f}")

# The same idea as a permutation test on scores: SciPy will do the shuffling for you
import numpy as np
with_tool    = np.array([88, 91, 79, 84, 95, 86, 62, 58, 71, 49, 66, 55])
without_tool = np.array([78, 84, 70, 75, 89, 62, 46, 53, 40, 58, 49, 37])
res = stats.permutation_test((with_tool, without_tool),
                             lambda a, b: a.mean() - b.mean(),
                             permutation_type="independent", n_resamples=10_000,
                             alternative="two-sided", random_state=20260703)
print(f"permutation test: gap {res.statistic:+.2f}, p = {res.pvalue:.3f}")

# The classical alternative, for comparison
t = stats.ttest_ind(with_tool, without_tool, equal_var=False)
print(f"Welch t-test: p = {t.pvalue:.3f}")

which prints:

Fisher exact p = 0.0005
permutation test: gap +11.92, p = 0.096
Welch t-test: p = 0.091

Where this leaves us

A p-value is the frequency with which a world where the null holds would produce data as extreme as yours. It is not the probability that the null holds, and it says nothing about how big the effect is. The 0.05 line is a shared convention for "surprising enough", set in advance so that it cannot be moved. Significant means surprising, not important, and not significant means not surprising enough, not absent. With that vocabulary in hand, the next part answers the question the p-value cannot: how big, and how sure.


Next: How Big, and How Sure: effect sizes in units that mean something, the confidence interval and the one thing it promises, and the bootstrap, which builds an interval by resampling the runs you have, one resample at a time.