Here is a sentence from the kind of write-up I used to produce, and that most of us produce the first time we test anything: "With the verification tool switched on, the agent scored 74 out of 100. Without it, 62. The tool helps."

Every number in that sentence is real, and the conclusion might even be right. But a reader who knows what to ask will not accept it, and the questions they will ask are the subject of this series. How many times did you run it? How much do the scores move when you run the same thing twice? If I shuffled which runs had the tool, how often would I get a gap of 12 by accident? How sure are you of the 12, and would you have written it up the same way if it had come out as 2?

I learned those questions the slow way, by writing two papers about AI coding agents and getting the first one politely taken apart. The first was 90 runs, one statistical test, and a lot of careful description. The second was 1,116 runs, six hypotheses fixed in writing before I looked at anything, a permutation test for each, a correction for having tested six things, and confidence intervals built by resampling. Three of the six hypotheses held up and three did not, and the paper says so in one table. This series explains every term in that last sentence, using those two studies as the worked examples, so that by the end you can read a methods section, or write one, without guessing what the words mean.

No maths beyond arithmetic is needed. Where a formula would help I will show one, but every idea here can be understood by shuffling cards, and most of the demos are exactly that.

The vocabulary, with a study attached

A study is built out of a small number of objects, and the words for them are used loosely in conversation and precisely in papers. Here they are, each attached to the second study so that they are concrete.

Word What it means In the verification study
Run One complete execution of the thing you are measuring, start to finish, producing one outcome One agent building one web application once, then being graded
Condition A setting you deliberately vary, the thing you are comparing Which verification tools the agent had: none, a boot probe, a shell, a screenshot tool, and so on. Eight in all
Outcome measure The number you write down at the end of each run A functional score from 0 to 100 against a rubric, plus survival (did the app start), tokens used, and a human interface score
Rubric The fixed list of things a run is scored on, with the points for each Weighted criteria written and frozen before any run existed
Replicate A repeat of the same run with nothing changed, to see how much the result moves on its own Every combination was run four times by default
Cell One combination of every factor you control: this model, this task, this condition One model building one app under one tool setting. Four runs per cell
Sample size, n How many runs sit behind a number 192 runs in each of the four core conditions, 1,116 in total
Factor Something else that varies and that you record, whether or not you are studying it Which of six models did the run, and which of seven tasks

The first study had the same objects and much smaller numbers: 90 runs, 48 of them from one model, three from another, two from a third. When you see a paper quote an average for a group, the first thing to look for is the n beside it. An average of 48 runs and an average of 2 runs are printed with the same number of decimal places and mean entirely different things.

The twenty-four runs

The whole series uses one small dataset, small enough to print and to shuffle by hand, shaped like a miniature of the second study. Two models, one strong and one weak. Two conditions, with a verification tool and without. Six runs in each of the four cells. Twenty-four functional scores on a 100-point rubric.

Model A (strong) Model B (weak)
With the tool 88, 91, 79, 84, 95, 86 62, 58, 71, 49, 66, 55
Without the tool 78, 84, 70, 75, 89, 62 46, 53, 40, 58, 49, 37

The numbers are invented, but they were invented to behave like the real ones: the strong model scores about 30 points above the weak one, the tool is worth about 12 points on average, and runs of the same model under the same condition still land 20 or 30 points apart. Add up the twelve runs with the tool and divide by twelve and you get 73.7. The twelve without come to 61.8. The gap is 11.9 points, which I will call 12 from here on.

020406080100 functional score, 0 to 100 with without mean 73.7 mean 61.8 Model A, strong Model B, weak
The twenty-four runs as dots. The means differ by 12 points, but look at how far the dots of a single colour in a single row spread. That spread is the whole problem.

The same experiment, run again

Look at the six runs of the strong model with the tool: 88, 91, 79, 84, 95, 86. Nothing changed between them. Same model, same task, same prompt, same tools. The scores still range over 16 points, because an agent's run is a long chain of choices and any one of them can send it down a different path. The weak model without the tool ranges from 37 to 58, a spread of 21 points, for the same reason.

This is not a defect of the experiment. It is a property of the thing being measured, and it has a name: variability, or in everyday language, noise. Every measurement that involves a language model has it, and so does nearly every measurement of anything, which is why the whole apparatus of statistics exists. In the first study I noticed it and gave it a section, "within-configuration variability", after repeating a few configurations and watching identical setups land in different places. In the second study I designed for it from the start by running every cell four times, so that no single lucky or unlucky run could stand in for a condition.

Now the uncomfortable consequence. If six runs of the same thing can differ by 20 points, then the average of six runs is also a noisy number. Draw six fresh runs and the average moves. How much it moves is the question that decides whether a 12-point gap between two conditions is a finding or a coincidence, and it is the question that the demo below lets you feel before the next parts put numbers on it.

Try it: run it again

The demo draws runs from two imaginary conditions whose true averages you set, adds noise of the size you set, and reports the gap between the two sample averages. Press the button a few times. Then change the number of runs per condition and press it again.

Three things to try. Leave the true gap at 12 and the noise at 15, which is roughly what the twenty-four runs look like, and press Draw 20 times with 6 runs per condition. The measured gap will have ranged from close to zero to over 20. Then set the true gap to 0 and draw 20 times: you will see gaps of 10 or more appear from nothing, which is what "the tool helps, it scored 12 higher" has to be defended against. Finally set runs per condition to 100 and watch the measured gaps huddle around the true one. Nothing about the tool changed between those three pictures. Only the amount of evidence did.

The three questions

Everything that follows in this series is a way of answering one of three questions about a gap like the 12 points above, and it helps to know which is which before the vocabulary arrives.

How big is it? Not the number itself, which is 12, but 12 in units that mean something: points on a rubric that a human can picture, or a multiple of the noise, or a rate that went from 28 per cent to 89 per cent. This is the effect size, and it is the question the first study answered well and the one most papers answer first. Part 2 covers the ways of summarising many runs into one number, and Part 6 covers effect sizes properly.

Could chance alone have done it? If the tool did nothing, how often would shuffling twenty-four runs into two piles of twelve give a gap of 12 or more? That question has an exact answer, obtained by actually doing the shuffling, and the answer is the p-value. Parts 3, 4 and 5 are about it, and Part 7 is about what happens to it when you ask six questions at once.

How sure are you of the size? A gap of 12 measured on six runs a side could really be 2 or 22. A confidence interval is the range of gaps that the data cannot rule out, and Part 6 shows how to build one by resampling. Part 8 is about the awkward case where the interval includes zero, which is to say a null result, and how to defend one.

Then there are two questions that are not about arithmetic at all, and that decide whether any of the above is believed: was the study designed so that the gap could not have been caused by something other than the tool, and were the analytic choices made before the results were in view? Those are Parts 9 and 10, and they are where the first study was weakest and the second study spent most of its effort.

In Python

Everything in this part fits in a dozen lines. The first block computes the gap. The second runs the experiment again, twenty times, at three sample sizes, which is the demo above in code.

import numpy as np

# The twenty-four runs: functional score out of 100
with_tool    = [88, 91, 79, 84, 95, 86,   # model A
                62, 58, 71, 49, 66, 55]   # model B
without_tool = [78, 84, 70, 75, 89, 62,   # model A
                46, 53, 40, 58, 49, 37]   # model B

gap = np.mean(with_tool) - np.mean(without_tool)
print(f"mean with {np.mean(with_tool):.1f}, without {np.mean(without_tool):.1f}, gap {gap:+.1f}")

# Run the same experiment again: draw fresh runs from a world where the
# tool is really worth 12 points and identical runs differ by about 15
rng = np.random.default_rng(1)
for n in (6, 24, 100):
    gaps = [rng.normal(74, 15, n).mean() - rng.normal(62, 15, n).mean() for _ in range(20)]
    print(f"{n:>3} runs a side: 20 measured gaps ranged from {min(gaps):+.1f} to {max(gaps):+.1f}")

which prints:

mean with 73.7, without 61.8, gap +11.9
  6 runs a side: 20 measured gaps ranged from -9.9 to +21.1
 24 runs a side: 20 measured gaps ranged from -0.2 to +18.8
100 runs a side: 20 measured gaps ranged from +7.3 to +15.0

Where this leaves us

A number from an experiment is a sample from a spread, and the spread is often wider than the gap you are trying to see. "The tool scored 12 points higher" is the start of an argument, not the end of one. The rest of the series turns it into one, using the same twenty-four runs at every step, so that by the last part you can look at the six-row results table from the second study, with its effects, intervals, corrected p-values and the words "supported" and "not supported", and know exactly what each entry cost to earn.


Next: Middles, Spreads and Rates: the mean, the median, why token costs need the median, spreads and box plots, and what a pass rate of 16 out of 18 does and does not tell you, including the interval the textbook formula gets wrong and the ones that get it right.