Everything in this series so far could be done perfectly and still produce a result nobody should believe. The permutation test can be exact, the intervals honest, the correction applied, the grader blind, and the finding still worthless, if the choices that led to it were made with the result in view. This last part is about timing: the difference between deciding how to analyse a study before the data exist and deciding afterwards, and the ways of proving which one you did.

The garden of forking paths

Every analysis is a sequence of small, reasonable decisions. Which outcome is the headline: the functional score or the interface score? Do the two models get pooled or reported separately? Does the run that crashed for an unrelated reason get dropped? Do you stop at 24 runs, as planned, or add a few more because the result was "nearly there"? Each choice is defensible on its own. Each has two or three options. And at every fork, an honest researcher who has already seen the data will tend, without noticing, to take the branch that leads somewhere.

Andrew Gelman and Eric Loken called this the garden of forking paths, and their point was that it needs no dishonesty. Nobody has to try twenty tests and report the best one. It is enough to make five reasonable choices, one at a time, each slightly informed by what the data looked like, to arrive at a p-value of 0.03 that would have been 0.30 down a different path. The paths that were not taken are invisible in the write-up, so the reader cannot correct for them, and Part 7 showed what uncorrected multiplicity does to the false-positive rate.

The demo makes the garden concrete. The data are pure noise: two conditions with identical true scores on every outcome. Walk the garden and see how often you can find a "finding".

Try it: walk the garden

Each checkbox is a choice you are allowed to make after looking at the data. With none ticked, the analysis is the one fixed in advance: functional score, both models, every run, all 24. With a box ticked, you get to pick whichever option gives the smaller p-value. The button analyses a hundred datasets of pure noise and counts how many yield a significant result somewhere in the garden.

Tick all four boxes and press the button: the grid goes a third to a half red, and every red square is a paper that could have been written, with a real p-value below 0.05, about a tool that does nothing. Then press Freeze every choice. The grid returns to a scatter of about five red squares. Nothing about the data changed. What changed is that the analyst no longer got to choose after looking.

The changelog of the study I am planning has the shape of a walk through this garden, and it is worth saying so plainly: a model roster that changed as providers released and withdrew models, a token budget raised twice, a sampling change in one version, and sets of early "smoke" runs used to check that the tasks behaved. The verification study had the same on a smaller scale, two dozen trial runs on two tasks before the real ones. All of that is legitimate. It is exploratory work, and the thing that makes it legitimate is that it was done on pilot data that never entered the analysis, and that the hypotheses were fixed before the confirmatory runs began. Which brings us to the ways of fixing them.

Five ways to fix your choices

The options below are in order of strength, and each is the previous one plus a witness.

Conventional, or post hoc. Run the experiment, look at the data, decide what to report, write it up. Fast and flexible, and every analytic choice is made with the results in view, which is why a reviewer of a null or mixed result can always ask: would you have reported it this way if it had come out differently? The first study is mostly this, and its results section is honest about which analyses were post hoc, the criterion-level and token breakdowns done after the headline was known. For a study whose most likely headline is "no effect", it is the weakest position to argue from.

Exploratory then confirmatory. Split the work in two. A pilot phase in which you are allowed to look, tinker and change your mind, and a confirmatory phase in which nothing changes and the pilot data are excluded. The line between them is the freeze: the moment the hypotheses, the design, the sample size, the outcomes, the exclusion rules and the analysis code stop moving, and after which the real runs begin. The verification study did this: the smoke runs were excluded by design, and the paper describes the six hypotheses as "decided in advance". A changelog that records every change and its date is good evidence of a freeze, and yours will be better evidence still if the changes stop before the confirmatory runs start.

Preregistration. The same freeze, with a third-party timestamp. You deposit the plan, hypotheses, design, sample size and its justification, outcomes, exclusion rules, analysis code, with a registry such as the Open Science Framework or AsPredicted, and the registry stamps it, so that nobody has to take your word for the date. It is not peer review and not a publication route. You still submit the paper wherever you like afterwards and cite the registration. OSF offers several templates, including one for studies of existing data, and lets you embargo the registration for up to four years if you do not want the design public before the paper. AsPredicted is the lighter option: nine questions, a one-page PDF, done in an afternoon.

The verification study did not do this, and says so in a sentence I would encourage anyone to copy: its hypotheses were "pre-specified rather than pre-registered in the external-registry sense". For that study it was a fair choice, because three of its six results were strong positives that no amount of after-the-fact tuning would have been needed to produce. For the study I am planning, whose likely result is a bounded null, I think registration is mandatory rather than optional, because the equivalence bounds of Part 8 and the Holm families of Part 7 are exactly the choices a sceptic will suspect were tuned once the interval was in view.

Registered Reports. The strongest form. You submit the plan itself for peer review before collecting data, as a Stage 1 manuscript: introduction, methods, analysis plan, power analysis, and no results. If it passes, you receive in-principle acceptance: the paper will be published whatever the results, provided you follow the protocol. Stage 2 adds the results and discussion, and reviewers check that you did what you said, not whether they like the answer. Peer Community In Registered Reports runs the process independently of any journal, and a list of friendly journals, including Royal Society Open Science and the PeerJ titles, commit to publishing a Stage 2 that PCI RR has recommended without further review. The costs are real. Stage 1 review adds months before the first run, and the target venue matters, since many engineering venues do not offer the track. The benefit is precisely suited to a null-result study: a mixed or null result is guaranteed a home, and the design gets adversarial scrutiny while it can still be changed cheaply.

Results-blind review. A halfway house some venues offer: the manuscript is reviewed with the results section redacted, but after the data exist. It removes outcome bias from the reviewers, who cannot favour a positive result they have not seen, but not from the author, whose analytic choices were still made with the results in view. Less common, and worth knowing the name.

When are the choices fixed? Who witnesses it? Publication guaranteed? Cost
Post hoc after the results nobody no none
Exploratory then confirmatory before the confirmatory runs you, and your changelog no a pilot phase, and the discipline to exclude it
Preregistration before the runs a registry, with a timestamp no an afternoon, and no changes afterwards without saying so
Registered Report before the runs peer reviewers yes, if the protocol is followed months of review before the first run, and a venue that offers it
Results-blind review after the data, before the reviewers see the results reviewers, partially no little

Preregistration in machine learning

It is still rare. NeurIPS ran a "pre-registration experiment" workshop in 2020 and 2021 that reviewed proposals before their results and published the accepted ones in the Proceedings of Machine Learning Research, which shows the community has tried the model, but it never became a mainstream track, and in practice most ML papers that preregister do so on OSF and submit conventionally. Two consequences for anyone writing an agent study. A preregistered one is still somewhat novel in this field, which is a selling point in the methods section. And reviewers may be unfamiliar with the conventions, so the paper should explain them in a paragraph: what was registered, when, where, and what changed afterwards, if anything. This series is, in part, that paragraph written long.

In Python

The garden, walked by a program: with one fixed, stratified path, noise reaches p < 0.05 about one time in twenty, and with twenty-four paths it does so about half the time. The last block is the cheapest possible freeze: write the plan as a file and fingerprint it, then deposit the fingerprint or the file with a registry before the first run.

import numpy as np, hashlib, json
from scipy import stats

# Walk the garden: pure noise, and four choices made after looking
rng = np.random.default_rng(20260703)
def one_dataset():
    model = np.repeat([82, 53], 12)                      # 12 runs per model, 6 per condition
    cond = np.tile(np.repeat([0, 1], 6), 2)              # 0 = with, 1 = without: no true effect anywhere
    functional = model + rng.normal(0, 15, 24)
    interface = 0.6 * functional + 0.4 * (model + rng.normal(0, 15, 24))
    order = rng.permutation(24)                          # the order the runs were collected in
    return cond[order], model[order], functional[order], interface[order]

def smallest_p(cond, model, functional, interface, open_choices):
    outcomes = [functional, interface] if open_choices else [functional]
    subgroups = [None, 82, 53] if open_choices else [None]
    stops = [24, 16] if open_choices else [24]
    drops = [False, True] if open_choices else [False]
    ps = []
    for outcome in outcomes:
        for sub in subgroups:
            for stop in stops:
                keep = (np.arange(24) < stop) & (True if sub is None else model == sub)
                centred = outcome - np.where(model == 82, 82, 53)   # stratify: remove each model's own level
                for drop in drops:
                    a, b = centred[keep & (cond == 0)], centred[keep & (cond == 1)]
                    if drop:                                    # drop the run that most hurts the case
                        a, b = (np.delete(a, a.argmin()), b) if a.mean() >= b.mean() else (a, np.delete(b, b.argmin()))
                    if len(a) >= 3 and len(b) >= 3:
                        ps.append(stats.ttest_ind(a, b).pvalue)
    return min(ps)

for open_choices in (False, True):
    found = sum(smallest_p(*one_dataset(), open_choices) < 0.05 for _ in range(500))
    print(f"{'24 paths' if open_choices else 'one fixed path'}: {found / 5:.0f}% of noise datasets yield p < 0.05")

# The freeze, as a file: write the plan down and fingerprint it before the data exist
plan = {"hypotheses": ["P1: shell vs none, functional score, higher"], "test": "stratified permutation, 10000 shuffles",
        "correction": "Holm over the primary family", "runs_per_cell": 4, "seed": 20260703}
text = json.dumps(plan, sort_keys=True, indent=2)
print("plan fingerprint:", hashlib.sha256(text.encode()).hexdigest()[:16], "(deposit this, or the file, with a registry)")

which prints:

one fixed path: 5% of noise datasets yield p < 0.05
24 paths: 52% of noise datasets yield p < 0.05
plan fingerprint: 637a2e71521d5180 (deposit this, or the file, with a registry)

Where this leaves us

The tests of Parts 4 to 8 measure how surprising a result would be under chance. They cannot measure how surprising it would be under an analyst who chose the path, and that second surprise can be all of the first. The fix is not more statistics. It is a date: the moment after which nothing changed, and a witness to it, whether that is your own changelog, a registry's timestamp, or a reviewer's in-principle acceptance. The stronger the claim of absence, the stronger the witness needs to be.

That closes the core of the series, and Part 1's sentence can now be rewritten the way a careful study would write it. Part 1 began with: the tool scored 12 points higher, so it helps. Here is the same claim as the verification study would make it, and every word has now been explained. Agents with a shell scored 12.3 points higher on the functional rubric than agents with no verification tools, 95 per cent bootstrap interval 7.0 to 17.5, stratified by model and task, Holm-corrected p = 0.0006 across six pre-specified hypotheses, condition-blind grading against a frozen rubric, every run counted, seed 20260703. Supported. And in the same table, three hypotheses that were not, reported the same way, which is how you know the first one means something.


Next: Resampling Beyond the Interval: the bootstrap as a test and how it differs from the permutation test, the paired bootstrap for conditions that ran on the same tasks, the bootstrap for anything you can compute, the jackknife, and where resampling lets you down. The first of five parts that widen the net.