Every term the series leans on, in alphabetical order, each with a link to the part that introduces it. If a word in a methods section is unfamiliar, it should be here.
- Adjusting for, controlling for
- Adding a known factor such as model or task to the regression so that its variation is removed from the noise before the effect of interest is judged. The regression form of stratifying. Part 13
- Alternative hypothesis
- The claim itself, once the null has been rejected: the tool changes the score. One-sided if it names a direction, two-sided if it does not. Part 3
- ANOVA, analysis of variance
- One test for whether any of three or more group means differ, before asking which. Followed by corrected pairwise comparisons. Kruskal-Wallis is its rank version. Part 12
- Bayes factor
- How many times more probable the data are under one hypothesis than another, each averaged over its prior. Can favour the null, which a p-value cannot. Sensitive to the priors. Part 15
- Benjamini-Hochberg, false discovery rate
- A correction for very many tests that limits the expected fraction of reported findings that are false, rather than the chance of any false finding at all. The tool for a thousand comparisons, not for six. Part 7
- Blinding
- Grading runs without knowing which condition they came from, enforced by stripping the labels and shuffling the order rather than by willpower. Removes the grader's expectations from the scores. Part 9
- Bonferroni correction
- Divide the 0.05 line by the number of tests, or multiply each p-value by it. Holds the family-wise error rate at 5 per cent, at some cost in power. Part 7
- Bootstrap
- Building a confidence interval by resampling the runs you have, with replacement, thousands of times, and reading the interval off the spread of the results. Stratified when the resampling stays inside each cell of the design. Part 6
- Bootstrap test
- A test that shifts each group onto the common mean, so the null is true, then resamples to see how often the real gap appears. Tests equal means while letting the spreads differ. Part 11
- Box plot
- A picture of a middle and a spread: the box is the middle half of the runs, the line is the median, the whiskers reach the rest, and lone dots are outliers. Part 2
- Cell
- One combination of everything the design controls: this model, this task, this condition. Replicates are runs within a cell. Part 1
- Chi-square test
- The large-sample approximation to Fisher's exact test for a table of counts. Needs an expected count of five or more in every cell. Part 12
- Cliff's delta
- An effect size with no units: how often a random run from one condition beats a random run from the other, minus the reverse. Part 6
- Clopper-Pearson interval
- The exact confidence interval for a rate, found by asking which true rates would make the observed count unsurprising. Never wrong, slightly wide. Part 2
- Cohen's d
- An effect size measured in units of the run-to-run spread: the gap divided by the standard deviation. Roughly, 0.2 is small, 0.5 medium, 0.8 large. Part 6
- Cohen's kappa, weighted kappa, intraclass correlation
- Measures of how closely two graders, or one grader on two occasions, agree, after discounting the agreement chance alone would give. Above 0.8 is excellent. Part 9
- Condition
- A setting the study deliberately varies and compares: with the tool and without, high reasoning effort and low. Part 1
- Confidence interval
- The range of effects the data cannot rule out. A 95 per cent interval comes from a procedure that captures the true effect in 95 of every 100 studies. Its width is what the study learned. Part 6
- Confirmatory analysis
- The phase in which the plan is frozen, nothing changes, and pilot data are excluded. The tests here are the ones that count. Part 10
- Confound
- Anything that varied along with the condition and could explain the gap instead: a model update, a slow week, a prompt rewording, which batch the runs were in. Randomisation is the cure. Part 9
- Correlation, Pearson's r, Spearman's rho
- How consistently two variables move together, from −1 to +1. Pearson measures a straight line, Spearman uses ranks. A third variable moving both is the usual reason they correlate. Part 13
- Credible interval
- The range holding 95 per cent of the posterior. It means what people think a confidence interval means: given the data and the prior, the effect is probably in there. Part 15
- Effect size
- How big the difference is, in the outcome's own units where it has them: points, percentage points, a multiplier. Not the p-value. Part 6
- Equivalence test, TOST
- The test that can support "no meaningful difference": two one-sided tests against a bound set in advance, or equivalently checking that the 90 per cent interval fits inside the band. Part 8
- Exchangeable
- The permutation test's one assumption: under the null, any run could have carried either label. True when conditions were assigned independently of how runs would score. Part 4
- Exploratory analysis
- Questions asked after seeing the data, or during a pilot. Reported uncorrected, labelled as such, and never presented as a claim that cleared the bar. Part 7, Part 10
- Family, family-wise error rate
- A set of hypotheses tested together, and the chance that at least one of them produces a false positive. Six tests at 0.05 give a family-wise rate of 26 per cent. Part 7
- Fisher's exact test
- The permutation test for counts in a two-by-two table, such as 16 of 18 against 5 of 18. Called exact because every possible deal is enumerated. Part 5
- Freeze
- The moment after which the hypotheses, design, sample size, outcomes, exclusion rules and analysis code stop changing, and before which no confirmatory run is made. Part 10
- Garden of forking paths
- The sequence of reasonable analytic choices that, made after seeing the data, lead an honest analyst to a significant result in noise. Fixed by deciding the path first. Part 10
- Geometric mean
- The average for things that multiply: average the logarithms and undo the log. The right middle for token costs, speedups and ratios. Part 2
- Holm correction
- A step-down version of Bonferroni: sort the p-values, compare the smallest to 0.05 divided by the number of tests, the next to 0.05 divided by one fewer, and stop at the first failure. Same protection, less waste. Part 7
- Hypothesis
- A bet with four parts: what is compared, on which outcome, in which direction, over what scope. Written before the data, or it is not a bet. Part 3
- Intention to treat
- Every run counts in the condition it was assigned to, including runs that ignored their tool and apps that never started. Excluding them rewards the condition that fails most. Part 9
- Interaction
- A term in a regression for whether one effect depends on another factor: the tool's gap on one task minus its gap on another. Needs several times the data of a plain effect. P5 was one. Part 13
- Interquartile range
- The stretch between the value a quarter of the way up the sorted runs and the value three quarters of the way up: the middle half. The spread that ignores outliers. Part 2
- Jackknife
- Leave each observation out in turn and recompute. Shows which single run is doing the work. The bootstrap's older relative. Part 11
- Linear regression, slope, r²
- The line closest to the points. The slope is the effect per unit of input, and r² is the share of the outcome's variation the line accounts for. With a yes-or-no input it is the t-test. Part 13
- Mann-Whitney test
- A permutation test on ranks rather than raw scores. Its natural effect size is Cliff's delta. Part 4
- McNemar's test
- The test for two models, or two conditions, on the same tasks with a pass or fail each: only the tasks they disagree on count, and the question is whether the split of those is a coin toss. Part 14
- Mean
- Add the values, divide by how many. The total shared out equally, and pulled hard by any extreme value. Part 2
- Median
- The middle value once sorted. Half the runs are below it. Robust to outliers, and the default for costs, latencies and anything with a long tail. Part 2
- Mixed-effects model
- A regression in which some factors, such as which model or task, are treated as random draws from a population and given their own baselines, while the effect of interest is fixed. The verification study's cross-check. Part 13
- Multiple comparisons
- The problem that asking several questions of noise gives noise several chances to answer one. Handled by a correction across the family. Part 7
- Newcombe's method
- A confidence interval for the difference between two rates, built from the Wilson interval of each. The honest interval for "16 of 18 against 5 of 18". Part 2
- Null distribution
- The hill of gaps that shuffling produces: what the test statistic looks like in a world where the null is true. The p-value is the share of it beyond the observed value. Part 4
- Null hypothesis
- The shadow of the claim: the comparison makes no difference, so the labels could be swapped. It is what gets tested, because only it says exactly what the data should look like. Part 3
- Observational study
- A study whose conditions were assigned or observed rather than randomised, so that other things may have travelled with them. The first agent study was one. Part 9
- Outcome measure
- The number written down at the end of each run: a rubric score, a pass or fail, a token count. Part 1
- Paired bootstrap, paired t-test
- For two conditions run on the same tasks or seeds: analyse the differences, one per pair, so that the pair's shared difficulty cancels. Turned 0.028 into 0.00001 on the same numbers. Part 11, Part 12
- pass@k
- The probability that at least one of k attempts at a task passes. Estimated from n samples with c correct as one minus C(n − c, k) / C(n, k), not by plugging a rate into a formula. Part 14
- Permutation test
- Shuffle the condition labels, recompute the gap, repeat thousands of times, and count how often shuffling alone beats the real gap. Assumes only exchangeability. Stratified when the shuffle stays inside each cell. Part 4
- Post hoc
- Decided after the results were in view. Not a sin, but a label a paper owes its reader. Part 10
- Power, power analysis
- The chance a study will detect an effect of a given size if it exists, usually aimed at 80 per cent. Power analysis runs the calculation backwards to say how many runs to buy. Part 8
- Preregistration
- Depositing the frozen plan with a registry such as OSF or AsPredicted before the runs, so that a third party's timestamp proves when the choices were made. Not peer review, not a publication route. Part 10
- Pre-specified
- Written down, with all four ingredients and the test to be used, before the analysis. Weaker than preregistered only in that nobody else witnessed the date. Part 3
- Primary and secondary contrasts
- The headline bets, corrected as a family, and the follow-up questions asked to understand them, reported uncorrected and labelled exploratory. Part 3
- Prior, likelihood, posterior
- What you believed before the data, what the data say at each possible value, and what you believe after: Bayes' theorem multiplies the first two to get the third. Part 15
- p-value
- The probability, if the null hypothesis were true, of a result at least as extreme as the one observed. Not the probability that the null is true, and not a measure of how big the effect is. Part 5
- Region of practical equivalence (ROPE)
- Part 8's band of no meaningful effect, read as the share of the posterior that falls inside it. A probability instead of a verdict. Part 15
- Registered Report
- The plan is peer-reviewed before the data are collected. If accepted in principle, the paper is published whatever the results, provided the protocol was followed. Part 10
- Replicate, replication
- A repeat of the same run with nothing changed, to measure the noise. Not to be confused with replicability, which is someone else reaching your conclusion with new data. Part 1, Part 9
- Reproducible
- Someone with your data and your code gets your numbers. Guaranteed by releasing both, with the random seed. Part 9
- Results-blind review
- The manuscript is reviewed with the results redacted, after the data exist. Removes the reviewers' outcome bias, not the author's. Part 10
- Rubric
- The fixed list of criteria and their weights against which every run is scored. Frozen before any run exists, or it drifts. Part 1, Part 9
- Run
- One complete execution of the thing being measured, producing one outcome. The unit that sample sizes count. Part 1
- Sample size, n
- How many runs sit behind a number. The single most important thing to read beside any average. Part 1
- Selection effect
- A bias created by which runs are kept: excluding crashes flatters the condition that crashes most. Part 9
- Shrinkage
- What an informative prior does to a small-study estimate: pulls it toward what was expected, in proportion to how noisy the estimate is. Usually wise, always to be declared. Part 15
- Smallest effect size of interest
- The bound, chosen before the data, inside which a difference does not matter. What "no meaningful effect" means, with a number in it. Part 8
- Standard deviation
- The typical distance of a run from the mean. About 15 points for the twenty-four runs. Pulled by outliers, like the mean. Part 2
- Standard error
- How much an average moves from sample to sample: the standard deviation divided by the square root of the number of runs. The formula-based confidence interval is the estimate plus or minus about two of these. Part 6
- Statistically significant
- Would be surprising under the null, by the 0.05 convention. Not the same as large, useful or important. Part 5
- Stratified
- Done within each cell of the design: a shuffle that never moves a score across models, a resample that draws from each cell separately. Removes the noise that the design already accounted for. Part 4, Part 6
- Test statistic
- The single number computed from the data that the test is about: the difference in means, the difference in medians, a count. Chosen before looking. Part 4
- Threats to validity
- The four ways a conclusion can be wrong: internal (a confound caused the gap), external (it does not generalise), construct (the measure is not the thing), statistical (the tests or sample were inadequate). Part 9
- Trimmed mean, weighted mean
- The mean with the extremes dropped, and the mean in which some items count more than others. A rubric total is a weighted mean. Part 2
- t-test
- The classical test for a difference in means, which computes the p-value from a formula that assumes bell-shaped noise. Agrees with the permutation test when the assumption holds. Part 4
- Tukey's HSD
- The classical correction for comparing every pair of groups after an ANOVA. Holm across the pairs does the same job. Part 12
- Two-sided, one-sided
- Whether a test counts surprise in both directions or only the expected one. Two-sided is the cautious default and the one the verification study used throughout. Part 3
- Variability, noise
- How much identical runs differ from each other. The reason one number is not a finding. Part 1
- Wald interval
- The textbook confidence interval for a rate, which fails at small counts and near 0 or 100 per cent, where it can claim a rate of 103 per cent. Worth recognising only to distrust. Part 2
- Welch's t-test
- The t-test that does not assume the two groups have the same spread. The default for two independent groups. Student's version assumes equal spreads and is what most software gives you unless asked. Part 12
- Wilcoxon signed-rank test
- The rank version of the paired t-test. With very few pairs it has a floor: seven pairs cannot give p below 0.0156. Part 12
- Wilson score interval
- The well-behaved confidence interval for a single rate, pulled toward 50 per cent and never past 0 or 100. The sensible default. Part 2