Home / Research / Research Basics
Research Basics
"The tool scored 12 points higher, so it helps." This series is about everything a careful reader would ask before believing that sentence, and everything a careful author would do before writing it. It starts with why one number is not a finding, then works through the averages and spreads, the hypothesis and its null, the permutation test done by shuffling cards, what a p-value does and does not mean, effect sizes and confidence intervals built by resampling, the correction for testing six things at once, how to defend a result of "no effect", the design choices that make a study believable, and the difference between deciding your analysis before the data and after. Five further parts widen the net: the bootstrap as a test and the paired bootstrap, the classical toolbox from Welch's t to ANOVA, correlation, regression and mixed models, how to read a leaderboard and pass@k, and the Bayesian alternative. Every term is illustrated with a worked example from two real studies of AI coding agents, one done the ordinary way and one done properly, nearly every part has a demo you can run in your browser, and every part ends with a short Python script that reproduces its numbers.
The two studies are Reasoning effort, not tool access, buys first-try reliability and The reach of a verification tool decides its value. Lost in the vocabulary? There is a glossary of every term the series uses.
-
Part 1
One Number Is Not a Finding
Why "the model scored 74" tells you almost nothing on its own, what runs, conditions and cells are, how much the same experiment moves when you simply run it again, and the three questions every result has to answer before it counts as a finding. With the twenty-four runs that the rest of the series will keep coming back to.
-
Part 2
Middles, Spreads and Rates
The mean and the median, and the runaway token cost that pulls them apart. Trimmed, weighted and geometric means and when each one is honest. Standard deviation, the interquartile range and how to read a box plot. Then rates, which are averages in disguise, and the four ways of putting an interval on "16 out of 18", one of which claims a pass rate of 103 per cent.
-
Part 3
The Claim and Its Shadow
What a hypothesis is and the four things it has to contain, why every hypothesis comes with a null hypothesis attached and why it is the null that gets tested, one-sided and two-sided claims, primary and secondary contrasts, and the six hypotheses from the verification study written out as bets, three of which lost.
-
Part 4
Shuffle the Labels
The permutation test, done by hand on the twenty-four runs. Take the null hypothesis at its word, swap the condition labels, and count how often shuffling alone produces a gap as big as the real one. Why the answer is one in eleven when you shuffle everything and one in seven hundred when you shuffle within each model, and what "stratified by model and task" in a paper means.
-
Part 5
What 0.0006 Means, and What 0.08 Does Not
The p-value defined once, correctly, and the four ways it is usually misread. Where the 0.05 line came from and what it is for. Why a tiny p-value can sit on a trivial effect and a large effect can fail to reach it. Fisher's exact test for counts, on the 16-of-18 result. And the six p-values from the verification study, read one at a time.
-
Part 6
How Big, and How Sure
Effect sizes in units a reader can picture, and the two standardised ones you will meet. The confidence interval, the one thing it promises and the thing it does not. The bootstrap, which builds an interval by resampling the runs you already have, done one resample at a time on the twenty-four runs. And why the verification study's intervals and p-values were made by the same shuffling.
-
Part 7
Six Hypotheses, One Bar
Why testing six hypotheses at the 0.05 line gives chance a one-in-four shot at fooling you. The Bonferroni fix, and Holm's step-down version worked line by line on the verification study's own six p-values, reproducing its table exactly. What a "family" of tests is, why secondary contrasts were reported uncorrected and labelled exploratory, and what false discovery rate means when there are hundreds.
-
Part 8
Defending a Null
Why the ordinary test can never say "no effect", the difference between absence of evidence and evidence of absence, and the equivalence test that can make the stronger claim. Choosing the smallest effect worth caring about before the data arrive, what statistical power is, how many runs it takes to bound a null within five points, and why the verification study's P5 could not.
-
Part 9
Designed to Be Believed
The arithmetic of the last five parts assumes the runs were assigned fairly and graded without favour. This part is about making that true. Assigned versus randomised conditions, confounds and the one the first study admitted, blinding the grader, freezing the rubric, the intention-to-treat rule, seeds and reproducibility, how to read a kappa of 0.973, and the four kinds of threat to validity.
-
Part 10
When You Decide Matters
The same data and the same tests give different evidence depending on when the analytic choices were made. The garden of forking paths, and a demo in which you walk it. Then the five ways of fixing the choices in advance, from a private freeze to a registered report, what each costs, what each buys, and which one a study built to defend a null should use.
-
Part 11
Resampling Beyond the Interval
The bootstrap can test as well as estimate, and the difference between a bootstrap test and a permutation test is a difference in what "nothing is going on" means. Then the paired bootstrap, which is how two conditions run on the same tasks should be compared, the bootstrap for anything you can compute (a ratio of medians, a correlation), the jackknife, and the cases where resampling lets you down.
-
Part 12
The Classical Toolbox
The tests that fill most results sections, each explained as a formula for a resampling test you have already seen. Student's t and Welch's t, and why Welch is the default. The paired t-test and why pairing changed a p-value from 0.03 to 0.00001. Mann-Whitney and Wilcoxon on ranks. Chi-square against Fisher's exact test. ANOVA and Kruskal-Wallis for three or more groups, and what to do after them. With a table that says which to reach for.
-
Part 13
Lines Through the Data
Correlation, what it measures and what it does not, and the trap in the twenty-four runs where tokens and scores rise together for a reason that has nothing to do with tokens. Fitting a line and reading its slope. Regression as the general form of every comparison so far, with the model added as a term. The mixed-effects model that gave each model its own baseline and cross-checked the verification study's permutation tests. And why P5 was an interaction, which is why its interval was so wide.
-
Part 14
Reading a Leaderboard
Everything in the series applied to the table most readers actually look at. A benchmark score is a rate and has an interval, which is wider than the gap between neighbours on most leaderboards. What pass@k means and the estimator that gets it right. Why temperature zero is not deterministic and what seeds are for. Comparing two models on the same tasks, where only the tasks they disagree on count. And what twenty benchmarks do to the word "wins".
-
Part 15
The Bayesian Alternative
The question nothing in the series has answered, how likely the hypothesis is given the data, and the framework that answers it. Priors, likelihoods and posteriors in words and in one line of arithmetic. Why a credible interval for a rate lands where Wilson did, how a sceptical prior shrinks a gap of 12 to 8, the Bayes factor and its scale, and the region of practical equivalence, which is Part 8's bound with a probability attached. And why the freeze of Part 10 applies here too.
-
Reference
Glossary
Every term the series uses, in one or two plain sentences each, with a link to the part that explains it.