Code · Systems · Ideas

Achint Mehta

Home / Research / Research Basics

Research Basics

"The tool scored 12 points higher, so it helps." This series is about everything a careful reader would ask before believing that sentence, and everything a careful author would do before writing it. It starts with why one number is not a finding, then works through the averages and spreads, the hypothesis and its null, the permutation test done by shuffling cards, what a p-value does and does not mean, effect sizes and confidence intervals built by resampling, the correction for testing six things at once, how to defend a result of "no effect", the design choices that make a study believable, and the difference between deciding your analysis before the data and after. Five further parts widen the net: the bootstrap as a test and the paired bootstrap, the classical toolbox from Welch's t to ANOVA, correlation, regression and mixed models, how to read a leaderboard and pass@k, and the Bayesian alternative. Every term is illustrated with a worked example from two real studies of AI coding agents, one done the ordinary way and one done properly, nearly every part has a demo you can run in your browser, and every part ends with a short Python script that reproduces its numbers.

The two studies are Reasoning effort, not tool access, buys first-try reliability and The reach of a verification tool decides its value. Lost in the vocabulary? There is a glossary of every term the series uses.