Every term the series leans on, in alphabetical order, each with a link to the part that introduces it. If a word in any part is unfamiliar, it should be here.

Accuracy
The fraction of examples a classifier gets right. Reported, but not trained on, because it changes in jumps and has no slope to follow. Part 13
Activation function
A fixed function (with no weights of its own) applied to a neuron's linear output to bend it. Without one, stacking neurons collapses back into a single straight line. Part 11
Backpropagation
The bookkeeping that works out each weight's share of the blame in a network with layers, passing the error backwards from the output. The slopes it produces feed ordinary gradient descent. Part 12
Batch, batch size
The handful of rows a single gradient descent step looks at. The batch size is how many. Smaller batches mean cheaper, noisier steps. Part 7
Bias
The number added after the weights have done their multiplying. For a straight line it is the intercept, where the line crosses the vertical axis. Part 2
Classification
A problem whose answer is one option from a fixed list (spam or not, which digit, which word next), as opposed to regression, whose answer is a number. Part 13
Closed-form solution
A formula that returns the best weights in one calculation, such as the Normal Equation for a straight line. Exact, but it stops being available once models get large or bendy. Part 3
Coefficient
An older word for a weight, common in statistics and in polynomial fitting. Part 8
Convergence
The point in training where the loss has stopped improving and the weights have stopped moving. The usual reason to stop. Part 5
Cross-entropy
The loss used for probabilities: minus the log of the probability the model gave to the true answer. Small when the model was confident and right, huge when it was confident and wrong. Part 13
Dataset
A collection of examples where both the feature and the target are known. The part used to fit the model is the training data. Part 1
Decision boundary
The line (or surface) where a classifier switches from one answer to another, where its probability is exactly 0.5. Part 13
Derivative
The slope of a curve at one point: how fast it is rising or falling right there. A partial derivative is the same thing for one knob at a time, with the others held still. Part 5
Divergence
Training blowing up: a learning rate so large that each step overshoots by more than the last, the loss climbs, and the weights fly off towards infinity. Part 6
Early stopping
Halting training when the loss on held-back data starts to rise even though the training loss is still falling, the sign that the model has begun memorising. Part 8
Epoch
One complete pass through the training data, however many iterations that takes. Part 7
Feature
An input: something you already know and use to make a prediction, like the parents' height or the hours studied. A model can have one or many. Part 1
Generalisation
How well a model performs on data it never trained on. The only score that ultimately matters. Part 8
Gradient
The slope of the loss with respect to a knob, telling you which way to nudge it and how urgently. With many knobs, the collection of all their slopes. Part 5
Gradient descent
The training method: measure the loss, find the slope for each knob, take a small step downhill, repeat. Part 5
Hyperparameter
A setting you choose before training rather than something training learns: the learning rate, batch size, number of epochs, how flexible the model is. Part 6
Iteration
One pass round the training loop: one batch through the model and one update of the weights. Also called a step. Part 5
Layer, hidden layer
A set of neurons that all act at the same stage of a network. A hidden layer is one that sits between the input and the output. Part 12
Learning rate
The small number the gradient is multiplied by to set the size of each step. Too small and training crawls; too large and it bounces or diverges. Part 6
Linear
A relationship where the output changes by a steady amount for each step in the input, so it draws a straight line. A model is "linear in the weights" when each weight is only ever multiplied by a number and added. Part 1, Part 3
Local minimum, global minimum
A local minimum is the bottom of one valley on the loss landscape; the global minimum is the lowest point of all. Gradient descent finds the nearest valley, which is usually fine. Part 9
Logistic regression
The simplest classifier: a linear model whose output is squashed by a sigmoid into a probability. One neuron with a sigmoid activation. Part 13
Loss, loss function
A single number scoring how badly one particular set of weights fits the data, and the formula that produces it. Training means making it smaller. Part 4
Loss curve, loss landscape
The loss plotted against one weight (a curve) or against weight and bias together (a landscape or surface). A map of how wrong every possible setting would be. Not to be confused with the training curve. Part 4, Part 9
MAE, MSE
Mean Absolute Error and Mean Squared Error, the two standard loss functions for regression. MAE averages the size of the errors; MSE averages their squares, which punishes big misses more and gives a smoother curve. Part 4
Model
The recipe that turns an input into a prediction: a formula with some numbers in it. Galton's formula is one; a language model is another. Part 1
Multivariate regression
Regression with several features, one weight per feature, sharing a single bias. Part 10
Neuron
A weight (or one per input), a bias, and an activation function. The basic building block of every neural network. Part 11
Normal Equation
The closed-form formula for the best straight line through a set of points. Also called ordinary least squares. Part 3
Overfitting, underfitting
Overfitting is memorising the training data, noise and all, and doing badly on new data. Underfitting is a model too simple to capture the pattern, which does badly everywhere. Part 8
Parameter
Any number inside the model that training sets: every weight and every bias. A modern language model has hundreds of billions. Part 1
Polynomial, degree
An equation built from powers of x, each with its own weight. Its degree is the highest power. More degree means more bend, and more risk of overfitting. Part 3
Regression
Predicting a number from inputs, and the machine learning word for prediction in general. Named by Galton after "regression to the mean". Part 1
Regularisation
A penalty added to the loss for large weights, a complexity tax that discourages the wild wiggles of an overfitted model. L2 (ridge) taxes squared weights, L1 (lasso) their sizes. Part 8
ReLU
Rectified Linear Unit, the most common activation function: zero for negative inputs, unchanged for positive ones. A hockey stick. Part 11
SGD, stochastic gradient descent
Gradient descent fed one batch at a time instead of the whole dataset. "Stochastic" means the randomness of which rows land in each batch. What almost all real training uses. Part 7
Sigmoid
The S-shaped function that squashes any number into the band between 0 and 1. An activation function, and the way a two-option classifier produces a probability. Part 11, Part 13
Softmax
Turns a list of scores into probabilities that add up to 1, by raising e to each score and dividing by the total. How a model chooses between many options. Part 13
Target
The thing the model is trying to predict, like the child's height or the exam score. Also called the label or the output. Part 1
Test set, validation set
Data held back from training. The validation set is checked during training to catch overfitting and to choose hyperparameters; the test set is used once, at the end, for the final score. Part 8
Token, tokenisation
Chopping text into pieces (roughly words or chunks of words) and giving each piece an ID number, so a model that only eats numbers can read text. Part 14
Training
Working out the parameters of a model from training data. In this series, gradient descent. Part 1, Part 5
Training curve
The loss plotted against time (iterations or epochs) for one training run. It falls and flattens; it has no U shape. Part 7
Weight
The number an input gets multiplied by. For a straight line it is the slope; in a network there is one on every connection. Part 2

← Back to the series