Every term the series leans on, in alphabetical order, each with a link to the part that introduces it. If a word in any part is unfamiliar, it should be here.
- Accuracy
- The fraction of examples a classifier gets right. Reported, but not trained on, because it changes in jumps and has no slope to follow. Part 13
- Activation function
- A fixed function (with no weights of its own) applied to a neuron's linear output to bend it. Without one, stacking neurons collapses back into a single straight line. Part 11
- Backpropagation
- The bookkeeping that works out each weight's share of the blame in a network with layers, passing the error backwards from the output. The slopes it produces feed ordinary gradient descent. Part 12
- Batch, batch size
- The handful of rows a single gradient descent step looks at. The batch size is how many. Smaller batches mean cheaper, noisier steps. Part 7
- Bias
- The number added after the weights have done their multiplying. For a straight line it is the intercept, where the line crosses the vertical axis. Part 2
- Classification
- A problem whose answer is one option from a fixed list (spam or not, which digit, which word next), as opposed to regression, whose answer is a number. Part 13
- Closed-form solution
- A formula that returns the best weights in one calculation, such as the Normal Equation for a straight line. Exact, but it stops being available once models get large or bendy. Part 3
- Coefficient
- An older word for a weight, common in statistics and in polynomial fitting. Part 8
- Convergence
- The point in training where the loss has stopped improving and the weights have stopped moving. The usual reason to stop. Part 5
- Cross-entropy
- The loss used for probabilities: minus the log of the probability the model gave to the true answer. Small when the model was confident and right, huge when it was confident and wrong. Part 13
- Dataset
- A collection of examples where both the feature and the target are known. The part used to fit the model is the training data. Part 1
- Decision boundary
- The line (or surface) where a classifier switches from one answer to another, where its probability is exactly 0.5. Part 13
- Derivative
- The slope of a curve at one point: how fast it is rising or falling right there. A partial derivative is the same thing for one knob at a time, with the others held still. Part 5
- Divergence
- Training blowing up: a learning rate so large that each step overshoots by more than the last, the loss climbs, and the weights fly off towards infinity. Part 6
- Early stopping
- Halting training when the loss on held-back data starts to rise even though the training loss is still falling, the sign that the model has begun memorising. Part 8
- Epoch
- One complete pass through the training data, however many iterations that takes. Part 7
- Feature
- An input: something you already know and use to make a prediction, like the parents' height or the hours studied. A model can have one or many. Part 1
- Generalisation
- How well a model performs on data it never trained on. The only score that ultimately matters. Part 8
- Gradient
- The slope of the loss with respect to a knob, telling you which way to nudge it and how urgently. With many knobs, the collection of all their slopes. Part 5
- Gradient descent
- The training method: measure the loss, find the slope for each knob, take a small step downhill, repeat. Part 5
- Hyperparameter
- A setting you choose before training rather than something training learns: the learning rate, batch size, number of epochs, how flexible the model is. Part 6
- Iteration
- One pass round the training loop: one batch through the model and one update of the weights. Also called a step. Part 5
- Layer, hidden layer
- A set of neurons that all act at the same stage of a network. A hidden layer is one that sits between the input and the output. Part 12
- Learning rate
- The small number the gradient is multiplied by to set the size of each step. Too small and training crawls; too large and it bounces or diverges. Part 6
- Linear
- A relationship where the output changes by a steady amount for each step in the input, so it draws a straight line. A model is "linear in the weights" when each weight is only ever multiplied by a number and added. Part 1, Part 3
- Local minimum, global minimum
- A local minimum is the bottom of one valley on the loss landscape; the global minimum is the lowest point of all. Gradient descent finds the nearest valley, which is usually fine. Part 9
- Logistic regression
- The simplest classifier: a linear model whose output is squashed by a sigmoid into a probability. One neuron with a sigmoid activation. Part 13
- Loss, loss function
- A single number scoring how badly one particular set of weights fits the data, and the formula that produces it. Training means making it smaller. Part 4
- Loss curve, loss landscape
- The loss plotted against one weight (a curve) or against weight and bias together (a landscape or surface). A map of how wrong every possible setting would be. Not to be confused with the training curve. Part 4, Part 9
- MAE, MSE
- Mean Absolute Error and Mean Squared Error, the two standard loss functions for regression. MAE averages the size of the errors; MSE averages their squares, which punishes big misses more and gives a smoother curve. Part 4
- Model
- The recipe that turns an input into a prediction: a formula with some numbers in it. Galton's formula is one; a language model is another. Part 1
- Multivariate regression
- Regression with several features, one weight per feature, sharing a single bias. Part 10
- Neuron
- A weight (or one per input), a bias, and an activation function. The basic building block of every neural network. Part 11
- Normal Equation
- The closed-form formula for the best straight line through a set of points. Also called ordinary least squares. Part 3
- Overfitting, underfitting
- Overfitting is memorising the training data, noise and all, and doing badly on new data. Underfitting is a model too simple to capture the pattern, which does badly everywhere. Part 8
- Parameter
- Any number inside the model that training sets: every weight and every bias. A modern language model has hundreds of billions. Part 1
- Polynomial, degree
- An equation built from powers of x, each with its own weight. Its degree is the highest power. More degree means more bend, and more risk of overfitting. Part 3
- Regression
- Predicting a number from inputs, and the machine learning word for prediction in general. Named by Galton after "regression to the mean". Part 1
- Regularisation
- A penalty added to the loss for large weights, a complexity tax that discourages the wild wiggles of an overfitted model. L2 (ridge) taxes squared weights, L1 (lasso) their sizes. Part 8
- ReLU
- Rectified Linear Unit, the most common activation function: zero for negative inputs, unchanged for positive ones. A hockey stick. Part 11
- SGD, stochastic gradient descent
- Gradient descent fed one batch at a time instead of the whole dataset. "Stochastic" means the randomness of which rows land in each batch. What almost all real training uses. Part 7
- Sigmoid
- The S-shaped function that squashes any number into the band between 0 and 1. An activation function, and the way a two-option classifier produces a probability. Part 11, Part 13
- Softmax
- Turns a list of scores into probabilities that add up to 1, by raising e to each score and dividing by the total. How a model chooses between many options. Part 13
- Target
- The thing the model is trying to predict, like the child's height or the exam score. Also called the label or the output. Part 1
- Test set, validation set
- Data held back from training. The validation set is checked during training to catch overfitting and to choose hyperparameters; the test set is used once, at the end, for the final score. Part 8
- Token, tokenisation
- Chopping text into pieces (roughly words or chunks of words) and giving each piece an ID number, so a model that only eats numbers can read text. Part 14
- Training
- Working out the parameters of a model from training data. In this series, gradient descent. Part 1, Part 5
- Training curve
- The loss plotted against time (iterations or epochs) for one training run. It falls and flattens; it has no U shape. Part 7
- Weight
- The number an input gets multiplied by. For a straight line it is the slope; in a network there is one on every connection. Part 2