Every model in this series has answered a "how much?" question. How tall will the child be, what score will the student get, what is the value of sin(x) here. The answer was always a number on a scale, and that kind of problem is what regression means.

Most of the models you actually meet answer a different kind of question: "which one?" Is this email spam or not? Which of ten digits is in this picture? Which of fifty thousand words comes next? These are classification problems: the answer is one option from a fixed list. The final part of this series is about the biggest classifier of all, and this part builds the two small additions that get us there. Neither changes the training loop at all.

A first attempt, and why it fails

Take the study-hours data one more time, but instead of the exam score, record only whether each student passed:

Hours studied (feature) Passed? (target)
1 no
2 no
3 yes
4 no
5 yes
6 yes

The obvious move is to write "no" as 0 and "yes" as 1 and fit a straight line as usual. It half works: the line slopes upward, so more hours means a higher number. But the number itself is meaningless. Feed in 10 hours and the line predicts 1.8; feed in 0 and it predicts something negative. There is no such thing as 180% passed. And there is no principled place to draw the line between "predict yes" and "predict no".

What we would like the model to output is not a number on an open-ended scale but a probability: a number between 0 and 1 that reads as "how sure am I that this student passes?" Zero means certain they fail, one means certain they pass, and 0.5 means no idea.

Squash it: sigmoid as a probability

You have already met the tool for this. The sigmoid function from Part 11 takes any number, however large or small, and squashes it into the band between 0 and 1, crossing 0.5 at zero. So keep the linear model exactly as it was, and pass its output through sigmoid:

z = w × hours + b            (a score, any size, like before)
p = sigmoid(z)               (a probability between 0 and 1)

That is the whole model, and it has a name: logistic regression. Despite the word "regression" in its name it is the simplest classifier there is, and it is also, in the notation of Part 11, exactly one neuron with a sigmoid activation.

Run some numbers through it. With w = 1.5 and b = −5.25, so z = 1.5 × (hours − 3.5):

Hours z p = sigmoid(z) Reads as
1 −3.75 0.02 almost certainly fails
2 −2.25 0.10 probably fails
3 −0.75 0.32 leaning fail
4 0.75 0.68 leaning pass
5 2.25 0.90 probably passes
6 3.75 0.98 almost certainly passes

To turn a probability into an actual answer, pick the more likely option: predict "pass" when p is above 0.5, "fail" when below. Since sigmoid crosses 0.5 exactly when z = 0, that rule has a clean geometric meaning: the model says pass whenever w × hours + b > 0, in other words whenever hours is past −b/w = 3.5. That threshold is called the decision boundary. (It is the same −b/w switch-on point that every ReLU neuron had in Part 12: the place where the raw score crosses zero.) With one feature the boundary is a point on a line. With two features it is a line across a plane, and the demo below lets you move it.

Scoring a probability: cross-entropy

The training loop needs a loss, and MSE is the wrong tool here. It can be made to work, but squared distance between a probability and a 0-or-1 label does not care about the right things: it punishes a confident wrong answer barely more than a hesitant one, and it gives gradient descent a flat, unhelpful slope exactly when the model is confidently wrong and most needs correcting.

The loss that fits probabilities asks a different question: how surprised was the model by what actually happened? If the model gave the true outcome a probability of 0.98, it was barely surprised. If it gave the true outcome 0.02, it was astonished, and should pay for it. The formula that turns "probability of the truth" into "surprise" is:

loss for one example = −log( probability the model gave to the true answer )

(log here is the natural logarithm; all you need to know is that −log(1) = 0, so a correct and fully confident prediction costs nothing, and −log(p) climbs without limit as p shrinks towards 0, so confident mistakes are punished very hard.) Averaged over all the examples, this is called cross-entropy loss, or log loss. Here it is on the six students, using the model from the table above:

Hours Truth Probability the model gave the truth Surprise, −log(p)
1 fail 0.98 (it said 0.02 for pass, so 0.98 for fail) 0.02
2 fail 0.90 0.10
3 pass 0.32 1.14
4 fail 0.32 1.14
5 pass 0.90 0.10
6 pass 0.98 0.02
average (cross-entropy) 0.42

Two students, the 3-hour pass and the 4-hour fail, supply almost all the loss, because the model leaned the wrong way on both. Everything else about training is untouched: this number is a loss for one particular w and b, a different w and b gives a different number, and gradient descent nudges them to make it smaller. The slope calculation changes because the formula changed, but the loop does not.

For the curious: the slope comes out unexpectedly tidy. Take the derivative of cross-entropy with respect to w, through the sigmoid, and the messy parts cancel, leaving (p − y) × x averaged over the examples: the gap between the predicted probability and the true 0 or 1, times the input. That is the same shape as the MSE gradient from Part 5, error times leverage, which is one reason this pairing of sigmoid and cross-entropy is so standard.

One more distinction that trips people up. Accuracy, the fraction of examples the model gets right, is what you report; cross-entropy is what you train on. Accuracy only changes when a prediction flips across 0.5, so it is a staircase, and a staircase has no slope for gradient descent to follow. Cross-entropy changes smoothly every time a probability moves, so it always says which way to go, and lowering it tends to raise accuracy as a side effect.

Try it: draw the boundary

Thirty-six students, each plotted by hours studied and hours slept the night before. Green passed, red failed. The model is one sigmoid neuron with two inputs, p = sigmoid(w1 × (study − 5) + w2 × (sleep − 6) + b), where study is measured from 5 hours and sleep from 6 so that the bias is simply the score of an average student. The shading shows the probability of passing the model assigns to every point on the plane, and the white line is the decision boundary, where that probability is exactly 0.5. Move the three sliders and watch the accuracy and the cross-entropy, then press Train it to let gradient descent take over from wherever the sliders are. Stop halts it, and the sliders stay live throughout.

accuracy - · cross-entropy -

A few things to notice. The boundary is a straight line because the model underneath is still linear: w1 and w2 set its tilt, b slides it. The shading fades from red to green across the line, and how quickly it fades is set by the size of the weights: double both weights and the boundary stays put but the model becomes more confident on either side of it. Training pushes the accuracy to 35 out of 36 within a few dozen steps, then keeps lowering the cross-entropy without changing the accuracy at all, which is the staircase-versus-slope point from above. The one stubborn red dot on the green side is a student who studied for six hours, slept a reasonable amount, and failed anyway. No straight line can be right about everyone, and the model has learned to give that student a 77% chance of passing rather than to bend around a single exception.

Many options: softmax

Yes-or-no covers a lot, but the questions that matter most have many answers. Ten digits. A thousand kinds of object in a photo. Fifty thousand possible next words. The pattern scales up in the most direct way possible: one score per option. Instead of one neuron producing one z, the model's final layer has one neuron per option, each producing its own score. For a next-word model with a fifty-thousand-word vocabulary, that is fifty thousand output neurons, fifty thousand scores.

Then those scores have to become probabilities that make sense together: each between 0 and 1, and all of them adding up to 1, because exactly one option is the answer. Sigmoid on each score separately would not do that; five sigmoids can happily all say 0.9. The function that does is softmax, and it is two steps:

  1. Raise e to the power of every score. This makes every number positive and stretches the gaps: a score 2 higher becomes about 7 times bigger.
  2. Divide each result by the total, so they sum to 1.

Here it is for five candidates to follow "It was a dark and":

Candidate Score e^score ÷ total Probability
stormy 4.0 54.6 54.6 / 76.2 0.72
quiet 2.5 12.2 12.2 / 76.2 0.16
cold 2.0 7.4 7.4 / 76.2 0.10
sunny 0.5 1.6 1.6 / 76.2 0.02
banana −1.0 0.4 0.4 / 76.2 0.005
total 76.2 1.00

The highest score gets the highest probability, every option gets some probability however small, and only the differences between scores matter: add 10 to every score and the probabilities do not move, because the extra factor cancels in the division. Softmax with two options collapses to exactly the sigmoid from earlier, so nothing new was really added, just the same idea extended to a list.

Cross-entropy carries over untouched: the loss for one example is still −log of the probability the model gave to the true answer. If the true next word was "stormy", the model above pays −log(0.72) = 0.33. If it was "cold", it pays −log(0.10) = 2.33. Training a next-word predictor means running this over billions of sentences and nudging billions of weights so that the true word's probability rises, one small step at a time, with the same gradient descent loop as always.

Try it: scores into probabilities

The five sliders are the five raw scores from the table. Move them and watch the probabilities respond; try adding the same amount to all five and see nothing change. The temperature slider divides every score by a number before softmax runs: below 1 the biggest score takes almost everything, above 1 the probabilities flatten out. It is a knob you will meet again when a language model picks its next word.

What changed, and what did not

Two small additions turned a regression model into a classifier: a squashing function on the output (sigmoid for two options, softmax for many) so that the model speaks in probabilities, and a loss that scores probabilities by how surprised the model was (cross-entropy). Everything upstream is unchanged. Weights, biases, neurons, layers, activation functions, gradient descent, learning rate, batches, held-back data: all of it works exactly as before, because none of it ever cared what the output meant, only how far it was from where it should be.

That is the last piece. A model that turns an input into a list of scores, softmaxes them into a probability for every option, and is trained by gradient descent to make the true option's probability as high as possible, is a complete description of what a language model is trained to do. The final part looks at what that means at scale.


Next: From Straight Lines to ChatGPT: the series opened with "LLMs are just predicting the next token." Time to close the loop and see how everything you have learned scales up to a language model.