This series opened with a dismissal you have probably heard: "LLMs are just predicting the next token." Thirteen parts later, you are in a position to see what that sentence actually describes, and why "just" is doing some very heavy lifting.
A language model is a prediction machine
Everything in this series has been about one shape of problem: given an input, predict an output. Hours studied in, exam score out. Parent height in, child height out.
A language model is the same shape of problem with a different pair of columns. The feature is the text so far; the target is whichever token comes next.
| Text so far (feature) | Next token (target) |
|---|---|
| It was a dark and | stormy |
| The cat sat on the | mat |
| To be or not to | be |
Train a model on billions of rows like these, harvested from books and the web, and you get something that can continue any text you hand it. Generate one token, append it, feed the longer text back in, and repeat: that loop, prediction after prediction, is a chatbot writing you a paragraph.
But models eat numbers, not words
Every model in this series took a number in. Text needs a conversion step, and that step is called tokenisation: chop the text into pieces (tokens, roughly words or chunks of words), and give every distinct piece an ID number from a big fixed dictionary.
"It was a dark and stormy"
→ ["It", " was", " a", " dark", " and", " stormy"]
→ [1122, 572, 264, 6319, 323, 13458]
Those IDs are what the network actually receives, and what it produces is exactly the many-option output from Part 13: one score for every token in its dictionary, pushed through softmax into probabilities. " night" 62%, " evening" 9%, " sea" 3%, and so on down the list, fifty thousand entries long and adding up to 100%. Predicting the next token really means ranking the whole dictionary by likelihood, and training means cross-entropy: pay −log of the probability given to the token that actually came next, and nudge the weights to make that cheaper.
Try it: chop a sentence into tokens
Type anything below. The splitter is a toy (a real tokenizer learns its splits from data and has about fifty thousand of them), but it shows the shape of the idea: text becomes a list of pieces, every piece becomes a number, and the same piece always gets the same number.
Notice two things a real tokenizer shares with this toy. A space is part of the token that follows it (the ␣ marks), so " dark" and "dark" are different tokens. And long words get split into pieces, which is why language models can handle words they have never seen, and also why they are famously bad at counting the letters in one: the model never sees letters, only these chunks.
Try it: a next-word predictor you can steer
The toy below has been "trained" on a few dozen sentences by simply counting which word follows which. Given the current word, it looks up every word that ever followed it and how often, exactly the probability ranking described above, just with counts instead of a neural network. (A table of counts is still a model with parameters, the counts, found from data. A neural language model learns to reproduce the same kind of table, compressed into weights, and can fill in the entries no table could hold.) Click a candidate to accept it (or let the model pick), and watch a sentence assemble itself one prediction at a time. The context switch sets how many previous words the model looks at.
the
The model's predictions for the next word (click one to accept it):
Two things are worth noticing. First, even this comically simple model produces mostly grammatical text, because "which word tends to follow which" already encodes a surprising amount of structure. Second, notice its limitation: with one word of context it happily wanders ("the cat chased the ball was red"), because after "ball" it has forgotten that a "chased" came before. Switch the context to two words and the wandering drops sharply: "chased the" is followed by different things than "on the", and the model can now tell them apart. It also becomes more repetitive, because far fewer two-word pairs were ever seen in the tiny training text, so there is less to choose from. A real LLM pushes this to the limit: it weighs the entire text so far, thousands of tokens, when ranking the dictionary, and it uses a trained neural network rather than a lookup table to do it. But the game being played, rank the possible next tokens and pick one, is the same one you just played.
Everything you learned, scaled up
Here is the payoff. Every concept in this series maps directly onto what happens inside a large language model:
| In this series | In an LLM |
|---|---|
| Feature → target (hours → score) | Text so far → next token |
| Weight and bias, 2 knobs | Weights and biases, hundreds of billions of knobs |
| Neurons in feedforward layers (LLM Basics Part 1) | The same, wider and deeper: feedforward networks sit inside every layer of an LLM |
| Softmax over a few options, cross-entropy loss (Part 13) | Softmax over the whole vocabulary, cross-entropy on the true next token: "how surprised was the model?" |
| Batches and epochs (Part 7) | Batches of thousands of sequences at a time, and rarely more than one epoch, because the training text is so large |
| Gradient descent + backpropagation | Exactly the same loop, run on thousands of GPUs |
| Learning rate | Still there, still fiddly, still tuned with care |
| Overfitting and held-back data | Still there: a model can memorise its training text, and evaluation still relies on data it has not seen |
| Local minima and "good enough" | Still there: nobody finds the global minimum of a trillion-dimensional loss surface, and nobody needs to |
There is new machinery in modern LLMs that this series has not covered, the transformer architecture and its attention mechanism (the trick for weighing all the earlier context at once), and embeddings (the learned number-lists that represent each token's meaning). Those stories now have a series of their own: LLM Basics picks up exactly here and walks, in plain language, from networks that read in order all the way to the architecture inside ChatGPT. But they are refinements on top of the engine you now understand, not replacements for it. In fact, roughly two thirds of the knobs in a modern LLM live in plain feedforward networks, the simple layered architecture that opens that series, stacked between the attention layers. Strip any LLM to its skeleton and you find neurons doing w·x + b, activation functions bending the results, a loss measuring the surprise, and gradient descent turning billions of knobs a tiny step at a time.
One last secret: base models and alignment
Everything described above, the giant network trained by gradient descent to predict the next token, produces what is called a base model (you will also hear pretrained model). And here is the thing: a base model is exactly what this article says it is, a next-token predictor, and nothing more. Hand it "It was a dark and stormy" and it will continue the story. But ask it "What is the capital of France?" and it may just as happily reply with more questions, because in its training data, one question is very often followed by another. It has no notion that it is supposed to be answering you. Left to roam, base models drift, repeat themselves, and sometimes produce output that reads as eerily incoherent, not because they lack knowledge, but because "continue this text plausibly" is all they were trained to do, much like our toy predictor wandering into "the cat chased the ball was red."
The step that turns that raw text-continuer into the helpful assistant you actually talk to is called alignment (or post-training). After pretraining, the model is additionally trained on examples of good conversations, question in, helpful answer out, and refined with human feedback on which of its responses people prefer. The machinery is the same as everything in this series, gradient descent nudging weights to reduce a loss, but the goal shifts: no longer "predict the next token of the internet," now "respond the way a helpful assistant would." That is what teaches the model to answer questions rather than continue them, to follow instructions, to talk in a particular manner. The raw capability comes from pretraining; the conversational character you experience is taught afterwards.
So the next time someone says a language model is "just predicting the next token," you can agree cheerfully, and then explain what it takes to predict one well, and what it took to make the predictor talk to you like an assistant.
Congratulations: you now understand the core building blocks of machine learning well enough to be the most dangerous person in the room at any party involving software engineers.