Throughout this series the target has depended on a single feature: one input x, one weight, one bias. Parents' height in, child's height out. Hours studied in, exam score out. Most real problems are not so tidy. Predicting a house price might draw on floor area, number of bedrooms, age, and location all at once, and no single one of those tells the whole story.

This part is short, because the answer to "what changes?" is mostly "nothing". But it is worth seeing exactly where nothing changes, because the same argument is what lets the models in the rest of the series grow from two knobs to billions.

One weight per feature

Having several features is called multivariate regression (several variables), as opposed to the univariate (one variable) regression we have used so far. The model simply grows more terms, one weight per feature, plus the single shared bias:

y = w1·x1 + w2·x2 + w3·x3 + ... + b

Each feature gets multiplied by its own weight, the results are added up, and the bias is added on. That is all. For a house with three features, floor area, number of bedrooms and age, the data might look like this:

Floor area (m²) Bedrooms Age (years) Price ($000s)
60 2 10 260
85 3 5 370
100 3 20 385
120 4 2 506
75 2 30 265

Three feature columns, one target column. Training finds one weight for each feature, plus a bias. For this data a fitted model comes out as:

price = 3 × area + 25 × bedrooms − 2 × age + 50

Each weight is a price sensitivity you can read straight off: about $3k for every extra square metre, $25k per bedroom, and $2k off for every year of age, on top of a $50k base. Check the first row: 3 × 60 + 25 × 2 − 2 × 10 + 50 = 180 + 50 − 20 + 50 = 260. Feed a new house's three numbers into that equation and out comes its predicted price.

Try it: what each weight contributes

Set the three features of a house below. The bars show how much each term of the formula adds to (or takes off) the price, and the total is the prediction. Notice that changing one feature moves only its own bar: the features do not interact, they just add up. That additive simplicity is what "linear" means when there are several inputs.

Training does not change

Nothing about the recipe from Parts 4 to 6 changes. You still make a prediction for every row, measure the loss, and work out which way each knob should move. The only difference is that "each knob" is now four knobs instead of two: the partial derivative of the loss with respect to w1 says which way to nudge w1, the one for w2 says which way to nudge w2, and so on. Every weight gets its own slope, and the same xᵢ leverage rule from Part 5 applies to each: the slope for the area weight uses each row's area, the slope for the bedrooms weight uses each row's bedroom count. Then all the knobs take their small step downhill together, exactly as weight and bias did on the map in Part 5.

The loss landscape from the last part grows too. With two knobs it was a surface over a floor; with four knobs it lives in a space with one axis per knob, which nobody can draw. But gradient descent never needed the picture. It only ever needed the slope in each direction, and that it can always compute, whether there are four directions or four billion.

For the curious: one practical wrinkle does appear with several features. Floor area runs into the hundreds while bedrooms run from 1 to 6, so a step of the same size means something very different to the two weights, and the loss landscape becomes a long thin valley of the kind you saw on the map in Part 5, only more so. Gradient descent still gets there, but slowly and with a lot of zig-zagging. The standard fix is to rescale each feature before training so they all run over a similar range (subtracting the average and dividing by the spread is the usual recipe). This is called feature scaling or normalisation, it is a routine first step in almost every real project, and it is purely a convenience for the optimiser: the model it finds is the same, just found faster.

Why this matters for what comes next

Look at the multi-feature formula once more:

y = w1·x1 + w2·x2 + w3·x3 + b

Several inputs, a weight on each, add them up, add a bias. Hold on to that shape. In the next part a single unit called a neuron is introduced with one input, to keep the pictures simple. Real neurons take many inputs, and the sum they compute is exactly this formula. So a neuron is a small multivariate linear model with one extra step bolted onto its output, and a neural network is a great many of them wired together. Everything you know about training the formula above transfers directly.


Next: Activation Functions: the trick that lets a model bend beyond straight lines to fit curves and other non-linear shapes.