What linear models actually do: classification and regression with a straight line
Long before neural networks, scientists needed to make predictions from data. Given a set of measurements — temperature readings, astronomical observations, economic indicators — they wanted a formula that could predict new values from known ones. The simplest such formula is a straight line:
Slope , intercept , one input, one output. With multiple inputs it becomes:
A weighted sum of inputs plus a bias. This is the equation at the core of every neuron in every neural network — but it started as a tool for fitting lines to data.
Fitting a line to data
In the early 1800s, Gauss and Legendre independently developed a method called least squares — a way to find the straight line that best fits a set of data points. “Best” means the line that minimizes the total squared distance between itself and every point:
where is the actual value and is what the line predicts. Squaring makes all errors positive and penalizes large errors more than small ones. This is mean squared error — the same loss function used to train neural networks on regression tasks today.
The beauty of least squares is that it has a closed-form solution. You don’t need to search or iterate — you plug the data into a formula and get the exact optimal and :
One calculation, exact answer. This made linear regression practical long before computers existed — scientists could compute it by hand.
Try it yourself. Drag the sliders to position the line manually — the red dashed lines show the error at each point. Then click Fit to snap to the exact least squares solution, or Gradient descent to watch an iterative algorithm find the same answer step by step:
The Fit button uses the closed-form formula — instant. Gradient descent computes and and nudges the parameters a little at a time. Both land on the same line. Gradient descent is slower for this simple case, but it generalizes to models where no closed-form solution exists — which is every neural network. This is the same algorithm covered in detail in the backpropagation article.
Linear regression was the workhorse of science and economics for two centuries. It’s fast, stable, interpretable — you can look at the slope and immediately understand how much changes when increases. And it’s surprisingly strong: even complex relationships often behave approximately linear in small ranges.
Same equation, different job
Regression and classification look different, but the model is identical — multiply inputs by weights, add a bias, get a number. The difference is what you do with that number:
- Regression: use it directly — . The linear function makes a prediction.
- Classification: threshold it — is ? The linear function draws a boundary.
The equation defines a line in 2D space — the decision boundary. Everything on one side belongs to one class, everything on the other to the other. The model is still a weighted sum; the threshold turns that one number into a decision.
The first machine to learn its own decision boundary from examples — rather than have a human compute the weights — was Rosenblatt’s perceptron in 1958, followed by Widrow and Hoff’s ADALINE in 1960. Both used the linear equation above; they differed only in the loss they trained against. The widget below shows a perceptron-style learner finding the boundary one correction at a time. Watch the yellow line rotate and shift until it separates the two classes:
The exact training rule the widget uses, why it works, and why ADALINE’s later refinement matters more than the perceptron version, are covered in the ADALINE article — including a physical 16-knob machine you can train by hand. For our purposes here, the point is just that the same linear equation underwrites both regression and classification; only the loss and the post-processing change.
This duality persists in modern neural networks. The MNIST article uses dense layers for classification (softmax output, cross-entropy loss). The backpropagation article uses the same layers for regression (linear output, MSE loss). The layers don’t know or care which task they’re serving — they just compute weighted sums and pass gradients.
The wall
The excitement lasted about a decade. In 1969, Marvin Minsky and Seymour Papert published Perceptrons, a book that proved a hard mathematical limit: there are simple patterns a single linear model cannot learn.
The most famous example is XOR — a function that outputs 1 when exactly one of two inputs is 1:
| XOR | ||
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
Plot these four points and try to draw a single straight line separating the 1s from the 0s. You can’t — the classes are interleaved diagonally. No values of , , and will ever make positive for and but negative for and . The data isn’t linearly separable, so the perceptron’s convergence guarantee doesn’t apply. It will bounce forever and never settle.
XOR is a toy example, but the limitation is general. Any pattern that requires a curve, a bend, or an interaction between variables is beyond reach. Click through the tabs below — each shows the best-fit line on a different kind of data:
The linear tab fits perfectly. The next three don’t — the line systematically misses the pattern. This is underfitting: the model is too simple for the data, and no amount of training will fix it.
Underfitting is only half the story. The opposite failure shows up the moment you hand the model too much flexibility. Switch to the overfit tab: the data are essentially a noisy straight line — the same shape as the first tab — but now they are fit with a flexible degree-9 polynomial instead of a line. The curve twists to pass through almost every point, driving the training error to nearly zero. That looks like a triumph until you remember the wiggles are tracking the noise, not the signal: feed this curve a new point drawn from the same line and it will miss badly. A straight line ignores the noise and generalizes; the squiggle memorizes it and doesn’t. The job is to land between the two extremes — too rigid on one side, too flexible on the other — which is the bias–variance tradeoff that the regularization toolkit takes up in full, with the Goldilocks-zone risk curve that makes the balance precise.
Minsky and Papert’s proof was devastating. Research funding dried up, and the field entered what’s now called the AI winter — a long stretch where few people worked on learning machines. The linear model was correct about the core idea (machines learning from data), but it was stuck behind a hard mathematical ceiling.
Past the wall, in one paragraph
Stacking linear layers doesn’t help — two linear transformations back to back, , collapse algebraically into a single linear transformation . Ten layers collapse into one. The escape is to put a nonlinear activation function between layers — ReLU, sigmoid, tanh — so the next layer is recombining bent versions of the previous layer’s outputs, not the raw linear projections. With activations between layers, the composition no longer collapses, and stacks of these layers can approximate any function (the universal approximation theorem). The linear model isn’t a stepping stone you leave behind — it’s the brick every modern network is built from.
The historical detour from “single linear model” to “stacked layers with nonlinearities” is its own story: Widrow and Hoff’s first attempt at multi-layer networks in 1962 (MADALINE), the awkward heuristic they used to train it (no chain rule yet for step functions), and the backpropagation paper by Rumelhart, Hinton, and Williams in 1986 that finally made deep nets trainable. The ADALINE article walks through the machine that first hit this wall and what came after; the backpropagation article covers the modern resolution.