Draft

The knob-turning machine: playing with ADALINE, the first trainable neuron

In 1960, a Stanford professor named Bernard Widrow and his graduate student Ted Hoff built a machine that could learn. It was not code. It was a cabinet of wires, switches and electrochemical cells — a physical device where each weight in the model was a real, tunable component, and where “training” meant small electric currents shifting copper atoms around inside tiny glass tubes. They called it ADALINE, short for Adaptive Linear Neuron.

ADALINE is the direct ancestor of every neural network running today. The summing junction, the learning rule, the idea that a machine can adjust its own parameters based on its error — all of it starts here. And because the original had only sixteen inputs and one output, it’s small enough that we can rebuild it in a single web page and watch it learn in front of us.

What can one neuron do?

Before we open the cabinet, it’s worth asking what to expect from it. ADALINE is a single neuron — one weighted sum, one threshold, nothing hidden. That sounds like almost nothing. How much can one neuron actually do?

More than you’d guess, but with a hard limit — and it’s worth knowing both halves before you start turning knobs.

The ceiling is real, and low. A single neuron computes one weighted sum and checks its sign, which geometrically is a single flat cut through the space of inputs — one line in 2D, one plane in 3D, one hyperplane in 16D. Any two groups of patterns you can separate with one straight cut, it can learn to tell apart. Any you can’t, it provably never will: the four points of XOR are enough to defeat it forever, and that exact failure is what stalled the whole field for a decade and forced the invention of layered networks. (We’ll drive the machine straight into that wall in exercise 5.)

And yet, under that ceiling, it does real work. One neuron can learn to tell a hand-drawn T from a J, by turning its weights into a template of the very thing it’s looking for. The same neuron, the same learning rule, with nothing changed but the names of its inputs and outputs, becomes an adaptive filter — and in that disguise it cancelled echoes on transcontinental phone calls, equalised modems, and quieted the noise in headphones, running in production hardware for decades. And the rule it learns by — measure the error, take a derivative, nudge each weight against the gradient — is, line for line, the one running inside every neural network today.

So one neuron is at once the smallest thing that can learn and a useful tool in its own right. But it’s also the atom of everything that came after — so before we start turning its knobs, it’s worth asking what happens when you have more than one.

What can a layer of them do?

If one neuron draws one cut, the obvious next move is to line several up side by side — each reading the same inputs, each with its own knobs, each drawing its own cut. That’s a layer, and it buys you exactly one thing: more lines at once. Where a single neuron can only answer “T or not-T,” a layer of twenty-six can answer “which letter is this” — one neuron per class, each firing for its own. Multi-class recognition, the workhorse of every classifier since, is just a row of ADALINEs scoring the input in parallel.

What a layer does not buy you is a way around the ceiling. Every cut it draws is still straight. Read its outputs directly and XOR is still impossible — many flat lines is not the same as one bent one. And you can’t cheat by stacking, either: a layer of linear neurons feeding another layer of linear neurons collapses, algebraically, right back into a single layer. Two flat cuts composed are still one flat cut.

The thing that finally bends the boundary is a layer feeding another layer with a non-linear step wedged in between — so the second layer gets to work on a warped version of the input the first one couldn’t separate. That single ingredient is the leap from ADALINE to MADALINE, and later to backprop; we reach it near the end. For now, hold the ladder in mind: one neuron, one cut; one layer, many cuts; stacked layers with a non-linearity, curved cuts. Everything in this article lives on the bottom rung — so that’s where we start.

But sixteen knobs is a lot to watch at once. Start with the smallest version that still has every essential part: two inputs, two knobs, one output. Each knob sets one weight; the meter shows the weighted sum z=w1x1+w2x2z = w_1 x_1 + w_2 x_2; the LED reports its sign. There is nothing else.

x₁x₂+1+1+0.00+0.000−+z = 0.00ŷ+1−1TWO-KNOB ADALINE
z = (+0.00)(+1) + (+0.00)(+1) = +0.00
target yerr +0.00

Grab either knob and turn it. The needle tracks your hand in real time, and the live equation underneath shows exactly why: you’re changing one number in a two-term sum, and that sum is the output. Turn a knob far enough and the needle crosses zero — the prediction LED flips from +1+1 to −1-1. That is the whole idea of a weight: a direct, physical handle on the answer.

So here is how you make the machine say what you want. Set the two switches to the input pattern you care about, pick a target with the yy buttons (the blue dashed line on the meter marks it), then dial the knobs by eye until the needle lands on it. No learning rule required — you just built the output you wanted by turning knobs. Now press Train step and watch the machine turn those same two knobs for you, nudging zz toward the target a little at a time. That automatic knob-turning is the LMS rule we’ll derive below — it does nothing you couldn’t do by hand; it just picks the direction and amount. The only thing that changes when we scale up is the number of handles.

Here is that full machine.

0−+z = 0.00PREDICTION ŷ−1+1TARGET y−1 (J)+1 (T)KNOBBY ADALINEstanford · 1960
step 0 · err +0.00

How to play: flip the switches on the right to draw a pattern, pick a target label (+1 or −1), and press Train step. Each press runs one pass of the Least Mean Squares rule — the same rule Widrow and Hoff derived in 1960 — and you’ll see the sixteen weight dials rotate. Do it enough times, with a few different patterns, and the machine learns to tell them apart.

The rest of this article walks through what you’re actually looking at — where the machine came from, what the dials really represented in the original hardware, why the learning rule works, and what survives of ADALINE inside a modern transformer.

Where it came from

Bernard Widrow arrived at Stanford from MIT in 1959, twenty-nine years old and newly hired as an assistant professor of electrical engineering. Frank Rosenblatt had just published the perceptron the year before — a single-layer learning machine that had drawn enormous press attention, including a famous New York Times piece predicting machines that would soon “walk, talk, see, write, reproduce itself and be conscious of its existence”. Widrow wanted to build something similar, but with a learning rule he could derive from first principles rather than argue for heuristically.

His first PhD student, Marcian “Ted” Hoff — who would later co-invent the Intel 4004 microprocessor — took on the hardware. The two of them published “Adaptive switching circuits” at the 1960 WESCON conference, introducing both the machine and the LMS rule in one paper. ADALINE stood for Adaptive Linear Neuron.

The memistor. In the original ADALINE, each weight was a memistor — a device Widrow invented specifically for this project. A memistor was a small glass tube containing a graphite rod suspended in copper sulphate solution. Running current one direction through the tube plated copper onto the rod, thickening it and lowering its resistance. Running current the other way deplated copper, raising the resistance. The rod’s resistance was the weight. No software state, no register in memory — each weight was a physical amount of copper on a metal rod, adjustable by electrolysis.

Training was electrochemistry. You set the input switches, let the summing junction compute zz, measured the error (y−z)(y - z), and drove a small current through each memistor proportional to η⋅(y−z)⋅xi\eta \cdot (y - z) \cdot x_i — the exact LMS update. A few milliseconds of current per pattern, a few thousand patterns, and the copper had redistributed itself into the weights that recognise T and reject J.

“Knobby ADALINE” — the machine in the photo every textbook uses — was a slightly later, teaching-oriented demonstrator. The memistors were replaced with physical potentiometers whose shafts you could grab and turn, so students could watch the learning rule at work and even set the weights by hand to see how predictions shifted. Our simulation imitates this version, because it’s the more legible one.

MADALINE. In 1962, Widrow and colleagues connected several ADALINEs together into a Multiple ADALINE — MADALINE. This machine solved its first real industrial problem shortly after: adaptive echo cancellation on long-distance telephone lines, which AT&T put into production and ran for decades. If you’ve ever had a clear phone call across an ocean, a direct descendant of ADALINE was in the circuit somewhere.

The rename. In 1969, Marvin Minsky and Seymour Papert published Perceptrons, a book-length proof that single-layer networks like ADALINE couldn’t learn simple functions such as XOR. Funding for neural networks collapsed. When Widrow wrote about his machines in the 1970s and 80s he began calling the acronym Adaptive Linear Element — same letters, but with the word “neuron” quietly removed. The brand had become a liability. It would take until the mid-1980s and the rediscovery of backpropagation (Rumelhart, Hinton, and Williams, 1986) for multi-layer networks to break past the XOR ceiling; “neural” became a safe word again only in the 1990s.

Widrow, now in his nineties, remains emeritus professor of electrical engineering at Stanford. The canonical retrospective he wrote — “30 Years of Adaptive Neural Networks: Perceptron, Madaline, and Backpropagation” (Widrow & Lehr, Proceedings of the IEEE, 1990) — is the one paper to read if you want the history from the source, with every diagram and every convergence curve.

What the dials are

The machine has three kinds of parts, and together they implement a single equation.

The switches — sixteen toggles arranged as a 4×4 grid on the right. Each switch stores one input value xix_i, either +1+1 (up, lit) or −1-1 (down, dim). You set them by hand; they’re how you draw the pattern you want the machine to look at. In Widrow’s original hardware, these were real electrical toggle switches wired into the rest of the circuit.

The knobs — sixteen cream dials arranged as a 4×4 grid on the left. Each knob stores one weight wiw_i in the range [−1,+1][-1, +1], represented physically as the rotation angle of its pointer. Pointing straight up means wi=0w_i = 0; full right is wi=+1w_i = +1; full left is wi=−1w_i = -1. You can drag any knob with the mouse to set its weight by hand. Along with a hidden bias term bb, these seventeen numbers are the entire model — every piece of information the machine has about the world. (Same idea as the “a neural network is just parameters” section of the earlier article: the weights are the network.)

The summing junction — the cream-faced meter with the black needle. This is the one piece of electronics that actually computes something. It takes every switch, multiplies it by its paired knob, adds up all sixteen products, and swings its needle to show the result:

z  =  ∑i=116wi xi  +  bz \;=\; \sum_{i=1}^{16} w_i \, x_i \;+\; b

When the needle is dead-centre, z=0z = 0. When it swings right, the sum is positive. Left, negative. This weighted-sum-plus-bias is the same equation discussed in the abstract in What linear models actually do — there used to fit lines to continuous data, here used to score whether a pattern matches a learned template.

The prediction. The machine doesn’t report zz directly — it reports its sign. If z≥0z \geq 0 it decides "+1+1" and lights the green LED. If z<0z < 0 it decides "−1-1" and lights the orange one. That final sign step is the threshold, or “decision”, of the classifier:

y^  =  sgn(z)\hat{y} \;=\; \mathrm{sgn}(z)

That’s the entire forward pass: switch times knob, sum everything, check the sign. No matrix multiplies, no softmax, no hidden layers. Just a weighted sum followed by a threshold — and yet this is the building block every modern network is an elaboration of.

The spatial trick

The knobs and switches are laid out in the same 4×4 shape on purpose. The knob at position (r,c)(r, c) multiplies with the switch at position (r,c)(r, c). This isn’t fundamental to the math — weights could be numbered any way — but it makes the learned structure visible. When the machine has learned the letter T, the knobs themselves spell out T: wherever T has an “on” pixel, the matching knob is turned right (positive weight → this pixel contributes toward +1+1); wherever T is off, the knob is turned left. The grid of weights becomes a picture of what the machine is looking for.

How it learns

The forward pass was easy: multiply and sum. The hard part is training — figuring out how to nudge each of the seventeen numbers so the predictions get better.

Widrow and Hoff’s answer, in 1960, was the Least Mean Squares rule (also called the delta rule, or the Widrow–Hoff rule). It’s a single equation per weight:

wi  ←  wi  +  η⋅(y−z)⋅xiw_i \;\leftarrow\; w_i \;+\; \eta \cdot (y - z) \cdot x_i

and one for the bias:

b  ←  b  +  η⋅(y−z)b \;\leftarrow\; b \;+\; \eta \cdot (y - z)

Here yy is the target label (+1+1 or −1-1), zz is the summing junction’s current output, and η\eta is a small number called the learning rate — in this machine η=0.02\eta = 0.02. The quantity (y−z)(y - z) is the error: how wrong the machine is right now, measured on the analog sum rather than the thresholded prediction.

Press Train step and the machine does exactly this: it reads zz, subtracts it from yy, multiplies by η\eta and by each input in turn, and adds the result to each weight. Every pointer rotates a little, the needle swings, the LED maybe flips. If you press it enough times on the same pattern, the needle will peg at the target value and the error drops to near zero — the machine has learned that one pattern. Press it on a second pattern with the opposite label and the knobs rearrange to fit both at once.

Why this rule, specifically

The short answer is calculus. Define the squared error of one prediction — the same MSE loss used in the earlier article:

E  =  12(y−z)2E \;=\; \tfrac{1}{2}(y - z)^2

Differentiate with respect to wiw_i, pushing the chain rule through z=∑jwjxj+bz = \sum_j w_j x_j + b:

∂E∂wi  =  −(y−z)⋅xi\frac{\partial E}{\partial w_i} \;=\; -(y - z) \cdot x_i

Step in the direction of the negative gradient, scaled by η\eta:

wi  ←  wi−η⋅∂E∂wi  =  wi+η(y−z)xiw_i \;\leftarrow\; w_i - \eta \cdot \frac{\partial E}{\partial w_i} \;=\; w_i + \eta(y - z) x_i

That’s the LMS rule. It is gradient descent on squared error, applied to a single neuron. The only real difference between this update and the SGD you’d run in PyTorch today is the size of the network and the shape of the loss. The move — compute error, take a derivative, nudge in the opposite direction — is identical.

Why the error uses the sum, not the decision

You might have expected the error to be y−y^y - \hat{y} — “how wrong is the final decision” — rather than y−zy - z, “how wrong is the analog sum”. There are two reasons, one technical and one practical.

Technical: the threshold y^=sgn(z)\hat{y} = \mathrm{sgn}(z) is a step function. Its derivative is zero everywhere except exactly at z=0z = 0, where it’s undefined. You cannot do gradient descent through something whose slope is zero — every update would be zero, no learning. LMS side-steps this by doing gradient descent on the smooth sum before the threshold. The threshold is still there at prediction time; it just isn’t in the training loop.

Practical: the continuous error gives a graded signal. Even when the decision is already correct — say z=0.1z = 0.1 and y=1y = 1 — there’s a meaningful error of 0.90.9, and LMS keeps pushing zz further from zero, making the prediction more confident. A post-threshold error would report 00 here and stop learning, leaving the machine one noise-bump away from misclassifying.

This trick — train on the pre-activation error — is still how every modern network works. Cross-entropy loss, MSE, everything; they all operate on a continuous pre-decision quantity, not on the discrete prediction.

The geometric view

There’s a clean geometric reading of everything you just saw — the forward pass and the learning rule together. Once you see it, the seventeen numbers stop being a list and become one arrow in space, with one flat boundary perpendicular to it.

We’re going to look at the same quantity zz in three representations, each unlocking a different intuition: the scalar sum we already have (z=∑iwixi+bz = \sum_i w_i x_i + b), the vector dot product (z=w⋅x+bz = \mathbf{w} \cdot \mathbf{x} + b), and the geometric reading (zz = signed distance from the input pattern to a hyperplane). All three compute the same number; what changes is what you can see about the machine.

The weights are an arrow. The sixteen wiw_i are the components of a single vector w\mathbf{w} in 16-dimensional space — one axis per input. Each input pattern is also a vector x\mathbf{x} in that same space: each switch contributes one coordinate, +1+1 or −1-1. The summing junction’s whole job is one operation on these two vectors:

z  =  w⋅x  +  bz \;=\; \mathbf{w} \cdot \mathbf{x} \;+\; b

That’s a dot product, plus an offset. And a dot product, geometrically, is the signed projection of x\mathbf{x} onto the w\mathbf{w} direction — a measure of how far along that direction the pattern lies.

The threshold is a hyperplane. The set of inputs where z=0z = 0 — the boundary between green LED and orange LED — is the flat slab perpendicular to w\mathbf{w}, offset from the origin by bb. In two dimensions that boundary is a line; in three it’s a plane; in sixteen it’s a fifteen-dimensional hyperplane. Same shape, just one dimension less than the space it lives in. Any pattern on the side w\mathbf{w} points toward gives z>0z > 0 and the LED lights green; any pattern on the opposite side gives z<0z < 0 and it lights orange. The machine’s entire opinion about any input is which side of one flat boundary it falls on, and how far.

This also tells you what zz‘s magnitude means: it’s (proportional to) the signed distance from x\mathbf{x} to that plane. A pattern far from the boundary on the w\mathbf{w} side gives a large positive zz — confident +1+1, needle pegged right. A pattern that lies exactly on the boundary gives z=0z = 0 — needle dead centre, the LED could go either way.

Training is rotating and sliding the plane. Each LMS update tweaks the weight components, and tweaking w\mathbf{w} rotates the plane that’s perpendicular to it. Each update to bb slides the plane without rotating it. The entire training procedure is a single flat boundary being shoved around in 16-D space until the +1+1 patterns end up on the w\mathbf{w} side and the −1-1 patterns on the other. The knobs, the needle, the LED — every visible element of the machine is a surface reading of that one rotating-and-sliding plane.

This gives the cleanest one-line description of what ADALINE actually does: finding the right w\mathbf{w} and b\mathbf{b} is finding the right hyperplane. The seventeen knobs and the boundary aren’t two separate things the machine is juggling — they’re the same object expressed two ways. Adjust the knobs and you’ve adjusted the boundary; pick a boundary and you’ve picked the knobs.

Many planes will do — that’s why “best” is a separate question. If the patterns are linearly separable at all, infinitely many planes separate them. LMS guarantees only that it finds one, not the one with the largest gap between clusters or the strongest generalisation to unseen inputs. Picking which separating plane — the one that maximises the margin between classes — is a different problem, and its answer (Vladimir Vapnik’s support vector machines, 1995) wouldn’t arrive for thirty-five years.

You can’t see this in the 16-knob widget. You cannot draw a 16-D arrow or a 15-D hyperplane. Two dimensions is the largest case where the geometry sits cleanly on a page — one input on the x-axis, one on the y-axis, w\mathbf{w} drawn as a literal arrow, and the decision boundary drawn as a literal line. The widget in exercise 5 is exactly that: a two-input ADALINE with the arrow and the line drawn for you. The line you’ll watch swing into place there is the same kind of plane that’s silently rotating inside the 16-D machine above — just one you can finally see.

Beyond one plane. A single flat boundary can only separate patterns that are already linearly separable in the input space — and most interesting datasets aren’t. The escape route isn’t a cleverer plane; it’s stacking layers so the network can bend the input space itself, warping non-separable clusters into a new geometry where one flat plane suffices. Chris Olah’s Neural Networks, Manifolds, and Topology is the canonical visual demonstration — sheets of input space being stretched, folded, and pulled apart by successive layers until a single hyperplane can finally make the cut. ADALINE is what happens when you have only one such cut and no warping. MADALINE, in the next section, is the first attempt at adding the warp.

What to try

A few things to play with. Each exercise teaches a different piece of the picture.

1. Teach it one letter

Press RESET to randomise all sixteen knobs. Load T. The target is already on +1+1. Press Train ×20.

Watch the cream knobs rotate into a pattern. By the end you’ll see something striking: the knobs that line up with T’s “on” pixels — the top row, and the vertical bar down the middle — have all swung to the right, toward positive weights. The knobs for pixels T doesn’t use are turned left. The grid of knobs has become a picture of T. The weights are a template the machine matches against.

2. Teach it two letters at once

Straight after step 1, Load J and flip Target to −1-1. Press Train ×20 again. Now the machine has been shown two patterns with opposite labels.

Flip back to Load T (target +1+1) without training. Where is the needle? It should be on the +1+1 side. Do the same for Load J — needle on −1-1. A single set of sixteen knobs has learned to classify two patterns. That’s memory, physical and adjustable.

3. Drive the output with one knob

Forget the learning rule for a moment and take the controls yourself. After step 2 the machine reads T as +1+1 — load T and the needle sits on the green side. Now pick one knob and slowly drag it. Watch the needle move with your hand: each degree you turn the pointer adds or subtracts that knob’s contribution from the sum, and the needle tracks it in real time. Keep turning until the needle crosses dead-centre and the LED flips from green to orange. You just changed the machine’s answer by adjusting a single weight — no training, no error signal, just one knob and the output you wanted.

That’s the whole point in miniature: a weight is a direct, physical handle on the output. The needle’s position is nothing but the sum of sixteen such handles, and moving any one of them moves the result. The learning rule from the next section isn’t doing anything more mysterious than this — it’s turning these same knobs, just choosing the direction and amount for you instead of leaving it to your eye. (This is exactly the hand-setting that perceptron labs in the late 1950s did before a learning rule existed; exercise 4 builds a whole template this way.)

Then hand it back: press Train step a few times with the current pattern loaded. The machine pushes that one knob back toward a value that fits all the patterns it knows. The others fidget slightly — with ±1\pm 1 inputs every weight updates every step — but they were already close to right, so they settle back quickly. You moved the output by hand; the rule moves it back on purpose.

4. Train it by hand

Press RESET so every knob is near zero. Without pressing Train, drag each knob yourself: turn the knobs for T’s “on” pixels all the way right, and the rest all the way left. Then Load T and check the needle. You just trained the machine by eyeball — which is literally what happened in the perceptron labs of the late 1950s, before Widrow’s rule existed. The learning rule doesn’t do anything you couldn’t do by hand; it just automates the eyeball.

5. Find its ceiling

Reset, then train T as +1+1, then L as −1-1, then J as −1-1 (twenty steps each). All three should now classify correctly — T goes positive, L and J both go negative.

Now try a harder set. Take T and manually flip one switch — say the top-left corner — making a “broken T”. Train this broken T to target −1-1 while still wanting the clean T to stay at +1+1. Keep cycling through all four patterns (clean T, broken T, L, J) and pressing Train ×20 on each.

Sometimes this converges; more often, the error just bounces around and never settles. This is ADALINE’s actual ceiling. A single layer of weights can only separate patterns with one straight hyperplane through the space of inputs. Any set of patterns that can’t be cut apart by a single flat plane is out of reach. This is exactly the limitation Minsky and Papert published in 1969 — the limitation that multi-layer networks and backpropagation were invented to overcome.

It’s hard to see why in 16 dimensions. In 2 dimensions you can see it directly. Below is a miniature ADALINE with only two inputs x1,x2x_1, x_2 — every pattern is now a single point on a plane, every weight vector is an arrow, and the “decision boundary” is a single line that rotates as the machine trains. Press TRAIN ×50 on the SEPARABLE dataset: the line swings into the gap between the two clusters and the error drops to near zero. Now switch to XOR: the same algorithm, running on four points that can’t be separated by any line, never settles. The line tilts one way, then the other, then back. That endless bouncing is exactly what happens in 16 dimensions when you ask ADALINE to do something linearly impossible — you just can’t see it as directly.

2-input ADALINE — the decision line
x₁x₂wstep 0|err| avg 0.00w = (0.00, 0.00)b = 0.00
dataset:

6. Watch the error

Through every exercise above, keep an eye on the err value in the bottom strip. Right after loading a fresh pattern, ∣err∣|\text{err}| is typically 11 to 22. After a handful of Train steps it drops close to zero and stays there. If it refuses to shrink, either you’ve hit a perfect fit (LED lit, needle pegged — you’re done) or you’ve hit the linear-separability ceiling from exercise 5, and there’s no single setting of sixteen knobs that can satisfy all the patterns you’ve asked for at once.

What came next: MADALINE

The XOR ceiling was a real problem, and Widrow knew it. In 1962, he and his next graduate student stacked several ADALINEs together into a Multiple ADALINE — MADALINE — and asked whether a network of ADALINEs could do what one alone couldn’t.

The architecture below is the simplest one that does: two inputs, three ADALINE units in a hidden layer, and a fixed majority vote as the combining layer. Each hidden unit is a miniature ADALINE, computing zj=wj1x1+wj2x2+bjz_j = w_{j1} x_1 + w_{j2} x_2 + b_j and emitting hj=sgn(zj)h_j = \mathrm{sgn}(z_j). The final prediction is the majority of h1,h2,h3h_1, h_2, h_3. The dataset is XOR — the exact task that sinks the single-layer machine.

INPUTHIDDEN LAYER — 3 ADALINEsMAJORITYOUTPUTMAJvote?ŷTARGET y−1+1XOR pattern 1/4(x₁, x₂) = (+1, +1)KNOBBY MADALINE
epoch 0 · wrong 0/4

Press TRAIN ×20 EPOCHS and watch it run. A few things to notice.

A flash tells you what MRI just picked. Whenever MRI updates a hidden unit, a cream-coloured ring flashes around it for a fraction of a second. That’s the unit the rule identified as “easiest to flip” — the one whose zjz_j sat closest to zero among the wrong-voting units.

There’s no gradient here. The hidden units go through a sgn\mathrm{sgn} function before the vote, and you can’t differentiate through a step. That’s precisely the problem Widrow and his students could not solve with calculus in 1962. MADALINE Rule I (MRI) — the training rule this widget implements — is the 1962 hack: if the final output is wrong, find one hidden ADALINE that’s currently voting the wrong way and whose pre-threshold sum is closest to the flip point, then do one step of LMS on just that unit to push its sum over zero. Repeat until the output is right or you give up. Re-present the pattern. Cycle through all four. It’s heuristic, it’s ugly, it sometimes gets stuck in configurations from which no single flip helps — and yet it works often enough to be useful.

Sometimes it solves XOR, sometimes it doesn’t. Depending on the random initialisation, you’ll see one of three outcomes: (a) the error reaches 0/4 within a few epochs and stays there — MADALINE has learned XOR, something ADALINE provably can’t; (b) it oscillates, with the error flipping between 2/4 and 0/4 as training fights itself; (c) it settles on some other wrong configuration. Press RESET and try again. The finickiness is the story, not a bug in the simulation.

What MADALINE did in production. Despite its training rule’s awkwardness, Widrow used MADALINE (and its successors MRII, MRIII) to solve real problems through the 1960s and into the 1980s — adaptive echo cancellation on long-distance telephone lines (put into production by AT&T), adaptive antenna beamforming, modem channel equalisation. These filters were neural networks in commercial deployment, running silently inside infrastructure, a full two decades before the “neural network revival” of the late 1980s made the word respectable again.

The real fix came in 1986. The awkwardness of MRI is a direct consequence of sgn\mathrm{sgn} having no useful gradient. Replace it with a smooth non-linearity (sigmoid, later tanh, later ReLU), and the chain rule starts working all the way down through the layers — which is exactly what backpropagation is. MADALINE’s hidden-unit problem isn’t solved by a cleverer heuristic; it’s dissolved by picking a different activation function, so that “which unit to update” stops being a search and becomes calculus.

The other framing: ADALINE as an adaptive filter

The MADALINE-on-AT&T story above is worth slowing down on, because it reveals something the “first neural network” framing tends to obscure: the LMS rule and the adaptive filter are the same algorithm, and the killer application of Widrow’s 1960 paper wasn’t pattern recognition at all — it was signal processing.

The problem. A long-distance phone call in the 1960s travelled over a four-wire trunk between cities and a two-wire local loop at each end, joined by a piece of analog hardware called a hybrid coupler. A perfect hybrid would route the far speaker’s voice into your earpiece and your voice out onto the trunk, with no leakage. Real hybrids weren’t perfect. A small fraction of your voice leaked back across the coupler at the far end, travelled all the way back across the country, and arrived in your earpiece a few hundred milliseconds later. That’s an echo — your own voice, delayed by long enough to be unmistakable, just quiet enough to be maddening. Conversation collapsed.

The trick. You can’t stop the leak. But if you know what you just sent and you know how the echo path will warp it (delay DD, attenuation α\alpha, possibly some smearing), you can predict the echo before it arrives back, and subtract your prediction from the return signal. What’s left is what the far speaker actually said. The whole job of the canceller is to learn the echo path’s impulse response well enough to cancel it out in real time.

The mapping to ADALINE is exact. Drop the picture of the 16-knob machine into a phone line and rename every part:

ADALINE partAdaptive filter equivalent
Knobs w1…w16w_1 \dots w_{16}Filter taps — the impulse response being learned
Switches x1…x16x_1 \dots x_{16}Sliding window of recent TX samples — what you just sent
Summing junction z=∑wixiz = \sum w_i x_iPredicted echo at this instant
Target yyReturn signal (RX) — what’s actually arriving back
Error y−zy - zResidual echo — what the listener actually hears
LMS update wi←wi+η(y−z)xiw_i \leftarrow w_i + \eta(y - z)x_iTap update — exactly identical

Same equations, same convergence guarantees. The only thing that changed is what the inputs and outputs mean.

ADALINE as an adaptive filter — echo cancellation
TXRXRESTAPStap index k →step 0residual RMS 0.00true echo: D=6 α=0.60learned taptrue coefficientscale: ±1.0color: + = cream, − = orange+1−1+1−1+1−1

What you’re looking at. The top trace (TX) is the outgoing signal — a synthetic mix of sine waves standing in for speech. The middle trace (RX) is what’s arriving back: the same signal, delayed by DD samples and scaled by α\alpha, plus a little ambient noise. The bottom trace (RES) is the residual: RX minus the filter’s prediction. The bar chart is the sixteen tap weights — start at zero, so the filter predicts nothing and the residual is the echo. Press TRAIN and the LMS rule starts adjusting the taps, one sample at a time, to drive the residual toward zero. Within a few seconds you should see one tap rise into a tall spike at index DD, the others stay near zero, and the residual flatline. The filter has discovered the echo’s impulse response by watching the error. That’s the entire trick.

Why it has to be adaptive. A fixed filter would do, in principle, if the echo path were fixed. It isn’t. Cable temperature changes, switching fabrics rearrange paths, hybrids age, customers pick up and hang up. Toggle DRIFT and the true delay and attenuation start jittering; the filter chases them. The filter is never done. It runs forever during the call, perpetually re-tuning. This is what “adaptive” means in adaptive filtering, and it’s exactly what online SGD does in modern neural networks: never freeze, always update.

Why this framing shipped while “neural network” didn’t. The 1969 publication of Minsky and Papert’s Perceptrons made “neural network” toxic for funding and respectability for a full decade. But the same math, presented as adaptive filtering in IEEE signal-processing journals, was uncontroversial — it solved real problems, it had clean convergence proofs, no biological metaphor required. So Widrow and his students kept publishing under the filter framing, kept landing real deployments (AT&T echo cancellation, modem equalisation, antenna beamforming, ANC for pilots’ helmets), and kept the LMS rule alive through what would otherwise have been a decade of silence. The neural-network revival of the late 1980s inherited a working algorithm that had been quietly running in production telecoms hardware the whole time.

Where it lives today. The list is long. Active noise cancellation in headphones (predict the noise from the outside microphone, play its inverse into the ear cup). Hearing aids (subtract whistling feedback from the microphone’s input). Modem and DSL line equalisation (cancel the channel’s frequency response). MIMO wireless beamforming. Even early-generation acoustic echo cancellation in speakerphones and Zoom is a direct intellectual descendant of the AT&T system — same LMS update, same sliding-window FIR, just running on a CPU instead of memistors. ADALINE didn’t disappear into deep learning alone. It also slipped sideways into DSP and never left.

What survived

If you squint at ADALINE’s update rule, every modern neural network is already inside it.

The update shape is identical. In PyTorch, with whatever architecture you like, training boils down to:

θ  ←  θ  −  η⋅∇θ L\theta \;\leftarrow\; \theta \;-\; \eta \cdot \nabla_\theta \, \mathcal{L}

That is exactly what ADALINE is doing, one weight at a time: take the gradient of the loss with respect to a parameter, multiply by a learning rate, subtract. The only difference between Widrow’s 1960 paper and the .step() method on a modern optimiser is that the loss is more elaborate and the parameter vector is a billion times longer. The move is the same.

The “smooth loss, discrete decision” pattern stuck. Every classifier you train today follows the split ADALINE introduced: a continuous loss for training, a discrete prediction at inference. Cross-entropy on softmax probabilities replaces squared error on a sign function, but the structure is identical — differentiable for the training loop, thresholded at the moment of prediction. LMS wasn’t just a learning rule; it was the demonstration that you need this separation if you want to do gradient descent at all.

Per-example updates started here. ADALINE updated its weights after every single pattern — what we now call online SGD, or batch size =1= 1. Larger mini-batches came later as a computational optimisation; the algorithmic essence is the same single-example update Widrow wrote down in 1960.

What had to be added later:

  • Hidden layers. ADALINE was one layer. Stacking neurons and pushing gradients back through them required backpropagation — the chain rule applied repeatedly down the stack — which Rumelhart, Hinton and Williams formalised in 1986. ADALINE can only draw one straight line through input space; backprop is what lets you draw curves.

  • Non-linear activations. ADALINE used sgn\mathrm{sgn} as its decision — a step function with no gradient. Replacing that step with a smooth non-linear activation (sigmoid, tanh, later ReLU) is what turns a stack of linear layers into a genuinely non-linear function approximator. Without it, a ten-layer network collapses mathematically to a single layer.

  • Automatic differentiation. Widrow derived LMS by hand — one paper, one rule. Every modern framework derives the equivalent rule automatically for whatever model you define. The math is identical; the human labour is completely different.

Also absent in ADALINE and added later: regularisation, batch normalisation, dropout, principled weight initialisation, learning-rate schedules, adaptive optimisers like Adam, attention, convolution, transformers, GPUs. All of these are elaborations on top of a core that ADALINE already had: a parametric model whose parameters adjust by following the gradient of a continuous loss toward a target.

The first trainable neural network was sixteen knobs, a meter, and a rule for turning the knobs. Everything since is an elaboration of those three parts.