In our article on weight matrices, we moved between two pictures of a neural network: neurons connected to other neurons, and a matrix transforming a vector. Once you understand the calculation, that seems like a small step. Historically, it required a substantial change in what researchers thought they were studying.

You can see how far that change had gone by opening Frank Rosenblatt’s Principles of Neurodynamics, published in 1962. On page 83, he defines an interaction matrix of connections between units. Nearby, the network’s adjustable parameters become coordinates in a Euclidean space. Learning can be described as movement through that space.

This is decades before modern deep learning software. The matrix is already there.

What makes the story interesting is how it got there, and why it eventually became the ordinary way to teach and program a neural network. The early researchers were trying to explain computation, recognition, and memory. Each problem asked something different of the model, and each made another part of the mathematics useful.

The trouble with a perfectly specified circuit

A neuron fires or it doesn’t. That observation gave Warren McCulloch and Walter Pitts a starting point for their 1943 paper. If neural activity could be idealized as an all-or-none event, perhaps networks of neurons could be studied through logic.

Their simplified units had thresholds and delays. Some inputs excited a unit; inhibitory inputs could prevent it from firing. With the network’s structure fixed, they could investigate how its activity expressed logical relationships, including relationships between events at different times.

The attraction of this approach is easy to see. Instead of confronting all the biological detail at once, you could specify a circuit and reason about what it computed. A drawing of connected cells became a mathematical object.

But specifying the circuit leaves out much of what makes learning interesting. An animal acquires behavior through experience. Something about the system must be able to change.

Donald Hebb’s The Organization of Behavior, published in 1949, placed that change in the effectiveness of connections. When one cell repeatedly helped cause another to fire, he proposed, a lasting change could make that influence stronger. This was a biological hypothesis, expressed in words. It gave later mathematical models a concrete place to put the effects of experience.

Rosenblatt wanted a quantitative account of how a system with partly random wiring could learn. In his 1958 perceptron paper, he pointed to a difficulty with the logical approach: the broad organization of a nervous system might be known while its precise connections remained unknown. He therefore chose probability theory as the basis for his model.

That choice changes the question. Given a precisely specified circuit, we can ask what it computes. Given a statistically described population of units, we can ask under what conditions it will acquire useful behavior.

Rosenblatt’s perceptron contained sensory, association, and response units, with randomness in parts of their connectivity. Learning altered how association activity influenced responses. It was a much richer proposal than the single threshold neuron that usually introduces perceptrons today.

Now there were patterns of activity to distinguish, connections to adjust, and learning behavior to predict. Logic had made neural computation tractable. Probability allowed Rosenblatt to investigate it without knowing every wire.

When a connection became a coordinate

Once connection strengths are adjustable numbers, a useful possibility appears: every possible setting of those numbers can be treated as a location. Learning becomes a path through the possible settings.

Bernard Widrow and Marcian Hoff made this view concrete in their 1960 work on ADALINE. Their adaptive device had coefficients that could be changed to reduce its errors. In their analysis, inputs were vectors and mean squared error formed a surface over the coefficient settings.

They measured that error before the final threshold converted the signal into a discrete decision. This gave them a numerical surface along which adaptation could proceed toward a minimum. The same coefficient was a setting in a device and a coordinate in an optimization problem. Our article on ADALINE follows the learning rule itself.

This brings us back to Rosenblatt’s book. His interaction matrix collected the coupling coefficients between units, using zero where no connection existed. His parameter space described the possible states of the network’s memory. He went on to use rank and singularity to analyze which classifications perceptrons could realize.

The matrix was doing real explanatory work. A network could be examined as a whole, and properties of that whole could tell you something that was difficult to see by inspecting its connections individually.

That power also worked against optimistic expectations. In Perceptrons, published in 1969, Marvin Minsky and Seymour Papert investigated systems that gathered partial information through feature detectors and combined their outputs in a weighted decision. Their subtitle, An Introduction to Computational Geometry, tells you how central the mathematical viewpoint had become.

Take parity: deciding whether an image contains an odd or even number of active pixels. They proved that a perceptron combining feature detectors through a single weighted threshold could not solve this for every input unless at least one detector depended on the entire image. If every detector saw only part of the image, no adjustment of the final weights could make the system succeed. The limitation came from the architecture and its features.

Learning, then, was more complicated than finding good connection strengths. The representation presented to those connections also mattered. If it did not support the required distinction, the model needed a different way to represent the input.

A matrix that could remember

Recognition was only one reason to care about changing connections. Memory posed another problem: how could an experience leave a trace that a later cue could recover?

Karl Steinbuch had approached this problem in Die Lernmatrix, published in 1961. The name means learning matrix. His proposed circuit had a learning phase, in which signals representing properties appeared alongside their associated meaning, and a recall phase, in which presenting one could recover the other. He proposed implementations using magnetic storage elements or electrochemical processes, and envisaged applications such as character recognition and information retrieval.

This made memory an engineering problem: how to change an array during learning so that signals passing through it later recovered an association.

Teuvo Kohonen took up that problem in Correlation Matrix Memories, published in 1972. He explicitly replaced Steinbuch’s switching matrix with a correlation matrix, studying the mathematics of associative recall while leaving its possible biological role outside the paper.

In the version that stored every pairwise product, each association paired a key vector with a data vector. Multiplying every data component by every key component formed an outer product: a matrix of pairwise products. The memory stored a scaled sum of these matrices. Applying it to a key then produced a recalled pattern.

The same mathematics explained why recall could go wrong. Similar keys could bring back unwanted contributions from other associations. Memories stored together could interfere with one another.

James Anderson’s work published that year approached the same product-based construction through local synaptic changes. A connection changed in proportion to the product of the activity in the sending cell and the activity in the receiving cell. With connections between every pair of cells in the two groups, these changes formed the same kind of outer product. Here a numerical Hebbian rule became a matrix update: local changes accumulated into a store from which an associated pattern could later be retrieved.

John Hopfield gave the memory problem a particularly recognizable form in his 1982 paper: imagine recovering a complete bibliographic reference from a fragment, perhaps even a fragment containing a misspelling.

His network fed activity back through recurrent connections. An incomplete pattern could evolve toward a stored one. The coupling matrix described how the units influenced one another, while a separate vector described their current states.

With symmetric connections, no self-connections, and asynchronous updates, Hopfield could show that an energy function would not increase as the states changed. That connected neural recall to the behavior of interacting physical systems. Stored patterns could act as attractors: states toward which nearby activity moved.

During recall, the connections stayed fixed. The activity settled. The coupling matrix and the nonlinear updates together explained how a partial cue could lead to a more complete pattern.

We have come a long way from collecting connection strengths in a table. Matrices could help explain how experience was stored, why memories interfered, and how a network recovered from an imperfect cue.

Learning what the next layer should see

The recognition problem still contained the difficulty exposed by perceptron analysis: useful decisions depended on useful features. Could the network acquire those features through experience as well?

One revealing result came from Erkki Oja in 1982. He studied a linear neuron with a Hebbian learning rule modified to control the growth of its weights. Under the conditions of his analysis, the weights approached a dominant eigenvector of the input’s second-moment matrix. For zero-mean inputs, that direction is the first principal component.

A rule expressed as local connection changes could therefore extract a known statistical structure. Linear algebra helped identify what the neuron was learning.

For a network with several stages, there was an additional challenge. Training examples could specify the answer, but they generally did not specify what each intermediate unit ought to detect. The desired behavior at the output had to guide changes deeper inside the network.

David Rumelhart, Geoffrey Hinton, and Ronald Williams made that possibility vivid in Learning Representations by Back-Propagating Errors, published in 1986. Their examples adjusted hidden connections as well as output connections. In a network trained to recognize mirror symmetry, they examined the learned weights to understand how the intermediate units contributed to the solution.

The derivative machinery had earlier roots in Seppo Linnainmaa’s work on reverse accumulation, originating in his 1970 thesis, and in Paul Werbos’s 1974 thesis, described in his later account. The 1986 demonstration made clear why this machinery mattered for neural learning: earlier stages could learn to supply the features that later stages needed.

Our article on backpropagation explains the calculation. For this story, the important change is what became adjustable. Training could shape the transformations that produced the internal representation, as well as the final decision made from it.

The picture of connected neurons still described the network. The picture of successive learned transformations described what the network was accomplishing.

How the diagram became a program

By this point, vectors and matrices could describe connections, memories, statistical features, and transformations between layers. Someone entering the subject needed a way to understand those descriptions together.

Michael I. Jordan supplied one in the first volume of Parallel Distributed Processing, published in 1986. His contribution was chapter 9, a tutorial on the linear algebra used to analyze the volume’s models.

He taught vector spaces, inner products, and linearity, showing how simple connectionist models corresponded to operations on vectors. The reader learned a mathematical language that could be carried from one model to another.

This gives us a concrete point at which to see the connection becoming part of the curriculum. Researchers had been using matrices for decades. Now a major presentation of connectionist models included an introduction designed to teach readers how to think in those terms.

That educational shift is visible again in Christopher Bishop’s Neural Networks for Pattern Recognition, published in 1995. In the foreword, Hinton presented Bishop’s book as filling a need for a clear account built on linear algebra, calculus, and probability.

Bishop built from simpler statistical models toward multilayer networks. A newcomer could learn neural networks as part of a mathematical treatment of estimation, optimization, and generalization. The journey from a biological analogy to numerical models was increasingly built into the way the subject was taught.

Software made that language practical. Ian Nabney’s NETLAB toolbox put the matrix description directly into code. Its forward-pass routine takes a matrix of input examples. A matrix multiplication, bias addition, and nonlinear activation compute the hidden layer. Another matrix multiplication takes its activity to the output stage.

Each row holds an example, so the same operations process a whole batch through both layers. The connections in the diagram have become the coefficients in the program.

For a student learning from these books and tools, the two descriptions arrive together. There is no separate moment of discovering that a layer can be a matrix: it is already how the layer is explained and implemented.

GPUs subsequently made the organization of those operations even more consequential. In their 2009 work on large-scale unsupervised learning, Rajat Raina, Anand Madhavan, and Andrew Ng expressed computations for deep belief networks through matrix operations that parallel numerical routines could execute efficiently. Batching work and limiting transfers between CPU and GPU memory were essential to obtaining the benefit. Our CPU and GPU comparison explores that practical issue.

The algebra also offered a way to discuss architectural choices. In the 2017 Transformer paper, the authors grouped queries, keys, and values into matrices and pointed to optimized matrix multiplication as a speed and memory advantage of dot-product attention. Some of the matrices now described interactions computed from the current input, extending the vocabulary beyond fixed connection strengths.

The two pictures we started with became familiar together because researchers found so much they could do by moving between them. Treating neural activity as numbers and connections as coefficients gave weighted summation an exact algebraic description, alongside the network’s nonlinear operations. That description helped explain learning and memory, then became the language of textbooks and software. When we study the subspaces of a weight matrix today, we inherit that history: a way of seeing a network that became natural through use.