Understanding CNNs — convolutions, feature maps, and pooling
In the previous articles, we trained a neural network on MNIST using dense layers — where every neuron in one layer connects to every neuron in the next, with weights updated via backpropagation. To work with dense layers, we had to flatten each 28×28 image into a 784-element vector and feed it into a layer of 128 neurons. That first layer alone had weights + biases = parameters — every input pixel connects to every neuron. It worked — we hit 97% accuracy — but this approach has two fundamental problems:
1. No built-in spatial structure. Flattening preserves every pixel and its order; we can reshape the vector back into the image. But a dense layer does not encode which pixels are neighbors or reuse a detector across positions. A shifted pattern reaches different weights, so learning it in one location does not automatically teach the network to recognize it elsewhere.
2. Too many parameters. With a 28×28 image, 100k parameters is manageable. But real images are much larger. A 224×224 RGB image (a common input size) has pixels color channels values. A single dense layer with 512 neurons would need million parameters — just for one layer. That’s expensive to train and easy to overfit.
How can parameter count affect overfitting?
Convolutional Neural Networks (CNNs) solve both problems at once. Instead of looking at all pixels at once, a CNN focuses on small local regions — like a 3×3 patch — and learns to recognize patterns within them. The fundamental difference: dense layers learn global patterns involving all pixels at once, whereas convolutional layers learn local patterns — and reuse the same detector at every position across the image.
Consider distinguishing squares from circles on a 28×28 grid. Local patches contain useful clues: a square has sharp corners and straight sides, while a circle has curved edges. The diagram illustrates candidate features; training determines which responses the model actually uses.
Using CNNs allows us to address each problem directly:
Convolution provides local connectivity and weight sharing. A 3×3 filter over one input channel has 9 weights plus a bias. The same weights are applied at every position, so the layer can respond to a learned pattern wherever it appears.
CNNs are used for image classification, object detection, and text recognition. The same sliding-filter operation also extends to one-dimensional signals, such as audio, and three-dimensional inputs, such as volumes.
Convolution in neural networks
Our CNN combines convolution, ReLU activation, pooling, and a dense classifier. Convolution computes local responses; ReLU makes the transformation nonlinear; pooling reduces the spatial dimensions. Here is the architecture in Keras notation:
model = keras.Sequential([
Input(...), # input image
Conv2D(...), # convolution
MaxPooling2D(...), # pooling
Conv2D(...), # convolution
MaxPooling2D(...), # pooling
Flatten(...), # flatten to 1D
Dense(...), # classification
])Convolution detects patterns, pooling compresses the result, and a dense layer makes the final classification. Here’s what the complete CNN pipeline looks like — from pixels to prediction:
The convolution and pooling layers produce features that the dense layer combines into a prediction. Training adjusts both the feature extractor and the classifier together.
The convolution step detects local patterns — it slides learned filters across the image and produces feature maps that highlight where each pattern appears. It doesn’t classify anything — it transforms the raw pixel grid into a richer representation. This is feature extraction, the same idea as feature engineering in traditional ML. Instead of feeding raw data into a classifier, you first transform it into more useful representations. The difference is that CNNs learn which features to extract automatically — the kernel values are discovered during training, not designed by hand.
The pooling step compresses each feature map spatially — it shrinks the width and height while keeping the strongest signals. This reduces the number of values the dense layer needs to process (from 28×28×2 = 1,568 down to 14×14×2 = 392 in our model) and makes the features more robust to small shifts in position.
We’ll implement these layers in NumPy and inspect the filters learned on a small shape-classification task.
Convolution as image transformation
Convolution is not an ML invention — it’s a fundamental technique from signal processing and image processing that existed long before deep learning. If you’ve ever applied a blur, sharpen, or edge detection filter in image editing software — that’s convolution. Under the hood, each filter is a small grid of numbers (called a kernel) that slides across the image. At each position, every kernel value is multiplied by the overlapping pixel, the products are summed, and the result is written into the output.
Switch between the filters below and compare the 3×3 kernel values with the output on the right:
Notice how the horizontal edge detector highlights the top and bottom edges of the shape — this is achieved by having negative values in the top row and positive in the bottom row of the kernel, so it responds where brightness changes vertically. The sharpen filter has a large positive center value and negative neighbors — it amplifies the difference between a pixel and its surroundings, making edges crisper.
These are classical, well-established kernels from image processing — hand-designed decades before deep learning existed. The description below explains what each kernel computes and why the output looks the way it does.
| Filter | How it works |
|---|---|
| Blur | Averages all 9 pixels equally — smooths out differences. Uniform area stays the same, but sharp transitions get diluted |
| Sharpen | Amplifies the center pixel relative to its neighbors. In smooth areas the neighbors cancel out; at edges the mismatch gets exaggerated |
| Horizontal edges | Subtracts top pixels from bottom — if they’re similar, the result is near zero. Large output only where brightness changes vertically, i.e. a horizontal edge |
| Vertical edges | Same idea rotated: subtracts left from right. Output is strong only where brightness varies horizontally, i.e. a vertical edge |
Same sliding operation, same input — different results. The 9 numbers in the kernel completely determine what the output looks like. For an excellent visual explanation of how convolutions work, see this video by 3Blue1Brown.
Convolution as pattern matching
A sliding window can also be used for pattern matching. Select a small fragment of the image and compare it with patches elsewhere. The next demo normalizes those comparisons to focus on the pattern rather than its overall brightness.
Click on any part of the shape below to cut out a 3×3 patch and use it as the kernel. For example, click on a vertical edge of the square to see it highlight all vertical edges, or click on the circle’s rounding to see where similar curves appear. The widget slides the kernel across the entire image, and the output lights up where similar patterns are found. You can also click Play to see the sliding mechanics step by step:
This demo uses cosine similarity: the dot product divided by the lengths of the kernel and patch vectors. It defines the score as zero when either vector is zero, and darkens scores below 0.85 for display. A CNN convolution does not perform this normalization: it computes a weighted sum plus a bias. The demo illustrates searching with a sliding patch, but its scores are not convolution outputs.
From hand-designed to learned filters
Both uses of convolution — image transformation and pattern matching — rely on the kernel values being chosen well. Image processing has used hand-designed kernels for decades (the blur, sharpen, and Sobel filters in the demo above). They’re useful, but choosing them by hand is tedious and limited.
The core idea of CNNs is: instead of hand-designing these values, let the network learn them automatically for the task at hand. The filter starts with random numbers and gets updated by gradient descent during training — the same way weights in dense layers are learned. The network discovers which patterns matter for the task. For our square vs circle task, it might learn filters that respond differently to straight edges vs curved edges — whatever helps separate the two classes.
The convolution operation
We’ve seen above how convolutions can help with image transformation and pattern matching. Now let’s look at the actual mechanics: what exactly happens at each position, and how does the output — called a feature map — get built up?
Why "feature map"?
The name uses “feature” in the same sense as in ML generally — a measurable property useful for the task. Just as a tabular model might use features like “age” or “income”, a CNN’s feature is a visual pattern like “vertical edge” or “curve”. The feature map is a spatial grid showing where that feature was detected across the image — high values where the pattern is present, low where it isn’t.
The core operation is straightforward: take a small matrix (the kernel), slide it across the input, and at each position compute the dot product — multiply every kernel value by the overlapping pixel, sum the products — and write the result into the feature map. The kernel can be any size (1×1, 5×5, 7×7) — older networks like AlexNet used 11×11 kernels — but 3×3 is the most common choice in modern architectures and what we’ll use throughout this article.
Let’s see it in practice. Say you click on the left edge of the square in the demo above — the kernel becomes a 3×3 patch with black (0) as background and white (255) where the edge line is. When this kernel slides to another left edge, the image patch looks the same — every position matches, so the products are large:
Now let’s apply the same kernel to the top edge. The white pixels are in different positions — the kernel’s white column overlaps mostly with black pixels:
A large dot product can come from aligned values, large magnitudes, or both. It is not a normalized measure of similarity: an all-white patch gets the same dot product as the matching edge here, because pixels under zero kernel weights do not affect the sum. Learned filters can also use negative weights to penalize unwanted structure.
The same math explains blur too. The blur kernel has every value set to ⅑, so the weighted sum is just the average of the 9 pixels. On a left edge patch:
At this one position, the output pixel is ~85 — a gray value between black (0) and white (255). The kernel keeps moving, processing the image position by position, row by row. Every position near an edge gets a similar mixed value, softening the sharp 0-to-255 jump into a gradual transition. On a flat region (all black or all white), every pixel in the patch is the same, so the average equals the original — nothing changes there.
Output size
When a filter slides across the input without any padding, the output becomes smaller than the input — this is called the border effect. Consider a 5×5 input with a 3×3 kernel. The kernel starts at the top-left corner, covering rows 0-2 and columns 0-2. It slides right one pixel at a time — but it can only go to column 0, 1, or 2. Starting at column 3 would make the kernel cover columns 3-5, and column 5 doesn’t exist. Same vertically. So there are only 3 valid positions per axis, producing a 3×3 output from a 5×5 input:
For one dimension, with input size , kernel size , padding on each side, stride , and no dilation, the output size is . Here and , giving .
Padding
To counteract the border effect, you can add rows and columns of zeros around the input — this is called padding. For a 3×3 kernel, you add 1 pixel of zeros on each side, so a 28×28 input becomes a 30×30 padded input, and the output is back to 28×28. For a 5×5 kernel, you’d add 2 pixels on each side.
We add zeros, and not say 1, because zero times any kernel value is zero — the padded pixels contribute nothing to the dot product. Only the real pixels that overlap with the kernel affect the result. It’s like saying “there’s nothing outside the image boundary.”
Stride
So far we’ve assumed the kernel moves one pixel at a time — stride 1. But the distance between successive kernel positions is a parameter you can change. For example, with stride 2, the kernel jumps 2 pixels at a time, skipping every other position — both horizontally and vertically. A larger stride means fewer positions, which means a smaller output.
With stride 2, the kernel moves two pixels at a time. In this 5×5 example with a 3×3 kernel and no padding, the output becomes 2×2 instead of 3×3.
Downsampling reduces the number of activations and the computation in later layers. It also reduces the parameter count of a following flattened dense layer. For a later convolution with fixed kernel size and channel counts, the number of parameters stays the same.
Our model downsamples with max pooling, a separate operation that takes the maximum in each local window. We’ll implement it below.
Inside a convolutional layer
So far we’ve looked at convolution as a standalone operation — one kernel sliding across an image, producing one feature map. Now let’s look at how this operation is packaged into a layer that the network can train. In Keras, a convolutional layer is defined like this:
layers.Conv2D(2, kernel_size=3, activation='relu', padding='same')The learned parameters are the kernel weights and biases. Feature maps are activations computed for each input, not learned parameters. During training, the implementation also caches inputs and activations needed for backpropagation.
Both dense and convolutional layers take the entire input, but process it differently. Take a simplified 3×3 grid with a 2×2 kernel. A dense neuron has a unique weight per pixel (9 weights for our simple 3x3 grid, or 784 weights for a 28×28 image), computes one weighted sum, and produces one value.
A conv kernel has only a few shared weights, but slides across every position in the image, producing one value per position — an entire feature map. At each position, the layer extracts a local patch from the input, multiplies it element-wise with the kernel, sums all the products, and adds a bias.
The dense neuron sees all 9 pixels at once with 9 unique weights, producing one output. The conv kernel sees 4 pixels at a time (2×2), but visits every position — same 4 weights, reused everywhere. Fewer parameters, but the full image is still covered.
There’s a useful way to read that sliding picture. Each position the kernel visits produces one output value — and you can think of each position as its own neuron. The kernel at the top-left corner is one neuron, looking only at the top-left patch; the kernel one step to the right is another neuron, looking at a patch shifted by one pixel; and so on. For the 3×3 input and 2×2 kernel above, the kernel lands in 4 positions — so this tiny layer is effectively 4 neurons, and their outputs assemble into the 2×2 feature map. Scale up and the count grows with the image: a 28×28 input with a 3×3 kernel becomes a layer of 26×26 = 676 neurons.
The small patch each neuron looks at has a name: its receptive field — its own region of interest in the input. The top-left neuron’s receptive field is the top-left 2×2 corner, and it responds only to those pixels, ignoring the rest of the image entirely. Every neuron has a different receptive field, and together they tile the whole input.
A dense neuron has its own weights over the entire input. A convolution output uses a local receptive field and shares weights with other positions in the same output channel. The four positions above are the same filter evaluated in four places; other output channels use their own filters.
Weight sharing makes stride-1 convolution translation equivariant, away from boundary effects: shifting the input shifts the feature map. That does not make the final classifier invariant. Pooling can reduce sensitivity to some small shifts, but crossing a pooling-window boundary can change the result.
And once layers stack, a neuron’s receptive field grows — a second-layer neuron reads a patch of first-layer outputs, each of which already summarized a patch of pixels, so it indirectly sees a wider region of the original image. That’s the “wider field of view” we’ll return to with stride and depth.
Here is a NumPy implementation for a single input channel, stride 1, and odd-sized kernels. It stores output as (channels, height, width). Keras uses (height, width, channels) by default. Start with import numpy as np:
import numpy as np
class Conv2D:
def __init__(self, num_filters, kernel_size, padding='same'):
if padding not in ('same', 'valid') or kernel_size % 2 == 0:
raise ValueError("Use an odd kernel size and 'same' or 'valid' padding")
self.kernels = np.random.randn(num_filters, kernel_size, kernel_size) * 0.1
self.biases = np.zeros(num_filters)
self.padding = (kernel_size - 1) // 2 if padding == 'same' else 0
def forward(self, input):
# Pad input with zeros if padding='same'
if self.padding > 0:
p = self.padding
input = np.pad(input, ((p, p), (p, p)), mode='constant')
self.input = input # save padded input for backward
h, w = input.shape
k = self.kernels.shape[1]
out_h, out_w = h - k + 1, w - k + 1
self.z = np.zeros((self.kernels.shape[0], out_h, out_w))
for f in range(len(self.kernels)): # each filter
for r in range(out_h): # each row
for c in range(out_w): # each column
patch = input[r:r+k, c:c+k] # extract local patch
self.z[f, r, c] = np.sum(patch * self.kernels[f]) + self.biases[f]
self.out = np.maximum(0, self.z) # ReLU activation
return self.out # shape: (num_filters, out_h, out_w)The weighted sum at each position, plus a bias term — that’s the entire operation. The bias (one per kernel) shifts the output up or down, giving the kernel a threshold for how strongly a pattern needs to match before it activates. It works exactly like a bias in a dense layer.
How the kernels learn: backpropagation in CNNs
Backpropagation computes gradients; the optimizer uses them to update weights. For one sample, a dense weight contributes to one output calculation. A convolution weight is reused at every spatial position, so its gradient sums contributions from all those positions. Both layers also accumulate gradients across samples in a mini-batch.
# For one sample and one dense weight:
dw = input_value * output_gradient
# For one convolution weight (kr, kc), sum over output positions:
dw = sum(
padded_input[r + kr, c + kc] * gradient[r, c]
for r in range(out_h)
for c in range(out_w)
)Here gradient[r, c] is the derivative of the loss with respect to the pre-activation at that output position, after applying the ReLU derivative. Each term is an input value multiplied by the corresponding output gradient.
Here’s how that accumulation looks in our Conv2D class:
class Conv2D:
# ... __init__ and forward from above ...
def backward(self, upstream_gradient):
# ReLU backward: zero gradient where activation was ≤ 0
grad = upstream_gradient * (self.z > 0)
k = self.kernels.shape[1]
out_h, out_w = grad.shape[1], grad.shape[2]
self.grad_biases = grad.sum(axis=(1, 2))
self.grad_kernels = np.zeros_like(self.kernels)
for f in range(len(self.kernels)):
for kr in range(k):
for kc in range(k):
for r in range(out_h): # sum over every position
for c in range(out_w):
self.grad_kernels[f][kr][kc] += self.input[r+kr][c+kc] * grad[f][r][c]
def update(self, lr):
self.kernels -= lr * self.grad_kernels
self.biases -= lr * self.grad_biasesThe bias gradient is the sum of grad over spatial positions. This first-layer implementation only computes parameter gradients; stacking trainable convolution layers would also require returning a gradient with respect to the input.
Multiple kernels, multiple feature maps
One filter produces one feature map: a map of responses to its weights. Different filters can respond to different structures, though a filter need not correspond to one neatly named pattern.
This necessitates using multiple kernels per layer to produce multiple feature maps — one per kernel, each detecting a different pattern. So each convolutional layer produces as many feature maps as it has kernels — 2 in our case, but this is intentionally minimal.
Using more filters increases the number of output channels and the model’s capacity. Training may produce complementary filters, but it can also produce redundant or inactive ones.
Here’s what we got after training our 2-filter CNN — the kernel values it converged on produce visibly different feature maps for squares vs circles. These are the feature maps from our single convolutional layer — in a deeper network, each layer would produce its own set of feature maps, with deeper layers capturing increasingly complex patterns:
These maps use the learned weights and biases from the training run below, with input pixels normalized to [0, 1] and ReLU applied. Each map’s brightness is scaled independently, so compare the spatial patterns, not absolute brightness across maps.
The dense layer receives the pooled responses from both maps. These responses can help separate the two classes without either filter being a dedicated “square detector” or “circle detector.”
Stacking layers: from edges to shapes
One convolutional layer with two filters is sufficient for this generated dataset. Stacking layers lets later filters combine earlier responses over larger receptive fields, which can help with more varied images.
The real power of CNNs comes from stacking convolutional layers. To see why, think about how you’d decompose a recognition problem by hand. If someone asked you to detect the digit “0”, you might break it into sub-problems: is there a top curve? A bottom curve? A left edge? A right edge? Each sub-network answers one question, and a final layer combines their outputs:
Each of those sub-problems can be decomposed further. “Top curve?” breaks down into: is there a top-left arc? A top-right arc? A horizontal edge at the peak? Each question gets simpler, closer to raw visual features:
Edges → curves → parts is a useful illustration of increasing receptive fields, not a guaranteed description of what each layer learns. Actual features depend on the data, architecture, and training.
And at each level, multiple kernels work in parallel — one kernel might learn to detect horizontal edges while another detects vertical edges, one might find right-angle corners while another finds smooth curves. This is why each layer has many kernels — the network needs to track many different patterns simultaneously at each stage of the hierarchy.
Multiple input channels
Everything above describes the first convolutional layer in the hierarchy, where each kernel receives a single 2D image as input. But deeper layers in the stack work with something different. The first convolutional layer takes a grayscale image — a single 2D grid 28×28. But the second convolutional layer’s input isn’t a flat 2D image anymore — it’s the output of the first layer, a stack of 2D feature maps, which makes it a 3D volume (28×28×2 in our case, or 28×28×32 in a typical network):
Each output feature map becomes an input channel for the next layer. Channels are parallel components at the same spatial locations. Their order must stay consistent with the next layer’s weights; swapping channels without swapping the corresponding weights changes the computation.
Since the input now has multiple stacked feature maps (multiple channels), the kernel needs to process them all. The kernel automatically mirrors the input structure — its depth (number of slices) is inferred to match the number of input channels. Each layer still can — and usually does — have multiple kernels, but now each kernel becomes 3D — its depth automatically matches the number of input channels.
With 2 input channels, each kernel becomes a stack of two 2×2 matrices, one per channel. At each position, the first matrix multiplies with channel 1, the second with channel 2, and all products get summed into a single one output value per position. Sliding this 3D kernel across all positions produces just one feature map as the output of one kernel — one value per position.
The diagram below shows two positions to illustrate how the same kernel produces different values that all contribute to that one feature map:
The operation is the same weighted sum we already know — just run once per channel, then sum the results. The diagram shows a 2×2 kernel sliding over a 3×3 input with 2 channels. At each position, we compute the weighted sum for each channel separately, then add the results:
For an ordinary convolution, each output filter spans all input channels. Its parameter count is kernel_height × kernel_width × input_channels + 1 for the bias. Depth alone does not make a kernel larger; changing the number of input channels does.
Pooling
After a convolutional layer detects features, we want to shrink each feature map — like downscaling an image, but instead of averaging pixels, we keep only the strongest signals. Each feature map is shrunk independently — pooling doesn’t merge channels. The depth stays the same: 28×28×2 becomes 14×14×2, not 14×14×1.
Pooling has no learned weights. It summarizes a local window independently in each channel. Max pooling keeps the largest value; it can tolerate some shifts within a window, but is not generally translation invariant.
There are different pooling strategies — max, average, and others — but the most common is max pooling: take a 2×2 window and slide it across the feature map with stride 2 — since the stride matches the window size, each window covers a fresh region with no overlap. At each position, keep only the maximum value.
Pooling might sound like strided convolution — both slide a window and downsample. But they differ in two important ways. First, nothing is learned — pooling has no weights, no kernel, no parameters that get updated during training. Second, the operation itself is different — instead of a weighted sum (multiply-and-sum), pooling simply picks the max.
For our model, pooling takes the 28×28 feature map down to 14×14:
If a response moves within the same pooling window and remains its maximum, that pooled value stays the same. Move it across a window boundary, and the output can change. Pooling offers limited shift tolerance, not a guarantee that a one-pixel shift leaves predictions unchanged.
Average pooling takes the mean of each window instead of the maximum. Global average pooling averages an entire feature map into one value per channel. A classifier can then use those values, often with a final dense layer, instead of flattening every spatial position.
Our MaxPool2D uses non-overlapping windows and drops trailing rows or columns that do not form a complete window. During backpropagation, each gradient goes to the window’s maximum; ties go to the first maximum. There are no parameters to update:
class MaxPool2D:
def __init__(self, size=2):
self.size = size
def forward(self, x):
self.input = x # save for backward
s = self.size
channels, h, w = x.shape
out = np.zeros((channels, h // s, w // s))
for c in range(channels):
for r in range(0, h - s + 1, s):
for col in range(0, w - s + 1, s):
out[c, r//s, col//s] = np.max(x[c, r:r+s, col:col+s])
self.output_shape = out.shape
return out
def backward(self, gradient):
s = self.size
out = np.zeros_like(self.input)
channels, h, w = self.input.shape
for c in range(channels):
for r in range(0, h - s + 1, s):
for col in range(0, w - s + 1, s):
patch = self.input[c, r:r+s, col:col+s]
max_idx = np.unravel_index(np.argmax(patch), patch.shape)
out[c, r+max_idx[0], col+max_idx[1]] = gradient[c, r//s, col//s]
return out
# No update() — pooling has no parametersFlatten and classify
Pooling processes channels independently; convolution can combine them. For this classifier, we flatten all pooled activations into one vector so the dense output connects to every spatial position in every channel. Keras Dense can accept higher-rank inputs, but then it operates on the last axis rather than flattening the input automatically.
Our NumPy arrays are channels-first: flattening reads the 196 values of channel 1, then the 196 values of channel 2. Keras defaults to channels-last, so its flattened ordering differs. Either convention works when the forward and backward passes use the same layout.
This is the same reshape we did in the dense-only MNIST model when we flattened 28×28 into 784. Flatten itself doesn’t know or care what it’s reshaping — the operation is identical. The difference is what the layers before it have done: in the dense-only approach, there were no preceding layers, so Flatten received raw pixels. Here, convolution and pooling have already extracted and compressed spatial patterns, so the dense layer receives meaningful features rather than raw pixel values.
Then the Dense layer acts as the classifier. Since we only have two classes (square vs circle), this is binary classification — we need just one neuron with sigmoid activation. It connects to all 392 values, multiplies by learned weights, sums, adds a bias, and outputs a single number between 0 and 1: the probability that the input is a circle. Close to 0 = square, close to 1 = circle. For multi-class problems like MNIST (10 digits), you’d use 10 neurons with softmax instead — one per class.
Here’s the DenseLayer class — the same structure as the HiddenLayer from the previous article:
class DenseLayer:
def __init__(self, n_inputs, n_outputs):
self.W = np.random.randn(n_outputs, n_inputs) * np.sqrt(2.0 / n_inputs)
self.b = np.zeros(n_outputs)
def forward(self, x):
self.x = x
return self.W @ x + self.b
def backward(self, grad_output):
self.grad_W = np.outer(grad_output, self.x)
self.grad_b = grad_output
return self.W.T @ grad_output
def update(self, lr):
self.W -= lr * self.grad_W
self.b -= lr * self.grad_bBuilding a CNN model for shapes
Now that we’ve covered each component, let’s build the model and training loop — the same way we did for the dense MNIST model, but now with convolutional layers.
We instantiate the layers we built throughout this article:
# 28×28 grayscale → 2 feature maps → pooling → flatten → 1 output
conv = Conv2D(num_filters=2, kernel_size=3) # 20 parameters (2×9 weights + 2 biases)
pool = MaxPool2D(size=2) # 0 parameters
dense = DenseLayer(n_inputs=392, n_outputs=1) # 393 parameters (392 weights + 1 bias)Preparing the data
We generate 4,000 images: 2,000 squares and 2,000 circles. Both classes use the same size range, border margin, line thickness, and edge smoothing rule, so rendering differences are less likely to provide shortcuts. This remains a deliberately simple synthetic task.
np.random.seed(42)
SIZE = 28
LINE_WIDTH = 2.0
rr, cc = np.mgrid[0:SIZE, 0:SIZE] + 0.5 # grid of pixel centers
def make_shape(kind):
radius = np.random.uniform(4.0, 10.0)
margin = radius + 2.0
cy, cx = np.random.uniform(margin, SIZE - margin, size=2)
dy, dx = np.abs(rr - cy), np.abs(cc - cx)
distance = np.maximum(dy, dx) if kind == 'square' else np.hypot(dy, dx)
edge_distance = np.abs(distance - radius)
coverage = np.clip((LINE_WIDTH / 2 + 0.5 - edge_distance) / 0.5, 0, 1)
return (255 * coverage).astype(np.float32)
N = 2000
squares = np.array([make_shape('square') for _ in range(N)])
circles = np.array([make_shape('circle') for _ in range(N)])Why use np.mgrid?
np.mgrid creates arrays of pixel coordinates. make_shape then computes distances and edge coverage for all pixels with NumPy array operations, avoiding a Python loop over individual pixels.Normalize pixels to [0, 1], then split each class separately: 1,440 images for training, 160 for validation, and 400 for testing. Shuffling within each subset gives balanced sets of 2,880, 320, and 800 images. Keep the test set aside until training is complete.
X = np.concatenate([squares, circles]) / 255.0 # normalize to [0, 1]
y = np.concatenate([np.zeros(N), np.ones(N)]) # 0 = square, 1 = circle
# Split each class separately, then shuffle each subset.
splits = [[], [], []]
for label in (0, 1):
class_indices = np.random.permutation(np.flatnonzero(y == label))
for target, part in zip(splits, np.split(class_indices, [1440, 1600])):
target.extend(part)
train_idx, val_idx, test_idx = [np.random.permutation(part) for part in splits]
X_train, y_train = X[train_idx], y[train_idx]
X_val, y_val = X[val_idx], y[val_idx]
X_test, y_test = X[test_idx], y[test_idx]Notice we don’t flatten the images — unlike the dense MNIST model where we reshaped to 784 vectors, the conv layer needs the 2D spatial structure intact.
The output activation
Our model outputs a single number — the probability that the input is a circle. To convert the dense layer’s raw output (which can be any value) into a probability between 0 and 1, we use sigmoid:
def sigmoid(x):
z = np.exp(-np.abs(x))
return np.where(np.asarray(x) >= 0, 1 / (1 + z), z / (1 + z))This is the same function we’d use in logistic regression — it squashes any value into the (0, 1) range. For multi-class problems like MNIST, we’d use softmax instead.
The loss function
To measure how wrong the prediction is, we use binary cross-entropy — the same idea as the cross-entropy loss from the MNIST article, but adapted for two classes instead of ten:
def binary_cross_entropy(predicted, label):
# Clip probabilities to keep log(0) out of the reported loss.
p = np.clip(predicted, 1e-10, 1 - 1e-10)
return -label * np.log(p) - (1 - label) * np.log(1 - p)For a circle, predicting 0.95 gives a smaller loss than predicting 0.3. Clipping keeps this reporting function finite at probabilities 0 and 1. The backward pass below uses the analytic sigmoid-plus-BCE gradient p - y, before clipping; production implementations usually compute BCE directly from logits for numerical stability.
The training loop
Recreate the model after generating and splitting the data, as in the notebook. The forward pass uses channels-first arrays throughout:
conv = Conv2D(num_filters=2, kernel_size=3)
pool = MaxPool2D(size=2)
dense = DenseLayer(n_inputs=392, n_outputs=1)
def forward(image):
activated = conv.forward(image) # (28, 28) → (2, 28, 28)
pooled = pool.forward(activated) # (2, 28, 28) → (2, 14, 14)
flat = pooled.reshape(-1) # (2, 14, 14) → (392,)
return sigmoid(dense.forward(flat))[0]The backward pass mirrors it in reverse — from the loss gradient back through dense, un-flatten, pooling, and conv. We’ll dive deep into why output - label is the correct starting gradient for sigmoid + binary cross-entropy in a following article:
def backward(output, label):
grad = np.array([output - label]) # sigmoid + BCE gradient
grad = dense.backward(grad) # dense layer
grad = grad.reshape(pool.output_shape) # un-flatten
grad = pool.backward(grad) # pooling: route to max positions
conv.backward(grad) # conv: ReLU + accumulate across positionsWe accumulate per-image gradients and average them before each SGD update. The notebook also evaluates validation loss after every epoch and evaluates the test set once at the end. Its vectorized convolution and pooling implement the same operations as the loops shown here.
def train(X_train, y_train, epochs=20, lr=0.25, batch_size=32):
for epoch in range(epochs):
indices = np.random.permutation(len(X_train))
for i in range(0, len(X_train), batch_size):
batch_idx = indices[i:i+batch_size]
bs = len(batch_idx)
# Accumulate gradients over the mini-batch
acc_conv_k = np.zeros_like(conv.kernels)
acc_conv_b = np.zeros_like(conv.biases)
acc_dense_W = np.zeros_like(dense.W)
acc_dense_b = np.zeros_like(dense.b)
for idx in batch_idx:
# 1. Forward pass
output = forward(X_train[idx])
# 2. Loss
loss = binary_cross_entropy(output, y_train[idx])
# 3. Backpropagation (computes per-sample gradients)
backward(output, y_train[idx])
# Accumulate
acc_conv_k += conv.grad_kernels
acc_conv_b += conv.grad_biases
acc_dense_W += dense.grad_W
acc_dense_b += dense.grad_b
# 4. Average gradients and update weights
conv.grad_kernels = acc_conv_k / bs
conv.grad_biases = acc_conv_b / bs
dense.grad_W = acc_dense_W / bs
dense.grad_b = acc_dense_b / bs
conv.update(lr)
dense.update(lr)Training results
The run below uses 2,880 training images and 320 validation images for 20 epochs. Training metrics are accumulated during each epoch; validation metrics use the weights at the end of that epoch. The separate 800-image test set is not used for updates or the validation curve.
The final held-out test accuracy is reported beneath the chart. The downloadable notebook exports both these metrics and the weights used by the feature-map demo. Results on these generated shapes do not establish robustness to other drawing styles or real images.
The same architecture in Keras uses channels-last input. This version uses Adam rather than the notebook’s SGD, so its training trajectory will differ:
from tensorflow import keras
from tensorflow.keras import layers
model = keras.Sequential([
layers.Input(shape=(28, 28, 1)),
layers.Conv2D(2, kernel_size=3, activation='relu', padding='same'),
layers.MaxPooling2D(pool_size=2),
layers.Flatten(),
layers.Dense(1, activation='sigmoid'),
])
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
model.fit(X_train[..., None], y_train, epochs=20, batch_size=32,
validation_data=(X_val[..., None], y_val))Scaling up: from shapes to MNIST
Our shape classifier has just 413 parameters because the task is simple. Real image classification needs more capacity. Here’s a typical CNN for MNIST digit recognition — the same principles, just bigger:
model = keras.Sequential([
layers.Input(shape=(28, 28, 1)),
layers.Conv2D(32, kernel_size=3, activation='relu'), # 28×28×1 → 26×26×32
layers.MaxPooling2D(pool_size=2), # 26×26×32 → 13×13×32
layers.Conv2D(64, kernel_size=3, activation='relu'), # 13×13×32 → 11×11×64
layers.MaxPooling2D(pool_size=2), # 11×11×64 → 5×5×64
layers.Flatten(), # 5×5×64 → 1600
layers.Dense(128, activation='relu'), # 1600 → 128
layers.Dense(10, activation='softmax'), # 128 → 10
])The differences from our toy model:
- 28×28 input same size — but grayscale handwritten digits instead of simple geometric shapes
- 32 and 64 filters instead of 2 — many more patterns to detect
- Two conv+pool blocks instead of one — hierarchical feature extraction
- 10-class softmax instead of binary sigmoid — 10 digits to distinguish
- ~225,000 parameters instead of 413 — much more capacity
But the building blocks are identical: slide small filters to build feature maps, pool to reduce size, flatten, classify. Here’s how training progresses over 5 epochs:
The saved MNIST run reaches about 99.08% validation accuracy after 5 epochs. These plotted values are validation metrics; they do not record the separate final test evaluation. A comparison with the earlier dense network requires evaluating both models on the same held-out test set.