- Represent an image as an (N, C, H, W) tensor and normalise it
- Compute a convolution by hand and find the output size and the number of parameters
- Build a small CNN for MNIST and trace the shapes layer by layer
- Explain transfer learning and adapt a pretrained model
Your phone camera recognises faces, programs that assist doctors find changes on X-ray images, cars read road signs. Most of these rely on convolutional neural networks (CNNs). In the previous lesson we saw that a fully connected layer for a 224 × 224 colour image needs over 150 million parameters. CNNs exploit two simple observations: meaningful patterns in an image are local (edges, corners, textures), and the same pattern can appear anywhere in the picture. So a small filter is “slid” over the whole image, and its weights are shared everywhere.
Images as tensors
A grayscale 28 × 28 image is a tensor of shape (1, 28, 28): one channel, 28 rows, 28 columns. A colour image has three channels — red, green, blue: (3, H, W). For batches PyTorch uses the (N, C, H, W) layout. Pixels are usually integers 0–255; transforms.ToTensor() scales them to [0, 1], and then each channel is normalised.
- xpixel value in [0, 1]
- μ, σmean and standard deviation of the channel over the training set (0.1307 and 0.3081 for MNIST)
After normalisation the inputs are roughly centred on zero with unit scale — this speeds up gradient descent.
The convolution operation
A filter (kernel) is a small matrix of weights, for example 3 × 3. It slides over the image, and at every position it is multiplied element-wise with the pixels beneath it and summed. The result is a feature map: it has large values wherever the pattern the filter “looks for” is present. A convolutional layer has several filters, each producing its own output channel; every filter spans all the input channels.
- Xinput image (or the previous layer's map)
- Ka K × K filter — learned weights
- bthe filter's bias
- Y[i, j]element (i, j) of the feature map
2D convolution for one channel. With several input channels, the sum also runs over the channels.
The code below computes a convolution from scratch with NumPy. The left half of the image is dark (0) and the right half bright (9); the filter responds to brightness increasing from left to right. Run the code and check the result yourself:
import numpy as np
image = np.array([
[0, 0, 0, 9, 9, 9],
[0, 0, 0, 9, 9, 9],
[0, 0, 0, 9, 9, 9],
[0, 0, 0, 9, 9, 9],
[0, 0, 0, 9, 9, 9],
])
kernel = np.array([
[-1, 0, 1],
[-1, 0, 1],
[-1, 0, 1],
])
K = kernel.shape[0]
H, W = image.shape
out = np.zeros((H - K + 1, W - K + 1), dtype=int)
for i in range(out.shape[0]):
for j in range(out.shape[1]):
out[i, j] = np.sum(image[i:i + K, j:j + K] * kernel)
print(out)▸ Expected output
[[ 0 27 27 0] [ 0 27 27 0] [ 0 27 27 0]]
In trained CNNs, filters in the first layers typically respond to edges and colour transitions, middle layers to textures and parts (eyes, wheels), and the last layers to whole objects. After each convolution and pooling layer, the region a neuron “sees” — its receptive field — grows.
- Winput width (or height), in pixels
- Kkernel size
- Ppadding: zero pixels added around the border
- Sstride: how many pixels the filter moves
- Ooutput width; ⌊ ⌋ means rounding down
The output-size formula; the same holds for the height. For pooling, usually P = 0 and S = K.
- Cin, Coutnumber of input and output channels
A convolutional layer's parameter count does not depend on the image size — the weights are shared across all positions.
For a 28 × 28 input, find the output size: a) K = 3, P = 1, S = 1; b) K = 5, P = 0, S = 1; c) K = 3, P = 1, S = 2. d) How many parameters does Conv2d(1, 16, 3) with 16 filters have?
Show solutionHide solution
b) (28 − 5 + 0) / 1 + 1 = 24.
c) ⌊(28 − 3 + 2) / 2⌋ + 1 = ⌊13.5⌋ + 1 = 14 — the size is halved.
d) (3 · 3 · 1 + 1) · 16 = 160. For comparison, a fully connected layer from a 28 × 28 image to a 16 × 28 × 28 output would have 784 · 12,544 ≈ 9.8 million weights!
import torch
from torch import nn
torch.manual_seed(42)
x = torch.randn(8, 1, 28, 28)
conv = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=1)
print(conv(x).shape)
print(nn.Conv2d(1, 16, kernel_size=5)(x).shape)
print(nn.Conv2d(1, 16, kernel_size=3, stride=2, padding=1)(x).shape)
print(nn.MaxPool2d(kernel_size=2)(conv(x)).shape)
print(conv.weight.shape, sum(p.numel() for p in conv.parameters()))torch.Size([8, 16, 28, 28]) torch.Size([8, 16, 24, 24]) torch.Size([8, 16, 14, 14]) torch.Size([8, 16, 14, 14]) torch.Size([16, 1, 3, 3]) 160
Pooling and a CNN for 28 × 28 images
Max pooling keeps only the largest value from each window (usually 2 × 2). It halves the size of the map, makes computation cheaper, adds robustness to small shifts and has no parameters at all. Average pooling takes the mean instead.
Apply 2 × 2 max and average pooling (S = 2) to the 4 × 4 map:
1 3 2 1
4 6 5 0
7 2 1 9
3 8 4 2
Show solutionHide solution
Max: [[6, 5], [8, 9]].
Average: [[3.5, 2], [5, 4]].
The output is 2 × 2: O = ⌊(4 − 2) / 2⌋ + 1 = 2.
Now let us build a small CNN for MNIST: two “convolution → ReLU → max pooling” blocks extract features, and at the end Flatten plus a fully connected layer produce logits for the 10 digits. We print the shape after every layer — the most reliable way to debug a CNN.
import torch
from torch import nn
class SmallCNN(nn.Module):
def __init__(self):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
)
self.classifier = nn.Sequential(nn.Flatten(), nn.Linear(32 * 7 * 7, 10))
def forward(self, x):
return self.classifier(self.features(x))
torch.manual_seed(42)
model = SmallCNN()
x = torch.randn(64, 1, 28, 28)
for layer in model.features:
x = layer(x)
print(f'{type(layer).__name__:9s} {tuple(x.shape)}')
print(model(torch.randn(64, 1, 28, 28)).shape)
print(sum(p.numel() for p in model.parameters()))Conv2d (64, 16, 28, 28) ReLU (64, 16, 28, 28) MaxPool2d (64, 16, 14, 14) Conv2d (64, 32, 14, 14) ReLU (64, 32, 14, 14) MaxPool2d (64, 32, 7, 7) torch.Size([64, 10]) 20490
Trace the shapes of SmallCNN and count its parameters by hand. Compare it with the MLP from lesson 4 (101,770 parameters).
Show solutionHide solution
conv1: (3 · 3 · 1 + 1) · 16 = 160.
conv2: (3 · 3 · 16 + 1) · 32 = 4,640.
fc: (1568 + 1) · 10 = 15,690.
Total: 20,490 — about 5 times fewer than the MLP, yet usually more accurate on MNIST because it exploits the structure of images.
MNIST with torchvision
The torchvision package (pip install torchvision) provides ready-made datasets (MNIST, CIFAR-10, the ImageNet format), image transforms (transforms) and pretrained models. MNIST consists of 70,000 images of handwritten digits: 60,000 for training and 10,000 for testing, each 28 × 28 grayscale. download=True downloads the files into the data folder the first time and reads them from disk afterwards.
import torch
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
torch.manual_seed(42)
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.1307,), (0.3081,)),
])
train_set = datasets.MNIST('data', train=True, download=True, transform=transform)
test_set = datasets.MNIST('data', train=False, download=True, transform=transform)
train_loader = DataLoader(train_set, batch_size=64, shuffle=True)
test_loader = DataLoader(test_set, batch_size=1000)
images, labels = next(iter(train_loader))
print(len(train_set), len(test_set))
print(images.shape, labels.shape)60000 10000 torch.Size([64, 1, 28, 28]) torch.Size([64])
model = SmallCNN().to(device). A small CNN like this usually reaches about 98–99% test accuracy within a few epochs.Transfer learning
Training a large CNN from scratch needs millions of images and powerful GPUs. But the edges, textures and shapes learned by the first layers of a network trained on ImageNet (1000 classes, over a million training images) are useful for almost any image task. Transfer learning works like this: take a pretrained model, freeze its layers, replace the last classification layer with a new one for your own classes and train only that.
import torch
from torch import nn
from torchvision import models
model = models.resnet18(weights=models.ResNet18_Weights.DEFAULT)
for p in model.parameters():
p.requires_grad = False
print(model.fc)
model.fc = nn.Linear(model.fc.in_features, 3)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(trainable, total)Linear(in_features=512, out_features=1000, bias=True) 1539 11178051
The new layer has (512 + 1) · 3 = 1,539 parameters. So few parameters can be trained quickly with a few hundred images, even on a CPU. With more data, in a second stage you can “unfreeze” the last few layers too and fine-tune the whole network with a very small learning rate. Give the optimizer only the trainable parameters: torch.optim.Adam(model.fc.parameters(), lr=1e-3).
Key points
- Images in PyTorch are (N, C, H, W) tensors; inputs are normalised with (x − μ) / σ.
- Convolution slides a small filter over the image; shared weights keep the parameter count small: (K·K·Cin + 1)·Cout.
- Output size: O = ⌊(W − K + 2P) / S⌋ + 1; a 3 × 3 filter with P = 1 keeps the size, S = 2 or 2 × 2 pooling halves it.
- A typical CNN: [Conv → ReLU → Pool] × n → Flatten → Linear; print the shapes after every layer.
- Transfer learning: freeze a pretrained model, replace the last layer and train only that.
Check yourself
10 questions. Every correct answer earns XP.