Skip to content
Educora
University25 min41 / 42

Convolutional neural networks (CNNs)

Images as tensors (C × H × W), convolution and the output-size formula, filters, pooling, a small CNN for 28 × 28 images with a shape walk-through, MNIST with torchvision and transfer learning.

Check yourself
In this lesson you will learn
  • Represent an image as an (N, C, H, W) tensor and normalise it
  • Compute a convolution by hand and find the output size and the number of parameters
  • Build a small CNN for MNIST and trace the shapes layer by layer
  • Explain transfer learning and adapt a pretrained model

Your phone camera recognises faces, programs that assist doctors find changes on X-ray images, cars read road signs. Most of these rely on convolutional neural networks (CNNs). In the previous lesson we saw that a fully connected layer for a 224 × 224 colour image needs over 150 million parameters. CNNs exploit two simple observations: meaningful patterns in an image are local (edges, corners, textures), and the same pattern can appear anywhere in the picture. So a small filter is “slid” over the whole image, and its weights are shared everywhere.

Images as tensors

A grayscale 28 × 28 image is a tensor of shape (1, 28, 28): one channel, 28 rows, 28 columns. A colour image has three channels — red, green, blue: (3, H, W). For batches PyTorch uses the (N, C, H, W) layout. Pixels are usually integers 0–255; transforms.ToTensor() scales them to [0, 1], and then each channel is normalised.

x′ = (x − μ) / σx′ = (x − μ) / σ
where:
  • xpixel value in [0, 1]
  • μ, σmean and standard deviation of the channel over the training set (0.1307 and 0.3081 for MNIST)

After normalisation the inputs are roughly centred on zero with unit scale — this speeds up gradient descent.

The convolution operation

A filter (kernel) is a small matrix of weights, for example 3 × 3. It slides over the image, and at every position it is multiplied element-wise with the pixels beneath it and summed. The result is a feature map: it has large values wherever the pattern the filter “looks for” is present. A convolutional layer has several filters, each producing its own output channel; every filter spans all the input channels.

Y[i, j] = ∑ₘ ∑ₙ X[i + m, j + n] · K[m, n] + b
where:
  • Xinput image (or the previous layer's map)
  • Ka K × K filter — learned weights
  • bthe filter's bias
  • Y[i, j]element (i, j) of the feature map

2D convolution for one channel. With several input channels, the sum also runs over the channels.

The code below computes a convolution from scratch with NumPy. The left half of the image is dark (0) and the right half bright (9); the filter responds to brightness increasing from left to right. Run the code and check the result yourself:

Python
import numpy as np

image = np.array([
    [0, 0, 0, 9, 9, 9],
    [0, 0, 0, 9, 9, 9],
    [0, 0, 0, 9, 9, 9],
    [0, 0, 0, 9, 9, 9],
    [0, 0, 0, 9, 9, 9],
])
kernel = np.array([
    [-1, 0, 1],
    [-1, 0, 1],
    [-1, 0, 1],
])
K = kernel.shape[0]
H, W = image.shape
out = np.zeros((H - K + 1, W - K + 1), dtype=int)
for i in range(out.shape[0]):
    for j in range(out.shape[1]):
        out[i, j] = np.sum(image[i:i + K, j:j + K] * kernel)
print(out)
▸ Expected output
[[ 0 27 27  0]
 [ 0 27 27  0]
 [ 0 27 27  0]]
The values 27 appear exactly where the vertical edge is: the filter is an edge detector. In a CNN such filters are not written by hand — gradient descent learns them.

In trained CNNs, filters in the first layers typically respond to edges and colour transitions, middle layers to textures and parts (eyes, wheels), and the last layers to whole objects. After each convolution and pooling layer, the region a neuron “sees” — its receptive field — grows.

O = ⌊(W − K + 2P) / S⌋ + 1O = ⌊(W − K + 2P) / S⌋ + 1
where:
  • Winput width (or height), in pixels
  • Kkernel size
  • Ppadding: zero pixels added around the border
  • Sstride: how many pixels the filter moves
  • Ooutput width; ⌊ ⌋ means rounding down

The output-size formula; the same holds for the height. For pooling, usually P = 0 and S = K.

params = (K · K · Cin + 1) · Cout
where:
  • Cin, Coutnumber of input and output channels

A convolutional layer's parameter count does not depend on the image size — the weights are shared across all positions.

Example 1: compute the output sizes

For a 28 × 28 input, find the output size: a) K = 3, P = 1, S = 1; b) K = 5, P = 0, S = 1; c) K = 3, P = 1, S = 2. d) How many parameters does Conv2d(1, 16, 3) with 16 filters have?

Show solution
a) (28 − 3 + 2) / 1 + 1 = 28 — the size is preserved.
b) (28 − 5 + 0) / 1 + 1 = 24.
c) ⌊(28 − 3 + 2) / 2⌋ + 1 = ⌊13.5⌋ + 1 = 14 — the size is halved.
d) (3 · 3 · 1 + 1) · 16 = 160. For comparison, a fully connected layer from a 28 × 28 image to a 16 × 28 × 28 output would have 784 · 12,544 ≈ 9.8 million weights!
Python
import torch
from torch import nn

torch.manual_seed(42)
x = torch.randn(8, 1, 28, 28)
conv = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=1)
print(conv(x).shape)
print(nn.Conv2d(1, 16, kernel_size=5)(x).shape)
print(nn.Conv2d(1, 16, kernel_size=3, stride=2, padding=1)(x).shape)
print(nn.MaxPool2d(kernel_size=2)(conv(x)).shape)
print(conv.weight.shape, sum(p.numel() for p in conv.parameters()))
Expected output
torch.Size([8, 16, 28, 28])
torch.Size([8, 16, 24, 24])
torch.Size([8, 16, 14, 14])
torch.Size([8, 16, 14, 14])
torch.Size([16, 1, 3, 3]) 160
PyTorch confirms the answers of Example 1. The filter weights have shape (Cout, Cin, K, K).

Pooling and a CNN for 28 × 28 images

Max pooling keeps only the largest value from each window (usually 2 × 2). It halves the size of the map, makes computation cheaper, adds robustness to small shifts and has no parameters at all. Average pooling takes the mean instead.

Example 2: pooling by hand

Apply 2 × 2 max and average pooling (S = 2) to the 4 × 4 map:
1 3 2 1
4 6 5 0
7 2 1 9
3 8 4 2

Show solution
Windows: top-left {1, 3, 4, 6}, top-right {2, 1, 5, 0}, bottom-left {7, 2, 3, 8}, bottom-right {1, 9, 4, 2}.
Max: [[6, 5], [8, 9]].
Average: [[3.5, 2], [5, 4]].
The output is 2 × 2: O = ⌊(4 − 2) / 2⌋ + 1 = 2.

Now let us build a small CNN for MNIST: two “convolution → ReLU → max pooling” blocks extract features, and at the end Flatten plus a fully connected layer produce logits for the 10 digits. We print the shape after every layer — the most reliable way to debug a CNN.

Python
import torch
from torch import nn

class SmallCNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
        )
        self.classifier = nn.Sequential(nn.Flatten(), nn.Linear(32 * 7 * 7, 10))

    def forward(self, x):
        return self.classifier(self.features(x))

torch.manual_seed(42)
model = SmallCNN()
x = torch.randn(64, 1, 28, 28)
for layer in model.features:
    x = layer(x)
    print(f'{type(layer).__name__:9s} {tuple(x.shape)}')
print(model(torch.randn(64, 1, 28, 28)).shape)
print(sum(p.numel() for p in model.parameters()))
Expected output
Conv2d    (64, 16, 28, 28)
ReLU      (64, 16, 28, 28)
MaxPool2d (64, 16, 14, 14)
Conv2d    (64, 32, 14, 14)
ReLU      (64, 32, 14, 14)
MaxPool2d (64, 32, 7, 7)
torch.Size([64, 10])
20490
Example 3: shapes and parameters

Trace the shapes of SmallCNN and count its parameters by hand. Compare it with the MLP from lesson 4 (101,770 parameters).

Show solution
Shapes: (1, 28, 28) → conv (P = 1): (16, 28, 28) → pool: (16, 14, 14) → conv: (32, 14, 14) → pool: (32, 7, 7) → flatten: 32 · 7 · 7 = 1568.
conv1: (3 · 3 · 1 + 1) · 16 = 160.
conv2: (3 · 3 · 16 + 1) · 32 = 4,640.
fc: (1568 + 1) · 10 = 15,690.
Total: 20,490 — about 5 times fewer than the MLP, yet usually more accurate on MNIST because it exploits the structure of images.

MNIST with torchvision

The torchvision package (pip install torchvision) provides ready-made datasets (MNIST, CIFAR-10, the ImageNet format), image transforms (transforms) and pretrained models. MNIST consists of 70,000 images of handwritten digits: 60,000 for training and 10,000 for testing, each 28 × 28 grayscale. download=True downloads the files into the data folder the first time and reads them from disk afterwards.

Python
import torch
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

torch.manual_seed(42)
transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.1307,), (0.3081,)),
])
train_set = datasets.MNIST('data', train=True, download=True, transform=transform)
test_set = datasets.MNIST('data', train=False, download=True, transform=transform)
train_loader = DataLoader(train_set, batch_size=64, shuffle=True)
test_loader = DataLoader(test_set, batch_size=1000)
images, labels = next(iter(train_loader))
print(len(train_set), len(test_set))
print(images.shape, labels.shape)
Expected output
60000 10000
torch.Size([64, 1, 28, 28]) torch.Size([64])
An internet connection is needed (the first time only). The training loop is the same as in the previous lesson — just use model = SmallCNN().to(device). A small CNN like this usually reaches about 98–99% test accuracy within a few epochs.

Transfer learning

Training a large CNN from scratch needs millions of images and powerful GPUs. But the edges, textures and shapes learned by the first layers of a network trained on ImageNet (1000 classes, over a million training images) are useful for almost any image task. Transfer learning works like this: take a pretrained model, freeze its layers, replace the last classification layer with a new one for your own classes and train only that.

Python
import torch
from torch import nn
from torchvision import models

model = models.resnet18(weights=models.ResNet18_Weights.DEFAULT)
for p in model.parameters():
    p.requires_grad = False
print(model.fc)
model.fc = nn.Linear(model.fc.in_features, 3)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(trainable, total)
Expected output
Linear(in_features=512, out_features=1000, bias=True)
1539 11178051
The first run downloads the ImageNet weights (about 45 MB). For a 3-class task (for example, three leaf diseases of a plant), only 1,539 of the 11.18 million parameters are trained.

The new layer has (512 + 1) · 3 = 1,539 parameters. So few parameters can be trained quickly with a few hundred images, even on a CPU. With more data, in a second stage you can “unfreeze” the last few layers too and fine-tune the whole network with a very small learning rate. Give the optimizer only the trainable parameters: torch.optim.Adam(model.fc.parameters(), lr=1e-3).

Key points

  • Images in PyTorch are (N, C, H, W) tensors; inputs are normalised with (x − μ) / σ.
  • Convolution slides a small filter over the image; shared weights keep the parameter count small: (K·K·Cin + 1)·Cout.
  • Output size: O = ⌊(W − K + 2P) / S⌋ + 1; a 3 × 3 filter with P = 1 keeps the size, S = 2 or 2 × 2 pooling halves it.
  • A typical CNN: [Conv → ReLU → Pool] × n → Flatten → Linear; print the shapes after every layer.
  • Transfer learning: freeze a pretrained model, replace the last layer and train only that.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
Input 32 × 32, K = 5, P = 0, S = 1. What is the output size?