- Write down the mathematical model of a neuron and of a fully connected layer
- Compute ReLU, sigmoid, tanh and softmax and know where each is used
- Build a network with
nn.Moduleandnn.Sequentialand count its parameters
We want to recognise a handwritten digit: the image is 28 × 28 = 784 pixels and the answer is one of ten digits from 0 to 9. What function can turn 784 numbers into ten probabilities? A neural network builds such a function from many simple pieces: linear transformations with non-linear “bends” between them. The name was inspired by the brain, but in fact it is a mathematical function whose parameters are chosen by gradient descent — exactly as in the previous lessons.
The artificial neuron
- xinput vector (n features)
- wweights — the importance of each input
- bbias — shifts the activation threshold
- φactivation function (ReLU, sigmoid…)
- athe neuron's output (activation)
A neuron computes a weighted sum of its inputs, adds the bias and passes the result through an activation function. In the example below z = 0.4 · 0.5 + 0.3 · (−1) + (−0.2) · 2 + 0.1 = −0.4: the sigmoid turns it into 0.4013 and ReLU into 0.
import torch
x = torch.tensor([0.5, -1.0, 2.0])
w = torch.tensor([0.4, 0.3, -0.2])
b = torch.tensor(0.1)
z = w @ x + b
print(round(z.item(), 4))
print(round(torch.sigmoid(z).item(), 4), torch.relu(z).item())-0.4 0.4013 0.0
Activation functions
Without activations the network would stay linear no matter how many layers we stack. The proof is short: applying two layers in a row gives W₂(W₁x + b₁) + b₂ = (W₂W₁)x + (W₂b₁ + b₂) — that is one linear layer with matrix W = W₂W₁ and bias b = W₂b₁ + b₂. A non-linear φ “bends” the layers and lets the network describe complex, curved boundaries.
- ReLUzeroes negative values; its derivative is 1 for z > 0 and 0 for z < 0
- σsigmoid: values in (0, 1), read as a probability
- tanhhyperbolic tangent: values in (−1, 1), symmetric around zero
- zraw scores (logits) for K classes
- softmax(z)ᵢprobability of class i; all are positive and sum to 1
Softmax turns the logits of the last layer into a probability distribution in multi-class classification.
| Function | Range | Where it is used |
|---|---|---|
| ReLU | [0; ∞) | the default in hidden layers: simple, fast, gradients do not vanish |
| Sigmoid | (0; 1) | output for binary classification, “gates” (LSTM) |
| tanh | (−1; 1) | recurrent networks, when a zero-centred output is needed |
| softmax | (0, 1), sums to 1 | last layer of multi-class classification |
ReLU has one weakness: if a neuron's input is negative for every example, its gradient is always zero and the neuron never learns again (a “dead ReLU”). LeakyReLU counters this by keeping a small slope on the negative side (for example 0.01z). Transformers widely use GELU, a smooth variant of ReLU. PyTorch has them all ready: nn.LeakyReLU(), nn.GELU(), nn.Tanh(), nn.Sigmoid().
import torch
z = torch.tensor([-2.0, 0.0, 1.0, 3.0])
print(torch.relu(z))
print(torch.sigmoid(z).round(decimals=4))
print(torch.tanh(z).round(decimals=4))
p = torch.softmax(z, dim=0)
print(p.round(decimals=4), round(p.sum().item(), 4))tensor([0., 0., 1., 3.]) tensor([0.1192, 0.5000, 0.7311, 0.9526]) tensor([-0.9640, 0.0000, 0.7616, 0.9951]) tensor([0.0057, 0.0418, 0.1135, 0.8390]) 1.0
A network outputs the logits z = (2, 1, 0) for three classes. Find the probabilities.
Show solutionHide solution
p ≈ (7.389 / 11.107, 2.718 / 11.107, 1 / 11.107) ≈ (0.665, 0.245, 0.090).
The logits differ by only 1, yet the probabilities differ by a factor of about e ≈ 2.72: softmax “amplifies” the largest score.
import torch
print(torch.softmax(torch.tensor([2.0, 1.0, 0.0]), dim=0).round(decimals=3))tensor([0.6650, 0.2450, 0.0900])
Layers: nn.Linear
When several neurons look at the same input, their weight vectors become the rows of one matrix W, and all outputs of the layer are computed with one matrix product. Inputs arriving as a batch have shape (N, nin), and outputs have shape (N, nout).
- Xbatch of inputs, shape (N, nin)
- Wweights, shape (nout, nin) —
layer.weight - bbiases, shape (nout,) —
layer.bias - Youtputs, shape (N, nout)
Each output neuron has nin weights and one bias, so a fully connected layer has (nin + 1) · nout parameters.
import torch
from torch import nn
torch.manual_seed(42)
layer = nn.Linear(in_features=3, out_features=2)
print(layer.weight.shape, layer.bias.shape)
x = torch.randn(4, 3)
print(layer(x).shape)
print(sum(p.numel() for p in layer.parameters()))torch.Size([2, 3]) torch.Size([2]) torch.Size([4, 2]) 8
Weights start as small random values: by default nn.Linear draws them uniformly from (−1/√nin, 1/√nin). Randomness matters — if all weights were equal, every neuron in the layer would get the same gradient and learn the same thing (the symmetry problem).
Networks with nn.Module and nn.Sequential
In PyTorch every model is a subclass of nn.Module. Layers are created in __init__ — their parameters are registered automatically — and the forward method describes the computation from input to output. We call the model as model(x), and PyTorch runs forward for us.
import torch
from torch import nn
class MLP(nn.Module):
def __init__(self, n_in, n_hidden, n_out):
super().__init__()
self.fc1 = nn.Linear(n_in, n_hidden)
self.fc2 = nn.Linear(n_hidden, n_out)
def forward(self, x):
h = torch.relu(self.fc1(x))
return self.fc2(h)
torch.manual_seed(42)
model = MLP(784, 128, 10)
print(model)
x = torch.randn(32, 784)
print(model(x).shape)
print(sum(p.numel() for p in model.parameters()))MLP( (fc1): Linear(in_features=784, out_features=128, bias=True) (fc2): Linear(in_features=128, out_features=10, bias=True) ) torch.Size([32, 10]) 101770
model.parameters()— all trainable parameters; we pass them to the optimizer.model.to(device)— moves all parameters from CPU to GPU (or back) with one call.model.train()andmodel.eval()— training and evaluation modes (they matter for Dropout and BatchNorm layers).model.state_dict()— a dictionary of parameters; we will use it in the next lesson to save the model to a file.
A network made of fully connected layers with activations between them is called a multilayer perceptron. The unnormalised outputs of its last layer — the logits — can be any real numbers. For a prediction we take the index of the largest logit (logits.argmax(dim=1)); for probabilities we use torch.softmax(logits, dim=1).
How do you choose an architecture? For tabular data, 1–3 hidden layers with 32–256 neurons are usually enough; start with a small model and grow it while the validation loss allows. For images, fully connected layers are far too expensive — going from a 224 × 224 colour image to a layer of 1000 neurons already needs over 150 million parameters. Convolutional networks solve this problem.
If the layers simply follow one another, there is no need to write a class: nn.Sequential applies them in order. Here the activation is written as an nn.ReLU() layer. named_parameters() shows the name and shape of every parameter — very useful for debugging:
import torch
from torch import nn
model = nn.Sequential(
nn.Linear(784, 128), nn.ReLU(),
nn.Linear(128, 64), nn.ReLU(),
nn.Linear(64, 10),
)
for name, p in model.named_parameters():
print(name, tuple(p.shape))
print(sum(p.numel() for p in model.parameters()))0.weight (128, 784) 0.bias (128,) 2.weight (64, 128) 2.bias (64,) 4.weight (10, 64) 4.bias (10,) 109386
Sequential; the ReLU layers at positions 1 and 3 have no parameters.How many parameters does the network 784 → 128 → 64 → 10 have? How much memory do they take in float32?
Show solutionHide solution
Layer 2: (128 + 1) · 64 = 8,256.
Layer 3: (64 + 1) · 10 = 650.
Total: 109,386 — the same as the code's output.
Memory: 109,386 · 4 bytes ≈ 0.44 MB. 92% of the parameters sit in the first layer: large inputs are expensive.
A 2 → 2 → 1 network: W₁ = [[1, −1], [0.5, 0.5]], b₁ = (0, −1), ReLU in the hidden layer, W₂ = [[2, −3]], b₂ = 1. Find the output for x = (1, 2).
Show solutionHide solution
ReLU: h = (0, 0.5) — the first neuron is “switched off”.
y = 2 · 0 + (−3) · 0.5 + 1 = −0.5.
Note: for this x the first neuron has no effect on the output, so in the backward pass the gradient of its weights will be zero as well — ReLU lets the gradient through active neurons only.
import torch
x = torch.tensor([1.0, 2.0])
W1 = torch.tensor([[1.0, -1.0], [0.5, 0.5]])
b1 = torch.tensor([0.0, -1.0])
W2 = torch.tensor([[2.0, -3.0]])
b2 = torch.tensor([1.0])
h = torch.relu(W1 @ x + b1)
y = W2 @ h + b2
print(h, y)tensor([0.0000, 0.5000]) tensor([-0.5000])
Key points
- A neuron computes z = w · x + b, then an activation a = φ(z); a layer does this in matrix form: Y = X · Wᵀ + b.
- Without activations a multilayer network equals one linear layer; hidden layers usually use ReLU.
- Softmax turns logits into probabilities that sum to 1;
nn.CrossEntropyLosstakes raw logits. - A model subclasses
nn.Module: layers in__init__, computation inforward; for a simple chainnn.Sequentialis enough. - A fully connected layer has (nin + 1) · nout parameters;
nn.Linear(n_in, n_out).weighthas shape (nout, nin).
Check yourself
10 questions. Every correct answer earns XP.
nn.Linear(10, 5) have?