Skip to content
Educora
University22 min39 / 42

PyTorch: building neural networks

From an artificial neuron to a multilayer network: z = w·x + b, activation functions (ReLU, sigmoid, tanh, softmax), nn.Linear layers, nn.Module and nn.Sequential, the forward pass and parameter counts.

Check yourself
In this lesson you will learn
  • Write down the mathematical model of a neuron and of a fully connected layer
  • Compute ReLU, sigmoid, tanh and softmax and know where each is used
  • Build a network with nn.Module and nn.Sequential and count its parameters

We want to recognise a handwritten digit: the image is 28 × 28 = 784 pixels and the answer is one of ten digits from 0 to 9. What function can turn 784 numbers into ten probabilities? A neural network builds such a function from many simple pieces: linear transformations with non-linear “bends” between them. The name was inspired by the brain, but in fact it is a mathematical function whose parameters are chosen by gradient descent — exactly as in the previous lessons.

The artificial neuron

z = w · x + b = ∑ᵢ₌₁ⁿ wᵢxᵢ + b a = φ(z)
where:
  • xinput vector (n features)
  • wweights — the importance of each input
  • bbias — shifts the activation threshold
  • φactivation function (ReLU, sigmoid…)
  • athe neuron's output (activation)

A neuron computes a weighted sum of its inputs, adds the bias and passes the result through an activation function. In the example below z = 0.4 · 0.5 + 0.3 · (−1) + (−0.2) · 2 + 0.1 = −0.4: the sigmoid turns it into 0.4013 and ReLU into 0.

Python
import torch

x = torch.tensor([0.5, -1.0, 2.0])
w = torch.tensor([0.4, 0.3, -0.2])
b = torch.tensor(0.1)
z = w @ x + b
print(round(z.item(), 4))
print(round(torch.sigmoid(z).item(), 4), torch.relu(z).item())
Expected output
-0.4
0.4013 0.0

Activation functions

Without activations the network would stay linear no matter how many layers we stack. The proof is short: applying two layers in a row gives W₂(W₁x + b₁) + b₂ = (W₂W₁)x + (W₂b₁ + b₂) — that is one linear layer with matrix W = W₂W₁ and bias b = W₂b₁ + b₂. A non-linear φ “bends” the layers and lets the network describe complex, curved boundaries.

ReLU(z) = max(0, z) σ(z) = 1 / (1 + e⁻ᶻ) tanh(z) = (eᶻ − e⁻ᶻ) / (eᶻ + e⁻ᶻ)ReLU(z) = max(0, z) σ(z) = 1 / (1 + e⁻ᶻ) tanh(z) = (eᶻ − e⁻ᶻ) / (eᶻ + e⁻ᶻ)
where:
  • ReLUzeroes negative values; its derivative is 1 for z > 0 and 0 for z < 0
  • σsigmoid: values in (0, 1), read as a probability
  • tanhhyperbolic tangent: values in (−1, 1), symmetric around zero
softmax(z)ᵢ = eᶻⁱ / ∑ⱼ₌₁ᴷ eᶻʲsoftmax(z)ᵢ = eᶻⁱ / ∑ⱼ₌₁ᴷ eᶻʲ
where:
  • zraw scores (logits) for K classes
  • softmax(z)ᵢprobability of class i; all are positive and sum to 1

Softmax turns the logits of the last layer into a probability distribution in multi-class classification.

FunctionRangeWhere it is used
ReLU[0; ∞)the default in hidden layers: simple, fast, gradients do not vanish
Sigmoid(0; 1)output for binary classification, “gates” (LSTM)
tanh(−1; 1)recurrent networks, when a zero-centred output is needed
softmax(0, 1), sums to 1last layer of multi-class classification

ReLU has one weakness: if a neuron's input is negative for every example, its gradient is always zero and the neuron never learns again (a “dead ReLU”). LeakyReLU counters this by keeping a small slope on the negative side (for example 0.01z). Transformers widely use GELU, a smooth variant of ReLU. PyTorch has them all ready: nn.LeakyReLU(), nn.GELU(), nn.Tanh(), nn.Sigmoid().

Python
import torch

z = torch.tensor([-2.0, 0.0, 1.0, 3.0])
print(torch.relu(z))
print(torch.sigmoid(z).round(decimals=4))
print(torch.tanh(z).round(decimals=4))
p = torch.softmax(z, dim=0)
print(p.round(decimals=4), round(p.sum().item(), 4))
Expected output
tensor([0., 0., 1., 3.])
tensor([0.1192, 0.5000, 0.7311, 0.9526])
tensor([-0.9640,  0.0000,  0.7616,  0.9951])
tensor([0.0057, 0.0418, 0.1135, 0.8390]) 1.0
Example 1: softmax by hand

A network outputs the logits z = (2, 1, 0) for three classes. Find the probabilities.

Show solution
e² ≈ 7.389, e¹ ≈ 2.718, e⁰ = 1. Sum ≈ 11.107.
p ≈ (7.389 / 11.107, 2.718 / 11.107, 1 / 11.107) ≈ (0.665, 0.245, 0.090).
The logits differ by only 1, yet the probabilities differ by a factor of about e ≈ 2.72: softmax “amplifies” the largest score.
Python
import torch

print(torch.softmax(torch.tensor([2.0, 1.0, 0.0]), dim=0).round(decimals=3))
Expected output
tensor([0.6650, 0.2450, 0.0900])

Layers: nn.Linear

When several neurons look at the same input, their weight vectors become the rows of one matrix W, and all outputs of the layer are computed with one matrix product. Inputs arriving as a batch have shape (N, nin), and outputs have shape (N, nout).

Y = X · Wᵀ + b params = (nin + 1) · nout
where:
  • Xbatch of inputs, shape (N, nin)
  • Wweights, shape (nout, nin) — layer.weight
  • bbiases, shape (nout,) — layer.bias
  • Youtputs, shape (N, nout)

Each output neuron has nin weights and one bias, so a fully connected layer has (nin + 1) · nout parameters.

Python
import torch
from torch import nn

torch.manual_seed(42)
layer = nn.Linear(in_features=3, out_features=2)
print(layer.weight.shape, layer.bias.shape)
x = torch.randn(4, 3)
print(layer(x).shape)
print(sum(p.numel() for p in layer.parameters()))
Expected output
torch.Size([2, 3]) torch.Size([2])
torch.Size([4, 2])
8

Weights start as small random values: by default nn.Linear draws them uniformly from (−1/√nin, 1/√nin). Randomness matters — if all weights were equal, every neuron in the layer would get the same gradient and learn the same thing (the symmetry problem).

Networks with nn.Module and nn.Sequential

In PyTorch every model is a subclass of nn.Module. Layers are created in __init__ — their parameters are registered automatically — and the forward method describes the computation from input to output. We call the model as model(x), and PyTorch runs forward for us.

Python
import torch
from torch import nn

class MLP(nn.Module):
    def __init__(self, n_in, n_hidden, n_out):
        super().__init__()
        self.fc1 = nn.Linear(n_in, n_hidden)
        self.fc2 = nn.Linear(n_hidden, n_out)

    def forward(self, x):
        h = torch.relu(self.fc1(x))
        return self.fc2(h)

torch.manual_seed(42)
model = MLP(784, 128, 10)
print(model)
x = torch.randn(32, 784)
print(model(x).shape)
print(sum(p.numel() for p in model.parameters()))
Expected output
MLP(
  (fc1): Linear(in_features=784, out_features=128, bias=True)
  (fc2): Linear(in_features=128, out_features=10, bias=True)
)
torch.Size([32, 10])
101770
A batch of 32 images (32, 784) → hidden layer (32, 128) → logits (32, 10). Parameters: 785 · 128 + 129 · 10 = 100,480 + 1,290 = 101,770.
  • model.parameters() — all trainable parameters; we pass them to the optimizer.
  • model.to(device) — moves all parameters from CPU to GPU (or back) with one call.
  • model.train() and model.eval() — training and evaluation modes (they matter for Dropout and BatchNorm layers).
  • model.state_dict() — a dictionary of parameters; we will use it in the next lesson to save the model to a file.
Definition
Multilayer perceptron (MLP) and logits

A network made of fully connected layers with activations between them is called a multilayer perceptron. The unnormalised outputs of its last layer — the logits — can be any real numbers. For a prediction we take the index of the largest logit (logits.argmax(dim=1)); for probabilities we use torch.softmax(logits, dim=1).

How do you choose an architecture? For tabular data, 1–3 hidden layers with 32–256 neurons are usually enough; start with a small model and grow it while the validation loss allows. For images, fully connected layers are far too expensive — going from a 224 × 224 colour image to a layer of 1000 neurons already needs over 150 million parameters. Convolutional networks solve this problem.

If the layers simply follow one another, there is no need to write a class: nn.Sequential applies them in order. Here the activation is written as an nn.ReLU() layer. named_parameters() shows the name and shape of every parameter — very useful for debugging:

Python
import torch
from torch import nn

model = nn.Sequential(
    nn.Linear(784, 128), nn.ReLU(),
    nn.Linear(128, 64), nn.ReLU(),
    nn.Linear(64, 10),
)
for name, p in model.named_parameters():
    print(name, tuple(p.shape))
print(sum(p.numel() for p in model.parameters()))
Expected output
0.weight (128, 784)
0.bias (128,)
2.weight (64, 128)
2.bias (64,)
4.weight (10, 64)
4.bias (10,)
109386
The 0, 2, 4 in the names are the layers' positions inside Sequential; the ReLU layers at positions 1 and 3 have no parameters.
Example 2: count the parameters by hand

How many parameters does the network 784 → 128 → 64 → 10 have? How much memory do they take in float32?

Show solution
Layer 1: (784 + 1) · 128 = 100,480.
Layer 2: (128 + 1) · 64 = 8,256.
Layer 3: (64 + 1) · 10 = 650.
Total: 109,386 — the same as the code's output.
Memory: 109,386 · 4 bytes ≈ 0.44 MB. 92% of the parameters sit in the first layer: large inputs are expensive.
Example 3: a forward pass by hand

A 2 → 2 → 1 network: W₁ = [[1, −1], [0.5, 0.5]], b₁ = (0, −1), ReLU in the hidden layer, W₂ = [[2, −3]], b₂ = 1. Find the output for x = (1, 2).

Show solution
W₁x + b₁ = (1 − 2 + 0, 0.5 + 1 − 1) = (−1, 0.5).
ReLU: h = (0, 0.5) — the first neuron is “switched off”.
y = 2 · 0 + (−3) · 0.5 + 1 = −0.5.
Note: for this x the first neuron has no effect on the output, so in the backward pass the gradient of its weights will be zero as well — ReLU lets the gradient through active neurons only.
Python
import torch

x = torch.tensor([1.0, 2.0])
W1 = torch.tensor([[1.0, -1.0], [0.5, 0.5]])
b1 = torch.tensor([0.0, -1.0])
W2 = torch.tensor([[2.0, -3.0]])
b2 = torch.tensor([1.0])
h = torch.relu(W1 @ x + b1)
y = W2 @ h + b2
print(h, y)
Expected output
tensor([0.0000, 0.5000]) tensor([-0.5000])

Key points

  • A neuron computes z = w · x + b, then an activation a = φ(z); a layer does this in matrix form: Y = X · Wᵀ + b.
  • Without activations a multilayer network equals one linear layer; hidden layers usually use ReLU.
  • Softmax turns logits into probabilities that sum to 1; nn.CrossEntropyLoss takes raw logits.
  • A model subclasses nn.Module: layers in __init__, computation in forward; for a simple chain nn.Sequential is enough.
  • A fully connected layer has (nin + 1) · nout parameters; nn.Linear(n_in, n_out).weight has shape (nout, nin).

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
How many parameters does nn.Linear(10, 5) have?