- Explain how the computational graph and the backward pass (backpropagation) work
- Compute gradients with
requires_grad,.backward()and.gradand check them by hand with the chain rule - Zero gradients and use
torch.no_grad()correctly - Write gradient descent from scratch with autograd
In the first lesson we derived the gradients of linear regression by hand — for just two parameters. Modern neural networks have millions of parameters, large language models billions; deriving each derivative on paper is impossible. PyTorch's autograd does this automatically: it records every operation you perform on tensors and then, with a single command, computes exact derivatives with respect to all parameters. This is neither approximate numerical differences nor symbolic algebra — it is automatic differentiation.
The computational graph
Every operation on a tensor with requires_grad=True adds a node to a directed acyclic graph (DAG). The graph's leaves are the inputs and parameters, and its root is the loss. During the forward pass the graph is built and values are computed; during the backward pass autograd walks from the root to the leaves and applies the chain rule at every node. In neural networks this process is called backpropagation.
| Node | Forward | Local derivative | grad_fn |
|---|---|---|---|
| u | u = w · x | ∂u/∂w = x | MulBackward0 |
| ŷ | ŷ = u + b | ∂ŷ/∂u = 1, ∂ŷ/∂b = 1 | AddBackward0 |
| e | e = ŷ − y | ∂e/∂ŷ = 1 | SubBackward0 |
| L | L = e² | ∂L/∂e = 2e | PowBackward0 |
In PyTorch the graph is dynamic (define-by-run): it is rebuilt on every forward pass. That is why you can freely use ordinary Python if and for statements inside a model — the graph matches whatever path was actually executed.
| Method | How it works | Drawback |
|---|---|---|
| Numerical differentiation | (f(θ + h) − f(θ − h)) / 2h | approximate, and needs two evaluations per parameter |
| Symbolic differentiation | transforms the formula with algebra rules (like SymPy) | expressions blow up for large models; loops and branches are hard |
| Automatic differentiation (autograd) | combines the local derivatives of the executed operations with the chain rule | needs extra memory to store intermediate values |
requires_grad, backward() and .grad
import torch
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2 + 2 * x + 1
print(y)
y.backward()
print(x.grad)tensor(16., grad_fn=<AddBackward0>) tensor(8.)
requires_grad=Truetells PyTorch: “I will need derivatives with respect to this tensor, record the operations”. Parameters (the weights ofnn.Linearand so on) do this automatically.grad_fnshows which operation created the tensor; it is its node in the graph.y.backward()computes dy/dx and stores it in the leaf's.grad: for y = x² + 2x + 1, dy/dx = 2x + 2 = 8 at x = 3.
When a parameter is a vector, .grad is a vector of the same shape: one partial derivative per element. For example, for f = ∑ vᵢ², ∂f/∂vᵢ = 2vᵢ:
import torch
v = torch.tensor([1.0, -2.0, 3.0], requires_grad=True)
f = (v ** 2).sum()
f.backward()
print(v.grad)tensor([ 2., -4., 6.])
The requires_grad flag can also be changed later: p.requires_grad_(False) “freezes” a parameter — its gradient is not computed and it does not change during training. This is used in transfer learning: most layers of a pretrained network are frozen and only the last layer is adapted to a new task. You will see this in the lesson on convolutional networks.
The chain rule
- Lthe loss (root of the graph)
- y, ŷintermediate values (inner nodes of the graph)
- x, wleaves: inputs and parameters
The derivative of a composite function is the product of local derivatives. If a variable affects the loss along several paths, the products along all paths are added: ∂L/∂x = ∑ₖ ∂L/∂yₖ · ∂yₖ/∂x.
ŷ = w · x + b, L = (ŷ − y)². For w = 2, b = 1, x = 3, y = 10, find ∂L/∂w and ∂L/∂b with the chain rule.
Show solutionHide solution
∂L/∂ŷ = 2(ŷ − y) = −6.
∂ŷ/∂w = x = 3 ⇒ ∂L/∂w = −6 · 3 = −18.
∂ŷ/∂b = 1 ⇒ ∂L/∂b = −6.
Both gradients are negative: increasing w and b lowers the loss — indeed ŷ = 7 is below the target 10.
import torch
w = torch.tensor(2.0, requires_grad=True)
b = torch.tensor(1.0, requires_grad=True)
x, y = torch.tensor(3.0), torch.tensor(10.0)
y_hat = w * x + b
loss = (y_hat - y) ** 2
print(type(loss.grad_fn).__name__, type(y_hat.grad_fn).__name__)
loss.backward()
print(loss.item(), w.grad.item(), b.grad.item())PowBackward0 AddBackward0 9.0 -18.0 -6.0
x and y have no requires_grad — they are data, not parameters.- σ(z)the sigmoid function, with values in (0, 1)
- σ′(z)its derivative; its maximum is 0.25 at z = 0
Proof: σ′(z) = e⁻ᶻ / (1 + e⁻ᶻ)² = σ(z) · e⁻ᶻ / (1 + e⁻ᶻ), and the last fraction equals 1 − σ(z).
y = σ(w · x), w = 0, x = 2. Find dy/dw.
Show solutionHide solution
dy/dw = σ′(z) · dz/dw = σ(z)(1 − σ(z)) · x = 0.5 · 0.5 · 2 = 0.5.
Consequence: since σ′ is never larger than 0.25, the gradient shrinks at least 4 times at every sigmoid layer it passes through. After 10 layers that is a factor of 4¹⁰ ≈ 10⁶ — the vanishing gradient problem. This is why ReLU dominates in deep networks.
import torch
w = torch.tensor(0.0, requires_grad=True)
x = torch.tensor(2.0)
y = torch.sigmoid(w * x)
y.backward()
print(y.item(), w.grad.item())0.5 0.5
f = x · y + x², x = 2, y = 3. Find ∂f/∂x and ∂f/∂y. What does autograd do here?
Show solutionHide solution
First path: ∂u/∂x = y = 3. Second path: ∂v/∂x = 2x = 4.
The paths are added: ∂f/∂x = 3 + 4 = 7; ∂f/∂y = x = 2.
Autograd does the same: it adds every gradient flow that reaches x into
.grad. This is exactly where gradient accumulation comes from.The power of the backward pass is its efficiency: one backward() call computes the gradient of the loss with respect to all parameters at a cost of the same order as the forward pass (a small constant factor). If we computed each parameter separately, a network with a million parameters would need a million forward passes.
Gradients accumulate: zeroing and torch.no_grad()
.backward() does not replace the old gradient — it adds to it. Below, the derivative of y = x³ at x = 2 is 3x² = 12, yet after three calls .grad is 36:
import torch
x = torch.tensor(2.0, requires_grad=True)
for i in range(3):
y = x ** 3
y.backward()
print(x.grad)
x.grad.zero_()
print(x.grad)
with torch.no_grad():
z = x * 2
print(z.requires_grad, x.detach().requires_grad)tensor(12.) tensor(24.) tensor(36.) tensor(0.) False False
x.grad.zero_()resets the gradient in place (a trailing underscore_means an in-place operation). In the next lessonsoptimizer.zero_grad()will do this for us.- Inside
with torch.no_grad():no graph is built: use it when updating parameters and when evaluating a model (inference); it saves memory and time. .detach()returns a tensor that shares the same data but is cut off from the graph — for example, to print a value or convert it to NumPy.
Gradient descent with autograd
Now let us put everything together and fit a linear model to 20 noisy points from the line y = 2x + 1. This time we write no derivative formulas — autograd computes them at every step. The five steps of this loop are the core of all deep learning: forward pass → loss → backward → update inside no_grad → zero the gradients.
import torch
torch.manual_seed(42)
X = torch.linspace(0, 1, 20)
Y = 2 * X + 1 + 0.1 * torch.randn(20)
w = torch.zeros(1, requires_grad=True)
b = torch.zeros(1, requires_grad=True)
lr = 0.5
for step in range(201):
loss = ((w * X + b - Y) ** 2).mean()
loss.backward()
with torch.no_grad():
w -= lr * w.grad
b -= lr * b.grad
w.grad.zero_()
b.grad.zero_()
if step % 50 == 0:
print(f'step {step:3d} loss {loss.item():.4f}')
print(f'w = {w.item():.3f}, b = {b.item():.3f}')step 0 loss 4.4302 step 50 loss 0.0085 step 100 loss 0.0084 step 150 loss 0.0084 step 200 loss 0.0084 w = 1.907, b = 1.068
Remember this loop well: in the next lessons we will shorten it, but its essence will not change. The step() method of the torch.optim.SGD optimizer performs exactly p -= lr * p.grad inside no_grad for every parameter, and zero_grad() clears the gradients. Layers such as nn.Linear create w and b themselves and switch on requires_grad for you. So the high-level code that looks “magical” is just a convenient wrapper around these five lines.
Key points
- Autograd builds a computational graph from tensor operations and
backward()applies the chain rule from the root to the leaves. - Gradients of leaves with
requires_grad=Trueare stored in.grad;backward()is called on a scalar. - Chain rule: dL/dx = dL/dy · dy/dx; for the sigmoid σ′ = σ(1 − σ).
- Gradients accumulate — zero them after every step; updates and evaluation go inside
torch.no_grad(). - The training loop: forward → loss → backward → update → zero.
Check yourself
10 questions. Every correct answer earns XP.
x.grad after y.backward()?