Skip to content
Educora
University22 min38 / 42

PyTorch: autograd and automatic differentiation

How PyTorch computes derivatives for you: the computational graph, requires_grad, backward() and .grad, the chain rule, gradient accumulation, torch.no_grad() and gradient descent with autograd.

Check yourself
In this lesson you will learn
  • Explain how the computational graph and the backward pass (backpropagation) work
  • Compute gradients with requires_grad, .backward() and .grad and check them by hand with the chain rule
  • Zero gradients and use torch.no_grad() correctly
  • Write gradient descent from scratch with autograd

In the first lesson we derived the gradients of linear regression by hand — for just two parameters. Modern neural networks have millions of parameters, large language models billions; deriving each derivative on paper is impossible. PyTorch's autograd does this automatically: it records every operation you perform on tensors and then, with a single command, computes exact derivatives with respect to all parameters. This is neither approximate numerical differences nor symbolic algebra — it is automatic differentiation.

The computational graph

Every operation on a tensor with requires_grad=True adds a node to a directed acyclic graph (DAG). The graph's leaves are the inputs and parameters, and its root is the loss. During the forward pass the graph is built and values are computed; during the backward pass autograd walks from the root to the leaves and applies the chain rule at every node. In neural networks this process is called backpropagation.

NodeForwardLocal derivativegrad_fn
uu = w · x∂u/∂w = xMulBackward0
ŷŷ = u + b∂ŷ/∂u = 1, ∂ŷ/∂b = 1AddBackward0
ee = ŷ − y∂e/∂ŷ = 1SubBackward0
LL = e²∂L/∂e = 2ePowBackward0
The graph of the loss L = (w · x + b − y)². In the backward pass the local derivatives are multiplied from the bottom up: ∂L/∂w = 2e · 1 · 1 · x.

In PyTorch the graph is dynamic (define-by-run): it is rebuilt on every forward pass. That is why you can freely use ordinary Python if and for statements inside a model — the graph matches whatever path was actually executed.

MethodHow it worksDrawback
Numerical differentiation(f(θ + h) − f(θ − h)) / 2happroximate, and needs two evaluations per parameter
Symbolic differentiationtransforms the formula with algebra rules (like SymPy)expressions blow up for large models; loops and branches are hard
Automatic differentiation (autograd)combines the local derivatives of the executed operations with the chain ruleneeds extra memory to store intermediate values

requires_grad, backward() and .grad

Python
import torch

x = torch.tensor(3.0, requires_grad=True)
y = x ** 2 + 2 * x + 1
print(y)
y.backward()
print(x.grad)
Expected output
tensor(16., grad_fn=<AddBackward0>)
tensor(8.)
  • requires_grad=True tells PyTorch: “I will need derivatives with respect to this tensor, record the operations”. Parameters (the weights of nn.Linear and so on) do this automatically.
  • grad_fn shows which operation created the tensor; it is its node in the graph.
  • y.backward() computes dy/dx and stores it in the leaf's .grad: for y = x² + 2x + 1, dy/dx = 2x + 2 = 8 at x = 3.

When a parameter is a vector, .grad is a vector of the same shape: one partial derivative per element. For example, for f = ∑ vᵢ², ∂f/∂vᵢ = 2vᵢ:

Python
import torch

v = torch.tensor([1.0, -2.0, 3.0], requires_grad=True)
f = (v ** 2).sum()
f.backward()
print(v.grad)
Expected output
tensor([ 2., -4.,  6.])

The requires_grad flag can also be changed later: p.requires_grad_(False) “freezes” a parameter — its gradient is not computed and it does not change during training. This is used in transfer learning: most layers of a pretrained network are frozen and only the last layer is adapted to a new task. You will see this in the lesson on convolutional networks.

The chain rule

dL/dx = dL/dy · dy/dx ∂L/∂w = ∂L/∂ŷ · ∂ŷ/∂wdL/dx = dL/dy · dy/dx ∂L/∂w = ∂L/∂ŷ · ∂ŷ/∂w
where:
  • Lthe loss (root of the graph)
  • y, ŷintermediate values (inner nodes of the graph)
  • x, wleaves: inputs and parameters

The derivative of a composite function is the product of local derivatives. If a variable affects the loss along several paths, the products along all paths are added: ∂L/∂x = ∑ₖ ∂L/∂yₖ · ∂yₖ/∂x.

Example 1: the gradient of a linear model

ŷ = w · x + b, L = (ŷ − y)². For w = 2, b = 1, x = 3, y = 10, find ∂L/∂w and ∂L/∂b with the chain rule.

Show solution
Forward: ŷ = 2 · 3 + 1 = 7, L = (7 − 10)² = 9.
∂L/∂ŷ = 2(ŷ − y) = −6.
∂ŷ/∂w = x = 3 ⇒ ∂L/∂w = −6 · 3 = −18.
∂ŷ/∂b = 1 ⇒ ∂L/∂b = −6.
Both gradients are negative: increasing w and b lowers the loss — indeed ŷ = 7 is below the target 10.
Python
import torch

w = torch.tensor(2.0, requires_grad=True)
b = torch.tensor(1.0, requires_grad=True)
x, y = torch.tensor(3.0), torch.tensor(10.0)
y_hat = w * x + b
loss = (y_hat - y) ** 2
print(type(loss.grad_fn).__name__, type(y_hat.grad_fn).__name__)
loss.backward()
print(loss.item(), w.grad.item(), b.grad.item())
Expected output
PowBackward0 AddBackward0
9.0 -18.0 -6.0
Autograd confirms the −18 and −6 we found by hand. x and y have no requires_grad — they are data, not parameters.
σ(z) = 1 / (1 + e⁻ᶻ) σ′(z) = σ(z) · (1 − σ(z))σ(z) = 1 / (1 + e⁻ᶻ) σ′(z) = σ(z) · (1 − σ(z))
where:
  • σ(z)the sigmoid function, with values in (0, 1)
  • σ′(z)its derivative; its maximum is 0.25 at z = 0

Proof: σ′(z) = e⁻ᶻ / (1 + e⁻ᶻ)² = σ(z) · e⁻ᶻ / (1 + e⁻ᶻ), and the last fraction equals 1 − σ(z).

Example 2: the gradient of a sigmoid neuron

y = σ(w · x), w = 0, x = 2. Find dy/dw.

Show solution
z = w · x = 0, y = σ(0) = 0.5.
dy/dw = σ′(z) · dz/dw = σ(z)(1 − σ(z)) · x = 0.5 · 0.5 · 2 = 0.5.
Consequence: since σ′ is never larger than 0.25, the gradient shrinks at least 4 times at every sigmoid layer it passes through. After 10 layers that is a factor of 4¹⁰ ≈ 10⁶ — the vanishing gradient problem. This is why ReLU dominates in deep networks.
Python
import torch

w = torch.tensor(0.0, requires_grad=True)
x = torch.tensor(2.0)
y = torch.sigmoid(w * x)
y.backward()
print(y.item(), w.grad.item())
Expected output
0.5 0.5
Example 3: a variable that acts along two paths

f = x · y + x², x = 2, y = 3. Find ∂f/∂x and ∂f/∂y. What does autograd do here?

Show solution
x enters two nodes of the graph: u = x · y and v = x².
First path: ∂u/∂x = y = 3. Second path: ∂v/∂x = 2x = 4.
The paths are added: ∂f/∂x = 3 + 4 = 7; ∂f/∂y = x = 2.
Autograd does the same: it adds every gradient flow that reaches x into .grad. This is exactly where gradient accumulation comes from.

The power of the backward pass is its efficiency: one backward() call computes the gradient of the loss with respect to all parameters at a cost of the same order as the forward pass (a small constant factor). If we computed each parameter separately, a network with a million parameters would need a million forward passes.

Gradients accumulate: zeroing and torch.no_grad()

.backward() does not replace the old gradient — it adds to it. Below, the derivative of y = x³ at x = 2 is 3x² = 12, yet after three calls .grad is 36:

Python
import torch

x = torch.tensor(2.0, requires_grad=True)
for i in range(3):
    y = x ** 3
    y.backward()
    print(x.grad)
x.grad.zero_()
print(x.grad)
with torch.no_grad():
    z = x * 2
print(z.requires_grad, x.detach().requires_grad)
Expected output
tensor(12.)
tensor(24.)
tensor(36.)
tensor(0.)
False False
  • x.grad.zero_() resets the gradient in place (a trailing underscore _ means an in-place operation). In the next lessons optimizer.zero_grad() will do this for us.
  • Inside with torch.no_grad(): no graph is built: use it when updating parameters and when evaluating a model (inference); it saves memory and time.
  • .detach() returns a tensor that shares the same data but is cut off from the graph — for example, to print a value or convert it to NumPy.

Gradient descent with autograd

Now let us put everything together and fit a linear model to 20 noisy points from the line y = 2x + 1. This time we write no derivative formulas — autograd computes them at every step. The five steps of this loop are the core of all deep learning: forward pass → loss → backward → update inside no_grad → zero the gradients.

Python
import torch

torch.manual_seed(42)
X = torch.linspace(0, 1, 20)
Y = 2 * X + 1 + 0.1 * torch.randn(20)
w = torch.zeros(1, requires_grad=True)
b = torch.zeros(1, requires_grad=True)
lr = 0.5
for step in range(201):
    loss = ((w * X + b - Y) ** 2).mean()
    loss.backward()
    with torch.no_grad():
        w -= lr * w.grad
        b -= lr * b.grad
    w.grad.zero_()
    b.grad.zero_()
    if step % 50 == 0:
        print(f'step {step:3d}  loss {loss.item():.4f}')
print(f'w = {w.item():.3f}, b = {b.item():.3f}')
Expected output
step   0  loss 4.4302
step  50  loss 0.0085
step 100  loss 0.0084
step 150  loss 0.0084
step 200  loss 0.0084
w = 1.907, b = 1.068
In just 50 steps the loss fell from 4.43 to 0.0085. The final loss is close to the noise variance (0.1² = 0.01) — going lower would mean overfitting. w and b are not exactly 2 and 1 because this is the best line for 20 noisy points.

Remember this loop well: in the next lessons we will shorten it, but its essence will not change. The step() method of the torch.optim.SGD optimizer performs exactly p -= lr * p.grad inside no_grad for every parameter, and zero_grad() clears the gradients. Layers such as nn.Linear create w and b themselves and switch on requires_grad for you. So the high-level code that looks “magical” is just a convenient wrapper around these five lines.

Key points

  • Autograd builds a computational graph from tensor operations and backward() applies the chain rule from the root to the leaves.
  • Gradients of leaves with requires_grad=True are stored in .grad; backward() is called on a scalar.
  • Chain rule: dL/dx = dL/dy · dy/dx; for the sigmoid σ′ = σ(1 − σ).
  • Gradients accumulate — zero them after every step; updates and evaluation go inside torch.no_grad().
  • The training loop: forward → loss → backward → update → zero.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
x = 1, y = x³ + 4x. What is x.grad after y.backward()?