- Install PyTorch in a virtual environment and tell the CPU and CUDA builds apart
- Create tensors and read and change their shape and dtype
- Apply broadcasting rules and matrix multiplication and predict the resulting shape
- Move tensors between CPU and GPU and convert them to and from NumPy arrays
NumPy arrays are great for numerical computing, but deep learning needs two things they lack: they do not run on a GPU, and they cannot compute derivatives automatically. PyTorch fills exactly these two gaps. Its core object — the tensor — looks very much like a NumPy array, but it can live on a graphics card and keep track of its gradients. In this lesson you will learn to work fluently with tensors; the next lesson covers automatic differentiation.
What PyTorch is and how to install it
PyTorch is an open-source deep learning library created at the AI research lab of Meta (formerly Facebook); since 2022 it has been governed by the PyTorch Foundation, part of the Linux Foundation. Researchers and companies use it widely to build, train and deploy neural networks. PyTorch does not run in the browser, so you run the PyTorch code of this module on your own computer; the output is shown under every example.
- 1Create a virtual environment
In your project folder run
python -m venv .venvand activate it, so PyTorch does not interfere with your other projects. - 2Choose the CPU or CUDA build
Without an NVIDIA graphics card, the CPU build is perfectly enough for this module. For a GPU (CUDA) build, copy the exact command from the selector on pytorch.org, choosing your operating system and CUDA version.
- 3Install and check
Run
pip install torch, thenimport torchin Python and print the version.
python -m venv .venv
# Windows: .venv\Scripts\activate macOS/Linux: source .venv/bin/activate
pip install torch numpy
# CPU-only build (smaller download):
pip install torch --index-url https://download.pytorch.org/whl/cpuimport torch
print(torch.__version__)
print(torch.cuda.is_available())2.14.0+cpu False
True.A multi-dimensional array of numbers of a single type. With 0 dimensions (ndim) it is a scalar, 1 — a vector, 2 — a matrix, 3 or more — a general tensor. Every tensor has a shape (the length along each dimension), a dtype (the element type) and a device (CPU or GPU).
| Data | shape | Meaning |
|---|---|---|
| a loss value | () | scalar, 0 dimensions |
| tabular data | (N, D) | N examples, D features |
| a colour image | (3, H, W) | RGB channels, height, width |
| a batch of images | (N, C, H, W) | the standard image layout in PyTorch |
| a batch of texts | (N, T, E) | T tokens, each an E-dimensional vector |
Creating tensors: shape and dtype
import torch
a = torch.tensor([[1, 2, 3], [4, 5, 6]])
print(a)
print(a.shape, a.ndim, a.dtype)
print(torch.tensor([1.5, 2.0]).dtype)
print(torch.zeros(2, 3))
print(torch.arange(0, 10, 3))
torch.manual_seed(42)
print(torch.rand(2, 2))
print(a.float().dtype, a.to(torch.float16).dtype)tensor([[1, 2, 3],
[4, 5, 6]])
torch.Size([2, 3]) 2 torch.int64
torch.float32
tensor([[0., 0., 0.],
[0., 0., 0.]])
tensor([0, 3, 6, 9])
tensor([[0.8823, 0.9150],
[0.3829, 0.9593]])
torch.float32 torch.float16torch.tensor(...)infers the type: integers →int64, decimals →float32. NumPy, by contrast, usesfloat64for decimals by default!torch.zeros,torch.ones,torch.arangeandtorch.linspacecreate ready-made tensors;torch.rand(uniform between 0 and 1) andtorch.randn(standard normal) create random ones.torch.manual_seed(42)fixes the random number generator, so results are reproducible..float(),.long()or.to(torch.float16)change the type and return a new tensor.
| dtype | Bytes | Typical use |
|---|---|---|
torch.float32 | 4 | weights, inputs, losses — the default |
torch.float16 / torch.bfloat16 | 2 | fast mixed-precision training on GPUs |
torch.int64 | 8 | class labels, indices, token ids |
torch.bool | 1 | masks (for example x > 0) |
- d₁ … dₖthe sizes in the shape
- numelnumber of elements (
x.numel()) - ssize of one element in bytes (
x.element_size())
A batch of 64 colour images of size 224 × 224 is stored as float32. What is the tensor's shape and how many bytes does it take? What about float16?
Show solutionHide solution
numel = 64 · 3 · 224 · 224 = 9,633,792.
Memory = 9,633,792 · 4 = 38,535,168 bytes ≈ 38.5 MB.
In
float16 each element takes 2 bytes, so the memory is halved: ≈ 19.3 MB. This is exactly why mixed precision saves GPU memory.Reshaping and indexing
import torch
m = torch.arange(12).reshape(3, 4)
print(m)
print(m[0], m[:, 1])
print(m[1, 2].item())
print(m[m > 8])
v = m.view(2, 6)
print(v.shape, m.T.shape)
print(m.unsqueeze(0).shape, m.flatten().shape)
v[0, 0] = 100
print(m[0, 0])tensor([[ 0, 1, 2, 3],
[ 4, 5, 6, 7],
[ 8, 9, 10, 11]])
tensor([0, 1, 2, 3]) tensor([1, 5, 9])
6
tensor([ 9, 10, 11])
torch.Size([2, 6]) torch.Size([4, 3])
torch.Size([1, 3, 4]) torch.Size([12])
tensor(100)- Indexing works as in NumPy:
m[0]is the first row,m[:, 1]the second column,m[m > 8]selects with a mask..item()turns a one-element tensor into a plain Python number. reshape(2, 6)andview(2, 6)change the shape without changing the number of elements; write-1for one dimension and PyTorch computes it:x.view(-1, 784).viewdoes not copy the data; it looks at the same memory through a different “window”. That is whyv[0, 0] = 100also changedm!unsqueeze(0)adds a new axis of size 1 (for example, to turn a single image into a batch),squeeze()removes such axes,.Ttransposes a matrix, andpermutereorders axes in any order.
Operations: broadcasting and matrix multiplication
The operators +, -, *, /, ** work element-wise. Reductions such as sum, mean and max work along an axis given by dim: a.mean(dim=0) is the mean of each column. Tensors with different shapes are matched by the broadcasting rules:
- Shapes are compared from right to left; the shorter shape is padded on the left with axes of size 1.
- Two sizes are compatible if they are equal or one of them is 1.
- An axis of size 1 is “stretched” to the other size (without copying data).
import torch
a = torch.tensor([[1., 2.], [3., 4.]])
b = torch.tensor([10., 20.])
print(a + b)
print(a * a)
print(a @ a)
print(a.sum(), a.mean(dim=0))
X = torch.ones(64, 3)
W = torch.ones(3, 5)
print((X @ W).shape)
print((torch.ones(4, 1) + torch.ones(3)).shape)tensor([[11., 22.],
[13., 24.]])
tensor([[ 1., 4.],
[ 9., 16.]])
tensor([[ 7., 10.],
[15., 22.]])
tensor(10.) tensor([2., 3.])
torch.Size([64, 5])
torch.Size([4, 3])- A, Bmatrices of sizes m × n and n × p
- nthe inner size — must be the same in both matrices
- Cᵢⱼdot product of row i of A and column j of B
Matrix multiplication is the core operation of neural networks: multiplying (N, D) inputs by a (D, K) weight matrix processes the whole batch at once. For tensors with more axes, @ works on the last two axes and treats the others as batch axes.
Find the shape of each result or say that it is an error: a) (64, 3) @ (3, 5); b) (4, 1) + (3,); c) (32, 10) + (10,); d) (2, 3) @ (2, 3); e) (8, 1, 6) + (7, 1).
Show solutionHide solution
b) (3,) → (1, 3); 4 with 1 and 1 with 3 are compatible: (4, 3).
c) (10,) → (1, 10) and is added to every row (this is how a bias vector works): (32, 10).
d) Inner sizes 3 ≠ 2 — error. What was needed is
A @ B.T: (2, 2).e) (7, 1) → (1, 7, 1); axis by axis: 8 and 1 → 8, 1 and 7 → 7, 6 and 1 → 6: (8, 7, 6).
Working with GPUs and NumPy
A graphics card has thousands of simple cores that compute matrix products in parallel; training large models on a GPU is often many times — even tens of times — faster than on a CPU. The standard pattern: choose a device at the start of the script, then move the model and the data there with .to(device). On Apple Silicon computers the GPU device is 'mps'.
import numpy as np
import torch
device = 'cuda' if torch.cuda.is_available() else 'cpu'
x = torch.ones(2, 2, device=device)
print(device, x.device)
arr = np.array([1.0, 2.0, 3.0])
t = torch.from_numpy(arr)
arr[0] = 99
print(t)
back = (t * 2).numpy()
print(back, type(back).__name__)cpu cpu tensor([99., 2., 3.], dtype=torch.float64) [198. 4. 6.] ndarray
cuda cuda:0.torch.from_numpy(arr) shares the array's memory: the change arr[0] = 99 is visible in the tensor. Note that the dtype stayed float64 — convert it to float32 with .float() before training. In the other direction, .numpy() only works for tensors on the CPU: bring a GPU tensor back with .cpu() first, and if it tracks gradients, call .detach() first: t.detach().cpu().numpy().
Key points
- A tensor is a multi-dimensional array with a shape, a dtype and a device; PyTorch stores images as (N, C, H, W).
- The default type for decimals is
float32; labels and indices useint64. viewshares memory and needs contiguous memory; when unsure, usereshape.- Broadcasting compares shapes from the right: sizes must be equal or one of them 1.
@is the matrix product: (m × n) @ (n × p) → (m × p). - Model and data are moved to the same device with
.to(device);from_numpyand.numpy()share memory.
Check yourself
10 questions. Every correct answer earns XP.
torch.tensor([1.5, 2.0]).dtype?