Skip to content
Educora
University22 min37 / 42

PyTorch: tensors

Install PyTorch and learn its core data type: creating tensors, dtype and shape, reshape/view, indexing, broadcasting, matrix multiplication with @, GPUs and NumPy interop.

Check yourself
In this lesson you will learn
  • Install PyTorch in a virtual environment and tell the CPU and CUDA builds apart
  • Create tensors and read and change their shape and dtype
  • Apply broadcasting rules and matrix multiplication and predict the resulting shape
  • Move tensors between CPU and GPU and convert them to and from NumPy arrays

NumPy arrays are great for numerical computing, but deep learning needs two things they lack: they do not run on a GPU, and they cannot compute derivatives automatically. PyTorch fills exactly these two gaps. Its core object — the tensor — looks very much like a NumPy array, but it can live on a graphics card and keep track of its gradients. In this lesson you will learn to work fluently with tensors; the next lesson covers automatic differentiation.

What PyTorch is and how to install it

PyTorch is an open-source deep learning library created at the AI research lab of Meta (formerly Facebook); since 2022 it has been governed by the PyTorch Foundation, part of the Linux Foundation. Researchers and companies use it widely to build, train and deploy neural networks. PyTorch does not run in the browser, so you run the PyTorch code of this module on your own computer; the output is shown under every example.

  1. 1
    Create a virtual environment

    In your project folder run python -m venv .venv and activate it, so PyTorch does not interfere with your other projects.

  2. 2
    Choose the CPU or CUDA build

    Without an NVIDIA graphics card, the CPU build is perfectly enough for this module. For a GPU (CUDA) build, copy the exact command from the selector on pytorch.org, choosing your operating system and CUDA version.

  3. 3
    Install and check

    Run pip install torch, then import torch in Python and print the version.

Terminal
python -m venv .venv
# Windows: .venv\Scripts\activate    macOS/Linux: source .venv/bin/activate
pip install torch numpy
# CPU-only build (smaller download):
pip install torch --index-url https://download.pytorch.org/whl/cpu
Python
import torch

print(torch.__version__)
print(torch.cuda.is_available())
Expected output
2.14.0+cpu
False
Your version number may differ. With an NVIDIA GPU and a CUDA build, the second line prints True.
Definition
Tensor

A multi-dimensional array of numbers of a single type. With 0 dimensions (ndim) it is a scalar, 1 — a vector, 2 — a matrix, 3 or more — a general tensor. Every tensor has a shape (the length along each dimension), a dtype (the element type) and a device (CPU or GPU).

DatashapeMeaning
a loss value()scalar, 0 dimensions
tabular data(N, D)N examples, D features
a colour image(3, H, W)RGB channels, height, width
a batch of images(N, C, H, W)the standard image layout in PyTorch
a batch of texts(N, T, E)T tokens, each an E-dimensional vector

Creating tensors: shape and dtype

Python
import torch

a = torch.tensor([[1, 2, 3], [4, 5, 6]])
print(a)
print(a.shape, a.ndim, a.dtype)
print(torch.tensor([1.5, 2.0]).dtype)
print(torch.zeros(2, 3))
print(torch.arange(0, 10, 3))
torch.manual_seed(42)
print(torch.rand(2, 2))
print(a.float().dtype, a.to(torch.float16).dtype)
Expected output
tensor([[1, 2, 3],
        [4, 5, 6]])
torch.Size([2, 3]) 2 torch.int64
torch.float32
tensor([[0., 0., 0.],
        [0., 0., 0.]])
tensor([0, 3, 6, 9])
tensor([[0.8823, 0.9150],
        [0.3829, 0.9593]])
torch.float32 torch.float16
  • torch.tensor(...) infers the type: integers → int64, decimals → float32. NumPy, by contrast, uses float64 for decimals by default!
  • torch.zeros, torch.ones, torch.arange and torch.linspace create ready-made tensors; torch.rand (uniform between 0 and 1) and torch.randn (standard normal) create random ones.
  • torch.manual_seed(42) fixes the random number generator, so results are reproducible.
  • .float(), .long() or .to(torch.float16) change the type and return a new tensor.
dtypeBytesTypical use
torch.float324weights, inputs, losses — the default
torch.float16 / torch.bfloat162fast mixed-precision training on GPUs
torch.int648class labels, indices, token ids
torch.bool1masks (for example x > 0)
numel = d₁ · d₂ · … · dₖ memory = numel · s
where:
  • d₁ … dₖthe sizes in the shape
  • numelnumber of elements (x.numel())
  • ssize of one element in bytes (x.element_size())
Example 1: how much memory does a batch of images take?

A batch of 64 colour images of size 224 × 224 is stored as float32. What is the tensor's shape and how many bytes does it take? What about float16?

Show solution
Shape: (64, 3, 224, 224).
numel = 64 · 3 · 224 · 224 = 9,633,792.
Memory = 9,633,792 · 4 = 38,535,168 bytes ≈ 38.5 MB.
In float16 each element takes 2 bytes, so the memory is halved: ≈ 19.3 MB. This is exactly why mixed precision saves GPU memory.

Reshaping and indexing

Python
import torch

m = torch.arange(12).reshape(3, 4)
print(m)
print(m[0], m[:, 1])
print(m[1, 2].item())
print(m[m > 8])
v = m.view(2, 6)
print(v.shape, m.T.shape)
print(m.unsqueeze(0).shape, m.flatten().shape)
v[0, 0] = 100
print(m[0, 0])
Expected output
tensor([[ 0,  1,  2,  3],
        [ 4,  5,  6,  7],
        [ 8,  9, 10, 11]])
tensor([0, 1, 2, 3]) tensor([1, 5, 9])
6
tensor([ 9, 10, 11])
torch.Size([2, 6]) torch.Size([4, 3])
torch.Size([1, 3, 4]) torch.Size([12])
tensor(100)
  • Indexing works as in NumPy: m[0] is the first row, m[:, 1] the second column, m[m > 8] selects with a mask. .item() turns a one-element tensor into a plain Python number.
  • reshape(2, 6) and view(2, 6) change the shape without changing the number of elements; write -1 for one dimension and PyTorch computes it: x.view(-1, 784).
  • view does not copy the data; it looks at the same memory through a different “window”. That is why v[0, 0] = 100 also changed m!
  • unsqueeze(0) adds a new axis of size 1 (for example, to turn a single image into a batch), squeeze() removes such axes, .T transposes a matrix, and permute reorders axes in any order.

Operations: broadcasting and matrix multiplication

The operators +, -, *, /, ** work element-wise. Reductions such as sum, mean and max work along an axis given by dim: a.mean(dim=0) is the mean of each column. Tensors with different shapes are matched by the broadcasting rules:

  1. Shapes are compared from right to left; the shorter shape is padded on the left with axes of size 1.
  2. Two sizes are compatible if they are equal or one of them is 1.
  3. An axis of size 1 is “stretched” to the other size (without copying data).
Python
import torch

a = torch.tensor([[1., 2.], [3., 4.]])
b = torch.tensor([10., 20.])
print(a + b)
print(a * a)
print(a @ a)
print(a.sum(), a.mean(dim=0))
X = torch.ones(64, 3)
W = torch.ones(3, 5)
print((X @ W).shape)
print((torch.ones(4, 1) + torch.ones(3)).shape)
Expected output
tensor([[11., 22.],
        [13., 24.]])
tensor([[ 1.,  4.],
        [ 9., 16.]])
tensor([[ 7., 10.],
        [15., 22.]])
tensor(10.) tensor([2., 3.])
torch.Size([64, 5])
torch.Size([4, 3])
C = A @ B, Cᵢⱼ = ∑ₖ Aᵢₖ · Bₖⱼ, (m × n) @ (n × p) → (m × p)
where:
  • A, Bmatrices of sizes m × n and n × p
  • nthe inner size — must be the same in both matrices
  • Cᵢⱼdot product of row i of A and column j of B

Matrix multiplication is the core operation of neural networks: multiplying (N, D) inputs by a (D, K) weight matrix processes the whole batch at once. For tensors with more axes, @ works on the last two axes and treats the others as batch axes.

Example 2: predict the result's shape

Find the shape of each result or say that it is an error: a) (64, 3) @ (3, 5); b) (4, 1) + (3,); c) (32, 10) + (10,); d) (2, 3) @ (2, 3); e) (8, 1, 6) + (7, 1).

Show solution
a) The inner sizes 3 = 3 “cancel”: (64, 5).
b) (3,) → (1, 3); 4 with 1 and 1 with 3 are compatible: (4, 3).
c) (10,) → (1, 10) and is added to every row (this is how a bias vector works): (32, 10).
d) Inner sizes 3 ≠ 2 — error. What was needed is A @ B.T: (2, 2).
e) (7, 1) → (1, 7, 1); axis by axis: 8 and 1 → 8, 1 and 7 → 7, 6 and 1 → 6: (8, 7, 6).

Working with GPUs and NumPy

A graphics card has thousands of simple cores that compute matrix products in parallel; training large models on a GPU is often many times — even tens of times — faster than on a CPU. The standard pattern: choose a device at the start of the script, then move the model and the data there with .to(device). On Apple Silicon computers the GPU device is 'mps'.

Python
import numpy as np
import torch

device = 'cuda' if torch.cuda.is_available() else 'cpu'
x = torch.ones(2, 2, device=device)
print(device, x.device)
arr = np.array([1.0, 2.0, 3.0])
t = torch.from_numpy(arr)
arr[0] = 99
print(t)
back = (t * 2).numpy()
print(back, type(back).__name__)
Expected output
cpu cpu
tensor([99.,  2.,  3.], dtype=torch.float64)
[198.   4.   6.] ndarray
This is the output on a computer without a GPU; with CUDA the first line is cuda cuda:0.

torch.from_numpy(arr) shares the array's memory: the change arr[0] = 99 is visible in the tensor. Note that the dtype stayed float64 — convert it to float32 with .float() before training. In the other direction, .numpy() only works for tensors on the CPU: bring a GPU tensor back with .cpu() first, and if it tracks gradients, call .detach() first: t.detach().cpu().numpy().

Key points

  • A tensor is a multi-dimensional array with a shape, a dtype and a device; PyTorch stores images as (N, C, H, W).
  • The default type for decimals is float32; labels and indices use int64.
  • view shares memory and needs contiguous memory; when unsure, use reshape.
  • Broadcasting compares shapes from the right: sizes must be equal or one of them 1. @ is the matrix product: (m × n) @ (n × p) → (m × p).
  • Model and data are moved to the same device with .to(device); from_numpy and .numpy() share memory.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
What is torch.tensor([1.5, 2.0]).dtype?