Skip to content
Educora

Formulas & shortcuts

Python · 97

Every formula in this course and the easiest ways to remember them, on one page.

1Functions and modulesIntermediate

Functions

Open lesson
F = C · 9/5 + 32F = C · 9/5 + 32
where:
  • Ftemperature in degrees Fahrenheit
  • Ctemperature in degrees Celsius

Let's turn this formula into a function

2Python in depthAdvanced

Iterators and generators

Open lesson

Decorators and closures

Open lesson

Context managers and the with statement

Open lesson

Type hints and dataclasses

Open lesson

Advanced OOP: special methods, properties and the MRO

Open lesson

Regular expressions with re

Open lesson

Asynchronous Python: async and await

Open lesson

Packages, modules and virtual environments

Open lesson

Performance, Big-O and the collections module

Open lesson

3Web and automationAdvanced

Working with JSON and CSV

Open lesson

Web APIs and the requests library

Open lesson

Web services with FastAPI

Open lesson

Automation scripts

Open lesson

4Data science: NumPy, pandas, matplotlibUniversity

NumPy basics: arrays

Open lesson
size = d₁ · d₂ · … · dₙ nbytes = size · itemsize
where:
  • nnumber of axes (ndim)
  • dₖlength along axis k (an element of shape)
  • itemsizesize of one element in bytes: 8 for float64 and int64, 4 for float32 and int32, 1 for bool and int8

For m with shape (2, 3): size = 2 · 3 = 6. For temps: nbytes = 4 · 8 = 32 bytes.

h = (stop − start) / (num − 1) n = ⌈(stop − start) / step⌉h = (stop − start) / (num − 1) n = ⌈(stop − start) / step⌉
where:
  • hthe spacing between linspace points
  • nnumber of elements produced by arange
  • ⌈ ⌉rounding up (ceiling)

np.linspace(0, 1, 5): h = 1 / 4 = 0.25. np.arange(0, 10, 2): n = ⌈10 / 2⌉ = 5 elements.

x̄ = (1/n) · ∑ xᵢ σ = √( ∑ (xᵢ − x̄)² / (n − ddof) )x̄ = (1/n) · ∑ xᵢ σ = √( ∑ (xᵢ − x̄)² / (n − ddof) )
where:
  • x̄arithmetic mean (mean)
  • nnumber of values
  • σstandard deviation (std): how far the values typically are from the mean
  • ddof0 — for a whole population (NumPy's default), 1 — for a sample (unbiased variance)

np.std divides by n by default, while statistics textbooks and pandas' Series.std divide by n − 1. Pass ddof=1 to get the same result.

NumPy: vectorization and linear algebra

Open lesson
zᵢⱼ = (xᵢⱼ − μⱼ) / σⱼzᵢⱼ = (xᵢⱼ − μⱼ) / σⱼ
where:
  • xᵢⱼfeature j of example i
  • μⱼ, σⱼmean and standard deviation of column j
  • zᵢⱼstandardized value (z-score): mean 0, standard deviation 1
cᵢⱼ = ∑ₖ₌₁ⁿ aᵢₖ · bₖⱼ (m × n) @ (n × p) → (m × p)
where:
  • aᵢₖelement in row i, column k of A
  • bₖⱼelement in row k, column j of B
  • nthe shared number of columns of A and rows of B
A = [[a, b], [c, d]]: det A = a·d − b·c A⁻¹ = (1 / det A) · [[d, −b], [−c, a]]A = [[a, b], [c, d]]: det A = a·d − b·c A⁻¹ = (1 / det A) · [[d, −b], [−c, a]]
where:
  • det Adeterminant; if it is 0, the matrix is singular and has no inverse
  • A⁻¹inverse matrix: A · A⁻¹ = I (identity matrix)
A · v = λ · v det(A − λ · I) = 0
where:
  • veigenvector (v ≠ 0)
  • λeigenvalue
  • Iidentity matrix

The second equation is the characteristic equation: the system (A − λI)v = 0 has a non-zero solution only when the determinant is 0.

w = (Xᵀ X)⁻¹ Xᵀ y
where:
  • Xn × d feature matrix (the first column is ones)
  • yvector of target values
  • wcoefficients that minimise the sum of squared errors

pandas basics: Series and DataFrame

Open lesson
h = (n − 1) · q Q(q) = x₍ₖ₎ + (h − k) · (x₍ₖ₊₁₎ − x₍ₖ₎), k = ⌊h⌋
where:
  • qthe level: 0.25 (first quartile), 0.5 (median), 0.75 (third quartile)
  • x₍ₖ₎the k-th value after sorting in ascending order (counting from 0)
  • hthe “fractional position” in the sorted data

By default pandas and NumPy compute quantiles by linear interpolation between two neighbouring values.

Data analysis with pandas

Open lesson
x̄_w = ∑ wᵢ · xᵢ / ∑ wᵢx̄_w = ∑ wᵢ · xᵢ / ∑ wᵢ
where:
  • xᵢthe mean of group i (for example, the average receipt)
  • wᵢthe weight of the group — usually the number of observations

The weighted mean. In pandas: np.average(x, weights=w), or simply divide the overall total by the overall count.

Δ% = (xₜ − xₜ₋₁) / xₜ₋₁ · 100%Δ% = (xₜ − xₜ₋₁) / xₜ₋₁ · 100%
where:
  • xₜthe value in the current period
  • xₜ₋₁the value in the previous period

The growth rate; in pandas pct_change() computes it (without multiplying by 100). The first period has no previous value, so it gets NaN.

Visualization with matplotlib

Open lesson
k = ⌈log₂ n⌉ + 1 h = 2 · IQR · n^(−1/3)
where:
  • knumber of bins (Sturges' rule)
  • hbin width (Freedman–Diaconis rule), in the data's units
  • nnumber of observations
  • IQRinterquartile range Q3 − Q1

Sturges suits roughly normal, not too large datasets; Freedman–Diaconis is robust to outliers. ax.hist(x, bins='fd') applies the second one automatically.

fᵢ = nᵢ / (n · h) ∑ fᵢ · h = 1fᵢ = nᵢ / (n · h) ∑ fᵢ · h = 1
where:
  • nᵢobservations in bin i
  • fᵢbar height with density=True (density)
  • hbin width

ax.hist(x, density=True) makes the bar areas add up to 1 — such a histogram can be compared with a probability density curve.

Statistics with Python

Open lesson
x̄ = (1/n) · ∑ xᵢ s² = ∑ (xᵢ − x̄)² / (n − 1) s = √s²x̄ = (1/n) · ∑ xᵢ s² = ∑ (xᵢ − x̄)² / (n − 1) s = √s²
where:
  • nsample size
  • s²sample variance (its unit is the square of the data's unit)
  • ssample standard deviation (same unit as the data); in NumPy std(ddof=1)
[Q1 − 1.5 · IQR, Q3 + 1.5 · IQR]
where:
  • Q1, Q3first and third quartile (25% and 75%)
  • IQRQ3 − Q1, the width of the middle half of the data

Tukey's rule: values outside this interval count as outliers. Here: 14.75 + 1.5 · 1.75 = 17.375, so 45 is an outlier.

f(x) = 1 / (σ · √(2π)) · e^(−(x − μ)² / (2σ²)) z = (x − μ) / σf(x) = 1 / (σ · √(2π)) · e^(−(x − μ)² / (2σ²)) z = (x − μ) / σ
where:
  • μmean of the distribution
  • σstandard deviation (σ > 0)
  • f(x)probability density; P(a < X < b) is the area under the curve between a and b
  • zz-score: the position in the standard normal distribution (μ = 0, σ = 1)
SE = σ / √n ≈ s / √nSE = σ / √n ≈ s / √n
where:
  • SEstandard error of the mean
  • σ, spopulation and sample standard deviation
x̄ ± t* · s / √n, t* = t₁₋α/₂(n − 1)x̄ ± t* · s / √n, t* = t₁₋α/₂(n − 1)
where:
  • t*critical value of the t-distribution; α = 0.05 for 95%
  • n − 1degrees of freedom

A 95% confidence interval for the mean. stats.t.ppf(0.975, df=n - 1) returns t*, and stats.t.interval returns the whole interval.

t = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂)t = (x̄₁ − x̄₂) / √(s₁²/n₁ + s₂²/n₂)
where:
  • x̄₁, x̄₂group means
  • s₁², s₂²sample variances of the groups
  • n₁, n₂group sizes

Welch's t statistic: the difference of the means divided by its standard error. The larger |t|, the smaller p.

r = ∑ (xᵢ − x̄)(yᵢ − ȳ) / √( ∑ (xᵢ − x̄)² · ∑ (yᵢ − ȳ)² )r = ∑ (xᵢ − x̄)(yᵢ − ȳ) / √( ∑ (xᵢ − x̄)² · ∑ (yᵢ − ȳ)² )
where:
  • r−1 ≤ r ≤ 1; the sign gives the direction, |r| the strength
  • x̄, ȳthe mean of each variable

Machine learning with scikit-learn

Open lesson
MSE = (1/n) · ∑ (yᵢ − ŷᵢ)² RMSE = √MSE R² = 1 − ∑ (yᵢ − ŷᵢ)² / ∑ (yᵢ − ȳ)²MSE = (1/n) · ∑ (yᵢ − ŷᵢ)² RMSE = √MSE R² = 1 − ∑ (yᵢ − ŷᵢ)² / ∑ (yᵢ − ȳ)²
where:
  • yᵢ, ŷᵢtrue value and prediction
  • RMSEtypical error in the target's unit (here thousands of manat)
  • R²share of variation explained: 1 is perfect, 0 is the level of a model that always predicts the mean

The model found the coefficients 1.48 and 11.06 — close to the true 1.5 and 12. RMSE ≈ 9.8 is the level of the noise we added (σ = 10): the model captured the signal completely, and the noise cannot be predicted.

p = σ(z) = 1 / (1 + e^(−z)), z = w · x + bp = σ(z) = 1 / (1 + e^(−z)), z = w · x + b
where:
  • σ(z)sigmoid: maps any number into (0, 1), σ(0) = 0.5
  • pprobability that the example belongs to the positive class; “1” is predicted when p ≥ 0.5
G = 1 − ∑ₖ pₖ² Gsplit = (nL / n) · GL + (nR / n) · GRG = 1 − ∑ₖ pₖ² Gsplit = (nL / n) · GL + (nR / n) · GR
where:
  • pₖshare of class k in the node
  • G0 means a pure node (one class); the maximum for two classes is 0.5
  • nL, nRnumber of examples in the left and right child nodes
Accuracy = (TP + TN) / N Precision = TP / (TP + FP) Recall = TP / (TP + FN) F1 = 2 · P · R / (P + R)Accuracy = (TP + TN) / N Precision = TP / (TP + FP) Recall = TP / (TP + FN) F1 = 2 · P · R / (P + R)
where:
  • Precisionwhat share of the messages we called spam really are spam
  • Recallwhat share of all spam we managed to catch
  • F1the harmonic mean of precision and recall
  • Ntotal number of examples
CV = (1/k) · ∑ᵢ₌₁ᵏ scoreᵢCV = (1/k) · ∑ᵢ₌₁ᵏ scoreᵢ
where:
  • knumber of folds, usually 5 or 10
  • scoreᵢthe score when fold i is the test set

5AI and deep learning with PyTorchUniversity

Machine learning concepts

Open lesson
MSE = (1/n) · ∑ᵢ₌₁ⁿ (yᵢ − ŷᵢ)²MSE = (1/n) · ∑ᵢ₌₁ⁿ (yᵢ − ŷᵢ)²
where:
  • nnumber of examples
  • yᵢtrue value (label) of example i
  • ŷᵢthe model's prediction

Mean squared error. Squaring punishes large errors much more; the unit is the square of the label's unit.

CE = −∑ₖ₌₁ᴷ yₖ · ln pₖ = −ln pₜ
where:
  • Knumber of classes
  • yₖ1 if class k is the correct one, otherwise 0 (one-hot)
  • pₖprobability the model gives to class k
  • pₜprobability of the true class

Cross-entropy. For two classes: BCE = −[y · ln p + (1 − y) · ln(1 − p)]. As the probability of the true class approaches 1, the loss approaches 0.

θ ← θ − η · ∇L(θ)
where:
  • θall parameters of the model (weights and biases)
  • ηlearning rate, usually 0.0001–0.1
  • ∇L(θ)gradient: the vector of partial derivatives of the loss with respect to each parameter

The gradient descent update rule. Each repetition is one step; one full pass over the training set is called an epoch.

∂L/∂w = (2/n) · ∑ᵢ (ŷᵢ − yᵢ) · xᵢ ∂L/∂b = (2/n) · ∑ᵢ (ŷᵢ − yᵢ)∂L/∂w = (2/n) · ∑ᵢ (ŷᵢ − yᵢ) · xᵢ ∂L/∂b = (2/n) · ∑ᵢ (ŷᵢ − yᵢ)
where:
  • w, bslope (weight) and intercept (bias) of the line
  • ŷᵢ − yᵢerror on example i
E[(y − f̂(x))²] = Bias[f̂(x)]² + Var[f̂(x)] + σ²
where:
  • f̂(x)prediction of a model trained on a random training set
  • Biasbias: the average error caused by the model's too-simple assumptions
  • Varvariance: how much the prediction depends on the particular training set
  • σ²irreducible noise in the data

The bias–variance decomposition of the MSE. As a model becomes more complex, bias falls and variance grows; the best model minimises the sum.

J(θ) = L(θ) + λ · ∑ⱼ θⱼ²
where:
  • λregularisation strength (a hyperparameter chosen on the validation set)
  • ∑ⱼ θⱼ²sum of squared weights — penalises large weights

L2 regularisation (weight decay): the model prefers “smoother” functions and is less inclined to memorise noise.

PyTorch: tensors

Open lesson
numel = d₁ · d₂ · … · dₖ memory = numel · s
where:
  • d₁ … dₖthe sizes in the shape
  • numelnumber of elements (x.numel())
  • ssize of one element in bytes (x.element_size())
C = A @ B, Cᵢⱼ = ∑ₖ Aᵢₖ · Bₖⱼ, (m × n) @ (n × p) → (m × p)
where:
  • A, Bmatrices of sizes m × n and n × p
  • nthe inner size — must be the same in both matrices
  • Cᵢⱼdot product of row i of A and column j of B

Matrix multiplication is the core operation of neural networks: multiplying (N, D) inputs by a (D, K) weight matrix processes the whole batch at once. For tensors with more axes, @ works on the last two axes and treats the others as batch axes.

PyTorch: autograd and automatic differentiation

Open lesson
dL/dx = dL/dy · dy/dx ∂L/∂w = ∂L/∂ŷ · ∂ŷ/∂wdL/dx = dL/dy · dy/dx ∂L/∂w = ∂L/∂ŷ · ∂ŷ/∂w
where:
  • Lthe loss (root of the graph)
  • y, ŷintermediate values (inner nodes of the graph)
  • x, wleaves: inputs and parameters

The derivative of a composite function is the product of local derivatives. If a variable affects the loss along several paths, the products along all paths are added: ∂L/∂x = ∑ₖ ∂L/∂yₖ · ∂yₖ/∂x.

σ(z) = 1 / (1 + e⁻ᶻ) σ′(z) = σ(z) · (1 − σ(z))σ(z) = 1 / (1 + e⁻ᶻ) σ′(z) = σ(z) · (1 − σ(z))
where:
  • σ(z)the sigmoid function, with values in (0, 1)
  • σ′(z)its derivative; its maximum is 0.25 at z = 0

Proof: σ′(z) = e⁻ᶻ / (1 + e⁻ᶻ)² = σ(z) · e⁻ᶻ / (1 + e⁻ᶻ), and the last fraction equals 1 − σ(z).

PyTorch: building neural networks

Open lesson
z = w · x + b = ∑ᵢ₌₁ⁿ wᵢxᵢ + b a = φ(z)
where:
  • xinput vector (n features)
  • wweights — the importance of each input
  • bbias — shifts the activation threshold
  • φactivation function (ReLU, sigmoid…)
  • athe neuron's output (activation)
ReLU(z) = max(0, z) σ(z) = 1 / (1 + e⁻ᶻ) tanh(z) = (eᶻ − e⁻ᶻ) / (eᶻ + e⁻ᶻ)ReLU(z) = max(0, z) σ(z) = 1 / (1 + e⁻ᶻ) tanh(z) = (eᶻ − e⁻ᶻ) / (eᶻ + e⁻ᶻ)
where:
  • ReLUzeroes negative values; its derivative is 1 for z > 0 and 0 for z < 0
  • σsigmoid: values in (0, 1), read as a probability
  • tanhhyperbolic tangent: values in (−1, 1), symmetric around zero
softmax(z)ᵢ = eᶻⁱ / ∑ⱼ₌₁ᴷ eᶻʲsoftmax(z)ᵢ = eᶻⁱ / ∑ⱼ₌₁ᴷ eᶻʲ
where:
  • zraw scores (logits) for K classes
  • softmax(z)ᵢprobability of class i; all are positive and sum to 1

Softmax turns the logits of the last layer into a probability distribution in multi-class classification.

Y = X · Wᵀ + b params = (nin + 1) · nout
where:
  • Xbatch of inputs, shape (N, nin)
  • Wweights, shape (nout, nin) — layer.weight
  • bbiases, shape (nout,) — layer.bias
  • Youtputs, shape (N, nout)

Each output neuron has nin weights and one bias, so a fully connected layer has (nin + 1) · nout parameters.

PyTorch: the training loop

Open lesson
steps per epoch = ⌈N / B⌉ updates = epochs · ⌈N / B⌉steps per epoch = ⌈N / B⌉ updates = epochs · ⌈N / B⌉
where:
  • Nnumber of training examples
  • Bbatch size
  • ⌈ ⌉rounding up: the last incomplete batch is a step too
L = −zₜ + ln ∑ⱼ₌₁ᴷ eᶻʲ
where:
  • zthe model's raw logits
  • tindex of the true class

nn.CrossEntropyLoss combines log-softmax and cross-entropy: −ln softmax(z)ₜ = −zₜ + ln ∑ eᶻʲ. For a batch, the losses are averaged.

mₜ = β₁·mₜ₋₁ + (1 − β₁)·gₜ vₜ = β₂·vₜ₋₁ + (1 − β₂)·gₜ² θₜ = θₜ₋₁ − η · m̂ₜ / (√v̂ₜ + ε)mₜ = β₁·mₜ₋₁ + (1 − β₁)·gₜ vₜ = β₂·vₜ₋₁ + (1 − β₂)·gₜ² θₜ = θₜ₋₁ − η · m̂ₜ / (√v̂ₜ + ε)
where:
  • gₜthe gradient at step t
  • mₜ, vₜrunning averages of the gradient and of its square (momentum and scale)
  • m̂ₜ, v̂ₜbias-corrected values: mₜ/(1 − β₁ᵗ), vₜ/(1 − β₂ᵗ)
  • β₁, β₂, ε, ηPyTorch defaults 0.9, 0.999, 10⁻⁸, 0.001

The Adam optimizer. Every parameter gets its own step size: a parameter with consistently large gradients is updated cautiously, one with small gradients more boldly. SGD with momentum works as v ← μ·v + g, θ ← θ − η·v.

accuracy = (1/N) · ∑ᵢ [argmax(zᵢ) = yᵢ]accuracy = (1/N) · ∑ᵢ [argmax(zᵢ) = yᵢ]
where:
  • argmax(zᵢ)the class with the largest logit for example i — the prediction
  • [ … ]1 if the condition holds, otherwise 0

Convolutional neural networks (CNNs)

Open lesson
x′ = (x − μ) / σx′ = (x − μ) / σ
where:
  • xpixel value in [0, 1]
  • μ, σmean and standard deviation of the channel over the training set (0.1307 and 0.3081 for MNIST)

After normalisation the inputs are roughly centred on zero with unit scale — this speeds up gradient descent.

Y[i, j] = ∑ₘ ∑ₙ X[i + m, j + n] · K[m, n] + b
where:
  • Xinput image (or the previous layer's map)
  • Ka K × K filter — learned weights
  • bthe filter's bias
  • Y[i, j]element (i, j) of the feature map

2D convolution for one channel. With several input channels, the sum also runs over the channels.

O = ⌊(W − K + 2P) / S⌋ + 1O = ⌊(W − K + 2P) / S⌋ + 1
where:
  • Winput width (or height), in pixels
  • Kkernel size
  • Ppadding: zero pixels added around the border
  • Sstride: how many pixels the filter moves
  • Ooutput width; ⌊ ⌋ means rounding down

The output-size formula; the same holds for the height. For pooling, usually P = 0 and S = K.

params = (K · K · Cin + 1) · Cout
where:
  • Cin, Coutnumber of input and output channels

A convolutional layer's parameter count does not depend on the image size — the weights are shared across all positions.

Transformers and large language models

Open lesson
Q = X · WQ K = X · WK V = X · WV
where:
  • Xtoken vectors of shape (T, d)
  • WQ, WK, WVlearned (d, dₖ) matrices
Attention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · VAttention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · V
where:
  • Q · Kᵀa (T, T) score matrix: similarity of every query with every key
  • √dₖscaling factor; dₖ is the key dimension
  • softmaxapplied row by row: attention weights are positive and sum to 1
  • Vthe values; the output is their weighted average

Scaled dot-product attention.

MultiHead(X) = Concat(head₁, …, headₕ) · WO, headᵢ = Attention(X·WQⁱ, X·WKⁱ, X·WVⁱ)
where:
  • hnumber of heads; each head has size d / h
  • WOa (d, d) matrix that mixes the concatenated heads

Each head can learn to attend to a different relationship: one to grammatical links, another to what a pronoun refers to. Parameters: 4 · d² weights and 4 · d biases for WQ, WK, WV, WO.

PE(pos, 2i) = sin(pos / 10000^(2i/d)) PE(pos, 2i + 1) = cos(pos / 10000^(2i/d))PE(pos, 2i) = sin(pos / 10000^(2i/d)) PE(pos, 2i + 1) = cos(pos / 10000^(2i/d))
where:
  • posposition of the token: 0, 1, 2, …
  • iindex of the pair of components; small i oscillates fast, large i slowly
  • dembedding dimension
L = −(1/T) · ∑ₜ₌₁ᵀ ln p(xₜ₊₁ | x₁, …, xₜ) PPL = eᴸL = −(1/T) · ∑ₜ₌₁ᵀ ln p(xₜ₊₁ | x₁, …, xₜ) PPL = eᴸ
where:
  • xₜthe t-th token of the text
  • p(xₜ₊₁ | x₁, …, xₜ)probability the model gives to the correct next token given the previous ones
  • PPLperplexity; the lower, the better the model
pᵢ = exp(zᵢ / τ) / ∑ⱼ exp(zⱼ / τ)pᵢ = exp(zᵢ / τ) / ∑ⱼ exp(zⱼ / τ)
where:
  • τtemperature: τ < 1 gives more confident, repetitive text, τ > 1 more varied and risky text

Sampling the next token with a temperature softmax.