Formulas & shortcuts
Python · 97
Every formula in this course and the easiest ways to remember them, on one page.
1Functions and modulesIntermediate
Functions
Open lesson- Ftemperature in degrees Fahrenheit
- Ctemperature in degrees Celsius
Let's turn this formula into a function
2Python in depthAdvanced
Iterators and generators
Open lessonDecorators and closures
Open lessonContext managers and the with statement
Open lessonType hints and dataclasses
Open lessonAdvanced OOP: special methods, properties and the MRO
Open lessonRegular expressions with re
Open lessonAsynchronous Python: async and await
Open lessonPackages, modules and virtual environments
Open lessonPerformance, Big-O and the collections module
Open lesson3Web and automationAdvanced
Working with JSON and CSV
Open lessonWeb APIs and the requests library
Open lessonWeb services with FastAPI
Open lessonAutomation scripts
Open lesson4Data science: NumPy, pandas, matplotlibUniversity
NumPy basics: arrays
Open lesson- nnumber of axes (
ndim) - dₖlength along axis k (an element of
shape) - itemsizesize of one element in bytes: 8 for
float64andint64, 4 forfloat32andint32, 1 forboolandint8
For m with shape (2, 3): size = 2 · 3 = 6. For temps: nbytes = 4 · 8 = 32 bytes.
- hthe spacing between
linspacepoints - nnumber of elements produced by
arange - ⌈ ⌉rounding up (ceiling)
np.linspace(0, 1, 5): h = 1 / 4 = 0.25. np.arange(0, 10, 2): n = ⌈10 / 2⌉ = 5 elements.
- x̄arithmetic mean (
mean) - nnumber of values
- σstandard deviation (
std): how far the values typically are from the mean - ddof0 — for a whole population (NumPy's default), 1 — for a sample (unbiased variance)
np.std divides by n by default, while statistics textbooks and pandas' Series.std divide by n − 1. Pass ddof=1 to get the same result.
NumPy: vectorization and linear algebra
Open lesson- xᵢⱼfeature j of example i
- μⱼ, σⱼmean and standard deviation of column j
- zᵢⱼstandardized value (z-score): mean 0, standard deviation 1
- aᵢₖelement in row i, column k of A
- bₖⱼelement in row k, column j of B
- nthe shared number of columns of A and rows of B
- det Adeterminant; if it is 0, the matrix is singular and has no inverse
- A⁻¹inverse matrix: A · A⁻¹ = I (identity matrix)
- veigenvector (v ≠ 0)
- λeigenvalue
- Iidentity matrix
The second equation is the characteristic equation: the system (A − λI)v = 0 has a non-zero solution only when the determinant is 0.
- Xn × d feature matrix (the first column is ones)
- yvector of target values
- wcoefficients that minimise the sum of squared errors
pandas basics: Series and DataFrame
Open lesson- qthe level: 0.25 (first quartile), 0.5 (median), 0.75 (third quartile)
- x₍ₖ₎the k-th value after sorting in ascending order (counting from 0)
- hthe “fractional position” in the sorted data
By default pandas and NumPy compute quantiles by linear interpolation between two neighbouring values.
Data analysis with pandas
Open lesson- xᵢthe mean of group i (for example, the average receipt)
- wᵢthe weight of the group — usually the number of observations
The weighted mean. In pandas: np.average(x, weights=w), or simply divide the overall total by the overall count.
- xₜthe value in the current period
- xₜ₋₁the value in the previous period
The growth rate; in pandas pct_change() computes it (without multiplying by 100). The first period has no previous value, so it gets NaN.
Visualization with matplotlib
Open lesson- knumber of bins (Sturges' rule)
- hbin width (Freedman–Diaconis rule), in the data's units
- nnumber of observations
- IQRinterquartile range Q3 − Q1
Sturges suits roughly normal, not too large datasets; Freedman–Diaconis is robust to outliers. ax.hist(x, bins='fd') applies the second one automatically.
- nᵢobservations in bin i
- fᵢbar height with
density=True(density) - hbin width
ax.hist(x, density=True) makes the bar areas add up to 1 — such a histogram can be compared with a probability density curve.
Statistics with Python
Open lesson- nsample size
- s²sample variance (its unit is the square of the data's unit)
- ssample standard deviation (same unit as the data); in NumPy
std(ddof=1)
- Q1, Q3first and third quartile (25% and 75%)
- IQRQ3 − Q1, the width of the middle half of the data
Tukey's rule: values outside this interval count as outliers. Here: 14.75 + 1.5 · 1.75 = 17.375, so 45 is an outlier.
- μmean of the distribution
- σstandard deviation (σ > 0)
- f(x)probability density; P(a < X < b) is the area under the curve between a and b
- zz-score: the position in the standard normal distribution (μ = 0, σ = 1)
- SEstandard error of the mean
- σ, spopulation and sample standard deviation
- t*critical value of the t-distribution; α = 0.05 for 95%
- n − 1degrees of freedom
A 95% confidence interval for the mean. stats.t.ppf(0.975, df=n - 1) returns t*, and stats.t.interval returns the whole interval.
- x̄₁, x̄₂group means
- s₁², s₂²sample variances of the groups
- n₁, n₂group sizes
Welch's t statistic: the difference of the means divided by its standard error. The larger |t|, the smaller p.
- r−1 ≤ r ≤ 1; the sign gives the direction, |r| the strength
- x̄, ȳthe mean of each variable
Machine learning with scikit-learn
Open lesson- yᵢ, ŷᵢtrue value and prediction
- RMSEtypical error in the target's unit (here thousands of manat)
- R²share of variation explained: 1 is perfect, 0 is the level of a model that always predicts the mean
The model found the coefficients 1.48 and 11.06 — close to the true 1.5 and 12. RMSE ≈ 9.8 is the level of the noise we added (σ = 10): the model captured the signal completely, and the noise cannot be predicted.
- σ(z)sigmoid: maps any number into (0, 1), σ(0) = 0.5
- pprobability that the example belongs to the positive class; “1” is predicted when p ≥ 0.5
- pₖshare of class k in the node
- G0 means a pure node (one class); the maximum for two classes is 0.5
- nL, nRnumber of examples in the left and right child nodes
- Precisionwhat share of the messages we called spam really are spam
- Recallwhat share of all spam we managed to catch
- F1the harmonic mean of precision and recall
- Ntotal number of examples
- knumber of folds, usually 5 or 10
- scoreᵢthe score when fold i is the test set
5AI and deep learning with PyTorchUniversity
Machine learning concepts
Open lesson- nnumber of examples
- yᵢtrue value (label) of example i
- ŷᵢthe model's prediction
Mean squared error. Squaring punishes large errors much more; the unit is the square of the label's unit.
- Knumber of classes
- yₖ1 if class k is the correct one, otherwise 0 (one-hot)
- pₖprobability the model gives to class k
- pₜprobability of the true class
Cross-entropy. For two classes: BCE = −[y · ln p + (1 − y) · ln(1 − p)]. As the probability of the true class approaches 1, the loss approaches 0.
- θall parameters of the model (weights and biases)
- ηlearning rate, usually 0.0001–0.1
- ∇L(θ)gradient: the vector of partial derivatives of the loss with respect to each parameter
The gradient descent update rule. Each repetition is one step; one full pass over the training set is called an epoch.
- w, bslope (weight) and intercept (bias) of the line
- ŷᵢ − yᵢerror on example i
- f̂(x)prediction of a model trained on a random training set
- Biasbias: the average error caused by the model's too-simple assumptions
- Varvariance: how much the prediction depends on the particular training set
- σ²irreducible noise in the data
The bias–variance decomposition of the MSE. As a model becomes more complex, bias falls and variance grows; the best model minimises the sum.
- λregularisation strength (a hyperparameter chosen on the validation set)
- ∑ⱼ θⱼ²sum of squared weights — penalises large weights
L2 regularisation (weight decay): the model prefers “smoother” functions and is less inclined to memorise noise.
PyTorch: tensors
Open lesson- d₁ … dₖthe sizes in the shape
- numelnumber of elements (
x.numel()) - ssize of one element in bytes (
x.element_size())
- A, Bmatrices of sizes m × n and n × p
- nthe inner size — must be the same in both matrices
- Cᵢⱼdot product of row i of A and column j of B
Matrix multiplication is the core operation of neural networks: multiplying (N, D) inputs by a (D, K) weight matrix processes the whole batch at once. For tensors with more axes, @ works on the last two axes and treats the others as batch axes.
PyTorch: autograd and automatic differentiation
Open lesson- Lthe loss (root of the graph)
- y, ŷintermediate values (inner nodes of the graph)
- x, wleaves: inputs and parameters
The derivative of a composite function is the product of local derivatives. If a variable affects the loss along several paths, the products along all paths are added: ∂L/∂x = ∑ₖ ∂L/∂yₖ · ∂yₖ/∂x.
- σ(z)the sigmoid function, with values in (0, 1)
- σ′(z)its derivative; its maximum is 0.25 at z = 0
Proof: σ′(z) = e⁻ᶻ / (1 + e⁻ᶻ)² = σ(z) · e⁻ᶻ / (1 + e⁻ᶻ), and the last fraction equals 1 − σ(z).
PyTorch: building neural networks
Open lesson- xinput vector (n features)
- wweights — the importance of each input
- bbias — shifts the activation threshold
- φactivation function (ReLU, sigmoid…)
- athe neuron's output (activation)
- ReLUzeroes negative values; its derivative is 1 for z > 0 and 0 for z < 0
- σsigmoid: values in (0, 1), read as a probability
- tanhhyperbolic tangent: values in (−1, 1), symmetric around zero
- zraw scores (logits) for K classes
- softmax(z)ᵢprobability of class i; all are positive and sum to 1
Softmax turns the logits of the last layer into a probability distribution in multi-class classification.
- Xbatch of inputs, shape (N, nin)
- Wweights, shape (nout, nin) —
layer.weight - bbiases, shape (nout,) —
layer.bias - Youtputs, shape (N, nout)
Each output neuron has nin weights and one bias, so a fully connected layer has (nin + 1) · nout parameters.
PyTorch: the training loop
Open lesson- Nnumber of training examples
- Bbatch size
- ⌈ ⌉rounding up: the last incomplete batch is a step too
- zthe model's raw logits
- tindex of the true class
nn.CrossEntropyLoss combines log-softmax and cross-entropy: −ln softmax(z)ₜ = −zₜ + ln ∑ eᶻʲ. For a batch, the losses are averaged.
- gₜthe gradient at step t
- mₜ, vₜrunning averages of the gradient and of its square (momentum and scale)
- m̂ₜ, v̂ₜbias-corrected values: mₜ/(1 − β₁ᵗ), vₜ/(1 − β₂ᵗ)
- β₁, β₂, ε, ηPyTorch defaults 0.9, 0.999, 10⁻⁸, 0.001
The Adam optimizer. Every parameter gets its own step size: a parameter with consistently large gradients is updated cautiously, one with small gradients more boldly. SGD with momentum works as v ← μ·v + g, θ ← θ − η·v.
- argmax(zᵢ)the class with the largest logit for example i — the prediction
- [ … ]1 if the condition holds, otherwise 0
Convolutional neural networks (CNNs)
Open lesson- xpixel value in [0, 1]
- μ, σmean and standard deviation of the channel over the training set (0.1307 and 0.3081 for MNIST)
After normalisation the inputs are roughly centred on zero with unit scale — this speeds up gradient descent.
- Xinput image (or the previous layer's map)
- Ka K × K filter — learned weights
- bthe filter's bias
- Y[i, j]element (i, j) of the feature map
2D convolution for one channel. With several input channels, the sum also runs over the channels.
- Winput width (or height), in pixels
- Kkernel size
- Ppadding: zero pixels added around the border
- Sstride: how many pixels the filter moves
- Ooutput width; ⌊ ⌋ means rounding down
The output-size formula; the same holds for the height. For pooling, usually P = 0 and S = K.
- Cin, Coutnumber of input and output channels
A convolutional layer's parameter count does not depend on the image size — the weights are shared across all positions.
Transformers and large language models
Open lesson- Xtoken vectors of shape (T, d)
- WQ, WK, WVlearned (d, dₖ) matrices
- Q · Kᵀa (T, T) score matrix: similarity of every query with every key
- √dₖscaling factor; dₖ is the key dimension
- softmaxapplied row by row: attention weights are positive and sum to 1
- Vthe values; the output is their weighted average
Scaled dot-product attention.
- hnumber of heads; each head has size d / h
- WOa (d, d) matrix that mixes the concatenated heads
Each head can learn to attend to a different relationship: one to grammatical links, another to what a pronoun refers to. Parameters: 4 · d² weights and 4 · d biases for WQ, WK, WV, WO.
- posposition of the token: 0, 1, 2, …
- iindex of the pair of components; small i oscillates fast, large i slowly
- dembedding dimension
- xₜthe t-th token of the text
- p(xₜ₊₁ | x₁, …, xₜ)probability the model gives to the correct next token given the previous ones
- PPLperplexity; the lower, the better the model
- τtemperature: τ < 1 gives more confident, repetitive text, τ > 1 more varied and risky text
Sampling the next token with a temperature softmax.