- Explain how text becomes tokens and embedding vectors
- Compute self-attention by formula and in code and know the role of the mask and √dₖ
- Explain how LLMs are trained (next token, fine-tuning) and what perplexity means
- Use pretrained models and recognise the risks addressed by responsible AI
Chat assistants, machine translation, tools that help you write code — at the heart of all of them is the transformer architecture. It was introduced in 2017 in the paper “Attention Is All You Need” (Vaswani et al.). A large language model (LLM) is essentially a very large transformer trained on one simple task: predicting the next token in a text. In this lesson we look inside this idea and compute the attention mechanism with our own hands.
Tokens and embeddings
A neural network works with numbers, not letters. So text is first split into tokens — words, parts of words or characters — and every token gets an id in a vocabulary. Modern models use subword tokenisation (for example BPE): frequent words become a single token, while rare words are split into pieces. Vocabularies usually hold tens of thousands of tokens; GPT-2, for example, has 50,257.
text = 'the cat sat on the mat'
words = text.split()
vocab = sorted(set(words))
stoi = {w: i for i, w in enumerate(vocab)}
ids = [stoi[w] for w in words]
print(vocab)
print(ids)▸ Expected output
['cat', 'mat', 'on', 'sat', 'the'] [4, 0, 3, 2, 4, 1]
A token id carries no meaning by itself, so each id is mapped to a learned embedding vector. This is a V × d table: row i is the vector of token i. During training, tokens with similar meanings end up with nearby vectors.
import torch
from torch import nn
torch.manual_seed(42)
emb = nn.Embedding(num_embeddings=5, embedding_dim=4)
ids = torch.tensor([4, 0, 3, 2, 4, 1])
vectors = emb(ids)
print(vectors.shape, emb.weight.shape)
print(torch.equal(vectors[0], vectors[4]))torch.Size([6, 4]) torch.Size([5, 4]) True
Self-attention
The word “bank” can mean a river bank or a financial institution — the context decides which. In self-attention every token “looks at” the other tokens and takes a weighted mix of their information. To do this, each token produces three vectors: a query (“what am I looking for?”), a key (“what do I contain?”) and a value (“what do I pass on?”). The larger the dot product of a query and a key, the more that token's value contributes to the result.
- Xtoken vectors of shape (T, d)
- WQ, WK, WVlearned (d, dₖ) matrices
- Q · Kᵀa (T, T) score matrix: similarity of every query with every key
- √dₖscaling factor; dₖ is the key dimension
- softmaxapplied row by row: attention weights are positive and sum to 1
- Vthe values; the output is their weighted average
Scaled dot-product attention.
Why divide by √dₖ? If the components of q and k are independent with mean 0 and variance 1, the sum q · k = ∑ qᵢkᵢ has variance dₖ. With dₖ = 64 the scores spread around ±8, softmax becomes almost “one-hot” and the gradients vanish. Dividing by √dₖ brings the variance back to 1.
dₖ = 4. Query q = (1, 0, 1, 0); keys k₁ = (1, 0, 1, 0), k₂ = (0, 1, 0, 1); values v₁ = (1, 0), v₂ = (0, 1). Find the attention output.
Show solutionHide solution
Divide by √dₖ = 2: (1, 0).
softmax: e / (e + 1) ≈ 0.731 and 1 / (e + 1) ≈ 0.269.
Output: 0.731 · v₁ + 0.269 · v₂ = (0.731, 0.269) — more information was taken from the first token, which matches the query better.
In text-generating models such as GPT, a token must not look into the future — otherwise it would “see” the answer it has to predict. So the scores above the diagonal are set to −∞ (a causal mask); after softmax their weight is 0. The code below computes self-attention for 4 tokens completely in NumPy:
import numpy as np
def softmax(z):
e = np.exp(z - z.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
rng = np.random.default_rng(4)
T, d = 4, 8
X = rng.normal(size=(T, d))
Wq, Wk, Wv = (rng.normal(size=(d, d)) / np.sqrt(d) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv
scores = Q @ K.T / np.sqrt(d)
mask = np.triu(np.ones((T, T), dtype=bool), k=1)
scores[mask] = -np.inf
A = softmax(scores)
print(np.round(A, 2))
print(A.sum(axis=1))
print((A @ V).shape)▸ Expected output
[[1. 0. 0. 0. ] [0.82 0.18 0. 0. ] [0.14 0.47 0.39 0. ] [0.23 0.66 0.03 0.08]] [1. 1. 1. 1.] (4, 8)
Multi-head attention, positional encoding and the transformer block
- hnumber of heads; each head has size d / h
- WOa (d, d) matrix that mixes the concatenated heads
Each head can learn to attend to a different relationship: one to grammatical links, another to what a pronoun refers to. Parameters: 4 · d² weights and 4 · d biases for WQ, WK, WV, WO.
import torch
from torch import nn
torch.manual_seed(42)
mha = nn.MultiheadAttention(embed_dim=64, num_heads=8, batch_first=True)
x = torch.randn(2, 10, 64)
out, weights = mha(x, x, x)
print(out.shape, weights.shape)
print(sum(p.numel() for p in mha.parameters()))torch.Size([2, 10, 64]) torch.Size([2, 10, 10]) 16640
Attention by itself does not know the order of tokens: “the dog chased the cat” and “the cat chased the dog” contain the same tokens. So a positional encoding is added to the embeddings. The original transformer used sines and cosines of different frequencies; models like GPT-2 learn position vectors, and many modern LLMs use rotary encoding (RoPE).
- posposition of the token: 0, 1, 2, …
- iindex of the pair of components; small i oscillates fast, large i slowly
- dembedding dimension
import numpy as np
def positional_encoding(T, d):
pos = np.arange(T)[:, None]
i = np.arange(0, d, 2)[None, :]
angles = pos / 10000 ** (i / d)
pe = np.zeros((T, d))
pe[:, 0::2] = np.sin(angles)
pe[:, 1::2] = np.cos(angles)
return pe
print(np.round(positional_encoding(T=5, d=4), 3))▸ Expected output
[[ 0. 1. 0. 1. ] [ 0.841 0.54 0.01 1. ] [ 0.909 -0.416 0.02 1. ] [ 0.141 -0.99 0.03 1. ] [-0.757 -0.654 0.04 0.999]]
A transformer block is built like this: x → LayerNorm → multi-head attention → residual connection (x + …) → LayerNorm → a two-layer MLP (usually 4d hidden neurons and GELU) → another residual connection. Residual connections help the gradient pass through dozens of blocks without vanishing. A model is N such blocks stacked. There are three main kinds:
| Architecture | Attention | Examples | Typical tasks |
|---|---|---|---|
| Encoder-only | bidirectional | BERT | text classification, search, named-entity recognition |
| Decoder-only | causal (masked) | GPT | text generation, chat, code |
| Encoder–decoder | bidirectional + causal + cross-attention | T5 | translation, summarisation |
How LLMs are trained
In pretraining, the model predicts the next token at every position of a huge text corpus. This is self-supervised learning: nobody writes labels, they are simply the next tokens of the text itself. The loss is the cross-entropy averaged over all positions; from it we get the perplexity — how many options, on average, the model “hesitates” between at each step.
- xₜthe t-th token of the text
- p(xₜ₊₁ | x₁, …, xₜ)probability the model gives to the correct next token given the previous ones
- PPLperplexity; the lower, the better the model
import torch
import torch.nn.functional as F
torch.manual_seed(42)
tokens = torch.tensor([4, 0, 3, 2, 4, 1])
inputs, targets = tokens[:-1], tokens[1:]
logits = torch.randn(len(inputs), 5)
loss = F.cross_entropy(logits, targets)
print(inputs.tolist(), targets.tolist())
print(round(loss.item(), 4), round(torch.exp(loss).item(), 2))
print(round(torch.log(torch.tensor(5.0)).item(), 4))[4, 0, 3, 2, 4] [0, 3, 2, 4, 1] 1.6522 5.22 1.6094
The vocabulary has 50,000 tokens. a) What are the loss and perplexity of a model that guesses blindly? b) A trained model has an average loss of 2.0. What is its perplexity and what does it mean?
Show solutionHide solution
b) PPL = e² ≈ 7.4: at each step the model is as uncertain as if it were choosing among about 7 equally likely options — a huge improvement over 50,000.
After pretraining the model can continue text, but it is not yet a good assistant. Next comes fine-tuning: supervised training on question–answer examples, followed by methods based on the answers people prefer (for example reinforcement learning from human feedback, RLHF). When adapting a model to your own task, you do not have to train all its billions of parameters: methods like LoRA train only small extra matrices. When generating text, the model outputs a distribution over the next token, a token is chosen, and the process repeats.
- τtemperature: τ < 1 gives more confident, repetitive text, τ > 1 more varied and risky text
Sampling the next token with a temperature softmax.
import numpy as np
z = np.array([2.0, 1.0, 0.0])
for T in [0.5, 1.0, 2.0]:
p = np.exp(z / T) / np.exp(z / T).sum()
print(T, np.round(p, 3))▸ Expected output
0.5 [0.867 0.117 0.016] 1.0 [0.665 0.245 0.09 ] 2.0 [0.506 0.307 0.186]
Pretrained models and responsible AI
In practice, instead of training models from scratch, people use pretrained ones. The Hugging Face Hub hosts thousands of open models, and the transformers library (pip install transformers) loads them in a few lines. This code does not run in the browser: the first run downloads the model weights (hundreds of megabytes).
from transformers import pipeline
classifier = pipeline('sentiment-analysis', model='distilbert/distilbert-base-uncased-finetuned-sst-2-english')
print(classifier('I love learning with Educora!'))
generator = pipeline('text-generation', model='openai-community/gpt2')
result = generator('Deep learning is', max_new_tokens=20)
print(result[0]['generated_text'])- Bias: models can repeat stereotypes present in their training texts — especially dangerous in decisions such as hiring or lending.
- Hallucinations: a model can invent convincing-sounding but false facts, quotes and sources.
- Privacy: do not type personal data, passwords or confidential documents into online AI services; personal data in training data must be protected too.
- Copyright and licences: every model and dataset has terms of use — read them.
- Energy and resources: training and running large models requires a lot of electricity and computing power.
Key points
- Text is split into tokens, each token becomes a learned embedding vector, and position information is added.
- Attention(Q, K, V) = softmax(QKᵀ/√dₖ)·V: each token takes a weighted average of the others' values; √dₖ keeps softmax from saturating.
- A causal mask hides future tokens in a decoder; multi-head attention learns different relationships in parallel.
- An LLM is trained to predict the next token: L = −(1/T)∑ ln p(xₜ₊₁ | x≤ₜ), PPL = eᴸ; then it is fine-tuned.
- When using pretrained models, keep bias, hallucinations, privacy and licences in mind.
Check yourself
10 questions. Every correct answer earns XP.