Skip to content
Educora
University25 min42 / 42

Transformers and large language models

Tokens and embeddings, self-attention Attention(Q, K, V) = softmax(QKᵀ/√dₖ)·V, multi-head attention, positional encoding, encoders and decoders, how LLMs learn by next-token prediction, fine-tuning, Hugging Face and responsible AI.

Check yourself
In this lesson you will learn
  • Explain how text becomes tokens and embedding vectors
  • Compute self-attention by formula and in code and know the role of the mask and √dₖ
  • Explain how LLMs are trained (next token, fine-tuning) and what perplexity means
  • Use pretrained models and recognise the risks addressed by responsible AI

Chat assistants, machine translation, tools that help you write code — at the heart of all of them is the transformer architecture. It was introduced in 2017 in the paper “Attention Is All You Need” (Vaswani et al.). A large language model (LLM) is essentially a very large transformer trained on one simple task: predicting the next token in a text. In this lesson we look inside this idea and compute the attention mechanism with our own hands.

Tokens and embeddings

A neural network works with numbers, not letters. So text is first split into tokens — words, parts of words or characters — and every token gets an id in a vocabulary. Modern models use subword tokenisation (for example BPE): frequent words become a single token, while rare words are split into pieces. Vocabularies usually hold tens of thousands of tokens; GPT-2, for example, has 50,257.

Python
text = 'the cat sat on the mat'
words = text.split()
vocab = sorted(set(words))
stoi = {w: i for i, w in enumerate(vocab)}
ids = [stoi[w] for w in words]
print(vocab)
print(ids)
▸ Expected output
['cat', 'mat', 'on', 'sat', 'the']
[4, 0, 3, 2, 4, 1]
A toy tokenizer: each word is one token. “the” appears twice and gets the same id (4) both times.

A token id carries no meaning by itself, so each id is mapped to a learned embedding vector. This is a V × d table: row i is the vector of token i. During training, tokens with similar meanings end up with nearby vectors.

Python
import torch
from torch import nn

torch.manual_seed(42)
emb = nn.Embedding(num_embeddings=5, embedding_dim=4)
ids = torch.tensor([4, 0, 3, 2, 4, 1])
vectors = emb(ids)
print(vectors.shape, emb.weight.shape)
print(torch.equal(vectors[0], vectors[4]))
Expected output
torch.Size([6, 4]) torch.Size([5, 4])
True
6 tokens → (6, 4) vectors. The same token (“the”) always gets the same vector — telling contexts apart is the job of the attention layers.

Self-attention

The word “bank” can mean a river bank or a financial institution — the context decides which. In self-attention every token “looks at” the other tokens and takes a weighted mix of their information. To do this, each token produces three vectors: a query (“what am I looking for?”), a key (“what do I contain?”) and a value (“what do I pass on?”). The larger the dot product of a query and a key, the more that token's value contributes to the result.

Q = X · WQ K = X · WK V = X · WV
where:
  • Xtoken vectors of shape (T, d)
  • WQ, WK, WVlearned (d, dₖ) matrices
Attention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · VAttention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · V
where:
  • Q · Kᵀa (T, T) score matrix: similarity of every query with every key
  • √dₖscaling factor; dₖ is the key dimension
  • softmaxapplied row by row: attention weights are positive and sum to 1
  • Vthe values; the output is their weighted average

Scaled dot-product attention.

Why divide by √dₖ? If the components of q and k are independent with mean 0 and variance 1, the sum q · k = ∑ qᵢkᵢ has variance dₖ. With dₖ = 64 the scores spread around ±8, softmax becomes almost “one-hot” and the gradients vanish. Dividing by √dₖ brings the variance back to 1.

Example 1: attention by hand

dₖ = 4. Query q = (1, 0, 1, 0); keys k₁ = (1, 0, 1, 0), k₂ = (0, 1, 0, 1); values v₁ = (1, 0), v₂ = (0, 1). Find the attention output.

Show solution
Scores: q · k₁ = 2, q · k₂ = 0.
Divide by √dₖ = 2: (1, 0).
softmax: e / (e + 1) ≈ 0.731 and 1 / (e + 1) ≈ 0.269.
Output: 0.731 · v₁ + 0.269 · v₂ = (0.731, 0.269) — more information was taken from the first token, which matches the query better.

In text-generating models such as GPT, a token must not look into the future — otherwise it would “see” the answer it has to predict. So the scores above the diagonal are set to −∞ (a causal mask); after softmax their weight is 0. The code below computes self-attention for 4 tokens completely in NumPy:

Python
import numpy as np

def softmax(z):
    e = np.exp(z - z.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

rng = np.random.default_rng(4)
T, d = 4, 8
X = rng.normal(size=(T, d))
Wq, Wk, Wv = (rng.normal(size=(d, d)) / np.sqrt(d) for _ in range(3))
Q, K, V = X @ Wq, X @ Wk, X @ Wv
scores = Q @ K.T / np.sqrt(d)
mask = np.triu(np.ones((T, T), dtype=bool), k=1)
scores[mask] = -np.inf
A = softmax(scores)
print(np.round(A, 2))
print(A.sum(axis=1))
print((A @ V).shape)
▸ Expected output
[[1.   0.   0.   0.  ]
 [0.82 0.18 0.   0.  ]
 [0.14 0.47 0.39 0.  ]
 [0.23 0.66 0.03 0.08]]
[1. 1. 1. 1.]
(4, 8)
The attention matrix is lower-triangular: the first token looks only at itself, the fourth at all four tokens. Every row sums to 1. Change the seed in the code and watch how the weights change.

Multi-head attention, positional encoding and the transformer block

MultiHead(X) = Concat(head₁, …, headₕ) · WO, headᵢ = Attention(X·WQⁱ, X·WKⁱ, X·WVⁱ)
where:
  • hnumber of heads; each head has size d / h
  • WOa (d, d) matrix that mixes the concatenated heads

Each head can learn to attend to a different relationship: one to grammatical links, another to what a pronoun refers to. Parameters: 4 · d² weights and 4 · d biases for WQ, WK, WV, WO.

Python
import torch
from torch import nn

torch.manual_seed(42)
mha = nn.MultiheadAttention(embed_dim=64, num_heads=8, batch_first=True)
x = torch.randn(2, 10, 64)
out, weights = mha(x, x, x)
print(out.shape, weights.shape)
print(sum(p.numel() for p in mha.parameters()))
Expected output
torch.Size([2, 10, 64]) torch.Size([2, 10, 10])
16640
2 sentences of 10 tokens each, d = 64, 8 heads (8 dimensions each). The output has the same shape as the input; the attention weights are (2, 10, 10) (averaged over heads by default). Parameters: 4 · 64² + 4 · 64 = 16,640.

Attention by itself does not know the order of tokens: “the dog chased the cat” and “the cat chased the dog” contain the same tokens. So a positional encoding is added to the embeddings. The original transformer used sines and cosines of different frequencies; models like GPT-2 learn position vectors, and many modern LLMs use rotary encoding (RoPE).

PE(pos, 2i) = sin(pos / 10000^(2i/d)) PE(pos, 2i + 1) = cos(pos / 10000^(2i/d))PE(pos, 2i) = sin(pos / 10000^(2i/d)) PE(pos, 2i + 1) = cos(pos / 10000^(2i/d))
where:
  • posposition of the token: 0, 1, 2, …
  • iindex of the pair of components; small i oscillates fast, large i slowly
  • dembedding dimension
Python
import numpy as np

def positional_encoding(T, d):
    pos = np.arange(T)[:, None]
    i = np.arange(0, d, 2)[None, :]
    angles = pos / 10000 ** (i / d)
    pe = np.zeros((T, d))
    pe[:, 0::2] = np.sin(angles)
    pe[:, 1::2] = np.cos(angles)
    return pe

print(np.round(positional_encoding(T=5, d=4), 3))
▸ Expected output
[[ 0.     1.     0.     1.   ]
 [ 0.841  0.54   0.01   1.   ]
 [ 0.909 -0.416  0.02   1.   ]
 [ 0.141 -0.99   0.03   1.   ]
 [-0.757 -0.654  0.04   0.999]]
Each row is one position. The first two columns change quickly, the last two very slowly — like the second and hour hands of a clock. Every position looks different, and the model can easily work out distances between them.

A transformer block is built like this: x → LayerNorm → multi-head attention → residual connection (x + …) → LayerNorm → a two-layer MLP (usually 4d hidden neurons and GELU) → another residual connection. Residual connections help the gradient pass through dozens of blocks without vanishing. A model is N such blocks stacked. There are three main kinds:

ArchitectureAttentionExamplesTypical tasks
Encoder-onlybidirectionalBERTtext classification, search, named-entity recognition
Decoder-onlycausal (masked)GPTtext generation, chat, code
Encoder–decoderbidirectional + causal + cross-attentionT5translation, summarisation

How LLMs are trained

In pretraining, the model predicts the next token at every position of a huge text corpus. This is self-supervised learning: nobody writes labels, they are simply the next tokens of the text itself. The loss is the cross-entropy averaged over all positions; from it we get the perplexity — how many options, on average, the model “hesitates” between at each step.

L = −(1/T) · ∑ₜ₌₁ᵀ ln p(xₜ₊₁ | x₁, …, xₜ) PPL = eᴸL = −(1/T) · ∑ₜ₌₁ᵀ ln p(xₜ₊₁ | x₁, …, xₜ) PPL = eᴸ
where:
  • xₜthe t-th token of the text
  • p(xₜ₊₁ | x₁, …, xₜ)probability the model gives to the correct next token given the previous ones
  • PPLperplexity; the lower, the better the model
Python
import torch
import torch.nn.functional as F

torch.manual_seed(42)
tokens = torch.tensor([4, 0, 3, 2, 4, 1])
inputs, targets = tokens[:-1], tokens[1:]
logits = torch.randn(len(inputs), 5)
loss = F.cross_entropy(logits, targets)
print(inputs.tolist(), targets.tolist())
print(round(loss.item(), 4), round(torch.exp(loss).item(), 2))
print(round(torch.log(torch.tensor(5.0)).item(), 4))
Expected output
[4, 0, 3, 2, 4] [0, 3, 2, 4, 1]
1.6522 5.22
1.6094
The targets are the inputs shifted by one position. A random “model” (random logits) over a 5-token vocabulary gives a loss of about ln 5 = 1.609 and a perplexity of about 5 — the level of blind guessing.
Example 2: interpreting perplexity

The vocabulary has 50,000 tokens. a) What are the loss and perplexity of a model that guesses blindly? b) A trained model has an average loss of 2.0. What is its perplexity and what does it mean?

Show solution
a) Probability 1/50,000 for every token: L = ln 50,000 ≈ 10.8, PPL = 50,000.
b) PPL = e² ≈ 7.4: at each step the model is as uncertain as if it were choosing among about 7 equally likely options — a huge improvement over 50,000.

After pretraining the model can continue text, but it is not yet a good assistant. Next comes fine-tuning: supervised training on question–answer examples, followed by methods based on the answers people prefer (for example reinforcement learning from human feedback, RLHF). When adapting a model to your own task, you do not have to train all its billions of parameters: methods like LoRA train only small extra matrices. When generating text, the model outputs a distribution over the next token, a token is chosen, and the process repeats.

pᵢ = exp(zᵢ / τ) / ∑ⱼ exp(zⱼ / τ)pᵢ = exp(zᵢ / τ) / ∑ⱼ exp(zⱼ / τ)
where:
  • τtemperature: τ < 1 gives more confident, repetitive text, τ > 1 more varied and risky text

Sampling the next token with a temperature softmax.

Python
import numpy as np

z = np.array([2.0, 1.0, 0.0])
for T in [0.5, 1.0, 2.0]:
    p = np.exp(z / T) / np.exp(z / T).sum()
    print(T, np.round(p, 3))
▸ Expected output
0.5 [0.867 0.117 0.016]
1.0 [0.665 0.245 0.09 ]
2.0 [0.506 0.307 0.186]
The same logits at three temperatures: τ = 0.5 sharpens towards the most likely token, τ = 2 flattens the distribution.

Pretrained models and responsible AI

In practice, instead of training models from scratch, people use pretrained ones. The Hugging Face Hub hosts thousands of open models, and the transformers library (pip install transformers) loads them in a few lines. This code does not run in the browser: the first run downloads the model weights (hundreds of megabytes).

Python
from transformers import pipeline

classifier = pipeline('sentiment-analysis', model='distilbert/distilbert-base-uncased-finetuned-sst-2-english')
print(classifier('I love learning with Educora!'))

generator = pipeline('text-generation', model='openai-community/gpt2')
result = generator('Deep learning is', max_new_tokens=20)
print(result[0]['generated_text'])
The first model classifies the sentiment of a text and returns a list of dictionaries with a label and a confidence score; the second continues a text. The generated text may differ on every run.
  • Bias: models can repeat stereotypes present in their training texts — especially dangerous in decisions such as hiring or lending.
  • Hallucinations: a model can invent convincing-sounding but false facts, quotes and sources.
  • Privacy: do not type personal data, passwords or confidential documents into online AI services; personal data in training data must be protected too.
  • Copyright and licences: every model and dataset has terms of use — read them.
  • Energy and resources: training and running large models requires a lot of electricity and computing power.

Key points

  • Text is split into tokens, each token becomes a learned embedding vector, and position information is added.
  • Attention(Q, K, V) = softmax(QKᵀ/√dₖ)·V: each token takes a weighted average of the others' values; √dₖ keeps softmax from saturating.
  • A causal mask hides future tokens in a decoder; multi-head attention learns different relationships in parallel.
  • An LLM is trained to predict the next token: L = −(1/T)∑ ln p(xₜ₊₁ | x≤ₜ), PPL = eᴸ; then it is fine-tuned.
  • When using pretrained models, keep bias, hallucinations, privacy and licences in mind.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
Why are the scores divided by √dₖ in the attention formula?