- Explain what a token is and how a model generates text
- Tell training apart from inference
- Understand the effect of the context window and the knowledge cutoff
- Explain why hallucinations happen and check for them
When you type a message on your phone, the keyboard suggests the next word: you write “See you”, and it offers “tomorrow”. A large language model is, at heart, a far more powerful version of the same idea. In this lesson you will look inside ChatGPT, Claude, Gemini, Copilot and similar assistants and understand both their strength and their typical mistakes.
A neural network with billions of parameters, trained on a huge amount of text to predict the next piece of text. Chat assistants are built on top of such models.
Tokens: how a model reads text
A model does not read letters or whole words but tokens — pieces of text. A token can be a common word, part of a longer word, a word with its leading space, or a punctuation mark. Each token is turned into a number, and the neural network works with these numbers.
Text: Tokenization is unbelievably useful!
Tokens: Token | ization | is | un | believ | ably | useful | !
→ 8 tokens for 4 words and 1 sign
IDs: each token becomes a number, e.g. 'useful' → 5505In English, one token is on average about three quarters of a word. Words in Azerbaijani, Russian and Turkish are often split into more pieces, so the same text in these languages usually takes more tokens. This matters because model limits and prices are measured in tokens.
One job: predict the next token
Everything a large language model does comes from one skill. Looking at the tokens so far, it calculates a probability for every possible next token, picks one, adds it to the text and repeats. Even a long, well-organised answer is produced piece by piece in this way.
| Possible next token | Probability |
|---|---|
| Baku | 92% |
| the | 4% |
| all other tokens | 4% |
The model does not always take the most probable token. A setting called temperature adds a controlled amount of randomness: low temperature gives predictable, repeatable answers; higher temperature gives more varied, creative ones. That is why the same question can get different answers.
text = 'the cat sat on the mat . the cat ate . the dog sat .'
words = text.split()
counts = {}
for current, following in zip(words, words[1:]):
counts.setdefault(current, {})
counts[current][following] = counts[current].get(following, 0) + 1
def next_word_probs(word):
options = counts[word]
total = sum(options.values())
return {w: c / total for w, c in options.items()}
for w, p in next_word_probs('the').items():
print(f'the -> {w}: {p:.0%}')
sentence = ['the']
for _ in range(5):
probs = next_word_probs(sentence[-1])
sentence.append(max(probs, key=probs.get))
print(' '.join(sentence))▸ Expected output
the -> cat: 50% the -> mat: 25% the -> dog: 25% the cat sat on the cat
Training and inference
| Stage | What happens | Parameters |
|---|---|---|
| Pre-training | The model reads an enormous amount of text — books, websites, code — and learns to predict the next token. This takes weeks or even months on large clusters of specialised chips. | Change |
| Fine-tuning | People show the model good examples and rate its answers (reinforcement learning from human feedback, RLHF). The model learns to follow instructions, be helpful and refuse harmful requests. | Change |
| Inference | You send a message; the finished model computes the answer token by token. | Do not change |
Two important consequences follow. First, a model's knowledge stops at its knowledge cutoff — the date its training data ends. Some assistants can search the web or read files you upload, which brings fresh information into the chat. Second, the model does not learn from you while you talk: its parameters stay the same. Whether conversations are later used to train future models depends on the service and your settings.
The context window
The context window is the maximum number of tokens a model can take into account at once: your messages, its answers and any documents you add. In current models it ranges from tens of thousands to around a million tokens. When a conversation grows longer than the window, the earliest parts drop out or are summarised, and the model effectively “forgets” them. A new chat starts empty, unless the service has a memory feature that saves notes and adds them to the context.
Why models hallucinate
A hallucination is a fluent, confident answer that is false: a book that does not exist, a wrong date, an invented quotation or reference. It happens because the model produces the most plausible continuation, not a checked fact. Even when its training data had little about a topic, the model still writes smooth text instead of saying “I don't know”. Rare facts, exact numbers, quotations, references and recent events are most at risk.
Prompt:
Give me 3 research articles about AI in
Azerbaijani schools, with authors and links.
Answer:
1. J. Smith, K. Lee (2019). AI in Baku
classrooms. Journal of Educational AI, 12(3).
https://doi.org/...
→ looks real, but may not exist at allPrompt:
Where can I search for research on AI in
Azerbaijani schools? Suggest databases and
search terms. Do not invent references;
say when you are not sure.
Answer:
Try Google Scholar with terms such as ...
→ you find and check the real sources yourselfPractice: test an assistant
- 1Ask about something you know well
For example, your city, school or favourite subject. Note every mistake, even small ones.
- 2Ask the same question twice
Open two new chats and compare the answers. Differences show that the text is generated, not taken from a ready database.
- 3Ask about a very recent event
See whether the assistant mentions its knowledge cutoff or searches the web.
- 4Ask for sources and check them
Search for every book or article it names. Mark which ones really exist.
- 5Write a short conclusion
Where was the assistant reliable, and where did it hallucinate?
Key points
- An LLM splits text into tokens and writes its answer by predicting the next token again and again.
- Parameters change during training; they do not change while you chat (inference).
- A model's knowledge ends at a cutoff date; web search and uploaded files can add fresh information.
- The context window is the token limit the model sees at once; whatever falls outside is “forgotten”.
- A hallucination is a plausible but false answer; always check references and numbers.
Check yourself
10 questions. Every correct answer earns XP.