Skip to content
Educora
Beginner16 min3 / 10

How large language models work

Tokens, next-token prediction, training versus inference, the context window and why models hallucinate — a simple look inside chat assistants.

Check yourself
In this lesson you will learn
  • Explain what a token is and how a model generates text
  • Tell training apart from inference
  • Understand the effect of the context window and the knowledge cutoff
  • Explain why hallucinations happen and check for them

When you type a message on your phone, the keyboard suggests the next word: you write “See you”, and it offers “tomorrow”. A large language model is, at heart, a far more powerful version of the same idea. In this lesson you will look inside ChatGPT, Claude, Gemini, Copilot and similar assistants and understand both their strength and their typical mistakes.

Definition
Large language model (LLM)

A neural network with billions of parameters, trained on a huge amount of text to predict the next piece of text. Chat assistants are built on top of such models.

Tokens: how a model reads text

A model does not read letters or whole words but tokens — pieces of text. A token can be a common word, part of a longer word, a word with its leading space, or a punctuation mark. Each token is turned into a number, and the neural network works with these numbers.

Text
Text:    Tokenization is unbelievably useful!
Tokens:  Token | ization | is | un | believ | ably | useful | !
         → 8 tokens for 4 words and 1 sign
IDs:     each token becomes a number, e.g. 'useful' → 5505
An illustrative split. Every model family has its own tokenizer, so the exact pieces and numbers differ.

In English, one token is on average about three quarters of a word. Words in Azerbaijani, Russian and Turkish are often split into more pieces, so the same text in these languages usually takes more tokens. This matters because model limits and prices are measured in tokens.

One job: predict the next token

Everything a large language model does comes from one skill. Looking at the tokens so far, it calculates a probability for every possible next token, picks one, adds it to the text and repeats. Even a long, well-organised answer is produced piece by piece in this way.

Possible next tokenProbability
Baku92%
the4%
all other tokens4%
Illustrative numbers for the text “The capital of Azerbaijan is …”.

The model does not always take the most probable token. A setting called temperature adds a controlled amount of randomness: low temperature gives predictable, repeatable answers; higher temperature gives more varied, creative ones. That is why the same question can get different answers.

Python
text = 'the cat sat on the mat . the cat ate . the dog sat .'
words = text.split()

counts = {}
for current, following in zip(words, words[1:]):
    counts.setdefault(current, {})
    counts[current][following] = counts[current].get(following, 0) + 1

def next_word_probs(word):
    options = counts[word]
    total = sum(options.values())
    return {w: c / total for w, c in options.items()}

for w, p in next_word_probs('the').items():
    print(f'the -> {w}: {p:.0%}')

sentence = ['the']
for _ in range(5):
    probs = next_word_probs(sentence[-1])
    sentence.append(max(probs, key=probs.get))
print(' '.join(sentence))
▸ Expected output
the -> cat: 50%
the -> mat: 25%
the -> dog: 25%
the cat sat on the cat
A tiny “language model” that only counts which word follows which. It looks at just one previous word, so it soon goes round in circles. A real LLM uses the transformer's attention mechanism to take thousands of previous tokens into account at once.

Training and inference

StageWhat happensParameters
Pre-trainingThe model reads an enormous amount of text — books, websites, code — and learns to predict the next token. This takes weeks or even months on large clusters of specialised chips.Change
Fine-tuningPeople show the model good examples and rate its answers (reinforcement learning from human feedback, RLHF). The model learns to follow instructions, be helpful and refuse harmful requests.Change
InferenceYou send a message; the finished model computes the answer token by token.Do not change

Two important consequences follow. First, a model's knowledge stops at its knowledge cutoff — the date its training data ends. Some assistants can search the web or read files you upload, which brings fresh information into the chat. Second, the model does not learn from you while you talk: its parameters stay the same. Whether conversations are later used to train future models depends on the service and your settings.

The context window

The context window is the maximum number of tokens a model can take into account at once: your messages, its answers and any documents you add. In current models it ranges from tens of thousands to around a million tokens. When a conversation grows longer than the window, the earliest parts drop out or are summarised, and the model effectively “forgets” them. A new chat starts empty, unless the service has a memory feature that saves notes and adds them to the context.

Why models hallucinate

A hallucination is a fluent, confident answer that is false: a book that does not exist, a wrong date, an invented quotation or reference. It happens because the model produces the most plausible continuation, not a checked fact. Even when its training data had little about a topic, the model still writes smooth text instead of saying “I don't know”. Rare facts, exact numbers, quotations, references and recent events are most at risk.

Invites hallucination
Prompt:
Give me 3 research articles about AI in
Azerbaijani schools, with authors and links.

Answer:
1. J. Smith, K. Lee (2019). AI in Baku
   classrooms. Journal of Educational AI, 12(3).
   https://doi.org/...

→ looks real, but may not exist at all
Safer request
Prompt:
Where can I search for research on AI in
Azerbaijani schools? Suggest databases and
search terms. Do not invent references;
say when you are not sure.

Answer:
Try Google Scholar with terms such as ...

→ you find and check the real sources yourself
Models most often invent details when asked for exact references. Ask how to find sources, then check them yourself.

Practice: test an assistant

  1. 1
    Ask about something you know well

    For example, your city, school or favourite subject. Note every mistake, even small ones.

  2. 2
    Ask the same question twice

    Open two new chats and compare the answers. Differences show that the text is generated, not taken from a ready database.

  3. 3
    Ask about a very recent event

    See whether the assistant mentions its knowledge cutoff or searches the web.

  4. 4
    Ask for sources and check them

    Search for every book or article it names. Mark which ones really exist.

  5. 5
    Write a short conclusion

    Where was the assistant reliable, and where did it hallucinate?

Key points

  • An LLM splits text into tokens and writes its answer by predicting the next token again and again.
  • Parameters change during training; they do not change while you chat (inference).
  • A model's knowledge ends at a cutoff date; web search and uploaded files can add fresh information.
  • The context window is the token limit the model sees at once; whatever falls outside is “forgotten”.
  • A hallucination is a plausible but false answer; always check references and numbers.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
What does a large language model do at every step when it writes an answer?