- Explain what a code table is and encode and decode words with ASCII codes
- Calculate the size of a text under the ASCII (1 byte) and UNICODE (2 bytes) conventions, counting spaces and punctuation too
- Find the power of an alphabet from the size of a document, and the number of pages from transfer data
When you type “Salam” on a keyboard, the computer stores not letters but numbers: 83, 97, 108, 97, 109. The screen shows letters again because a program draws the sign that matches each number. Sometimes, though, a text from a friend shows strange signs like “É™” instead of the Azerbaijani letter ə. In this lesson we will see how a computer encodes text and where such errors come from.
In the lesson “Measuring information: bits, bytes and units” we learned the formula I = K · i; now we will apply it to real texts. The topic comes up often in the entrance exam: in the four entrance exams of 2025–2026, five tasks were about encoding text (ASCII and UNICODE volumes, the power of an alphabet, words in a 64-symbol alphabet, a UNICODE text sent over a channel), and one text-processing task chose words by their UNICODE volume.
Code tables: a number for every symbol
Turning information from one system of signs into another by a fixed rule. Encoding a text means turning every symbol into a binary code; the reverse operation is called decoding.
A table that gives every symbol (letter, digit, punctuation mark, space, special sign) a number — its code. The programs that write and read a text must use the same table.
ASCII: one symbol, one byte
ASCII (American Standard Code for Information Interchange) was created in 1963. It has 128 codes (0–127): 0–31 are control characters (for example, a new line), 32 is the space, 48–57 are the digits, 65–90 the capital and 97–122 the small Latin letters. In a computer every ASCII symbol takes 1 byte (8 bits). Codes 128–255 were used for national letters: tables such as Windows-1251 and KOI8-R were made for Cyrillic and Windows-1254 for Turkish.
| Symbol | Decimal code | Binary code (1 byte) |
|---|---|---|
| space | 32 | 00100000 |
| 0 | 48 | 00110000 |
| 9 | 57 | 00111001 |
| A | 65 | 01000001 |
| B | 66 | 01000010 |
| Z | 90 | 01011010 |
| a | 97 | 01100001 |
| z | 122 | 01111010 |
1) Find the word whose ASCII codes are 73 78 70 79.
2) The code of “A” is 65. What is the code of “K”?
3) Which word do the binary codes 01000010 01001001 01010100 stand for?
4) The code of “d” is 100. Find the codes of “D” and “h”.
Show solutionHide solution
2) K is the 11th letter of the English alphabet: 65 + 10 = 75.
3) 01000010 = 64 + 2 = 66 ⇒ B; 01001001 = 64 + 8 + 1 = 73 ⇒ I; 01010100 = 64 + 16 + 4 = 84 ⇒ T. The word is BIT.
4) “D” = 100 − 32 = 68; “h” comes 4 letters after “d”: 100 + 4 = 104.
word = "INFO"
print([ord(c) for c in word])
print(chr(66) + chr(73) + chr(84))
print(ord("a") - ord("A"), ord("ə"))▸ Expected output
[73, 78, 70, 79] BIT 32 601
ord() gives the code of a symbol and chr() gives the symbol for a code. The code of ə is 601: it is not in ASCII, only in UNICODE.UNICODE: all languages in one table
The 8-bit tables of different countries did not agree: the same code meant one letter in one table and another letter in another. So in 1991 the UNICODE standard was created. It gives a separate number to the letters of all the world’s scripts, to math signs and even to emoji — more than 150,000 symbols today. The first 128 codes are the same as in ASCII: “A” is still 65, and the number of ə is U+0259, that is, 601.
In school problems and DİM tasks, “UNICODE” means 16 bits (2 bytes) per symbol: 2¹⁶ = 65,536 different symbols. This matches the basic case of the UTF-16 encoding. Real files and the web mostly use UTF-8, where a symbol takes from 1 to 4 bytes: Latin letters take 1, the special Azerbaijani letters (ə, ş, ç, ğ, ı, ö, ü) and Cyrillic letters 2, the € sign and most Chinese characters 3, and emoji 4 bytes.
| Encoding | One symbol | Number of symbols |
|---|---|---|
| ASCII | 8 bits = 1 byte | 2⁸ = 256 (standard ASCII — 128) |
| UNICODE (in school problems) | 16 bits = 2 bytes | 2¹⁶ = 65,536 |
| UTF-8 | 1–4 bytes | all UNICODE symbols |
- Ithe information volume of the text
- Kthe number of all symbols in the text: letters, digits, spaces, punctuation marks
This is a special case of I = K · i: i = 8 in ASCII and i = 16 in UNICODE. The same text takes twice as much space in UNICODE as in ASCII.
- 1Count the letters
Count the letters of every word and add them up.
- 2Add the spaces
If there is one space between words, the number of spaces = the number of words − 1.
- 3Do not forget punctuation
Full stops, commas, question and exclamation marks, quotation marks, dashes and digits are all symbols.
- 4Multiply and convert
Multiply K by 8 (ASCII) or 16 (UNICODE) bits, then convert to the unit you need.
Match the encoding systems with the information volume of the sentence
Small steps lead to big results.
(there is one space between words).
1. UNICODE
2. ASCII
a. 32 bytes
b. 2⁹ bits
c. 64 bits
d. 2⁸ bits
e. 64 bytes
Show solutionHide solution
1. UNICODE: 32 · 2 = 64 bytes = 512 bits = 2⁹ bits ⇒ b, e.
2. ASCII: 32 · 1 = 32 bytes = 256 bits = 2⁸ bits ⇒ a, d.
“64 bits” (c) matches neither: it is the answer of someone who mixes up 64 bytes and 64 bits.
Answer: 1 – b, e; 2 – a, d.
In a word processor the sentence “Small steps lead to big results.” is typed, and the symbols are encoded in UNICODE. Aysel makes every word with a volume of 80 bits bold and the word with a volume of 64 bits italic. Which words become bold and which becomes italic? (The full stop does not belong to a word.)
Show solutionHide solution
80 bits ÷ 16 = 5 letters ⇒ Small and steps become bold.
64 bits ÷ 16 = 4 letters ⇒ lead becomes italic.
The other words (to — 32 bits, big — 48 bits, results — 112 bits) stay as they are.
Match the words encoded in a 32-symbol alphabet with their information volumes.
1. data
2. python
3. internet
a. 5 bytes
b. 30 bits
c. 20 bits
d. 2.5 bytes
e. 40 bits
Show solutionHide solution
1. data: 4 · 5 = 20 bits = 2.5 bytes ⇒ c, d.
2. python: 6 · 5 = 30 bits ⇒ b.
3. internet: 8 · 5 = 40 bits = 5 bytes ⇒ a, e.
Answer: 1 – c, d; 2 – b; 3 – a, e.
s = "Small steps lead to big results."
print(len(s)) # spaces and the dot included
print(len(s) * 8, len(s) * 16) # bits in ASCII and in UNICODE
for w in ["Salam", "Gəncə", "Баку", "€"]:
print(w, len(w.encode("utf-8"))) # real size in UTF-8, bytes▸ Expected output
32 256 512 Salam 5 Gəncə 7 Баку 8 € 3
len() counts the symbols, spaces and the full stop included. The last lines show the real size in UTF-8: the word Gəncə is 7 bytes because each of its two ə letters takes 2 bytes.Documents, the power of an alphabet and transmission
In a document of several pages, the number of symbols is K = pages · lines · symbols per line. In exam problems the numbers are usually powers of two (32 lines, 64 symbols), so it is convenient to work with exponents. The other way round, if the volume is known, you can find the bits per symbol and the power of the alphabet.
1) A 16-page text has 40 lines per page and 64 symbols per line, encoded in UNICODE. How many KB is it?
2) A text was converted from UNICODE to ASCII, and its volume fell by 480 bytes. How many symbols does the text have?
Show solutionHide solution
Shortcut: 16 · 64 = 2¹⁰, so I = 40 · 2¹⁰ · 2 bytes = 80 KB.
2) Each symbol went from 2 bytes to 1 byte, so it lost 1 byte: 480 bytes ÷ 1 byte = 480 symbols.
- ithe number of bits per symbol
- Ithe volume of the text part (bits); subtract the pictures first
- Kthe number of symbols
- Nthe power of the alphabet
The key to reverse problems: first separate out the volume of the text, then divide by the number of symbols.
In a 20-page document, each of the first 4 pages holds only a 16 KB picture, and each of the other pages holds text of 32 lines with 64 symbols per line. The volume of the document is 88 KB. Find the power of the alphabet the text is written in.
A) 32 B) 64 C) 6 D) 128 E) 256
Show solutionHide solution
Text pages: 20 − 4 = 16 = 2⁴; K = 2⁴ · 2⁵ · 2⁶ = 2¹⁵ symbols.
i = 24 · 2¹³ ÷ 2¹⁵ = 24 ÷ 4 = 6 bits ⇒ N = 2⁶ = 64 (B).
Option C is i itself: the question asks for N, not i.
When a text is sent over a communication channel, the formula I = v · t from the lesson “Measuring information: bits, bytes and units” applies: the volume of the text in bits = rate · time.
1) A UNICODE text with 60 symbols per page is sent over a 48 bit/s channel in 5 minutes. How many pages does it have?
2) A 2-page ASCII text has 30 lines per page and 64 symbols per line. How many minutes does it take to send it at 256 bit/s?
Show solutionHide solution
One page: 60 · 16 = 960 bits ⇒ 14,400 ÷ 960 = 15 pages.
2) I = 2 · 30 · 64 · 8 = 30,720 bits; t = 30,720 ÷ 256 = 120 s = 2 min.
The variable s holds a sentence. Print its information volume in bits in UNICODE (16 bits per symbol).
s = "Bakı is the capital of Azerbaijan."
# print the UNICODE volume in bits▸ Expected output
544
In the next lesson, “Computer graphics: encoding raster and vector images”, we will see how pictures are encoded; the formula N = 2ⁱ plays the key role there too.
Key points
- When a text is encoded, every symbol is replaced by its number from a code table; in ASCII “A” = 65, “a” = 97, the space = 32.
- A symbol takes 1 byte (8 bits) in ASCII and, in school problems, 2 bytes (16 bits) in UNICODE.
- When you calculate the size of a text, spaces, punctuation marks and digits count too.
- In an N-symbol alphabet one symbol is i bits (N = 2ⁱ): 64 symbols — 6 bits, 32 symbols — 5 bits.
- In reverse problems, first subtract the pictures, then find i = I / K and compute N = 2ⁱ.
- Real files use UTF-8: a Latin letter is 1 byte,
əand Cyrillic letters 2 bytes, emoji 4 bytes.
Check yourself
12 questions. Every correct answer earns XP.