- Identify the features and the label in a dataset
- Tell supervised, unsupervised and reinforcement learning apart
- Explain how a model learns by reducing its error
- Recognise the danger of poor data and overfitting
How did you learn to tell a cat from a dog? Nobody gave you a list of rules like “if the ears are pointed and the whiskers are long, it is a cat”. You simply saw many cats and dogs, someone told you which was which, and your brain found the differences itself. Machine learning works in a surprisingly similar way: instead of writing rules, we show the computer examples and let it find the pattern.
A branch of AI in which a program gets better at a task by learning from data (examples) instead of following rules written by a programmer. The result of learning is a model — a function with adjustable settings called parameters.
Data: features and labels
Suppose we want to predict the price of a taxi ride in Baku. Each past trip is an example. What we know before the trip — distance, time of day, day of the week — are the features. What we want to predict — the fare — is the label. A table of such examples is called training data.
| Trip | Distance, km (feature) | Fare, manat (label) |
|---|---|---|
| 1 | 1 | 2.0 |
| 2 | 2 | 2.6 |
| 3 | 3 | 3.0 |
| 4 | 4 | 3.4 |
| 5 | 5 | 4.0 |
Three types of learning
| Type | How it learns | Example |
|---|---|---|
| Supervised learning | From examples with correct answers (labels) | Predicting fares, recognising spam, reading handwriting |
| Unsupervised learning | From examples without answers; it finds groups and structure itself | Grouping shop customers by their buying habits |
| Reinforcement learning | By trial and error, getting rewards for good actions | Game-playing programs such as AlphaGo, robot control |
Most of the AI you meet every day uses supervised learning. Chat assistants combine several methods, as you will see in the next lesson.
How a model learns: guess, check, adjust
Our model is a straight line: fare = w · km + b. Here w is the price per kilometre and b is the starting price — the two parameters. At first the model knows nothing: w = 0 and b = 0, and all its predictions are wrong. Learning is a loop:
- 1Guess
The model predicts the fare for every trip in the training data.
- 2Check
It compares each prediction with the real fare and calculates the average error (here: the mean of the squared differences).
- 3Adjust
It changes w and b a tiny bit in the direction that makes the error smaller. The size of this step is the learning rate.
- 4Repeat
It repeats the loop hundreds or thousands of times until the error stops falling.
km = [1, 2, 3, 4, 5]
fare = [2.0, 2.6, 3.0, 3.4, 4.0]
w, b = 0.0, 0.0 # the model: fare = w * km + b
lr = 0.02 # learning rate: the size of each small step
for step in range(2001):
grad_w = grad_b = 0.0
for x, y in zip(km, fare):
error = (w * x + b) - y
grad_w += 2 * error * x / len(km)
grad_b += 2 * error / len(km)
w -= lr * grad_w
b -= lr * grad_b
if step % 500 == 0:
loss = sum((w * x + b - y) ** 2 for x, y in zip(km, fare)) / len(km)
print(f'step {step:4d}: w = {w:.3f}, b = {b:.3f}, error = {loss:.4f}')
print(f'Prediction for 8 km: {w * 8 + b:.2f} manat')▸ Expected output
step 0: w = 0.398, b = 0.120, error = 2.8551 step 500: w = 0.492, b = 1.516, error = 0.0036 step 1000: w = 0.480, b = 1.559, error = 0.0032 step 1500: w = 0.480, b = 1.560, error = 0.0032 step 2000: w = 0.480, b = 1.560, error = 0.0032 Prediction for 8 km: 5.40 manat
The model has discovered the rule hidden in the data: about 0.48 manat per kilometre, plus 1.56 manat to start. Nobody wrote this rule; it came from the examples. Now the model can predict a trip it has never seen: 8 km → 5.40 manat. Modern AI, from spam filters to chat assistants, learns with this same loop, only with far more parameters and data.
Neural networks and deep learning
A straight line has two parameters: enough for taxi fares, but not for recognising a face or writing a sentence. For such tasks we use neural networks: many simple units (artificial neurons) arranged in layers. Each unit multiplies its inputs by weights, adds them up and passes the result on. A network with many layers is called deep, hence the term deep learning. The weights are learned with the same guess–check–adjust loop.
Data quality and overfitting
A model can only be as good as its data. If all the example trips happened in the daytime, the model will not know that night fares may be different. If the data contains mistakes, or mostly one kind of example, the model learns those mistakes too.
Training photos for a 'cat' detector:
20 photos
all of orange cats
all taken indoors, on a sofa
→ the model learns: orange + sofa = cat
→ a grey cat in a garden: 'not a cat'Training photos for a 'cat' detector:
20 000 photos
cats of every colour and breed
indoors and outdoors, day and night
+ many photos without cats
→ the model learns ears, eyes, whiskers, shape
→ a grey cat in a garden: 'cat'Another trap is overfitting: the model memorises its training examples instead of learning the general pattern — like a student who memorises last year's test answers and then fails a new test. To catch this, we always keep part of the data aside as a test set and check the model on examples it has never seen.
Practice
A school wants to predict whether a student will pass the final exam. It has data from past years: attendance, homework scores, hours of study per week and whether each student passed. Name the features, the label and the learning type.
Show solutionHide solution
Label: passed / did not pass.
Learning type: supervised (the past data contains the correct answers). The label is a category, not a number, so this is a classification task.
Important: the prediction is a probability, not a verdict — teachers should use it to offer help in time, not to label students.
The line fare = 0.5 · km + 1.5 is a guess made by a person. Calculate its average error (the mean of the squared differences) on the five trips and print it with 4 decimals. Is it better or worse than the learned model's 0.0032?
km = [1, 2, 3, 4, 5]
fare = [2.0, 2.6, 3.0, 3.4, 4.0]
w, b = 0.5, 1.5
total = 0
# add the squared error of every trip to total,
# then print total / len(km) with 4 decimals▸ Expected output
0.0040
Key points
- In machine learning the program finds the rule in examples itself; the result is a model with parameters.
- Features are what we know in advance; the label is the answer we want to predict.
- The three types of learning are supervised, unsupervised and reinforcement learning.
- A model learns by reducing its error in the loop guess → check → adjust → repeat.
- A model is only as good as its data; always test it on data it has never seen.
Check yourself
10 questions. Every correct answer earns XP.