Skip to content
Educora
Beginner16 min2 / 10

How machine learning learns from data

Features and labels, the three types of learning, the guess–check–adjust loop, neural networks and data quality — without heavy maths, with a tiny runnable Python demo.

Check yourself
In this lesson you will learn
  • Identify the features and the label in a dataset
  • Tell supervised, unsupervised and reinforcement learning apart
  • Explain how a model learns by reducing its error
  • Recognise the danger of poor data and overfitting

How did you learn to tell a cat from a dog? Nobody gave you a list of rules like “if the ears are pointed and the whiskers are long, it is a cat”. You simply saw many cats and dogs, someone told you which was which, and your brain found the differences itself. Machine learning works in a surprisingly similar way: instead of writing rules, we show the computer examples and let it find the pattern.

Definition
Machine learning

A branch of AI in which a program gets better at a task by learning from data (examples) instead of following rules written by a programmer. The result of learning is a model — a function with adjustable settings called parameters.

Data: features and labels

Suppose we want to predict the price of a taxi ride in Baku. Each past trip is an example. What we know before the trip — distance, time of day, day of the week — are the features. What we want to predict — the fare — is the label. A table of such examples is called training data.

TripDistance, km (feature)Fare, manat (label)
112.0
222.6
333.0
443.4
554.0
The five trips used in the demo below. Real systems use thousands of examples and many features.

Three types of learning

TypeHow it learnsExample
Supervised learningFrom examples with correct answers (labels)Predicting fares, recognising spam, reading handwriting
Unsupervised learningFrom examples without answers; it finds groups and structure itselfGrouping shop customers by their buying habits
Reinforcement learningBy trial and error, getting rewards for good actionsGame-playing programs such as AlphaGo, robot control

Most of the AI you meet every day uses supervised learning. Chat assistants combine several methods, as you will see in the next lesson.

How a model learns: guess, check, adjust

Our model is a straight line: fare = w · km + b. Here w is the price per kilometre and b is the starting price — the two parameters. At first the model knows nothing: w = 0 and b = 0, and all its predictions are wrong. Learning is a loop:

  1. 1
    Guess

    The model predicts the fare for every trip in the training data.

  2. 2
    Check

    It compares each prediction with the real fare and calculates the average error (here: the mean of the squared differences).

  3. 3
    Adjust

    It changes w and b a tiny bit in the direction that makes the error smaller. The size of this step is the learning rate.

  4. 4
    Repeat

    It repeats the loop hundreds or thousands of times until the error stops falling.

Python
km = [1, 2, 3, 4, 5]
fare = [2.0, 2.6, 3.0, 3.4, 4.0]

w, b = 0.0, 0.0   # the model: fare = w * km + b
lr = 0.02         # learning rate: the size of each small step

for step in range(2001):
    grad_w = grad_b = 0.0
    for x, y in zip(km, fare):
        error = (w * x + b) - y
        grad_w += 2 * error * x / len(km)
        grad_b += 2 * error / len(km)
    w -= lr * grad_w
    b -= lr * grad_b
    if step % 500 == 0:
        loss = sum((w * x + b - y) ** 2 for x, y in zip(km, fare)) / len(km)
        print(f'step {step:4d}: w = {w:.3f}, b = {b:.3f}, error = {loss:.4f}')

print(f'Prediction for 8 km: {w * 8 + b:.2f} manat')
▸ Expected output
step    0: w = 0.398, b = 0.120, error = 2.8551
step  500: w = 0.492, b = 1.516, error = 0.0036
step 1000: w = 0.480, b = 1.559, error = 0.0032
step 1500: w = 0.480, b = 1.560, error = 0.0032
step 2000: w = 0.480, b = 1.560, error = 0.0032
Prediction for 8 km: 5.40 manat
Run the code. The error falls from 2.86 to 0.0032, and the model finds w ≈ 0.48 and b ≈ 1.56.

The model has discovered the rule hidden in the data: about 0.48 manat per kilometre, plus 1.56 manat to start. Nobody wrote this rule; it came from the examples. Now the model can predict a trip it has never seen: 8 km → 5.40 manat. Modern AI, from spam filters to chat assistants, learns with this same loop, only with far more parameters and data.

Neural networks and deep learning

A straight line has two parameters: enough for taxi fares, but not for recognising a face or writing a sentence. For such tasks we use neural networks: many simple units (artificial neurons) arranged in layers. Each unit multiplies its inputs by weights, adds them up and passes the result on. A network with many layers is called deep, hence the term deep learning. The weights are learned with the same guess–check–adjust loop.

Data quality and overfitting

A model can only be as good as its data. If all the example trips happened in the daytime, the model will not know that night fares may be different. If the data contains mistakes, or mostly one kind of example, the model learns those mistakes too.

Weak data
Training photos for a 'cat' detector:

20 photos
all of orange cats
all taken indoors, on a sofa

→ the model learns: orange + sofa = cat
→ a grey cat in a garden: 'not a cat'
Good data
Training photos for a 'cat' detector:

20 000 photos
cats of every colour and breed
indoors and outdoors, day and night
+ many photos without cats

→ the model learns ears, eyes, whiskers, shape
→ a grey cat in a garden: 'cat'
A model learns whatever the data shows, including accidental patterns such as the sofa.

Another trap is overfitting: the model memorises its training examples instead of learning the general pattern — like a student who memorises last year's test answers and then fails a new test. To catch this, we always keep part of the data aside as a test set and check the model on examples it has never seen.

Practice

Features, label, learning type

A school wants to predict whether a student will pass the final exam. It has data from past years: attendance, homework scores, hours of study per week and whether each student passed. Name the features, the label and the learning type.

Show solution
Features: attendance, homework scores, hours of study.
Label: passed / did not pass.
Learning type: supervised (the past data contains the correct answers). The label is a category, not a number, so this is a classification task.
Important: the prediction is a probability, not a verdict — teachers should use it to offer help in time, not to label students.
Exercise

The line fare = 0.5 · km + 1.5 is a guess made by a person. Calculate its average error (the mean of the squared differences) on the five trips and print it with 4 decimals. Is it better or worse than the learned model's 0.0032?

Exercise · Python
km = [1, 2, 3, 4, 5]
fare = [2.0, 2.6, 3.0, 3.4, 4.0]
w, b = 0.5, 1.5
total = 0
# add the squared error of every trip to total,
# then print total / len(km) with 4 decimals
▸ Expected output
0.0040

Key points

  • In machine learning the program finds the rule in examples itself; the result is a model with parameters.
  • Features are what we know in advance; the label is the answer we want to predict.
  • The three types of learning are supervised, unsupervised and reinforcement learning.
  • A model learns by reducing its error in the loop guess → check → adjust → repeat.
  • A model is only as good as its data; always test it on data it has never seen.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
In the taxi example, what is the label?