- Organise raw data in a frequency table and choose a suitable chart (bar, pie, line, histogram)
- Find the mean (including the weighted mean), the median, the mode and the range
- Compute the variance and the standard deviation of a small data set and explain what they mean
- Decide when the median describes data better than the mean and spot misleading graphs
Leyla and her classmates ran a survey about themselves: height, favourite sport, the number of books read in the summer. They ended up with hundreds of numbers and words, and in that jumbled form the answers say nothing. Statistics teaches you to collect data, arrange it in tables and charts, and describe it with a few numbers: the average, the median, the spread. It also helps you read the charts in the news critically. The step from data to chance is taken in the lesson “Basics of probability”.
Collecting data and frequency tables
Data come from surveys, observations and measurements. Categorical (qualitative) data are names or labels: a sport, a city, a colour. Numerical (quantitative) data are numbers you can calculate with: a score, a height, the number of siblings. The number of values in a data set is written n.
The frequency f of a value is how many times it occurs in the data. Its relative frequency w = f/n is its share of all the data (as a fraction or a percentage).
- fᵢthe frequency of the i-th value
- wᵢthe relative frequency of the i-th value
- nthe number of all values
- kthe number of different values
The two sums are a ready-made check: if the frequencies do not add up to n, a value was missed or counted twice.
- 1List the values
Write all possible values in a column (numbers in increasing order).
- 2Tally
Go through the data once: for each answer put a stroke next to its value and cross the answer out.
- 3Count
Count the strokes: these are the frequencies. Check that they add up to n.
- 4Find the shares
Divide each frequency by n; multiply by 100 for a percentage. The shares must add up to 1 (100%).
20 pupils were asked how many books they read in the summer: 2, 1, 3, 0, 2, 2, 4, 1, 2, 3, 1, 2, 0, 3, 2, 1, 4, 2, 3, 1. Make a frequency table with relative frequencies.
Show solutionHide solution
0 books: f = 2, w = 2/20 = 0.10 (10%)
1 book: f = 5, w = 0.25 (25%)
2 books: f = 7, w = 0.35 (35%)
3 books: f = 4, w = 0.20 (20%)
4 books: f = 2, w = 0.10 (10%)
Check: 2 + 5 + 7 + 4 + 2 = 20 and 10% + 25% + 35% + 20% + 10% = 100%.
Charts: which picture for which data
| Chart | Best for | Example |
|---|---|---|
| Bar chart | comparing categories or frequencies; the bars stand apart | books read |
| Pie chart | parts of one whole (100%) | favourite sports |
| Line chart | change over time | the temperature during a day |
| Histogram | numerical data grouped into equal intervals; the bars touch | pupils’ heights |
- αthe angle of the category’s sector
- fthe frequency of the category
- nthe number of all answers
A pie chart is a frequency table drawn as a circle: each sector’s angle is its share of 360°, and all the angles add up to 360°.
40 pupils named their favourite sport: football 14, basketball 8, volleyball 6, swimming 7, chess 5. Find each sector’s angle and percentage.
Show solutionHide solution
Football: 14 · 9° = 126° (35%)
Basketball: 8 · 9° = 72° (20%)
Volleyball: 6 · 9° = 54° (15%)
Swimming: 7 · 9° = 63° (17.5%)
Chess: 5 · 9° = 45° (12.5%)
Check: 126° + 72° + 54° + 63° + 45° = 360°.
A histogram is used for numerical data with many different values, such as heights. The values are grouped into equal intervals, a value on a boundary goes into the interval on its right, and each bar’s height is the frequency of its interval. Unlike a bar chart, a histogram’s bars are drawn touching, because the intervals follow one another without gaps.
Using the histogram above: 1) how many pupils are shorter than 155 cm? 2) what percentage of the pupils are at least 165 cm tall? 3) which interval contains the most pupils?
Show solutionHide solution
2) The last two columns: 4 + 2 = 6; 6/30 = 0.2 = 20%.
3) The tallest column: [155, 160), 9 pupils.
Measures of centre: mean, median and mode
- x̄the mean (read “x bar”)
- x₁, …, xₙthe data values
- nthe number of values
The mean is the “fair share”: if all the values were pooled and shared out equally, each would get x̄. So the sum of all values is n · x̄ — the key to problems with a missing value.
1) Leyla scored 9, 6, 8, 10, 9, 5, 9 in 7 tests. Find the mean.
2) Murad’s mean score in 4 tests is 7. What must he score in the fifth test to make the mean of all 5 tests 7.4?
3) The mean of 5 numbers is 12. One of them, 20, is removed. What is the mean of the rest?
Show solutionHide solution
2) The total needed: 5 · 7.4 = 37. Already scored: 4 · 7 = 28. The fifth test: 37 − 28 = 9 points.
3) The sum is 5 · 12 = 60; without 20 it is 40 for 4 numbers: 40 ÷ 4 = 10.
- xᵢthe different values
- fᵢtheir frequencies or weights
The weighted mean is used for a frequency table and whenever values have different importance. If the weights are percentages adding up to 100%, just add the products x · w: x̄ = x₁w₁ + … + xₖwₖ.
1) Use the frequency table of books to find the mean number of books a pupil read in the summer.
2) A term score is made up of homework (average 9) with weight 20%, quizzes (average 7) with 30% and a final test (8) with 50%. What is the term score?
3) One class has 20 pupils with a mean score of 80, another has 30 pupils with a mean of 70. What is the mean score of both classes together?
Show solutionHide solution
2) 0.2 · 9 + 0.3 · 7 + 0.5 · 8 = 1.8 + 2.1 + 4 = 7.9.
3) The sum of all scores: 20 · 80 + 30 · 70 = 1600 + 2100 = 3700; 50 pupils: 3700 ÷ 50 = 74. (The plain mean of 80 and 70, 75, would be wrong: the second class is larger.)
The middle value of the data sorted in increasing order: half of the values are not greater than it and half are not smaller.
- x₁ ≤ x₂ ≤ … ≤ xₙthe data sorted in increasing order
- nthe number of values: odd (2k + 1) or even (2k)
Odd n: the value in position (n + 1)/2. Even n: the mean of the two middle values. Always sort first!
The value that occurs most often. A data set can have two or more modes, or none (when all values occur equally often). Of the three measures, only the mode also works for categorical data.
Find the median or the mode:
1) the median of 1, 4, 6, 9, 20;
2) the median of 3, 8, 5, 10, 7, 2;
3) the median and the mode of the books frequency table;
4) the mode of 2, 3, 3, 5, 7, 7 and the mode of the sports survey.
Show solutionHide solution
2) Sort: 2, 3, 5, 7, 8, 10; n = 6 is even: median = (5 + 7) ÷ 2 = 6.
3) n = 20, so we need the 10th and 11th values. Running totals: 0 books fill places 1–2, 1 book places 3–7, 2 books places 8–14. Both are 2: median = 2. The mode is also 2 (f = 7).
4) 3 and 7 each occur twice: two modes, 3 and 7. The mode of the survey is football — a mean or a median makes no sense here.
Spread: range, variance and standard deviation
Two basketball players average 12 points per game. Murad’s last five games: 10, 12, 14, 12, 12; Elvin’s: 4, 20, 12, 6, 18. Their means are equal, but Murad is steady while Elvin swings between high and low. Measures of centre cannot see this difference — measures of spread can.
- Rthe range
- xₘₐₓ, xₘᵢₙthe largest and the smallest value
The range is quick to find, but it depends only on the two extreme values, so a single outlier changes it a lot.
- σ²the variance: the mean of the squared deviations
- σthe standard deviation
- xᵢ − x̄the deviation of a value from the mean
You cannot just add the deviations — their sum is always 0 — so we square them. σ has the same unit as the data and shows the typical distance of the values from the mean. For sample data, university statistics divides by n − 1 instead of n.
- 1Find the mean
Compute x̄.
- 2Write the deviations
For each value, xᵢ − x̄; they must add up to 0 (a check).
- 3Square them
Square each deviation and add the squares.
- 4Divide by n, take the root
Sum ÷ n = the variance; σ is its square root.
Find the range, the variance and the standard deviation of Murad’s points (10, 12, 14, 12, 12) and Elvin’s points (4, 20, 12, 6, 18).
Show solutionHide solution
Murad: R = 14 − 10 = 4. Deviations −2, 0, 2, 0, 0; squares 4, 0, 4, 0, 0, sum 8. σ² = 8 ÷ 5 = 1.6, σ = √1.6 ≈ 1.26.
Elvin: R = 20 − 4 = 16. Deviations −8, 8, 0, −6, 6; squares 64, 64, 0, 36, 36, sum 200. σ² = 200 ÷ 5 = 40, σ = √40 ≈ 6.32.
Conclusion: Murad’s points are usually 1–2 points away from the mean, Elvin’s about 6 points.
Outliers, mean or median, and misleading graphs
A value far away from most of the others. It may be a mistake (a typo, a faulty measurement) or a real but unusual case — check before you delete it.
A small firm pays its 5 employees 700, 700, 800, 900 and 4400 manat a month (the last one is the director). Find the mean and the median, then recompute them without the director.
Show solutionHide solution
Without the director: mean = 3100 ÷ 4 = 775, median = (700 + 800) ÷ 2 = 750.
One outlier moved the mean by 725 manat but the median by only 50. The median describes a “typical” salary better: four of the five employees earn far less than 1500 manat.
| Measure | When to use it | Examples |
|---|---|---|
| Mean | numerical data, roughly symmetric, without outliers | test scores, heights, temperatures |
| Median | skewed data or data with outliers | salaries, flat prices, waiting times |
| Mode | categorical data or “the most popular” questions | favourite sport, the shoe size a shop needs most |
A chart can be correct in its numbers and still leave a false impression. The main tricks of misleading graphs are:
- Truncated axis: the vertical axis of a bar chart does not start at 0, so small differences look huge.
- Uneven scale: equal steps on an axis stand for different amounts, or histogram intervals have different widths.
- Pictures scaled in two directions: a figure twice as tall and twice as wide has 4 times the area, so a doubling looks like a fourfold increase.
- A chosen time window, or no units and no source: only the years that support a claim are shown, or the numbers cannot be checked.
- 1.The mean of 3, 5 and 10 is .
- 2.The median of 1, 2, 8, 9 is .
- 3.The mode of 4, 7, 7, 2, 4, 7 is .
- 4.The range of 12, 5, 19, 8 is .
- 5.A category with 25% of the answers gets a sector of ° in a pie chart.
Key points
- In a frequency table the frequencies add up to n and the relative frequencies to 1 (100%).
- Bar chart — categories; pie chart — parts of a whole (α = 360° · f/n); line chart — change over time; histogram — grouped numerical data.
- Mean = sum ÷ n; weighted mean = Σ x·f ÷ Σ f; the sum of all values is n · x̄.
- Median — the middle of the sorted data (for even n, the mean of the two middle values); mode — the most frequent value; range = max − min.
- The variance is the mean of the squared deviations; σ, its square root, is the typical distance of the values from the mean.
- With outliers or skewed data use the median; check the axes, scales and areas before trusting a chart.
Check yourself
12 questions. Every correct answer earns XP.