- Build a chart the object-oriented way with
fig, ax = plt.subplots()and add a title, axis labels and a legend - Choose the chart type that fits the question: line, bar, scatter or histogram
- Justify the number of histogram bins with the Sturges and Freedman–Diaconis rules
- Recognise misleading charts and draw honest, readable ones
In 1973 the statistician Francis Anscombe published four small datasets whose means, variances, correlation coefficient and regression line are almost identical. Yet on a chart one is scattered around a straight line, one follows a curve, in the third a single outlier changes everything, and in the fourth all points but one lie on a vertical line. The lesson: before trusting the numbers, look at the data. In Python the main tool for this is matplotlib.
Figure and Axes: how a chart is built
The Figure is the whole image, the “canvas”. An Axes is one plotting area inside it: the axes, lines, title and legend. A Figure can contain several Axes. Most methods belong to the Axes: ax.plot, ax.set_title, ax.legend.
matplotlib has two styles: the short “pyplot” style such as plt.plot(...) and the object-oriented style. Always start with the second: fig, ax = plt.subplots() — such code stays clear even with several charts. In the browser the image appears under the code automatically; on your computer plt.show() opens a window and fig.savefig('chart.png', dpi=200) saves a file.
import matplotlib.pyplot as plt
months = ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun']
baku = [279, 153, 426, 390, 455, 510]
ganja = [150, 306, 168, 210, 240, 265]
fig, ax = plt.subplots(figsize=(7, 4))
ax.plot(months, baku, marker='o', label='Baku')
ax.plot(months, ganja, marker='s', linestyle='--', label='Ganja')
ax.set_title('Monthly revenue by branch, 2026')
ax.set_xlabel('Month')
ax.set_ylabel('Revenue, AZN')
ax.grid(alpha=0.3)
ax.legend()
plt.show()- Title — what the chart shows, where and when.
- Axis labels with units: “Revenue, AZN”, “Height, cm”.
- Legend — whenever there is more than one series.
- A faint grid (
grid(alpha=0.3)) — makes values easier to read without distracting.
Choosing the chart type
| Question | Chart | matplotlib |
|---|---|---|
| How does it change over time? | line chart | ax.plot |
| How do categories compare? | bar chart | ax.bar, ax.barh |
| Are two numeric variables related? | scatter plot | ax.scatter |
| How are the values distributed? | histogram, box plot | ax.hist, ax.boxplot |
import numpy as np
import matplotlib.pyplot as plt
products = np.array(['Coffee', 'Cake', 'Tea', 'Juice', 'Sandwich'])
revenue = np.array([564, 540, 378, 212, 305])
order = np.argsort(revenue)
fig, ax = plt.subplots(figsize=(6, 3.5))
bars = ax.barh(products[order], revenue[order], color='#4C72B0')
ax.bar_label(bars, fmt='%d AZN', padding=3)
ax.set_xlim(0, 680)
ax.set_xlabel('Revenue, AZN')
ax.set_title('Revenue by product, Q1 2026')
print(products[order][-1], revenue.max())
plt.show()▸ Expected output
Coffee 564
A scatter plot shows the relationship between two numeric variables. Colour can carry a third variable (c=, cmap=), and np.polyfit(x, y, 1) finds the trend line by least squares. The data below is invented: the area, number of rooms and price of flats.
import numpy as np
import matplotlib.pyplot as plt
rng = np.random.default_rng(3)
area = rng.uniform(40, 140, 60)
rooms = np.clip(area // 30, 1, 4).astype(int)
price = 1.5 * area + 10 * rooms + rng.normal(0, 15, 60)
slope, intercept = np.polyfit(area, price, 1)
print(f'slope: {slope:.2f} thousand AZN per m2')
print(f'correlation r = {np.corrcoef(area, price)[0, 1]:.3f}')
fig, ax = plt.subplots(figsize=(6, 4))
points = ax.scatter(area, price, c=rooms, cmap='viridis', alpha=0.8)
xs = np.array([40, 140])
ax.plot(xs, slope * xs + intercept, color='red', label='trend line')
fig.colorbar(points, ax=ax, label='rooms')
ax.set_xlabel('Area, m²')
ax.set_ylabel('Price, thousand AZN')
ax.legend()
plt.show()▸ Expected output
slope: 2.00 thousand AZN per m2 correlation r = 0.960
viridis colour map stays readable in black-and-white print and for people with colour-vision deficiency. The correlation coefficient r is explained in detail in the next lesson.Histograms and the number of bins
A histogram divides the value axis into equal intervals (bins) and counts the observations that fall into each one. With too few bins the shape of the distribution is lost; with too many the chart turns into a noisy “comb”. Two well-known rules give a starting point:
- knumber of bins (Sturges' rule)
- hbin width (Freedman–Diaconis rule), in the data's units
- nnumber of observations
- IQRinterquartile range Q3 − Q1
Sturges suits roughly normal, not too large datasets; Freedman–Diaconis is robust to outliers. ax.hist(x, bins='fd') applies the second one automatically.
import numpy as np
import matplotlib.pyplot as plt
rng = np.random.default_rng(42)
heights = rng.normal(170, 8, 500)
k = int(np.ceil(np.log2(heights.size))) + 1
print('Sturges bins:', k)
print(f'mean = {heights.mean():.1f} cm, std = {heights.std(ddof=1):.1f} cm')
fig, ax = plt.subplots(figsize=(6, 4))
ax.hist(heights, bins=k, edgecolor='white')
ax.axvline(heights.mean(), color='red', linestyle='--', label='mean')
ax.set_xlabel('Height, cm')
ax.set_ylabel('Number of people')
ax.legend()
plt.show()▸ Expected output
Sturges bins: 10 mean = 169.9 cm, std = 7.7 cm
n = 500, the heights are roughly normal with σ ≈ 8 cm. Compute the bins with the Sturges and Freedman–Diaconis rules. For a normal distribution IQR ≈ 1.35σ.
Show solutionHide solution
Freedman–Diaconis: IQR ≈ 1.35 · 8 = 10.8 cm; n^(1/3) = 500^(1/3) ≈ 7.94.
h = 2 · 10.8 / 7.94 ≈ 2.7 cm.
The range is about 170 ± 3σ, i.e. ~48 cm → 48 / 2.7 ≈ 18 bins.
FD shows finer detail; as n grows, Sturges adds bins very slowly.
- nᵢobservations in bin i
- fᵢbar height with
density=True(density) - hbin width
ax.hist(x, density=True) makes the bar areas add up to 1 — such a histogram can be compared with a probability density curve.
200 measurements are split into bins of width h = 2. One bin holds 30 measurements. What is the height of this bar with density=True? What does its area show?
Show solutionHide solution
Area = f · h = 0.075 · 2 = 0.15, so 15% of the measurements are in this bin (30 / 200).
The areas of all bars add up to 1 (100%).
Several charts in one figure
plt.subplots(2, 2) returns a 2 × 2 array of Axes; you reach each chart with axes[row, col]. layout='constrained' keeps titles and axes from overlapping, and sharex=True makes the axes shared.
import numpy as np
import matplotlib.pyplot as plt
rng = np.random.default_rng(0)
x = np.linspace(0, 2 * np.pi, 100)
fig, axes = plt.subplots(2, 2, figsize=(8, 6), layout='constrained')
axes[0, 0].plot(x, np.sin(x))
axes[0, 0].set_title('Line: sin x')
axes[0, 1].bar(['A', 'B', 'C'], [5, 3, 7])
axes[0, 1].set_title('Bar')
axes[1, 0].scatter(rng.random(30), rng.random(30))
axes[1, 0].set_title('Scatter')
axes[1, 1].hist(rng.normal(size=300), bins=15)
axes[1, 1].set_title('Histogram')
fig.suptitle('Four basic chart types')
print(axes.shape)
plt.show()▸ Expected output
(2, 2)
pandas has its own .plot() method built on matplotlib: each DataFrame column becomes a series and the index goes on the X axis. The method returns an Axes, so you can finish the chart with ordinary matplotlib commands.
import pandas as pd
import matplotlib.pyplot as plt
df = pd.DataFrame({'month': ['Jan', 'Feb', 'Mar'],
'Baku': [279, 153, 426],
'Ganja': [150, 306, 168]}).set_index('month')
ax = df.plot(kind='bar', figsize=(6, 3.5), rot=0, title='Revenue by month, AZN')
ax.set_ylabel('AZN')
print(df.sum())
plt.show()▸ Expected output
Baku 858 Ganja 624 dtype: int64
Rules of a good chart
- One chart, one idea. Put the conclusion in the title: “Baku branch leads in March”.
- On a bar chart the Y axis must start at 0, because the eye compares bar lengths.
- Make colours meaningful: give the series you want to highlight a bright colour and the rest grey.
- Avoid 3D effects, shadows and pie charts with many slices — they distort values.
import matplotlib.pyplot as plt
branches = ['Baku', 'Ganja']
satisfied = [96, 98]
fig, (bad, good) = plt.subplots(1, 2, figsize=(8, 3.5), layout='constrained')
bad.bar(branches, satisfied, color=['gray', 'orange'])
bad.set_ylim(95, 98.5)
bad.set_title('Misleading: axis starts at 95')
good.bar(branches, satisfied, color=['gray', 'orange'])
good.set_ylim(0, 100)
good.set_title('Honest: axis starts at 0')
for ax in (bad, good):
ax.set_ylabel('Satisfied customers, %')
plt.show()Build a horizontal bar chart (ax.barh) from the revenue dictionary, set the title to 'Revenue by product' and label the X axis 'AZN'. Then print on separate lines: 1) the chart's title (ax.get_title()); 2) the number of bars on the chart (len(ax.patches)); 3) the product with the highest revenue.
import matplotlib.pyplot as plt
revenue = {'Coffee': 564, 'Cake': 540, 'Tea': 378, 'Juice': 212, 'Sandwich': 305}
fig, ax = plt.subplots()
# draw the bars, set the title and the X label, then print the three results▸ Expected output
Revenue by product 5 Coffee
Build a histogram of the test scores with 5 bins. ax.hist returns three things: the counts, the bin edges and the bars. Print on separate lines: 1) the count in each bin as a list of integers; 2) the width of one bin rounded to 2 decimal places.
import matplotlib.pyplot as plt
scores = [55, 62, 68, 70, 71, 75, 78, 80, 82, 85, 88, 90, 94, 97]
fig, ax = plt.subplots()
# counts, edges, _ = ax.hist(...)▸ Expected output
[2, 3, 3, 3, 3] 8.4
Key points
- The Figure is the whole image and an Axes is a chart inside it; start with
fig, ax = plt.subplots(). - Time → line, categories → bars, two numeric variables → scatter, distribution → histogram.
- Every chart needs a title, axis labels with units and, when needed, a legend.
- Bins: Sturges k = ⌈log₂ n⌉ + 1, Freedman–Diaconis h = 2 · IQR · n^(−1/3).
- A bar chart's axis starts at 0; choose
viridisfor colour maps.
Check yourself
10 questions. Every correct answer earns XP.