- Compute unbiased estimates of the mean and variance from a sample
- Build a confidence interval for the mean and interpret it correctly
- Carry out a z- or t-test and read the p-value correctly
- Compute the correlation coefficient and the least-squares regression line
Before an election a polling firm asks 1000 people and reports “52% ± 3%”. A tea factory weighs 30 packets rather than each of the millions it produces. A teacher wonders whether a new method really raised the scores or the class was just lucky. Statistical inference answers such questions: how to draw reliable conclusions about a whole population from a random sample, and how sure we can be.
Population, sample and estimators
The population is the whole set of objects we are interested in (all voters, all packets). The sample is the part we actually measure; it should be random, so that every object has the same chance of being chosen. Numbers that describe the population (μ, σ, p) are parameters; numbers computed from the sample (x̄, s, p̂) are statistics, and they estimate the parameters.
- x̄the sample mean, which estimates μ
- s²the sample variance, which estimates σ²
- sthe sample standard deviation
- nthe sample size
Both estimators are unbiased: averaged over all possible samples, E(x̄) = μ and E(s²) = σ².
Why n − 1? The deviations are measured from x̄, which was fitted to the same data, so on average they come out slightly too small. The deviations always add up to 0, so only n − 1 of them are free (the degrees of freedom). Dividing by n − 1 instead of n corrects this bias exactly.
A packet should hold 100 g of tea. Five packets weigh 98, 102, 101, 97 and 102 g. Find x̄, s² and s.
Show solutionHide solution
Deviations: −2, 2, 1, −3, 2 (sum 0 ✓); squares: 4, 4, 1, 9, 4, sum 22.
s² = 22/(5 − 1) = 5.5 g², s = √5.5 ≈ 2.35 g.
Dividing by 5 would give 4.4 g² and underestimate the spread.
Confidence intervals for the mean
A single number x̄ never hits μ exactly. A confidence interval is a range built so that in 95% of samples it contains μ. By the central limit theorem x̄ ≈ N(μ, σ²/n), so with probability 0.95, μ lies within 1.96 standard errors of x̄.
- z*the critical value: 1.645 for 90%, 1.96 for 95%, 2.576 for 99%
- σ/√nσ/√nthe standard error of the mean
- σthe known population standard deviation
Used when σ is known (or when n is large; then s may replace σ).
- t*the critical value of Student's t distribution with n − 1 degrees of freedom
- sthe sample standard deviation
Used when σ is unknown, especially for small samples: the t distribution has heavier tails than the normal one, so the interval is wider.
| Confidence level | z* | t* (df = 4) | t* (df = 9) | t* (df = 29) |
|---|---|---|---|---|
| 90% | 1.645 | 2.132 | 1.833 | 1.699 |
| 95% | 1.960 | 2.776 | 2.262 | 2.045 |
| 99% | 2.576 | 4.604 | 3.250 | 2.756 |
In a random sample of 36 students, the average daily study time is x̄ = 2.4 hours; from earlier surveys σ = 0.9 h. Build a 95% confidence interval for the mean μ.
Show solutionHide solution
Margin of error: 1.96 · 0.15 ≈ 0.29 h.
Interval: 2.4 ± 0.29, that is [2.11, 2.69] hours.
Interpretation: the method used here captures the true mean in 95% of samples.
Build a 95% confidence interval for the mean mass of the tea packets (n = 5, x̄ = 100 g, s ≈ 2.35 g).
Show solutionHide solution
Standard error: s/√n = 2.345/√5 ≈ 1.049 g.
Margin: 2.776 · 1.049 ≈ 2.91 g.
Interval: [97.09, 102.91] g. It contains 100 g, so the data do not contradict the label, but the interval is wide because n is small.
Hypothesis testing
- 1State the hypotheses
The null hypothesis H₀ is the “nothing special” claim (μ = 500 g, the method has no effect). The alternative H₁ is what we suspect (μ < 500 g, the method helps).
- 2Choose the significance level
α is the risk of wrongly rejecting a true H₀ that we accept; usually α = 0.05 (sometimes 0.01).
- 3Compute the test statistic
It measures how far the data are from H₀ in units of the standard error, for example z or t.
- 4Find the p-value
The probability, assuming H₀ is true, of a result at least as extreme as the one observed.
- 5Decide
If p ≤ α, reject H₀ (the result is statistically significant). If p > α, do not reject H₀: the data are not strong enough evidence, which is not a proof that H₀ is true.
- μ₀the mean claimed by H₀
- zthe test statistic when σ is known (normal distribution)
- twhen σ is unknown: t distribution with n − 1 degrees of freedom
A bakery claims that its loaves weigh 500 g on average, with σ = 12 g. An inspector weighs 36 loaves and finds x̄ = 495 g. Test at α = 0.05 whether the loaves are lighter than claimed.
Show solutionHide solution
z = (495 − 500)/(12/√36) = −5/2 = −2.5.
p-value = P(Z ≤ −2.5) = 1 − Φ(2.5) = 1 − 0.9938 ≈ 0.006.
p ≈ 0.006 < 0.05, so we reject H₀: the data are strong evidence that the loaves weigh less than 500 g on average.
Note: 5 g is a small difference for one loaf, but with 36 loaves the standard error is only 2 g.
| Decision | H₀ is true | H₀ is false |
|---|---|---|
| Reject H₀ | Type I error (probability α) | Correct decision (power 1 − β) |
| Do not reject H₀ | Correct decision | Type II error (probability β) |
The t-test. When σ is unknown we replace it with s and compare t with Student's t distribution (df = n − 1). A two-sample t-test compares the means of two groups, for example classes taught with a new and an old method, by dividing the difference of the means by its standard error. The t distribution was published in 1908 by William Gosset, who worked at the Guinness brewery in Dublin and wrote under the pen name “Student”.
Correlation and linear regression
- rPearson's correlation coefficient, −1 ≤ r ≤ 1
- r ≈ ±1the points lie almost on a straight line
- r ≈ 0no linear relationship
The regression line ŷ = b₀ + b₁x is chosen by the least squares method: it minimises the sum of squared vertical errors S(b₀, b₁) = ∑(yᵢ − b₀ − b₁xᵢ)². Setting the partial derivatives ∂S/∂b₀ and ∂S/∂b₁ equal to zero gives two linear equations (the normal equations), whose solution is:
- b₁the slope: the average change in y when x grows by 1
- b₀the intercept (ŷ when x = 0)
- ŷthe predicted value
The line always passes through the point (x̄, ȳ), and r² is the share of the variation in y explained by the line.
Five students: hours of study x = 1, 2, 3, 4, 5 and test scores y = 52, 58, 65, 70, 80. Find the regression line and r, and predict the score for 6 hours.
Show solutionHide solution
x − x̄: −2, −1, 0, 1, 2; y − ȳ: −13, −7, 0, 5, 15.
∑(x − x̄)(y − ȳ) = 26 + 7 + 0 + 5 + 30 = 68; ∑(x − x̄)² = 10; ∑(y − ȳ)² = 169 + 49 + 0 + 25 + 225 = 468.
b₁ = 68/10 = 6.8 points per hour, b₀ = 65 − 6.8 · 3 = 44.6, so ŷ = 44.6 + 6.8x.
r = 68/√(10 · 468) ≈ 68/68.41 ≈ 0.994, r² ≈ 0.99.
Prediction: ŷ(6) = 44.6 + 40.8 = 85.4 points.
import numpy as np
from scipy import stats
tea = np.array([98, 102, 101, 97, 102])
se = tea.std(ddof=1) / np.sqrt(len(tea))
low, high = stats.t.interval(0.95, df=4, loc=tea.mean(), scale=se)
print(round(low, 2), round(high, 2))
t, p = stats.ttest_1samp(tea, 103)
print(round(t, 3), round(p, 3))
hours = [1, 2, 3, 4, 5]
score = [52, 58, 65, 70, 80]
b1, b0 = np.polyfit(hours, score, 1)
print(round(b1, 2), round(b0, 2), round(np.corrcoef(hours, score)[0, 1], 3))▸ Expected output
97.09 102.91 -2.86 0.046 6.8 44.6 0.994
Key points
- x̄ and s² = ∑(xᵢ − x̄)²/(n − 1) are unbiased estimators of μ and σ²; the sample must be random.
- 95% confidence interval: x̄ ± 1.96σ/√n, or x̄ ± t*·s/√n when σ is unknown.
- Reject H₀ when p ≤ α; the p-value is not the probability that H₀ is true.
- Type I error: rejecting a true H₀ (α); type II error: keeping a false H₀ (β).
- Regression: b₁ = ∑(x − x̄)(y − ȳ)/∑(x − x̄)², b₀ = ȳ − b₁x̄; the line passes through (x̄, ȳ).
- r measures only linear association, and correlation does not imply causation.
Check yourself
10 questions. Every correct answer earns XP.