Skip to content
Educora
University32 min81 / 82

Statistical inference: estimation and hypothesis testing

From a sample to the population: estimators, confidence intervals for the mean, hypothesis testing with p-values, significance levels and type I/II errors, the t-test, correlation and least-squares linear regression.

Check yourself
In this lesson you will learn
  • Compute unbiased estimates of the mean and variance from a sample
  • Build a confidence interval for the mean and interpret it correctly
  • Carry out a z- or t-test and read the p-value correctly
  • Compute the correlation coefficient and the least-squares regression line

Before an election a polling firm asks 1000 people and reports “52% ± 3%”. A tea factory weighs 30 packets rather than each of the millions it produces. A teacher wonders whether a new method really raised the scores or the class was just lucky. Statistical inference answers such questions: how to draw reliable conclusions about a whole population from a random sample, and how sure we can be.

Population, sample and estimators

Definition
Population and sample

The population is the whole set of objects we are interested in (all voters, all packets). The sample is the part we actually measure; it should be random, so that every object has the same chance of being chosen. Numbers that describe the population (μ, σ, p) are parameters; numbers computed from the sample (x̄, s, p̂) are statistics, and they estimate the parameters.

x̄ = (x₁ + x₂ + … + xₙ) / n, s² = ∑(xᵢ − x̄)² / (n − 1)x̄ = (x₁ + x₂ + … + xₙ) / n, s² = ∑(xᵢ − x̄)² / (n − 1)
where:
  • x̄the sample mean, which estimates μ
  • s²the sample variance, which estimates σ²
  • sthe sample standard deviation
  • nthe sample size

Both estimators are unbiased: averaged over all possible samples, E(x̄) = μ and E(s²) = σ².

Why n − 1? The deviations are measured from x̄, which was fitted to the same data, so on average they come out slightly too small. The deviations always add up to 0, so only n − 1 of them are free (the degrees of freedom). Dividing by n − 1 instead of n corrects this bias exactly.

Five packets of tea

A packet should hold 100 g of tea. Five packets weigh 98, 102, 101, 97 and 102 g. Find x̄, s² and s.

Show solution
x̄ = (98 + 102 + 101 + 97 + 102)/5 = 500/5 = 100 g.
Deviations: −2, 2, 1, −3, 2 (sum 0 ✓); squares: 4, 4, 1, 9, 4, sum 22.
s² = 22/(5 − 1) = 5.5 g², s = √5.5 ≈ 2.35 g.
Dividing by 5 would give 4.4 g² and underestimate the spread.

Confidence intervals for the mean

A single number x̄ never hits μ exactly. A confidence interval is a range built so that in 95% of samples it contains μ. By the central limit theorem x̄ ≈ N(μ, σ²/n), so with probability 0.95, μ lies within 1.96 standard errors of x̄.

x̄ ± z* · σ/√nx̄ ± z* · σ/√n
where:
  • z*the critical value: 1.645 for 90%, 1.96 for 95%, 2.576 for 99%
  • σ/√nσ/√nthe standard error of the mean
  • σthe known population standard deviation

Used when σ is known (or when n is large; then s may replace σ).

x̄ ± t* · s/√n, df = n − 1x̄ ± t* · s/√n, df = n − 1
where:
  • t*the critical value of Student's t distribution with n − 1 degrees of freedom
  • sthe sample standard deviation

Used when σ is unknown, especially for small samples: the t distribution has heavier tails than the normal one, so the interval is wider.

Confidence levelz*t* (df = 4)t* (df = 9)t* (df = 29)
90%1.6452.1321.8331.699
95%1.9602.7762.2622.045
99%2.5764.6043.2502.756
Critical values for two-sided intervals. As n grows, t* approaches z*.
Daily study time

In a random sample of 36 students, the average daily study time is x̄ = 2.4 hours; from earlier surveys σ = 0.9 h. Build a 95% confidence interval for the mean μ.

Show solution
Standard error: σ/√n = 0.9/√36 = 0.9/6 = 0.15 h.
Margin of error: 1.96 · 0.15 ≈ 0.29 h.
Interval: 2.4 ± 0.29, that is [2.11, 2.69] hours.
Interpretation: the method used here captures the true mean in 95% of samples.
A small sample: the t interval

Build a 95% confidence interval for the mean mass of the tea packets (n = 5, x̄ = 100 g, s ≈ 2.35 g).

Show solution
σ is unknown and n is small, so we use t with df = 4: t* = 2.776.
Standard error: s/√n = 2.345/√5 ≈ 1.049 g.
Margin: 2.776 · 1.049 ≈ 2.91 g.
Interval: [97.09, 102.91] g. It contains 100 g, so the data do not contradict the label, but the interval is wide because n is small.

Hypothesis testing

  1. 1
    State the hypotheses

    The null hypothesis H₀ is the “nothing special” claim (μ = 500 g, the method has no effect). The alternative H₁ is what we suspect (μ < 500 g, the method helps).

  2. 2
    Choose the significance level

    α is the risk of wrongly rejecting a true H₀ that we accept; usually α = 0.05 (sometimes 0.01).

  3. 3
    Compute the test statistic

    It measures how far the data are from H₀ in units of the standard error, for example z or t.

  4. 4
    Find the p-value

    The probability, assuming H₀ is true, of a result at least as extreme as the one observed.

  5. 5
    Decide

    If p ≤ α, reject H₀ (the result is statistically significant). If p > α, do not reject H₀: the data are not strong enough evidence, which is not a proof that H₀ is true.

z = (x̄ − μ₀) / (σ/√n), t = (x̄ − μ₀) / (s/√n)z = (x̄ − μ₀) / (σ/√n), t = (x̄ − μ₀) / (s/√n)
where:
  • μ₀the mean claimed by H₀
  • zthe test statistic when σ is known (normal distribution)
  • twhen σ is unknown: t distribution with n − 1 degrees of freedom
Are the loaves too light?

A bakery claims that its loaves weigh 500 g on average, with σ = 12 g. An inspector weighs 36 loaves and finds x̄ = 495 g. Test at α = 0.05 whether the loaves are lighter than claimed.

Show solution
H₀: μ = 500 g; H₁: μ < 500 g (a one-sided test).
z = (495 − 500)/(12/√36) = −5/2 = −2.5.
p-value = P(Z ≤ −2.5) = 1 − Φ(2.5) = 1 − 0.9938 ≈ 0.006.
p ≈ 0.006 < 0.05, so we reject H₀: the data are strong evidence that the loaves weigh less than 500 g on average.
Note: 5 g is a small difference for one loaf, but with 36 loaves the standard error is only 2 g.
DecisionH₀ is trueH₀ is false
Reject H₀Type I error (probability α)Correct decision (power 1 − β)
Do not reject H₀Correct decisionType II error (probability β)
A smaller α means fewer false alarms but more missed effects; a larger sample reduces both errors.

The t-test. When σ is unknown we replace it with s and compare t with Student's t distribution (df = n − 1). A two-sample t-test compares the means of two groups, for example classes taught with a new and an old method, by dividing the difference of the means by its standard error. The t distribution was published in 1908 by William Gosset, who worked at the Guinness brewery in Dublin and wrote under the pen name “Student”.

Correlation and linear regression

r = ∑(xᵢ − x̄)(yᵢ − ȳ) / √(∑(xᵢ − x̄)² · ∑(yᵢ − ȳ)²)r = ∑(xᵢ − x̄)(yᵢ − ȳ) / √(∑(xᵢ − x̄)² · ∑(yᵢ − ȳ)²)
where:
  • rPearson's correlation coefficient, −1 ≤ r ≤ 1
  • r ≈ ±1the points lie almost on a straight line
  • r ≈ 0no linear relationship

The regression line ŷ = b₀ + b₁x is chosen by the least squares method: it minimises the sum of squared vertical errors S(b₀, b₁) = ∑(yᵢ − b₀ − b₁xᵢ)². Setting the partial derivatives ∂S/∂b₀ and ∂S/∂b₁ equal to zero gives two linear equations (the normal equations), whose solution is:

b₁ = ∑(xᵢ − x̄)(yᵢ − ȳ) / ∑(xᵢ − x̄)², b₀ = ȳ − b₁ · x̄b₁ = ∑(xᵢ − x̄)(yᵢ − ȳ) / ∑(xᵢ − x̄)², b₀ = ȳ − b₁ · x̄
where:
  • b₁the slope: the average change in y when x grows by 1
  • b₀the intercept (ŷ when x = 0)
  • ŷthe predicted value

The line always passes through the point (x̄, ȳ), and r² is the share of the variation in y explained by the line.

Study hours and scores

Five students: hours of study x = 1, 2, 3, 4, 5 and test scores y = 52, 58, 65, 70, 80. Find the regression line and r, and predict the score for 6 hours.

Show solution
x̄ = 3, ȳ = 325/5 = 65.
x − x̄: −2, −1, 0, 1, 2; y − ȳ: −13, −7, 0, 5, 15.
∑(x − x̄)(y − ȳ) = 26 + 7 + 0 + 5 + 30 = 68; ∑(x − x̄)² = 10; ∑(y − ȳ)² = 169 + 49 + 0 + 25 + 225 = 468.
b₁ = 68/10 = 6.8 points per hour, b₀ = 65 − 6.8 · 3 = 44.6, so ŷ = 44.6 + 6.8x.
r = 68/√(10 · 468) ≈ 68/68.41 ≈ 0.994, r² ≈ 0.99.
Prediction: ŷ(6) = 44.6 + 40.8 = 85.4 points.
Interactive
Loading simulation…
The least-squares line ŷ = 44.6 + 6.8x (a = b₀, b = b₁). It passes close to the points (1, 52), (2, 58), (3, 65), (4, 70), (5, 80); any other a or b makes the sum of squared errors larger than 5.6.
Python
import numpy as np
from scipy import stats

tea = np.array([98, 102, 101, 97, 102])
se = tea.std(ddof=1) / np.sqrt(len(tea))
low, high = stats.t.interval(0.95, df=4, loc=tea.mean(), scale=se)
print(round(low, 2), round(high, 2))

t, p = stats.ttest_1samp(tea, 103)
print(round(t, 3), round(p, 3))

hours = [1, 2, 3, 4, 5]
score = [52, 58, 65, 70, 80]
b1, b0 = np.polyfit(hours, score, 1)
print(round(b1, 2), round(b0, 2), round(np.corrcoef(hours, score)[0, 1], 3))
▸ Expected output
97.09 102.91
-2.86 0.046
6.8 44.6 0.994
scipy checks the examples: the t interval for the tea packets, a t-test of H₀: μ = 103 g (t = −2.86, p = 0.046: rejected at α = 0.05 but not at α = 0.01) and the regression line with r.

Key points

  • x̄ and s² = ∑(xᵢ − x̄)²/(n − 1) are unbiased estimators of μ and σ²; the sample must be random.
  • 95% confidence interval: x̄ ± 1.96σ/√n, or x̄ ± t*·s/√n when σ is unknown.
  • Reject H₀ when p ≤ α; the p-value is not the probability that H₀ is true.
  • Type I error: rejecting a true H₀ (α); type II error: keeping a false H₀ (β).
  • Regression: b₁ = ∑(x − x̄)(y − ȳ)/∑(x − x̄)², b₀ = ȳ − b₁x̄; the line passes through (x̄, ȳ).
  • r measures only linear association, and correlation does not imply causation.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
n = 100, x̄ = 50, σ = 10. What is the 95% confidence interval for μ?