Skip to content
Educora
Intermediate28 min23 / 68

The normal distribution

Discrete random variables and their distribution tables, expectation and variance (including the binomial distribution), the normal curve N(μ, σ²) and how μ and σ shape it, the 68–95–99.7 rule, probabilities by symmetry, the 3σ principle and standardisation.

Check yourself
In this lesson you will learn
  • Build a distribution table and compute E(X), D(X), E(aX + b) and D(aX + b), including the binomial case.
  • Read μ and σ from N(μ, σ²) and describe how they shape the normal curve.
  • Find normal probabilities with symmetry and the 68–95–99.7 rule, and apply the 3σ principle.
  • Standardise a normal variable with z = (x − μ)/σ and compare results from different distributions.

A machine fills bags of rice labelled «1 kg». Weigh 1,000 bags and you will find many very close to 1,000 g, fewer at 990 g or 1,010 g and almost none at 970 g. Draw the histogram of the masses with narrower and narrower classes, and its outline turns into a smooth, symmetric bell — the normal curve. Heights, measurement errors and exam scores behave the same way. The CSCA syllabus asks for the «basic concepts of the normal distribution»; to use them we first need random variables with their expectation and variance — the probability versions of the mean and variance from «Numerical characteristics of data». Counting outcomes is covered in «Classical probability».

Random variables, expectation and variance

Definition
Discrete random variable and its distribution (分布列)

A random variable X (随机变量) is a quantity whose value is decided by the outcome of a random experiment. It is discrete if its values can be listed, e.g. the number of heads when two coins are tossed. Its distribution table lists every value xᵢ with its probability pᵢ; always pᵢ ≥ 0 and p₁ + p₂ + … + pₙ = 1.

X012
P1/41/21/4
X = the number of heads when two fair coins are tossed: the outcomes HH, HT, TH, TT are equally likely.
E(X) = x₁p₁ + x₂p₂ + … + xₙpₙ, D(X) = (x₁ − E(X))²p₁ + … + (xₙ − E(X))²pₙ = E(X²) − [E(X)]², σ(X) = √D(X)
where:
  • E(X)expectation (mean value, 数学期望)
  • D(X)variance (方差); English books often write Var(X)
  • σ(X)standard deviation
  • E(X²)x₁²p₁ + x₂²p₂ + … + xₙ²pₙ

E(X) is the long-run average value of X and D(X) measures how far X scatters around it. These are the formulas of the data lesson with the relative frequencies replaced by probabilities. Chinese textbooks write E(X) and D(X).

Worked examples: expectation and variance

1) Find E(X) and D(X) for the number of heads in two coin tosses (table above).
2) X takes the values −1, 0, 1, 2 with probabilities 0.1, a, 0.3, 0.2. Find a, E(X) and D(X).
3) A lottery ticket costs 5 yuan. It wins 20 yuan with probability 0.1 and 50 yuan with probability 0.02; otherwise it wins nothing. What is the expected profit per ticket?

Show solution
1) E(X) = 0 · ¼ + 1 · ½ + 2 · ¼ = 1; E(X²) = 0 + ½ + 4 · ¼ = 1.5; D(X) = 1.5 − 1² = 0.5.
2) 0.1 + a + 0.3 + 0.2 = 1 ⇒ a = 0.4. E(X) = −0.1 + 0 + 0.3 + 0.4 = 0.6; E(X²) = 0.1 + 0 + 0.3 + 0.8 = 1.2; D(X) = 1.2 − 0.36 = 0.84.
3) Expected prize: 20 · 0.1 + 50 · 0.02 = 2 + 1 = 3 yuan; expected profit: 3 − 5 = −2 yuan — on average the buyer loses 2 yuan per ticket.
E(aX + b) = a·E(X) + b, D(aX + b) = a²·D(X)
where:
  • a, bconstants

The same rules as for data (see «Numerical characteristics of data»): the shift b moves the expectation but never changes the variance.

Worked examples: E(aX + b) and D(aX + b)

1) E(X) = 3 and D(X) = 2. Find E(2X − 1) and D(2X − 1).
2) E(X) = 1 and D(X) = 0.5. Find E(4 − 3X) and D(4 − 3X).

Show solution
1) E = 2 · 3 − 1 = 5; D = 2² · 2 = 8.
2) E = 4 − 3 · 1 = 1; D = (−3)² · 0.5 = 4.5 — the minus sign disappears when squared.
CSCA-style item: D(2X + 1)

X takes the values 1, 2, 3 with probabilities 0.3, 0.4, 0.3. What is D(2X + 1)?
A) 1.2 B) 2.4 C) 3.4 D) 0.6

Show solution
E(X) = 0.3 + 0.8 + 0.9 = 2; E(X²) = 0.3 + 1.6 + 2.7 = 4.6; D(X) = 4.6 − 4 = 0.6. D(2X + 1) = 2² · 0.6 = 2.4 (B).
A forgets to square the 2, C also adds the 1, and D is D(X) itself.

Two special cases appear again and again. In the two-point distribution (两点分布) X is 1 («success») with probability p and 0 otherwise. The binomial distribution B(n, p) (二项分布) counts the successes in n independent repetitions of the same trial, e.g. the hits in 10 shots; its probabilities P(X = k) = Cₙᵏpᵏ(1 − p)ⁿ⁻ᵏ come from the counting of «Classical probability». For both, E and D have ready-made formulas.

X ~ B(1, p): E(X) = p, D(X) = p(1 − p); X ~ B(n, p): E(X) = np, D(X) = np(1 − p)
where:
  • nnumber of independent trials
  • pprobability of success in one trial
  • 1 − pprobability of failure in one trial

B(1, p) is the two-point distribution. For a binomial variable D/E = 1 − p — a quick way to recover p and n.

Worked examples: the binomial distribution

1) A basketball player scores a free throw with probability 0.8. She takes 10 independent throws; X is the number of throws she scores. Find E(X) and D(X).
2) X ~ B(n, p) with E(X) = 6 and D(X) = 2.4. Find n and p.
3) A shot hits the target with probability 0.7; X = 1 for a hit and X = 0 for a miss. Find E(X) and D(X).

Show solution
1) E(X) = 10 · 0.8 = 8; D(X) = 10 · 0.8 · 0.2 = 1.6.
2) 1 − p = D/E = 2.4/6 = 0.4, so p = 0.6 and n = 6/0.6 = 10.
3) A two-point distribution: E(X) = 0.7, D(X) = 0.7 · 0.3 = 0.21.

The normal curve N(μ, σ²)

Mass, height or time can take any value in an interval, so a single value has probability 0 and we work with areas instead: P(a < X < b) is the area under a density curve between x = a and x = b, and the total area under the curve is 1. For the rice bags this curve is the normal curve.

Definition
Normal distribution (正态分布)

X follows the normal distribution with parameters μ and σ (σ > 0), written X ~ N(μ, σ²), if its density curve is given by the formula below. Then E(X) = μ and D(X) = σ². N(0, 1) is the standard normal distribution.

f(x) = 1/(σ√(2π)) · e^(−(x − μ)²/(2σ²)), x ∈ ℝf(x) = 1/(σ√(2π)) · e^(−(x − μ)²/(2σ²)), x ∈ ℝ
where:
  • μthe mean: the axis of symmetry of the curve and the position of its peak
  • σthe standard deviation: the width of the bell; in N(μ, σ²) the second number is σ², not σ
  • 1/(σ√(2π))1/(σ√(2π))the height of the peak

In the CSCA you never integrate this function; you need its shape and a few areas.

  • The curve lies above the x-axis and is symmetric about the line x = μ; so P(X < μ) = P(X > μ) = 0.5.
  • The peak is at x = μ, with height 1/(σ√(2π)); moving away from μ on either side, the curve approaches the x-axis without touching it.
  • The total area under the curve is 1.
  • μ moves the curve left or right without changing its shape.
  • σ changes the shape: a small σ gives a tall, narrow curve (the data are concentrated), a large σ a low, wide one (the data are scattered).
  • P(X = a) = 0, so P(X < a) = P(X ≤ a): whether the ends are included does not matter.
Interactive
Loading simulation…
The normal density with m = μ and s = σ; the shaded area from a to b is P(a < X < b). Start with m = 0, s = 1, a = −1, b = 1: the area is about 0.683. Set a = −2, b = 2 (≈ 0.954), then a = −3, b = 3 (≈ 0.997). Now change m and s and move a and b to m − s and m + s: the area is again ≈ 0.683.
Worked examples: reading N(μ, σ²)

1) X ~ N(5, 9). Give the axis of symmetry, σ and the height of the peak.
2) A normal density is f(x) = 1/(2√(2π)) · e^(−(x + 1)²/8). Write it as N(μ, σ²).
3) Compare the curves of N(0, 1), N(0, 4) and N(2, 1).

Show solution
1) Axis x = 5; σ² = 9, so σ = 3; peak 1/(3√(2π)).
2) (x + 1)² = (x − (−1))², so μ = −1; 2σ² = 8 gives σ² = 4 (σ = 2, which matches the factor 1/(2√(2π))): N(−1, 4).
3) N(0, 1) and N(2, 1) have the same shape; the second is shifted 2 units to the right. N(0, 4) has the same axis as N(0, 1) but σ = 2: it is lower and wider, with half the peak height.
CSCA-style item: three curves

The density curves of X ~ N(μ₁, σ₁²), Y ~ N(μ₂, σ₂²) and Z ~ N(μ₃, σ₃²) are drawn together. The curves of X and Y have the same axis of symmetry x = 1, and the curve of Y is lower and wider; the curve of Z has the same shape as that of X, but its axis is x = 3. Which statement is correct?
A) σ₁ > σ₂ B) μ₁ = μ₂ < μ₃ and σ₁ = σ₃ < σ₂ C) μ₃ < μ₁ D) P(Y < 1) > P(X < 1)

Show solution
Same axis → μ₁ = μ₂ = 1 < μ₃ = 3; same shape → σ₁ = σ₃; lower and wider → larger σ, so σ₂ > σ₁: B.
A reverses the effect of σ; C reverses the order of the axes; D is false — both probabilities are 0.5.

The 68–95–99.7 rule and symmetry

P(μ − σ ≤ X ≤ μ + σ) ≈ 0.6827, P(μ − 2σ ≤ X ≤ μ + 2σ) ≈ 0.9545, P(μ − 3σ ≤ X ≤ μ + 3σ) ≈ 0.9973
where:
  • μ ± kσthe interval of k standard deviations around the mean
  • μ, σthe parameters of the normal distribution

Chinese textbooks give exactly these values (older books: 0.6826, 0.9544, 0.9974); CSCA items usually print them in the question. They hold for every normal distribution, whatever μ and σ are.

34.1%34.1%13.6%13.6%2.1%2.1%μ−3σμ−2σμ−σμμ+σμ+2σμ+3σ68.27%95.45%99.73%
Only about 0.27% of the values lie outside μ ± 3σ (about 0.13% in each tail).
  1. 1
    Find the axis

    Read μ from N(μ, σ²): the curve is symmetric about the line x = μ.

  2. 2
    Mirror

    Points at equal distances from μ have equal tails: P(X < μ − a) = P(X > μ + a). If P(X < c) = P(X > d), then μ = (c + d)/2.

  3. 3
    Use 1 and 0.5

    The whole area is 1 and each half is 0.5: P(X > c) = 1 − P(X < c), P(μ < X < μ + a) = 0.5 − P(X > μ + a).

  4. 4
    Sketch

    A quick sketch of the bell with the given numbers prevents most sign mistakes.

Worked examples: symmetry

1) X ~ N(1, σ²) and P(X < −2) = 0.15. Find P(1 < X < 4).
2) X ~ N(μ, σ²) and P(X < 2) = P(X > 8). Find μ.
3) X ~ N(4, σ²) and P(X < a) = P(X > a + 2). Find a.

Show solution
1) −2 and 4 are both 3 units from μ = 1, so P(X > 4) = 0.15 and P(1 < X < 4) = 0.5 − 0.15 = 0.35.
2) μ = (2 + 8)/2 = 5.
3) The points a and a + 2 are symmetric about 4: (a + a + 2)/2 = 4, so a = 3.
CSCA-style item: from one given piece

X ~ N(2, σ²) and P(0 < X < 2) = 0.3. What is P(X > 4)?
A) 0.2 B) 0.3 C) 0.7 D) 0.8

Show solution
By symmetry P(2 < X < 4) = 0.3, and the right half has area 0.5: P(X > 4) = 0.5 − 0.3 = 0.2 (A).
B is P(2 < X < 4) itself, C = 1 − 0.3 forgets the half, and D = 0.5 + 0.3 is P(X < 4).
Worked examples: the 68–95–99.7 rule

X ~ N(50, 25).
1) Find P(45 < X < 55) and P(40 < X < 60).
2) Find P(X > 60).
3) Find P(45 < X < 60).
4) The scores of 1,000 students follow this distribution. About how many scored more than 60?

Show solution
σ = √25 = 5.
1) 45–55 is μ ± σ: ≈ 0.6827; 40–60 is μ ± 2σ: ≈ 0.9545.
2) The two tails outside μ ± 2σ share 1 − 0.9545 = 0.0455, so P(X > 60) ≈ 0.0228 (0.02275).
3) From μ − σ to μ: 0.6827/2; from μ to μ + 2σ: 0.9545/2; together ≈ 0.3413 + 0.4773 = 0.8186.
4) 1000 · 0.0228 ≈ 23 students.

The 3σ principle (3σ原则): a value outside (μ − 3σ, μ + 3σ) occurs with probability only about 0.0027, so in a single observation it is treated as practically impossible. Factories use this for quality control: if a part made by a machine falls outside this interval, the machine is stopped and checked.

CSCA-style item: the 3σ principle

The lengths (in mm) of nails made by a machine follow N(50, 0.09). Four nails from one hour are measured. Which reading suggests that the machine is out of order?
A) 49.3 mm B) 50.8 mm C) 49.6 mm D) 51.0 mm

Show solution
σ = √0.09 = 0.3, so μ ± 3σ = (49.1, 50.9). Only 51.0 mm lies outside: D.
The other three lie inside (49.1, 50.9), so they are normal readings. Taking 0.09 as σ would make all four nails look faulty.

Standardisation

Every normal variable can be turned into the standard one by measuring its distance from μ in units of σ. Then one table of Φ (or a few given values) serves all normal distributions, and results from different distributions can be compared.

Z = (X − μ)/σ ~ N(0, 1), P(X < x) = Φ((x − μ)/σ), Φ(−z) = 1 − Φ(z), Φ(0) = 0.5Z = (X − μ)/σ ~ N(0, 1), P(X < x) = Φ((x − μ)/σ), Φ(−z) = 1 − Φ(z), Φ(0) = 0.5
where:
  • Zthe standardised variable (z-score): how many σ a value lies from μ
  • Φ(z)P(Z < z), the distribution function of N(0, 1)

A z-score of 2 means «two standard deviations above the mean», so the 68–95–99.7 rule is really a statement about Z: P(|Z| < 1) ≈ 0.6827, P(|Z| < 2) ≈ 0.9545, P(|Z| < 3) ≈ 0.9973.

Worked examples: z-scores

1) X ~ N(70, 25) and Φ(2) ≈ 0.9772. Find P(X < 80) and P(X < 60).
2) X ~ N(60, 16) and Φ(1.5) ≈ 0.9332. Find P(54 < X < 66).
3) Murad scored 84 in mathematics, where the scores follow N(72, 36), and 81 in physics, where they follow N(66, 100). In which subject did he do relatively better?

Show solution
1) z = (80 − 70)/5 = 2: P(X < 80) = Φ(2) ≈ 0.9772; z = −2: P(X < 60) = 1 − 0.9772 = 0.0228 (the same as the 2σ-tail found above).
2) z runs from (54 − 60)/4 = −1.5 to 1.5: P = Φ(1.5) − Φ(−1.5) = 2Φ(1.5) − 1 ≈ 0.8664.
3) Mathematics: z = (84 − 72)/6 = 2; physics: z = (81 − 66)/10 = 1.5. Compared with the other candidates Murad did better in mathematics, although the raw scores are close.

How CSCA asks about this

  • Distribution tables: a missing probability (the sum is 1), then E(X) and D(X) = E(X²) − [E(X)]²; the expected profit of a game.
  • Rules: E(aX + b) and D(aX + b); binomial E = np, D = np(1 − p), also backwards (from E and D to n and p).
  • Reading N(μ, σ²): μ, σ (the second number is σ²!), the axis, the effect of μ and σ on the curve, comparing two or three curves.
  • Symmetry: P(X < μ) = 0.5, equal tails, «P(X < c) = P(X > d) ⇒ μ = (c + d)/2», the probability between two points from one given piece.
  • 68–95–99.7 and 3σ: probabilities of σ-intervals, numbers of people (N · P), quality control, standardising with a given value of Φ.
Term中文Pinyin
random variable随机变量suíjī biànliàng
distribution (table)分布列fēnbùliè
expectation (mean)数学期望(均值)shùxué qīwàng (jūnzhí)
variance方差fāngchā
standard deviation标准差biāozhǔnchā
two-point distribution两点分布liǎngdiǎn fēnbù
binomial distribution二项分布èrxiàng fēnbù
normal distribution正态分布zhèngtài fēnbù
normal curve正态曲线zhèngtài qūxiàn
probability density概率密度gàilǜ mìdù
standard normal distribution标准正态分布biāozhǔn zhèngtài fēnbù
axis of symmetry对称轴duìchènzhóu
3σ principle3σ原则3σ yuánzé
standardisation标准化biāozhǔnhuà
The CSCA can be taken in English or Chinese, so learn both names.

Key points

  • Distribution table: pᵢ ≥ 0, Σpᵢ = 1; E(X) = Σxᵢpᵢ, D(X) = E(X²) − [E(X)]².
  • E(aX + b) = aE(X) + b, D(aX + b) = a²D(X); B(n, p): E = np, D = np(1 − p).
  • X ~ N(μ, σ²): the curve is symmetric about x = μ, E(X) = μ, D(X) = σ²; a larger σ means a lower, wider curve.
  • P(X < μ) = 0.5; P(X < μ − a) = P(X > μ + a); P(X < c) = P(X > d) ⇒ μ = (c + d)/2.
  • ≈ 68.27%, 95.45% and 99.73% of the values lie within μ ± σ, μ ± 2σ and μ ± 3σ; outside μ ± 3σ is practically impossible.
  • Z = (X − μ)/σ ~ N(0, 1), P(X < x) = Φ((x − μ)/σ), Φ(−z) = 1 − Φ(z).

Check yourself

12 questions. Every correct answer earns XP.

1 / 12
X takes the values 1, 2, 3 with probabilities 0.2, 0.5 and a. What is a?

Topic test: 20 questions · 25 min

Finished the lesson? Check yourself with a timed test on this topic.

Start the test