- Explain the ndarray attributes (
ndim,shape,size,dtype) and calculate how much memory an array takes - Create arrays with
arange,linspace,zerosand a seededdefault_rng - Select data with indexes, slices and boolean masks, and tell a view from a copy
- Compute statistics along rows and columns with the
axisparameter
A weather station measures the temperature every hour for a whole year — that is 8,760 numbers. You could find their average, the hottest days and the monthly means with ordinary Python lists and loops, but slowly and with long code. NumPy (Numerical Python) does the same job in one line and many times faster. pandas, matplotlib, SciPy, scikit-learn — every data science tool in Python is built on NumPy arrays, so this module starts right here.
NumPy's core object: a fixed-size, multi-dimensional grid of elements that all have the same type. The elements sit next to each other in memory, and operations run over the whole array at once in compiled C code.
The ndarray: shape, size and type
A Python list keeps every element as a separate object and itself holds only pointers (addresses) to them, while an array stores the raw numbers in one block. That is why an array takes less memory and is processed faster. Every array has four key attributes: ndim — the number of dimensions (axes), shape — a tuple with the length along each axis, size — the total number of elements, and dtype — the element type.
import numpy as np
temps = np.array([12.5, 14.0, 9.8, 11.2])
print(temps)
print(temps.ndim, temps.shape, temps.size, temps.dtype)
print(temps.itemsize, temps.nbytes)
m = np.array([[1, 2, 3], [4, 5, 6]])
print(m)
print(m.ndim, m.shape, m.size)
print(m * 10)▸ Expected output
[12.5 14. 9.8 11.2] 1 (4,) 4 float64 8 32 [[1 2 3] [4 5 6]] 2 (2, 3) 6 [[10 20 30] [40 50 60]]
m * 10 works without a loop: the operation is applied to every element. With a list, [1, 2] * 10 would repeat the list 10 times instead.- nnumber of axes (
ndim) - dₖlength along axis k (an element of
shape) - itemsizesize of one element in bytes: 8 for
float64andint64, 4 forfloat32andint32, 1 forboolandint8
For m with shape (2, 3): size = 2 · 3 = 6. For temps: nbytes = 4 · 8 = 32 bytes.
| dtype | Bytes | Range / precision | Typical use |
|---|---|---|---|
bool | 1 | True / False | masks, flags |
uint8 | 1 | 0 … 255 | image pixels |
int32 | 4 | about ±2.1 · 10⁹ | counters, indexes |
int64 | 8 | about ±9.2 · 10¹⁸ | large integers |
float32 | 4 | ~7 significant digits | neural networks, GPUs |
float64 | 8 | ~15–16 significant digits | the default for real numbers |
All elements of an array must have the same type. When types are mixed, NumPy converts everything to the “widest” type: integers and floats together give float64, numbers and strings together give strings. astype changes the type explicitly; when you convert floats to integers, the fractional part is not rounded but simply dropped.
import numpy as np
print(np.array([1, 2, 3.5]))
print(np.array([1, 2, '3']))
print(np.array([1.9, -1.9]).astype(np.int32))
small = np.array([100, 120], dtype=np.int8)
print(small + small)▸ Expected output
[1. 2. 3.5] ['1' '2' '3'] [ 1 -1] [-56 -16]
Creating arrays
| Function | What it creates |
|---|---|
np.array(data) | an array from a list or nested lists |
np.zeros(shape), np.ones(shape), np.full(shape, v) | an array filled with 0, 1 or v |
np.arange(start, stop, step) | a sequence with a step, like range; stop is not included |
np.linspace(start, stop, num) | num evenly spaced points; both ends included |
np.eye(n) | the n × n identity matrix |
np.random.default_rng(seed) | a random number generator |
import numpy as np
print(np.zeros((2, 3)))
print(np.arange(0, 10, 2))
print(np.linspace(0, 1, 5))
print(np.full(3, 7.5))
print(np.eye(3))▸ Expected output
[[0. 0. 0.] [0. 0. 0.]] [0 2 4 6 8] [0. 0.25 0.5 0.75 1. ] [7.5 7.5 7.5] [[1. 0. 0.] [0. 1. 0.] [0. 0. 1.]]
- hthe spacing between
linspacepoints - nnumber of elements produced by
arange - ⌈ ⌉rounding up (ceiling)
np.linspace(0, 1, 5): h = 1 / 4 = 0.25. np.arange(0, 10, 2): n = ⌈10 / 2⌉ = 5 elements.
Be careful with arange and a fractional step: 0.1 cannot be represented exactly in binary, and the rounding error can change whether the last element ends up in the array. Below, 1.3 is included even though stop = 1.3:
import numpy as np
print(np.arange(0, 1, 0.1).size)
print(np.arange(1, 1.3, 0.1))
print(np.linspace(1, 1.3, 4))▸ Expected output
10 [1. 1.1 1.2 1.3] [1. 1.1 1.2 1.3]
The modern way to get random numbers is to create a generator: rng = np.random.default_rng(42). Here 42 is the seed: with the same seed the generator produces the same sequence every time. That matters in science — a colleague must be able to run your code and get the same result. integers(a, b) excludes the upper bound, normal(loc, scale) draws from a normal distribution, and random() returns numbers in [0, 1). Instead of np.random.seed and np.random.rand, which you will see in older code, use a generator in new code.
import numpy as np
rng = np.random.default_rng(42)
print(rng.integers(1, 7, size=10))
print(rng.normal(loc=170, scale=8, size=4).round(1))
print(rng.random((2, 3)).round(3))▸ Expected output
[1 5 4 3 3 6 1 5 2 1] [159.6 171. 167.5 169.9] [[0.45 0.371 0.927] [0.644 0.823 0.443]]
Indexing, slicing and boolean masks
A one-dimensional array is indexed like a list: a[0], a[-1], a[2:5], a[::2]. In a two-dimensional array the indexes are separated by a comma: m[row, col]. A colon : means “everything along this axis”: m[:, 1] is the second column and m[1] is the second row. Slices work on each axis separately, and a negative step reverses the order.
import numpy as np
m = np.arange(1, 13).reshape(3, 4)
print(m)
print(m[1, 2], m[-1, -1])
print(m[:, 1])
print(m[0:2, 1:3])
print(m[::2, ::-1])▸ Expected output
[[ 1 2 3 4] [ 5 6 7 8] [ 9 10 11 12]] 7 12 [ 2 6 10] [[2 3] [6 7]] [[ 4 3 2 1] [12 11 10 9]]
reshape(3, 4) arranges 12 elements into 3 rows and 4 columns (more in the next lesson). m[::2, ::-1] takes every second row with the columns in reverse order.import numpy as np
m = np.arange(1, 13).reshape(3, 4)
row = m[0]
row[0] = 100
print(m[0])
safe = m[1].copy()
safe[0] = -1
print(m[1], safe)▸ Expected output
[100 2 3 4] [5 6 7 8] [-1 6 7 8]
A comparison applied to an array returns True or False for every element — this is a boolean mask. If you use the mask as an index, only the elements where it is True are selected. Because True counts as 1 and False as 0, mask.sum() gives the number of matching elements and mask.mean() gives their share. np.where(cond, a, b) makes an “if…, then a, else b” choice for every element in one operation. Here are the midday temperatures of one week:
import numpy as np
temps = np.array([21.4, 25.8, 31.2, 28.9, 33.5, 26.1, 19.7])
hot = temps > 28
print(hot)
print(temps[hot])
print(np.flatnonzero(hot))
print(hot.sum(), 'hot days,', f'{hot.mean():.0%} of the week')
print(temps[(temps > 20) & (temps < 27)])
print(np.where(temps > 30, 30.0, temps))▸ Expected output
[False False True True True False False] [31.2 28.9 33.5] [2 3 4] 3 hot days, 43% of the week [21.4 25.8 26.1] [21.4 25.8 30. 28.9 30. 26.1 19.7]
np.flatnonzero returns the indexes of the True elements: days 2, 3 and 4 counting from 0. The last line replaces values above 30 with 30.For a = np.arange(12).reshape(3, 4), find without running the code: a) a[1:, ::2]; b) a.sum(axis=0) and a.sum(axis=1) and their shapes; c) (a % 5 == 0).sum().
Show solutionHide solution
a)
1: takes rows 1 and 2, ::2 takes columns 0 and 2: [[4 6], [8 10]].b) axis=0 “squeezes” the rows and leaves the sum of each column: [0+4+8, 1+5+9, 2+6+10, 3+7+11] = [12 15 18 21], shape (4,).
axis=1 gives the sum of each row: [6 22 38], shape (3,).
c) The elements divisible by 5 are 0, 5 and 10, so the mask has 3
True values: the answer is 3.Aggregation and the axis parameter
Functions such as sum, mean, min, max, std, argmax and cumsum “collapse” many numbers into one number or into a smaller array — this is called aggregation. Without axis, the whole array is used. axis=0 works along the first axis (down the rows) and gives one result per column; axis=1 works along the columns and gives one result per row. Below are the scores of 4 students (rows) in maths, physics and chemistry (columns):
import numpy as np
scores = np.array([[85, 92, 78],
[67, 74, 81],
[93, 88, 95],
[72, 65, 70]])
print(scores.shape, scores.sum())
print(scores.mean(axis=0))
print(scores.mean(axis=1).round(2))
print(scores.max(axis=0), scores.argmax(axis=0))
print(scores.std(axis=0).round(2))
print(scores.std(axis=0, ddof=1).round(2))
print(scores.cumsum(axis=1)[0])▸ Expected output
(4, 3) 960 [79.25 79.75 81. ] [85. 74. 92. 69.] [93 92 95] [2 0 2] [10.3 10.83 9.03] [11.9 12.5 10.42] [ 85 177 255]
argmax(axis=0) gives the index of the best student in each subject: student 2 in maths and chemistry, student 0 in physics.- x̄arithmetic mean (
mean) - nnumber of values
- σstandard deviation (
std): how far the values typically are from the mean - ddof0 — for a whole population (NumPy's default), 1 — for a sample (unbiased variance)
np.std divides by n by default, while statistics textbooks and pandas' Series.std divide by n − 1. Pass ddof=1 to get the same result.
x = [2, 4, 4, 4, 5, 5, 7, 9]. What will np.std(x) and np.std(x, ddof=1) return?
Show solutionHide solution
Squared deviations from the mean: 9, 1, 1, 1, 0, 0, 4, 16; their sum is 32.
ddof = 0: σ² = 32 / 8 = 4, σ = 2.
ddof = 1: s² = 32 / 7 ≈ 4.571, s ≈ 2.138.
As n grows, the difference between the two answers shrinks.
Where will you use this? The columns of pandas tables are NumPy arrays; scikit-learn takes data as an array of shape (n × d); image processing, signals, financial time series and physics models are all operations on arrays. In the next lesson you will learn to replace loops completely with vectorized code and to do linear algebra.
The sales array holds 4 weeks of sales of a small café (rows are weeks, columns are the days from Monday to Sunday). Print on separate lines: 1) the total sales of each week; 2) the index of the weekday with the highest sales over the 4 weeks (0 = Monday); 3) the number of days with sales above 100. Do not use a loop.
import numpy as np
sales = np.array([[80, 95, 110, 90, 130, 150, 70],
[85, 100, 105, 95, 140, 160, 75],
[90, 92, 120, 88, 125, 155, 80],
[78, 99, 115, 97, 135, 165, 72]])
# 1) total per week
# 2) index of the best weekday
# 3) number of days above 100▸ Expected output
[725 760 750 761] 5 12
Negative values in the sensor readings are errors. 1) Replace the negative values with 0.0 using np.where and print the new array; 2) print the mean of the readings in the original array that are greater than 10, rounded to 1 decimal place.
import numpy as np
readings = np.array([12.5, -3.0, 8.4, 15.1, -1.2, 20.3, 9.9, 11.0])
# 1) negatives -> 0.0
# 2) mean of readings > 10, rounded to 1 decimal▸ Expected output
[12.5 0. 8.4 15.1 0. 20.3 9.9 11. ] 14.7
Key points
- An ndarray stores same-type elements contiguously in memory; its key attributes are
ndim,shape,sizeanddtype, and nbytes = size · itemsize. arangeexcludesstop; for fractional stepslinspace, which includes both ends, is safer.- Use a generator
np.random.default_rng(seed)for reproducible random numbers. - A slice is a view: changing it changes the original; use
.copy()for an independent copy. - Combine masks with
&,|,~and parentheses;mask.sum()counts andmask.mean()gives the share. axis=kremoves axis k from the result's shape;np.stdusesddof=0by default.
Check yourself
10 questions. Every correct answer earns XP.
np.zeros((3, 4)).sum(axis=0)?