Skip to content
Educora
University25 min29 / 42

NumPy basics: arrays

Meet the NumPy array (ndarray): type and shape, creating arrays, reproducible random numbers, indexing, slicing, boolean masks and aggregation along axes.

Check yourself
In this lesson you will learn
  • Explain the ndarray attributes (ndim, shape, size, dtype) and calculate how much memory an array takes
  • Create arrays with arange, linspace, zeros and a seeded default_rng
  • Select data with indexes, slices and boolean masks, and tell a view from a copy
  • Compute statistics along rows and columns with the axis parameter

A weather station measures the temperature every hour for a whole year — that is 8,760 numbers. You could find their average, the hottest days and the monthly means with ordinary Python lists and loops, but slowly and with long code. NumPy (Numerical Python) does the same job in one line and many times faster. pandas, matplotlib, SciPy, scikit-learn — every data science tool in Python is built on NumPy arrays, so this module starts right here.

Definition
ndarray (N-dimensional array)

NumPy's core object: a fixed-size, multi-dimensional grid of elements that all have the same type. The elements sit next to each other in memory, and operations run over the whole array at once in compiled C code.

The ndarray: shape, size and type

A Python list keeps every element as a separate object and itself holds only pointers (addresses) to them, while an array stores the raw numbers in one block. That is why an array takes less memory and is processed faster. Every array has four key attributes: ndim — the number of dimensions (axes), shape — a tuple with the length along each axis, size — the total number of elements, and dtype — the element type.

Python
import numpy as np

temps = np.array([12.5, 14.0, 9.8, 11.2])
print(temps)
print(temps.ndim, temps.shape, temps.size, temps.dtype)
print(temps.itemsize, temps.nbytes)

m = np.array([[1, 2, 3], [4, 5, 6]])
print(m)
print(m.ndim, m.shape, m.size)
print(m * 10)
▸ Expected output
[12.5 14.   9.8 11.2]
1 (4,) 4 float64
8 32
[[1 2 3]
 [4 5 6]]
2 (2, 3) 6
[[10 20 30]
 [40 50 60]]
m * 10 works without a loop: the operation is applied to every element. With a list, [1, 2] * 10 would repeat the list 10 times instead.
size = d₁ · d₂ · … · dₙ nbytes = size · itemsize
where:
  • nnumber of axes (ndim)
  • dₖlength along axis k (an element of shape)
  • itemsizesize of one element in bytes: 8 for float64 and int64, 4 for float32 and int32, 1 for bool and int8

For m with shape (2, 3): size = 2 · 3 = 6. For temps: nbytes = 4 · 8 = 32 bytes.

dtypeBytesRange / precisionTypical use
bool1True / Falsemasks, flags
uint810 … 255image pixels
int324about ±2.1 · 10⁹counters, indexes
int648about ±9.2 · 10¹⁸large integers
float324~7 significant digitsneural networks, GPUs
float648~15–16 significant digitsthe default for real numbers

All elements of an array must have the same type. When types are mixed, NumPy converts everything to the “widest” type: integers and floats together give float64, numbers and strings together give strings. astype changes the type explicitly; when you convert floats to integers, the fractional part is not rounded but simply dropped.

Python
import numpy as np

print(np.array([1, 2, 3.5]))
print(np.array([1, 2, '3']))
print(np.array([1.9, -1.9]).astype(np.int32))

small = np.array([100, 120], dtype=np.int8)
print(small + small)
▸ Expected output
[1.  2.  3.5]
['1' '2' '3']
[ 1 -1]
[-56 -16]

Creating arrays

FunctionWhat it creates
np.array(data)an array from a list or nested lists
np.zeros(shape), np.ones(shape), np.full(shape, v)an array filled with 0, 1 or v
np.arange(start, stop, step)a sequence with a step, like range; stop is not included
np.linspace(start, stop, num)num evenly spaced points; both ends included
np.eye(n)the n × n identity matrix
np.random.default_rng(seed)a random number generator
Python
import numpy as np

print(np.zeros((2, 3)))
print(np.arange(0, 10, 2))
print(np.linspace(0, 1, 5))
print(np.full(3, 7.5))
print(np.eye(3))
▸ Expected output
[[0. 0. 0.]
 [0. 0. 0.]]
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]
[7.5 7.5 7.5]
[[1. 0. 0.]
 [0. 1. 0.]
 [0. 0. 1.]]
h = (stop − start) / (num − 1) n = ⌈(stop − start) / step⌉h = (stop − start) / (num − 1) n = ⌈(stop − start) / step⌉
where:
  • hthe spacing between linspace points
  • nnumber of elements produced by arange
  • ⌈ ⌉rounding up (ceiling)

np.linspace(0, 1, 5): h = 1 / 4 = 0.25. np.arange(0, 10, 2): n = ⌈10 / 2⌉ = 5 elements.

Be careful with arange and a fractional step: 0.1 cannot be represented exactly in binary, and the rounding error can change whether the last element ends up in the array. Below, 1.3 is included even though stop = 1.3:

Python
import numpy as np

print(np.arange(0, 1, 0.1).size)
print(np.arange(1, 1.3, 0.1))
print(np.linspace(1, 1.3, 4))
▸ Expected output
10
[1.  1.1 1.2 1.3]
[1.  1.1 1.2 1.3]

The modern way to get random numbers is to create a generator: rng = np.random.default_rng(42). Here 42 is the seed: with the same seed the generator produces the same sequence every time. That matters in science — a colleague must be able to run your code and get the same result. integers(a, b) excludes the upper bound, normal(loc, scale) draws from a normal distribution, and random() returns numbers in [0, 1). Instead of np.random.seed and np.random.rand, which you will see in older code, use a generator in new code.

Python
import numpy as np

rng = np.random.default_rng(42)
print(rng.integers(1, 7, size=10))
print(rng.normal(loc=170, scale=8, size=4).round(1))
print(rng.random((2, 3)).round(3))
▸ Expected output
[1 5 4 3 3 6 1 5 2 1]
[159.6 171.  167.5 169.9]
[[0.45  0.371 0.927]
 [0.644 0.823 0.443]]
Ten dice rolls, the random heights of four people with an average height of 170 cm, and a random 2 × 3 matrix. Run the code again and the result stays the same.

Indexing, slicing and boolean masks

A one-dimensional array is indexed like a list: a[0], a[-1], a[2:5], a[::2]. In a two-dimensional array the indexes are separated by a comma: m[row, col]. A colon : means “everything along this axis”: m[:, 1] is the second column and m[1] is the second row. Slices work on each axis separately, and a negative step reverses the order.

Python
import numpy as np

m = np.arange(1, 13).reshape(3, 4)
print(m)
print(m[1, 2], m[-1, -1])
print(m[:, 1])
print(m[0:2, 1:3])
print(m[::2, ::-1])
▸ Expected output
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
7 12
[ 2  6 10]
[[2 3]
 [6 7]]
[[ 4  3  2  1]
 [12 11 10  9]]
reshape(3, 4) arranges 12 elements into 3 rows and 4 columns (more in the next lesson). m[::2, ::-1] takes every second row with the columns in reverse order.
Python
import numpy as np

m = np.arange(1, 13).reshape(3, 4)
row = m[0]
row[0] = 100
print(m[0])

safe = m[1].copy()
safe[0] = -1
print(m[1], safe)
▸ Expected output
[100   2   3   4]
[5 6 7 8] [-1  6  7  8]

A comparison applied to an array returns True or False for every element — this is a boolean mask. If you use the mask as an index, only the elements where it is True are selected. Because True counts as 1 and False as 0, mask.sum() gives the number of matching elements and mask.mean() gives their share. np.where(cond, a, b) makes an “if…, then a, else b” choice for every element in one operation. Here are the midday temperatures of one week:

Python
import numpy as np

temps = np.array([21.4, 25.8, 31.2, 28.9, 33.5, 26.1, 19.7])
hot = temps > 28
print(hot)
print(temps[hot])
print(np.flatnonzero(hot))
print(hot.sum(), 'hot days,', f'{hot.mean():.0%} of the week')
print(temps[(temps > 20) & (temps < 27)])
print(np.where(temps > 30, 30.0, temps))
▸ Expected output
[False False  True  True  True False False]
[31.2 28.9 33.5]
[2 3 4]
3 hot days, 43% of the week
[21.4 25.8 26.1]
[21.4 25.8 30.  28.9 30.  26.1 19.7]
np.flatnonzero returns the indexes of the True elements: days 2, 3 and 4 counting from 0. The last line replaces values above 30 with 30.
Example 1: indexes and axes by hand

For a = np.arange(12).reshape(3, 4), find without running the code: a) a[1:, ::2]; b) a.sum(axis=0) and a.sum(axis=1) and their shapes; c) (a % 5 == 0).sum().

Show solution
a = [[0 1 2 3], [4 5 6 7], [8 9 10 11]].
a) 1: takes rows 1 and 2, ::2 takes columns 0 and 2: [[4 6], [8 10]].
b) axis=0 “squeezes” the rows and leaves the sum of each column: [0+4+8, 1+5+9, 2+6+10, 3+7+11] = [12 15 18 21], shape (4,).
axis=1 gives the sum of each row: [6 22 38], shape (3,).
c) The elements divisible by 5 are 0, 5 and 10, so the mask has 3 True values: the answer is 3.

Aggregation and the axis parameter

Functions such as sum, mean, min, max, std, argmax and cumsum “collapse” many numbers into one number or into a smaller array — this is called aggregation. Without axis, the whole array is used. axis=0 works along the first axis (down the rows) and gives one result per column; axis=1 works along the columns and gives one result per row. Below are the scores of 4 students (rows) in maths, physics and chemistry (columns):

Python
import numpy as np

scores = np.array([[85, 92, 78],
                   [67, 74, 81],
                   [93, 88, 95],
                   [72, 65, 70]])
print(scores.shape, scores.sum())
print(scores.mean(axis=0))
print(scores.mean(axis=1).round(2))
print(scores.max(axis=0), scores.argmax(axis=0))
print(scores.std(axis=0).round(2))
print(scores.std(axis=0, ddof=1).round(2))
print(scores.cumsum(axis=1)[0])
▸ Expected output
(4, 3) 960
[79.25 79.75 81.  ]
[85. 74. 92. 69.]
[93 92 95] [2 0 2]
[10.3  10.83  9.03]
[11.9  12.5  10.42]
[ 85 177 255]
The subject averages are 79.25, 79.75 and 81; the student averages are 85, 74, 92 and 69. argmax(axis=0) gives the index of the best student in each subject: student 2 in maths and chemistry, student 0 in physics.
x̄ = (1/n) · ∑ xᵢ σ = √( ∑ (xᵢ − x̄)² / (n − ddof) )x̄ = (1/n) · ∑ xᵢ σ = √( ∑ (xᵢ − x̄)² / (n − ddof) )
where:
  • x̄arithmetic mean (mean)
  • nnumber of values
  • σstandard deviation (std): how far the values typically are from the mean
  • ddof0 — for a whole population (NumPy's default), 1 — for a sample (unbiased variance)

np.std divides by n by default, while statistics textbooks and pandas' Series.std divide by n − 1. Pass ddof=1 to get the same result.

Example 2: standard deviation by hand

x = [2, 4, 4, 4, 5, 5, 7, 9]. What will np.std(x) and np.std(x, ddof=1) return?

Show solution
x̄ = 40 / 8 = 5.
Squared deviations from the mean: 9, 1, 1, 1, 0, 0, 4, 16; their sum is 32.
ddof = 0: σ² = 32 / 8 = 4, σ = 2.
ddof = 1: s² = 32 / 7 ≈ 4.571, s ≈ 2.138.
As n grows, the difference between the two answers shrinks.

Where will you use this? The columns of pandas tables are NumPy arrays; scikit-learn takes data as an array of shape (n × d); image processing, signals, financial time series and physics models are all operations on arrays. In the next lesson you will learn to replace loops completely with vectorized code and to do linear algebra.

Exercise

The sales array holds 4 weeks of sales of a small café (rows are weeks, columns are the days from Monday to Sunday). Print on separate lines: 1) the total sales of each week; 2) the index of the weekday with the highest sales over the 4 weeks (0 = Monday); 3) the number of days with sales above 100. Do not use a loop.

Exercise · Python
import numpy as np

sales = np.array([[80, 95, 110, 90, 130, 150, 70],
                  [85, 100, 105, 95, 140, 160, 75],
                  [90, 92, 120, 88, 125, 155, 80],
                  [78, 99, 115, 97, 135, 165, 72]])
# 1) total per week
# 2) index of the best weekday
# 3) number of days above 100
▸ Expected output
[725 760 750 761]
5
12
Exercise

Negative values in the sensor readings are errors. 1) Replace the negative values with 0.0 using np.where and print the new array; 2) print the mean of the readings in the original array that are greater than 10, rounded to 1 decimal place.

Exercise · Python
import numpy as np

readings = np.array([12.5, -3.0, 8.4, 15.1, -1.2, 20.3, 9.9, 11.0])
# 1) negatives -> 0.0
# 2) mean of readings > 10, rounded to 1 decimal
▸ Expected output
[12.5  0.   8.4 15.1  0.  20.3  9.9 11. ]
14.7

Key points

  • An ndarray stores same-type elements contiguously in memory; its key attributes are ndim, shape, size and dtype, and nbytes = size · itemsize.
  • arange excludes stop; for fractional steps linspace, which includes both ends, is safer.
  • Use a generator np.random.default_rng(seed) for reproducible random numbers.
  • A slice is a view: changing it changes the original; use .copy() for an independent copy.
  • Combine masks with &, |, ~ and parentheses; mask.sum() counts and mask.mean() gives the share.
  • axis=k removes axis k from the result's shape; np.std uses ddof=0 by default.

Check yourself

10 questions. Every correct answer earns XP.

1 / 10
What is the shape of np.zeros((3, 4)).sum(axis=0)?