Definition
Adding two columns of a thousand numbers with a Python loop works fine; over a million numbers, the same loop turns slow. NumPy adds to Python what was missing for that kind of work: a numeric array, homogeneous and of fixed size, whose values all share the same type and sit in a single block of memory. One operation then applies to millions of items at once, without writing a single loop.
import numpy as np
celsius = np.array([18.2, 21.5, 19.8, 23.1])
fahrenheit = celsius * 9 / 5 + 32
print(fahrenheit)
# [64.76 70.7 67.64 73.58]The conversion line holds for four values as well as for four million: written once, it applies to the whole array. That style of writing has a name, vectorisation, the reason the library exists.
NumPy is not part of the standard library. It is installed with pip, preferably inside a virtual environment belonging to the project.
pip install numpyThe import under the nickname np, with the as keyword, is a convention so universal that every published example follows it. Writing it any other way breaks nothing, but forces anyone reading the code to relearn a shortcut everyone else already knows.
The problem it solves
A Python list accepts anything: integers, strings, other lists too. Each number in it is a complete object sitting somewhere in memory, and the list only keeps its address. Adding two lists of a million values therefore forces the interpreter to follow two million addresses, check the type of every object it meets, then build a brand new object for each result: three steps, repeated a million times, for a single addition.
A NumPy array removes those three costs at once. The type is decided once and for all at creation, the values sit next to each other in a contiguous block, and the loop doing the real work runs in C, outside the Python interpreter. On a million items, the gap with a loop written in Python is measured in tens of times. That is not a comfort detail: it is the difference between a result that appears right away and a calculation started before going off to get a coffee.
Against the Python list
The honest comparison names no winner: the two structures answer different needs, and the array loses what made the list convenient. The detail, point by point:
| On this point | Python list | NumPy array |
|---|---|---|
| Contents | Any objects at all, mixed | One single type, fixed at creation |
| The a + b operation | Joins the two sequences end to end | Adds them term by term |
| Adding an item | Immediate, the list grows | Copies the whole array |
| Computing on a million values | An interpreted loop | A compiled operation |
| Two dimensions | Lists of lists, handled by hand | A native shape, rows and columns |
The second row is the one that surprises people most: on lists the plus sign joins end to end, on arrays it adds term by term. The same symbol, two unrelated operations. It is the first thing to check when a result comes out twice too long, and the length measured with len says straight away which of the two behaviours applied.
The first-day trap
Slicing an array does not copy it: where slicing a list hands back a brand new list, slicing an array hands back a view, a window opened onto the same data. Changing the view therefore changes the original.
import numpy as np
readings = np.array([10, 20, 30, 40])
extract = readings[1:3]
extract[0] = 999 # changes readings too, not a copy
print(readings)
# [ 10 999 30 40]A function that receives a slice of an array and changes it "to work on it" is actually changing the original data. The bug shows up further along, wherever the original gets read again, not here.
The behaviour is deliberate: it saves duplicating the memory of a huge array just to work on one portion of it. The copy method gives a real copy, the reflex to build before changing an extract.
The second stumble comes from the fixed type: an array of integers stays an array of integers, and storing 3.7 in it raises no ValueError, the value is silently truncated to 3. Building the array from decimal numbers saves a long hunt for results that are wrong yet perfectly plausible.
When it is of no use
On a few hundred values, converting to an array costs more than the calculation it saves: an ordinary loop is enough and reads better. The library brings nothing either on heterogeneous data, such as a file whose every line mixes a name, a date and an amount. For that kind of data, pandas is the right door: the columns of a pandas data frame are NumPy arrays in disguise.
Computing is not showing, either. NumPy draws no curve: that is the job of matplotlib, which takes NumPy arrays directly and forms with it the most common pairing in Python's scientific computing. On text, nested dictionaries or a response coming from a web interface, the standard library stays more direct. The classic mistake is installing NumPy by reflex, only to end up storing strings inside an array built for numbers.
Frequently asked questions
Should NumPy be learnt before pandas?
Not necessarily, but a few hours pay for themselves quickly. Array, shape, fixed type, selection by condition: these notions come back almost unchanged in pandas, on named columns.
Why is walking an array item by item slower than walking a list?
Because every access rebuilds a Python object around a raw value, work the list never has to do since it already stored objects. The answer is therefore not to move the loop but to delete it, by writing the operation on the whole array.
Does NumPy replace the math module of the standard library?
No, the two aim at different scales. The math module works on one number at a time and stays the fastest in that precise case. NumPy functions often carry the same names, but apply to a whole array: needlessly heavy on a single isolated value.