LANG·I Language Chapter 6 of 65
Lists
A seismologist’s night shift: the USGS catalog of 19,073 earthquakes and the strongest quake in fourteen years. We learn to keep thousands of values under one name, slice them and sort them, and to stay out of the trap of two names for one list.
Language
Builds on: 05 · Words of your own
What you will take away
- keep thousands of values under one name: indices, slices, adding and removing items
- find the maximum, the mean, the median and the top ten in a data set without a spreadsheet
- tell when two names point to one list, and make independent copies
The last chapter ended with a question. A function can check whether one number is prime, or run one card number through Luhn’s test. What if there are a thousand numbers? Tonight there will be nineteen thousand, and behind each one is an earthquake that seismographs recorded.
Imagine the night shift at a seismic monitoring center. There are dozens of such centers around the world. Around the clock they take in seismograph records, work out where the ground shook and how hard, and send out bulletins. The seismologist on duty lives on Coordinated Universal Time, UTC: an earthquake doesn’t care about the time zone of the nearest town, and bulletins from different countries have to merge into a single feed. Our shift is the night of July 29–30, 2025, UTC. What happens at 23:24 you don’t know yet.
On the desk is the catalog of the United States Geological Survey (USGS): every earthquake of magnitude 5 and above from 2015 to 2025, 19,073 rows. It sits in the course’s sandbox, and every cell on this page can read it. Over the shift we’ll answer the questions people ask the duty seismologist—what happened, where, how strong—and on the way pick up the data structure that hardly any Python program can do without.
22:00. Taking over the shift
The outgoing seismologist has left the log for July 29: nine quakes. Three weak ones near Macquarie Island, between Australia and Antarctica, one in the Red Sea, one near the Andaman Islands, two deep under Fiji, and two in Guatemala half an hour ago. Their magnitudes in order of time: 5.4, 5.1, 5.1, 5.2, 5.0, 5.8, 6.6, 5.6, 5.7.
With what you know so far, you would need nine variables to hold them: m1 = 5.4, m2 = 5.1 and so on up to m9. The mean would be a sum of nine names written out by hand. Tomorrow there will be a different number of quakes, and the program will have to be rewritten. And a loop from Chapter 4 can’t walk through such variables: to Python, m1 and m2 are two separate names with nothing linking them.
We need one thing that holds all the values in order. That thing is a list: values separated by commas, in square brackets.
The values in a list are its items, and an item’s number is its index. Indices in Python start at zero: the first quake is mags[0], the seventh mags[6]. It helps to read an index as the distance from the start: the first item is zero steps away, the seventh six. In Chapter 14 you’ll see that this is how the machine counts: an item’s address in memory is the start of the list plus the index times the size of a cell.
Starting from zero leads to a rule that every beginner trips over: in a list of nine items, the last index is eight. Asking for mags[9] ends in an IndexError with the explanation list index out of range. Negative indices count from the end: mags[-1] is the last item, mags[-2] the one before it. So every item has two addresses.
Slices
Someone asks the duty seismologist: “What were the last three reports?” You could write out mags[6], mags[7], mags[8], but for pieces of a list there is a shorter notation, the slice. mags[2:7] is the items from index 2 up to index 7, not including index 7. It is the same half-open interval as in range(2, 7): the length of the piece is the difference between the bounds, and two neighboring slices, mags[:5] and mags[5:], join with no gap and no overlap. A missing start means “from the beginning,” a missing stop “to the end”: the last three are mags[-3:]. A third number sets the step: mags[::2] is every other item.
A slice doesn’t fail because of its bounds: if they run past the list, Python quietly trims them to its edges, and if the start is to the right of the stop, it returns an empty list, []. One more property will come in handy toward morning: a slice always builds a new list and leaves the original alone.
23:05. The log grows
At 23:05 a report comes in from Guatemala: another quake, M 5.2. A list can grow: mags.append(5.2) adds a value at the end. The dot notation (a value, a dot, the name of a function) means “ask the list itself to do this.” Functions attached to the type of a value in this way are called methods. Chapter 12 covers them in detail; for now it is enough to know that a list has about a dozen.
At 23:24:52 comes the event that night shifts exist for. The first USGS estimate is magnitude 8.0, off the east coast of Kamchatka. Once records arrive from stations all over the world, USGS will raise it to 8.8. The seismologist corrects the entry in place: assigning to an index replaces an item.
The last line asks whether a value is in the list: the operator in answers True or False. To answer, Python goes through the list from the start, item by item, until it finds the value. With nine items that takes no time at all; with a million it becomes noticeable. Chapter 8 brings a structure where in doesn’t go through the items and is as fast on a million as on ten, and Chapter 16 explains how it manages that.
Sometimes a report arrives late, or twice, from two agencies about one event. For that a list has insert(i, x), which puts x in so that it gets index i, and pop(i), which takes out the item at index i and returns it; pop() with no argument takes out the last one. Watch the neighbors: on an insert, every item to the right shifts along, and its index goes up by one.
i), type x and press an action. The items that moved are highlighted: compare insert(0, x) with append(x).Inserting at the front makes the whole list move; appending at the end moves nobody. With ten items you can’t see the difference. In Chapter 14 we’ll measure it on a million items and find out why append is almost always fast.
And remember: a list changes in place. After append it is the same object, only longer. Numbers and strings don’t work like that: they can’t be changed, only replaced by others. That is how, in Chapter 2, Grace unexpectedly ended up in two crews at once. Objects that can change like this are called mutable, and toward morning this property will get us into trouble.
23:24. The quake
In Kamchatka it is 11:24 in the morning. The magnitude 8.8 quake is the strongest on Earth since 2011, when a magnitude 9.1 earthquake off Japan set off a tsunami and the accident at the Fukushima Daiichi nuclear plant. By USGS estimates it is tied for sixth place among all earthquakes ever recorded by instruments. The tsunami spread across the whole Pacific but turned out weaker than feared, and the damage was small for a quake of this size: Kamchatka builds with earthquakes in mind. In 1952 there was an earthquake of about magnitude 9 in almost the same place.
In the first hour after the main shock, the catalog gained more than thirty records above magnitude 5 near Kamchatka. These are aftershocks, the repeat quakes that a ruptured plate goes on producing for months. In the first day after the main shock there were more than a hundred and fifty of them, and July 30, 2025, UTC, became the busiest day in the whole catalog: 143 records. Nobody keeps series like that by hand, but while the log holds eleven values, we can answer the first questions.
Four built-in functions work on a whole list at once: len gives its length, max and min the largest and smallest items, sum the total. The loop for m in mags walks through the items one by one: at each step the name m is tied to the next value. Inside it is the accumulator you know from Chapter 4. Press “Steps” and watch total grow. It grows unevenly: ten of the eleven quakes add almost nothing.
Ninety-nine point nine percent of the day’s energy came from one quake. The history of the scale explains why the formula is 10 ** (1.5 * m) and not something else.
On a logarithmic scale, adding one means multiplying. A magnitude 6 swings a seismograph’s pen ten times as far as a 5, and a 7 a hundred times as far. For energy the gap is wider still: each unit of magnitude means about 32 times more energy, because $10^{1.5} \approx 31.6$. Gutenberg and Richter tied the two together with the formula $\lg E = 1.5\,M + 4.8$, where $E$ is the energy in joules and lg is the base-10 logarithm. That is why energy in our program is proportional to $10^{1.5 M}$. The numbers 8.8 and 5.8 look one and a half times apart; in energy the 8.8 is $10^{4.5}$ times stronger, about 30,000 times.
How many magnitude 6 quakes does it take to release as much energy as one magnitude 8 quake?
A difference of two units of magnitude is $10^{1.5 \cdot 2} = 10^3$: a thousand quakes of magnitude 6. On a seismograph’s trace, though, the difference is only a hundredfold. Logarithmic scales fool the intuition: “8 against 6” sounds modest.
00:30. The catalog
Kamchatka keeps producing a quake every two or three minutes, and the seismologist gives up on the log: from here on we work with the catalog. It lives in the file /data/quakes.csv. The CSV format (comma-separated values) is the simplest way to store a table as text: one line per record, fields separated by commas, and the column names in the first line.
Time, latitude, longitude, depth in kilometers, magnitude, magnitude type and a description of the place. The string method split can cut a line at the commas. But a description of a place sometimes has a comma of its own, and then it falls apart into two fields. That’s why CSV puts such fields in quotes, but a cut at the commas knows nothing about quotes.
The first call cut at every comma, and the place fell apart. In the second we asked for no more than six cuts: the first six fields came off, and everything else, the comma inside the quotes included, stayed together as the seventh piece. We’re lucky the place is the last column. In general, quotes can turn up in any field, and Python has the csv module for that; strings get proper attention in the next chapter. Here is a function that reads the whole catalog.
Line by line: open opens the file, readline swallows the header line, and the loop for line in file goes through the rest. The method strip cuts off the newline at the end, split(",", 6) cuts the line into seven pieces, and float turns text like "8.8" into a number. Each line becomes one record, and the record is written in parentheses.
The question marks in the name of the Iranian town aren’t a bug in our program. When the catalog was exported, letters with diacritics, like the ī and ā in Fīrūzābād, turned into “?”. Live data is almost always a little dirty, and the first thing an analyst does is eyeball it. Why letters get lost in transit will become clear in Chapter 28.
Tuples
Values in parentheses, separated by commas, make a tuple. A tuple is like a list: it has a length, indices and slices, and you can loop over it. There is one difference: a tuple can’t be changed. You can’t add an item to it, remove one or replace one.
The first line is shorter now: we put the same function, load_quakes, into a course module so as not to rewrite it in every cell. Next comes unpacking: on the right, a tuple of seven values; on the left, seven names, and each name is tied to its value. You saw unpacking in Chapter 2, in the swap a, b = b, a: there too the right-hand side was a tuple, only without parentheses.
Tuples and lists differ in purpose. A list is a collection of similar things that grows and shrinks, like the night’s quakes. A tuple is one record with a fixed set of fields: an earthquake always has seven fields, and the fifth is always the magnitude. There is nothing in it to change, and immutability protects it from accidental damage; besides, as Chapter 8 will show, a tuple can serve as a dictionary key. The drawback is plain at once: q[4] gives no hint that it is a magnitude. Fields with names arrive in Chapter 8; for now, keep in mind that 0 is the time, 1 the latitude, 2 the longitude, 3 the depth, 4 the magnitude and 6 the place.
Slices work on the catalog the same way they did on the log.
A map made of dots
Every quake has a latitude and a longitude, and that is enough to draw a map. We collect them into two lists and hand them to the plotting library matplotlib, which you saw in Chapter 0: scatter puts a dot at each pair of coordinates.
We didn’t draw a single continent, only the dots of the epicenters. And yet the map of the world emerges by itself: the west coasts of both Americas, the arcs of the Aleutian and Kuril Islands, Japan, Indonesia, a winding seam down the middle of the Atlantic. Earthquakes follow the boundaries of tectonic plates, where plates pull apart, collide or grind past each other. The ring around the Pacific is the Ring of Fire, where most strong quakes happen.
Below is the same map, alive. Its data loads straight into your browser, so the sliders respond instantly. Color shows the depth of the source, size the magnitude.
Leave only the deep quakes, those below 300 kilometers. There are fewer of them than you might expect, and they lie in bands set back from the ocean trenches: under the Sea of Japan and the Sea of Okhotsk, under Fiji, under the middle of South America. There an oceanic plate dives under its neighbor and keeps breaking as it sinks, down to 700 kilometers. The deepest quake in the catalog struck 670.8 kilometers under Fiji in September 2018. Move the lower magnitude bound up to 7, and about a hundred and fifty dots remain, almost all of them on the Ring of Fire.
The map also keeps an eye on this chapter’s cells. Every quake a cell below prints, as a whole tuple or at least with its time, lights up on the map, and a link to it appears under the cell.
01:30. The strongest
A newsroom calls: is it true that this is the strongest earthquake in ten years? The answer is in the catalog. We find the strongest quake the way we found a maximum in Chapter 4: keep the “best so far” and compare each next quake with it.
The loop found the Kamchatka quake, 8.8, the strongest in all eleven years. But max(quakes) returned something odd: a magnitude 6 off Japan on December 31, 2025. The built-in max doesn’t know that we care about magnitude. It compares tuples item by item, the way a dictionary orders words: first the first items, and if they are equal, the second, and so on. The first item is the time, time strings are compared alphabetically, and since our format is “year-month-day,” the latest date comes last in the alphabet. There is no error here. The answer is correct, only to a different question. How to tell max which field to compare you’ll learn in Chapter 10.
The list pairs holds three tuples: (5.2, "Guatemala"), (8.8, "Kamchatka") and (8.8, "Alaska"). What does max(pairs) return?
The first items, 8.8 and 8.8, are equal, so the second ones decide: the string “Kamchatka” is greater than “Alaska,” because K comes later in the alphabet. This dictionary-style comparison is called lexicographic.
The strongest in eleven years, then, but not the strongest in history.
How does Valdivia compare with our whole catalog? Add up the energy of all 19,073 quakes over eleven years and divide it by the energy of the single earthquake of 1960.
All the planet’s earthquakes above magnitude 5 over eleven years, taken together, released less than a third of what Valdivia released in ten minutes. And Kamchatka 2025 by itself accounts for more than a quarter of the catalog’s energy. That is how a logarithmic scale works: thousands of middling quakes weigh almost nothing next to one giant. The summary above the map shows it too: move the magnitude slider to 6, and there are thirteen times fewer quakes, while the energy drops only from 32 to 31 percent of Valdivia.
02:30. The middle and the top ten
In the morning the boss will have two questions: what does a “typical” quake look like, and which are the ten strongest? The arithmetic mean is a poor answer to the first, because rare giants drag it upward. The median is more reliable: the value that lands in the middle when you line everything up in order. The function sorted does the lining up.
The mean is 5.33 and the median 5.2: more than half the quakes in the catalog sit right at its threshold. The last line shows that the original list hasn’t changed: sorted built a new one. Lists also have a second way to sort, the method mags.sort(). It sorts the list itself, in place, and returns nothing. That is why beginners often fall for ordered = mags.sort(): what lands in ordered is None, “nothing.” It’s easy to remember: sorted(x) says “give me a sorted copy,” x.sort() says “sort yourself.” Either one accepts reverse=True to sort in descending order.
Now for the ten strongest. Sorting the catalog’s tuples as they are won’t do: sorted, like max, will compare them by time. But we already know how tuples are compared: by the first item first. So it is enough to repack each record into a new tuple with the magnitude in front.
Unpacking works right in the loop header: for time, lat, … in lays each tuple out into names. After sorting in descending order, all that’s left is to take the slice [:10]. A link has appeared under the cell: all ten quakes are already marked on the epicenter map. Five of the ten are off the coasts of the Americas, from Alaska to Chile; the rest are in the western Pacific, except for one in the South Atlantic.
One more question seismologists put to a catalog before any other: how often do quakes of different strengths happen? We set up a list of four counters, for magnitudes 5, 6, 7 and 8, and use the magnitude as an index: int(5.8) drops the fractional part and gives 5, and subtracting five makes it index 0.
Almost eighteen thousand fives, over a thousand sixes, a hundred and fifty sevens and eight eights: with every unit of magnitude there are roughly ten times fewer quakes. Gutenberg and Richter described this pattern in 1944, and it bears their names. The ratios wander from 9 to 18 because there are only eight eights, and with numbers that small, chance shows. Put this together with the energy count from the last section and you get the conclusion engineering seismology rests on: the frequent weak quakes weigh almost nothing, and it is the rare strong ones you have to reckon with.
04:00. The relief
Toward morning your relief arrives to prepare the summary, which needs the night’s quakes in order of strength. Your relief takes the shift log and writes three lines.
The report is sorted, and the log, which was supposed to stay in order of time, is ruined. You met this in Chapter 2: the assignment report = night copies nothing. It hangs a second label on the same list, and sort, called through either name, changes the only list there is. The operator is confirms it: this is one object. Press “Steps”: the arrows from both names lead to the same place.
A copy is a new list with the same items. You get one from the full slice night[:], from the function list(night) or from the method night.copy(); all three do the same thing.
“Is this the same object?” and “Do they have the same contents?” are different questions. is asks the first, == the second. Right after copying, report == night would have been True, and report is night False.
Functions are the sneakiest case. A function’s parameter is a label too, and at the call it is hung on the same object as the argument. A function that sorts the list it was given “for its own use” ruins that list for whoever called it.
The cure is not to change what you were given: sorted(mags, reverse=True)[:3]. A good function returns a result and leaves its arguments alone; in Chapter 5 we called such functions pure. Below are five scenarios on one diagram. Before you see the output, predict it.
print, the diagram asks what will be printed.An assignment with = moves a label and never copies anything. Methods like append, sort and pop change the list itself, and the change shows through every label hanging on it. If you need an independent copy, make one explicitly: a[:].
05:00. Where the ground shakes
The last question of the shift: where on Earth does the ground shake most often? Divide the planet into cells of 10 degrees: 18 bands of latitude, from north to south, and 36 of longitude, from west to east. Each cell needs a counter, so the counters form an 18-by-36 table. A table is easy to keep as a list of rows in which each row is a list too. A list’s item can be anything, another list included: grid[row] is a row of the table, grid[row][col] a cell in it.
Latitude 90 (the North Pole) gives row 0, latitude −90 row 17; longitude −180 gives column 0. The expression [0] * 36 builds a list of 36 zeros: multiplying a list by a number repeats its contents. The picture comes from the function show_grid in the course module, with one character per cell. The redder the cell, the more quakes in it. The Ring of Fire shows here too, if coarsely: a 10-degree cell at the equator is a square more than a thousand kilometers on a side. And more than half the cells, 339 of 648, didn’t get a single quake above magnitude 5 in eleven years.
All that’s left is to find the hottest cell. It is the maximum over the table, and we search for it with two nested loops.
Over eleven years, 1,054 quakes struck the square between 20 and 30 degrees south, by the date line. Here the Tonga and Kermadec trenches meet, and the Pacific Plate dives under the Australian Plate. At the northern end of the Tonga Trench, GPS measurements published in 1995 showed the plates closing at 24 centimeters a year, the fastest rate measured anywhere on the planet. On the epicenter map above, this cell is a dense scatter of dots north of New Zealand.
The multiplication trap
Building the table with a loop is tedious. Since [0] * 36 repeats zeros, why not repeat a row as well: [[0] * 36] * 18? One line instead of three, and this is where everything we have learned about labels catches up with us.
We added one to a single cell, and it showed up in all three rows. Multiplying a list repeats its contents, and what the outer list contains is a label on a row. The result is three labels on one and the same row. The loop with append computed [0] * 36 afresh on every pass and created a new row each time; the multiplication created the row once. The zeros inside a row don’t have this problem: a number can’t be changed, so += 1 hangs a new object on the cell, and the neighboring cells don’t see it.
A world map built with the trap turns into vertical stripes: every cell holds the total of its whole column, and hundreds of earthquakes “happen” at the North Pole. This mistake produces no traceback, only a wrong answer that Python keeps quiet about, as in the last entry of Chapter 1.
Copying a table is a trap too
The slice grid[:] copies only the outer list: the new list holds the same labels on the same rows. A copy like that is called shallow. Change a cell in the copy and it changes in the original as well, because the row is shared. A full, deep copy, with all the nested lists, comes from the function copy.deepcopy in the copy module, or from a loop that copies each row: new.append(row[:]).
06:00. Handing over the shift
The shift is over, and the seismologist hands in the report. Five tasks: the same questions that came in during the night, now written as functions so that they can be put to any catalog. The tests also check the edges: an empty list, repeats, huge inputs, and whether the function damages a list that isn’t its own.
Write a function strongest(quakes). It takes a list of quakes, tuples of the same shape as in the catalog (time, latitude, longitude, depth, magnitude, type, place), and returns the tuple of the strongest one. If several are equally strong, return the one that comes first in the list. For an empty list, return None.
The built-in max(quakes) won’t do: it compares the tuples by their first field, the time.
What’s wrong with starting from best_mag = 0? The magnitudes of very weak quakes can be negative: the Richter scale is logarithmic, and zero on it isn’t “nothing” but a definite, tiny trace amplitude. Start from the first item of the list.
The strict > keeps the first of several equals: a later quake with the same magnitude won’t push it out. Starting from the first item is safer than starting from some invented “very small” number: the first item is certainly one of the candidates.
Write a function mean_mag_depth(quakes) that returns a tuple of two numbers: the mean magnitude and the mean depth of the quakes in the list. No rounding. For an empty list, return None.
You need two accumulators: the sum of the magnitudes (index 4) and the sum of the depths (index 3). Divide each by len(quakes).
Returning a tuple means writing two values separated by a comma: return a, b. The parentheses are optional.
if not quakes is a short test for emptiness: an empty list, like an empty string and zero, counts as false. For the whole catalog the mean magnitude comes out at about 5.33 and the mean depth at about 52 km: most earthquakes are shallow.
Write a function top_k(numbers, k) that returns a new list of the k largest numbers in descending order. If there are fewer than k numbers, return them all. One condition: the list numbers you were given must stay as it was.
The starter gives the right answer. So what’s wrong with it? Run it on a list and print that list after the call.
sorted builds a new list; .sort() changes the one it was called on.
The slice handles the edges by itself: if k is larger than the length, [:k] returns everything, and [:0] returns an empty list. Sorting the whole list to get the top ten is wasteful; Chapter 18 brings a faster way, the heap.
Shifts rotate in a circle. Write a function rotate(items, k) that returns a new list rotated to the right by k positions: rotate([1, 2, 3, 4, 5], 2) → [4, 5, 1, 2, 3]. A negative k rotates to the left, and k may be larger than the length of the list. The original list must not change.
Rotating to the right by k gives the last k items followed by all the rest. Both pieces are slices.
Rotating by the length of the list brings it back to where it started, so k can be replaced by the remainder k % len(items); in Python it is non-negative even for a negative k. And what if the list is empty?
Check the edge case k = 0: items[-0:] is items[0:], the whole list, and items[:-0] is items[:0], which is empty. Together they make the whole list, as they should. Without the check for an empty list, k % 0 would crash with a ZeroDivisionError. The plus glues two lists into a new one and leaves both of them unchanged.
Write a function second_largest(numbers) that returns the second largest distinct value: for [5.2, 8.8, 6.3, 8.8] that is 6.3, not 8.8. If there are fewer than two distinct values, return None. The list may be very long, and you must not change it.
The starter stumbles on repeats and on short lists. Try it on [8.8, 8.8] and on [7.0].
You can go through the list once, keeping two numbers: the best and the second best. A new number larger than the best pushes the best down to second place. One smaller than the best but larger than the second becomes the second. One equal to the best is skipped.
One pass over the list is faster than sorting, and nothing gets copied. Starting with None instead of a “very small number” means “nothing seen yet.” The line first, second = x, first is the swap you know from Chapter 2: the right-hand side is worked out in full before any label moves.
What next
The shift is handed over, but one question is still open: which country shook most often? The answer is in the catalog, in the seventh field: “125 km SE of Petropavlovsk-Kamchatsky, Russia”. Only it is written as text, and the country is the piece after the last comma. To cut it out, you need to know how to work with strings. In the next chapter we take strings on in earnest and build a conversation partner out of them: a program that people in 1966 confided their most personal thoughts to, and that alarmed its own author.