LANG·I Language Chapter 11 of 65
The inquiry report
On June 4, 1996, Ariane 5 broke up on its very first flight, about 39 seconds after main engine ignition, because one number did not fit into sixteen bits. This chapter is laid out like the inquiry board’s report, and each of its recommendations becomes a tool of yours: exceptions, checks, tests, debugging.
Language
- 01 First program
- 02 Variables
- 03 Conditions
- 04 Loops
- 05 Functions
- 06 Lists
- 07 Strings
- 08 Dictionaries
- 09 Recursion
- 10 Functions as values
- 11 Errors and tests you are here
- 12 Objects
Builds on: 05 · Words of your own
What you will take away
- catch and handle errors with try, except, else and finally, and raise your own with raise
- write down what your code assumes as checks, and write tests that catch real bugs
- hunt for a bug methodically: read the traceback, print telemetry, split the suspects in half
The last chapter ended with a warning: the shorter a line of code and the more it does, the easier it is for a bug to hide in it, and the more that bug can cost. One of the best-known failures in the history of programming shows how much. This chapter is laid out like the report of the board that investigated it, with the same sections: the circumstances, the chain of events, the conclusions, the recommendations. Each of the board’s recommendations will become a tool you can use in any program you write.
Foreword
The Lions report is a model of how to investigate a failure. It is short enough to read in half an hour, and it looks for causes rather than culprits. The board worked backward in time from the explosion, link by link, until it reached the first cause, and closed with a list of fourteen recommendations. We will follow the same path.
1. The failure
A rocket has two “inner ears”: its inertial reference systems, called SRI in the report (from the French Système de Référence Inertielle). Each is a box with laser gyroscopes, accelerometers and a computer of its own, which works out from their readings which way the rocket is pointing and how fast it is moving. The on-board computer receives these data over a bus and steers the engine nozzles by them. So that the failure of one box doesn’t doom the flight, there are two: one active, the other in hot standby. If the on-board computer sees that the active system has failed, it switches to the backup at once.
Below is flight 501 second by second, as the board reconstructed it; you can move the time back and forth. Each switch underneath asks “what if…?” and tests one of the board’s recommendations.
At 36.7 seconds after H0 the backup system, SRI 1, declared itself faulty and fell silent. About 0.05 seconds later the active SRI 2 fell silent for the same reason. By then the backup had been silent since the previous data cycle (a cycle lasts 72 milliseconds), so there was nothing left to switch to. Worse, before shutting down, SRI 2 sent a diagnostic message about its own failure onto the bus, and the on-board computer took those bits for flight data. According to those bits the rocket had swung far off course, so the computer ordered a correction for a deviation that wasn’t there: it turned the booster nozzles as far as they would go, and a moment later the main engine’s nozzle too. The rocket turned abruptly. Its angle of attack, the angle between its axis and the oncoming air, went past twenty degrees, the aerodynamic loads began to tear it apart, and about 39 seconds after H0 the self-destruct system fired.
2. The chain of events
The board traced the chain from the end back to the beginning: the explosion came from the angle of attack; the angle of attack from the turned nozzles; the nozzles from wrong data; the wrong data from the failure of both SRIs; the failure from an exception in the software. You read a traceback from Chapter 1 the same way, from the bottom up: from what happened to where it happened.
The first link turned out to be a function that had no business running in flight at all. Before launch an inertial system has to be aligned: it has to work out which way is down and which way is north. That is the job of the alignment function. On earlier Arianes it was left running for fifty seconds after the SRI switched to flight mode, which covered the first few dozen seconds of flight as well. If the countdown was stopped in its last seconds, it could be resumed without waiting for a fresh alignment, which takes 45 minutes or more. Over the years the option was used once, in 1989, on flight 33. Ariane 5 had no use for it at all, but the SRI software was carried over from Ariane 4 almost unchanged: why touch something that works well?
Among other things, the alignment computed a quantity called BH, the horizontal bias, which is tied to the rocket’s horizontal velocity. At one point BH, computed as a 64-bit floating-point number, had to be converted into a 16-bit signed integer. In its first seconds Ariane 5 follows a different trajectory: its horizontal velocity builds up five times faster than Ariane 4’s. At 36.7 seconds after H0, BH became larger than sixteen bits can hold.
Sixteen bits
Sixteen binary digits give $2^{16} = 65\,536$ different combinations. Give half of them to the negative numbers and the other half to zero and the positive ones, and you get a range from $-32\,768$ to $32\,767$. How the sign is written in bits is the subject of Chapter 28. For now it is enough that 40,000 doesn’t fit in such a cell. The result is an overflow.
Python’s integers never overflow: 2 ** 1000 from Chapter 0 came out as a 302-digit number, and Python printed every digit. But a processor’s registers have a fixed width, and almost every other language, along with the libraries for fast numerical work, uses them directly. So does numpy, the library that scientific computing in Python is built on. It has a type int16, the same as the one in the SRI.
Each line behaves in its own way. In the first, 32,767 plus 1 gives −32,768: the cell rolled over like the odometer of an old car and came round to the far end of its range, while numpy issued a warning and kept counting. In the second, converting an array to int16 silently turned 40,000 into −25,536 and 70,000 into 4,464. In the third, numpy refused to make an int16 out of a Python number that was too big and raised OverflowError. It has done so since version 2.0; the 1.x versions wrapped the number around (the last of them with a warning). Quietly spoil the number or loudly refuse: on overflow every machine picks one of the two. Try the register for yourself.
The SRI software was written in Ada, and the oversized number was not spoiled silently: the conversion stopped with an exception, which the report calls an Operand Error. In that respect everything worked as it should. The machine noticed that the number didn’t fit and said so. The question is who was listening.
The exception
When a function can’t do its job, it raises an exception instead of returning a value: an object that describes what went wrong. The function stops right there, and the exception flies to the function that called it. If that one doesn’t catch it, the exception flies on to its caller, and so on up the call stack from Chapter 5. If it reaches the top and nobody has caught it, the program stops and prints a traceback, the path the exception took. Here is a scaled-down copy of the SRI. Its numbers are made up, since the public report gives no values of BH.
The word raise throws an exception. The traceback shows all three frames it flew through: flight_cycle, alignment, to_int16. The program’s last line never ran. Press “Steps”: at second 37 the frames disappear one after another.
To catch an exception, you use try:
Python runs the try block. If an exception named in except happens there, control jumps to that block and the program carries on. The else block runs only if there was no exception. The finally block runs every time, even when the exception flies on; it is the place for whatever must be done no matter what, such as closing a file or writing to a log. In the cell, finally printed its line even after return.
Exceptions come in kinds, and the kinds form a family tree. ZeroDivisionError and OverflowError are special cases of ArithmeticError, KeyError and IndexError are special cases of LookupError, and all of them descend from Exception. A handler except ArithmeticError catches division by zero and overflow alike.
The form except LookupError as e puts the caught exception into the variable e, and you can ask it for its type and message. Where an exception ends up depends on where the handlers are and what they catch. Test yourself.
main → flight_cycle → alignment → to_int16. Choose what goes wrong at the bottom and which handlers sit in each function, then raise the exception. On the right is the same program in Python; the “Check in Python” button runs it on the server.What the handler did
The exception in the SRI was caught. There was a handler, and it did everything the specification asked of it: on any exception, report the failure on the bus, store the context in non-volatile memory (it was read out once the boxes had been recovered from the debris) and shut down the processor. In Python it would look roughly like this:
The report traces the decision to shut the processor down to “the culture within the Ariane programme of only addressing random hardware failures.” If a failure is random, say a chip burns out, it makes sense to switch the sick unit off and go over to the backup. A software bug is not random. Both SRIs ran the same program on the same data, so they failed in the same way, five hundredths of a second apart. Two sound boxes switched themselves off because of a function that wasn’t needed in flight. The board summed it up like this: the view had been taken “that software should be considered correct until it is shown to be at fault,” and the board was “in favour of the opposite view, that software should be assumed to be faulty until applying the currently accepted best practice methods can demonstrate that it is correct.”
Replace return None in the cell with return attitude, and both systems go on sending the right attitude; only the alignment fails. One of the board’s recommendations says to do that; we will come to it shortly.
Why the tests missed it
Rockets are tested for years, and Ariane 5 was no exception. But the SRI was tested on the bench as a box, against every external influence, and with margins beyond what was required. Running it together with the whole flight control system along the Ariane 5 trajectory was not done: by common agreement, the trajectory had been left out of the SRI requirements. In the large flight simulations, which used the on-board hardware itself, both SRIs were replaced by software models, because in 1992 it had been decided that the boxes themselves were already proven. The decision had its reasons, and the board called them “technically valid.” But, as it also wrote, “Had such a test been performed by the supplier or as part of the acceptance test, the failure mechanism would have been exposed.” After the accident the SRI software was run on a computer with the recorded trajectory of flight 501, and the simulation faithfully reproduced the chain of events that led to the failure of both SRIs.
One more detail. The investigators found seven variables in the alignment code that could overflow on conversion to an integer. Four were protected and three were not: the SRI computer was not to be loaded beyond 80%, and those three quantities, by the engineers’ reckoning, were either physically limited or had a large margin. For BH the reckoning was wrong. The code itself contained no reference to the reasoning behind this decision.
3. Conclusions
In the board’s words, “The failure of the Ariane 501 was caused by the complete loss of guidance and attitude information 37 seconds after start of the main engine ignition sequence (30 seconds after lift-off). This loss of information was due to specification and design errors in the software of the inertial reference system.” The reviews and tests carried out over the years of development, it goes on, did not include adequate analysis and testing of the inertial system or of the complete flight control system, which could have revealed the failure in advance.
For the accident to happen, many things had to coincide: a useless function ran in flight; a variable was left unprotected; the trajectory differed from the old one; the handler shut the processor down; the two systems were identical; in the simulations, models stood in for the boxes. Remove any link and the chain breaks. That is why the report has so many recommendations.
4. Recommendations
Of the board’s fourteen recommendations, we will take the ones that apply to any program, not only a rocket’s, and turn each into a tool. In the boxes below are the report’s own words.
R1. Nothing you don’t need
The cheapest code is the code that isn’t there: nothing in it can break. Every line that runs “just in case” is a place for a bug to hide. Before you carry a piece of an old program over into a new one, ask what it is doing there.
R3. A best estimate instead of silence
An exception handler decides what happens next, and the decision can go different ways. Sometimes the right move is to fail loudly: better to stop the program than to go on computing with corrupted data. Sometimes the right move is to carry on with a fallback value, as R3 advises. The only wrong move is not to think about it at all. A good handler follows two rules. First, catch only the exception you expect and know how to deal with: except ZeroDivisionError, not except Exception. A broad handler also swallows errors you never thought of, a misspelled variable name, for example. Second, keep the try as small as you can, around only the lines where you expect the error.
Write a function safe_divide(a, b, default=None) that returns a / b, or the value default when it would divide by zero. It must not swallow other errors: safe_divide("1", 2) has to fail with TypeError, as ordinary division does, or else a mistake in the data would go unnoticed.
Wrap the division in try. Which exception does 1 / 0 raise? The answer is in the last line of the traceback.
except ZeroDivisionError: return default. That type and no other: with Exception, a TypeError would turn into None as well.
You could also check beforehand: if b == 0: return default. Both styles are common. Python code leans toward “try it and catch the failure,” a habit known as “easier to ask forgiveness than permission,” because when you check beforehand it is easy to forget some case. Here both work, since 0 and 0.0 are both equal to zero. But except Exception, or a bare except:, fails the test with the string. A mistake in the data quietly becomes None and surfaces somewhere far away, like the SRI diagnostic bits that the on-board computer took for flight data.
R5. Name your assumptions
Every line of code assumes something. int(text) assumes the string holds a number. sum(mags) / len(mags) assumes the list isn’t empty. Converting BH to 16 bits assumed that BH stayed below 32,768, and that assumption lived in the documents but not in the code. R5 says: drag your assumptions into the light. Python gives you two tools for it.
The first is raise with an exception of your own, one with a clear name. If a function receives something it can’t work with, it is better for it to say so at once, plainly, than to return garbage. You declare your own exception as a descendant of a suitable built-in one. We will take the word class apart in the next chapter; for now it is a two-line recipe:
The second is assert. It is a check that stays silent while its condition holds and raises AssertionError when it doesn’t. You use it to write down what you, the programmer, believe to be impossible: if it happens anyway, the program has a bug.
The typo is caught the moment it enters the function. Without the check it would surface a thousand lines later, when a mean magnitude of 46.6 ruined a report. So how do the two tools differ? raise is for errors that can happen in a correct program: a person typed letters instead of an age, a file wasn’t found. assert is for errors in the program itself. Python can be run with the -O switch, which skips every assert, so they must never be used to check what a user typed.
What people type is the main source of errors that happen in a correct program. The program below catches the ValueError that int raises and asks again until it gets a number. Run it and try answering “forty”. In the task after it you will raise an exception yourself, one of your own with a clear name.
Declare an exception AgeError that descends from ValueError, and write a function parse_age(text). It receives a string a person typed and returns the age as an integer. Spaces and newlines around it don’t matter: parse_age(" 42\n") == 42. If the string doesn’t hold a whole number, or the age is outside the range from 0 to 150 inclusive, the function raises AgeError with a clear message. It must be AgeError, not a plain ValueError.
int("abc") raises ValueError on its own. Catch it and raise yours in its place: inside except ValueError: write raise AgeError(f'"{text}" is not a number').
int accepts "-5" too, so check the range separately, after the conversion. int("42.5") also raises ValueError, so fractions are weeded out by themselves.
Whoever calls parse_age can catch AgeError without knowing which exceptions int raises inside: the details are hidden. And because AgeError descends from ValueError, old code with except ValueError keeps working. The tail from None is optional: it tells Python not to show int’s original error in the traceback, since we have replaced it with our own, clearer one. Incidentally, int ignores spaces around a number by itself, but strip makes the intention explicit and the message cleaner.
The next task is about the SRI register itself and the number that wouldn’t fit in it. There are no exceptions in it, only arithmetic.
Write two functions. wrap16(n) takes an integer of any size and returns what ends up in a signed 16-bit cell if you write $n$ into it the way C’s integer arithmetic and numpy do it, wrapping around. For example, wrap16(32768) == -32768, wrap16(40000) == -25536, wrap16(-1) == -1. clamp16(x) is the R3 strategy, the “best estimate”: drop the fractional part of the number, as int does, and if the result doesn’t fit, return the nearest limit: clamp16(40000.7) == 32767, clamp16(-2.7) == -2. Infinity has to be handled too: clamp16(float("inf")) == 32767.
Wrapping around means working modulo $2^{16} = 65\,536$. In Python the remainder n % 65536 always lies between 0 and 65,535, even for negative $n$. All that is left is to move the upper half down.
For wrap16: (n + 32768) % 65536 - 32768. For clamp16, try int(float("inf")) and you get OverflowError. So the limits have to be checked before int is called.
Shifting by 32,768 moves the range $[-32\,768, 32\,767]$ to $[0, 65\,535]$, the remainder wraps the number around, and shifting back restores the sign. In clamp16 the comparisons come before int: infinity compares with numbers without any trouble but can’t be turned into an integer. Instead of x >= 32767 you could have written > 32767: for 32,767.5 both give 32,767, so here the choice makes no difference. In the task further down, spots like this are where the mutants hide.
R2, R10, R11. Test as you fly
In Chapter 1 we got into the habit of checking an answer against an example we knew in advance. Here the habit becomes a tool. A unit test is a small function that calls the code under test on a known input and compares the result with the expected one using assert. You write the tests once and run them after every change to the program: if something breaks, they tell you straight away.
The last lines are a tiny test runner. It rests on the idea from Chapter 10 that functions are ordinary values: globals() returns a dictionary of every name in the program, and the functions in it can be looped over and called. Larger projects use a ready-made runner, most often pytest: it finds the test_… functions in all the files by itself, runs them and prints a report, and for checking exceptions it offers pytest.raises.
The ordinary case works anyway, so a good test checks the edges. Programs break most often on edge cases: an empty list, a single element, zero, the largest and smallest allowed values and their neighbors on the far side of the line, repeated values, a huge input. test_edges holds the edges themselves, 32,767 and −32,768; test_overflow holds their neighbors outside. The Ariane 5 trajectory was an edge of this kind, and nobody had thought of it.
Remember the docstring from Chapter 5? We put the example “digit_sum(1843) == 16” in it. Write it as a Python session instead, and the doctest module will check it automatically:
doctest finds the lines in the docstring that begin with >>>, runs them and compares each result with the line below. Change 16 to 17 and run it again to see what a failure looks like. Documentation that checks itself can’t drift out of date unnoticed. That is the spirit of R12: “Give the justification documents the same attention as code.”
Who tests the tests
The tests pass. Does that mean there are no bugs? Edsger Dijkstra answered that in 1970: “Program testing can be used to show the presence of bugs, but never to show their absence!” And tests can be weak, too, checking only what never breaks anyway.
To find out what your tests are worth, break the program on purpose and see whether they raise the alarm. Take to_int16 and make several mutants of it, copies that each carry one small bug: > replaced by >=, a check left out, int swapped for round. If even one test fails on a mutant, the mutant is “killed.” If all the tests pass, the mutant survives: your tests would have let that bug through. Richard Lipton proposed mutation testing as a student in 1971, and in 1978 Richard DeMillo, Lipton and Frederick Sayward described it in a paper titled “Hints on Test Data Selection: Help for the Practicing Programmer.”
to_int16 and runs your tests from the cell above. Add tests to the cell and release the mutants again. Tap a survivor to see what is broken in it. One of the mutants has a catch.The two tests in the cell kill nobody: 42 and −100 are far from every edge, and whole numbers are left alone by both round and wrapping around. Add tests until all five are caught. Mutants are used in industry as well: at Google, according to the company’s engineers (2018), surviving mutants are shown to the authors of a change during code review. Mutants are small changes, and most bugs in the wild are small too, so a test that kills mutants also catches live typos.
Write tests for the function to_int16 from this chapter: functions whose names start with test_ and which check its behavior with assert. The server will run your tests first on the correct to_int16, where they must all pass, and then on the five mutants from the zoo above, leaving out the twin. Each mutant must be killed by at least one of your tests.
Mutants hide at the edges. Check the edges themselves, 32,767 and −32,768, and their neighbors outside, which must raise OverflowError. Do it on both sides, above and below.
Two more mutants have to do with the fractional part. What should to_int16(2.7) return? And is to_int16(32767.5) an overflow or not?
To check that an exception is raised, call the function inside try, put return in except OverflowError, and after the block write assert False, "no exception".
Every line here kills somebody. 32767 kills the mutant with n >= 32767. -40000.0 kills the one without the lower check. 2.7 kills the one with round. 32767.5 kills the one that checks the limits before dropping the fraction. And any overflow kills the mutant that silently wraps the number around instead of raising. Nobody kills the twin: after int, n is a whole number, and for whole numbers n > 32767 and n >= 32768 are the same condition, so it is the same program written another way. Such equivalent mutants are the method’s perennial headache: no program can tell a twin from a surviving mutant in general, and Chapter 56 explains why.
R7. More telemetry
The board was lucky: both SRIs turned up among the debris, and their memory could be read. Otherwise part of the chain would have stayed a guess. Your program, too, is a rocket already in flight, and its telemetry is whatever it prints. Debugging starts when you stop guessing and see what is going on inside. Here are four techniques, from the plain to the crafty.
Read the traceback. From the bottom up: the type of error, the message, the line, the path through the functions. Many bugs end right there.
Print. The oldest technique, and still the most useful, is a print in a suspicious place. f-strings have a handy form for it: print(f"{bh=}") prints both the name and the value, bh=40000.0. Print the values you have doubts about and compare them with what you expect; a line that only says “here” tells you little. The “Steps” button in this course’s cells is the same printing, done for you at every step.
Split in half. The program works on a small file and fails on a big one. Take half the file: if the program fails, the bug is in that half; if not, it is in the other. Each run throws out half the suspects, so among a million lines the culprit is found in twenty runs: the binary search from Chapter 0, working as a debugger. The earthquake catalog, it turns out, hides a culprit of this kind too. Here is a computation that takes the logarithm of each depth and fails on the whole catalog. We will find the culprit without reading the traceback:
The culprit has a depth of 0.0 km, and zero has no logarithm. Nineteen thousand records took fourteen runs. The same method works on code. Yesterday the program worked, today it doesn’t, and a hundred changes were made in between: check the middle of the change history, then the middle of the half… The version control system git does this for you with the command git bisect; git comes up in Chapter 40.
Tell the duck. The Pragmatic Programmer by Andrew Hunt and David Thomas (1999) tells of a programmer who kept a rubber duck on his desk and explained his code to it, line by line. It sounds like a joke, but it works more often than you would think: to explain, you have to say every assumption out loud (“the list isn’t empty here, because…”), and on one of them you will stumble. It is R5 again, spoken aloud. The duck can be replaced by a colleague, a cat or a chat message you never end up sending.
What next
Four years after the accident, four new Cluster satellites went into space on Russian Soyuz rockets, two at a time, and worked in orbit for many years. Ariane 5, once modified, flew until 2023 and among other things launched the James Webb Space Telescope. Bugs didn’t go away, but since then people have hunted them differently.
You now have a toolbox for reliability: exceptions, checks, tests, debugging. Meanwhile the programs we write are growing longer and more crowded. A bank account, a game character, a rabbit on an island: each has its own state and its own behavior, and there are thousands of them. Dictionaries and loose functions are too cramped for them. You have already written the mysterious word class twice, declaring your own exceptions. In Chapter 12 we will find out what it means, and we will populate an island with rabbits and foxes to watch their numbers rise and fall year after year.