LANG·I Language Chapter 2 of 65

Names and values

A chapter staged as a magic show: eight tricks with variables. Before each one you place a bet, then run the code and find out how the trick works. And at the end, a story in which the same arithmetic stops being a game: in 1991 a tenth of a second, stored imprecisely, made a missile defense system lose its target.

From zero 45 minutes Python Programming

Builds on: 01 · First program, first bug

What you will take away

  • keep values under names, and see a name as a label on an object rather than a box
  • tell int, float, str and bool apart, convert one into another, and ask a person for data with input()
  • never compare fractions with ==, and count money and time in whole numbers

Chapter 1 ended with a barrel organ: our programs play the same tune every time. They don’t remember what they computed a line earlier, and they can’t ask a person anything. To remember a number, you have to give it a name. That sounds simple, but names in Python have secrets of their own, so this chapter is staged like a magic show. Before each trick you bet on what the program will print, then you run it, and we reveal how the trick was done.

There are no false bottoms here: every trick lives in the language itself, and each one has bitten somebody in real code. The show ends before the chapter does. After it comes a story in which the same mistake with fractions cost twenty-eight people their lives.

Trick one. x = x + 1

A mathematician who sees the line x = x + 1 will tell you it’s impossible: no number equals itself plus one. Here is a program with that line in it. Don’t run it yet; place your bet first.

What does the program print?

Twelve. In Python the sign = isn’t an equation but an order: “work out what’s on the right and give the result the name on the left.” The second line takes the current x (five), adds one and hangs the name x on the six. The third does the same with multiplication: $6 \cdot 2 = 12$.

Run it and check your bet. The whole trick is in the equals sign. In mathematics $x = x + 1$ is a statement, and a false one. In Python x = x + 1 is a command, and its order is strict: first the right-hand side is worked out with the old value of x, then the name on the left is tied to the result. A command like this is called an assignment, and a name that can be tied to different values is a variable.

Now a program can remember. Here it works out how many seconds there are in a day, so we don’t have to keep the intermediate numbers in our heads:

A name can be made of letters, digits and the underscore, but it can’t start with a digit. Upper and lower case are different: Day and day are two names. Python even allows names in other alphabets (Δt = 5 works), but the convention is plain Latin letters in lower case, with words separated by underscores: seconds_per_day. A good name says what it stands for. A month from now you’ll understand seconds_per_day at a glance, while s will make you think.

And don’t confuse a name with text: print(seconds_per_day) prints the number the name points to, while print("seconds_per_day") prints those fifteen characters themselves.

Trick two. Where did the ten go?

The next trick is shorter, but it hides the chapter’s main secret.

What does the program print?

11 10. The line b = a ran once, when a was ten, and nothing ties the two names together after that. It is a one-off command: “let b point wherever a points right now.”

There are two pictures of what a variable is. School textbooks often draw a box: a is a box with 10 inside, and b = a puts a copy of the ten into box b. Python works differently. The number 10 lives in memory on its own, like a thing on a shelf: it is an object. The name a is a label on a piece of cord, tied to that object. The command b = a copies nothing: it ties a second label to the same ten. And a = a + 1 creates a new object, 11, and moves label a onto it. The ten stays where it was, with b hanging on it.

The diagram below shows this step by step, and its pictures aren’t drawn in advance: the code goes to the server, runs one line at a time, and after each line the server records which name is tied to which object.

On the left is the program, on the right memory: the name labels and the objects they point to. Step with the arrows. The tabs at the top are the chapter’s other tricks. The “Your code” button lets you type in any program: it runs on the server, and the diagram is built as it runs.

Every object has a number, which the function id() gives you. If two names have the same number, they point to one and the same object. A shorter check is the operator is: a is b means “a and b are one object.”

With numbers, boxes and labels give the same answer, and labels look like a needless complication. You can’t see the difference because a number can’t be changed: you can only compute a new one and move the label onto it, so you’ll never notice that two names are looking at the same ten. In the eighth trick we’ll meet an object that can change, and the box picture will break down. Meanwhile, one more act with labels.

Encore: the swap

You need to swap the values of two variables, so that a ends up with what b had, and vice versa. The first thing that comes to mind:

What does the program print?

“coffee coffee.” After a = b both labels hang on the string "coffee", and not one is left on "tea". An object with no labels is out of the program’s reach: no name points to it. When the next line runs b = a, label b moves to wherever a points, which is the coffee, and the tea is lost.

Open the “The swap” tab of the diagram and find the moment when the tea is left without a label. The classic cure is a third name to hold the tea while the labels move: tmp = a, a = b, b = tmp. Python has a shorter way. To the right of = you may write several values separated by commas, and on the left the same number of names:

The same rule of assignment is at work. First the right-hand side is worked out, giving a pair (the object b points to now and the object a points to), and only then do the labels move. While the right-hand side is being worked out, no label has moved yet, so there is nothing to lose. The “One-line swap” tab shows it step by step.

Trick three. Grains on a chessboard

According to an old legend, the inventor of chess asked the ruler for wheat as his reward: one grain on the first square of the board, two on the second, four on the third, and on every square twice as many as on the one before. The board ends up holding $2^{64} - 1$ grains (why that many is explained in the chapter on sequences of our math course, “Mathematics, the Queen of the Sciences”). Python can count them for us.

What will Python print?

All the digits: 18446744073709551615. Whole numbers in Python have no upper limit; they grow as far as memory allows.

Betting on “error” or “wrap around” is far from silly. In most languages a whole number takes a fixed amount of memory, usually 32 or 64 binary digits, and the largest number that fits in 64 digits is $2^{64} - 1$, the legend’s reward down to the last grain. One more grain, and a 64-bit counter like that goes back to zero, like a car’s odometer after 999,999 miles. Chapter 28 explains why, and what came of it in practice. Python, in that situation, takes more memory: try typing 2 ** 1000 or 10 ** 100 into the cell.

To get a feel for this number, imagine someone counting the grains one per second, without sleep or days off.

Five hundred and eighty-five billion years, forty-two times as long as the universe has existed. Python ignores the underscores inside 13_800_000_000; they are there so that people don’t lose count of the zeros.

Look closely at the output: grains has no decimal point, while years has one, with a fractional part after it. They are values of different types. A whole number has the type int (from integer), a fractional one float (from floating point). The function type() tells you the type of any value:

Division with / always gives a float, even when the division is exact: 6 / 2 is 3.0, not 3. A whole quotient comes from //. And float, unlike int, does have limits, and more than one. Python showed 2.0 ** 64 approximately: 1.8446744073709552e+19 means $1.8446744073709552 \cdot 10^{19}$, and the last digits are lost. A fractional number keeps only about sixteen significant digits. And 2.0 ** 1024 already raises OverflowError: the largest float is about $1.8 \cdot 10^{308}$.

Trick four. 0.1 + 0.2

Every programmer knows this trick, because every programmer has fallen for it at some point.

The double sign == asks “Are they equal?” What will Python answer?

False. The next cell shows what the sum is.

0.30000000000000004. And ten tenths make 0.9999999999999999, a shade under one. Python isn’t broken: almost every programming language and almost every processor compute this way. The explanation lies in how the machine writes fractions down.

We write fractions in the decimal system: $0.1$ is one tenth. The machine keeps numbers in binary, where a fraction is a sum of halves, quarters, eighths and so on. One quarter has an exact binary form: 0.01. One tenth doesn’t: its binary form goes on forever, 0.000110011001100…, with the tail 0011 repeating endlessly. In the same way one third goes on forever in decimal, $0.333\ldots$ Why some fractions have a tail that ends and others one that repeats forever is explained in the chapter on fractions of the math course. An endless tail doesn’t fit in memory, so the machine cuts it off: a float keeps 53 significant binary digits. So what sits in memory in place of 0.1 is the binary number closest to one tenth. Here it is under the magnifier.

Type a decimal fraction. The machine converts it to binary by doubling: every doubling pushes out one binary digit. Below is what the float ends up holding, to the last digit, and a check of the addition.

You can see the exact value stored in place of 0.1 in Python too. The decimal module can show a float without rounding:

Normally Python prints a fraction in the shortest form that turns back into the same number. That is why print(0.1) shows 0.1 and creates an illusion of precision. But 0.1 is stored as a number slightly greater than one tenth, 0.2 as one slightly greater than two tenths, and their sum overshoots further still: it turns out to be a different binary number from the one stored in place of 0.3. The difference is in the seventeenth decimal place, and == compares numbers down to the last binary digit.

Don’t compare fractions with ==. Compare them with a tolerance, “equal to within a billionth”: abs(x - y) < 1e-9. The math module has a ready-made function for this, math.isclose(x, y). And money and time, where every cent and every fraction of a second counts, are best kept in whole numbers: cents instead of dollars, milliseconds instead of seconds.

Two more tricks with rounding

round(2.5) gives 2, and round(3.5) gives 4: Python rounds halves to the nearest even number. This rule is sometimes called bankers’ rounding: with it, rounding errors in long sums don’t pile up in one direction. And round(2.675, 2) gives 2.67, although by the rule you learned at school it should be 2.68: in memory 2.675 is stored as 2.67499999999999982236431605997495353221893310546875 (check it with the magnifier above). How a float is built, with its sign, exponent and mantissa, is taken apart in Chapter 28.

Trick five. Two and two make twenty-two

What does the first line of the program print?

22. The twos in quotes are strings, type str. A plus glues strings together, and an asterisk with a number repeats a string: "2" * 3 is "222".

The same sign does different things depending on the type. 2 + 2 adds numbers, "2" + "2" glues strings together. And "2" + 2, as we saw in Chapter 1, is a TypeError: Python won’t try to guess what you meant. The function len() counts the characters in a string, and a space is a character too: “Ada Lovelace” has twelve.

Neither a number nor a string

One more type is left, the smallest of all: it has only two values.

A comparison gives a Boolean value, type bool: True or False. In the next chapter such values will decide which way a program goes. And True + True equals two for historical reasons: the separate type bool appeared in Python only in version 2.3 (2003); before that, true and false were written as one and zero, and so that old programs keep working, True still behaves like 1.

Trick six. The machine reads your mind

A classic trick: “Think of a number, do a few things to it, tell me the result, and I’ll name the number you thought of.” Run the program and it will ask you right under the code. But first, a bet.

You thought of 7 and honestly answered 76. What happens?

An error. The function input() always returns a string, even if you typed nothing but digits. Given "76", the program tries to subtract six from a piece of text, and Python refuses.

The function input() prints a question, waits until the person types an answer and presses Enter, and returns what was typed, as a string. The machine has no way of knowing you meant “76” as a number: you could just as well have answered “seventy-six” or “not telling.” Turning the string into a number is your job, and the function int() does it:

Now the machine “reads minds.” Behind it is school algebra: if you thought of $n$, you said $(5n + 3) \cdot 2 = 10n + 6$. Subtract six, divide by ten, and there’s $n$. Try answering something odd: “76.5,” “seventy-six,” an empty line. Each time Python will answer ValueError: the type is right, a string, but no whole number can be made from that value.

Four functions do conversions, one per type: int(), float(), str() and bool(). In places they don’t behave the way intuition suggests. Test yours: first guess what a cell will show, then click it.

The type customs: rows are values, columns are what they are converted into. Click a cell to see Python’s answer and an explanation.

Three things from the customs are worth remembering. int() doesn’t round; it drops the fractional part: int(3.99) is 3, and int(-3.99) is minus 3. int() doesn’t understand a string with a decimal point: int("3.5") is an error, so you need float() first. And bool() counts everything non-empty as true: bool("0") is True, because the one-character string “0” isn’t empty. Only zero, the empty string and a few other “empty” values are false; we’ll come back to them in Chapter 3.

Trick seven. A string that does its own sums

What does the third line of the program print, the one without the letter f?

The curly braces stay as they are. To Python this is an ordinary string, and {name} in it is six characters, nothing more. Substitution works only in a string with the letter f before the quote.

A string with the letter f before the quote is an f-string (from “formatted”). Inside the curly braces you can write any expression: a name, arithmetic, a function call. Python works it out and puts the result into the text. That is much handier than gluing pieces together with pluses and making sure no spaces go missing and every number has been turned into a string. F-strings arrived in Python 3.6 (2016), proposed by Eric V. Smith, and have since become the usual way to build text.

After a colon inside the braces you can say how to show the value:

:, splits a number into groups of three digits, :.2f keeps two digits after the decimal point (rounding), :.1% shows a share as a percentage. The last line is a gift for debugging: an = at the end of the braces prints both the expression and its value.

Trick eight. A list with two names

Here is the promised trick that breaks the box picture. It contains something we haven’t covered yet, a list, the subject of Chapter 6. For now two hints are enough. Square brackets create a list, several values under one name. And crew.append("Grace") adds one more value to the end of the list.

We added Grace to crew. What is in team?

['Ada', 'Grace']. There is one list with two labels on it. Change the list through one name, and the change shows through the other.

If variables were boxes, crew = team would have put a copy of the list into crew, and Grace would have ended up only in the copy. Labels explain what happened: crew = team hung a second label on the same list, and append changes that list without creating a new object. Open the “The list” tab of the diagram: the arrows from both names lead to the same point.

This never happens with numbers and strings: they can’t be changed, only replaced by others. A list can change, and without the picture of labels you can’t make sense of it. This trap is one of the most common bugs in Python programs. In Chapter 6 we’ll learn to make independent copies, and in Chapter 12 we’ll see that every object you create yourself works the same way.

Dhahran, February 25, 1991

A year later the US General Accounting Office (GAO) published a report on the investigation. Here is what it says, put into the language of this chapter.

The computer’s clock counted time in tenths of a second, as a whole number of ticks: 1, 2, 3… To get seconds, the program multiplied the number of ticks by one tenth. But the computer’s registers, its working cells, were 24 bits wide, and one tenth was kept in them as a binary fraction with 23 digits after the binary point: 0.00011001100110011001100. The endless tail 1100… was cut off, as under our magnifier, only much sooner than in a float. With every tick the clock fell short by about 0.000000095 seconds. Here is what that grows into.

In the second line int() drops the fractional part, the way the 24-bit register cut off the tail. In a hundred hours the clock fell behind by thirty-four hundredths of a second. In that time a Scud flying at about 1.7 kilometers a second (the report says about five times the speed of sound) moves more than half a kilometer.

The radar doesn’t follow a target continuously. Once it has spotted one, it calculates where the target will be next and looks only there, inside a range “gate.” The calculation rests on the target’s speed and the time of the last detection. If the time is wrong, the gate is in the wrong place, and the longer the system runs without a restart, the farther off it drifts.

The battery’s clock. The slider sets how many hours the system has been running without a restart. The clock error is computed as in the cell above; the gate shift comes from the table in the GAO report. The report doesn’t give the gate’s width: we estimated it from the Israeli data (a 55-meter shift after 8 hours is 20 % of the gate).

The lagging clock alone doesn’t explain the miss, as Robert Skeel, a numerical analyst, pointed out when he examined the report for SIAM News. Aiming doesn’t need absolute time; it needs the difference between two moments: when the target was here and when it was there. If both moments carry the same error, the errors almost cancel when you subtract. But the program, written in assembly language some twenty years before the war, had been modified several times to handle fast ballistic missiles it was never designed for, and one of those modifications added a more precise conversion of ticks into seconds, but not in every place that needed it. One of the two times came out precise, the other with its tail cut off, and their difference no longer canceled the error.

The timeline in the report makes for grim reading. On February 11 the Army received data from the Israelis: after eight hours of continuous operation the radar’s gate had shifted by 20 %. The remedy was simple: restart the system every few hours. A restart took a minute to a minute and a half and reset the clock. On February 21 the batteries were sent a message that “very long” run times could shift the gate, without saying how long “very long” was. By the evening of February 25 Alpha Battery had been running without a break for more than a hundred hours. The corrected software arrived in Dhahran the next day, February 26.

Two rules follow from this story, and they matter more than anything else in the chapter. You know the first already: fractions in a machine are approximate, and you can’t compare them with ==. The second: wherever you can, count in whole numbers, and turn them into fractions once, at the end. Compare:

Ten ticks added up a tenth at a time give 0.9999999999999999. The same ten ticks counted as a whole number and divided once give exactly 1.0.

Your turn

Four tasks: two small calculator programs with input, a swap and a trap with fractions. In each, the program asks for data with input(): when it is checked, the server supplies the answers, and when you run it yourself, the program asks you.

The program asks how old you are in whole years and prints how many seconds that is, if every year has 365 days. The last line of the output must be a single number with no spaces inside; for 15 years, say, 473040000.

So far the starter counts days, and counts them wrong. Run it and answer “2”. Why did it print a long string of twos instead of 730?

input() returns a string, and a string multiplied by a number is repeated. Turn the answer into a whole number, then multiply the days by hours, minutes and seconds.

The starter repeated the string "2" 365 times. Wrap the input in int() and carry the count on to seconds:

Fifteen years is 473 million seconds, and you turn a billion seconds old at about 31 years and 8 months.

A group of friends had dinner at a café. The program asks for the amount of the bill (it may have cents), the tip as a percentage (a whole number) and how many people are splitting the bill. Everyone pays the same, and the tip is split equally too. The last line of the output is one person’s share with two digits after the decimal point, in this form:

Each pays: 275.00

For a bill of 1000, a 10 % tip and four people: $1000 + 100 = 1100$, divided by four, 275.

The bill may have cents. Which function turns “1234.5” into a number? The percentage and the number of people are whole numbers.

The tip is bill * tip / 100. The total is divided by people.

An f-string gives two digits after the decimal point: f"{x:.2f}".

The :.2f format rounds as well: 1000 split three ways without a tip is 333.33. Where the last cent went is for you and your friends to settle.

The program reads two words, one per line, and prints them in reverse order separated by a space. No workarounds, though: swap the values of the variables a and b without introducing a third variable, and only then print a and b. The two lines with input() at the start and the line print(a, b) at the end of the program must stay as they are.

Remember the encore to the second trick: what happened when we wrote a = b and then b = a?

To the right of = you may write several values separated by commas, and on the left several names.

The right-hand side b, a is worked out in full before a single label moves, so no value gets lost. In languages without this kind of assignment you need a third variable: tmp = a; a = b; b = tmp.

A cash-register program gets a price in dollars and cents, as a string like 19.99, and must print the price in cents as a whole number: 1999. The starter looks right, but for some prices it is off by a cent. Find such a price and fix the program. The last line of the output is a single whole number.

Run the starter with the price 19.99. Then print float(price) * 100 without the int().

19.99 is stored as slightly less than 19.99, so the multiplication gives 1998.9999999999998. And int() doesn’t round; it drops the fractional part.

round() without a second argument rounds to the nearest whole number and returns an int. The floating-point error here is less than a trillionth of a cent, and rounding removes it. More reliable still is never to let the price pass through a float at all: split the string at the decimal point and build the cents from two whole numbers. You’ll be able to do that after Chapter 7.

What next

Our programs have learned to remember and to ask, but they are still as straight as a railroad track without a single switch: they run from top to bottom, line by line, the same way every time. The mind reader can’t say, “If you thought of more than fifty, multiply by two, and if less, by three.” The bill splitter can’t warn you that there are zero people and nobody to divide by. For a program to choose what to do, it needs forks in the road: if this, then that, else something else. Every game and every input check rests on forks, and so does all the logic inside a processor. That is Chapter 3, where you’ll write a text game of your own.