NET·VI Networks Chapter 43 of 65
Anatomy of this page
You typed an address, and half a second later the page was in front of you. This chapter dissects the page itself: its name in DNS, the server’s response headers, the HTML your browser grew a tree from, and the load timings your browser wrote down a minute ago. Along the way: Tim Berners-Lee’s 1989 memo, on which his boss wrote “vague but exciting,” and a web server of our own, built from sockets.
Networks
- 41 Networks
- 42 TCP/IP
- 43 The web you are here
- 44 Distributed
Builds on: 42 · Inventing a protocol 07 · A conversation made of strings
What you will take away
- trace the path from the address in the browser bar to the finished page (DNS, connection, request, response, parsing and drawing) and find where the time goes
- write and read HTTP by hand: requests, status codes, headers, cookies, caching; and run your own server and JSON API in Python
- understand how a browser turns HTML into a tree and why it forgives mistakes that Python never would
The last chapter ended with a question: you type legost.in into your browser and press Enter. What happens? You can answer it from a textbook, or from the evidence. We have the evidence: you are reading this page, so everything the question asked about has already happened, a few seconds or minutes ago, in your browser. And the browser kept a record. It wrote down when it started looking for the server’s address, when it opened the connection, when the first byte of the response arrived and when it finished drawing. Here is that record.
The record has a dozen or so rows, each a separate event. Find where the server lives. Agree with it on a connection and on a cipher. Ask for the page. Wait while the server puts it together. Receive it, parse it, style it, draw it. For some readers the first row takes hundreds of milliseconds, for others zero, and why that happens will be clear by the middle of the chapter. We’ll do the autopsy by the book, from the outside in, from what you see in the address bar to what you see on the screen, and at the end we’ll come back to this timeline, when every bar on it makes sense.
Vague but exciting
To make the web work, Berners-Lee needed three inventions, and you used all three to open this page. The first is an address that names a document unambiguously anywhere on the network. The second is a protocol by which the browser asks a server for a document and the server hands it over. The third is a markup language to write the document in: where the heading is, where a paragraph is, where a link leads. The network to carry all this already existed: the packets of Chapter 41 and the reliable connections of Chapter 42. In Weaving the Web (1999) Berners-Lee recalled suggesting more than once, to the hypertext people and the internet people alike, that the two technologies be married; when nobody took the job, he did it himself. His three inventions are the specimens we’ll dissect, and a fourth will squeeze in between them, one that nobody had to invent in 1989: the server’s name.
Specimen one: the address
The line you see above the page is called a URL. It has a strict structure, and Python takes it apart with one function from the urllib.parse module. Here is the address of this chapter with parameters and an anchor tacked on, the way links look when people share them.
The scheme https says which language to speak with the server. The server name says whom to look for. The port says which “door” of the machine to knock on: we met ports in Chapter 41, and the web has its customary numbers, 80 for plain HTTP and 443 for the encrypted kind. The path /en/cs/basics/web is the server’s to interpret as it likes: once it was the path to a file on the server’s disk, and today it is more often a key by which the server’s program decides what to show. The parameters after the question mark are “name=value” pairs joined by &; parse_qs turns them into a dictionary of lists, because a name can repeat.
The last two lines answer a question that sooner or later occurs to everyone who has copied a link to the Wikipedia article on Pokémon: where does the %C3%A9 come from? A URL may contain only Latin letters, digits and a few symbols. Everything else is written as bytes in the UTF-8 encoding, and each byte as a percent sign and two hex digits. The letter é in UTF-8 is two bytes, C3 A9, hence %C3%A9. The browser shows you a tidy “Pokémon,” and the network carries the percent form.
You opened the link https://legost.in/en/cs/basics/web?from=reddit#dns. Which part of the address never reaches the server?
The browser keeps the fragment after # to itself: it is the address of a spot inside a page it already has, the section with id="dns". That’s why jumping around this chapter’s table of contents takes no requests to the server at all, and the server never learns which section you jumped to.
Specimen two: the name
The name legost.in is handy for people, but the packets of Chapter 41 travel by IP address, that is, by number. So the first thing the browser does is turn the name into a number. It looks in its own memory first, then asks the operating system, which checks a small text file. The course sandbox has that file too.
The file /etc/hosts holds pairs of address and name. The system found localhost in it and returned 127.0.0.1, “this same machine.” But legost.in isn’t in the file, and the sandbox gives up with the same error as at the end of the last chapter: it has no network access and nobody to ask. Your computer, in its place, would ask your internet provider’s name server. But first, why does this file exist at all?
DNS is built like a tree from Chapter 17. Read a name from right to left: legost.in is the node legost inside the node in, which sits inside the nameless root. Strictly speaking, every name ends in an invisible dot, the root: legost.in. A node of the tree together with everything below it is called a domain. Mockapetris proposed handing each subtree to whoever owns it. The root knows only who is responsible for in, com, uk and nearly fifteen hundred other top-level domains (1,437 in October 2026). The servers of the in zone, run by the Indian registry, know who is responsible for legost.in. And the servers of legost.in (for this site, the name servers of the hosting company DigitalOcean) know the address itself.
The dig utility with the +trace option walks this path itself, starting from the root. Below is its output, shortened, captured on October 2, 2026.
Each line is a record: a name, a lifetime in seconds, a type and a value. A record of type NS (name server) says “these servers are responsible for this subtree,” and an A record says “this name has this IPv4 address.” There are others: AAAA is an IPv6 address, CNAME says “this name is an alias for another one,” MX says where to deliver mail, and TXT holds arbitrary text, such as proof to a search engine that the site belongs to its owner. Walk the tree yourself.
Whoever walks the tree on your behalf is called a resolver. Usually it is your provider’s server or a public service like 1.1.1.1 or 8.8.8.8. The browser asks it one question, “what is the address of legost.in?” and waits for a finished answer, while the resolver does the rest: it asks the root, gets sent to in, asks there, gets sent to DigitalOcean, asks there. Three trips halfway around the world for one number. If it happened every time, the internet would choke on names alone.
The cache saves it, the same idea as in Chapter 34. Every record arrives with a lifetime, and until the lifetime runs out the resolver answers from memory. The referral to the in zone lives for two days (172,800 seconds), so the resolver rarely has to go all the way up to the root. The address of legost.in lives for ten minutes: the site’s owner chose a short lifetime so that when the site moves to another server, the world learns the new address quickly. There are caches on every floor: in the browser, in the operating system, at the resolver. That is why the “DNS” row of the timeline is zero for many readers: the name was already in memory. The flip side is the one every cache has: change the address, and for ten more minutes part of the world keeps going to the old one.
DNS is a distributed database with no owner. The root, the top-level zones and the sites’ zones are run by different organizations on different continents, and the speed comes from caches with lifetimes. There are only 13 root names, a through m, but behind them stand about two thousand servers around the world: one address is announced from many places at once, and a packet reaches the nearest copy.
DNS questions and answers are short UDP packets, with no connection: ask, get the answer, and if it got lost, ask again. They are simple enough to take apart by hand. Below are two answers caught on October 2, 2026: one from a root server and one from the server for the legost.in zone. A name in a packet is written as labels, a length byte followed by the letters: legost.in is 06 l e g o s t 02 i n 00. And to avoid repeating the same name ten times, a packet may point back to a name written earlier: the byte C0 and the position of the byte to read from.
The root didn’t answer the question about legost.in: the “answer” section is empty. Instead it gave a referral, four servers of the in zone, and in the “glue” section it attached their addresses right away, so that the resolver doesn’t have to look up the name servers’ own names by the same long road. The “authoritative” flag is off at the root and on at the DigitalOcean server: that server owns the zone, and its word is final. Its answer takes 43 bytes: a 12-byte header, the question, and one record with the address 9d e6 10 a2, that is, 157.230.16.162, and a lifetime of 600 seconds. The query id 11111 was made up by whoever asked, and the server sent it back so that the answer can be matched to its question: UDP guarantees nothing, and answers to different questions can arrive in any order.
The connection and the padlock
The address is found: 157.230.16.162. Next the browser opens a TCP connection to port 443, with the handshake from Chapter 42, one round trip of packets there and back. Then comes a second conversation on top of the first, in which the two sides agree on a cipher. This is TLS, and with it HTTP becomes HTTPS, the padlock next to the address bar. How two computers that have never met agree on a secret key while every router along the way hears every word they say is a big question, and Chapter 60 is given over to it. Here we care about the price: in the current version, TLS 1.3, the agreement costs one more round trip, and in the old one, 1.2, two. On the timeline these are the “TCP” and “TLS” rows. If the server is across an ocean and a round trip takes 150 milliseconds, half a second passes before the first byte of the page arrives: half a second in which not one byte of the page itself was sent.
Specimen three: the conversation
The connection is open, and the browser speaks its first sentence to the server in HTTP. In version 1.1, which still runs everywhere, it is plain text you can read with your own eyes. Browsers talking to big sites moved to HTTP/2 and HTTP/3 long ago: there the same messages are packed into binary frames, but they mean the same things, and tools always show them as text. Here is a conversation with this site’s server, captured with curl on October 2, 2026: the request and the response to it, without the body of the page that followed.
A request begins with a line of three words: the method, the path and the protocol version. The method HEAD means “send only the headers, not the page itself”; that’s how curl with the -I option peeks at a response without downloading it. Then come the headers, one per line: Host says which site is wanted (hundreds of sites live at one address, and this header is the only way the server tells them apart), User-Agent says who is asking, and Accept says what the asker is ready to accept. An empty line means “end of headers.”
The response mirrors the request. Its first line is the version and the status code: 200, all is well. Then the response headers. server gives away the server program: it is nginx, which accepts connections and passes the requests on to the site’s PHP program. content-type says that HTML in UTF-8 has arrived; without this header the browser wouldn’t know whether to show the bytes as a page or as a picture, or to offer to save them as a file. date is the time by the server’s clock. The last three headers are protection: the page may not be embedded in someone else’s site, its type may not be second-guessed, and for a year ahead this address may be visited only over HTTPS. We’ll come back to cache-control and set-cookie in a couple of sections.
Since HTTP is text, we can write a server for it ourselves, from the sockets we used in Chapter 42. The server waits for a connection, reads bytes up to the empty line, parses the first line and the headers, decides what to answer and writes the answer in the same format. The client in the same cell is a socket too, and it writes its request by hand. The server runs in a separate thread, as in Chapter 39: otherwise it would wait for the client while the client waited for it.
Under fifty lines, and it’s a working web server. Lines in the protocol end with the pair of characters \r\n, as on a teletype, which first returned the carriage and then fed the line. The length of the body goes in the Content-Length header, hence the comment next to body.encode(): the length must be in bytes. The body looks like plain English, but the curly quotes around “Hello” take three bytes each in UTF-8, so its 33 characters are 37 bytes. Get this wrong and the client either cuts the page off or waits forever for bytes that never come. The Connection: close header promises that the server will close the connection after the answer, so the client reads until recv returns nothing.
The server puts the request headers into a dictionary, as in Chapter 8, and lowercases the keys: HTTP header names ignore case, and Content-Length and content-length are the same header. Our parsing is still very trusting. A line without a colon will bring the server down, and it doesn’t read the request body at all. Careful parsing, with a body and with errors, is the chapter’s first task, “Parse a request.”
Methods and codes
GET and HEAD aren’t the only words a request can begin with. The method says what the client wants. GET fetches a document. POST sends data, such as a form, a comment or a program to run; the data goes in the request body after the empty line. PUT puts a document at an address, and DELETE deletes one. The difference is more than a formality. GET promises to change nothing on the server, so a browser repeats such a request without a second thought, search robots follow such links, and caches keep the responses. A browser won’t repeat a POST without asking you: what if it was paying for an order?
Status codes are organized by their first digit, and the first digit tells you the meaning even of a code you’ve never seen. Python keeps them all in one enumeration.
The 200s are success; 204 means “done, nothing to show.” The 300s are redirects: 301 means “moved for good, here is the new address” in the Location header, and 304 “you already have a fresh copy.” The 400s are the client’s fault: 400 “I can’t make sense of the request,” 401 “identify yourself,” 403 “you may not,” 404 “there’s no such thing,” 405 “not that way at this address,” 429 “too often.” The 500s are the server’s fault: 500 means the site’s program crashed, and 502 and 503 mean that a go-between like nginx couldn’t reach the program or the program is overloaded. Code 418, “I’m a teapot,” is an April Fools’ joke from 1998 (RFC 2324), but Python knows it, like plenty of other software. Telling the 400s from the 500s is the first thing anyone reading a server log does: a spike of 404s means broken links, a spike of 500s means a fire. That’s the chapter’s third task, “The server log.”
Now try speaking HTTP yourself. The request builder below gives you two servers to talk to. The first is a server in the course sandbox, like ours but with a few more addresses; you can send it any bytes at all, nonsense included, and see what it says. The second is this site’s server, and to it the browser lets through only polite GET and HEAD requests.
Host, write the method in lowercase, give the wrong body length. “This site”: the browser sends the request to legost.in and shows the status code and the response headers.The forgetful server and cookies
HTTP has a property that at first looks like a flaw: every request stands on its own. The server answers and forgets. To the server, the next request from the same computer comes from a stranger: it has nothing to recognize a returning visitor by. This is deliberate. A server that remembers nothing about its clients doesn’t care which of a thousand requests it serves first, and it can crash, restart and multiply without losing anyone’s conversation. But then how does a site remember that you’ve logged in, or what you put in your cart?
In June 1994 Lou Montulli of Netscape (then still called Mosaic Communications) was puzzling over this: the company was building an online store for a client who didn’t want to keep half-finished orders on its own servers. Montulli borrowed a trick Unix programmers had long known as a magic cookie, a scrap of data that a program receives and later hands back unchanged. The server puts a Set-Cookie: name=value header in its response; the browser remembers it and from then on attaches Cookie: name=value to every request to the same site. The Netscape browser learned to handle cookies in October of that year. Here is a server that counts visits, and two clients: one remembers nothing, the other keeps a cookie jar, as a browser does.
This server is no longer built from bare sockets. The standard library’s http.server module reads the first line and the headers by itself and calls the do_GET method, and ThreadingHTTPServer gives every visitor a thread of its own. The first client got a new number three times: the server didn’t know it was the same visitor. The second sent its number back in the Cookie header, and the server recognized it. Logging in to any site works this way: after the right password, the server issues a cookie with a long random session number and notes on its side whose number it is. Whoever knows the number is you, as far as the server is concerned, which is why a session number is guarded like a password.
The HttpOnly flag in our server forbids the page’s scripts to read this cookie: only the server sees it. In the legost.in headers above, the legostin_session cookie has that flag and XSRF-TOKEN doesn’t. The second cookie is left open to scripts on purpose: the page’s script attaches it to its own requests, and that lets the server tell a request from its own page from one forged by another site. Check it in your browser.
document.cookie. The session cookie isn’t on the list, though the browser stores and sends it: it has the HttpOnly flag.Don’t ask twice
The cache-control: no-cache, private header in the legost.in response is an instruction to the browser and to every go-between. private says the page is personal and must not be kept in shared caches between you and the server: your name might be in it. no-cache, despite its name, allows keeping a copy but forbids showing it without checking with the server. For pictures and programs that never change, sites write the opposite: max-age=31536000, “don’t ask for a year.” This is the browser cache from Chapter 34: on the timeline’s “Files” tab, many of this page’s files most likely never went to the network at all.
To find out whether a copy is still fresh without downloading it again, the browser sends a conditional request. The best place to show one is the world’s first website. Here is its response, captured on the same day, October 2, 2026.
The CERN site still answers in plain HTTP/1.1, unencrypted: it has HTTPS too, but doesn’t redirect you there. Its home page hasn’t changed since February 2014, as Last-Modified tells us. ETag is a version tag, a fingerprint of the content. A browser that already has a copy asks If-None-Match: "286-4f1aadb3105c0", “send it only if the tag is different.” If nothing has changed, the server answers 304 Not Modified with no body: two hundred bytes instead of the whole page. The request builder has a preset that asks the sandbox server this question.
A thousand visitors at once
The server in server.py serves visitors one at a time: while it reads a request and writes the answer, everyone else waits at accept. While an answer takes microseconds, nobody notices. But a working site, to answer, goes to its database, reads files and computes, and spends tens of milliseconds on a request during which the processor is mostly idle. Below, a loner, the one-at-a-time server from http.server, races a multithreaded one: six visitors arrive at once, and every answer “thinks” for 0.3 seconds.
The loner answers the six in almost two seconds: the last in line waited for five others. The multithreaded server takes a third of a second, as if there were only one. This is the scheduling problem of Chapter 37: while one request waits for the database, the processor goes to another. Servers on the internet do the same with threads, with processes or with an event loop: the nginx that answered us above holds tens of thousands of connections in a few processes, switching among them as asyncio does. And as soon as visitors are served at the same time, the races of Chapter 39 are back: the visits dictionary in the cookie cell is shared by all the server’s threads, and under serious load the visit counter ought to be guarded by a lock.
Specimen four: the body
After the headers and the empty line comes what all of this was for: the text of the page in HTML. It is ordinary text with markers called tags in angle brackets: <h1> opens a heading, </h1> closes it, <a href="…"> makes a link. Tags nest inside one another like brackets, which means the whole document is a tree. You can build it with the stack that checked brackets in Chapter 15: push each opening tag, pop on each closing one. The html.parser module knows how to cut text into tags and pieces of text.
The first tree is what the browser builds from any page and calls the DOM, the Document Object Model: the root html, inside it head and body, then headings, paragraphs and lists, with text on the leaves. Everything else works with this tree: styles, scripts, drawing.
The second tree came out wrong. In HTML a paragraph can’t sit inside a paragraph: the second <p> silently closes the first, and the <b> and <i> tags are closed crosswise. Our builder obediently nested one paragraph inside the other and complained about the brackets. Python would answer a program like that with a SyntaxError, as in Chapter 1. A browser never complains: the HTML standard spells out how to repair every mistake, and all browsers repair them the same way. That’s why a page from 1992 opens in a browser of 2026, though it contains tags that left the language long ago. In that copy, the window title and the housekeeping tag NEXTID are wrapped in HEADER, which by its meaning is what we write today as head. Today’s HTML has a header too, but it means a header strip inside the page, and the browser puts it in body. The widget below runs your own browser’s parser, so this is easy to check.
info.cern.ch page in its 1992 copy, and “This page” is the tree of the chapter you are reading. The “Keep the browser busy” button loads the page’s main thread with work for a second and a half.A tree isn’t a picture yet. The look is set by a separate language, CSS, Cascading Style Sheets. Håkon Wium Lie, who worked with Berners-Lee at CERN, proposed it on October 10, 1994, and the first version of the standard came out in December 1996. A CSS rule picks out nodes of the tree and gives them properties: h1 { color: navy } makes every top-level heading dark blue. Behavior is set by a third language, JavaScript. Brendan Eich, who joined Netscape in April 1995, wrote its first version in ten days; the language was briefly called Mocha, then LiveScript, before it got its name in December 1995, an echo of Java, the fashionable language of the day. Scripts read and change the DOM tree: every widget in this course is a script that adds to the chapter’s tree right there in your browser.
From here the browser works like an assembly line. It parses the HTML into a tree, parses the CSS into rules and computes the styles of every node. Then comes layout: how wide each rectangle is and where it stands. The width of a paragraph depends on the window, its height on how many lines came out, and the position of the next paragraph on the height of the one before. Then painting: rectangles become pixels. When the window or the tree changes, the line runs again for the part that changed. On the timeline these steps are the rows “HTML parsing” and “Scripts” and the marks “first text on screen” and “main content on screen.”
Nearly all of this is done by one thread, the page’s main thread. It works like the event loop of Chapter 37: it takes the next job from its queue (a tap, a server response, an animation frame, a piece of script) and runs it to the end. While a script is computing, the page responds to nothing: it doesn’t scroll, its buttons don’t press, the animation freezes. The “Keep the browser busy” button in the widget above shows this for a second and a half. That is why the course’s heavy computing happens on the sandbox server. And where a browser does have a lot of computing to do, the work is cut into pieces, and between the pieces control goes back to the loop.
A page no person reads
A server doesn’t have to answer with pages. Press “Run” under any cell in this chapter, and the browser sends the course’s server a request with no HTML in it at all. Here it is, slightly shortened, as your browser’s developer tools show it on the “Network” tab.
The request body is JSON from Chapter 8: a dictionary with the mode, the text of the program and the language. The response is JSON too, but it arrives line by line, as the program prints: the NDJSON format, one object per line. The page’s script reads the lines and adds the output under the cell. A set of addresses that programs send requests to and get data from is called an API, an application programming interface. Nearly the whole web of today works like this: a phone app, a map, a weather forecast are scripts and programs that talk to servers in HTTP and JSON. Here is an API of our own: the earthquake catalog from Chapter 6, answering questions asked by address.
From 2015 through 2025 the catalog holds eight quakes of magnitude 8 or more, and the strongest was off Kamchatka in July 2025. The server answered the question asked by the address with code 200 and JSON, and the malformed questions with codes 400 and 404, again with JSON that explains. That is good manners for an API: the HTTP codes tell a program what happened, and the body explains it to the person who will have to fix that program. Whoever calls the API writes three lines: the request, a check of the code, json.loads. On the command line, curl does the same: curl 'https://…/api/quakes?min=8'. The programs you write can talk this way to any service that has an API.
Reassembly: the timeline again
The specimens are dissected, and we can put them back together and see the timeline from the start of the chapter with new eyes. A redirect: the server answered 301 and sent the browser to a new address, a whole extra conversation. DNS: the resolver’s trip down the tree, or an instant answer from a cache. TCP and TLS: a round trip each. “Waiting”: one more round trip, plus the time the server’s program took to put the page together. HTML download: how many bytes, over what kind of link. Then the browser’s own work: the tree, styles, scripts, layout, painting. And behind it all, the “Files” tab: the page asked for dozens more files, and each one has a little timeline of its own.
curl can show the same breakdown: its -w option prints the moment each step finished, counted from the start. Here is the site’s home page, loaded from a home computer on October 2, 2026.
The name was found in three milliseconds: it was in a cache. The TCP handshake took 141 ms, which is one round trip to the server in Frankfurt and back, so the route here is a long one. TLS took another 180 ms. From the request to the first byte took 306 ms: a round trip plus about 160 ms of work by the site’s program. And the last 166 ms were the bytes themselves. Of eight hundred milliseconds, more than half went on bare round trips, in which not one byte of the page was transferred.
A page of HTML and 30 files of 50 KB each; a round trip takes 80 ms; the link is 50 Mbit/s; all the files go over one HTTP/1.1 connection, one after another. What speeds up loading more?
HTTP/2, several times over: with the narrow link and one connection taking turns, the page loads in about three seconds, with HTTP/2 in about 0.65 s, while the wide link without HTTP/2 would barely help. The arithmetic is in the cell below.
A rough model settles it. The HTML costs four round trips (DNS, TCP, TLS, the request) plus transfer time; after that, the files travel in one of four ways, the ways the web has gone through in its history. HTTP/1.0 opened a new connection for every file. HTTP/1.1 started keeping the connection open and sending requests over it one after another. Browsers open up to six such connections to one server at once. And HTTP/2 sends all the requests at once over a single connection.
The model is crude: it knows nothing about TCP slow start or about files that depend on one another. But its main conclusion holds, and for anyone who will ever build websites it is the main lesson of the chapter: load time is set by the number of round trips times the length of a round trip, far more than by the width of the link. We saw the same difference between latency and bandwidth in Chapter 41. The length of a round trip is bounded by the speed of light in fiber, and no money buys it down: from New York to London and back takes at least 56 milliseconds on any plan. So every way of speeding up the web is a way to make round trips fewer, shorter or lighter. Caches with lifetimes, in DNS and in the browser. One connection for many requests. Compression: this page arrived compressed (the response headers say content-encoding), and the timeline shows how many bytes that saved. The content delivery network from Chapter 34, which keeps a copy of a file in a nearby city. And fewer files, the ones a page can do without.
When a site is slow, look at the timeline first: the “Network” tab in your browser’s developer tools, or curl -w. If most of the time is “waiting,” the server’s program is slow. If DNS, TCP and TLS are long, the server is far away or nothing is cached. If the download is long, the files are heavy. If the network finished long ago and the page still isn’t ready, it’s the browser’s work: scripts and layout.
Tasks
All three tasks come from everyday work with the web: take apart someone else’s request, answer it, and read the server log when something has gone wrong. The tests for the second task start your server in the sandbox and talk to it over sockets, as a browser would.
Write parse_request(data). It takes the bytes of an HTTP/1.1 request and returns a dictionary with the keys method, path, version (strings), headers (a dictionary) and body (bytes). Header names are lowercase and values have no spaces at the ends; if a header repeats, its values are joined with ", ". The body is as many bytes as Content-Length says, and without that header there is no body: whatever came after the empty line belongs to the next request. If the request is broken (there is no empty line after the headers, the first line doesn’t have three parts, a header has no colon, Content-Length isn’t a non-negative whole number, or the body is shorter than promised), raise ValueError.
For example, parse_request(b"GET / HTTP/1.1\r\nHost: legost.in\r\n\r\n") returns {"method": "GET", "path": "/", "version": "HTTP/1.1", "headers": {"host": "legost.in"}, "body": b""}.
The starter is the parsing from the server.py cell, and it breaks at every step. split(b"\r\n\r\n") cuts up the body if the body has an empty line in it too; split(":") cuts Host: localhost:8000 into three pieces. Cut only at the first occurrence: data.partition(b"\r\n\r\n") and line.partition(":") return a triple of “before, separator, after,” and when there is no separator it comes back empty, which makes errors easy to catch.
Take the body’s length from the header, int(headers.get("content-length", "0")), and the body as the slice rest[:length]. Before calling int, check that the value is all digits: "10".isdigit() is true, while "-1".isdigit() and "ten".isdigit() are false. And compare the length of what is left with the promised one.
Don’t turn the body into a string: it may be a picture, and Content-Length counts bytes. Only the headers need decoding.
The first idea of the solution is partition: cut at the first separator, because everything after it is someone else’s data, where the separator may turn up any number of times. The second is that the body is defined by what was promised, not by what “arrived”: on one HTTP/1.1 connection requests follow one another, and only Content-Length says where this one ends and the next begins. Full-grown servers also understand Transfer-Encoding: chunked, a body in pieces with a length before each, for when the length isn’t known in advance, but that is beyond this task.
Write serve(listener), a web server on sockets. It receives a listening socket that is already open and accepts connections in an endless loop: on each one it reads one request, answers it and closes the connection. The addresses:
GET /: code 200, bodyHome — welcome!;GET /hello?name=Zoë: 200,Hello, Zoë!; without the parameter,Hello, stranger!(the parameter arrives percent-encoded, so decode it);POST /echo: 200, the request body unchanged, with the request’sContent-Type, orapplication/octet-streamif the request has none;- any other path: 404, body
Not found: path(the path without parameters); - a known path with the wrong method: 405 and an
Allowheader with the right method (GETorPOST); HEADwhereGETis allowed: the same headers as forGET, but no body;- a request that can’t be parsed: 400, and the server keeps running.
Text responses have Content-Type: text/plain; charset=utf-8, and every response has a correct Content-Length and Connection: close. The tests start your server in a separate thread and send it requests over a socket, as a browser does. The starter already reads a request and sends an answer, with one bug.
The starter’s bug is in respond: len(body) of “Home — welcome!” is 15 characters, but in UTF-8 it is 17 bytes, because the dash takes three. The client reads 15 bytes and the word breaks off: “Home — welcom”. Turn the body into bytes before you measure it. While you’re at it, let respond accept bytes too: the echo sends back something that isn’t text.
urlsplit(target) takes apart a path with parameters: .path is the path and .query the parameters. parse_qs(query) returns a dictionary of lists and decodes the percent form by itself: parse_qs("name=Zo%C3%AB") is {'name': ['Zoë']}.
The order of checks: first, is the path known (if not, 404); then, does the method fit (if not, 405 with Allow); then the answer itself. For HEAD, compute the body as for GET and put its length in the header, but don’t send it. For garbage, wrap read_request in try: a ValueError means answer 400. And close the connection in finally, whatever happens.
The ROUTES dictionary is a routing table, as in any web framework: path → allowed method. Frameworks like Flask or Django do the same, except that the table holds a function that prepares the answer instead of a method. Code 400 can’t be skipped: a server on the internet receives garbage all the time (scanners, broken clients, TLS handshakes sent to an unencrypted port), and a server that falls over on one such request is down for good. This server serves visitors one at a time; to keep a slow client from holding up the rest, handle runs in a separate thread for every connection, as in ThreadingHTTPServer.
nginx writes every request as a line in its log, in this format:
The client’s address, two dashes, the time in square brackets, the request in quotes, the status code, the size of the response, where the visitor came from, and their User-Agent, also in quotes. Write status_report(lines): given a list of log lines, return a dictionary with the keys "2xx", "3xx", "4xx", "5xx" (how many responses of each class; codes outside 200–599 aren’t counted) and "top404", up to three of the most frequent paths answered with exactly 404, as a list of (path, count) tuples by count descending, ties broken alphabetically. A path is taken without the parameters after ?. Skip lines that don’t look like log entries. Logs contain all sorts of things: a "-" request, the bytes of a TLS handshake instead of a request, digits and quotes in the User-Agent.
The starter takes the ninth word of the line. That works while the request has three words. The request "-" is a single word, and the code ends up somewhere else. Look for the code right after the closing quote of the request: the time in brackets, then the request in quotes, then the code.
A regular expression anchored to the start of the line is the most convenient: re.compile(r'^\S+ \S+ \S+ \[[^\]]*\] "([^"]*)" (\d{3}) '). \S+ is a word without spaces, [^\]]* is everything up to the closing bracket, [^"]* everything up to a quote. If it doesn’t match, the line isn’t from the log. Regular expressions are covered properly in Chapter 54; this pattern is all you need here.
For a 404, take the second word of the request and cut off everything after ?: path.split("?")[0]. The path counters are a dictionary, and sorting “by count descending, then alphabetically” takes the key lambda kv: (-kv[1], kv[0]), as in Chapter 10.
Spaces won’t do for parsing a log: the quoted fields can contain spaces themselves, and nobody promises how many words a request will have. A pattern anchored to the start of the line finds the code right after the request field, and digits in a User-Agent won’t fool it. On a live site, a report like this is the first thing people check after rolling out a new version: did the 500s go up, did new broken links appear? And the 404 list will almost always include /wp-login.php and /.env: robots comb the whole internet around the clock for forgotten passwords and vulnerable software.
What next
Count how many machines worked so that you could see this page. Your provider’s resolver. A root server, one of about two thousand copies. A server of the in zone. DigitalOcean’s name server. The site’s own server. The servers of the delivery networks that the fonts and libraries came from: open the “Files” tab of the timeline to see how many different names this page’s sources have. None of these machines is in charge. The root doesn’t know the site’s address, the site’s server knows nothing about the root, and if one copy of the root goes down, the packets quietly go to a neighbor.
But the chain has a weak link, in plain sight: in DNS, legost.in has a single address. Behind it is one machine in Frankfurt. If it goes down, there is no page, however many copies of the root are running. Big services don’t live like that: the data of your mail, your bank or a search engine sits on hundreds or thousands of machines, and machines break every day; in a large data center something, somewhere, is always breaking. While the copies are identical, that’s no trouble. But copies have to change, and the messages between them, as we know from Chapter 41, get lost.
Three copies of one account, four operations, and every message lost with a probability of one in five. In the end the copies diverge, and to the question “how much money is in the account?” the system has three different answers. Resending messages, as TCP does, isn’t enough: a copy may do worse than lose a message: it may die holding it, or fall silent for a long time, and nobody can tell whether it is dead or merely slow. There are thousands of servers, and they fail. How do they agree? How do you make a thousand unreliable machines behave like one reliable machine when none of them is in charge? The answer was found in 1989, and it was told as a parable about the parliament of a Greek island that keeps passing laws although its legislators keep wandering off to the market. That parable is the next chapter.