Lab notebook
How the Mathematics textbook works, and how we check every formula
Sixty-one chapters, from notches on a bone to Gödel’s theorems. Underneath the text run a mathematical engine of its own, a kit of live figures and seven layers of checks. Here is how it all works, and the problems that had to be solved before the textbook could be trusted: a single error costs it the reader’s trust in everything else.
- chapters in 10 parts, in two languages
- 61
- thousand words of English text, formulas not counted
- ≈ 350
- theorems, lemmas and corollaries in their own blocks
- 694
- proofs drawn as figures, step by step
- 313
- chapter widgets, 59 trainers, 24 solvers
- 306
- automated tests, all green
- 678
01
What we built
Mathematics, the Queen of the Sciences, is a free interactive textbook on this site. It starts with how people learned to count and goes all the way to what mathematicians work on today: measure theory, functional analysis, Gödel’s theorems, topology, chaos. The levels run from age eleven to the later years of university. For the university part, the syllabuses of MIT, Cambridge, Moscow State’s Faculty of Mechanics and Mathematics and the Independent University of Moscow served as the benchmark.
The chapters are strung on a single thread. Each one opens with a wall — a problem the tools so far can’t solve — and ends with a new wall that the next chapter breaks through. Seven big questions run through the whole course. Why does minus times minus make plus? Can you cut a disc into pieces and reassemble them into a square? Why is there no formula for the roots of a quintic? Are there more whole numbers or points on a line segment? How does the cipher that protects your bank work? Why can’t the weather be forecast a month ahead? Can a statement be true but unprovable?
And yet every chapter has a shape of its own. Prime numbers is a hunt, Root two is a detective story, Zero and minus is an argument with a seventeenth-century sceptic, The triangle is a lab of ten experiments, Counting is a museum tour. The formulas can be touched: tap a part of one and it lights up and explains itself. Proofs are drawn step by step, and you can drag a triangle’s vertices with a finger and watch the argument hold.
Map of the course
61 stations, 10 lines
Lines are the parts of the course, stations are chapters, interchanges are places where an idea from one part works in another. Every station is a link.
- INumbers
- IIAlgebra
- IIIGeometry
- IVCalculus
- VLinear algebra
- VIStructures
- VIIChance & data
- VIIIFoundations
- IXHorizons






02
Four floors
From the outside the textbook looks like pages of text and pictures. Inside it has four floors, each standing on the ones below. Break the engine and the solvers, the trainers and the widgets all start telling lies. That is why the checks described later start at the bottom.
-
4
Laravel · PHP 8.3
Platform
Addresses in two languages, the metro map, a glossary of 593 terms, the topic catalogue, a page with every trainer.
-
3
HTML + JS modules
Content
61 chapters, 24 solvers, 59 trainers. A chapter’s widgets load only on its own page.
-
2
Custom Elements · KaTeX
The parts kit
Annotated formula, proof-as-a-figure, answer box, quiz, trainer, graphs, geometry board, 3D.
-
1
≈ 22,000 lines of JS, zero dependencies
Engine
Exact arithmetic, algebra, equations, derivatives, integrals, limits, matrices, number theory — with steps in two languages.
Floor one: an engine that counts in fractions
A computer stores ordinary numbers in binary, with a limited number of digits. One tenth in binary is an endless fraction, like 1/3 in decimal, so it goes into memory slightly rounded. Ask almost any programming language what 0.1 + 0.2 is and it answers 0.30000000000000004.
Add two decimals
The computer’s usual arithmetic
0.30000000000000004
The textbook’s engine: fractions of whole numbers
3/10 = 0.3
And this is what memory actually holds instead of 0.1: 0.1000000000000000055511151231257827021181583404541015625.
On a graph an error like that is invisible. In a textbook it is dangerous: a trainer would compare the reader’s correct answer, 0.3, with 0.30000000000000004 and mark it wrong. So the engine keeps numbers as fractions whose numerator and denominator are whole numbers of any length (JavaScript has the BigInt type for this). A third stays a third, and 2100 keeps all thirty-one of its digits.
The engine treats formulas as trees. It can walk a tree by the rules, the way you would on paper, and it records each move as a step in two languages: which rule, what it was before, what it became and why that is allowed. Those steps become the worked solutions in the solvers and the explanations in the trainers.
The engine is forgiving about how people write: it understands “2x”, “sin 2x”, “2(x + 1)”, a decimal comma as well as a point, the symbols √ and π, and |x| for the absolute value. It solves equations and inequalities with their domains in mind, and systems of them; it takes derivatives, integrals (with a cascade of 18 methods) and limits; it works with matrices over the rationals, tests primality with Miller–Rabin and factors numbers with Pollard’s rho. All of that is about 22,000 lines of JavaScript without a single outside library.
Floor two: the parts kit
On top of the engine sits a kit of about fifteen parts that chapters are assembled from. The two most visible are the annotated formula and the proof drawn as a figure. The ones below are real, taken from chapter 17, The triangle. Tap a part of the formula; in the proof, press “Next” and drag the vertices.
Theorem The angle sum of a triangle
The angles of any triangle add up to 180°.
The trick is to gather all three angles at one vertex.
The other parts are simpler, but they lean on the engine too. The “Try it yourself” box checks an answer by meaning: it accepts both x + 1 and 1 + x, and if a fraction in lowest terms was asked for, 6/8 gets “Right, but the fraction can be reduced”. In quizzes a wrong option can carry its own explanation of why people pick it. Trainers serve an endless stream of problems at several levels, with hints and worked solutions. There are also graphs with draggable handles, a geometry board and 3D scenes.
Floor three: chapters, solvers, trainers
A chapter is an HTML file with a header: number, part, level, tags, and which chapters to read first. A solver first reads the problem the way a person would type it — “x^2 - 5x + 6 = 0”, “derivative of sin 2x”, “20% of 350” — and shows how it understood it: “Read as: …”. Then it solves step by step and explains why each step is legitimate. There are 24 solvers, from fractions and percentages to differential equations, and 59 trainers, each with its own problem generator, answer checker and worked solutions.
Floor four: the platform
At the top sits a Laravel module of the site. It turns internal links like ch:derivative#rules into addresses in the right language, builds a glossary from the definitions in the chapters (593 terms, each defined exactly once), builds the topic catalogue from tags and draws the metro map. The routes of the lines are laid out by hand on a 64 × 36 grid, and a program places the station labels: it knows the width of every letter in the font and tries positions until no label runs into another one or into a line.
03
The hard parts
On a diagram it all looks tidy. The trouble starts where mathematics meets a computer, a reader’s finger and two languages. Here are the problems that took the most effort, and how they were solved.
An answer in any form
People write maths however it comes out, and that’s fine. The engine has to understand “2x”, “sin 2x” and “2(x + 1)” alike. Then the subtleties begin.
- The comma. In Russian it is the decimal separator: 2,5. In English a comma separates thousands and list items. So in Russian mode roots are listed with semicolons, “x₁ = −2; x₂ = 3”, and in English with commas.
- School conventions. “1/2x” means (1/2)·x, as in an exercise book; “log x” is the common logarithm; and two numbers in a row, “2 3”, count as an error rather than something to guess at. By the Russian school rule a fractional power is defined only for x ≥ 0, so x2/3 = 4 has one root, 8, while ∛(x²) = 4 has two: ±8.
- A right answer in an unexpected shape. The checker compares answers by meaning, not by spelling: x + 1 and 1 + x are the same to it, and so are 0.5 and 1/2. If a fraction in lowest terms was asked for, 6/8 gets “Right, but the fraction can be reduced”. The base-conversion trainer recognises digits written in reverse order and explains that mistake separately, and forgives a Cyrillic “с” typed instead of a Latin “c” in a hexadecimal number.
Calculus: where it’s easy to lie
- Integrals. No single algorithm takes every integral in the textbooks. The engine tries 18 methods: the table, substitution, integration by parts, partial fractions, trigonometric substitutions and more. It differentiates every answer it finds back and compares it with the original function. It can get stuck on trifles: after the substitution t = tan(x/2) the integral of 1/(1 + cos x) tripped over a nested 1 in the denominator until the result of the substitution was brought to lowest terms.
- Integrals that barely converge. The tail of ∫0∞ sin x/x oscillates forever, and an ordinary numerical method doesn’t know where to stop. The engine cuts the tail at the zeros of the sine into “lobes”, adds them up, speeding up the convergence with Euler averaging, and gets 1.570796…, which is π/2. If the lobes don’t shrink, as with ∫0∞ sin x dx, the integral diverges and the engine says so. A singularity inside the interval is split off and handled on its own: ∫02 dx/√|x − 1| = 4.
- L’Hôpital’s rule. People love to apply it without looking. The limit of (x + sin x)/x as x → ∞ is 1, but the ratio of the derivatives, 1 + cos x, tends to nothing at all, and a naive program would declare that there is no limit. The engine applies the rule only when the limit of the ratio of derivatives exists.
- The edge of the domain. To the left of zero x·ln x has no values at all, and at first the engine couldn’t find the limit at zero. Now it notices that the function lives on one side of the point only, takes the one-sided limit, gets 0 and explains in a separate step why that’s allowed.
Numbers that don’t fit
- Millions of candidates. The rational roots of a polynomial are sought among fractions p/q, where p divides the constant term and q the leading coefficient. The number 735,134,400 has 1,344 divisors, so for 735,134,400x⁵ − x − 735,134,400 = 0 there are more than a million pairs p/q. Trying them all took 6.7 seconds. Now the engine first finds the roots approximately and checks exactly only the candidates near them: 3 milliseconds.
- Big numbers. An attempt to factor 2128 + 1 takes more than a second, and the page would freeze all that time. So numbers above 1014 are factored in a background thread of the browser. Primality is tested with Miller–Rabin: below 3.3·1024 its answer is exact, above that it is probabilistic, and the solver says so plainly.
Pictures versus floating point
- The galaxy that swallows the gap. In chapter 0 a rope made one metre longer is wrapped round the Earth, a tennis ball and the Milky Way, and the gap is about 16 cm every time. But the Galaxy’s radius is about 4.7·1020 m, and in floating point R + 0.16 m is exactly R. The widget measures heights from the surface; otherwise the Galaxy’s gap would simply vanish.
- The golden angle that must not become a fraction. The sunflower in chapter 0 labels which fraction of a turn the angle resembles: 90° gives 4 rays, 144° gives 5. The golden angle is approximated by fractions worse than any other number, and it has no rays. Yet it differs from 21/55 by only about 0.00015 of a turn, and with a generous tolerance the widget would draw 55 rays that aren’t there. The tolerance had to be narrowed to 1.5·10−6.
The golden angle, 137.508°
21/55 of a turn, 137.455°
- A proof you can drag. A proof figure is a function of how far each step has been drawn. So “Back”, jumping to any step and dragging vertices in the middle of an animation all work by themselves, with no special code for each case. On a phone a finger on the figure scrolls the page unless it lands on a vertex, and a vertex moves only after 6 pixels of travel; otherwise a stray touch would spoil the picture.
- Heavy pages. The triangle chapter has about fifty live elements, 27 of them proof figures. On a phone, opening the chapter used to create 37 drawing canvases at once; now it creates none, and each one appears only as it scrolls near the screen.
One book, two languages
- Sixty-one chapters as one book. The chapters are linked by the chain of walls, follow a shared table of notation and refer to one another. Each of the 593 terms is defined exactly once, in the chapter where it first appears. The textbook has 694 theorems, lemmas and corollaries, 26 axioms and 643 proofs: 313 as figures that build up step by step and 330 as step-by-step text where no honest picture exists. “We take this without proof” is allowed only for results that need tools beyond the chapter, such as Fermat’s Last Theorem. Then the text explains the idea of the proof, what it rests on and why it doesn’t fit.
- Two school traditions. A Russian textbook writes tg and ctg, an English one tan and cot. An interval is (a; b) in Russian and (a, b) in English. Even the general solution of sin x = a looks different: Russian schools use a single formula with (−1)n, English ones two families of roots. The engine writes each solution in the reader’s tradition.
- Solvers that understand English. Every example in the English articles is run through the solver itself. To make them pass, the solvers learned phrasings such as “sketch the curve” and “a 20% discount”, where the word after the number sets the direction.
04
Seven layers of checks
In aviation and medicine, safety is explained with the Swiss cheese model proposed by the British psychologist James Reason. Every defence is like a slice of cheese: it has holes. Stack several slices and the holes almost never line up into a tunnel, so an error gets stuck in one of the layers. The picture at the top of this page is about exactly that.
The textbook’s checks work the same way. We don’t trust any single one of them completely, but there are seven, and they look at an error from different sides. The main principle is a second opinion. Checking an answer the same way you got it is useless: the error repeats and confirms itself. So almost every check reaches the answer by a different road.
- 1Rules before the first line
- 2The engine checks itself
- 3Tests: a second opinion
- 4Generator versus checker
- 5A figure must not lie
- 6Scientific review
- 7Structure and language
Rules before the first line
The first defence works before a word is written. The chapter drafts were prepared with the help of language models, so the text gets no trust by default: the course has a 1,213-line plan that sets out each chapter’s wall, shape, formulas, widgets and checked facts, and rules for the author and the reviewer. The rules begin with a portrait of the reader:
They read voluntarily. If it gets boring, they close the tab. If they find an error, they stop trusting everything else.
from the chapter author’s brief (translated)
Further on, the rules say:
- recompute every numerical example with the engine or a separate script; round honestly: “≈ 1.414”, not “= 1.414”;
- you may simplify the presentation, never the statement: every condition of a theorem stays in place, be it a ≠ 0, continuity on an interval or coprimality;
- for every formula, say when it holds and how it breaks if a condition fails;
- call legends legends, quote only what can be checked, and flag what is disputed.
Every chapter comes with a list of the facts its author isn’t sure about. For the quadratic equations chapter that list included the exact number of problems on the Babylonian tablet BM 13901; the text keeps a careful “more than two dozen”. The story of Grothendieck naming 57 as a prime is told in chapter 3 as an anecdote, flagged “according to a well-known anecdote”. And about Hippasus, supposedly drowned for discovering irrationality, chapter 6 says plainly: “It’s a legend: there is no reliable evidence for it, and many historians consider it a later invention.”
The engine checks itself as it answers
Every answer the engine shows a reader it first checks itself, and by another method.
- It computes a derivative twice: with the tree of rules that produces the solution steps, and with a separate fast calculation. The results are compared, then checked against numbers at six points. The code states the rule in a comment: never show a result that disagrees with the independent computation.
- It differentiates an antiderivative back and compares it with the original function. If they disagree there is no answer, and the reader sees: “The candidate failed the check by differentiation, so no answer is shown.” Subtler still: the antiderivative must be defined everywhere the function is. So for 1/x the answer ln x is rejected; it has to be ln |x|.
- Roots are substituted into the original equation and checked against the domain. Extraneous roots don’t vanish silently: the solver shows where they came from and why they were thrown out.
- A limit is compared with a table of values of the function ever closer to the point. If the numbers and the answer disagree, the answer isn’t shown.
- When the engine can’t do something, it says so. The integral of ex² can’t be written in elementary functions, and the engine names the special function erfi instead of inventing an answer. A numerical root search marks what it finds with “≈” and never claims “no roots” just because it didn’t find any.

Tests: a second opinion on every answer
The textbook’s code is covered by 678 automated tests. Almost all of them follow one principle: the engine’s answer is compared with something computed in a completely different way.
There are hooligan tests too. They throw fifteen hundred random strings at the solvers, from empty to nonsense, and demand one thing: an answer or a polite error message, never a crash. Others make sure that simplifying twice changes nothing: once simplified, there is nothing left to simplify. On the night of the build that rule failed on about two per cent of random expressions. After the fix, not once in 51,000.
The most visual of these tricks is checking a derivative with a secant. Take two nearby points, x − h and x + h, and find the slope of the segment between them on the graph. The smaller h, the closer the slope gets to the derivative. Up to a point.
A magnifying glass for the checker
The engine, symbolically
f′(x) = …
f′(1.3) = 3.8808561912
- secant slope
- 7.5986568494
- error
- 3.7
The checker needs checking too. On the night of the build the engine’s own numerical derivative of this very function at 1.3 returned 6.17 instead of 3.88: its first step jumped across almost a whole wave. The method had even estimated its own error at 1.6, far too large to trust the answer. Now it notices such an estimate and starts over with a step 8, 64, … times smaller, keeping the best result.
Trainers: generator versus checker
A trainer has two halves: the generator invents problems and the checker marks answers. A bug in either half hurts the reader: either the problem is broken or a correct answer is marked wrong. So in the tests each of the 59 trainers produces 300 problems at every level, 55,800 problems per run. The checker must accept the generator’s own answer, and every formula must render without errors.
But the checker will always accept the generator’s own answer, even a wrong one: it compares the answer with itself. That is why a second look is needed as well: recomputing the problem independently from its statement. The shared tests do this for the fractions and linear equations trainers, among others, and chapter authors ran their trainers through such recomputations by the thousand. Try it yourself: here is the real fractions trainer from chapter 5.
900 problems in a second
The generator produces 300 problems at each of its three levels. Every one is checked twice: does the checker accept the generator’s answer, and does that answer match an exact recomputation of the statement?
- the checker accepted the generator’s answer
- —
- matched the independent recomputation
- —
Tap a cell to see its problem.
That is exactly how we found that about 6% of the linear equations trainer’s level 2 and 3 problems had the root 0, and that the common denominator could reach 1,638. Now the root is never zero and denominators stay at 20 or below. The answer checker itself had a bug that rejected correct answers: while splitting a list, “3+sqrt(2)” lost a bracket and turned into “3+sqrt(2”.
Pictures: a figure must not lie
A picture is harder to check than a number: it can be right in one position and lying in another.
- The parts that proof figures are drawn with are run through every moment of the animation in the tests, and NaN, “not a number”, must never appear. On top of that, chapter authors ran their figures through dozens of random handle positions at every step.
- Cut-and-rearrange proofs are guarded by a test of their own: the part that moves a piece of a figure keeps its lengths and angles at every moment of the animation. Otherwise the picture would “prove” things by a quiet stretch.
- A script photographs every page at 360 and 1280 pixels wide, in the light and the dark theme. It checks that the console is free of errors and that the page doesn’t scroll sideways, and it can press buttons and drag handles.
- The reviewer drags handles to the extremes. That is how we found the “all sums” strip in the coin problem for coins of 11 and 15: it stopped at 120, although the largest impossible sum is 11·15 − 11 − 15 = 139. And a division-with-remainder figure that pushed everything interesting off the edge when the dividend was negative.
Scientific review
A finished chapter is read by a scientific reviewer, with a brief of its own and a different role: “a nit-picking scientific editor with a mathematical education at the level of a good university”. Its job is to leave not a single error in the chapter.
The reviewer checks that every theorem has all its conditions. It recomputes every number and every “≈”, looks for holes in proofs, checks names, years and quotations. It drags widgets to their extremes and reads their computational code, and runs trainers through hundreds of problems. It fixes clear errors itself, leaves style alone, and lists its doubts in the report. A typical outcome for three chapters: no outright mathematical errors, eleven precision fixes, every number recomputed, four trainers passing an independent checker at 600 problems per level.
On top of the reviews comes a read-through: chapters 12 to 21 were read in full, the rest were fact-checked. A script pulled the dates, names and numbers out of the text, and each one was checked. The fresh events in the last chapter, about the frontier of mathematics in 2026, were checked against sources with a web search.
Structure and language
The last layer keeps the textbook in one piece. The structure test checks that every chapter has a cover and a “What’s next” section, that links lead to chapters and anchors that exist, that each term is defined exactly once, that every widget in the text exists in the code, that every trainer has a generator, and that the Russian and English editions use the same widgets and sections.
The Russian text is checked by a style script. It hunts for the stock phrases of machine-written Russian and for bureaucratese, the local cousins of “it is worth noting”, “let’s dive in”, “not just X but Y”, “plays a key role” and “the amazing world of”, and it points at every hit. Above it stands a rule from the brief: read the paragraph aloud; if people don’t talk like that, rewrite it.
Translation into English turned out to be another check: a translator reads every sentence more slowly than any editor. That is how inaccuracies turned up in the Russian original. For instance, Terence Tao’s result on the Collatz conjecture holds for “almost all” numbers in the sense of logarithmic density, and that qualification was added to the Russian text.
05
The file of caught errors
Here is what the checks caught in the first day. Not all of it, just the most telling. Struck through is how it was, below is how it is now. The text quotes are translated from the Russian edition.
Text · chapter 3, Prime numbers
They can’t be any closer: of two neighbouring numbers, one is even.
…of two neighbouring numbers one is even, and the only even prime is 2, so the only neighbours are 2 and 3.
This is about twin primes: 3 and 5, 11 and 13, 41 and 43.
Layer 6 · scientific reviewer
Text · chapter 3
With twenty digits in the numerator and denominator, factoring becomes slow work even for a computer.
For twenty-digit numbers trial division stretches to billions of divisions, and numbers several hundred digits long can’t be factored in reasonable time even by a computer.
The textbook’s own engine factors the product of two 13-digit primes in about 0.3 seconds.
Layer 6 · scientific reviewer
Text · chapter 3
The ratio of π(x) to x/ln x tends to one, slowly but steadily.
Slowly and not monotonically: from 102 to 103 it even goes up.
The text contradicted a table in the same chapter: 25/21.7 ≈ 1.151, while 168/144.8 ≈ 1.161.
Layer 6 · scientific reviewer
Text · chapter 5, Fractions
How to find the shortest decomposition is, in general, unknown.
The shortest decomposition can be found by exhaustive search, but no fast method for it is known.
“Unknown how” and “unknown how to do it fast” are different claims.
Layer 6 · scientific reviewer
Engine · numerical integral
∫1∞ dx/x = 37.29, error estimate 0.07.
The integral diverges, and now the engine says so: it recognises a slow tail, a pole inside the interval and oscillations that never die down.
Layer 3 · while building the calculus engine
Engine · numerical derivative
(sin(cos(tan x)))′ at 1.3 = 6.17.
3.8808…: when its error estimate is large, the method starts over with a smaller step.
Layer 6 · engine review
Engine · limits
lim x·ln x as x → 0: no answer.
0, with a step of its own explaining why the limit is taken from the right: to the left of zero the logarithm is undefined.
Layer 6 · solver review
Engine · simplification
(1/|x|)·(x/|x|) = x/x²
= 1/x. Simplifying twice changes nothing more: 0 failures in 51,000 random expressions instead of about 2%.
Layer 3 · hooligan test
Picture · sieve of Eratosthenes
Tap 7: “7 is prime: none of 2, 3, 5, 7 divides it.”
“7 is prime: the sieve reached it and nothing had crossed it out.”
Layer 6 · scientific reviewer
Picture · Frobenius coins
The “all sums” strip goes up to 120.
The strip goes up to 160: for coins of 11 and 15 the largest impossible sum is 139, and it has to be visible.
Layer 6 · scientific reviewer
Trainer · linear equations
≈ 6% of problems with root 0, common denominators up to 1,638.
The root is never zero, denominators stay at 20 or below. A run of 9,000 problems finds no degenerate ones.
Layer 4 · independent run
Trainer · answer checker
“3+sqrt(2)” → “3+sqrt(2” → wrong.
A list of answers is split with the bracket depth in mind, and “(1 ± √5)/2” reads correctly too.
Layer 4 · trainer check
06
What the checks don’t promise
Seven layers don’t make a textbook error-free. They make an error unlikely, and every layer has known holes.
- A check at points is not a proof. The engine decides exactly whether two expressions are equivalent only for polynomials. For everything else it compares values at 24 points to within 10−9. Agreement at 24 points almost certainly means an identity, but the “almost” stays.
- Some pretty answers the engine recognises by their numbers. It notices that arctan(2 − √3) = π/12 because the values agree to ten decimal places, not because it derived it symbolically. It does check the other way round: the tangent of the angle it found must match 2 − √3 as well. Without that, arctan 1010 would be “recognised” as π/2.
- Tests catch only what we thought to ask about. An error nobody imagined will slip past every test.
- History is the weakest link. You can’t recompute a date in node. Hence the rule: only checkable quotations, legends called legends, disputes flagged.
That is why the most valuable layer is the eighth. Found an error? Press “Feedback” on the course page.
Where to start
Chapter 0: five short stories
A cat under a rope, a sunflower that can count, a dragon made of a paper strip, an ant that builds a highway, and forty primes in a row. Each comes with a live picture you can play with.
“Mathematics is the queen of the sciences, and arithmetic the queen of mathematics.”
The demos on this page run on the same engine and the same parts as the textbook. Under the hood: Laravel, PHP 8.3, Vite, JavaScript web components, KaTeX 0.16, node:test, Pest and Playwright.