

areas of knowledge - Maths
Methods and tools of Maths
How does mathematical knowledge get produced - and what does the method guarantee?
The proof that changed what proof means
THE PROVOCATION
The Scope page established that mathematics produces knowledge of a distinctive kind: certain, provable, and independent of any particular observer. The Perspectives page asked where perspective enters that picture despite the certainty. This page asks something more fundamental: how does mathematical knowledge get produced at all, and how secure is the method that produces it?
The Four Colour Theorem
The video works through the four-colour problem from first principles. Starting from a map-maker's practical task of colouring countries so no two adjacent ones share a colour, it builds to the theorem through a sequence of examples, including one map that appears to need five colors until a smarter strategy cuts it back to four. The video also notes the theorem's restrictions and ends with two open challenges: how many colours are needed for a globe, and how many for a doughnut?
In October 1976, two mathematicians at the University of Illinois announced that they had proved the four-colour theorem: that any map, however complex, can be coloured using only four colours so that no two adjacent regions share a colour. Cartographers had known this empirically - simply by making maps - for a long time. The theorem had been conjectured in 1852, when a student named Francis Guthrie noticed while colouring a map of England that four colours were always enough. In the 124 years between the conjecture and the proof, no mathematician had ever found a map that required five.
Kenneth Appel and Wolfgang Haken proved it. Their proof was valid - mathematicians checked it and found no error. But the proof was unlike any that had appeared before in mathematics. To establish that four colours were sufficient, Appel and Haken had to show that every possible map could be reduced to one of 1,936 configurations, and that four colours were sufficient for each. A human mathematician checking each configuration in turn, working at full pace, would take years. They used a computer. The computation ran for more than a thousand hours. No human being has ever verified the four-color theorem step by step.
The proof stands. The logic of the reduction is sound, the program has been examined, and the result has been confirmed by later independent proofs using similar methods. But the core of it - the verification of 1,936 cases - exists as a computer output that no mind has worked through from end to end.
For most of the history of mathematics, a proof was something a mathematician could follow: a chain of deductive steps, each warranted by the previous one, that any trained reader could trace to the conclusion. The four-colour proof is not that. It raises a question that cuts to the foundations of mathematical method: what exactly is a proof, and what does having one guarantee?
Big idea 1 - Every proof rests on assumptions it cannot justify.
In your Maths class you use the theorem that the angles of a triangle sum to 180°. That theorem follows necessarily from Euclid’s fifth postulate. This Big Idea explains what that means, and what happens to triangle geometry when the postulate changes.
Around 300 BCE, Euclid of Alexandria organised the known results of Greek geometry into a system built entirely on deductive reasoning. Starting from five postulates – basic assumptions stated clearly at the outset – he derived over 400 theorems by logic alone. Each result followed necessarily from previous ones, and all of them from the same five starting points. The method was the innovation as much as the mathematics.
Four of the postulates are simple and immediate: a straight line can be drawn between any two points; all right angles are equal; a finite line can be extended; a circle can be drawn from any centre with any radius. The fifth is different in character. It states, roughly, that if a straight line crosses two other lines and the interior angles on one side sum to less than 180 degrees, those two lines will eventually meet on that side. In its most familiar form: through any point not on a given line, exactly one parallel line can be drawn.
For two thousand years, mathematicians were uneasy with the fifth postulate. It is noticeably more complex than the other four and lacks their immediate self-evidence. Many attempted to derive it from the remaining postulates – to show that it followed necessarily from them. None succeeded.
In the early nineteenth century, the suspicion became a discovery. The Hungarian mathematician János Bolyai and the Russian mathematician Nikolai Lobachevsky independently showed that you could replace the fifth postulate – allowing more than one parallel line through a given point – and derive a complete, consistent geometry. No contradiction emerged. The logic held throughout. Bernhard Riemann showed the same was true in the opposite direction: allow no parallel lines at all, and you get a different but equally rigorous system.
These were not mathematical curiosities. They were complete geometries, as internally valid as Euclid’s. What they demonstrated was something precise: deductive proof guarantees that conclusions follow from starting assumptions. It cannot, and does not, guarantee that the starting assumptions describe anything real.
The consequences of replacing that one postulate are visible in a direct comparison. Each of the three geometries below is internally consistent – the rules of logic apply throughout. What changes is only the starting assumption about parallel lines.
When Einstein needed a geometry to describe the curvature of spacetime near a massive object, he reached for Riemann’s. His general theory of relativity (1915) required a geometry in which space curves around mass. Euclid’s version describes flat space accurately and remains the foundation of everyday engineering and architecture – it is the geometry of the world at ordinary scales, where curvature is negligible.
What the non-Euclidean revolution established is a distinction that runs through this course: the difference between validity and truth. A proof is valid when its conclusions follow necessarily from its premises. It is sound when the premises are also true. Euclid’s proofs were always valid. Whether they were sound depended on something no proof could settle – whether the universe actually works the way the axioms say it does. That turned out to be an empirical question, not a mathematical one.
The IB maths syllabus states that ‘different statistical techniques require justification and the identification of their limitations and validity.’ The validity/soundness distinction made here gives you precise vocabulary for exactly that: a technique can be valid (correctly applied) without its results being true, if the underlying model assumptions do not hold.
The four-colour proof returns here with a different force. Appel and Haken’s proof is valid: the computer checked every configuration correctly. Whether it is a proof in the traditional sense – whether validity achieved by machine carries the same epistemic weight as validity verified by a human mathematician – is the question the provocation raised. What Euclid’s axiom story shows is that the question of certainty in mathematics has always been more complicated than it appeared.
Big idea 2 - Whose mathematics counts?
By 1900, the mathematician David Hilbert had identified what he believed was the central task of twentieth-century mathematics: to place the entire subject on a rigorous formal foundation. Every branch of mathematics - arithmetic, geometry, analysis - should be derived from a finite set of axioms using explicit rules of inference. The system should be complete (every true statement should be provable within it), consistent (no statement and its negation could both be proved), and decidable (there should be a mechanical procedure for determining whether any statement is provable). Hilbert's programme was an attempt to make the certainty that mathematics appeared to offer into something that could be verified rather than assumed.
For three decades, some of the best mathematicians in the world worked toward this goal. Then, in 1931, a twenty-five-year-old Austrian mathematician named Kurt Gödel published a paper that ended it.
Gödel's Incompleteness Theorem - du Sautoy
The video opens with the liar's paradox and uses it to show how Gödel translated mathematical statements into numbers - allowing mathematics to refer to itself. The key moment is the construction of a statement that asserts its own unprovability. Pay attention to what follows from that construction: if the statement is false, the system is inconsistent; if it is true, it is unprovable. That is the incompleteness theorem in its essential form, and it is exactly the result that ended Hilbert's program.
By 1900, the mathematician David Hilbert had identified what he believed was the central task of twentieth-century mathematics: to place the entire subject on a rigorous formal foundation. Every branch of mathematics - arithmetic, geometry, analysis - should be derived from a finite set of axioms using explicit rules of inference. The system should be complete (every true statement should be provable within it), consistent (no statement and its negation could both be proved), and decidable (there should be a mechanical procedure for determining whether any statement is provable). Hilbert's programme was an attempt to make the certainty that mathematics appeared to offer into something that could be verified rather than assumed.
For three decades, some of the best mathematicians in the world worked toward this goal. Then, in 1931, a twenty-five-year-old Austrian mathematician named Kurt Gödel published a paper that ended it.
What Gödel showed was precise: in any consistent formal system powerful enough to express basic arithmetic, there exist true statements the system cannot prove. The proof works by encoding mathematical statements as numbers - a technique now called Gödel numbering - then constructing a statement that says, in effect, "this statement cannot be proved in this system." If the system is consistent, the statement must be true. If it is true, it cannot be proved. Adding new axioms to close the gap does not help: the extended system will contain its own unprovable truths. Incompleteness is a structural feature of any formal system of sufficient power, not a problem that better axiom choices can eliminate.
The consequences for Hilbert's programme were direct. The completeness requirement failed by Gödel's first theorem. The consistency requirement failed by his second: no sufficiently powerful formal system can prove its own consistency using only the resources available within that system.
The result leaves the ordinary practice of mathematics intact. Mathematicians continue to prove theorems, and those proofs remain valid. The unprovable statements Gödel constructed are specific and somewhat artificial. But the result does clarify what proof can and cannot guarantee. A proof establishes that a conclusion follows from the axioms of a given system. It cannot establish that the system itself captures all mathematical truth. The four-colour theorem is provable within standard mathematics. What Gödel's theorem shows is that no set of axioms, however carefully chosen, can in principle be made to contain every truth.

Big idea 3 - The answer arrives before the proof.
Mathematics is defined by what it can prove. But mathematical discovery, the moment when a new result becomes visible, rarely begins with a proof. Henri Poincaré, one of the most important mathematicians of the late nineteenth century, wrote about this with unusual candour in a 1908 lecture called "Mathematical Creation." He had been working for weeks on a problem in complex analysis without progress. Then, stepping onto a bus at Caen while thinking about something else entirely, the solution appeared in his mind complete. He did not work it out on the bus. He returned to his desk and wrote it up. The verification was almost mechanical.
Poincaré argued that mathematical creation involves two distinct processes: unconscious work generates and combines mathematical ideas in ways that cannot be directly observed, while conscious verification checks the results against the standards of formal proof. The architecture he describes from introspection is what Kahneman maps psychologically as System 1 and System 2, and what Eagleman's neuroscience explains as the brain completing its work before consciousness receives the summary. (See core lesson 6) What mathematicians call discovery is the moment when unconscious work surfaces. The creative act is largely invisible; the proof is its residue.
Poincaré is not an isolated case. James Clerk Maxwell, whose equations unified electricity and magnetism in 1862, made the same admission on his deathbed. As Eagleman records in Incognito, Maxwell declared that "something within him" had discovered the equations; he had no idea how the ideas had come to him. They simply had. What neuroscience describes as the brain running its processing outside conscious access is what mathematicians, viewed from the outside, tend to call genius.
The same pattern appears more starkly in the case of Srinivasa Ramanujan. Ramanujan grew up in southern India in the early twentieth century with almost no formal mathematical training beyond a single textbook. He filled notebooks with extraordinary results in number theory, infinite series, and continued fractions. When he wrote to the Cambridge mathematician G.H. Hardy in 1913, the letter contained over a hundred results - some already known to Cambridge mathematicians, some apparently new, and some so unusual that Hardy could not determine whether they were true at all. Hardy recognized an extraordinary mathematical mind behind all three categories, and invited Ramanujan to Cambridge. Ramanujan arrived in Cambridge in 1914 and collaborated with Hardy until his death in 1920 at the age of thirty-two. He attributed his mathematical knowledge to his family goddess, who he said revealed results to him in dreams.
1729 and Taxi Cabs - Numberphile
The video tells the hospital story directly: Hardy arrives in a taxi numbered 1729, apologises to the dying Ramanujan that it seems a dull number, and Ramanujan identifies it as the smallest number expressible as the sum of two cubes in two different ways. Pay attention to Grime's footnote near the end. Hardy's own reading of the story was that Ramanujan had already encountered 1729 in his research and was retrieving something he knew. The knowledge was there before the question arose.
Whether Ramanujan's explanation of his process is taken literally or not, what it points to is real: his route to mathematical truth was entirely unlike formal deduction. The results came first. The proofs, where they exist at all, came from others, sometimes generations later. His case makes visible a distinction that Poincaré's account implies but does not state so sharply: mathematical knowledge and the formal proof of that knowledge are not the same thing. The knowledge can be present, in some meaningful sense, before the proof exists to certify it.
This is not unusual in the history of mathematics. Newton and Leibniz developed calculus in the seventeenth century and used it to produce results that transformed physics and geometry. But their underlying notion of infinitesimals, quantities that are somehow both zero and non-zero, was intuitive rather than rigorous. George Berkeley famously mocked "the ghosts of departed quantities." It took nearly two hundred years for Cauchy and Weierstrass to supply the epsilon-delta definition of a limit that placed the whole edifice on solid ground. The concept of a limit you encounter in calculus is the answer to a question that Newton was solving problems with long before anyone had stated the question precisely.
George Pólya's work approached the same gap from a different direction. His 1945 book "How to Solve It" set out to codify the strategies mathematicians actually use when working on a problem: guess and check, work backwards, find a simpler related problem, consider an extreme case. Pólya called these heuristics: productive ways of navigating mathematical space before the destination is known, not logical steps that guarantee progress. His claim was that mathematical discovery is a learnable skill, and what you learn is not more logic but better instinct about where logic might usefully apply. When the IB asks you to select and justify a mathematical method in your exploration, that initial act of selection, before you know whether the method will work, is exactly the kind of reasoning Pólya was trying to describe.
The standard of formal proof survives these accounts intact. None of them suggest that rigorous proof is dispensable, only that it tends to follow discovery rather than precede it. Gödel's incompleteness theorem, which required the insight that mathematical statements could encode statements about themselves, was itself an act of exactly this kind of imagination. The formal proof followed an insight that no formal procedure could have generated. Where mathematical knowledge comes from, before proof exists to certify it, remains an open question.

The Methods and Tools of Maths compared
The method that closes arguments cannot choose where they begin.
Mathematics is often described as the only area of knowledge where correct method guarantees correct conclusions. A valid deductive proof does not merely make a conclusion probable or well-supported; it makes it necessary. If the axioms hold and the logic is followed correctly, the conclusion cannot be otherwise. This is what distinguishes mathematical knowledge from knowledge produced by any other method - or so it appears.
The comparison with natural science clarifies what this means in practice. The natural sciences generate knowledge through experiment and observation: evidence accumulates, hypotheses are tested, theories are revised. The method is designed to converge on better and better approximations while remaining open to revision if new evidence demands it. Even the most successful scientific theory, confirmed by thousands of experiments, remains in principle defeasible. Mathematical proof works differently. The Pythagorean theorem was not confirmed by measuring triangles in multiple environments; it was proved from axioms by logical deduction, and no measurement could refute it. The method produces a different kind of conclusion.
What the natural sciences have that mathematics lacks is a direct connection to the world. The non-Euclidean geometries explored earlier in this lesson show that which geometry governs spacetime is not determined by mathematics - it is an empirical question, answered by observation. Mathematics can describe any consistent set of axioms; which description applies to the actual universe requires the scientist's method, not the mathematician's. The two methods are complementary for this reason: mathematics guarantees logical necessity, natural science guarantees empirical contact.
History presents a different contrast. The historian works from documents, artefacts, and material remains toward conclusions about what happened and why. The method produces probable rather than necessary conclusions - the best available explanation, which remains open to revision as new evidence emerges or interpretive frameworks change. Mathematical proof aims to close questions permanently; historical argument keeps them open in principle. But the comparison cuts the other way too. The historian's evidence, however fragmentary, connects directly to what actually occurred. The mathematician's proof, however airtight, only establishes what follows from the chosen axioms. Which axioms to start from - which geometry to adopt, which set-theoretic foundations to accept - is not itself settled by proof. It involves something closer to the historian's judgement: a decision about which starting assumptions best serve the inquiry at hand.
The Arts suggest a more unexpected parallel. Poincaré's account of mathematical discovery - the unconscious work, the sudden illumination, the verification that follows - describes a creative process that artists and composers have recognised in their own practice. Maxwell's deathbed admission that something within him had discovered his equations is structurally the same as the experience Coleridge described in writing Kubla Khan. The moment of discovery, in mathematics as in art, often arrives before deliberate reasoning has caught up with it. What differs radically is what follows. In the arts, the result of creative work is evaluated by audiences, critics, and communities of practice - by aesthetic response, cultural significance, and enduring resonance. In mathematics, it is evaluated by proof: a standard that, correctly applied, settles the question definitively. The creative sources of mathematical and artistic knowledge may be similar. The methods for validating them are not.
Human Science’s methods depend on something proof does not have to worry about: a subject who notices. Methods and Tools in Human Science argues that a person being studied can change their behaviour simply because they know they are being studied, which has no equivalent in mathematics. A theorem does not behave differently once it learns a mathematician is checking it.
What this comparison shows is that mathematical method is distinctive in a specific and limited sense. Proof does guarantee conclusions from axioms. But this lesson has shown three qualifications on that guarantee. The axioms themselves are not guaranteed: they are chosen, and different choices produce different geometries, all equally valid. The proof cannot establish the consistency of the system it operates within: Gödel showed that any sufficiently powerful formal system contains truths it cannot prove from the inside. And the process by which mathematicians find what to prove involves intuition, unconscious work, and heuristics that are closer to artistic imagination than to formal deduction. The guarantee that mathematical method offers is real. What it covers is narrower than it first appears.
Next: Ethics in Maths
Think further: questions and resources
-
The four-colour theorem was proved using a computer to check 1,936 configurations - a process no human being has verified in full. Mathematicians eventually accepted it, but not without resistance. Does a proof that cannot be humanly read in its entirety carry the same epistemic weight as one that can?
-
Non-Euclidean geometry shows that mutually contradictory systems can each be internally consistent and logically valid. Einstein found that Riemannian geometry describes actual spacetime; Euclidean geometry remains the foundation of everyday engineering. Does this mean mathematics discovers truths about the world, or invents structures that may or may not turn out to apply to it?
-
Gödel's theorem is a result about formal systems - precisely defined structures of axioms and inference rules. Mathematics as actually practised by working mathematicians is not quite the same thing. Does incompleteness affect the knowledge that mathematicians produce from day to day, or only the foundations they rarely examine?
-
Hilbert's programme aimed to place all of mathematics on a foundation that was provably consistent and complete. Its failure does not mean mathematics stopped working - theorems continued to be proved and applied. Does the failure of Hilbert's programme make mathematics less certain than it appeared, or does it only require us to revise our account of what mathematical certainty actually is?
-
Pólya argues that mathematical discovery is a learnable skill - that the heuristics experts deploy can be made explicit and taught. Ramanujan arrived at results through a process he could not explain and that others could not replicate. What does the difference between these two cases suggest about the extent to which mathematical creativity can be taught, or whether it involves something that teaching cannot reach?
-
The axioms a mathematician starts from are chosen rather than derived - different choices produce different but equally valid geometries. Yet mathematicians do not treat axiom choice as arbitrary: some systems are preferred for their elegance, fertility, or applicability. On what grounds is one set of axioms chosen over another, and who has the authority to make that choice?
Films
For more see my 10 films for the TOK journey page.
🎬 WATCH — Gödel's Incompleteness Theorem (2017)
Du Sautoy opens with what drew him to Gödel personally: the unsettling possibility that a conjecture he might spend his career on simply has no proof. The video covers the self-referential mechanism but its real value is the final section, which asks whether Goldbach's conjecture or the Riemann hypothesis might be true but unprovable within any axiomatic system we can construct. That is where incompleteness stops being a logical curiosity and becomes a live worry for practising mathematicians.
🎬 WATCH — The story of maths (1996)
BBC Open University
Du Sautoy's four-part history of mathematics concludes with the twentieth century's most unsettling discoveries: Cantor on infinite sets, Gödel on incompleteness, and Paul Cohen's proof that some questions in mathematics can be neither proved nor disproved within standard axioms - leaving open the possibility of genuinely different, conflicting mathematics. Episode 4 is the most directly relevant to this page, though the series rewards watching from the beginning. Du Sautoy is already the voice behind the Big Idea 2 video, which makes this a natural extension. The question the episode leaves: if mathematics can branch into conflicting systems, in what sense is there one mathematics?
Further reading
📚 READ - A Mathematician's Apology by G.H. Hardy (1940)
Hardy's short book is part defence of mathematics, part elegy for a mathematician who knows his best work is behind him. The sections on proof as method - and on why beauty is not a luxury but a diagnostic tool for whether an argument is right - sit directly behind the ideas on this page. The reductio ad absurdum passage alone is worth the price.
📚 READ - Gödel, Escher, Bach: An Eternal Golden Braid by Douglas Hofstadter (1979)
Long, demanding, and like nothing else. Hofstadter uses the interplay of Bach's fugues, Escher's impossible staircases, and Gödel's incompleteness theorems to argue that self-reference is fundamental to mind, meaning, and mathematics. Chapter I is the place to start; the rest rewards patience.
📚 READ - Fermat's Last Theorem by Simon Singh (1997)
The story of a problem that took 350 years to prove - and of what it cost Andrew Wiles to prove it. Singh makes accessible not just the mathematics but the question this page keeps returning to: what does it actually mean to demonstrate that something must be true?
📚 READ - Proof: The Uncertain Science of Certainty by Adam Kucharski (2024)
Kucharski traces what "proof" means across mathematics, law, medicine, and computing. The opening chapter, on Lincoln teaching himself to argue by working through all of Euclid's Elements, is a striking reminder that the axiomatic method was never just a mathematical tool - it was a model for how to build any claim that would hold.
📚 READ - Incognito: The Secret Lives of the Brain by David Eagleman (2011)
Eagleman's central argument - that "most of what we do and think and feel is not under our conscious control" - arrives here with specific force. The sections on creativity and unconscious processing read as neuroscience for what Poincaré was describing from the inside: the conscious mind as late-arriving editor of a process it did not initiate.
📚 READ - How Not to Be Wrong: The Power of Mathematical Thinking by Jordan Ellenberg (2014)
Ellenberg's argument is that mathematical thinking is rigorous common sense - and that the failure to apply it explains most of the errors in public reasoning about statistics, probability, and risk. A good companion to the methods question: not just how mathematicians justify claims, but what happens when non-mathematicians try.