Web Analytics
top of page
areas of knowledge - Maths

Ethics in Maths

Can a proof be innocent?

THE PROVOCATION

A harmless and innocent occupation?

In 1940, G.H. Hardy published A Mathematician's Apology - the book whose arguments appear on the Scope and Perspectives pages of this site. Near its end, he addressed the question that had become unavoidable: did mathematics bear any responsibility for modern war? His answer was direct. "Real mathematics has no effects on war. No one has yet discovered any warlike purpose to be served by the theory of numbers or relativity, and it seems very unlikely that anyone will do so for many years." A real mathematician, Hardy concluded, has "his conscience clear." Mathematics was "a harmless and innocent occupation." He wrote this in 1940. The Manhattan Project began the following year.

Oppenheimer: "I am become death"

J. Robert Oppenheimer, scientific director of the Manhattan Project, is interviewed by NBC in 1965, two years before his death. He describes watching the detonation of the first atomic bomb at the Trinity test site in New Mexico on 16 July 1945. He recalls a line from the Bhagavad Gita that passed through his mind at that moment: "Now I am become Death, the destroyer of worlds." He says he supposed they all thought that, in one way or another.

The bomb required mathematics: isotope separation, neutron cross-sections, implosion lens geometry. Hardy's beloved number theory, which he had celebrated as gloriously impractical, now underlies RSA encryption - the system securing military communications and every financial transaction on the internet. 

Hardy claimed his work was harmless. The pages that follow test that claim on ground he did not consider: not the physics of weapons, but the statistics of human populations, the training data behind facial recognition systems, and the algorithms that evaluate teachers. In each case, the harm does not arrive through a misuse of the mathematics. It is present in the assumptions embedded before the calculation runs.

Big Idea 1 - The statistics of human worth

Eugenics and Francis Galton

Francis Galton's statistical work on heredity, twins, fingerprints and human measurement shaped modern science. But Galton also coined “eugenics,” arguing that desirable human traits should be encouraged through breeding. His ideas influenced later social Darwinists, IQ testing, sterilisation laws and racist policies. The video presents Galton as both scientifically and mathematically  innovative and morally dangerous: measurement became a tool for judging human worth

Francis Galton coined the word "eugenics" in 1883, from the Greek for "well-born." His project was to use scientific measurement to identify which people were fit to reproduce and which were not. Galton was Charles Darwin's cousin and took natural selection as his premise: if animals could be bred for desirable traits, so could people. In 1904 he convinced University College London to establish the world's first Eugenics Record Office. His close collaborator Karl Pearson became the first Professor of National Eugenics at UCL when Galton died in 1911.

What Galton and Pearson built was also statistics. Galton named regression - what he called "regression to mediocrity," observing that tall parents tend to have children closer to the population average, and applying the same logic to the inheritance of intelligence. Pearson developed the correlation coefficient that still bears his name, along with chi-square and standard deviation. These are tools in the IB Statistics syllabus. The r value you calculate when running a bivariate analysis in your IA was designed at a laboratory whose full name was the Galton Laboratory for National Eugenics. As the archivist of Galton's papers at UCL told Angela Saini in Superior: "Pearson's greatest contribution, the thing that people remember him for, is founding the discipline of statistics. A lot of work on that was done with Galton."

The laboratory was eventually renamed. It is now UCL's Department of Genetics, Evolution and Environment. A Galton Professor of Genetics remains, funded by money Galton left in his will. The Eugenics Society became the Galton Institute in 1989. The statistical tools kept only their names.

Hardy wrote A Mathematician's Apology in 1940. He was drawing a distinction between pure mathematics, which he believed served no harmful purpose, and applied mathematics, which could. His examples of dangerous applied mathematics were ballistics and aerodynamics. Galton and Pearson were his contemporaries, working a few miles from his Cambridge office. The bell curve Galton used obsessively - to argue that heritable intelligence was normally distributed and that eugenic intervention should target its tails - is the same normal distribution in Topic 4 of the Maths syllabus you are being examined on.

Gould's analysis in The Mismeasure of Man identifies what made this possible. Numbers carry a specific kind of authority. A skull measurement, an IQ score, a correlation coefficient - each appears to report a fact about the world rather than an argument about it. Gould's term for the error is reification: treating a number as if it were the thing it measures, rather than an abstraction from it. A correlation between inherited traits is a statistical relationship in a dataset. It does not establish causation, fix what the traits mean, or justify the choice of which population to measure. But once the number exists, it acquires a weight the underlying judgment did not have. Gould's observation is precise: "Numbers and graphs do not gain authority from increasing precision of measurement, sample size, or complexity in manipulation. Basic experimental designs may be flawed and not subject to correction by extended repetition."

The question is not whether Galton and Pearson were wrong - they were - but what their work reveals about mathematical tools more broadly. A tool developed for a specific purpose does not shed that purpose when moved to new contexts. It carries it in the questions it makes natural to ask and the questions it makes harder to think of. Statistical methods designed to rank human populations continue to shape how researchers think about human difference. Whether that framing is neutral is a question the mathematics itself does not raise.

Karl Pearson - AOK Key Thinker.webp

Big idea 2 - Defaults are not neutral

Joy Buolamwini is a computer scientist, founder of the Algorithmic Justice League, and researcher at MIT. In 2015, working on an art project that used facial recognition software to project different faces onto her own reflection, she made a discovery. The software could not detect her face. She tested a hand-drawn smiley face on her palm - detected. She held a white Halloween mask over her face - detected. When she removed the mask, her dark-skinned face went undetected again. She describes the experience as encountering the "coded gaze": the structuring of a system around an assumed default that is neither named nor questioned. Her book Unmasking AI (Random House, 2023) follows the research that began that night.

Decoding Algorithmic Bias

Joy Buolamwini explains how an art project led her to discover racial and gender bias in facial detection software. Her face was not recognised until she wore a white mask, prompting research into AI discrimination. Her Gender Shades project found systems performed worst on darker-skinned women, sometimes close to random guessing. The research helped pressure companies to stop selling facial recognition to law enforcement. Buolamwini argues AI is not just technical: society must decide which technologies it wants.

The study she subsequently published as Gender Shades (2018) tested three major commercial facial recognition systems - IBM, Microsoft, and Face++ - on a dataset she built herself, the Pilot Parliaments Benchmark. She chose parliament members from countries with high female representation: from Iceland, Finland, and Sweden on the lighter-skinned end; from Rwanda, Senegal, and South Africa on the darker-skinned end. The benchmark was small by industry standards - 1,270 images - but its demographic composition was deliberate. The datasets most commonly used to train and test facial recognition systems at the time were drawn predominantly from lighter-skinned male subjects.

The overall accuracy figures looked acceptable: IBM 87.9%, Microsoft 93.7%, Face++ 90%. These are the kind of aggregate performance metrics that appear in technical reports and sales materials. Buolamwini's contribution was to show that overall accuracy was the wrong question. She broke the dataset into four intersectional groups - lighter-skinned males, lighter-skinned females, darker-skinned males, darker-skinned females - and found something the headline figures concealed. Microsoft achieved 100% accuracy on lighter-skinned males and 79.2% on darker-skinned females. IBM's gap was larger still: the best-performing group was lighter-skinned males; the worst was darker-skinned females at 65.3%. The largest single gap, 34.4 percentage points between lighter-skinned males and darker-skinned females on the IBM classifier, was, as Buolamwini writes, "not captured in the aggregate performance of IBM at 87.9 percent accuracy on the entire benchmark."

What the intersectional analysis revealed was an unequal distribution of errors. In Topic 4 of your statistics course, false positives and false negatives appear as Type I and Type II errors - the two kinds of mistake a statistical test can make. The Gender Shades study shows that in a facial recognition classifier, these errors do not fall randomly. They fall along demographic lines. Darker-skinned women were significantly more likely to be misclassified: their faces falsely rejected, or falsely attributed wrong attributes. The mathematics of the classifier does not contain this bias in any explicit form. The errors are not a flaw in the algorithm's logic. They are a consequence of what the algorithm was trained on.

This is Buolamwini's diagnosis: the problem precedes the mathematics. If the datasets used to train a system are predominantly lighter-skinned and male, the system learns to recognise predominantly lighter-skinned male faces. Those faces become the de facto standard against which accuracy is measured, and errors that fall on other groups become invisible in the aggregate figure. The chapter title in Unmasking AI names this precisely: "Defaults Are Not Neutral." Every choice about what data to collect, whose faces to include, and how to define a correct classification embeds assumptions that the finished system treats as objective. Buolamwini argues that "the burden of proof of performance needs to be placed on the people developing systems, not those who are impacted by their use."

The same principle applies when you design your IA. Reporting a headline accuracy figure or an overall correlation coefficient without examining whether the result holds across different subgroups is exactly the limitation Buolamwini identifies in the industry benchmarks she audited. The technical validity of a statistical method and the validity of its conclusions are different questions - and the second depends on what the data actually represents.

Where Big Idea 1 showed that statistical tools carry their ideological origins into subsequent use, Buolamwini shows that contemporary systems embed their assumptions before the mathematics begins. The number produced by the classifier appears to be objective.  What it has actually processed is a particular version of the world - one shaped by whose data was collected, and whose performance anyone thought to measure. Big Idea 3 provides another example of this. 

Big idea 3 - The authority of the opaque

Cathy O'Neil, whose Weapons of Math Destruction (2016) appeared in the Ethics in Technology section of this site, trained as a mathematician before working at a hedge fund and later in data science. Her book's central argument is that mathematical models acquire authority from their complexity - they make consequential decisions while appearing too objective and too technical to question. The case she opens with is drawn from education: a teacher's career and an algorithm that could not be held to account.

In 2009, Washington DC's school chancellor Michelle Rhee introduced a teacher assessment system called IMPACT. One component was a value-added model (VAM): a statistical formula built by a Princeton-based consultancy, Mathematica Policy Research, to measure how much each teacher had contributed to students' test score progress over the year. The formula accounted for half of each teacher's overall evaluation score.

Sarah Wysocki was a fifth-grade teacher with two years at MacFarland Middle School. Her principal praised her. Parents praised her. One evaluation called her "one of the best teachers I've ever come into contact with." At the end of the 2010-11 school year, her IMPACT score fell below the minimum threshold, and she was fired along with 205 other teachers.

She tried to find out why. She was told the algorithm was too complex to explain. A colleague pressed the district administrator for months and eventually asked directly: "How do you justify evaluating people by a measure for which you are unable to provide explanation?" She was told to wait for a technical report. What O'Neil identifies here is not an administrative failure - the opacity is built into how the model functions. A judgment that cannot be explained is one that cannot be appealed.

Later investigations by the Washington Post and USA Today found a high rate of erased and corrected answers on standardised tests at forty-one schools in the district, including the school where many of Wysocki's students had come from. Her incoming fifth graders had been recorded as reading at five times the district average, yet when classes started many struggled with simple sentences. The most plausible explanation is that fourth-grade teachers at the feeder school had altered their students' answer sheets. Wysocki's students then scored normally in fifth grade, and the algorithm recorded a decline and attributed it to her. The model could not account for what had actually happened, had no mechanism for receiving that information, and produced its verdict regardless. O'Neil's observation is exact: "Instead of searching for the truth, the score comes to embody it."

Wysocki was out of a job briefly, landed a position at a school in an affluent Virginia district - one that did not use algorithmic evaluation - and the poor DC school lost a good teacher to a rich one that recruited by more traditional means. The algorithm had not identified a bad teacher. It had transferred a good one.

The question this raises connects directly to what the Scope and Perspectives pages explored: what mathematics can and cannot represent. A human assessment of a teacher's performance can be examined - its assumptions can be identified, its reasoning challenged, its evidence weighed. A formula processes the same inputs and produces a number that appears to have bypassed all of that. When the complexity of the model is also the reason it resists inspection, the mathematics has not replaced human judgment. It has made human judgment invisible.

The ethics of Mathematics compared

A tool that carries its purpose long after the purpose is forgotten

Natural Science’s central ethical failures, on its own Ethics page, happen at the point of production: an experiment performed on a subject who did not consent, a virus engineered before anyone has decided who is accountable for it. Mathematics’s harm arrives later. Galton and Pearson built regression and the correlation coefficient to rank human populations by heredity. Nobody is harmed in that room in 1904. The harm arrives a century later, when the same formula sits inside an IB syllabus, or inside a value-added model that ends a teacher’s career, carrying an intention nobody examining the output can see.

Human Science’s Ethics page turns on subjects who can read what has been written about them and answer back. The people processed by a facial recognition classifier or a teacher-evaluation algorithm mostly cannot. Sarah Wysocki never got an explanation she could contest, because the model that scored her was built to produce a number, not a reason. Where Human Science studies a subject who notices and reacts, Mathematics here studies a subject who is simply computed.

History shares something with this page that is easy to miss: both disciplines can be weaponised by people who were never the ones who built the original tool or told the original story. A national myth outlives the historian who first wrote a milder version of it. A correlation coefficient outlives the eugenicist who designed it. The tool or the story survives; the purpose behind it has to be argued back into view by someone willing to ask why it exists.

Ethics in the Arts asks whether a work’s own success can be part of what makes it wrong to have made. This page asks something structurally close: whether a model’s own technical accuracy, its 87.9 percent, can be part of what makes it wrong to have used. In both cases the thing that makes the object impressive on its own terms is exactly what makes it dangerous once it is put to use.

Think further: questions and resources

  • Galton and Pearson developed correlation, regression, and the chi-square test as instruments for measuring hereditary traits they believed demonstrated racial hierarchy. Those methods have since been applied across medicine, social science, and your own IB statistics course, shorn of their original purpose. If the mathematical validity of a method is independent of the theory that motivated it, does knowledge of that theory change how the method should be used?

  • Buolamwini found that IBM's facial recognition classifier achieved 87.9% overall accuracy on her benchmark but only 65.3% on darker-skinned female faces - a gap of more than 34 percentage points that the headline figure concealed. If a mathematical system distributes its errors unevenly across demographic groups while performing well in aggregate, which figure constitutes an accurate description of its performance?

  • Buolamwini argues that the burden of demonstrating a system's performance should rest with those who build it, not those it affects. If mathematical complexity makes it difficult for those subject to a system to assess its workings independently, does that complexity change where the burden of proof should sit?

  • O'Neil identifies a feedback loop at the heart of predictive models: a system trained on historical data encodes the patterns of past decisions, including those shaped by discrimination, then generates predictions that produce new data appearing to confirm it. If the training data reflects past discrimination, what would it mean for the model's predictions to be accurate?

  • O'Neil argues that the opacity of a mathematical model is not incidental to its authority but constitutive of it: the score is trusted precisely because it cannot be questioned. If the appearance of objectivity depends on the impossibility of scrutiny, what kind of knowledge does the score represent?

  • Hardy maintained that pure mathematics was ethically harmless because it had no practical application, and that whatever moral weight followed belonged to the applied scientist. If the harm arises not from application but from assumptions embedded in the tools before they are applied, where should the ethical question be located?

Films
For more see my 10 films for the TOK journey page.

🎬  WATCH — Oppenheimer (2023)

Christopher Nolan

Christopher Nolan's three-hour account of the Manhattan Project and its aftermath opens on this page's provocation: the physicist who described watching the first atomic bomb detonate as the moment he became death. The film is less interested in the physics than in the moral reckoning that followed - the security hearings, the political manoeuvring, the question of whether Oppenheimer's conscience caught up with him too late to matter. Hardy's claim that the pure mathematician has their conscience clear rests on a distinction between the work and its uses. The film examines what that distinction costs when the uses become visible. It also shows the institutional pressures that shape what scientists pursue and what they suppress, which connects to the argument in BI3 about who controls the framing of mathematical outputs and who bears the consequences.. My students can watch the film here.

🎬  WATCH — Gattaca (1997)

Andrew Niccol

Set in a near future where genetic sequencing at birth determines life outcomes - career, insurance, social standing - Gattaca imagines a society that has made statistical prediction total. The protagonist Vincent Freeman is classified at birth as having a 99% probability of heart failure before 30; his brother Anton, with a superior genetic profile, follows the path the numbers prescribed. The film is the fiction version of BI1's argument: what happens when a mathematical ranking of human potential becomes the basis for all institutional decisions, and the score replaces the person it was supposed to describe. Galton believed that hereditary measurement would allow society to identify its most capable individuals. Gattaca shows that world built out, and what it costs the people the numbers place at the bottom. My students can watch the film here

🎬  WATCH — Coded Bias (2020)

Shalini Kantayya

The documentary that follows Buolamwini's Gender Shades research from the MIT lab where the white mask discovery happened through to congressional testimony and the campaign for legislation regulating facial recognition. The film is the closest companion to Big Idea 2 on this page: it shows the research being conducted, the industry responses when the accuracy gaps were published, and what happens when someone asks a technology company to explain why its system fails on particular faces. The second half broadens from facial recognition to predictive policing and benefits algorithms, which connects the argument in BI2 to the opacity and accountability questions BI3 raises. Buolamwini appears throughout as both researcher and activist, which makes visible something the page also examines: that identifying a flaw in a mathematical system and persuading institutions to act on that identification are entirely different problems. My students can watch the film here

🎬  WATCH — Eugenics: Science's Greatest Scandal (2019)

BBC Two / Angela Saini and Adam Rutherford

A two-part BBC documentary presented by Angela Saini - whose book Superior appears in the further reading below - and geneticist Adam Rutherford. It covers the origins of the eugenics movement in Britain, the role of Galton and Pearson, the spread of eugenics policy across Europe and North America in the early twentieth century, and the question of how far the science has actually been left behind. The documentary is the most direct companion to BI1 on this page: it puts faces and institutions to the argument that the statistical tools developed in this period were not politically neutral, and it examines the persistence of eugenicist thinking in contemporary genetics research. Saini's presence as presenter makes the connection to her written argument explicit.

Further reading

📚 READ - Weapons of Math Destruction, Cathy O'Neil, 2016

O'Neil's book is the fullest account of BI3's argument. The reader extract covers the teacher evaluation case from Chapter 1; read Chapter 2 ("Shell Shocked") for O'Neil's account of the 2008 financial crisis and the role of mathematical models in producing it, and Chapter 5 ("Civilian Casualties") for her analysis of predictive policing algorithms. The concept of the Weapon of Math Destruction - a model that is opaque, widely used, and damaging - is defined in the introduction and is worth reading before the extract. In the library in TOK Books > Mathematics.

📚 READ - Unmasking AI, Joy Buolamwini, 2023

The reader extract comes from Chapter 11, which covers the Gender Shades study. Chapter 5 ("Defaults Are Not Neutral") is the chapter that names the central argument and is essential reading alongside it - it explains how the assumptions embedded in training data become invisible in the finished system. The introduction describes the white mask moment in full. In the library in TOK Books > Mathematics.

📚 READ - The Mismeasure of Man, Stephen Jay Gould, 1981 (revised 1996)

Gould's account of the history of attempts to quantify human intelligence and rank human groups by innate ability. The reader extract comes from Part II, which covers Galton and the origins of correlation and regression in the eugenics programme. Chapter 6 ("The Real Error of Cyril Burt") extends the argument into the twentieth century and shows how fabricated data can circulate for decades inside a scientific community when it confirms what the community already believes. In the library in TOK Books > Mathematics.

📚 READ - Superior: The Return of Race Science, Angela Saini, 2019

Saini's investigation into contemporary race science and its connections to the eugenics tradition. The reader extract covers Galton and Pearson at UCL. Chapter 7 ("Caste") examines how categories of human classification developed in one context migrate into scientific research conducted in entirely different contexts, which is the argument BI1 makes about statistical tools. Chapter 11 ("The Silence") addresses the question of why scientists working in fields adjacent to eugenics often avoid naming the connection. In the library in TOK Books > Mathematics.

📚 READ - Hello World: How to Be Human in the Age of the Machine, Hannah Fry, 2018

Fry's chapter on criminal justice (Chapter 3) covers the COMPAS algorithm and the mathematical impossibility theorem at the heart of predictive sentencing: it is not possible for a risk score to be simultaneously fair to defendants and accurate in aggregate. Her account is more measured than O'Neil's - she acknowledges the cases where algorithmic decision-making outperforms human judgment and asks what the comparison should be. Read it alongside O'Neil for a more complete picture of what the debate about algorithmic fairness actually involves. In the library in TOK Books > Mathematics.

bottom of page