

areas of knowledge - NAtural science
Methods and Tools of Natural Science
How does natural science produce knowledge that compels agreement - and what does it mean when the machinery fails?
THE PROVOCATION
The waiter sets out to prove Tyson wrong. He gets the same result.
Neil Degrasse Tyson
Neil deGrasse Tyson is telling a story about a cup of hot chocolate. The argument he arrives at describes something that runs through every page of this site: a single researcher's result is not yet a scientific truth. What it takes to get there is the question this page is built around.
Â
Watch the clip before reading on, and notice what Tyson says has to happen - and who has to do it.
Neil deGrasse Tyson ordered hot chocolate with whipped cream. The waiter told him it had sunk to the bottom. Tyson pointed out that whipped cream, being less dense than hot chocolate, cannot sink. The waiter, determined to prove Tyson wrong, went into the kitchen, scooped fresh whipped cream into the mug, and watched it float. The result was identical to what Tyson had predicted. That, Tyson argues, is how science works: a provisional claim, verified by someone with a stake in disconfirmation, converging on consensus through repeated independent test. Not one researcher's result, but a competitor's attempt to disprove it that gets the same answer.
​
In Core Lesson 5 we established the logical structure of scientific argument: Popper's asymmetry between verification and falsification, Kuhn's account of how paradigms determine what counts as a problem worth solving. This page takes up the question that follows: what does implementing that logic look like across thousands of researchers, decades of work, and billions in funding? And what happens when the institutional machinery that is supposed to carry out that logic has structural failures of its own?
​
Three sections follow. The first examines replication - the human institution that gives Popper's logic its practical form. The second looks at the instruments through which science reaches the world, and asks whether those instruments are as neutral as they appear. The third examines a case where the method eventually worked but the system delayed it for decades.
Big idea 1 -Â The logic of science needs an institution to carry it
Friends - Ross, Phoebe and Evolution
Ross Geller, palaeontologist, arrives with a briefcase of fossil evidence to convince Phoebe that evolution is not just one of many possibilities. Phoebe responds not with counter-evidence but with a rhetorical move: science has been wrong before - the flat earth, the indivisible atom. She asks Ross to admit there might be a teeny tiny possibility he is wrong. Watch what happens when he agrees. Then watch Phoebe's reaction. Her response turns out to be the more interesting epistemological position.
Karl Popper, introduced in Core Lesson 5, showed that scientific knowledge advances through falsification rather than confirmation. Any number of white swans does not prove all swans are white; one black swan falsifies the claim. The asymmetry is decisive: verification accumulates, but a single falsification refutes. This is what distinguishes scientific claims from claims that can always be reinterpreted to accommodate contrary evidence.
​
Popper described a logical structure. He did not describe a social institution. Tyson's waiter fills that gap. The waiter did not set out to confirm Tyson's claim about whipped cream; he set out to prove Tyson wrong, and got the same result. That is what makes the result epistemically significant: it came from a competitor, someone with a direct stake in disconfirmation. Peer review and independent replication are the institutional forms of this. A finding becomes a scientific result not when one researcher obtains it, but when researchers working independently - ideally, researchers who would prefer a different answer - obtain the same thing.
​
The Friends clip puts pressure on this from a different angle. Phoebe Buffay invokes scientific fallibilism - science has been wrong before, the earth was once thought flat, the atom was once thought indivisible - to argue that evolution might be "one of many possibilities." Ross, the palaeontologist, eventually concedes there might be "a teeny tiny possibility" he is wrong. Phoebe immediately loses all respect for him. She did not want him to change his mind. She wanted him to cave. Ross's capitulation identifies exactly the failure that peer review is designed to prevent. He revises his position not because Phoebe produces a falsifying result but because the argument became uncomfortable. The institutional structure of science - the requirement that position changes track evidence rather than social dynamics - is a structural response to that temptation.
Â
The epidemiologist John Ioannidis published a paper in 2005 titled "Why Most Published Research Findings Are False." The argument was not that scientists are dishonest. It was that the institutional machinery of science has structural failure modes. Small sample sizes reduce the probability that a positive result reflects a real effect rather than chance. Publication bias - the tendency of journals to publish positive findings and to reject null results - means the published literature is not a representative sample of all experiments run. Flexible analysis practices allow researchers, sometimes consciously and sometimes not, to adjust how data is handled until a result crosses the threshold of statistical significance. Each failure is individually understandable; together they produce a literature where, in some fields, the majority of published findings cannot be reproduced.
​
This 'replication crisis' is evidence that the machinery through which scientific logic gets implemented is imperfect - and that this imperfection can be identified, measured, and corrected. Ioannidis's paper, published in a peer-reviewed journal and subsequently examined by the Open Science Collaboration in 2015, is the method working on itself. Phoebe is right that science has been wrong, and right that specific published findings have failed to replicate at scale. Her conclusion - that evolution is therefore "one of many possibilities" - does not follow, because the replication crisis does not show that scientific claims are as uncertain as non-scientific ones. It shows that the machinery is imperfect; the mechanism that identifies that imperfection is itself scientific.

Big idea 2 - Scientific instruments do not simply observe - they intervene
In the Optional Theme on Language, we saw that writing is not a neutral carrier of thought: it restructures what kinds of knowledge are possible. Abstract argument, systematic comparison across time, and certain forms of reasoning that oral culture cannot sustain all depend on the tool itself. The instrument does not merely assist thinking; it changes what can be thought.
​
Scientific instruments work analogously, but outward - toward the world - rather than inward toward cognition. The philosopher Ian Hacking, in Representing and Intervening (1983), made the argument precise. Experiments in natural science do not simply observe what is already there, waiting to be measured. They intervene: they create the conditions under which something with specific properties becomes detectable, measurable, and therefore real as an object of knowledge. The electron microscope did not reveal mycorrhizal networks - the underground fungal connections between trees examined in the Perspectives page - as if they had been waiting for a better lens. It made them countable. The instrument produced a category of object that prior frameworks had no access to.
​
The same logic appears at more familiar scales. Galileo's telescope made the moons of Jupiter observable in 1610, and in doing so produced facts that the Ptolemaic system had no framework to accommodate: four satellites in predictable orbits around a body that was not Earth. The instrument did not extend vision and leave the existing framework intact; it generated objects whose existence forced a reckoning with the framework itself. The thermometer worked differently but analogously. It did not measure something already quantified in nature - it created temperature as a precise, numerical, communicable object. Before the instrument existed, hot and cold were qualitative and personal; after it, they were data points that could be compared between observers, replicated across contexts, and entered into equations. The instrument brought a new category of scientific object into existence.


(Left) Galileo's telescope in the Museo Galileo in Florence, with Mr Jones-Nerzic reflected in the glass. (Right) Moon engravings in the book the telescope made possible.
The particle accelerator at CERN follows the same logic on a larger scale. The Higgs boson was not sitting somewhere in space waiting to be detected. The conditions under which something with its predicted properties became observable required 27 kilometres of superconducting magnets, decades of engineering, and a threshold of statistical confidence - five sigma, a result that would occur by chance less than once in 3.5 million trials - that the collaboration set in advance as what would count as a confirmed result. The instrument, the method, and the evidentiary standard were constructed together. None was independent of the theoretical framework they were built to test.
​
Hacking's argument has a specific consequence for replication. When an instrument produces a surprising finding, two explanations are always in competition: either something has been discovered, or the instrument is producing an artefact - a result that reflects how the instrument works rather than how the world is. Distinguishing between these requires triangulation: when multiple independent methods using different instruments converge on the same phenomenon, the artefact explanation becomes difficult to sustain. This is why scientific consensus requires not just replication but independent replication using different experimental approaches.
​
IB Chemistry laboratory design makes this concrete. Choosing which instrument to use, which variables to control, and which units to report results in are not neutral preliminaries to the real experiment. Every choice embeds a prior commitment about what counts as precision, what counts as a measurable quantity, and which features of the situation are relevant. The design is already a theoretical claim about what the instrument can reach - and what it cannot.
Big idea 3Â -Â The method can work and the system can still delay it
Joanna Moncrieff
Joanna Moncrieff is a professor of psychiatry at University College London and one of the leading critics of how antidepressants are understood and prescribed. This Dr. Josef video outlines her central argument: that the chemical imbalance theory of depression was never robustly established, and that a drug-centred model offers a more accurate account of what these drugs do. Watch for the distinction she draws between what a drug is claimed to fix and what it does. That distinction drives the analysis below.
In 2022, a team led by the psychiatrist Joanna Moncrieff published a systematic review examining the evidence base for the chemical imbalance theory of depression: the claim, widely communicated to patients and embedded in pharmaceutical marketing for three decades, that depression is caused by low serotonin levels in the brain. The review found no consistent evidence that people with depression have lower serotonin than people without it. The theory had not been robustly established before it became a clinical consensus.
​
The systematic review - a method that synthesises evidence across multiple existing studies rather than relying on any single trial - was not invented in 2022. The tools needed to interrogate the serotonin hypothesis had existed for decades. What had not happened, until Moncrieff's team did it, was a comprehensive and independent application of those tools to this specific question.
​
This is a different kind of failure from the one Ioannidis describes in Big Idea 1. The replication crisis is a failure of the machinery: too many small studies, too much publication bias, too much flexibility in analysis. The serotonin story is a failure of incentives: the systematic review was not conducted because conducting it was not in the interests of the parties with resources to fund it. Pharmaceutical companies that had profitable drugs built around the serotonin hypothesis did not commission research designed to test whether the hypothesis was correct.Â
​
Kuhn's account of normal science, introduced in Core Lesson 5, helps explain the delay. Once a paradigm is established - once the serotonin hypothesis had been absorbed into clinical training, prescribing guidelines, and patient communication - anomalies tend to get explained away rather than examined seriously. Researchers working within the paradigm are not acting in bad faith; they are doing what researchers trained in any framework do: fitting new data into the existing structure rather than questioning the structure itself. The paradigm shifts eventually, but not on the schedule that the evidence alone would warrant.
Â
The serotonin case shows that the method does not implement itself. The tools exist and the logic is sound, but when they get applied - and to which questions - depends on institutional conditions that the method alone does not control: funding structures, publishing incentives, commercial interests, and the social momentum of an established consensus.
The methods and tools of Natural Science compared
How does a result become a fact - and what happens when the machinery fails?
Natural science occupies a specific position among the areas of knowledge in relation to method. Its institutional structure builds competitor falsification directly into the process: peer review requires that findings survive scrutiny from researchers who have every incentive to find them wrong. The knowledge that emerges carries a particular kind of authority - not certainty, but documented survival under adversarial test. What this looks like across the other areas of knowledge is revealing by contrast.
Â
Mathematics does not replicate - it proves. A mathematical result, once established through valid proof, does not need to be run again in a different institution by a different researcher. The guarantee is deductive: if the premises are correct and the reasoning is valid, the conclusion follows necessarily. This is a different kind of guarantee from what natural science can offer, and in some respects a stronger one - but it operates only within a formal system, and it has nothing to say about whether that system accurately describes any feature of the physical world. What Gödel showed - examined in the Methods and Tools in Mathematics page - is that even this internal guarantee has limits from the inside.
​
History cannot run the experiment again. The events historians study are fixed, the sources are fragmentary, and no amount of methodological rigour can recreate the conditions under which they occurred. Where natural science isolates variables and replicates results, history reconstructs causes from incomplete evidence and argues for the most defensible interpretation. Claims must survive competitor scrutiny, but the object of that scrutiny is an argument about evidence rather than a reproducible result.
​
The human sciences apply the experimental method to subjects who know they are being studied. The Hawthorne effect - the tendency of people to modify their behaviour under observation - means that the controlled conditions natural science depends on are, in the human case, partly constituted by the act of investigation itself. The replication crisis that Ioannidis identified in medicine has proved even more acute in psychology, where sample sizes tend to be smaller, effects tend to be more context-dependent, and the gap between laboratory conditions and real-world behaviour tends to be larger.
​
The arts produce knowledge without the falsification criterion. A poem is not wrong. A painting does not replicate. What it means to produce and evaluate knowledge in the absence of competitor falsification is a question this site returns to in the Methods and Tools in the Arts page.
Think further: questions and resources
-
John Ioannidis showed in 2005 that most published findings in some fields do not replicate, partly because journals are more likely to publish positive results and partly because small samples inflate apparent effect sizes. Is this a failure of the scientific method, or evidence that the method is working - since the replication crisis was itself identified, quantified, and published?
-
Ian Hacking argues in Representing and Intervening that experiments do not reveal pre-existing phenomena but create the conditions under which certain features of the world become measurable. If the instrument partly constitutes what it measures, what does independent replication actually establish - that the phenomenon is real, or that the instrument is consistent?
-
Goldacre documents that trials with positive results are roughly twice as likely to be published as trials with negative results. The serotonin hypothesis shaped prescribing for thirty years before a systematic review found no consistent evidence for it. Does this undermine the claim that evidence-based medicine is evidence-based, or only the claim that current practice lives up to that standard?
-
Oreskes and Conway in Merchants of Doubt show how a small group of scientists manufactured public uncertainty about tobacco harm and climate change - not by producing new evidence but by making existing consensus appear contested. What distinguishes genuine scientific disagreement from manufactured doubt, and who is responsible for that distinction being legible to the public?
-
The CERN collaboration set five sigma as the threshold for announcing the Higgs boson: a result that would occur by chance less than once in 3.5 million trials. In most medical research the threshold is p < 0.05, meaning one in twenty. Who should set the evidentiary standard for what counts as a scientific result, and on what grounds should that standard be set?
-
Ross abandons his position on evolution not because Phoebe produces counter-evidence but because the conversation becomes uncomfortable. Kuhn showed that scientific communities also resist change through social mechanisms rather than purely evidential ones. Is the social dimension of science a corruption of its method, or part of how the method actually works?
Films
For more see my 10 films for the TOK journey page.
🎬  WATCH — Particle Fever (2013)
Mark Levinson
A documentary following physicists at CERN in the years surrounding the search for the Higgs boson. What makes it directly relevant to this page is not the physics but the methodology: the scale of the instrument required, the evidentiary threshold the collaboration set in advance, and the moment when the data crossed it. The film makes visible what accounts of scientific discovery usually skip over - the extended period of null results, the debates about what would count as confirmation before any result arrived, and the collective process through which a statistical signal becomes a scientific fact. The instrument here required decades of engineering and a standard of evidence agreed upon before the data arrived; neither the tool nor the threshold was independent of the theoretical framework they were constructed to test. My students can watch the film here.Â
🎬  WATCH — Icarus (2017)
Bryan Fogel begins with a personal experiment: can he pass a professional cycling drug test while using the same doping protocols as the professionals, with medical supervision? The experiment draws in Grigory Rodchenkov, then director of Russia's national anti-doping laboratory, who reveals that a state-sponsored system had been systematically manipulating the drug testing protocol for years. The film is relevant here because the testing system - the institutional machinery designed to produce reliable results through independent verification - had been engineered to fail from the inside. The method was sound; the institution was not. The case sits alongside the serotonin story as an example of what happens when the conditions the scientific method requires are systematically compromised by the institutions that are supposed to implement it. My students can watch the film here.Â
Further reading
📚 READ - Bad Science, Ben Goldacre, 2008
Goldacre is a physician and epidemiologist who writes about how evidence works and how it gets distorted. Chapters 5 (The Placebo Effect) and 14 (Bad Stats) are the most directly relevant to this page. Chapter 5 examines the randomised controlled trial: what it is designed to establish, why it became the standard it is, and what earlier forms of medical evidence failed to show. Chapter 14 covers the statistical machinery of published research, including statistical significance, confidence intervals, and the kinds of analytical flexibility that allow researchers to extract a positive result from data that would not otherwise yield one. These two chapters together establish what Ioannidis's critique is actually aimed at: not science but a specific set of institutional practices that have accumulated around it. In the library in TOK Books > Science.
​
📚 READ - Bad Pharma, Ben Goldacre, 2012
Where Bad Science examines how evidence works and gets distorted in general, Bad Pharma focuses on the pharmaceutical industry's relationship with clinical research. Chapter 1 (Missing Data) documents the systematic suppression of negative trial results: companies run multiple trials, publish the ones that support the drug, and register the others in ways that make them effectively invisible to the doctors and regulators who depend on them. This is the institutional pattern behind the serotonin story - not one act of suppression but a structural arrangement across dozens of drugs and decades of trials. In the library in TOK Books > Science.
​
📚 READ - Why Trust Science?, Naomi Oreskes, 2019
Oreskes argues that science is the most reliable system of knowledge production humans have developed, and that this reliability comes from its social structure: the requirement that claims survive criticism from a community of diverse, independent investigators. Chapter 1 directly addresses the question the Phoebe and Ross clip opens - what distinguishes genuine scientific fallibilism from the manufactured uncertainty that Oreskes and Conway document in Merchants of Doubt. Chapter 2 (Science Awry) examines historical cases where the scientific community failed, and asks what structural features were absent in those cases that the method requires to work. In the library in TOK Books > Science.
​
📚 READ - Conjectures and Refutations, Karl Popper, 1963
Popper is the AOK Key Thinker for this page, and Chapter 1 - "Science: Conjectures and Refutations" - is the essay from which the reader extract is drawn. It is worth reading in full: Popper works through the contrast between Adler and Einstein slowly, and the seven numbered theses at pp.47-48 are more carefully argued than any summary can convey. He is also worth reading against Kuhn: Popper thought science advances through individual acts of falsification, Kuhn thought it advances through community-level paradigm shifts, and the disagreement between them is itself a Methods and Tools question. In the library in TOK Books > General TOK books.
​
📚 READ - The Science Delusion, Rupert Sheldrake, 2012
Sheldrake is a Cambridge-trained biochemist who argues that science has allowed certain metaphysical assumptions to harden into unexamined dogmas. The Introduction ("The Ten Dogmas of Modern Science") sets out his list of ten beliefs that most scientists absorb without critical scrutiny - the starting point for the reader extract. Whether you accept each item on his list is itself a question the book is designed to open, not close. Read the Introduction alongside Popper's Chapter 1: Popper asks what distinguishes a scientific claim from a pseudo-scientific one; Sheldrake asks whether the scientific community applies that standard to its own foundational commitments. Note that Sheldrake's specific hypothesis of morphic resonance is not accepted in mainstream science; the epistemological argument in the Introduction does not depend on it. In the library in TOK Books > Science.