Web Analytics
top of page
areas of knowledge - human science

Ethics of Human Science

Why is the mildest of these three studies banned exactly as completely as the worst?

Nobody was hurt. But now it's a study you can't run

THE PROVOCATION
Solomon Asch

This clip reconstructs Asch's 1951 procedure as the provocation describes it: confederates giving a wrong answer in turn, a naive participant near the end of the line, and a task simple enough that the correct answer is never in doubt. Watching someone contradict their own eyes to agree with the group makes the provocation's opening concrete before the argument about why. The mildest of the three studies here is also the hardest of the three bans to explain by harm alone.

In 1951, the psychologist Solomon Asch gathered groups of students and gave them a simple task: look at one line, then say which of three others matched it in length. The answer was obvious. Only one person in each group was a real participant, sat near the end of the line of respondents; everyone else had been told in advance to call out the same wrong answer, on certain trials, before the real participant had to speak. Across the study's critical trials, about three-quarters of participants went along with an answer they could see was wrong at least once, and the larger the group giving that wrong answer, the more likely they were to do it too.

​

Nobody in Asch's study was deceived about anything painful, and nobody was asked to hurt anyone else. The trials lasted seconds each, and the worst a participant experienced was the discomfort of disagreeing with a room full of people, or of going along with something in front of them that plainly was not true. Set next to a shock generator or a mock prison, a line-judging task is about as harmless as psychological research gets.

​

And yet, immediately after setting out when deception is allowed in student research, no harm caused, full debriefing given, the IB Psychology guide draws one line that ignores those conditions entirely: "conformity or obedience studies are not permitted under any circumstances," however mild the version, however thorough the debrief. Milgram's obedience experiment sits in the same banned category as Asch's lines. So does Zimbardo's Stanford Prison Experiment, a study that put people in far more genuine distress than Asch ever did.

​

Ethics in the human sciences works on more than one level, and this page moves through three of them in order. First, the ethics of what a study does to the person inside it while it is happening, consent, deception, harm. Second, the ethics of what the resulting knowledge does once it leaves the lab and enters institutions that use it, whether or not that knowledge turns out to be true. Third, the ethics of the stance a researcher takes while producing knowledge about people in the first place, since describing human behaviour is never quite as neutral as describing a chemical reaction.

Big idea 1 - Human subjects are not objects.

Milgram's obedience experiment, run in 1961, told participants they were taking part in a study of memory and learning, and placed them in the role of "teacher," responsible for shocking a "learner" in another room whenever a question was answered wrong. Ten years later, Philip Zimbardo's Stanford Prison Experiment asked ordinary student volunteers to spend two weeks living as either guards or prisoners in a mock prison built in a university psychology department. Both studies convinced their participants that a constructed situation was real, the way Asch's had, and both went a great deal further with what that conviction was then used to do. The full procedures are set out below.

Both figures above describe a result, not just a procedure. What each study produced, once the deception worked, was real distress that the participants had no way to stop or opt out of. Neither offered anything resembling the consent, debriefing, or right to withdraw that the field now requires as standard. Set against those requirements, in the terms your own Human Sciences lesson uses, the ESRC's own six objectives for ethical research, both studies would be refused outright by any ethics board operating today.

This is not only a historical problem. In 2014, researchers at Facebook and Cornell published a study in the Proceedings of the National Academy of Sciences describing an experiment run on 689,003 Facebook users, whose news feeds had been altered for a week to show more positive or more negative posts, without their knowledge, to see whether emotional states spread through a social network the way they do face to face. Facebook's position was that its terms of service, agreed to by every user on signing up, constituted consent. No user was told they were part of an experiment, and none was offered the chance to opt out. The methods had changed completely, a lab coat and a shock generator had become a line in a terms-of-service agreement nobody reads, but the underlying question Milgram's experiment raised in 1961 (what is owed to a person before you study them) had not gone away. It had simply moved onto a platform where consent could be claimed to exist without anyone experiencing it as consent.

Stanley Milgram - AOK Key Thinker.webp

Big idea 2 - Knowledge about behaviour becomes power. 

So far, the question has been what a study does to the person inside it. This Big Idea asks what happens once the resulting finding is loose in the world, whether or not it holds up.

​

Zimbardo's mock prison was dismantled once the Stanford Prison Experiment ended. What the study produced did not stay contained the same way. In 2019, the researcher Thibault Le Texier obtained previously unheard recordings of Zimbardo directing the study while it was running, coaching the "guards" on how to behave and suggesting specific methods of humiliating the "prisoners," including denying them access to the toilet. The behaviour Zimbardo later described as having emerged naturally from the prison situation had, in large part, been instructed. The psychologist and science writer Stuart Ritchie, reviewing Le Texier's findings, concludes that "the 'results' of the Stanford Prison Experiment, such as they are, are scientifically meaningless." None of that stopped the study from being used. Zimbardo testified as an expert witness at the trial of US military guards charged with abusing prisoners at Abu Ghraib, arguing that the situation, not the individuals, explained their behaviour. A finding that was largely staged still carried enough institutional weight to help decide how blame was assigned at an actual trial. Knowledge about behaviour did not need to be true to function as power there. It only needed to be believed, at the moment it mattered, by the people deciding who was responsible.

Zimbardo - AOK Key Thinker.webp

This is another version of a pattern that runs through the human sciences more broadly. A classification, once it exists, a diagnosis, a label, a score, tends to stay attached to the people it describes, and to change what happens to them next.

The philosopher Ian Hacking calls this the looping effect of human kinds. As he puts it, "what was known about people of a kind may become false because people of that kind have changed in virtue of how they have been classified, what they believe about themselves, or because of how they have been treated as so classified." Hacking's own examples are deliberately varied: the Romantic-era idea of the genius reshaped the behaviour of people who came to see themselves as geniuses, and the category of "the child viewer of television" produced a kind of person, and an entire field of research to match. Children who watched television were reconceived as a population at risk, exposed to on-screen violence, groomed into consumption, and drawn away from sport and schoolwork, and by 1997 the category had its own World Congress, with researchers travelling from as far as Chile and Tunisia to discuss it. The V-chip, a device built into television sets so parents could block whatever "the child viewer" was assumed to need protecting from, is a piece of hardware that exists only because the category did first. This example makes the mechanism easiest to see as a cycle rather than a single event, set out below.

You have already met a version of this twice. Ainsworth's Strange Situation, on the Methods and Tools page, classified infants as securely or insecurely attached, and Heidi Keller's work with Nso caregivers showed what happens when that classification meets a culture it was never built to describe. Showalter's account of hysteria, on the Perspective page, traces a diagnosis that shaped how an entire category of patient, mostly women, was treated by doctors who believed the label described something fixed rather than something the diagnosis itself was helping to produce. Both are the looping effect, applied before Hacking had a name for it.

As we saw in Core Lesson 4 the philosopher Michel Foucault made a more profound claim about what classification does inside institutions specifically. An examination, whether medical, psychological, or educational, does more than record a person: Foucault argues it is "at the centre of the procedures that constitute the individual as effect and object of power, as effect and object of knowledge," fixing each person against a norm and turning them into what he calls a "case." His example of Bentham's Panopticon prison design, where a prisoner can never verify whether the central watchtower is occupied and so behaves as though it always is, makes the mechanism vivid: "a real subjection is born mechanically from a fictitious relation." The file, the report card, the diagnosis do the same work outside any prison. They do not just describe a person to the institution holding the file, they make that person easier to control.

​

A classification raises a moral question the moment it begins to shape a life, whether or not it later turns out to be accurate. The Mathematics and Technology pages on this site make a version of this same argument through predictive sentencing and behavioural data, run by an algorithm rather than a diagnosis. Here the machinery is older and needs no computer at all: a file, a label, a case, doing the same sorting by hand.

Big idea 3 -  Description is never completely innocent. 

So far, this page has asked what a study does to the people inside it, and what happens to a finding once it leaves the lab. This last Big Idea asks about the stance a researcher takes while producing the finding in the first place: whether studying people can ever be purely descriptive, or whether some judgement is built in from the start.

Esther Duflo - TED

In this TEDGlobal talk, Esther Duflo asks why disaster aid flows freely, Haiti's earthquake, while a comparable death toll from preventable disease, every eight days, draws far less urgency, then argues we cannot know whether traditional aid works because there is no comparison group: only one Africa exists. Her answer is to break poverty into smaller, testable questions using randomised trials, illustrated with immunisation incentives, bed net subsidies, and deworming, each returning results startlingly different from intuition.

The economist Esther Duflo, who shared the Nobel Prize in economics in 2019 with Abhijit Banerjee and Michael Kremer for pioneering randomised controlled trials in development economics, has built her career on a particular answer to that question. Rather than asking why poverty exists, or arguing for a particular economic system, Duflo and her colleagues break the problem into pieces small enough to test: whether a specific subsidy raises school attendance, whether a specific reminder raises vaccination rates, whether a specific loan structure changes how a family saves. Duflo has described the aim of this approach as keeping the fight against poverty tied to evidence rather than ideology, description and measurement, not argument.

​

That looks like a way of staying out of the judgement business entirely: measure the intervention, report the result, leave the bigger questions to someone else. But the choice of what to measure carries a judgement of its own. Deciding that the right question is whether a subsidy raises attendance, rather than why the school lacks resources in the first place, or why the family is poor at all, already decides which causes are worth investigating and which are somebody else's problem. The Perspective page on this site showed how a measure as basic as GDP quietly builds in a judgement about what counts as economic activity. Testing small, answerable interventions is the same kind of choice: a judgement about the right scale to work at, arrived at before a single result comes in.

​

Geertz, on the Methods and Tools page, approaches the same question by the opposite method. He argues explicitly that understanding a culture requires interpreting the meaning behind a gesture, not just recording it, so judgement is a stated part of his approach from the first observation, not something he is trying to avoid. Duflo tries to keep judgement out of her method entirely. But the conclusion above applies to her just as much as it applies to Geertz: a method built to avoid judgement ends up making one anyway, in the choice of what scale to work at. Two methods that could not be more different, one narrowly quantitative, one openly interpretive, arrive at the same place. Neither removes the human scientist from the job of deciding what mattered enough to study.

​

Whether a human scientist's role is only to describe what is the case, or also to judge what should be the case, does not have a clean answer on either method. The choice of what to study, and at what scale, has already made part of that judgement before the research begins.

The ethics of Human Science compared

What happens when the thing you are studying can read what you wrote about it?

A natural scientist studying a chemical reaction, or a mathematician working through a proof, is not accountable to their subject matter the way a human scientist is. A compound does not need to consent to being tested, and a number does not change its behaviour once it learns how it has been classified. Every Big Idea on this page depends on the fact that human subjects can. Milgram and Zimbardo's participants had to be deceived precisely because a fully informed subject might behave differently. Hacking's looping effect exists only because a person classified as a certain kind can read about that classification and respond to it, something no rock formation or chemical compound has ever done. Even Duflo's attempt to stay purely empirical runs into the same wall: the family whose behaviour is being measured is not inert data, and the choice of what to measure is a decision about people who could, in principle, object to how they are being studied.

​

The Mathematics and Technology pages on this site examine what happens once numbers about people are put to use, in sentencing, in surveillance. This page sits one step earlier. It asks about the ethics of producing knowledge about a subject who is aware they are being studied, and who does not hold still while you study them.

​

Natural Science and History sit on either side of this problem. Ethics of Natural Science describes harm that happens at the point of production, an experiment performed on a subject who cannot consent. Ethics of History describes an obligation that runs the other way, toward subjects who are mostly already dead and can never object to anything written about them. Human Science sits between the two: its subjects are alive, and unlike a chemical compound or a historical archive, they can read what has been written about them and answer back.

​

Ethics in the Arts asks whether a work’s own success can be part of what makes it wrong to have made. Human Science’s Milgram and Zimbardo cases ask something adjacent: whether producing genuinely valuable knowledge about obedience or conformity can be part of what makes an experiment wrong to have run, independent of whether the knowledge produced was worth having.

Think further: questions and resources

  • Milgram and Asch both deceived participants about the situation they were actually in, but only Milgram's study caused real distress. Milgram argued that what his experiment revealed justified the cost. Is a study's value ever a legitimate reason to deceive or distress the people inside it, or does that judgement only look reasonable once the finding is already famous?

  • Thibault Le Texier found that Zimbardo had scripted much of the Stanford Prison Experiment's cruelty rather than simply observing it emerge. Zimbardo went on defending the study's central claims until his death in 2024. If a finding turns out to have been staged, does a researcher's continued belief in it change how responsible they are for what it was later used to justify?

  • Ian Hacking argues that a classification does not only describe people, it changes them, through what he calls the looping effect of human kinds. If every classification risks producing the very behaviour it claims only to observe, is there any way for a human scientist to describe a group of people without also altering them?

  • Michel Foucault claims that an examination, medical, psychological, or educational, turns a person into what he calls a "case," something an institution can then manage. Is a diagnosis or a file ever just a neutral record, or does naming someone this way always carry some form of power over them?

  • Esther Duflo built her career on breaking large questions about poverty into small, testable interventions rather than making broader claims about economic systems. Is choosing to ask only small, answerable questions a way of avoiding judgement, or a judgement in its own right about which causes are worth investigating?

  • The IB Psychology guide bans conformity and obedience studies outright, with no exception for how mild the version or how thorough the debrief, even though Asch's line-judging task caused nobody any real harm. If the ban is not really about how much harm a study causes, what is it actually protecting?

Films
For more see my 10 films for the TOK journey page.

🎬  WATCH — Three Identical Strangers (2018)

Tim Wardle

This documentary follows three triplets separated at birth in 1961 by a New York adoption agency and placed with different families, who found each other by chance in 1980. Their separation was arranged, without any family's knowledge, so the psychiatrist Peter Neubauer could secretly study identical siblings raised apart, one case among at least eleven twin pairs and one set of triplets. The data was never published, and the records remain sealed at Yale until 2065. The film connects to Big Idea 1, since nobody studied ever consented, and to BI2, since the knowledge produced has stayed entirely in the institution's control, kept even from the people it was about. My students can watch the complete film here. 

🎬  WATCH — Conformity - Mind Field  (2017)

Vsauce 

This episode from Vsauce recreates Asch's line-judging experiment with a modern participant, then goes further, testing whether people will laugh along at a nonsensical joke simply because everyone else is laughing. Both versions confirm Asch's finding: conformity comes from politeness and a fear of looking foolish, more than from genuine confusion. Mind Field runs to three seasons of similar experiments and is worth exploring further as a series for TOK, since most episodes re-test a classic study with a modern participant. This episode looks at the Stanford Prison Experiment. 

Further reading

📚 READ - Obedience to Authority: An Experimental View, Stanley Milgram, 1974.

Read "Appendix I: Problems of Ethics in Research" (pp.199-201), where Milgram makes his own case for why the participant, not the external critic, must be the ultimate judge of whether a study was acceptable, illustrated with a thought experiment about a snipped-off finger and a real exchange of letters with a fully obedient participant from a 1964 replication who later cited the experiment when applying for conscientious objector status. Relevant because BI1 argues from the outside, what the study did to people who could not consent to it, and this is Milgram's own answer from the inside, written by the person BI1 is arguing against. In the library in TOK Books > Human science.

​

📚 READ - The Lucifer Effect: Understanding How Good People Turn Evil, Philip Zimbardo, 2007.

Read pp.55-56, "Saturday's Guard Orientation," where Zimbardo prints, in his own words, the speech he gave his guards the day before the experiment began, telling them exactly what atmosphere of powerlessness he wanted them to create. Relevant because BI2's whole argument about Zimbardo's finding rests on this passage: it is Zimbardo's own account, published decades after the study, of instructions he spent forty years telling interviewers the guards had invented themselves. In the library in TOK Books > Human science.

​

📚 READ - Science Fictions: Exposing Fraud, Bias, Negligence and Hype in Science, Stuart Ritchie, 2020.

Read the Preface, which sets the Stanford Prison Experiment's reputation, including its use as expert testimony at the Abu Ghraib trials, against what Thibault Le Texier's 2019 examination of Zimbardo's own archives actually found. Relevant because Ritchie's verdict, that the study's results are scientifically meaningless, is the sharpest single sentence behind BI2's argument that a finding does not need to be true to function as power. In the library in TOK Books > Human science.

​

📚 READ - The Social Construction of What?, Ian Hacking, 1999.

Read "Why Ask What?" (pp.25-26) for the child-viewer-of-television example BI2's infographic only has room to summarise, including the 1997 world congress where researchers from Chile and Tunisia joined a field of study that had not existed a generation earlier, then "Madness: Biological or Constructed?" (p.104) for the single sentence that defines the looping effect itself. Relevant because both passages show the same mechanism working at two different scales, a whole field of research, and one sentence of theory. In the library in TOK Books > Human science.

​

📚 READ - Discipline and Punish: The Birth of the Prison, Michel Foucault, 1975 (trans. Alan Sheridan, 1977).

Read "Docile Bodies" for the argument that an examination turns a person into a "case," and "Panopticism" for Bentham's prison design, where a prisoner who can never verify whether the watchtower is occupied behaves as though it always is. Relevant because BI2 compresses both chapters into one paragraph; reading them separately shows the file and the tower are the same argument made twice, once about the record kept on a person and once about the architecture built to watch them. In the library in TOK Books > Human science.

​

📚 READ - Poor Economics: A Radical Rethinking of the Way to Fight Global Poverty, Abhijit Banerjee and Esther Duflo, 2011. Read Chapter 3, "Low-Hanging Fruit for Better (Global) Health?", which works through the bed net pricing question in far more depth than the TED talk in BI3 has time for: whether a $10 insecticide-treated net should be given free, subsidized, or sold at full price, and why the answer changes depending on what you assume about why poor people don't already have one. Relevant because it shows the method BI3 describes actually being applied to a single, specific decision, not just described in the abstract. In the library in TOK Books > Human science.

bottom of page