Web Analytics
top of page
areas of knowledge - human science

Methods and Tools of Human Science

Can you trust a method that was built to find what it was looking for?

The perfect balanced sample.

THE PROVOCATION
Yes, Prime Minister

Yes, Prime Minister (1986-88) is a BBC sitcom, my favourite. It follows Jim Hacker, newly made prime minister, and the two civil servants who actually run his office: Sir Humphrey Appleby, the Cabinet Secretary, endlessly skilled at redirecting decisions toward whatever the civil service prefers, and Bernard Woolley, his junior, caught between loyalty to his minister and his boss. The show satirises the gap between elected politicians and the permanent officials who advise them, and is still quoted in British political journalism today.

In a 1986 episode of Yes, Prime Minister, the civil servant Sir Humphrey Appleby shows his colleague Bernard Woolley how to get whichever survey result you want on the same policy question. Asked first about crime, discipline, and the value of structure in young people's lives, Bernard ends up saying yes to reintroducing National Service. Asked first about the danger of war and the wrongness of forcing people to fight, he ends up saying no to the same policy. Humphrey calls it "the perfect balanced sample." Nothing about Bernard's actual opinion changed between the two runs. Only the questions that came before it did.

​

You have probably taken a survey that worked this way on you, at school, in a customer feedback form, in a political poll shared on social media, without noticing the questions ahead of the one that mattered. This page asks what it takes to trust a piece of human-science evidence once you know how easily the method itself can produce the answer it was built to find.

Big idea 1 - Whose attachment counts as secure?

Reliability asks whether a method gets the same result under the same conditions. Validity asks whether it measures what it claims to measure. Representativeness asks whether the people you tested can stand in for the people you're making claims about. Developmental psychology's most famous test of infant attachment turns out to fail one of these three questions once you take it outside the culture it was built in.

Reliabilty, validity and representativeness.webp

In 1970, the psychologist Mary Ainsworth developed the Strange Situation: a short, filmed procedure in which a baby is left alone with a stranger, then reunited with its mother, and rated on how it responds. A baby who is upset by the separation but calms quickly and returns to play once reunited is classified as securely attached, the healthiest outcome on Ainsworth's scale. This classification has been used in clinical assessments, family court decisions, and parenting research for over fifty years, treated as a measure of something universal about how infants bond.

Heidi Keller, a developmental psychologist who has spent decades studying the Nso farming communities of Northwest Cameroon, argues that the Strange Situation measures something considerably narrower than that. Ainsworth's test assumes a specific model of the infant: an independent individual whose signals a single caregiver should notice and respond to, face to face, on the baby's own schedule. Among the Nso, and in many farming communities Keller has studied, infants are cared for by several people at once, kept in close physical contact for most of the day, and soothed through touch rather than eye contact or verbal exchange. Keller calls this the proximal style, against the distal, face-to-face style assumed by Ainsworth's test and standard in the Western middle-class households the original research was built on. As she puts it, "distal styles are associated with accentuation of individual uniqueness and sticking out whereas the proximal style promotes fitting into the social system." A method built around one culture's idea of a healthy bond, applied to a community that bonds a different way, will not simply get a different result. It risks calling a difference a deficiency.

​

This is both an issue of validity and reliability.  The validity failure is the clearer of the two. The Strange Situation claims to measure something universal, whether an infant is securely attached. What it actually measures is whether an infant responds to brief separation and reunion the way a distally-parented, Western infant would: distress, followed by quick comfort through eye contact and verbal reassurance. A Nso infant, soothed through touch rather than eye contact, and rarely left alone with a stranger in the first place, can fail to produce that exact pattern while still being securely attached by the standard that actually governs Nso caregiving. The test measures one culture's expression of attachment and reports it as attachment in general.

 

The reliability failure sits underneath that one and is easy to miss. Reliability depends on holding conditions constant, "under the same conditions" is the whole test of it. But a stranger, a strange room, and a brief separation are not the same experience for an infant raised inside a Baltimore-style dyadic household as they are for an infant raised inside a Nso compound with several regular caregivers. The procedure looks identical from the researcher's side of the camera. It is not administering an equivalent condition from the infant's side of it. A method can be run exactly the same way twice and still not be reliable, if "the same way" only means the same for the people who designed it.

Big idea 2 - If there is a Hawthorne effect, Hawthorne didn't prove it. 

Hawthorne Effect.webp

Between 1924 and 1927, engineers at Western Electric's Hawthorne Works in Cicero, Illinois, ran a simple experiment: change the lighting in a workroom and measure what happens to output. The result became one of the most repeated stories in social science. Every change increased productivity, brighter light, dimmer light, even light turned down to roughly the level of moonlight. The conclusion drawn from this, and taught for decades afterward under the name the Hawthorne effect, was that people work differently simply because they know they are being watched. Elton Mayo's follow-up studies at the same plant, running until 1932, appeared to confirm it: whatever the researchers changed, rest breaks, working hours, pay, output kept rising.

​

There is an irony built into this story that took eighty years to surface. The original illumination data were never properly analysed at the time and were long believed destroyed. In 2009, the economists Steven Levitt and John List tracked the surviving records down to two library archives and ran the numbers for the first time. What they found did not match the version everyone had been taught. Levitt and List concluded that the dramatic patterns described in textbooks "prove to be entirely fictional." Output did rise over the course of the study, but so did output across the rest of the plant during the same years, and much of the apparent effect disappeared once ordinary factors, the season, the day of the week, whether a holiday had just passed, were properly accounted for. Only a faint, much more modest trace of anything resembling a Hawthorne effect survived the reanalysis.

​

The lesson usually drawn from the Hawthorne studies is that people behave differently once they know they're being watched. The lesson the studies actually teach, once you check the primary data, is a different one: a finding can be taught as settled fact for eighty years without anyone going back to see whether it holds up. That is a reliability failure of its own, not in the workers being studied, but in the field studying them. A method's reliability is not established by how often a finding gets repeated. It is established by whether anyone goes back to the original data and checks.

Big idea 3 -  Blinking or winking?

Reliability and validity, so far, have been questions you could ask of surveys, tests, and experiments: numbers, procedures, results that can in principle be rerun. Observation and interview, the other two methods in your own comparison table, produce a different kind of data, and a different kind of problem.

​

The anthropologist Clifford Geertz, borrowing an example from the philosopher Gilbert Ryle, asks you to picture two boys rapidly contracting the eyelid of one eye. In one boy this is an involuntary twitch. In the other it is a wink aimed at a friend, deliberate, meant for someone in particular, carrying a message under a code both of them understand. Filmed or photographed, the two movements are identical. There is no measurement you could take of the eyelid itself that would tell you which one you were looking at.

​

Geertz uses the example to make a claim about what observation alone can and cannot give a researcher. A human scientist watching a village ritual, a classroom, or a family at dinner is recording a stream of movements that, described on their own, could mean several entirely different things, a twitch, a wink, or even a third boy's parody of somebody else's clumsy attempt at winking, contracting the same eyelid to mean something different again. Getting the description right requires knowing the code the people in front of you are using, not just watching more closely. Geertz calls the method built to handle this "thick description," description that includes the layers of meaning a culture has built around a gesture, not just the gesture itself.

This puts observation and interview in a different relationship to reliability than a survey or an experiment. A thermometer does not need to understand the room it is measuring. A person recording human behaviour does, and two careful observers who disagree about what a gesture means are not necessarily making an error. They may be reading two different codes, or reading the same code with different fluency.

​

Observation carries one more complication, and it connects back to the Hawthorne effect. A visible observer risks the same problem as Mayo's researchers: people act differently once they know they are being watched. The obvious fix is to make the watching invisible, a hidden camera, an observer who blends in rather than announces themselves. But removing the observer's presence to protect the reliability of the data means recording people without their knowledge. Solving the Hawthorne problem this way trades a reliability question for an ethical one. Methods and Tools can tell you a method is unreliable or invalid. It cannot, on its own, tell you whether fixing that is worth doing at the cost of someone's consent.

Is TOK an effective IB subject.webp

The methods and tools of Human Science compared

What happens when the thing you are measuring can tell it is being measured?

A natural scientist's instruments extend what a human eye or hand cannot do alone, a particle accelerator resolves matter, a telescope resolves distance, and none of what they observe changes its behaviour because it has noticed the instrument. A historian's sources, a letter, a court record, an artefact, are already fixed by the time the historian arrives; the evidence cannot react to being read. Human science does not have this luxury. Its subjects notice, and noticing changes them: a survey question can produce the very opinion it claims to record, and a workroom under observation can raise its own output simply because it is being watched, whether or not that particular finding survives a recheck decades later.

​

This is what separates Methods and Tools in Human Science from its counterpart in Mathematics, where the question is what a proof can guarantee once no human is left to check it line by line, and from Natural Science, where the question is what an instrument can extend or distort about a world that does not know it is being measured. Here the question is different again: what a method can still tell you once the person being measured has become, however briefly, a participant in their own measurement, and what you are willing to give up, visibility, consent, or the naturalness of the setting, to get an answer anyway.

​

​Methods and Tools in the Arts raises a related problem from the opposite direction. There, the tools used to examine a work, an X-ray, a cleaning solvent, can multiply the number of things that were once true about it rather than settle a single fact. Human Science’s instruments have the reverse problem: the tool does not multiply the truth, the subject’s awareness of the tool changes what there is to find.

Think further: questions and resources

  • Heidi Keller argues that the Strange Situation measures a specific, Western style of caregiving rather than attachment security itself. If every method for studying humans is built inside some particular culture's assumptions, is a genuinely culture-neutral test of anything psychological even possible in principle, or only a matter of degree?

  • Steven Levitt and John List found that the original Hawthorne data, once finally analysed, barely supported the effect that made the study famous. If a finding this widely taught turned out to rest on an unchecked assumption, how many other results assumed to be "settled science" have never actually been reread against their own original data?

  • Clifford Geertz argues that a wink and a twitch are indistinguishable to an outside observer without knowing the code behind them. Can an observer from outside a culture ever truly learn its codes well enough to interpret behaviour accurately, or does every outside observation stay to some degree a guess?

  • Mary Ainsworth's Strange Situation has been used in family court decisions and clinical assessments for decades, built from a sample of Baltimore families in the 1960s. Should a fifty-year-old classification system still carry this much legal and clinical weight today, or does a method need to be reverified before it keeps being used this way?

  • Hiding an observer, or filming participants without telling them, can remove the Hawthorne effect from a study. Should a method ever be allowed to become more reliable at the direct cost of the people being studied not knowing they are being studied?

  • Sir Humphrey Appleby gets Bernard to give opposite answers to the same underlying question by changing what comes before it. If the order and wording of questions can determine the result this precisely, what would a survey have to look like for you to actually trust its conclusion?

Films
For more see my 10 films for the TOK journey page.

🎬  WATCH — Trobriand Cricket: An Ingenious Response to Colonialism (1976)

Gary Kildea and Jerry Leach

British missionaries introduced cricket to the Trobriand Islands in Papua New Guinea to replace tribal warfare with orderly sport. The islanders kept the bat and ball and changed everything else: war paint, chanting, magic rituals for the bowlers, and matches used to build a leader's political reputation rather than to decide a winner by the rules. Filmed in the early 1970s by an anthropologist and a filmmaker working together, it is close to a documentary illustration of Big Idea 3, the same physical actions meaning something completely different depending on the code the players and their audience share, exactly Geertz's point about the wink and the twitch. Watch the complete film here.

🎬  WATCH — Persona (2021)

Tim Travers Hawkins

This documentary traces the Myers-Briggs Type Indicator from its origins to its current use in hiring decisions, online dating, and workplace management, despite psychologists' long-standing doubts about its scientific basis. One study cited in the film found that almost half of people who retake the test five weeks later get a different result, a reliability failure playing out in real institutions right now, not eighty years in the past. It makes the same point Big Idea 2 makes with the Hawthorne effect: a measure can be trusted for decades, built into real decisions about real people's jobs and lives, without ever being properly checked. My students can watch the complete film here.

🎬  WATCH — The Beginning of Life (2016)

Estela Renner

Filmed across nine countries, including Kenya and Brazil alongside the US, France, and China, this documentary follows early childhood development and the wide range of caregiving practices that shape it, with input from developmental researchers and economists including Nobel laureate James Heckman. It doesn't set out to test Heidi Keller's specific distal/proximal distinction the way a narrower film might, but it shows the same underlying point at a larger scale: attachment and bonding look different across cultures long before any test is applied to them, which is exactly what makes a single culture's model of secure attachment risky to generalise.

Further reading

📚 READ - The Myth of Attachment Theory: A Critical Understanding for Multicultural Societies, Heidi Keller, 2018.

Read the chapter "Cultural blindness of attachment theory" (p.70), where Keller sets out the distal/proximal distinction used in BI1, and stay with the passage describing Manuela Lavelli's study comparing Italian, Nso, and West African migrant mother-infant pairs, which shows the same distinction holding up across a third group rather than just two. Relevant because it's the primary source behind BI1's whole argument, not a supporting reference to it. In the library in TOK Books > Human science.

​

📚 READ - The Interpretation of Cultures, Clifford Geertz, 1973.

Read the opening essay, "Thick Description: Toward an Interpretive Theory of Culture" (p.7), for the full version of the wink/twitch example BI3 only has space to summarise, including the third boy's parody of a wink and the would-be satirist rehearsing in front of a mirror, each adding a further layer of interpretation the page's shorter version leaves out. In the library in TOK Books > Human science.

​

📚 READ - Social Research Methods, Bryman, Clark, and Foster, most recent edition.

Read Chapter 7 in full. The reliability section walks through stability and the test-retest method in more depth than the page has room for, including why researchers rarely actually run these tests despite treating reliability as settled. The validity section extends the weighing-scales example into face validity, criterion validity, and construct validity, the full toolkit BI1's validity argument only samples one piece of. The chapter's discussion of representative sampling in opinion polling connects directly to a case that isn't in this book: the 1936 Literary Digest poll, which surveyed 2.4 million people and still called the US presidential election wrong, while George Gallup's much smaller but properly constructed sample got it right. In the library in TOK Books > Human science.

bottom of page