How Do You Know? Journal Contact

By Manthan Jha · Free to read

How Do You Know?

An interactive book on critical thinking. Fourteen chapters on belief, evidence, numbers, probability, bias, trust and decisions — and every chapter asks you to predict, commit, write and test ideas against your own life.

1 · PredictGuess before you read, and put a confidence number on it.
2 · ReadTrue stories first, then the idea underneath.
3 · AnswerQuizzes explain every option, right or wrong.
4 · JournalWrite real beliefs down, revisit, and update them.
5 · Teach backExplain it simply. That's how you'll know you know it.
Private by design. Everything you write — answers, reflections, your journal — is saved only in this browser, on this device. Nothing is sent to anyone, including the author. Browsers can clear saved data, so use Backup in the menu now and then; Restore brings it back, or moves it to another device. Privacy details.

Questions, corrections or a workshop?

Talk to Manthan directly — on WhatsApp or by email. Sessions for schools, colleges and teams can be arranged.

Chapter 1 · Part I — The Mindset

Belief Isn’t the Enemy

Humans run on inherited belief — and that's how we took over the planet. The real skill isn't believing less. It's knowing which beliefs deserve your trust.

This book begins by questioning the idea that started it. The idea goes like this: the world runs on belief, not critical thinking. Beliefs passed down through generations, like Chinese whispers, have damaged people's ability to think. Thinking is what separates humans from animals, and we need to get back to it.

Part of that is exactly right, and part of it is almost exactly backwards. Sorting out which part is which is the best possible first lesson, because it's the whole skill in miniature: take a belief you care about, and look at it honestly.

Predict Before you read

Researchers gave the same set of thinking tests — memory, spatial reasoning, counting, understanding cause and effect — to adult chimpanzees and to human toddlers about 2½ years old. On problems about the physical world, how did the toddlers do?

How sure are you?60%

About the same. In a well-known 2007 study led by Esther Herrmann and Michael Tomasello, over a hundred chimpanzees, a group of orangutans and around a hundred 2½-year-old children took a big battery of cognitive tests. On the physical-world problems — space, quantities, causality — the toddlers and chimps performed roughly equally.

But on social learning — watching someone solve a problem and copying it, following where someone points, understanding what someone intends — the toddlers were far ahead. Hold on to that. It's the key to this chapter.

1.1The explorers who starved in a pantry

True story · Australia, 1861

In 1860, the Burke and Wills expedition set out to cross Australia from south to north. They were well funded, well equipped and led by confident, educated men. On the way back, low on supplies, they got stranded near a place called Cooper Creek.

The area wasn't empty. The Yandruwandha people lived there, and they lived well. They shared fish and cakes made from nardoo, a seed-bearing fern. The explorers watched, learned to collect the seeds, and made their own flour.

They ate plenty of it. They starved anyway. Burke and Wills died within weeks. The one man who survived, John King, lived because the Yandruwandha took him in and he ate the way they did.

Researchers now think the problem was the preparation. Raw nardoo contains an enzyme that destroys vitamin B1 in the body. The Yandruwandha ground the seeds with water and prepared them in ways that neutralised it. The explorers copied the what — collect, grind, bake — but skipped steps whose purpose they couldn't see. They ate full meals while their bodies ran out of an essential vitamin.

These were intelligent, educated men who reasoned their way to the dangerous version. What they lacked was not intelligence. It was inherited knowledge — the kind nobody could have worked out in a few weeks by thinking hard.

1.2A recipe smarter than the cook

The anthropologist Joseph Henrich, whose book The Secret of Our Success this chapter leans on, gives a sharper example: manioc (cassava), a staple root in the Amazon.

Bitter manioc contains compounds that release cyanide. Traditional processing is long and tedious — scraping, grating, washing, boiling, then leaving the pulp to sit for days before cooking. Now imagine a sensible, independent-minded cook who notices that boiling alone removes the bitter taste. Why waste two days? She simplifies the recipe. Nothing bad happens. Not this week, not this year.

The damage comes slowly: chronic cyanide poisoning builds up over years. By then, nobody can connect the illness to the shortcut. The link between cause and effect is simply too long and too hidden for one person to spot by thinking.

This really happened on a large scale. When the Portuguese carried manioc from South America to Africa, the plant travelled but the full processing tradition often didn't. Henrich notes that chronic cyanide poisoning from poorly processed cassava is still a health problem in parts of Africa, centuries later.

Closer to home

Your grandmother insists rajma must be soaked overnight and cooked thoroughly — never quick-simmered. She's right, and probably couldn't tell you the chemistry. Raw and undercooked red kidney beans contain a lectin (phytohaemagglutinin) that can cause severe vomiting. Proper soaking and a hard boil destroy it. Food-safety agencies warn that slow cooking at low temperature can make undercooked beans worse.

The kitchen rule carried the knowledge. The reason got lost along the way. The rule still works.

Key idea: A tradition can be right for reasons that nobody alive can explain. Practices survive when the reasons behind them are forgotten — so “I don't see why we do this” is not the same as “there's no reason to do this.”

Quiz

Manioc reached Africa, but the full processing tradition often didn't. What's the most important lesson?

1.3Not the smartest animal — the best copier

Go back to your prediction. Toddlers and chimpanzees are roughly matched at solving problems about the physical world. Where human children win by a mile is in learning from others: watching, imitating, trusting, absorbing what the group already knows.

Henrich's argument is that this — not raw individual intelligence — is our species' real advantage. No single human could reinvent fire-making, boat-building, which plants are poisonous and how to make them safe, or how to hunt seals on Arctic ice. Those things were built by thousands of people over thousands of years, each one adding a little and passing it on. He calls this our collective brain.

That forces a correction to the thesis this book started with:

“The reason humans grew over other animals is the ability to think.”— The starting thesis

A better version: humans won because we are extraordinarily good at believing what others tell us and passing it on. Inherited belief isn't a bug in human thinking. It's the engine that built civilisation. You believe the Earth goes round the Sun, that germs cause disease, that your medicine contains what the label says — and almost none of that did you check yourself. You couldn't. Nobody can.

1.4So where does it go wrong?

The Chinese whispers part of the thesis is real. The same machinery that carries wisdom also carries noise, and it can't tell the difference by itself.

True story · 21 September 1995

Before dawn at a temple in New Delhi, word spread that an idol of Ganesha was drinking milk offered on a spoon. By noon, temples in the UK, Canada, the UAE and Nepal were reporting the same thing. Queues stretched for over a mile. Milk sales in Delhi reportedly jumped by more than 30%.

A team from India's Ministry of Science and Technology went to a temple and added food colouring to the milk. The coloured milk didn't disappear into the idol — it was drawn off the spoon by surface tension and coated the stone, running down its front. Capillary action. The test took minutes.

The milk miracle was mostly harmless. The next example wasn't.

True story · India, 2017–2020

Messages spread on WhatsApp warning that child kidnappers and organ harvesters were roaming villages, often with local details added to make them feel real. Mobs attacked strangers they believed matched the description. At least 23 people were killed. In Rainpada, in Maharashtra's Dhule district, five men from a nomadic community were beaten to death on 1 July 2018.

In almost every place where these killings happened, police found that no child abductions had been recorded in the previous three months. WhatsApp responded by labelling forwarded messages and limiting how many chats a message could be forwarded to at once.

Notice what these stories have in common, and how they differ from manioc. In both, the claim was checkable, cheaply, and the stakes were high — yet people acted on it at many hops away from any actual evidence. Nobody in the chain had seen a kidnapping. Most people who saw the milk miracle had only heard about it.

Key idea: There are two opposite failures, not one. Trusting where you should check (the milk miracle, the lynching rumours). And checking-and-discarding where you should trust (the manioc shortcut, the explorers). Critical thinking isn't “believe less”. It's choosing the right mode for each belief.

1.5Chesterton's fence

In 1929, the writer G. K. Chesterton described a reformer who comes across a fence or gate built across a road for no obvious reason. The modern reformer says, “I don't see the use of this; let us clear it away.” Chesterton's reply:

“If you don't see the use of it, I certainly won't let you clear it away. Go away and think. Then, when you can come back and tell me that you do see the use of it, I may allow you to destroy it.”— G. K. Chesterton, The Thing (1929)

This is not an argument for keeping every fence. It's an argument for finding out why before you remove it. Sometimes you'll find the reason is dead — a rule made for a world that no longer exists — and you can remove the fence with a clear conscience. Sometimes you'll find the fence was holding back a flood.

You meet fences at work all the time. An old procedure at a property that everyone calls pointless paperwork. A supplier nobody is allowed to switch. A check-in step that seems to waste guests' time. Before you tear one out, find the person who built it, or the incident that caused it. It might be a legal requirement, or the scar of a past disaster.

1.6The tool: Belief Triage

Here's the first tool in your kit. Hospitals use triage to decide who needs attention first. You can do the same with beliefs, because you can't examine them all.

Tool 1

Belief Triage

  1. Stakes — What happens if this is wrong? Who gets hurt, and how badly?
  2. Hops — How many people sit between me and whoever actually saw the evidence? Did anyone in the chain see it at all?
  3. Cost of checking — Can I test it cheaply (the coloured-milk test)? Or would the effects take years to show (manioc)?

Then choose a mode. High stakes + cheap to check → check before you act or forward. High stakes + slow to check + a long-surviving tradition → Chesterton: find out why first; lean on the tradition meanwhile. Low stakes → hold it lightly and move on.

Cheap to checkSlow or costly to check
High stakesCheck before acting or sharing
(kidnapper rumours, miracle cures, “RBI announces…” forwards)
Chesterton's fence: find out why; don't casually discard
(food preparation, safety rules, long-standing SOPs)
Low stakesCheck if you're curious
(trivia, “did you know…” facts)
Hold lightly; don't pick a fight
(harmless customs, superstitions with no victims)
Triage Scenario 1

A forward in your family group: “Doctors are hiding this — warm water with lemon every morning cures cancer. Share with everyone you love!”

Triage Scenario 2

A new chef at your property wants to scrap the kitchen's rule that rajma is always soaked overnight and pressure-cooked fully. “It's just an old habit — a slow simmer is fine and saves time.”

Triage Scenario 3

An older relative says you must never cut your nails after sunset.

1.7Where this leaves the thesis

The original thesis had a real insight and a wrong villain. The insight: much of what people believe arrives through long chains of retelling, and those chains distort. The wrong villain: belief itself.

A sharper diagnosis is this. We evolved to trust what our community passes down, in small groups where the chain was short and people could see who knew what. Today the chains are enormous, fast and anonymous. A WhatsApp forward can travel a thousand hops in a day, arriving with the same feeling of “someone I trust sent this” as your grandmother's recipe. Our trust instincts weren't built for that.

So the mission isn't to replace belief with thinking. It's to help people become better at deciding what to trust: which beliefs to hold firmly, which to hold lightly, which to check, and which fences to understand before tearing down.

That version is more accurate. It's also far easier to teach. You can't win many people over by telling them their family's beliefs are corrupted. You can by showing them how to sort beliefs into the right boxes — and letting them do the sorting themselves.

Reflect Rewrite the thesis

Here's the thesis this book started with: “The world runs on belief, not critical thinking… belief passed down like Chinese whispers has messed up people's capability to think.” Rewrite it in your own words, in 2–4 sentences, so it survives this chapter.

Compare with one possible version

“Humans became powerful by trusting and passing on what others know — inherited belief is how civilisation was built. But our trust instincts evolved for small communities with short chains, and today beliefs travel through huge, fast, anonymous chains that carry noise as easily as wisdom.”

“The problem isn't that people believe; it's that most of us were never taught how to decide what deserves belief. I want to teach that.”

Apply Your first triage

Pick one belief you inherited from family or community and have never checked. Run the Belief Triage: stakes, hops, cost of checking. Which mode does it land in?

Pick something that affects real decisions — money, health, career, how you treat people — not trivia.

See a worked example

Belief: “A government job is the only truly safe career.”

Stakes: High — it shapes career choices, and family pressure around them.

Hops: Several. It was based on my parents' generation's experience, in a very different economy. I've never looked at current evidence myself.

Cost of checking: Moderate — data on job security, pay growth and exam odds is public, and I can talk to people in both paths.

Mode: Check — but with Chesterton in mind. The reason behind the belief (security in a scarce economy) was real. The question is whether it still holds, and for whom.

Teach it back

Explain to a 15-year-old, in four sentences or fewer, why “just question everything” is bad advice — and what to do instead.

If you can't say it simply, that's where your understanding still has gaps. Rewrite until a teenager would get it.

Compare with a model answer

“You can't check everything — nobody can — and a lot of what you were taught is true for reasons nobody told you. But some of what reaches you is just rumour that got passed along. So before you trust or share something, ask three things: what happens if it's wrong, did anyone in the chain actually see it, and how easy is it to check? Check the risky, checkable stuff; respect old rules until you know why they exist; and don't waste energy fighting about harmless ones.”

Field work This week

Takeaways

  • Humans didn't win by out-reasoning other animals. We won by learning from each other and passing it on. Inherited belief is the engine, not the bug.
  • Traditions can be smarter than the people who carry them. Practices survive after their reasons are forgotten.
  • The same chains carry noise, and modern channels move it faster and further than any check.
  • There are two failures: trusting what you should check, and discarding what you should understand first. The skill is choosing the right mode.
  • Belief Triage: stakes, hops, cost of checking.

Sources and further reading

  • Joseph Henrich, The Secret of Our Success (Princeton, 2015) — Burke and Wills, manioc processing, the collective brain.
  • Herrmann, Call, Hernández-Lloreda, Hare & Tomasello, “Humans have evolved specialized skills of social cognition,” Science (2007).
  • G. K. Chesterton, The Thing (1929), chapter “The Drift from Domesticity.”
  • “Ganesha drinking milk miracle” — Wikipedia summary of 1995 reporting, including the Ministry of Science and Technology coloured-milk test.
  • “Indian WhatsApp lynchings” — Wikipedia summary of 2017–2020 reporting, including the Rainpada (Dhule) killings.
  • US FDA, Bad Bug Book — entry on phytohaemagglutinin in red kidney beans.

Chapter 2 · Part I — The Mindset

Scouts and Soldiers

Intelligent people don't usually believe wrong things because they can't reason. They believe them because part of their reasoning is quietly working for the other side.

Chapter 1 was about beliefs that reach you from outside. This chapter turns the lens inward. Even with perfect information, your own mind has a motive — and it isn't always accuracy.

This chapter follows Julia Galef's The Scout Mindset (2021). Her central image: when you reason, you are sometimes a soldier, defending a position against threats, and sometimes a scout, whose only job is to map the terrain as it really is.

Predict Before you read

In a classic 1981 study, American drivers were asked whether they were more skilful than the median driver in the group. Roughly what share said yes?

How sure are you?60%

Over 90%. In Ola Svenson's study, about 93% of the American participants rated themselves as more skilful than the median — something at most half of them can be. (A Swedish group was more modest, at around 69%. Still too high.)

Nobody in that study felt biased. Each person felt they were simply reporting the truth about their driving. That's the point of this chapter: feeling objective is not evidence that you are.

2.1The officers who knew they were right

True story · Paris, 1894

French military intelligence found a torn-up memo in the wastebasket of the German embassy. Someone in the French army was offering military secrets to Germany. Suspicion quickly fell on Captain Alfred Dreyfus, an artillery officer on the General Staff. He was Jewish, in an army where antisemitism was common, and colleagues said he was unpleasant and aloof.

His handwriting was compared with the memo. Experts disagreed, and some pointed out clear differences. One argued the differences proved Dreyfus had disguised his handwriting on purpose. Investigators searched his home and found nothing incriminating. They concluded he must have hidden the evidence carefully — which, to them, was what a traitor would do.

Dreyfus was convicted, publicly stripped of his rank, and sent to Devil's Island, a prison colony off South America.

In 1896, the new head of counter-intelligence, Colonel Georges Picquart, came across evidence pointing to a different officer, Major Esterhazy, whose handwriting matched the memo. Picquart was no admirer of Dreyfus, and he shared the prejudices of his class. But he followed the evidence. His superiors told him to drop it. He was sent away to North Africa, and later arrested and imprisoned.

Dreyfus was finally cleared in 1906. Picquart was reinstated and went on to become France's Minister of War.

What's striking is not that the officers lied. Most of them believed their reasoning. Every piece of evidence — a matching handwriting, a non-matching handwriting, an empty house — got interpreted as more proof of guilt. From the inside, it felt exactly like thinking.

Psychologist Thomas Gilovich describes the switch this way. When we want to believe something, we ask, “Can I believe this?” and look for a single reason that permits it. When we don't want to believe something, we ask, “Must I believe this?” and look for a single reason to escape. Same mind, two different standards of evidence, and no sense of switching between them.

Key idea: Soldier mindset is reasoning as defence — searching for support for what you already want to be true. Scout mindset is the motivation to see things as they are, not as you wish they were. Everyone does both. The goal is to notice which one is driving.

2.2Why the soldier is so attractive

If soldier mindset leads us wrong, why is it everywhere? Because it pays — immediately. Galef lists six benefits it gives us:

Emotional benefitsSocial benefits
Comfort — avoiding painful truths (“the business is fine”)Persuasion — it's easier to convince others if you've convinced yourself first
Self-esteem — protecting the story that you're capable and rightImage — looking confident, loyal, successful
Morale — staying motivated by believing success is certainBelonging — holding the beliefs that keep you inside your family, community or team

Every one of these is real. That's why lecturing people about their biases rarely works: you're asking them to give up comfort, pride and belonging in exchange for an abstract thing called accuracy.

Galef's counter-argument has two parts. First, the benefits of soldier mindset arrive now, but the costs — bad decisions, missed warnings, wasted years — arrive later, so we systematically underrate them. Second, you can usually get the same benefits honestly. You can find comfort in having a plan, rather than in pretending there's no problem. You can build morale by knowing a bet is worth taking even at 40% odds, rather than by believing the odds are 100%.

That second point matters for your teaching. You won't change minds by attacking the benefit someone gets from a belief. You change minds by offering another way to get it.

2.3Signs you're actually a scout

Almost everyone believes they're a scout. So Galef suggests looking at behaviour, not feelings. Be honest — nobody else will see this list.

Self-audit In the last few months, have you…

Few ticks isn't a failure — it's an honest starting point. The people who tick all six easily are sometimes the ones to worry about.

Quiz

Which of these is the strongest evidence that someone has scout mindset?

2.4Five tests that catch the soldier

Motivated reasoning is invisible from the inside, so you need tricks that make it visible. Galef's five tests all do the same thing: they change one detail of the situation and check whether your judgement changes too. If it does, the detail was doing the reasoning, not the evidence.

1 · The double standard test

Are you judging this person, or group, by a standard you wouldn't apply to others? When your manager is late, he's “busy with important things”. When a junior staff member is late, they're “careless”. Swap the people and see whether your verdict survives.

2 · The outsider test

Imagine someone with no emotional stake looking at your situation. What would they do? In 1985, Intel was losing its memory-chip business to Japanese competitors. Memory chips had been Intel's identity. CEO Andy Grove asked his chairman, Gordon Moore: if we got kicked out and the board brought in a new CEO, what would he do? Moore answered at once: he'd get us out of memories. Grove's reply was, in effect: then why don't we walk out the door, come back in, and do it ourselves? They did, and Intel became the microprocessor company.

3 · The conformity test

Imagine the people whose opinion you respect suddenly announced the opposite view. Would you still hold yours? If your confidence would collapse, it may have been resting on them, not on the evidence.

4 · The selective sceptic test

Imagine the evidence pointed the other way. A study says your preferred diet works, and you share it. How closely would you have checked the same study if it had said the diet didn't work? Apply the scrutiny you'd use for unwelcome evidence to welcome evidence too.

5 · The status quo bias test

Imagine your current situation weren't the default. Would you actively choose it? If you weren't already running this service, living in this city, or keeping this supplier — would you start today? If the answer is no, “it's how things are” may be doing the work.

Tool 2

The Five Bias Tests

  1. Double standard — Would I judge it the same way if it were the other person, party or group?
  2. Outsider — What would someone with no stake in this do?
  3. Conformity — Would I still believe this if the people I respect believed the opposite?
  4. Selective sceptic — How hard would I check this evidence if it said the opposite?
  5. Status quo — If this weren't already the default, would I choose it?

Each test changes one detail and checks whether your judgement moves. If it does, the detail was doing the thinking. (Julia Galef, The Scout Mindset)

Which test? 1 of 4

A property has lost money for three years. You ask yourself: “If I didn't already own this, would I buy it today at this price?”

Which test? 2 of 4

A news outlet publishes an investigation criticising a leader you support. You dismiss it as biased. Last month you shared an investigation from the same outlet that criticised the other side.

Which test? 3 of 4

When a senior colleague misses a deadline, you think “she's overloaded”. When a new hire misses one, you think “he's not serious”.

Which test? 4 of 4

Everyone in your friend circle is sure a particular stock can only go up, and so are you. You imagine they all sold out tomorrow and said it was overhyped.

2.5How sure are you, really?

Scouts think in shades of grey. Instead of “this will work”, a scout says “I'd put this at about 70%”. But most of us have never practised putting numbers on beliefs, and our words — “definitely”, “probably”, “I'm sure” — hide a lot.

Galef suggests the equivalent bet test. Take a claim you feel sure about. Imagine two bets, each paying the same prize. Bet A wins if your claim turns out true. Bet B wins if you draw a red ball from a bag. Change the number of red balls until you can't decide between the two bets. That number is your real confidence.

Try it The equivalent bet

Imagine you win ₹10,000 either way. Bet A: you win if your claim comes true. Bet B: you win if you draw a red ball from this bag. Move the slider until you honestly can't choose between them.

Red balls

If you'd still rather take Bet A with 9 red balls in the bag, you're above 90% sure. If you'd switch to the bag at 6, you're nearer 60%. Most people who say “definitely” find they switch somewhere around 7 or 8.

Tool 3

The Equivalent Bet

  1. State the claim precisely (what, by when).
  2. Bet A: win a prize if the claim is true. Bet B: win the same prize by drawing red from a bag of 10.
  3. Adjust the red balls until you're torn. That's your confidence.
  4. Write the number down. Later, check how often your “80%” beliefs came true.

Turns vague certainty into a number you can be scored on. Your journal's calibration table does the scoring.

2.6Hold your identity lightly

The beliefs that are hardest to examine are the ones that have become part of who you are: your politics, your religion, your profession's way of doing things, even “I'm a critical thinker.” Once a belief is part of your identity, evidence against it feels like an attack on you, and the soldier takes over.

The essayist Paul Graham suggested keeping your identity small. Galef's gentler version is to hold identities lightly: you can still be an entrepreneur, a Hindu, a liberal or a conservative, but treat these as facts about you rather than flags to defend. The test is simple: can you hear a good argument against your side and feel curious rather than threatened?

One practical tool is the Ideological Turing Test, a term coined by the economist Bryan Caplan. Can you explain a view you disagree with so well that people who hold it would say, “Yes, that's exactly why I believe it”? If you can only produce a caricature, you don't yet understand what you're rejecting.

Tool 4

The Ideological Turing Test

  1. Pick a view you disagree with.
  2. Write its strongest case, in the words its best defenders would use.
  3. Ask: would a thoughtful believer recognise their view in this — and nod?
  4. Only then write your disagreement.

Also the most persuasive thing you can do: people listen to those who clearly understand them.

Try it Ideological Turing Test

Choose a view you disagree with and care about. Write the strongest case for it in one paragraph — as its best defenders would put it, not as you'd summarise it to mock it.

Why this is hard, and what good looks like

A weak attempt describes the other side's motives (“they believe this because they're scared / stupid / paid”). A strong one describes their reasons and values: what they're protecting, what evidence they find convincing, and what they fear would happen if they're wrong.

A good check: if you showed your paragraph to someone who holds the view, would they want to share it? If you can't imagine that, keep revising.

Apply Run the five tests

Take one belief you hold strongly — about your business, a person, or the world. Run all five tests on it and note where your judgement wobbled.

See a worked example

Belief: “Our guests choose us mainly for the location.”

Double standard: When a competitor near us does well, I credit their marketing, not their location. Wobble.

Outsider: A new manager would look at reviews and booking sources before agreeing. I haven't.

Conformity: The whole team says it, which is partly why I believe it.

Selective sceptic: I'd read a review mentioning the location as proof; I skip reviews that praise the food.

Status quo: It's the story we've always told.

Result: Confidence drops from 90% to about 60%. Next step: tag the last 100 reviews by what guests actually praise.

Teach it back

Explain scout versus soldier mindset using a cricket analogy, in five sentences or fewer.

Compare with a model answer

“Think of a captain with a limited number of DRS reviews. The soldier captain reviews every decision that goes against his team, because he wants it to be not-out — and he wastes his reviews. The scout captain reviews only when he genuinely thinks the umpire got it wrong, even if it's his own star batter who's out. Both are loyal to the team, but only one is trying to see what actually happened. Over a season, the scout wins more reviews — and more matches.”

Field work This week

Takeaways

  • Soldier mindset defends what you want to be true; scout mindset tries to see what is true. Everyone does both.
  • Soldier mindset persists because its benefits (comfort, pride, belonging) arrive now and its costs arrive later.
  • Feeling objective proves nothing. Look at behaviour: when did you last change your mind?
  • The five tests work by changing one detail and seeing if your judgement moves.
  • Put numbers on beliefs, and hold your identity lightly enough that evidence doesn't feel like an attack.

Sources and further reading

  • Julia Galef, The Scout Mindset: Why Some People See Things Clearly and Others Don't (Portfolio, 2021) — the Dreyfus story, the six benefits, the five tests, the equivalent bet.
  • Thomas Gilovich, How We Know What Isn't So (1991) — “Can I believe it?” versus “Must I believe it?”
  • Ola Svenson, “Are we all less risky and more skillful than our fellow drivers?” Acta Psychologica (1981).
  • Andrew S. Grove, Only the Paranoid Survive (1996) — the Intel memory-chip decision.
  • Paul Graham, “Keep Your Identity Small” (essay, 2009); Bryan Caplan, “The Ideological Turing Test” (2011).

Chapter 3 · Part I — The Mindset

Anatomy of a Claim

Before you can judge whether something is true, you have to know exactly what is being said, and why. Most bad reasoning is never refuted. It just never gets taken apart.

Chapters 1 and 2 were about attitude: when to trust, and noticing when you want something to be true. This chapter gives you the basic anatomy you'll use in every chapter that follows: claims, arguments, assumptions, and how to tell whether reasons actually support a conclusion.

Predict Before you read

Is this argument logically valid — does the conclusion have to follow from the premises?

“All flowers need water. Roses need water. Therefore, roses are flowers.”

How sure are you?70%

Not valid. The conclusion is true — roses are flowers — but it doesn't follow from the premises. Here's the same structure with different words: “All cats need water. Fish need water. Therefore, fish are cats.” Needing water doesn't make you a flower any more than it makes you a cat.

If you said “valid”, you have plenty of company. Psychologists call this belief bias: in classic experiments by Jonathan Evans and colleagues, people accept invalid arguments far more often when the conclusion sounds believable. We judge the conclusion and assume the reasoning was fine. This chapter trains you to look at the two separately.

3.1What kind of claim is it?

A claim is any statement that could be true or false, or better or worse supported. Claims come in different kinds, and each kind is tested differently. Mixing them up is one of the most common sources of pointless arguments.

KindExampleHow it's settled
Factual“Bookings fell 20% in August.”Evidence and measurement
Causal“The price rise caused the drop.”Evidence, plus ruling out other causes (Chapter 7)
Prediction“Bookings will recover by December.”Waiting and checking — so write it down with a date
Value“Cheap stays are bad for the brand.”Reasoning about what matters, and to whom
Policy“We should reverse the price rise.”Facts and values together

Notice how a single business conversation can move through all five in two minutes. When two people seem to disagree about a policy, check whether they actually disagree about a fact, a cause, a prediction or a value. Often they agree on the facts and are arguing about values without realising it, or the other way round.

Quiz

Two partners argue for an hour about whether to open a second property. One says “We'll fill it within a year.” The other says “Even if we fill it, I don't want to spend my weekends managing staff.” What's actually going on?

3.2Pin the claim down

Vague claims are almost impossible to refute, which is exactly why they spread. “This oil is good for your health.” Good how? For whom? Compared with which other oil? How much of it? A claim that can't be wrong can't tell you anything either.

Before you evaluate a claim, pin it down. You'll often find that the pinned version is either obviously checkable or quietly much weaker than it sounded.

Tool 5

Pin the Claim

  • What exactly? Replace vague words (“good”, “ruining”, “works”) with something observable.
  • For whom? Everyone, most people, some people, people like me?
  • How much? A 1% change or a 50% change?
  • Compared to what? Compared with doing nothing, or with the alternative?
  • By when? Especially for predictions.
  • How would we know? What observation would show it's wrong?

If nobody can say what would show the claim is wrong, it isn't really saying anything yet.

Try it Pin it down

Pin down this claim: “Social media is ruining children.” Rewrite it as one or two specific, checkable claims.

Compare with some pinned versions

“Teenagers who spend more than three hours a day on social media report more anxiety and depression than those who spend under one hour.” (Factual — and checkable, though it says nothing yet about cause.)

“For 13–16-year-olds, cutting social media use to under an hour a day for a month improves self-reported mood compared with a group that doesn't cut back.” (Causal — much harder to establish, but now you know what evidence would count.)

Notice how different these are. The original sentence could mean either — or anything. Much of the public argument about this topic is people defending different pinned versions without noticing.

3.3Arguments: reasons that point to a conclusion

An argument is a set of statements where some — the premises — are offered as reasons to accept another — the conclusion. Words like so, therefore, which means, that's why usually introduce a conclusion. Words like because, since, given that usually introduce a premise.

In real life, arguments come out messy and half-stated. The first job is to reconstruct them: write out the premises and conclusion as clearly as you can. Here's an example you might hear in any business.

Example · A team meeting

“Our bookings dropped 20% in August right after we raised prices. The price rise killed demand. We need to reverse it.”

Reconstructed, it looks like this:

  1. P1. Bookings fell 20% in August.
  2. P2. The price rise came just before the drop.
  3. C1. So the price rise caused the drop.
  4. C2. So we should reverse the price rise.

Written out like this, you can already see the gaps. P1 and P2 might both be true, and C1 still doesn't have to follow. And C2 needs more than C1.

3.4Hidden assumptions

Almost every real argument leans on premises nobody says out loud. These hidden assumptions are where most weak reasoning lives, because nobody examines what nobody states.

To get from P1 and P2 to C1, the argument quietly assumes:

  • Nothing else changed in August — no monsoon, no new competitor, no bad review, no change in how the listing appeared on booking sites.
  • A 20% drop is unusual — bigger than the normal month-to-month swing. (Was August last year also down?)

And to get from C1 to C2, it also assumes:

  • The bookings lost cost more than the extra margin earned on the bookings kept. A 20% drop in bookings with a 30% higher price can still mean more profit.

The test for a hidden assumption is simple: if this were false, would the conclusion still stand? If not, the argument depends on it, so it deserves to be checked.

Quiz Find the hidden assumption

“He graduated from a top engineering college, so he'll be an excellent product manager.” Which hidden assumption does this argument need most?

3.5Valid, sound, strong

Three words let you say precisely what's wrong with an argument.

  • Valid: if the premises were true, the conclusion would have to be true. Validity is about structure only — the roses argument failed here.
  • Sound: valid and the premises actually are true. This is what a deductive argument needs to prove its conclusion.
  • Strong: most everyday arguments aren't meant to be airtight. They're inductive — the premises make the conclusion probable. We call these strong or weak rather than valid or invalid.

So there are two separate ways to attack any argument: challenge a premise (“that isn't true”) or challenge the link (“even if that's true, it doesn't follow”). Keeping them apart stops you from arguing past each other.

Quiz 1 of 3

“If it rained last night, the road would be wet. The road is wet. So it rained last night.”

Quiz 2 of 3

“All politicians are corrupt. Meera is a politician. So Meera is corrupt.”

Quiz 3 of 3

“Most graduates from this college get strong job offers. Ravi graduated from this college. So Ravi probably got a strong job offer.”

3.6Fallacies are questions, not labels

You've probably seen lists of logical fallacies. They're useful, but they're easy to misuse. Shouting “ad hominem!” wins nothing if you can't explain why the reasoning fails. And spotting a fallacy doesn't prove the conclusion is false — a bad argument can have a true conclusion, as the roses showed. Treating a fallacy as proof the other side is wrong is itself a fallacy, sometimes called the fallacy fallacy.

A better approach is to turn each fallacy into the question it's warning you to ask.

Tool 6

The Questions Behind the Fallacies

  • Ad hominem — Is the attack on the person actually relevant to whether the claim is true?
  • Straw man — Is this the view they actually hold, or a weaker version that's easier to knock down?
  • False dilemma — Are there really only two options?
  • Appeal to authority — Is this person an expert on this specific question?
  • Bandwagon — Does the number of people who believe it tell us whether it's true?
  • Post hoc — Did it happen because of this, or just after it?
  • Hasty generalisation — How many cases is this based on, and are they typical?
  • Slippery slope — Is each step in the chain actually likely?

Ask the question; don't just shout the label. A flawed argument can still reach a true conclusion.

Quiz Spot the question

An ad shows India's most popular cricketer saying: “I start every day with this health drink. That's my secret.” Which question matters most?

Quiz Spot the question

In a family discussion: “Either you take the government job, or you'll be struggling your whole life.”

3.7Steelman before you judge

In Chapter 2 you met the Ideological Turing Test. Its everyday cousin is the steelman: before you criticise an argument, rebuild it in its strongest reasonable form. Fill in the most plausible hidden assumptions, not the silliest ones. Then judge that version.

There are two reasons to do this. First, if you can defeat the strongest version, you've actually learned something; defeating the weakest version teaches you nothing. Second, people notice when they're being understood. Nobody changes their mind for someone who has just misrepresented them.

3.8Putting it together: the Claim X-Ray

Tool 7

The Claim X-Ray

  1. Pin the claim — What exactly is being said, and what kind of claim is it?
  2. Find the reasons — What's offered in support? Write out premises and conclusion.
  3. Surface the hidden assumptions — What must also be true? Would the conclusion survive if each were false?
  4. Check the link — Even if the reasons are true, do they get you to the conclusion? Valid, strong, or weak?
  5. Steelman, then judge — Build the best version, then decide. Put a confidence number on your verdict.

Works on forwards, ads, headlines, pitches, and your own plans.

Try it on the kind of message that arrives in family groups every week. This one is a composite of several that circulate widely.

Forwarded many times

🚨 SHOCKING!! A study by top scientists has PROVEN that people who sleep with their mobile phone near their head are 5 TIMES more likely to get cancer! 5G radiation is the reason. That's why cancer cases in India are rising every year. Doctors don't want you to know this. Keep your phone in another room and FORWARD to everyone you love 🙏

Claim X-Ray Your turn
Confidence the main claim is true50%
Compare with a model X-ray

1 · Claims. (a) Factual/causal: sleeping with a phone near your head makes you five times more likely to get cancer. (b) Causal: 5G radiation is the mechanism. (c) Factual/causal: this is why cancer cases in India are rising. (d) A conspiracy claim: doctors are hiding it. “Cancer” is vague — which cancer? Five times more likely than whom?

2 · Reasons. An unnamed study by unnamed “top scientists”, and the fact that recorded cancer cases are rising.

3 · Hidden assumptions. The study exists and measured this. Rising cancer cases aren't explained by other things — a growing and ageing population, and better diagnosis and record-keeping (more cancer is found when more people are screened). Phones and 5G arrived before the rise, and the timing fits. Thousands of doctors worldwide are keeping a secret.

4 · The link. Weak. “Cases rose after phones spread” is post hoc reasoning — many things rose in the same period. An unnamed study can't be checked. And “doctors don't want you to know” makes the claim immune to evidence: any expert who disagrees becomes part of the cover-up.

5 · Steelman. It's reasonable to want evidence on the health effects of something people keep near their heads for hours a day. In 2011 the WHO's cancer agency (IARC) classified radio-frequency fields as “possibly carcinogenic” (Group 2B) because the evidence was limited and uncertain. In 2024, a WHO-commissioned systematic review of studies up to 2022 (Karipidis and colleagues) found no link between mobile phone use and brain cancer. Some scientists dispute that review, so the honest summary is: the question has been studied a lot, and there's no good evidence for the dramatic claim.

Verdict. The specific “five times more cancer, caused by 5G” claim is unsupported and very likely false. Confidence it's true: around 2–5%. Don't forward. If a family member is worried, the steelman gives you a respectful way to respond.

Teach it back

Explain the difference between “the conclusion is true” and “the argument is good” to someone who's never studied logic. Use your own example.

Compare with a model answer

“My uncle says, ‘India will win the match because I'm wearing my lucky shirt.’ India wins. His conclusion was true — but his reason was nonsense, and it'll let him down next time. A good argument is one where the reasons actually lead to the conclusion. A true conclusion can come from bad reasoning by luck, and a well-reasoned conclusion can still turn out wrong. So judge the reasons, not just whether the answer matches what you already believe.”

Field work This week

Takeaways

  • Identify the kind of claim first: factual, causal, prediction, value, or policy. Each is settled differently.
  • Pin vague claims down until you know what would show them wrong.
  • Reconstruct arguments into premises and conclusion; most weakness hides in unstated assumptions.
  • Two separate attacks: challenge a premise, or challenge the link. A true conclusion doesn't make an argument good.
  • Use fallacies as questions, steelman before judging, and run the Claim X-Ray on anything that matters.

Sources and further reading

  • Evans, Barston & Pollard, “On the conflict between logic and belief in syllogistic reasoning,” Memory & Cognition (1983) — belief bias.
  • Anthony Weston, A Rulebook for Arguments — a short, practical guide to building and reconstructing arguments.
  • IARC (WHO), press release on radio-frequency electromagnetic fields, Group 2B classification (2011).
  • Karipidis et al., systematic review on mobile phone use and cancer, Environment International (2024), commissioned by the WHO.

Chapter 4 · Part II — Reasoning & Evidence

Why We Reason

If reasoning evolved to find the truth, it's strangely bad at the job. But if it evolved for arguing with other people, it's surprisingly good. That difference changes how you should learn — and how you should teach.

Part I was about attitude. Part II is about machinery: how arguments, evidence, numbers and causes actually work. It starts with the machine itself. What is human reasoning for?

This chapter follows The Enigma of Reason (2017) by the cognitive scientists Hugo Mercier and Dan Sperber. Their answer is surprising, and it's the single most useful idea in this book for anyone who wants to teach.

Predict Try the puzzle

Four cards each have a letter on one side and a number on the other. You can see:
E  K  4  7
Rule: “If a card has a vowel on one side, it has an even number on the other.” Which cards must you turn over to find out whether the rule is being broken?

How sure are you?60%

E and 7. E might have an odd number behind it, which would break the rule. 7 might have a vowel behind it, which would also break it. The 4 doesn't matter: the rule says nothing about what must be behind an even number. K doesn't matter either.

This is the Wason selection task, one of the most studied puzzles in psychology. When people try it alone, fewer than 1 in 10 usually get it right. Most choose “E and 4” — they look for cards that could confirm the rule instead of cards that could break it.

Now the twist. In a 1998 study by David Moshman and Molly Geil, 9% of people working alone solved it. But 70% of small groups that discussed it solved it — and in groups formed from people who had already tried it alone, 80% succeeded, including groups in which nobody had solved it individually. Same people, same puzzle. What changed was the argument.

4.1The enigma

For centuries, reason was treated as our noblest faculty — the thing that lets us rise above instinct and see the truth. But experiments since the 1960s show that reasoning is full of flaws. We look for evidence that supports what we already think. We find it easy to invent justifications for our choices after the fact. We're persuaded more by conclusions we like than by logic, as you saw with the roses in Chapter 3.

The most consistent flaw has a name: myside bias. When we reason, we overwhelmingly produce arguments for our own side. Ask someone to think about a question and they'll mostly list reasons supporting their first instinct, and very few against.

If reasoning were designed to help lone individuals discover the truth, this would be a terrible design. A truth-seeking tool should look hardest for the reasons you're wrong. So either evolution did a remarkably bad job, or we've misunderstood what reasoning is for.

4.2Reasoning is for arguing

Mercier and Sperber's answer is the argumentative theory of reasoning. Reasoning didn't evolve to help us think alone. It evolved for social life: to produce reasons that justify ourselves and convince others, and to evaluate the reasons other people give us.

Seen that way, the flaws start to make sense:

  • We're biased and lazy when producing arguments. If your job is to convince someone, looking for reasons you're wrong is wasted effort — the other person will supply those. So you produce reasons for your side, quickly and cheaply.
  • We're sharp when evaluating other people's arguments. If you accepted every argument you heard, anyone could manipulate you. So when someone else argues, you check carefully.
Experiment · 2016

Emmanuel Trouche, Hugo Mercier and colleagues asked people to solve logic problems and write down their reasoning. Later, participants were shown the same problems with answers and arguments supposedly written by other people, and asked whether they wanted to change their answer.

Here's the trick: one of those “other people's” arguments was actually the participant's own, with the answer swapped in. About half the participants didn't notice. And those who didn't notice rejected their own argument more than half the time — the moment it looked like someone else's, it suddenly seemed unconvincing.

The researchers called this the selective laziness of reasoning. We aren't bad at spotting weak arguments. We're bad at spotting weak arguments when they're ours.

Key idea: Reasoning is a social tool. Alone, it mostly defends what you already think. In a real exchange — where people who disagree evaluate each other's reasons — it's powerful. The unit of good reasoning is not the individual. It's the conversation.

Quiz

According to the argumentative theory, why are we lazy about checking our own arguments?

4.3When groups get smarter

Go back to the card puzzle. Why did groups do so much better? Because when someone in the group says “E and 4”, someone else asks, “Why the 4? What would it prove?” The person who spotted the 7 can explain. Good arguments are easy to recognise once someone says them out loud, even if you couldn't have produced them yourself.

This is a division of cognitive labour. Each person does the cheap part — offering reasons for their view. Everyone does the valuable part — checking each other's. For problems where the right answer can be demonstrated, the truth tends to win.

But this only works under certain conditions:

  • There has to be real disagreement. If everyone starts with the same view, nobody produces the counter-arguments.
  • People have to share the goal of getting it right, rather than winning, saving face or pleasing the boss.
  • It has to be safe to say “you're wrong”, including to someone senior.

4.4When groups get dumber

Remove those conditions and group reasoning goes wrong in predictable ways.

Experiment · Colorado, 2005

The researchers David Schkade, Cass Sunstein and Reid Hastie brought together citizens from two cities: Boulder, which leans liberal, and Colorado Springs, which leans conservative. They put people into small groups with others from their own city and asked them to discuss contested issues including climate change, affirmative action and civil unions for same-sex couples.

After discussion, the liberal groups had become more liberal and the conservative groups more conservative. Each group also became more uniform: the moderate voices inside them faded. Talking made everyone more certain, not more accurate.

This is group polarisation. When like-minded people discuss, each person hears mostly arguments for the side they already lean towards, plus the social reward of agreeing. Nobody plays the evaluator. Think of a family or community WhatsApp group where everyone already agrees: every day of discussion pushes it further in the same direction.

Two other failures are common in workplaces:

  • Groupthink — the psychologist Irving Janis's term for cohesive teams that suppress doubts to keep harmony. His classic case was the US government's 1961 Bay of Pigs invasion, a plan whose obvious flaws advisers kept to themselves.
  • Deference to rank — when the most senior person speaks first, everyone else anchors on their view or stays quiet. The group ends up with one person's reasoning and the illusion of a discussion.
Quiz Design the meeting

Your team must decide whether to open a second property. Which meeting design is most likely to reach a good decision?

Quiz

A family WhatsApp group shares daily messages about a political topic that everyone in it already agrees on. What does the research predict?

4.5Rules for reasoning together

If good reasoning happens between people, then the most practical skill isn't thinking harder on your own. It's designing conversations where the evaluator in everyone gets switched on.

Tool 8

Rules for Reasoning Together

  1. Think alone first. Everyone writes their view and confidence before discussion, so nobody anchors on the first speaker.
  2. Guarantee disagreement. Include people who see it differently — or assign someone to argue the other side.
  3. Lowest rank speaks first. Seniors speak last, so their view doesn't silence others.
  4. Argue about reasons, not people. “What makes you think that?” instead of “That's ridiculous.”
  5. Make changing your mind a win. Praise it out loud when someone updates.
  6. Know how you'll decide before you start arguing: who decides, and by what criteria.

Groups beat individuals only when disagreement is real, safe, and aimed at getting it right.

4.6What this means for you

As a learner

You can't reliably be a scout alone. Your myside bias will quietly build the case for whatever you already believe, and it will feel like careful thought. What you need are sparring partners: people who will disagree with you in good faith and whose arguments you'll actually evaluate.

AI assistants can help — but with a warning. AI models tend to agree with the person they're talking to. If you ask “Isn't my plan great?”, you'll probably hear yes. Ask instead: “Give me the strongest case against this,” or “Argue the other side as well as its best defender would.” Then evaluate that case as hard as you'd evaluate anyone else's.

As a teacher

This is the part to underline. If people evaluate other people's arguments with suspicion, then lecturing someone about why they're wrong triggers the evaluator against you. Every point you make gets the “Must I believe this?” treatment.

But when people reach a conclusion through their own evaluation — because they worked through a puzzle, argued with peers, or answered a good question — they own it. That's why the card puzzle is so much more convincing when a group solves it than when a teacher explains it. It's also why the most effective critical thinking teaching, as you'll see in Chapter 14, is built on dialogue rather than lectures.

Apply Find your evaluator

Pick a belief you've mostly discussed with people who agree with you. Who do you know who sees it differently? Write down the strongest argument you think they'd make — then plan to actually ask them.

See a worked example

Belief: “Online travel agencies are bad for us — we should push all guests to book directly.”

Person who disagrees: A friend who runs revenue for a larger hotel chain.

Their strongest argument: “For a small property, the agency is your marketing department. The commission buys you visibility you couldn't afford otherwise, and guests who first find you there often book direct next time. Cutting them could shrink demand more than it saves in commission.”

Plan: Ask her this week, and bring our actual numbers on repeat guests.

Teach it back

Explain to a friend, in five sentences or fewer, why arguing with people who disagree makes us smarter, while discussing with people who agree can make us dumber.

Compare with a model answer

“Our brains are good at finding holes in other people's arguments and bad at finding holes in our own. So when people who disagree argue honestly, each side checks the other, and the weak arguments get thrown out. That's why groups solve logic puzzles that most individuals fail. But when everyone agrees, nobody checks anything — you just hear more and more reasons for what you already believed. So seek out good-faith disagreement, and be suspicious of any group where everyone always agrees.”

Field work This week

Takeaways

  • Reasoning evolved mainly for social life: to justify ourselves, persuade others, and evaluate what others tell us.
  • We're lazy and one-sided when producing arguments, and sharp when evaluating other people's — even our own arguments, when we think they're someone else's.
  • Groups that genuinely disagree and aim at the truth outperform individuals. Like-minded groups polarise.
  • Design conversations for good reasoning: think alone first, guarantee disagreement, seniors speak last.
  • For teaching: people change their minds through their own evaluation, in dialogue — rarely through lectures.

Sources and further reading

  • Hugo Mercier & Dan Sperber, The Enigma of Reason (Harvard, 2017); and their “Why do humans reason?” Behavioral and Brain Sciences (2011).
  • David Moshman & Molly Geil, “Collaborative reasoning: Evidence for collective rationality,” Thinking & Reasoning (1998).
  • Trouche, Johansson, Hall & Mercier, “The selective laziness of reasoning,” Cognitive Science (2016).
  • Schkade, Sunstein & Hastie, “What happened on Deliberation Day?” (2007) — the Colorado experiment.
  • Irving Janis, Victims of Groupthink (1972).

Chapter 5 · Part II — Reasoning & Evidence

Evidence and Sources

Not all information deserves equal weight, even when it's technically true. This chapter is about judging evidence and sources quickly — the way professionals actually do it, which isn't the way most of us were taught.

In Chapter 1 you learned to count the hops between you and the original evidence. Now we look at the evidence itself, and at the people and pages that deliver it. You can't check everything — so the goal is to get good at checking fast, and at knowing what kind of evidence could actually settle a question.

Predict Before you read

Stanford researchers asked three groups to judge the reliability of real websites and online claims: 10 PhD historians, 10 professional fact-checkers, and 25 Stanford undergraduates. Which group was fastest and most accurate?

How sure are you?60%

The fact-checkers — by a wide margin. In Sam Wineburg and Sarah McGrew's study, one task compared an article from the American Academy of Pediatrics, a large professional body, with one from the American College of Pediatricians, a much smaller advocacy group with a similar name. Most students rated the smaller group as more reliable. Many historians judged both reliable. The fact-checkers quickly worked out who was behind the second site.

The historians were brilliant readers — but they read vertically, staying on the page and studying it closely. The fact-checkers barely read the page at all. They opened new tabs and looked up what others said about the source. The next section explains why that works.

5.1What kind of evidence is this?

Different kinds of evidence answer different questions. A lot of confusion comes from using one kind to answer a question only another kind can settle.

Kind of evidenceWhat it can showWhat it can't show
Anecdote / testimonial
“It worked for my aunt.”
That something is possible; that it happened at least onceHow often it happens, or whether the thing caused it
Expert opinionA good summary, if the expert knows this specific field and agrees with other expertsMuch, if it's a lone expert outside their field, or someone with an incentive
Observational data
surveys, records, trends
Patterns: what goes with what, how common things areCause and effect, on its own (Chapter 7)
Experiments / randomised trialsCause and effect, in the conditions testedWhether it applies to very different people or settings
Systematic reviews
summaries of all studies on a question
The overall weight of evidenceMuch, if the underlying studies are poor

Anecdotes deserve a special word. They're the most persuasive kind of evidence for humans — vivid, personal, easy to remember — and the weakest for establishing how common something is. A single story can never tell you whether you're looking at the rule or the exception.

The discrimination question

The single most useful question to ask about any piece of evidence is this: would I expect to see this evidence if the claim were false?

A product has hundreds of five-star reviews. Would you expect that if the product were bad? Unfortunately, yes — reviews can be bought or filtered. So the reviews don't discriminate much between “good product” and “bad product with a marketing budget”. Now suppose an independent testing lab bought the product anonymously and it passed. You'd be much less likely to see that if the product were bad. That's strong evidence.

Tool 9

The Discrimination Question

  1. What would I expect to see if the claim is true?
  2. What would I expect to see if the claim is false?
  3. Is this evidence much more likely in one case than the other?

If you'd see the same evidence either way, it's weak — however impressive it looks. Evidence is strong only when it would be surprising if the claim were false. (Chapter 8 turns this into numbers.)

Quiz

“My neighbour's blood sugar came down after he started taking this herbal powder.” This is best treated as evidence that…

5.2Read sideways, not down

Most of us were taught to evaluate a source by studying it closely: Is it well written? Does it have references? Does it look professional? Is it a .org? What does its About page say?

The problem is that all of these are easy to fake. Anyone can buy a professional design, a .org address and an impressive name. A page can't be trusted to describe itself honestly, any more than a stranger's business card can.

Fact-checkers do something different, which Wineburg and McGrew called lateral reading. Within seconds of landing on an unfamiliar source, they leave it. They open new tabs and search for the organisation or author: Who funds them? What do independent sources — news reports, encyclopedias, other experts — say about them? They also showed what the researchers called click restraint: scanning the search results before clicking, instead of clicking the first link.

Key idea: You learn more about a source by reading what others say about it than by reading what it says about itself.

5.3SIFT: four moves in a minute

The digital literacy researcher Mike Caulfield turned the fact-checkers' habits into four quick moves, known as SIFT. They're designed to take about a minute, not an afternoon.

Tool 10

SIFT

  1. Stop. Notice your emotional reaction. Do you know this source? If it makes you angry, delighted or scared, slow down.
  2. Investigate the source. Leave the page. Search the author or organisation. What's their expertise, and their agenda?
  3. Find better coverage. Ignore this source for a moment. Is the claim being reported by sources you already know to be reliable? What do fact-checkers say?
  4. Trace to the original. Follow quotes, statistics, photos and videos back to where they first appeared. Was the context the same?

Method by Mike Caulfield; see Verified (Caulfield & Wineburg, 2023).

Quiz

You land on a slick website with a .org address, a medical-sounding name, and a long About page describing its mission. It makes a surprising health claim. What's the best first move?

5.4Ten forwards, one source

Hearing a claim from many people feels like strong evidence. But evidence only adds up when the sources are independent — when each one checked the facts for itself. Ten people forwarding the same message is not ten sources. It's one source, copied ten times.

A hoax that keeps coming back

For well over a decade, messages have circulated claiming that UNESCO declared “Jana Gana Mana” the best national anthem in the world. UNESCO doesn't rank national anthems, and fact-checkers in India and abroad have debunked the claim many times. It keeps returning anyway — each new wave feeling fresh, and each forward arriving from someone you know.

It's a pleasant claim, which is part of why it spreads: it passes the “Can I believe this?” test from Chapter 2 easily. But however many times you receive it, the number of independent sources behind it stays the same: zero.

The same problem affects professional news. When one outlet reports something and others rewrite that report without checking, you see the claim in ten places. It looks confirmed. It's still one source. This is why “Trace to the original” matters: follow the chain back and see how many links actually touched the evidence.

Quiz

The same surprising claim reaches you from eight different relatives in one day. Compared with hearing it once, how much stronger is the evidence?

5.5Who benefits?

Every source has incentives. That doesn't mean everyone with an interest is lying — the manufacturer of a product may know it best. But incentives tell you which errors to expect.

  • Sponsored content and “advertorials” look like journalism but are paid for by the subject.
  • “Top 10” and “best of” lists often earn a commission on every product linked, which rewards enthusiasm over accuracy.
  • Press releases are frequently rewritten as news with little change.
  • Reviews can be bought, filtered, or written by the seller's friends — something anyone in hospitality knows well.
  • Viral content is rewarded for being shared, not for being true. Outrage and wonder travel furthest.

The question isn't “does this source have an interest?” — almost everyone does. It's “does this interest push them towards a particular answer, and is there independent evidence that doesn't share that push?”

5.6Photos, videos and AI

A photo feels like direct evidence: you're seeing it with your own eyes. But the most common trick isn't a faked image. It's a real image in the wrong context — a photo from a different year, a different country, or a different event, captioned to fit today's story. During floods, riots and elections, old images reliably resurface with new captions.

The fix is the T in SIFT: trace it. A reverse image search (Google Lens or similar) often shows where and when an image first appeared within seconds.

AI adds two new problems. First, fluent text and realistic images are now cheap. Polished writing used to be a weak signal of effort and expertise; now it's no signal at all. AI chatbots can state false things confidently, so check any factual claim from an AI that matters to a decision. Second, there's the liar's dividend: once people know fakes exist, anyone caught on real video can claim it's fake. Both make SIFT more important, not less.

5.7When “do your own research” backfires

Experiment · published in Nature, 2024

Kevin Aslett and colleagues ran a series of experiments in which people were shown news articles — some true, some false — and some were encouraged to search online before deciding whether they were true. You'd expect searching to help. Instead, searching increased the chance that people rated false stories as true, by roughly 20% in their experiments.

The effect was concentrated among people whose searches returned low-quality results. When you search the exact words of a false claim, you mostly find other pages repeating it — a “data void” where reliable sources haven't written anything.

This is why SIFT tells you to search the source and look for better coverage, rather than typing the claim itself into a search box. Searching the headline finds people who agree with the headline. Searching “who is [this outlet]” or “[claim] fact check” finds people who evaluated it.

Apply SIFT something real

Pick a claim, article or forward you saw this week. Run SIFT on it and write down what each step found. Time yourself — it should take a few minutes, not an hour.

See a worked example

Claim: A forward says a new government rule will fine anyone ₹10,000 for forwarding political messages.

Stop: I felt alarmed and wanted to warn my family. That's the moment to slow down.

Investigate the source: No source named — just “news”. That alone is a warning sign.

Find better coverage: Searched “fine for forwarding political messages fact check”. Two fact-checking sites had already covered it: no such rule exists.

Trace: The fact-checkers traced a version of it back several years; it resurfaces around elections.

Verdict: False. Confidence it's true: about 2%. Time taken: four minutes.

Teach it back

Explain “lateral reading” to a parent or older relative who gets a lot of forwards, in three or four sentences they'd actually remember.

Compare with a model answer

“When a message or website tells you something surprising, don't judge it by how official it looks — anyone can make things look official. Instead, leave it and look up who's behind it and whether any newspaper or fact-checker you already trust is saying the same thing. If nobody reliable is reporting it, wait before believing or forwarding it. It's like checking a stranger's reference before trusting their business card.”

Field work This week

Takeaways

  • Match the evidence to the question: anecdotes show possibility, observations show patterns, experiments show causes.
  • Ask the discrimination question: would I see this evidence even if the claim were false?
  • Read laterally. Judge a source by what independent sources say about it, not by how it presents itself.
  • SIFT: Stop, Investigate the source, Find better coverage, Trace to the original.
  • Many copies of a claim are one source. Searching a claim's own words can make you more confident in a falsehood.

Sources and further reading

  • Sam Wineburg & Sarah McGrew, “Lateral Reading and the Nature of Expertise,” Teachers College Record (2019); Stanford Report, “Fact checkers outperform historians when evaluating online information” (2017).
  • Mike Caulfield & Sam Wineburg, Verified: How to Think Straight, Get Duped Less, and Make Better Decisions about What to Believe Online (Chicago, 2023).
  • Aslett et al., “Online searches to evaluate misinformation can increase its perceived veracity,” Nature (2024).
  • Snopes and Factly fact-checks of the “UNESCO best national anthem” claim.

Chapter 6 · Part II — Reasoning & Evidence

Numbers That Lie

Numbers feel objective, which is exactly what makes them persuasive. Most misleading numbers aren't false. They're true numbers with the context removed.

This chapter draws on Calling Bullshit (2020) by Carl Bergstrom and Jevin West, who teach a university course by that name. Their key insight is that you don't need advanced statistics to catch most number tricks. You need a handful of questions, asked every time.

Predict Before you read

The WHO's cancer agency reported that eating 50 g of processed meat every day raises the risk of bowel cancer by 18%. If someone's lifetime risk without processed meat is about 8 in 100, roughly what is it if they eat 50 g a day?

How sure are you?60%

About 9 in 100. The 18% is a relative increase — 18% more than the starting risk, not 18 extra percentage points. For Australians, the Union for International Cancer Control worked it out as roughly 7.9% lifetime risk for non-eaters versus about 9.3% for daily 50 g eaters.

That's still a real increase, and across millions of people it matters. But “18% higher risk” and “from about 8 in 100 to about 9 in 100” feel completely different — and headlines in 2015 mostly used the first.

6.1Of what? Relative, absolute, and percentage points

Every percentage is a percentage of something. Change the something, and the same fact sounds bigger or smaller.

  • Relative change compares the new number with the old one: from 8 to 9 is a 12.5% increase.
  • Absolute change is the plain difference: from 8 in 100 to 9 in 100 is 1 extra case per 100 people.
  • Percentage points are the absolute difference between two percentages. If a loan rate goes from 2% to 3%, it rose by one percentage point — which is also a 50% increase.

People who want a number to sound big use relative change: “risk doubles!” (from 1 in 10,000 to 2 in 10,000). People who want it to sound small use absolute change. Both are true. You need both to understand it.

Try it Risk translator

Type a starting risk and a relative increase from a headline, and see what it means in people.

Quiz

Your occupancy rose from 60% last year to 66% this year. Which description is accurate?

6.2Out of how many? The denominator

A raw count without a denominator tells you almost nothing. Big places have more of everything — more crime, more hospitals, more accidents, more good news.

  • “Most road accidents happen within a few kilometres of home.” Of course — most driving happens near home.
  • “More people died on the roads in State A than State B.” How many people live there? How many vehicles? How many kilometres are driven?
  • “Complaints doubled this month!” From 2 to 4? Out of 50 guests or 5,000?

Rates — per 1,000 people, per 100 guests, per crore rupees — make comparisons fair. But rates can hide scale, too. India's Ministry of Road Transport reported 1,68,491 road deaths in 2022: about 19 people every hour. A per-capita rate can sound small next to other countries while the absolute number is a national emergency. Good reasoning looks at both.

Quiz

District A reported 500 dengue cases this season; District B reported 100. Where is dengue worse?

6.3What is the chart hiding?

Charts are arguments in picture form. The same numbers can tell opposite stories depending on the choices made in drawing them.

01009699 Us: 97Rival: 98Us: 97Rival: 98 Axis starts at 0Axis starts at 96
Same two numbers. On the right, a one-point difference looks like the rival scores twice as high.

Common chart tricks:

  • A truncated axis that doesn't start at zero, making small differences look huge (fine for line charts tracking change; misleading for bar charts comparing size).
  • A cherry-picked time window — start the chart at the lowest point and any trend looks like a recovery.
  • Cumulative totals, which can only go up, making a slowing trend look like steady growth.
  • Two different y-axes on one chart, which can make any two lines appear to move together.
  • 3D effects and pie charts that distort how big the slices look.

6.4Who got counted? Selection and survivors

A number can be perfectly accurate for the people it counted and completely misleading about everyone else. The question is always: how did these cases get into the data, and who's missing?

Survivorship bias is the most famous version. You see the survivors, not the failures, and draw lessons from a skewed sample. College dropouts who became billionaires are famous; the millions of dropouts who didn't are invisible. Investment ads show the funds that did well; funds that did badly were often closed or merged and quietly disappear from the record.

Closer to home

Every results season, coaching institutes run full-page ads featuring toppers. Sometimes the same topper appears in ads for several institutes, because a student who attended a short test series or a crash course at one institute can be counted by it. The ads show who succeeded. They don't show how many students enrolled, how many didn't succeed, or how much of the success came from the student rather than the institute.

Selection effects are everywhere in business, too. Online reviews come mostly from people who were delighted or furious. A survey of your current guests leaves out everyone who looked at your listing and booked elsewhere — often the people you most need to hear from.

Quiz

A coaching institute's ad: “12 of the top 50 rankers are our students!” What's the most important missing number?

6.5When a number becomes a target

The economist Charles Goodhart gave his name to a principle usually stated like this: when a measure becomes a target, it ceases to be a good measure. People start optimising the number instead of the thing it was meant to track.

True story · Hanoi, 1902

During a plague scare, French colonial authorities in Hanoi paid a bounty for every rat tail handed in. The number of tails rose impressively. But people soon noticed rats running around without tails: catchers were cutting off tails and releasing the rats to breed. Some reportedly farmed rats for their tails. The bounty was cancelled. The historian Michael Vann has written about the episode in detail.

You'll see Goodhart's law wherever numbers are rewarded. Staff paid on review scores start pressuring guests for five stars. Schools judged on exam results teach to the test. Call centres measured on call length hang up early. When someone shows you an impressive metric, ask whether anyone was rewarded for moving it.

6.6Big data, confident garbage

Algorithms and AI can make weak data look authoritative. But you rarely need to understand the algorithm to spot the problem. You need to ask where the data came from.

Case study · 2016

Two researchers posted a paper claiming a machine-learning system could tell criminals from non-criminals using only a photo of their face, with about 90% accuracy. Bergstrom and West took it apart without touching the code. The “criminal” photos were ID photos supplied by police. The “non-criminal” photos were taken from professional websites, where people present themselves well.

The authors' own composite images showed the “non-criminals” faintly smiling and the “criminals” frowning. The system had most likely learned to detect smiles and photo context — not criminality.

Garbage in, garbage out: an algorithm trained on biased data produces biased results with extra confidence. The question “how was the data collected?” beats any amount of technical detail.

6.7The bullshit asymmetry

A programmer named Alberto Brandolini observed that the energy needed to refute bullshit is an order of magnitude bigger than the energy needed to produce it. A false statistic takes seconds to invent and hours to debunk. That has two lessons:

  • You can't debunk everything. Use Belief Triage (Chapter 1) to decide what's worth your time.
  • Prevention beats cure. Teaching people the questions in this chapter protects them against thousands of claims you'll never see. That's the logic behind “prebunking”, which you'll meet in Chapter 14.

And one rule of thumb from Bergstrom and West: if a claim seems too good or too bad to be true, it probably is — or at least it needs checking before you share it.

Tool 11

The Number Sanity Check

  1. Of what? Relative or absolute? Percent or percentage points?
  2. Out of how many? What's the denominator? Is this a count or a rate?
  3. Compared to what? Last year? A similar place? Doing nothing? Why this time window?
  4. Who got counted — and who's missing? Survivors only? Only people who replied?
  5. What is the chart doing? Where does the axis start? Is it cumulative?
  6. Was anyone rewarded for this number? (Goodhart)
  7. Too good or too bad to be true? Then check before sharing.

Based on Bergstrom & West, Calling Bullshit.

Apply Check a real number

Find a number in an ad, a news story, or your own business reports this week. Run the Number Sanity Check on it. Which questions changed how you read it?

See a worked example

The number: “Our average review score rose from 4.2 to 4.6 this quarter.”

Out of how many: 12 reviews this quarter versus 60 last quarter. With 12 reviews, a couple of very happy guests move the average a lot.

Who's missing: Guests who didn't review. We also started asking only guests who seemed happy at checkout — a selection effect.

Target: The team got a bonus tied to the score. Goodhart alert.

What it means: We can't tell if guests are happier. Next step: ask every guest, and track the review rate as well as the score.

Teach it back

Explain the difference between relative and absolute risk to a family member, using an example from a health headline. Keep it to four sentences.

Compare with a model answer

“When a headline says something ‘raises your risk by 18%’, it means 18% more than your starting risk — not an 18% chance. If your risk was 8 in 100, an 18% rise takes it to about 9 in 100. That's worth knowing, but it's very different from what the headline makes you feel. So whenever you see ‘risk rises by X%’, ask: from what, to what?”

Field work This week

Takeaways

  • Most misleading numbers are true numbers with missing context.
  • Always get both relative and absolute figures, and know percent from percentage points.
  • Counts need denominators; rates need scale. Ask: out of how many, compared to what?
  • Charts are arguments: check the axis and the time window.
  • Ask who got counted and who's missing; ask whether anyone was rewarded for the number.

Sources and further reading

  • Carl T. Bergstrom & Jevin D. West, Calling Bullshit: The Art of Skepticism in a Data-Driven World (Random House, 2020); callingbullshit.org case study, “Criminal machine learning.”
  • IARC Monographs, red and processed meat (2015); UICC, “How to interpret IARC findings on red and processed meat as cancer risk factors.”
  • Ministry of Road Transport and Highways, Road Accidents in India 2022.
  • Michael G. Vann, “Of Rats, Rice, and Race: The Great Hanoi Rat Massacre” (2003).
  • Darrell Huff, How to Lie with Statistics (1954).

Chapter 7 · Part II — Reasoning & Evidence

Cause and Correlation

Two things happening together is the most common evidence in the world, and the most commonly misread. This chapter is about telling “goes with” from “causes” — in the news, in health claims, and in your own business.

Almost every decision rests on a causal belief: if we do X, Y will happen. Raise prices and bookings fall. Run ads and occupancy rises. Take this supplement and feel better. Most of those beliefs come from noticing that two things occurred together. Sometimes that's because one caused the other. Often it isn't.

Predict Before you read

Flight instructors noticed a pattern. When they praised a trainee pilot for an excellent landing, the next landing was usually worse. When they shouted at a trainee for a bad landing, the next one was usually better. What's the best explanation?

How sure are you?60%

It's mostly luck evening out. Daniel Kahneman tells this story from his time teaching Israeli Air Force flight instructors. An experienced instructor insisted that praise made cadets worse and shouting made them better. Kahneman realised the pattern would appear even if praise and shouting had no effect at all.

Every landing is a mix of skill and luck. An exceptionally good landing usually involves some good luck, which probably won't repeat, so the next one tends to be less exceptional. The same happens in reverse after a terrible one. This is called regression to the mean, and it fools experts every day. You'll see it in action in section 7.4.

7.1Four explanations for any correlation

When A and B go together, there are always at least four possibilities:

  1. A causes B. Maybe it really is the cause.
  2. B causes A. The arrow points the other way — reverse causation.
  3. Something else causes both. A third factor, called a confounder, drives both A and B.
  4. Chance. Look at enough pairs of things and some will line up by coincidence.

The classic example: ice cream sales and drownings rise and fall together. Ice cream doesn't cause drowning. Hot weather drives both. Obvious when it's ice cream — much harder when it's a claim you want to believe.

Quiz

At your property, guests who use the spa give much higher ratings than guests who don't. Your manager says: “The spa makes guests happy. Let's push it to everyone.” What's the strongest alternative explanation?

7.2Confounders: the hidden third factor

Confounders are the most common reason correlations mislead, because the people who choose to do something are usually different from those who don't, in many ways at once.

  • Children with more books at home do better at school. Is it the books, or the parents' education, income and attention that also bring the books?
  • Students with private tuition score higher. Is it the tuition, or families that can afford tuition? (And here's a twist: weaker students are also more likely to be sent for tuition, which pushes the other way. Confounders can hide effects as well as invent them.)
True story · medicine, 1980s–2002

For years, observational studies found that women who took hormone replacement therapy (HRT) after menopause had less heart disease. It became common to prescribe it partly for heart protection.

Then a large randomised trial, the Women's Health Initiative, tested one common form of HRT directly. In 2002 it reported that the women given the hormones had more heart disease and strokes, not less. One major explanation for the earlier finding: the women who chose HRT tended to be healthier, wealthier and more health-conscious to begin with. That was the confounder.

The story didn't end there. Later analyses suggested that age and the timing of starting treatment matter, and the debate continues among specialists. The lasting lesson is about method: patterns among people who choose a treatment can't tell you what the treatment does.

7.3Reverse causation

Sometimes the arrow runs the other way. People who use walking sticks fall more often — but the sticks aren't causing falls. Properties that spend the most on advertising sometimes have lower occupancy — because low occupancy is why they're advertising.

A quick test: could B have come first? If you can tell a plausible story in which the “effect” produced the “cause”, you need evidence about timing before you can conclude anything.

7.4Regression to the mean

Whenever a result involves luck — sales in a month, a batter's score, a patient's symptoms, a landing — extreme results tend to be followed by less extreme ones. Not because anything changed, but because the luck that made them extreme doesn't repeat.

Try it. Below, thirty staff members each have a fixed level of skill. Each month, their result is their skill plus some random luck. Pick out the five best and five worst performers in Month 1, and see what happens to those same people in Month 2 — when nothing about them has changed.

Try it Regression simulator

Run it several times. The Month 1 stars almost always slip, and the Month 1 strugglers almost always improve — with no praise, no shouting, no training, and no change in skill.

Now imagine a manager who gives the bottom five a stern talk after Month 1. Next month, they improve. The talk “worked”. Or someone who introduces a new policy in the worst month in three years; the next month is better, and the policy gets the credit. Many treatments for things that come and go, like back pain or colds, seem to work partly because people seek treatment when they feel worst — the point from which they'd usually improve anyway.

Quiz

Occupancy hit a three-year low in July, so you introduced a new staff incentive in August. August occupancy rose. How strong is the evidence that the incentive worked?

7.5Why randomising works

The cleanest way to find out whether A causes B is to decide who gets A by chance. Flip a coin: heads, you get the treatment; tails, you don't. Because a coin doesn't care about your income, health, mood or motivation, the two groups end up similar in every way — including ways nobody thought to measure. If the groups then differ in outcome, the treatment is the most likely reason.

That's a randomised controlled trial. It's how medicines are tested, and increasingly how social programmes are, too. In 2019, Abhijit Banerjee, Esther Duflo and Michael Kremer won the Nobel Prize in economics for their experimental approach to fighting poverty, much of it tested through randomised trials in India and elsewhere.

Businesses can do the same thing at small scale. Show half your website visitors one version of a page and half another, chosen at random. Try a new check-in process at randomly chosen properties or on randomly chosen days. It's called an A/B test, and it beats arguing about what “obviously” works.

Randomisation has limits. You can't ethically randomise people to smoke. A trial in one population may not apply to another. Small trials can be fooled by chance. But when it's possible, it answers causal questions that no amount of observational data can.

7.6When you can't randomise

True story · Britain, 1950s

Nobody could randomly assign people to smoke for thirty years. Yet by the 1960s, scientists were confident that smoking caused lung cancer. Richard Doll and Austin Bradford Hill built the case from several directions. One study followed tens of thousands of British doctors over years. The heavier they smoked, the higher their rate of lung cancer; those who quit saw their risk fall. The same pattern appeared in study after study, in different countries and different groups.

In 1965, Bradford Hill listed the kinds of evidence that make a causal link more convincing when you can't experiment. Paraphrased:

  • Strength — a big association is harder to explain away than a tiny one.
  • Consistency — it appears in different studies, places and populations.
  • Timing — the cause comes before the effect.
  • Dose-response — more of the cause, more of the effect.
  • Plausibility — there's a believable mechanism.
  • Experiment — removing the cause reduces the effect.

Researchers also look for natural experiments: situations where chance or an arbitrary rule assigned people to different conditions — a policy that applied only to people born after a certain date, or a lottery for school places.

Tool 12

The Cause Checklist

  1. Chance? How many things were compared before this pattern was found?
  2. Reverse? Could the “effect” have caused the “cause”?
  3. Confounder? What else differs between the groups?
  4. Regression? Did it start at an extreme high or low?
  5. Selection? Who chose to be in each group, and why?
  6. Hill's signs: Strong? Consistent? Right timing? Dose-response? Plausible mechanism?
  7. Has anyone randomised — or found a natural experiment?

Before you act on “X causes Y”, run the list. Better still, test it: try X in some places and not others, chosen at random.

Quiz

A report says: “Hotels that spend more on online ads have lower occupancy.” The best first question is…

Apply Test a causal belief

Write down one causal belief you act on in your business or life — “X leads to Y”. Run the Cause Checklist. Then design the simplest fair test you could actually run.

See a worked example

Belief: “Instagram posts bring us direct bookings.”

Checklist: We post more in peak season, when bookings rise anyway — a confounder. We also post more when we're busy with events, which might be the real draw. We noticed it after one viral post in a record month — possible regression and selection.

Hill's signs: No clear dose-response; we've never tracked bookings by source.

Fair test: Add a booking code to posts. For eight weeks, post on randomly chosen weeks only, and compare direct bookings in posting weeks versus non-posting weeks.

Teach it back

Explain regression to the mean to a cricket fan, in four sentences or fewer, using a batter who scores a century and then fails in the next match.

Compare with a model answer

“A century needs good form and some luck — a dropped catch, a close lbw that went your way. The luck part usually doesn't repeat, so the next innings tends to be closer to the batter's normal average. That's not ‘pressure’ or ‘complacency’; it would happen even if nothing about the batter changed. So be careful when you credit a new coach, or blame a hairstyle, for a change that comes right after an extreme performance.”

Field work This week

Takeaways

  • Every correlation has at least four explanations: cause, reverse cause, confounder, chance.
  • People who choose something differ from people who don't in many ways; that's why confounders are everywhere.
  • Extreme results regress toward the average. Whatever you did right after an extreme will get undeserved credit or blame.
  • Randomising breaks confounding. Where you can, test rather than argue.
  • Where you can't randomise, look for strength, consistency, timing, dose-response and mechanism.

Sources and further reading

  • Daniel Kahneman, Thinking, Fast and Slow (2011), chapter 17 — the flight instructors.
  • Writing Group for the Women's Health Initiative Investigators, JAMA (2002); later reappraisals of the WHI hormone trial.
  • Doll & Hill, British doctors study (from 1951); Austin Bradford Hill, “The Environment and Disease: Association or Causation?” (1965).
  • The Nobel Prize in Economic Sciences 2019 — Banerjee, Duflo and Kremer.
  • Judea Pearl & Dana Mackenzie, The Book of Why (2018).

Chapter 8 · Part III — Uncertainty

Thinking in Probabilities

Most of life isn't true or false. It's more likely or less likely. This chapter replaces yes-or-no thinking with a few simple habits for reasoning about chance — no formulas required.

Part III is about uncertainty: how to think clearly when you can't be sure, which is almost always. You've already started. In Chapter 2 you put a number on your confidence. In Chapter 5 you asked whether evidence would be surprising if a claim were false. This chapter joins those ideas into one way of thinking.

Predict Before you read

A disease affects 1 in 1,000 people. A test for it catches 99% of people who have it. But it also wrongly flags 5% of healthy people. Your test comes back positive. Roughly how likely is it that you have the disease?

How sure are you?60%

About 2%. Picture 1,000 people. One has the disease, and the test almost certainly catches them. Of the 999 healthy people, 5% — about 50 — also test positive. So about 51 people test positive, and only one of them is ill. One in 51 is about 2%.

If you guessed 95%, you're in large company. In studies by Gerd Gigerenzer and colleagues, many doctors struggled with questions like this when the numbers were given as percentages — and did far better when the same information was given as counts of people. That trick is the first tool in this chapter.

8.1Say it in numbers

Words like “likely”, “possible”, “rare” and “almost certain” feel precise, but different people mean very different things by them.

True story · CIA, 1951

A US intelligence estimate concluded that a Soviet-bloc attack on Yugoslavia “should be considered a serious possibility.” Afterwards, the analyst Sherman Kent asked the colleagues who had agreed on that wording what odds they'd had in mind. The answers ranged from about 20% to about 80%. The same phrase had meant “probably not” to some and “probably” to others — and nobody had noticed, because they all agreed on the words.

Kent spent years arguing that analysts should attach numbers to their judgements. The same applies to you. “The new property will probably be profitable” hides whether you mean 55% or 90% — and those call for very different decisions. Saying a number feels uncomfortable because it can be wrong. That's the point: a number can be checked, so you can learn.

Quiz

A doctor tells you a medicine's side effect is “rare”. What's the most useful follow-up question?

8.2Start with the base rate

The disease question fooled people because they jumped straight to the evidence (a positive test) and forgot to ask the first question: before any evidence, how common is this? That's the base rate. When something is rare, even good evidence produces lots of false alarms.

Base rates are everywhere once you look for them:

  • A founder is sure their new restaurant will succeed. How many new restaurants in that area are still open after three years?
  • A quiet, bookish person who loves poetry — more likely to be a librarian or to work in a bank? The description fits a librarian, but there are vastly more bank employees than librarians, so a bank job may well be more likely.
  • A security system flags a guest as suspicious. How often are flagged guests actually a problem?
Tool 13

Think in Frequencies

  1. Imagine a concrete crowd — 100 or 1,000 people.
  2. Split it by the base rate: how many have the thing, how many don't?
  3. Apply the evidence to each group separately: how many of each would show this sign?
  4. Compare: of everyone showing the sign, how many actually have the thing?

Percentages confuse; counts of people make the answer visible. Based on the work of Gerd Gigerenzer.

Try it Base-rate explorer

Change the numbers and watch what happens to the 1,000 people. Try making the disease more common (say 100 in 1,000), or the test more accurate (1% false alarms).

ill, test positive ill, test missed healthy, false alarm healthy, test negative

8.3Updating: Bayes without the maths

The Reverend Thomas Bayes, an 18th-century English minister, gave his name to the rule for how evidence should change your beliefs. You don't need the formula. You need three steps:

  1. Start with a prior — your best estimate before the new evidence, usually anchored on a base rate.
  2. Weigh the evidence — how much more likely is this evidence if the claim is true than if it's false? That's the discrimination question from Chapter 5, turned into a ratio.
  3. Update in proportion — strong evidence moves you a lot, weak evidence a little. Neither should take you straight to 0% or 100%.
Worked example · hiring

From experience, about 1 in 5 applicants for a manager role turns out to be a strong performer. That's your prior: 20%, or odds of 1 to 4.

A candidate gives a brilliant interview. How strong is that evidence? Suppose strong performers give brilliant interviews three times as often as weak ones (some weak candidates interview well; some strong ones don't). Multiply the odds by 3: 1-to-4 becomes 3-to-4. That's 3 out of 7 — about 43%.

So the brilliant interview more than doubled your confidence, from 20% to 43%. But it didn't make the candidate a sure thing. A reference check that's also three times likelier for strong performers would take you from 3-to-4 to 9-to-4 — about 69%. Each independent piece of evidence moves you further.

Tool 14

Update Like Bayes

  1. Prior: Before this evidence, how likely was it? (Start from the base rate.)
  2. Strength: How many times more likely is this evidence if true than if false?
  3. Update: Multiply your odds by that number. Strong evidence (×10) moves you a lot; weak evidence (×1.2) barely moves you.
  4. Repeat with each independent piece of evidence — and remember that ten copies of one source count once.

If the evidence is equally likely either way (×1), it shouldn't move you at all — however dramatic it feels.

A tragic example: confusing the two questions

True story · England, 1999–2003

Sally Clark, a solicitor, lost two baby sons, each apparently to sudden infant death syndrome (SIDS). She was charged with murdering them. At her trial, a prominent paediatrician told the court that the chance of two SIDS deaths in a family like hers was about 1 in 73 million. She was convicted in 1999.

The statistic was badly wrong in two ways. First, it assumed the two deaths were independent, like two coin tosses — ignoring genetic and environmental factors that can make a second SIDS death in the same family more likely. Second, and more fundamentally, it answered the wrong question. The chance of two SIDS deaths is small, but so is the chance of a mother murdering two babies. The court needed to compare the two rare explanations with each other, not treat the rarity of one as proof of the other.

The Royal Statistical Society publicly criticised the misuse of statistics in the case. Clark's conviction was overturned in 2003, after other evidence emerged. She died in 2007.

This error — confusing “how likely is this evidence if she's innocent?” with “how likely is she innocent, given this evidence?” — is called the prosecutor's fallacy. It's the same mistake as the disease test: a 5% false-alarm rate doesn't mean a 95% chance you're ill.

Quiz

A fraud-detection system flags 1% of honest bookings as suspicious and catches 90% of fraudulent ones. About 1 booking in 1,000 is fraudulent. A booking gets flagged. Roughly how likely is it to be fraud?

8.4Expected value — and the risk of ruin

Probabilities matter for decisions because outcomes have different sizes. The expected value of a choice is what you'd get on average if you could make it many times: each outcome's value multiplied by its probability, added up.

Worked example · a new site

You're considering a new glamping site. Over three years, you estimate: a 30% chance it makes ₹40 lakh, a 50% chance it makes ₹5 lakh, and a 20% chance it loses ₹30 lakh.

Expected value = (0.3 × 40) + (0.5 × 5) + (0.2 × −30) = 12 + 2.5 − 6 = ₹8.5 lakh. On average, a good bet.

But you don't get to make this decision a hundred times. If losing ₹30 lakh would sink the whole business, a positive average doesn't protect you. That's the risk of ruin: never take a bet whose worst case you can't survive, however good its average. This is also why buying insurance can be sensible even though its expected value is negative.

Tool 15

Expected Value, with a Ruin Check

  1. List the main outcomes (3 is usually enough: good, middling, bad).
  2. Give each a probability (they should add to 100%) and a value.
  3. Multiply and add: that's the expected value.
  4. Ruin check: can I survive the worst case? If not, change the bet — make it smaller, share the risk, or insure it.

The numbers will be rough. Writing them down still beats a gut feeling, because it shows you which estimate the decision depends on.

Quiz

A lottery ticket costs ₹100. There's a 1 in 10 lakh (1,000,000) chance of winning ₹50 lakh, and no other prizes. What's the expected value of buying one ticket?

8.5Chance has no memory

In 1913, at the Monte Carlo Casino, the roulette ball reportedly landed on black 26 times in a row. Gamblers lost fortunes betting on red, convinced it was “due”. But a roulette wheel has no memory: each spin is independent. This is the gambler's fallacy.

The same mistake appears whenever people expect random things to “even out” in the short run: a toss that's been lost five times in a row, a month that “must” be better because the last three were bad. If the events are genuinely independent, the next one doesn't know about the others. (If they're not independent — if a bad run reflects a real change, like a new competitor — then the streak is information, and you should update on it.)

Apply Put a number on it

Pick a real decision or prediction you're facing. Write down the base rate (how often do things like this work out?), your prior, the main piece of evidence and how strong it is, and your updated number.

See a worked example

Prediction: “A new corporate client who stayed with us once will book a second offsite within 12 months.”

Base rate: Of our last 20 one-time corporate clients, 5 came back within a year — 25%.

Evidence: Their HR head emailed to say the team loved it. Clients who send unsolicited praise seem maybe twice as likely to return (×2).

Update: Odds 1-to-3 × 2 = 2-to-3, so about 40%. Not the “definitely” I felt after reading the email. Worth a follow-up call, not worth holding dates for.

Teach it back

Explain to a friend why a positive result on a fairly accurate test for a rare condition often turns out to be a false alarm. Use a crowd of people, not percentages.

Compare with a model answer

“Imagine 1,000 people get tested for something only 1 of them has. The test catches that one person. But it also wrongly flags a few percent of the healthy people — say 50 of them. So 51 people get a positive result, and only 1 is actually ill. A positive result means ‘worth a second test’, not ‘you have it’ — because the rarer the condition, the more of the positives are false alarms.”

Field work This week

Takeaways

  • Words hide probabilities. Use numbers, so your judgements can be checked.
  • Start with the base rate. When something is rare, most alarms are false.
  • Think in frequencies: imagine 1,000 people, split them, and count.
  • Update like Bayes: prior, strength of evidence, proportional update. Never jump straight to certainty.
  • Expected value guides bets; the ruin check keeps you in the game. Chance has no memory.

Sources and further reading

  • Gerd Gigerenzer, Calculated Risks / Reckoning with Risk (2002); Hoffrage & Gigerenzer on natural frequencies in medical diagnosis (1998).
  • Sherman Kent, “Words of Estimative Probability,” Studies in Intelligence (1964).
  • Royal Statistical Society statement on the Sally Clark case (2001); R v Clark, Court of Appeal (2003).
  • Tim Harford, The Data Detective (2021).

Chapter 9 · Part III — Uncertainty

Calibration and Forecasting

Predicting the future is hard, and most experts are worse at it than they sound. But a small group of ordinary people turned out to be remarkably good — and what they do can be learned.

Every plan is a forecast. Opening a property, hiring someone, launching a course: each assumes something about what will happen. This chapter follows the political scientist Philip Tetlock, who spent decades measuring who forecasts well and why, told in Superforecasting (2015, with Dan Gardner).

Predict Before you read

Over nearly two decades, Tetlock collected tens of thousands of predictions from hundreds of experts — academics, government advisers, journalists — about politics and economics. On average, how accurate were they?

How sure are you?60%

Barely better than chance. In Expert Political Judgment (2005), Tetlock found that the average expert did only a little better than random guessing on many questions, and was often beaten by simple rules like “assume things stay the same”. It became famous through his quip about a dart-throwing chimpanzee.

But the average hid a crucial difference. Some experts did much better than others — and what separated them wasn't their field, their politics or their credentials. It was how they thought.

9.1Foxes and hedgehogs

Tetlock borrowed an ancient line made famous by the philosopher Isaiah Berlin: “The fox knows many things, but the hedgehog knows one big thing.”

  • Hedgehogs organise their thinking around one big idea — a theory, an ideology, a favourite explanation — and fit everything into it. They speak with confidence, use words like “certainly” and “impossible”, and rarely change their minds.
  • Foxes draw on many ideas, are comfortable with doubt and complexity, and adjust their views when facts change. They say “however” and “on the other hand” a lot.

The foxes forecast better. And Tetlock found an uncomfortable twist: the experts most in demand by the media tended to be hedgehogs. A confident, simple story makes better television than “it's about 60–40, and here's what would change my mind.” The qualities that make someone sound like an expert can be the opposite of the ones that make them right.

9.2The superforecasters

True story · 2011–2015

After intelligence failures such as the belief that Iraq had weapons of mass destruction, the US intelligence community's research agency (IARPA) ran a forecasting tournament. Teams competed to answer hundreds of questions like “Will this country's leader still be in office by the end of the year?” Tetlock and Barbara Mellers entered the Good Judgment Project, built from thousands of volunteers — ordinary people with an interest in the world.

It won decisively. Its best performers, the top 2% or so, were dubbed superforecasters. According to Good Judgment's reports, they outperformed a prediction market run inside the intelligence community — whose participants had access to classified information — by roughly 25–30%.

Two other findings matter for you. A training module of about an hour, teaching basic probabilistic reasoning, improved forecasters' accuracy — by roughly 10% in the published research — and the benefit lasted about a year. And forecasters working in teams did better than those working alone.

Key idea: Good judgement about the future isn't a gift. It's a set of habits — and a short, focused training measurably improved it. That's encouraging for you as a learner, and even more for you as a teacher.

9.3What superforecasters do

Tetlock summarised the habits he saw. Paraphrased and condensed:

  • Pin the question. Vague questions can't be scored. “Will the business do well?” becomes “Will monthly revenue exceed ₹15 lakh by March?”
  • Break it into parts. A hard question becomes several easier ones.
  • Outside view first. Start with the base rate for things like this, then adjust for what's special about this case.
  • Update often, in small steps. React to new information, but don't lurch.
  • Use fine-grained numbers. Superforecasters distinguish 60% from 65%, and that precision proved meaningful.
  • Look for the forces pushing both ways. What makes it more likely? Less likely?
  • Keep score, and study your misses — including the ones you got right for the wrong reasons.

9.4The outside view

True story · Israel, 1970s

Daniel Kahneman was on a team writing a new school curriculum and textbook. About a year in, he asked everyone to estimate how long it would take to finish. The estimates clustered around two years: roughly 18 to 30 months.

Then he asked the curriculum expert on the team, Seymour Fox, a different question: how long had other teams like this taken? Fox thought, and admitted that about 40% of such teams never finished at all. Of those that did, he couldn't think of one that had taken less than seven years, or more than ten. He also rated their team as slightly below average.

The team carried on anyway. The book took eight years to finish, and was never used.

The team's own estimate was the inside view: imagining the specific project and how it would go. Fox's answer was the outside view: what happened to similar projects. The inside view feels more relevant, because it's about your case. But it misses all the unexpected delays that reliably happen to projects in general, which is why projects so often run late and over budget. Psychologists call this the planning fallacy.

The fix is to start from the outside view, then adjust. How long do renovations like this usually take? What share of new properties in this area break even in year one? Your specific reasons can move you from the base rate — but they should move you from there, not replace it.

Quiz

Your contractor says the renovation will take 8 weeks. What's the best way to form your own estimate?

9.5Fermi estimation: guessing well

The physicist Enrico Fermi was famous for making surprisingly good estimates of things nobody could look up, by breaking them into parts he could roughly guess. The individual guesses are rough, but their errors tend to partly cancel out — and, more importantly, the method shows you exactly which assumptions matter.

Example · a Fermi estimate

Question: How many cups of tea do roadside tea stalls in a city of 1 crore people sell each day?

  • Adults who buy tea outside home on a typical day: maybe a quarter to a third of 1 crore → roughly 25–30 lakh people.
  • Cups per buyer: most buy one or two → say 1.5.
  • Share bought from stalls rather than offices, restaurants or cafés: maybe half.
  • Estimate: roughly 30 lakh × 1.5 × 0.5 ≈ about 20 lakh cups a day.

Every number here is an assumption, stated openly. That's the point: if someone disagrees, you can argue about the specific assumption — “I think far more people drink outside home” — rather than about a gut feeling.

Tool 16

Fermi Estimation

  1. Break the unknown into factors you can roughly estimate.
  2. Give each a rough value or range. Write your assumptions down.
  3. Multiply through. Round sensibly.
  4. Ask which assumption the answer is most sensitive to — that's what to check first.

Useful for market sizing, budgets, and any claim that sounds too big or too small. Also a quick check on other people's numbers.

9.6Keeping score

Forecasters get better only if their forecasts are scored. The standard measure is the Brier score: for each forecast, take the gap between the probability you gave and what happened (1 if it happened, 0 if not), square it, and average over all your forecasts.

  • 0 is perfect: 100% on everything that happened, 0% on everything that didn't.
  • 0.25 is what you'd get by saying 50% about everything.
  • 1 is the worst: fully confident and always wrong.

The Brier score rewards two things at once: being calibrated (your 70% predictions come true about 70% of the time) and being decisive (saying 90% or 10% when the evidence justifies it, rather than hiding at 50%).

Try it Score a forecast
Your forecast that it happens80%
Pick an outcome to see the score.

Your Thinking Journal already keeps score for you: every Predict question you lock in and every journal prediction you resolve feeds the calibration table. The more you use it, the more it tells you.

Tool 17

The Forecasting Recipe

  1. Pin the question — precise outcome, clear deadline.
  2. Outside view — how often do things like this happen? That's your starting number.
  3. Inside view — what's special about this case? Adjust, usually modestly.
  4. Commit — write a specific probability in your journal.
  5. Update — small, frequent adjustments as evidence arrives.
  6. Score — resolve it, and ask what you'd do differently.

After Tetlock & Gardner, Superforecasting.

Quiz

You forecast 70% that a new corporate client will sign by month's end. Midway through, their finance head asks for a revised quote with a smaller group. What's the best response?

Apply Make three forecasts

Write three pinned forecasts about your own business or life for the next 1–3 months, each with an outside view and a probability. Add each to the journal with a review date.

See an example

1. Will weekend occupancy in November exceed 80%? Outside view: last three Novembers were 72%, 78%, 85%. New factor: a highway closure nearby. My number: 45%.

2. Will the new booking engine be live by 15 December? Outside view: our last two tech projects ran 6 and 8 weeks late. My number: 30%.

3. Will at least 10 people sign up for my first critical-thinking workshop by 31 December? Outside view: none yet — I'll learn from this one. My number: 55%.

Teach it back

Explain the outside view to a friend who's sure their home renovation will finish on time. Four sentences.

Compare with a model answer

“Your plan tells you how long the renovation takes if nothing goes wrong — but something always does, and you can't list those things in advance. So instead of starting from your plan, start from how long renovations like yours actually took for people you know. If they typically ran a month late, assume yours will too, and then adjust if you have a genuinely good reason. It feels pessimistic, but it's just honest.”

Field work This week

Takeaways

  • Most expert forecasts are barely better than chance; confident hedgehogs do worse than flexible foxes.
  • Superforecasters are made, not born: short training, teamwork and practice with scoring improve accuracy.
  • Start from the outside view (base rates for similar cases), then adjust for specifics.
  • Break hard questions into guessable parts; state your assumptions.
  • Commit to numbers, update in small steps, and keep score.

Sources and further reading

  • Philip E. Tetlock, Expert Political Judgment (Princeton, 2005).
  • Philip E. Tetlock & Dan Gardner, Superforecasting: The Art and Science of Prediction (2015).
  • Mellers et al., “Psychological strategies for winning a geopolitical forecasting tournament,” Psychological Science (2014); Good Judgment Inc., reports on superforecaster performance.
  • Daniel Kahneman, Thinking, Fast and Slow (2011), chapter 23 — the curriculum story and the outside view.

Chapter 10 · Part III — Uncertainty

The Mind's Shortcuts

Your brain uses shortcuts to make fast judgements. Usually they work. In predictable situations they misfire — and knowing their names, it turns out, barely protects you. Here's what does.

This is the chapter most critical thinking courses start with: the list of cognitive biases. It comes late in this book on purpose. On its own, a list of biases mostly teaches people to spot them in others. You now have the tools to use it better — and to be honest about which famous findings held up and which didn't.

Predict Before you read

In a classic experiment, people watched a wheel of fortune spin and stop at either 10 or 65 (it was rigged to stop at one of those). Then they estimated the percentage of African countries in the United Nations. What happened?

How sure are you?60%

Much higher. In Amos Tversky and Daniel Kahneman's 1974 study, the median estimate was 25% for people who saw 10, and 45% for people who saw 65 — from a number everyone knew was random.

This is anchoring: an initial number pulls later estimates towards it, even when it's irrelevant. Unlike some famous findings you'll meet later in this chapter, anchoring has been replicated many times.

10.1Shortcuts, not stupidity

Kahneman popularised a useful picture of two modes of thinking. System 1 is fast, automatic and effortless — recognising a face, reading a mood, getting a feel for a price. System 2 is slow, deliberate and effortful — working out 17 × 24, checking an argument. Treat these as a metaphor for two styles of thinking, not two parts of the brain.

Most of the time, System 1's shortcuts — heuristics — work brilliantly. The psychologist Gerd Gigerenzer has shown that simple rules of thumb often beat complex analysis in real, messy environments. You couldn't get through a day without them.

A bias is what happens when a shortcut is used in a situation it wasn't built for, and misfires in a consistent direction. The goal isn't to stop using shortcuts. It's to recognise the situations where they reliably fail.

10.2A field guide

Here are the shortcuts that matter most in daily life and business, with a countermeasure for each.

ShortcutHow it shows upCountermeasure
Anchoring
first number pulls you
A crossed-out MRP makes the sale price feel cheap. The first offer in a negotiation sets the range. Last year's budget shapes this year's.Form your own estimate first, from the outside view, before you see theirs.
Availability
easy to recall = common
People fear plane crashes more than road travel, yet India recorded 1,68,491 road deaths in 2022 alone. Your last bad review feels like “all the reviews”.Ask for the base rate. Count, don't recall.
Framing
same facts, different choice
In a 1982 study, people — including doctors — were more willing to choose surgery when outcomes were described as survival rates rather than death rates, though the numbers were equivalent.Restate it the other way: “90% survive” = “10% die”. Do you still feel the same?
Loss aversion
losses loom larger
Guests react more to a ₹500 surcharge than to a ₹500 discount off a higher price. Owners hold failing assets to avoid “booking the loss”.Compare end states, not gains and losses from where you stand now.
Sunk cost
throwing good after bad
“We've already spent ₹20 lakh on this renovation, we can't stop now.”Status quo test (Chapter 2): knowing what you know now, would you start this today?
Hindsight
“I knew it all along”
After an outcome, it seems obvious — so we judge past decisions unfairly and learn the wrong lessons.Write predictions down before the outcome. Your journal is the cure.
OverconfidencePlanning fallacy (Chapter 9); 93% of drivers above the median (Chapter 2).Outside view; equivalent bet; keep score.
Attribution error
blame the person, not the situation
A staff member is late: “careless”. You're late: “traffic”.Double standard test: what situation would explain this if it were you?
Halo effectA charming or famous person is assumed to be competent at everything — the cricketer selling health drinks (Chapter 3).Judge each quality separately. Competent at what?
Which shortcut? 1 of 3

“We've spent three years and ₹40 lakh building this property's brand. We can't sell it now, even though it's losing money every month.”

Which shortcut? 2 of 3

After a guest's valuables were stolen at a nearby resort — widely shared on social media — your manager wants to spend heavily on new security, although your property has had no incidents in ten years.

Which shortcut? 3 of 3

A supplier's first quote is ₹5 lakh. You negotiate hard and get it down to ₹4 lakh, feeling pleased — though a fair market price would have been about ₹3 lakh.

10.3Why knowing about biases barely helps

You might hope that learning this list makes you less biased. The evidence says: not much, on its own.

In 2002, Emily Pronin, Daniel Lin and Lee Ross described the bias blind spot: people readily see biases in others, but rate themselves as less biased than average. We judge others by their behaviour, but ourselves by our intentions — and our intentions always feel reasonable. A later study by Richard West, Russell Meserve and Keith Stanovich found that people with stronger reasoning skills didn't have a smaller blind spot; if anything, it was slightly larger.

This fits everything in this book so far. Biases work below awareness (Chapter 2). We're sharp at evaluating other people's reasoning and lazy about our own (Chapter 4). A list of biases, learned alone, mostly becomes a new weapon for criticising others.

Key idea: You can't reliably catch biases by introspection. You catch them with processes — written predictions, checklists, outside views, other people — that work even when you can't feel the bias happening.

10.4Which findings survived?

Honesty requires a detour. Some of the most famous findings in popular psychology books have not held up.

The replication crisis

In 2015, a large collaboration of researchers, the Open Science Collaboration, tried to repeat 100 published psychology studies. Only about 36% of the repeats produced statistically significant results, against 97% of the originals, and the effects they found were on average about half as large.

Some celebrated ideas fared badly. “Ego depletion” — the idea that willpower is a resource that runs down with use — was tested in a coordinated replication across 23 labs in 2016, which found an effect close to zero. A famous “priming” study, in which reading words about old age supposedly made people walk more slowly, failed to replicate when tested carefully in 2012.

Daniel Kahneman himself had featured social priming studies in Thinking, Fast and Slow. In 2012 he warned priming researchers of a “train wreck looming”. In 2017 he wrote publicly that he had “placed too much faith in underpowered studies” in that chapter.

That last detail is a model of scout mindset from one of the field's most famous figures. And the broad picture is reassuring in a different way: anchoring, framing, availability and loss aversion have been widely replicated, though researchers still debate how large some effects are and when they apply. The lesson for you: treat any single striking study as a lead, not a law — especially if it appears in a bestseller. Chapter 11 shows you how to read studies with this in mind.

10.5What actually helps

If knowing isn't enough, what works? The evidence points to changing the situation and the process rather than relying on willpower in the moment.

  • Checklists. In a 2009 study across eight hospitals in different countries, introducing a simple surgical safety checklist was followed by a fall in deaths and complications. Checklists don't make surgeons less biased; they make it harder for bias to matter.
  • Write it down first. Predictions and estimates recorded before the outcome defeat hindsight bias and make calibration possible.
  • Consider the opposite. A 1984 study by Charles Lord, Mark Lepper and Elizabeth Preston found that asking people to consider how they'd judge the evidence if it had pointed the other way reduced biased thinking. (It's the selective sceptic test from Chapter 2.)
  • Practice with feedback. In 2015, Carey Morewedge and colleagues found that a single training session reduced several biases. An interactive game that gave personal feedback worked better than an instructional video — cutting bias by roughly a third immediately, with much of the benefit still there two to three months later.
  • Other people. Structured disagreement (Chapter 4) catches what introspection can't.
Tool 18

Process Beats Willpower

  • Anchoring → make your own estimate first, from the outside view.
  • Availability → look up the base rate; count cases.
  • Framing → restate the options the opposite way.
  • Sunk cost → “Would I start this today?”
  • Hindsight → record predictions and reasons before the outcome.
  • Overconfidence → equivalent bet, outside view, keep score.
  • Confirmation → consider the opposite; assign a devil's advocate.
  • Recurring decisions → turn them into a checklist.

Don't try to feel less biased. Build routines that work even when you can't feel it.

Apply Build one process

Pick a decision you make repeatedly — pricing, hiring, choosing suppliers, responding to reviews. Which shortcut is most likely to misfire there? Design a short process (three to five steps) that would catch it.

See a worked example

Recurring decision: Hiring front-desk staff.

Likely shortcuts: Halo effect (a confident, polished candidate seems good at everything) and anchoring on the first candidate interviewed.

Process: 1. Write the four qualities that matter before any interview. 2. Ask every candidate the same questions, including a role-play with a difficult guest. 3. Score each quality separately, right after each interview, before discussing. 4. Check one reference for each finalist. 5. Record our prediction of how they'll do, and review at three months.

Teach it back

Explain to a colleague why reading a list of cognitive biases won't make them much less biased — and what will. Four sentences.

Compare with a model answer

“Biases happen below awareness, so from the inside your judgement always feels reasonable — which is why people spot biases in others much more easily than in themselves. Knowing the names mostly makes us better at criticising other people. What works is building habits that catch biases whether or not we feel them: writing predictions down, checking base rates, using checklists and asking someone to argue the other side. Think of it like a seatbelt: you don't wear it because you feel a crash coming.”

Field work This week

Takeaways

  • Heuristics are useful shortcuts; biases are shortcuts misfiring in predictable situations.
  • Key ones: anchoring, availability, framing, loss aversion, sunk cost, hindsight, overconfidence, attribution error, halo.
  • Knowing about biases barely protects you — the bias blind spot is real, even for skilled reasoners.
  • Some famous findings (ego depletion, social priming) didn't replicate; anchoring, framing and loss aversion broadly did.
  • Processes beat willpower: write it down, check base rates, use checklists, bring in other people.

Sources and further reading

  • Tversky & Kahneman, “Judgment under Uncertainty: Heuristics and Biases,” Science (1974).
  • McNeil, Pauker, Sox & Tversky, “On the elicitation of preferences for alternative therapies,” NEJM (1982).
  • Pronin, Lin & Ross, “The Bias Blind Spot,” PSPB (2002); West, Meserve & Stanovich, “Cognitive sophistication does not attenuate the bias blind spot,” JPSP (2012).
  • Open Science Collaboration, “Estimating the reproducibility of psychological science,” Science (2015); Hagger et al., ego-depletion registered replication (2016); Doyen et al., priming replication (2012); Kahneman's comments (2012, 2017).
  • Haynes et al., surgical safety checklist, NEJM (2009); Lord, Lepper & Preston (1984); Morewedge et al., “Debiasing Decisions” (2015).
  • Gerd Gigerenzer, Gut Feelings (2007); Stuart Ritchie, Science Fictions (2020).

Chapter 11 · Part IV — Judgment & Action

Science as a Way of Knowing

Science isn't a pile of facts or a group of people in lab coats. It's a set of habits for catching our own mistakes — slow, imperfect, and still the best error-correcting system humans have built.

Part IV turns from reasoning to action: how knowledge gets made, whom to trust, and how to decide. We start with science, because it's the method this whole book has been borrowing from — prediction, measurement, testing, updating — organised on a large scale.

Predict Before you read

In 1847, a doctor in Vienna showed that when doctors washed their hands in a chlorine solution before examining women in childbirth, deaths from “childbed fever” in his clinic fell dramatically. How did most of the medical establishment respond?

How sure are you?60%

They largely resisted it. Ignaz Semmelweis noticed that the ward staffed by doctors, who came straight from performing autopsies, had far more deaths than the ward staffed by midwives. After he required handwashing with chlorinated lime in 1847, deaths in the doctors' ward fell from around 18% in the worst month to around 2%.

Yet many doctors rejected his conclusion. There was no germ theory yet to explain it; doctors were offended at being told they were carrying death on their hands; and Semmelweis himself was slow to publish and increasingly combative. He died in an asylum in 1865. Handwashing became standard only after Pasteur's and Lister's work on germs made the mechanism clear. Evidence eventually won — but not quickly, and not by itself.

11.1What makes science different

The physicist Richard Feynman put the heart of it in one line, in a 1974 speech to graduating students: “The first principle is that you must not fool yourself — and you are the easiest person to fool.”

Science is less a body of facts than a set of procedures designed to stop people fooling themselves — and each other. Individual scientists are as prone to soldier mindset as anyone (Chapter 2). What's different is the system: claims are made public, exposed to criticism, tested by people who'd love to prove them wrong, and revised.

Risky predictions

The philosopher Karl Popper argued that what makes a theory scientific is that it could be proven wrong — it forbids certain things from happening. A theory that fits every possible outcome tells you nothing.

His favourite example was Einstein. General relativity predicted that the Sun's gravity would bend starlight by a specific amount. In 1919, expeditions led by Arthur Eddington photographed stars during a solar eclipse to check. If the light hadn't bent as predicted, the theory would have been in serious trouble. It passed — and it was the risk of failing that made passing meaningful.

A test in India · 2008

The astrophysicist Jayant Narlikar and colleagues in Pune designed a test of astrology that astrologers themselves could accept. They collected birth charts of 200 children: 100 academically bright students and 100 children with intellectual disabilities. Astrologers were invited to examine a randomly chosen set of 40 charts each and say which group each child belonged to.

Guessing randomly would get about 20 of 40 right. The 27 astrologers who submitted answers averaged about 17. An astrological institute that assessed all 200 charts got 102 right — almost exactly what chance predicts.

The point isn't to mock anyone. It's that a clear, fair test with a prediction that could have succeeded tells you far more than thousands of satisfied customers.

11.2The toolkit against fooling ourselves

Each standard part of scientific method exists to block a specific way of fooling yourself:

ToolWhat it protects against
Control groupThings that would have happened anyway — regression to the mean, natural recovery (Chapter 7)
RandomisationConfounders — differences between people who choose a treatment and people who don't
PlaceboImprovement from expecting to improve
BlindingResearchers and participants unconsciously seeing what they hope to see
Pre-registrationChanging the question or the analysis after seeing the data until something looks significant
Peer reviewObvious errors — a filter, not a guarantee
ReplicationFlukes and errors in any single study (Chapter 10)
Systematic reviewsCherry-picking — looking only at the studies you like

11.3Significant isn't the same as important

News stories often report that a study found a “significant” effect. In statistics, significant means roughly: this result would be surprising if there were no real effect at all. It does not mean the effect is large, important, or even certainly real.

  • With a big enough sample, tiny effects become “significant”. A supplement that raises a 100-point memory score by half a point can be statistically significant and practically useless. Always ask how big the effect was.
  • If researchers test many outcomes or try many analyses, some will look significant by chance. Statisticians call this the garden of forking paths, and deliberately fishing for significance is called p-hacking. Pre-registration exists to stop it.
  • One significant study is a lead. Several independent ones pointing the same way are evidence.

11.4How to read a study without a PhD

You don't need to understand the statistics to ask good questions. Most misleading science news fails on the simplest ones.

Tool 19

The Study Reader's Checklist

  1. In whom? Cells in a dish, mice, 20 students, or 50,000 adults?
  2. What was measured? The thing you care about, or a stand-in for it?
  3. Compared with what? Was there a control group? Was it randomised and blinded?
  4. How big was the effect? In absolute terms, not just “significant”.
  5. Is it one study or many? Has it been replicated? What do systematic reviews say?
  6. Who paid, and was it pre-registered?
  7. Does the headline claim more than the study found?

Most science headlines fail on questions 1, 4 or 7.

Quiz

Headline: “Turmeric compound destroys cancer cells, scientists find.” The study exposed cancer cells in a laboratory dish to a concentrated compound from turmeric. What's the biggest gap between study and headline?

11.5Consensus — and when it was wrong

When most experts in a field agree after reviewing the evidence, that consensus deserves heavy weight. You'll almost never be in a position to outthink an entire field on its own subject.

But consensus has been wrong, and it's worth knowing how it got corrected.

  • In 1912, Alfred Wegener proposed that continents drift. Most geologists rejected it, partly because he had no convincing mechanism. By the 1960s, new evidence from the ocean floor led to plate tectonics, and the field changed its mind.
  • In the early 1980s, Barry Marshall and Robin Warren argued that most stomach ulcers were caused by a bacterium, not stress or acid. Many doctors were sceptical. Marshall famously drank a culture of the bacterium and developed gastritis. The evidence accumulated, antibiotics became standard treatment, and in 2005 they won the Nobel Prize.

Notice how these corrections happened: through better evidence, presented within science's own channels, which eventually persuaded the field. Not through popularity, and not by the outsiders being loud.

That matters because of a common argument: “They laughed at Galileo, and they're laughing at me.” Being rejected doesn't make someone right. For every Wegener, there were many rejected ideas that were simply wrong. What made Wegener and Marshall different was evidence that others could check.

11.6Tradition as hypothesis, science as test

Chapter 1 argued that traditions can carry real knowledge. Science is how you find out which ones do.

True story · China, 1970s

During a secret government programme to find new malaria treatments, the chemist Tu Youyou and her team searched ancient Chinese medical texts. A 4th-century text described using sweet wormwood soaked in cold water. The cold-water detail was a clue: heating might destroy the active ingredient. Her team developed a low-temperature extraction, isolated artemisinin, and tested it. Artemisinin-based treatments went on to save millions of lives. Tu Youyou received the Nobel Prize in 2015.

This is the healthy relationship between tradition and science. Tradition supplied a hypothesis worth testing. Science supplied the test — and, just as importantly, sorted this remedy from the many traditional remedies that didn't work. Neither “it's ancient, so it's true” nor “it's traditional, so it's nonsense” would have found artemisinin.

Warning signs of pseudoscience

  • Claims that can't fail — every outcome is explained after the fact.
  • Testimonials instead of controlled tests.
  • No change in the claims over decades, whatever new evidence appears.
  • Science-sounding words used as decoration: “quantum”, “energy”, “detox”, “frequency”.
  • Claims of suppression: “doctors don't want you to know”.
  • Something for sale, with the seller as the main source of evidence.
Quiz

A wellness brand says its product works through “quantum energy alignment”, has thousands of happy customers, and that “mainstream doctors are threatened by it.” How many warning signs is that?

Apply Read a real study

Find a science or health headline from this week. Track down the study it's based on (or at least a detailed report of it) and run the Study Reader's Checklist. Does the study support the headline?

See a worked example

Headline: “Drinking coffee helps you live longer.”

In whom: About 4 lakh adults in a long-running observational study.

Compared with what: Coffee drinkers versus non-drinkers — not randomised. Coffee drinkers may differ in income, health and habits.

How big: A modestly lower death rate over the study period.

Verdict: The headline turns an association into a cause. Fair summary: “Moderate coffee drinking is not linked to harm and may be linked to slightly lower mortality; it's not clear coffee is the cause.”

Teach it back

Explain to a relative who trusts traditional remedies how tradition and science can work together, using the artemisinin story. Four or five sentences, and no mockery.

Compare with a model answer

“Old traditions often contain real wisdom — one of the most important malaria medicines in the world came from a 1,600-year-old Chinese text. But the scientist who found it didn't just trust the text; she tested it carefully, and that's how she discovered exactly what worked and how to prepare it. The same testing also showed that many other traditional remedies didn't work. So the question isn't ‘traditional or modern?’ — it's ‘has anyone tested this fairly?’ If they have, we can trust the result, whichever tradition it came from.”

Field work This week

Takeaways

  • Science is a system for catching mistakes, built because individuals fool themselves.
  • Good theories make risky predictions that could fail. Claims that fit every outcome say nothing.
  • Each method — controls, randomisation, blinding, pre-registration, replication — blocks a specific error.
  • “Significant” isn't “important”: always ask how big the effect is, in whom, compared with what.
  • Weigh consensus heavily; it changes through checkable evidence. Traditions are hypotheses; science is the test.

Sources and further reading

  • On Semmelweis: “The Semmelweis Effect,” Significance magazine; Semmelweis, The Etiology, Concept and Prophylaxis of Childbed Fever (1861).
  • Richard Feynman, “Cargo Cult Science,” Caltech commencement address (1974); Karl Popper, Conjectures and Refutations (1963).
  • Narlikar et al., “A statistical test of astrology,” Current Science (2009); “An Indian Test of Indian Astrology,” Skeptical Inquirer (2013).
  • The Nobel Prize in Physiology or Medicine 2005 (Marshall & Warren) and 2015 (Tu Youyou).
  • Carl Sagan, The Demon-Haunted World (1995); Ben Goldacre, Bad Science (2008).

Chapter 12 · Part IV — Judgment & Action

Whom to Trust

Chapter 1 ended with a puzzle: you can't check everything, so almost everything you know comes from trusting someone. This chapter is the answer — how to choose whom to trust, on what, and how much.

Critical thinking is often presented as doubting everything. That's impossible, and Chapter 1 showed why it would be a disaster: humans thrive by learning from others. The real skill is calibrated trust — trusting the right people about the right things, to the right degree.

Predict Before you read

Researchers Joshua Kalla and David Broockman analysed 49 field experiments on political campaign contact — ads, mail, phone calls, door-to-door canvassing — in US general elections. On average, how much did this contact change people's votes?

How sure are you?60%

About zero. Their 2018 paper, titled “The Minimal Persuasive Effects of Campaign Contact in General Elections,” found that the average effect on vote choice was essentially nil. (Campaigns can still matter in other ways, such as getting supporters to turn out, and effects were sometimes larger early on or in primaries.)

This fits a broader argument by Hugo Mercier, whom you met in Chapter 4, in his book Not Born Yesterday: people are not nearly as gullible as we assume. We're “open but vigilant” — hard to persuade of things that clash with what we already believe, and easy to persuade of things we already wanted to hear. That changes how you should think about misinformation, and about trust.

12.1We're vigilant, not gullible

If people were simply gullible, the fix for misinformation would be to make them more sceptical. Mercier argues that the bigger problem is often the reverse: people are sceptical of the wrong things, and credulous about claims that fit what they already believe or want.

Remember the lynching rumours in Chapter 1. People didn't believe them because they'd believe anything. They believed them because the messages arrived from trusted people, matched existing fears about strangers, and asked them to protect children — something they already cared about. The trust was misplaced, but it wasn't random.

Mercier describes several checks we naturally use when someone tells us something:

  • Plausibility — does it fit what I already know?
  • Arguments — are the reasons they give good? (Chapter 4: we're good at judging others' arguments.)
  • Competence — are they in a position to know?
  • Benevolence — are their interests aligned with mine, or might they want to mislead me?
  • Accountability — would they pay a price, in reputation, if they turned out to be wrong?

These are good instincts. They go wrong in predictable ways — especially when we mix up two of them.

12.2Kind isn't the same as knowledgeable

The most common trust mistake is treating benevolence as if it were competence. Your uncle loves you and would never try to mislead you. That makes him trustworthy about his intentions. It doesn't make him knowledgeable about vaccines, share prices or constitutional law.

Family WhatsApp groups run on benevolence: messages are trusted because of who sent them, not because the sender knows the subject. A forward from a cousin feels safer than the same text from a stranger — even though the cousin is usually just passing along something from a stranger.

Key idea: Ask two separate questions. “Do they mean well?” and “Are they in a position to know?” Most misplaced trust comes from answering the first and assuming the second.

Quiz

Your most honest, loyal employee — who has never lied to you — strongly recommends investing the business's reserves in a new cryptocurrency a friend told them about. How should you weigh this?

12.3When to trust expert intuition

Not all expertise is equal. Some experts have gut feelings you should trust; others have confidence you shouldn't.

In 2009, Daniel Kahneman — a sceptic of expert intuition — and Gary Klein — who had spent his career studying how firefighters and nurses make fast, accurate calls — wrote a joint paper about where they agreed. Their conclusion: intuitive expertise is real when two conditions hold.

  1. The environment is regular enough to be predictable — patterns really do repeat.
  2. The expert has had lots of practice with quick, clear feedback about whether they were right.

Firefighters, chess players, experienced nurses, and a front-desk manager who has watched thousands of guests check in: their intuitions are trained by constant feedback. Long-range political pundits and people picking individual stocks work in environments that are either too unpredictable or give feedback too slowly and noisily. They may feel just as confident — confidence is not the signal (Chapter 9).

Quiz

Whose gut feeling deserves the most weight?

12.4Follow the incentives

The novelist Upton Sinclair put it memorably in 1935: “It is difficult to get a man to understand something, when his salary depends upon his not understanding it.”

True story · tobacco, 1950s–1990s

As evidence linking smoking to cancer mounted (Chapter 7), tobacco companies didn't need to prove smoking was safe. They only needed to keep the question looking unsettled. A 1969 internal memo from one company put it bluntly: “Doubt is our product.” The industry funded friendly research and experts, and emphasised uncertainty for decades. The strategy has since been copied in other fields.

This doesn't mean anyone with an interest is lying. It means asking: does this source's incentive push towards a particular answer — and do independent experts, without that incentive, agree? A drug company's trial is more credible when independent researchers get the same result. An industry-funded study that contradicts every independent study deserves extra scrutiny.

12.5When experts disagree

Experts disagree all the time, and that confuses people into thinking nobody knows anything. Some questions help:

  • Is the disagreement about the core or the details? Scientists argue about how much and how fast, while agreeing on whether. News coverage often blurs the two.
  • Is it one maverick against a field, or a field genuinely split? A lone dissenter is sometimes right (Chapter 11), but the base rate favours the field.
  • Are the experts independent? Twenty experts all funded by the same interest are closer to one voice than twenty.
  • What do bodies that review all the evidence say? Systematic reviews and national academies exist precisely to weigh disagreement.

12.6“Do your own research” — carefully

In Chapter 5 you met the 2024 finding that searching online to check false news made people more likely to believe it. Doing “your own research” on a technical question usually means searching for the claim, finding pages that repeat it, and feeling confirmed.

A better version is: research the sources, not the subject. You probably can't evaluate the chemistry of a vaccine or the economics of a policy. You can evaluate who's making the claim, what their track record and incentives are, and whether independent experts agree. That's the research that's actually within reach.

Tool 20

The Trust Test

  1. Competence: Are they in a position to know this specific thing?
  2. Feedback: Does their field give them quick, clear feedback on whether they're right?
  3. Benevolence ≠ competence: Do they mean well? (Separate question.)
  4. Incentives: Does anything push them towards this answer?
  5. Independence: Do experts without that incentive agree?
  6. Transparency: Do they show sources, admit uncertainty and correct their mistakes?

Use with Belief Triage (Chapter 1): decide what's worth checking, then decide whom to trust on it.

Apply Map your trust network

For each area below, write down the person or source you rely on most. Then run the Trust Test briefly on each. Where is your trust mostly benevolence rather than competence?

See a worked example

Health: Our family doctor — competent, gets feedback, no obvious incentive. Good. But I also follow a fitness influencer who sells supplements — incentive problem.

Money: A friend who “did well in the market” — benevolent, but no track record I've ever seen scored. Thin.

News: Mostly forwards from two family groups. Benevolent, not competent. Should switch to two named outlets and one fact-checker.

Business: My own experience plus our booking data — decent, but I should add an outside view from other operators.

Where trust is misplaced: Money and news. Both rely on people who mean well but aren't in a position to know.

Teach it back

Explain the difference between trusting someone's intentions and trusting their knowledge, to a relative who forwards health advice from friends. Keep it warm, and under five sentences.

Compare with a model answer

“When someone we love sends us health advice, we trust it because we trust them — and they really do mean well. But meaning well and knowing the answer are two different things. Your friend would never lie to you, but she probably got the message from someone she's never met. So it's worth asking: is the person at the start of this chain actually a doctor or scientist who knows this subject? Loving the sender and double-checking the message can go together.”

Field work This week

Takeaways

  • You can't avoid trusting; the skill is calibrating whom you trust, on what, and how much.
  • People are vigilant rather than gullible — but credulous about claims that fit what they already believe.
  • Separate benevolence (do they mean well?) from competence (are they in a position to know?).
  • Trust intuition only from experts in regular environments with fast, clear feedback.
  • Follow incentives, look for independent agreement, and research the sources rather than the subject.

Sources and further reading

  • Joshua Kalla & David Broockman, “The Minimal Persuasive Effects of Campaign Contact in General Elections: Evidence from 49 Field Experiments,” American Political Science Review (2018).
  • Hugo Mercier, Not Born Yesterday: The Science of Who We Trust and What We Believe (Princeton, 2020).
  • Daniel Kahneman & Gary Klein, “Conditions for intuitive expertise: A failure to disagree,” American Psychologist (2009).
  • Upton Sinclair, I, Candidate for Governor: And How I Got Licked (1935); Naomi Oreskes & Erik Conway, Merchants of Doubt (2010).

Chapter 13 · Part IV — Judgment & Action

Deciding Under Uncertainty

Everything in this book ends up here: a choice, made without knowing how it will turn out. Good decisions don't guarantee good outcomes — but good decision habits, repeated, tilt the odds.

You've learned to pin claims, weigh evidence, think in probabilities and catch your own biases. A decision combines all of these with one more ingredient: values — what you actually want. This chapter gives you a practical routine for decisions that matter.

Predict Before you read

In a 1989 study, people were asked to explain a future event. Some were told to imagine it might happen; others to imagine it had already happened. How did imagining it as certain affect the number of reasons people came up with?

How sure are you?60%

About 30% more reasons. Deborah Mitchell, Jay Russo and Nancy Pennington called this prospective hindsight: imagining an outcome as a fact makes it easier to generate explanations for it. The psychologist Gary Klein built a decision technique on this idea, the pre-mortem, which you'll use in section 13.6.

A fair caveat, in the spirit of Chapter 10: the study counted how many reasons people produced, not how good the reasons were. The pre-mortem is widely used and makes sense, but its benefits are better supported by practice than by large trials.

13.1What a decision is made of

Every decision has three parts, and confusion usually comes from mixing them:

  • Options — what you could do.
  • Beliefs — what you think will happen if you do each (predictions, with probabilities).
  • Values — how much you care about each possible outcome.

Chapter 3 showed how two partners can argue for an hour because one is making a prediction and the other stating a value. Separating the three lets you argue about each in the right way: evidence for beliefs, honest conversation for values, and creativity for options.

13.2Don't judge decisions by results alone

The former professional poker player Annie Duke calls it resulting: judging the quality of a decision by how it turned out. It's natural, and it's one of the main reasons people learn the wrong lessons from experience.

True story · Johannesburg, 2007

In the final of the first T20 World Cup, Pakistan needed 13 runs from the last over with one wicket left, and Misbah-ul-Haq was well set. MS Dhoni handed the ball to Joginder Sharma, a relatively inexperienced medium-pacer. Misbah tried a scoop shot, and Sreesanth took the catch at short fine leg. India won by five runs, and Dhoni's gamble was hailed as inspired.

Now imagine the ball had cleared Sreesanth by a metre. The same decision, made for the same reasons, with the same information, would have been called one of the worst in Indian cricket history.

Whether the decision was good depends on what Dhoni knew and what his alternatives were at that moment — not on where the ball landed. Outcomes mix decision quality with luck. Duke's four boxes make this explicit:

Good outcomeBad outcome
Good decisionDeserved successBad luck — don't change the process
Bad decisionDumb luck — dangerous, because it teaches the wrong lessonJust deserts
Quiz

You skipped the insurance renewal for your property to save money. Nothing went wrong that year. Which box is this most likely in?

13.3Widen your options

Many bad decisions are made between too few options — often just one: “Should we do this, yes or no?” The management researcher Paul Nutt studied hundreds of decisions in organisations and found that about half failed. Failures were linked to a narrow search for alternatives: decision-makers seized on one idea early and spent their energy defending it.

Chip and Dan Heath's book Decisive suggests two quick fixes:

  • The vanishing options test. Imagine none of your current options were possible. What would you do instead? This forces new ideas onto the table.
  • “And”, not “or”. Could you do more than one, at smaller scale? Test two marketing channels at once instead of betting on one.

13.4One-way and two-way doors

In his 2015 letter to shareholders, Amazon's founder Jeff Bezos distinguished two kinds of decision. Some are one-way doors: hard or impossible to reverse, so they deserve slow, careful thought. Most are two-way doors: if you don't like what's on the other side, you can walk back through. Those should be made quickly, by the people closest to them, and adjusted as you learn.

Signing a ten-year lease, taking on large debt, or firing a senior person are one-way doors. Trying a new menu, a new price for a month, or a new booking channel are usually two-way doors. A common mistake is treating two-way doors with one-way caution — and endlessly debating things that a small experiment (Chapter 7) would settle faster.

Quiz

Which of these deserves the slowest, most careful process?

13.5Putting options side by side

For decisions with several options and several things you care about, a simple decision matrix helps. List options as columns and criteria as rows, give each criterion a weight for how much it matters, and score each option.

A warning: the matrix is a tool for thinking and discussing, not a machine that produces the answer. The weights are your values, stated honestly. The scores are predictions, which could be wrong. If the result surprises you, don't simply obey it — ask which weight or score you disagree with. That's where the real decision is.

Try it Decision matrix

Rename the options and criteria for a real decision. Weights: how much each criterion matters (1–5). Scores: how well each option does on it (1–5).

13.6The pre-mortem

A post-mortem asks why something died. A pre-mortem, Gary Klein's technique, asks the same question before you start, while you can still change course.

Tool 21

The Pre-mortem

  1. Gather the people involved, once the plan is nearly final.
  2. Say: “Imagine it's a year from now. We went ahead — and it failed badly. What happened?”
  3. Everyone writes their reasons privately for a few minutes (independence — Chapter 4).
  4. Go round the room, one reason each, until they're all listed. The most senior person goes last.
  5. Pick the most likely or most damaging failures and change the plan to guard against them.

It makes doubt socially safe: people aren't being negative, they're doing the exercise.

13.7Keep a decision journal

Because of hindsight bias (Chapter 10), you won't remember what you knew or expected when you made a decision. So write it down at the time: the options you considered, what you expected to happen and how confident you were, the key reasons, and how you felt. Months later, compare with what actually happened.

The point isn't to feel good or bad about outcomes. It's to learn which parts of your process are reliable. Your Thinking Journal has a “Decision” entry type for exactly this.

Tool 22

The Decision Checklist

  1. Frame: What exactly am I deciding? Separate options, beliefs and values.
  2. Door type: One-way or two-way? Match the effort to the reversibility.
  3. Widen: At least three real options. Try the vanishing options test.
  4. Reality-test: Outside view and base rates. Can I run a small test first?
  5. Expected value and ruin check (Chapter 8).
  6. Pre-mortem: It failed — why?
  7. Record: Decision, expectations with numbers, reasons. Set a review date.

For two-way doors, steps 1, 2 and 7 are often enough. For one-way doors, do them all.

Apply Run a real decision

Take a real decision you're facing. Work through the Decision Checklist in writing, then add it to your journal as a Decision entry with a review date.

See a worked example

Decision: Whether to launch a paid critical-thinking workshop in January.

Door type: Two-way — a single workshop is cheap to try and easy to stop.

Options: (a) a paid public workshop; (b) a free pilot for 10 people first; (c) a workshop for one company's staff; (d) and-not-or: run the free pilot in December, then the paid one in January.

Outside view: First-time workshops typically draw mostly friends and friends-of-friends.

Pre-mortem: It failed because nobody came (no audience yet); because the content was too abstract; because I lectured instead of running exercises.

Choice: Option (d). Prediction: 8+ people at the pilot (70%); at least half would recommend it (60%). Review 15 January.

Teach it back

Explain “resulting” to a friend who is beating themselves up over a decision that turned out badly. Be kind, and honest.

Compare with a model answer

“How something turned out isn't the same as whether it was a good decision — luck plays a big part in outcomes. The fair question is: given what you knew then and the options you had, was it a reasonable choice? If yes, this was bad luck, and changing your approach because of it could make future decisions worse. If there's something you'd genuinely do differently with the same information, that's worth learning — but learn that, not just ‘it went badly’. It's what poker players and good investors do.”

Field work This week

Takeaways

  • A decision = options + beliefs + values. Argue about each in its own way.
  • Don't judge decisions only by outcomes. Luck is real; learn from process.
  • Widen options before choosing; “whether or not” decisions are a warning sign.
  • Match effort to reversibility: slow for one-way doors, fast experiments for two-way doors.
  • Use the matrix to discuss, the pre-mortem to find failure modes, and the journal to learn.

Sources and further reading

  • Mitchell, Russo & Pennington, “Back to the future: Temporal perspective in the explanation of events,” Journal of Behavioral Decision Making (1989); Gary Klein, “Performing a Project Premortem,” Harvard Business Review (2007).
  • Annie Duke, Thinking in Bets (2018) and How to Decide (2020).
  • Paul C. Nutt, “Surprising but true: Half the decisions in organizations fail,” Academy of Management Executive (1999); Chip & Dan Heath, Decisive (2013).
  • Jeff Bezos, Letter to Amazon shareholders (2015).

Chapter 14 · Part V — Teaching

Teaching Critical Thinking

If you want to pass these ideas on, start here. The research is clear on one thing: telling people to think critically almost never works. Here's what does, and how to design your first session.

Everything in this book so far has been preparation for this chapter. You now know how reasoning works (Chapter 4), why people resist (Chapter 2), whom they trust (Chapter 12), and why knowing about biases isn't enough (Chapter 10). Each of those findings turns out to be a teaching principle.

Predict Before you read

Before India's 2019 election, researchers gave 1,224 people in Bihar an hour-long, in-person training on spotting misinformation. What happened to their ability to tell true stories from false ones?

How sure are you?60%

No significant improvement. In Sumitra Badrinathan's study, published in 2021, the hour-long training didn't improve people's ability to identify misinformation on average. Supporters of the ruling party who received it actually became less able to spot false stories that favoured their own side. Motivated reasoning (Chapter 2) beat a one-off lesson.

Now the hopeful half. In a later study published in 2025, Badrinathan and colleagues worked with about 13,500 teenagers in 583 villages in Bihar. Students took four interactive 90-minute sessions over about 14 weeks, focused mostly on health information. Their ability to tell true from false improved by about a third of a standard deviation — a meaningful effect — and it lasted at least four months. It even carried over to political news that was never discussed in class. And the students' parents improved too.

Same state, same problem, opposite results. The differences — more time, repeated practice, interaction, younger learners, a less politically charged topic — are exactly what this chapter is about.

14.1What the research says works

In 2015, Philip Abrami and colleagues published a large review of studies on teaching critical thinking: 341 effect sizes from many classrooms. The average effect was positive but modest (an effect size of about 0.3). Three instructional ingredients stood out:

  • Dialogue — discussion, questioning and debate, not just listening.
  • Authentic problems — real, engaging problems relevant to the learners' lives (the researchers call this “anchored instruction”).
  • Mentoring — one-to-one guidance and feedback from someone more experienced.

They worked best in combination. That's also a fair description of a good workshop: people wrestling with real problems, arguing about them, with someone guiding and giving feedback.

Why transfer is hard

The psychologist Daniel Willingham argued in a widely read 2007 article that critical thinking isn't a general skill like riding a bicycle, which, once learned, works everywhere. People who reason brilliantly about cricket statistics can reason badly about health claims. Thinking well depends on knowing the domain — and on recognising that a new problem has the same deep structure as one you've seen before. That recognition is what usually fails.

So teaching for transfer means:

  • Name the tool explicitly — “this is a denominator problem”, “this is regression to the mean”. A named pattern is easier to spot again.
  • Use many examples of the same idea in different settings — health, business, cricket, family, news.
  • Put two examples side by side and ask what they share. Research by Dedre Gentner and colleagues found that comparing two cases helps people extract the underlying principle better than studying the cases one at a time.
  • Practise over time. One session rarely sticks; four spaced sessions did in Bihar.

14.2Don't make it a fight

You might fear that correcting people makes them dig in — the so-called backfire effect. The good news: in large studies, genuine backfire is rare. Thomas Wood and Ethan Porter tested corrections on 52 issues with over 10,000 participants and found that, on average, people moved towards the facts, even on politically charged topics — though often modestly.

The bad news: people move far less when a belief is tied to their identity, and a one-off correction rarely changes deeper attitudes. What does?

Experiment · Miami, 2016

David Broockman and Joshua Kalla studied a method called deep canvassing. Canvassers knocked on doors and had conversations of around ten minutes about prejudice against transgender people. They didn't lecture. They asked people about their own experiences of being judged, listened without judging, and shared stories. Those conversations reduced prejudice measurably, and the change lasted at least three months — longer than most persuasion effects.

The pattern matches everything in this book. People change their minds through their own reasoning (Chapter 4), when they don't feel attacked (Chapter 2), in conversation with someone who has shown they understand (the Ideological Turing Test, Chapter 2).

Key idea: Start with low-stakes material, not identity. Teach the tools on health scams, ads, business claims and puzzles. People then carry the tools into harder topics themselves. Their conclusions will be theirs.

14.3Prebunking: vaccinate, don't just cure

Remember Brandolini's law (Chapter 6): refuting nonsense takes far more effort than producing it. You can't debunk every false claim. But you can teach people to recognise the techniques behind them.

The psychologists Jon Roozenbeek and Sander van der Linden call this inoculation or prebunking: exposing people to a weakened dose of a manipulation technique, and showing how it works, so they resist it later. Their online game Bad News puts players in the role of a fake-news producer, using techniques like impersonation, emotional language, polarisation and conspiracy. Playing it improved people's ability to spot those techniques. In a 2022 study, short prebunking videos about manipulation techniques, shown as YouTube ads in a large field test, also improved people's recognition of those techniques.

For your teaching, the lesson is to teach patterns rather than individual claims: “Here's how an emotional hook works. Here's what a false dilemma sounds like. Here's the ‘doctors don't want you to know’ move.” Then let people find the patterns in real examples themselves.

Quiz

You have 90 minutes with a group of parents. Which opening is most likely to work?

14.4Teach with questions

The oldest method of teaching critical thinking is also one of the best: Socratic questioning. Instead of telling people what to think, you ask questions that help them examine what they already think. It works because it uses the evaluator in each person on their own ideas — the one thing they usually skip.

Tool 23

The Socratic Question Bank

  • Clarify: “What exactly do you mean by…?” “Can you give an example?”
  • Assumptions: “What are we taking for granted here?” “What if that weren't true?”
  • Evidence: “How do we know?” “What would we expect to see if it were false?”
  • Source: “Who first said this, and how did they know?”
  • Alternatives: “What else could explain it?” “How would someone who disagrees see it?”
  • Implications: “If this is true, what else would have to be true?”
  • Confidence: “How sure are you, as a number?” “What would change your mind?”

Ask with curiosity, not as a trap. If the person feels cross-examined, you've become the soldier.

14.5Designing your first session

Here's a blueprint that combines everything above. It's built for about 90 minutes with 8–20 people, but it scales down to a conversation with one friend.

Tool 24

The Lesson Blueprint

  1. Hook (10 min). A surprising Predict question. Everyone commits to an answer and a confidence number, in writing.
  2. Argue (10 min). Pairs compare answers and try to persuade each other. Then reveal.
  3. Story (10 min). A true story showing the idea at work — local where possible.
  4. Tool (10 min). One named tool, three to five steps. Just one.
  5. Practice (25 min). Small groups apply the tool to real, low-stakes examples: ads, forwards, business claims. Rotate who leads.
  6. Journal (10 min). Each person writes one belief of their own, with a confidence number and what would change their mind.
  7. Teach back (10 min). Each person explains the tool to a partner in three sentences.
  8. Next time (5 min). One field-work task before the next session.

Every chapter of this book follows this pattern. Four spaced sessions will do far more than one long one.

Measure whether it worked

Don't rely on people saying they enjoyed it. At the start and end of a course, give participants a short set of the same kinds of tasks — spot the flaw, judge a source, answer a base-rate question — with confidence ratings. Compare accuracy and calibration. Better still, check again a month later. That's how you'll know whether you're changing thinking, and it's the honest way to describe your results to anyone who pays for your teaching.

14.6The ethics of teaching thinking

  • Teach methods, not conclusions. The moment a critical thinking course becomes a way to bring people round to the teacher's views, it has become the thing it warns against. Your learners should end up able to disagree with you well.
  • Show your own mistakes. Share predictions you got wrong, beliefs you changed. It's the most convincing demonstration of scout mindset, and it makes it safe for others.
  • No contempt. Words like “sheep” or “idiots” for people who believe false things are soldier mindset in scout's clothing. Chapter 1 showed why believing what trusted people tell you is normal and often wise.
  • Tools are not weapons. Teach people to use these tools on their own beliefs first, and on others' only with care.
Apply Design your first session

Design a real 60–90 minute session for a specific audience you could reach within the next two months. Use the Lesson Blueprint. Who is it for, what's the hook, which tool, which examples, and how will you measure it?

See one worked example

Audience: 12 small-business owners from my network. Session title: “Decisions that don't lie to you.”

Hook: The flight-instructor question (Chapter 7), with confidence ratings.

Story: A version of “we introduced an incentive after the worst month and things improved” from a real business.

Tool: The Cause Checklist (Tool 12), cut to five questions.

Practice: Each person brings one “X caused Y” belief about their business; small groups run the checklist and design a fair test.

Journal: Log the belief, a confidence number, and the test.

Measure: Five short scenario questions with confidence ratings at the start and end; a follow-up message after a month asking whether anyone ran their test.

14.7Where this leaves you

In Chapter 1, you rewrote the book's opening thesis about the world running on belief. Here's your version, if you saved it:

Your Chapter 1 thesis

You haven't written this yet. Go back to Chapter 1, section 1.7.

You've since learned that humans run on trust, not doubt; that reasoning works best between people who disagree; that numbers, causes and studies can mislead in predictable ways; that confidence should be a number, and that the number should be scored; that knowing about biases isn't enough, but processes are; and that people change their minds when they reason their own way there.

That last point is the one to carry into your teaching. You can't think for anyone else. But you can give people good questions, safe disagreement, and a few tools — and then let them do the most human thing there is: learn from each other.

Final reflection Your next step

Write down what you want to do with critical thinking from here — for yourself, your family, your work or others — and how you'll know it's working. Compare it with your Chapter 1 thesis above. What changed?

Teach it back The last one

In three sentences, explain what critical thinking is to someone who has never heard the term.

Compare with a model answer

“Critical thinking isn't doubting everything — it's getting better at deciding what to believe and how strongly. It means asking what exactly is being claimed, how anyone could know it, what else could explain it, and how sure you should be. And it works best when you do it with other people who see things differently, because we're all much better at spotting mistakes in each other's thinking than in our own.”

Field work Your next month

Takeaways

  • One-off lessons rarely work; sustained, interactive practice can — with lasting effects.
  • What works: dialogue, authentic problems and mentoring, combined. Name tools and practise across contexts for transfer.
  • Corrections rarely backfire, but identity threat blocks change. Start low-stakes; listen before persuading.
  • Prebunk patterns rather than debunking claims one by one.
  • Teach with questions, measure honestly, show your own mistakes, and teach methods — never conclusions.

Sources and further reading

  • Sumitra Badrinathan, “Educative Interventions to Combat Misinformation: Evidence from a Field Experiment in India,” American Political Science Review (2021); Badrinathan et al., “Countering Misinformation Early: Evidence from a Classroom-Based Field Experiment in India,” APSR (2025).
  • Abrami et al., “Strategies for Teaching Students to Think Critically: A Meta-Analysis,” Review of Educational Research (2015).
  • Daniel Willingham, “Critical Thinking: Why Is It So Hard to Teach?” American Educator (2007); Gentner, Loewenstein & Thompson, “Learning and transfer: A general role for analogical encoding” (2003).
  • Wood & Porter, “The Elusive Backfire Effect,” Political Behavior (2019); Broockman & Kalla, “Durably reducing transphobia,” Science (2016).
  • Roozenbeek & van der Linden, “Fake news game confers psychological resistance against online misinformation,” Palgrave Communications (2019); Roozenbeek et al., “Psychological inoculation improves resilience against misinformation on social media,” Science Advances (2022).
  • Guess et al., “A digital media literacy intervention increases discernment between mainstream and false news in the United States and India,” PNAS (2020).

Reference

Your Toolkit

Every tool from the book in one place — 24 of them. Keep these handy; they're also what you'd teach.

Practice · Private to this device

Thinking Journal

Write down what you believe, how sure you are, and what would change your mind. Come back later and update it. Over time this becomes a record of how your thinking moved. Pick your pen and paper while you write.

70%
Calibration

When you say 80%, are you right about 80% of the time?

Built from the Predict questions you've locked in inside chapters, plus journal predictions you've resolved. It needs 20+ data points before it means much.

About & contact

About this book

A free, interactive book for anyone who wants to get better at deciding what to believe — and help others do the same.

How Do You Know? started from a simple worry: most of what people believe arrives through long chains of retelling, and very few of us were ever taught how to decide which beliefs deserve our trust. The book draws on research in psychology, statistics, science and education, with examples from India and around the world. Every chapter lists its sources.

It's designed to be used, not just read. Each chapter asks you to predict before you read, answer questions that explain every option, write in a private journal, and teach the idea back in your own words.

The author

Manthan Jha is based in Ahmedabad. He built and grew two businesses before starting GinWorks, which builds custom systems that businesses own outright. This book is his attempt to learn critical thinking properly — and to make it easier for others to learn it too.

Get in touch

Questions, corrections, ideas — or a session for your school, college or team. Messages are read by Manthan himself. English, Hindi or Gujarati.

Instagram @ginworks.ai

Your privacy

There are no accounts, no ads and no tracking cookies. Your answers and journal never leave your device. Visits are counted anonymously, without cookies. Read the full privacy note.

A note on the content

This book is for learning. It isn't medical, legal or financial advice. Examples are simplified, and research moves on — if you spot an error, please say so; corrections are welcome and will be made.

Typefaces: Fraunces, Inter, Atkinson Hyperlegible, Kalam, Caveat, Patrick Hand and Courier Prime, all under the SIL Open Font License and served from this site.