# Lesson 1: Introduction and Causal Concepts

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson one, Introduction and Causal Concepts.

**Sarah:** It's meant for after you've finished the lesson. We'll add some perspective and some critique, spend time on the questions students tend to find thorny in this material, and work through a few extra examples, including some that are harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. What did John Snow actually show about cholera? If we know the average effect of a treatment, do we know who it helped? When does adjusting for a variable make an estimate worse? And how can the attributable fractions for one disease add up to more than one hundred percent? We'll finish with the Bradford Hill viewpoints, where Kiffer and I see things a little differently.

**Kiffer:** Let's start with Snow.

**Sarah:** Question one. What did John Snow actually show? The version most people remember is the pump. He mapped cholera deaths around the Broad Street pump in 1854, the handle came off, and the outbreak ended.

**Kiffer:** The handle did come off, on the eighth of September 1854. Snow himself wrote, though, that the attacks had already fallen off so far before the water was stopped that it was impossible to say whether the well was still dangerous. Many residents had also fled the neighbourhood. So the pump handle makes a memorable story, and on its own it's weak evidence.

**Sarah:** Then where's the strong evidence?

**Kiffer:** In what Snow called a grand experiment. Two water companies, the Southwark and Vauxhall Company and the Lambeth Company, piped water into the same districts. In 1852, Lambeth moved its intake upriver, to a supply free of the city's sewage. Southwark and Vauxhall kept supplying water that contained it. Snow wrote that the pipes of both companies went down the same streets, and that in many cases a single house had a different supplier from the houses on either side.

**Sarah:** And he counted deaths by company.

**Kiffer:** Yes. For the first seven weeks of the 1854 epidemic, houses supplied by Southwark and Vauxhall had three hundred and fifteen cholera deaths for every ten thousand houses. Houses supplied by Lambeth had thirty-seven. So here's a question for everyone. Roughly how many times higher was the death rate in the Southwark and Vauxhall houses? Take a few seconds.

*(Pause)*

**Sarah:** Three hundred and fifteen divided by thirty-seven is about eight and a half. So the Southwark and Vauxhall houses had roughly eight and a half times as many deaths per house.

**Kiffer:** Right. Snow put it at between eight and nine times as great. The ratio is large, and the design is what makes it convincing.

**Sarah:** Because the two groups of houses were mixed together on the same streets.

**Kiffer:** Yes. Snow wrote that each company supplied rich and poor, large houses and small, and that no fewer than three hundred thousand people had been divided into two groups without their choice, and in most cases without their knowledge.

**Sarah:** So in the terms of this lesson, he was claiming exchangeability. Had the Lambeth households received the other company's water, their death rate would have looked like their neighbours'. I want to push on that, though. Somebody chose the water company at some point.

**Kiffer:** That's a fair challenge. Snow says the supplier had been chosen by the owner or occupier back when the companies were competing for customers, so this was no coin toss. The case for exchangeability rests on that old choice having nothing to do with cholera risk. There's a second weakness. His denominator was houses, because a return made to Parliament gave the number of houses each company supplied, and houses hold very different numbers of people.

**Sarah:** So a modern reviewer would ask whether the Southwark and Vauxhall houses were more crowded.

**Kiffer:** That's the right instinct. Crowding would have to differ a great deal between the two groups to explain a ratio of eight or nine, and Snow's description of how mixed the supply was gives little reason to expect that.

**Sarah:** Is there anything else in this history that students tend to get wrong?

**Kiffer:** One thing. You'll sometimes hear that Robert Koch discovered the cholera organism. Filippo Pacini described it in Florence in 1854, the same year as Broad Street, and his discovery was largely overlooked. Koch isolated it in pure culture in Calcutta in 1884, about thirty years after Snow's work. The organism's formal name, Vibrio cholerae Pacini 1854, credits Pacini.

**Sarah:** So Snow and Pacini had pieces of the same answer in the same year, from opposite directions.

**Kiffer:** Right. One worked from the pattern of deaths across a population, and the other worked at a microscope. That's the tension between population thinking and mechanism that runs through the whole history of the field.

**Sarah:** Question two. If we know the average effect of a treatment, do we know who it helped?

**Kiffer:** Let's use made-up numbers. Picture a large, well-run randomized trial of a new drug, scaled down to one thousand people. It tells us that if nobody got the drug, fifteen percent would die within a year, and if everybody got it, ten percent would.

**Sarah:** So the average treatment effect is ten percent minus fifteen percent, a drop of five percentage points.

**Kiffer:** Right. Here's the question. Does that mean exactly fifty of those thousand people were saved by the drug? Take a few seconds.

*(Pause)*

**Sarah:** My first instinct is yes. One hundred and fifty deaths without the drug, one hundred with it, so fifty people saved.

**Kiffer:** That's the usual answer, and fifty is only the smallest number consistent with the totals. Think about four kinds of people. Some would die either way, and some would survive either way. Some would die only without the drug, so the drug saves them. And some would die only with it, say from a rare side effect.

**Sarah:** And we can never tell which kind a given person is, because we only see one of their two potential outcomes. Someone who dies after taking the drug might have died anyway, or might have been killed by it, and the data look the same.

**Kiffer:** That's the fundamental problem of causal inference. So try this version. The drug saves one hundred people and kills fifty. Another fifty would die either way. Count the deaths if nobody gets the drug.

**Sarah:** The fifty who die either way plus the hundred the drug would have saved. That's one hundred and fifty. If everybody gets it, it's the fifty who die either way plus the fifty it kills, so one hundred. Those are exactly the same totals.

**Kiffer:** Exactly. The totals pin down the net effect, fifty more people helped than harmed per thousand. The number helped could be anywhere from fifty to one hundred and fifty.

**Sarah:** So the vaccine thought experiment from the lesson, where the vaccine causes the disease in one person, describes one of those harmed people.

**Kiffer:** Yes. A treatment that helps on average can still harm some individuals, and comparing group averages can't show you which ones. That's why a trial result supports a decision about a population, such as whether to recommend a drug. A clinician can tell a patient what tends to happen on average to people like them, and nobody can promise that this patient will be one of the people the drug saves.

**Sarah:** And all of this depends on exchangeability. A trial's two observed risks stand in for those two totals only if the treated and untreated groups would have had the same risk had they swapped places.

**Kiffer:** Right, which is what randomization buys you, and what Snow had to argue for.

**Sarah:** Question three. When does adjusting for a variable make an estimate worse? A lot of students start out believing that adjusting for more variables is always safer. That belief works for confounders. In a directed acyclic graph, a confounder sits on a fork, as a common cause of the exposure and the outcome, and adjusting for it removes a distortion.

**Kiffer:** Right, and the belief fails for the other two building blocks. If you want the total effect, adjusting for a mediator on a chain removes part of the effect you're trying to measure. Adjusting for a collider, where two arrows meet, can create an association that isn't there.

**Sarah:** Colliders are the ones I find hardest to picture. Can we put numbers on one?

**Kiffer:** Sure, made-up ones. A town has one thousand adults. Twenty percent have chronic back pain and twenty percent have high blood pressure, and in this town the two are completely unrelated. A research team recruits from a clinic. Every adult in town with at least one of the two conditions attends it, and nobody else does.

**Sarah:** So attending the clinic is the collider. Back pain points into it, and high blood pressure points into it.

**Kiffer:** Right. Here's the question. Among clinic patients who don't have back pain, what share have high blood pressure? Take a few seconds.

*(Pause)*

**Sarah:** All of them. Everyone at the clinic has at least one of the two conditions, so anyone there without back pain must have high blood pressure.

**Kiffer:** And among clinic patients with back pain?

**Sarah:** Let me count. Twenty percent of twenty percent of a thousand is forty people with both. A hundred and sixty have back pain only, and a hundred and sixty have high blood pressure only. So two hundred clinic patients have back pain, and forty of them have high blood pressure. That's twenty percent.

**Kiffer:** So inside the clinic, it's twenty percent with back pain compared with one hundred percent without it, and back pain looks strongly protective against high blood pressure. In the town as a whole, it's twenty percent in both groups. Selecting people through the clinic created the association.

**Sarah:** That's a made-up clinic. Does this happen in real research?

**Kiffer:** It does. A well-known case comes from infant health. In 1991 data from the United States, babies born to mothers who smoked were more likely to have low birth weight and more likely to die in infancy. Yet among low birth weight babies, those born to smokers had lower infant mortality, with a relative rate of 0.79.

**Sarah:** Which makes it look as if smoking protects small babies.

**Kiffer:** That's how it looks. A 2006 paper in the American Journal of Epidemiology used causal diagrams to explain it. Smoking lowers birth weight. Other causes of low birth weight, such as birth defects, also raise the risk of death. So birth weight is a collider between smoking and those other causes. Among small babies, one whose mother didn't smoke is more likely to be small for a more dangerous reason.

**Sarah:** But birth weight is also a mediator. Smoking affects it, and it affects survival.

**Kiffer:** Yes, and that's what makes it thorny. The same variable can sit on a chain and be a collider on another path. The authors concluded that, under realistic causal diagrams, adjusting for birth weight is unwarranted when the goal is the overall effect of a prenatal exposure, such as smoking, on infant death.

**Sarah:** So the rule is to draw the directed acyclic graph first, and to be wary of adjusting for anything the exposure itself affects, including the way people got into the sample.

**Kiffer:** That's the rule I'd use. Selecting a sample is a form of conditioning, so a sample drawn from a clinic, a hospital or a list of volunteers deserves the same scrutiny as any variable you plan to adjust for.

**Sarah:** Question four. How can the attributable fractions for one disease add up to more than one hundred percent? I'd like to try this one myself.

**Kiffer:** Okay. Here's a made-up town of ten thousand people. Half have exposure A and half have exposure B, independently, so there are two thousand five hundred people in each of the four combinations. Everyone has a one percent risk from causes that don't involve A or B. People with both A and B have a ten percent risk, because A and B are components of the same sufficient cause, along with other components we haven't measured. People with only one of them stay at one percent.

**Sarah:** So neither A nor B does anything on its own.

**Kiffer:** Right. Using Levin's formula from the lesson, what's the population attributable fraction for A? Take a few seconds.

*(Pause)*

**Sarah:** I'll start with the risk ratio. Among people with A, the two thousand five hundred with both are at ten percent, which is two hundred and fifty cases, and the two thousand five hundred with A only are at one percent, which is twenty-five. That's two hundred and seventy-five cases among five thousand people, or 5.5 percent. Among people without A, it's twenty-five plus twenty-five, fifty cases among five thousand, or one percent. So the risk ratio is 5.5.

**Kiffer:** Good. Now the formula.

**Sarah:** The prevalence of exposure times the risk ratio minus one, divided by that same quantity plus one. The risk ratio minus one is 4.5, and one half times 4.5 is 2.25. Then 2.25 divided by 3.25 is about 0.69. So about sixty-nine percent of cases are attributable to A.

**Kiffer:** And B?

**Sarah:** It's symmetric, so B is also sixty-nine percent. Which means removing both would prevent a hundred and thirty-eight percent of cases. Okay, that can't be right. You can't prevent more cases than there are.

**Kiffer:** Good catch. So work out what actually happens if you remove both.

**Sarah:** Then everyone is at the one percent background risk, so there are one hundred cases. Before, there were three hundred and twenty-five. So we prevent two hundred and twenty-five, which is still about sixty-nine percent.

**Kiffer:** Exactly. Removing A prevents those two hundred and twenty-five cases, and removing B prevents the same two hundred and twenty-five. Each fraction answers its own question, what would happen if we removed this one exposure. Those extra cases all occur in people with both A and B and need both, so each one is counted once in A's fraction and again in B's. That's why adding the fractions can't tell you the effect of removing both.

**Sarah:** So a sum above one hundred percent is a sign that the causes work together, and adding the fractions is where people go wrong.

**Kiffer:** Right. One more turn of the same example connects to causal complements. Keep A at fifty percent, and suppose only ten percent of people have B. What happens to the risk ratio for A?

**Sarah:** Among people with A, the ninety percent without B are at one percent risk, which contributes 0.9 percent, and the ten percent with B are at ten percent risk, which contributes one percent. That's 1.9 percent. People without A are still at one percent. So the risk ratio for A drops from 5.5 to 1.9.

**Kiffer:** And nothing about A's biology changed. In the respiratory example from the lesson, a co-factor becoming more common lowered the risk ratio, because that co-factor also caused disease without the exposure. Here B works only alongside A, so a rarer B makes A look weaker. Both cases show the same thing. The strength of an association depends on how common its causal complements are.

**Sarah:** Is there a real example of two causes working together like this?

**Kiffer:** Radon and smoking are a Canadian one. Health Canada links about sixteen percent of lung cancer deaths in Canada to radon, and the Canadian Cancer Society attributes about seventy-two percent of lung cancer cases to smoking. Health Canada calls radon and smoking a dangerous combination, and puts the risk of lung cancer for a smoker with long-term exposure to high radon at about one in three.

**Sarah:** So a lung cancer in a smoker who lived with high radon could sit inside both of those numbers.

**Kiffer:** Under the component-cause model, yes. So you can't add those two figures, and you can't find one by subtracting the other from one hundred. And for cancers that need both exposures, removing either one would prevent them.

**Sarah:** Okay. Now the part where we disagree. The Bradford Hill viewpoints.

**Kiffer:** Go ahead.

**Sarah:** I think several of them mislead students. Take strength. According to the 2006 United States Surgeon General's report, secondhand smoke raises a non-smoker's risk of lung cancer by about twenty to thirty percent, and it's accepted as a cause. A strong association, meanwhile, can come entirely from confounding. Specificity is worse. Smoking causes many diseases, so the idea that one cause should have one effect fails on the best-known causal question in epidemiology.

**Kiffer:** Those are fair points, and Hill anticipated some of them. He wrote that none of his nine viewpoints could bring indisputable evidence for or against cause and effect, and that none could be required as a sine qua non, meaning an absolute requirement. He described them as viewpoints from which we should study association before we cry causation.

**Sarah:** Yet they're often used as a checklist, as if meeting six of the nine settled the question. Rothman and Greenland, among others, have criticized that use. And plausibility depends on what people already believe. Snow's water theory looked implausible to the officials of his day. After the outbreak they put the pump handle back, and the Board of Health blamed the epidemic on miasma.

**Kiffer:** I agree that a score out of nine is a misuse. Where I'd disagree is on dropping them. A causal diagram helps you design and analyse one study. It doesn't tell you how to weigh thirty studies of different designs, some agreeing and some not. Hill's viewpoints are questions for that job. Does the exposure come before the outcome? Is there a gradient? Do studies with different designs, and so different biases, point the same way?

**Sarah:** Couldn't you get there by asking, for each study, what chance, bias or confounding could explain its result?

**Kiffer:** You could, and that's the right question for each study. I'd still keep the viewpoints, because some things, like consistency across designs with different weaknesses, only show up when you look across the whole body of evidence. And temporality is the one viewpoint later writers treat as a firm requirement, because a cause has to come before its effect.

**Sarah:** I'd still give them less weight than Kiffer would, but we agree on the practical rule. Treat temporality as the one requirement, use the other eight as prompts to look for chance, bias and confounding, and never add them up into a score.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, a causal claim compares what happened with what would have happened otherwise. Snow's comparison persuades because the two groups of households were plausibly exchangeable.

**Kiffer:** Second, an average effect is a net effect. It tells you how many more people were helped than harmed, and it can't tell you which individuals those were.

**Sarah:** And third, think through the causal structure before you adjust for anything or add anything up. Colliders and shared sufficient causes are where the intuitive answer goes wrong.

**Kiffer:** If you'd like more practice, rework the attributable fraction example from this episode with different prevalences for A and B, and check every answer by counting the cases.

**Sarah:** Next time, it's Lesson two, Surveillance and Sampling, where we look at how public health systems detect disease and how investigators sample when they can't study everyone.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson two.
