# Lesson 7: Designing Against Bias: Validity and Confounding in Study Protocols

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson seven, Designing Against Bias: Validity and Confounding in Study Protocols.

**Sarah:** It's meant for after you've finished the lesson. We'll bring in some perspective and some critique, work through the questions students tend to find thorny with this material, and add a few worked examples, including some that are harder than the ones on the lesson page.

**Kiffer:** At a few points we'll ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. If a variable is associated with both the exposure and the outcome, is it a confounder? Does a low response rate mean a study is biased? If misclassification is non-differential, is the estimate safely conservative? And what can a study say about a confounder it never measured? That last one brings in the E-value, which Kiffer and I see a little differently.

**Kiffer:** Let's start with the first one.

**Sarah:** In the pre-course survey, about one student in four described a variable on the causal pathway as a confounder. I understand why. If a variable is related to both the exposure and the outcome, it feels like something you should adjust for.

**Kiffer:** Being related to both is true of confounders, mediators and colliders alike. What separates them is the direction of the arrows. A confounder causes the exposure and the outcome. A mediator is caused by the exposure and goes on to cause the outcome. A collider is caused by both, or by their causes. So the question to ask about any variable is what causes it and what it causes, and the answer comes from subject-matter knowledge. The data can show that a variable is associated with both, and they can't show which way the arrows point.

**Sarah:** And adjusting for a mediator removes part of the effect you're trying to estimate. Is that the worst that can happen?

**Kiffer:** It can be worse, and there's a well-known real example. In United States birth records from 1991, babies born to mothers who smoked had higher risks of low birth weight and of dying in their first year. Among low birth weight babies only, though, the babies of smokers had lower infant mortality, with a relative rate of 0.79.

**Sarah:** Which would mean smoking protects small babies. That can't be right.

**Kiffer:** In 2006, Hernández-Díaz and colleagues used causal diagrams to explain it. Birth weight is a mediator, because smoking lowers it. It also shares causes with death in infancy, such as birth defects, which lower birth weight and raise mortality. Among low birth weight babies, a baby whose mother didn't smoke is more likely to be small for one of those more dangerous reasons. Restricting to small babies links smoking with the absence of those other causes, so smoking looks protective.

**Sarah:** So when a mediator shares causes with the outcome, it's a collider as well. Adjusting for it removes part of the effect and also opens a non-causal path between the exposure and the outcome.

**Kiffer:** Right. The authors concluded that adjusting for birth weight is unwarranted when the goal is the overall effect of a prenatal exposure on infant mortality.

**Sarah:** Couldn't the data settle it, though? You adjust, and if the odds ratio moves by more than ten percent on the log scale, you call the variable a confounder.

**Kiffer:** Adjusting for a mediator moves the estimate too, and so does adjusting for birth weight in that example. And when the outcome is common, a variable that affects the outcome can shift an odds ratio even with no link to the exposure, because the odds ratio is non-collapsible. A change in the estimate shows that something moved. It can't show the direction of the arrows, which is why the diagram comes first and the change-in-estimate check only supports it.

**Sarah:** Let's give everyone one to try. Here's a made-up study of whether living within one hundred and fifty metres of a major highway raises the risk of asthma in children. The investigators believe that household income at the child's birth affects where a family can afford to live, and also affects asthma through things like housing quality. They also believe that traffic pollution makes respiratory infections in the first two years of life more common, and that those infections raise the risk of asthma. Which of these two variables belongs in the adjustment set? Take a few seconds.

*(Pause)*

**Kiffer:** Household income belongs in the set. It's measured at birth, before the exposure, and under these assumptions it causes both where the family lives and the child's asthma risk, so it opens a backdoor path. Early respiratory infections are a mediator. Pollution makes them more common and they raise asthma risk, so for the total effect of living near the highway they stay out of the set.

**Sarah:** And if someone adjusted for infections anyway, the estimate would shrink, and the change-in-estimate check would make infections look like a confounder.

**Kiffer:** Which is the mistake the check can't catch on its own. The protocol should say when each covariate is measured, and the arrows should be drawn before anyone looks at the data.

**Sarah:** Question two. Does a low response rate mean a study is biased?

**Kiffer:** It depends on which estimate you mean. Groves and Peytcheva pooled fifty-nine studies that were designed to measure non-response bias, and in 2008 they reported very little correlation between a survey's response rate and the size of its bias. Groves and colleagues also noted that non-response bias usually varies more between different estimates from the same survey than between surveys with different response rates.

**Sarah:** There's a Canadian example of the worry about response. In 2011, the voluntary National Household Survey replaced the mandatory long-form census. Its response rate was 68.6 percent, and Statistics Canada didn't publish its standard estimates for areas where the global non-response rate was fifty percent or higher. The mandatory long form came back in 2016.

**Kiffer:** The survey's published figures are descriptive estimates, such as the share of households in a small town with low income. For a proportion like that, response that depends on the characteristic itself biases the estimate. For an odds ratio, the condition is narrower. Selection biases it only when response depends on the exposure and the outcome jointly.

**Sarah:** Can one survey show both?

**Kiffer:** Here's a made-up example that does. A telephone survey of adults over sixty-five asks about hearing loss and about whether they ever worked in a noisy job. People with hearing loss find phone interviews hard, so suppose they respond at a rate of 0.20, while people without hearing loss respond at 0.40, whatever their work history. The true prevalence of hearing loss is twenty-five percent.

**Sarah:** So here's the question for everyone. Which estimate is biased, the prevalence of hearing loss, the odds ratio for noisy work, or both? Take a few seconds.

*(Pause)*

**Kiffer:** Only the prevalence. Picture one thousand people. Two hundred and fifty have hearing loss, and a fifth of them respond, which is fifty. Seven hundred and fifty don't, and four in ten of them respond, which is three hundred. The survey sees fifty people with hearing loss among three hundred and fifty respondents, about fourteen percent, against a true twenty-five percent.

**Sarah:** And the odds ratio is fine, because the sampling fractions are 0.20 for both groups of cases and 0.40 for both groups of non-cases. The odds ratio of the sampling fractions is 0.20 times 0.40, divided by 0.20 times 0.40, which is exactly one.

**Kiffer:** So the same survey gives a badly biased prevalence and an unbiased odds ratio. Now change one thing. Suppose the invitation says it's a study of noise at work and hearing, and people who worked in noisy jobs and now have hearing loss find the topic relevant, so they respond at 0.25. Everyone else stays the same. The survey observes an odds ratio of 2.0. Sarah, what would it have been without the selection?

**Sarah:** Okay. Exposed cases are 0.25, unexposed cases 0.20, and both groups of non-cases 0.40. The odds ratio of the sampling fractions is 0.25 times 0.40, which is 0.10, divided by 0.20 times 0.40, which is 0.08. That's 1.25. So the true odds ratio is 2.0 times 1.25, which is 2.5.

**Kiffer:** Before you settle on that, which group is over-represented, and which way should that push the observed odds ratio?

**Sarah:** Exposed cases are the group most likely to respond, so there are too many of them in the data, and that makes the observed odds ratio too big. So the truth should be smaller than 2.0, and I made it bigger. The observed odds ratio equals the true one times 1.25, so the true one is the observed divided by 1.25. Two divided by 1.25 is 1.6.

**Kiffer:** That's it. Checking the direction before you trust the arithmetic catches that mistake. And notice how little it took. Raising one group's response from a fifth to a quarter inflated the odds ratio by twenty-five percent.

**Sarah:** Which is why the invitation matters so much. A neutral description, something like a survey of health and daily life, keeps the topic from sorting people by exposure and outcome together.

**Kiffer:** And the protocol should plan the bounding calculation we just did, over a range of assumed fractions, because the true ones are almost never known.

**Sarah:** Question three. If misclassification is non-differential, is the estimate safely conservative? This argument turns up often in discussion sections. Any error in measuring the exposure would have been the same in cases and controls, so the true effect is probably larger.

**Kiffer:** It's a tempting argument, and it has three weak points. First, toward the null describes what happens on average over many repetitions of a study. In 2005, Jurek, Greenland and colleagues used simulations to show that a single study with non-differential misclassification can still overestimate the true effect by chance. Second, the tendency toward the null holds for a binary exposure. With three or more categories it can fail. Third, it gives little comfort to a null finding, since misclassification can pull a real effect close to one.

**Sarah:** Let's do the categories one, because I find it surprising.

**Kiffer:** Here's a made-up case-control study of alcohol use and a disease, with three levels of drinking: none, moderate and heavy. In truth, among the cases, one hundred drink none, one hundred drink moderately and three hundred drink heavily. Among the controls, it's three hundred at each level. Compared with no drinking, the true odds ratio is 1.0 for moderate drinking and 3.0 for heavy drinking. So in this made-up example there's a threshold. Moderate drinking has no effect, and heavy drinking triples the odds.

**Sarah:** And the misclassification?

**Kiffer:** Drinking is self-reported, and heavy drinkers sometimes under-report. Suppose a fifth of true heavy drinkers say they drink moderately, and that this happens equally among cases and controls, so it's non-differential. What happens to the odds ratio for moderate drinking? Take a few seconds.

*(Pause)*

**Sarah:** A fifth of the three hundred heavy-drinking cases is sixty, so the moderate group among cases grows from one hundred to one hundred and sixty, and the heavy group falls to two hundred and forty. Among controls, sixty also move, so moderate becomes three hundred and sixty and heavy becomes two hundred and forty. For moderate drinking, the odds among cases are one hundred and sixty over one hundred, which is 1.6. Among controls it's three hundred and sixty over three hundred, which is 1.2. And 1.6 divided by 1.2 is about 1.33.

**Kiffer:** So the moderate category moved from 1.0 to 1.33, away from the null. The sixty heavy drinkers raise the moderate group by sixty percent among cases and by twenty percent among controls, because among the cases there are three heavy drinkers for every moderate drinker, and among the controls it's one for one.

**Sarah:** And heavy drinking stays at 3.0, because the same fifth was taken from cases and controls. Among cases, two hundred and forty over one hundred is 2.4, among controls two hundred and forty over three hundred is 0.8, and 2.4 divided by 0.8 is three.

**Kiffer:** Right. So the observed pattern, 1.33 for moderate and 3.0 for heavy, looks like a dose-response gradient. Verkerk and Buitendijk described this in 1992. When people under-report a behaviour, a true threshold can look like a dose-response relationship. And Dosemeci and colleagues showed in 1990 that with non-differential misclassification across several categories, the odds ratios for middle categories can be biased away from the null or even change direction.

**Sarah:** So a credible version of the conservative argument names the variable, the direction in which people are likely to be misclassified, and the number of categories.

**Kiffer:** And it's backed by a bias analysis with plausible values for sensitivity and specificity. If I were reviewing a protocol, I'd want to see that calculation planned for the main exposure, and I'd be cautious about a discussion section that uses non-differential misclassification to call its estimate a minimum.

**Sarah:** Question four. What can a study say about a confounder it never measured? This is the part where Kiffer and I disagree, about the E-value.

**Kiffer:** Let's start with numbers. Here's a made-up cohort of teenagers. Those who sleep less than seven hours on school nights have 1.8 times the risk of a depression diagnosis over the next three years, compared with those who sleep more, after adjustment for the measured confounders. The ninety-five percent confidence interval runs from 1.3 to 2.5. The E-value from the lesson is the risk ratio, plus the square root of the product of the risk ratio and the risk ratio minus one. What's the E-value for the estimate of 1.8? Take a few seconds.

*(Pause)*

**Sarah:** The product is 1.8 times 0.8, which is 1.44. Its square root is 1.2, and 1.8 plus 1.2 gives an E-value of 3.0.

**Kiffer:** And for the confidence interval, we use the limit closer to one, which is 1.3. Here the product is 1.3 times 0.3, which is 0.39, its square root is about 0.62, and so the E-value is about 1.92. An unmeasured confounder would need a risk ratio of at least three with both short sleep and depression, beyond the measured confounders, to explain the estimate away. To move the interval to include one, it would need about 1.9 with both.

**Sarah:** Here's my concern. In 2019, Ioannidis, Tan and Blum argued that the E-value follows almost directly from the estimate, so it adds no information beyond the estimate itself, and that there's no general rule for when an E-value is small enough to worry about. The next year, the same authors reviewed eighty-seven papers that reported E-values. Only fourteen concluded that unmeasured confounding was likely to threaten at least some of their main conclusions, and only nineteen compared their E-values with the strength of the confounders expected in their field. My worry is that a large E-value gets read as reassurance, and the thinking about confounding stops there.

**Kiffer:** Those are fair criticisms of how it's been used, and I'd still defend the tool itself. The same review matched those papers with papers from the same journal issues that didn't report E-values, and fifty-two of the sixty-nine comparison papers didn't comment on unmeasured confounding at all. The E-value gives every study a common scale for that vulnerability. And at the planning stage it tells investigators how strong a confounder they need to worry about before they collect anything.

**Sarah:** Then show me how it should be used.

**Kiffer:** Put it beside the confounders you can name. Suppose family conflict at home is the main worry, and with made-up values, it's present for thirty percent of short sleepers and twenty percent of other teenagers, and it multiplies the risk of depression by 2.5. For the bias factor from the lesson, the confounder's risk ratio minus one is 1.5. On top, 0.30 times 1.5, plus one, is 1.45. On the bottom, 0.20 times 1.5, plus one, is 1.30. Then 1.45 divided by 1.30 is about 1.115. Divide 1.8 by 1.115, and the adjusted risk ratio is about 1.61.

**Sarah:** So a confounder with a risk ratio of 2.5 for depression only brought the estimate down from 1.8 to about 1.6. That's mostly because its link with short sleep is weak. Thirty percent against twenty percent is a ratio of 1.5, far short of three.

**Kiffer:** Exactly. The E-value states the strength a confounder would need. The bias factor shows what a named confounder, with strengths you can defend, would actually do. Reported together, they show a reader how strong a confounder would have to be and how strong the likely ones are.

**Sarah:** I'm still more skeptical of the E-value than Kiffer is. We agree on the practical rule, though. Report the E-value for the estimate and for the confidence limit closer to one, name the confounders you're worried about, and show what plausible strengths for them would do to the estimate. A large E-value on its own never shows that confounding is absent.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, a variable's role comes from the direction of its arrows. Associations with the exposure and the outcome fit confounders, mediators and colliders alike, and adjusting for a mediator that shares causes with the outcome can create bias, as the birth weight example shows.

**Kiffer:** Second, work out the direction of a bias before you rely on a rule of thumb. Selection biases an odds ratio only when response depends on exposure and outcome jointly, and non-differential misclassification pulls toward the null on average, for a binary variable.

**Sarah:** And third, for a confounder you couldn't measure, report the E-value alongside the confounders you can name and the strengths that are plausible for them.

**Kiffer:** If you'd like more practice, rework the hearing loss telephone survey with your own response rates. Find the prevalence the survey would report, then the odds ratio of the sampling fractions, and check the direction of the bias before you divide.

**Sarah:** Next time, it's Lesson eight, Time-to-Event Data, where follow-up is incomplete and a study has to define time zero, the event and censoring.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson eight.
