# Lesson 6: Measures of Association

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson six, Measures of Association.

**Sarah:** It's for listening once you've worked through the lesson. We'll add some perspective on how these measures get used, and some critique of how they get misread. We'll also work through the questions students tend to find thorny with this material, and do some extra worked examples, a few of them harder than the ones on the lesson page.

**Kiffer:** At a few points we'll ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. Why can't a case-control study give you a risk ratio? When does an odds ratio make an effect sound bigger than it is? How can the same drug be clearly worth taking for one group of people and barely worth it for another? And how can two causes each account for most of the same cases? We'll finish with confidence intervals and the word significant, where Kiffer and I don't fully agree.

**Kiffer:** Let's start with case-control studies.

**Sarah:** Question one. Why can't a case-control study give you a risk ratio? I know the short answer is that the investigator chooses how many cases and controls to collect. I'd like to see what that does to the numbers.

**Kiffer:** Let's build it with made-up numbers. Imagine a cohort of fifty thousand people. Ten thousand are exposed to something, and one hundred of them develop a disease. Forty thousand are unexposed, and one hundred of them develop it too.

**Sarah:** So the risk in the exposed is one hundred out of ten thousand, which is one percent. The risk in the unexposed is one hundred out of forty thousand, which is a quarter of a percent. The risk ratio is four.

**Kiffer:** Now suppose nobody ran that cohort. A researcher finds all two hundred cases and picks two hundred controls at random from the people without the disease. About one in five of those people is exposed, so about forty controls are exposed and one hundred and sixty are unexposed.

**Sarah:** If I treat that table as though it were a cohort, the exposed group is one hundred cases plus forty controls, so the so-called risk is one hundred out of one hundred and forty, about 0.71. In the unexposed group it's one hundred out of two hundred and sixty, about 0.38. Dividing gives about 1.86. The odds ratio is one hundred times one hundred and sixty, divided by one hundred times forty, which is four.

**Kiffer:** So here's a question for everyone. Suppose the researcher had collected twice as many controls, four hundred in all. Which of those two numbers would change, and what would it become? Take a few seconds.

*(Pause)*

**Sarah:** With four hundred controls, about eighty are exposed and three hundred and twenty are unexposed. The pretend risk ratio becomes one hundred out of one hundred and eighty, divided by one hundred out of four hundred and twenty, which is about 2.33. The odds ratio is one hundred times three hundred and twenty, divided by one hundred times eighty, and that's still four.

**Kiffer:** Right. The so-called risks in a case-control table depend on how many controls the researcher decided to collect, so their ratio has no meaning. The odds ratio holds steady because doubling the controls doubles both control cells, and that cancels out in the cross-product. And because of its symmetry, comparing the odds of exposure in cases and controls gives the same number as comparing the odds of disease in the exposed and unexposed.

**Sarah:** And the odds ratio of four matches the risk ratio of four from the cohort because the disease is rare.

**Kiffer:** Yes. In the full cohort, even the exposed group's risk is only one percent, well under the five percent rule of thumb, so the odds and the risks are nearly the same. The cohort odds ratio is about 4.03. It's also worth noticing what the design bought us. The case-control study got the same answer from four hundred people that the cohort needed fifty thousand people to produce. That efficiency is why case-control studies are the usual choice for rare diseases.

**Sarah:** And with density sampling, rarity doesn't even matter.

**Kiffer:** That's right, as long as it's the rate ratio you want. If each control is picked from the people still at risk at the moment a case occurs, the odds ratio estimates the rate ratio directly, whether the disease is rare or common.

**Sarah:** Question two. When does an odds ratio make an effect sound bigger than it is?

**Kiffer:** When the outcome is common, and there's a well-known real case. In 1999, the New England Journal of Medicine published a study in which seven hundred and twenty physicians each watched a recorded interview with an actor playing a patient with chest pain, and then decided whether to recommend cardiac catheterization. The study reported an odds ratio of 0.6 for referral of Black patients compared with white patients.

**Sarah:** And the news coverage said Black patients were forty percent less likely to be referred.

**Kiffer:** That's how it was widely reported. But referral was common. It was recommended for 84.7 percent of the Black patients in the videos and 90.6 percent of the white patients. Can you work out the risk ratio and the odds ratio from those?

**Sarah:** The risk ratio is 0.847 divided by 0.906, which is about 0.93. So Black patients were about seven percent less likely to be referred. For the odds, 0.847 divided by 0.153 is about 5.5, and 0.906 divided by 0.094 is about 9.6. Dividing those gives an odds ratio of about 0.57, close to the 0.6 that was reported.

**Kiffer:** So the odds ratio sat further from one than the risk ratio, as it always does, and with an outcome this common the gap is large. Reading 0.6 as forty percent less likely turned a difference of about six percentage points into something that sounded far bigger. Later that year, researchers at Dartmouth made this point in the same journal, and the journal's editors wrote that they shouldn't have allowed odds ratios in the abstract.

**Sarah:** Does that mean there was no real difference?

**Kiffer:** The study did find a difference. The Dartmouth reanalysis found it was concentrated among Black women, who were referred less often than any other group, and the authors traced much of that gap to the referrals for one older Black actress. They said the finding raised questions about the treatment of Black women that deserved a second look. The problem was the size the headline gave it.

**Sarah:** So the practical rule is that you can read an odds ratio as a risk ratio only when the outcome is rare, under about five percent. Above that, an odds ratio read as a risk ratio overstates the association, on either side of one.

**Kiffer:** And if the data come from a cohort or a trial, I'd report the risk ratio, since you can calculate it directly. The odds ratio is essential in a case-control study. In a cohort with a common outcome, it mostly creates room for misreading.

**Sarah:** Question three. How can the same drug be clearly worth taking for one group and barely worth it for another?

**Kiffer:** Here's a made-up example. A drug lowers the five-year risk of stroke from two percent to one and a half percent among people at low risk. That's a risk ratio of 0.75, or a relative reduction of twenty-five percent.

**Sarah:** I'll work out the number needed to treat.

**Kiffer:** Let's give everyone a chance first. How many low-risk people need to take this drug for five years to prevent one stroke? Take a few seconds.

*(Pause)*

**Sarah:** The drug cuts the risk by twenty-five percent, so the number needed to treat is one divided by 0.25, which is four.

**Kiffer:** Check that against the risks.

**Sarah:** If treating four people prevents one stroke, then treating a hundred people would prevent twenty-five strokes. But only two people in a hundred were going to have a stroke at all. That's impossible. I used the relative reduction, and the formula needs the risk difference. Two percent minus one and a half percent is half a percentage point, or 0.005, and one divided by 0.005 is two hundred.

**Kiffer:** So two hundred people take the drug for five years to prevent one stroke. Now take a high-risk group, where the same drug lowers the five-year risk from sixteen percent to twelve percent.

**Sarah:** The risk ratio is twelve divided by sixteen, which is 0.75 again. The risk difference is 0.04, and one divided by 0.04 is twenty-five. So the relative effect is identical, and the number needed to treat is eight times smaller.

**Kiffer:** Now add a harm, also made up. Suppose the drug causes a serious side effect in one person out of every hundred who take it, whatever their stroke risk. In the low-risk group, treating two hundred people prevents one stroke and causes two serious side effects. In the high-risk group, treating a hundred people prevents four strokes and causes one side effect.

**Sarah:** So a headline saying the drug cuts stroke risk by a quarter is true for both groups, and it tells you almost nothing about whether either group should take it.

**Kiffer:** That's the main reason to pair every ratio with a difference. The ratio describes the strength of the effect. The difference, and the number needed to treat that comes from it, tells you how much good the drug does in a particular group, and that's the figure you weigh against the harms. Always give the time frame as well. Two hundred over five years is a different claim from two hundred over one year.

**Sarah:** Question four. How can two causes each account for most of the same cases? Before we get there, I'd like to test the rule of thumb that a common, weak risk factor matters more to a population than a rare, strong one.

**Kiffer:** Here's a made-up pair. Factor X has a risk ratio of ten, and two percent of the population is exposed. Factor Y has a risk ratio of 1.3, and half the population is exposed. Which one accounts for the larger share of cases in the population? Take a few seconds.

*(Pause)*

**Kiffer:** Let's use the population attributable fraction formula. For factor X, the exposure prevalence times the risk ratio minus one is 0.02 times nine, which is 0.18. Divide that by 1.18 and you get about fifteen percent. For factor Y, it's 0.5 times 0.3, which is 0.15, divided by 1.15, about thirteen percent.

**Sarah:** So the rare, strong factor comes out slightly ahead. That's the reverse of what I expected.

**Kiffer:** The rule of thumb describes a tendency. What settles it is the product of the exposure prevalence and the risk ratio minus one, and whichever factor has the larger product has the larger population attributable fraction. If factor Y's risk ratio were 1.5, its product would be 0.25, and it would account for twenty percent.

**Sarah:** Now for a harder one, with real figures. Health Canada publishes estimated lifetime risks of lung cancer by radon level and smoking status. For a non-smoker exposed only to low, outdoor levels of radon, the lifetime risk is about one percent. At a high radon level, four times Canada's guideline for homes, it's about five percent. For a smoker, it's about twelve percent from smoking alone, and about thirty percent at that high radon level.

**Kiffer:** Let's compare radon's effect in the two groups on both scales. Start with non-smokers.

**Sarah:** For non-smokers, the risk ratio for high radon is five divided by one, which is five, and the risk difference is four percentage points. For smokers, the risk ratio is thirty divided by twelve, which is 2.5, and the risk difference is eighteen points.

**Kiffer:** So which group is radon worse for?

**Sarah:** On the ratio scale, it looks twice as strong in non-smokers. On the difference scale, it causes four and a half times as many extra cancers per hundred smokers as per hundred non-smokers.

**Kiffer:** Both statements are correct, and they answer different questions. If you want to know how many cancers fixing high-radon homes would prevent, the eighteen extra cases per hundred smokers matter more than the risk ratio does. Together, smoking and high radon raise the risk by twenty-nine points above the one percent baseline. That's nearly twice the fifteen points you'd get by adding their separate effects of eleven and four. So the two exposures modify each other's effects on the additive scale.

**Sarah:** Now the attributable fractions. Take smokers living with high radon, whose lifetime risk is thirty percent. With low radon, their risk would have been twelve percent, so the fraction of their cancers attributable to radon is eighteen divided by thirty, which is sixty percent. If they had never smoked, at the same radon level, their risk would have been five percent. So the fraction attributable to smoking is twenty-five divided by thirty, about eighty-three percent.

**Kiffer:** And sixty plus eighty-three is one hundred and forty-three percent.

**Sarah:** That looks impossible.

**Kiffer:** It looks impossible only if each case has a single cause. Think back to the causal pies from Lesson one. Many of these cancers need both smoking and radon in the same pie. Remove either one and that case doesn't happen, so the same case is counted in both fractions. Attributable fractions for different causes overlap, and they often add up to more than one hundred percent.

**Sarah:** So an attributable fraction answers one specific question. What share of the cases would disappear if this one exposure were removed and everything else stayed the same?

**Kiffer:** Yes, and it answers that only if the association is causal and free of confounding, which is where Lesson seven picks up. When a report lists the share of a disease due to each of several risk factors, read each share on its own, and don't add them up.

**Sarah:** Okay, now the part where we disagree. I'd like to start with a puzzle. A study reports an odds ratio with a ninety-five percent confidence interval from 1.25 to 3.2, and the point estimate has been cut off the page. What was the point estimate? Take a few seconds.

*(Pause)*

**Kiffer:** The tempting answer is the halfway point, about 2.2. But a ratio interval is built on the log scale, so the lower limit is the estimate divided by some factor, and the upper limit is the estimate times that same factor. That makes the estimate the square root of the two limits multiplied together. 1.25 times 3.2 is four, and the square root of four is two. You can check it. Two divided by 1.25 is 1.6, and 3.2 divided by two is also 1.6.

**Sarah:** And from that same interval you can recover the p-value.

**Kiffer:** You can, as long as it was calculated in the usual way. The natural log of 1.6 is about 0.47, and that distance is 1.96 standard errors, so the standard error on the log scale is about 0.24. The log of two is about 0.69, which puts the estimate about 2.9 standard errors from the null. That gives a two-sided p-value of about 0.004.

**Sarah:** So the interval holds everything the p-value does, and more. That brings us to the word significant.

**Kiffer:** My view is that we should stop using statistically significant as a verdict. In 2016, the American Statistical Association said that scientific conclusions and policy decisions shouldn't rest only on whether a p-value passes a threshold. In 2019, a comment in Nature with more than eight hundred signatories called for the concept of statistical significance to be abandoned. Its authors said they weren't calling for a ban on p-values, and they suggested calling confidence intervals compatibility intervals. I find that wording helpful.

**Sarah:** Why call them that?

**Kiffer:** Because it describes what the interval shows. Take a made-up odds ratio of 1.8 with an interval from 0.9 to 3.6. Calling that non-significant invites people to read it as no effect. Saying the data are compatible with anything from ten percent lower odds to about three and a half times the odds describes the evidence far better.

**Sarah:** I agree with all of that, and I'd still keep thresholds. When a decision has to be made, such as whether a trial shows a drug works well enough to approve, someone has to draw a line, and it's better drawn before the data come in. Without a line set in advance, it's easy to talk up whatever you happen to find. In 2021, a task force set up by the president of the American Statistical Association said that p-values and significance tests, properly applied and interpreted, increase the rigour of conclusions, and that thresholds can help when actions are required.

**Kiffer:** That's a fair point, and the same statement said a threshold should be chosen with the study's goals and the cost of a wrong decision in mind. Some researchers want a stricter line. In 2017, seventy-two of them proposed lowering the default threshold for claims of new discoveries from 0.05 to 0.005.

**Sarah:** So we're closer than it sounds. My concern is about decisions, and yours is about the word.

**Kiffer:** I think that's right. Where we still differ is that I'd drop the label from the results sections of most epidemiology papers, and you'd keep it wherever a threshold was set in advance.

**Sarah:** And we agree on the practical rule. Report the estimate, its confidence interval and the exact p-value. If you use a threshold, set it before you see the data and say why you chose it. And never read an interval that includes the null as proof that there's no effect.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, the design decides the measure. A case-control study gives you an odds ratio, and you can read it as a risk ratio only when the outcome is rare. When the outcome is common, the odds ratio sits further from one, and reading it as a risk ratio overstates the effect.

**Kiffer:** Second, pair every ratio with a difference. The same risk ratio of 0.75 meant two hundred people treated per stroke prevented in one group and twenty-five in another. And radon's effect was larger among non-smokers on the ratio scale and larger among smokers on the difference scale.

**Sarah:** And third, an attributable fraction answers a what-if question about removing one exposure. For the whole population, it depends on the exposure prevalence times the risk ratio minus one, and the fractions for different causes overlap, so they can't be added together.

**Kiffer:** If you'd like more practice, rework the radon and smoking problem using Health Canada's figures for homes at the guideline level. There, the lifetime risk is about two percent for non-smokers and about seventeen percent for smokers. Compare radon's effect on both scales again, and work out the attributable fraction for each group.

**Sarah:** Next time, it's Lesson seven, Designing Against Bias: Validity and Confounding in Study Protocols. We'll look at building protection against bias into a study's design, including the confounding we set aside in this episode.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson seven.
