# Lesson 4: Generalized Linear Models

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson four, Generalized Linear Models.

**Sarah:** It's meant for after you've worked through the lesson. We'll add some perspective on how these models are used and some critique of how their results get misread, work through the questions students tend to find thorny with this material, and do some extra worked examples, including a couple that are harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. Why doesn't an odds ratio of 0.5 mean half as likely? What does a relative risk ratio actually compare? How can leaving out an offset turn a result upside down? And where does overdispersion come from? We'll finish with a question Kiffer and I see a little differently, which is whether it's reasonable to treat a five-point rating as a number.

**Kiffer:** Let's start with the first one.

**Sarah:** Some perspective first. Self-rated health is one of the most familiar ordinal outcomes in Canadian data. Statistics Canada asks it in the Canadian Community Health Survey, and it usually reports the share of adults who rate their health as very good or excellent. That share was about sixty-one percent in 2020 and about fifty-two percent in 2023.

**Kiffer:** And that headline figure collapses five ordered categories into two. For a single number in a national report, that's a sensible choice, because it's easy to read. When the question is what predicts self-rated health, the ordinal model keeps all five categories and gives each predictor one odds ratio.

**Sarah:** That one odds ratio is where I'd like to try a problem, because I think it's easy to get wrong. These are made-up numbers. In a hypothetical survey, forty percent of food-secure fifty-year-olds rate their health as very good or excellent. An ordinal model gives an odds ratio of 0.5 for food-insecure adults, compared with food-secure adults of the same age. What share of food-insecure fifty-year-olds would the model put at very good or excellent? I'll try it myself, but take a few seconds with it first.

*(Pause)*

**Sarah:** Okay. The odds ratio is 0.5, so food-insecure adults are half as likely to be at very good or excellent. Half of forty percent is twenty percent.

**Kiffer:** Let's check that from the other end. The ordinal model can be read in either direction, so an odds ratio of 0.5 for being in a higher category is an odds ratio of two for being in a lower one. Apply your method below the split.

**Sarah:** Below the split, food-secure adults are at sixty percent. If I double that, food-insecure adults would be at a hundred and twenty percent. That's impossible, so my method is wrong. I treated the odds ratio as if it multiplied the probability.

**Kiffer:** Right. It multiplies the odds. So go through the odds.

**Sarah:** For food-secure adults, the odds are forty divided by sixty, which is two thirds. Half of that is one third. To turn odds back into a probability, I divide the odds by one plus the odds, so one third divided by one and a third, which is one quarter. So the model puts twenty-five percent of food-insecure fifty-year-olds at very good or excellent.

**Kiffer:** Now the check from the other end works. Below the split, the odds for food-secure adults are sixty divided by forty, which is 1.5. Doubled, that's three, and three divided by four is seventy-five percent. Twenty-five and seventy-five add up to one hundred, which is what we need.

**Sarah:** So my first answer was off by five points. That's close enough that I wouldn't have noticed without the check.

**Kiffer:** Halving the probability works reasonably well when the probability is small, because then the odds and the probability are close. It gets worse as the probability rises, and that matters here, because the same odds ratio applies at every split. Suppose ninety-two percent of food-secure fifty-year-olds rate their health as fair or better. What does the model give for food-insecure adults at that split?

**Sarah:** Halving would give forty-six percent, which would put more than half of them at poor. Going through the odds, ninety-two divided by eight is 11.5. Half of that is 5.75. Then 5.75 divided by 6.75 is about 0.85. So about eighty-five percent of food-insecure fifty-year-olds are at fair or better, and the shortcut would have been far off.

**Kiffer:** So at the split for very good or better, the gap is fifteen percentage points, forty percent compared with twenty-five, and at the split for fair or better, it's about seven, ninety-two compared with eighty-five. The odds ratio is the same at both splits, and the difference in probability is about twice as large at one as at the other. Measured as odds, the effect is constant. Measured as a probability, it depends on where you start.

**Sarah:** Which is a good reason to report predicted probabilities beside the odds ratio. In R, the predict function gives them for any combination of predictors.

**Kiffer:** It is. This behaviour comes from the link function. The logit link puts the model on the log-odds scale, where effects add, so each exponentiated coefficient multiplies the odds.

**Sarah:** That idea, that the link sets the scale for every coefficient, is what ties the lesson together, and it has an interesting origin. In 1972, John Nelder and Robert Wedderburn, two statisticians at Rothamsted, an agricultural research station in England, showed that linear, logistic and Poisson regression could all be treated as members of one family of models.

**Kiffer:** Which is why a single R function, glm, fits both logistic and Poisson regression, with the family setting the distribution.

**Sarah:** Question two. What does a relative risk ratio actually compare? Some perspective first. Statistics Canada reported that in 2021, about 4.7 million Canadians, or about fourteen percent, didn't have a regular health care provider. Where people go for care, and whether they have anywhere regular to go, is a nominal outcome. A family doctor, a walk-in clinic and an emergency department are different kinds of places, and none ranks above another.

**Kiffer:** And the multinomial model compares each category with a reference category. The trouble is the name, which makes it sound like a ratio of two probabilities. Here's a made-up example. A survey asks adults where they got their most recent flu shot, at a pharmacy, a family doctor's office or a public health clinic, with pharmacy as the reference. In the North region, fifty percent used a pharmacy, thirty percent a doctor's office and twenty percent a public health clinic. In the South, twenty percent used a pharmacy, fifteen percent a doctor's office and sixty-five percent a public health clinic. For the South compared with the North, what's the relative risk ratio for a doctor's office versus a pharmacy? Take a few seconds.

*(Pause)*

**Sarah:** In the North, doctor's office to pharmacy is thirty divided by fifty, which is 0.6. In the South, it's fifteen divided by twenty, which is 0.75. So the relative risk ratio is 0.75 divided by 0.6, which is 1.25.

**Kiffer:** Right. So is a person in the South more likely to get their shot at a doctor's office?

**Sarah:** No. The share fell from thirty percent to fifteen percent, so it halved, and the relative risk ratio is still above one. That's because pharmacy use fell even further, from fifty to twenty percent. The bottom of the ratio shrank faster than the top.

**Kiffer:** Exactly. The ratio of doctor's office to pharmacy went up while doctor's office use went down. A relative risk ratio is a ratio of two ratios, and the reference category is always part of it.

**Sarah:** So the sentence "people in the South were 1.25 times as likely to use a doctor's office" would be wrong.

**Kiffer:** It would. A correct sentence names both categories. The ratio of doctor's office use to pharmacy use was 1.25 times as high in the South as in the North. If the question is simply who uses doctor's offices more, the predicted probabilities answer it directly.

**Sarah:** Would a different reference category fix it?

**Kiffer:** It changes which comparisons you see. With doctor's office as the reference, the South's relative risk ratio is 0.8 for a pharmacy and 6.5 for a public health clinic. The fitted model and the predicted probabilities stay the same. Each ratio is correct for its own comparison, and none of them tells you on its own how the share in one category changed.

**Sarah:** The smoking result for emergency departments is a milder version of the same thing. The relative risk ratio is about three, while a thirty-year-old smoker's predicted probability of naming an emergency department is about double a non-smoker's.

**Kiffer:** Right, because smokers are also less often attached to a family doctor, which is the reference. Both numbers are correct, and they answer different questions.

**Sarah:** Question three. How can leaving out an offset turn a result upside down? This is the harder example.

**Kiffer:** Here's a made-up study. A clinic follows children with asthma and counts flare-ups that needed an unscheduled visit. It compares children who use a daily controller inhaler with children who don't. Enrolment ran over four years and the study ended on a single date. Children on the controller tended to join early, so they were followed for three years on average, and the others joined late and were followed for one year on average. There are a hundred children in each group. The controller group had a hundred and fifty flare-ups, and the other group had eighty.

**Sarah:** So the controller group averages 1.5 flare-ups per child, and the other group averages 0.8.

**Kiffer:** Here's the question. Which group has the higher rate of flare-ups? Take a few seconds.

*(Pause)*

**Sarah:** You have to divide by the time. The controller group is a hundred children followed for three years, which is three hundred child-years. A hundred and fifty divided by three hundred is 0.5 flare-ups per child-year. The other group has a hundred child-years, so eighty divided by a hundred is 0.8 per child-year. The controller group has the lower rate.

**Kiffer:** So the rate ratio for the controller group is 0.5 divided by 0.8, which is 0.625, a rate about thirty-seven and a half percent lower. Now, what would a Poisson model without the offset report?

**Sarah:** It would compare the average counts per child, 1.5 against 0.8, which gives 1.875. So it would say children on the controller have nearly twice as many flare-ups.

**Kiffer:** The same data give a ratio of about 1.9 without the offset and about 0.6 with it. The direction reverses because follow-up time is tied to the exposure. The offset adds the log of each child's follow-up time with its coefficient fixed at one, which turns the model into a model for flare-ups per child-year. With only that one predictor, the model's rate ratio is exactly the 0.625 we just worked out.

**Sarah:** I can see how someone would miss this. Every child has a count, the counts look like any other counts, and the model runs without complaint.

**Kiffer:** That's why I'd make it a habit to ask, before fitting any count model, whether everyone had the same time at risk, or the same population size. In the urgent care example, everyone reported over the same twelve months, so no offset was needed. Here the offset decides the answer. And when follow-up varies from person to person, a missing offset can make a Poisson model look overdispersed, because the people followed longest have larger counts that the model can't explain.

**Sarah:** Could you put follow-up time into the model as an ordinary predictor?

**Kiffer:** You can, and it's a reasonable check. The offset assumes that flare-ups build up in proportion to time, so twice the follow-up means twice the expected count. If you enter the log of follow-up time as a predictor and its coefficient comes out close to one, that assumption looks fine. The offset is the usual choice because it gives you a rate directly.

**Sarah:** Question four. Where does overdispersion come from? I think students learn the check, a dispersion ratio well above one, without a picture of what causes it.

**Kiffer:** Let's build one. Here's a made-up population. Half the people are frequent users of urgent care, who average 1.4 visits a year, and half are rare users, who average 0.2. Within each half, visits follow a Poisson distribution exactly. Across everyone, the average is 0.8 visits.

**Sarah:** The same average as the urgent care visits in the lesson.

**Kiffer:** On purpose. Now compare this population with a single Poisson group whose average is also 0.8. Which one has more people with zero visits? Take a few seconds.

*(Pause)*

**Sarah:** I'd say the mixed population, because the rare users mostly have zero visits.

**Kiffer:** That's right. For a Poisson count, the chance of a zero is e to the minus the mean. For a single group with a mean of 0.8, that's about 0.45, so forty-five percent zeros. Can you work out the mixed population?

**Sarah:** Rare users have a zero with probability e to the minus 0.2, which is about 0.82. Frequent users have e to the minus 1.4, which is about 0.25. Half the people are in each group, so I average the two, and that gives about 0.53. So fifty-three percent zeros, against forty-five percent for a single Poisson group with the same average.

**Kiffer:** And the variance tells the same story.

**Sarah:** Where does the extra variance come from?

**Kiffer:** It comes from two sources. Within each half, the variance equals that half's mean, and those average out to 0.8. Between the halves, each group's mean sits 0.6 away from the overall mean, and 0.6 squared is 0.36. So the total variance is 0.8 plus 0.36, which is 1.16. There's a variance above the mean and extra zeros, even though each half is perfectly Poisson.

**Sarah:** So overdispersion is what you see when people differ in ways the model hasn't captured.

**Kiffer:** That's the most common source. If the survey recorded something that separates the frequent users from the rare ones, say a chronic condition, putting it in the model would remove much of the overdispersion. When the difference can't be measured, the negative binomial model allows for it. Its extra parameter, theta, describes how much people's underlying rates vary, and a smaller theta means more variation.

**Sarah:** Does real health care use look like this?

**Kiffer:** It does. A study of emergency department visits in British Columbia, published in CMAJ Open in 2021, defined frequent users as people with three or more visits a year. They were about fourteen to fifteen percent of patients, and in 2015 to 2016 they made about forty percent of all visits. A population like that is a mixture of very different users, which is what produces overdispersion.

**Sarah:** And what does it do to the conclusions?

**Kiffer:** A rough guide is that when the dispersion ratio is about two, the Poisson standard errors are too small by a factor of about the square root of two, which is about 1.4. In the urgent care model, the dispersion ratio was 2.05, and the negative binomial standard error for smoking was about one and a half times the Poisson one. That's why the interval widened and the p-value rose from 0.0001 to 0.012, while the rate ratio stayed at 1.44.

**Sarah:** Okay. Now the part where we disagree.

**Kiffer:** Five-point ratings.

**Sarah:** Coding a five-point rating as the numbers one to five and fitting a linear regression is usually presented as a shortcut to avoid. I think that's often too strict. Geoff Norman, a methods researcher at McMaster University in Hamilton, reviewed this question in a 2010 paper. Drawing on studies going back to the 1930s, he concluded that parametric methods, the family that includes t-tests and linear regression, can be used with Likert-type data without concern for getting the wrong answer. And a mean score is easy for readers to follow.

**Kiffer:** For a single item with five categories, I'd still start with the ordinal model. The numbers one to five assume that every step is the same size, and the data give no way to check that. A 2018 paper by Torrin Liddell and John Kruschke showed that treating ordinal ratings as numbers can produce false alarms, missed effects, and even effects in the wrong direction.

**Sarah:** Showing that something can go wrong tells us it's possible. It doesn't tell us how often it happens in typical survey data. And the ordinal model has its own assumption, proportional odds, which also has to hold.

**Kiffer:** That's fair, and in many datasets the two approaches will point the same way. My concern is that you can't tell which situation you're in from the linear model alone, and Liddell and Kruschke make that point too. They also report that averaging several ordinal items doesn't remove the problem. Where I agree with you is that a mean is easier to read. The ordinal model can give some of that back through predicted probabilities, such as the share of each group in the top two categories.

**Sarah:** So where do we land?

**Kiffer:** For a single ordered rating, fit the ordinal model as the main analysis and show the percentage of each group in each category. If you also report a mean score because readers expect one, check that it points in the same direction as the ordinal model, and say that you checked.

**Sarah:** I'd lean on the mean more often than Kiffer would, but I agree with that rule. And if the two approaches disagree, the disagreement is worth reporting.

**Kiffer:** Okay, let's pull it together. Here are three things to take away.

**Sarah:** First, an exponentiated coefficient multiplies odds or rates, and turning it into a statement about probabilities takes another step. Go through the odds, or use predicted probabilities.

**Kiffer:** Second, every ratio has a comparison built into it. A relative risk ratio always involves the reference category, and a count model compares rates only when everyone had the same time at risk or an offset accounts for the difference.

**Sarah:** And third, a failed check tells you something about the data. A dispersion ratio well above one usually means people differ in ways the model hasn't captured, and the response is a better predictor or a model that allows for the extra spread.

**Kiffer:** If you'd like more practice, try the asthma flare-up problem from this episode again with different follow-up times, and work out how different the two groups' follow-up has to be before the direction of the result flips.

**Sarah:** Next time, it's Lesson five, Modelling Dependent Data, where we take up the assumption that every model in this lesson shares, that the observations are independent.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson five.
