# Lesson 5: Modelling Dependent Data

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson five, Modelling Dependent Data.

**Sarah:** It's best heard once you've finished the lesson. We'll bring in some perspective on where these methods come from and how they're used, offer some critique, work through the questions students tend to find thorny, and add a few worked examples, including some that are harder than the ones on the lesson page.

**Kiffer:** There are a few places where we'll ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. If the ICC is tiny, can you ignore the clustering? If a clustered study is short of information, should you add more people to each cluster or add more clusters? Why do a mixed model and GEE give different odds ratios from the same data, and which one should you report? Kiffer and I don't fully agree on that one. And finally, does a mixed model really take care of missed visits?

**Kiffer:** Let's start with the first one.

**Sarah:** Question one. If the ICC is tiny, can you ignore the clustering? When I see an ICC of 0.02, my instinct is to call it zero, treat the people as independent, and fit an ordinary regression.

**Kiffer:** It's a natural instinct, and real ICCs often are small. An analysis of the Health Survey for England 1994, published in 1999, estimated ICCs for lifestyle risk factors and health outcomes at several levels. Most ICCs for large health districts were below 0.01, and most for postal code areas were below 0.05. The same authors also found that the design effects could be substantial when the clusters were large.

**Sarah:** Because the design effect depends on the cluster size as well as the ICC.

**Kiffer:** Right, so let's try one. Imagine a hypothetical provincial survey of student wellbeing, with sixty schools and one hundred students surveyed in each school. The made-up ICC for the outcome is 0.02. What's the design effect? Take a few seconds.

*(Pause)*

**Sarah:** The design effect is one plus the cluster size minus one, times the ICC. So that's one plus ninety-nine times 0.02. Ninety-nine times 0.02 is 1.98, so the design effect is 2.98, or about three.

**Kiffer:** That's right. An ICC that sounds like nothing has roughly tripled the variance of an estimate such as the overall average. Now, what's the effective sample size?

**Sarah:** Sixty schools times one hundred students is six thousand students. Six thousand times 2.98 is seventeen thousand eight hundred and eighty. Okay, wait. That says six thousand students carry nearly eighteen thousand students' worth of information. That can't be right. Students in the same school partly repeat each other's information, so the answer has to be smaller than six thousand.

**Kiffer:** That sanity check is worth doing every time. It's an easy slip, because when you plan a study, you multiply the sample size by the design effect to find how many people to recruit. When the ICC is above zero, the effective sample size is always smaller than the number of people you actually measured.

**Sarah:** So I divide. Six thousand divided by 2.98 is about two thousand and thirteen. The six thousand students carry about as much information as two thousand independent students.

**Kiffer:** That holds for some questions. The figure applies to estimating an overall average and to questions about the schools themselves, such as whether schools with a daily physical activity policy have more active students. For a person-level predictor, like a student's age, which varies inside every school, the loss is much smaller, and allowing for the schools can even make the estimate more precise.

**Sarah:** So a small ICC can still matter a lot when the clusters are big. Is there a case where it does very little harm?

**Kiffer:** It does little harm when the clusters are small. With clusters of three people and an ICC of 0.02, the design effect is one plus two times 0.02, which is 1.04. Even then, I'd allow for the clustering, because the study design tells you it's there, and a mixed model or GEE costs very little when the clustering turns out to be weak.

**Sarah:** Here's a related point that trips people up. In the clinic data, the ICC went up from 0.175 to 0.24 once age, gender, smoking and urban location were in the model. So the ICC is a property of the model as well as the data.

**Kiffer:** Right. Predictors that vary among patients shrink the within-clinic variance, so the clinic share of what's left goes up. In that model the between-clinic variance barely moved, but a clinic-level predictor that explained why some clinics run higher than others would shrink it, and the ICC would fall. Whenever you report an ICC, say which model it came from.

**Sarah:** Question two. If a clustered study doesn't have enough information, is it better to add more people to each cluster, or to add more clusters?

**Kiffer:** Let's work it through, because this one takes a few more steps than the examples in the lesson. A health authority wants to compare six clinics that have adopted a team-based care model with six clinics that haven't. These are made-up numbers. There are two hundred patients in each clinic, and the ICC for the outcome is 0.03. Sarah, what's the effective sample size?

**Sarah:** Twelve clinics times two hundred is two thousand four hundred patients. The design effect is one plus one hundred and ninety-nine times 0.03, which is one plus 5.97, so 6.97. Two thousand four hundred divided by 6.97 is about three hundred and forty-four.

**Kiffer:** Good. Now suppose the health authority can afford to survey twice as many patients. If we double the number of patients in each of the same twelve clinics, to four hundred, roughly how much does the effective sample size go up? Take a few seconds.

*(Pause)*

**Sarah:** The design effect becomes one plus three hundred and ninety-nine times 0.03, which is 12.97. Four thousand eight hundred patients divided by 12.97 is about three hundred and seventy. So doubling the patients adds only about twenty-six effective people.

**Kiffer:** And if we kept two hundred patients per clinic and doubled the number of clinics to twenty-four?

**Sarah:** Then the design effect stays at 6.97, and four thousand eight hundred divided by 6.97 is about six hundred and eighty-nine. That's twice the effective sample size, for the same four thousand eight hundred patients.

**Kiffer:** There's a ceiling hiding in the formula. When the clusters get very large, the design effect is close to the cluster size times the ICC. What does that do to the effective sample size?

**Sarah:** The number of patients is the number of clusters times the cluster size, and the design effect is roughly the cluster size times the ICC, so the cluster size cancels out. The effective sample size gets close to the number of clusters divided by the ICC. With twelve clinics, that's twelve divided by 0.03, which is four hundred.

**Kiffer:** And it can never go above that, however many patients you add to those twelve clinics.

**Sarah:** So adding people to the same clinics helps less and less, while adding clinics keeps helping.

**Kiffer:** And twelve clinics cause a second problem that the effective sample size hides. Team-based care is a cluster-level predictor, so the comparison rests on twelve values, six against six. That's below the twenty to thirty clusters a mixed model usually needs, and well below the thirty to forty that the sandwich standard errors of GEE need.

**Sarah:** So three hundred and forty-four effective people sounds comfortable, but the twelve clinics are the real limit.

**Kiffer:** Exactly. Corrections exist for GEE when there are few clusters, and a 2023 review of cluster randomized trials in the International Journal of Epidemiology recommends using one when there are fewer than about forty clusters. The same review notes that none of these corrections performs well in every setting. In my view, the more dependable fix is at the design stage, with more clusters.

**Sarah:** That connects to cluster randomized trials, where whole groups are randomized together.

**Kiffer:** Yes, and there's a famous warning about them. In 1978, the statistician Jerome Cornfield wrote that randomizing by group and then analysing the data as if individuals had been randomized is "an exercise in self-deception." Allan Donner, of Western University in London, Ontario, co-wrote a textbook on the design and analysis of cluster randomization trials with Neil Klar, published in 2000.

**Sarah:** Is there a Canadian trial that shows what this looks like?

**Kiffer:** A good one was led by a McMaster University researcher and published in the Journal of the American Medical Association in 2010. Forty-nine Hutterite colonies in Alberta, Saskatchewan and Manitoba were randomized. Children aged three to fifteen were offered either a flu vaccine or, as the control, a hepatitis A vaccine, depending on their colony's arm. The question was whether vaccinating children protected the other people in the colony. Among residents who weren't vaccinated, the protective effectiveness was sixty-one percent, which means confirmed flu was about sixty percent less common among them than among unvaccinated residents of the control colonies.

**Sarah:** And the vaccine arm is a cluster-level predictor, because the whole colony is in the same arm.

**Kiffer:** Right, so the information about the vaccine comes from forty-nine colonies. The published ninety-five percent confidence interval runs from eight to eighty-three percent. If you take the published counts and treat every unvaccinated resident as independent, a quick calculation gives an interval of roughly forty-one to seventy-two percent. The published interval is much wider, which is what we'd expect when the information comes from colonies.

**Sarah:** And the trial was designed to find out whether vaccinating children cut the spread of flu to other people in the colony. So in this study, the dependence between people was the very thing the researchers set out to measure.

**Kiffer:** That's a good reminder that clustering is sometimes the subject of the research as well as a problem for the analysis.

**Sarah:** Question three. In the lesson, the GLMM gave an odds ratio of 2.31 for urban clinics and GEE gave 2.04. Students often assume one of them must be wrong.

**Kiffer:** Let's build a small made-up example where we know the truth exactly. There are two kinds of clinics, equal in number and size, and half the patients in every clinic smoke. In the low-referral clinics, twenty percent of non-smokers and fifty percent of smokers are referred. In the high-referral clinics, fifty percent of non-smokers and eighty percent of smokers are referred.

**Sarah:** So in the low clinics, the odds for non-smokers are twenty over eighty, which is 0.25, and for smokers fifty over fifty, which is one. That's an odds ratio of four. In the high clinics, the odds are one for non-smokers and eighty over twenty, which is four, for smokers. That's also an odds ratio of four.

**Kiffer:** So within every clinic, smoking multiplies the odds of referral by four. Here's the question. If we pool all the patients and compare all smokers with all non-smokers, will the odds ratio be four, more than four, or less than four? Take a few seconds.

*(Pause)*

**Sarah:** My first guess was four, because nothing is confounded. Smoking is equally common in both kinds of clinic.

**Kiffer:** That's a common first guess. Let's check it.

**Sarah:** Across all clinics, thirty-five percent of non-smokers are referred, halfway between twenty and fifty, and sixty-five percent of smokers, halfway between fifty and eighty. The odds for smokers are sixty-five over thirty-five, about 1.86, and for non-smokers thirty-five over sixty-five, about 0.54. So the odds ratio is about 3.4.

**Kiffer:** So the within-clinic odds ratio is four, and the population-averaged odds ratio is about 3.4, from the same patients, with no confounding and no sampling error. The first is the kind of odds ratio a GLMM estimates, and the second is the kind GEE estimates. Averaging over clinics with different baselines pulls the odds ratio toward one.

**Sarah:** What if I'd compared percentages?

**Kiffer:** Then nothing changes. The difference is thirty percentage points in each kind of clinic, and thirty points overall. It's the binary version of the point that with an identity link the two kinds of estimate agree. The odds ratio changes because the population average is taken over probabilities, and an odds ratio doesn't survive that averaging unchanged.

**Sarah:** Which brings us to where we disagree. If I'm advising a provincial health ministry, I think GEE should usually be the main analysis.

**Kiffer:** Make the case.

**Sarah:** The ministry wants to know what happens across the population, and that's what a population-averaged odds ratio describes. GEE also makes no assumption about the shape of the distribution of clinic effects, and with thirty clinics I can't check that assumption well. Hubbard and colleagues made this argument in the journal Epidemiology in 2010. They wrote that mixed models involve assumptions that can't be verified, and that population-averaged models give a more useful approximation of the truth.

**Kiffer:** That paper drew a reply in the same issue from Subramanian and O'Malley, and I find myself closer to them. They argued that comparing the two approaches in general is futile, because they answer different questions, so you settle the question first. They also pointed out that variation between neighbourhoods is often something we want to understand. In the clinic data, the variance in referral between clinics is itself a finding, which a quality improvement team would want to know about.

**Sarah:** A quality improvement team has a different question from the ministry, though.

**Kiffer:** Agreed, and for the ministry's question I'd accept the population-averaged answer. My hesitation is practical. GEE needs more clusters, and a public health study can easily end up with fewer than thirty. In repeated measures, which we'll come to next, a mixed model is valid under a weaker condition about missed visits. So I'd choose the mixed model as the main analysis more often than Sarah would.

**Sarah:** I'm still more inclined toward GEE when there are enough clusters and the question is about the population. We agree on the practical rule, though.

**Kiffer:** Write down the question before you choose the model. Say whether your odds ratio is cluster-specific or population-averaged, report how many clusters you had, and give the other method as a sensitivity analysis.

**Sarah:** Question four. Does a mixed model take care of missed visits? Students often come away thinking that a mixed model simply solves the problem of missing data.

**Kiffer:** It handles one kind of missingness well, and a made-up example shows why. Picture one arm of a trial with one hundred people. At the six-month visit, forty of them had high blood pressure and sixty had lower blood pressure. At twelve months, the high group would average one hundred and forty millimetres of mercury and the lower group one hundred and twenty, if everyone came. All sixty in the lower group come to the twelve-month visit, but half of the high group miss it, and the ones who miss are otherwise just like the ones who come.

**Sarah:** So we only see eighty people at twelve months.

**Kiffer:** Right. Here's the question. Is the average blood pressure among the people who came too high or too low, and by how much? Take a few seconds.

*(Pause)*

**Sarah:** It'll be too low, because the people who stayed away were all from the high group. Among the people who came, there are twenty at one hundred and forty and sixty at one hundred and twenty. That's two thousand eight hundred plus seven thousand two hundred, which is ten thousand, divided by eighty, so one hundred and twenty-five. If everyone had come, it would be five thousand six hundred plus seven thousand two hundred, divided by one hundred, which is one hundred and twenty-eight.

**Kiffer:** So the simple average is three millimetres of mercury too low. That's about the size of the program effect in the lesson's wellness trial, which shows how easily missed visits could create or hide a result.

**Sarah:** And a mixed model would get it right?

**Kiffer:** In this example, yes, provided the model is otherwise correct. The chance of missing depends only on the six-month reading, which was recorded and which the model uses. That's what missing at random means. Because the model uses each person's earlier readings, it knows that the people who stayed away at twelve months had high readings at six months, and it allows for that when it estimates the twelve-month average.

**Sarah:** And it would fail if people stayed away because their blood pressure was high on the day of the twelve-month visit, in a way their earlier readings couldn't predict.

**Kiffer:** Right. That's missing not at random, and a mixed model fitted to the observed visits can't correct for it. The data can't tell you which situation you're in, so the usual steps are to include the measured characteristics that predict missed visits and to state the assumption clearly.

**Sarah:** What about GEE in the same example?

**Kiffer:** GEE needs a stronger condition. Its results are valid when missing a visit depends only on the predictors in the GEE model, which in a trial are usually arm and visit. Here, missing depends on the earlier blood pressure, which is outside that model, so a standard GEE could be biased. That's why a mixed model is usually preferred for repeated measures when missed visits are likely to depend on earlier results.

**Sarah:** So the practical step is to count missed visits by visit and by arm before fitting anything, and to compare the people who missed visits with those who didn't at their first visit. If missed visits are common or differ between arms, that belongs in the results, because the conclusion depends on the missing at random assumption.

**Kiffer:** I agree. Let's pull it together with three things to take away.

**Sarah:** First, judge clustering by the design effect, which depends on both the ICC and the cluster size. A small ICC in large clusters can still remove much of the information about cluster-level questions.

**Kiffer:** Second, the number of clusters sets the limit for cluster-level predictors. Adding people to the same clusters helps less and less, and both mixed models and GEE need enough clusters to work well.

**Sarah:** And third, know what kind of estimate you're reporting. For a binary outcome, a cluster-specific odds ratio from a GLMM is usually further from one than a population-averaged odds ratio from GEE, and the two methods make different assumptions about missed visits.

**Kiffer:** If you'd like more practice, rework the problem from this episode about adding patients or adding clinics, this time with an ICC of 0.1, and find the ceiling on the effective sample size for twelve clinics.

**Sarah:** Next time, it's Lesson six, Exploratory Data Analysis and Visualization, where we look at data before modelling them and turn what we see into figures that other people can read.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson six.
