Summary
The episode works through four questions that students commonly find difficult when modelling dependent data. It shows that a small intracluster correlation coefficient can still produce a large design effect when clusters are large, and that the number of clusters sets a ceiling on the information available for cluster-level predictors. A Canadian cluster randomized trial of influenza vaccination in Hutterite colonies illustrates how allowing for clustering widens the confidence interval for an intervention assigned to whole communities. A constructed example explains why cluster-specific odds ratios from a mixed model differ from population-averaged odds ratios from generalized estimating equations, and the hosts discuss a published debate about which approach should be the main analysis. The episode closes by showing how missed visits can bias a simple average and why mixed models and generalized estimating equations rest on different assumptions about missing data.