# Lesson 1: A Structured Approach to Data Analysis

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson one, A Structured Approach to Data Analysis.

**Sarah:** It's meant for after you've finished the lesson. We'll add some perspective on how these habits play out in real studies and some critique of their limits, work through the questions students tend to find thorny with this material, and do some extra worked examples, including a few that are harder than the ones on the lesson page.

**Kiffer:** In a few places we'll ask you to work something out before we give the answer. When that happens, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. If you've measured a variable, shouldn't you adjust for it? If the causal diagram decides what goes in the model, what are the crude associations for? What does a missing-value code do to your numbers? And why check the shape of an outcome before any model is fitted? We'll finish with cut-points, where Kiffer and I see things a little differently.

**Kiffer:** Let's start with the first one.

**Sarah:** Question one. If I've measured a variable, shouldn't I adjust for it? It feels safer to put everything in the model.

**Kiffer:** It's one of the most common instincts in data analysis, and the causal diagram is there to check it. Adjusting for a confounder removes bias. If you want the total effect, adjusting for a mediator removes part of the effect you're trying to estimate. And adjusting for a collider adds bias that wasn't there before.

**Sarah:** The mediator part I find easy to accept. If smoking raises blood pressure and high blood pressure raises the risk of a heart attack, then holding blood pressure fixed hides part of what smoking does.

**Kiffer:** Right. The collider is the one people struggle with, because it can enter the analysis without anyone typing it into a model. Choosing who is in the study can do it. Here's a made-up example. A specialist clinic takes anyone referred with either of two conditions, which we'll call condition A and condition B. The town it serves has ten thousand adults. Twenty percent have condition A, ten percent have condition B, and in this made-up town the two are completely unrelated. Everyone with either condition gets referred.

**Sarah:** So in the town, ten percent of the people with condition A have condition B, and ten percent of the people without it do too.

**Kiffer:** Exactly. Now think only about the clinic's patients. Among the patients who don't have condition A, what share have condition B? Take a few seconds.

*(Pause)*

**Sarah:** All of them. If you don't have condition A, the only way into the clinic is to have condition B.

**Kiffer:** That's it. Can you count out the whole clinic?

**Sarah:** Two thousand people have condition A, and because the conditions are unrelated, ten percent of them, two hundred, also have B. Another eight hundred have B without A. So the clinic sees two thousand eight hundred patients. Among the two thousand with A, two hundred have B, which is ten percent. Among the eight hundred without A, every one has B. So a study run in that clinic would find that condition A seems to protect strongly against condition B, when in the town there's no association at all.

**Kiffer:** And every number in the clinic is correct. The bias comes entirely from who was let in. This is often called Berkson's bias, after Joseph Berkson, a statistician at the Mayo Clinic who described the problem for hospital data in 1946.

**Sarah:** Is there a recent example?

**Kiffer:** COVID-19 gave a well-known one. In 2020, a team at a large university hospital in Paris reported that only about four to six percent of their COVID-19 patients smoked daily, compared with about twenty-four percent of French adults. The researchers raised the possibility that nicotine might protect against the virus, and a trial of nicotine patches was planned.

**Sarah:** But those patients were selected. You had to be sick enough to be tested or admitted to be counted.

**Kiffer:** Right. Later that year, Griffith and colleagues showed in Nature Communications that selection of this kind could create a sizeable protective association between smoking and COVID-19 even if none existed, and that the people tested for COVID-19 in the UK Biobank cohort differed from the rest of the cohort on a wide range of traits. A 2022 study in the UK Biobank cohort, using both observational and genetic analyses, concluded that smoking raises the risk of severe COVID-19.

**Sarah:** So a study's recruitment rule belongs on the causal diagram, the same as any measured variable.

**Kiffer:** Yes. My advice is to ask what decided who is in the data, and to draw it.

**Sarah:** Question two. If the diagram chooses the adjustment set, what are the crude associations for? This one often feels like a contradiction. You run the unconditional associations before any model, and then you don't use them to decide what goes in it.

**Kiffer:** They have different jobs. The diagram decides which variables you adjust for. The crude associations check the data and the plan. An association in the wrong direction, or far stronger or weaker than anyone expects, can mean a variable was coded backwards or a missing-value code slipped through. A scatterplot shows whether a relationship is straight or curved. A cross-tabulation shows sparse cells that will make a model unstable. And the correlations among predictors warn you about collinearity.

**Sarah:** Here's the case students find hardest. Your diagram says age is a confounder, and in your data age shows no crude association with the outcome. Surely you can drop it?

**Kiffer:** Age stays in the model, and the reason is worth working through. Imagine the exposure is a medication that lowers the risk of the outcome, and older adults are more likely to take it. Age raises the risk directly, and through the medication it lowers it. Those two paths can cancel in a crude table, so age looks unrelated to the outcome.

**Sarah:** And yet age still affects both who takes the medication and who has the outcome.

**Kiffer:** Right, so leaving it out still biases the estimate for the medication. There's a second problem with letting the data choose. Whether a crude association reaches statistical significance depends heavily on the sample size. The amount of bias a confounder causes doesn't.

**Sarah:** And a third, I think. A mediator is associated with both the exposure and the outcome. So is a collider. If my rule is to adjust for anything associated with both, I'll pick up exactly the variables we said to leave out in question one.

**Kiffer:** That's the strongest reason of all. An association between two variables in the data fits several different arrangements of arrows, and only subject-matter knowledge can choose among them.

**Sarah:** Let me push back, though. The arrows are assumptions. Isn't the diagram just my opinion written down?

**Kiffer:** In a sense, yes. It's your assumptions written down, and that's what makes it useful. Once they're on paper, a colleague or a reviewer can dispute a specific arrow. When you genuinely aren't sure about an arrow, I'd draw both versions, run the analysis under each, and report both, with the reason recorded in the analysis log.

**Sarah:** Question three. What does a missing-value code do to your numbers?

**Kiffer:** Software treats it as a real number until it's told the value is a code. And large surveys are full of codes that look like ordinary values. The data dictionaries for Statistics Canada's Canadian Community Health Survey, for example, reserve six, seven, eight and nine in one-digit variables for not applicable, don't know, refusal and not stated, and 996 to 999 in three-digit variables.

**Sarah:** So a seven could be a real answer in one variable and a don't know in the next.

**Kiffer:** Exactly, which is why every code belongs in the codebook and in the cleaning script. Here's a made-up example. A survey of two hundred adults records smoking as one for yes, zero for no, and nine for not stated. Fifty people said yes, one hundred and forty said no, and ten didn't answer. What proportion of these adults smoke? Take a few seconds.

*(Pause)*

**Sarah:** I'll try it. The mean of a zero-one variable is the proportion of ones, so I can ask R for the mean. I add up all the values and divide by two hundred. The total is one hundred and forty, and one hundred and forty divided by two hundred is 0.70. So seventy percent smoke.

**Kiffer:** Does seventy percent make sense here?

**Sarah:** No. Only fifty of the two hundred said yes, so the answer should be somewhere around a quarter, and seventy percent is nearly three times that. Oh, I added the nines. The ten nines contributed ninety to the total.

**Kiffer:** Right. It's the same trap as the minus 999 in the lesson, hiding inside a variable that's supposed to hold only zeros and ones. So what's the fix?

**Sarah:** Convert the nines to missing first. Then it's fifty smokers out of the one hundred and ninety people who answered, which is about 0.263, so about twenty-six percent.

**Kiffer:** Good. Dividing by one hundred and ninety is itself a decision. It assumes the ten who didn't answer smoke at the same rate as everyone else. If that's wrong, what are the lowest and highest values the true proportion could take? Take a few seconds.

*(Pause)*

**Sarah:** If none of the ten smoke, it's fifty out of two hundred, which is twenty-five percent. If all ten smoke, it's sixty out of two hundred, which is thirty percent.

**Kiffer:** So the true value lies between twenty-five and thirty percent, and the complete-case estimate of about twenty-six percent sits inside that range. For a yes-or-no answer like this, the width of the range always equals the share of people who are missing, here five percentage points. If half the sample hadn't answered, the range would be fifty percentage points wide.

**Sarah:** So the number of records a complete-case decision removes belongs in the log, because it tells a reader how far the answer could move.

**Kiffer:** Yes. The same reasoning applies to the thirty-six participants lost to follow-up in the lesson's smoking cohort. If the people who dropped out were sicker than the rest, the complete-case risk is too low.

**Sarah:** Software can also change values without being asked. In 2016, Ziemann and colleagues reported in Genome Biology that about one in five papers with gene lists in Excel supplementary files contained gene names that the spreadsheet had converted into dates or numbers. A gene called MARCH1, for example, became the first of March. In 2020, the committee that names human genes renamed some of them, so MARCH1 is now MARCHF1.

**Kiffer:** It's a good argument for keeping the raw file untouched and reading it with a script, where the type of every column is reported and can be checked.

**Sarah:** Question four. Why check the shape of the outcome before any model is fitted, when the assumption is really about the residuals?

**Kiffer:** Because the raw outcome is an early warning, though an imperfect one. If an outcome is badly skewed and there are no strong predictors, the residuals will usually be skewed too, so the warning is worth heeding. But a strong predictor can make a raw outcome look odd when the residuals are fine. Picture a made-up sample of six-year-olds and twelve-year-olds, with height as the outcome. The histogram has two humps, one for each age, and once age group is in the model, the residuals can be bell-shaped.

**Sarah:** And counts work the same way?

**Kiffer:** They do, and this is the harder example for the episode. Here are made-up numbers. Two clinics each have one hundred patients. At clinic A, patients average one visit a year, and the variance of their visit counts is also about one. At clinic B, patients average five visits, and the variance is about five.

**Sarah:** So each clinic on its own looks just the way a Poisson model expects. The mean and the variance match.

**Kiffer:** Right. Now an analyst pools all two hundred patients and compares the mean with the variance. The pooled mean is three. Will the pooled variance be close to three? Take a few seconds.

*(Pause)*

**Sarah:** I'd guess it's bigger. The patients at clinic A are bunched around one and the patients at clinic B around five, so around the overall mean of three there's extra spread.

**Kiffer:** That's the idea. The pooled variance has two parts. The first is the average spread within each clinic, which is the average of one and five, so three. The second is how far each clinic's mean sits from the overall mean, squared and averaged, the same way a variance is built. Can you work that part out?

**Sarah:** Both clinics are two visits away from three, and two squared is four. So the pooled variance is three plus four, about seven. That's more than twice the mean, which looks like overdispersion, even though neither clinic has any.

**Kiffer:** Exactly. The extra variation comes from the clinics being different, which is the multilevel structure showing up in the count check. The Poisson assumption applies to the counts once the predictors are taken into account. Put clinic in the model, and in this example the problem goes away.

**Sarah:** And an analyst who looked only at the pooled counts might switch to a negative binomial model when the real issue is a missing variable.

**Kiffer:** That's one risk. The other is that patients in the same clinic resemble one another in ways no variable captures, which is the clustering problem. A model that treats them as independent gives standard errors that are too small.

**Sarah:** Can you give a feel for how much too small?

**Kiffer:** Here's an extreme, made-up case. One hundred people each have their blood pressure measured five times, and each person's five readings are nearly identical. The file has five hundred rows, but it holds about as much information as one hundred readings. Standard errors shrink with the square root of the sample size. So how far off are the standard errors from a model that treats the five hundred rows as five hundred people?

**Sarah:** It behaves as if the sample were five times bigger than the information supports, so its standard errors are too small by a factor of the square root of five, which is about 2.2. And real repeated measurements sit somewhere between that extreme and full independence.

**Kiffer:** Right. That's why counting measurements per person and people per centre comes before any model. Mixed models, later in the course, handle the clustering properly.

**Sarah:** Okay. Now the part where we disagree. Cut-points.

**Kiffer:** Turning a continuous measure into categories.

**Sarah:** I'm wary of them. When you split systolic pressure at 140, someone at 139 and someone at 141 land in different groups, while someone at 141 and someone at 185 land in the same one. Information is thrown away. In 2006, Royston and colleagues published a paper in Statistics in Medicine with the title Dichotomizing continuous predictors in multiple regression: a bad idea. They argued that it costs statistical power and can leave confounding behind.

**Kiffer:** That's well established, and I agree about the loss of power.

**Sarah:** And the cut-point itself is a value judgement. In November 2017, the American College of Cardiology and the American Heart Association lowered the threshold for high blood pressure from 140 over 90 to 130 over 80. The estimated share of American adults with high blood pressure went from thirty-two percent to forty-six percent, even though nobody's blood pressure had changed.

**Kiffer:** Canada went the other way. In December 2017, Hypertension Canada said it would not change its guidelines in response. Statistics Canada later applied both rules to the Canadian Health Measures Survey for 2012 to 2015. Counting people on blood-pressure medication as hypertensive, twenty-four percent of men aged twenty to seventy-nine had hypertension under the 140 over 90 rule, and forty percent under the 130 over 80 rule.

**Sarah:** So the same measurements give two different prevalences, depending on which committee you follow.

**Kiffer:** Yes. Where I'd defend categories is in what they're for. Clinical decisions are made at thresholds, and a health authority can plan services around the number of people above a treatment threshold. Interest holders can act on a prevalence more easily than on a regression slope. Categories also let you see the shape of a relationship early without assuming a straight line, which is what age bands do.

**Sarah:** I'd still keep the continuous measure in any model and use the categories to describe the results.

**Kiffer:** For the main model, I mostly agree. My position is that categories have a legitimate place in description and communication, provided the cut-point is chosen for a reason before anyone looks at the outcome. The worst version is searching the data for the cut-point that gives the smallest p-value, which Royston and colleagues warned leads to serious bias.

**Sarah:** So the shared rule might be this. Keep the continuous variable in the model where you can, and use categories for description and for checking shape. Choose any cut-point in advance from a guideline, write it in the codebook, and report what happens under a second cut-point.

**Kiffer:** I agree with all of that.

**Sarah:** Let's pull it together.

**Kiffer:** First, the adjustment set comes from the causal diagram, and the diagram should include whatever decided who is in the data. Crude associations are for checking the data and the plan.

**Sarah:** Second, software treats a missing-value code as a real number until the cleaning script converts it. Convert every code before computing anything, and record how many records a complete-case decision removes, because that number sets how far the answer could move.

**Kiffer:** And third, the raw outcome is an early warning. Mixing groups can give an outcome two humps or make counts look overdispersed when the real issue is a missing variable or clustering, so understand the structure of the data before choosing the model.

**Sarah:** If you'd like more practice, rework the two-clinic example from this episode with your own numbers. Try clinics that average two and six visits, and predict the pooled mean and variance before you work them out.

**Kiffer:** Next time, it's Lesson two, Data Cleaning and Descriptive Analyses, where we take on outliers, missing values and the descriptive table that opens most papers.

**Sarah:** Take care, everyone.

**Kiffer:** See you in Lesson two.
