# Lesson 6: Exploratory Data Analysis and Visualization

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson six, Exploratory Data Analysis and Visualization.

**Sarah:** It's for after you've worked through the lesson. We'll add some perspective and some critique, work through the questions that students tend to find thorny with this material, and do some extra worked examples, a couple of them harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. When is a point beyond the whiskers of a boxplot really an outlier? How can a few skipped questions remove almost a quarter of a sample? What can a correlation of minus 0.40 tell you, and where can it mislead? And is the age pattern in the loneliness data really about age? We'll finish with a question that Kiffer and I see a little differently, which is whether looking at the data can go too far.

**Kiffer:** Let's start with the boxplot.

**Sarah:** Question one. When is a point beyond the whiskers really an outlier? I think a lot of people read those separate dots as a list of problem cases.

**Kiffer:** That's a common reading, and it's worth taking apart. The whisker rule is a convention that John Tukey proposed as a quick screen. The whiskers reach the most extreme values within one and a half interquartile ranges of the box, and anything further out gets its own dot. The rule looks only at where a value sits relative to the box. It knows nothing about whether the value is possible.

**Sarah:** Let's try one with made-up numbers, then. Imagine a community survey of weekly hours of physical activity. The lower quartile is two hours, the median is four, and the upper quartile is seven. Above what value would a person be drawn as a separate dot? Take a few seconds.

*(Pause)*

**Kiffer:** The interquartile range is seven minus two, which is five hours. One and a half times five is 7.5. Add that to the upper quartile, and the cut-off is seven plus 7.5, or 14.5 hours. Anyone above 14.5 hours a week gets a dot.

**Sarah:** And at the bottom, two minus 7.5 is minus 5.5, which nobody can report. So in this example there can't be any dots below the box at all.

**Kiffer:** Right. Now think about who sits above 14.5 hours. Fifteen or twenty hours a week is about two to three hours a day. That's high, but it's possible for someone with a physically demanding job or a serious training routine. These made-up data are skewed to the right, with the median closer to the bottom of the box, so the rule gives the long upper tail its own dots.

**Sarah:** Whereas the thirteen people in the hours-worked question who reported more than one hundred and twelve hours a week are a different kind of problem.

**Kiffer:** Yes. Those values are implausible on their face, and you'd find them with a histogram and a check against the possible range, whatever the box looked like. Whether a value is beyond the whiskers and whether it's believable are two separate questions, and I find it helps to keep them apart.

**Sarah:** Does the rule ever flag values when the data are perfectly well behaved?

**Kiffer:** It does. A 2014 article on boxplots in the journal Nature Methods notes that for data with a normal, bell-shaped distribution, the whiskers cover about 99.3 percent of the values. So about seven values in every thousand fall beyond them by chance. If a variable in a file the size of the briefing file, with three thousand and eighty-three people, had a perfect bell shape, you'd expect a little over twenty dots with every value correct.

**Sarah:** So a dot is a prompt to look again at that value, and the decision about it comes from the codebook and the context.

**Kiffer:** That's how I'd use it.

**Sarah:** Question two. The 2021 wave of the Canadian Social Connection Survey had four thousand and forty-five respondents, and the briefing file has three thousand and eighty-three. How can a few skipped questions remove almost a quarter of the people?

**Kiffer:** Because na.omit drops anyone who is missing any one of the briefing variables, so the losses add up across questions. Here's a made-up version to think through. Suppose a file needs answers to nine questions, each question is skipped by three percent of respondents, and skipping one question has nothing to do with skipping another. What share of respondents would answer all nine? Take a few seconds.

*(Pause)*

**Sarah:** A quick guess would be nine times three percent, so about twenty-seven percent lost. To be exact, each person answers a given question with a probability of 0.97, so the chance of answering all nine is 0.97 multiplied by itself nine times. That's about 0.76. So about seventy-six percent are kept, and about twenty-four percent are lost.

**Kiffer:** Good. The quick guess overstates the loss a little, because it counts a person who skipped two questions twice. Twenty-four percent also happens to be about what the briefing file lost. We made up the three percent, though, so treat the match as a coincidence of the example.

**Sarah:** How would the real pattern differ?

**Kiffer:** In online surveys, one common source of missing answers is breakoff, where someone starts the questionnaire and stops before the end. That person misses every question after that point, so the gaps pile up in the same people. When the gaps fall on the same people, fewer people are lost overall for the same rate per question, and the people who are lost form a distinct group who may differ from those who stay.

**Sarah:** So what's the exploratory step?

**Kiffer:** I'd do two things. First, count the missing values for each variable, which takes one line of R. That tells you whether one question is doing most of the damage. Second, compare the people who were dropped with the people who were kept, on the questions they did answer, such as age and gender. If the dropped group is older, for example, the briefing describes a younger group than the survey as a whole, and the text beside the figure should say so.

**Sarah:** Question three is about the heatmap. Loneliness and social support have a correlation of minus 0.40. Suppose a member of the working group says that social support explains forty percent of loneliness. Why is that wrong?

**Kiffer:** Because a correlation measures how closely the points follow a straight line, on a scale from minus one to one. To get a share, you square it. In a simple straight-line regression of one variable on the other, the R squared is the correlation squared, and it's the share of the variance in one variable that the line accounts for.

**Sarah:** Then here's one for everyone. What share of the variance in loneliness scores does a straight line through social support account for? Take a few seconds.

*(Pause)*

**Kiffer:** Minus 0.40 squared is 0.16, so a straight line through support accounts for about sixteen percent of the variance in loneliness scores in this sample. The minus sign disappears when you square. The same calculation on Anscombe's quartet is worth doing. The correlation was about 0.82 in every set, and 0.82 squared is about 0.67, so the line accounts for about two thirds of the variance in every set, including set four, where a single point creates the whole relationship.

**Sarah:** So the squared value carries every weakness the correlation has.

**Kiffer:** It does, because it's the same number squared. It still measures fit to a straight line, one point can still create it, and it says nothing about cause. What squaring does give you is the right scale for talking about a share, so nobody reads minus 0.40 as forty percent of anything.

**Sarah:** The other trap is the one with age. The correlation between age and loneliness was 0.07, and the loess curve still showed a clear pattern. Is there a real-world case where that kind of thing matters?

**Kiffer:** Sleep and mortality is a good one. In 2010, Francesco Cappuccio and colleagues pooled sixteen prospective studies with nearly 1.4 million participants. Compared with people in the middle of the range, usually seven hours a night, short sleepers had a relative risk of death of 1.12, and long sleepers had a relative risk of 1.30. The risk rises at both ends.

**Sarah:** So a single correlation between hours of sleep and death would average the two ends together, and it could come out close to zero.

**Kiffer:** Right, and someone who stopped at that number might conclude that sleep doesn't matter. A scatterplot or a smoother would show the U shape straight away. There's a second point in that study, too. The authors cautioned that the link with long sleep might partly reflect existing illness and confounding that the studies could not fully remove. The plot shows the shape, and the study design decides what the shape can tell you about cause.

**Sarah:** Question four. Is the age pattern in the briefing really about age? The two older groups have a median loneliness score one point higher than the two younger groups. But when I look at the table of counts behind the faceted grid, women make up a much larger share of the older groups.

**Kiffer:** And women in the briefing file have a higher mean loneliness score than men, 5.74 against 5.36. So it's fair to ask whether the gender mix explains part of the age difference. Take the youngest group and the forty-five to sixty-four group. Their mean scores are 5.49 and 6.01, a gap of 0.52 points.

**Sarah:** And what does the table of counts say about those two groups?

**Kiffer:** The row for women has one thousand five hundred and eighty people in all, with four hundred and ninety in the youngest group and three hundred and sixty-two in the forty-five to sixty-four group. The two age groups have one thousand and seventy-nine and five hundred and seventy-eight people altogether. Using those counts and the gap between women and men, about how much of that 0.52 could the gender mix produce? Sarah will try it aloud in a moment. Take a few seconds.

*(Pause)*

**Sarah:** Okay. First I need the share of women in each group. In the forty-five to sixty-four group, there are three hundred and sixty-two women. There are one thousand five hundred and eighty women in the file, and three hundred and sixty-two divided by one thousand five hundred and eighty is about 0.23. So that group is about twenty-three percent women.

**Kiffer:** Check that against the column for that age group.

**Sarah:** The column has three hundred and sixty-two women and two hundred and four men, so women are the clear majority, and twenty-three percent can't be right. I divided by all the women in the file, which tells me what share of women are aged forty-five to sixty-four. I need to divide by everyone in the age group, which is five hundred and seventy-eight people. That gives about 0.63, so a little under sixty-three percent women.

**Kiffer:** That's the right denominator. And the youngest group?

**Sarah:** Four hundred and ninety women out of one thousand and seventy-nine people, which is about 0.45, so about forty-five percent. The older group has about seventeen percentage points more women.

**Kiffer:** Now bring in the gap between women and men.

**Sarah:** Women score 0.38 points higher than men on average, so each woman in place of a man adds about 0.38 to the expected score. With seventeen more women in every hundred people, the group mean goes up by about 0.17 times 0.38, which is about 0.065. So the gender mix could produce a gap of less than a tenth of a point. The actual gap is 0.52, so the gender mix accounts for about an eighth of it.

**Kiffer:** That's right. I left the non-binary respondents out to keep it simple, but they make up about two percent of both groups, so they barely change the answer. A full weighted average with all three genders gives 0.064. The general habit is to check, before you read a difference between groups as a feature of the grouping variable, whether the groups differ in their mix of something else that's related to the outcome.

**Sarah:** Which is what the grid with gender in the rows lets you see.

**Kiffer:** Yes, and a model with both age group and gender would do the same thing more formally. Our rough check also assumes that the gap between women and men is the same in every age group, and the facets let you look at that directly. And part of the overall gap of 0.38 may itself come from the larger share of women in the older groups, so the gender mix could account for a little less than an eighth.

**Sarah:** There's a second reason to be careful with that peak in middle age. The survey recruited people online. Has anyone looked at loneliness by age in a survey that selected people at random?

**Kiffer:** Statistics Canada did, in the Canadian Social Survey, which was collected in August and September 2021 from randomly selected households. Thirteen percent of people aged fifteen and older in the ten provinces said they always or often felt lonely. The figure was twenty-three percent at ages fifteen to twenty-four, about ten percent in each age group from thirty-five to sixty-four, nine percent at sixty-five to seventy-four, and fourteen percent at seventy-five and older. Women reported it more often than men, fifteen percent against eleven.

**Sarah:** So in both surveys the youngest people report the most loneliness, although in ours that rests on the few respondents under twenty, and the peak in the fifties doesn't appear in the Statistics Canada figures.

**Kiffer:** That's how I read it, with a caution. The two surveys asked different questions, a single question about how often people feel lonely and a three-item scale, so the numbers can't be compared directly, although the shapes of the two patterns can. One possible explanation is that the middle-aged people who joined an online survey about social connection were lonelier than middle-aged people in general. An exploratory plot describes the people in the file, and a comparison with a randomly sampled survey is a good way to judge which features might generalize.

**Sarah:** Okay. The last one is the question where Kiffer and I differ. Can looking at the data go too far?

**Kiffer:** Go ahead and make the case.

**Sarah:** Almost everything we've praised today involves looking at the data and then changing the analysis. We saw the age curve, so now we'd enter age as groups or as a curve. Andrew Gelman and Eric Loken called this the garden of forking paths, in a 2013 paper. Each choice made after seeing the data opens a different path, and the p-value at the end assumes you'd have taken the same path whatever the data showed.

**Kiffer:** And you think that makes exploration risky.

**Sarah:** I think it makes the final p-values too optimistic. Joseph Simmons and two colleagues showed it with simulations in 2011. They combined four common choices, such as measuring two outcomes and reporting whichever worked, or adding more participants when the first test missed significance. Each choice on its own raised the false-positive rate from five percent to somewhere between about eight and thirteen percent. All four together raised it to about sixty-one percent.

**Kiffer:** That's a real problem, and I agree with the general point. Where I'd push back is on what those choices were. They were choices about which outcome to report, when to stop collecting data, which variable to adjust for and which groups to compare. Checking that a household can't hold one hundred and eighty people, or seeing that loneliness has a ceiling, doesn't open a path of that kind, if you decide how to handle those values before looking at the main result. Skipping the look also has a cost. You'd fit a straight line to a curve or keep impossible values, and the p-value from that model would be wrong in its own way.

**Sarah:** But the age curve is a different case. That's the outcome plotted against an exposure, and the shape of the curve then decides the model.

**Kiffer:** Agreed, and that's where I'd draw the line. Looking at each variable on its own, for quality and shape, carries much less of that risk, and it's necessary. Looking at the main relationship you plan to test is where the paths start to fork, and that's where I'd want the plan written down first.

**Sarah:** I'd still go further than Kiffer. I'd treat any pattern found by exploring as a hypothesis until it's been seen in new data.

**Kiffer:** I'd accept that for a surprising pattern. An obvious data problem can be fixed and reported without new data. But we agree on the practical rule.

**Sarah:** Here it is. Write down the main question and the planned model before you plot the main relationship. Explore each variable freely to check its quality and shape. If a plot changes the model, report the change, show both results, and describe any pattern you found by exploring as a question for the next study.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, the dots beyond a boxplot's whiskers and the spikes in a histogram are prompts to check. The decision about a value comes from its possible range, the codebook and the context.

**Kiffer:** Second, a single summary number can hide how it was produced. Square a correlation before you talk about variance, plot the pair before you trust the correlation, and count who was dropped before you trust a complete-case file.

**Sarah:** And third, before reading a group difference as a finding, check the mix of people in each group and how they came to be in the survey, and label the patterns you found by exploring as questions for further work.

**Kiffer:** If you'd like more practice, try the gender-mix check from Question four with the thirty to forty-four group and the sixty-five and older group. You'll get a different answer from ours, and it's worth thinking about why.

**Sarah:** Next time, it's Lesson seven, Measurement and Psychometrics, where we ask whether scale scores like the loneliness score measure what they're meant to measure, and how reliably they do it.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson seven.
