# Lesson 3: Questionnaire Design

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson three, Questionnaire Design.

**Sarah:** It's meant for after you've finished the lesson. We'll add some perspective on how these ideas play out in real surveys and some critique of the rules themselves, we'll work through the questions students tend to find thorny, and we'll do some extra worked examples, including one that's harder than anything on the lesson page.

**Kiffer:** In a few places we'll ask you to work something out before we give the answer. When that happens, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. Does a low response rate mean a survey is biased? Why does random error in a question make associations look weaker, and what does adding more items do about it? How can the response options change the answers people give? And we'll finish with Likert scales, because Kiffer and I don't fully agree about whether you can take their mean.

**Kiffer:** Let's start with response rates.

**Sarah:** Question one. I think most students carry a mental scale where seventy percent is good, fifty percent is worrying and twenty percent is a disaster. Is that the right way to judge a survey?

**Kiffer:** It's a fair rough guide, but it leaves out the part that matters most. Non-response bias depends on two things: how many people didn't respond, and how different they are from the people who did. A low response rate tells you there's room for bias. It doesn't tell you how much bias there is.

**Sarah:** And there's a simple way to write that down.

**Kiffer:** There is. For a proportion, the bias equals the share of people who didn't respond, times the gap between responders and non-responders. If everyone responds, the first part is zero, so the bias is zero. If non-responders are just like responders, the second part is zero, and the bias is zero again, whatever the response rate.

**Sarah:** Okay, here's one for everyone listening, with made-up numbers. Two surveys estimate smoking in the same province. In Survey A, the response rate is thirty percent, and the share of smokers among responders is two percentage points lower than among non-responders. In Survey B, the response rate is seventy percent, and the gap is fifteen points in the same direction. Which survey's estimate is further from the truth? Take a few seconds.

*(Pause)*

**Kiffer:** Let's take Survey A first. Seventy percent didn't respond, so the bias is 0.7 times two points, which is 1.4 percentage points. In Survey B, thirty percent didn't respond, so the bias is 0.3 times fifteen points, which is 4.5 points. The survey with the much better response rate is more than three times as far off.

**Sarah:** And both are off in the same direction. Responders smoke less in both surveys, so both estimates come out too low.

**Kiffer:** Right. The published evidence points the same way. In 2008, Robert Groves, a survey researcher who later directed the United States Census Bureau, pooled fifty-nine studies with a colleague. Each study had been designed to measure non-response bias. They found very little correlation between a survey's non-response rate and the size of its bias.

**Sarah:** I'd still say response rates matter, though, because a low rate leaves so much room. Canada has a well-known example. For 2011, the mandatory long-form census was replaced by the voluntary National Household Survey. The mandatory long form had drawn responses from well over ninety percent of households in 2006. The voluntary survey got 68.6 percent.

**Kiffer:** And the concern was exactly the one in our formula. Statistics Canada noted that the survey's voluntary nature carried a risk of non-response bias. So they used follow-up of non-responders, which is the most consistently effective strategy we have. In mid-July 2011, they chose a subsample of 400,000 of the 1.2 million dwellings that hadn't yet responded and concentrated follow-up on them.

**Sarah:** Why only a subsample?

**Kiffer:** Because intensive follow-up is expensive, and a random subsample of non-responders can tell you about the whole non-responding group. The households reached in that subsample are given extra weight to stand in for the ones that weren't followed up. Then the long form was made mandatory again for the 2016 census, and its response rate was 97.8 percent.

**Sarah:** So what should a reader do with a response rate in a published paper?

**Kiffer:** My own view is that the response rate should always be reported, alongside whatever evidence the authors have about who didn't respond. Did they compare responders with the sampling frame on age, sex or region? Did the late responders, the ones who needed reminders, look different from the early ones? A fifty percent response rate with that kind of checking can be more trustworthy than an eighty percent rate with none.

**Sarah:** Question two. Why does random error in a question make an association look weaker?

**Kiffer:** Students often expect random error to be harmless on average. Some people answer a little high, some a little low, and it cancels out. For the average of the variable, that's roughly true. For an association between two variables, it fails.

**Sarah:** Because the noise spreads out the exposure without moving the outcome along with it.

**Kiffer:** Exactly. Picture the scatterplot. The noise slides each point left or right at random, the cloud stretches along the exposure axis, and the line through it gets flatter. That's the attenuation you can watch in the reliability simulator. For a straight-line regression on one exposure, with purely random error in that exposure, the observed slope is expected to equal the true slope times the reliability.

**Sarah:** Can I try a harder one?

**Kiffer:** Go ahead. The numbers are made up. A study asks one question, how lonely have you felt in the past month, on a scale from zero to ten. A test-retest study puts its reliability at 0.4. Suppose that in truth, each one-point increase in loneliness goes with two more points on a depression symptom score. What slope will the study report?

**Sarah:** The true slope times the reliability, so two times 0.4, which is 0.8. The study would report less than half the real association.

**Kiffer:** Now the harder part. The researchers replace the single question with the average of four similar loneliness items, each with a reliability of 0.4. What slope should they expect now? Even without a formula, you can work out the range the answer has to fall in. Take a few seconds.

*(Pause)*

**Sarah:** Okay. Four items, each with reliability 0.4, so the reliability of the average is four times 0.4, which is 1.6. Then the slope is two times 1.6, which is 3.2. Hmm. That's bigger than the true slope.

**Kiffer:** Can that happen?

**Sarah:** No. Reliability is the share of the variation that comes from the true score, so it can't go above one. And random noise should only shrink the slope, so the answer has to land between 0.8 and two. I can't just add the reliabilities together.

**Kiffer:** Right. Averaging helps because the random errors in different items partly cancel, while the true signal they share stays put. The Spearman-Brown formula, which drives the simulator's multi-item mode, keeps track of that. If r is the reliability of one item, the reliability of the average of k items is k times r, divided by one plus r times the quantity k minus one.

**Sarah:** So that's four times 0.4, which is 1.6, divided by one plus three times 0.4, which is 2.2. 1.6 divided by 2.2 is about 0.727. And the slope is two times that, about 1.45.

**Kiffer:** That's it. The slope goes from 0.8 to about 1.45, much closer to the true value of two. With six items, it's six times 0.4, which is 2.4, divided by one plus five times 0.4, which is three. That gives exactly 0.8, and a slope of 1.6. Notice the diminishing returns. Going from one item to four added about 0.33 to the reliability, and the next two items added less than 0.08.

**Sarah:** What does the rule depend on?

**Kiffer:** It depends on three things. The errors are random, unrelated to the true score, to the outcome and to each other. The items measure the same thing equally well. And the error is in the exposure. Random error in the outcome of a linear regression mostly costs you precision, and the slope stays centred on the true value. Error that differs between groups, like recall bias between cases and controls, can push an estimate in either direction.

**Sarah:** And then there's error that leans one way for everybody. Statistics Canada has a good example. In the 2005 Canadian Community Health Survey, about forty-six hundred respondents reported their height and weight in an interview and were then measured. On average, women under-reported their weight by 2.5 kilograms and men by 1.8, and both slightly over-reported their height. Obesity came out at 15.2 percent from self-report and 22.6 percent from measurement.

**Kiffer:** And adding more self-report items wouldn't fix that, because the error has a direction. Averaging cancels random error and leaves systematic error where it was. Finding systematic error is the job of validation against a gold standard, which in that study meant measured height and weight.

**Sarah:** Question three. How can the response options change the answers people give? I think students treat the options as a neutral container for an answer the respondent already has.

**Kiffer:** The cognitive model from the lesson predicts something different. Respondents use everything on the page to work out what you mean, and that includes the options. A classic example comes from Howard Schuman and Stanley Presser in 1981. People were asked what they considered the most important thing to prepare children for life. When the option, to think for themselves, was offered on a list, 61.5 percent chose it. When the question was open, only 4.6 percent gave an answer that fit that category.

**Sarah:** That's more than a thirteen-fold difference. Which one is right?

**Kiffer:** Each one captures part of the picture. The list reminds people of an idea they may hold but wouldn't have put into words, and it also signals what the researcher considers a sensible answer. The open question avoids that signal, but people tend to leave out things that seem too obvious to mention.

**Sarah:** The same thing happens with frequency scales, and this one's worth trying. In a study by Norbert Schwarz and a colleague, patients were asked how often they experienced various physical symptoms. One group got a scale running from twice a month or less up to several times a day. The other group got a scale running from never up to more than twice a month. In which group did more patients report symptoms more than twice a month? Take a few seconds.

*(Pause)*

**Kiffer:** It was the first group. With the scale that ran up to several times a day, sixty-two percent reported symptoms more than twice a month. With the scale that stopped at more than twice a month, thirty-nine percent did. The question was the same in both groups, and only the scale changed.

**Sarah:** So people read the scale as a hint about what's normal. If the middle of the scale sits at a fairly high frequency, someone who isn't sure assumes they're about average and answers near the middle.

**Kiffer:** That's the usual explanation. The scale can also change what counts as a symptom. A scale that runs up to several times a day suggests you mean minor complaints, and one that tops out at twice a month suggests you mean serious ones. The effect is largest when the answer is hard to recall or the symptom is vaguely defined, because that's when respondents lean most on the options. The primacy and recency effects from the lesson are another case of the format shaping the answer.

**Sarah:** What can a designer do with that?

**Kiffer:** A few things. For counts, the fill-in-the-blank number is already the preferred format, and it removes the scale's hint. For opinions, think-aloud pre-testing shows you what respondents infer from the options. And when you borrow an item from a survey like the Canadian Community Health Survey so you can compare your sample with national figures, keep its wording, its response options and, as far as you can, its place in the questionnaire. If you change the options, your figure and the national one are no longer measuring the same thing.

**Sarah:** Okay, now for our disagreement. Can you take the mean of Likert responses? Before we argue, there's a coding trap we both agree on.

**Kiffer:** Go ahead and pose it.

**Sarah:** A made-up satisfaction item runs from one, strongly disagree, to five, strongly agree, with a separate don't-know option. A hundred people answer. Eighty give answers on the scale, and those average exactly three. Twenty say don't know, and the online survey exports those as a six. Nobody declares six as missing. What mean do you get if you average all hundred answers? Take a few seconds.

*(Pause)*

**Kiffer:** Eighty times three is two hundred and forty. Twenty times six is one hundred and twenty. That adds to three hundred and sixty, and divided by a hundred it gives 3.6.

**Sarah:** And that's the nasty part. It's a believable number. It sits inside the one-to-five range, so a glance at the mean won't flag anything. You catch it by running a frequency table for the item, or checking its highest value, and noticing a six on a five-point scale.

**Kiffer:** Which is why reserved codes go in the codebook and get declared missing before any analysis. Now the argument.

**Sarah:** My position is that a single Likert item is ordinal. We know agree is above neutral, but nothing tells us that the distance from neutral to agree equals the distance from agree to strongly agree. Susan Jamieson made this case in the journal Medical Education in 2004, and she quoted a good line from earlier critics: the average of fair and good is not fair-and-a-half.

**Kiffer:** It's a good line, and for a single item I find it hard to argue with.

**Sarah:** And here's a made-up example of what a mean can hide. Clinic A surveys a hundred patients. Forty strongly disagree, twenty are neutral and forty strongly agree. Clinic B has ten who disagree, eighty who are neutral and ten who agree. Scoring strongly disagree as one up to strongly agree as five, Clinic A's total is forty plus sixty plus two hundred, which is three hundred, so its mean is three. Clinic B's total is twenty plus two hundred and forty plus forty, which is also three hundred, so its mean is also three. One clinic is split down the middle and the other is calm.

**Kiffer:** I agree the mean hides that, and so does the median. Both clinics have a median of three. The split only shows up in the distribution. In Clinic A, eighty percent of patients sit at the two extremes, and in Clinic B nobody does. So for a single item, I'd lead with the percentage in each category.

**Sarah:** So where do we disagree?

**Kiffer:** We disagree about scales. Strictly speaking, a Likert scale is the total of several Likert items, and that's how Likert built it in 1932. He combined each person's answers across many statements into one score. A score built from, say, eight items has many possible values and behaves much more like a continuous measure. Geoff Norman of McMaster University reviewed the evidence in 2010. He pointed to studies going back to the 1930s showing that analysis of variance, regression and correlation hold up well when data are skewed or ordinal.

**Sarah:** Holding up well statistically doesn't settle what the number means, though. If I report that one group's mean is 3.4 and another's is 3.1, a reader takes that 0.3 as a real distance. If the categories are unevenly spaced, part of that distance could come from the numbers we chose to assign.

**Kiffer:** That's a fair point for single items, and it's the strongest version of your argument. For a multi-item scale whose items each meet the lesson's minimum of five points, I think the mean is defensible, and it's much easier to work with in a regression that includes other variables. I'd still want the distribution shown alongside it.

**Sarah:** I'm more cautious than Kiffer. For single items I'd report the distribution and perhaps the median, and I'd move to means only for a scale with evidence that it works as a scale.

**Kiffer:** Then I think we agree on the practical rule. Show the distribution of every item first. Keep don't-know answers out of the mean, and report how many there were. Use means for multi-item scales with at least five points per item, and check that an ordinal method leads to the same conclusion.

**Sarah:** That works for me.

**Kiffer:** Let's pull it together with three takeaways.

**Sarah:** First, a response rate tells you how much room there is for non-response bias. The size of the bias depends on how different the non-responders are, so look for evidence about who is missing.

**Kiffer:** Second, random error in an exposure flattens associations toward zero, and averaging several good items is the main remedy. Systematic error, like under-reported weight, needs validation against direct measurement.

**Sarah:** And third, the format of a question affects the answers it collects. Response options, scale ranges and codes all shape the numbers you analyse, so check them with think-aloud pre-tests and frequency tables before you trust a mean.

**Kiffer:** If you'd like more practice, rework the four-item loneliness problem from this episode with a single-item reliability of 0.5. Find how many items you need to reach a reliability of 0.8, and check at each step that the reliability stays below one.

**Sarah:** Next time, it's Lesson four, Measures of Disease Frequency, where the answers we've been collecting become prevalence, incidence and mortality rates.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson four.
