# Lesson 2: Surveillance and Sampling

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson two, Surveillance and Sampling.

**Sarah:** It's meant for listening once you've worked through the lesson. We'll add some perspective and some critique, work through the questions students most often find thorny in this material, and do some extra worked examples, including one that's harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here's what we'll cover. When reported cases go up, has the disease become more common? Why does a surveillance system that catches nearly every outbreak raise so many false alarms? Why can survey weights change an estimate so much? And how many people does a survey really need once it samples whole classrooms? We'll finish with non-probability samples, because Kiffer and I see them a little differently.

**Kiffer:** Let's take the first one.

**Sarah:** I see this kind of headline all the time. Reported cases of some infection rose by a quarter this year. My first instinct is that more people are getting infected.

**Kiffer:** That's the natural reading, and sometimes it's correct. A count from passive surveillance is the end of a chain, though. Someone has to seek care or ask to be tested, a clinician has to order the test, the test has to come back positive, and the result has to be reported. A change at any link in that chain changes the count, even when the amount of infection in the community stays the same.

**Sarah:** So more testing on its own could push the count up.

**Kiffer:** Yes. Let's try it with made-up numbers. In one year, a health region runs twenty-four thousand chlamydia tests, and twelve hundred come back positive. The next year, after a campaign encouraging young adults to get tested, it runs thirty-seven thousand five hundred tests, and fifteen hundred come back positive. Reported cases are up by twenty-five percent. So here's a question for everyone listening. What share of tests came back positive in each year, and what does that suggest about whether infection became more common? Take a few seconds.

*(Pause)*

**Sarah:** In the first year, twelve hundred divided by twenty-four thousand is 0.05, or five percent. In the second year, fifteen hundred divided by thirty-seven thousand five hundred is 0.04, or four percent. Testing went up by more than half, reported cases went up by a quarter, and the share of tests that were positive went down. So the rise in cases could come entirely from finding infections that were already there.

**Kiffer:** It could, and the falling positivity fits that explanation. I'd be careful with positivity as well, though, because it depends on who gets tested. If the campaign brought in mostly lower-risk people, positivity would fall even if infection were rising among the people at highest risk. Neither number settles the question by itself. I'd want to know who was tested, and whether the test or the case definition changed between the two years.

**Sarah:** Which is why the second step of the outbreak framework includes ruling out artefacts, like a new lab test or a change in reporting policy.

**Kiffer:** Right. And the chain can push the count down as well. During the Omicron wave, testing in British Columbia reached capacity, and by early January 2022 the province acknowledged that confirmed case numbers no longer reflected the true spread of the virus. By late January, most British Columbians no longer qualified for publicly funded testing, which was recommended only for people with symptoms in specific higher-risk groups and settings.

**Sarah:** So from then on, a falling case count could simply mean fewer tests. And the signals that don't depend on who gets tested, like wastewater, become more useful.

**Kiffer:** Yes, along with hospital admissions, although those lag behind infections.

**Sarah:** There's an equity side to this as well. Testing is easier to reach for some people than for others. If a clinic is close by and you can take time off work, you're more likely to be counted. So a map of reported cases can partly be a map of where testing is easy to get, and a rise in one neighbourhood can reflect a new testing site that opened there.

**Kiffer:** That's the non-random under-reporting in passive surveillance at work. The numerator holds only the cases that made it through every link of the chain, and who makes it through depends on severity, access to care and testing policy. So when a passive count changes, the first step is to check whether the chain that produced it changed.

**Sarah:** Question two. Why does a surveillance system that catches nearly every outbreak raise so many false alarms?

**Kiffer:** This one is about syndromic surveillance and two of the five quality dimensions, sensitivity and predictive value positive. Sensitivity asks what share of real outbreaks the system flags. Predictive value positive asks what share of the system's flags turn out to be real outbreaks.

**Sarah:** Those sound like the same question asked two ways.

**Kiffer:** They have different denominators, and that's where students tend to get tripped up. Here's a made-up example. A region monitors emergency department visits for gastrointestinal symptoms. Over one year, the system raises fifty alerts. Staff investigate each one and find that five of them were real outbreaks. Other sources show that the region had six real outbreaks of that kind during the year. What are the sensitivity and the predictive value positive of this system? Take a few seconds.

*(Pause)*

**Sarah:** Sensitivity uses the real outbreaks as the denominator. The system caught five of the six, and five divided by six is about eighty-three percent. Predictive value positive uses the alerts as the denominator. Five real outbreaks out of fifty alerts is ten percent.

**Kiffer:** Right. The system catches most outbreaks, and nine out of ten alerts are false alarms, each of which costs someone time to check. Now suppose the team re-runs the same year with a higher threshold, so it takes a bigger jump in visits to trigger an alert. The system would have raised twelve alerts, and four of them would have been real.

**Sarah:** Then sensitivity is four out of six, about sixty-seven percent, and predictive value positive is four out of twelve, about thirty-three percent. The team would investigate far fewer false alarms, and it would miss one more outbreak.

**Kiffer:** Moving the threshold generally trades one measure for the other. Where to set it depends on what a missed outbreak costs compared with a false alarm, and that's a judgement about values as well as statistics. The predictive value positive also depends on how often real outbreaks happen. A threshold that works well in a large city may produce mostly false alarms in a small region where outbreaks are rare.

**Sarah:** That sounds a lot like screening.

**Kiffer:** It's the same arithmetic. You'll meet it again in Lesson five, Screening and Diagnostic Tests, where the predictive value of a test depends on prevalence in the same way.

**Sarah:** Is there any way to improve both at once?

**Kiffer:** Better information in each signal can do it. Laboratory surveillance with genetic sequencing is the clearest case. In the 2017 to 2018 romaine lettuce outbreak, forty-two Canadians fell ill across five eastern provinces, and sequencing showed that bacteria from patients in Canada and the United States were closely related genetically. Sequencing can link cases that are too scattered to stand out on their own, and a cluster of genetically matching cases is much more likely to be a real outbreak than a jump in emergency visits is. So it raises sensitivity and predictive value positive together.

**Sarah:** Question three is about survey weights. I'll admit I used to think of weighting as a technical correction that nudged the numbers a little.

**Kiffer:** It can move them a long way. Here's a made-up provincial survey of daily smoking. Rural residents are twenty percent of adults in the province, and urban residents are eighty percent. The survey team wants a precise estimate for rural areas, so it samples five hundred rural adults and five hundred urban adults. That's oversampling the rural stratum. In the sample, eighteen percent of rural respondents smoke daily, and eight percent of urban respondents do. What's the best estimate of daily smoking for the province as a whole? Take a few seconds.

*(Pause)*

**Sarah:** The tempting answer is to pool everyone. Ninety rural smokers plus forty urban smokers is a hundred and thirty, out of a thousand respondents, so thirteen percent.

**Kiffer:** And why is that too high?

**Sarah:** Because rural adults make up half the sample but only a fifth of the province, and they smoke more. So each group gets weighted by its share of the population. 0.2 times 0.18 is 0.036, and 0.8 times 0.08 is 0.064. Together that's 0.1, or ten percent.

**Kiffer:** Right. Can you put that in terms of sampling weights?

**Sarah:** A weight is one divided by the probability of selection. Say the province has two hundred thousand rural adults. Then each rural respondent had a one in four hundred chance of selection and gets a weight of four hundred. Each urban respondent was drawn from eight hundred thousand adults and gets a weight of sixteen hundred. Each person stands in for that many people in the province, and the weighted estimate comes out at ten percent again.

**Kiffer:** Exactly. And a bigger sample wouldn't rescue the unweighted figure. With fifty thousand people in each stratum, the unweighted estimate would still be about thirteen percent, now with a narrow confidence interval around the wrong value. Sample size shrinks random error. It does nothing about a sample whose make-up differs from the population's. That's why the Canadian Community Health Survey has to be analysed with its weights, however large the file is.

**Sarah:** Is there a cost to weighting?

**Kiffer:** There usually is a cost in precision. When weights vary a lot, a few heavily weighted people carry more of the estimate, and the standard error tends to grow. The weighted analysis gives an unbiased estimate with a wider interval, and that wider interval reflects what the sample can actually tell you.

**Sarah:** Question four. How many people does a survey really need once it samples whole classrooms? I'd like to try this one myself.

**Kiffer:** Go ahead. Here's the set-up, with made-up values. A health region wants to estimate the share of Grade ten students who vaped in the past month. Earlier surveys suggest about twenty percent. The region wants the estimate within plus or minus four percentage points, with ninety-five percent confidence. It will sample whole classes. Classes average twenty-five students, and about eighty percent of students in a class usually take part, so about twenty per class provide data. Assume the intracluster correlation for vaping within a class is 0.05. How many students should the region invite? Take a few seconds.

*(Pause)*

**Sarah:** Okay. Step one is the formula for a proportion, 1.96 squared times p times q, divided by L squared. 1.96 squared is 3.8416. P is 0.2 and q is 0.8, so p times q is 0.16. L is 0.04, and 0.04 squared is 0.0016. Conveniently, 0.16 divided by 0.0016 is exactly one hundred, so the answer is 384.16, which rounds up to three hundred and eighty-five students.

**Kiffer:** Good. Now the clustering.

**Sarah:** The design effect is one plus rho times the quantity m minus one. M is the number of people per cluster who actually provide data, so twenty. That's one plus 0.05 times nineteen, which is one plus 0.95, or 1.95. Three hundred and eighty-five times 1.95 is 750.75, so we need seven hundred and fifty-one students who take part.

**Kiffer:** Good. So what's the next step after the design effect?

**Sarah:** Then comes non-response. Twenty percent won't take part, so I add twenty percent. Seven hundred and fifty-one times 1.2 is 901.2, so the region invites nine hundred and two students.

**Kiffer:** Try a sanity check. If you invite nine hundred and two students and eighty percent of them take part, how many do you end up with?

**Sarah:** Nine hundred and two times 0.8 is 721.6, so about seven hundred and twenty-two. That's short of seven hundred and fifty-one. Adding twenty percent doesn't make up for losing twenty percent, because the twenty percent I lose is taken from the larger number.

**Kiffer:** Right. So what's the fix?

**Sarah:** Divide by the share who take part. Seven hundred and fifty-one divided by 0.8 is 938.75, so nine hundred and thirty-nine students. And since the region samples whole classes, nine hundred and thirty-nine divided by twenty-five is about 37.6, so thirty-eight classes, which means inviting nine hundred and fifty students. For comparison, a simple random sample of individual students would have needed three hundred and eighty-five divided by 0.8, or four hundred and eighty-two invitations. Sampling whole classes nearly doubles that.

**Kiffer:** That's the answer. A buffer of ten to twenty-five percent is the usual rule of thumb, and when you have an expected response rate, dividing by it gives the number directly.

**Sarah:** And the clustering has to be carried into the analysis as well. If the region analysed its students as though each one had been sampled independently, the standard errors would be too small by a factor of the square root of 1.95, which is about 1.4.

**Kiffer:** So the reported confidence interval would be only about seventy percent as wide as it should be, and differences between schools or between years would look more certain than they are.

**Sarah:** And the whole calculation rests on that intracluster correlation of 0.05, which is really an educated guess.

**Kiffer:** It is, so it's worth checking how much the answer moves. If the intracluster correlation were 0.1, the design effect would be one plus 0.1 times nineteen, or 2.9. Three hundred and eighty-five times 2.9 is 1,116.5, so about eleven hundred and seventeen students taking part, or fifty-six classes. I'd run the calculation for a few plausible values, choose one, and write down why.

**Sarah:** Okay. Now the part where we disagree.

**Kiffer:** This one is about non-probability samples.

**Sarah:** I think the rule that non-probability samples are unsuitable for estimating a prevalence is too strict for the way surveys work now. Response rates to probability surveys have been falling. Statistics Canada's Labour Force Survey had a response rate of eighty-seven percent in 2019 and about seventy percent by September 2023. Once a large share of the people you selected don't respond, the estimate depends on weighting adjustments that rest on assumptions about the non-respondents. A weighted online panel rests on assumptions of the same kind.

**Kiffer:** That's a fair point, and Canada has a well-known example of how much response matters. In 2010, the federal government replaced the mandatory long-form questionnaire for the 2011 census with the voluntary National Household Survey. It went to about four and a half million dwellings, slightly under a third of all private dwellings in the country, and the response rate was 68.6 percent. The mandatory long form in 2006 had a response rate of 93.5 percent.

**Sarah:** So it was an enormous probability sample, and nearly a third of households still didn't answer.

**Kiffer:** Yes. Here's where I'd hold the line, though. Statistics Canada could measure that problem. It knew the non-response rate for every area, it drew a follow-up subsample of four hundred thousand dwellings that hadn't responded, and it withheld its standard estimates for areas where the non-response rate was fifty percent or higher. With an opt-in online panel, nobody knows who could have joined and didn't, so there's no response rate to report and no way to follow up the people who were missed.

**Sarah:** But for some questions there's no frame at all. If you want a prevalence among people who use drugs, or among gay and bisexual men, there's no list to draw from.

**Kiffer:** Agreed, and that's where respondent-driven sampling and time-location sampling are worth their extra effort. My concern is the general opt-in panel used for an ordinary prevalence. In an experiment the Pew Research Center reported in 2024, twelve percent of opt-in respondents under thirty said they were licensed to operate a nuclear submarine. The real share of Americans with that licence rounds to zero. Pew suggests that some respondents say yes to get past screening questions and collect a reward, and that can inflate the prevalence of anything rare.

**Sarah:** That's a data-quality problem, and panels can screen out some of those respondents with quality checks. I'd still say a careful non-probability sample, weighted to the census, can give a usable estimate when no probability survey exists or one would take too long.

**Kiffer:** Those checks help. In an earlier Pew study, though, most of the respondents giving bogus answers still passed a simple attention check. I'd accept a careful panel estimate as a rough guide, especially when it agrees with a probability source. I wouldn't let it stand alone as the official figure for a prevalence. For what it's worth, the mandatory long form came back for the 2016 census, and its response rate was 97.8 percent.

**Sarah:** I'm more open to non-probability samples than Kiffer is, but we agree on the practical rule. For a prevalence, use a probability sample whenever a frame exists, and report the response rate. When no frame exists, describe exactly how people were recruited, weight to known population figures such as the census, and check the estimate against a probability source wherever one is available.

**Kiffer:** And whichever design you use, compare the people who took part with the population you're describing. That's the representativeness question from the surveillance material, applied to a survey.

**Sarah:** Let's pull it together with three things to take away.

**Kiffer:** First, a surveillance count is the end of a chain of care-seeking, testing and reporting. When the count changes, check whether the chain changed before concluding that the disease did.

**Sarah:** Second, sensitivity and predictive value positive have different denominators, and moving an alert threshold trades one for the other. Where you set it depends on the cost of a missed outbreak compared with the cost of a false alarm.

**Kiffer:** And third, the sampling design has to be carried through to the planning and the analysis. Weights correct for unequal chances of selection, the design effect enlarges the sample for clustering, and the target is divided by the expected response rate. A larger sample shrinks random error and leaves bias where it was.

**Sarah:** If you'd like more practice, rework the Grade ten vaping problem from this episode with a different intracluster correlation or response rate. Then do the sanity check at the end: multiply the number you invite by the response rate and confirm that you reach your target.

**Kiffer:** Next time, it's Lesson three, Questionnaire Design, where the question becomes what to ask the people you've sampled and how to ask it.

**Sarah:** Take care, everyone.

**Kiffer:** See you in Lesson three.
