Surveillance and Sampling
Fundamental Epidemiological Concepts and Approaches
Learning objectives for this lesson:
- Distinguish passive, active, sentinel, and syndromic surveillance, identify Canadian examples of each, and trace a notifiable disease report from the clinic to the Public Health Agency of Canada
- Navigate the major federal and BC surveillance products (CNDSS, FluWatch, CCDSS, BCCDC dashboards, CVSD) and the registry and vital-statistics infrastructure beneath them, and interrogate any source on its timeliness, completeness, representativeness, sensitivity, and predictive value positive
- Apply the CDC 10-step outbreak investigation framework and the Canadian FIORP to a real Canadian outbreak, and compute attack rates and risk ratios from a line list
- Explain why an investigation must sample when the population at risk cannot be enumerated, and describe the target, source, and study population hierarchy and the sampling frame
- Compare probability sampling methods (simple random, systematic, stratified, cluster, multistage, targeted) with non-probability methods (judgement, convenience, purposive, chain-referral) and choose among them for a given question
- Use probability distributions, the standard error, and the central limit theorem to justify inference from a sample to a population, and account for weights, stratification, clustering, and the design effect in analysis
- Explain Type I and Type II errors and statistical power, and compute required sample sizes for common descriptive and analytic objectives, including adjustments for clustering, attrition, and finite populations
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.
Glossary: Key Terms, People & Concepts
📚 Reference page, available throughout the lesson
This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.
Surveillance Systems and Canadian Data Sources
Introduction and Overview
You are the duty epidemiologist at a regional health authority on a Tuesday afternoon. Three minutes ago, your phone buzzed: a paediatrician at a community clinic just called to say she has seen four children from the same school presenting with bloody diarrhoea over two days. She is requesting stool cultures and wants to know whether you have seen anything similar elsewhere. By the time you put the phone down, you will need to know, quickly, whether four cases is unusual for this organism in this catchment, whether a notifiable disease report has already been filed, who else needs to be looped in, and what data sources you can pull within the next hour. That sequence of questions is what this lesson is about. The infrastructure that lets you answer them is called public-health surveillance, and the structured response that follows is called outbreak investigation.
This section covers that infrastructure. It defines surveillance, describes the four types of surveillance system and the Canadian reporting flow, introduces the federal and provincial surveillance products and the data infrastructure beneath them, and sets out five dimensions on which the quality of any surveillance source can be judged.
Learning Objectives
- State Langmuir's working definition of surveillance and explain what it means to "close the loop."
- Distinguish passive, active, sentinel, and syndromic surveillance, give Canadian examples of each, and explain why under-reporting in passive systems is non-random.
- Trace the notifiable-disease reporting flow from clinician to medical health officer, province, and PHAC.
- Identify the major federal and BC surveillance products (CNDSS, FluWatch, CCDSS, CVSD, BCCDC dashboards) and the long-running data infrastructure they read from.
- Apply the five dimensions of surveillance data quality (timeliness, completeness, representativeness, sensitivity, and predictive value positive) to any data source.
What Public-Health Surveillance Is, and What It Is For
The classic working definition, attributed to Langmuir (1963) and refined by Thacker & Berkelman (1988), is the ongoing, systematic collection, analysis, interpretation, and dissemination of health-related data for the planning, implementation, and evaluation of public-health practice. Each of those verbs is doing real work. Ongoing distinguishes surveillance from a one-time study. Systematic rules out anecdote. Analysis and interpretation rule out a passive data warehouse. Dissemination for action rules out research that is never returned to the people who can act on it. A famous shorthand attributed to Langmuir (EIS founder) is that surveillance only counts if it “closes the loop”: information must come back out as decisions, alerts, or programs, otherwise the system is just bookkeeping.
The five purposes of surveillance
When you read a published surveillance report, the authors are usually doing one or more of these five things: (1) detect outbreaks, clusters, and unusual events early; (2) characterize who is getting sick, where, and why (descriptive epi by person, place, and time); (3) monitor trends in incidence, prevalence, and risk factors over time; (4) evaluate the effect of interventions and programs; and (5) plan resource allocation and policy. A single dataset can serve more than one purpose, but the design tradeoffs are different for each.
The Surveillance Action Loop
It helps to picture surveillance as a closed loop with four moving parts: data sources (clinical reports, lab results, vital records, administrative data), data systems (the registries, dashboards, and notifiable-disease platforms that ingest and store data), analysis and interpretation (the epidemiologists who turn case counts into rates, anomalies, and narratives), and action (the case finding, control measures, public communications, and policy decisions that flow from the analysis). Each handoff between these stages is also a place where the system can leak: a clinician who never reports a case, a province that does not share data with the federal level, an analyst who does not see a signal in time, or a recommendation that never reaches a decision-maker.
For the rest of this section we focus on the data-system layer because it determines what kinds of signals you can detect at all. The four conventional system types differ on a single axis: who is doing the work of finding cases.
The Four Surveillance System Types
Most public-health surveillance systems can be sorted into one of four types; the syndromic category is the newest, codified by the CDC framework of Buehler and colleagues (2004). They are not mutually exclusive, since a single disease can be tracked by several at once, but they have different strengths, costs, and biases.
| Type | Who initiates the report | Strengths | Limitations | Canadian example |
|---|---|---|---|---|
| Passive | Clinicians and labs report when they encounter a notifiable disease. | Cheap, broad coverage, mandated by law, runs continuously. | Under-reporting (often substantial), variable timeliness, completeness depends on clinician burden. | The Canadian Notifiable Disease Surveillance System (CNDSS), aggregated case counts for ~50 reportable conditions submitted by provinces to PHAC. |
| Active | Public-health staff actively contact providers, labs, or households to find cases. | Higher case ascertainment, better data quality, useful in outbreak investigations. | Resource-intensive, narrow scope, hard to sustain over time. | The 100-clinician active-search component of FluWatch, and contact tracing during COVID-19 case investigations. |
| Sentinel | A small, designated network of providers reports systematically, trading breadth for depth. | High data quality, manageable cost, can collect richer data than passive systems. | Not population-representative; trends but not absolute counts. | FluWatch sentinel practitioners (general practitioners reporting weekly influenza-like-illness rates) and the Canadian Paediatric Surveillance Program (CPSP). |
| Syndromic | Real-time signals (chief-complaint codes, EMS calls, OTC drug sales, school absences) flag clusters before lab confirmation. | Fast, and can detect events before a definitive diagnosis. Useful for emerging or rare events. | Low specificity (lots of false alarms); validation is hard; needs analytic infrastructure. | BC's Acute and Communicable Disease Prevention ED chief-complaint monitoring; PHAC's pandemic-era wastewater surveillance dashboards. |
A fifth category, laboratory-based surveillance, is sometimes broken out separately because it sits beside (rather than under) the clinician's desk: provincial public-health labs aggregate isolates from clinical labs, perform serotyping or whole-genome sequencing, and feed the results into both passive and active systems. PulseNet Canada is the best-known example, the network that enabled the 2008 Maple Leaf listeriosis outbreak to be linked across provinces by genome (Gilmour et al., 2010).
Why “passive” under-reporting is not random
A common student misconception is that under-reporting in passive surveillance just makes case counts smaller. It does, but it also biases who is counted. Cases that present to the health system, get tested, and produce a positive lab result are systematically over-represented. The result is that severe cases, urban cases, and cases in well-insured populations are more visible in the data than mild, rural, or uninsured ones. When you read a CNDSS rate, the denominator is the catchment population, but the numerator is selected.
The Notifiable Disease Reporting Flow in Canada
For passive surveillance to function, every link in the chain has to do its part. The Canadian flow looks like this:
- Clinician or laboratory identifies a case of a notifiable disease (e.g., a positive shiga-toxin-producing E. coli culture or a clinical diagnosis of measles). Each province publishes its own list of conditions reportable under public-health legislation.
- Local Medical Officer of Health (MOH) or regional health authority receives the report (typically within 24 hours for urgent conditions, longer windows for routine ones). The MOH may immediately initiate case investigation, contact tracing, or public-health control measures.
- Provincial / territorial public-health authority (e.g., BCCDC, Public Health Ontario, the Quebec INSPQ) aggregates reports from the regional level, performs initial analysis, and shares anonymized aggregate data with the federal level.
- Public Health Agency of Canada (PHAC) publishes nationally aggregated counts in CNDSS and feeds disease-specific products like FluWatch and the CCDSS. PHAC also reports onward under the WHO International Health Regulations (2005) for events of international concern; the post-SARS rationale for that regime is laid out by Heymann & Rodier (2004).
Two features of this flow deserve emphasis. First, public-health legislation is provincial in Canada, so the list of notifiable conditions and reporting timelines differ across provinces. A condition might be urgent-reportable in one province and routine in another. Second, the federal level receives aggregated, de-identified data only; PHAC cannot pull individual records. This federal-provincial division is one reason that the COVID-19 pandemic exposed gaps in real-time data sharing, because the legal architecture was not designed for the latency that a respiratory pandemic demands.
The Federal Layer
PHAC operates several headline products. They are aggregated and curated, the agency cannot dispense individual-level data, and they each have a different cadence, scope, and intended audience.
The flagship passive system. Provinces submit weekly aggregated counts for ~50 nationally notifiable conditions (the list is harmonized but not identical to provincial lists). CNDSS feeds the agency's annual Notifiable Diseases Online reports and the underlying open-data tables on Open.canada.ca. Useful for long-run trends and inter-provincial comparisons; less useful for real-time outbreak detection because of the multi-week lag from clinic to PHAC.
A hybrid system: a sentinel network of ~150 family-medicine clinicians reports influenza-like-illness rates each week, lab partners across the country submit subtyped influenza and (post-2020) SARS-CoV-2 results, and provincial outbreak counts feed the national picture. The weekly FluWatch report is what most public-health communicators reach for during respiratory season.
A federal-provincial collaboration that uses validated case definitions on top of health-administrative records (physician billing claims, hospital discharge abstracts) to estimate prevalence and incidence of chronic conditions like diabetes, hypertension, asthma, dementia, and ischaemic heart disease. CCDSS is the largest single source of population-level chronic-disease data in Canada and powers the agency's chronic-disease infobase.
Not strictly a surveillance product but the spine of mortality-based surveillance. Statistics Canada compiles every death registered in Canada, with cause coded to ICD-10. CVSD enables life-expectancy estimates, cause-specific mortality trends, and excess-mortality analyses (most prominently used during COVID-19 to estimate pandemic burden).
PHAC also runs targeted systems for HIV (including the Canadian Perinatal Surveillance System branch), tuberculosis (CTBRS), antimicrobial resistance (CIPARS, CARSS), opioid-related harms, and others, alongside Canada's contribution to global digital surveillance, the Global Public Health Intelligence Network (GPHIN), described by Mykhalovskiy & Weir (2006) and complemented internationally by HealthMap (Brownstein, Freifeld, & Madoff, 2009). Each lives at canada.ca/en/public-health and is worth knowing exists.
The Provincial Layer (with BC examples)
The federal products are valuable for the country-level view, but most case-level work happens provincially. In British Columbia, the BC Centre for Disease Control (BCCDC) is the analytic and operational arm of provincial public health, and most of its products are publicly available.
A short tour of BCCDC dashboards worth bookmarking
- BCCDC respiratory pathogens dashboards: weekly influenza, RSV, and SARS-CoV-2 surveillance with regional breakdowns.
- BCCDC enteric pathogens dashboards: Salmonella, Campylobacter, STEC, Listeria; updated weekly.
- BC Vital Statistics overdose deaths: the unrestricted-toxic-drug-supply mortality reports that the BC Coroners Service and BCCDC co-produce.
- BC Cancer Registry and surveillance: cancer-incidence dashboards by health-service-delivery area.
- BC Sexually Transmitted Infection Quarterly: gonorrhea, chlamydia, syphilis, congenital syphilis, and infectious syphilis trends.
Behind these dashboards sit two case-management platforms: IRIS (Integrated Reporting Information System, BCCDC's communicable-disease platform) and Panorama (a multi-province public-health platform used in BC for outbreak management and immunization records).
The Long-Running Data Infrastructure
Most surveillance products do not generate their own data; they read from infrastructure that exists for other reasons.
- Vital statistics (births, deaths, marriages, divorces) are the oldest population health data in Canada, with continuous coverage since the late 19th century in some provinces. They are the denominator for life expectancy and the numerator for mortality surveillance.
- Cancer registries: the Canadian Cancer Registry aggregates provincial cancer registries and is one of the few systems with active follow-up and complete case ascertainment. The BC Cancer Registry is the provincial source.
- Health-administrative data: the Discharge Abstract Database (DAD, hospital admissions), the National Ambulatory Care Reporting System (NACRS, ED visits), and provincial Medical Services Plan billing claims (MSP in BC). CCDSS and many academic studies are built on top of these.
- Population health surveys: the Canadian Community Health Survey (CCHS) is the workhorse cross-sectional survey for self-reported health behaviours and outcomes; the Canadian Health Measures Survey (CHMS) layers in physical measurement.
- Wastewater-based surveillance: a rapidly maturing infrastructure since 2020. PHAC and many provinces now sample wastewater for pathogens (SARS-CoV-2, influenza, mpox) as a non-clinical signal that is independent of who seeks testing.
Data Quality: Five Dimensions to Question Every Source
Every surveillance product makes tradeoffs across these five dimensions. When you read a public-health report, the report's author has implicitly resolved each one:
- Timeliness: how long from event to data product? Wastewater can be days. CNDSS can be months.
- Completeness: what fraction of true events are captured? STIs are notoriously incomplete; cancer registries are nearly complete.
- Representativeness: do the captured cases reflect the affected population? Sentinel networks usually do not.
- Sensitivity: will the system detect a true outbreak when it occurs? Syndromic surveillance is highly sensitive, passive surveillance often is not.
- Predictive value positive: when the system flags an event, is it real? It tends to trade off against sensitivity as the alert threshold moves, especially for syndromic systems, and it also depends on how often true events occur.
Pick one of the BCCDC dashboards listed above (or a comparable PHAC product) and spend ten minutes with it. Locate (a) the most recent week's case count, (b) the underlying case definition, and (c) any explicit data-quality caveats the dashboard publishes. Note the date of the most recent update versus today's date: how big is the lag?
Key Takeaways
- Surveillance is the ongoing, systematic collection, analysis, interpretation, and dissemination of health data for action; collection without action does not count.
- Passive, active, sentinel, and syndromic systems trade coverage, cost, depth, and timeliness, and passive under-reporting falls most heavily on mild, rural, and under-served cases.
- Notifiable-disease lists are set by the provinces and territories; PHAC aggregates provincial data into national products such as CNDSS, FluWatch, and CCDSS.
- Most surveillance products read from long-running infrastructure (vital statistics, registries, health-administrative data, national surveys, wastewater).
- Timeliness, completeness, representativeness, sensitivity, and predictive value positive determine what any source can and cannot answer.
Reflection
You are asked by a journalist for the “most accurate” current count of chlamydia cases in BC. Three sources give different numbers: the Canadian Notifiable Disease Surveillance System (CNDSS), which compiles laboratory-confirmed cases that clinicians and laboratories report through provincial public health to the Public Health Agency of Canada, with a reporting lag; the BCCDC quarterly STI report, which tabulates confirmed cases reported within the province closer to real time but still counts only people who were tested; and a recent academic estimate based on self-reported diagnoses in the Canadian Community Health Survey (CCHS), a population sample survey run by Statistics Canada that reaches untested people but depends on recall and willingness to disclose. Surveillance data can be judged on five quality dimensions: timeliness, completeness, representativeness, sensitivity, and predictive value positive. How would you explain the discrepancy among the three numbers, using these dimensions, without making any of the three systems sound discredited?
Minimum 20 characters required.
Question 1: A provincial laboratory operates a network of 80 family-medicine clinics across Canada that submit weekly counts of patients presenting with influenza-like illness, alongside the results of throat-swab subtyping. This is best described as:
Question 2: Which of the following is the strongest reason that passive surveillance under-counts cases?
Question 3: Which Canadian surveillance product is built on top of physician billing claims and hospital discharge abstracts to estimate the prevalence of conditions like diabetes and hypertension?
Detecting and Investigating Outbreaks
Introduction and Overview
Surveillance is meant to detect signals, and outbreak investigation is the structured response when a signal turns into a problem. This section sets out what counts as an outbreak, the CDC 10-step framework that organizes an investigation, and the Canadian protocol for foodborne outbreaks. It then follows the 2017–18 romaine lettuce E. coli O157:H7 outbreak through the same ten steps and closes with the line-list analysis that sits at the centre of most investigations.
Learning Objectives
- Distinguish cluster, outbreak, epidemic, and pandemic, and explain why declaring an outbreak is partly a statistical and partly an operational judgement.
- Describe the CDC 10-step outbreak investigation framework and identify which steps typically run in parallel.
- Describe the Canadian Foodborne Illness Outbreak Response Protocol (FIORP) and the roles of PHAC, CFIA, and Health Canada.
- Follow the 2017–18 romaine lettuce outbreak through the ten steps and explain the roles of whole-genome sequencing and the case-case design.
- Compute attack rates and risk ratios from a line list and explain how stratification separates two suspect foods.
What Counts as an Outbreak?
The textbook definition is the occurrence of more cases of a disease than expected in a given population, place, or time. Each italicized phrase is doing work. “More than expected” presupposes a baseline: the local 5-year average for influenza in the same week, or the seasonal threshold from a regression model. “Population, place, time” says outbreaks are local: 30 cases of Campylobacter across Canada in a week is not unusual; 30 cases at one wedding is.
Three working terms you need to use precisely
- Cluster: an aggregation of cases in time and/or space that may or may not be statistically unusual. Clusters are flagged by surveillance and triaged for further investigation.
- Outbreak: a cluster judged to exceed the expected baseline. The threshold is decided by the public-health authority, not by the data alone.
- Epidemic: a large-scale outbreak, often across multiple jurisdictions; in international usage often synonymous with outbreak.
- Pandemic: an epidemic with worldwide geographic spread. The WHO declares pandemics; severity is a separate dimension and is not part of the definition.
A common student error is to treat “pandemic” as “severe outbreak”: it is a geographic claim, not a severity claim.
The decision to call something an outbreak is partly statistical and partly operational. Statistically, you can compare current counts to a baseline distribution and apply a threshold (e.g., the upper 95% confidence limit of the 5-year mean). Operationally, an outbreak declaration mobilizes resources, triggers a coordinated response, and may require public communication, and authorities are reasonably cautious about both over- and under-calling.
The CDC 10-Step Outbreak Investigation Framework
The reference framework most North American epidemiologists learn is the CDC's 10-step process, articulated in the modern era by Reingold (1998) and elaborated in the CDC Field Epidemiology Manual. The steps look orderly on paper but in practice you often work several at once and revisit earlier ones as new information arrives.
Walk through the most famous post-war outbreak investigation, scene by scene. Next ▶ advances.
An 8-scene retelling of the 1976 American Legion convention outbreak in Philadelphia, illustrating the CDC's 10-step framework: detection, case definition, descriptive epi, hypothesis generation, analytic study, environmental investigation, agent identification (a brand-new bacterium), and control measures.
Before you arrive, you confirm authority and roles, brief on the suspected etiology, gather supplies (case-report forms, lab kits, PPE), and identify local liaisons (MOH, environmental health, lab). Preparation is the step new investigators most often skip and most often regret.
Compare current counts to a baseline. If the baseline is unstable (small denominators, seasonal variation), state how you constructed it. Ruling out artefact, whether a new lab test, a reporting policy change, or a clinician on a reporting kick, is part of this step.
Talk to clinicians, review charts, check that lab results are correctly attributed. A pseudo-outbreak driven by a contaminated lab reagent or a misclassification is not unheard of.
A case definition has three parts: clinical criteria (symptoms, signs, lab tests), person/place/time restrictions (e.g., attendees of the August 12 potluck), and a level of certainty (suspect, probable, confirmed). You will revise it as the investigation evolves, which is normal.
Active case finding (chart review, asking clinicians, contacting attendees of the implicated event) yields a line list, with one row per case carrying demographic, clinical, exposure, and outcome variables. The line list is the working dataset for everything that follows.
The classic person, place, time triad. The flagship visualization is the epidemic curve (epi curve), a histogram of case counts by date of symptom onset. Its shape (point-source vs propagated) constrains your hypotheses about the exposure window.
From the descriptive epi (and from open-ended interviews of cases) you generate plausible exposures: a specific food, a specific water source, a specific event, a specific procedure. Good hypotheses are testable with the data you have or can collect.
The two workhorse designs in outbreak settings are the retrospective cohort (when you can enumerate everyone who attended an event, e.g., the wedding-guest list) and the case-control study (when you cannot). The attack rates, 2×2 tables, and risk ratios used here are developed formally in Lesson 4 (measures of disease frequency) and Lesson 6 (measures of association), and the line-list analysis later in this section shows how they are used in an investigation.
The first analytic pass often points to several plausible exposures. Environmental sampling, traceback investigations (where did the implicated food come from?), and lab characterization (genome-typing of isolates) sharpen the inference.
Control measures (recalling a product, closing a venue, prophylaxis, isolation) often happen before Step 8; the precautionary principle does not require you to wait for a p-value. Communication runs throughout: with the public, with affected communities, with policymakers, and through the final outbreak report.
The Canadian FIORP and Its Multi-Jurisdictional Structure
For foodborne outbreaks specifically, Canada operates under the Foodborne Illness Outbreak Response Protocol (FIORP), summarised by Vik & Hexemer (2014), which formalizes the roles of the federal partners and provincial/territorial public-health authorities. Three federal partners share the load:
- PHAC: epidemiology and surveillance lead, including PulseNet Canada (the lab network that does whole-genome sequencing of bacterial isolates).
- Canadian Food Inspection Agency (CFIA): food-safety investigations, traceback, and product recalls.
- Health Canada: risk assessment of contaminated products and health-impact guidance.
FIORP defines escalation triggers (when a multi-jurisdictional outbreak exists), establishes an Outbreak Investigation Coordinating Committee (OICC) for inter-provincial events, and lays out communication protocols. The 2008 Maple Leaf listeriosis outbreak (the deli-meat-associated Listeria monocytogenes outbreak that killed 22 Canadians) is the case that drove the modern revisions to FIORP and to the supporting infrastructure of PulseNet.
Case Study: The 2017–18 Romaine Lettuce E. coli O157:H7 Outbreak
The ten steps are easier to follow in a real investigation. The boxes below follow one Canadian outbreak through the framework.
Between November 2017 and February 2018, PHAC and US CDC investigated a multi-jurisdictional outbreak of E. coli O157:H7 infections that ultimately involved 42 confirmed cases across 5 Canadian provinces (and 25 cases across 15 US states). All five affected Canadian provinces were in eastern Canada (Ontario, Quebec, New Brunswick, Nova Scotia, and Newfoundland and Labrador); western provinces were not involved. One death was reported among the 42 Canadian cases, a case fatality rate of about 2% (roughly 2 in every 100 diagnosed cases died). The outbreak is a useful teaching case because it shows the full FIORP machinery in motion and makes the role of whole-genome sequencing visible.
Step 2–3: Establishing the outbreak and verifying the diagnosis
PulseNet Canada flagged a cluster of E. coli O157:H7 isolates with matching pulsed-field gel electrophoresis (PFGE) patterns, later confirmed by whole-genome sequencing (WGS) to share a tight phylogenetic neighbourhood. The genomic signal is what made this an investigable outbreak rather than a scattered set of unrelated cases; without WGS, the same cases would have been distributed across the routine STEC surveillance baseline.
Steps 4–6: Case definition and descriptive epidemiology
Confirmed cases were Canadian residents with WGS-matched E. coli O157:H7 isolates and symptom onset between mid-November 2017 and the closing date. Probable cases were household contacts of a confirmed case with compatible symptoms. The descriptive analysis (epi curve, geographic distribution, age and sex distribution) showed a polymorphic temporal pattern with several waves, typical of a continuous common-source outbreak driven by an ongoing contaminated supply chain, not a single point exposure.
Steps 7–8: Hypothesis development and the case-case analysis
Standard hypothesis-generating interviews were conducted with confirmed cases using PHAC's Hypothesis Generating Questionnaire (HGQ) for STEC, which lists hundreds of food and exposure variables. Cases reported leafy greens consumption substantially more often than expected based on national consumption patterns. A case-case analysis comparing outbreak cases to historical sporadic STEC cases pointed strongly to romaine lettuce. CFIA conducted parallel traceback investigations from grocery purchases reported by cases.
Steps 9–10: Refinement, control, and communication
Joint PHAC–CFIA–US-CDC coordination produced a public advisory in late December 2017 urging Canadians in the affected provinces to avoid romaine lettuce. The Canadian outbreak was declared over in mid-January 2018. Importantly, despite intensive traceback, no specific grower or facility could be definitively implicated, a sobering reminder that even good investigations sometimes end without the closure of a single confirmed source.
Three teaching points are worth pulling out of this case. First, the outbreak would not have been detected at all without genomic surveillance; the case counts in any one province in any one week looked like background noise. Second, the case-case study design is a pragmatic alternative to a traditional case-control study when controls are hard to recruit during a fast-moving foodborne investigation; you trade some validity for considerable speed. Third, control measures (the public advisory) were issued before the source was definitively confirmed; precaution is a defensible public-health stance when downside asymmetry favours acting early.
From Line List to Risk Ratio
Most investigations rest on a line list, which holds one row per person, with columns for illness, date of onset, and exposures. Three quantities come directly from it. The attack rate is the proportion of an exposed group who became ill. For each suspect food, a 2×2 table compares the attack rate among people who ate the food with the attack rate among people who did not, and the ratio of the two attack rates is the risk ratio. A risk ratio near 1 suggests that the food is unrelated to illness, and a large risk ratio marks the food as a suspect.
When many people ate two foods together, both foods can show a high crude risk ratio, because each food carries part of the other's association. Stratification separates them: the risk ratio for one food is computed separately among people who did and did not eat the other. The reflection at the end of this section applies this step to a potluck outbreak. Lesson 4 (measures of disease frequency) and Lesson 6 (measures of association) develop the attack rate and the risk ratio formally, and Lesson 7 develops stratification as Mantel-Haenszel adjustment for confounding.
Real-Time vs Retrospective: A Standing Tension
An outbreak investigation is run under two competing pressures. Speed, since every day of delay can mean more illness, pushes you toward early hypotheses and precautionary control measures. Accuracy, since falsely accusing a food product or a venue has real costs, pulls in the other direction. Experienced investigators learn to act on confident-enough evidence, communicate uncertainty honestly, and revise control measures as data evolve. The skill is partly statistical and partly ethical: who bears the cost of being wrong in either direction is rarely symmetric.
Equity in surveillance and outbreak response
Surveillance systems do not see all populations equally. Data quality, case ascertainment, and willingness to be tested all vary by social position; outbreak investigators are increasingly expected to ask whose communities are over- or under-represented in the line list and how to adjust the response accordingly. The COVID-19 pandemic made this question impossible to ignore in Canada: differential burdens by neighbourhood income, racialized status, and Indigenous identity were visible in surveillance data once collected, and absent when not.
The romaine investigation compared outbreak cases with earlier sporadic cases because a list of healthy people from the same population could not be assembled quickly. The same problem arises whenever a population at risk cannot be counted in full: the investigation has to draw a sample. Section 3 sets out how samples are drawn, and Sections 4 and 5 explain how they are analysed and how large they need to be.
Key Takeaways
- A cluster is a grouping of cases that may or may not be unusual; an outbreak is a cluster that the public-health authority judges to exceed the expected baseline; a pandemic is defined by geographic spread, independent of severity.
- The CDC 10-step framework runs from preparation and confirming the outbreak, through case definition, case finding, and descriptive epidemiology, to hypothesis testing, control, and communication; several steps run in parallel.
- Under FIORP, PHAC leads epidemiology and surveillance, CFIA leads food-safety investigation, traceback, and recalls, and Health Canada leads risk assessment.
- In the 2017–18 romaine outbreak, whole-genome sequencing made a scattered set of cases visible as one cluster, a case-case analysis implicated romaine lettuce, and the public advisory preceded confirmation of a source.
- Attack rates and risk ratios from a line list identify suspect foods, and stratification separates foods that were eaten together.
Reflection
Suppose a line-list analysis of a potluck outbreak attended by 200 guests produces a ranked attack-rate table with two foods at the top: chicken (attack rate 35% among guests who ate it and 8% among guests who did not, risk ratio about 4.6) and the sandwiches (attack rate 29% and 16%, risk ratio about 1.8). The attack rate is the proportion of people who ate a food who became ill, and the risk ratio is the attack rate among those who ate the food divided by the attack rate among those who did not. Fifty-seven guests ate both foods, so the two exposures overlap. What additional analytic step would you take to disentangle the two foods? (Consider stratifying on one food while looking at the other, for example by comparing chicken eaters with non-eaters separately among guests who ate sandwiches and among guests who did not. Lesson 7 formalizes this technique as confounding control.) Briefly describe what evidence would show which of the two foods is the source.
Minimum 20 characters required.
Question 1: A working case definition for an outbreak typically includes:
Question 2: In the 2017–18 romaine lettuce E. coli O157:H7 outbreak, what was the role of whole-genome sequencing?
Question 3: An outbreak attack-rate table shows that 60% of attendees who ate chicken got ill, vs 15% of those who did not. The crude risk ratio is approximately:
Drawing a Sample
⏱ Estimated reading time: 35 minutes
Introduction and Overview
When a population cannot be measured in full, a study draws a sample, and the way the sample is drawn determines much of what the results can show. This section distinguishes a census from a sample, describes the nested populations that a study touches and the sampling frame that links them, and works through the probability sampling designs. It ends with the non-probability designs used when no sampling frame exists.
Learning Objectives
- Distinguish a census from a sample and the sources of error that each involves.
- Describe the target, source, and study populations and the sampling frame that links the source population to the sample.
- Define a probability sample and describe simple random, systematic, stratified, cluster, multistage, and targeted sampling, with the advantages and limitations of each.
- Describe judgement, convenience, purposive, and chain-referral samples, and explain why non-probability samples are generally unsuitable for estimating prevalence.
Census vs. Sample
When we conduct research, we need data from either all individuals in a population or a subset of them. The process of obtaining this data is called measurement.
In a census, every individual in the population is evaluated. In a sample, data are collected from only a subset. Sampling is generally more convenient and less costly than conducting a full census. Interestingly, even a census can be viewed as a kind of sample: it captures the population at one point in time, making it a "sample" of the population over time.
Key Distinction
In a census, the only source of error is the measurement itself. With a sample, you contend with both measurement error and sampling error. However, a well-planned sample can provide virtually the same information as a census at a fraction of the cost.
Canadian Examples: Census vs. National Health Surveys
Canada runs both kinds of data collection at the population scale, and you will encounter all of them in public health practice:
- Census of Population (Statistics Canada, every 5 years; most recent 2021). A near-complete enumeration of every household in Canada. The short-form census goes to all households; the long-form census goes to a 25% mandatory sample. Provides the denominators behind almost every population health rate you will calculate.
- Canadian Community Health Survey (CCHS): Statistics Canada / Health Canada / PHAC. A continuous cross-sectional sample survey (~65,000 respondents per cycle) covering self-reported health, behaviours, and health-care use. The flagship descriptive survey for population health surveillance.
- Canadian Health Measures Survey (CHMS): Statistics Canada / Health Canada / PHAC. A multi-stage sample survey that adds direct physical measurements (blood pressure, biomarkers, fitness) and a biobank to self-report data. Smaller (~5,700 respondents per cycle) but anchors objective measurement of population health.
- National Population Health Survey (NPHS): the longitudinal predecessor (1994–2011) to the CCHS, still used for life-course research.
The Census gives you a denominator and demographic context; CCHS and CHMS give you population estimates of health states with sampling error attached. Choosing among them is the first applied sampling decision a public-health analyst makes.
Hierarchy of Populations
Understanding the different populations involved in a study is essential for evaluating validity. There are three key populations to consider, each nested inside the next. The diagram and accordion below define them in turn; the labels that appear next to the arrows (external validity, sampling frame, internal validity) are the technical vocabulary you will use to talk about how a sample inherits or loses information from the broader population it was meant to represent.
Figure 2.1. Hierarchy of populations in epidemiologic research. The target population is the broadest; the source population is the accessible subset; the study sample consists of those who actually participate.
The target population is the population to which you want to extrapolate your results. It is often not clearly defined and may vary depending on the perspective of the person interpreting the study. For example, researchers studying rainwater cisterns in Pernambuco State, Brazil might define the target as that state, while someone else may want to generalize the findings to all semi-arid regions of Brazil.
The source population is the population from which study subjects are actually drawn. All units in the source population should be "listable" and have a non-zero probability of being included in the study. For example, in a diarrhea study in Brazil, the source population included families from households participating in the One Million Cisterns Project (OMCP).
The study sample (or study group) consists of the individuals who actually end up in the study. It is typically a subset drawn from the source population. Researchers determine the necessary sample size, draw their sample, collect data from eligible subjects, and the final study sample consists of those who agreed to participate and whose data met quality requirements.
Two kinds of validity attach to this hierarchy. Internal validity concerns whether the results are correct for the source population, and external validity concerns whether they can be generalized from the source population to the target population. HSCI 230 introduced both ideas, and Lesson 7 of this course returns to them with the selection, information, and confounding biases that threaten internal validity.
The Sampling Frame
The sampling frame is the list of all sampling units in the source population. Sampling units are the basic elements that will be sampled (e.g., households, individuals). A complete list of all sampling units is required for drawing a simple random sample, though some other methods do not require such a complete listing.
Example: Brazil Diarrhea Study
In a study of water cisterns and diarrhea in Brazil, a suitable sampling frame was the list of all households eligible for the One Million Cisterns Project. Once households were selected, a separate strategy was used for selecting individuals within each household.
Canadian Examples: Sampling Frames You Will Actually Use
National-scale Canadian surveys rarely have a single tidy list of every person. Instead, they assemble a frame from administrative listings:
- Statistics Canada Address Register (AR): the dwelling-level frame used by the Census and many StatCan household surveys.
- Labour Force Survey (LFS) area frame: CCHS draws part of its sample from the LFS area frame (which is itself based on Census enumeration areas).
- Provincial health insurance registries (e.g., the BC Medical Services Plan client registry), close to a population census of residents and the backbone of administrative-data research at Population Data BC (PopData BC).
- Disease registries such as the Canadian Cancer Registry or provincial reportable-disease lists serve as case frames for surveillance.
Notice how each frame has different coverage: a registry-based frame misses people without provincial coverage; an LFS area frame excludes residents on First Nations reserves and in institutions. The frame, not the questionnaire, is usually where exclusions and selection bias enter.
What Is a Probability Sample?
A probability sample is one in which every element in the population has a known, non-zero probability of being included. This implies that a formal process of random selection has been applied to the sampling frame. The key advantage is that probability samples allow for valid statistical inferences about the source population.
Random ≠ Haphazard
Random selection uses a formal, reproducible process (e.g., computer-generated random numbers, random number tables); it is not the same as selecting participants haphazardly or arbitrarily.
Types of Probability Sampling
Simple Random Sample
In a simple random sample, every study subject in the source population has an equal probability of being included. A complete list of the source population is required, and a formal random process is used to select individuals.
Example: To study wait times in a hospital emergency room, you need 1,000 records from 13,000 admissions over the past year. You randomly generate 1,000 numbers between 1 and 13,000 and pull those records.
Advantage: Conceptually simple; all standard statistical analyses apply directly.
Limitation: Requires a complete list of the entire source population.
Systematic Random Sample
In a systematic random sample, a complete list is not required; you only need an estimate of the total population and sequential access to individuals. The sampling interval (j) is computed as the population size divided by the desired sample size.
How it works: Randomly pick a starting point between 1 and j, then select every jth subject after that.
Example: To sample 1,000 from 13,000 emergency patients, the sampling interval is 13. Randomly pick a number between 1 and 13 for your starting patient, then select every 13th patient thereafter. So if your random start happens to be 7, you would sample patients 7, 20, 33, 46, and so on down the list.
Caution: Bias may occur if the factor you are studying is related to the sampling interval (e.g., periodic patterns in admissions).
Stratified Random Sample
The population is divided into mutually exclusive strata based on factors likely to affect the outcome. Then, within each stratum, a simple or systematic random sample is chosen. The mathematical foundations of stratification, including the now-standard optimum (Neyman) allocation rule for assigning sample size across strata, were laid out by Neyman (1934) in his landmark Royal Statistical Society paper that effectively founded probability sampling theory.
In proportional stratified sampling, the number sampled from each stratum is proportional to that stratum's share of the total population.
Three key advantages:
- Ensures all strata are represented in the sample.
- Can produce more precise overall estimates than a simple random sample because between-strata variation is removed.
- Allows estimation of stratum-specific outcomes.
Example: If hospital wait times differ between males and females, stratify records by sex and randomly sample within each group.
Cluster Sampling
A cluster is a natural grouping of study subjects with one or more common characteristics (e.g., a household is a cluster of people; a classroom is a cluster of students; a clinic is a cluster of patients).
In cluster sampling, the primary sampling unit (PSU) is the cluster itself, and it is often larger than the unit of concern. Every individual within a selected cluster is included in the sample.
Example: To estimate smoking prevalence among Grade 12 students, randomly select 10 of 47 Grade 12 classes and survey all students in those 10 classes.
Advantage: Easier when getting a list of clusters is simpler than listing all individuals. Often cheaper to visit fewer locations.
Limitation: Individuals within a cluster tend to be more alike, increasing sampling variation for a given sample size compared to SRS.
Important: A sample is only a "cluster sample" if the group is the sampling unit and the individuals within it are the unit of concern. If the group itself is the unit of concern (e.g., "does anyone in the household smoke indoors?"), it is not a cluster sample.
Multistage Sampling
Multistage sampling is similar to cluster sampling, except that after selecting primary sampling units (PSUs), a sample of secondary sampling units (individuals) is drawn within each PSU rather than surveying everyone.
Example: To study smoking among students, first randomly select 10 classes (PSUs), then randomly select 5 students from each class rather than surveying all students in every class. Within-household selection in face-to-face surveys is most often done using a Kish grid, the objective respondent-selection procedure introduced by Kish (1949).
To ensure all individuals have the same probability of being selected, either choose PSUs with probability proportional to their size (PPS) and take a fixed number of individuals from each, or choose PSUs with equal probability and use a constant sampling proportion within each PSU.
The number of individuals per cluster (ni) can be optimized by balancing within-cluster and between-cluster variance against the costs of sampling groups versus individuals.
Targeted (Risk-Based) Sampling
Targeted sampling stratifies the source population based on characteristics associated with the probability of disease occurrence, then focuses sampling on strata where disease is most likely to be found.
Individuals are assigned point values based on their probability of having the disease of interest, and sampling proceeds until a predetermined number of points have been sampled. This is an unequal probability sampling strategy; some individuals may even have a zero probability of inclusion.
Advantage: Requires a much smaller sample to detect rare diseases when key risk characteristics can be identified.
Limitation: Key epidemiological parameters (e.g., risk ratios) may not be known for the study population and must be estimated from other evidence.
Comparison of Sampling Methods
| Method | Requires Complete List? | Key Advantage | Key Limitation |
|---|---|---|---|
| Simple Random | Yes | Simple; all standard analyses apply | Needs complete population list |
| Systematic | No (needs estimate) | Practical; easy to implement | Periodic bias if factor linked to interval |
| Stratified | Yes (within strata) | More precise; ensures representation | Needs to know stratum membership |
| Cluster | List of clusters only | Cheaper; no need to list individuals | Higher variance than SRS for same n |
| Multistage | List of PSUs only | Flexible; cost-effective | Complex design; needs more subjects |
| Targeted | No (risk-based) | Efficient for rare diseases | Needs prior knowledge of risk factors |
Worked Example: How the CCHS Combines These Methods
The Canadian Community Health Survey illustrates a real multistage probability design in action:
- Stratification: the population is first stratified by health region (about 110 health regions across Canada), and a target sample size is allocated to each so that every region produces stable estimates.
- Clustering: within each health region, dwellings are sampled from the LFS area frame (groups of dwellings that share a geographic boundary). This is the cluster stage.
- Selection within cluster: one person is randomly selected from each chosen household to complete the interview.
- Top-up samples: an RDD (random digit dialling) telephone frame fills in coverage for areas where the area frame is sparse.
The result is a probability sample where every Canadian resident has a known, non-zero chance of selection, but the selection probability differs by region, household size, and frame. That is why CCHS data must be analysed with survey weights and bootstrap replicate weights (covered in Section 4).
Probability designs require a sampling frame. When no frame exists, as for many hidden or hard-to-reach populations, investigators turn to non-probability designs, with predictable consequences for inference.
Non-Probability Sampling
Samples drawn without an explicit method for determining each individual's probability of selection are known as non-probability samples. Whenever there is no formal process for random selection, the sample should be considered non-probability. Sample selection that is unrelated to the outcome of interest leaves inference intact, but selection that depends on unmeasured determinants of the outcome produces selection bias, a form of specification error formalised by Heckman (1979) and reviewed for hidden populations by Sudman & Kalton (1986). There are three main types:
Click each card to learn more:
SampleClick to learn more
SampleClick to learn more
SampleClick to learn more
Important Limitation
Non-probability samples are generally inappropriate for descriptive studies because you cannot generalize prevalence estimates to the source population without knowing each individual's probability of being included. However, non-probability methods are commonly used in analytical studies where comparing exposure groups is the priority.
Chain-Referral and Hybrid Designs
Snowball sampling, first formalised by Goodman (1961), recruits hidden-population members through peer referrals and is widely used when no sampling frame exists. Two newer hybrids partially recover probability-style inference: respondent-driven sampling (RDS), introduced by Heckathorn (1997) and extended with unbiased estimators by Salganik & Heckathorn (2004); and time-location (venue-based) sampling, applied at national scale for HIV behavioural surveillance by MacKellar and colleagues (2007). Magnani, Sabin, Saidel, & Heckathorn (2005) review when each design is appropriate for hard-to-reach populations.
Key Takeaways
- A census measures every member of a population; a sample measures a subset and adds sampling error to measurement error.
- The target, source, and study populations form a hierarchy, and the sampling frame links the source population to the sample; people missing from the frame cannot be selected.
- A probability sample gives every unit in the frame a known, non-zero chance of selection through a formal random process.
- Stratification improves precision and guarantees representation; cluster and multistage designs are cheaper but less precise for the same number of people; targeted sampling is efficient for rare outcomes.
- Judgement, convenience, purposive, and chain-referral samples do not allow each participant's selection probability to be known, so they are generally unsuitable for estimating prevalence.
Reflection
Think of a health research question you are interested in. The probability sampling designs are simple random sampling (every unit in the sampling frame has an equal chance of selection), systematic sampling (every k-th unit from a random start), stratified random sampling (the population is divided into strata and a random sample is drawn within each), cluster sampling (naturally occurring groups such as schools or clinics are randomly selected and every unit within a selected cluster is included), and multistage sampling (clusters are selected first, then units are sampled within them). Which sampling design would be most appropriate for your question, and why? What practical constraints (cost, time, the availability of a complete list of units) would influence your choice?
Minimum 20 characters required.
1. What defines a probability sample?
2. In cluster sampling, why is sampling variation typically greater than in simple random sampling for the same sample size?
3. Why are non-probability samples generally inappropriate for descriptive studies?
✦ Complete the reflection and pass the knowledge check with 100% to continue
Sampling Distributions and Survey Analysis
⏱ Estimated reading time: 40 minutes
Introduction and Overview
A probability sample makes it possible to attach a measure of uncertainty to every estimate, and this section explains why. It introduces random variables, expected values, and variances, surveys the probability distributions that describe most public-health data, and states the central limit theorem, which describes how sample means behave. It then shows how the stratification, unequal selection probabilities, and clustering of a complex survey enter the analysis, and how the design effect summarizes their combined effect on precision.
Learning Objectives
- Define a random variable, its expected value, and its variance, and explain why spread controls the precision of an estimate.
- Recognize common probability distributions (Bernoulli, Binomial, Poisson, Uniform, Normal, Exponential, log-normal) and the public-health phenomena they describe.
- Explain the sampling distribution of the mean, calculate a standard error, and state the central limit theorem in plain language.
- Explain how stratification, sampling weights, and clustering affect the analysis of survey data, and interpret a design effect and the finite population correction.
The designs in Section 3 decide who can be selected. A further question is why a small random sample can describe a whole population at all, and the answer comes from probability theory.
Probability Theory: Why Sampling Works
Every quantity we estimate from a sample, whether a prevalence, a mean, or a risk ratio, is the value of a random variable. If we drew a different sample tomorrow, the value would be slightly different. Probability theory is the formal language we use to describe how those values vary, and it is what lets us turn a single sample into a defensible statement about a population.
Three concepts do most of the work in introductory biostatistics:
Three Foundational Ideas
1. A random variable is a numerical outcome of a random process. For example, the number of new TB cases reported in a health region next week, or the systolic blood pressure of the next adult who walks into a clinic.
2. The expected value (also called the mean, written μ or E[X]) is the long-run average of a random variable across many repetitions of the random process. It is the value we are usually trying to estimate.
3. The variance (σ2) and its square root, the standard deviation (σ), measure how spread out the random variable is around its mean. Spread, not the mean alone, is what controls how precisely we can estimate things from a sample.
A second idea worth naming is independence. Two observations are independent when knowing the value of one tells you nothing about the value of the other. Independence is the assumption that lets a small random sample stand in for a much larger population, and it is the assumption most often violated in real public-health data (clustered households, repeated measures on the same person, contagion in infectious disease).
Why You Should Care
Whenever you compute a confidence interval, run a hypothesis test, or quote a margin of error, you are doing arithmetic on a probability distribution, usually a Normal distribution, that describes what would happen if you repeated the study many times. If you don’t know what distribution your statistic comes from, you can’t honestly attach uncertainty to it.
Types of Probability Distributions
A probability distribution describes the values a random variable can take and how likely each value is. Distributions split first into discrete (countable outcomes such as 0, 1, 2 cases) and continuous (any value on a range, such as height or blood pressure, time). The handful below are the ones you will encounter again and again in public health.
Discrete distributions
| Distribution | What it models | Public-health example |
|---|---|---|
| Bernoulli(p) | A single yes/no trial with success probability p. Mean = p, variance = p(1−p). | Whether one randomly chosen adult currently smokes. |
| Binomial(n, p) | Number of "successes" in n independent Bernoulli trials. Mean = np, variance = np(1−p). | Number of smokers in a CCHS sample of 1,000 adults. |
| Poisson(λ) | Number of rare events in a fixed interval of time, area, or person-time. Mean = variance = λ. | New cases of measles per week in a public-health unit; ER visits per hour. |
Continuous distributions
| Distribution | What it models | Public-health example |
|---|---|---|
| Uniform(a, b) | Every value between a and b is equally likely. The "default ignorance" distribution. | Random-digit dialling within an area code; random selection from a list. |
| Normal(μ, σ2) | The classic bell curve: symmetric, with most mass within 2σ of the mean. The default for many continuous biological measurements and, importantly, for sample means (see CLT below). | Adult height; systolic blood pressure; standardized test scores. |
| Exponential(λ) | Time between independent events occurring at a constant rate λ. Right-skewed, memoryless. Mean = 1/λ. | Time between successive ED arrivals; survival time under a constant hazard. |
| Log-normal / Right-skewed | Variables that are positive and span several orders of magnitude. The log of the variable is approximately Normal. | Household income; hospital length of stay; viral loads. |
How to Read a Distribution
Every distribution is summarized by two things: a shape (symmetric? right-skewed? bimodal?) and a small set of parameters that control its location and spread. When you see "BMI ~ Normal(27, 42)", read it as: BMI is approximately Normal, centred at 27 kg/m2, with a standard deviation of 4. About 95% of the population falls within 2 SDs of the mean, roughly 19 to 35.
Figure 2.2. Four common distributions you will meet in public-health data. Discrete distributions assign probability to whole-number outcomes (left two); continuous distributions describe a smooth curve over a real-valued measurement (right two).
🔥 Try it Yourself: Distribution Simulator
What you'll do: use the simulator below to play with each distribution's parameters and watch the shape change in real time, then run the Central Limit Theorem demo to see why sample means from any population eventually look Normal. What to take away: distributions describe specific public-health phenomena rather than serving as textbook abstractions, and the CLT is what lets us trust confidence intervals and power calculations even when the underlying data are skewed.
Aim for at least 10–15 minutes of play; this is the kind of intuition you will draw on for every confidence interval and power calculation in the rest of the course. Use the tabs to switch between exploring single distributions, the CLT demonstration, and a side-by-side comparison.
How to use this
Choose any of the six common distributions on the left. Adjust its parameters to see how the shape, mean, and spread respond. Then click "Draw a sample" to take a random draw and watch the empirical histogram converge on the theoretical curve as your sample grows.
Distribution
Parameters
Sampling
About Binomial: n = 20, p = 0.30
Binomial(n, p) counts the number of successes in n independent Bernoulli trials with the same success probability p.
Mean = np = 6.00, SD = √np(1−p) = 2.05.
- For large n and moderate p, the binomial looks Normal; this is one of the oldest examples of the CLT.
- Used for sample-size formulas around proportions and prevalence estimates.
- Assumes independence and constant p; both can fail in clustered data (households, schools).
Central Limit Theorem demonstration
Pick a population shape (the heavily skewed ones show the effect most clearly). Set a sample size n, then click Draw 1,000 sample means. Each sample of size n is summarised by its mean, and we plot those means. Watch the histogram of means converge on a Normal curve as n grows, even when the underlying population is wildly non-Normal.
Population shape
Sample size
Side-by-side: shapes at a glance
Six distributions, plotted on the same axes. Use this view to remember which shape goes with which name, and to compare how parameters change appearance. Hover for tooltips.
Display options
How to choose a distribution in practice
- Binary outcome (yes/no for one person)? Bernoulli. Count of "yes"s in a fixed sample? Binomial.
- Counting rare events in time/space (cases per week, ER visits per hour)? Poisson.
- Continuous, symmetric biological measurement (BP, height, lab values)? Normal, or check first whether the variable is approximately Normal.
- Time until next event under a constant hazard? Exponential. Survival under varying hazard? Weibull or Gamma (advanced).
- Strictly positive, right-skewed, multiplicative process (income, length of stay, viral load)? Log-normal; analyze on the log scale.
- No prior information about likely values? Uniform on a sensible range.
Each distribution above describes the underlying behaviour of a single random variable. But when we draw a sample from a population, the quantity we typically care about is a summary of that sample: a mean, proportion, or rate. To make inferences from such summaries, we need one more layer of theory.
Sampling Distributions and the Central Limit Theorem
So far we have talked about distributions of individuals in a population. Now consider the distribution of a statistic, for example the mean of a random sample of n people. Because each sample of size n would give a slightly different mean, the mean itself has a distribution. We call it the sampling distribution of the mean.
Two facts about that sampling distribution drive almost all of frequentist inference:
The Two Pillars
1. Standard error. If individual observations have standard deviation σ, then sample means of size n have standard deviation σ/√n. This quantity, the SD of a statistic, is called the standard error (SE). Quadrupling your sample size halves the standard error. Keep the two ideas apart: the standard deviation tells you how much individual people differ from one another, while the standard error tells you how much a summary such as the mean would bounce around if you repeated the whole study.
2. The Central Limit Theorem (CLT). For sufficiently large n, the sampling distribution of the mean is approximately Normal, no matter what the underlying population looks like, even if the population is heavily skewed, bimodal, or discrete. "Sufficiently large" is often around n = 30 for moderately skewed distributions, and much smaller for symmetric ones.
The CLT is the hidden engine behind the bell curve that shows up everywhere in statistics. It is the reason a 95% confidence interval can be written as estimate ± 1.96 · SE: the 1.96 comes from the Normal distribution that the CLT promises us applies to the estimate, even when we have no idea what shape the underlying population has. The theorem traces from de Moivre (1733) through Laplace's 1812 binomial approximation to Lyapunov's 1901 general proof (Central limit theorem, Wikipedia).
Figure 2.3. The Central Limit Theorem in action. The population can be wildly skewed, but as n grows the sampling distribution of the mean becomes increasingly Normal and increasingly narrow (its SD = σ/√n).
Caveat: The CLT Is About Means, Not Individuals
A common student misconception is that “a large enough sample makes the data Normal.” It does not. Income data with 1,000 observations is still right-skewed. What becomes Normal is the distribution of the sample mean across hypothetical repeated samples; that is what we use to construct confidence intervals around the mean. Inference for medians, proportions, ratios, or extreme values relies on the CLT or its analogues in different ways and may need larger n or different methods (bootstrapping, exact methods).
The results above assume a simple random sample, in which every observation is independent and has the same chance of selection. Real surveys such as the CCHS combine stratification, unequal selection probabilities, and clustering, and each of these features changes the analysis.
Analysing Complex Survey Data
When data come from a complex sampling design (involving stratification, weighting, or clustering), the analysis must account for these features. Ignoring them can lead to incorrect point estimates and underestimated standard errors.
Accounting for Stratification
If the population was divided into strata before sampling, this must be reflected in the analysis. Stratification provides stratum-specific estimates and can reduce the standard error of the overall estimate if the stratifying variable is related to the outcome.
However, stratification alone does not change the overall point estimate; it primarily affects precision. The total population size in each stratum must be known to compute appropriate sampling weights.
Sampling Weights
Not all individuals in a probability sample necessarily have the same probability of selection. The sampling weight for each individual is the inverse of their overall selection probability; this inverse-probability weighting underlies the Horvitz-Thompson estimator introduced by Horvitz & Thompson (1952), which produces unbiased totals and means from any probability sample with known inclusion probabilities.
The probability of selection depends on multiple stages. For example, in a household survey:
p(selection) = (n/N) × (m/M)
where n = households in sample, N = households in source population, m = individuals selected per household, and M = total people in that household.
Multiplying the two stages captures a simple idea: your overall chance of being sampled is the chance your household is chosen times the chance you are then picked within it.
The sampling weight = 1/p(selection). This weight reflects how many people in the source population each sampled individual "represents." Incorporating weights may change both the point estimate and the standard error.
Accounting for Clustering
In cluster and multistage sampling, individuals within groups are usually more alike than randomly chosen individuals. This means observations are not independent, and standard errors must be adjusted upward.
The most common approach is to identify the primary sampling unit (PSU) and adjust all standard error calculations for clustering at that level. The technique called variance linearisation is widely used for this purpose and requires a large number of PSUs to be reliable.
The Design Effect (deff)
The design effect (deff) summarizes the overall impact of the sampling plan on precision. It is the ratio of the variance from the complex sampling design to the variance that would have been obtained from a simple random sample of the same size. The concept and the term were coined by Leslie Kish in his classic textbook Survey Sampling (Kish, 1965) and remain the standard summary statistic for complex-design efficiency.
Interpreting the Design Effect
A deff > 1 means the complex design produces less precise (larger variance) estimates than a simple random sample would. For example, in the Brazil diarrhea study, the deff was 4.43, meaning the variance of the incidence estimate was 4.43 times larger than what a simple random sample of the same size would have produced.
Example: Impact of Survey Design on Estimates
| Type of Analysis | Incidence Estimate | SE |
|---|---|---|
| Simple random sample (assumed) | 0.1462 | 0.0061 |
| + Stratification | 0.1462 | 0.0059 |
| + Stratification + Weights | 0.1751 | 0.0091 |
| + Clustering | 0.1462 | 0.0088 |
| All features combined | 0.1751 | 0.0128 |
Notice how incorporating all features of the sampling plan changes both the point estimate (from 14.62% to 17.51%) and dramatically increases the standard error (from 0.0061 to 0.0128). Ignoring the sampling design would give a misleadingly precise, and potentially incorrect, result.
Canadian Practice: Bootstrap Weights for the CCHS and CHMS
Statistics Canada distributes the CCHS and CHMS with a set of 500 bootstrap replicate weights rather than releasing the underlying cluster identifiers (which would risk re-identification). The rescaling-bootstrap method that produces these weights was developed by Rao & Wu (1988). To get correct standard errors you re-run your analysis 500 times, once with each replicate weight, and combine the results.
Most statistical packages include survey-analysis procedures that support replicate weights and combine the 500 results automatically. If you ignore the bootstrap weights and just analyse the CCHS as if it were a simple random sample, your standard errors will typically be about 25–35% too small (more for some estimates), and your confidence intervals and p-values become meaningless.
Finite Population Correction (FPC)
When the proportion of the population sampled is relatively large (>10%), precision improves beyond what would be expected from an "infinite" population. The finite population correction adjusts the estimated variance downward:
FPC Formula
FPC = (N − n) / (N − 1)
where N is the population size and n is the sample size. The FPC should not be applied in multistage sampling even if the number of PSUs sampled exceeds 10% of the total PSUs. It is only applicable to descriptive studies using simple or stratified random sampling.
The intuition is straightforward: once you have already measured a large share of the population, little of it is left to be uncertain about, so the estimate is more precise than the standard (infinite-population) formula assumes.
Key Takeaways
- Random variables, expected values, and variances describe how sample-based estimates vary; spread, as well as location, controls how precisely a quantity can be estimated.
- A small set of distributions (Bernoulli, Binomial, Poisson, Uniform, Normal, Exponential, log-normal) describes most public-health data, and the choice among them follows the mechanism that generates the data.
- The standard error of the mean is σ/√n, and by the central limit theorem the sampling distribution of the mean is approximately Normal for sufficiently large n.
- Complex survey analyses must use the stratification, sampling weights, and clustering of the design; the design effect is the ratio of the design's variance to that of a simple random sample of the same size.
- The finite population correction reduces the variance when more than about 10% of a population is sampled by simple or stratified random sampling.
Reflection
A provincial analyst estimates mean annual household income and the prevalence of daily smoking from a Canadian Community Health Survey file of 4,000 respondents. Household income is strongly right-skewed. The survey has a stratified multistage design, and Statistics Canada supplies a final sampling weight and 500 bootstrap replicate weights for each respondent. A colleague suggests (a) that the income analysis is invalid because income is not Normally distributed, and (b) that the weights can be ignored because 4,000 is a large sample. Respond to both suggestions, drawing on the central limit theorem and on what sampling weights and bootstrap replicate weights do.
Minimum 20 characters required.
1. The Central Limit Theorem says that, for sufficiently large n:
2. If a measurement has population standard deviation σ = 12, and you draw a random sample of n = 144, what is the standard error of the sample mean?
3. Sampling weights are computed as:
✦ Complete the reflection and pass the knowledge check with 100% to continue
Planning Sample Size
⏱ Estimated reading time: 40 minutes
Introduction and Overview
The last section concerns the planning stage of a study. A study that is too small may miss a real effect, and a study that is too large wastes resources and burdens participants. This section defines the two types of statistical error and statistical power, sets out the four statistical inputs to a sample-size calculation, and provides formulae and calculators for estimating a proportion or a mean and for comparing two groups, with the adjustments for a small population, for clustering, and for non-response.
Learning Objectives
- Explain Type I and Type II errors and statistical power, and explain why a negative result from an underpowered study is uninformative.
- Describe how precision, expected variation, confidence level, and power determine the required sample size.
- Calculate sample sizes for estimating a proportion or a mean and for comparing two proportions or two means.
- Adjust a sample size for a finite population, for clustering through the design effect, and for expected non-response.
Types of Error
In any study based on a sample, the variability of the outcome, measurement error, and sample-to-sample variability all affect results. When making inferences based on sample data, they are subject to error. Within hypothesis testing in analytical studies, there are two key types of error:
Table 2.1. Types of Error
| Conclusion of Analysis | Effect Truly Present | Effect Truly Absent |
|---|---|---|
| Effect present (reject null) | Correct | Type I (α) error |
| No effect (accept null) | Type II (β) error | Correct |
A Type I error occurs when you conclude that the outcomes in the groups are different (i.e., that an association exists), when in fact they are not. In other words, you falsely reject the null hypothesis. The probability of a Type I error is denoted α.
Statistical tests are aimed at disproving the null hypothesis (that there is no difference between groups). When P ≤ 0.05, we are "reasonably sure" that any detected effect is not due to chance, but there remains a 5% chance of making a Type I error.
A Type II error occurs when you conclude that there is no association between the exposure and outcome, when in fact there is. You fail to reject the null hypothesis when you should have. The probability of a Type II error is denoted β.
Reasons a study might fail to find a real effect include: the exposure truly had no effect, the study design was inappropriate, the sample size was too small (low power), or simply bad luck.
Power is the probability that you will find a statistically significant difference when a real difference of a defined magnitude exists. Mathematically, power = 1 − β.
For example, if a study has 80% power, it has an 80% chance of detecting a true effect of the specified size. To increase power, you need to increase the sample size. So-called negative findings (failure to find a difference) are less commonly reported in the literature, partly because many studies lack adequate power.
🎲 Interactive: Sample Size & the Law of Large Numbers
What you'll do: pick a population, set a sample size n, draw repeated samples, and watch the distribution of sample means concentrate around the true population mean as n grows. What to take away: the standard error shrinks as 1/√n; this is the formal reason why “more data” means better precision and higher statistical power. The intuition you build here drives every confidence interval and sample-size calculation in the rest of the course.
Population Distribution
The "true" distribution we're sampling from. Red line = true mean (μ).
Sampling Distribution of the Mean
Each bar is the count of sample means falling in that range. Yellow = most recent sample's mean.
Power, precision, and confidence are design choices that are fixed before any data are collected. The rest of this section shows how they determine the number of people a study needs.
Sample-Size Determination
Choosing the right sample size involves both statistical and non-statistical considerations. Non-statistical factors include available resources (time, money, personnel) and the nature of the sampling frame. Statistical considerations include:
The more precise you need your estimate to be, the larger the sample you need. If you want to know diarrhea prevalence within ±5%, you need more subjects than if ±10% is acceptable. Precision is denoted L (the "allowable error" or half the desired confidence interval width).
For proportions, variance = p × q (where q = 1 − p). You need a rough estimate of the proportion to calculate the required sample size. For continuous variables like BMI, you need an estimate of the population variance (σ²). One approach: estimate the range that covers 95% of values, divide by 4 to get σ, then square it for σ².
The confidence level (typically 95%) determines how sure you want to be that the confidence interval includes the true population value. This is linked to the Z-value: for 95% confidence, Zα = 1.96. Higher confidence requires a larger sample.
In analytical studies, you also need to specify the desired power (often 80%). Power determines the sample size needed to detect a specific effect size. For 80% power, Zβ = −0.84. Greater power requires a larger sample.
Key Sample-Size Formulae
The four statistical inputs above (precision, expected variation, confidence, and power) combine into a handful of formulae, one for each planning objective. Choose the objective that matches your study, read the formula and its key, then enter your own assumptions in the calculator. Colours link every symbol to its meaning: hover over any term in the key, in the formula, or in the calculator to trace it. Teal is always the answer, orange is always Zα, and violet is always Zβ.
What the Z-values mean
Z is a position on the standard normal curve (mean 0, standard deviation 1), measured in standard deviations from the centre. Two Z-values appear in the formulae, and both are read off the same curve.
Zα is the confidence multiplier. It is the number of standard errors on either side of an estimate that captures the chosen confidence level. For a two-sided 95% confidence interval, 95% of the curve lies between −1.96 and +1.96, so Zα = 1.96. A higher confidence level means a larger Z and a larger sample.
Zβ is the power multiplier, used only when comparing groups. It is the point that cuts off the lower β of the curve, where β is the Type II error rate (1 − power). For 80% power, β = 0.20 and Zβ = −0.84. Following the textbook convention, Zβ is written with its negative sign, which is why the formulae subtract it: subtracting a negative number adds to the numerator, so more power means a larger sample.
Zα: confidence level
| Confidence | α | Zα |
|---|---|---|
| 90% | 0.10 | 1.645 |
| 95% | 0.05 | 1.96 |
| 99% | 0.01 | 2.576 |
Zβ: statistical power
| Power | β | Zβ |
|---|---|---|
| 80% | 0.20 | −0.84 |
| 90% | 0.10 | −1.28 |
| 95% | 0.05 | −1.645 |
Every formula gives a minimum. Always round n up to the next whole number, treat it as the number of people with usable data, and add a buffer (typically 10 to 25%) for non-response and drop-out.
Estimate a proportion
Objective. Report how common something is in one population (a prevalence, a proportion, a percentage) with a confidence interval no wider than you can tolerate. This is the formula for a descriptive survey with a yes/no outcome: What proportion of households have a rainwater cistern? What share of adults had diarrhea in the past month?
Before you start you need a rough guess of the proportion (from a pilot study or an earlier survey; if you have no idea, use 0.5, which gives the largest possible n), the amount of error you can accept on either side of the estimate, and a confidence level.
- nRequired sample size: the number of people you need to measure.
- ZαConfidence multiplier from the standard normal curve.1.96 for 95% confidence (1.645 for 90%, 2.576 for 99%).
- p, qp is the proportion you expect to find, as a decimal; q = 1 − p is its complement. Their product pq is the variance of a proportion, largest when p = 0.5.15% expected prevalence: p = 0.15, q = 0.85.
- LAllowable error: half the width of the confidence interval you want, in the same decimal units as p.Plus or minus 5 percentage points: L = 0.05. Halving L quadruples n.
Calculator
Estimate a mean
Objective. Estimate the average of a continuous measurement in one population (mean BMI, mean systolic blood pressure, mean days of illness) to within a chosen margin. This is the formula for a descriptive study with a measured outcome.
Before you start you need an estimate of how spread out the measurement is in the population (its standard deviation σ; if all you know is a plausible range that covers about 95% of people, divide that range by 4 to get σ), the margin of error you can accept in the same units as the measurement, and a confidence level.
- nRequired sample size: the number of people you need to measure.
- ZαConfidence multiplier from the standard normal curve.1.96 for 95% confidence (1.645 for 90%, 2.576 for 99%).
- σ²Population variance: the standard deviation of the measurement, squared. The more people vary, the more of them you need.Systolic blood pressure with SD 14 mmHg: σ² = 196.
- LAllowable error: half the width of the confidence interval you want, in the units of the measurement.A mean within ±2 mmHg: L = 2.
Calculator
Compare two proportions
Objective. Test whether the proportion with an outcome differs between two groups (exposed and unexposed, intervention and control), with enough people to detect a difference of a stated size if it is real. This is the formula for an analytic study with a yes/no outcome: a cohort study, a comparative cross-sectional survey, or a trial. The result is the number needed in each group.
Before you start you need the proportion you expect in each group (the gap between them is the smallest difference worth detecting), a confidence level (which fixes α, the false-positive rate you will accept), and a power (the probability of detecting the difference if it exists).
- nRequired sample size per group. Double it for the whole study.
- ZαConfidence multiplier, which sets α (the Type I error rate).1.96 for α = 0.05, two-sided.
- ZβPower multiplier, which sets β (the Type II error rate). Written with its negative sign, so subtracting it adds to the bracket.−0.84 for 80% power (−1.28 for 90%, −1.645 for 95%).
- p1, q1Expected proportion with the outcome in group 1, and its complement q1 = 1 − p1.Diarrhea in households without cisterns: p₁ = 0.15.
- p2, q2Expected proportion in group 2, and its complement. The difference p1 − p2 in the denominator is the effect you want to detect; smaller differences need far more people.Households with cisterns: p₂ = 0.10.
- p̄, q̄Pooled proportion, the average of the two groups: p̄ = (p1 + p2) ÷ 2, and q̄ = 1 − p̄. It describes the variability when the two groups are assumed alike (the null hypothesis).(0.15 + 0.10) ÷ 2 = 0.125.
Calculator
Compare two means
Objective. Test whether the average of a continuous measurement differs between two groups, with enough people per group to detect a difference of a stated size. This is the formula for an analytic study with a measured outcome, such as a trial comparing mean blood pressure between two arms. The result is the number needed in each group.
Before you start you need the standard deviation of the measurement (assumed to be the same in both groups), the smallest difference between the two means that would matter, a confidence level, and a power.
- nRequired sample size per group. The leading 2 is there because two groups each contribute sampling variability to the comparison.
- ZαConfidence multiplier, which sets α (the Type I error rate).1.96 for α = 0.05, two-sided.
- ZβPower multiplier, which sets β (the Type II error rate). Written with its negative sign, so (Zα − Zβ) = 1.96 + 0.84 = 2.80 at 80% power.−0.84 for 80% power (−1.28 for 90%, −1.645 for 95%).
- σ²Population variance of the measurement: the standard deviation, squared, assumed equal in both groups.SD 14 mmHg: σ² = 196.
- μ1 − μ2The difference in means you want to be able to detect, in the units of the measurement. Halving it quadruples n.A 5 mmHg difference in systolic blood pressure.
Calculator
Finite population correction (FPC) adjustment
Objective. Reduce a sample size you have already calculated when the population you are sampling from is small enough that your sample will be a sizeable share of it (more than about 10%). Once you have measured a large fraction of a finite population, little of it is left to be uncertain about, so fewer people are needed for the same precision. Use it only for descriptive estimates from simple or stratified random samples; it does not apply to multistage or cluster designs.
Before you start you need the sample size from one of the formulae above and the size of the population you will sample from (the number of units in your sampling frame).
- n′Corrected sample size. It is always smaller than n, and the smaller the population, the bigger the reduction.
- nStarting sample size: the number from the proportion or mean formula, calculated as if the population were infinite.196 from the proportion example.
- NPopulation size: how many people (or units) are in the population you are sampling from. The ratio n ÷ N is the sampling fraction.A workplace of 1,000 employees: N = 1,000.
Calculator
Clustering adjustment
Objective. Inflate a sample size you have already calculated when the design samples clusters (households, schools, villages, clinics) and takes several people from each. People in the same cluster tend to resemble one another, so each additional person within a cluster adds less new information than an independent draw. Multiplying by the design effect gives the number needed for the clustered study to reach the same precision or power as a simple random sample.
Before you start you need the sample size from one of the formulae above, an estimate of how similar people within a cluster are (the intracluster correlation, from earlier studies of the same outcome and cluster type), and the average number of people you will sample per cluster.
- n′Cluster-adjusted sample size. If the starting n was per group, this is also per group.
- nStarting sample size: the number from one of the formulae above, calculated as if every person were sampled independently.685 per group from the two-proportion example.
- ρIntracluster correlation coefficient (rho): how alike people in the same cluster are for this outcome, from 0 (no more alike than strangers) to 1 (identical). Community surveys often see 0.01 to 0.05; outcomes shared within households can be much higher.Diarrhea within households: ρ = 0.45.
- mAverage number of people sampled per cluster. The bracket 1 + ρ(m − 1) is the design effect (deff): with m = 1 there is no clustering and no inflation.Average household size: m = 6.
Calculator
Worked Example: Comparing Two Proportions
Suppose you want to determine if rainwater cisterns reduce the monthly risk of diarrhea from 15% to 10%. With 95% confidence and 80% power:
p1 = 0.15, p2 = 0.10, p = 0.125, q = 0.875
Applying the formula (with Zα = 1.96 and Zβ = −0.84) yields n = 684.8, rounded up to 685 per group, so you would need 1,370 total individuals (685 with cisterns, 685 without). Enter these values in the two-proportion calculator above to see each step.
If the outcome is clustered within households (ρ = 0.45, average household size m = 6), the clustering adjustment multiplies this by 3.25: 685 × 3.25 = 2,226.25, rounded up to 2,227 per group, more than triple the unadjusted estimate.
That multiplier, 1 + ρ(m - 1) = 3.25, is itself the design effect for this clustered design: because responses within a household are correlated, each additional person in a cluster carries less new information than an independent draw, so a larger overall sample is needed to reach the same precision.
Key Takeaways
- A Type I (α) error is a false positive and a Type II (β) error is a missed real effect; power, 1 − β, is the probability of detecting an effect of a stated size.
- Sample size depends on the desired precision or effect size, the expected variation, the confidence level, and, for comparisons, the power.
- Halving the allowable error, or halving the difference to be detected, roughly quadruples the required sample size.
- Clustering multiplies the required sample size by the design effect, 1 + ρ(m − 1); the finite population correction reduces it when a large share of a small population is sampled.
- Every calculated sample size is a minimum: it is rounded up and increased for expected non-response and drop-out.
Reflection
When observations are clustered (for example, students within schools or patients within clinics), members of the same cluster tend to resemble one another, and the required sample size must be inflated by the design effect, deff = 1 + ρ(m − 1), where ρ is the intracluster correlation and m is the average cluster size. Why is it important to account for clustering when determining sample size? What would happen to your standard errors, confidence intervals, and study conclusions if you calculated the sample size and analysed the data as though every observation were independent?
Minimum 20 characters required.
1. A Type I (α) error occurs when you:
2. Statistical power is defined as:
3. Which of the following increases the required sample size?
✦ Complete the reflection and pass the knowledge check with 100% to continue
Final Assessment
Bringing It All Together
This lesson covered two subjects that sit at the front of the course for the same reason: everything that follows depends on having data, and on knowing what those data can and cannot say about the population they came from. Sections 1 and 2 set out the institutional and methodological vocabulary of surveillance and outbreak investigation: the four system types, the Canadian product layer, the five dimensions of data quality, the CDC 10-step framework, FIORP, and a Canadian outbreak followed step by step. The duty epidemiologist who answered the telephone at the start of the lesson was reasoning in the same way as the rest of the course, under time pressure and inside a multi-jurisdictional structure of legislation, agencies, and dashboards.
Sections 3 to 5 turned to sampling, to which the outbreak workflow already pointed. When the population at risk cannot be enumerated, an investigation samples, and the questions of these sections then apply: which population is the target and which is the source, which frame was used, which design selected the participants, what sampling error and design effect the estimates carry, and how large a sample the question requires. The population hierarchy, the probability and non-probability designs, the central limit theorem, Type I and Type II error and power, and the sample-size formulae are the means by which a sample stands in for the population from which it was drawn.
The final assessment below draws on all five sections. Its questions ask for recognition of surveillance system types from descriptions, the steps and design choices of an investigation, the definition of populations and frames, the match between a sampling design and a question, and reasoning about error, weights, design effects, and sample size. Lesson 3, Questionnaire Design, takes up the next link in the design chain, and later lessons develop the measures of disease frequency and association that this lesson used in the outbreak setting.
Key Takeaways from this lesson
- Surveillance is defined by closing the loop: ongoing data collection counts as surveillance only when analysis and dissemination produce decisions, alerts, or programs.
- The four system types (passive, active, sentinel, syndromic) trade off coverage, cost, depth, and timeliness, and almost every modern surveillance product is a hybrid.
- Canadian surveillance is multi-layered and provincially anchored, and every source can be judged on five dimensions of data quality: timeliness, completeness, representativeness, sensitivity, and predictive value positive.
- Outbreak investigation follows the CDC 10-step framework within FIORP's federal and provincial division of labour, and it uses cohort, case-control, and case-case designs, attack rates, and risk ratios under time pressure.
- Sampling links a research question to feasible data collection: the target, source, and study populations and the sampling frame determine who can be selected and to whom the results apply.
- Probability designs (simple, systematic, stratified, cluster, multistage, targeted) trade off precision, cost, and feasibility, and non-probability designs are defensible only for specific purposes.
- The central limit theorem makes inference from a sample to a population possible, and complex designs require weights and a design effect in the analysis.
- Type I error, Type II error, and power are set deliberately at the design stage, and a sample-size calculation documents its assumptions about precision or effect size, variance, confidence, clustering, attrition, and population size.
Core Concepts Reviewed
Section 1: Langmuir's working definition of surveillance and the action loop; the five purposes of surveillance; the four system types (passive, active, sentinel, syndromic) with Canadian examples; the notifiable-disease reporting flow from clinician to medical health officer, province, PHAC, and WHO; the federal products (CNDSS, FluWatch, CCDSS, CVSD); the BC provincial layer (BCCDC dashboards, IRIS, Panorama); the long-running data infrastructure (vital statistics, cancer registries, DAD/NACRS, CCHS, wastewater); and the five dimensions of data quality.
Section 2: Cluster, outbreak, epidemic, and pandemic; the CDC 10-step investigation framework and epidemic-curve shapes; FIORP and its three federal partners (PHAC, CFIA, Health Canada); the 2017–18 romaine lettuce E. coli O157:H7 outbreak, whole-genome sequencing, and the case-case design; attack rates, risk ratios, and stratification in a line-list analysis; and the tension between speed and accuracy, with its equity dimension.
Section 3: Census and sample; the target, source, and study population hierarchy; the sampling frame and Canadian frames; probability sampling (simple random, systematic, stratified, cluster, multistage, targeted) and the CCHS design; and non-probability sampling (judgement, convenience, purposive, snowball, respondent-driven, and time-location sampling).
Section 4: Random variables, expected value, and variance; the Bernoulli, Binomial, Poisson, Uniform, Normal, Exponential, and log-normal distributions; the sampling distribution of the mean, the standard error, and the central limit theorem; and the analysis of complex survey data with stratification, weights, clustering, the design effect, the finite population correction, and bootstrap replicate weights.
Section 5: Type I and Type II error and statistical power; the inputs to a sample-size calculation; the formulae for estimating a proportion or a mean and for comparing two proportions or two means; and the adjustments for a finite population, for clustering, and for non-response.
The final reflection below asks for the two subjects of the lesson to be combined in a single plan. There is no single right answer; the aim is a reasoned account of how surveillance signals and sampling design fit together.
Reflection
You are the duty epidemiologist at a regional health authority. At the start of this lesson, a paediatrician called you to report four children from the same school presenting with bloody diarrhoea over two days. Suppose that stool cultures identified a notifiable bacterial enteric pathogen, the investigation traced the cases to the school cafeteria, and the outbreak has now been declared over. The regional health authority wants to know how common the same pathogen is across the whole region, in and out of outbreak settings, over the coming year. Use the following information in your answer. The surveillance signals available to you are passive notifiable-disease reports (clinicians and laboratories report confirmed cases, which flow through the province to the Canadian Notifiable Disease Surveillance System; these counts under-report because mild cases are never tested), whole-genome sequencing of isolates through PulseNet Canada (which shows whether cases belong to the same strain), and syndromic or wastewater signals (which indicate where under-detection is likely). The target population is the group to which you want to generalise; the source population is the group from which participants are actually drawn, through a sampling frame; and the study population is the people who actually provide data. The probability sampling designs are simple random, systematic, stratified, cluster, and multistage sampling. A sample-size calculation for a proportion needs an expected proportion, a margin of error, and a confidence level, and must be inflated by the design effect, deff = 1 + ρ(m − 1), when clusters of average size m are sampled and the intracluster correlation is ρ. Describe how you would move from the outbreak to a defensible sampling plan for the year-ahead question: which surveillance signals you would draw on and for what purpose, how you would define the target, source, and study populations, which sampling design you would use and why, and what values would drive your sample-size calculation (including the design effect). Name one limitation of your plan that you would acknowledge in the report.
Minimum 20 characters required.
Final Knowledge Assessment
Complete the following 15-question assessment, which draws on all five sections of the lesson. A score of 100% is required to complete the lesson. You may retake the assessment as many times as needed.
Question 1: Which of the following is the strongest argument that surveillance does analytical work beyond “data collection”?
Question 2: A regional health authority sets up a system in which paramedics flag the chief complaint of every 911 call into a real-time dashboard, with anomaly-detection algorithms triggering review when call volume for a syndrome exceeds the baseline. This is best described as:
Question 3: A surveillance dashboard reports a weekly count of laboratory-confirmed STEC infections in BC. The lag from symptom onset to appearance on the dashboard is roughly 14 days. Which dimension of surveillance data quality is most directly described by this lag?
Question 4: Which of the following best distinguishes an outbreak from a cluster?
Question 5: The shape of an epi curve where cases rise sharply, peak briefly, and fall away over a period roughly equal to one incubation period is most consistent with:
Question 6: In a closed-population outbreak (e.g., a wedding), the appropriate analytic study design is typically:
Question 7: The source population is best described as:
Question 8: A convenience sample is characterized by:
Question 9: A key advantage of stratified random sampling is:
Question 10: Which distribution would best describe the number of new measles cases reported per week in a public-health unit, where cases are rare and arrive roughly independently?
Question 11: What does a design effect (deff) of 4.43 indicate?
Question 12: An analyst estimates the prevalence of daily smoking from the CCHS and treats the file as a simple random sample, ignoring the bootstrap replicate weights. Compared with a design-based analysis, the reported 95% confidence interval will most likely be:
Question 13: A Type II (β) error occurs when:
Question 14: The clustering adjustment formula n′ = n[1 + ρ(m−1)] shows that the required sample size increases when:
Question 15: A prevalence survey needs n = 196 people under the usual infinite-population formula. The survey will be run in a workplace of N = 1,000 employees using simple random sampling. After the finite population correction, n′ = 1 ÷ (1/n + 1/N), the required sample is closest to: