HSCI 230, Lesson 5

Cohort
Studies

Evaluating Epidemiological Research

This lesson follows exposed and unexposed people forward in time. It covers the design, analysis, and reporting of cohort studies, and then randomised controlled trials and other experimental designs, in which the investigator assigns the exposure that a cohort study can only observe.

Learning objectives for this lesson:

  • Distinguish between open and closed source populations as they relate to cohort study design
  • Describe the major design features of risk-based and rate-based cohort studies
  • Identify hypotheses and population types consistent with risk-based and rate-based cohort studies
  • Elaborate the principles used to select and measure the exposure in cohort studies
  • Design and implement a valid cohort study to investigate a specific hypothesis
  • Design a randomised controlled trial that produces a valid and efficient evaluation of an intervention: state its objectives, specify the target and source populations, and describe the phases of clinical research from Phase 0 through Phase IV
  • Allocate subjects using simple, stratified, cross-over, factorial, cluster, and split-plot randomisation; distinguish single, double, and triple blinding and the bias each prevents; and compute sample size requirements, including the inflation factor for cluster randomised trials
  • Compare intention-to-treat and per-protocol analyses, define direct, indirect, and total vaccine efficacy, and apply the CONSORT 2025 reporting standards to plan and report a randomised trial

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Key Concepts & Ideas
Cohort A defined group of people followed over time. In a cohort study, the cohort is classified by exposure status at baseline (or during follow-up) and then watched for outcomes.
Closed (Fixed) Cohort A cohort with fixed membership, all of whom are followed (or attempted to be followed) for a defined window. The natural setting for risk-based analyses.
Open (Dynamic) Cohort A cohort whose membership changes over time as people enter and leave. Person-time is the appropriate denominator and incidence rates the natural measure.
Prospective Cohort A cohort study in which exposure is measured at the start, and outcomes are then watched for as time passes. Best protection against differential measurement of exposure.
Retrospective (Historical) Cohort A cohort study assembled using existing records about exposure that occurred in the past, with outcomes also already observed. Faster than prospective designs but constrained by available data quality.
Ambidirectional Cohort A cohort study that uses both retrospective and prospective elements, e.g., reconstructing past exposures from records, then continuing prospective follow-up for new outcomes.
Exposure A factor whose effect on a health outcome is being investigated. In cohort studies, exposure status is fixed (or measured longitudinally) before the outcome occurs.
Outcome The health state or event whose occurrence the cohort is being followed for. Must be defined operationally and measured consistently across exposure groups.
Person-Time The sum of time each individual is observed and at risk for the outcome, e.g., person-years. Denominator for incidence-rate calculations in open cohorts.
Loss to Follow-Up Participants who can no longer be observed before the study ends. Threatens validity if loss is differential by exposure or outcome.
Censoring When follow-up ends before the outcome is observed. Right-censoring (most common) is handled in survival analysis; informative censoring biases estimates.
Risk Ratio (RR, Relative Risk) The cumulative incidence in the exposed divided by that in the unexposed. The natural effect measure from a closed cohort.
Incidence Rate Ratio (IRR) The incidence rate in the exposed divided by the incidence rate in the unexposed. The natural effect measure from an open cohort with person-time follow-up.
Hazard Ratio (HR) A ratio of instantaneous failure rates, the effect measure produced by Cox proportional-hazards models. Often interpreted similarly to an IRR when the proportional-hazards assumption holds.
Risk Difference (Attributable Risk) Cumulative incidence in the exposed minus that in the unexposed. Captures the absolute, not relative, public-health impact of exposure.
Population Attributable Fraction (PAF) The proportion of disease in the population that would be eliminated if the exposure were removed (assuming causality). Combines effect size with how common the exposure is. HSCI 341 writes it as AFp.
Confounding A distortion of the exposure–outcome association by a confounder: a common cause of the exposure and the outcome, or a proxy for one, that is not on the causal pathway. Association with both the exposure and the outcome follows from this definition but is not sufficient, because a collider or a mediator can show both. Cohort studies handle it through restriction, stratification, and multivariable adjustment.
Healthy-Worker Effect A specific selection bias in occupational cohorts: people who are employed are systematically healthier than the general population, biasing comparisons against external referents.
Immortal Time Bias A specific bias arising when, by design, members of one exposure group cannot have the outcome during a stretch of follow-up, e.g., classifying treatment status using post-baseline information.
External Comparison Group A reference group drawn from outside the cohort (e.g., national rates) when an internal unexposed group is unavailable. Used in occupational cohorts; vulnerable to the healthy-worker effect.
Internal Comparison Group An unexposed (or differently exposed) reference group sampled from the same source population as the exposed. Generally less biased than external comparisons.
STROBE Checklist Reporting guideline (Strengthening the Reporting of Observational Studies in Epidemiology) with specific items for cohort designs, sampling, exposure measurement, follow-up, statistical methods (von Elm et al, 2007).
Clinical Equipoise Genuine uncertainty in the expert clinical community about which arm of a trial is superior; this is the ethical precondition for randomising participants to one of several treatments (Freedman, 1987).
Allocation Concealment Procedures (e.g., central randomisation, sequentially numbered opaque sealed envelopes) that prevent investigators or participants from knowing the upcoming allocation before consent and enrolment. Distinct from blinding; protects against selection bias at entry (Schulz & Grimes, 2002b).
Blinding (Masking) Concealment of allocation status from participants, providers, outcome assessors, and/or analysts after randomisation. Reduces performance bias and detection bias. Single-, double-, and triple-blind designs differ in who is masked (Schulz & Grimes, 2002c).
Intention-to-Treat (ITT) Analysis principle in which participants are analysed in the group to which they were randomised, regardless of adherence, switching, or loss to follow-up. Preserves the benefits of randomisation and provides a pragmatic effect estimate (Hernán & Hernández-Díaz, 2012).
Per-Protocol Analysis Analysis restricted to participants who adhered to the assigned intervention as specified by the protocol. Estimates an idealised efficacy but is vulnerable to selection bias.
Efficacy vs Effectiveness Efficacy: the effect of an intervention under ideal/controlled conditions (explanatory trials). Effectiveness: the effect under usual real-world conditions (pragmatic trials). The PRECIS-2 tool maps trials along this continuum across nine design domains (Loudon et al., 2015).
Placebo An inert intervention indistinguishable from the active treatment, used to control for the placebo effect, regression to the mean, and natural history of the condition.
Contamination When control-arm participants inadvertently receive the intervention (or vice versa). Dilutes the estimated effect; cluster designs reduce this risk for community-level interventions.
Compliance / Adherence The extent to which participants follow the assigned regimen. Imperfect compliance attenuates ITT estimates and motivates per-protocol or instrumental-variable analyses.
Design Effect and Intracluster Correlation Coefficient The intracluster correlation coefficient (ICC, ρ) measures how alike people in the same cluster are on the outcome. In a cluster randomised trial the required sample size is multiplied by the design effect, DEFF = 1 + (m − 1) × ICC, where m is the average cluster size, so even a small ICC inflates the sample substantially when clusters are large. The intraclass correlation coefficient used to assess reliability is a different application of a similar statistic.
Direct, Indirect, and Total Vaccine Efficacy Direct: protection of the vaccinated individual. Indirect: protection of unvaccinated people in vaccinated communities (herd effect). Total: combined direct + indirect protection within a vaccinated community. These vaccine terms describe herd effects and differ in meaning from the direct, indirect and total effects of mediation analysis.
CONSORT 2025 Consolidated Standards of Reporting Trials: the 30-item checklist and flow diagram used to standardise transparent reporting of parallel-group RCTs (Hopewell et al., 2025). CONSORT 2025 updates CONSORT 2010 (Schulz, Altman, & Moher, 2010), which had 25 items, and adds a section on open science.
Methods & Study Designs
Cohort Study An observational design that classifies individuals by exposure and follows them over time to compare outcome occurrence. The closest observational analogue to an experiment.
Risk-Based Cohort Set in a closed cohort with full follow-up. Reports cumulative incidence and the risk ratio.
Rate-Based Cohort Set in an open or dynamic cohort with person-time follow-up. Reports incidence rates and the incidence-rate ratio, well suited to long follow-up with entries, exits, and censoring.
Kaplan–Meier Estimator A nonparametric estimator of the survival function from time-to-event data, accommodating right-censoring. Visualized as the familiar “step” survival curve.
Cox Proportional-Hazards Model A semi-parametric regression for time-to-event data that estimates hazard ratios while leaving the baseline hazard unspecified. Workhorse of modern cohort analysis.
Randomised Controlled Trial (RCT) A prospective experimental study in which the investigator assigns the intervention to participants using a chance-based mechanism, then compares outcomes between groups. The reference standard for causal inference about treatment effects.
Phases of Clinical Research Pre-clinical laboratory and animal work, then Phase 0 (first-in-human, sub-therapeutic doses, very small N), Phase I (safety, dose-finding, small N), Phase II (preliminary efficacy and side-effect profile), Phase III (definitive efficacy vs comparator, large RCTs), and Phase IV (post-marketing surveillance).
Simple Randomisation Each participant has an independent probability of being allocated to each arm (analogous to a coin toss). Easy but may yield unequal arm sizes in small trials.
Block Randomisation Allocation within sequential blocks of fixed size to maintain balanced arm sizes throughout the trial. Variable block lengths help conceal sequence in unblinded settings.
Stratified Randomisation Randomisation performed separately within strata defined by prognostic factors (e.g., site, age, severity), ensuring balance on those covariates.
Cluster Randomised Trial A trial in which intact groups (clinics, schools, villages) rather than individuals are randomised. Required when the intervention is delivered at the cluster level or contamination is unavoidable; analysis must account for intracluster correlation (Campbell, Elbourne, & Altman, 2004).
Crossover Trial Each participant receives both interventions in a randomly assigned order, separated by a washout period. Each person serves as their own control. Suitable only for stable, chronic conditions with reversible outcomes.
Factorial Design Two or more interventions tested simultaneously in a single trial (e.g., 2×2). Efficient when interventions act independently; allows tests of interaction.
Split-Plot Design A trial design that applies one intervention at the cluster level (the “whole plot”) and another at the individual level within those clusters (the “split plot”), so that one trial can answer two questions. The analysis must account for the different degrees of freedom at each level.
Non-Inferiority & Equivalence Trials Trials designed to show that a new intervention is not unacceptably worse (non-inferiority) or is essentially equivalent (equivalence) to an active comparator within a pre-specified margin.
Sample Size & Power Calculation Pre-trial calculation of the number of participants needed to detect a clinically meaningful effect of a given size with specified type-I error (alpha) and power (1 - beta).
Data Safety Monitoring Board (DSMB) An independent committee that periodically reviews accumulating trial data and recommends continuation, modification, or early stopping for benefit, harm, or futility.
Quasi-Experimental Study An intervention study in which assignment is not randomised, using methods like interrupted time series, regression discontinuity, or natural experiments to approximate causal inference.
Hybrid Effectiveness-Implementation Design A design that simultaneously evaluates the clinical effectiveness of an intervention and an implementation strategy, commonly classified as Type 1, 2, or 3 depending on the relative emphasis.
Key People & Cohorts
Framingham Heart Study (1948–) A landmark prospective cohort begun in Framingham, Massachusetts that gave epidemiology the term “risk factor” (Kannel et al, 1961) and identified hypertension, smoking, cholesterol, and diabetes as major drivers of cardiovascular disease, see Dawber, Meadors, & Moore (1951) for the original design paper and Mahmood et al (2014) for a historical overview.
British Doctors Study (1951–2001) Doll & Hill's (1954) prospective cohort of UK physicians that established smoking as a cause of lung cancer. Followed for half a century with extraordinary retention; the 50-year follow-up appears in Doll et al (2004).
Nurses' Health Study (1976–) A massive prospective cohort of US nurses that has produced foundational evidence on diet, hormones, and chronic disease (Colditz, Manson, & Hankinson, 1997).
Richard Doll (1912–2005) British epidemiologist whose case-control and cohort work with Bradford Hill established the smoking–lung cancer link and modeled rigorous long-term cohort follow-up.
Austin Bradford Hill (1897–1991) British statistician and epidemiologist who designed the 1948 Medical Research Council streptomycin trial (Medical Research Council, 1948), the first published modern RCT to use formal random allocation; co-author with Doll of the British Doctors Study, and author of the “Hill viewpoints” for assessing causality from observational data (Hill, 1965).
David Cox (1924–2022) British statistician who introduced the proportional-hazards model that bears his name (Cox, 1972); one of the most cited statistical methods in cohort analysis.
Ronald A. Fisher (1890–1962) Statistician who introduced randomisation as a foundational principle of experimental design in his agricultural work at Rothamsted, providing the mathematical basis for RCT inference.
Archie Cochrane (1909–1988) Scottish epidemiologist whose advocacy for evidence-based medicine and systematic appraisal of RCT evidence inspired the Cochrane Collaboration.
No matching entries. Try a different search term.
Section 1

Introduction & Cohort Study Design

⏱ Estimated reading time: 20 minutes

Section 1 of 7

Introduction & Cohort Study Design

The logic of the design, the study-group choices, and the open versus closed population question.

The core logic

What a cohort study is

Follow disease-free subjects from exposure to outcome, then compare disease frequency between exposed and non-exposed groups.

Source population (disease-free) Exposed Non-exposed Follow forward → Compare incidence
Landmarks

The studies that defined the method

Framingham Heart Study (1948–)

Enrolled residents of one Massachusetts town; coined the term “risk factor” (Kannel et al, 1961). Three generations now under study.

British Doctors Study (1951–2001)

Doll & Hill followed United Kingdom physicians for fifty years. Established smoking as a cause of lung cancer with extraordinary retention across half a century.

Study-sample selection

Three structures

Two-cohort design

Exposure status known in advance. Recruit an exposed and a non-exposed group separately.

Single (longitudinal) cohort

Exposure unknown at recruitment. Select one group with a range of exposures, classify later.

Virtual cohort

Assembled from existing records. McCartney et al (2010) used deli purchase data to trace an E. coli O157 outbreak.

Timing

Prospective vs. retrospective

Prospective

Disease has not yet occurred. Exposure measured at baseline; investigators follow forward in time. Richer data, slower, more costly.

Retrospective

Follow-up has already ended. Both exposure and outcome reconstructed from existing records. Faster and cheaper, but constrained by data quality.

A third option, the ambidirectional cohort, reconstructs past exposure from records and then continues prospective follow-up for new outcomes.

Population structure

Open vs. closed source populations

FeatureClosedOpen
MembershipFixed at startEnters and leaves
Follow-upFull risk periodVariable per subject
Best forShort risk periodsChronic disease
MeasureRisk (cumulative incidence)Rate (incidence density)
Carry forward

Into the next section

  • Cohort studies follow disease-free subjects from exposure to outcome and compare incidence directly.
  • The design closely resembles a trial; exposure is classified, not randomised.
  • Three selection structures (two-cohort, single longitudinal, virtual) depend on what data you already have.
  • Open vs. closed population dictates whether the denominator is persons or person-time, and which measure of disease frequency is valid.

Introduction and Overview

An earlier lesson ended with a promise: cohort studies invert the case-control logic by sampling on the exposure rather than the disease, and that inversion lets us measure incidence directly without the rare-disease assumption that complicated odds-ratio interpretation. This lesson cashes that promise. Across seven content sections we walk from the basic logic of cohort design (this section), to the choice between risk-based and rate-based flavors that should now feel familiar from a later section, to the surprisingly difficult problem of measuring exposure in a longitudinal setting (a later section), to the practical questions of comparability, follow-up, outcome ascertainment, and analysis (a later section), and finally to the randomised controlled trial and other experimental designs, in which the investigator assigns the exposure (the last three sections). The unified-design discipline from an earlier lesson still applies; the lessons of an earlier lesson about pre-specified analysis plans apply with extra force, because cohort studies often run for decades and offer many opportunities for selective reporting.

Learning Objectives

  • Describe the fundamental logic of the cohort study design.
  • Distinguish between open and closed source populations.
  • Differentiate between prospective and retrospective cohort designs.
  • Recognise how cohort studies relate to controlled trials.
  • Compare cohort and case-control designs, and identify Canadian record-linkage resources that support cohort studies.

What Is a Cohort Study?

The word cohort denotes a group of study subjects that has a defined characteristic in common. In epidemiological study design, that characteristic is usually exposure status. In a cohort study, we follow subjects from exposure to outcome (Grimes and Schulz, 2002).

▸ INTERACTIVE STORY, THE TOWN THAT WAS FOLLOWED Open full screen ↗

Walk through the Framingham Heart Study from 1948 enrollment to three generations of follow-up. Next ▶ advances scenes.

A 7-scene retelling of the most famous cohort study ever launched: town enrollment (Dawber, Meadors, & Moore, 1951), baseline measurements, decades of follow-up, incidence comparisons, the birth of the term "risk factor" (Kannel et al, 1961), and the three generations still under study today (Mahmood et al, 2014).

Key Idea

A cohort study closely resembles a controlled trial, without the randomisation of exposure. We start with subjects who do not yet have the disease, classify them by exposure, follow them forward in time, and compare the frequency of the outcome between exposure groups.

Most frequently, the outcome is the occurrence of a specific disease, but cohort studies can also examine outcomes such as birth weight, body mass index, blood pressure, or quality of life. Subjects are usually individuals, but can also be groups (e.g., families).

The cohort design's modern reputation rests on a small number of landmark studies that the rest of this lesson will return to repeatedly: the Framingham Heart Study (Dawber, Meadors, & Moore, 1951), which gave epidemiology the term “risk factor” (Kannel et al, 1961); the British Doctors Study (Doll & Hill, 1954; Doll et al, 2004), which followed UK physicians for 50 years and pinned down the smoking–lung cancer link; the Whitehall II civil-servant cohort (Marmot et al, 1991), which exposed a graded socioeconomic gradient in chronic disease; the Nurses' Health Study (Colditz, Manson, & Hankinson, 1997); the multinational EPIC cohort (Riboli et al, 2002); and the recent generation of population biobanks, UK Biobank (Sudlow et al, 2015) and the Canadian Longitudinal Study on Aging (Raina et al, 2019).

Source Population (disease-free) Exposed cohort Non-exposed cohort Follow forward in time → Diseased / Non-diseased Diseased / Non-diseased Compare disease frequency

Figure, The logic of cohort design: classify disease-free subjects by exposure, follow them forward, compare the disease frequency between groups.

Cohort and Case-Control Designs Compared

An earlier lesson built the case-control design, which starts from disease status and looks back to exposure. The table below sets the two designs side by side; read it as a guide to which direction of inquiry suits a given question.

FeatureCohort StudyCase-Control Study
Direction of inquiryExposure → OutcomeOutcome → Exposure
Starting pointDefined by exposure statusDefined by disease status
Measures incidence directly?YesNo
Primary measure of associationRisk ratio, rate ratioOdds ratio
Best suited forRare exposures, multiple outcomesRare diseases, multiple exposures
Temporal sequenceClearly establishedRelies on retrospective data
Cost and timeOften expensive and lengthyGenerally less expensive and faster
Key biasesLoss to follow-up, selective attritionRecall bias, selection bias in controls

Why Cohort Studies Are Well Suited to Causal Questions

The ability to measure incidence directly is the central strength of the cohort design. It establishes temporal sequence (exposure precedes outcome), allows calculation of multiple measures of association, and can study multiple outcomes associated with a single exposure. For rare exposures, cohort studies are particularly efficient because you can intentionally over-sample exposed individuals.

Choosing the Right Design: A Practical Example

Suppose you want to study whether exposure to a specific industrial solvent increases the risk of a rare liver cancer. A prospective cohort study would require following thousands of exposed and unexposed workers for decades, which is extremely expensive and slow. A case-control study, by contrast, could identify 200 liver cancer cases from a cancer registry, select 400 matched controls, and assess past occupational exposure through interviews and employment records, producing results in months rather than years.

For rare diseases, case-control studies are usually the design of choice because they can efficiently identify enough cases to detect meaningful associations.

Selecting the Study Sample (Participants)

How we select the cohort depends on what we know in advance. The three flip cards below name the standard choices, click each one and notice that the choice flows from the data already in hand, not from any abstract preference for one design over another.

Two-Cohort Design
Click to learn more
Single (Longitudinal) Cohort
Click to learn more
Virtual Cohort
Click to learn more

In both two-cohort and single-cohort designs, after selecting subjects we (1) verify they meet inclusion criteria, (2) confirm exposure status, (3) ensure they do not yet have the outcome, then (4) follow them for a defined period and compare incidence between exposure groups.

Whichever cohort structure you pick, the next decision is whether the follow-up has already happened or whether you will be doing it as the study runs.

Prospective vs. Retrospective Designs

Cohort studies can be conducted either way, depending on whether suitable records already exist (Euser et al, 2009). The two tabs below put the trade-offs side by side.

In a prospective cohort study, the disease has not yet occurred when the study begins. Subjects are recruited, exposure is assessed at baseline, and they are followed forward in time as outcomes develop.

Advantages: Allows more detailed information-gathering and careful recording of exposure, confounders, and outcome timing (see Examples 8.6, 8.7, and 8.9).

Disadvantages: Time-consuming and expensive; vulnerable to losses to follow-up over long study periods.

In a retrospective cohort study, the follow-up period has already ended and the disease event has already occurred when subjects are selected (Hudson et al, 2005). Investigators reconstruct exposure and outcome from existing records.

Advantages: Faster and cheaper; useful when good historical records exist (Examples 8.1, 8.4, 8.5).

Disadvantages: Requires suitable existing databases; depth of information is limited to what was recorded.

Beyond the timing question is a structural one about the population itself: does its membership stay fixed for the duration of follow-up, or do people enter and leave? You met this distinction in an earlier lesson; it returns here as a more central concern, because cohort follow-up is what makes it operational.

Open vs. Closed Source Populations

The nature of the source population determines the appropriate design. This is a critical decision that affects everything from sample-size calculations to the choice of analytic methods. Read the table below as a checklist for matching disease type to design type, chronic outcomes almost always require open-population, rate-based handling, and a later section builds out exactly what that requires.

FeatureClosed PopulationOpen Population
MembershipFixed at start of studySubjects can enter and leave
Follow-upAll subjects observed for full risk periodVariable time-at-risk per subject
Best disease typeShort risk period (e.g., outbreaks)Long or chronic risk period (e.g., cancers)
Disease frequencyRisk (cumulative incidence)Rate (incidence density)
Acceptable lossesFew or none preferred (<10%)Time-at-risk accounted for explicitly

Key Examples

Three published examples bring the design choices we have just enumerated into one place. We will refer back to these throughout the lesson by number, so it is worth pausing on each one to identify which boxes the investigators ticked: prospective vs. retrospective, two-cohort vs. single, open vs. closed, risk-based vs. rate-based.

Example 8.1, Retrospective Risk-Based (Discharge Against Medical Advice) ▼

Choi et al (2011) conducted a hospital-based cohort study where the exposure was discharge against medical advice (DAMA) versus discharged with medical advice (DWMA). The outcome was readmission within 14 days. Each DAMA patient was matched with one DWMA patient by 10-year age group, gender, and clinical characteristics. Because all patients were observable for the full 14-day risk period, this is a classic risk-based design. Conditional logistic regression accounted for the matching. Result: 26% of DAMA patients were readmitted within 14 days versus only 3% of DWMA patients.

Example 8.2, Continuous-Scale Outcome (Environmental Tobacco Smoke) ▼

Crane et al (2011) conducted a retrospective cohort study based on interviews with 11,000+ women who gave birth in two Canadian provinces (2001–2009). Eleven per cent self-declared exposure to environmental tobacco smoke. Outcomes included infant body dimensions, Apgar scores, respiratory distress syndrome, and stillbirth. Multiple regression was used to control for confounders. Tobacco smoke was associated with lower birth weight, smaller body size, and increased stillbirths. Note: when outcomes are on a continuous scale (e.g., birth weight), the cohort design still applies; we just use linear rather than logistic regression.

Example 8.3, Propensity-Score Matching (Antipsychotics & Falls) ▼

Mehta et al (2010) used a population-based retrospective cohort to investigate falls and fractures in adults ≥50 years. The exposure was atypical versus typical antipsychotic agents. More than 60 covariates were combined into a propensity score, and the “Greedy 5-1 matching technique” was used to match subjects with similar scores. Each exposure group contained 5,580 people. While the hazard ratio did not differ significantly between drug classes, taking any antipsychotic for >90 days was associated with HR = 1.8 for falls or fractures.

The propensity score, the probability of exposure given the measured covariates, was introduced in Lesson 3, Section 1, as part of Rubin's design-before-data step.

Record-Linkage Infrastructure for Canadian Cohort Studies

The retrospective and virtual cohorts described above depend on records that someone else has already collected. Most epidemiological work in Canada does not involve enrolling a fresh cohort. Instead, researchers reuse data that have already been collected through health-system encounters, surveys, environmental monitoring, and registries. Three pieces of national infrastructure show up repeatedly:

Population Data BC (PopData BC)

A platform that links de-identified individual-level administrative data across the BC Ministry of Health, Vital Statistics, PharmaNet (every dispensed prescription), the BC Cancer Registry, MSP physician billings, hospital discharges (DAD), and education and social-services data. Researchers obtain a study-specific extract under a data access agreement.

Designs supported: retrospective cohorts such as those in this lesson; nested case-control and case-crossover designs (see the earlier lesson Case-Control Studies); ecological and spatial analyses (see the later lesson Ecological and Group-Level Studies); and intervention evaluations using natural experiments. Other provinces have analogous systems: ICES (Ontario), MCHP (Manitoba), HDNS (Nova Scotia), IRSPUM (Quebec).

Health Data Research Network Canada (HDRN Canada / SPOR DSP)

A federation of provincial data centres (PopData BC, ICES, MCHP, etc.) that supports multi-jurisdictional studies under a single application. Each centre runs the analysis behind its own firewall and only summary results are shared, so individual-level data never crosses provincial lines. Useful when you need national power or generalisability across health systems.

CANUE: Canadian Urban Environmental Health Research Consortium

A national repository of standardised, postal-code- and DA-level environmental exposures: air pollution (NO2, PM2.5, O3), greenness (NDVI), walkability, noise, climate, neighbourhood SES indices. CANUE indicators can be linked to any cohort with postal codes (including PopData BC extracts and CCHS shared files), turning subject-level health data into ecological or multilevel exposure–outcome studies.

Worked Example: A PopData BC + CANUE Cohort Study

To estimate the effect of long-term PM2.5 exposure on incident cardiovascular disease in BC adults, a researcher could:

  1. Define the cohort using the PopData BC Consolidation File, that is, everyone with active MSP coverage on 1 Jan 2010 (a near-complete census of BC residents).
  2. Pull baseline covariates from MSP and DAD (chronic conditions, comorbidity scores).
  3. Link each person's six-character postal code to CANUE annual PM2.5 estimates to assign exposure.
  4. Follow forward in DAD/Vital Stats to ascertain incident MI, stroke, and CV death (the outcomes).
  5. Apply Cox regression with a time-varying exposure, adjusting for area-level deprivation (also from CANUE).

This is a retrospective cohort with an environmental exposure, sitting at the boundary of cohort and ecological designs. The data come from three different stewards but no new participant was ever recruited.

Stating the Study Objective

Each study should clearly specify:

  • The target population (to which inferences will be made)
  • The source population (from which the study sample, the participants, will be drawn)
  • The unit of observation (individuals or groups)
  • The exposure, the disease, and the follow-up period
  • The setting (context or venue) of interest
  • If biology is known: the amount or duration of exposure thought to cause disease, and the relevant time window for exposure (current vs. lifetime vs. historical)

Key Takeaways

  • Cohort studies follow disease-free subjects from exposure forward to outcome.
  • The design parallels a controlled trial, minus randomisation.
  • Two-cohort designs select by exposure status; longitudinal designs select a single group with a range of exposures.
  • Studies can be prospective or retrospective; the difference is timing relative to outcome occurrence.
  • Closed populations call for risk-based designs; open populations call for rate-based designs.
  • Compared with case-control studies, cohort studies measure incidence directly and suit rare exposures and multiple outcomes; linked administrative data, such as PopData BC records joined to CANUE exposures, make large retrospective cohorts possible without recruiting anyone new.

The takeaways above name what changed conceptually compared with case-control designs. The worked example that follows makes the change concrete: because we sampled on exposure, the same kind of 2×2 table you met in earlier lessons now yields a risk ratio and an incidence rate ratio directly, no rare-disease assumption required.

Worked example: risk ratio and incidence rate ratio from cohort data

A hypothetical cohort enrols 1,000 exposed and 1,000 unexposed people who are free of the outcome and follows them for up to five years. Some participants develop the outcome or leave the study early, so the two groups contribute different amounts of person-time.

GroupPeopleNew casesPerson-yearsRisk (cases / people)Rate per 1,000 person-years
Exposed1,000804,5000.08017.78
Unexposed1,000304,9000.0306.12

The risk ratio is 0.080 / 0.030 = 2.67, so over five years exposed people had 2.67 times the cumulative risk of the outcome. The incidence rate ratio is 17.78 / 6.12 = 2.90. It is slightly larger than the risk ratio because the exposed group contributed less person-time (4,500 against 4,900 person-years), through earlier events or earlier loss to follow-up, and a smaller denominator raises the rate.

The risk ratio answers how many times more likely an exposed person is to develop the outcome over the follow-up window. The rate ratio answers how much more frequent the outcome is per unit of person-time. In an open cohort with variable follow-up, the rate-based answer is usually the appropriate one. Both measures come directly from the cohort because it was sampled on exposure; a case-control design could only approximate them through the odds ratio.

The reflection below asks you to use the timing distinction in a concrete research scenario. After working through it and the knowledge check, a later section returns to the risk-vs.-rate split that the worked example just previewed and shows what each design buys, costs, and assumes.

Reflection

Reflection

Think of a health question you find compelling. Would you address it with a prospective or retrospective cohort design? What records or recruitment infrastructure would you need? What might be lost or gained by each choice in your specific case?

Model answerFor a fast-moving exposure–outcome (e.g., antibiotic prescribing patterns and 30-day Clostridioides difficile infection), a retrospective cohort assembled from administrative data is faster, cheaper, and feasible, the records already exist, the inclusion window can be defined by ICD-coded events, and follow-up is short enough that linkage is reliable. Trade-offs: you inherit the data quality of the chart, miss any exposure not coded, and have no biomarker corroboration. For a slower-moving exposure (e.g., air pollution and dementia), a prospective cohort is the right choice: you can pre-specify the exposure assessment (personal monitors, residential geocoding), measure covariates before the outcome, and avoid recall bias, at the cost of decades of follow-up and study budget. The infrastructure you need is different too: administrative-data prospective cohorts (UK Biobank, Sudlow et al, 2015; the Canadian Longitudinal Study on Aging, Raina et al, 2019) sit between the two.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. The fundamental logic of a cohort study is to:

Correct answer: B. In a cohort study we begin with disease-free subjects classified by exposure, follow them forward in time, and compare the frequency of the outcome (usually disease) between the exposed and non-exposed groups.

2. A cohort study most closely resembles which of the following designs?

Correct answer: C. Grimes and Schulz (2002) describe cohort studies as closely resembling controlled trials, the key difference is that exposure is not randomly assigned. This similarity is often cited as an advantage for causal inference.

3. In a retrospective cohort study, when has the outcome event occurred relative to the start of the study?

Correct answer: A. Retrospective cohort studies use existing records: by the time the investigators begin, the follow-up period has ended and outcome events have already occurred. Prospective designs are the opposite, the outcome has not yet occurred at study initiation.

4. Which of the following best describes a closed source population?

Correct answer: D. A closed population has fixed membership at the start of the study and all subjects are observed for the full risk period. This makes it appropriate for risk-based (cumulative incidence) designs and works best when the risk period is short and losses are few.
Section 2

Risk-Based & Rate-Based Designs

⏱ Estimated reading time: 15 minutes

Section 2 of 7

Risk-Based & Rate-Based Designs

Two denominators, two effect measures, and the logic that chooses between them.

Risk-based design

Closed cohort, cumulative incidence

Eq 8.1 · Risk and risk ratio
\[ \color{#0B7B6B}{R_1} = \frac{\color{#C2410C}{a_1}}{\color{#1D4ED8}{n_1}} \qquad \color{#6D28D9}{R_0} = \frac{\color{#C2410C}{a_0}}{\color{#1D4ED8}{n_0}} \qquad \color{#BE185D}{\text{RR}} = \frac{\color{#0B7B6B}{R_1}}{\color{#6D28D9}{R_0}} \]
R1 risk in exposedR0 risk in unexposeda new casesn people at riskRR risk ratio

Valid only when the cohort is closed and all subjects are observable for the full risk period. Losses should be few (some authors use <10% as a threshold).

Best for: acute outbreaks, surgical complications, short-window outcomes.

Rate-based design

Open cohort, incidence density

Eq 8.2 · Rate and rate ratio
\[ \color{#0B7B6B}{I_1} = \frac{\color{#C2410C}{a_1}}{\color{#1D4ED8}{t_1}} \qquad \color{#6D28D9}{I_0} = \frac{\color{#C2410C}{a_0}}{\color{#1D4ED8}{t_0}} \qquad \color{#BE185D}{\text{IRR}} = \frac{\color{#0B7B6B}{I_1}}{\color{#6D28D9}{I_0}} \]
I1 rate in exposedI0 rate in unexposeda new casest person-timeIRR rate ratio

Each subject contributes person-time until the event, censoring, or study end. Accommodates variable follow-up and dynamic populations.

Best for: chronic disease, long follow-up, open cohorts (e.g., 10-year breast cancer study, Example 8.7).

Analytic choice

Poisson vs. Cox

Poisson regression

Person-time as the offset. Use when the incidence rate is reasonably constant over follow-up. Produces the incidence rate ratio directly.

Example: rugby injury rates over one season (Example 8.6).

Cox proportional hazards

Semi-parametric. Use when the rate changes substantially over follow-up. The workhorse of modern cohort analysis.

Example: smoking and invasive breast cancer over 10 years (Example 8.7).

Sample size

Planning the study

Initial sample-size estimates typically assume a risk-based design even when a rate-based analysis is planned. That gives a workable ballpark for early planning.

Modern software extends to: unequal group sizes, survival-time outcomes (Matsui, 2005), strata-matched designs (Mazumdar et al, 2006), and time-varying exposures (Basagana et al, 2011).

For rate-based designs, the person-time denominator can be derived from the expected rate and the follow-up window once preliminary estimates exist.

Carry forward

Into the next section

  • Risk-based: closed cohort, subjects as denominator, cumulative incidence and risk ratio.
  • Rate-based: open cohort or variable follow-up, person-time as denominator, incidence rates and rate ratio.
  • Poisson regression suits constant rates; Cox proportional hazards suits long follow-up with changing rates.
  • Risk ratio and rate ratio are distinct quantities. They converge only in a closed cohort with no censoring.

Introduction and Overview

An earlier section established what a cohort study is and named the design choices its investigators have to make. This section drills into the most consequential of those choices: whether to count events per person (risk) or events per unit of person-time (rate). The two designs share a 2×2 layout but differ in what they assume about the population and what they let you say about disease frequency. Sample-size planning, surprisingly, is a useful place to start, because the calculation is the same for both designs even when the analysis ends up being different.

Learning Objectives

  • Describe the design and assumptions of risk-based (cumulative incidence) cohort studies.
  • Describe the design and analysis of rate-based (incidence density) cohort studies.
  • Identify hypotheses and population types appropriate for each design.
  • Calculate and interpret the basic measures of disease frequency for each design.

Sample Size

Initial sample-size estimates are usually performed assuming an equal number of exposed and non-exposed subjects, and assuming the disease is measured by risk (Section 8.2.2). This approach is often sufficient for initial planning even if the population is open and a rate-based design must ultimately be used.

Modern Sample-Size Software

Recent software allows for unequal sample sizes, repeated measures, multivariable regression models, and proportional hazards models. Specialised methods exist for competing risks (Latouche and Porcher, 2007), survival-time outcomes (Matsui, 2005), strata-matched designs (Mazumdar et al, 2006), and time-varying exposures (Basagana et al, 2011).

Risk-Based (Cumulative Incidence) Designs

This is the simplest form of cohort study, but several assumptions must hold:

  • Exposure groups are defined at the start of the study and remain unchanged (fixed cohorts).
  • The study groups are closed, all subjects must be observed for the full risk period.
  • There should be few or no losses (some authors use >10% losses as a cut-point that casts doubt on validity).

When Risk-Based Designs Work Best

Risk-based designs work best for diseases with a relatively short risk period (e.g., acute infections, post-surgical complications). For chronic diseases such as many cancers, where the risk period is lifelong and often longer than feasible follow-up, a rate-based design is preferred.

2×2 Table: Risk-Based Cohort Design

ExposedNon-exposedTotal
Diseaseda1a0m1
Non-diseasedb1b0m0
Totaln1n0n

We select n1 exposed and n0 non-exposed individuals (free of disease) from the source population, follow them for the full follow-up period, and observe a1 exposed cases and a0 non-exposed cases. The two risks of interest are:

Eq 8.1
\[ \color{#0B7B6B}{R_1} = \frac{\color{#C2410C}{a_1}}{\color{#1D4ED8}{n_1}} \qquad \color{#6D28D9}{R_0} = \frac{\color{#C2410C}{a_0}}{\color{#1D4ED8}{n_0}} \]
The risk in the exposed and the risk in the unexposed are each the new cases in that group divided by the number of people at risk in it.

The Denominator

In risk-based designs, the denominator is the number of subjects in each exposure category. This is only valid because every subject is observed for the full risk period, otherwise, who you count and who you don’t would depend on follow-up time.

HSCI 341 Lesson 4, Section 2 (Incidence: Risk and Rate) develops incidence risk with formula calculators, worked examples and the assumption of a closed population.

The two risks can be compared on two scales. The risk ratio divides one risk by the other; the risk difference (also called the attributable risk) subtracts them, giving the absolute difference in incidence between the exposed and unexposed groups:

RD
\[ \color{#BE185D}{\text{RD}} = \color{#0B7B6B}{R_1} - \color{#6D28D9}{R_0} = \frac{\color{#C2410C}{a_1}}{\color{#1D4ED8}{n_1}} - \frac{\color{#C2410C}{a_0}}{\color{#1D4ED8}{n_0}} \]
The risk difference is the risk in the exposed minus the risk in the unexposed, each being the new cases divided by the number of people at risk. It reports the absolute excess risk attributable to the exposure, in the same units as the risks themselves, whereas the risk ratio reports how many times larger the exposed risk is.

In the hypothetical cohort from the worked example in an earlier section, RD = 0.080 − 0.030 = 0.050, or 50 extra cases per 1,000 exposed people over five years, alongside a risk ratio of 2.67. The absolute figure is the one that tells a health planner how many cases the exposure adds.

Risk-based designs are conceptually clean but operationally fragile, their assumptions break the moment people leave the cohort or the risk period extends beyond a few months. The rate-based alternative was developed precisely to handle the populations where those assumptions do not hold.

Rate-Based (Incidence Density) Designs

In many cohort studies, not every subject is under observation for the full risk period, especially when:

  • The source population is dynamic (subjects enter and leave).
  • The follow-up period is long.
  • Subjects are added part-way through the biological risk period.
  • A significant proportion of subjects withdraw from the study.
  • Exposure status itself changes during the study.

In these situations, we cannot just count exposed and non-exposed subjects. Instead, we accumulate the amount of ‘at-risk time’ contributed by each subject in each exposure category. The denominator becomes person-time, not persons. For example, 10 people each followed for 2 years and 20 people each followed for 1 year both contribute 20 person-years, so this denominator credits every subject for exactly as long as we actually observed them.

2×2 Table: Rate-Based Cohort Design

ExposedNon-exposedTotal
Diseaseda1a0m1
Person-time at riskt1t0T

Each subject contributes ‘at-risk’ time until they develop the disease, are lost to follow-up, or the study ends. The two rates of interest are:

Eq 8.2
\[ \color{#0B7B6B}{I_1} = \frac{\color{#C2410C}{a_1}}{\color{#1D4ED8}{t_1}} \qquad \color{#6D28D9}{I_0} = \frac{\color{#C2410C}{a_0}}{\color{#1D4ED8}{t_0}} \]
The incidence rate in the exposed and the rate in the unexposed are each the new cases divided by the person-time at risk accumulated in that group.

Choice of Analysis

If follow-up is relatively short and rates are reasonably constant, Poisson models are appropriate. If follow-up is long and the assumption of a constant rate is not tenable, survival analysis (e.g., Cox proportional hazards) is preferred (Cox, 1972; see Chapter 19).

You have now seen both designs from the inside. The next subsection puts them side by side; read it as a decision aid for matching design to research situation, not as a statement that one design is generally better than the other.

Comparing the Two Designs

Risk-based designs are best when:

  • The population is closed (fixed cohort).
  • The risk period is short (so all subjects can be observed for the full period).
  • Losses to follow-up are minimal (under ~10%).
  • Examples: acute outbreaks, surgical complications within 30 days, hospital readmissions within 14 days (Example 8.1).

Rate-based designs are best when:

  • The population is open (dynamic).
  • Follow-up is long or the risk period is chronic.
  • Subjects enter or leave the study at different times.
  • Exposure status may change during follow-up.
  • Examples: rugby injury rates over a season (Example 8.6), invasive breast cancer over 10+ years (Example 8.7), fracture incidence over decades (Example 8.8).

Risk denominator: the number of subjects in each exposure category. Counts people.

Rate denominator: the cumulative person-time at risk in each exposure category. Counts time.

This means risk is dimensionless (a proportion between 0 and 1), while a rate has units of cases per person-time (e.g., 4.0 per 1,000 person-years).

HSCI 341 Lesson 4, Section 2 develops exact and approximate person-time and the conversion between risk and rate.

Key Examples

The four examples below illustrate the design choice with real published studies. The first two are risk-based; the last two are rate-based. As you expand each, ask yourself why the investigators chose what they chose, in every case the source population's behaviour and the follow-up window's length will be doing most of the work.

Example 8.4, Risk-Based (Time-of-Day & Surgical Complications) ▼

Kelz et al (2009) compared morbidity and mortality following 56,000+ general and vascular surgical procedures (2001–2004). Time of operation was grouped into seven 2-hour periods. Risk of mortality within 30 days had a moderately strong association with start times after 9:30 pm (OR = 1.22), and morbidity had OR = 1.32 for late-night surgeries. However, when emergency cases were excluded, no odds ratios were significant. The excess crude risk was largely explained by the nature of the clinical cases, an important reminder about confounding by indication.

Example 8.5, Risk-Based (Cervical Screening in HIV-Positive Women) ▼

Leece et al (2010) followed approximately 250 HIV-positive women receiving care at the Ottawa Hospital General Campus Immunodeficiency Clinic (2002–2005). The outcome was undergoing cervical screening; predictors included demographics, HIV status, and primary care provider status. Analysis combined χ2 tests with logistic regression. The 12 women without a primary-care provider were less likely to undergo screening (RR = 1.6) than the 84 women with providers. The authors noted that abnormal screening results were common and that recent low CD4 cell count was the only significant predictor.

Example 8.6, Rate-Based (Rugby Injury Rates) ▼

Chalmers et al (2011) followed 704 male amateur rugby players (aged 13+) over a season. The ‘time’ component was a game, with a total of 6,263 player-games of follow-up. Exposures included age, ethnicity, experience, BMI, smoking, previous injury, training, weather, ground conditions, foul play, and protective equipment. Because rates were reasonably constant over the period, Poisson regression was appropriate. Notable findings: Pacific Island vs. Maori ethnicity (IR = 1.5), ≥40 hours of strenuous activity weekly (IR = 1.5), playing while injured (IR = 1.5), foul play (IR = 1.9), and headgear use (IR = 1.2).

Example 8.7, Rate-Based (Smoking & Breast Cancer) ▼

Luo et al (2011) drew on the Women’s Health Initiative Observational Study: 90,000+ women aged 50–79 followed across 40 US clinical centres. Smoking exposure was characterised in detail (status, age started, age quit, cigarettes/day, pack-years). Over an average of 10.3 years of follow-up, 3,520 incident invasive breast cancers were identified. Because of the long follow-up, Cox proportional hazards models (rather than Poisson) were used. Findings: HR = 1.09 for former smokers and HR = 1.16 for current smokers. Among lifetime non-smokers, only those with the highest passive-smoke exposure had increased risk; no significant dose-response trend was seen.

Key Takeaways

  • Risk-based designs use number of subjects as the denominator and require a closed cohort followed for the full risk period.
  • Rate-based designs use person-time as the denominator and accommodate dynamic populations and variable follow-up.
  • The choice between Poisson and Cox proportional hazards depends on whether the rate is reasonably constant or changes substantially over follow-up.
  • Initial sample-size calculations can be done assuming a risk-based design even when the analysis will ultimately be rate-based.

The reflection below asks you to apply the choice from this section to a specific long-running occupational cohort. After working through it, a later section turns to a problem that stays mostly hidden in textbook treatments: how do you actually measure exposure when it can change over years or decades of follow-up?

Reflection

Reflection

Suppose you are studying the effect of a workplace exposure (e.g., shift work) on cardiovascular disease over 20 years. Workers can join or leave the company at any time. Which design (risk-based or rate-based) is more appropriate, and why? What practical issues would arise that wouldn’t arise in a 14-day hospital readmission study?

Model answerRate-based is appropriate. Workers entering and leaving the company across a 20-y window create an open cohort where person-time, not headcount, is the natural denominator. Practical issues: (a) defining exposure time, how to handle workers who move in and out of shift work; (b) healthy-worker effect, long-tenured workers are healthier than the general working population, biasing toward null; (c) healthy-worker survivor effect, those who keep doing shift work are those who tolerate it, dynamically depleting the susceptible from the exposed group; (d) time-varying confounders like BMI or hypertension that may be affected by past shift work; (e) loss-to-follow-up when workers leave the company. None of these arise in a 14-day readmission study because the window is too short for healthy-worker dynamics to matter, the cohort stays effectively closed, and confounder profiles are fixed at index admission. Methods: standardised mortality ratios with internal comparators, g-methods for time-varying exposure, sensitivity for unmeasured occupational confounders.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. Which of the following is a key requirement of a risk-based cohort study?

Correct answer: B. Risk-based designs require a closed cohort with all subjects observed for the full risk period; without this, the risk-as-proportion calculation is not valid. Few or no losses are preferred (some authors use >10% losses as a cut-point indicating doubt about validity).

2. In a rate-based cohort study, the denominator of the rate is:

Correct answer: D. Rate-based designs use person-time as the denominator. Each subject contributes time-at-risk until they develop the disease, are lost to follow-up, or the study ends. This accommodates variable follow-up and dynamic populations.

3. Which study design is most appropriate when the source population is dynamic and follow-up is long?

Correct answer: C. Rate-based (incidence density) designs handle dynamic populations and long follow-up periods because they explicitly account for variable time-at-risk per subject. Risk-based designs require all subjects to be observed for the full risk period.

4. When follow-up is long and the assumption of a constant rate is not tenable, which analysis is preferred?

Correct answer: A. When the constant-rate assumption is suspect, survival models such as Cox proportional hazards are the preferred approach. Poisson works when rates are reasonably stable across the follow-up period (e.g., the rugby season in Example 8.6); Cox is preferred for very long follow-up like the 10-year breast cancer study (Example 8.7).
Section 3

The Exposure

⏱ Estimated reading time: 15 minutes

Section 3 of 7

The Exposure

Measurement scales, permanence, induction periods, and changing status.

Measurement scales

Four types of exposure variable

Dichotomous

Exposed vs. non-exposed. Simple, but can lose information when the true exposure is graded.

Continuous

Exact dose (mg, ppm, MET-hours). Preserves the most information; model directly when possible.

Ordinal

Ordered categories (never / former / current). Balances information with interpretability.

Compound

Combines intensity and duration. Pack-years = years smoked × cigarettes/day ÷ 20.

Permanence

Permanent vs. non-permanent exposures

Permanent exposures

Time-invariant: sex, race, one-time vaccination. Easier to measure, but operational definitions still matter. Age at exposure and threshold dose may be important.

Non-permanent exposures

Change over time: diet, lifestyle, environmental factors. Need timing, intensity, and duration data. More complexity, more credible inference.

Induction period

Time before disease can arise

Exposure occurring Induction period At-risk period Exposure complete At-risk begins

During the induction period, exposed subjects' time-at-risk is assigned to the non-exposed category, or discarded.

Changing status

Contributing to both categories

Exposed (smoking) Lag Non-exposed (quit) Quits smoking

If the outcome occurs, the subject is assigned to whichever category they occupied at that time. Lost subjects contribute time until their last known date.

Carry forward

Into the next section

  • Exposure scales range from dichotomous to compound; store data at the highest resolution available.
  • Non-permanent exposures need timing, intensity, and duration to support credible causal inference.
  • Time during the induction period should be counted as non-exposed or discarded.
  • Subjects can contribute time-at-risk to both categories when status changes; outcome category is assigned at the time of the event.
  • Disease itself can be an exposure when studying downstream outcomes.

Introduction and Overview

Earlier sections set up the architecture of the cohort study and the choice of risk-vs.-rate design. Both treated “exposed” as if it were a fixed property of each subject. In real cohorts that is rarely true. People start smoking, quit, switch jobs, change diets; doses accumulate. This section is about how exposure is actually measured and handled when it can vary across years or decades of follow-up, the technical problem that distinguishes a working cohort study from a textbook one.

Learning Objectives

  • Elaborate the principles used to select and measure exposure in cohort studies.
  • Distinguish between permanent and non-permanent exposures.
  • Describe the concept of an induction period and its role in time-at-risk calculations.
  • Identify how dichotomous, ordinal, and continuous exposure scales differ in measurement and analysis.

Why Exposure Measurement Matters

In cohort studies, the objective is to identify the consequences of a specific exposure factor. Exposures can range from study-subject characteristics (sex, age) to infectious or noxious agents, environmental exposures, or food-related factors. Exposures that can be manipulated are of special interest because they lead more directly to disease control.

Measurement Is Not Trivial

Although measuring exposure might seem simple at first glance, careful thought should be given to how it is measured and expressed. Each study should specify what constitutes exposure and, when possible, the ‘induction period’, how long after exposure is reached before disease might reasonably arise. The more complex the exposure, the more important it is to validate the assessment.

Scales of Exposure Measurement

Exposure status can be measured on different scales, each with implications for design and analysis. The four flip cards below run from the simplest binary contrast up to compound measures that combine intensity and duration. The pattern to take away: more granular measurement makes the design better suited to detecting dose-response, but only if the underlying biology really has a graded effect.

Dichotomous
Click to learn more
Ordinal
Click to learn more
Continuous
Click to learn more
Compound
Click to learn more

The choice of measurement scale is independent of a second question: does the exposure stay fixed for the rest of the subject's life, or can it change?

Permanent vs. Non-Permanent Exposures

Permanent exposures are time-invariant, factors such as sex, race, or one-time exposures such as a vaccination. These are relatively easy to measure, but a moment’s thought reveals subtleties:

  • Age at exposure may matter (e.g., age at vaccination, age at smoking initiation).
  • Even ‘one-time’ exposures may have a threshold or dosage requirement to count as ‘exposed’.
  • If the disease event occurs before exposure is completed, it should not be counted as an outcome event, the exposure could not have caused it.

In early studies, the goal may be to determine the threshold at which exposure becomes biologically meaningful (Rohan et al, 2007).

Non-permanent exposures change over time: food intake, lifestyle factors, environmental exposures, or any exposure where the timing matters. These add complexity:

  • When did exposure start? (e.g., age started smoking)
  • When did exposure stop? (e.g., age quit smoking)
  • How intense was exposure? (cigarettes per day)
  • How long did it last? (years smoked)

Sometimes a simple summary will suffice (e.g., years smoked); sometimes a compound measure (e.g., pack-years) is needed. The more information collected, the more credibly causal relationships can be inferred and the more useful the findings are for prevention.

Both permanent and non-permanent exposures share a complication that becomes visible only when you start counting person-time: the gap between when exposure happens and when disease can plausibly result.

The Induction Period & Time-at-Risk

An important concept in cohort design is the induction period, the time after exposure is completed before disease can reasonably arise.

time Exposure occurring Induction period At-risk period Life experience Exposure complete At-risk period begins (induction period over)

Figure 8.1, Life experience: exposure, induction period, and time-at-risk.

Handling the Induction Period

If there is a known induction period following completion of exposure, then until that period is over, the time-at-risk of ‘exposed’ individuals should be added to the non-exposed group. Some researchers prefer to discard disease experience during the induction period because of uncertainty about its duration; this is often the safest choice provided sufficient time-at-risk remains in the exposed group to maintain precision.

Changing Exposure Status

If exposure status changes during follow-up, an individual subject can contribute time-at-risk to both exposed and non-exposed categories:

  • Previously non-exposed subjects contribute to the exposed category after the exposure threshold is reached.
  • Previously exposed subjects contribute to the non-exposed category after any lag effects have ended.
  • If a subject develops the disease, the exposure category assigned is the one they were in at the time the outcome occurred.

Losses to Follow-Up

For subjects lost to follow-up, time-at-risk accumulates until the last date their exposure status is known. If the precise time of loss is unknown, the midpoint of the last known exposure period is conventionally used.

One last twist on the meaning of “exposure” is worth flagging before we move on, because it shows up surprisingly often in chronic-disease cohorts.

Disease as Exposure

Disease itself can serve as an exposure for other outcomes such as additional diseases, mortality, or quality-of-life measures. Lazo et al (2011) followed 11,000+ adults for 18 years using non-alcoholic fatty-liver disease (NAFLD) as the exposure. Those with NAFLD (whether or not they had elevated liver enzymes) had a similar mortality hazard ratio as those without NAFLD, an interesting null result.

Key Examples

The three published cohorts below illustrate the full range of exposure handling we just covered: a continuous diet exposure validated against multiple recalls, a multi-exposure design with biospecimens, and a multinational study that pushes from estimated relative risk to population-attributable fraction. Expand each one and notice which exposure-measurement decisions shape the rest of the design.

Example 8.8, Continuous Exposure (Calcium Intake & Fracture) ▼

Warensjo et al (2011) studied women in the Swedish Mammography Cohort. Calcium intake from diet, supplements (1 dose = 500 mg), and multivitamins (1 dose = 120 mg) was the major exposure. Total calcium intake correlated well (r = 0.77) between food-frequency questionnaire and 14 repeated 24-hour recalls, an example of careful exposure validation. Cumulative dietary calcium intake (in quintiles) was related to fracture and osteoporosis incidence using Cox proportional hazards and logistic models. Findings: a chronic, low-dietary calcium intake was associated with increased fracture and osteoporosis. Above the base level, only minor differences were observed; in the highest-intake group, hip fracture rate was somewhat increased.

Example 8.9, Multiple Exposures (Canadian Diet, Lifestyle & Health Study) ▼

Rohan et al (2007) recruited alumni from three Ontario universities (1995–1998). The major outcome was new cancer incidence. Participants completed lifestyle and food-frequency questionnaires, measured waist and hip circumferences, and provided hair and toenail specimens (for trace element and DNA analysis). Of the 73,000+ recruits, 97% provided biological specimens. Exposures included exercise, lifestyle factors, molecular markers, and dietary characteristics. The paper includes a particularly good discussion of creating compound nutritional variables from food-frequency data and verifying them.

Example 8.10, Population Attributable Fraction (Alcohol & Cancer) ▼

Schutze et al (2011) reported on the European Prospective Investigation into Cancer and Nutrition (EPIC; Riboli et al, 2002), a multicentre prospective cohort that recruited 520,000 men and women aged 35–70 from 10 European countries (1992–2000). Alcohol consumption was measured at recruitment in grams/day and classified as never, former, or lifetime consumer. Cancer incidence was identified through cancer registries (follow-up ended 2002–2005). The analysis combined regression coefficients with population prevalence of consumption above recommended levels. If causality is assumed, alcohol consumption beyond recommended levels was responsible for an estimated 10% of all cancers in men and 3% of all cancers in women.

These figures are population attributable fractions (PAF; HSCI 341 writes the measure as AFp), the share of all cases in the population that would not occur if the exposure were removed. HSCI 341 Lesson 1, Section 3 (Seeking Causes and Models of Causation) gives Levin's formula for computing it from the prevalence of exposure and the relative risk.

Key Takeaways

  • Exposure can be dichotomous, ordinal, continuous, or compound, choose the scale that best captures the biology.
  • Permanent exposures (sex, race, vaccination) are easier but still require careful operational definitions.
  • Non-permanent exposures require timing, intensity, and duration data; the more detail, the more credible the inferences.
  • Time before the induction period ends should not be counted as ‘exposed’ time-at-risk.
  • Subjects can contribute time-at-risk to multiple exposure categories if their status changes.
  • Disease itself can serve as an exposure for downstream outcomes.

The reflection below puts the measurement-scale decision into a real research scenario. After working through it and the knowledge check, a later section closes the lesson by addressing the four practical questions that any cohort investigator faces once exposure is settled: keeping the groups comparable, managing the follow-up period, ascertaining outcomes, and analysing the data.

Reflection

Reflection

Imagine designing a cohort study of physical activity and cardiovascular disease. How would you measure activity, dichotomous (active vs. inactive), ordinal (low/moderate/high), or continuous (MET-hours/week)? What might be lost or gained at each level of measurement? Consider how you would handle people whose activity changes substantially during follow-up.

Model answerContinuous (MET-hours/week) is the most informative measurement and the one to default to: it preserves dose-response information, can be analysed flexibly (linear, splines, categorical), and avoids losing power to arbitrary cut-offs. A dichotomous active/inactive split throws away most of the signal and depends on the cut-off you choose; an ordinal three-level scheme is a reasonable compromise for reporting. Whatever you record, store it continuous and categorise only for presentation. People whose activity changes substantially: model it as a time-varying exposure (each interval gets its own MET-hours value), and pre-specify how you'll handle measurement protocols (e.g., re-administered every 2 years). If reverse causation is a worry (people reduce activity because they're getting sick), use a lag (e.g., assign exposure status from 2 years before the outcome window) and run sensitivity analyses for different lag lengths.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. The induction period refers to:

Correct answer: C. The induction period is the time after exposure is completed before disease can reasonably arise from that exposure. During this period, the time-at-risk of ‘exposed’ individuals should be added to the non-exposed group, or this experience may be discarded.

2. Which exposure measurement is the example of a compound variable?

Correct answer: B. Pack-years is a compound variable that combines duration and intensity of exposure. Luo et al (2011) used this measure in their study of smoking and breast cancer (Example 8.7).

3. If a subject’s exposure status changes during the study, how is their time-at-risk handled?

Correct answer: D. When exposure status changes, the subject contributes time-at-risk to both categories. Time before the change goes to the original category; after the change (allowing for any lag effects) goes to the new category. If they develop the disease, they are assigned to whichever category they were in at the time of the outcome.

4. Why is categorising a continuous exposure variable often discouraged in analysis?

Correct answer: A. Categorising a continuous exposure variable usually results in loss of information. When possible, modelling the exposure-outcome relationship on the continuous scale (e.g., dose-response) preserves more of the underlying information.
Section 4

Comparability, Follow-up, Outcomes & Analysis

⏱ Estimated reading time: 18 minutes

Section 4 of 7

Comparability, Follow-up, Outcomes & Analysis

Making groups comparable, tracking follow-up, measuring outcomes, and analysing the data.

Comparability

Three tools for achieving exchangeability

Restriction

Applied before selection. Limit eligibility to one level of the confounder. Prevents confounding but reduces sample size and generalisability.

Matching

Applied at selection. Select non-exposed subjects with the same confounder value. Propensity scores (Example 8.3; introduced in Lesson 3) handle many confounders at once.

Analytic control

Applied during analysis. Measure confounders and adjust via stratification or regression. Flexible but only works for measured confounders.

Follow-up

Complete and unbiased

Blinded follow-up

Outcome assessors should be unaware of exposure status. In retrospective designs, record reviewers should be kept blind where possible.

Active vs. passive surveillance

Active surveillance (regular examinations) gives more accurate event timing. Passive surveillance identifies cases only when they present for care.

Loss to follow-up is a perennial problem. Differential loss by exposure group biases the result; shared parameter models (Chang et al, 2009) can reduce that bias.

Outcome ascertainment

Why incidence, not prevalence

Two examinations needed

Baseline exam confirms disease-free status. Follow-up exam detects whether and when disease developed. If screening intervals are used, timing goes at the midpoint.

Why not prevalence?

Prevalent cases may have changed their exposure because of their disease (reverse causation). They also reflect duration and survival, not incidence.

Analysis

Choosing the right model

Risk-based studies

Logistic regression (odds ratios). Log-binomial or modified Poisson for risk ratios directly.

Rate-based studies

Poisson with person-time offset (constant rates). Cox proportional hazards (varying rates). Hernán (2010) cautions that average hazard ratios shift with follow-up duration.

Cox model and hazard ratio
\[ \color{#0B7B6B}{h(t)} = \color{#6D28D9}{h_0(t)} \exp\!\bigl(\color{#1D4ED8}{\beta_1 x_1} + \cdots + \beta_p x_p\bigr) \quad \color{#BE185D}{\text{HR}} = e^{\color{#1D4ED8}{\hat{\beta}}} \]
h(t) hazard at time th0(t) baseline hazardβ log hazard ratioHR hazard ratio
Wrapping up

STROBE and the lesson arc

  • Restriction, matching, analytic control: three tools for comparability, each applied at a different stage.
  • Blinded, active surveillance gives the most accurate event timing.
  • Always measure incidence to avoid reverse causation and duration bias.
  • Risk-based studies use logistic regression; rate-based studies use Poisson or Cox. Hazard ratios shift with follow-up (Hernan, 2010).
  • STROBE is both a reporting standard and a reading checklist for published cohorts.

Introduction and Overview

Earlier sections walked through the design choices: cohort architecture, risk-vs.-rate handling, and exposure measurement. With those locked in, the remaining work is operational. Four practical questions structure this section in turn: how do you make sure the exposed and non-exposed groups are comparable on the variables that aren't your exposure of interest, how long should you follow them, how do you confirm the outcome happened, and how do you analyse what you collect?

Learning Objectives

  • Identify the three approaches to ensuring comparability of exposed and non-exposed groups.
  • Describe principles for unbiased follow-up and outcome ascertainment.
  • Recognise appropriate analytic approaches for risk-based and rate-based cohort designs.
  • Apply STROBE reporting guidelines to a cohort study.

Ensuring Exposed & Non-Exposed Groups Are Comparable

If exposed and non-exposed groups are not comparable (that is, not exchangeable: you could not swap the exposed and non-exposed labels and still expect the same pattern of disease) with respect to factors related to both exposure and outcome, a biased (confounded) assessment results (Klein-Geltink et al, 2007).

The Achilles Heel of Observational Studies

As Hernan (2012) notes, this is a key reason to prefer randomised experiments, exchangeability is expected when exposure is randomised. In observational studies, investigators must use expert knowledge to identify and measure all potential confounders in hopes of achieving exchangeability conditional on the measured covariates. Unfortunately, exchangeability cannot be empirically tested, so we never know with certainty whether we have succeeded.

Three approaches help ensure comparability:

Restriction
Click to learn more
Matching
Click to learn more
Analytic Control
Click to learn more

Comparability addresses the cross-sectional baseline. The next question is what happens over time, specifically, how long the follow-up should last and what counts as time-at-risk for each subject.

Follow-Up Period

To enhance validity, the follow-up process must be as complete as possible and unbiased with respect to exposure. Achieving unbiased follow-up may require some form of blinding to exposure status:

  • In prospective studies: assign follow-up tasks to researchers unaware of exposure status.
  • In retrospective studies: keep record reviewers unaware of exposure status when possible.

Active vs. Passive Surveillance

With passive surveillance, cases are identified when they present (e.g., date of first symptoms or physician examination). With active surveillance and regular evaluation of subjects, more accurate timing of outcome occurrence is feasible. Tooth et al (2005) recommend enumerating the at-risk population at specified times during the study period.

Losses to follow-up are a perennial concern. Chang et al (2009) describe shared parameter models to reduce bias when losses are not randomly distributed.

Even careful follow-up only matters if the outcome itself is measured well. Cohort studies tend to be ambitious about exposure characterization but surprisingly casual about outcome ascertainment, and the next subsection is the corrective.

Measuring the Outcome

Although the most frequent outcome in cohort studies is the occurrence of a specific disease (measured as risk or rate), outcomes can also be:

  • Ordinal: e.g., none, mild, moderate, severe.
  • Continuous: e.g., birth weight, blood pressure, BMI, quality of life index (Example 8.2).

Harley et al (2011), for instance, examined polybrominated diphenyl ethers (PBDEs (a flame retardant) and infant birth weight, length, and head circumference) both exposure and outcome on continuous scales.

Diagnostic Criteria

Each study needs explicit protocols for determining outcome occurrence and timing. Clear diagnostic criteria minimise diagnostic errors. In retrospective studies, this can be challenging when only summary diagnostic information is available.

Incidence Requires Two Examinations

Strictly, measuring incidence requires:

  1. An examination at the start of follow-up to ensure subjects do not yet have the disease.
  2. A second examination to determine whether (and when) the disease developed.

Why Incidence, Not Prevalence?

Including only new disease events avoids the reverse-causation problem from measuring prevalence and ensures that associations are not biased by duration-of-disease effects and survival bias (see Chapter 12). In retrospective studies, freedom from disease at the start of follow-up often must be assumed; in prospective studies, it should be formally verified.

If clinical diagnostic data are used, the incident date is usually based on the time of diagnosis, not the time of underlying disease occurrence, an important caveat for diseases with long subclinical phases. If subjects are screened at regular intervals, the time of disease occurrence is conventionally placed at the midpoint between examinations.

Multiple Outcomes

The Multiple-Comparisons Problem

One advantage of cohort studies is the ability to assess multiple outcomes from a given exposure. However, with multiple outcomes, some may be statistically significant by chance alone. The remedy is either to consider the study as hypothesis-generating rather than hypothesis-testing, or to apply a penalty to the P-value threshold, unless outcomes were specified a priori.

Comparability, follow-up, and outcome ascertainment are all design-stage decisions. The analysis stage is where those decisions translate into a quantitative result, and where pre-specified analysis plans (the EGAP-style discipline from an earlier lesson) earn their keep, because cohort data tempt almost limitless slicing.

Analysis

For closed populations, average risk and survival times can be measured during follow-up.

  • Bivariable: methods in Chapter 6.
  • Stratified analysis to control confounding: Chapter 13.
  • Multivariable: traditionally logistic regression, which uses odds ratios as the base measure.
  • For estimating risk ratios directly in multivariable settings: see Section 18.4.1.
  • For risk differences as the association measure: linear regression as described by Cheung (2007).
  • For population attributable fractions: log-linear models (Cox, 2006) or various models including logistic, log-linear, and Poisson (Greenland, 2004; Example 8.10).

For open populations, rates measure disease frequency. The choice depends on follow-up length:

  • If the rate is reasonably constant over follow-up: Poisson regression is appropriate. Subject time-at-risk is included as the offset (a fixed term carrying each group’s observation time, so the model returns a rate instead of a raw case count).
  • If the rate varies substantially over follow-up: Cox proportional hazards models are preferred (most rate-based cohort analyses in the medical literature use these).
  • For grouped data with multivariable analysis: Poisson with offsets gives direct incidence rate ratio estimates.
  • If the time of occurrence matters and not only the fact of the outcome: survival models (Case et al, 2002).

Callas et al (1998) compared proportional hazards, Poisson, and logistic models, concluding that the first two are preferable to logistic for cohort data, a finding confirmed by Greenland (2004).

Hernan (2010) describes two drawbacks to hazard ratios (HRs):

  1. Average HRs can change over time. The reported average HR depends on the duration of follow-up.
  2. Period-specific HRs have a built-in bias. The HR at time t is conditional on not having developed the outcome before time t. As follow-up lengthens, susceptible people in the exposed group are progressively depleted, so the apparent risk in the exposed group decreases relative to the non-exposed group.

Hernan describes how to circumvent these problems using covariate-adjusted survival curves. Hernan et al (2008) propose subdividing follow-up into shorter intervals and treating each as a ‘trial’, an analytic strategy that more closely approximates a randomised trial. Danaei et al (2011) elaborate this approach.

Time-Dependent Confounders

Gran et al (2010) describe a sequential Cox technique for data with time-dependent confounders, covariates that are affected by past exposure and predict future exposure and outcome (e.g., CD4 count when assessing HIV treatment effects).

Repeated Measurements

Xue et al (2010) explain how marginal and mixed-effect models can handle cohort data with repeated measurements of exposure and covariates, contrasting these with logistic and proportional hazards approaches.

The last operational concern, as in earlier lessons, is reporting. STROBE returns one more time, now with the items that matter specifically for cohort designs.

Reporting of Cohort Studies (STROBE)

The STROBE statement (von Elm et al, 2007) provides reporting guidelines for observational studies. Tooth et al (2005) elaborate criteria specific to cohort studies. These should be used both to plan and report studies, and to assess the validity of published work.

Table 8.1, Selected Criteria for Assessing Cohort Studies ▼

Study definition:

  • Are objectives or hypotheses stated?
  • Are the target population, sampling frame, and study population defined?
  • Are the setting, geographic location, and dates stated?
  • Are eligibility criteria stated?
  • Is the number of participants justified?

Recruitment & participation:

  • Are numbers meeting and not meeting eligibility criteria stated?
  • Are reasons for ineligibility and refusal given?
  • Were responders compared with non-responders?

Measurement:

  • Are methods of data collection stated?
  • Was the reliability and validity of measurement methods mentioned?
  • Were any confounders mentioned?

Follow-up & analysis:

  • Was the number of participants at each stage specified?
  • Were reasons for loss to follow-up quantified?
  • Was the type of analysis stated, including longitudinal methods?
  • Were absolute and relative effect sizes reported?
  • Were confounders and missing data accounted for in analyses?

Discussion:

  • Was the impact of biases assessed (qualitatively or quantitatively)?
  • Did authors relate results to a target population?
  • Was generalisability discussed?

Reporting Study Designs Before Results

Tooth et al (2005) and others note an increased frequency of reporting on the design and implementation of proposed or early-stage studies (e.g., Gern et al, 2009; Hermsen et al, 2011; Origasa et al, 2011; Poulos et al, 2011; Schuz et al, 2011). This is a positive trend; it allows assessment of strengths and weaknesses without the bias that comes from already knowing the results.

Appraising a Cohort Study: A Worked CASP Example

The comparability, follow-up and outcome principles of this section are what an appraisal checklist asks about. The CASP cohort checklist (Critical Appraisal Skills Programme, 2024) has 12 items: a focused issue (item 1), acceptable recruitment of the cohort (2), accurate measurement of the exposure and the outcome (3 and 4), identification of confounders and allowance for them (5a and 5b), follow-up that is complete and long enough (6a and 6b), the results and their precision (7 and 8), whether the results are believable (9), and whether they apply locally, fit other evidence and have implications for practice (10 to 12). The worked example below applies the items that carry most of the weight to a fictional excerpt, including a calculation of loss to follow-up.

Fictional excerpt: living near a major road and adult-onset asthma

Investigators invited adults aged 40 to 69 from the patient lists of eight family practices in one province in 2015 and enrolled 2,000 who had no history of asthma. Exposure was classified once, from the residential address at enrolment: 500 participants lived within 200 metres of a major road (exposed) and 1,500 lived farther away (unexposed). New physician-diagnosed asthma was identified from provincial billing records over eight years. By the end of follow-up, 85 exposed and 120 unexposed participants had left the province or withdrawn and could not be traced. Among those followed to the end, asthma developed in 24 of 415 exposed participants (5.8%) and 51 of 1,380 unexposed participants (3.7%), a crude risk ratio of 1.6. The risk ratio adjusted for age, sex and smoking was 1.5 (95% CI 0.9 to 2.4). Household income was not recorded.

Open each item to compare your own answer with a model answer.

Item 2: Was the cohort recruited in an acceptable way? ▼

Largely yes, with a limit on generalisation. Everyone was free of asthma at entry, so entry into the cohort could not depend on the outcome itself. Recruiting from family-practice lists leaves out adults without a family doctor, which affects how far the result applies to the province (external validity); it would bias the comparison only if being on a practice list depended jointly on road proximity and later asthma risk.

Items 3 and 4: Were the exposure and the outcome accurately measured? ▼

Partly. Exposure was taken from a single address at enrolment, and over eight years some people move toward or away from major roads. Because these errors are unrelated to who later develops asthma, they are probably non-differential, which for a two-category exposure tends to bias the risk ratio toward 1. The outcome came from billing records, which capture only people who see a physician. If people living near major roads, often in urban areas, visit physicians more often, asthma would be detected more completely in the exposed group, a differential error that would inflate the risk ratio.

Items 5a and 5b: Were confounders identified and taken into account? ▼

Partly. Age, sex and smoking were adjusted for. Household income is the main gap: socioeconomic position plausibly influences both where people can afford to live and their risk of asthma (through housing quality and occupational exposures), so it is a likely confounder, and its omission probably leaves residual confounding that inflates the risk ratio.

Items 6a and 6b: Was follow-up complete enough, and long enough? ▼

Calculate the losses. Overall, 205 of 2,000 participants were lost (10.3%), close to the 10% that some authors use as a threshold. The losses differ by exposure: 85 of 500 exposed (17%) against 120 of 1,500 unexposed (8%). Differential loss matters when the people lost differ in their risk of the outcome from those who stayed, for example if exposed people with early respiratory symptoms moved away from busy roads.

Check how sensitive the result is. If the 85 exposed people who were lost had developed asthma at twice the rate of those followed (11.6%), about 9.8 more cases would have occurred, the exposed risk would be (24 + 9.8) / 500 = 6.8%, and the crude risk ratio would be about 1.8 in place of 1.6. Item 6a is therefore answered “no”, or at best “can't tell”. Eight years is long enough to observe new adult-onset asthma, so item 6b is answered “yes”.

Items 7 to 9: What are the results, how precise are they, and do you believe them? ▼

The adjusted risk ratio of 1.5 means that exposed participants had about 1.5 times the risk of developing asthma over eight years. The 95% confidence interval of 0.9 to 2.4 includes 1, so the data are compatible with no effect and with more than a doubling of risk. The biases identified above pull in different directions: exposure misclassification toward 1, differential detection and confounding by income away from 1, and the losses to follow-up in a direction that cannot be known. The study raises a question worth testing with better exposure data; it does not settle it.

Items 10 to 12 call for the wider literature and the local context, which Lesson 2 showed how to find and appraise. The JBI critical appraisal checklist for cohort studies (Moola et al., 2020) covers the same ground in eleven items, including whether participants were free of the outcome at the start, whether follow-up was complete, and whether strategies were used to address incomplete follow-up.

Key Takeaways

  • Three tools for comparability: restriction (before selection), matching (at selection), analytic control (during analysis).
  • Exchangeability cannot be empirically tested in observational studies; we always work in some uncertainty.
  • Blinded follow-up reduces bias; active surveillance gives more accurate event timing than passive.
  • Always measure incidence (not prevalence) to avoid reverse-causation, duration, and survival biases.
  • Risk-based studies are typically analysed with logistic or log-linear regression; rate-based studies use Poisson or Cox proportional hazards.
  • Hernan (2010) cautions about hazard ratios changing over time and being conditional on prior survival.
  • STROBE provides a checklist for both planning and assessing cohort studies.
  • An appraisal checklist such as CASP asks about recruitment, measurement, confounding and follow-up; calculating loss to follow-up by exposure group, and testing how sensitive the result is to it, is part of the appraisal.

The reflection below is the section's exit ticket, a realistic review-the-paper prompt that requires you to use everything from this section. After working through it and the knowledge check, the next section turns from observational cohorts to the randomised controlled trial, the design that a cohort study most closely resembles.

Reflection

Reflection

You are reviewing a published cohort study reporting that a dietary supplement reduced cardiovascular events with HR = 0.75. Drawing on this section, what specific methodological aspects would you scrutinise to decide whether you trust this result? Consider comparability of groups, follow-up completeness, outcome ascertainment, choice of analysis, and STROBE reporting.

Model answerHR = 0.75 from an observational cohort needs four scrutiny points. (a) Comparability of groups: was the analysis a new-user / active-comparator design with adjustment for indication, or a prevalent-user design that conflates initiation with healthy adherence? (b) Follow-up completeness: what was the rate of loss-to-follow-up and was it differential by exposure? Substantial differential loss would bias the HR (typically toward stronger apparent benefit if dropouts are non-responders). (c) Outcome ascertainment: validated endpoints adjudicated blind to exposure, or self-report? Misclassification at the outcome end usually dilutes effects, so an HR away from null with poor measurement is suspicious. (d) Analytic choice: was a Cox model with proportional hazards checked, and were the covariates time-fixed or time-varying? (e) STROBE reporting: pre-specified protocol, registry entry, full reporting of attrition, sensitivity analyses, and confounder lists, if any are missing the result is provisional.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. Which approach to ensuring comparability is applied before subject selection?

Correct answer: B. Restriction (also called restricted sampling or exclusion) is applied before subject selection; we limit the study to subjects with one level of the confounder. Matching is applied at selection; analytic control is applied during analysis.

2. Why are incident cases preferred over prevalent cases for measuring outcomes?

Correct answer: D. Measuring incidence (new events) instead of prevalence avoids the reverse-causation problem and ensures that associations are not biased by duration-of-disease effects and survival bias. Prevalent cases include only those who have lived long enough with the disease to be counted, which distorts the picture.

3. According to Hernan (2010), which is a documented ‘hazard’ of the hazard ratio?

Correct answer: C. Hernan (2010) describes two drawbacks: (1) the average HR is dependent on the duration of follow-up, and (2) period-specific HRs are conditional on the subject not having developed the outcome before time t, which introduces a built-in bias as susceptible people in the exposed group are depleted.

4. According to STROBE-style criteria, which of the following should be reported in a cohort study?

Correct answer: A. STROBE-style criteria (Tooth et al, 2005) emphasise comprehensive reporting: reasons for loss to follow-up, methods of data collection, reliability and validity of measurements, how confounders and missing data were accounted for, and discussion of biases and generalisability.
Section 5

Randomised Controlled Trials: Foundations & Trial Set-Up

⏱ Estimated reading time: 18 minutes

Section 5 of 7

Foundations of RCTs & Trial Set-Up

The experiments of epidemiology, where the investigator controls the exposure and random allocation breaks confounding by design.

Definition

What makes a trial randomised

A randomised controlled trial allocates participants to interventions using a formal chance-based mechanism, then follows them for outcomes.

Random allocation produces baseline comparability on both measured and unmeasured factors. That is the property no observational design can match.

Portrait photograph of Austin Bradford Hill, who designed the 1948 MRC streptomycin trial.
Austin Bradford Hill (1897–1991), designer of the first modern RCT. Public domain, via Wikimedia Commons.
Research phases

From pre-clinical to post-marketing

Pre-clinical → Phase 0 → Phase I

Lab and animal models → first-in-human (sub-therapeutic, n < 15) → safety in 20–100 healthy volunteers.

Phase II → Phase III → Phase IV

Efficacy signal in 100–300 patients → pivotal efficacy in 300–3,000+ → post-marketing surveillance over large populations.

Pre-clinical Phase 0 Phase I Phase II Phase III / IV
Populations

Three nested groups

Target Population (who results should apply to) Source Population (eligible & reachable) Study Sample (eligible + consenting)
Trial objectives & comparators

Stating what you are testing

Good objective

Names the intervention, allocation design, and primary outcome. Limited in number to protect power and compliance.

Clinical equipoise

Randomisation is ethical only when genuine uncertainty exists. Where an effective standard treatment exists, it becomes the comparator.

Non-inferiority trials ask: is the new intervention no worse than the standard by more than a specified margin, delta? Choosing delta is one of the most consequential design decisions.

Specifying the intervention

Precise enough to replicate

  • State type, dose, route, timing, and duration with enough detail that another investigator can replicate the work.
  • A fixed protocol suits Phase III trials; a flexible protocol suits interventions with accumulated clinical experience.
  • Monitor compliance: pill counts, biological assays, direct observation.
  • Keep treatment assignment masked from clinical decision-makers wherever possible.
Carry forward

Next: the central design choices

  • Randomisation balances measured and unmeasured confounders, conferring causal authority no observational design can match.
  • Phased research establishes safety and efficacy incrementally before large-scale deployment.
  • Populations and eligibility criteria set the generalisability of the trial before enrolment opens.

Introduction and Overview

The opening section of this lesson described the cohort study as something close to a controlled trial without the randomisation of exposure, and the section on comparability quoted Hernan's observation that exchangeability is expected only when exposure is randomised. Every design in the first four sections measures exposure as it occurs naturally, so each must defend itself against confounding through restriction, matching, or analytic control. This section and the two that follow turn to the experimental alternative: when the investigator controls who is exposed through random assignment, the design itself breaks the link between exposure and confounders. The three sections move through a randomised controlled trial (RCT) in the order an investigator would build one: foundations and set-up (this section), the central design choices around outcome measurement, sample size, allocation, and blinding (the next section), and the conduct, analysis, and reporting of the trial (the section after that).

Learning Objectives

  • Define a randomised controlled trial and explain why RCTs are considered the gold standard for evaluating interventions.
  • Describe the phases of clinical research, from pre-clinical work and Phase 0 through Phase IV post-marketing surveillance.
  • Distinguish between target population, source population, and study sample (participants).
  • Develop appropriate eligibility criteria that balance internal validity with generalisability.
  • Specify an intervention with the precision needed for replication.

From Cohort Study to Randomised Trial

A randomised controlled trial (RCT) is a planned experiment in which the investigator deliberately allocates participants to one or more interventions, then follows them to observe outcomes. Because the allocation is determined by the investigator (rather than by self-selection or by clinical circumstance), randomisation produces groups that are comparable on both measured and unmeasured factors at baseline. This is the property that gives the RCT its inferential power.

▸ INTERACTIVE STORY: STREPTOMYCIN 1948 Open full screen ↗

Walk through the first modern RCT, the trial that turned scarcity into science. Next ▶ advances scenes.

A 7-scene retelling of the 1948 MRC streptomycin trial: postwar TB epidemic, the scarce U.S. drug shipment, Bradford Hill's elegant solution (random allocation), the slot-machine randomisation, dramatic 6-month outcomes, BMJ publication, and the rise of the RCT as gold standard.

Throughout this and the next two sections we follow the textbook's convention of using RCT to describe any planned experiment evaluating products or procedures outside the laboratory. The terms clinical trial (often restricted to therapeutic products in clinical settings) and field trial (carried out in general population settings) are used interchangeably with RCT. The factor under investigation is called the intervention; the effect of interest is called the outcome; people or groups participating are called subjects or participants. The lineage of the controlled comparison reaches back to James Lind's 1753 shipboard trial of citrus for scurvy, but it was the 1948 MRC streptomycin trial that introduced formal random allocation as the design's defining feature.

Why RCTs Are the Gold Standard

RCTs allow much better control of potential confounders than observational studies and reduce bias from selection and misinformation. By randomly assigning the intervention, the investigator breaks any link between exposure and unmeasured confounders, a feat that no observational design can match. As Lavori and Kelsey put it, the RCT is at present the unchallenged source of the highest standard of evidence used to guide clinical decision-making. That said, a single RCT is rarely sufficient to answer questions about complex interventions, and concerns persist that some trials lack relevance to real-world practice.

An earlier lesson (Introduction to Observational Studies) placed experimental studies beside observational ones and ranked designs in a hierarchy of evidence. Randomisation leaves some biases untouched, and a later lesson (Design-Specific and Temporal Biases) examines what can go wrong after allocation, including failures of concealment and blinding, placebo and Hawthorne effects, non-adherence, and contamination.

CONSORT and Trial Registration

The Consolidated Standards of Reporting Trials (CONSORT) statement was developed to improve the quality of trial reporting; it was first published in 1996 and updated in 2001, 2010 (Schulz, Altman, & Moher, 2010) and 2025 (Hopewell et al., 2025), and a later section of this lesson sets out its checklist. Since 2005 the International Committee of Medical Journal Editors has required investigators to register their trials before participant enrolment as a precondition for publishing in member journals, and the WHO operates an International Clinical Trials Registry Platform (ICTRP). A later lesson (Integrated Appraisal of Epidemiological Research) uses CONSORT as a framework for appraising published trials, and an earlier lesson (Foundations of Epidemiology) discusses registration and preregistration as research-integrity reforms.

The Phases of Clinical Research

While controlled trials are valuable for assessing a wide range of factors affecting health, one of their most common uses is to evaluate pharmacological products. Before any trial in humans, extensive pre-clinical studies are conducted in vitro (test tube or cell culture) and in vivo (animal) using wide-ranging doses to obtain preliminary efficacy, toxicity, and pharmacokinetic information.

Click any phase to see its purpose, typical sample size, and key design features.

Phase 0
First-in-Human
Click to learn more
Phase I
Safety
Click to learn more
Phase II
Efficacy Signal
Click to learn more
Phase III
Pivotal Efficacy
Click to learn more
Phase IV
Post-Marketing
Click to learn more

Objectives, Comparators, and Clinical Equipoise

The objectives of a trial must be stated clearly and succinctly. A good objective describes the intervention, the allocation design (parallel, factorial, cross-over, etc.), and the primary outcome(s). Each trial should have a limited number of objectives plus, if needed, a small number of secondary outcomes. Increasing the number of objectives complicates the protocol, jeopardises compliance, and sacrifices statistical power.

Most trials contrast two groups, intervention and comparison, and are sometimes referred to as two-arm studies. Trials with more arms can be efficient when factorial designs are used (see the next section). The comparison group might receive a placebo, no treatment, the usual treatment, or a different dose of the same product.

Choosing the Comparator: A Consequential Decision

Placebos are ideal when there is no established alternative intervention; where possible, a placebo is preferred to “no treatment.” However, when an effective standard treatment exists, withholding it from the comparison arm may be unethical; randomisation is ethically defensible only when there is genuine uncertainty in the expert community about which arm is superior, a state Freedman (1987) called clinical equipoise. In these settings, the standard of care serves as the comparator, and the trial may take the form of a non-inferiority trial, which aims to show the new intervention is no worse than the existing standard by more than a clinically unimportant margin (delta). Determining the appropriate value of delta is one of the most consequential design decisions in a non-inferiority trial.

Example 12.1: A Trial of Prostate Cancer Screening

The Andriole and colleagues (2012) trial randomised 38,340 men aged 55–74 to a screening intervention and 38,345 to usual care across 10 USA screening centres between 1993 and 2001. Men in the intervention arm were offered annual PSA tests for six years and digital rectal examination for four years. Follow-up extended through 2009 or 13 years from trial entry. The primary analysis was an intention-to-screen comparison of prostate cancer-specific mortality. The trial illustrates how a clearly stated objective (mortality reduction from screening) can be embedded in a simple two-arm parallel design that can nonetheless run for two decades.

Notice that the comparator was “usual care,” which sometimes included opportunistic screening. This pragmatic choice (Tunis, Stryer, & Clancy, 2003; Loudon et al., 2015) makes the results applicable to the real US health system but complicates interpretation of the “true” effect of organised screening.

This trial is the prostate component of the PLCO (Prostate, Lung, Colorectal and Ovarian) Cancer Screening Trial. A later lesson (Information Bias and Data Quality) returns to it to show how screening in the usual-care arm (contamination) narrowed the contrast between the two arms.

Participants: Target Population, Source Population, and Study Sample

Three nested populations must be distinguished in any trial:

Target Population (who results should apply to) Source Population (eligible & reachable) Study Sample (Participants) (eligible + consenting)

Figure 12.1. The target population is the group to which results should generalise. The source population is the subset that is eligible and reachable. The study sample (the participants) is the smaller subset that meets eligibility criteria and consents to participate.

Defining the Three Populations

These are the same target and source populations that the opening section of this lesson asked every cohort study to name when stating its objective; the figure above shows how they nest. What is distinctive in a trial is the study sample: it consists of the participants who meet the inclusion and exclusion criteria and agree to participate, so every participant is a volunteer. How well those volunteers represent the source and target populations must be considered when extrapolating results.

Unit of Concern: Individuals or Clusters?

An early decision is the level at which the intervention will be applied. Some interventions can only be delivered to groups (e.g., medication added to drinking water; a school-wide curriculum). When the intervention is applied at the group level and the outcome is measured at the group level, this is a group-level study. When the outcome is measured on individuals within those groups, this is a cluster randomised study, which the next section discusses. A later lesson (Ecological and Group-Level Studies) returns to group-level designs in observational research.

Eligibility Criteria

Document the subject's relevant past history ▼

Adequate records should be available to verify previous health history, prior treatments, and any condition relevant to the trial. This is especially important when the intervention may interact with prior therapies.

Use clear case definitions for therapeutic trials ▼

For trials of therapeutic agents, a precise definition of the disease being treated is essential. Subjects who do not actually have the condition will dilute any treatment effect and reduce statistical power.

For prophylactic trials, document baseline health ▼

For trials of preventive products, healthy subjects are required, and procedures must be in place to confirm and document health status at the start of the trial.

Subjects should be able to benefit; avoid high adverse-effect risk ▼

Restricting a trial to subjects most likely to benefit increases statistical power but may limit generalisability. Subjects at high risk for adverse effects should generally be excluded both for ethical reasons and to protect the validity of the safety assessment.

The Width of Eligibility Criteria: A Trade-off

A narrow set of eligibility criteria yields a more homogeneous response and increases statistical power, but reduces generalisability. A broad set increases the pool of potential participants and can reveal subgroup variation, but introduces greater background variability that may hurt overall power. The recommended balance: use criteria reflecting the breadth of subjects who would receive the intervention in real-world practice if it proves effective.

Specifying the Intervention

The nature of the intervention and how it is administered must be clearly defined, with enough detail that another investigator could replicate it. Interventions vary widely:

  • Medical interventions (e.g., creatine supplementation, antibiotic regimens)
  • Surgical techniques (e.g., single-layer vs. double-layer uterine closure)
  • Devices or instruments (e.g., progressive lenses for presbyopia)
  • Screening programmes (e.g., PSA testing for prostate cancer)
  • Behavioural programmes (e.g., adolescent smoking cessation)

A fixed intervention (one with no flexibility) is appropriate for assessing new products in Phase III trials. A more flexible protocol is appropriate for products that have been in use long enough that some clinical judgment has accumulated. Whenever possible, the initial treatment assignment should remain masked so clinical decisions are not influenced by knowledge of group allocation. Clear instructions are critical when participants administer some or all of the intervention themselves (such as instructions for taking medication). A monitoring system should be in place to verify that the intervention is delivered as planned.

Example 12.2: A Sequential Trial of Creatine in ALS

Groeneveld and colleagues (2003) recruited patients with amyotrophic lateral sclerosis (ALS) from neuromuscular outpatient clinics in Utrecht and Amsterdam. The two interventions were creatine monohydrate and a matching placebo, designated A and B. An independent physician, masked to assignment, instructed the research pharmacist which medication to dispense. Patients were seen at 1 month, 2 months, and every 4 months thereafter. Reasons for withdrawal (serious adverse events, withdrawal of consent) were documented; patients who stopped trial medication remained in the intent-to-treat analysis. The careful specification of this masking-and-allocation chain, from the masked physician through the pharmacist to the patient, illustrates how an unambiguous intervention specification protects the integrity of the trial.

Key Takeaways

  • RCTs are the gold standard for evaluating interventions because random allocation breaks the link between exposure and unmeasured confounders.
  • Clinical research progresses through pre-clinical studies, then Phase 0 (first-in-human), Phase I (safety), Phase II (efficacy signal), Phase III (pivotal efficacy), and Phase IV (post-marketing surveillance).
  • The CONSORT 2025 statement structures both trial design and reporting; trial registration is required by major journals.
  • Three nested populations matter: the target population (who results should apply to), the source population (eligible and reachable), and the study sample, or participants (eligible + consenting).
  • Eligibility criteria balance internal validity against generalisability. Recommend using criteria that mirror who would receive the intervention in real practice.
  • Interventions must be specified with the precision needed for replication, and a monitoring system should verify delivery.

The reflection below asks you to set up a trial of your own using the decisions from this section. After working through it and the knowledge check, the next section turns to the central design choices: how the outcome is measured, how large the trial must be, how participants are allocated, and who is masked.

Reflection

Reflection

Choose an intervention you would like to see evaluated in a randomised trial. Describe its target population, source population, and study sample (participants), and name one eligibility criterion you would use. Then name the comparator that would respect clinical equipoise, and state whether a superiority or a non-inferiority design would be appropriate, explaining why.

Model answerSuppose the intervention is a five-day course of amoxicillin for uncomplicated community-acquired pneumonia in young children, compared with the usual ten-day course. (a) Populations: the target population is children treated for uncomplicated pneumonia in outpatient care in Canada; the source population is children aged 6 months to 10 years who present to the emergency departments of participating hospitals during the recruitment period; the study sample (participants) is those who meet a clear case definition, have no severe disease, chronic lung condition, or recent antibiotic use, and whose parents consent. Excluding severe cases protects the children and the safety assessment, but it narrows generalisability, so the remaining criteria should mirror the children who would actually receive the shorter course in practice. (b) Comparator: an effective standard treatment exists, so a placebo-only or no-treatment arm would violate clinical equipoise. The comparator is the standard ten-day course, and the shorter-course arm receives five days of amoxicillin followed by five days of matching placebo so that families and clinicians cannot tell the arms apart. (c) Design: the shorter course is expected to cure about as many children as the standard course while causing fewer side effects, costing less, and exerting less selective pressure for antimicrobial resistance, so a non-inferiority design is appropriate. The protocol must pre-specify the margin (delta) by which the cure rate could fall before the shorter course is judged unacceptable, and choosing and justifying that margin with clinicians and families would be one of the most important design decisions. A superiority design would suit an intervention expected to improve outcomes over the comparator, such as a new vaccine tested against placebo for a disease with no existing vaccine.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. Phase II clinical trials are primarily designed to:

Correct answer: C. Phase II is the first evaluation of efficacy in 100–300 patients with the target condition. Phase 0 is the first-in-human study in a handful of subjects (A), Phase I focuses on safety and pharmacodynamics in healthy volunteers (B), and Phase IV is post-marketing surveillance (D).

2. The target population in an RCT refers to:

Correct answer: A. The target population is the broadest population to which results should generalise. The source population is the eligible and reachable subset; the study sample (participants) is the further subset that consents.

3. A non-inferiority trial differs from a standard superiority RCT in that the null hypothesis is that:

Correct answer: B. In a non-inferiority trial, the goal is to show the new intervention is not worse than the standard by more than a clinically tolerable amount. Choosing the value of delta is a critical and sometimes contentious design decision.

4. The defining feature that distinguishes a randomised controlled trial from an observational analytic study is that:

Correct answer: C. The defining feature of an RCT is investigator-controlled allocation of the intervention. Random assignment is what produces baseline comparability on both measured and unmeasured factors, the inferential strength that observational designs cannot match.
Section 6

Randomised Controlled Trials: Outcomes, Sample Size, Allocation & Masking

⏱ Estimated reading time: 22 minutes

Section 6 of 7

Allocation, Outcomes & Sample Size

Outcome measurement, sample size, allocation strategies, and masking.

Outcome measurement

Three scales, one principle

Dichotomous

Yes / no: disease occurred, death. Most common type. Requires larger n per participant.

Continuous

Numeric scale: blood pressure, FEV1, quality-of-life score. More efficient when variance is manageable.

Time-to-event

Uses all timing information. Often the most powerful. Analysed by survival methods.

Use clinically relevant endpoints, not just surrogates. Report dichotomous results in both absolute (risk difference) and relative (risk ratio) terms.

Sample size

The design effect for cluster trials

Design effect (cluster inflation)
\[ \color{#0B7B6B}{\text{DEFF}} = 1 + (\color{#1D4ED8}{m} - 1) \times \color{#C2410C}{\text{ICC}} \]
DEFF design effect (sample-size inflation)ICC intracluster correlation coefficient (ρ)m average cluster size

Where ICC (\(\rho\)) is the intracluster correlation coefficient and \(m\) is average cluster size. With ICC = 0.05 and \(m = 41\): DEFF = 1 + 40 × 0.05 = 3.0, tripling the required sample size.

Adding more clusters is usually more efficient than adding more individuals to existing clusters, because the power plateau occurs near 1/ρ participants per cluster.

Allocation designs

Six strategies

Simple

Independent random assignment for each participant.

Stratified

Randomise within strata of key covariates to ensure baseline balance.

Cross-over

Each participant receives both interventions in random sequence. Each serves as their own control.

Factorial

All combinations of two or more interventions tested simultaneously.

Cluster

Whole groups allocated together. Required when contamination between individuals would occur.

Split-plot

One intervention at the cluster level, another at the individual level.

Masking

Three nested layers of blinding

Triple-blind: + analysts blinded (prevents analytic bias) Double-blind: + clinicians & outcome assessors blinded Single-blind: participant unaware (reduces response bias, placebo effect) equalises attrition and co-intervention
Sequential & adaptive designs

When the plan can change mid-trial

Sequential design

Pre-specified stopping rules allow early termination for efficacy, harm, or futility. Efficient but stopping early for benefit can overestimate the effect.

Adaptive design

The protocol itself may change: dropping arms, re-sizing, switching hypotheses. Requires careful pre-specification. Platform trials are one important variant.

Carry forward

Next: the trial in motion

  • Outcome scale, sample size, allocation design, and masking each either preserve or undermine the advantage randomisation provides.
  • The design effect must be applied at the sample-size stage, not corrected after the fact.
  • Masking success should be evaluated and reported, not assumed.

Introduction and Overview

The previous section covered the up-front decisions: what an RCT is, what phase of clinical research it falls under, what the trial design is meant to test, and who gets included. This section turns to the four central design decisions that follow: how the outcome is measured, how large the sample needs to be to detect the effect, how participants are allocated to groups, and who is masked from the allocation. Each of these is a place where a poorly run RCT can lose its inferential advantage.

Learning Objectives

  • Identify primary and secondary outcomes that are clinically relevant and measurable.
  • Understand the inputs to sample size calculations and how cluster randomisation, sequential designs, and adaptive designs change them.
  • Distinguish simple, stratified, cross-over, factorial, cluster, split-plot, and multicentre allocation strategies.
  • Compare single, double, and triple blinding and identify the bias each is designed to prevent.

Measuring the Outcome

A controlled trial should be limited to one or two primary outcomes and a small number (one to three) of secondary outcomes. Too many outcomes create the multiple-comparisons problem introduced under “Multiple Outcomes” in an earlier section and inflate the false-positive rate; the next section shows how trials adjust for it. Composite outcomes, which combine several events into a single measure, are sometimes used but remain controversial; for our purposes a limited number of primary and secondary hypotheses is preferred.

Outcome Scales

Dichotomous Outcomes

The outcome is yes/no: occurrence of disease, death, recovery. This is the most common type in medical trials. Dichotomous outcomes generally require larger sample sizes than continuous outcomes for the same effect size, because they convey less information per subject.

Results should be reported in both absolute terms (risk difference, number needed to treat) and relative terms (risk ratio). The number needed to treat is 1 divided by the absolute risk difference: the number of people who must receive the intervention for one additional person to benefit. HSCI 341 Lesson 6, Section 2 (Measures of Effect in the Exposed Group) works through the calculation.

Continuous Outcomes

The outcome is measured on a numeric scale: blood pressure, FEV1, quality-of-life score. Continuous outcomes are often more statistically efficient than dichotomous ones because each subject contributes more information.

Adjustment for baseline (pre-intervention) values can substantially improve precision when the correlation between baseline and follow-up exceeds 0.5.

Time-to-Event Outcomes

Survival analysis, time to disease occurrence, time to relapse. These designs can be more powerful than simple occurrence-or-not in a defined follow-up window because they use all of the timing information. The accuracy of the actual time of event matters, although Korn and colleagues note that non-differential errors in event timing rarely have a major impact on treatment-effect estimation.

Choosing Clinically Relevant Outcomes

Outcomes that can be assessed objectively are preferred, but sometimes subjective outcomes are unavoidable (e.g., self-reported symptoms). When the outcome is not assessed by a near-gold-standard procedure, the impact of the intervention on the true outcome may differ from the impact on the surrogate. Intermediate outcomes (e.g., antibody titres in a vaccine trial) can illuminate mechanism but should not replace clinically relevant primary endpoints. Clinically relevant outcomes typically include:

  • Diagnosis of a particular disease: requires a clear case definition
  • Mortality: objective but still requires criteria for cause and time of death
  • Severity scores: difficult to develop reliably
  • Objective clinical measures (rectal temperature, blood biomarkers)
  • Quality-of-life and other patient-reported outcome measures

An earlier lesson (Introduction to Observational Studies) urged the same care when a surrogate stands in for a clinically important outcome in observational designs.

Sample Size

The size of a trial is determined through formal sample-size calculations that incorporate the estimated intervention effect, Type I error rate, and Type II error rate. Power is conventionally set to 90 percent. The sample sizes do not need to be equal in both arms. An earlier section discussed sample-size planning for cohort studies, and a later lesson (Confounding and Statistical Inference) explains Type I and Type II errors and statistical power in more detail.

Sample Size for Cluster Randomised Trials

When subjects are randomised in clusters (families, schools, clinics), the analysis must account for within-cluster similarity, summarised by the intracluster correlation coefficient (ICC, ρ) and the cluster size (m). The required sample size is inflated by the design effect:

Design effect
\[ \color{#0B7B6B}{\text{DEFF}} = 1 + (\color{#1D4ED8}{m} - 1) \times \color{#C2410C}{\text{ICC}} \]
The design effect is the factor by which the required sample size grows under cluster randomisation; it rises with the intracluster correlation coefficient and the cluster size, so even a small ICC inflates the sample when clusters are large.

For example, with ICC = 0.05 and m = 41, the design effect is 1 + 40 × 0.05 = 3.0, which triples the required sample size.

The design effect tells you how many times larger a clustered sample must be to carry the same information as a simple random sample of unrelated individuals. The reason is that people in the same cluster tend to resemble one another, so a second person from a cluster you have already sampled adds less that is new than a wholly independent person would.

Even when rho is small, large clusters cause substantial inflation. Notably, the power of a cluster trial does not increase appreciably once the number of subjects per cluster exceeds 1/rho, so adding more individuals to existing clusters yields diminishing returns. Adding more clusters is often more efficient than adding more individuals per cluster.

The same design effect applies to surveys that sample clusters. HSCI 341 Lesson 2, Sections 4 and 5 (Sampling Distributions and Survey Analysis; Planning Sample Size), applies it to the analysis and the sample size of complex surveys.

Sequential and Adaptive Designs

Sequential designs ▼

A sequential design (sometimes called a monitored study) allows hypothesis tests to be conducted on a number of occasions as data accumulate. Sample size is not fixed; instead, prespecified stopping rules halt the trial when efficacy, harm, or futility becomes clear. Sequential designs can be efficient but tend to lack power on a per-subject basis. Stopping early for benefit can produce overestimates of treatment effect, although the bias is often modest. Interim analyses should not be conducted unless the trial design accommodates them.

Adaptive designs ▼

Adaptive designs allow the trial design to change as the study progresses. The most common adaptation is modifying the second-stage sample size based on first-stage power. Other adaptations include dropping or adding treatment arms, changing the primary endpoint, or even switching from non-inferiority to superiority. Outcome-adaptive designs use accumulating evidence to assign more subjects to the better-performing intervention (e.g., “play-the-winner”); these are appropriate only when the result of the intervention is identifiable shortly after treatment. The platform trial, an extension that evaluates multiple treatments simultaneously against a shared control under a master protocol, adding and dropping arms over time, has become an important variant in oncology and infectious-disease research (Berry, Connor, & Lewis, 2015).

Other sample-size considerations ▼

Recruitment time matters. If season influences treatment response, the recruitment window should span a full calendar year. Loss to follow-up, non-compliance, and competing risks should be anticipated and the sample size adjusted upward to preserve power.

Allocation of Study Subjects

Once enrolled, subjects must be allocated to interventions. Formal randomisation is the strongest method; without it, bias is very likely to distort findings (Schulz & Grimes, 2002a). Random allocation should occur as close to the start of the intervention as possible to minimise withdrawals between assignment and treatment, and concealment of the upcoming allocation from those enrolling participants is essential to prevent selection bias (Schulz & Grimes, 2002b).

Background: random sampling and random allocation

The two procedures share the word “random” but do different jobs. Random sampling decides who enters the study, and it supports generalisation from the participants to the source population. Random allocation decides which arm each participant receives, and it makes the arms comparable at the start, so that differences in outcome can be credited to the intervention. A study can use either, both or neither:

Random allocation usedNo random allocation
Random sampling usedA survey experiment: a random sample of residents is randomly assigned one of two versions of a question.A population survey that selects people at random and assigns nobody to anything.
No random samplingMost randomised controlled trials, which enrol eligible volunteers and then allocate them by chance.A cross-sectional survey of patients who happen to attend one clinic during one week.

HSCI 207 Lesson 7, Section 2 (Probability and Non-Probability Samples, Recruitment and Participant Flow) develops this distinction and is optional reading.

Retrieval question. A trial enrols 200 adults with high blood pressure who answer a newspaper advertisement and uses a computer to assign half of them to a new exercise program. Which random procedure does it use, and what does that procedure protect?

Show answer ▼

It uses random allocation only. Allocation makes the two arms comparable at the start, which protects the internal validity of the comparison. Because the volunteers were not randomly sampled, how far the result applies beyond them is a separate question of external validity.

Concealment protects the moment of assignment, and blinding, discussed below, protects what happens after it; a later lesson (Design-Specific and Temporal Biases) examines both, along with the biases that arise when either fails.

Alternatives to Randomisation

Historical Controls and Systematic Assignment

Historical control trials compare outcomes after an intervention with outcomes from a pre-intervention period. For validity, four conditions must hold: predictable outcome, complete and accurate databases, constant diagnostic criteria, and no environmental changes for subjects. Rarely are all four met. Blinding is impossible in this design.

Systematic assignment (every other subject) can be reasonable in field settings (e.g., a vaccine clinic) and is often as effective as randomisation when outcome assessment is blinded. Randomise the very first subject's assignment to avoid predictability. Never give the intervention to the first half of subjects and the comparison to the second half, which introduces secular confounding.

The Family of Random Allocation Designs

Random allocation does not mean haphazard allocation. A formal process (a computer-based random number generator, sealed-envelope randomisation, or even a coin toss) must be used (Schulz & Grimes, 2002a). Click any design to see its features and an example.

Simple
Randomisation
Click to learn more
Stratified
Randomisation
Click to learn more
Cross-Over
Design
Click to learn more
Factorial
Design
Click to learn more
Cluster
Randomisation
Click to learn more
Split-Plot
Design
Click to learn more

Cluster Randomisation: Why It Is Less Efficient

Cluster randomised trials are statistically less efficient than individually randomised trials of the same total sample size (Campbell, Elbourne, & Altman, 2004). The clustering of subjects within groups must be accounted for in the analysis. The best follow-up scenario is to monitor all individuals for the duration of the study; if not, following a randomly selected cohort is the next most powerful approach. In some settings, repeated cross-sectional samples within each cluster must be used. Matched-cluster designs may be appropriate when the number of clusters is small, although “breaking the matches” can sometimes improve statistical efficiency.

Example 12.3: A Cluster RCT of Adolescent Smoking Cessation

Dalum and colleagues (2012) randomised 22 continuation schools in Denmark by coin toss to deliver a smoking cessation intervention or to act as control. The randomisation was blocked so that each county contained both intervention and control schools, balanced across school types (commercial vs. social-and-health). Smoking status was self-reported in surveys at baseline, week 11/2005 (short-term), and week 11/2006 (long-term). Analyses were intent-to-treat, with school as a random factor in logistic regression.

Why was cluster randomisation appropriate here? Because the intervention was delivered at the school level, individual randomisation would have introduced contamination, since intervention and control students in the same hallway would influence each other. A later lesson (Design-Specific and Temporal Biases) shows how contamination affected the community-level COMMIT smoking-cessation trial.

Example 12.4: A 2×2×2 Factorial Caesarean-Section Trial

The CAESAR trial (CAESAR study collaborative group, 2010) randomised women aged >15 undergoing their first Caesarean section to three independent factors: single- vs. double-layer uterine closure, closure vs. non-closure of the peritoneum, and liberal vs. restricted use of a subsheath drain. Telephone randomisation with a minimisation algorithm balanced participating centre, labour status, and pregnancy multiplicity. The primary outcome was maternal infectious morbidity (any of: antibiotic use for febrile morbidity, endometritis, or treated wound infection). The 3,500 women required to detect a 12% to 9% reduction with 80% power illustrate how factorial designs efficiently address multiple research questions in one trial.

Example 12.5: A Split-Plot Shoulder-Pain Trial

Watson and colleagues (2008) conducted a pragmatic split-plot trial across UK general practices. Physicians in 91 practices (the whole plot) were randomised to additional training in shoulder injection or to no additional training. Within the practices, 215 patients with acute shoulder pain were then randomised to receive either a corticosteroid or a lignocaine injection (the split plot). The main outcome was the British Shoulder Disability Questionnaire score. Notice how this design lets the trial answer two questions simultaneously: does training help, and does corticosteroid outperform lignocaine?

Example 12.6: A Multicentre Cross-Over Trial of Progressive Lenses

Boutron and colleagues (2008) compared two generations of progressive lenses for presbyopia at five primary-care optical dispensaries. 127 patients aged 43–60 were randomised to wear one lens for four weeks, then cross over to the other for four weeks, blinded to the lens sequence. Patients and the statistical analyst were both blinded; all equipment was assembled in one laboratory to ensure consistency. The primary outcome was patient preference at week 8. The cross-over design is appropriate here because lens preference is reversible and short-acting; conditions are stable, and switching has no carry-over.

Multicentre Trials

If an adequate sample is not available at one site, a multicentre trial is required. Within-centre and between-centre variances must be accounted for in design and analysis. Multicentre trials enhance generalisability (because of the broader geographic and clinical reach) and create opportunities to detect interaction effects across sites. For statistical efficiency, the number of subjects per centre should be approximately equal. Example 12.6 was conducted across five centres.

Masking (Blinding)

Blinding (or masking) refers both to the methodological principle of withholding information from individuals to prevent bias and to the specific procedures used to do so (Schulz & Grimes, 2002c). Terms can be used inconsistently in the literature, so the specific masking mechanisms always need to be described, and ideally pilot-tested in larger trials.

Triple-blind: + analysts blinded Double-blind: + clinicians/outcome assessors blinded Single-blind: participant unaware of allocation reduces response bias; equalises placebo effects

Figure 12.2. The three nested layers of blinding. Each additional layer is added on top of the prior layers, prevents an additional source of bias, and is harder to achieve in practice.

What Each Level of Blinding Prevents

A later lesson (Design-Specific and Temporal Biases) summarises the evidence that trials with inadequate blinding tend to report larger effects. The tabs below set out who is masked at each level and which biases each layer guards against.

Single-Blind: Participant Unaware

In a single-blind study the participant does not know which intervention they are receiving. This helps:

  • Reduce response bias (subjects reporting symptoms differently based on what they think they are getting)
  • Control for the placebo effect, which then occurs equally in both arms
  • Equalise differential attrition and non-compliance
  • Reduce co-intervention bias and follow-up bias

Double-Blind: Participant + Treatment/Outcome Personnel Unaware

In a double-blind study, both the participants and the people administering the intervention or assessing the outcome are unaware of allocation. This adds protection against:

  • Patient–provider interaction effects on the placebo response
  • Differential attrition, non-compliance, or co-intervention driven by clinician knowledge
  • Selective decisions and referrals based on treatment knowledge
  • Observer bias and diagnostic bias when assessing outcomes

Triple-Blind: + Data Analysts Unaware

In a triple-blind study, the people analysing the data are also unaware of group identity (often coded as A vs. B). This is designed to ensure unbiased analytic decisions (choice of subgroups, handling of outliers, modelling decisions) that might otherwise be subtly influenced by knowledge of which group is the “new” treatment.

The success of blinding should be evaluated rather than assumed. Methods for assessing blinding success have been published.

The Role of Placebos

A placebo is a product indistinguishable from the active intervention, administered to the comparison group. In drug trials, the placebo is often the vehicle without the active ingredient. Even apparently inert placebos can have positive or negative effects (e.g., a placebo vaccine without antigen can still induce some immunity through adjuvant). These issues should be addressed before the trial begins. In some situations, blinding cannot be achieved with placebos alone, but masking should be implemented wherever feasible. The same later lesson describes placebo and nocebo effects and how large they can be.

The Take-Away on Blinding

Each additional layer of blinding prevents a different bias, but each is harder to achieve in practice. Whenever feasible, design the trial so that as many sources of error as possible are prevented from the start; this also reduces the impact of any differential errors that do arise. The goal is not maximum blinding for its own sake, but rigorous masking matched to the biases that most threaten the specific trial.

Key Takeaways

  • Limit a trial to 1–2 primary outcomes and 1–3 secondary outcomes; report effects in both absolute and relative terms when outcomes are dichotomous.
  • Sample size depends on the expected effect, Type I and Type II error rates, and the outcome scale. Cluster randomisation requires inflation by the design effect, DEFF = 1 + (m − 1) × ICC.
  • Sequential designs allow stopping for efficacy, harm, or futility but tend to lack power per subject; adaptive designs offer flexibility but require careful pre-specification.
  • Random allocation is the strongest assignment method. The major design types are simple, stratified, cross-over, factorial, cluster, and split-plot, plus multicentre extensions.
  • Single-blind protects against participant response bias; double-blind adds clinician and outcome-assessor blinding; triple-blind adds analyst blinding. Each layer prevents a different category of bias.

The reflection below asks you to make these choices for a school-based trial. After working through it and the knowledge check, the next section follows the trial into motion: follow-up, analysis, vaccine trials, and reporting.

Reflection

Reflection

A school board plans to test a classroom physical-activity programme that teachers deliver during class time. Explain whether you would randomise individual students, classes, or whole schools, and why. If classes of 30 students are the clusters and the intracluster correlation coefficient (ICC) for the outcome is 0.03, calculate the design effect and explain what it does to your sample size. Finally, say which layers of blinding (participants, teachers and outcome assessors, analysts) would be feasible in this trial.

Model answer(a) Unit of randomisation: the teacher delivers the programme to the whole class, so students in the same room cannot be given different programmes, and individual randomisation would also invite contamination. Classes are the smallest workable unit; randomising whole schools further reduces contamination between classes that share gyms, playgrounds, and teachers, at the cost of fewer and larger clusters. (b) Design effect: DEFF = 1 + (30 − 1) × 0.03 = 1 + 0.87 = 1.87, so the trial needs about 1.9 times as many students as an individually randomised trial would. If 400 students would suffice with individual randomisation, roughly 748 (about 25 classes) are needed. Because power rises little once clusters exceed about 1/ρ ≈ 33 students, recruiting more classes is more efficient than enlarging them, and the inflation must be built into the sample-size calculation before recruitment. If whole schools are randomised, m becomes the number of students measured per school and the design effect grows accordingly. (c) Blinding: students and teachers will know which programme they are in, so participant and provider masking are not feasible. The staff who measure the outcome (for example, fitness tests or accelerometer data) can be kept unaware of allocation, and analysts can work with groups coded A and B, so outcome-assessor and analyst masking are both achievable. Allocation of schools or classes should also be concealed until each has enrolled.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. In a cluster randomised trial with intracluster correlation coefficient ICC = 0.05 and average cluster size m = 41, the design effect (sample-size inflation factor) is:

Correct answer: C. Design effect = 1 + (m − 1) × ICC = 1 + 40 × 0.05 = 1 + 2 = 3.0. Even a small intracluster correlation coefficient produces substantial inflation when cluster size is modest.

2. A cross-over design is appropriate when:

Correct answer: A. Cross-over designs require a stable condition and a short-acting intervention so that each subject can experience both treatments without carry-over. Mortality and irreversible disease progression rule out cross-over by definition.

3. The chief additional bias prevented by triple-blinding (over and above double-blinding) is:

Correct answer: D. Single-blind addresses participant response bias; double-blind adds clinician and outcome-assessor bias. Triple-blind extends masking to data analysts so that analytic choices (subgroup definitions, outlier handling, model specification) are not subtly influenced by knowledge of which group is which.

4. A sequential trial stops at a pre-specified interim analysis because the new treatment shows clear benefit. Which statement about the reported treatment effect is correct?

Correct answer: B. Pre-specified stopping rules let a trial halt early for benefit, harm, or futility, and pre-specification protects the false-positive rate. A trial that stops when the accumulating data look most favourable still tends to overstate the effect, so early-stopped results deserve some caution. Ad-hoc interim looks are worse, because they also inflate the false-positive rate.
Section 7

Randomised Controlled Trials: Conduct, Analysis, Vaccine Trials & Reporting

⏱ Estimated reading time: 22 minutes

Section 7 of 7

Conduct, Analysis & Special Topics

Follow-up, intent-to-treat, vaccine efficacy, and CONSORT 2025 reporting.

Follow-up & compliance

Keeping participants and data

Enrolled (n = N) Intervention arm Control arm Completed follow-up Completed follow-up Lost / withdrew (reasons) Lost / withdrew (reasons)
Analysis

Intent-to-treat vs. per-protocol

Intent-to-treat (ITT) principle
\[ \text{Analyse all participants as originally assigned, regardless of compliance.} \]

ITT (primary)

All randomised participants analysed as assigned. Conservative; reflects real-world effectiveness. Preserves randomisation.

Per-protocol (secondary)

Compliant completers only. Estimates efficacy under ideal compliance. Likely biased because non-compliance is not random.

Multiple comparisons

The Bonferroni correction

Bonferroni-adjusted significance threshold
\[ \color{#0B7B6B}{\alpha_{\text{adjusted}}} = \frac{\color{#C2410C}{\alpha_{\text{nominal}}}}{\color{#1D4ED8}{k}} \]
αadjusted per-test thresholdαnominal overall significance levelk number of comparisons

Where \(k\) is the number of comparisons. With \(k = 4\) and \(\alpha_{\text{nominal}} = 0.05\): each test must reach \(p < 0.0125\) to be declared significant.

Subgroup analyses must be pre-specified. Test interaction, not separate sub-analyses. Detecting interactions requires roughly four times the sample size of the main effect.

Vaccine efficacy

Direct, indirect, and total effects

Direct efficacy (within a single population)
\[ \color{#0B7B6B}{VE_d} = \frac{\color{#C2410C}{I_{nv}} - \color{#1D4ED8}{I_v}}{\color{#C2410C}{I_{nv}}} \]
VEd direct efficacyInv incidence, unvaccinatedIv incidence, vaccinated
Indirect efficacy (herd immunity effect)
\[ \color{#0B7B6B}{VE_{\text{ind}}} = \frac{\color{#C2410C}{I_{nvB}} - \color{#1D4ED8}{I_{nvA}}}{\color{#C2410C}{I_{nvB}}} \]
VEind indirect (herd-immunity) effectInvB unvaccinated, low-coverage areaInvA unvaccinated, high-coverage area
Total vaccine effect (population-level)
\[ \color{#0B7B6B}{VE_{\text{tot}}} = \frac{\color{#C2410C}{I_B} - \color{#1D4ED8}{I_A}}{\color{#C2410C}{I_B}} \]
VEtot total effectIB overall incidence, low-coverageIA overall incidence, high-coverage

Subscripts: v = vaccinated, nv = unvaccinated; A = higher-coverage population, B = lower-coverage population. Estimating all three requires data from at least two populations with different coverage.

Reporting

CONSORT 2025 key items

Design & methods

Trial design and ratio; eligibility; intervention detail; outcome definitions; sample-size justification; randomisation mechanics; blinding description.

Results & discussion

Participant flow diagram; baseline table; effect sizes with 95% CIs in both absolute and relative terms; harms; limitations; generalisability.

Extensions for cluster, non-inferiority, and pragmatic trials were developed for CONSORT 2010 and are used alongside the current statement.

Wrapping up

Three ideas into the final assessment

  • ITT and CONSORT are both commitments to transparency. Treat them as design tools, not write-up checklists.
  • When participants are not independent (as in vaccine or cluster trials), standard effect measures can mislead.
  • Consulting CONSORT at the design stage is what makes transparent reporting achievable, not an afterthought.

Introduction and Overview

Earlier sections settled the design. This section walks through what happens once the trial is running: how to track participants through follow-up, how to analyse the data (including the intention-to-treat principle and the special case of vaccine efficacy), and how to report the trial honestly via the CONSORT checklist. The reporting framework here connects back to the research-integrity material in an earlier lesson (Foundations of Epidemiology), including trial registration and selective outcome reporting.

Learning Objectives

  • Implement effective follow-up and compliance monitoring during a trial.
  • Distinguish intent-to-treat from per-protocol analysis and choose the appropriate approach.
  • Identify the sources of multiple comparisons in RCTs and apply the Bonferroni adjustment.
  • Compute and interpret direct, indirect, and total vaccine efficacy.
  • Use the CONSORT 2025 checklist to plan and report a randomised trial.

Follow-Up and Compliance

One of the most important practical issues is ensuring that all groups are followed rigorously and equally. The follow-up period must be long enough to capture all outcomes of interest. Some loss is inevitable through drop-out or non-compliance; for trials with long follow-up, the status of all subjects should be ascertained at regular intervals. The CONSORT statement strongly recommends a flow diagram showing participant numbers at allocation, intended intervention, protocol completion, and outcome assessment. The principles of complete, unbiased follow-up set out for cohort studies in an earlier section apply equally to trials, and a later lesson (Sampling, Selection Processes and External Validity) examines attrition bias in longitudinal studies.

Strategies to Minimise Loss and Maximise Compliance

Maintain regular communication with all participants ▼

Frequent contact, such as reminder messages, newsletters, and study updates, reduces attrition. Incentives may be provided, including study-related information that participants would not otherwise have, or public recognition of their contribution (subject to confidentiality).

Capture data on dropouts whenever possible ▼

For participants who drop out, information may still be available through routine databases if the participant consents. Documenting reasons for withdrawal allows comparison of withdrawn and remaining subjects, helping characterise potential bias.

Verify compliance directly and indirectly ▼

Compliance can be assessed through interviews, biological samples (drug or metabolite levels), or indirect indicators such as collecting empty pill containers, vials, and packaging. Compliance data are essential for interpreting the difference between intent-to-treat and per-protocol results.

Statistical Methods and Analysis

Outcomes might be analysed on a continuous scale, as categorical (often dichotomous) data, or as time-to-event measurements. Time-to-event analyses can have greater power than simple occurrence-or-not in a defined window. Whatever the analysis, results should report both the effect size and its precision (typically a 95 percent confidence interval), and dichotomous outcomes should appear in both absolute (risk difference) and relative (risk ratio) terms.

Intention-to-Treat and Per-Protocol Analysis

An intention-to-treat (ITT) analysis keeps every randomised participant in the group to which they were assigned, whatever they actually received, so it preserves the comparability that randomisation created and estimates the effect to expect when the intervention is used in practice. A per-protocol (PP) analysis is restricted to participants who complied with the protocol; it estimates the effect under ideal adherence, but it is likely to be biased because people who do not comply usually differ from those who do. ITT is therefore the primary analysis and PP a secondary one. A later lesson (Design-Specific and Temporal Biases) develops both analyses in its discussion of compliance and adherence bias.

Stating Numbers and Compliance Is Essential

Whichever analysis is primary, the number of subjects in each group, and whether or not they complied, must be reported. If there are considerable losses or major adherence problems, Hernán and Hernández-Díaz (2012) suggest using inverse probability weighting to reduce potential bias.

Baseline Comparison and Covariate Adjustment

Analysis usually starts with a baseline comparison of group characteristics as a check on randomisation. This is an assessment of comparability rather than a statistical significance test. Differences between groups, even if not statistically significant, should be noted and may justify covariate adjustment.

For dichotomous outcomes, adjustment for covariates is recommended, ideally for strong predictors identified a priori; failing that, for variables predictive of the outcome in the trial data. Adjustment can substantially increase power or reduce required sample size. Adjustment for non-confounders does little harm provided they are not intervening variables (mediators).

For continuous outcomes, controlling for baseline (pre-intervention) values can substantially improve precision. This can be done either by analysing the change score (post minus pre) or by including baseline as a covariate. Either approach gains power, particularly when the baseline-to-follow-up correlation exceeds 0.5.

The Multiple Comparisons Problem

Multiple comparisons in RCTs arise from three sources: examining multiple outcomes, examining multiple subgroups, and performing periodic interim analyses during the trial. The problem is that the experiment-wise (family-wise) error rate is much larger than the per-test rate, making spurious “significant” findings increasingly likely as the number of tests grows. The same problem arose for cohort studies with many outcomes in an earlier section; trials add subgroups and interim looks as further sources.

Bonferroni
\[ \color{#0B7B6B}{\alpha_{\text{adjusted}}} = \frac{\color{#C2410C}{\alpha_{\text{nominal}}}}{\color{#1D4ED8}{k}} \]
The per-test threshold is the overall (family-wise) significance level divided by the number of comparisons; with five comparisons each test must reach p < 0.01.

The Bonferroni adjustment is the simplest fix: divide the desired experiment-wise alpha by the number of comparisons. It is conservative; less conservative procedures (Holm, Hochberg, Benjamini-Hochberg) are available in standard texts.

The Subgroup Analysis Trap

It is tempting to evaluate many subgroups to see whether the intervention works in any of them. This should be avoided: only subgroup analyses planned a priori should be carried out; data-driven subgroup analyses generate spurious associations at alarming rates. The recommended way to test whether an intervention's effect varies by subgroup is a single overall interaction test. Note that detecting interactions reliably typically requires a sample size at least four times larger than detecting the overall main effect; effect sizes for interactions need to be roughly twice the magnitude of the main effect to be detected with similar power.

Sequential Analyses: When to Stop

Sequential designs, introduced in the previous section, plan periodic analyses throughout the trial to allow early stopping for one of three reasons:

  • Clear (and statistically significant) evidence of the superiority of one intervention over the other
  • Convincing evidence of harm from the intervention (regardless of statistical significance)
  • Little likelihood that the trial will produce evidence of an effect even if completed (futility)

Interim analyses must be pre-specified in the trial design; ad-hoc “peeking” inflates the false-positive rate and is regarded as a serious methodological breach.

Vaccine Trials: Direct, Indirect, and Total Efficacy

Standard RCT designs need modification when the intervention is a prophylactic against a communicable organism. The reason: an effective vaccine has effects on the vaccinated and on the unvaccinated, because vaccination reduces transmission. This means study subjects are not independent, an effect Hudgens and Halloran (2008) call interference. A later lesson (Ecological and Group-Level Studies) presents herd immunity as a property of populations that no individual possesses; vaccine trials need designs that can measure it.

In vaccine trials, direct, indirect and total effects describe herd effects: protection of the vaccinated person, protection that spreads to others, and their combination. The same three terms mean something different in mediation analysis (Lesson 7 and HSCI 410 Lesson 8), where they divide an exposure's effect into the part that runs through a mediator and the part that does not.

Why Standard Direct Efficacy Estimates Are Misleading

The standard individually randomised, placebo-controlled trial of a vaccine yields an estimate of vaccine efficacy that is confounded by the proportion of the study population vaccinated. Two identical trials in populations with different transmission levels, or different vaccination coverage levels, will report different vaccine efficacy estimates even if the underlying biological protection is the same.

Three Measures of Vaccine Efficacy

To get a fuller picture, epidemiologists distinguish three measures, each requiring information from at least two populations with different vaccination coverage. Use Iv for the incidence in the vaccinated and Inv for the incidence in the unvaccinated; subscript A denotes the higher-coverage population, B the lower-coverage population.

Direct VE
\[ \color{#0B7B6B}{VE_d} = \frac{\color{#C2410C}{I_{nv}} - \color{#1D4ED8}{I_v}}{\color{#C2410C}{I_{nv}}} \]
Direct efficacy compares the incidence in the unvaccinated with the incidence in the vaccinated inside one population. It equals one minus the ratio of vaccinated to unvaccinated incidence, a quantity known as the relative risk reduction.
Indirect VE
\[ \color{#0B7B6B}{VE_{\text{ind}}} = \frac{\color{#C2410C}{I_{nvB}} - \color{#1D4ED8}{I_{nvA}}}{\color{#C2410C}{I_{nvB}}} \]
The indirect (herd-immunity) effect compares the unvaccinated in the lower-coverage area with the unvaccinated in the higher-coverage area.
Total VE
\[ \color{#0B7B6B}{VE_{\text{tot}}} = \frac{\color{#C2410C}{I_B} - \color{#1D4ED8}{I_A}}{\color{#C2410C}{I_B}} \]
The total effect compares the overall incidence in the lower-coverage population with that in the higher-coverage population.

Example 12.7: Cholera Vaccine Effects in Bangladesh

Ali and colleagues (2005) and Hudgens and Halloran (2008) analysed an individually randomised, placebo-controlled trial of killed oral cholera vaccines in residential areas (baris) in Bangladesh. Two groups were compared: Group A (more than 50% coverage) and Group B (less than 28% coverage). The first-year risks per 1,000 were: RnvB = 7.01, RvB = 2.66, RnvA = 1.47, RvA = 1.27, RB = 4.13, RA = 1.34. Here each first-year risk R plays the role of the incidence I in the equations above, so the two symbols refer to the same quantity.

Direct effect in the high-coverage group A: VEd = (1.47 − 1.27) / 1.47 = 0.14. Looking at A alone you might conclude the vaccine has little effect.

Direct effect in the low-coverage group B: VEd = (7.01 − 2.66) / 7.01 = 0.62, so the vaccine reduces risk by 62% in the unprotected population.

Indirect effect (in the unvaccinated): (7.01 − 1.47) / 7.01 = 0.79, so herd immunity reduced the unvaccinated risk by 79%.

Total relative effect: (7.01 − 1.27) / 7.01 = 0.82. Overall effect: (4.13 − 1.34) / 4.13 = 0.68.

The lesson is striking: limiting analysis to the high-coverage population alone (Group A) would have suggested the vaccine barely worked. Looking across populations with different coverage reveals the dominant role of indirect effects.

Bar chart of four efficacy numbers from the Bangladesh cholera vaccine trial: direct efficacy 14% in the high-coverage group, 62% in the low-coverage group, indirect (herd-immunity) effect 79%, and total effect 68%.
Figure 12.3. The same trial yields four different numbers. Reading only the direct effect in the high-coverage area (14%) would understate the vaccine, because in that setting most of the protection arrives as herd immunity.

Designing Vaccine Trials to Estimate All Three Measures

To estimate VEd, VEind, and VEtot, the design must include at least two comparable populations with different vaccination coverage. The recommended approach: in population A, randomise individuals to vaccine or placebo. In population B, leave everyone unvaccinated (or assign a lower coverage). The two populations must be (a) comparable in characteristics that affect the outcome, especially baseline transmission level, and (b) physically separated so subjects do not intermix.

An alternative when finding two similar populations is impractical is to exploit natural clustering, for example randomly assigning vaccination to half the children in a geographic area, then study spread within schools, recording the proportion of children at each school who were vaccinated. Riggs and Koopman (2005) and Longini and colleagues (1998, 2002) describe how to design and analyse such trials.

Reporting: The CONSORT 2025 Checklist

A later lesson (Integrated Appraisal of Epidemiological Research) introduces CONSORT's key requirements and its flow diagram as a framework for appraising published trials. CONSORT 2025 (Hopewell et al., 2025) is the dominant reporting guideline for parallel-group randomised trials. It updates CONSORT 2010, replacing its 25 items with 30 and adding a section on open science (registration, protocol, data sharing, funding and conflicts of interest). Its items align with the design and conduct topics covered in this and the previous two sections; the table below highlights those most relevant to the topics we have studied, numbered as in the 2025 checklist.

Section / ItemWhat to Report
1a TitleIdentify the study as a randomised trial in the title (improves “searchability”).
2–5 Open scienceTrial registration; where the protocol and statistical analysis plan can be accessed; how data and code can be accessed; funding and conflicts of interest.
6–7 Background & objectivesScientific rationale and specific objectives about benefits and harms.
8 Patient and public involvementHow patients or the public were involved in the design, conduct and reporting of the trial.
9 Trial designType of design (parallel, factorial, cross-over, cluster, split-plot), allocation ratio, and framework (superiority, equivalence, non-inferiority).
11–12 Setting & eligibilitySettings and locations; eligibility criteria for participants and, if applicable, for sites and those delivering the interventions.
13 Intervention & comparatorSufficient detail for replication, including how and when administered.
14–15 Outcomes & harmsPre-specified primary and secondary outcomes and how they were assessed; how harms were defined and assessed.
16a–b Sample sizeHow sample size was determined; explanation of any interim analyses or stopping guidelines.
17–19 RandomisationHow the allocation sequence was generated, what type of restriction (blocking) was used, the mechanism for concealing allocation, and whether those enrolling participants had access to the sequence.
20a–b BlindingWho was blinded after assignment (participants, providers, outcome assessors); how blinding was achieved and the similarity of interventions.
21a–d Statistical methodsMethods for primary and secondary outcomes and harms; who was included in each analysis; how missing data were handled; subgroup and sensitivity analyses.
22a–b Participant flowFor each group, numbers randomly assigned, receiving intended treatment, and analysed for the primary outcome; losses and exclusions with reasons. A flow diagram is strongly recommended.
23a–b RecruitmentDates of recruitment and follow-up; reasons the trial ended or was stopped.
24a–b Intervention deliveryThe intervention and comparator as actually delivered, including adherence, and any concomitant care.
25 Baseline dataTable of baseline demographic and clinical characteristics for each group.
26 Numbers analysed, outcomes & estimationFor each group, the numbers analysed; effect size and precision (e.g., 95% CI); for binary outcomes, both absolute and relative effect sizes.
27 HarmsAll harms or unintended events in each group.
28 Ancillary analysesOther analyses (subgroup, sensitivity), distinguishing pre-specified from post hoc.
29–30 DiscussionInterpretation balanced against benefits and harms and other evidence; limitations (bias, imprecision, generalisability, multiplicity).

CONSORT extensions exist for cluster randomised trials, non-inferiority and equivalence trials, non-pharmacological treatments, herbal interventions, and pragmatic trials. They were developed for CONSORT 2010 and are used alongside the current statement until they are updated. Up-to-date references are available at www.consort-statement.org.

Key Takeaways

  • Rigorous and equal follow-up of all groups is essential. Document compliance and reasons for withdrawal; a CONSORT flow diagram is strongly recommended.
  • Intent-to-treat analysis preserves randomisation and gives a conservative, real-world estimate; per-protocol analysis estimates the effect under ideal compliance and should be reported alongside ITT, not in place of it.
  • Multiple comparisons inflate the experiment-wise error rate. The Bonferroni adjustment (α / k) is the simplest correction. Subgroup analyses should be pre-specified and tested via interaction terms; data-driven subgroup analyses generate spurious findings.
  • For prophylactic vaccines, three measures matter: direct (VEd), indirect or herd-immunity (VEind), and total (VEtot) efficacy. Estimating all three requires data from at least two populations with different coverage.
  • The CONSORT 2025 checklist of 30 items, which updates CONSORT 2010, structures both the planning and the reporting of randomised trials. Following CONSORT from the design stage through to write-up is what makes transparent reporting achievable.

The reflection below asks you to design a trial of your own, drawing on all three trial sections. After working through it and the knowledge check, the final assessment draws on all seven sections of the lesson.

Reflection

Reflection

Imagine you have been funded to design a randomised controlled trial of a new mobile-app-based behavioural intervention to reduce daily smoking among young adults aged 18–25 living in rural communities. Briefly outline how you would address: (1) the target population (the group to which you want to generalise), the source population (the accessible group from which participants are recruited), and the study sample (the participants actually enrolled); (2) the choice of allocation strategy (individual randomisation; cluster randomisation, in which whole communities or clinics are assigned together; or a factorial design, which tests two interventions at once); (3) the choice of comparator (no app, a sham app, or existing standard care); and (4) how you would handle the analysis if compliance with the app turned out to be poor, noting that intention-to-treat analysis keeps participants in the group to which they were assigned, while per-protocol analysis includes only those who complied. Identify at least one trade-off you would face and explain how you would resolve it.

Model answer(1) Populations: target = all 18–25-year-old smokers in rural Canadian communities; source = those with smartphone access in the 4–6 sampled rural regions; study sample = consenting volunteers from the source who pass screening (smoking ≥ 5/day, no current treatment). (2) Allocation: individual randomisation 1:1 stratified by community to balance regional effects, with permuted blocks of 4 within each stratum. Cluster randomisation isn't warranted because the intervention is delivered to the individual; factorial would dilute power unless a second clearly orthogonal arm is needed. (3) Comparator: sham app (matched look-and-feel, generic health content) rather than no-app control; this controls for app-engagement and Hawthorne effects. (4) Outcome: 7-day point-prevalence abstinence at 6 months, biochemically validated (cotinine in saliva or eCO breath test); secondary measure of cigarettes/day at 1, 3, 6 months. (5) Sample size: calculate from expected 25% abstinence in intervention vs. 12% in control, α=0.05, power=0.80, accounting for ~30% loss-to-follow-up. (6) Threats: selection (rural digital access varies), differential attrition (relapsers may drop out), contamination (intervention participants share app content with controls in same community), Hawthorne (knowing they're in a trial changes behaviour).

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

1. The major reason that intent-to-treat analysis is preferred as the primary analysis is that:

Correct answer: A. ITT analyses include all randomised subjects in their assigned group regardless of compliance, preserving the comparability that randomisation generated. Because real-world deployment will involve some non-compliance, ITT gives a realistic estimate of the effectiveness clinicians and policymakers can expect.

2. In a vaccine trial, the indirect vaccine efficacy (VEind) is estimated by:

Correct answer: C. VEind = (InvB − InvA) / InvB. This is the herd-immunity effect, that is, how much lower the disease risk is among unvaccinated people in the high-coverage area than among unvaccinated people in the low-coverage area.

3. With four pre-specified primary outcome comparisons and a desired family-wise error rate of 0.05, the Bonferroni-adjusted significance threshold for each comparison is:

Correct answer: B. Bonferroni-adjusted alpha = 0.05 / 4 = 0.0125. Each comparison must reach p < 0.0125 to be declared significant at the experiment-wise 0.05 level.

4. The CONSORT 2025 statement is:

Correct answer: B. CONSORT (Consolidated Standards of Reporting Trials) is a reporting guideline first published in 1996 and updated in 2001, 2010 and 2025. The 30-item CONSORT 2025 checklist covers all aspects of trial design and reporting, and the extensions for special trial types, developed for CONSORT 2010, are used alongside it. It is endorsed by major journals but is not a regulatory mandate.
Section 8

Final Review & Assessment

⏱ Estimated time: 30 minutes

Bringing It All Together

Where an earlier lesson looked backward from outcome to exposure, this lesson followed the arrow forward. An earlier section framed the cohort study as something close to a controlled trial without randomisation: pick a source population, classify exposure, follow people, and watch incidence accumulate. An earlier section then forced the central design choice, risk-based (cumulative incidence) for closed populations followed for a fixed time, or rate-based (incidence density) for open populations where person-time is the natural denominator, and connected each to its analytic toolkit (binomial / log-binomial vs. Poisson / survival).

Earlier sections zoomed in on the operational details that decide whether a cohort study survives appraisal. How exposure is scaled (dichotomous, ordinal, continuous, compound), whether it changes over time, how the induction period is handled, and how comparability is engineered through restriction, matching, and analytic control all determine whether the rate ratio you eventually report is estimating what you think it is. Blinded outcome ascertainment, careful handling of loss to follow-up, and STROBE-aligned reporting then turn a defensible design into a study other researchers can actually use.

The last three sections crossed from observational into experimental epidemiology. They set out the rationale for randomised controlled trials, the phases of clinical research, the mechanics of allocation and masking, sample size with the inflation factor for cluster trials, and the analytic and reporting standards (intention-to-treat, CONSORT 2025) that govern modern trial conduct. The shift is conceptually small, since the investigator now assigns the exposure that a cohort study can only classify, but methodologically large, because randomisation is what gives trials their causal authority.

The final reflection asks you to put the entire arc to work as a brief cohort proposal of your own; the 24-question assessment then checks the conceptual content of all seven sections directly. A later lesson will pull the camera back further to look at ecological and group-level designs (where the unit of analysis stops being the person) and the comparability and inference logic you just built will keep paying off there.

Key Takeaways from this lesson

  • Cohort studies follow exposed and unexposed people forward in time, giving you direct access to incidence, something case-control studies cannot deliver.
  • The choice between risk-based (closed cohort, cumulative incidence) and rate-based (open cohort, incidence density) designs is dictated by the source population and the follow-up structure, not by preference.
  • Person-time is the right denominator whenever follow-up is uneven or membership is dynamic; it makes Poisson and survival analyses possible.
  • Exposure must be measured on a meaningful scale, with explicit handling of induction periods and time-varying status, misclassification here flows through the whole analysis.
  • Comparability is engineered through restriction, matching, and analytic control; blinded outcome ascertainment and rigorous tracking of loss-to-follow-up protect internal validity.
  • STROBE-aligned reporting closes the loop: a cohort study is only as useful as its design, conduct, and analysis are visible to the reader.
  • Randomisation is what makes the RCT the gold standard for causal inference about interventions: it balances measured and unmeasured confounders in expectation, and trial design begins with explicit target, source, and study populations and a precisely specified intervention.
  • Allocation strategies (simple, stratified, cross-over, factorial, cluster, split-plot, multicentre) and masking (single, double, triple) each prevent specific biases; cluster trials must inflate sample size by the design effect, DEFF = 1 + (m − 1) × ICC.
  • Intention-to-treat is the primary analysis because it preserves randomisation; vaccine trials distinguish direct, indirect, and total efficacy; and the CONSORT 2025 checklist structures both planning and reporting.

Reflection

Design a brief cohort study proposal for a health question of your choice. Specify: (1) the research question and hypothesis, (2) whether the source population is open or closed, (3) whether you would use a risk-based or rate-based design and why, (4) how you would define and measure exposure (and on what scale), (5) how you would ensure comparability of groups, and (6) what analytic approach you would use.

Model answer(1) Question/hypothesis: Among adults 18–45, does sustained intake of ultra-processed food (NOVA-4) > 30% of total energy increase 10-year incidence of metabolic syndrome? (2) Source population: open cohort, community-recruited Vancouver-area adults; movers can be retained. (3) Design: rate-based with risk-set methods (Cox PH), to use person-time and handle censoring/loss cleanly. (4) Exposure measurement: 24-h dietary recalls (3 per year) coded under NOVA, then collapsed to %energy from NOVA-4; analysed continuously with restricted cubic splines, plus a pre-specified clinical threshold at 30%. (5) Comparability: baseline equivalence on income, education, ethnicity, physical activity, family history; restriction to participants without metabolic syndrome at baseline; DAG-guided adjustment for SES and physical activity (NOT for BMI, which is a mediator). (6) Analysis: Cox PH with the continuous exposure, multiple imputation for missing covariates, sensitivity analyses lagging exposure 2 years to address reverse causation, and pre-registration on OSF.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

This assessment covers all sections of this lesson. You must score 100% to complete the lesson. Review the feedback after each attempt.

Final Assessment, Cohort Studies (24 Questions)

1. The fundamental logic of a cohort study is to:

Correct answer: B. The cohort design follows disease-free subjects classified by exposure forward in time and compares disease frequency between exposed and non-exposed groups (Grimes and Schulz, 2002).

2. A cohort study most resembles which other design?

Correct answer: C. Cohort studies closely resemble controlled trials except that exposure is not randomly assigned. This similarity is often cited as an advantage for causal inference.

3. In a closed source population, all subjects:

Correct answer: A. A closed (fixed) cohort requires that all subjects be observable for the full risk period. This is the assumption that makes risk-based (cumulative incidence) designs valid.

4. In a risk-based cohort study, what is the denominator of the risk?

Correct answer: D. In risk-based designs, R1 = a1/n1 and R0 = a0/n0 (Eq 8.1). The denominator is the number of subjects, not person-time. This is only valid because every subject is observable for the full risk period.

5. In a rate-based cohort study, the denominator is:

Correct answer: B. Rate-based designs accumulate person-time at risk: I1 = a1/t1 and I0 = a0/t0 (Eq 8.2). Each subject contributes time-at-risk until they develop the disease, are lost, or the study ends.

6. Which design is best suited for studying a chronic disease with a long, lifelong risk period?

Correct answer: A. For chronic diseases like many cancers, where the risk period is lifelong and often longer than feasible follow-up, a rate-based design is preferred. Risk-based designs require all subjects to be observable for the full risk period.

7. Which is an example of a compound exposure variable?

Correct answer: D. Pack-years is a compound variable that combines duration and intensity of cigarette exposure into a single cumulative-dose measure. Luo et al (2011) used this in their breast cancer study (Example 8.7).

8. The induction period is:

Correct answer: B. The induction period is the time after exposure is completed before disease might reasonably arise. During the induction period, time-at-risk of exposed individuals should be assigned to the non-exposed group, or that experience may be discarded altogether.

9. If a subject’s exposure status changes during follow-up:

Correct answer: C. When exposure status changes, time-at-risk before the change is assigned to one category and time after the change (allowing for any lag) to the other. If they develop the disease, they are assigned to the category they were in at the time the outcome occurred.

10. Which approach to comparability is applied during analysis rather than design?

Correct answer: A. Analytic control identifies and measures confounders, then uses statistical control (Mantel-Haenszel stratification through to multivariable regression) during analysis. Restriction is applied before subject selection; matching is applied at the time of selection.

11. Which statement about exchangeability in observational studies is correct?

Correct answer: B. As Hernan (2012) notes, exchangeability cannot be empirically tested in observational studies, so we never know with certainty whether we have achieved it. This is a key reason to prefer randomised experiments when feasible.

12. Why are incident cases preferred over prevalent cases for cohort outcomes?

Correct answer: D. Including only new disease events (incidence) circumvents the reverse-causation problem from measuring prevalence and ensures that associations are not biased by duration-of-disease effects and survival bias.

13. For a rate-based cohort study where rates can be assumed reasonably constant, which model is appropriate?

Correct answer: A. Poisson regression with person-time as the offset is appropriate when rates can be assumed reasonably constant over follow-up. The Poisson coefficients give direct estimates of the incidence rate ratio. When constant rates are not tenable (e.g., long follow-up), Cox proportional hazards models are preferred.

14. Hernan (2010) describes which problem with the average hazard ratio?

Correct answer: C. Hernan (2010) points out that the average HR depends on the duration of follow-up. Period-specific HRs are also conditional on the subject not developing the outcome before time t, introducing a built-in bias as susceptible people in the exposed group are progressively depleted.

15. According to STROBE-style criteria for cohort studies, which item should be reported?

Correct answer: B. STROBE-style criteria (Tooth et al, 2005) require comprehensive reporting including numbers of participants at each stage, reasons for loss to follow-up, methods of data collection, validity of measurements, and how confounders were accounted for.

16. Randomising participants between a new treatment and a comparator is ethically defensible only when:

Correct answer: B. Freedman (1987) called this state clinical equipoise. When an effective standard treatment already exists, withholding it is hard to justify, so the standard of care becomes the comparator and the trial often takes a non-inferiority form.

17. A drug has been licensed and is now being monitored in very large patient populations over long periods to detect rare or slow-to-appear adverse effects. This stage of clinical research is:

Correct answer: D. Phase IV covers post-marketing surveillance and other post-registration studies of how a product performs in everyday practice. Phase III is the large pivotal efficacy trial that precedes licensing, Phase II gives the first efficacy signal in patients, and Phase 0 is the first-in-human study at sub-therapeutic doses.

18. Compared with broad eligibility criteria, narrow eligibility criteria in a trial generally:

Correct answer: B. A narrow set of criteria yields a more homogeneous group and more power, at the cost of how widely the results apply. The recommended balance is to use criteria that reflect the breadth of people who would receive the intervention in practice if it proves effective.

19. In a cluster randomised trial of a school-based health programme with intracluster correlation coefficient ICC = 0.02 and an average cluster size of 51 students per school, the sample-size inflation factor is:

Correct answer: D. Design effect = 1 + (m − 1) × ICC = 1 + 50 × 0.02 = 1 + 1 = 2.0. The trial requires twice the sample size of an individually randomised trial.

20. Triple-blinding adds which additional layer of masking compared with double-blinding?

Correct answer: C. In triple-blinded studies, the people analysing the data are also masked to allocation, preventing analytic decisions (subgroup definitions, outlier handling, model specification) from being subtly influenced by knowledge of which group is the new treatment.

21. The CAESAR trial randomised women having a first Caesarean section to single- or double-layer uterine closure, to closure or non-closure of the peritoneum, and to liberal or restricted use of a drain, with every combination possible. This allocation design is a:

Correct answer: C. A factorial design evaluates two or more interventions at once by assigning every combination of them; CAESAR was a 2×2×2 factorial trial. A cross-over design gives each participant both interventions in sequence, a split-plot design applies one intervention at the cluster level and another at the individual level, and a cluster design allocates whole groups.

22. The recommended primary analysis for a randomised controlled trial is:

Correct answer: A. ITT preserves the benefits of randomisation by analysing subjects in their originally assigned groups regardless of compliance. It gives a conservative, real-world estimate of effectiveness and is the recommended primary analysis. Per-protocol may be reported as a secondary analysis.

23. In the Bangladesh cholera vaccine trial, direct vaccine efficacy was only 0.14 in the high-coverage population but 0.62 in the low-coverage population. The best explanation is that:

Correct answer: B. Vaccination reduces transmission, so the unvaccinated in a high-coverage area are partly protected (interference). Their risk (1.47 per 1,000) was already far below that of the unvaccinated in the low-coverage area (7.01 per 1,000), an indirect effect of 0.79. A direct comparison inside one population is confounded by coverage, which is why vaccine trials estimate direct, indirect, and total efficacy across populations with different coverage.

24. Investigators want to know whether a trial's treatment effect differs between men and women. The recommended approach is to:

Correct answer: B. Only subgroup analyses planned a priori should be carried out, and whether an effect differs across subgroups is assessed with one overall interaction test. Data-driven subgroup searches generate spurious findings, and detecting an interaction reliably needs a sample roughly four times larger than detecting the main effect.
✦ Complete every section knowledge check and every reflection before submitting.

🎉 Congratulations!

You have completed this lesson: Cohort Studies.

You now understand the design, implementation, analysis, and reporting of cohort studies, including risk-based and rate-based designs, exposure measurement principles, comparability strategies, and STROBE reporting guidelines. You can also set up a randomised controlled trial, choose its allocation and masking strategies, analyse it by intention-to-treat, estimate vaccine efficacy, and report it with CONSORT 2025.

Earlier lessons covered the three workhorse observational designs at the level of the individual: cross-sectional, case-control, and cohort. A later lesson changes the unit of analysis. Ecological and Group-Level Studies uses populations rather than individuals as the unit of observation, a strategy that opens up routinely-collected data for epidemiology but introduces a famous interpretive trap (the ecological fallacy) you will need to recognize.