Randomized Designs for Health Services Interventions
Program Planning & Evaluation
Learning objectives for this lesson:
- Explain why random allocation supports a causal claim about a program, using the ideas of potential outcomes and exchangeability.
- Distinguish explanatory from pragmatic trials and use the PRECIS-2 tool to locate a trial design on the continuum between them.
- Define the intracluster correlation coefficient and calculate the design effect, the effective sample size and the number of clusters required for a cluster randomized trial.
- Interpret a linear mixed model analysis of a simulated cluster randomized trial, and explain why an analysis that ignores clustering understates uncertainty.
- Describe stepped-wedge, waitlist, factorial, sequential multiple assignment and non-inferiority designs, and the question each one answers.
- Assess the ethics of randomizing access to a program using clinical equipoise, the Ottawa Statement and TCPS 2.
- Distinguish intention-to-treat, per-protocol and complier average causal effect estimates, and identify the reporting guidelines that apply to a randomized evaluation.
- Assess whether a randomized design is feasible for a program and, if it is, specify the unit and scheme of randomization.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Eldridge, S., & Kerry, S. (2012). A Practical Guide to Cluster Randomised Trials in Health Services Research. Wiley; and Hayes, R. J., & Moulton, L. H. (2017). Cluster Randomised Trials (2nd ed.). CRC Press.
Explanatory and Pragmatic Trials
Learning Objectives for this section
- Explain, using the language of potential outcomes, why random allocation supports a causal claim about a program.
- Distinguish random allocation from random sampling, and describe allocation concealment and the main randomization schemes.
- Distinguish explanatory from pragmatic trials by the question each one answers and the decision each one informs.
- Use the nine domains of the PRECIS-2 tool to locate a proposed trial on the explanatory-pragmatic continuum.
1.1 The Causal Question in Program Evaluation
By this point in the course, an evaluation plan contains a program description, a needs statement, a logic model, a set of prioritized evaluation questions and an evaluation matrix. Lessons 6, 7 and 8 turn to the evaluation question that decision-makers often ask first and that is the hardest to answer well, which is whether the program caused the change that was observed. This lesson covers designs that answer the question by random allocation. Lesson 7 covers comparison-group designs, difference-in-differences and matching, and Lesson 8 covers interrupted time series, regression discontinuity and natural experiments. Those designs are the tools an evaluator uses when randomization is impossible or unacceptable.
Background: randomized trial basics
Sections 1.1 and 1.2 restate the basics of randomized trials (potential outcomes, exchangeability, random allocation, allocation concealment and the common randomization schemes) so that students who come to this course from other programs have them in one place. Students who took HSCI 230 met this material in Lesson 5, Sections 5 to 7 (Cohort Studies), which is optional review. The development beyond that treatment includes Fisher's argument for randomization as a basis for inference, the limits of the guarantee that randomization provides, minimization and constrained randomization, the PRECIS-2 tool in Sections 1.3 and 1.4, and the complier average causal effect in Section 4.
The running example is the fictional Cedar Valley Connector program, which runs through this course. The fictional Cedar Valley Health Authority in British Columbia invites primary care clinicians to refer adults aged 65 and older who screen as socially isolated or lonely to a community connector. The connector meets each person up to six times over twelve weeks, co-develops a plan with them, and links them to community groups, volunteer roles, transportation help and services. The program launched in 12 of the region's 24 primary care clinics in a first wave, and the remaining 12 clinics are scheduled to join a year later. The steering committee wants to know whether the program reduces loneliness, measured with the three-item UCLA Loneliness Scale (scored from 3 to 9, with higher scores indicating greater loneliness), and whether it changes social participation, self-rated health, and the use of emergency departments and primary care. All Cedar Valley figures in this lesson are illustrative.
Background: potential outcomes and the counterfactual
Causal questions are easiest to state in the potential outcomes framework associated with Neyman (1923) and Rubin (1974). Each older adult referred to the program has two potential six-month loneliness scores: Y(1), the score the person would have if they received the program, and Y(0), the score the same person would have if they did not. The individual causal effect is the difference Y(1) − Y(0). Holland (1986) called the impossibility of observing both quantities for the same person at the same time the fundamental problem of causal inference, because each person reveals only the potential outcome that corresponds to the condition they actually experienced. An evaluation therefore targets an average causal effect, such as the mean of Y(1) − Y(0) across the people the program serves, and estimates it by comparing groups.
Students who took HSCI 341 met these ideas in Lesson 1, Section 4 (Introduction and Causal Concepts), which is optional review; students from other programs will find a fuller introduction to the counterfactual there.
A comparison of groups estimates the average causal effect only when the groups would have had the same average outcome if neither had received the program. This condition is called exchangeability. When exchangeability fails, the observed difference between groups combines the program effect with the difference that would have existed between the groups in the absence of the program. Program evaluation texts often call this difference selection bias, and epidemiology calls it confounding; Lesson 7, Section 1 explains the two uses of the term.
Background: confounding and backdoor paths
A confounder is a common cause of receiving the program and of the outcome, or a proxy for one, that is not on the causal pathway from the program to the outcome. If clinicians refer patients whom they expect to engage, and if those patients would also have become less lonely without the program, then the clinician's judgement confounds a comparison of referred and unreferred patients. The causal diagram below draws this example, with the letters used across the course series in brackets: X for the exposure (here the program), Y for the outcome and C for the confounder.
A backdoor path is a non-causal path from the exposure to the outcome that begins with an arrow into the exposure. Here the path from referral back to the clinician's judgement and on to loneliness (X ← C → Y) is a backdoor path, and while it is open, a comparison of referred and unreferred patients mixes the program effect with the effect of the judgement. Exchangeability that holds within levels of measured covariates is called conditional exchangeability, and it corresponds to blocking every backdoor path by conditioning on measured variables. Observational analyses do this through stratification, regression, matching or weighting (Lesson 7 develops propensity scores), and the adjusted estimate is unbiased only when every backdoor path has been blocked by variables that were measured accurately. A judgement that was never recorded, such as a clinician's sense of which patients will follow through, cannot be adjusted for, so its backdoor path stays open.
HSCI 230 Lesson 7, Section 2 (Conceptualization, Measurement and Causal Specification), introduces causal diagrams, and HSCI 341 Lesson 7, Section 3 (Designing Against Bias: Validity and Confounding in Study Protocols), states the backdoor criterion in full. Both are optional reading.
Suppose the health authority chose the 12 first-wave clinics because their managers volunteered, because they had a room the connector could use, and because their electronic medical records could produce a list of older patients who had screened as lonely. Each of these reasons could plausibly be related to loneliness outcomes. Clinics with engaged managers may already run more group activities, clinics with spare rooms may sit in town centres with better transit, and clinics with stronger data systems may serve more affluent neighbourhoods. A comparison of six-month loneliness between first-wave and second-wave clinics would then mix any program effect with these differences, and no statistical adjustment could remove the parts that were never measured.
What random allocation does
Random allocation breaks the link between the factors that influence the outcome and the condition each unit receives. When a random number generator decides which clinics start first, a manager's enthusiasm, the availability of a room and the quality of a data system can no longer influence the allocation. The arms are then exchangeable in expectation: across all the allocations the randomization could have produced, the arms have the same distribution of every characteristic, measured or unmeasured, including the potential outcomes themselves. HSCI 230 Lesson 5, Section 5 (Cohort Studies), and HSCI 207 Lesson 7, Section 2 (Sampling, Recruitment and Measurement), introduce random allocation and the comparable groups it creates and are optional reading; the rest of this subsection develops Fisher's argument and the limits of the guarantee. Fisher (1935) added a second argument for randomization. Because the allocation mechanism is known, it provides a basis for inference, and a p-value can be read as a statement about how unusual the observed difference would be across the allocations that could have occurred if the program had no effect.
Three limits of this guarantee matter in program evaluation. First, the guarantee holds in expectation, so any single allocation can be imbalanced by chance. The risk of chance imbalance falls as the number of randomized units grows, which is why a trial that randomizes 24 clinics is more exposed to it than a trial that randomizes 960 people, a point that Section 2 develops. Second, randomization protects the comparison at the moment of allocation only. Events after allocation, such as dropout that differs between arms, outcome measurement that differs between arms, or the recruitment of participants by staff who already know their clinic's allocation, can reintroduce bias. Third, randomization establishes whether the program caused a difference in the people and settings where the trial was run, and whether the result applies to other people and settings remains a separate question of generalizability.
1.2 Random Allocation in Practice
Random allocation and random sampling
Random allocation and random sampling are often confused. Random sampling selects units from a defined population so that the sample represents that population, and it supports external validity. Random allocation assigns units that are already in a study to conditions, and it supports internal validity. A study can have either property without the other. Most trials randomly allocate volunteers who are a convenience sample of the eligible population, and most population surveys sample randomly without allocating anyone to anything.
Allocation concealment and blinding
Allocation concealment means that the people who enrol participants cannot know or predict the next allocation before a participant is enrolled. It differs from blinding, which concerns whether participants, staff and outcome assessors know the allocation after it has been made. Concealment can be achieved in almost every trial, whereas blinding often cannot. Schulz and colleagues (1995) found that trials with inadequate or unclear allocation concealment reported larger treatment effects than trials with adequate concealment, which is consistent with enrolment decisions being influenced by knowledge of the upcoming assignment. In program evaluation, concealment is usually achieved by having an independent statistician generate the allocation with a documented, seeded script and release each allocation only after the unit has been enrolled. In a cluster trial, the equivalent safeguard is that clinics, schools or communities commit to participate before they learn their allocation.
A community connector program cannot be delivered blind, because older adults know whether they are meeting a connector and connectors know whom they serve. Outcome assessment can still be blinded. In the Cedar Valley example, six-month loneliness could be collected by telephone interviewers who work from a central list and are not told which clinic a participant attends.
Randomization schemes
The scheme describes how the allocation sequence is generated. The choice depends on the number of units, the strength of known prognostic factors and the way units enter the trial.
Background: three basic randomization schemes
| Scheme | How it works | When it helps | Caution |
|---|---|---|---|
| Simple randomization | Each unit is allocated independently, as if by a coin toss. | Large trials with many units. | Small trials can end with unequal arms or chance imbalance. |
| Permuted blocks | Units are allocated within blocks (for example, of four or six) that contain equal numbers of each arm. | Trials that enrol people one at a time over months. | Fixed, small blocks let staff predict the last allocation in a block, so block sizes are usually varied and concealed. |
| Stratified randomization | A separate allocation sequence is used within each stratum of an important prognostic factor. | Trials with few units and one or two strong prognostic factors. | Too many strata leave each stratum nearly empty. |
HSCI 230 Lesson 5, Section 6 (Cohort Studies), introduces these schemes and is optional reading.
Trials with few units, and cluster trials in particular, often need more control over balance than these schemes provide. Three further schemes serve that purpose.
| Scheme | How it works | When it helps | Caution |
|---|---|---|---|
| Minimization | Each new unit is allocated to the arm that reduces imbalance on several factors, usually with a random element (Taves, 1974; Pocock & Simon, 1975). | Small trials that must balance several factors. | A deterministic version weakens concealment, so a random element is added. |
| Pair matching | Clusters are paired on similar characteristics, and one member of each pair is randomized to each arm. | Cluster trials with very few clusters. | The loss of one cluster removes the information from its pair. |
| Constrained randomization | Many possible allocations are generated, those that are balanced on chosen cluster characteristics are retained, and one is selected at random (Moulton, 2004). | Cluster trials in which all clusters are known before allocation. | The balance criteria should be specified in advance and reported. |
Suppose 8 of the 24 Cedar Valley clinics are rural and 16 are urban, and the clinics range from small two-physician practices to large team-based clinics. All 24 clinics are known before allocation, and the health authority wants exactly 12 clinics in each wave. Decide which scheme you would use and why.
A strong answer stratifies by rurality, randomizing 4 rural and 8 urban clinics to each wave, so that the rural-urban mix cannot differ by chance. An equally defensible answer uses constrained randomization that balances rurality, clinic size and the proportion of patients aged 75 and older, because all clusters are known in advance and three balancing factors are more than stratification can handle with 24 units.
1.3 Explanatory and Pragmatic Trials
Schwartz and Lellouch (1967) observed that trialists approach a comparison with one of two attitudes. An explanatory trial asks whether an intervention can work under ideal conditions, often in order to test a mechanism. A pragmatic trial asks whether an intervention works when it is delivered in the conditions of usual practice, in order to inform a choice between options. The distinction corresponds closely to the difference between efficacy and effectiveness, which Lesson 9 develops alongside implementation.
The two attitudes lead to different design choices. Explanatory trials select participants who are likely to respond and adhere, use expert staff and close monitoring, control co-interventions, compare the intervention with a placebo or attention control, and often measure mechanistic or intermediate outcomes. Pragmatic trials enrol the people who would receive the intervention in practice, deliver it through the usual staff and organizations, compare it with usual care, measure outcomes that matter to participants and decision-makers, and analyze everyone as randomized. Thorpe and colleagues (2009) argued that real trials sit on a continuum between the two attitudes, and that a trial can be pragmatic in some respects and explanatory in others.
For the Cedar Valley Connector program, an explanatory trial might be run by a university team that trains its own connectors, enrols older adults without cognitive impairment who are fluent in English, telephones participants between meetings to support attendance, and measures loneliness monthly. A pragmatic trial would use the health authority's own connectors, accept every older adult whom clinicians refer, compare the program with usual primary care, and measure loneliness at six months with a short telephone survey supplemented by administrative records of emergency department and primary care visits. The steering committee's decision is whether to fund the program across the region, so the pragmatic version answers the question it needs answered.
1.4 Locating a Trial with PRECIS-2
The Pragmatic Explanatory Continuum Indicator Summary (PRECIS) was introduced by Thorpe and colleagues (2009) to help trialists match design decisions to the purpose of a trial. Loudon and colleagues (2015) revised it as PRECIS-2 after consultation with trialists and testing in practice. PRECIS-2 has nine domains, each scored from 1 (very explanatory) to 5 (very pragmatic). The scores are plotted on a wheel, so that a trial whose points lie near the hub is explanatory and a trial whose points lie near the rim is pragmatic.
The tool is meant for the design stage. The trial team, ideally with the people who will use the results, scores each domain and records the reason for the score. The scores are meant to be read domain by domain, because summing them into one number would hide the pattern that the wheel is designed to show. A wheel is also neither a quality score nor a ranking. An explanatory trial is the appropriate design for a question about whether an intervention can work, and the PRECIS-2 exercise asks only whether each design choice fits the question the trial is meant to answer.
The domain asks how similar participants are to the people who would receive the program in usual care. The Cedar Valley trial enrols every adult aged 65 and older whom a clinician refers after a positive isolation or loneliness screen, and excludes only people who cannot complete a short telephone survey even with help from a family member or interpreter. The exclusion keeps the score below 5.
The domain asks how much extra effort goes into recruitment beyond what would happen in usual care. Clinicians refer during routine visits, as they would once the program is established, and a research coordinator then telephones referred patients to seek consent and collect baseline data. The added contact is modest.
The domain asks how different the trial settings are from the settings where the program would be used. The trial runs in all 24 of the region's primary care clinics, which are the clinics that would deliver the program at scale.
The domain asks how the resources, staff expertise and organization of care compare with usual care. Connectors are health authority employees with the program's standard training. The trial adds a short fidelity log that connectors complete after each meeting, which is a small addition to usual practice.
The domain asks how much flexibility connectors have in delivering the program compared with usual care. Connectors tailor each plan and offer between one and six meetings according to need, guided by the program manual. The manual sets some limits, which keeps the score from 5.
The domain asks what is done to encourage participants to adhere. The trial adds nothing beyond the program's usual reminder calls, and older adults who stop attending are not pursued for research purposes.
The domain asks how the intensity of measurement and follow-up compares with usual care. Research staff conduct baseline and six-month telephone surveys, which usual care would not include. The surveys are short, and health service use comes from administrative data, so the score sits in the middle.
The domain asks how relevant the primary outcome is to participants. Loneliness is the problem the program addresses, and the steering committee, which includes older adults with lived experience, judged a change in loneliness to be meaningful. The use of a research instrument at a fixed time point keeps the score below 5.
The domain asks to what extent all data are included in the primary analysis. The protocol specifies an intention-to-treat analysis that includes every enrolled older adult in the arm of their clinic, whether or not they engaged with a connector.
The two wheels in Figure 1.2 describe trials of the same program that would answer different questions. The explanatory trial, with university-trained connectors, narrow eligibility, adherence support and monthly follow-up, would show whether the connector model can reduce loneliness when delivered under favourable conditions. The pragmatic trial would show whether offering the program through Cedar Valley's clinics reduces loneliness among the older adults those clinics actually refer. A health authority deciding on regional funding needs the second answer. A research team testing whether social connection mediates the effect of the program on self-rated health might reasonably prefer the first.
During protocol development, a team member proposes two changes to the pragmatic Cedar Valley trial: excluding older adults who do not speak English, and telephoning every participant every two weeks to encourage attendance at connector meetings. Identify which PRECIS-2 domains each change affects and the direction in which the scores would move.
The language exclusion lowers the Eligibility score, because it removes people who would be referred in practice, and it would also reduce the relevance of the results for a region with linguistically diverse older adults. The fortnightly calls lower the Flexibility in adherence score, because they add an adherence-promoting measure that usual care would not provide, and they also lower the Follow-up score, because they add research contact. Both changes move the wheel toward the hub and make the trial less able to answer the steering committee's question.
1.5 From Individuals to Clinics
Pragmatic trials of health services programs face a recurring design problem. Programs such as the Cedar Valley Connector are delivered by and through organizations. Connectors are based in clinics, clinicians decide whom to refer, and clinic routines change once a connector is on site. If older adults within the same clinic were randomized individually, clinicians who had learned about community resources through the program would be likely to share that knowledge with patients in the usual-care arm, and patients in the same waiting room would talk with one another. This leakage, called contamination, would shrink the observed difference between arms. The usual solution is to randomize whole clinics, which is the subject of Section 2.
The CONSORT extension for pragmatic trials (Zwarenstein et al., 2008) asks authors to describe the setting, the usual-care comparator and the people and organizations involved, so that readers can judge whether the results apply to their own context. Section 4 returns to reporting and introduces the current CONSORT statement and its extensions.
Reflection
A provincial health ministry is choosing between two trial designs for a telephone befriending program for adults aged 75 and older who live alone. In Design A, a university team recruits 200 volunteers who pass a cognitive screen and speak English, trains its own befrienders, telephones participants weekly to encourage them to keep their befriending calls, measures loneliness monthly for six months, and analyzes only participants who completed at least eight calls. In Design B, the ministry's existing befriending service accepts every referral from family physicians in three health regions, delivers calls with its usual volunteers and schedule, compares the program with usual care, measures loneliness once at six months by telephone, and analyzes all randomized participants. The ministry must decide whether to fund the service across the province. PRECIS-2 scores nine domains (eligibility, recruitment, setting, organization, flexibility in delivery, flexibility in adherence, follow-up, primary outcome and primary analysis) from 1, very explanatory, to 5, very pragmatic. Score three domains for each design, explain each score, and state which design better fits the ministry's decision and why.
For eligibility, Design A scores about 1 or 2, because it enrols volunteers who pass a cognitive screen and speak English, whereas Design B scores about 5, because it accepts every physician referral, which is the population the service would serve. For flexibility in adherence, Design A scores about 1, because weekly encouragement calls are an adherence measure that the service would never provide, while Design B scores 5, because nothing is added beyond usual practice. For primary analysis, Design A scores 1, because it analyzes only participants who completed eight calls, and Design B scores 5, because it analyzes everyone as randomized.
Design B fits the ministry's decision. The ministry needs to know what happens when the existing service is offered to the people physicians refer, delivered as it would be in practice. Design A answers a different question, whether befriending can reduce loneliness among adherent, cognitively intact volunteers under favourable conditions, and its estimate would probably overstate the effect the province would see. A strong answer could also score setting (A about 2, B 5) or follow-up (A about 1 or 2 with monthly measurement, B about 4), with the same conclusion.
Minimum 20 characters required.
Question 1: Why does random allocation of clinics to the Cedar Valley program's two waves support a causal comparison?
Question 2: A trial team has an independent statistician generate a seeded allocation sequence, and each clinic learns its allocation only after signing its participation agreement. Which safeguard does this procedure provide?
Question 3: Which description best fits a pragmatic trial of the Cedar Valley Connector program?
Question 4: In PRECIS-2, a trial that telephones participants every two weeks to encourage attendance at program sessions would score toward the explanatory end of which domain?
Cluster Randomized Trials
Learning Objectives for this section
- Explain why health services programs are often randomized by clinic, school or community, and describe what cluster randomization costs.
- Define the intracluster correlation coefficient and calculate the design effect and the effective sample size.
- Calculate the number of clusters needed for a cluster randomized trial with a continuous outcome.
- Interpret a linear mixed model for a cluster trial, and explain why a model that ignores clustering understates the standard error of the program effect.
- Compare cluster-level, mixed-model and generalized estimating equation analyses, and identify the reporting items in the CONSORT extension for cluster trials.
2.1 Why Programs Are Randomized by Cluster
In a cluster randomized trial, intact groups of people, such as clinics, schools, workplaces, long-term care homes or communities, are randomized, and outcomes are measured on the individuals within them. The design is also called a group-randomized trial (Murray, 1998). Standard texts include Donner and Klar (2000), Eldridge and Kerry (2012), and Hayes and Moulton (2017). Health services programs are often evaluated with this design because the program is delivered through an organization, and randomizing individuals within that organization would be impractical or would contaminate the comparison.
Cluster randomization has costs. People in the same cluster tend to resemble one another, so each additional person in a cluster adds less new information than an independently sampled person, and the trial needs more participants in total. With few clusters, chance imbalance between arms is more likely. When participants are identified and recruited after clusters learn their allocation, staff in intervention clusters may recruit different people from staff in control clusters, a problem called identification or recruitment bias that Puffer, Torgerson and Watson (2003) documented in published cluster trials. Cluster trials also raise distinctive questions about consent, which Section 3 discusses. Most importantly for analysis, the clustering must be reflected in the statistical model. Cornfield (1978) described randomization by cluster followed by an analysis that treats individuals as independent as a form of self-deception, because it produces confidence intervals that are too narrow and p-values that are too small.
2.2 The Intracluster Correlation Coefficient
Older adults who attend the same clinic share a catchment area, a set of local community resources, the same transit routes, and clinicians with the same referral habits. Their loneliness scores are therefore likely to be more similar to one another than to the scores of older adults at other clinics. The intracluster correlation coefficient (ICC, often written with the Greek letter rho) measures this similarity.
Intracluster correlation coefficient
ICC = σ2between / (σ2between + σ2within)
Here σ2between is the variance of the true cluster means around the overall mean, and σ2within is the variance of individual outcomes around their own cluster's mean. The ICC is the proportion of the total variance in the outcome that lies between clusters. It can equivalently be read as the correlation between the outcomes of two randomly chosen individuals from the same cluster.
An ICC of 0 means that cluster membership tells us nothing about an individual's outcome, and an ICC of 1 means that everyone in a cluster has the same outcome. ICCs for patient-level outcomes in primary care are usually small (Adams et al., 2004), often below 0.05, although they vary by outcome and by the type of cluster. Process-of-care measures, which depend directly on clinic practices, tend to show larger ICCs than outcomes such as loneliness or self-rated health, which depend mostly on the individual. The simulated Cedar Valley data used later in this section have an ICC near 0.06, toward the upper end of what is typical for a patient-reported outcome, which makes the consequences of clustering easy to see.
Planners take the ICC from a pilot study, from routine data on the same clusters, or from published trials with similar clusters and outcomes. An ICC estimated from few clusters is imprecise, so calculations should be checked across a plausible range, and trial reports should state the ICC they observed so that later trials can plan with better information.
2.3 The Design Effect and the Effective Sample Size
The design effect is the factor by which the variance of an estimate from a cluster design exceeds the variance from an individually randomized design with the same total number of participants. For clusters of equal size m, the design effect is given by the formula below, which follows Kish (1965) in survey sampling and Donner, Birkett and Buck (1981) for cluster trials. HSCI 230 Lesson 5, Section 6 (Cohort Studies), and HSCI 410 Lesson 5, Section 1 (Modelling Dependent Data), introduce the design effect at undergraduate level and are optional reading.
Design effect and effective sample size
DEFF = 1 + (m − 1) × ICC
Effective sample size = N / DEFF
With unequal cluster sizes, Eldridge, Ashby and Kerry (2006) give DEFF = 1 + ((CV2 + 1) × m − 1) × ICC, where m is the mean cluster size and CV is the coefficient of variation of cluster size (the standard deviation of cluster sizes divided by their mean).
Notation: DEFF is the design effect, m the cluster size (the mean size when sizes vary), N the total sample size and ICC (also written ρ) the intracluster correlation coefficient. Reliability studies use the intraclass correlation coefficient, a different application of the same kind of statistic (HSCI 410 Lesson 7).
The effective sample size is the number of independent observations that would give the same precision as the clustered sample. Suppose the Cedar Valley trial enrols 40 older adults in each of 24 clinics (N = 960) and the ICC is 0.06. The design effect is 1 + (40 − 1) × 0.06 = 1 + 2.34 = 3.34, and the effective sample size is 960 / 3.34 = 287.4. Nine hundred and sixty clustered observations carry about as much information about the program effect as 287 independent ones.
Two features of the formula drive most practical decisions. First, the design effect depends on the product of the ICC and the cluster size, so an ICC that looks negligible can still produce a large design effect when clusters are large. An ICC of 0.01 with 200 participants per cluster gives a design effect of 2.99. Second, the design effect grows in a straight line with cluster size, so each additional participant per cluster buys less precision than the one before, as Figure 2.1 shows. Variation in cluster size raises the design effect further, because large clusters contribute disproportionately to the estimate.
2.4 Sample Size for a Cluster Randomized Trial
The sample size for a cluster trial with a continuous outcome is calculated in three steps. The first step calculates the number of participants per arm that an individually randomized trial would need. The second step multiplies that number by the design effect. The third step divides the inflated number by the cluster size to obtain the number of clusters per arm, rounding up at each step.
Sample size for a two-arm cluster trial, continuous outcome
nindividual per arm = 2 × (z1−α/2 + z1−β)2 × SD2 / Δ2
ncluster trial per arm = nindividual × DEFF
Clusters per arm = ncluster trial / m, rounded up. Here Δ is the smallest difference worth detecting, SD is the standard deviation of the outcome, α is the two-sided significance level and 1 − β is the power.
For Cedar Valley, suppose the steering committee judges that a difference of 0.5 points on the three-item UCLA Loneliness Scale would be large enough to justify regional funding, and that the standard deviation of six-month scores is about 1.5 points. With a two-sided α of 0.05 (z = 1.960) and 80 percent power (z = 0.842), an individually randomized trial needs 2 × (1.960 + 0.842)2 × 1.52 / 0.52 = 141.3 participants per arm, which rounds up to 142. With 40 participants per clinic and an ICC of 0.06, the design effect is 3.34, so the cluster trial needs 141.3 × 3.34 = 471.9, or 472 participants per arm. Dividing by 40 gives 11.8, which rounds up to 12 clinics per arm. Under these assumptions, the 24 clinics in the Cedar Valley region are just enough.
The number of clusters matters more than the number of participants per cluster. As the cluster size grows without limit, the number of clusters per arm approaches nindividual × ICC (Hemming et al., 2011). For Cedar Valley, 141.3 × 0.06 = 8.5, so at least 9 clinics per arm would be needed even if every clinic could enrol thousands of older adults. If fewer clusters are available than this minimum, no amount of recruitment within clusters will deliver the planned power, and the evaluator must accept a larger detectable difference or choose another design.
Several practical adjustments follow. Planners inflate the sample size for expected loss to follow-up and for the possibility that a cluster withdraws. Because a trial with few clusters is analyzed with a t distribution with limited degrees of freedom, a calculation based on normal quantiles slightly understates the clusters required, and planners often add at least one cluster per arm. Adjustment for a baseline measure of the outcome usually increases power. The calculation should also be repeated across a plausible range of ICCs and cluster sizes, as the worked example in Section 2.5 does.
2.5 Worked Example: Analyzing the Simulated Cedar Valley Cluster Trial
The worked example uses a simulated dataset that represents a hypothetical randomized version of the Cedar Valley rollout. Imagine that, before launch, the health authority had randomized the 24 clinics so that 12 joined the program in wave one and 12 continued usual care during year one before joining in wave two. Older adults referred at each clinic completed the UCLA Loneliness Scale at enrolment and six months later. The program indicator, called arm, equals 1 for wave-one clinics and 0 for wave-two clinics. The dataset also records whether each older adult in a wave-one clinic engaged with the program (attended at least two connector meetings), which Section 4 uses. The data are simulated, so the results describe this teaching dataset only.
Background: reading a regression coefficient
The worked example compares regression models, so it helps to recall how one is read. A linear regression of the six-month loneliness score on the binary program indicator, arm, predicts each person's score as an intercept plus a coefficient times arm. The intercept is the predicted mean score in the usual-care arm (arm = 0), and the coefficient for arm is the difference in mean score between the Connector arm and the usual-care arm, so a coefficient of −0.43 means that Connector-arm participants scored 0.43 points lower on average. The standard error of the coefficient describes how much the estimate would vary from sample to sample, and the 95 percent confidence interval is approximately the estimate plus or minus 1.96 standard errors, or plus or minus a t quantile times the standard error when the degrees of freedom are few. An interval that excludes zero corresponds to a two-sided p-value below 0.05. Adding the baseline loneliness score as a second predictor adjusts the comparison, so that the coefficient for arm compares participants with the same baseline score. Randomization has already balanced the arms in expectation, so the main gain from baseline adjustment in a trial is precision: the baseline score explains part of the variation in six-month scores, which usually narrows the confidence interval, and it also corrects for any chance imbalance at baseline.
HSCI 410 Lesson 3 (Linear and Logistic Regression) covers linear regression more fully and is optional reading.
Background: random intercepts and the empty mixed model
A linear mixed model (also called a multilevel or random-effects model) adds a random intercept for each cluster to an ordinary regression. Each clinic's mean score may sit above or below the overall mean by an amount that the model treats as a draw from a normal distribution with its own variance. The model therefore describes two sources of variation: the spread of clinic means around the overall mean (the between-clinic variance, σ2between) and the spread of individuals around their own clinic's mean (the within-clinic variance, σ2within). An empty model, also called a null or intercept-only model, contains only the overall intercept and the random intercept, so it splits the total variance of the outcome into these two parts without explaining any of it. The ICC is the between-clinic share, σ2between / (σ2between + σ2within), which is the formula in Section 2.2. When predictors such as arm or baseline loneliness are added, the variances that remain are called residual variances, and the ICC computed from them is a residual ICC.
HSCI 410 Lesson 5, Sections 2 and 3 (Modelling Dependent Data), covers random-intercept models and their extension to binary outcomes and is optional reading.
The clusters. The trial has 12 clinics in each arm, with 485 older adults in the usual-care arm and 475 in the Connector arm, for a total of 960. Clinics enrolled between 32 and 50 older adults, with a mean of 40 and a coefficient of variation of 0.164. Baseline mean scores are almost identical in the two arms (6.136 in the usual-care arm and 6.124 in the Connector arm), as randomization leads us to expect, while the six-month mean is 6.056 in the usual-care arm and 5.621 in the Connector arm.
The intracluster correlation. An empty linear mixed model, which contains only an intercept and a random intercept for clinic, splits the variance of the six-month score into a between-clinic part and a within-clinic part. The between-clinic variance is 0.1387 and the within-clinic variance is 2.2105, so the ICC is 0.1387 / (0.1387 + 2.2105) = 0.0590. About 6 percent of the variation in six-month loneliness lies between clinics. Because wave-one and wave-two clinics differ in their program status, part of that between-clinic variation is the program effect itself. Once arm is added to the model, the ICC falls to 0.0422. For planning a future trial, the ICC from the usual-care arm, from baseline data or from a model that includes arm is usually preferred, because it describes the clustering that exists before the program creates any difference.
The design effect and the effective sample size. With a mean of 40 older adults per clinic and an ICC of 0.0590, the design effect is 3.30 if cluster sizes were equal and 3.37 once their variation (a coefficient of variation of 0.164) is taken into account. The 960 older adults therefore carry about as much information as 285 independent observations. Variation in cluster size of this magnitude adds little to the design effect. The adjustment matters more when cluster sizes vary widely, for example when a few large clinics enrol most of the participants.
Sample size. The planning values are a smallest important difference of 0.5 points, an ICC of 0.06 (the estimate above, rounded) and a standard deviation of 1.5, which rounds up the standard deviation of 1.415 observed in the usual-care arm to be conservative. The three-step calculation from Section 2.4 gives 142 per arm for an individually randomized trial, a design effect of 3.34, 472 per arm for the cluster trial and 12 clinics per arm. A coefficient of variation of 0.15 in clinic size raises the design effect to 3.39 and the total to 480 per arm, still with 12 clinics. The table below repeats the calculation of clinics needed per arm across four ICCs and four cluster sizes, keeping the same difference, standard deviation and power.
| ICC | 10 per clinic | 20 per clinic | 40 per clinic | 80 per clinic |
|---|---|---|---|---|
| 0.01 | 16 | 9 | 5 | 4 |
| 0.03 | 18 | 12 | 8 | 6 |
| 0.05 | 21 | 14 | 11 | 9 |
| 0.08 | 25 | 18 | 15 | 13 |
The table shows the diminishing return from larger clusters. With an ICC of 0.05, doubling cluster size from 40 to 80 reduces the clinics needed per arm from 11 to 9, and with an ICC of 0.08 it reduces them from 15 to 13, while the total number of participants rises sharply.
Ignoring clustering compared with modelling it. Two models regress the six-month score on arm and the baseline score. The naive model, an ordinary linear regression, treats all 960 older adults as independent. The linear mixed model adds a random intercept for clinic, and its confidence interval and p-value use a t distribution with 22 degrees of freedom, which is the 24 clinics minus the 2 cluster-level parameters (the between-within approach).
| Model | Program effect (points) | Standard error | 95% confidence interval |
|---|---|---|---|
| Naive linear regression | −0.428 | 0.084 | −0.594 to −0.263 |
| Linear mixed model with a random intercept for clinic | −0.420 | 0.136 | −0.701 to −0.139 |
The two models agree closely about the size of the program effect. The naive model estimates that older adults in Connector clinics scored 0.428 points lower at six months, and the mixed model estimates 0.420 points lower, after adjustment for baseline loneliness. The models disagree about precision. The mixed-model standard error is 1.609 times the naive standard error, and the mixed-model p-value is 0.0052 on 22 degrees of freedom.
The standard errors differ because the program indicator is constant within each clinic. Every older adult in a clinic has the same value of arm, so the information about the program effect comes from comparing 12 clinics with 12 clinics, and older adults in the same clinic share a clinic-level deviation that does not average out within the clinic. The naive model assumes that the 960 residuals are independent and therefore credits the data with far more independent information than they contain. The mixed model estimates the between-clinic variance (0.0685) separately from the within-clinic variance (1.6418), and its standard error for arm reflects the number of clinics as well as the number of people. The squared ratio of the standard errors (2.588) is close to the design effect implied by the residual ICC of 0.040 (2.605), which confirms that the design effect is the quantity the naive model ignores. The residual ICC is smaller than the 0.0590 from the empty model because baseline loneliness and arm explain part of the difference between clinics, which is one reason baseline adjustment improves the precision of cluster trials.
A clustered standard error. An alternative keeps the naive point estimate and replaces its standard error with a sandwich estimator that allows residuals to be correlated within clinics. The clustered sandwich standard error is 0.1313, close to the mixed-model value of 0.136 and far from the naive 0.084. When its p-value is computed from a t distribution with the 957 residual degrees of freedom of the linear model, it is 0.0011, which overstates the information available. With 22 degrees of freedom the p-value is 0.0036. Sandwich estimators also tend to underestimate the variance when there are few clusters, so trials with few clusters (a common rule of thumb is fewer than about 40) should use a small-sample correction such as the one proposed by Mancl and DeRouen (2001). The clustered sandwich approach is closely related to the generalized estimating equations described next.
2.6 Choosing an Analysis for a Cluster Trial
Three families of analysis respect clustering. Each is valid when used appropriately, and the choice depends on the number of clusters, the type of outcome and whether the target is a cluster-specific or a population-averaged effect.
Background: cluster-specific and population-averaged effects
A mixed model estimates a cluster-specific effect, which compares the expected outcome with and without the program for people in the same clinic, or in clinics with the same random intercept. Generalized estimating equations estimate a population-averaged effect, which compares the average outcome across the whole population of clinics with and without the program. For a continuous outcome in a linear model the two coincide, whereas for a binary outcome analyzed on the odds ratio scale the cluster-specific effect lies further from the null. Both approaches need a standard error that respects clustering. A sandwich standard error is computed from the observed spread of the residuals within clusters: the model supplies the estimate, and the variance is then calculated from how much the clusters' residuals actually vary, so it remains valid even when the assumed correlation within clusters is wrong. Each cluster contributes one piece of information about that spread, so the sandwich estimator needs an adequate number of clusters, and with fewer than about 40 it tends to underestimate the variance unless a small-sample correction is applied.
HSCI 410 Lesson 5, Section 3 (Modelling Dependent Data), compares mixed models and generalized estimating equations and is optional reading.
The simplest valid analysis computes one summary per cluster, such as each clinic's mean six-month score, and compares the 12 Connector means with the 12 usual-care means using a t-test with 22 degrees of freedom. Covariates are handled in two stages: an individual-level regression without the arm indicator produces residuals, which are averaged within clusters and compared between arms (Hayes & Moulton, 2017). Cluster-level methods perform well with few clusters and are easy to explain to decision-makers.
Linear mixed models, such as the mixed model in the worked example in Section 2.5, include a random effect for each cluster and estimate a cluster-specific program effect. Generalized linear mixed models extend the approach to binary and count outcomes, such as an emergency department visit within six months. For a continuous outcome, cluster-specific and population-averaged effects coincide, but for a binary outcome analyzed with a logistic model the cluster-specific odds ratio lies further from 1 than the population-averaged one. With few clusters, tests should use small-sample degrees of freedom, such as the Satterthwaite or Kenward and Roger (1997) approximations, or the between-within approach used in the worked example.
Generalized estimating equations (Liang & Zeger, 1986) estimate a population-averaged effect, which answers the policy question of how much the program changes outcomes across the region. The analyst specifies a working correlation structure, usually exchangeable within clusters, and the sandwich standard errors remain valid even if that structure is misspecified, provided there are enough clusters. With few clusters, a bias-corrected sandwich estimator (Mancl & DeRouen, 2001) and a t reference distribution are advisable.
Many health services trials randomize fewer than 40 clusters. Leyrat and colleagues (2018) compared analysis methods in this situation. Methods that rely on large-sample approximations could produce inflated type I error rates, whereas cluster-level analyses, and mixed models or generalized estimating equations with small-sample corrections, generally kept the error rate close to its nominal level. The method and its correction should be specified in the protocol before outcome data are seen.
2.7 Reporting a Cluster Trial
The CONSORT extension for cluster randomized trials (Campbell et al., 2012) adds cluster-specific items to the CONSORT checklist. Section 4 introduces the current CONSORT statement, and the table below lists the cluster-specific requirements that bear most directly on sample size, randomization and analysis.
| Report element | What the cluster extension asks for | Cedar Valley illustration |
|---|---|---|
| Title and rationale | Identification of the trial as cluster randomized, and the reason for using a cluster design. | The connector is based in the clinic and clinicians refer, so individual randomization would risk contamination. |
| Eligibility | Eligibility criteria for clusters and for individuals. | All 24 primary care clinics; referred adults aged 65 and older who screen as isolated or lonely. |
| Sample size | The method of calculation, including the assumed ICC, cluster size and any allowance for variation in cluster size. | ICC 0.06, 40 per clinic, smallest important difference 0.5 points, 12 clinics per arm. |
| Randomization and recruitment | The unit and method of randomization, and whether individuals were identified and recruited before or after clusters were randomized, and by whom. | Clinics randomized by an independent statistician; referred older adults approached by a research coordinator. |
| Consent | From whom consent was sought (representatives of the cluster, individual members, or both) and whether it was sought before or after randomization. | Clinic agreement before randomization; individual consent for the research surveys. |
| Analysis and results | How clustering was accounted for, the number of clusters and individuals at each stage, and the estimated ICC for each primary outcome. | Linear mixed model with a random clinic intercept; ICC reported with the primary result. |
Section 3 extends the logic of cluster randomization over time. When a program will reach every clinic eventually, the order in which clinics start can itself be randomized, which is the idea behind the stepped-wedge design.
Reflection
A health authority plans a cluster randomized trial of a pharmacist-led medication review for older adults, randomizing community pharmacies. The primary outcome is a continuous medication appropriateness score with a standard deviation of 2.0 points, and the smallest difference worth detecting is 0.6 points. The trial will use a two-sided significance level of 0.05 (z = 1.96) and 80 percent power (z = 0.84). Each pharmacy is expected to enrol 25 older adults, and a pilot study suggests an intracluster correlation coefficient (ICC) of 0.02. Use these formulas: participants per arm for an individually randomized trial = 2 × (1.96 + 0.84)2 × SD2 / difference2; design effect = 1 + (m − 1) × ICC, where m is the number per cluster; pharmacies per arm = (participants per arm × design effect) / m, rounded up. (a) Calculate the design effect, the number of older adults per arm and the number of pharmacies per arm. (b) A colleague proposes analyzing the trial with ordinary linear regression of the score on the arm indicator. Explain what would go wrong and what you would do instead.
(a) An individually randomized trial would need 2 × (2.80)2 × 2.02 / 0.62 = 2 × 7.84 × 4 / 0.36 = 174.2, or 175 older adults per arm. The design effect is 1 + (25 − 1) × 0.02 = 1 + 0.48 = 1.48. The cluster trial therefore needs 174.2 × 1.48 = 257.8, or 258 older adults per arm, and 258 / 25 = 10.3, so 11 pharmacies per arm (22 in total). Using the rounded 175 gives 259 per arm and the same 11 pharmacies.
(b) The arm indicator is constant within each pharmacy, so the information about the effect comes from 22 pharmacies, and older adults in the same pharmacy share a pharmacy-level deviation. Ordinary regression would treat about 550 older adults as independent, so its standard error would be too small by a factor of roughly the square root of 1.48, about 1.22. Confidence intervals would be too narrow and the type I error rate would be inflated. I would instead fit a linear mixed model with a random pharmacy intercept and small-sample degrees of freedom (about 20), or compare pharmacy-level means with a t-test, and I would report the ICC observed in the trial.
Minimum 20 characters required.
Question 1: In the simulated Cedar Valley trial, the empty mixed model gave a between-clinic variance of 0.1387 and a within-clinic variance of 2.2105. What is the intracluster correlation coefficient?
Question 2: A cluster trial will enrol 30 participants in each cluster, and the intracluster correlation coefficient is 0.05. What is the design effect?
Question 3: An individually randomized trial would need 141.3 participants per arm, and the intracluster correlation coefficient is 0.06. About how many clusters per arm would a cluster trial need even if every cluster could enrol an unlimited number of participants?
Question 4: In the worked example in Section 2.5, the naive linear model gave a standard error of 0.084 for the program effect and the mixed model gave 0.136. Which explanation is correct?
Stepped-Wedge, Factorial and Adaptive Designs
Learning Objectives for this section
- Describe the stepped-wedge cluster randomized design and explain why its analysis must model time trends.
- Show how a staggered rollout, such as the Cedar Valley program's two waves, can be randomized as a parallel waitlist comparison or as a stepped-wedge trial.
- Describe factorial designs and sequential multiple assignment randomized trials, and the questions each one answers.
- Explain the logic of a non-inferiority trial, including the choice of margin and the role of the analysis population.
- Assess the ethics of randomizing access to a program using clinical equipoise, the Ottawa Statement and TCPS 2.
3.1 Randomizing a Staggered Rollout
Many programs cannot start everywhere at once. Staff must be hired and trained, budgets are released in stages, and managers prefer to learn from early sites before expanding. When the order in which sites start is not fixed by need or readiness, randomizing that order costs the program very little and produces a randomized comparison. Evaluators who join a program early can often negotiate this, and it is one of the most practical routes to a randomized design in health services.
The fictional Cedar Valley Connector program, which runs through this course, has exactly this structure: 12 clinics in a first wave and 12 a year later. Figure 3.1 shows two ways the rollout could have been randomized. In panel A, the health authority randomizes which 12 clinics join in wave one. During year one, the wave-two clinics act as a waitlist control (also called a delayed-intervention control), and the comparison between the two sets of clinics during that year is a parallel cluster randomized trial. This is the design analyzed in the Section 2 worked example. In panel B, the rollout is spread across the first year in four steps of six clinics each, with the order of the steps randomized. This is a stepped-wedge design.
If outcome data are also collected before wave one and during year two, the two-wave design in panel A becomes a design with two sequences and three periods, in which every clinic is observed both without and with the program. With only two sequences, most of the information about the program effect still comes from the year-one comparison between arms, but the before-and-after data within clinics add information, and they also bring the problem of time trends that dominates the analysis of stepped-wedge trials.
3.2 The Stepped-Wedge Cluster Randomized Design
In a stepped-wedge design, all clusters begin in the control condition. At regular intervals, called steps, a randomly selected group of clusters crosses over to the intervention, and by the end of the trial every cluster has received it. Outcomes are measured in every cluster in every period (Hussey & Hughes, 2007; Hemming et al., 2015). Each cluster therefore contributes observations under both conditions, and the program effect is estimated from comparisons between clusters within periods and from comparisons within clusters over time. Copas and colleagues (2015) distinguish designs by how participants are recruited and followed: a closed cohort identified at the start and followed throughout, an open cohort in which people join and leave over time, and continuous recruitment of new participants who are each exposed for a short time. A Cedar Valley stepped-wedge trial that enrols the older adults referred in each period would recruit continuously, with each person contributing outcomes to the period in which they were referred.
The design is attractive for health services programs for several reasons. Every cluster eventually receives the program, which makes the design acceptable to managers and communities who would refuse to be a permanent control. The phased rollout matches the way programs are often implemented. Because every cluster is observed before and after it starts, cluster-level differences contribute less to the error of the estimate. Whether a stepped-wedge design needs fewer clusters than a parallel design depends on the ICC, the cluster size and the number of steps, so the comparison should be made for the specific trial using appropriate methods (Hussey & Hughes, 2007; Hooper et al., 2016) or simulation.
Time trends and the analysis model
The defining analytical problem of the stepped-wedge design is that exposure to the program is confounded with calendar time. In the first period no clusters are exposed, and in the last period all are. Any change in the outcome over the course of the trial that is unrelated to the program will therefore look like a program effect unless the analysis separates the two. Loneliness among older adults could change over a year because of the seasons, a provincial initiative on social isolation, a change in transit service or a public health emergency. The standard analysis, proposed by Hussey and Hughes (2007), includes a fixed effect for each period alongside the program indicator and a random effect for each cluster.
The Hussey and Hughes model for a stepped-wedge trial
Yijt = μ + βt + θ Xjt + uj + eijt
Yijt is the outcome for individual i in cluster j in period t, βt is the fixed effect of period t, Xjt equals 1 if cluster j has started the program by period t and 0 otherwise, θ is the program effect, uj is a random cluster effect and eijt is individual error. The period effects absorb secular trends that are common to all clusters, so θ is estimated from the contrast between exposed and unexposed clusters within the same periods.
The model makes three assumptions that deserve scrutiny in a program evaluation. It assumes that the secular trend is the same in every cluster, so a trend that differs between rural and urban clinics would bias the estimate. It assumes that the program effect is immediate and constant, whereas a connector program may take months to build caseloads and community partnerships, and recent methodological work has shown that assuming an immediate, constant effect can bias the estimate when the effect actually changes with time since the program started. It also assumes, in its simplest form, that the correlation between observations in the same cluster is the same whether they come from the same period or from periods a year apart, and extensions allow that correlation to weaken over time. Each assumption can be relaxed, at a cost in precision, and the chosen model should be specified in the protocol.
Two practical risks also arise. Clusters may not start on the randomized date, because hiring or training is delayed, which blurs the contrast between periods. Measurement must also continue in every cluster in every period, which is more burdensome than a parallel design. The CONSORT extension for stepped-wedge trials (Hemming et al., 2018) asks authors to explain why the design was chosen, to show the schedule of sequences and periods in a diagram, and to describe how time effects were modelled in the analysis.
A non-randomized staggered rollout raises the same analytical issue in a sharper form. Lesson 7 shows that two-way fixed effects models applied to staggered adoption can give misleading estimates when program effects vary across groups or over time. The stepped-wedge design shares the staggered structure but adds randomization of the order, which protects the comparison from selection on readiness.
Clinics are randomized to start immediately or after a fixed delay, and the arms are compared during the delay. For Cedar Valley, 12 clinics would start in wave one and 12 would wait a year. The comparison is simple to analyze with the methods of Section 2 and is not confounded with time, because both arms are observed over the same year. Its limitation is that the comparison ends when the waitlist arm starts, so it cannot estimate effects beyond the delay period.
Clinics are randomized to the order in which they cross over at several steps, and every clinic is measured in every period. For Cedar Valley, four sequences of six clinics would start at three-month intervals. The design uses within-clinic as well as between-clinic comparisons and gives every clinic the program within the trial period, but the analysis must model time trends, and the results depend on assumptions about how the effect changes with exposure time.
The health authority chooses the order of rollout on readiness or need. The evaluator can still compare early and late clinics using the difference-in-differences and event-study methods of Lesson 7, but the credibility of the estimate then rests on the assumption that early and late clinics would have followed parallel trends without the program, which randomization would have made unnecessary.
3.3 Waitlist Controls
A waitlist control can be used with individuals or with clusters. It is common when a program has more applicants than places, because people on the waitlist would be waiting in any case and the trial simply determines the order of service at random. It is also acceptable to participants and partners, since everyone is promised the program.
The design has three limitations. First, it can estimate effects only for the length of the delay, so a six-month waitlist cannot show whether benefits last a year. Second, people who know they will receive the program soon may behave differently from people receiving usual care with no prospect of the program. They may postpone seeking other help or report disappointment at being made to wait, either of which can make the program look more effective than it would against usual care. Third, people may leave the waitlist before follow-up, and if those who leave differ from those who stay, the comparison is biased. Evaluators can reduce these problems by measuring outcomes for everyone regardless of whether they remain on the list, by describing to waitlisted participants what usual care includes, and by keeping the delay as short as the program's capacity allows.
3.4 Factorial Designs
A factorial design randomizes participants or clusters to two or more interventions at once, so that every combination is represented. In a 2 × 2 factorial design, each unit is randomized to receive or not receive intervention A and, independently, to receive or not receive intervention B. Suppose the Cedar Valley Health Authority wants to test both the connector program and a transportation voucher for older adults who live far from community activities.
| No transportation voucher | Transportation voucher | |
|---|---|---|
| Usual care | Cell 1: neither component | Cell 2: voucher only |
| Connector program | Cell 3: connector only | Cell 4: connector and voucher |
The main effect of the connector program compares cells 3 and 4 with cells 1 and 2, and the main effect of the voucher compares cells 2 and 4 with cells 1 and 3. If the two components do not interact, every participant contributes to both comparisons, so the factorial trial answers two questions with roughly the sample size needed for one. The efficiency depends on the absence of an important interaction. If vouchers help only when a connector has first identified an activity worth travelling to, the effect of the voucher depends on the presence of the connector, and the main effect averaged over both connector conditions becomes harder to interpret. Detecting an interaction as large as a main effect requires about four times the sample size needed to detect the main effect, so most factorial trials are powered to estimate main effects and can only describe interactions approximately. Factorial experiments are also central to the multiphase optimization strategy, in which investigators screen several candidate components of a program before assembling and testing the final package (Collins et al., 2007).
3.5 Adaptive Designs and Sequential Multiple Assignment Randomized Trials
An adaptive design allows pre-specified changes to a trial in response to accumulating data, without undermining its validity (Pallmann et al., 2018). Common adaptations include stopping early for benefit or futility at planned interim analyses, re-estimating the sample size, dropping arms that are performing poorly, and changing allocation ratios to favour better-performing arms. Platform trials extend the idea by adding and removing arms under a single master protocol. Adaptive designs require the adaptation rules, the interim analysis schedule and the methods that control the type I error rate to be specified in advance, and they usually need an independent data monitoring committee. The Adaptive designs CONSORT Extension (ACE; Dimairo et al., 2020) describes how to report them.
A sequential multiple assignment randomized trial (SMART) answers a different kind of question. Many programs are adaptive by nature: staff change what they offer according to how a person responds. A SMART randomizes participants at more than one decision point so that investigators can construct and compare adaptive interventions, which are sequences of decision rules that specify what to offer, to whom and when (Murphy, 2005; Collins et al., 2007).
In the hypothetical SMART in Figure 3.2, referred older adults are first randomized to in-person or telephone connector meetings. At week six, the connector records whether each person is engaged, defined as having attended at least two meetings and taken a first step on their plan. Engaged participants continue with their assigned format. Participants who are not engaged are randomized a second time to receive either a peer volunteer companion or transportation help. The design embeds four adaptive interventions, such as "start with telephone meetings and add a peer volunteer for those not engaged by week six", and allows them to be compared. Typical primary aims compare the two first-stage options, or compare the two second-stage options among people who were not engaged.
3.6 Non-Inferiority Trials
Most trials ask whether a new option is better than an existing one. A non-inferiority trial asks whether a new option that is cheaper, easier to scale or more acceptable is not worse than the existing option by more than a pre-specified amount, called the non-inferiority margin. For Cedar Valley, the question might be whether connector meetings delivered by telephone, which would let one connector serve more clinics, are non-inferior to in-person meetings for six-month loneliness.
The margin is the largest loss of effect that decision-makers would accept in exchange for the advantages of the new option. It should be justified on clinical grounds and on evidence about how much the existing option improves outcomes compared with no program. The margin must be smaller than that established effect, or the trial could declare non-inferior an option that is no better than usual care. In the simulated Section 2 trial, in-person delivery lowered six-month scores by about 0.42 points relative to usual care. A margin of 0.5 points would therefore be indefensible, while a margin of 0.2 points would preserve roughly half of the established benefit. The trial concludes non-inferiority if the upper limit of the confidence interval for the difference (telephone minus in-person, where higher scores are worse) lies below the margin.
Two further features distinguish non-inferiority trials. First, an intention-to-treat analysis is not conservative here. When participants in both arms fail to engage, the arms look more alike, which favours a conclusion of non-inferiority. The CONSORT extension for non-inferiority and equivalence trials (Piaggio et al., 2012) therefore recommends reporting both intention-to-treat and per-protocol analyses, and a conclusion of non-inferiority is more credible when both support it. Second, the trial cannot show that the existing option worked in this particular study, so it relies on an assumption that the effect of in-person delivery is similar to its effect in the earlier trial that established it. Because margins are small relative to the variability of the outcome, non-inferiority trials usually need larger samples than superiority trials of the same outcome.
3.7 The Ethics of Randomizing Access to a Program
Randomizing access to a health program requires an ethical justification. Freedman (1987) located the justification for clinical trials in clinical equipoise, a state of honest professional disagreement in the expert community about the comparative merits of the options being tested. For health services programs, the corresponding condition is genuine uncertainty, shared by informed people, about whether the program improves outcomes relative to usual care, or about whether it does so at an acceptable cost.
Scarcity strengthens the case for randomization. When a program cannot reach everyone at once, some people will wait whether or not there is a trial, and a lottery gives each eligible person or clinic an equal chance of being served first. Randomizing the order of a rollout that must be staggered anyway withholds nothing that would otherwise have been provided. The argument weakens if the program is already known to be effective, or if the trial delays the rollout beyond what capacity requires.
Cluster trials raise further questions, which the Ottawa Statement on the ethical design and conduct of cluster randomized trials addresses (Weijer et al., 2012). The Statement asks investigators to identify who the research participants are, a group that can include people who receive the intervention, people whose environment or care is deliberately altered, people with whom investigators interact, and people about whom identifiable data are collected. It asks that informed consent be sought from research participants unless a research ethics board approves a waiver or alteration, which generally requires that the research be impracticable without it and involve no more than minimal risk. It distinguishes gatekeepers, such as a clinic manager or a health authority executive, who may grant permission for a cluster to take part, from individuals, on whose behalf gatekeepers cannot consent. In Canada, TCPS 2 sets out the conditions for altering consent requirements and, in Chapter 9, the obligations for research involving First Nations, Inuit and Métis peoples.
The evaluator should document the evidence on similar programs and the views of clinicians, partners and older adults. Evidence that social prescribing programs vary widely in their effects can support a claim of uncertainty, while a well-replicated effect in comparable settings would weaken it.
The Cedar Valley program has funding for 12 clinics in its first year. Randomizing which clinics start first changes who waits, and it does not change how many wait.
Older adults who complete the surveys are research participants. Clinicians whose referral practices the program changes, and connectors whose meetings are logged for the fidelity assessment, may also be research participants under the Ottawa Statement's definition, and the research ethics board will need to decide what consent each group requires.
Clinic managers can agree to their clinic's participation before randomization. Older adults can consent to the research surveys after referral. Whether the program itself, offered as a health service, requires research consent is a question for the research ethics board, which may approve an alteration of consent if the conditions in TCPS 2 are met.
The Cedar Valley program has an Indigenous health partnership with local First Nations. Decisions about randomizing clinics that serve First Nations communities, and about the collection and governance of data from First Nations participants, belong with that partnership under TCPS 2 Chapter 9 and the OCAP® principles discussed in Lesson 4.
Some older adults referred for loneliness will have cognitive impairment. The protocol should describe how capacity is assessed, when an authorized third party may give consent, and how the person's own wishes are respected.
The waitlist or later-wave clinics should receive the program as promised, and the results should be shared with participating clinics, older adults and partners in an accessible form.
At a public meeting, a municipal councillor argues that the clinics with the longest waitlists for social services should receive the Cedar Valley program first, and that a lottery is unfair to them. The evaluator responds that the health authority could group clinics into strata of need and randomize within each stratum, using a higher probability of starting first in the highest-need stratum (for example, six of eight high-need clinics in wave one). Need would then shape the rollout, the comparison would remain randomized within each stratum, and the analysis would compare arms within strata. If the steering committee instead decides that need alone must determine the order, the evaluator will plan a quasi-experimental evaluation, such as the difference-in-differences and regression discontinuity designs in Lessons 7 and 8, and will document that choice and its consequences for the strength of the evidence.
Section 4 turns to the analysis and reporting of randomized evaluations, and to the structured judgement an evaluator makes about whether randomization is feasible for a program at all.
Reflection
A regional health authority will introduce a falls-prevention exercise program into 16 long-term care homes. Training capacity allows the program to start in 4 homes every 4 months, so the full rollout will take 16 months. Falls in long-term care vary with the season, and a provincial falls-prevention campaign is scheduled to begin partway through the rollout. All 16 homes are known in advance and have not yet been told when they will start. Propose a randomized design that fits this rollout, describe how the analysis would handle time trends, and identify two ethical or practical considerations the evaluator should raise with the health authority.
The rollout fits a stepped-wedge cluster randomized design. After a four-month baseline period in which no home has the program, the 16 homes would be randomized into four sequences of four homes, and one sequence would start at each four-month step, giving five periods in all. Falls would be measured in every home in every period, ideally from incident reports that homes already collect. Randomization could be stratified by home size so that each sequence contains a similar mix.
Because the proportion of homes with the program rises over time, seasonal variation and the provincial campaign would look like program effects unless the analysis separates them. The analysis would therefore include a fixed effect for each period, the program indicator and a random effect for each home, following Hussey and Hughes. I would also examine whether the effect grows with time since a home started, because exercise programs take time to build participation.
Two considerations are consent and fidelity of the schedule. Many residents have cognitive impairment, administrators can permit their home's participation but cannot consent for residents, and the research ethics board may need to consider an alteration of consent for the use of routine falls data. Homes may also start late if training is delayed, which would blur the comparison, so the plan should record actual start dates.
Minimum 20 characters required.
Question 1: Why must the analysis of a stepped-wedge trial include period effects?
Question 2: A 2 × 2 factorial trial tests a connector program and a transportation voucher. What is the main advantage of the design if the two components do not interact?
Question 3: A non-inferiority trial will compare telephone with in-person connector meetings. In-person meetings lower six-month loneliness scores by about 0.42 points compared with usual care. Which non-inferiority margin is most defensible?
Question 4: According to the Ottawa Statement, what may a clinic manager acting as a gatekeeper do in a cluster randomized trial?
Analysis, Reporting and Feasibility
Learning Objectives for this section
- Distinguish intention-to-treat, per-protocol and as-treated analyses, and explain what each one estimates.
- Calculate a complier average causal effect under one-sided noncompliance and state the assumptions it requires.
- Describe how the current CONSORT statement and its extensions for cluster and stepped-wedge trials govern the reporting of a randomized evaluation.
- Apply a structured set of questions to judge whether randomization is feasible for a program.
- Specify the unit and scheme of randomization for a program evaluation.
4.1 When People Do Not Take Up the Program
Randomization determines what each unit is offered. In a program evaluation, many people who are offered a program do not take it up, and some who are not offered it find a similar service elsewhere. In the simulated Cedar Valley trial from Section 2, some older adults referred at wave-one clinics never met a connector or stopped after one meeting. The analysis has to decide how to treat them, and the decision changes the question the analysis answers.
| Analysis | Who is compared | What it estimates | Main weakness |
|---|---|---|---|
| Intention-to-treat | Everyone, in the arm to which they (or their cluster) were randomized, regardless of what they received. | The effect of offering the program, which is the effect of the policy decision. | The effect among people who take up the program is diluted by those who do not. |
| Per-protocol | Only people who followed the protocol in each arm, such as engaged participants in the Connector arm and all participants in the usual-care arm. | An effect among adherent participants, if adherence were unrelated to prognosis. | Adherence is usually related to prognosis, so the comparison is no longer randomized. |
| As-treated | People grouped by what they actually received, regardless of assignment. | An effect of receipt, if receipt were unrelated to prognosis. | The comparison is observational and open to confounding. |
The intention-to-treat principle analyzes every randomized participant in the arm to which they were allocated. It preserves the comparison that randomization created and answers the question a health authority asks when it decides whether to offer a program: what happens to the population that is offered it, given that not everyone will participate. The intention-to-treat estimate is the primary analysis in almost all pragmatic trials.
The ICH E9(R1) addendum on estimands (International Council for Harmonisation, 2019) asks trialists to define the target of estimation precisely, including how events that occur after randomization, such as non-engagement, are handled. In its terms, an intention-to-treat analysis follows a treatment-policy strategy, in which non-engagement is part of what is being evaluated. The addendum describes other strategies, including the principal stratum strategy that underlies the complier average causal effect described below. Writing the estimand down before the trial begins keeps the evaluator and the steering committee clear about which question the primary analysis answers.
Intention-to-treat analysis also requires outcome data on everyone randomized. When some six-month surveys are missing, the evaluator should report how many are missing in each arm, use a principled method such as multiple imputation that respects the clustering, and run sensitivity analyses that test whether plausible departures from the imputation assumptions would change the conclusion.
4.2 The Complier Average Causal Effect
Decision-makers often want a second number: the effect of the program on the people who actually take part. The per-protocol analysis does not supply it, because engaged and non-engaged participants differ in ways that also affect outcomes. Angrist, Imbens and Rubin (1996) showed how to estimate an effect among participants who take up a program while keeping the protection of randomization. Their approach classifies participants by how they would behave under each possible assignment.
The groups are called principal strata. Membership is unobserved for most individuals: an older adult in a usual-care clinic could be a complier or a never-taker, and nothing in the data reveals which. Because randomization balances the strata between arms, their proportions in the usual-care arm are expected to equal their proportions in the Connector arm. When usual-care participants cannot obtain the program, a situation called one-sided noncompliance, the proportion of compliers equals the proportion of Connector-arm participants who engaged.
Complier average causal effect with one-sided noncompliance
CACE = ITT effect / proportion of compliers
The estimator requires four assumptions: random assignment; the exclusion restriction, under which assignment affects outcomes only through engagement; monotonicity, under which there are no defiers; and no interference between participants. Under these assumptions, never-takers contribute nothing to the intention-to-treat difference, so the whole difference is produced by compliers.
The logic is easiest to see with round numbers. Suppose the intention-to-treat effect is −0.40 points and half of the Connector-arm participants engaged. If never-takers experience no effect, the −0.40 average must come from the engaged half, which implies an effect of −0.80 among compliers. The CACE is the instrumental variable estimate with randomization as the instrument, an idea that Lesson 8 develops in the context of natural experiments.
The exclusion restriction deserves attention in a cluster trial. If clinicians in Connector clinics begin to ask every older patient about loneliness, or if the clinic starts hosting community group activities, older adults who never meet a connector may still be affected by their clinic's allocation. Never-takers would then contribute to the intention-to-treat difference, and dividing that difference by the proportion engaged would attribute their change to compliers, overstating the complier effect.
In the simulated trial from Section 2.5, older adults in the Connector arm who engaged were less lonely at enrolment (mean 5.887) than those who did not (mean 6.573), so engagement was selective. The proportion engaged in the Connector arm is 0.655. The intention-to-treat estimate, from the linear mixed model in Section 2.5, is −0.420 points. A per-protocol mixed model, which drops the Connector-arm participants who did not engage, gives −0.772 points. The CACE divides the intention-to-treat estimate by the proportion engaged (−0.420 / 0.655, computed from the unrounded values) and equals −0.642 points.
Because the data are simulated, the true values are known. The simulation built in a reduction of about 0.62 points for older adults who engaged and no effect for those who did not, and it made engagement more likely among people with lower baseline loneliness and greater mobility, a characteristic that the dataset does not record. The CACE recovers the built-in effect closely. The per-protocol estimate overstates it, even after adjustment for baseline loneliness, because engaged participants differ from the whole usual-care arm in mobility, which also lowers loneliness. The CACE is a point estimate here; its confidence interval requires instrumental variable methods or a bootstrap that resamples clinics.
For the Cedar Valley steering committee, the intention-to-treat estimate is the primary result, because the decision is whether to offer the program. The CACE is a useful secondary result for program managers who want to know what engagement achieves, provided its assumptions are stated. The per-protocol estimate belongs, at most, among the sensitivity analyses, except in a non-inferiority trial, where Section 3 explained why it is reported alongside the intention-to-treat analysis.
4.3 Reporting Randomized Evaluations
The CONSORT (Consolidated Standards of Reporting Trials) statement sets out the minimum information a report of a randomized trial should contain, in the form of a checklist and a flow diagram. CONSORT 2025 (Hopewell et al., 2025) is the current statement and updates CONSORT 2010. Design-specific extensions add items for particular designs. The extensions listed below were developed for CONSORT 2010 and are used alongside the current statement for design-specific items, and authors should check whether an updated version of a relevant extension has been published. Protocols are reported with the SPIRIT 2025 guideline (Chan et al., 2025), which updated SPIRIT 2013, trials should be registered in a public registry such as ClinicalTrials.gov or the ISRCTN registry before the first participant is enrolled, and Lesson 10 introduces TIDieR for describing the intervention itself.
| Design feature | Reporting guideline | What it adds |
|---|---|---|
| Any randomized trial | CONSORT 2025 (Hopewell et al., 2025) | The core checklist and the participant flow diagram. |
| Clusters randomized | Cluster extension (Campbell et al., 2012) | The rationale for clustering, clustering in sample size and analysis, flow of clusters and individuals, and the ICC. |
| Staggered crossover of clusters | Stepped-wedge extension (Hemming et al., 2018) | A diagram of sequences and periods, and the handling of time effects. |
| Usual-care setting and comparator | Pragmatic trials extension (Zwarenstein et al., 2008) | A description of the setting, participants and comparator that lets decision-makers judge applicability. |
| Non-inferiority question | Non-inferiority and equivalence extension (Piaggio et al., 2012) | The margin and its justification, and both intention-to-treat and per-protocol results. |
| Pre-planned adaptations | ACE (Dimairo et al., 2020) | The adaptation rules, interim analyses and methods for controlling error rates. |
| Behavioural or service intervention | Nonpharmacologic treatments extension (Boutron et al., 2017) | Details of intervention delivery, providers and centres. |
A cluster trial's flow diagram reports two levels, clusters and the individuals within them, at each stage from allocation to analysis. Figure 4.2 shows the structure for the simulated Cedar Valley trial, whose follow-up was complete by construction.
4.4 Deciding Whether Randomization Is Feasible
Randomization gives the strongest protection against confounding, but an evaluator recommends it only after judging that it is feasible, ethical and useful for the decision at hand. The questions below structure that judgement. They build on the evaluability assessment introduced in Lesson 4 and lead, when the answers are unfavourable, to the quasi-experimental designs of Lessons 7 and 8.
A decision about whether to fund, expand or end a program usually needs an estimate of its effect. A decision about how to improve delivery may be better served by process evaluation (Lesson 5) or quality improvement methods (Lesson 8).
Randomization is justified when informed people disagree about whether the program improves outcomes relative to usual care. If the effect is already well established in similar settings, a trial adds little and may be difficult to defend ethically.
Timing is the most common barrier. Once a program has been offered to everyone, or once sites have been chosen, the opportunity to randomize who receives it first has passed. Evaluators who are involved before launch have far more options.
Staggered rollouts, oversubscribed programs, new funding that cannot cover every site, and new components added to an existing program all create points at which allocation can be randomized at little cost.
The unit of randomization should be the smallest unit that avoids contamination and matches how the program is delivered. Individuals, clinicians, clinics and communities are all possible units.
A cluster trial needs enough clusters to deliver adequate power and to make chance imbalance unlikely. With three clusters per arm, there are only 20 possible allocations, so a randomization test can never produce a two-sided p-value below 0.10. The minimum number of clusters per arm, nindividual × ICC, sets a floor that larger clusters cannot lower.
Managers, clinicians, Indigenous partners, older adults and elected officials must find the allocation fair. Randomizing the order of a rollout, stratifying by need and promising the program to every site usually increase acceptance.
The same instruments, schedule and data sources must apply in every arm, and results must arrive before the decision they are meant to inform. Administrative data can reduce burden and cost when they capture the outcomes of interest.
The research ethics board, data stewards and Indigenous partners must approve the consent approach, data flows and governance arrangements. Section 3 listed the specific questions for cluster trials.
A randomized evaluation costs more to plan and manage than a routine monitoring report. The evaluator should name the quasi-experimental design that would be used if randomization proves infeasible, so the decision is made with a clear sense of what is gained.
When randomizing the program itself is not possible, related randomized designs often remain available. An encouragement design randomizes invitations, reminders or help to enrol in a program that is open to everyone, and the CACE then estimates the effect among people whom the encouragement moved to participate. Evaluators can also randomize the timing of access, an enhancement to the program, such as the transportation voucher in Section 3, or the order of service among applicants who exceed program capacity.
4.5 Worked Example: The Cedar Valley Randomization Assessment
A randomization feasibility assessment judges whether a randomized design is feasible for a program and, if it is, specifies the unit and scheme of randomization. The worked example below shows what that assessment could look like for the fictional Cedar Valley Connector program, written as if the evaluator joined the planning team six months before wave one, when the wave-one clinics had not yet been chosen.
Decision and question. The steering committee will decide, after the first year, whether to continue the program and how to configure it for the region. The primary evaluation question is whether offering the program to referred older adults reduces loneliness on the three-item UCLA Loneliness Scale at six months, compared with usual primary care.
Feasibility judgement. Randomization is feasible. Evidence on community connector programs is mixed, so informed people disagree about the likely effect. Funding covers 12 clinics in year one, so half of the clinics will wait regardless of the evaluation, and the order of entry has not been fixed. The program is delivered through clinics, so individual randomization would risk contamination through clinicians and waiting rooms. With 24 clinics of about 40 referred older adults each and an ICC of 0.06, a parallel comparison has 80 percent power to detect a 0.5-point difference (Section 2.4). Outcomes can be collected identically in all clinics by central telephone interviewers within the first year.
Unit and scheme. The unit of randomization is the clinic. The 24 clinics will be allocated 1:1 to wave one or wave two, stratified by rurality (8 rural and 16 urban clinics, giving 4 rural and 8 urban clinics per wave). An independent statistician outside the program team will generate the allocation with a seeded computer script after all 24 clinics have signed participation agreements, and the allocation will be revealed at a steering committee meeting. To prevent recruitment bias, the loneliness screen will be introduced in all 24 clinics before randomization, and a research coordinator who works from referral lists will approach eligible older adults in both arms in the same way.
Analysis and reporting. The primary analysis will be an intention-to-treat linear mixed model with a random clinic intercept, adjusted for baseline loneliness and the rurality stratum, with 21 degrees of freedom (24 clinics minus the intercept, the arm indicator and the stratum indicator). A CACE for engagement, defined as attending at least two connector meetings, will be a secondary analysis. The report will follow CONSORT 2025 and the cluster extension.
Ethics and partners. The Indigenous health partnership will review the randomization plan before it is finalized, including whether clinics that serve First Nations communities are included in the lottery, and the data governance arrangements for First Nations participants will follow the partnership's agreements. Clinic managers will give permission for their clinics' participation, and older adults will give consent for the research surveys. The research ethics board will decide whether the program itself requires research consent.
If wave one had already been chosen. If the health authority had already selected the wave-one clinics, the parallel comparison would no longer be randomized. The evaluator could still randomize the start dates of the 12 wave-two clinics across year two, in three steps of four clinics, which would create a small stepped-wedge trial within wave two, and would analyze the comparison between wave-one and wave-two clinics with the difference-in-differences methods of Lesson 7.
What a randomization feasibility assessment contains
A feasibility assessment works through the ten questions in Section 4.4, states whether a randomized design is feasible and justifies the judgement. If it is feasible, the assessment specifies the unit of randomization, the allocation ratio, the randomization scheme (including any stratification or constraint), how allocation will be concealed, the number of units and the planning values behind it, and the primary analysis. If it is not feasible, the assessment names the barrier and the randomized or quasi-experimental alternative, the designs that Lessons 7 and 8 describe.
A sound assessment lets the judgement follow from the program's decision context, timing and delivery structure, and it matches the unit of randomization to the level at which the program is delivered so that contamination is addressed. It shows any sample size or cluster calculation, with sourced planning values for the ICC and the standard deviation, and it addresses ethical and partner considerations, including Indigenous governance where relevant, specifically. It is written clearly enough for a steering committee to act on.
Reflection
A cluster randomized trial of a peer-support program for family caregivers randomized 20 community centres, 10 to the program and 10 to usual services. Caregivers in usual-service centres had no access to peer support. At six months, the intention-to-treat estimate of the program effect on a caregiver burden score (higher scores mean greater burden) was −2.4 points. In program centres, 60 percent of caregivers attended at least three peer-support sessions. A per-protocol analysis comparing attending caregivers in program centres with all caregivers in usual-service centres gave −5.1 points. As part of the program, staff at program centres also received training on recognizing caregiver burden. The complier average causal effect (CACE) with one-sided noncompliance equals the intention-to-treat effect divided by the proportion of compliers, and it assumes that assignment affects outcomes only through attendance. (a) Calculate the CACE. (b) Explain why the per-protocol estimate differs from it. (c) Explain how the staff training could threaten an assumption behind the CACE. (d) State which estimate should be primary for a decision about funding the program, and why.
(a) The CACE is −2.4 / 0.60 = −4.0 points, the estimated effect among caregivers who would attend if offered the program.
(b) The per-protocol estimate of −5.1 compares attenders with all usual-service caregivers, including those who would not have attended had they been offered the program. Caregivers who attend are likely to differ from non-attenders, for example in their available time, social support or initial burden, and these differences also affect burden at six months. The per-protocol comparison is therefore no longer randomized and is likely to exaggerate the effect, while the CACE keeps the randomized comparison.
(c) The CACE assumes the exclusion restriction: assignment to a program centre affects burden only through attending sessions. If trained staff recognize and respond to burden among all caregivers, including those who never attend, part of the −2.4 difference comes from non-attenders. Dividing by 0.60 then attributes that change to attenders, and the CACE overstates their benefit.
(d) The intention-to-treat estimate of −2.4 should be primary, because the funding decision concerns offering the program to all caregivers at a centre, including those who will not attend, and it is the estimate that randomization fully protects. The CACE can be reported as a secondary result with its assumptions stated.
Minimum 20 characters required.
Question 1: In the simulated Cedar Valley trial, the intention-to-treat estimate was −0.420 points and 65.5 percent of Connector-arm participants engaged. Under one-sided noncompliance, what is the complier average causal effect, to two decimal places?
Question 2: Why did the per-protocol estimate (−0.772) overstate the effect among engaged participants in the simulated trial?
Question 3: Which approach to reporting a stepped-wedge evaluation of the Cedar Valley program is correct?
Question 4: A health authority has already offered a new program in every clinic in its region. Which feasibility question most clearly rules out randomizing access to the program itself?
Final Assessment
Bringing It All Together
This lesson has examined how random allocation supports causal claims about health services programs and how the design of a randomized evaluation must follow the way a program is delivered. Section 1 showed that randomization makes the arms exchangeable in expectation and that pragmatic trials, located with the PRECIS-2 wheel, answer the questions decision-makers ask about programs in usual care. Section 2 showed why programs such as the fictional Cedar Valley Connector program are usually randomized by clinic, how the intracluster correlation coefficient and the design effect determine the number of clusters needed, and why an analysis that ignores clustering reports more precision than the data contain.
Section 3 turned staggered rollouts into randomized designs, including waitlist and stepped-wedge trials, and introduced factorial, adaptive, SMART and non-inferiority designs together with the ethics of randomizing access. Section 4 distinguished the effect of offering a program from the effect of taking part in it, introduced the current CONSORT statement and its extensions, and set out the questions an evaluator asks before recommending randomization. The worked example in Section 4.5 applies those questions to the Cedar Valley Connector program.
Key Takeaways from this lesson
- Random allocation makes the arms exchangeable in expectation, which protects a comparison from measured and unmeasured confounders at the moment of allocation.
- Allocation concealment protects enrolment from knowledge of upcoming assignments and is achievable even when blinding is not.
- Pragmatic trials answer whether offering a program in usual practice improves outcomes, and PRECIS-2 helps a team check that each design choice fits that purpose.
- Programs delivered through clinics, schools or communities are usually randomized by cluster to avoid contamination and to match how the program is delivered.
- The design effect, 1 + (m − 1) × ICC, shows how much clustering inflates the variance, and the number of clusters matters more than the number of participants per cluster.
- In the simulated Cedar Valley trial, a naive model and a mixed model gave similar effect estimates, but the naive standard error (0.084) was much smaller than the mixed-model standard error (0.136).
- Stepped-wedge designs randomize the order of a phased rollout, and their analysis must include period effects because exposure is confounded with calendar time.
- Factorial designs, SMARTs, adaptive designs and non-inferiority trials each answer a distinct question, and a non-inferiority margin must be smaller than the established effect of the standard option.
- Intention-to-treat estimates the effect of offering a program, while the complier average causal effect estimates the effect among those who take part under stated assumptions, and per-protocol comparisons are prone to selection bias.
- The feasibility of randomization depends most often on timing, the number of units available and the acceptability of the allocation to partners and communities.
Core Concepts Reviewed
Section 1: potential outcomes, exchangeability, random allocation versus random sampling, allocation concealment, randomization schemes, explanatory and pragmatic trials, and the PRECIS-2 wheel.
Section 2: cluster randomization and contamination, the intracluster correlation coefficient, the design effect and effective sample size, sample size for cluster trials, mixed models, generalized estimating equations and the CONSORT extension for cluster trials.
Section 3: randomized staggered rollouts, stepped-wedge designs and time trends, waitlist controls, factorial designs, adaptive designs and SMARTs, non-inferiority margins, and the ethics of randomizing access.
Section 4: intention-to-treat, per-protocol and as-treated analyses, principal strata and the complier average causal effect, CONSORT 2025 and its extensions, feasibility questions and the randomization feasibility assessment.
The final reflection asks you to bring the lesson together by recommending a randomized design for a program with a phased rollout.
Reflection
You are the evaluator for a new provincial program that offers home-based social prescribing to adults aged 70 and older after discharge from hospital. Because staff must be hired and trained, the program will start in 30 hospitals over 18 months, with 10 hospitals starting every 6 months. The hospital discharge teams make the referrals. Each hospital discharges about 50 eligible older adults in each 6-month period. Pilot data suggest a standard deviation of 1.5 points on the three-item UCLA Loneliness Scale (scored 3 to 9) and an intracluster correlation coefficient of about 0.03 among patients of the same hospital. The ministry wants to know whether offering the program reduces loneliness three months after discharge, and it will decide on permanent funding once the rollout is complete. The order in which hospitals start has not yet been set. In 200 to 300 words, recommend a design. State whether randomization is feasible and why, the unit and scheme of randomization, how the analysis will address clustering and time, which estimate will be primary and which secondary, and which reporting guidelines apply.
Randomization is feasible. The rollout must be staggered because of hiring, the order of hospitals has not been set, informed people are uncertain whether the program reduces loneliness, and the ministry needs an estimate of the effect of offering it. Randomizing the order changes who waits and does not change how many wait.
The unit of randomization should be the hospital, because discharge teams make the referrals and could not offer the program to some patients and withhold it from others without contamination. I would use a stepped-wedge design with a six-month baseline period and three steps, randomizing the 30 hospitals into three sequences of 10, stratified by hospital size or health authority. An independent statistician would generate the allocation after all hospitals agree to take part. With about 50 patients per hospital per period and an ICC of 0.03, the design effect within a period is about 1 + 49 × 0.03 = 2.47, so power should be calculated with a stepped-wedge method or by simulation.
The analysis would be a linear mixed model with a random hospital effect, fixed period effects to absorb secular trends, the program indicator and baseline loneliness, using small-sample degrees of freedom, with a sensitivity analysis for an effect that changes with time since a hospital started. The intention-to-treat estimate would be primary, and a complier average causal effect for patients who engage would be secondary. Reporting would follow CONSORT 2025 and the stepped-wedge extension.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: Which pairing of a design feature and its purpose is correct?
Question 2: A cluster trial was planned with 40 participants per cluster and an intracluster correlation coefficient of 0.06. If the investigators instead enrolled 80 participants per cluster, what would happen to the design effect?
Question 3: An evaluator wants to know whether offering the Cedar Valley program in usual care reduces loneliness. Which combination fits that purpose?
Question 4: In the Section 2 worked example, the squared ratio of the mixed-model and naive standard errors (2.588) was close to which quantity?
Question 5: Why might a stepped-wedge design be chosen over a parallel cluster trial for a program that every site will eventually receive?
Question 6: Which situation would violate the exclusion restriction needed for a complier average causal effect in a cluster trial?
Question 7: With three clusters per arm, what is the smallest two-sided p-value that a randomization test can produce?
Question 8: Which analysis of a cluster trial with 24 clusters and 960 participants is least defensible?
Question 9: What distinguishes a sequential multiple assignment randomized trial from a standard two-arm trial?
Question 10: In a non-inferiority trial, why is an intention-to-treat analysis on its own insufficient?
Question 11: The fictional Cedar Valley Health Authority wants the highest-need clinics to start the program first but also wants a randomized comparison. Which approach meets both aims?
Question 12: Which statement about the intracluster correlation coefficients in the Section 2 worked example is correct?
Question 13: A program has 600 eligible applicants for 300 places. Which randomized design is the most natural fit?
Question 14: Which pairing of a design and its reporting guideline is correct?
Question 15: An evaluator is assessing a program that will be delivered by school nurses in 10 schools and has not yet started. Which specification best follows the lesson's guidance?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
This glossary defines the terms, tools and people introduced in Lesson 6, in the order of the lesson's main themes.