HSCI 826 · Lesson 6

Randomized Designs for Health Services Interventions

Program Planning & Evaluation

Learning objectives for this lesson:

  • Explain why random allocation supports a causal claim about a program, using the ideas of potential outcomes and exchangeability.
  • Distinguish explanatory from pragmatic trials and use the PRECIS-2 tool to locate a trial design on the continuum between them.
  • Define the intracluster correlation coefficient and calculate the design effect, the effective sample size and the number of clusters required for a cluster randomized trial.
  • Interpret a linear mixed model analysis of a simulated cluster randomized trial, and explain why an analysis that ignores clustering understates uncertainty.
  • Describe stepped-wedge, waitlist, factorial, sequential multiple assignment and non-inferiority designs, and the question each one answers.
  • Assess the ethics of randomizing access to a program using clinical equipoise, the Ottawa Statement and TCPS 2.
  • Distinguish intention-to-treat, per-protocol and complier average causal effect estimates, and identify the reporting guidelines that apply to a randomized evaluation.
  • Assess whether a randomized design is feasible for a program and, if it is, specify the unit and scheme of randomization.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Eldridge, S., & Kerry, S. (2012). A Practical Guide to Cluster Randomised Trials in Health Services Research. Wiley; and Hayes, R. J., & Moulton, L. H. (2017). Cluster Randomised Trials (2nd ed.). CRC Press.

Lesson 6 · HSCI 826

Randomized Designs for Health Services Interventions

A short guided orientation before you work through the lesson at your own pace.

Program Planning & Evaluation
The running case

The Cedar Valley Connector program

The fictional Cedar Valley Health Authority refers lonely or isolated adults aged 65 and older from primary care to a community connector.

24 clinics

Twelve clinics start in wave one and twelve join a year later.

Up to six meetings

Connectors co-develop a plan and link people to groups, volunteering and transport.

Loneliness outcome

The three-item UCLA Loneliness Scale is scored from 3 to 9.

The road map

Four sections

1 · Explanatory and pragmatic trials

Why randomization supports causal claims, and how PRECIS-2 locates a trial.

2 · Cluster randomized trials

The intracluster correlation, the design effect, sample size and analysis.

3 · Stepped-wedge and other designs

Staggered rollouts, factorial, adaptive and non-inferiority designs, and ethics.

4 · Analysis, reporting and feasibility

Intention-to-treat, complier effects, CONSORT and the decision to randomize.

What you will work on

Two worked examples

  • You will follow the analysis of a simulated cluster trial of 960 older adults in 24 clinics.
  • You will see how the intracluster correlation, the design effect and the clusters needed for a trial are estimated.
  • You will compare a naive regression with a linear mixed model.
  • Section 4 shows a feasibility assessment that says whether the Cedar Valley program can be randomized and, if so, how.
How to use this lesson

Listen, read, then check

Each section opens with a narrated walkthrough like this one, followed by the full reading, a reflection and a knowledge check.

  • Use the play and arrow controls to move through the slides.
  • Open the transcript whenever you would rather read than listen.
Section 1 of 5

Explanatory and Pragmatic Trials

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Explanatory and Pragmatic Trials

Why random allocation supports a causal claim, and how to match a trial to the decision it informs.

The counterfactual

Two potential outcomes, one observed

Individual causal effect
\[ Y_i(1) - Y_i(0) \]

Each referred older adult has a loneliness score with the program and a score without it. Only one of them is ever observed, so evaluations estimate average effects by comparing groups.

What randomization does

Breaking the link between readiness and allocation

Chosen by readiness

Manager enthusiasm, space and data systems decide which clinics start first, and each may also affect loneliness.

Chosen at random

A random process decides, so the arms are exchangeable in expectation on measured and unmeasured factors.

Chance imbalance, events after allocation and generalizability remain open questions.

Allocation in practice

Concealment and randomization schemes

Allocation concealment keeps the next assignment hidden until a unit is enrolled, and it is achievable even when blinding is not.

Simple
Permuted blocks
Stratified
Minimization
Pair matching
Constrained
Two attitudes

Explanatory and pragmatic trials

Explanatory

The trial asks whether the program can work under ideal conditions, with selected participants, expert staff and close monitoring.

Pragmatic

The trial asks whether the program works in usual practice, with broad eligibility, usual staff, a usual-care comparator and analysis as randomized.

Schwartz & Lellouch (1967); Thorpe et al. (2009)

PRECIS-2

Nine domains on a wheel

Eligibility
Recruitment
Setting
Organization
Flexibility in delivery
Flexibility in adherence
Follow-up
Primary outcome
Primary analysis

Each domain is scored from 1 (very explanatory, near the hub) to 5 (very pragmatic, near the rim).

Carry forward

From individuals to clinics

  • Randomization protects the comparison at the moment of allocation.
  • A pragmatic trial answers the question that a funding decision asks.
  • Programs delivered through clinics invite contamination if individuals within a clinic are randomized.

Learning Objectives for this section

  • Explain, using the language of potential outcomes, why random allocation supports a causal claim about a program.
  • Distinguish random allocation from random sampling, and describe allocation concealment and the main randomization schemes.
  • Distinguish explanatory from pragmatic trials by the question each one answers and the decision each one informs.
  • Use the nine domains of the PRECIS-2 tool to locate a proposed trial on the explanatory-pragmatic continuum.

1.1 The Causal Question in Program Evaluation

By this point in the course, an evaluation plan contains a program description, a needs statement, a logic model, a set of prioritized evaluation questions and an evaluation matrix. Lessons 6, 7 and 8 turn to the evaluation question that decision-makers often ask first and that is the hardest to answer well, which is whether the program caused the change that was observed. This lesson covers designs that answer the question by random allocation. Lesson 7 covers comparison-group designs, difference-in-differences and matching, and Lesson 8 covers interrupted time series, regression discontinuity and natural experiments. Those designs are the tools an evaluator uses when randomization is impossible or unacceptable.

Background: randomized trial basics

Sections 1.1 and 1.2 restate the basics of randomized trials (potential outcomes, exchangeability, random allocation, allocation concealment and the common randomization schemes) so that students who come to this course from other programs have them in one place. Students who took HSCI 230 met this material in Lesson 5, Sections 5 to 7 (Cohort Studies), which is optional review. The development beyond that treatment includes Fisher's argument for randomization as a basis for inference, the limits of the guarantee that randomization provides, minimization and constrained randomization, the PRECIS-2 tool in Sections 1.3 and 1.4, and the complier average causal effect in Section 4.

The running example is the fictional Cedar Valley Connector program, which runs through this course. The fictional Cedar Valley Health Authority in British Columbia invites primary care clinicians to refer adults aged 65 and older who screen as socially isolated or lonely to a community connector. The connector meets each person up to six times over twelve weeks, co-develops a plan with them, and links them to community groups, volunteer roles, transportation help and services. The program launched in 12 of the region's 24 primary care clinics in a first wave, and the remaining 12 clinics are scheduled to join a year later. The steering committee wants to know whether the program reduces loneliness, measured with the three-item UCLA Loneliness Scale (scored from 3 to 9, with higher scores indicating greater loneliness), and whether it changes social participation, self-rated health, and the use of emergency departments and primary care. All Cedar Valley figures in this lesson are illustrative.

Background: potential outcomes and the counterfactual

Causal questions are easiest to state in the potential outcomes framework associated with Neyman (1923) and Rubin (1974). Each older adult referred to the program has two potential six-month loneliness scores: Y(1), the score the person would have if they received the program, and Y(0), the score the same person would have if they did not. The individual causal effect is the difference Y(1) − Y(0). Holland (1986) called the impossibility of observing both quantities for the same person at the same time the fundamental problem of causal inference, because each person reveals only the potential outcome that corresponds to the condition they actually experienced. An evaluation therefore targets an average causal effect, such as the mean of Y(1) − Y(0) across the people the program serves, and estimates it by comparing groups.

Students who took HSCI 341 met these ideas in Lesson 1, Section 4 (Introduction and Causal Concepts), which is optional review; students from other programs will find a fuller introduction to the counterfactual there.

A comparison of groups estimates the average causal effect only when the groups would have had the same average outcome if neither had received the program. This condition is called exchangeability. When exchangeability fails, the observed difference between groups combines the program effect with the difference that would have existed between the groups in the absence of the program. Program evaluation texts often call this difference selection bias, and epidemiology calls it confounding; Lesson 7, Section 1 explains the two uses of the term.

Background: confounding and backdoor paths

A confounder is a common cause of receiving the program and of the outcome, or a proxy for one, that is not on the causal pathway from the program to the outcome. If clinicians refer patients whom they expect to engage, and if those patients would also have become less lonely without the program, then the clinician's judgement confounds a comparison of referred and unreferred patients. The causal diagram below draws this example, with the letters used across the course series in brackets: X for the exposure (here the program), Y for the outcome and C for the confounder.

Clinician's judgement of likely engagement (C) Referral to the connector program (X) Six-month loneliness (Y)

A backdoor path is a non-causal path from the exposure to the outcome that begins with an arrow into the exposure. Here the path from referral back to the clinician's judgement and on to loneliness (X ← C → Y) is a backdoor path, and while it is open, a comparison of referred and unreferred patients mixes the program effect with the effect of the judgement. Exchangeability that holds within levels of measured covariates is called conditional exchangeability, and it corresponds to blocking every backdoor path by conditioning on measured variables. Observational analyses do this through stratification, regression, matching or weighting (Lesson 7 develops propensity scores), and the adjusted estimate is unbiased only when every backdoor path has been blocked by variables that were measured accurately. A judgement that was never recorded, such as a clinician's sense of which patients will follow through, cannot be adjusted for, so its backdoor path stays open.

HSCI 230 Lesson 7, Section 2 (Conceptualization, Measurement and Causal Specification), introduces causal diagrams, and HSCI 341 Lesson 7, Section 3 (Designing Against Bias: Validity and Confounding in Study Protocols), states the backdoor criterion in full. Both are optional reading.

Case: How the first wave might have been chosen

Suppose the health authority chose the 12 first-wave clinics because their managers volunteered, because they had a room the connector could use, and because their electronic medical records could produce a list of older patients who had screened as lonely. Each of these reasons could plausibly be related to loneliness outcomes. Clinics with engaged managers may already run more group activities, clinics with spare rooms may sit in town centres with better transit, and clinics with stronger data systems may serve more affluent neighbourhoods. A comparison of six-month loneliness between first-wave and second-wave clinics would then mix any program effect with these differences, and no statistical adjustment could remove the parts that were never measured.

A. Wave one chosen by readiness Clinic readiness and referral practices Wave one Connector program Six-month loneliness B. Wave one chosen at random Clinic readiness and referral practices Wave one Connector program Six-month loneliness Readiness opens a backdoor path, so the comparison is confounded. R random allocation Readiness no longer decides who gets the program, so the backdoor is closed.
Figure 1.1. When clinic readiness decides which clinics start first (panel A), readiness is a common cause of the program and the outcome; when a random process decides (panel B), the path from readiness to the program is removed.

What random allocation does

Random allocation breaks the link between the factors that influence the outcome and the condition each unit receives. When a random number generator decides which clinics start first, a manager's enthusiasm, the availability of a room and the quality of a data system can no longer influence the allocation. The arms are then exchangeable in expectation: across all the allocations the randomization could have produced, the arms have the same distribution of every characteristic, measured or unmeasured, including the potential outcomes themselves. HSCI 230 Lesson 5, Section 5 (Cohort Studies), and HSCI 207 Lesson 7, Section 2 (Sampling, Recruitment and Measurement), introduce random allocation and the comparable groups it creates and are optional reading; the rest of this subsection develops Fisher's argument and the limits of the guarantee. Fisher (1935) added a second argument for randomization. Because the allocation mechanism is known, it provides a basis for inference, and a p-value can be read as a statement about how unusual the observed difference would be across the allocations that could have occurred if the program had no effect.

Three limits of this guarantee matter in program evaluation. First, the guarantee holds in expectation, so any single allocation can be imbalanced by chance. The risk of chance imbalance falls as the number of randomized units grows, which is why a trial that randomizes 24 clinics is more exposed to it than a trial that randomizes 960 people, a point that Section 2 develops. Second, randomization protects the comparison at the moment of allocation only. Events after allocation, such as dropout that differs between arms, outcome measurement that differs between arms, or the recruitment of participants by staff who already know their clinic's allocation, can reintroduce bias. Third, randomization establishes whether the program caused a difference in the people and settings where the trial was run, and whether the result applies to other people and settings remains a separate question of generalizability.

1.2 Random Allocation in Practice

Random allocation and random sampling

Random allocation and random sampling are often confused. Random sampling selects units from a defined population so that the sample represents that population, and it supports external validity. Random allocation assigns units that are already in a study to conditions, and it supports internal validity. A study can have either property without the other. Most trials randomly allocate volunteers who are a convenience sample of the eligible population, and most population surveys sample randomly without allocating anyone to anything.

Allocation concealment and blinding

Allocation concealment means that the people who enrol participants cannot know or predict the next allocation before a participant is enrolled. It differs from blinding, which concerns whether participants, staff and outcome assessors know the allocation after it has been made. Concealment can be achieved in almost every trial, whereas blinding often cannot. Schulz and colleagues (1995) found that trials with inadequate or unclear allocation concealment reported larger treatment effects than trials with adequate concealment, which is consistent with enrolment decisions being influenced by knowledge of the upcoming assignment. In program evaluation, concealment is usually achieved by having an independent statistician generate the allocation with a documented, seeded script and release each allocation only after the unit has been enrolled. In a cluster trial, the equivalent safeguard is that clinics, schools or communities commit to participate before they learn their allocation.

A community connector program cannot be delivered blind, because older adults know whether they are meeting a connector and connectors know whom they serve. Outcome assessment can still be blinded. In the Cedar Valley example, six-month loneliness could be collected by telephone interviewers who work from a central list and are not told which clinic a participant attends.

Randomization schemes

The scheme describes how the allocation sequence is generated. The choice depends on the number of units, the strength of known prognostic factors and the way units enter the trial.

Background: three basic randomization schemes

SchemeHow it worksWhen it helpsCaution
Simple randomizationEach unit is allocated independently, as if by a coin toss.Large trials with many units.Small trials can end with unequal arms or chance imbalance.
Permuted blocksUnits are allocated within blocks (for example, of four or six) that contain equal numbers of each arm.Trials that enrol people one at a time over months.Fixed, small blocks let staff predict the last allocation in a block, so block sizes are usually varied and concealed.
Stratified randomizationA separate allocation sequence is used within each stratum of an important prognostic factor.Trials with few units and one or two strong prognostic factors.Too many strata leave each stratum nearly empty.

HSCI 230 Lesson 5, Section 6 (Cohort Studies), introduces these schemes and is optional reading.

Trials with few units, and cluster trials in particular, often need more control over balance than these schemes provide. Three further schemes serve that purpose.

SchemeHow it worksWhen it helpsCaution
MinimizationEach new unit is allocated to the arm that reduces imbalance on several factors, usually with a random element (Taves, 1974; Pocock & Simon, 1975).Small trials that must balance several factors.A deterministic version weakens concealment, so a random element is added.
Pair matchingClusters are paired on similar characteristics, and one member of each pair is randomized to each arm.Cluster trials with very few clusters.The loss of one cluster removes the information from its pair.
Constrained randomizationMany possible allocations are generated, those that are balanced on chosen cluster characteristics are retained, and one is selected at random (Moulton, 2004).Cluster trials in which all clusters are known before allocation.The balance criteria should be specified in advance and reported.
Try it: Choose a scheme for Cedar Valley

Suppose 8 of the 24 Cedar Valley clinics are rural and 16 are urban, and the clinics range from small two-physician practices to large team-based clinics. All 24 clinics are known before allocation, and the health authority wants exactly 12 clinics in each wave. Decide which scheme you would use and why.

A strong answer stratifies by rurality, randomizing 4 rural and 8 urban clinics to each wave, so that the rural-urban mix cannot differ by chance. An equally defensible answer uses constrained randomization that balances rurality, clinic size and the proportion of patients aged 75 and older, because all clusters are known in advance and three balancing factors are more than stratification can handle with 24 units.

1.3 Explanatory and Pragmatic Trials

Schwartz and Lellouch (1967) observed that trialists approach a comparison with one of two attitudes. An explanatory trial asks whether an intervention can work under ideal conditions, often in order to test a mechanism. A pragmatic trial asks whether an intervention works when it is delivered in the conditions of usual practice, in order to inform a choice between options. The distinction corresponds closely to the difference between efficacy and effectiveness, which Lesson 9 develops alongside implementation.

The two attitudes lead to different design choices. Explanatory trials select participants who are likely to respond and adhere, use expert staff and close monitoring, control co-interventions, compare the intervention with a placebo or attention control, and often measure mechanistic or intermediate outcomes. Pragmatic trials enrol the people who would receive the intervention in practice, deliver it through the usual staff and organizations, compare it with usual care, measure outcomes that matter to participants and decision-makers, and analyze everyone as randomized. Thorpe and colleagues (2009) argued that real trials sit on a continuum between the two attitudes, and that a trial can be pragmatic in some respects and explanatory in others.

The questionClick to explore
ParticipantsClick to explore
Setting and deliveryClick to explore
ComparatorClick to explore
OutcomesClick to explore
AnalysisClick to explore

For the Cedar Valley Connector program, an explanatory trial might be run by a university team that trains its own connectors, enrols older adults without cognitive impairment who are fluent in English, telephones participants between meetings to support attendance, and measures loneliness monthly. A pragmatic trial would use the health authority's own connectors, accept every older adult whom clinicians refer, compare the program with usual primary care, and measure loneliness at six months with a short telephone survey supplemented by administrative records of emergency department and primary care visits. The steering committee's decision is whether to fund the program across the region, so the pragmatic version answers the question it needs answered.

1.4 Locating a Trial with PRECIS-2

The Pragmatic Explanatory Continuum Indicator Summary (PRECIS) was introduced by Thorpe and colleagues (2009) to help trialists match design decisions to the purpose of a trial. Loudon and colleagues (2015) revised it as PRECIS-2 after consultation with trialists and testing in practice. PRECIS-2 has nine domains, each scored from 1 (very explanatory) to 5 (very pragmatic). The scores are plotted on a wheel, so that a trial whose points lie near the hub is explanatory and a trial whose points lie near the rim is pragmatic.

The tool is meant for the design stage. The trial team, ideally with the people who will use the results, scores each domain and records the reason for the score. The scores are meant to be read domain by domain, because summing them into one number would hide the pattern that the wheel is designed to show. A wheel is also neither a quality score nor a ranking. An explanatory trial is the appropriate design for a question about whether an intervention can work, and the PRECIS-2 exercise asks only whether each design choice fits the question the trial is meant to answer.

1. Eligibility (Cedar Valley score 4)v

The domain asks how similar participants are to the people who would receive the program in usual care. The Cedar Valley trial enrols every adult aged 65 and older whom a clinician refers after a positive isolation or loneliness screen, and excludes only people who cannot complete a short telephone survey even with help from a family member or interpreter. The exclusion keeps the score below 5.

2. Recruitment (score 4)v

The domain asks how much extra effort goes into recruitment beyond what would happen in usual care. Clinicians refer during routine visits, as they would once the program is established, and a research coordinator then telephones referred patients to seek consent and collect baseline data. The added contact is modest.

3. Setting (score 5)v

The domain asks how different the trial settings are from the settings where the program would be used. The trial runs in all 24 of the region's primary care clinics, which are the clinics that would deliver the program at scale.

4. Organization (score 4)v

The domain asks how the resources, staff expertise and organization of care compare with usual care. Connectors are health authority employees with the program's standard training. The trial adds a short fidelity log that connectors complete after each meeting, which is a small addition to usual practice.

5. Flexibility in delivery (score 4)v

The domain asks how much flexibility connectors have in delivering the program compared with usual care. Connectors tailor each plan and offer between one and six meetings according to need, guided by the program manual. The manual sets some limits, which keeps the score from 5.

6. Flexibility in adherence (score 5)v

The domain asks what is done to encourage participants to adhere. The trial adds nothing beyond the program's usual reminder calls, and older adults who stop attending are not pursued for research purposes.

7. Follow-up (score 3)v

The domain asks how the intensity of measurement and follow-up compares with usual care. Research staff conduct baseline and six-month telephone surveys, which usual care would not include. The surveys are short, and health service use comes from administrative data, so the score sits in the middle.

8. Primary outcome (score 4)v

The domain asks how relevant the primary outcome is to participants. Loneliness is the problem the program addresses, and the steering committee, which includes older adults with lived experience, judged a change in loneliness to be meaningful. The use of a research instrument at a fixed time point keeps the score below 5.

9. Primary analysis (score 5)v

The domain asks to what extent all data are included in the primary analysis. The protocol specifies an intention-to-treat analysis that includes every enrolled older adult in the arm of their clinic, whether or not they engaged with a connector.

Eligibility Recruitment Setting Organization Flexibility (delivery) Flexibility (adherence) Follow-up Primary outcome Primary analysis hub = 1, very explanatory Pragmatic Cedar Valley trial (clinic randomized) Hypothetical explanatory efficacy trial rim = 5, very pragmatic
Figure 1.2. PRECIS-2 wheel for the pragmatic Cedar Valley trial (solid teal) and a hypothetical explanatory efficacy trial of the same program (dashed red). Scores are illustrative and follow the domain descriptions above.

The two wheels in Figure 1.2 describe trials of the same program that would answer different questions. The explanatory trial, with university-trained connectors, narrow eligibility, adherence support and monthly follow-up, would show whether the connector model can reduce loneliness when delivered under favourable conditions. The pragmatic trial would show whether offering the program through Cedar Valley's clinics reduces loneliness among the older adults those clinics actually refer. A health authority deciding on regional funding needs the second answer. A research team testing whether social connection mediates the effect of the program on self-rated health might reasonably prefer the first.

Try it: Score a proposed change

During protocol development, a team member proposes two changes to the pragmatic Cedar Valley trial: excluding older adults who do not speak English, and telephoning every participant every two weeks to encourage attendance at connector meetings. Identify which PRECIS-2 domains each change affects and the direction in which the scores would move.

The language exclusion lowers the Eligibility score, because it removes people who would be referred in practice, and it would also reduce the relevance of the results for a region with linguistically diverse older adults. The fortnightly calls lower the Flexibility in adherence score, because they add an adherence-promoting measure that usual care would not provide, and they also lower the Follow-up score, because they add research contact. Both changes move the wheel toward the hub and make the trial less able to answer the steering committee's question.

1.5 From Individuals to Clinics

Pragmatic trials of health services programs face a recurring design problem. Programs such as the Cedar Valley Connector are delivered by and through organizations. Connectors are based in clinics, clinicians decide whom to refer, and clinic routines change once a connector is on site. If older adults within the same clinic were randomized individually, clinicians who had learned about community resources through the program would be likely to share that knowledge with patients in the usual-care arm, and patients in the same waiting room would talk with one another. This leakage, called contamination, would shrink the observed difference between arms. The usual solution is to randomize whole clinics, which is the subject of Section 2.

The CONSORT extension for pragmatic trials (Zwarenstein et al., 2008) asks authors to describe the setting, the usual-care comparator and the people and organizations involved, so that readers can judge whether the results apply to their own context. Section 4 returns to reporting and introduces the current CONSORT statement and its extensions.

Reflection

A provincial health ministry is choosing between two trial designs for a telephone befriending program for adults aged 75 and older who live alone. In Design A, a university team recruits 200 volunteers who pass a cognitive screen and speak English, trains its own befrienders, telephones participants weekly to encourage them to keep their befriending calls, measures loneliness monthly for six months, and analyzes only participants who completed at least eight calls. In Design B, the ministry's existing befriending service accepts every referral from family physicians in three health regions, delivers calls with its usual volunteers and schedule, compares the program with usual care, measures loneliness once at six months by telephone, and analyzes all randomized participants. The ministry must decide whether to fund the service across the province. PRECIS-2 scores nine domains (eligibility, recruitment, setting, organization, flexibility in delivery, flexibility in adherence, follow-up, primary outcome and primary analysis) from 1, very explanatory, to 5, very pragmatic. Score three domains for each design, explain each score, and state which design better fits the ministry's decision and why.

Model answer

For eligibility, Design A scores about 1 or 2, because it enrols volunteers who pass a cognitive screen and speak English, whereas Design B scores about 5, because it accepts every physician referral, which is the population the service would serve. For flexibility in adherence, Design A scores about 1, because weekly encouragement calls are an adherence measure that the service would never provide, while Design B scores 5, because nothing is added beyond usual practice. For primary analysis, Design A scores 1, because it analyzes only participants who completed eight calls, and Design B scores 5, because it analyzes everyone as randomized.

Design B fits the ministry's decision. The ministry needs to know what happens when the existing service is offered to the people physicians refer, delivered as it would be in practice. Design A answers a different question, whether befriending can reduce loneliness among adherent, cognitively intact volunteers under favourable conditions, and its estimate would probably overstate the effect the province would see. A strong answer could also score setting (A about 2, B 5) or follow-up (A about 1 or 2 with monthly measurement, B about 4), with the same conclusion.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Why does random allocation of clinics to the Cedar Valley program's two waves support a causal comparison?

Randomization makes the allocation independent of every characteristic, measured or unmeasured, so the arms are exchangeable in expectation. It does not guarantee identical arms in any single allocation (option a), because chance imbalance remains possible, and representativeness (option c) is a property of random sampling.

Question 2: A trial team has an independent statistician generate a seeded allocation sequence, and each clinic learns its allocation only after signing its participation agreement. Which safeguard does this procedure provide?

Allocation concealment prevents the people enrolling units from knowing the next allocation before enrolment. Blinding (option a) concerns knowledge of the allocation after it has been made, which this procedure does not address.

Question 3: Which description best fits a pragmatic trial of the Cedar Valley Connector program?

A pragmatic trial uses usual staff and settings, broad eligibility, a usual-care comparator and an intention-to-treat analysis. Options b, c and d describe explanatory features: narrow eligibility and adherence support, analysis of completers only, and an attention control.

Question 4: In PRECIS-2, a trial that telephones participants every two weeks to encourage attendance at program sessions would score toward the explanatory end of which domain?

Measures that promote adherence beyond usual practice move the flexibility in adherence domain toward the explanatory end. The calls would also add research contact, which affects follow-up. Flexibility in delivery (option d) concerns how staff deliver the program, and the calls do not change the organization of care or the outcome.
Section 2 of 5

Cluster Randomized Trials

⏱ Estimated reading time: 50 minutes
Section 2 of 5

Cluster Randomized Trials

Randomizing clinics, measuring people, and accounting for the similarity of people in the same clinic.

Why clusters

Reasons to randomize clinics

The program acts on the clinic

Referral pathways and on-site connectors change care for every patient.

Contamination

Staff and patients share information across arms within one clinic.

Logistics

Health authorities hire, train and budget by site.

Acceptability

Clinics accept a randomized start date more readily than a split caseload.

Similarity within clusters

The intracluster correlation coefficient

ICC
\[ \rho = \frac{\sigma^2_{between}}{\sigma^2_{between} + \sigma^2_{within}} \]

The ICC is the share of outcome variance that lies between clusters. In the simulated Cedar Valley data it is 0.059.

The design effect

How much clustering costs

Design effect and effective sample size
\[ \text{DEFF} = 1 + (m - 1)\,\rho \qquad n_{eff} = \frac{N}{\text{DEFF}} \]
3.34Design effect with 40 per clinic and an ICC of 0.06
287Effective sample size of 960 clustered older adults
Sample size

The number of clusters sets the limit

  • An individually randomized trial would need 142 older adults per arm to detect 0.5 points with 80% power.
  • Multiplying by the design effect of 3.34 gives 472 per arm, which is 12 clinics of 40.
  • However large the clinics, at least 9 clinics per arm would be needed, because 141.3 × 0.06 = 8.5.
The worked example

Naive model versus mixed model

0.084Naive standard error for the program effect (estimate −0.428)
0.136Mixed-model standard error (estimate −0.420)

The squared ratio of the standard errors (2.588) is close to the design effect from the residual ICC (2.605).

Analysis and reporting

Valid analyses and the cluster extension

Cluster-level summaries

Clinic means are compared with a t-test.

Mixed models

Random cluster effects give cluster-specific estimates.

Estimating equations

Sandwich standard errors give population-averaged estimates.

With few clusters, small-sample corrections are needed. The CONSORT cluster extension (Campbell et al., 2012) asks for the rationale, the ICC and a two-level flow diagram.

Carry forward

From clusters to time

  • The number of clusters drives the power of a cluster trial.
  • An analysis that ignores clustering reports more precision than the data contain.
  • When every clinic will eventually receive a program, the order of rollout can itself be randomized.

Learning Objectives for this section

  • Explain why health services programs are often randomized by clinic, school or community, and describe what cluster randomization costs.
  • Define the intracluster correlation coefficient and calculate the design effect and the effective sample size.
  • Calculate the number of clusters needed for a cluster randomized trial with a continuous outcome.
  • Interpret a linear mixed model for a cluster trial, and explain why a model that ignores clustering understates the standard error of the program effect.
  • Compare cluster-level, mixed-model and generalized estimating equation analyses, and identify the reporting items in the CONSORT extension for cluster trials.

2.1 Why Programs Are Randomized by Cluster

In a cluster randomized trial, intact groups of people, such as clinics, schools, workplaces, long-term care homes or communities, are randomized, and outcomes are measured on the individuals within them. The design is also called a group-randomized trial (Murray, 1998). Standard texts include Donner and Klar (2000), Eldridge and Kerry (2012), and Hayes and Moulton (2017). Health services programs are often evaluated with this design because the program is delivered through an organization, and randomizing individuals within that organization would be impractical or would contaminate the comparison.

The program acts on the groupClick to explore
ContaminationClick to explore
Logistics and administrationClick to explore
AcceptabilityClick to explore
Group-level effectsClick to explore

Cluster randomization has costs. People in the same cluster tend to resemble one another, so each additional person in a cluster adds less new information than an independently sampled person, and the trial needs more participants in total. With few clusters, chance imbalance between arms is more likely. When participants are identified and recruited after clusters learn their allocation, staff in intervention clusters may recruit different people from staff in control clusters, a problem called identification or recruitment bias that Puffer, Torgerson and Watson (2003) documented in published cluster trials. Cluster trials also raise distinctive questions about consent, which Section 3 discusses. Most importantly for analysis, the clustering must be reflected in the statistical model. Cornfield (1978) described randomization by cluster followed by an analysis that treats individuals as independent as a form of self-deception, because it produces confidence intervals that are too narrow and p-values that are too small.

2.2 The Intracluster Correlation Coefficient

Older adults who attend the same clinic share a catchment area, a set of local community resources, the same transit routes, and clinicians with the same referral habits. Their loneliness scores are therefore likely to be more similar to one another than to the scores of older adults at other clinics. The intracluster correlation coefficient (ICC, often written with the Greek letter rho) measures this similarity.

Intracluster correlation coefficient

ICC = σ2between / (σ2between + σ2within)

Here σ2between is the variance of the true cluster means around the overall mean, and σ2within is the variance of individual outcomes around their own cluster's mean. The ICC is the proportion of the total variance in the outcome that lies between clusters. It can equivalently be read as the correlation between the outcomes of two randomly chosen individuals from the same cluster.

An ICC of 0 means that cluster membership tells us nothing about an individual's outcome, and an ICC of 1 means that everyone in a cluster has the same outcome. ICCs for patient-level outcomes in primary care are usually small (Adams et al., 2004), often below 0.05, although they vary by outcome and by the type of cluster. Process-of-care measures, which depend directly on clinic practices, tend to show larger ICCs than outcomes such as loneliness or self-rated health, which depend mostly on the individual. The simulated Cedar Valley data used later in this section have an ICC near 0.06, toward the upper end of what is typical for a patient-reported outcome, which makes the consequences of clustering easy to see.

Planners take the ICC from a pilot study, from routine data on the same clusters, or from published trials with similar clusters and outcomes. An ICC estimated from few clusters is imprecise, so calculations should be checked across a plausible range, and trial reports should state the ICC they observed so that later trials can plan with better information.

2.3 The Design Effect and the Effective Sample Size

The design effect is the factor by which the variance of an estimate from a cluster design exceeds the variance from an individually randomized design with the same total number of participants. For clusters of equal size m, the design effect is given by the formula below, which follows Kish (1965) in survey sampling and Donner, Birkett and Buck (1981) for cluster trials. HSCI 230 Lesson 5, Section 6 (Cohort Studies), and HSCI 410 Lesson 5, Section 1 (Modelling Dependent Data), introduce the design effect at undergraduate level and are optional reading.

Design effect and effective sample size

DEFF = 1 + (m − 1) × ICC

Effective sample size = N / DEFF

With unequal cluster sizes, Eldridge, Ashby and Kerry (2006) give DEFF = 1 + ((CV2 + 1) × m − 1) × ICC, where m is the mean cluster size and CV is the coefficient of variation of cluster size (the standard deviation of cluster sizes divided by their mean).

Notation: DEFF is the design effect, m the cluster size (the mean size when sizes vary), N the total sample size and ICC (also written ρ) the intracluster correlation coefficient. Reliability studies use the intraclass correlation coefficient, a different application of the same kind of statistic (HSCI 410 Lesson 7).

The effective sample size is the number of independent observations that would give the same precision as the clustered sample. Suppose the Cedar Valley trial enrols 40 older adults in each of 24 clinics (N = 960) and the ICC is 0.06. The design effect is 1 + (40 − 1) × 0.06 = 1 + 2.34 = 3.34, and the effective sample size is 960 / 3.34 = 287.4. Nine hundred and sixty clustered observations carry about as much information about the program effect as 287 independent ones.

Two features of the formula drive most practical decisions. First, the design effect depends on the product of the ICC and the cluster size, so an ICC that looks negligible can still produce a large design effect when clusters are large. An ICC of 0.01 with 200 participants per cluster gives a design effect of 2.99. Second, the design effect grows in a straight line with cluster size, so each additional participant per cluster buys less precision than the one before, as Figure 2.1 shows. Variation in cluster size raises the design effect further, because large clusters contribute disproportionately to the estimate.

1 2 3 4 5 6 7 8 9 0 20 40 60 80 100 Mean cluster size, m (older adults per clinic) Design effect ICC 0.01 ICC 0.03 ICC 0.05 ICC 0.08 Cedar Valley planning values: m = 40, ICC = 0.06, DEFF = 3.34
Figure 2.1. The design effect rises linearly with mean cluster size, and more steeply for larger intracluster correlations. The red point marks the Cedar Valley planning values used in Section 2.4.

2.4 Sample Size for a Cluster Randomized Trial

The sample size for a cluster trial with a continuous outcome is calculated in three steps. The first step calculates the number of participants per arm that an individually randomized trial would need. The second step multiplies that number by the design effect. The third step divides the inflated number by the cluster size to obtain the number of clusters per arm, rounding up at each step.

Sample size for a two-arm cluster trial, continuous outcome

nindividual per arm = 2 × (z1−α/2 + z1−β)2 × SD2 / Δ2

ncluster trial per arm = nindividual × DEFF

Clusters per arm = ncluster trial / m, rounded up. Here Δ is the smallest difference worth detecting, SD is the standard deviation of the outcome, α is the two-sided significance level and 1 − β is the power.

For Cedar Valley, suppose the steering committee judges that a difference of 0.5 points on the three-item UCLA Loneliness Scale would be large enough to justify regional funding, and that the standard deviation of six-month scores is about 1.5 points. With a two-sided α of 0.05 (z = 1.960) and 80 percent power (z = 0.842), an individually randomized trial needs 2 × (1.960 + 0.842)2 × 1.52 / 0.52 = 141.3 participants per arm, which rounds up to 142. With 40 participants per clinic and an ICC of 0.06, the design effect is 3.34, so the cluster trial needs 141.3 × 3.34 = 471.9, or 472 participants per arm. Dividing by 40 gives 11.8, which rounds up to 12 clinics per arm. Under these assumptions, the 24 clinics in the Cedar Valley region are just enough.

The number of clusters matters more than the number of participants per cluster. As the cluster size grows without limit, the number of clusters per arm approaches nindividual × ICC (Hemming et al., 2011). For Cedar Valley, 141.3 × 0.06 = 8.5, so at least 9 clinics per arm would be needed even if every clinic could enrol thousands of older adults. If fewer clusters are available than this minimum, no amount of recruitment within clusters will deliver the planned power, and the evaluator must accept a larger detectable difference or choose another design.

Several practical adjustments follow. Planners inflate the sample size for expected loss to follow-up and for the possibility that a cluster withdraws. Because a trial with few clusters is analyzed with a t distribution with limited degrees of freedom, a calculation based on normal quantiles slightly understates the clusters required, and planners often add at least one cluster per arm. Adjustment for a baseline measure of the outcome usually increases power. The calculation should also be repeated across a plausible range of ICCs and cluster sizes, as the worked example in Section 2.5 does.

2.5 Worked Example: Analyzing the Simulated Cedar Valley Cluster Trial

The worked example uses a simulated dataset that represents a hypothetical randomized version of the Cedar Valley rollout. Imagine that, before launch, the health authority had randomized the 24 clinics so that 12 joined the program in wave one and 12 continued usual care during year one before joining in wave two. Older adults referred at each clinic completed the UCLA Loneliness Scale at enrolment and six months later. The program indicator, called arm, equals 1 for wave-one clinics and 0 for wave-two clinics. The dataset also records whether each older adult in a wave-one clinic engaged with the program (attended at least two connector meetings), which Section 4 uses. The data are simulated, so the results describe this teaching dataset only.

Background: reading a regression coefficient

The worked example compares regression models, so it helps to recall how one is read. A linear regression of the six-month loneliness score on the binary program indicator, arm, predicts each person's score as an intercept plus a coefficient times arm. The intercept is the predicted mean score in the usual-care arm (arm = 0), and the coefficient for arm is the difference in mean score between the Connector arm and the usual-care arm, so a coefficient of −0.43 means that Connector-arm participants scored 0.43 points lower on average. The standard error of the coefficient describes how much the estimate would vary from sample to sample, and the 95 percent confidence interval is approximately the estimate plus or minus 1.96 standard errors, or plus or minus a t quantile times the standard error when the degrees of freedom are few. An interval that excludes zero corresponds to a two-sided p-value below 0.05. Adding the baseline loneliness score as a second predictor adjusts the comparison, so that the coefficient for arm compares participants with the same baseline score. Randomization has already balanced the arms in expectation, so the main gain from baseline adjustment in a trial is precision: the baseline score explains part of the variation in six-month scores, which usually narrows the confidence interval, and it also corrects for any chance imbalance at baseline.

HSCI 410 Lesson 3 (Linear and Logistic Regression) covers linear regression more fully and is optional reading.

Background: random intercepts and the empty mixed model

A linear mixed model (also called a multilevel or random-effects model) adds a random intercept for each cluster to an ordinary regression. Each clinic's mean score may sit above or below the overall mean by an amount that the model treats as a draw from a normal distribution with its own variance. The model therefore describes two sources of variation: the spread of clinic means around the overall mean (the between-clinic variance, σ2between) and the spread of individuals around their own clinic's mean (the within-clinic variance, σ2within). An empty model, also called a null or intercept-only model, contains only the overall intercept and the random intercept, so it splits the total variance of the outcome into these two parts without explaining any of it. The ICC is the between-clinic share, σ2between / (σ2between + σ2within), which is the formula in Section 2.2. When predictors such as arm or baseline loneliness are added, the variances that remain are called residual variances, and the ICC computed from them is a residual ICC.

HSCI 410 Lesson 5, Sections 2 and 3 (Modelling Dependent Data), covers random-intercept models and their extension to binary outcomes and is optional reading.

Worked example: Analyzing the simulated Cedar Valley cluster trial

The clusters. The trial has 12 clinics in each arm, with 485 older adults in the usual-care arm and 475 in the Connector arm, for a total of 960. Clinics enrolled between 32 and 50 older adults, with a mean of 40 and a coefficient of variation of 0.164. Baseline mean scores are almost identical in the two arms (6.136 in the usual-care arm and 6.124 in the Connector arm), as randomization leads us to expect, while the six-month mean is 6.056 in the usual-care arm and 5.621 in the Connector arm.

The intracluster correlation. An empty linear mixed model, which contains only an intercept and a random intercept for clinic, splits the variance of the six-month score into a between-clinic part and a within-clinic part. The between-clinic variance is 0.1387 and the within-clinic variance is 2.2105, so the ICC is 0.1387 / (0.1387 + 2.2105) = 0.0590. About 6 percent of the variation in six-month loneliness lies between clinics. Because wave-one and wave-two clinics differ in their program status, part of that between-clinic variation is the program effect itself. Once arm is added to the model, the ICC falls to 0.0422. For planning a future trial, the ICC from the usual-care arm, from baseline data or from a model that includes arm is usually preferred, because it describes the clustering that exists before the program creates any difference.

The design effect and the effective sample size. With a mean of 40 older adults per clinic and an ICC of 0.0590, the design effect is 3.30 if cluster sizes were equal and 3.37 once their variation (a coefficient of variation of 0.164) is taken into account. The 960 older adults therefore carry about as much information as 285 independent observations. Variation in cluster size of this magnitude adds little to the design effect. The adjustment matters more when cluster sizes vary widely, for example when a few large clinics enrol most of the participants.

Sample size. The planning values are a smallest important difference of 0.5 points, an ICC of 0.06 (the estimate above, rounded) and a standard deviation of 1.5, which rounds up the standard deviation of 1.415 observed in the usual-care arm to be conservative. The three-step calculation from Section 2.4 gives 142 per arm for an individually randomized trial, a design effect of 3.34, 472 per arm for the cluster trial and 12 clinics per arm. A coefficient of variation of 0.15 in clinic size raises the design effect to 3.39 and the total to 480 per arm, still with 12 clinics. The table below repeats the calculation of clinics needed per arm across four ICCs and four cluster sizes, keeping the same difference, standard deviation and power.

ICC10 per clinic20 per clinic40 per clinic80 per clinic
0.0116954
0.03181286
0.052114119
0.0825181513

The table shows the diminishing return from larger clusters. With an ICC of 0.05, doubling cluster size from 40 to 80 reduces the clinics needed per arm from 11 to 9, and with an ICC of 0.08 it reduces them from 15 to 13, while the total number of participants rises sharply.

Ignoring clustering compared with modelling it. Two models regress the six-month score on arm and the baseline score. The naive model, an ordinary linear regression, treats all 960 older adults as independent. The linear mixed model adds a random intercept for clinic, and its confidence interval and p-value use a t distribution with 22 degrees of freedom, which is the 24 clinics minus the 2 cluster-level parameters (the between-within approach).

ModelProgram effect (points)Standard error95% confidence interval
Naive linear regression−0.4280.084−0.594 to −0.263
Linear mixed model with a random intercept for clinic−0.4200.136−0.701 to −0.139

The two models agree closely about the size of the program effect. The naive model estimates that older adults in Connector clinics scored 0.428 points lower at six months, and the mixed model estimates 0.420 points lower, after adjustment for baseline loneliness. The models disagree about precision. The mixed-model standard error is 1.609 times the naive standard error, and the mixed-model p-value is 0.0052 on 22 degrees of freedom.

The standard errors differ because the program indicator is constant within each clinic. Every older adult in a clinic has the same value of arm, so the information about the program effect comes from comparing 12 clinics with 12 clinics, and older adults in the same clinic share a clinic-level deviation that does not average out within the clinic. The naive model assumes that the 960 residuals are independent and therefore credits the data with far more independent information than they contain. The mixed model estimates the between-clinic variance (0.0685) separately from the within-clinic variance (1.6418), and its standard error for arm reflects the number of clinics as well as the number of people. The squared ratio of the standard errors (2.588) is close to the design effect implied by the residual ICC of 0.040 (2.605), which confirms that the design effect is the quantity the naive model ignores. The residual ICC is smaller than the 0.0590 from the empty model because baseline loneliness and arm explain part of the difference between clinics, which is one reason baseline adjustment improves the precision of cluster trials.

A clustered standard error. An alternative keeps the naive point estimate and replaces its standard error with a sandwich estimator that allows residuals to be correlated within clinics. The clustered sandwich standard error is 0.1313, close to the mixed-model value of 0.136 and far from the naive 0.084. When its p-value is computed from a t distribution with the 957 residual degrees of freedom of the linear model, it is 0.0011, which overstates the information available. With 22 degrees of freedom the p-value is 0.0036. Sandwich estimators also tend to underestimate the variance when there are few clusters, so trials with few clusters (a common rule of thumb is fewer than about 40) should use a small-sample correction such as the one proposed by Mancl and DeRouen (2001). The clustered sandwich approach is closely related to the generalized estimating equations described next.

2.6 Choosing an Analysis for a Cluster Trial

Three families of analysis respect clustering. Each is valid when used appropriately, and the choice depends on the number of clusters, the type of outcome and whether the target is a cluster-specific or a population-averaged effect.

Background: cluster-specific and population-averaged effects

A mixed model estimates a cluster-specific effect, which compares the expected outcome with and without the program for people in the same clinic, or in clinics with the same random intercept. Generalized estimating equations estimate a population-averaged effect, which compares the average outcome across the whole population of clinics with and without the program. For a continuous outcome in a linear model the two coincide, whereas for a binary outcome analyzed on the odds ratio scale the cluster-specific effect lies further from the null. Both approaches need a standard error that respects clustering. A sandwich standard error is computed from the observed spread of the residuals within clusters: the model supplies the estimate, and the variance is then calculated from how much the clusters' residuals actually vary, so it remains valid even when the assumed correlation within clusters is wrong. Each cluster contributes one piece of information about that spread, so the sandwich estimator needs an adequate number of clusters, and with fewer than about 40 it tends to underestimate the variance unless a small-sample correction is applied.

HSCI 410 Lesson 5, Section 3 (Modelling Dependent Data), compares mixed models and generalized estimating equations and is optional reading.

The simplest valid analysis computes one summary per cluster, such as each clinic's mean six-month score, and compares the 12 Connector means with the 12 usual-care means using a t-test with 22 degrees of freedom. Covariates are handled in two stages: an individual-level regression without the arm indicator produces residuals, which are averaged within clusters and compared between arms (Hayes & Moulton, 2017). Cluster-level methods perform well with few clusters and are easy to explain to decision-makers.

Linear mixed models, such as the mixed model in the worked example in Section 2.5, include a random effect for each cluster and estimate a cluster-specific program effect. Generalized linear mixed models extend the approach to binary and count outcomes, such as an emergency department visit within six months. For a continuous outcome, cluster-specific and population-averaged effects coincide, but for a binary outcome analyzed with a logistic model the cluster-specific odds ratio lies further from 1 than the population-averaged one. With few clusters, tests should use small-sample degrees of freedom, such as the Satterthwaite or Kenward and Roger (1997) approximations, or the between-within approach used in the worked example.

Generalized estimating equations (Liang & Zeger, 1986) estimate a population-averaged effect, which answers the policy question of how much the program changes outcomes across the region. The analyst specifies a working correlation structure, usually exchangeable within clusters, and the sandwich standard errors remain valid even if that structure is misspecified, provided there are enough clusters. With few clusters, a bias-corrected sandwich estimator (Mancl & DeRouen, 2001) and a t reference distribution are advisable.

Many health services trials randomize fewer than 40 clusters. Leyrat and colleagues (2018) compared analysis methods in this situation. Methods that rely on large-sample approximations could produce inflated type I error rates, whereas cluster-level analyses, and mixed models or generalized estimating equations with small-sample corrections, generally kept the error rate close to its nominal level. The method and its correction should be specified in the protocol before outcome data are seen.

2.7 Reporting a Cluster Trial

The CONSORT extension for cluster randomized trials (Campbell et al., 2012) adds cluster-specific items to the CONSORT checklist. Section 4 introduces the current CONSORT statement, and the table below lists the cluster-specific requirements that bear most directly on sample size, randomization and analysis.

Report elementWhat the cluster extension asks forCedar Valley illustration
Title and rationaleIdentification of the trial as cluster randomized, and the reason for using a cluster design.The connector is based in the clinic and clinicians refer, so individual randomization would risk contamination.
EligibilityEligibility criteria for clusters and for individuals.All 24 primary care clinics; referred adults aged 65 and older who screen as isolated or lonely.
Sample sizeThe method of calculation, including the assumed ICC, cluster size and any allowance for variation in cluster size.ICC 0.06, 40 per clinic, smallest important difference 0.5 points, 12 clinics per arm.
Randomization and recruitmentThe unit and method of randomization, and whether individuals were identified and recruited before or after clusters were randomized, and by whom.Clinics randomized by an independent statistician; referred older adults approached by a research coordinator.
ConsentFrom whom consent was sought (representatives of the cluster, individual members, or both) and whether it was sought before or after randomization.Clinic agreement before randomization; individual consent for the research surveys.
Analysis and resultsHow clustering was accounted for, the number of clusters and individuals at each stage, and the estimated ICC for each primary outcome.Linear mixed model with a random clinic intercept; ICC reported with the primary result.

Section 3 extends the logic of cluster randomization over time. When a program will reach every clinic eventually, the order in which clinics start can itself be randomized, which is the idea behind the stepped-wedge design.

Reflection

A health authority plans a cluster randomized trial of a pharmacist-led medication review for older adults, randomizing community pharmacies. The primary outcome is a continuous medication appropriateness score with a standard deviation of 2.0 points, and the smallest difference worth detecting is 0.6 points. The trial will use a two-sided significance level of 0.05 (z = 1.96) and 80 percent power (z = 0.84). Each pharmacy is expected to enrol 25 older adults, and a pilot study suggests an intracluster correlation coefficient (ICC) of 0.02. Use these formulas: participants per arm for an individually randomized trial = 2 × (1.96 + 0.84)2 × SD2 / difference2; design effect = 1 + (m − 1) × ICC, where m is the number per cluster; pharmacies per arm = (participants per arm × design effect) / m, rounded up. (a) Calculate the design effect, the number of older adults per arm and the number of pharmacies per arm. (b) A colleague proposes analyzing the trial with ordinary linear regression of the score on the arm indicator. Explain what would go wrong and what you would do instead.

Model answer

(a) An individually randomized trial would need 2 × (2.80)2 × 2.02 / 0.62 = 2 × 7.84 × 4 / 0.36 = 174.2, or 175 older adults per arm. The design effect is 1 + (25 − 1) × 0.02 = 1 + 0.48 = 1.48. The cluster trial therefore needs 174.2 × 1.48 = 257.8, or 258 older adults per arm, and 258 / 25 = 10.3, so 11 pharmacies per arm (22 in total). Using the rounded 175 gives 259 per arm and the same 11 pharmacies.

(b) The arm indicator is constant within each pharmacy, so the information about the effect comes from 22 pharmacies, and older adults in the same pharmacy share a pharmacy-level deviation. Ordinary regression would treat about 550 older adults as independent, so its standard error would be too small by a factor of roughly the square root of 1.48, about 1.22. Confidence intervals would be too narrow and the type I error rate would be inflated. I would instead fit a linear mixed model with a random pharmacy intercept and small-sample degrees of freedom (about 20), or compare pharmacy-level means with a t-test, and I would report the ICC observed in the trial.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the simulated Cedar Valley trial, the empty mixed model gave a between-clinic variance of 0.1387 and a within-clinic variance of 2.2105. What is the intracluster correlation coefficient?

The ICC is the between-cluster variance divided by the total variance: 0.1387 / (0.1387 + 2.2105) = 0.059. Dividing by the within-cluster variance alone gives 0.063 (option a), and 0.941 is the within-cluster share of the variance.

Question 2: A cluster trial will enrol 30 participants in each cluster, and the intracluster correlation coefficient is 0.05. What is the design effect?

The design effect is 1 + (m − 1) × ICC = 1 + 29 × 0.05 = 2.45. Option b uses m instead of m − 1, which slightly overstates the design effect.

Question 3: An individually randomized trial would need 141.3 participants per arm, and the intracluster correlation coefficient is 0.06. About how many clusters per arm would a cluster trial need even if every cluster could enrol an unlimited number of participants?

As cluster size grows without limit, the clusters needed per arm approach nindividual × ICC = 141.3 × 0.06 = 8.5, which rounds up to 9. Twelve clinics per arm (option b) is the requirement with 40 participants per clinic, which is above the floor.

Question 4: In the worked example in Section 2.5, the naive linear model gave a standard error of 0.084 for the program effect and the mixed model gave 0.136. Which explanation is correct?

Because every older adult in a clinic shares the same allocation and a clinic-level deviation, information about the program effect comes from 24 clinics. The naive model treats 960 residuals as independent and understates the standard error. Both models used all 960 participants and both adjusted for baseline loneliness, and the point estimates were almost identical (−0.428 and −0.420).
Section 3 of 5

Stepped-Wedge, Factorial and Adaptive Designs

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Stepped-Wedge, Factorial and Adaptive Designs

Randomizing a rollout, testing several components, adapting as data arrive, and the ethics of randomizing access.

Staggered rollouts

Two ways to randomize the Cedar Valley rollout

Parallel waitlist trial

Twelve clinics are randomized to wave one, and the wave-two clinics serve as a waitlist control during year one.

Stepped-wedge trial

Four randomized sequences of six clinics start at three-month steps, and every clinic is measured in every period.

Time trends

Exposure rises with calendar time

Hussey and Hughes model
\[ Y_{ijt} = \mu + \beta_t + \theta X_{jt} + u_j + e_{ijt} \]

Period effects absorb trends common to all clusters. The model assumes a common trend, an immediate and constant effect, and a simple correlation structure, and each assumption can be relaxed.

Several questions at once

Factorial designs and waitlist controls

2 × 2 factorial

Connector and transportation voucher are randomized independently, so each participant informs both main effects.

Waitlist control

The comparison lasts only as long as the delay, and waiting can change behaviour.

Detecting an interaction as large as a main effect needs about four times the sample size.

Adapting as data arrive

Adaptive designs and SMARTs

  • Adaptive designs allow pre-specified changes, such as stopping early or dropping arms (Pallmann et al., 2018).
  • A SMART randomizes people at more than one decision point to build adaptive interventions (Murphy, 2005).
  • In a Cedar Valley SMART, older adults not engaged at week six could be re-randomized to a peer volunteer or to transportation help.
Not worse by too much

Non-inferiority and the margin

0.42Established effect of in-person meetings in the simulated trial (points)
0.20A defensible margin that preserves about half of that effect

Non-inferiority is shown when the upper confidence limit lies below the margin. Both intention-to-treat and per-protocol results are reported.

Ethics

Randomizing access to a program

Equipoise

Informed people genuinely disagree about whether the program helps (Freedman, 1987).

Scarcity

When capacity is limited, a lottery allocates waiting fairly.

Ottawa Statement

Gatekeepers may permit participation but cannot consent for individuals (Weijer et al., 2012).

Carry forward

From design to analysis

  • A staggered rollout can be randomized as a waitlist trial or a stepped-wedge trial.
  • Stepped-wedge analyses must model time.
  • Each alternative design answers a distinct question, and randomizing access must be ethically justified.

Learning Objectives for this section

  • Describe the stepped-wedge cluster randomized design and explain why its analysis must model time trends.
  • Show how a staggered rollout, such as the Cedar Valley program's two waves, can be randomized as a parallel waitlist comparison or as a stepped-wedge trial.
  • Describe factorial designs and sequential multiple assignment randomized trials, and the questions each one answers.
  • Explain the logic of a non-inferiority trial, including the choice of margin and the role of the analysis population.
  • Assess the ethics of randomizing access to a program using clinical equipoise, the Ottawa Statement and TCPS 2.

3.1 Randomizing a Staggered Rollout

Many programs cannot start everywhere at once. Staff must be hired and trained, budgets are released in stages, and managers prefer to learn from early sites before expanding. When the order in which sites start is not fixed by need or readiness, randomizing that order costs the program very little and produces a randomized comparison. Evaluators who join a program early can often negotiate this, and it is one of the most practical routes to a randomized design in health services.

The fictional Cedar Valley Connector program, which runs through this course, has exactly this structure: 12 clinics in a first wave and 12 a year later. Figure 3.1 shows two ways the rollout could have been randomized. In panel A, the health authority randomizes which 12 clinics join in wave one. During year one, the wave-two clinics act as a waitlist control (also called a delayed-intervention control), and the comparison between the two sets of clinics during that year is a parallel cluster randomized trial. This is the design analyzed in the Section 2 worked example. In panel B, the rollout is spread across the first year in four steps of six clinics each, with the order of the steps randomized. This is a stepped-wedge design.

A. Randomized two-wave rollout (parallel trial with a waitlist control in year one) Baseline Year one Year two Wave one (12 clinics) Usual care Connector Connector Wave two (12 clinics) Usual care Usual care Connector Randomized comparison during year one B. Stepped-wedge alternative (four sequences of six clinics, three-month steps) Baseline Step 1 Step 2 Step 3 Step 4 Sequence 1 (6 clinics) Sequence 2 (6 clinics) Sequence 3 (6 clinics) Sequence 4 (6 clinics) Connector program delivered Usual care (not yet started)
Figure 3.1. Two randomized versions of a staggered rollout of the Cedar Valley Connector program. Panel A randomizes clinics to two waves a year apart, and panel B randomizes four sequences of six clinics that start at three-month intervals.

If outcome data are also collected before wave one and during year two, the two-wave design in panel A becomes a design with two sequences and three periods, in which every clinic is observed both without and with the program. With only two sequences, most of the information about the program effect still comes from the year-one comparison between arms, but the before-and-after data within clinics add information, and they also bring the problem of time trends that dominates the analysis of stepped-wedge trials.

3.2 The Stepped-Wedge Cluster Randomized Design

In a stepped-wedge design, all clusters begin in the control condition. At regular intervals, called steps, a randomly selected group of clusters crosses over to the intervention, and by the end of the trial every cluster has received it. Outcomes are measured in every cluster in every period (Hussey & Hughes, 2007; Hemming et al., 2015). Each cluster therefore contributes observations under both conditions, and the program effect is estimated from comparisons between clusters within periods and from comparisons within clusters over time. Copas and colleagues (2015) distinguish designs by how participants are recruited and followed: a closed cohort identified at the start and followed throughout, an open cohort in which people join and leave over time, and continuous recruitment of new participants who are each exposed for a short time. A Cedar Valley stepped-wedge trial that enrols the older adults referred in each period would recruit continuously, with each person contributing outcomes to the period in which they were referred.

The design is attractive for health services programs for several reasons. Every cluster eventually receives the program, which makes the design acceptable to managers and communities who would refuse to be a permanent control. The phased rollout matches the way programs are often implemented. Because every cluster is observed before and after it starts, cluster-level differences contribute less to the error of the estimate. Whether a stepped-wedge design needs fewer clusters than a parallel design depends on the ICC, the cluster size and the number of steps, so the comparison should be made for the specific trial using appropriate methods (Hussey & Hughes, 2007; Hooper et al., 2016) or simulation.

Time trends and the analysis model

The defining analytical problem of the stepped-wedge design is that exposure to the program is confounded with calendar time. In the first period no clusters are exposed, and in the last period all are. Any change in the outcome over the course of the trial that is unrelated to the program will therefore look like a program effect unless the analysis separates the two. Loneliness among older adults could change over a year because of the seasons, a provincial initiative on social isolation, a change in transit service or a public health emergency. The standard analysis, proposed by Hussey and Hughes (2007), includes a fixed effect for each period alongside the program indicator and a random effect for each cluster.

The Hussey and Hughes model for a stepped-wedge trial

Yijt = μ + βt + θ Xjt + uj + eijt

Yijt is the outcome for individual i in cluster j in period t, βt is the fixed effect of period t, Xjt equals 1 if cluster j has started the program by period t and 0 otherwise, θ is the program effect, uj is a random cluster effect and eijt is individual error. The period effects absorb secular trends that are common to all clusters, so θ is estimated from the contrast between exposed and unexposed clusters within the same periods.

The model makes three assumptions that deserve scrutiny in a program evaluation. It assumes that the secular trend is the same in every cluster, so a trend that differs between rural and urban clinics would bias the estimate. It assumes that the program effect is immediate and constant, whereas a connector program may take months to build caseloads and community partnerships, and recent methodological work has shown that assuming an immediate, constant effect can bias the estimate when the effect actually changes with time since the program started. It also assumes, in its simplest form, that the correlation between observations in the same cluster is the same whether they come from the same period or from periods a year apart, and extensions allow that correlation to weaken over time. Each assumption can be relaxed, at a cost in precision, and the chosen model should be specified in the protocol.

Two practical risks also arise. Clusters may not start on the randomized date, because hiring or training is delayed, which blurs the contrast between periods. Measurement must also continue in every cluster in every period, which is more burdensome than a parallel design. The CONSORT extension for stepped-wedge trials (Hemming et al., 2018) asks authors to explain why the design was chosen, to show the schedule of sequences and periods in a diagram, and to describe how time effects were modelled in the analysis.

A non-randomized staggered rollout raises the same analytical issue in a sharper form. Lesson 7 shows that two-way fixed effects models applied to staggered adoption can give misleading estimates when program effects vary across groups or over time. The stepped-wedge design shares the staggered structure but adds randomization of the order, which protects the comparison from selection on readiness.

Clinics are randomized to start immediately or after a fixed delay, and the arms are compared during the delay. For Cedar Valley, 12 clinics would start in wave one and 12 would wait a year. The comparison is simple to analyze with the methods of Section 2 and is not confounded with time, because both arms are observed over the same year. Its limitation is that the comparison ends when the waitlist arm starts, so it cannot estimate effects beyond the delay period.

Clinics are randomized to the order in which they cross over at several steps, and every clinic is measured in every period. For Cedar Valley, four sequences of six clinics would start at three-month intervals. The design uses within-clinic as well as between-clinic comparisons and gives every clinic the program within the trial period, but the analysis must model time trends, and the results depend on assumptions about how the effect changes with exposure time.

The health authority chooses the order of rollout on readiness or need. The evaluator can still compare early and late clinics using the difference-in-differences and event-study methods of Lesson 7, but the credibility of the estimate then rests on the assumption that early and late clinics would have followed parallel trends without the program, which randomization would have made unnecessary.

3.3 Waitlist Controls

A waitlist control can be used with individuals or with clusters. It is common when a program has more applicants than places, because people on the waitlist would be waiting in any case and the trial simply determines the order of service at random. It is also acceptable to participants and partners, since everyone is promised the program.

The design has three limitations. First, it can estimate effects only for the length of the delay, so a six-month waitlist cannot show whether benefits last a year. Second, people who know they will receive the program soon may behave differently from people receiving usual care with no prospect of the program. They may postpone seeking other help or report disappointment at being made to wait, either of which can make the program look more effective than it would against usual care. Third, people may leave the waitlist before follow-up, and if those who leave differ from those who stay, the comparison is biased. Evaluators can reduce these problems by measuring outcomes for everyone regardless of whether they remain on the list, by describing to waitlisted participants what usual care includes, and by keeping the delay as short as the program's capacity allows.

3.4 Factorial Designs

A factorial design randomizes participants or clusters to two or more interventions at once, so that every combination is represented. In a 2 × 2 factorial design, each unit is randomized to receive or not receive intervention A and, independently, to receive or not receive intervention B. Suppose the Cedar Valley Health Authority wants to test both the connector program and a transportation voucher for older adults who live far from community activities.

No transportation voucherTransportation voucher
Usual careCell 1: neither componentCell 2: voucher only
Connector programCell 3: connector onlyCell 4: connector and voucher

The main effect of the connector program compares cells 3 and 4 with cells 1 and 2, and the main effect of the voucher compares cells 2 and 4 with cells 1 and 3. If the two components do not interact, every participant contributes to both comparisons, so the factorial trial answers two questions with roughly the sample size needed for one. The efficiency depends on the absence of an important interaction. If vouchers help only when a connector has first identified an activity worth travelling to, the effect of the voucher depends on the presence of the connector, and the main effect averaged over both connector conditions becomes harder to interpret. Detecting an interaction as large as a main effect requires about four times the sample size needed to detect the main effect, so most factorial trials are powered to estimate main effects and can only describe interactions approximately. Factorial experiments are also central to the multiphase optimization strategy, in which investigators screen several candidate components of a program before assembling and testing the final package (Collins et al., 2007).

3.5 Adaptive Designs and Sequential Multiple Assignment Randomized Trials

An adaptive design allows pre-specified changes to a trial in response to accumulating data, without undermining its validity (Pallmann et al., 2018). Common adaptations include stopping early for benefit or futility at planned interim analyses, re-estimating the sample size, dropping arms that are performing poorly, and changing allocation ratios to favour better-performing arms. Platform trials extend the idea by adding and removing arms under a single master protocol. Adaptive designs require the adaptation rules, the interim analysis schedule and the methods that control the type I error rate to be specified in advance, and they usually need an independent data monitoring committee. The Adaptive designs CONSORT Extension (ACE; Dimairo et al., 2020) describes how to report them.

A sequential multiple assignment randomized trial (SMART) answers a different kind of question. Many programs are adaptive by nature: staff change what they offer according to how a person responds. A SMART randomizes participants at more than one decision point so that investigators can construct and compare adaptive interventions, which are sequences of decision rules that specify what to offer, to whom and when (Murphy, 2005; Collins et al., 2007).

Referred older adults R In-person connector meetings Telephone connector meetings Week six: engagement assessed Engaged: continue Not engaged R Engaged: continue Not engaged R Add peer volunteer Add transportation help Add peer volunteer Add transportation help
Figure 3.2. A hypothetical SMART for the Cedar Valley Connector program. The circled R marks each point of randomization. Older adults who are not engaged at week six are randomized a second time.

In the hypothetical SMART in Figure 3.2, referred older adults are first randomized to in-person or telephone connector meetings. At week six, the connector records whether each person is engaged, defined as having attended at least two meetings and taken a first step on their plan. Engaged participants continue with their assigned format. Participants who are not engaged are randomized a second time to receive either a peer volunteer companion or transportation help. The design embeds four adaptive interventions, such as "start with telephone meetings and add a peer volunteer for those not engaged by week six", and allows them to be compared. Typical primary aims compare the two first-stage options, or compare the two second-stage options among people who were not engaged.

3.6 Non-Inferiority Trials

Most trials ask whether a new option is better than an existing one. A non-inferiority trial asks whether a new option that is cheaper, easier to scale or more acceptable is not worse than the existing option by more than a pre-specified amount, called the non-inferiority margin. For Cedar Valley, the question might be whether connector meetings delivered by telephone, which would let one connector serve more clinics, are non-inferior to in-person meetings for six-month loneliness.

The margin is the largest loss of effect that decision-makers would accept in exchange for the advantages of the new option. It should be justified on clinical grounds and on evidence about how much the existing option improves outcomes compared with no program. The margin must be smaller than that established effect, or the trial could declare non-inferior an option that is no better than usual care. In the simulated Section 2 trial, in-person delivery lowered six-month scores by about 0.42 points relative to usual care. A margin of 0.5 points would therefore be indefensible, while a margin of 0.2 points would preserve roughly half of the established benefit. The trial concludes non-inferiority if the upper limit of the confidence interval for the difference (telephone minus in-person, where higher scores are worse) lies below the margin.

Telephone worse by more than the margin A Non-inferior B Non-inferior and superior C Non-inferior but worse than in-person D Inconclusive E Inferior −0.4 −0.2 0.0 0.2 0.4 0.6 Difference in six-month UCLA score, telephone minus in-person (points) margin = 0.2 no difference
Figure 3.3. Five hypothetical results of a non-inferiority trial with a margin of 0.2 points. Non-inferiority is shown when the whole confidence interval lies below the margin; result C shows that an option can be non-inferior and still statistically worse than the standard.

Two further features distinguish non-inferiority trials. First, an intention-to-treat analysis is not conservative here. When participants in both arms fail to engage, the arms look more alike, which favours a conclusion of non-inferiority. The CONSORT extension for non-inferiority and equivalence trials (Piaggio et al., 2012) therefore recommends reporting both intention-to-treat and per-protocol analyses, and a conclusion of non-inferiority is more credible when both support it. Second, the trial cannot show that the existing option worked in this particular study, so it relies on an assumption that the effect of in-person delivery is similar to its effect in the earlier trial that established it. Because margins are small relative to the variability of the outcome, non-inferiority trials usually need larger samples than superiority trials of the same outcome.

3.7 The Ethics of Randomizing Access to a Program

Randomizing access to a health program requires an ethical justification. Freedman (1987) located the justification for clinical trials in clinical equipoise, a state of honest professional disagreement in the expert community about the comparative merits of the options being tested. For health services programs, the corresponding condition is genuine uncertainty, shared by informed people, about whether the program improves outcomes relative to usual care, or about whether it does so at an acceptable cost.

Scarcity strengthens the case for randomization. When a program cannot reach everyone at once, some people will wait whether or not there is a trial, and a lottery gives each eligible person or clinic an equal chance of being served first. Randomizing the order of a rollout that must be staggered anyway withholds nothing that would otherwise have been provided. The argument weakens if the program is already known to be effective, or if the trial delays the rollout beyond what capacity requires.

Cluster trials raise further questions, which the Ottawa Statement on the ethical design and conduct of cluster randomized trials addresses (Weijer et al., 2012). The Statement asks investigators to identify who the research participants are, a group that can include people who receive the intervention, people whose environment or care is deliberately altered, people with whom investigators interact, and people about whom identifiable data are collected. It asks that informed consent be sought from research participants unless a research ethics board approves a waiver or alteration, which generally requires that the research be impracticable without it and involve no more than minimal risk. It distinguishes gatekeepers, such as a clinic manager or a health authority executive, who may grant permission for a cluster to take part, from individuals, on whose behalf gatekeepers cannot consent. In Canada, TCPS 2 sets out the conditions for altering consent requirements and, in Chapter 9, the obligations for research involving First Nations, Inuit and Métis peoples.

Is there genuine uncertainty about the program?v

The evaluator should document the evidence on similar programs and the views of clinicians, partners and older adults. Evidence that social prescribing programs vary widely in their effects can support a claim of uncertainty, while a well-replicated effect in comparable settings would weaken it.

Would access be limited without the trial?v

The Cedar Valley program has funding for 12 clinics in its first year. Randomizing which clinics start first changes who waits, and it does not change how many wait.

Who are the research participants?v

Older adults who complete the surveys are research participants. Clinicians whose referral practices the program changes, and connectors whose meetings are logged for the fidelity assessment, may also be research participants under the Ottawa Statement's definition, and the research ethics board will need to decide what consent each group requires.

From whom is consent needed, and when?v

Clinic managers can agree to their clinic's participation before randomization. Older adults can consent to the research surveys after referral. Whether the program itself, offered as a health service, requires research consent is a question for the research ethics board, which may approve an alteration of consent if the conditions in TCPS 2 are met.

How are Indigenous partners involved?v

The Cedar Valley program has an Indigenous health partnership with local First Nations. Decisions about randomizing clinics that serve First Nations communities, and about the collection and governance of data from First Nations participants, belong with that partnership under TCPS 2 Chapter 9 and the OCAP® principles discussed in Lesson 4.

What happens to participants with diminished capacity?v

Some older adults referred for loneliness will have cognitive impairment. The protocol should describe how capacity is assessed, when an authorized third party may give consent, and how the person's own wishes are respected.

What happens after the trial?v

The waitlist or later-wave clinics should receive the program as promised, and the results should be shared with participating clinics, older adults and partners in an accessible form.

Case: A councillor objects to the lottery

At a public meeting, a municipal councillor argues that the clinics with the longest waitlists for social services should receive the Cedar Valley program first, and that a lottery is unfair to them. The evaluator responds that the health authority could group clinics into strata of need and randomize within each stratum, using a higher probability of starting first in the highest-need stratum (for example, six of eight high-need clinics in wave one). Need would then shape the rollout, the comparison would remain randomized within each stratum, and the analysis would compare arms within strata. If the steering committee instead decides that need alone must determine the order, the evaluator will plan a quasi-experimental evaluation, such as the difference-in-differences and regression discontinuity designs in Lessons 7 and 8, and will document that choice and its consequences for the strength of the evidence.

Section 4 turns to the analysis and reporting of randomized evaluations, and to the structured judgement an evaluator makes about whether randomization is feasible for a program at all.

Reflection

A regional health authority will introduce a falls-prevention exercise program into 16 long-term care homes. Training capacity allows the program to start in 4 homes every 4 months, so the full rollout will take 16 months. Falls in long-term care vary with the season, and a provincial falls-prevention campaign is scheduled to begin partway through the rollout. All 16 homes are known in advance and have not yet been told when they will start. Propose a randomized design that fits this rollout, describe how the analysis would handle time trends, and identify two ethical or practical considerations the evaluator should raise with the health authority.

Model answer

The rollout fits a stepped-wedge cluster randomized design. After a four-month baseline period in which no home has the program, the 16 homes would be randomized into four sequences of four homes, and one sequence would start at each four-month step, giving five periods in all. Falls would be measured in every home in every period, ideally from incident reports that homes already collect. Randomization could be stratified by home size so that each sequence contains a similar mix.

Because the proportion of homes with the program rises over time, seasonal variation and the provincial campaign would look like program effects unless the analysis separates them. The analysis would therefore include a fixed effect for each period, the program indicator and a random effect for each home, following Hussey and Hughes. I would also examine whether the effect grows with time since a home started, because exercise programs take time to build participation.

Two considerations are consent and fidelity of the schedule. Many residents have cognitive impairment, administrators can permit their home's participation but cannot consent for residents, and the research ethics board may need to consider an alteration of consent for the use of routine falls data. Homes may also start late if training is delayed, which would blur the comparison, so the plan should record actual start dates.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Why must the analysis of a stepped-wedge trial include period effects?

In a stepped-wedge trial, no clusters are exposed in the first period and all are exposed in the last, so exposure is confounded with time. Period effects absorb trends common to all clusters. Clusters are randomized once to a sequence (option a), and the random cluster effect is still needed (option c).

Question 2: A 2 × 2 factorial trial tests a connector program and a transportation voucher. What is the main advantage of the design if the two components do not interact?

Without an interaction, each participant contributes to the comparison for each component, so the trial answers two questions with roughly the sample needed for one. Estimating an interaction as large as a main effect needs about four times the sample size (option a), and dropping arms at interim analyses is a feature of adaptive designs (option c).

Question 3: A non-inferiority trial will compare telephone with in-person connector meetings. In-person meetings lower six-month loneliness scores by about 0.42 points compared with usual care. Which non-inferiority margin is most defensible?

The margin must be smaller than the established effect of the standard option, or the trial could declare non-inferior an option no better than usual care. A margin of 0.50 (option a) exceeds the 0.42-point effect, and a margin of zero (option d) turns the question into a superiority test.

Question 4: According to the Ottawa Statement, what may a clinic manager acting as a gatekeeper do in a cluster randomized trial?

The Ottawa Statement distinguishes gatekeepers, who may permit a cluster's participation, from research participants, for whom gatekeepers cannot consent. Waivers or alterations of consent are decided by a research ethics board, and allocation must remain random.
Section 4 of 5

Analysis, Reporting and Feasibility

⏱ Estimated reading time: 40 minutes
Section 4 of 5

Analysis, Reporting and Feasibility

The effect of offering a program, the effect of taking part, how to report a trial, and when to recommend randomization.

Offer versus receipt

Three ways to analyze

Intention-to-treat

Everyone is analyzed as randomized, which estimates the effect of offering the program.

Per-protocol

Only adherent participants are compared, which breaks the randomized comparison.

As-treated

People are grouped by what they received, which makes the comparison observational.

The effect of taking part

The complier average causal effect

One-sided noncompliance
\[ \text{CACE} = \frac{\text{ITT effect}}{\text{proportion of compliers}} \]

The estimate assumes random assignment, the exclusion restriction, monotonicity and no interference (Angrist, Imbens & Rubin, 1996).

The simulated results

Three estimates from one trial

−0.420Intention-to-treat
−0.642Complier average causal effect (65.5% engaged)
−0.772Per-protocol

The simulation built in an effect of about −0.62 points among engaged older adults.

Reporting

CONSORT 2025 and its extensions

  • CONSORT 2025 (Hopewell et al., 2025) is the current statement for reporting randomized trials.
  • Design-specific extensions add items for cluster trials (2012), stepped-wedge trials (2018), pragmatic trials (2008), non-inferiority trials (2012) and adaptive designs (2020).
  • Protocols follow SPIRIT 2025 (Chan et al., 2025), and trials are registered before the first participant is enrolled.
Feasibility

Ten questions before recommending randomization

Decision needs a causal estimate
Genuine uncertainty
Allocation still open
Natural point of control
Level of delivery
Enough units
Acceptability
Equal measurement
Ethics and governance
Cost and fallback
Worked example

A randomization feasibility assessment

A feasibility assessment judges whether a program can be randomized and, if it can, specifies the unit and scheme.

  • The Cedar Valley example randomizes 24 clinics 1:1 to two waves, stratified by rurality.
  • An independent statistician generates the allocation after all clinics sign on.
  • The loneliness screen starts in every clinic before randomization to prevent recruitment bias.
Carry forward

On to the final assessment

  • The intention-to-treat estimate answers the funding question.
  • The complier average causal effect describes the effect of taking part under stated assumptions.
  • Feasibility depends most often on timing, the number of units and acceptability.

Learning Objectives for this section

  • Distinguish intention-to-treat, per-protocol and as-treated analyses, and explain what each one estimates.
  • Calculate a complier average causal effect under one-sided noncompliance and state the assumptions it requires.
  • Describe how the current CONSORT statement and its extensions for cluster and stepped-wedge trials govern the reporting of a randomized evaluation.
  • Apply a structured set of questions to judge whether randomization is feasible for a program.
  • Specify the unit and scheme of randomization for a program evaluation.

4.1 When People Do Not Take Up the Program

Randomization determines what each unit is offered. In a program evaluation, many people who are offered a program do not take it up, and some who are not offered it find a similar service elsewhere. In the simulated Cedar Valley trial from Section 2, some older adults referred at wave-one clinics never met a connector or stopped after one meeting. The analysis has to decide how to treat them, and the decision changes the question the analysis answers.

AnalysisWho is comparedWhat it estimatesMain weakness
Intention-to-treatEveryone, in the arm to which they (or their cluster) were randomized, regardless of what they received.The effect of offering the program, which is the effect of the policy decision.The effect among people who take up the program is diluted by those who do not.
Per-protocolOnly people who followed the protocol in each arm, such as engaged participants in the Connector arm and all participants in the usual-care arm.An effect among adherent participants, if adherence were unrelated to prognosis.Adherence is usually related to prognosis, so the comparison is no longer randomized.
As-treatedPeople grouped by what they actually received, regardless of assignment.An effect of receipt, if receipt were unrelated to prognosis.The comparison is observational and open to confounding.

The intention-to-treat principle analyzes every randomized participant in the arm to which they were allocated. It preserves the comparison that randomization created and answers the question a health authority asks when it decides whether to offer a program: what happens to the population that is offered it, given that not everyone will participate. The intention-to-treat estimate is the primary analysis in almost all pragmatic trials.

The ICH E9(R1) addendum on estimands (International Council for Harmonisation, 2019) asks trialists to define the target of estimation precisely, including how events that occur after randomization, such as non-engagement, are handled. In its terms, an intention-to-treat analysis follows a treatment-policy strategy, in which non-engagement is part of what is being evaluated. The addendum describes other strategies, including the principal stratum strategy that underlies the complier average causal effect described below. Writing the estimand down before the trial begins keeps the evaluator and the steering committee clear about which question the primary analysis answers.

Intention-to-treat analysis also requires outcome data on everyone randomized. When some six-month surveys are missing, the evaluator should report how many are missing in each arm, use a principled method such as multiple imputation that respects the clustering, and run sensitivity analyses that test whether plausible departures from the imputation assumptions would change the conclusion.

4.2 The Complier Average Causal Effect

Decision-makers often want a second number: the effect of the program on the people who actually take part. The per-protocol analysis does not supply it, because engaged and non-engaged participants differ in ways that also affect outcomes. Angrist, Imbens and Rubin (1996) showed how to estimate an effect among participants who take up a program while keeping the protection of randomization. Their approach classifies participants by how they would behave under each possible assignment.

CompliersClick to explore
Never-takersClick to explore
Always-takersClick to explore
DefiersClick to explore

The groups are called principal strata. Membership is unobserved for most individuals: an older adult in a usual-care clinic could be a complier or a never-taker, and nothing in the data reveals which. Because randomization balances the strata between arms, their proportions in the usual-care arm are expected to equal their proportions in the Connector arm. When usual-care participants cannot obtain the program, a situation called one-sided noncompliance, the proportion of compliers equals the proportion of Connector-arm participants who engaged.

Complier average causal effect with one-sided noncompliance

CACE = ITT effect / proportion of compliers

The estimator requires four assumptions: random assignment; the exclusion restriction, under which assignment affects outcomes only through engagement; monotonicity, under which there are no defiers; and no interference between participants. Under these assumptions, never-takers contribute nothing to the intention-to-treat difference, so the whole difference is produced by compliers.

The logic is easiest to see with round numbers. Suppose the intention-to-treat effect is −0.40 points and half of the Connector-arm participants engaged. If never-takers experience no effect, the −0.40 average must come from the engaged half, which implies an effect of −0.80 among compliers. The CACE is the instrumental variable estimate with randomization as the instrument, an idea that Lesson 8 develops in the context of natural experiments.

The exclusion restriction deserves attention in a cluster trial. If clinicians in Connector clinics begin to ask every older patient about loneliness, or if the clinic starts hosting community group activities, older adults who never meet a connector may still be affected by their clinic's allocation. Never-takers would then contribute to the intention-to-treat difference, and dividing that difference by the proportion engaged would attribute their change to compliers, overstating the complier effect.

Worked example: Three estimates from the simulated Cedar Valley trial

In the simulated trial from Section 2.5, older adults in the Connector arm who engaged were less lonely at enrolment (mean 5.887) than those who did not (mean 6.573), so engagement was selective. The proportion engaged in the Connector arm is 0.655. The intention-to-treat estimate, from the linear mixed model in Section 2.5, is −0.420 points. A per-protocol mixed model, which drops the Connector-arm participants who did not engage, gives −0.772 points. The CACE divides the intention-to-treat estimate by the proportion engaged (−0.420 / 0.655, computed from the unrounded values) and equals −0.642 points.

Because the data are simulated, the true values are known. The simulation built in a reduction of about 0.62 points for older adults who engaged and no effect for those who did not, and it made engagement more likely among people with lower baseline loneliness and greater mobility, a characteristic that the dataset does not record. The CACE recovers the built-in effect closely. The per-protocol estimate overstates it, even after adjustment for baseline loneliness, because engaged participants differ from the whole usual-care arm in mobility, which also lowers loneliness. The CACE is a point estimate here; its confidence interval requires instrumental variable methods or a bootstrap that resamples clinics.

simulated effect among engaged participants (about −0.62) Intention-to-treat −0.420 Complier average causal effect −0.642 Per-protocol −0.772 −1.0 −0.8 −0.6 −0.4 −0.2 0.0 0.2 Estimated difference in six-month UCLA score (points; negative favours the program)
Figure 4.1. Three estimates from the simulated Cedar Valley trial. The intention-to-treat estimate targets the effect of offering the program, and the other two target the effect among participants who engaged, which the simulation set at about −0.62 points.

For the Cedar Valley steering committee, the intention-to-treat estimate is the primary result, because the decision is whether to offer the program. The CACE is a useful secondary result for program managers who want to know what engagement achieves, provided its assumptions are stated. The per-protocol estimate belongs, at most, among the sensitivity analyses, except in a non-inferiority trial, where Section 3 explained why it is reported alongside the intention-to-treat analysis.

4.3 Reporting Randomized Evaluations

The CONSORT (Consolidated Standards of Reporting Trials) statement sets out the minimum information a report of a randomized trial should contain, in the form of a checklist and a flow diagram. CONSORT 2025 (Hopewell et al., 2025) is the current statement and updates CONSORT 2010. Design-specific extensions add items for particular designs. The extensions listed below were developed for CONSORT 2010 and are used alongside the current statement for design-specific items, and authors should check whether an updated version of a relevant extension has been published. Protocols are reported with the SPIRIT 2025 guideline (Chan et al., 2025), which updated SPIRIT 2013, trials should be registered in a public registry such as ClinicalTrials.gov or the ISRCTN registry before the first participant is enrolled, and Lesson 10 introduces TIDieR for describing the intervention itself.

Design featureReporting guidelineWhat it adds
Any randomized trialCONSORT 2025 (Hopewell et al., 2025)The core checklist and the participant flow diagram.
Clusters randomizedCluster extension (Campbell et al., 2012)The rationale for clustering, clustering in sample size and analysis, flow of clusters and individuals, and the ICC.
Staggered crossover of clustersStepped-wedge extension (Hemming et al., 2018)A diagram of sequences and periods, and the handling of time effects.
Usual-care setting and comparatorPragmatic trials extension (Zwarenstein et al., 2008)A description of the setting, participants and comparator that lets decision-makers judge applicability.
Non-inferiority questionNon-inferiority and equivalence extension (Piaggio et al., 2012)The margin and its justification, and both intention-to-treat and per-protocol results.
Pre-planned adaptationsACE (Dimairo et al., 2020)The adaptation rules, interim analyses and methods for controlling error rates.
Behavioural or service interventionNonpharmacologic treatments extension (Boutron et al., 2017)Details of intervention delivery, providers and centres.

A cluster trial's flow diagram reports two levels, clusters and the individuals within them, at each stage from allocation to analysis. Figure 4.2 shows the structure for the simulated Cedar Valley trial, whose follow-up was complete by construction.

Clinics randomized 24 primary care clinics Allocation Wave one (Connector) 12 clinics 475 older adults enrolled Wave two (usual care in year one) 12 clinics 485 older adults enrolled Follow-up Six-month follow-up Clinics lost: 0 Older adults lost: 0 (simulation) Six-month follow-up Clinics lost: 0 Older adults lost: 0 (simulation) Analysis Analyzed (intention-to-treat) 12 clinics 475 older adults Analyzed (intention-to-treat) 12 clinics 485 older adults A real trial would report clinics and individuals excluded, lost and analyzed at each stage, with reasons, and the number of individuals per cluster (Campbell et al., 2012).
Figure 4.2. A two-level flow diagram for the simulated Cedar Valley cluster trial, using the clinic and participant counts reported in the worked example in Section 2.5.

4.4 Deciding Whether Randomization Is Feasible

Randomization gives the strongest protection against confounding, but an evaluator recommends it only after judging that it is feasible, ethical and useful for the decision at hand. The questions below structure that judgement. They build on the evaluability assessment introduced in Lesson 4 and lead, when the answers are unfavourable, to the quasi-experimental designs of Lessons 7 and 8.

1. What decision will the evaluation inform, and does it need a causal estimate?v

A decision about whether to fund, expand or end a program usually needs an estimate of its effect. A decision about how to improve delivery may be better served by process evaluation (Lesson 5) or quality improvement methods (Lesson 8).

2. Is there genuine uncertainty about the program's effect?v

Randomization is justified when informed people disagree about whether the program improves outcomes relative to usual care. If the effect is already well established in similar settings, a trial adds little and may be difficult to defend ethically.

3. Can allocation still be influenced?v

Timing is the most common barrier. Once a program has been offered to everyone, or once sites have been chosen, the opportunity to randomize who receives it first has passed. Evaluators who are involved before launch have far more options.

4. Is there a natural point of control?v

Staggered rollouts, oversubscribed programs, new funding that cannot cover every site, and new components added to an existing program all create points at which allocation can be randomized at little cost.

5. At what level is the program delivered?v

The unit of randomization should be the smallest unit that avoids contamination and matches how the program is delivered. Individuals, clinicians, clinics and communities are all possible units.

6. Are there enough units?v

A cluster trial needs enough clusters to deliver adequate power and to make chance imbalance unlikely. With three clusters per arm, there are only 20 possible allocations, so a randomization test can never produce a two-sided p-value below 0.10. The minimum number of clusters per arm, nindividual × ICC, sets a floor that larger clusters cannot lower.

7. Will partners and communities accept randomization?v

Managers, clinicians, Indigenous partners, older adults and elected officials must find the allocation fair. Randomizing the order of a rollout, stratifying by need and promising the program to every site usually increase acceptance.

8. Can outcomes be measured equally in all arms within the decision timeline?v

The same instruments, schedule and data sources must apply in every arm, and results must arrive before the decision they are meant to inform. Administrative data can reduce burden and cost when they capture the outcomes of interest.

9. What do ethics review and governance require?v

The research ethics board, data stewards and Indigenous partners must approve the consent approach, data flows and governance arrangements. Section 3 listed the specific questions for cluster trials.

10. What will it cost, and what is the fallback design?v

A randomized evaluation costs more to plan and manage than a routine monitoring report. The evaluator should name the quasi-experimental design that would be used if randomization proves infeasible, so the decision is made with a clear sense of what is gained.

When randomizing the program itself is not possible, related randomized designs often remain available. An encouragement design randomizes invitations, reminders or help to enrol in a program that is open to everyone, and the CACE then estimates the effect among people whom the encouragement moved to participate. Evaluators can also randomize the timing of access, an enhancement to the program, such as the transportation voucher in Section 3, or the order of service among applicants who exceed program capacity.

4.5 Worked Example: The Cedar Valley Randomization Assessment

A randomization feasibility assessment judges whether a randomized design is feasible for a program and, if it is, specifies the unit and scheme of randomization. The worked example below shows what that assessment could look like for the fictional Cedar Valley Connector program, written as if the evaluator joined the planning team six months before wave one, when the wave-one clinics had not yet been chosen.

Worked example: Feasibility assessment for the Cedar Valley Connector program

Decision and question. The steering committee will decide, after the first year, whether to continue the program and how to configure it for the region. The primary evaluation question is whether offering the program to referred older adults reduces loneliness on the three-item UCLA Loneliness Scale at six months, compared with usual primary care.

Feasibility judgement. Randomization is feasible. Evidence on community connector programs is mixed, so informed people disagree about the likely effect. Funding covers 12 clinics in year one, so half of the clinics will wait regardless of the evaluation, and the order of entry has not been fixed. The program is delivered through clinics, so individual randomization would risk contamination through clinicians and waiting rooms. With 24 clinics of about 40 referred older adults each and an ICC of 0.06, a parallel comparison has 80 percent power to detect a 0.5-point difference (Section 2.4). Outcomes can be collected identically in all clinics by central telephone interviewers within the first year.

Unit and scheme. The unit of randomization is the clinic. The 24 clinics will be allocated 1:1 to wave one or wave two, stratified by rurality (8 rural and 16 urban clinics, giving 4 rural and 8 urban clinics per wave). An independent statistician outside the program team will generate the allocation with a seeded computer script after all 24 clinics have signed participation agreements, and the allocation will be revealed at a steering committee meeting. To prevent recruitment bias, the loneliness screen will be introduced in all 24 clinics before randomization, and a research coordinator who works from referral lists will approach eligible older adults in both arms in the same way.

Analysis and reporting. The primary analysis will be an intention-to-treat linear mixed model with a random clinic intercept, adjusted for baseline loneliness and the rurality stratum, with 21 degrees of freedom (24 clinics minus the intercept, the arm indicator and the stratum indicator). A CACE for engagement, defined as attending at least two connector meetings, will be a secondary analysis. The report will follow CONSORT 2025 and the cluster extension.

Ethics and partners. The Indigenous health partnership will review the randomization plan before it is finalized, including whether clinics that serve First Nations communities are included in the lottery, and the data governance arrangements for First Nations participants will follow the partnership's agreements. Clinic managers will give permission for their clinics' participation, and older adults will give consent for the research surveys. The research ethics board will decide whether the program itself requires research consent.

If wave one had already been chosen. If the health authority had already selected the wave-one clinics, the parallel comparison would no longer be randomized. The evaluator could still randomize the start dates of the 12 wave-two clinics across year two, in three steps of four clinics, which would create a small stepped-wedge trial within wave two, and would analyze the comparison between wave-one and wave-two clinics with the difference-in-differences methods of Lesson 7.

What a randomization feasibility assessment contains

A feasibility assessment works through the ten questions in Section 4.4, states whether a randomized design is feasible and justifies the judgement. If it is feasible, the assessment specifies the unit of randomization, the allocation ratio, the randomization scheme (including any stratification or constraint), how allocation will be concealed, the number of units and the planning values behind it, and the primary analysis. If it is not feasible, the assessment names the barrier and the randomized or quasi-experimental alternative, the designs that Lessons 7 and 8 describe.

A sound assessment lets the judgement follow from the program's decision context, timing and delivery structure, and it matches the unit of randomization to the level at which the program is delivered so that contamination is addressed. It shows any sample size or cluster calculation, with sourced planning values for the ICC and the standard deviation, and it addresses ethical and partner considerations, including Indigenous governance where relevant, specifically. It is written clearly enough for a steering committee to act on.

Reflection

A cluster randomized trial of a peer-support program for family caregivers randomized 20 community centres, 10 to the program and 10 to usual services. Caregivers in usual-service centres had no access to peer support. At six months, the intention-to-treat estimate of the program effect on a caregiver burden score (higher scores mean greater burden) was −2.4 points. In program centres, 60 percent of caregivers attended at least three peer-support sessions. A per-protocol analysis comparing attending caregivers in program centres with all caregivers in usual-service centres gave −5.1 points. As part of the program, staff at program centres also received training on recognizing caregiver burden. The complier average causal effect (CACE) with one-sided noncompliance equals the intention-to-treat effect divided by the proportion of compliers, and it assumes that assignment affects outcomes only through attendance. (a) Calculate the CACE. (b) Explain why the per-protocol estimate differs from it. (c) Explain how the staff training could threaten an assumption behind the CACE. (d) State which estimate should be primary for a decision about funding the program, and why.

Model answer

(a) The CACE is −2.4 / 0.60 = −4.0 points, the estimated effect among caregivers who would attend if offered the program.

(b) The per-protocol estimate of −5.1 compares attenders with all usual-service caregivers, including those who would not have attended had they been offered the program. Caregivers who attend are likely to differ from non-attenders, for example in their available time, social support or initial burden, and these differences also affect burden at six months. The per-protocol comparison is therefore no longer randomized and is likely to exaggerate the effect, while the CACE keeps the randomized comparison.

(c) The CACE assumes the exclusion restriction: assignment to a program centre affects burden only through attending sessions. If trained staff recognize and respond to burden among all caregivers, including those who never attend, part of the −2.4 difference comes from non-attenders. Dividing by 0.60 then attributes that change to attenders, and the CACE overstates their benefit.

(d) The intention-to-treat estimate of −2.4 should be primary, because the funding decision concerns offering the program to all caregivers at a centre, including those who will not attend, and it is the estimate that randomization fully protects. The CACE can be reported as a secondary result with its assumptions stated.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the simulated Cedar Valley trial, the intention-to-treat estimate was −0.420 points and 65.5 percent of Connector-arm participants engaged. Under one-sided noncompliance, what is the complier average causal effect, to two decimal places?

The CACE is the intention-to-treat effect divided by the proportion of compliers: −0.420 / 0.655 = −0.64. Multiplying instead of dividing gives −0.28 (option b), and −0.77 is the per-protocol estimate.

Question 2: Why did the per-protocol estimate (−0.772) overstate the effect among engaged participants in the simulated trial?

Engagement was selective: engaged older adults were less lonely at baseline and more mobile, and mobility was not recorded. Comparing them with the whole usual-care arm mixes the program effect with these differences. The per-protocol model kept all 24 clinics and included a random clinic intercept.

Question 3: Which approach to reporting a stepped-wedge evaluation of the Cedar Valley program is correct?

CONSORT 2025 is the current statement, and design-specific items come from extensions such as the stepped-wedge extension (Hemming et al., 2018), which builds on the cluster extension. SPIRIT (option c) governs trial protocols.

Question 4: A health authority has already offered a new program in every clinic in its region. Which feasibility question most clearly rules out randomizing access to the program itself?

Once everyone has been offered the program, there is no allocation left to randomize. The evaluator might still randomize an enhancement or encouragement, or use a quasi-experimental design from Lessons 7 and 8.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson has examined how random allocation supports causal claims about health services programs and how the design of a randomized evaluation must follow the way a program is delivered. Section 1 showed that randomization makes the arms exchangeable in expectation and that pragmatic trials, located with the PRECIS-2 wheel, answer the questions decision-makers ask about programs in usual care. Section 2 showed why programs such as the fictional Cedar Valley Connector program are usually randomized by clinic, how the intracluster correlation coefficient and the design effect determine the number of clusters needed, and why an analysis that ignores clustering reports more precision than the data contain.

Section 3 turned staggered rollouts into randomized designs, including waitlist and stepped-wedge trials, and introduced factorial, adaptive, SMART and non-inferiority designs together with the ethics of randomizing access. Section 4 distinguished the effect of offering a program from the effect of taking part in it, introduced the current CONSORT statement and its extensions, and set out the questions an evaluator asks before recommending randomization. The worked example in Section 4.5 applies those questions to the Cedar Valley Connector program.

Key Takeaways from this lesson

  • Random allocation makes the arms exchangeable in expectation, which protects a comparison from measured and unmeasured confounders at the moment of allocation.
  • Allocation concealment protects enrolment from knowledge of upcoming assignments and is achievable even when blinding is not.
  • Pragmatic trials answer whether offering a program in usual practice improves outcomes, and PRECIS-2 helps a team check that each design choice fits that purpose.
  • Programs delivered through clinics, schools or communities are usually randomized by cluster to avoid contamination and to match how the program is delivered.
  • The design effect, 1 + (m − 1) × ICC, shows how much clustering inflates the variance, and the number of clusters matters more than the number of participants per cluster.
  • In the simulated Cedar Valley trial, a naive model and a mixed model gave similar effect estimates, but the naive standard error (0.084) was much smaller than the mixed-model standard error (0.136).
  • Stepped-wedge designs randomize the order of a phased rollout, and their analysis must include period effects because exposure is confounded with calendar time.
  • Factorial designs, SMARTs, adaptive designs and non-inferiority trials each answer a distinct question, and a non-inferiority margin must be smaller than the established effect of the standard option.
  • Intention-to-treat estimates the effect of offering a program, while the complier average causal effect estimates the effect among those who take part under stated assumptions, and per-protocol comparisons are prone to selection bias.
  • The feasibility of randomization depends most often on timing, the number of units available and the acceptability of the allocation to partners and communities.

Core Concepts Reviewed

Section 1: potential outcomes, exchangeability, random allocation versus random sampling, allocation concealment, randomization schemes, explanatory and pragmatic trials, and the PRECIS-2 wheel.

Section 2: cluster randomization and contamination, the intracluster correlation coefficient, the design effect and effective sample size, sample size for cluster trials, mixed models, generalized estimating equations and the CONSORT extension for cluster trials.

Section 3: randomized staggered rollouts, stepped-wedge designs and time trends, waitlist controls, factorial designs, adaptive designs and SMARTs, non-inferiority margins, and the ethics of randomizing access.

Section 4: intention-to-treat, per-protocol and as-treated analyses, principal strata and the complier average causal effect, CONSORT 2025 and its extensions, feasibility questions and the randomization feasibility assessment.

The final reflection asks you to bring the lesson together by recommending a randomized design for a program with a phased rollout.

Reflection

You are the evaluator for a new provincial program that offers home-based social prescribing to adults aged 70 and older after discharge from hospital. Because staff must be hired and trained, the program will start in 30 hospitals over 18 months, with 10 hospitals starting every 6 months. The hospital discharge teams make the referrals. Each hospital discharges about 50 eligible older adults in each 6-month period. Pilot data suggest a standard deviation of 1.5 points on the three-item UCLA Loneliness Scale (scored 3 to 9) and an intracluster correlation coefficient of about 0.03 among patients of the same hospital. The ministry wants to know whether offering the program reduces loneliness three months after discharge, and it will decide on permanent funding once the rollout is complete. The order in which hospitals start has not yet been set. In 200 to 300 words, recommend a design. State whether randomization is feasible and why, the unit and scheme of randomization, how the analysis will address clustering and time, which estimate will be primary and which secondary, and which reporting guidelines apply.

Model answer

Randomization is feasible. The rollout must be staggered because of hiring, the order of hospitals has not been set, informed people are uncertain whether the program reduces loneliness, and the ministry needs an estimate of the effect of offering it. Randomizing the order changes who waits and does not change how many wait.

The unit of randomization should be the hospital, because discharge teams make the referrals and could not offer the program to some patients and withhold it from others without contamination. I would use a stepped-wedge design with a six-month baseline period and three steps, randomizing the 30 hospitals into three sequences of 10, stratified by hospital size or health authority. An independent statistician would generate the allocation after all hospitals agree to take part. With about 50 patients per hospital per period and an ICC of 0.03, the design effect within a period is about 1 + 49 × 0.03 = 2.47, so power should be calculated with a stepped-wedge method or by simulation.

The analysis would be a linear mixed model with a random hospital effect, fixed period effects to absorb secular trends, the program indicator and baseline loneliness, using small-sample degrees of freedom, with a sensitivity analysis for an effect that changes with time since a hospital started. The intention-to-treat estimate would be primary, and a complier average causal effect for patients who engage would be secondary. Reporting would follow CONSORT 2025 and the stepped-wedge extension.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Randomized Designs for Health Services Interventions (15 Questions)

Question 1: Which pairing of a design feature and its purpose is correct?

Allocation concealment protects enrolment from knowledge of upcoming assignments. Random sampling supports generalizability, blinding concerns knowledge after allocation, and stratification improves balance without changing how clustering must be analyzed.

Question 2: A cluster trial was planned with 40 participants per cluster and an intracluster correlation coefficient of 0.06. If the investigators instead enrolled 80 participants per cluster, what would happen to the design effect?

The design effect is 1 + (m − 1) × ICC, so it rises from 1 + 39 × 0.06 = 3.34 to 1 + 79 × 0.06 = 5.74. It does not double (option d), because the 1 in the formula does not scale with m.

Question 3: An evaluator wants to know whether offering the Cedar Valley program in usual care reduces loneliness. Which combination fits that purpose?

The question concerns offering the program in practice, which calls for pragmatic choices: broad eligibility, usual settings and an intention-to-treat analysis. The other options combine explanatory features that answer whether the program can work under favourable conditions.

Question 4: In the Section 2 worked example, the squared ratio of the mixed-model and naive standard errors (2.588) was close to which quantity?

The squared ratio of standard errors approximates the design effect. With the residual ICC of 0.040 and the observed cluster sizes, the design effect is 2.605, close to 2.588. This shows that the naive model fails by ignoring the design effect.

Question 5: Why might a stepped-wedge design be chosen over a parallel cluster trial for a program that every site will eventually receive?

Stepped-wedge designs give every cluster the program within the trial, which suits phased rollouts and improves acceptability. Their analysis must model time trends (option a), and whether they need fewer clusters depends on the ICC, cluster size and number of steps (option c).

Question 6: Which situation would violate the exclusion restriction needed for a complier average causal effect in a cluster trial?

The exclusion restriction requires assignment to affect outcomes only through participation. If clinicians change their care for every older patient, never-takers in intervention clinics are also affected, and the CACE overstates the complier effect. Non-attendance (option a) is what the CACE is designed to handle.

Question 7: With three clusters per arm, what is the smallest two-sided p-value that a randomization test can produce?

Six clusters can be split into two arms of three in 20 ways. The observed allocation can be at most the most extreme of the 20 in one direction, and a two-sided test counts both extremes, so the smallest p-value is 2 / 20 = 0.10.

Question 8: Which analysis of a cluster trial with 24 clusters and 960 participants is least defensible?

Ordinary regression ignores the correlation within clusters and understates the standard error of a cluster-level effect, as Cornfield (1978) warned. The other three options are valid analyses that respect clustering, with small-sample corrections where needed.

Question 9: What distinguishes a sequential multiple assignment randomized trial from a standard two-arm trial?

A SMART re-randomizes participants, often according to their response, so that sequences of decision rules can be compared. Option b describes a stepped-wedge design, option c an adaptive design and option d a factorial design.

Question 10: In a non-inferiority trial, why is an intention-to-treat analysis on its own insufficient?

When participants in either arm fail to engage, the arms converge, which pushes the result toward non-inferiority. For this reason both intention-to-treat and per-protocol results are reported, and non-inferiority is more credible when both support it.

Question 11: The fictional Cedar Valley Health Authority wants the highest-need clinics to start the program first but also wants a randomized comparison. Which approach meets both aims?

Randomizing within strata, with unequal probabilities by stratum, lets need shape the rollout while keeping a randomized comparison within each stratum. Option a makes need determine allocation, so the comparison is confounded, and option b allows selection by managers.

Question 12: Which statement about the intracluster correlation coefficients in the Section 2 worked example is correct?

Arm varies between clinics, so in the empty model the program effect appears as between-clinic variance. Including arm removes it and lowers the ICC from 0.0590 to 0.0422. ICCs can be estimated with unequal cluster sizes (option d).

Question 13: A program has 600 eligible applicants for 300 places. Which randomized design is the most natural fit?

Oversubscription creates a natural point of control: a lottery allocates scarce places fairly and creates a randomized waitlist comparison. A comparison of applicants with non-applicants (option c) is not randomized, and self-selected components (option d) are not randomized either.

Question 14: Which pairing of a design and its reporting guideline is correct?

The Adaptive designs CONSORT Extension (ACE) covers adaptive designs. Campbell et al. (2012) wrote the cluster extension, Piaggio et al. (2012) the non-inferiority extension, Zwarenstein et al. (2008) the pragmatic trials extension and Hemming et al. (2018) the stepped-wedge extension.

Question 15: An evaluator is assessing a program that will be delivered by school nurses in 10 schools and has not yet started. Which specification best follows the lesson's guidance?

Nurses deliver the program at the school level, so schools are the natural unit, and with only 10 schools the power and chance imbalance must be checked. Randomizing students invites contamination (option a), and adding students per school cannot overcome the floor set by too few clusters (option b).
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Randomized Designs for Health Services Interventions.

You can now explain why randomization supports causal claims, locate a trial on the explanatory-pragmatic continuum, calculate the sample size for a cluster randomized trial, interpret a mixed-model analysis of one, choose among stepped-wedge, waitlist, factorial, adaptive and non-inferiority designs, distinguish intention-to-treat from complier effects, and judge whether randomization is feasible for a program.

Lesson 7 turns to quasi-experimental designs for programs that cannot be randomized. It covers threats to validity, comparison group designs and difference-in-differences, propensity scores and synthetic control, and it returns to the Cedar Valley Connector program's two waves to ask what can be learned when the order of rollout was not randomized.

Continue to Lesson 7 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

This glossary defines the terms, tools and people introduced in Lesson 6, in the order of the lesson's main themes.

Core Concepts
Potential outcomes The outcomes a unit would have under each condition being compared, such as Y(1) with the program and Y(0) without it. Only one is observed for each unit, so causal effects are estimated as averages across groups.
Exchangeability The condition that the groups being compared would have had the same average outcome had they received the same condition. Randomization makes the arms exchangeable in expectation.
Allocation concealment Procedures that stop the people enrolling participants or clusters from knowing or predicting the next allocation before enrolment is complete.
Explanatory trial A trial that asks whether an intervention can work under ideal conditions, typically with selected participants, expert delivery and close monitoring.
Pragmatic trial A trial that asks whether an intervention works when delivered in usual practice, with broad eligibility, usual staff and settings, and an intention-to-treat analysis.
Cluster randomized trial A trial in which intact groups, such as clinics, schools or communities, are randomized and outcomes are measured on individuals within them.
Contamination The spread of an intervention, or of its information and practices, to participants in the comparison arm, which shrinks the observed difference between arms.
Intracluster correlation coefficient The proportion of the total variance in an outcome that lies between clusters, equal to the between-cluster variance divided by the sum of the between- and within-cluster variances.
Design effect The factor (DEFF) by which clustering inflates the variance of an estimate relative to individual randomization, DEFF = 1 + (m − 1) × ICC for clusters of equal size m, where ICC is the intracluster correlation coefficient, with an adjustment for unequal cluster sizes.
Effective sample size The number of independent observations that would give the same precision as a clustered sample, equal to the total sample size divided by the design effect.
Identification and recruitment bias Bias that arises when participants are identified or recruited after clusters know their allocation, so that the arms enrol different kinds of people.
Stepped-wedge design A cluster design in which all clusters start in the control condition and cross over to the intervention at randomized steps until every cluster has received it.
Secular trend A change in the outcome over calendar time that is unrelated to the intervention, such as a seasonal pattern or a new provincial policy.
Waitlist control A comparison group that receives the program after a delay, so that the arms can be compared until the delayed group starts.
Factorial design A design that randomizes units to two or more interventions at once so that every combination is represented, allowing several main effects to be estimated from one sample.
Non-inferiority margin The largest loss of effect, specified in advance, that decision-makers would accept in exchange for the advantages of a new option.
Clinical equipoise Honest disagreement among informed experts about the comparative merits of the options being tested, proposed by Freedman (1987) as the ethical basis for randomization.
Intention-to-treat An analysis that includes every randomized participant in the arm to which they were allocated, which estimates the effect of offering the intervention.
Per-protocol analysis An analysis restricted to participants who followed the protocol, which no longer compares randomized groups and is prone to selection bias.
Complier average causal effect The average effect among participants who would take up the intervention if offered it and not otherwise, estimated under one-sided noncompliance as the intention-to-treat effect divided by the proportion of compliers.
Frameworks & Tools
PRECIS-2 A tool with nine domains, each scored from 1 (very explanatory) to 5 (very pragmatic) and plotted on a wheel, used to match trial design choices to the trial's purpose (Loudon et al., 2015).
Stratified randomization Randomization carried out separately within strata of an important prognostic factor so that the arms are balanced on that factor.
Constrained randomization A method for cluster trials that generates many possible allocations, keeps those balanced on chosen cluster characteristics and selects one of them at random (Moulton, 2004).
Linear mixed model A regression model with fixed effects and random effects, such as a random intercept for each clinic, that separates between-cluster from within-cluster variation.
Generalized estimating equations A method that estimates population-averaged effects from clustered data using a working correlation structure and sandwich standard errors (Liang & Zeger, 1986).
Sequential multiple assignment randomized trial A trial in which participants can be randomized at several decision points so that adaptive interventions, which are sequences of decision rules, can be built and compared (Murphy, 2005).
Adaptive design A trial design that allows pre-specified changes, such as early stopping or dropping arms, in response to accumulating data, reported with the ACE guideline (Dimairo et al., 2020).
CONSORT 2025 and its extensions The current CONSORT statement for reporting randomized trials (Hopewell et al., 2025), used with design-specific extensions such as those for cluster trials (Campbell et al., 2012) and stepped-wedge trials (Hemming et al., 2018).
Ottawa Statement Recommendations on the ethical design and conduct of cluster randomized trials, covering who counts as a research participant, consent and its waiver, and the role of gatekeepers (Weijer et al., 2012).
Key People
Ronald A. Fisher British statistician and geneticist who set out the principles of randomization, replication and blocking in experimental design, notably in The Design of Experiments (1935).
Donald B. Rubin American statistician whose 1974 work formalized the potential outcomes approach to causal inference, and who developed the complier average causal effect with Angrist and Imbens (1996).
Joshua Angrist and Guido Imbens Economists who, with Rubin, showed how random assignment can serve as an instrument for estimating effects among compliers; they shared the 2021 Nobel Memorial Prize in Economic Sciences for contributions to the analysis of causal relationships.
Daniel Schwartz and Joseph Lellouch French statisticians who, in 1967, distinguished explanatory from pragmatic attitudes in therapeutic trials.
Allan Donner Canadian biostatistician at Western University whose work, including a 1981 paper with Birkett and Buck and a 2000 text with Klar, established methods for the sample size and analysis of cluster randomization trials.
Merrick Zwarenstein Health services researcher based in Canada and a co-developer of the PRECIS tool and of the CONSORT extension for pragmatic trials.
Karla Hemming Biostatistician known for methods for stepped-wedge and cluster randomized trials, and lead author of the 2018 CONSORT extension for stepped-wedge trials.
Susan A. Murphy Statistician who proposed the sequential multiple assignment randomized trial for developing adaptive interventions.
Charles Weijer Canadian bioethicist at Western University who led the development of the Ottawa Statement on the ethical design and conduct of cluster randomized trials.
Benjamin Freedman Bioethicist at McGill University who introduced the concept of clinical equipoise in 1987.
No matching entries. Try a different search term.