# Lesson 6: Randomized Designs for Health Services Interventions

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we start the design block of the course, which runs across Lessons six, seven and eight. This lesson is about randomized designs for health services interventions.

**Sarah:** So far the course has been about planning: describing a program, building a logic model, writing evaluation questions and filling in an evaluation matrix. What changes now?

**Kiffer:** Now we take the hardest question in the matrix, which is usually whether the program caused the change we see, and ask what kind of study can answer it. This lesson covers studies that answer it by randomizing. The next two lessons cover what to do when randomizing is not possible.

**Sarah:** And we're still working with the Cedar Valley Connector program.

**Kiffer:** We are. For anyone joining now, it is a fictional community connector program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged sixty-five and older who screen as lonely or isolated to a connector, who meets them up to six times over twelve weeks and links them to groups, volunteering and transportation help. The program starts in twelve of the region's twenty-four clinics, and the other twelve join a year later.

**Sarah:** And the main outcome is loneliness.

**Kiffer:** Yes. It is measured with the three-item University of California, Los Angeles Loneliness Scale, which is scored from three to nine, with higher scores meaning more loneliness. All the Cedar Valley numbers in this episode are illustrative.

**Sarah:** Let's start where the lesson starts. Why does randomization matter so much? People compare programs without it all the time.

**Kiffer:** They do, and sometimes that is the best available option. The problem is that a comparison between people who got a program and people who did not tells you the program's effect only if the two groups would have had the same outcomes without it. The lesson uses the potential outcomes framework to say this precisely. Every older adult has a loneliness score they would have with the program and a score they would have without it. We only ever see one of those two.

**Sarah:** That's what Holland called the fundamental problem of causal inference.

**Kiffer:** That's right. So we compare groups and estimate an average effect. The comparison is fair only when the groups are exchangeable, meaning they would have had the same average outcome had neither received the program.

**Sarah:** Give me a Cedar Valley example of where that breaks down.

**Kiffer:** Suppose the health authority picked the first-wave clinics because their managers volunteered, they had a spare room and their records could produce a list of lonely patients. Each of those features could also be related to loneliness. A clinic with an enthusiastic manager may already run group activities. A clinic with space may be downtown, near better transit. So if wave-one clinics end up with lower loneliness, you cannot tell how much of that is the program and how much is the kind of clinic that volunteered.

**Sarah:** Couldn't we just adjust for those differences in a regression?

**Kiffer:** We could for the ones we measured. That is what Lesson seven does with matching and weighting. But some of the differences, like a manager's enthusiasm or a clinician's sense of which patients will follow through, are never recorded. You cannot adjust for something you do not have.

**Sarah:** And randomization handles the ones we didn't measure.

**Kiffer:** It does, in expectation. When a random number generator decides which clinics go first, enthusiasm and spare rooms cannot influence the allocation. Across all the allocations that could have happened, the arms have the same distribution of every characteristic, measured or not. Fisher also pointed out that because we know the allocation mechanism, it gives us a basis for inference.

**Sarah:** You keep saying "in expectation." What's the catch?

**Kiffer:** There are three. Any one allocation can be unbalanced by chance, and that risk is larger when you randomize twenty-four clinics than when you randomize nine hundred people. Randomization only protects the comparison at the moment of allocation, so differential dropout or biased recruitment afterwards can undo it. And it tells you what happened in the trial's setting, which may or may not generalize.

**Sarah:** The reading is careful to separate random allocation from random sampling. Why do students mix those up?

**Kiffer:** They mix them up because both use the word random. Random sampling picks people from a population so the sample represents it, which supports generalizability. Random allocation assigns people who are already in the study to conditions, which supports the causal comparison. Most trials do the second without the first.

**Sarah:** Then there's allocation concealment, which sounds like blinding but apparently isn't.

**Kiffer:** Concealment means the people enrolling participants cannot know the next allocation before someone is enrolled. Blinding is about who knows the allocation afterwards. You cannot blind a connector program, because people know whether they are meeting a connector. But you can conceal allocation, for instance by having an independent statistician release each clinic's allocation only after the clinic has signed on. Schulz and colleagues found that trials with poor concealment tended to report larger effects.

**Sarah:** And with only twenty-four clinics, how would you actually generate the allocation?

**Kiffer:** Simple randomization could easily put most of the rural clinics in one wave. So I would stratify by rurality, or use constrained randomization, where you generate many possible allocations, keep the ones that are balanced on a few clinic characteristics, and pick one of those at random.

**Sarah:** Let's move to explanatory and pragmatic trials. That distinction goes back to the nineteen sixties.

**Kiffer:** It goes back to Schwartz and Lellouch in nineteen sixty-seven. An explanatory trial asks whether an intervention can work under ideal conditions. A pragmatic trial asks whether it works when delivered in ordinary practice, so that someone can choose between options.

**Sarah:** Isn't an explanatory trial just more rigorous, though? Tighter control, better adherence.

**Kiffer:** It is more controlled, and that is appropriate for its question. The issue is whether the design fits the decision. The Cedar Valley steering committee has to decide whether to fund the program across the region. An explanatory trial run by a university team with its own trained connectors, cognitively intact volunteers and fortnightly reminder calls would tell them whether the connector model can work under favourable conditions. It would not tell them what happens when their own staff serve everyone their clinicians refer.

**Sarah:** And that's where the Pragmatic Explanatory Continuum Indicator Summary comes in, the tool known as PRECIS two.

**Kiffer:** Yes. Loudon and colleagues published PRECIS two in twenty fifteen, building on the original version by Thorpe and colleagues. It has nine domains, such as eligibility, setting, flexibility in adherence and primary analysis. Each is scored from one, very explanatory, to five, very pragmatic, and the scores are plotted on a wheel.

**Sarah:** Where does the pragmatic Cedar Valley trial land?

**Kiffer:** It lands near the rim on most domains. Setting scores five because the trial runs in all the region's clinics. Primary analysis scores five because everyone is analyzed in their clinic's arm. Follow-up scores three, because research staff add telephone surveys that usual care would not include.

**Sarah:** Would you use the wheel as a quality score?

**Kiffer:** No, and the reading is explicit about that. The wheel shows whether each design choice fits the purpose. A team that scores the domains together often discovers it disagreed about what the trial was for.

**Sarah:** The section ends by pointing toward clinics. Why?

**Kiffer:** The reason is that the Cedar Valley program works through clinics. The connector is on site and clinicians do the referring. If you randomized individual patients within a clinic, clinicians would carry what they learned into conversations with usual-care patients, and patients would talk in the waiting room. That leakage is contamination, and it shrinks the difference you are trying to measure.

**Sarah:** So we randomize whole clinics. That's Section two, cluster randomized trials.

**Kiffer:** That's right. In a cluster randomized trial, intact groups are randomized and outcomes are measured on the people within them. Besides contamination, the reasons include that the program acts on the whole clinic, that health authorities hire and budget by site, and that managers accept a randomized start date more readily than having half their patients turned away.

**Sarah:** What does it cost?

**Kiffer:** It costs information. Older adults at the same clinic share a neighbourhood, transit routes and clinicians, so their loneliness scores are a bit alike. Each extra person from the same clinic tells you less than a person from a new clinic would. The intracluster correlation coefficient measures that likeness. It is the share of the total variance in the outcome that sits between clusters.

**Sarah:** And in the simulated Cedar Valley data it's about six percent.

**Kiffer:** The empty mixed model puts it at zero point zero five nine, so about six percent of the variation in six-month loneliness lies between clinics.

**Sarah:** Six percent sounds small enough to ignore.

**Kiffer:** That is the most common mistake in this area. The correlation gets multiplied by the cluster size. The design effect is one plus the cluster size minus one, times the correlation. With forty older adults per clinic and a correlation of zero point zero six, that is one plus thirty-nine times zero point zero six, which is three point three four.

**Sarah:** What does that mean, concretely?

**Kiffer:** The variance of the program effect is more than three times what it would be if those same people had been randomized individually. Nine hundred and sixty clustered older adults carry about as much information as two hundred and eighty-seven independent ones.

**Sarah:** That is a large loss of information.

**Kiffer:** It is, and it gets worse as clusters get bigger. With a correlation of only zero point zero one and two hundred people per cluster, the design effect is nearly three.

**Sarah:** Walk me through the sample size for Cedar Valley.

**Kiffer:** Suppose the steering committee decides that half a point on the loneliness scale is the smallest difference worth detecting, and the standard deviation is about one and a half points. For an individually randomized trial with eighty percent power and a five percent significance level, you need about one hundred and forty-two people per arm, or one hundred and forty-one point three before rounding up. Multiply that unrounded figure by the design effect of three point three four and you need four hundred and seventy-two per arm. At forty per clinic, that is twelve clinics per arm.

**Sarah:** That is exactly what Cedar Valley has.

**Kiffer:** Under those assumptions, the region's twenty-four clinics are just enough. And there is a floor that surprises people. However many people each clinic enrols, you need at least the individual sample size times the correlation in clinics per arm. Here that is one hundred and forty-one point three times zero point zero six, about eight and a half, so nine clinics per arm.

**Sarah:** So you can't recruit your way out of having too few clinics.

**Kiffer:** That is correct. The worked example shows it with a table. With a correlation of zero point zero eight, doubling clinic size from forty to eighty only cuts the clinics needed from fifteen to thirteen.

**Sarah:** Let's talk about the worked example's main comparison. A naive model and a mixed model.

**Kiffer:** Both regress the six-month score on the program indicator and baseline loneliness. The naive model, an ordinary linear regression, treats all nine hundred and sixty older adults as independent. The mixed model adds a random intercept for each clinic. Both estimate that people in Connector clinics scored about zero point four two points lower.

**Sarah:** If the estimates agree, why should anyone care which model they use?

**Kiffer:** They matter because the models disagree about precision. The naive standard error is zero point zero eight four. The mixed-model standard error is zero point one three six, about one point six times larger. The naive confidence interval runs from about minus zero point five nine to minus zero point two six, and the mixed-model interval from about minus zero point seven zero to minus zero point one four.

**Sarah:** Both exclude zero, though.

**Kiffer:** Here they do. With a smaller true effect, the naive analysis could declare a significant result that the correct analysis would not support. Cornfield made this point in nineteen seventy-eight, describing randomization by cluster followed by an individual-level analysis as a form of self-deception.

**Sarah:** Explain in plain terms why the naive standard error is too small.

**Kiffer:** Everyone in a clinic has the same value of the program indicator, and they share a clinic-level deviation. So the information about the program comes from comparing twelve clinics with twelve clinics. The naive model behaves as if it had nine hundred and sixty independent pieces of evidence. And there is a nice check in the worked example. The squared ratio of the two standard errors is two point five nine, almost exactly the design effect implied by the residual correlation in the mixed model, which is two point six one.

**Sarah:** The worked example also shows sandwich standard errors.

**Kiffer:** It shows them as an option. You keep the ordinary regression estimate and replace its standard error with one that allows correlation within clinics. It comes out at zero point one three one, close to the mixed model. The trap is that by default the test still uses nine hundred and fifty-seven degrees of freedom, as if there were hundreds of independent units. With twenty-four clusters you should use about twenty-two.

**Sarah:** How do you choose between a mixed model and generalized estimating equations?

**Kiffer:** Mixed models estimate cluster-specific effects, and generalized estimating equations estimate population-averaged effects. For a continuous outcome like ours they coincide. For a binary outcome, such as any emergency department visit, the cluster-specific odds ratio is further from one, so you must be clear which you are reporting. With few clusters, a common rule of thumb being fewer than about forty, both need small-sample corrections, and a simple comparison of clinic means is also a perfectly valid option.

**Sarah:** What about reporting?

**Kiffer:** The standard is the cluster extension of the Consolidated Standards of Reporting Trials, known as CONSORT, from Campbell and colleagues in twenty twelve. It asks why you randomized clusters, how clustering entered your sample size and analysis, who consented at the cluster and individual levels, the flow of both clinics and people, and the intracluster correlation you observed.

**Sarah:** Section three takes the two-wave rollout seriously. What's the idea?

**Kiffer:** Many programs cannot start everywhere at once. When the order is not fixed by need, randomizing it costs almost nothing. For Cedar Valley, if the health authority randomizes which twelve clinics go first, the wave-two clinics serve as a waitlist control during year one. That is the trial we analyzed in the worked example.

**Sarah:** What does the stepped-wedge version look like?

**Kiffer:** Spread the rollout over the first year instead, in four steps of six clinics, three months apart, with the order randomized. Every clinic starts in usual care, every clinic eventually gets the program, and outcomes are measured everywhere in every period. Hussey and Hughes set out the standard design and analysis in two thousand seven, and Hemming and colleagues have written extensively on it.

**Sarah:** That sounds like the best of both worlds. Why not always use it?

**Kiffer:** The reason is that exposure is tied to calendar time. In the first period no clinics have the program and in the last all of them do. If loneliness changes over the year because of winter, a provincial initiative or a transit strike, that change will look like a program effect unless you model it. So the analysis includes a fixed effect for each period.

**Sarah:** Does that solve it?

**Kiffer:** It handles trends that are common to all clinics. The standard model also assumes the program effect is immediate and constant. A connector program that takes months to build caseloads and community partnerships may not behave that way, and assuming a constant effect when it actually grows can bias the estimate. Clinics may also start late if hiring is delayed. So a stepped-wedge trial needs careful planning, and it is not automatically more efficient than a parallel design.

**Sarah:** Lesson seven deals with staggered rollouts too.

**Kiffer:** It does, for the non-randomized case. Lesson seven shows how two-way fixed effects models can mislead when adoption is staggered. The stepped-wedge design has the same staggered structure, plus randomization of the order.

**Sarah:** What about waitlist controls on their own?

**Kiffer:** They are common when a program has more applicants than places, and they are acceptable because everyone is promised the program. Their limits are that the comparison only lasts as long as the delay, and people who know help is coming may postpone seeking other support, which can flatter the program.

**Sarah:** The section then moves through factorial designs. Give me the Cedar Valley version.

**Kiffer:** Suppose the health authority also wants to test a transportation voucher. In a two by two factorial design, each person is randomized to the connector or not, and separately to the voucher or not. If the two do not interact, every participant informs both comparisons, so one sample answers two questions.

**Sarah:** What happens if they do interact?

**Kiffer:** Then the main effects are harder to interpret, and estimating the interaction properly takes about four times the sample size needed for a main effect of the same size. Most factorial trials are powered for main effects only.

**Sarah:** Adaptive designs and sequential multiple assignment trials sound similar. Are they?

**Kiffer:** They answer different questions. An adaptive design allows pre-planned changes as data accumulate, like stopping early or dropping an arm, with the error rates controlled. A sequential multiple assignment randomized trial, or SMART, randomizes people more than once so you can build adaptive interventions. In a Cedar Valley version, people would first be randomized to in-person or telephone meetings, and those not engaged by week six would be randomized again to a peer volunteer or transportation help.

**Sarah:** Non-inferiority trials seem to run backwards.

**Kiffer:** They ask whether a cheaper or more scalable option is not worse than the standard by more than a set margin. Telephone meetings would let one connector cover more clinics, so the question is whether they are acceptably close to in-person meetings.

**Sarah:** How do you pick the margin?

**Kiffer:** It has to be smaller than the standard option's established benefit. In our simulated trial, in-person meetings lowered loneliness by about zero point four two points. A margin of half a point would let you declare telephone delivery non-inferior even if it were worse than usual care. A margin of zero point two preserves about half the benefit.

**Sarah:** And intention-to-treat isn't the safe choice there.

**Kiffer:** That's right. When people in both arms fail to engage, the arms look more alike, which pushes toward a finding of non-inferiority. So both intention-to-treat and per-protocol results are reported.

**Sarah:** Let's talk ethics. Some people find it uncomfortable to randomize access to something that might help.

**Kiffer:** That discomfort deserves a real answer. Freedman's idea of clinical equipoise locates the justification in genuine disagreement among informed experts about whether an option is better. Scarcity adds to it. Cedar Valley can only fund twelve clinics in year one, so half will wait regardless. Randomizing the order changes who waits and leaves the number who wait unchanged.

**Sarah:** Where does the Ottawa Statement fit?

**Kiffer:** It deals with cluster trials specifically, and it was led by Charles Weijer and colleagues in twenty twelve. It asks you to identify who the research participants are, which may include clinicians whose practice the program changes. It sets out when consent can be waived or altered. And it says gatekeepers like clinic managers can permit their clinic to take part but cannot consent on behalf of individual patients.

**Sarah:** Where does the Indigenous health partnership come in?

**Kiffer:** It comes in at the start. Decisions about randomizing clinics that serve First Nations communities, and about governing data from First Nations participants, belong with that partnership, under Chapter nine of the Tri-Council Policy Statement and the principles we covered in Lesson four.

**Sarah:** The reading has a councillor who objects that a lottery is unfair to high-need clinics.

**Kiffer:** And the evaluator's response is a good one. Group the clinics by need and randomize within each group, with a higher chance of starting first in the highest-need group. Need shapes the rollout, and the comparison stays randomized within each group. If the committee insists that need alone decides, the evaluator plans a quasi-experimental design and is honest about what that costs.

**Sarah:** Section four starts from the fact that randomization decides what people are offered, and people then decide what to do with the offer.

**Kiffer:** That is the central issue. In the simulated trial, about a third of older adults at Connector clinics did not engage, meaning they attended fewer than two meetings. An intention-to-treat analysis keeps everyone in their randomized arm. It estimates the effect of offering the program, which is what the funding decision is about.

**Sarah:** But managers want to know what happens to the people who actually take part.

**Kiffer:** They do, and the tempting answer is a per-protocol analysis that drops the non-engagers. The trouble is that engagers differ from non-engagers in ways that matter. In our data, people who engaged were less lonely at baseline, about five point nine compared with six point six.

**Sarah:** So what's the alternative?

**Kiffer:** The alternative is the complier average causal effect, from Angrist, Imbens and Rubin in nineteen ninety-six. When usual-care clinics have no access to the program, it equals the intention-to-treat effect divided by the proportion who engaged. Here that is minus zero point four two divided by zero point six five five, which gives about minus zero point six four.

**Sarah:** What was the per-protocol estimate?

**Kiffer:** It is minus zero point seven seven. Because the data are simulated, we know the program was built to lower scores by about zero point six two among people who engaged. The complier estimate lands close to that. The per-protocol estimate overshoots, because engagers were also more mobile, which the dataset never recorded, and mobility lowers loneliness on its own.

**Sarah:** What assumptions does the complier estimate need?

**Kiffer:** It needs random assignment, no defiers, no interference between participants, and the exclusion restriction, which says that being in a Connector clinic affects outcomes only through engagement. That last one can fail in a cluster trial. If clinicians at Connector clinics start asking every older patient about loneliness, even non-engagers are affected, and the complier estimate will be too large.

**Sarah:** On reporting, the lesson mentions CONSORT twenty twenty-five.

**Kiffer:** That is the current Consolidated Standards of Reporting Trials statement, from Hopewell and colleagues, and it updates the twenty ten version. The design-specific extensions, for cluster trials, stepped-wedge trials, pragmatic trials and others, were written for the twenty ten statement and are used alongside the current one. Protocols follow the Standard Protocol Items, Recommendations for Interventional Trials guideline, known as SPIRIT, and trials should be registered before the first participant enrols.

**Sarah:** Then the section turns to feasibility. How does an evaluator decide whether to recommend randomization at all?

**Kiffer:** The lesson gives ten questions. They start with whether the decision needs a causal estimate and whether there is genuine uncertainty. The question that most often decides things is timing: whether allocation can still be influenced, or whether the program has already started. Then come natural points of control, the level of delivery, the number of units, acceptability, equal measurement, ethics, and finally cost together with the fallback design.

**Sarah:** The number of units question has a striking example.

**Kiffer:** With three clusters per arm there are only twenty ways to split six clusters into two groups of three. So a randomization test can never give a two-sided p-value below zero point one, however large the effect. That is a hard limit, and it is worth checking before anyone promises a trial.

**Sarah:** So what does an evaluator write down after working through those questions?

**Kiffer:** A feasibility assessment, usually a short memo for the steering committee. It should work through the ten questions and say whether a randomized design is feasible. If it is, the memo should specify the unit of randomization, the allocation ratio, the scheme, how allocation will be concealed, the number of units and where the planning values came from, and the primary analysis.

**Sarah:** What if it isn't feasible?

**Kiffer:** Then the memo should name the barrier and the alternative design, which is where Lessons seven and eight come in. A clear statement that randomization is not possible, with good reasons, is still a useful result for the committee.

**Sarah:** The reading has a Cedar Valley worked example. What does the evaluator recommend?

**Kiffer:** The evaluator recommends randomizing the twenty-four clinics one to one between the two waves, stratified by rurality, with an independent statistician generating the allocation after every clinic has signed on. The loneliness screen is introduced in all clinics before randomization, so recruitment cannot depend on allocation. The primary analysis is an intention-to-treat mixed model, with a complier analysis as a secondary result.

**Sarah:** Any last advice for someone writing an assessment like that?

**Kiffer:** Start from the decision the program faces and from its delivery structure, and let those drive the unit of randomization. And show the numbers. If you claim a cluster trial is feasible, show how many clusters you need and where your intracluster correlation came from.

**Sarah:** Next week is Lesson seven, quasi-experimental designs.

**Kiffer:** Yes. We will look at threats to validity, difference-in-differences, propensity scores and synthetic control, and we will return to Cedar Valley's two waves to ask what we can learn when the order was not randomized.

**Sarah:** Thanks, Kiffer.

**Kiffer:** Thanks, Sarah, and thanks to everyone listening. I'll see you next week.
