# Lesson 7: Quasi-Experimental Designs I: Comparison Groups, Difference-in-Differences and Matching

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week is Lesson seven, the first of two lessons on quasi-experimental designs. We cover comparison groups, difference-in-differences and matching.

**Sarah:** Last week was all about randomized trials. So this week we're in the world where randomization didn't happen.

**Kiffer:** Last week we imagined that the Cedar Valley Health Authority had randomized its clinics to two waves. This week we look at what actually happened in our story.

**Sarah:** Remind listeners what Cedar Valley is.

**Kiffer:** The Cedar Valley Connector program is a fictional community connector program, run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged sixty-five and older who screen as lonely or isolated to a community connector. A connector meets them up to six times over twelve weeks and links them to community supports. It launched in twelve of the region's twenty-four clinics, and the other twelve are due to join a year later. All the numbers we use are illustrative.

**Sarah:** How were the first twelve clinics chosen?

**Kiffer:** They were chosen for readiness. The managers volunteered, the clinics had a room the connector could use, and their electronic medical records could produce a list of older patients who had screened as lonely.

**Sarah:** Those sound like perfectly sensible reasons.

**Kiffer:** They are sensible reasons to start somewhere. The trouble for an evaluator is that each of them could also be related to the outcomes. A clinic with an engaged manager might already run more group activities. So when we compare first-wave and second-wave clinics, we have to worry that they would have differed anyway.

**Sarah:** Let's start with Section one, then. The lesson calls it the counterfactual problem. We covered that last week, didn't we?

**Kiffer:** We did, and everything this week depends on it. Every older adult has two potential loneliness scores, one with the program and one without it, and we only ever see one. For a program that people are chosen for, the quantity we usually want is the average treatment effect on the treated, which is the average effect among the people who actually took part.

**Sarah:** And the score they would have had without the program is the counterfactual.

**Kiffer:** That's exactly right. The lesson writes out a simple decomposition. If you compare the average outcome of participants with the average outcome of a comparison group, that difference equals the effect on the treated plus a second term. The second term is the difference the two groups would have shown if neither had received the program. That's selection bias.

**Sarah:** So randomization makes that second term zero, on average.

**Kiffer:** Yes. Without randomization, every design in this lesson is a different strategy for making that term small or removing it. Difference-in-differences uses another group's change. Matching uses people who look similar. A synthetic control uses a weighted blend of other places.

**Sarah:** The section then brings in Shadish, Cook and Campbell. Who are they, for anyone who hasn't met them?

**Kiffer:** Donald Campbell was a psychologist and methodologist who, with Julian Stanley in nineteen sixty-three, wrote a short book cataloguing research designs and the reasons they can mislead. Thomas Cook and Campbell extended it in nineteen seventy-nine, and William Shadish, Cook and Campbell consolidated it in two thousand and two. They distinguish four kinds of validity. Statistical conclusion validity asks whether the program and outcome covary. Internal validity asks whether that covariation is causal. Construct validity asks whether we've labelled the program and outcomes correctly. External validity asks whether the result holds for other people, places and times.

**Sarah:** Which one matters most?

**Kiffer:** For a causal question, they give internal validity priority, because there's no point generalizing an effect that may not exist. In Cedar Valley, though, the readiness criterion affects two of them at once. It threatens internal validity, and it also means the effect in ready clinics may be larger than what the second wave will see.

**Sarah:** Let's go through the threats to internal validity.

**Kiffer:** The lesson focuses on seven. History is an outside event during the follow-up period, like a new seniors' centre opening or a bad respiratory virus season. Maturation is change that would have happened anyway, like loneliness easing after a bereavement. Selection is a difference between the groups before the program. Regression to the mean affects people chosen for extreme scores. Instrumentation is a change in how the outcome is measured. Testing is an effect of taking the pretest itself. And attrition is selective loss to follow-up. In Cedar Valley, fifty-three of the two hundred and forty-one people who attended a first meeting have no follow-up score, which is about twenty-two percent.

**Sarah:** The lesson makes a point about interactions too.

**Kiffer:** Adding a comparison group deals with a lot of these threats, because history and maturation would affect both groups. What survives are the interactions of selection with the other threats. A selection-history interaction is an event that hits one group only, like an urgent care centre opening near several first-wave clinics. A selection-maturation interaction is when the groups would have changed at different rates anyway. Those come back in Section two as the parallel trends assumption.

**Sarah:** The regression to the mean example really struck me. Can you walk through it?

**Kiffer:** The first Cedar Valley report showed that among one hundred and eighty-eight participants with both measurements, mean loneliness fell from seven point one to six point three on the three-item University of California, Los Angeles Loneliness Scale, which runs from three to nine. People were referred because they scored six or higher, so the group was selected on high scores.

**Sarah:** And some of those high scores were partly bad luck on the day.

**Kiffer:** That's the idea. Any measure that isn't perfectly reliable has a chance component. When we measure again, the chance part is drawn fresh, and the group mean moves back toward the population mean. The lesson gives a formula: the expected follow-up mean equals the population mean plus the test-retest correlation times the distance between the group's baseline mean and the population mean.

**Sarah:** What happens with some plausible numbers?

**Kiffer:** If the regional mean is four point six and the correlation is zero point seven, the expected follow-up mean with no program effect at all is six point three five. The observed mean was six point three. So, under those assumptions, regression to the mean alone could account for about three quarters of a point of the eight-tenths of a point drop. Under those assumptions, that is almost all of the drop. The honest conclusion is that the pre-post change can't tell the two explanations apart. The remedies are design remedies. You take a second baseline measurement at the first connector meeting, and you recruit a comparison group by the same rule, so that both groups regress by the same amount.

**Sarah:** The section ends with design notation. Is that just a shorthand, or does it do real work?

**Kiffer:** It does real work, because it forces you to say exactly what's observed, when and for whom. X is the program, O is an observation, and N R means nonrandom assignment. A one-group pre-post design is O, X, O. Adding a comparison group adds a second row.

**Sarah:** And the Cedar Valley design has a special name.

**Kiffer:** It's a switching replications design, because the comparison clinics get the program a year later. That means the effect can, in principle, be seen twice, at two different times, which makes a coincidental history event much less plausible. And because the health authority has quarterly records, we get eight observations before launch instead of one.

**Sarah:** Which brings us to Section two and difference-in-differences.

**Kiffer:** Right. Section two uses a simulated panel of twenty-four clinics observed for twelve quarters each, with emergency department visits per thousand enrolled patients aged sixty-five and older. The program starts in quarter nine in the first-wave clinics.

**Sarah:** And the lesson starts with two weak estimates.

**Kiffer:** If you compare first-wave and second-wave clinics after launch, you get minus seventeen point one visits per thousand per quarter. That looks like a big effect. But the first-wave clinics already had rates nearly ten visits lower before the program started. So the posttest-only comparison is mostly the pre-existing gap. The other weak estimate compares first-wave clinics before and after launch, which gives minus three point seven. That one understates the effect, because the comparison clinics' rates rose by about three and a half visits over the same period.

**Sarah:** And difference-in-differences combines them.

**Kiffer:** It takes the program group's change, minus three point seven, and subtracts the comparison group's change, plus three point six. That gives minus seven point three visits per thousand per quarter. The comparison group's change is standing in for the change the program clinics would have had without the program. That needs the parallel trends assumption. Without the program, the program group's average would have changed by the same amount as the comparison group's average. Notice what it allows. The groups can differ in level, so ready clinics can have lower rates. They have to be on the same path, though.

**Sarah:** Are there other assumptions hiding in there?

**Kiffer:** There are a few. No anticipation means first-wave staff didn't start changing things before launch. Stable composition means the program didn't change who's being counted. There should be no spillover between the groups. And the scale matters, because parallel trends in rates and parallel trends in logarithms are different assumptions, so you should choose the scale before you look at the results.

**Sarah:** Let's talk about the worked example. What does it show?

**Kiffer:** It starts with the panel. It plots the mean rate in each group by quarter and computes the difference-in-differences by hand from the four group means. Then it fits a regression with an interaction between first-wave status and the post period, and the interaction coefficient reproduces the hand calculation exactly. Next it fits a two-way fixed effects model, with a separate intercept for every clinic and every quarter. Because the panel is balanced and every first-wave clinic starts at once, the estimate is identical, minus seven point three. What changes is the standard error.

**Sarah:** This was the part about clustering.

**Kiffer:** Yes. A conventional standard error treats all two hundred and eighty-eight clinic-quarters as independent, and it comes out at about two point one. But a clinic that has a high rate one quarter tends to have a high rate the next. Clustering by clinic allows for that, and the standard error rises to about three point three, which is roughly one and a half times larger.

**Sarah:** Does that change the conclusion?

**Kiffer:** It changes how confident we are. The ninety-five percent confidence interval runs from about minus fourteen to about minus zero point four, so it still excludes zero, but only just. Bertrand, Duflo and Mullainathan showed in two thousand and four that ignoring this kind of serial correlation makes difference-in-differences studies reject true null hypotheses far too often. With only twenty-four clinics, the tests use a t distribution with twenty-three degrees of freedom.

**Sarah:** Then comes the event study. Explain what it adds.

**Kiffer:** The two-by-two estimate averages over all the quarters. An event study estimates a separate coefficient for each quarter relative to launch, with the quarter just before launch as the reference. So you can see whether the groups moved together before the program, and how the effect develops afterwards.

**Sarah:** What did it show in the Cedar Valley data?

**Kiffer:** The seven pre-launch coefficients sit near zero, between about minus four and zero, and a joint test of all seven gives a p-value of about zero point nine nine. After launch, the coefficients are roughly minus three, minus eight, minus fourteen and minus eight, which looks like an effect that builds as connectors' caseloads grow.

**Sarah:** So parallel trends is confirmed.

**Kiffer:** That's the claim I'd push back on. The assumption is about what would have happened after launch without the program, which we never see. The pre-launch coefficients only show that the groups moved together when neither had the program. And each coefficient has a standard error of about five and a half, so the test has very little power. A real divergence of a few visits could easily pass.

**Sarah:** So what do you do instead?

**Kiffer:** You still report the event study, because it's the best picture of the evidence. Jonathan Roth showed in two thousand and twenty-two that reporting results only when a pre-trend test passes can actually make bias worse. Rambachan and Roth proposed a sensitivity analysis that asks how big a violation of parallel trends would need to be, relative to the movements before launch, to overturn your conclusion.

**Sarah:** The last part of Section two was about staggered adoption, which I found the hardest part.

**Kiffer:** Once the second wave launches, every clinic has the program, and the rollout is staggered. A common approach has been to fit the same two-way fixed effects model to the whole period, switching the program indicator on for each clinic at its own launch date.

**Sarah:** How does that go wrong?

**Kiffer:** Andrew Goodman-Bacon showed that the estimate from that model is a weighted average of every two-group, two-period comparison in the data. Some of those comparisons use clinics that haven't started yet as controls, which is fine. Others use clinics that have already started as controls for the later ones.

**Sarah:** Why is that a problem, if trends are parallel?

**Kiffer:** If the earlier group's effect is still growing, its change over that period includes its own growing effect, and subtracting that change biases the comparison. The lesson has a toy example with no noise at all, in which the effect deepens by one visit per quarter of exposure. The true average effect is about minus three point eight, but the fixed-effects estimate is about minus one point two.

**Sarah:** That's a big miss, with perfect parallel trends.

**Kiffer:** It is. The clean comparison, wave one against a wave two that hasn't started yet, gives minus two point five. The forbidden comparison, wave two against an already-treated wave one, gives plus one point five, which looks harmful. The model weights them two thirds and one third, and you get minus one point two.

**Sarah:** What's the fix?

**Kiffer:** You use estimators that only compare treated units with units that haven't been treated yet or never will be. Callaway and Sant'Anna's group-time approach is one. Sun and Abraham, and Borusyak, Jaravel and Spiess, proposed related approaches. For Cedar Valley, the practical advice is to report the first-year effect with the second wave as the comparison and to avoid one pooled estimate over both years.

**Sarah:** Section three moves from clinics to people.

**Kiffer:** Yes. The steering committee wants to know whether the program reduced loneliness among participants. The comparison group is older adults screened at second-wave clinics during the same year. The simulated data have four hundred and twenty-three participants and eight hundred and seventy-seven comparisons. The groups differ by design. Participants were referred because they were lonely, and they chose to attend. At baseline they were lonelier, about six point three against five point five, and about half lived alone, compared with about a third of the comparison group.

**Sarah:** The lesson starts that section with a refresher on confounding.

**Kiffer:** A confounder is a common cause of getting the program and of the outcome. When clinicians refer the loneliest patients, baseline loneliness drives both referral and later loneliness, which is confounding by indication. You remove it by randomizing, or by conditioning on the confounders through stratification, regression, matching or weighting, and conditioning only works for the confounders you measured.

**Sarah:** What did the unadjusted comparison show?

**Kiffer:** It showed almost nothing. Six-month scores were about the same in both groups, a difference of about minus zero point zero four. If you stopped there, you'd conclude the program did nothing. But participants started out lonelier, so being equal at six months actually suggests they improved more.

**Sarah:** So where does the propensity score come in?

**Kiffer:** The propensity score is the probability of receiving the program given measured characteristics before the program. Paul Rosenbaum and Donald Rubin showed in nineteen eighty-three that it's a balancing score. Among people with the same score, the measured covariates have the same distribution in both groups. So instead of matching on six variables at once, you can match on one number.

**Sarah:** What assumptions does it rest on?

**Kiffer:** The big one is no unmeasured confounding, sometimes called conditional exchangeability. Then there's positivity, which means everyone has some chance of being in either group, and in practice you check it by looking at overlap. You also need a well-defined program, no interference, meaning one person's outcome doesn't depend on whether someone else got the program, and a model that actually balances the covariates.

**Sarah:** How do you decide what goes into the model?

**Kiffer:** You include variables measured before the program that affect the outcome, and usually the reasons people enter the program. You never include anything the program could have changed, like the number of connector meetings someone attended. The baseline loneliness score is the most important one, partly because it means both groups regress toward the mean by similar amounts.

**Sarah:** Tell me about the matching step.

**Kiffer:** In the worked example, logistic regression estimates each person's score. Each extra point of baseline loneliness multiplies the odds of enrolling by about one and a half, and living alone multiplies them by about one point eight. Then nearest-neighbour matching pairs each participant with the comparison whose score is closest, one to one, without replacement, with a caliper.

**Sarah:** Define caliper.

**Kiffer:** A caliper is the largest acceptable distance between matched scores. Peter Austin, a biostatistician in Toronto, recommended zero point two standard deviations of the logit of the propensity score. With that caliper, we get four hundred and ten pairs, and thirteen participants have no match.

**Sarah:** Who were the thirteen?

**Kiffer:** They were the people most likely to enrol, all with scores of about zero point six six or higher, where there were very few similar comparisons. Dropping them changes what we're estimating. The result now applies to the matched participants, which here is ninety-seven percent of them, and the report should say so.

**Sarah:** Then you check balance.

**Kiffer:** Balance is checked with standardized mean differences, which are the difference in means divided by a pooled standard deviation. Before matching, baseline loneliness had a standardized difference of about zero point six three and living alone about zero point three four. After matching, every covariate was below zero point zero five. The usual threshold is zero point one.

**Sarah:** Why not just use significance tests?

**Kiffer:** P-values depend on sample size, and matching changes the sample size. You can make a p-value bigger just by throwing people away. Imai, King and Stuart made this point in two thousand and eight. Standardized differences measure the size of the imbalance directly.

**Sarah:** And what was the effect estimate?

**Kiffer:** In the matched sample, participants' six-month loneliness was about zero point six eight points lower than their matches'. The confidence interval ran from about minus zero point eight two to minus zero point five four. Inverse probability weighting gave about minus zero point seven one, and regression adjustment gave about minus zero point seven.

**Sarah:** Quickly, what's inverse probability weighting?

**Kiffer:** Instead of discarding unmatched people, you weight everyone by the inverse of the probability of the group they're in. For the effect on participants, participants get a weight of one and comparisons get their score divided by one minus their score, which makes the comparison group look like the participants.

**Sarah:** So all the methods agree, the balance is excellent, and the interval is narrow. That sounds like a strong result.

**Kiffer:** This is the most important moment in the lesson. The simulation that generated the data included a trait that was never recorded, readiness to engage. It made people more likely to enrol and less lonely later. The true effect built into the simulated data is about minus zero point four two. So the estimate of minus zero point six eight overstates the benefit by about a quarter of a point, and the confidence interval doesn't even contain the true value. No balance diagnostic could have shown this, because balance can only be checked on what was measured. All the methods agreed because they all handled the measured covariates the same way.

**Sarah:** That's sobering. What can an evaluator actually do?

**Kiffer:** There are four things. Measure the likely reasons for selection at baseline, such as motivation and social support, in both groups. Combine matching with difference-in-differences when you have outcomes from before the program, so stable unmeasured differences drop out. Use a negative control outcome, something the program can't affect but that shares the confounding. And run a sensitivity analysis, such as Rosenbaum's bounds or the E-value from VanderWeele and Ding.

**Sarah:** Section four is synthetic control. When would I use that?

**Kiffer:** You'd use it when there's only one treated unit. A province changes a policy, or a city introduces free transit for seniors. You have one treated place and many potential comparisons, and picking any one of them is hard to defend. Alberto Abadie developed it with colleagues. Abadie and Gardeazabal used it in two thousand and three to study the economic costs of conflict in the Basque Country, and Abadie, Diamond and Hainmueller used it in two thousand and ten for California's nineteen eighty-eight tobacco control program. The idea is to build a synthetic version of the treated unit as a weighted average of untreated units from a donor pool.

**Sarah:** What kind of weights does it use?

**Kiffer:** The weights are zero or positive and add up to one, and they're chosen so that the weighted average tracks the treated unit before the intervention. After the intervention, the synthetic version's path is the estimate of what would have happened without it.

**Sarah:** The lesson has a Cedar Valley example.

**Kiffer:** It imagines the whole region as the treated unit, with annual emergency visit rates for Cedar Valley and fourteen other health service areas. The optimization puts half the weight on Area A, three tenths on Area B and one fifth on Area C. Before the program, the synthetic region stays within about two visits of the real one every year. After the program starts, Cedar Valley falls six visits below its synthetic control in year one and almost nine in year two.

**Sarah:** How do you know that isn't chance, when there's only one treated unit?

**Kiffer:** You use placebo tests. You apply the method to each donor area as if it had been treated and see how unusual the real result is. With fourteen donors, the smallest possible p-value is one in fifteen, about zero point zero six seven.

**Sarah:** Are there cautions for that example?

**Kiffer:** There's a big one. Only a small fraction of the region's older residents are in the program in any year, so a region-wide drop of nearly nine visits per thousand implies a large effect per participant. Before believing it, you'd check whether the program's reach could plausibly produce it, and look for region-wide events. A synthetic control can't rule out something that happened to Cedar Valley alone.

**Sarah:** The section closes by comparing all the designs.

**Kiffer:** Yes. Difference-in-differences uses the comparison group's change and assumes parallel trends. Propensity score methods use similar individuals and assume no unmeasured confounding. Synthetic control uses weighted donors and assumes that a good fit before the intervention would have continued. The choice mostly follows from the data you have.

**Sarah:** Give me the rule of thumb.

**Kiffer:** Many units observed over time suggest difference-in-differences with an event study. Individuals with rich baseline data suggest matching or weighting, with a sensitivity analysis. A single jurisdiction with a long series suggests a synthetic control. And if there's an eligibility threshold, or nothing to compare with, the designs in Lesson eight come in. You can also combine designs, and you should when you can. Lawlor, Tilling and Davey Smith call this triangulation. If designs with different weaknesses point the same way, you can be more confident. If they disagree, that tells you where to look.

**Sarah:** Let's finish with the comparison group memo. What does an evaluator write at this stage?

**Kiffer:** A memo that identifies a candidate comparison group for the program and describes the threats to validity it leaves open. It defines the group precisely, explains why it's the best available option, writes the design in notation, and gives a table of threats, how each could operate, and what the plan will do about it.

**Sarah:** What does that look like for Cedar Valley?

**Kiffer:** The comparison group is the twelve second-wave clinics during the year before they start, plus screened older adults in those clinics for the loneliness question. The memo lists the threats: ready clinics on a different path, local events that affect one group, anticipation and contamination, measurement by connectors, unmeasured motivation, attrition and limited generalizability. For each threat there's a concrete response. There's an event-study plot and a sensitivity analysis for different paths. There's a log of local service changes and a negative control population, patients aged fifty to sixty-four at the same clinics who aren't eligible. Research staff, never connectors, collect the follow-up scores. A short readiness measure is added at baseline. And then the memo makes a plain statement of what remains open.

**Sarah:** What's the most common mistake you see in these memos?

**Kiffer:** Writers often list threats generically, as if copying them from the textbook. The strong memos explain how each threat would actually operate in the program and what data would show whether it did. The other common mistake is promising mitigation that the evaluation can't afford or can't get the data for. My final advice is to be honest about what your comparison can and can't do. Every design leaves something open, and decision-makers trust an evaluation more when it names its weaknesses than when it hides them.

**Sarah:** That's Lesson seven. Next week, Lesson eight continues with interrupted time series, regression discontinuity and natural experiments.

**Kiffer:** And that's where we choose a design for the Cedar Valley evaluation and justify it. Thanks for listening.

**Sarah:** Thanks, everyone. See you next week.
