HSCI 826 · Lesson 8

Quasi-Experimental Designs II: Time Series, Discontinuities and Natural Experiments

Program Planning & Evaluation

Learning objectives for this lesson:

  • Explain how an interrupted time series projects a counterfactual from the pre-intervention level and trend, and identify the threats to validity that the design leaves open.
  • Interpret a segmented regression with level and slope changes, and explain how residual autocorrelation, Newey-West standard errors and a comparison series affect its conclusions.
  • Judge whether a time series has enough time points and enough sensitivity for an evaluation question, considering seasonality, autocorrelation and the program's reach.
  • Distinguish sharp from fuzzy regression discontinuity designs, calculate a fuzzy estimate with the Wald estimator, and apply checks for manipulation, covariate balance and bandwidth sensitivity.
  • Summarize the Medical Research Council guidance on natural experiments and describe Canadian provincial policies that have been studied as natural experiments.
  • State the assumptions of an instrumental variable analysis, calculate a Wald estimate from a binary instrument and interpret it as a local average treatment effect among compliers.
  • Use Plan-Do-Study-Act cycles, run charts and statistical process control charts appropriately, and identify the reporting elements of SQUIRE 2.0.
  • Compare quality improvement designs with the randomized and quasi-experimental designs of Lessons 6 to 8 by the question each answers and the assumption each relies on.
  • Choose a design for a program evaluation and justify it in writing, naming its assumptions and how they will be checked.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.

Lesson 8 · HSCI 826

Quasi-Experimental Designs II

Time series, discontinuities and natural experiments, with a short look at quality improvement designs.

Program Planning and Evaluation
Why these designs

Using how a program was delivered

Many programs start on a known date, follow an eligibility rule, or are shaped by events that evaluators do not control.

The designs in this lesson turn those features into comparisons. Each one rests on a specific assumption, and much of the lesson is about checking those assumptions.

The road map

Four sections

1 · Interrupted time series

Segmented regression estimates level and slope changes, with checks for seasonality and autocorrelation.

2 · Regression discontinuity

Sharp and fuzzy designs compare people just above and below an eligibility threshold.

3 · Natural experiments

Instrumental variables and Canadian policy examples show how outside events can be used.

4 · Quality improvement

Plan-Do-Study-Act cycles, run charts and control charts are compared with the evaluation designs.

Running case

Three opportunities in Cedar Valley

A monthly series

Emergency visits per 1,000 older adults are available for 48 months before launch and 24 after.

A referral rule

Older adults scoring 6 or higher on the three-item UCLA scale are referred, with room for judgement.

Limited capacity

Full connector caseloads in year two delayed some starts, arguably for reasons unrelated to the person.

The Cedar Valley Connector program is fictional, and its numbers are illustrative.

What you will do

Analysis and a design justification

  • You will follow a segmented regression analysis that checks its residuals, applies Newey-West standard errors and adds a comparison region.
  • A shorter worked example applies a fuzzy regression discontinuity to the referral threshold.
  • Section 4 ends with a design justification for the Cedar Valley evaluation.
Section 1 of 5

Interrupted Time Series

⏱ Estimated reading time: 50 minutes
Section 1 of 5

Interrupted Time Series

Projecting a counterfactual from the past, and testing whether a program changed the series.

The logic

A trend as the counterfactual

O1 O2 O3 O4 O5 X O6 O7 O8 O9 O10

What it controls

The pre-launch trend accounts for secular change and maturation.

What it leaves open

History, instrumentation and changes in the population remain possible explanations.

Segmented regression

Level change and slope change

Segmented regression
\[ Y_t = \beta_0 + \beta_1\,\text{time}_t + \beta_2\,\text{post}_t + \beta_3\,\text{time\_after}_t + e_t \]

Level change (β2)

This term is the step at the launch month relative to the projected trend.

Slope change (β3)

This term is the difference between the slopes after and before the launch.

Two problems of time

Seasonality and autocorrelation

Seasonality

Winter peaks can mimic a level change, so the model includes month indicators or Fourier terms.

Autocorrelation

Adjacent months resemble each other, which makes ordinary standard errors too small.

The autocorrelation function and the Durbin-Watson test detect the problem, and Newey-West standard errors or autoregressive models correct it.

Enough data?

Time points, power and dilution

  • A common working minimum for monthly data is twelve points before and twelve after the launch.
  • Power depends on variability, autocorrelation, effect size and where the launch falls, and it can be estimated by simulation.
  • A region-wide rate dilutes the effect of a program that reaches only a small share of residents.
Controlled design

Adding a comparison series

Emergency visit rates in Cedar Valley and Juniper Ridge with fitted segments
Both simulated regions drop at month 49; only Cedar Valley's slope changes afterward.
The worked example results

What the simulated data show

1.26Durbin-Watson statistic, showing positive autocorrelation
−2.03Single-series effect in month 72, per 1,000
−1.38Controlled effect in month 72, per 1,000

A plausibility check shows that even the controlled effect is larger than the program's reach could produce.

Carry forward

From dates to thresholds

  • An interrupted time series uses time as the basis for comparison.
  • Its main threat is an event that coincides with the launch.
  • A comparison series and a plausibility check strengthen the conclusion.

Learning Objectives for this section

  • Explain how an interrupted time series uses the pre-intervention level and trend to project a counterfactual, and name the threats to validity that the design leaves open.
  • Specify a segmented regression model with a level change and a slope change, and interpret each coefficient.
  • Describe how seasonality and autocorrelation affect estimates and standard errors, and how each is detected and handled.
  • Judge whether a series has enough time points and enough sensitivity to detect the effect a program could plausibly produce.
  • Interpret the results of a single and a controlled interrupted time series, including the residual checks and a plot of the counterfactual.

1.1 The Logic of an Interrupted Time Series

An interrupted time series measures an outcome at regular intervals before and after an intervention begins at a known point in time. The observations before the interruption establish the level of the outcome, its trend and its ordinary variation, and the design asks whether the series changed at the interruption by more than that pattern would predict. In the design notation of Shadish, Cook and Campbell (2002), a simple series with five observations on each side is written O1 O2 O3 O4 O5 X O6 O7 O8 O9 O10, where each O is a measurement and X marks the start of the program. Campbell and Stanley (1963) included this "time-series experiment" among their quasi-experimental designs and identified history as its main weakness, and it remains one of the most common designs for evaluating health policies and programs introduced to a whole population at once (Lopez Bernal et al., 2017).

Lesson 7 showed that a single pre-test and post-test cannot separate a program effect from a secular trend, maturation or regression to the mean. A long series addresses these threats directly. If emergency visits were already rising by about half a visit per 1,000 older adults each year, the counterfactual is the continuation of that rise, and comparing the post-program observations with the projected trend removes the part of the change that the trend explains. The pre-program points also show how much the series varies from month to month.

The design leaves other threats open. The most important is history: an event at about the same time as the program, such as a new provincial policy or an unusual respiratory virus season, can produce a change at the interruption. Instrumentation is a second threat, because a change in how the outcome is recorded, such as a new emergency department information system, can create a step in the series. Changes in the composition of the population can alter a rate without any change in individual risk, and regression to the mean reappears if a program was launched in response to an unusually high recent value.

Case: Cedar Valley's emergency department series

The Cedar Valley Connector program is a fictional community connector (social prescribing) program run by the fictional Cedar Valley Health Authority in British Columbia. Primary care clinicians refer adults aged 65 and older who score 6 or higher on the three-item UCLA Loneliness Scale (scored 3 to 9), or whom they judge to be isolated, to a community connector who meets them up to six times over twelve weeks. The program's planners expect fewer emergency department visits, although Lesson 1 recorded this as an expectation that had not been tested. The health authority's analyst can extract monthly emergency department visits per 1,000 residents aged 65 and older for 72 months: 48 months before launch and 24 months after. Month 1 is a January. Wave one (the 12 clinics chosen for readiness, as Lesson 7 described) began at month 49, and wave two (the other 12 clinics) began at month 61. A fictional neighbouring region, Juniper Ridge, has no connector program and supplies a comparison series. All Cedar Valley and Juniper Ridge numbers in this lesson are illustrative and come from simulated data.

Background: regression on time

A linear regression with time as its only predictor, Yt = β0 + β1 × timet + et, fits a straight line through a series: β0 is the fitted value at time zero, β1 is the change in the outcome per month, and et is the residual, the distance of each month's observation from the line. Segmented regression adds two predictors that switch on at the launch. An indicator that equals 0 before the launch and 1 from the launch month onward lets the line jump at the launch, and its coefficient is the level change. A counter of months since the launch, which is 0 before it and 0, 1, 2 and so on from the launch month, lets the line bend, and its coefficient is the slope change, the difference between the post-launch and pre-launch slopes. Ordinary standard errors assume that the residuals are independent, whereas in a monthly series a month above the line tends to be followed by another month above it. This positive autocorrelation means that the series holds less independent information than its number of months suggests, so ordinary standard errors are too small, and Newey-West standard errors recompute them from the observed correlation between nearby residuals.

HSCI 410 Lesson 3 (Linear and Logistic Regression) covers linear regression more fully and is optional reading.

1.2 Segmented Regression

The standard analysis of an interrupted time series is segmented regression, which fits a separate straight line to the periods before and after the interruption and estimates how the level and the slope changed (Wagner et al., 2002). With one interruption, the model is:

Segmented regression model

Yt = β0 + β1 × timet + β2 × postt + β3 × time_aftert + et

Here timet counts months from the start of the series (1 to 72), postt equals 0 before the launch and 1 from the launch month onward, and time_aftert equals 0 before the launch and then 0, 1, 2 and so on from the launch month. β0 is the level at time zero, β1 is the pre-launch slope, β2 is the level change at the launch month relative to the projected trend, and β3 is the slope change. The post-launch slope is β1 + β3, and the estimated effect k months after the launch month is β2 + β3 × k.

Time (months) Outcome rate Program launch Level change (β2) Effect at k = β2 + β3 × k Pre-launch slope (β1) Counterfactual Post-launch slope (β1 + β3)
The dashed counterfactual continues the pre-launch level and slope; the level change is the step at the launch month, and the effect at a later month adds the accumulated slope change.

The coding of time_after determines what β2 means. With time_after equal to 0 at the launch month, β2 is the difference between the fitted and projected values in that month; if time_after starts at 1, the level change refers to an earlier point and changes slightly while the slope change does not. The evaluation plan should state the coding. Because the effect at a later month combines two coefficients, its confidence interval needs both variances and their covariance, as the worked example in Section 1.6 shows.

Specifying the impact model in advance

Lopez Bernal and colleagues (2017) recommend that the evaluator state the expected shape of the effect, the impact model, before analyzing the data, because a shape chosen after looking at the series can fit random fluctuation. The impact model should follow from the program theory developed in Lesson 3.

An immediate, persistent step is expected when a policy changes conditions for everyone at once, such as a new fee or a ban. For Cedar Valley, a region-wide step at launch would be surprising, because only a few hundred older adults are referred in the first months, and it would prompt a search for co-occurring events.

A gradual divergence from the projected trend is expected when a program's reach accumulates over time. The Connector program enrols older adults continuously and adds the second wave of clinics at month 61, so its program theory predicts a slope change in emergency visits as the number of past participants grows.

Estimating both terms is the usual default, because a program can produce an early step, for example from changes in clinic practice, followed by a growing effect. The evaluator then reports the effect at specified months after launch.

Some effects begin after a delay, such as the twelve weeks a participant spends with a connector, and some fade. A lag is handled by excluding a transition period, and a temporary effect needs an additional segment; both choices should be fixed before the analysis.

1.3 Seasonality and Autocorrelation

Emergency department visits by older adults usually rise in winter with the respiratory virus season. If the program launches near a seasonal peak or trough, an unmodelled seasonal cycle can masquerade as a level change, and seasonal swings reduce precision. The model can include indicators for calendar month (eleven terms for monthly data) or Fourier terms, which are pairs of sine and cosine functions with a twelve-month period; one pair, sin(2π × month / 12) and cos(2π × month / 12), describes a smooth annual cycle with two parameters, which matters in short series. A second pair can be added if the seasonal shape is not a single smooth wave.

Autocorrelation means that observations close together in time resemble each other more than observations far apart. In monthly health data, a busy month tends to be followed by another busy month, so the residuals of a segmented regression are often positively correlated. Ordinary least squares still gives unbiased coefficients when the model is correctly specified, but with positive autocorrelation its standard errors are usually too small, so confidence intervals are too narrow. The evaluator checks by plotting the residuals against time, by examining the autocorrelation function (ACF), which shows the correlation between residuals one, two or more months apart, and by formal tests. The Durbin-Watson test (Durbin and Watson, 1950) examines first-order autocorrelation: its statistic is close to 2 when there is none and falls toward 0 as positive autocorrelation increases. The Breusch-Godfrey test examines several lags at once.

Newey-West standard errors (Newey and West, 1987) keep the ordinary least squares estimates and replace the standard errors with ones that remain valid under autocorrelation and unequal variance up to a chosen lag. Generalized least squares with first-order autoregressive errors, such as the Prais-Winsten estimator, models the autocorrelation directly, and autoregressive integrated moving average (ARIMA) models handle more complex dependence in long series. When the outcome is a count, a Poisson or negative binomial model with the logarithm of the population as an offset estimates a rate ratio (Lopez Bernal et al., 2017).

Plot the series before modellingv

A plot of the raw series shows the trend, the seasonal pattern, outliers and data problems, such as a month with missing records that appears as a sudden drop.

Look for one-off eventsv

Record the dates of events that could affect the outcome, such as a pandemic wave, a hospital closure or a change in coding rules. Real series covering 2020 to 2022 need explicit treatment of the COVID-19 pandemic, which disrupted emergency department use across Canada.

Confirm the denominatorv

A rate per 1,000 residents depends on population estimates, which are often interpolated between censuses, and a change in the estimation method can create an artificial step.

1.4 How Many Time Points?

There is no single rule for the length of a series. The Cochrane Effective Practice and Organisation of Care group requires at least three points before and three after for inclusion as an interrupted time series in its reviews, which is a minimum for inclusion. For monthly data, twelve observations before and twelve after is often cited as a working minimum (Wagner et al., 2002), partly because a full year on each side helps separate seasonality from the effect. More pre-launch points sharpen the counterfactual trend, and more post-launch points allow a slope change to be detected and its persistence assessed.

Power depends on the number of points, the residual variability, the autocorrelation, the size and shape of the expected effect and the position of the interruption. Zhang, Wagner and Ross-Degnan (2011) showed how to estimate power by simulation: the evaluator generates many series with plausible features and an assumed effect, fits the planned model to each, and counts how often the effect is detected.

The level of aggregation also matters. An outcome measured for a whole region dilutes the effect of a program that reaches a small share of its population. Before relying on a population-level series, the evaluator should estimate the change the program could plausibly produce, given its reach, and compare it with the month-to-month variability of the series. The worked example in Section 1.6 ends with such a check.

1.5 Controlled Interrupted Time Series

A single series cannot rule out history, because any event that coincides with the launch changes the series in the same way as the program. A controlled interrupted time series adds a comparison series that should be exposed to the same co-occurring events but not to the program (Lopez Bernal et al., 2018). The key assumption is that, without the program, the difference between the two series would have continued along its pre-launch path. A controlled interrupted time series is closely related to the difference-in-differences designs of Lesson 7, with many periods and an explicitly modelled trend on each side of the interruption.

The analysis can stack both series and add a group indicator and its interactions with each term of the segmented model: Yt = β0 + β1time + β2post + β3time_after + β4group + β5group × time + β6group × post + β7group × time_after. The coefficients β6 and β7 estimate how much more the level and the slope changed in the program series than in the comparison series. An equivalent approach, used in the worked example in Section 1.6, fits the segmented model to the difference between the two series month by month, which removes shared seasonality and shared shocks.

Comparison populationClick to explore
Comparison age groupClick to explore
Comparison outcomeClick to explore

1.6 Worked Example: An Interrupted Time Series for Cedar Valley

The worked example uses a simulated monthly series of emergency department visits per 1,000 residents aged 65 and older for Cedar Valley and Juniper Ridge over 72 months. The simulation includes a winter cycle, a gentle upward trend, first-order autocorrelation and a small drop in both regions in the launch month that stands in for a hypothetical province-wide change. Its rates, about 0.5 visits per resident per year, are set somewhat higher than the illustrative annual rates in Lessons 2 and 7 (about 0.4). The worked example concerns the changes in the series, and its results describe this teaching dataset only.

Worked example: An interrupted time series for Cedar Valley

The naive comparison. Each region has 48 pre-launch and 24 post-launch months. The mean rate fell slightly in Cedar Valley, from 41.59 to 41.36 visits per 1,000, and rose in Juniper Ridge, from 44.57 to 44.98. This comparison ignores the upward trend under way before the launch.

The segmented regression. The model for Cedar Valley contains time, post and time_after, coded as in Section 1.2, and one pair of Fourier terms for the winter cycle, as described in Section 1.3. The table gives the main estimates with ordinary standard errors. The coefficients of the sine and cosine terms are 1.5316 and 1.9442.

TermEstimateStandard error
Intercept40.68420.1527
time (pre-launch slope per month)0.03700.0054
post (level change at launch)−1.09130.2617
time_after (slope change per month)−0.04080.0163

Before the launch, the rate rose by 0.0370 visits per 1,000 per month, about 0.44 per year. At the launch month the level fell by 1.0913 visits per 1,000 relative to the projected trend, and the slope changed by −0.0408 per month, so the post-launch slope was 0.0370 − 0.0408 = −0.0038, which is close to flat. The two Fourier terms capture the winter peak, and the model explains 92.8 percent of the variance (R2 = 0.928). These ordinary standard errors assume independent residuals, an assumption checked next.

Residual autocorrelation. The lag-one autocorrelation of the residuals is 0.369, above the approximate bound of 0.236, while the autocorrelations at lags 2 to 6 lie within it. The Durbin-Watson statistic is 1.2551 (p = 5.257e-05), well below 2, which confirms positive first-order autocorrelation.

Bar chart of residual autocorrelations at lags 1 to 12, with only lag 1 exceeding the dashed 95 percent bounds
Residual autocorrelation for the Cedar Valley segmented model; only the lag-one value lies outside the approximate 95 percent bounds.

Newey-West standard errors and the effect at month 72. The standard errors were recomputed with a Newey-West estimator using three lags, and the effect was estimated for month 72, the 24th month of the program, when time_after equals 23.

CoefficientOrdinary standard errorNewey-West standard errorp (Newey-West)
Level change (post)0.26170.35750.0033
Slope change (time_after)0.01630.02090.0555

Allowing for autocorrelation increases the standard error of the level change from 0.2617 to 0.3575 and that of the slope change from 0.0163 to 0.0209. The level change remains clearly different from zero (p = 0.0033), while the p-value for the slope change rises to 0.0555. The estimated effect in month 72 is −2.029 visits per 1,000 (95 percent confidence interval −2.755 to −1.303), against a counterfactual of 43.350, a relative reduction of 4.7 percent.

The fitted segments and the counterfactual. The figure shows the fitted lines with the Fourier terms set to zero, their average over a year, together with the counterfactual, which sets post and time_after to zero in the post-launch months.

Monthly emergency visit rates in Cedar Valley over 72 months with a strong winter cycle, a rising fitted line before month 49, a lower and nearly flat fitted line afterward, and a dashed red counterfactual continuing the earlier rise
Cedar Valley's simulated series with deseasonalized fitted segments and the counterfactual projection of the pre-launch trend.

A controlled interrupted time series. The same model was first fitted to Juniper Ridge, and a segmented regression was then fitted to the month-by-month difference between the regions, in which shared seasonality and shared shocks cancel.

Juniper Ridge, which has no program, also dropped at the launch month, by 0.7687 visits per 1,000 (p = 0.0075), with no change in slope (0.0051, p = 0.716). This pattern points to a co-occurring event that affected both regions, which in the simulation represents a hypothetical province-wide change. In the difference model, Cedar Valley's level change relative to Juniper Ridge is −0.3989 (p = 0.38185) and its slope change is −0.0426 per month (p = 0.08917). The controlled effect in month 72 is −1.378 visits per 1,000 (95 percent confidence interval −2.136 to −0.620), about a third smaller than the single-series estimate of −2.029. Most of the single-series level change was shared with the comparison region, and the remaining evidence for an effect rests on the divergence in slopes.

Monthly emergency visit rates in Cedar Valley and Juniper Ridge with fitted segments; both regions step down at month 49, and only Cedar Valley's post-launch line flattens
Both regions drop at the launch month, which the comparison series attributes to a shared event; only Cedar Valley's slope changes afterward.

A plausibility check on the size of the effect

Suppose, for illustration, that about 5 percent of the region's residents aged 65 and older had attended the program by month 72. This is a generous assumption, because the attendance figures reported in this course imply a share closer to 3 percent of the region's 46,000 older adults, which would make the required effect per participant larger still. A region-wide reduction of 1.378 visits per 1,000 residents per month would then require a reduction of about 27.6 visits per 1,000 participants per month (1.378 / 0.05), or about 0.33 visits per participant per year. The counterfactual rate of 43.350 per 1,000 per month corresponds to about 0.52 visits per resident per year, so participants would have to avoid about 64 percent of the emergency visits an average older resident makes, which is difficult to believe for a twelve-week program. The simulated effect was set large so that the methods are easy to see. In a real evaluation, this arithmetic would mark the region-wide rate as an insensitive outcome, and the plan would add a series restricted to referred older adults using linked records.

Reflection

A health authority introduced a pharmacist-led medication review for adults aged 65 and older in month 37 of a 60-month monthly series of falls-related emergency department visits per 10,000 residents aged 65 and older. The analyst fitted a segmented regression in which time counts months from 1 to 60, post equals 1 from month 37 onward, and time_after equals 0 before month 37 and then 0, 1, 2 and so on from month 37. The estimates are a pre-intervention slope of 0.05 visits per month, a level change of −1.2 (ordinary standard error 0.40, Newey-West standard error 0.55) and a slope change of −0.08 per month (ordinary standard error 0.03, Newey-West standard error 0.04). The Durbin-Watson statistic is 1.31, and the projected counterfactual rate in month 60 is 32.0 per 10,000. In the month the review began, the province also changed the emergency department triage codes used to classify falls. (a) Calculate the estimated effect in month 60 and express it as a percentage of the counterfactual. (b) State which standard errors you would report and why. (c) Explain the threat that the coding change creates and propose one comparison series that would address it.

Model answer

(a) Month 60 is the 24th month of the intervention, so time_after equals 23. The estimated effect is −1.2 + 23 × (−0.08) = −1.2 − 1.84 = −3.04 visits per 10,000. Against a counterfactual of 32.0, this is a reduction of 3.04 / 32.0 = 9.5 percent.

(b) A Durbin-Watson statistic of 1.31 is well below 2 and indicates positive first-order autocorrelation, so the ordinary standard errors are likely too small. I would report the Newey-West standard errors. With them, the level change is −1.2 / 0.55 = −2.18 standard errors from zero and the slope change is −0.08 / 0.04 = −2.0, so both are near conventional significance (two-sided p of about 0.03 and 0.05), and the evidence is weaker than the ordinary standard errors suggest. A confidence interval for the month-60 effect would also require the covariance of the two coefficients.

(c) The coding change is an instrumentation threat: if the new triage codes classify fewer visits as falls, the series would step down at month 37 with no change in actual falls, and the level change would be wrongly credited to the review. A good comparison series is falls-related emergency visits among residents aged 50 to 64 in the same province, who were not offered the review but were coded under the same new system. A drop of similar size in that series would point to the coding change.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a segmented regression in which time_after equals 0 before the launch and 0, 1, 2 and so on from the launch month, what does the coefficient on post estimate?

With this coding, post shifts the line at the launch month, so its coefficient is the step relative to the projected trend. The slope difference (option c) is the coefficient on time_after, and the effect in the final month (option d) combines both coefficients.

Question 2: A segmented regression of monthly data has a Durbin-Watson statistic of 1.26 and a lag-one residual autocorrelation of 0.37. What is the most likely consequence of reporting ordinary least squares standard errors?

Positive autocorrelation leaves the coefficients unbiased when the model is correctly specified but makes ordinary standard errors too small, so intervals are too narrow. Option c describes the opposite problem; in the worked example, Newey-West standard errors were larger than the ordinary ones.

Question 3: In the Cedar Valley worked example, the Juniper Ridge series dropped by 0.7687 visits per 1,000 at the launch month with no change in slope. What does this finding most directly suggest?

A region without the program changed at the same month, which points to a shared event, such as the hypothetical province-wide change in the simulation. A difference in average level (option c) does not disqualify a comparison series, because the controlled analysis models the difference between the series.

Question 4: Why is a region-wide emergency department rate a potentially insensitive outcome for the Cedar Valley Connector program?

Dilution means that even a meaningful effect among participants produces a small change in a population rate. The plausibility check in Section 1.6 showed that the simulated region-wide effect would require participants to avoid an implausibly large share of their visits.
Section 2 of 5

Regression Discontinuity

⏱ Estimated reading time: 45 minutes
Section 2 of 5

Regression Discontinuity

Comparing people just above and just below an eligibility threshold.

The idea

Assignment by a threshold

The running variable decides eligibility at a threshold, and people near the threshold are alike apart from eligibility.

Origins

Thistlethwaite and Campbell introduced the design in 1960 to study merit awards.

Health examples

Age 65, clinical cut-offs, birth-date rules and provincial drinking ages have all been used.

Two kinds of design

Sharp and fuzzy

Sharp

The probability of receiving the program jumps from 0 to 1 at the threshold.

Fuzzy

The probability jumps by less than 1, because of discretion, refusal or exceptions.

The Cedar Valley rule is fuzzy, because clinicians may refer older adults below 6 and some above 6 are never referred.

The Wald estimator

Scaling the jump in the outcome

Fuzzy regression discontinuity
\[ \tau = \frac{\text{jump in mean outcome at the threshold}}{\text{jump in } P(\text{program}) \text{ at the threshold}} = \frac{-0.1985}{0.4089} = -0.485 \]

In the simulated data, referral lowered the six-month score by about half a point for older adults at the margin of eligibility.

Estimation

Bandwidth and local linear regression

  • A narrow bandwidth reduces bias but adds noise, and a wide one does the reverse.
  • Data-driven bandwidths and bias-corrected intervals are available in the rdrobust package.
  • High-order polynomials across the whole range give unstable estimates.
  • A score with only seven values requires extrapolation and caution.
Checks

Is the comparison clean?

Density

An excess of people just above the threshold suggests manipulation of the score.

Covariates

Characteristics fixed before screening, such as age, should not jump.

Placebo thresholds

No jump should appear at values where no rule applies.

Bandwidths

The estimate should be similar across reasonable windows.

Interpretation

A local effect for a local decision

Probability of referral and mean six-month score by screening score, with a jump at 6
The first stage jumps sharply at 6; the reduced form steps down slightly.
Carry forward

From thresholds to instruments

  • A fuzzy design treats crossing the threshold as an instrument for receiving the program.
  • The estimate applies to compliers at the threshold.
  • Section 3 generalizes this logic to other instruments and natural experiments.

Learning Objectives for this section

  • Explain the logic of a regression discontinuity design and why units close to an eligibility threshold provide a credible comparison.
  • Distinguish sharp from fuzzy regression discontinuity designs, and calculate a fuzzy estimate with the Wald estimator.
  • Describe local linear estimation, bandwidth choice and the problems created by a running variable with few distinct values.
  • Apply the standard checks for manipulation of the running variable, covariate balance and placebo thresholds.
  • Interpret a regression discontinuity estimate as a local effect and judge its relevance to a program decision.

2.1 Assignment by a Threshold

Many health programs decide eligibility with a rule applied to a measured score: adults become eligible at a given age, patients receive a treatment when a laboratory value crosses a cut-off, families qualify for a subsidy below an income limit, and older adults are referred to a connector when a screening score reaches a set value. A regression discontinuity design uses such a rule to estimate the program's effect. The score that determines eligibility is called the running variable (also the forcing or assignment variable), and the cut-off is the threshold. People just below and just above the threshold are similar in everything except eligibility, so a jump in the outcome at the threshold, beyond what the smooth relationship between the score and the outcome would predict, estimates the effect of the program for people near the threshold.

Thistlethwaite and Campbell (1960) introduced the design to study the effect of merit awards given to students whose test scores exceeded a cut-off. In the notation of Shadish, Cook and Campbell (2002), the design is written as two rows, OA C X O and OA C O, where OA is the pre-assignment measure of the running variable, C indicates that assignment was made by cut-off, X is the program and O is the outcome. The formal condition for identification is continuity: the average outcome that people would have without the program must change smoothly with the running variable as it crosses the threshold (Hahn, Todd and Van der Klaauw, 2001). Lee (2008) showed that when people cannot precisely control their own score, assignment near the threshold behaves almost like a randomized experiment, which is why the design is often described as having a "local randomization" interpretation.

Running variable (screening score) Later outcome Threshold Jump at the threshold estimates the effect Not eligible Eligible Projection without program
In a sharp design, the outcome line shifts at the threshold; the size of the shift, compared with the projection from below, estimates the program's effect for people with scores near the threshold.

Eligibility thresholds are common in health programs, and evaluators have used them to answer questions that randomization could not. The cards below give examples, including two from Canadian provinces.

Age thresholdsClick to explore
Clinical thresholdsClick to explore
Birth-date thresholdsClick to explore
Legal age thresholdsClick to explore
Case: the Cedar Valley referral threshold

In the fictional Cedar Valley Connector program, clinicians at wave-one clinics screen older adults with the three-item UCLA Loneliness Scale and refer those who score 6 or higher, or whom they judge to be isolated, to a community connector. In the program's first twelve months, the 12 wave-one clinics screened 2,600 older adults, of whom 705 (27.1 percent) scored 6 or higher, and 616 were referred. With 24 further older adults referred on clinician judgement without a recorded screening score, the clinics made 640 referrals in the year, 312 of them in the first six months. The clinics repeat the three-item screen at a routine visit about six months later. The steering committee asks whether referral reduces loneliness for older adults at the margin of eligibility, and whether the threshold should be lowered to 5.

2.2 Sharp and Fuzzy Designs

In a sharp regression discontinuity design, the rule is followed exactly: everyone at or above the threshold receives the program and no one below it does, so the probability of receiving the program jumps from 0 to 1. The effect is estimated as the difference between the limit of the average outcome as the score approaches the threshold from above and the limit as it approaches from below. In a fuzzy design, the probability of receiving the program jumps at the threshold by less than 1, because some eligible people are not treated and some ineligible people are. Clinician discretion, patient refusal, waiting lists and exceptions all make designs fuzzy. The Cedar Valley rule is fuzzy by construction, because clinician judgement can bring a person with a score below 6 into the program, and some older adults who score 6 or higher are never referred.

Sharp design Fuzzy design 0 1 0 1 Jump less than 1 Score (threshold dashed) Score (threshold dashed) P(program) P(program)
In a sharp design the threshold determines receipt completely; in a fuzzy design it changes the probability of receipt, and the size of that change becomes the denominator of the estimate.

In a fuzzy design, the jump in the outcome at the threshold is diluted, because only some of the people who cross the threshold change their program status. This jump is the reduced form, and it is the regression discontinuity analogue of an intention-to-treat effect. The jump in the probability of receiving the program is the first stage. Dividing the reduced form by the first stage gives the Wald estimator (Hahn, Todd and Van der Klaauw, 2001):

The fuzzy regression discontinuity (Wald) estimator

Effect at the threshold = (jump in the average outcome at the threshold) / (jump in the probability of receiving the program at the threshold)

For example, if the probability of referral jumps by 0.40 at a score of 6 and the average six-month loneliness score drops by 0.20 points, the estimated effect of referral is −0.20 / 0.40 = −0.50 points for the older adults whose referral depended on the threshold.

The fuzzy design is an instrumental variable analysis in which being above the threshold is the instrument, a connection that Section 3 develops. The estimate applies to compliers at the threshold: people who would be referred if their score were 6 and would not be referred if it were 5. Two further assumptions are needed. Crossing the threshold must affect the outcome only through receipt of the program (the exclusion restriction), and no one may be less likely to receive the program because their score crossed the threshold (monotonicity). Lesson 6 met the same logic when it divided an intention-to-treat effect by the difference in program uptake to obtain a complier average causal effect.

2.3 Estimation and Bandwidth

Because the design compares people near the threshold, the analysis concentrates on observations within a window, called the bandwidth, on either side of it. The standard estimator is local linear regression: a straight line is fitted to the outcome on each side of the threshold within the bandwidth, and the effect is the difference between the two fitted lines at the threshold (Imbens and Lemieux, 2008). Imbens and Lemieux suggested equal weights within the bandwidth (a rectangular kernel) for simplicity, and current software usually applies weights that decline with distance from the threshold (a triangular kernel). Centring the running variable at the threshold makes the coefficient on the threshold indicator equal to that difference. Covariates are unnecessary for identification, although adjusting for variables measured before assignment can improve precision.

The choice of bandwidth is a trade-off between bias and variance. A narrow bandwidth compares people who are very similar, which reduces bias from curvature in the relationship between score and outcome, but it uses fewer observations and gives a noisier estimate. A wide bandwidth adds observations but relies more heavily on the straight-line approximation. Imbens and Kalyanaraman (2012) derived a bandwidth that minimizes mean squared error, and Calonico, Cattaneo and Titiunik (2014) showed that confidence intervals built around such bandwidths need a bias correction to achieve their stated coverage. Their methods are implemented in the rdrobust package for R and Stata. Good practice reports the main estimate and shows how it changes across a range of bandwidths. Gelman and Imbens (2019) argued against fitting high-order polynomials to the whole range of the running variable, because such fits give noisy estimates that depend heavily on points far from the threshold.

The three-item UCLA scale creates a specific difficulty. The running variable takes only seven values (3 to 9), so there are no observations between 5 and 6, and the fitted line from below must be extrapolated one full point to reach the threshold. The bandwidth cannot shrink below one point, and the analysis rests on the assumption that the relationship between score and outcome is close to linear over the range used. Lee and Card (2008) and Kolesár and Rothe (2018) discuss inference when the running variable is discrete. In practice, the evaluator reports results for more than one window, as the worked example in Section 2.6 does, and is cautious about precise confidence intervals.

Decisions to state in the analysis planv

The plan should name the running variable and threshold, the estimator (local linear regression is the usual choice), the bandwidth rule and the alternative bandwidths to be reported, the kernel, any covariates, how standard errors account for clustering (for example, by clinic), and the outcomes and follow-up times.

Regression discontinuity in timev

When the running variable is calendar time and the threshold is a policy start date, the design resembles an interrupted time series. Analyses of this kind compare observations just before and just after the start date, and they face the time-series threats discussed in Section 1, such as autocorrelation and co-occurring events.

Multiple thresholds and multiple programsv

If other programs use the same threshold, the jump reflects all of them. Age 65, for example, triggers several benefits at once. In Cedar Valley, the evaluator should confirm that no other service uses a UCLA score of 6 to determine eligibility.

2.4 Checking the Design

The main threat to a regression discontinuity design is manipulation of the running variable. If people can control their score precisely, those who end up just above the threshold may differ from those just below. A clinician who wants to secure a referral might record a borderline score of 5 as 6, or an older adult who does not want a referral might understate loneliness. Such sorting usually leaves a trace in the distribution of the running variable: an excess of people just above the threshold (heaping) and a deficit just below. McCrary (2008) proposed a formal test for a discontinuity in the density of the running variable, and Cattaneo, Jansson and Ma (2020) developed a local polynomial density test implemented in the rddensity package. With a discrete score, the evaluator compares the frequencies at 5 and 6 with the smooth pattern across the other values.

A second check examines covariates fixed before assignment, such as age, sex, living alone and emergency visits in the previous year. If the design is valid, these should not jump at the threshold, and the same local regression used for the outcome is applied to each covariate. A third check uses placebo thresholds: the analysis is repeated at values where no rule applies (for example, a score of 4 or 8), where no jump should appear. Outcomes that the program cannot affect should also show no discontinuity. When heaping is found, a donut design that excludes observations at the threshold can test whether the result depends on the people who may have been sorted.

CheckWhat it examinesA worrying result
Density of the running variableWhether people sort themselves or are sorted around the thresholdToo many observations just above the threshold and too few just below
Covariate continuityWhether people just above and below are similar before assignmentA jump in age, prior health care use or living arrangements
Placebo thresholdsWhether jumps appear where no rule appliesJumps of similar size at values other than the real threshold
Bandwidth sensitivityWhether the estimate depends on the windowEstimates that change sign or size substantially across reasonable bandwidths

2.5 Interpreting a Local Effect

A regression discontinuity estimate describes the effect for people near the threshold, and in a fuzzy design only for compliers near the threshold. It does not describe the effect for older adults who score 9, who may be lonelier and may respond differently, or for those who would be referred by clinician judgement regardless of their score. This limitation matters less than it might appear when the decision concerns the threshold itself. The Cedar Valley steering committee's question, whether to lower the threshold from 6 to 5, concerns exactly the people at the margin, and a credible local effect speaks directly to it. A decision about whether to fund the program for all lonely older adults would need evidence from other designs as well, such as the comparison of wave-one and wave-two clinics in Lesson 7.

2.6 Worked Example: A Fuzzy Regression Discontinuity at the Referral Threshold

This short worked example uses a simulated dataset of the 2,600 older adults screened at the wave-one clinics in the program's first year. It shows the jump in the probability of referral at a score of 6, the jump in the six-month loneliness score, and the Wald estimator, each estimated with simple linear models. The data are simulated, and the results describe this teaching dataset only.

Worked example: A fuzzy regression discontinuity at the referral threshold

The running variable and the referral rule. The table gives, for each screening score, the number of older adults screened, the proportion referred and the mean six-month score.

Screening scoreOlder adultsProportion referredMean six-month score
38430.0333.622
46060.0814.226
54460.1594.868
63310.6225.302
72100.6676.071
81040.7406.933
9600.7507.850

The frequencies decline smoothly from 843 at a score of 3 to 60 at 9, with 446 at 5 and 331 at 6 and no excess just above the threshold, so the distribution gives no sign of sorting. The proportion referred rises slowly below the threshold (0.033, 0.081, 0.159), reflecting clinician judgement, and then jumps to 0.622 at a score of 6. The mean six-month score rises with the screening score on both sides, as expected, with a smaller step between 5 and 6 (4.868 to 5.302) than between 6 and 7 (5.302 to 6.071).

First stage, reduced form and a covariate check. The score is centred at the threshold, and each model fits a separate line on each side over all scores from 3 to 9, so the estimated jump is the difference between the two lines at a score of 6. The first stage has a standard error of 0.0280 and the reduced form a standard error of 0.0837.

Crossing the threshold raises the probability of referral by 0.4089 (the first stage) and lowers the mean six-month score by 0.1985 points (the reduced form, p = 0.0178). Age, which cannot be affected by the referral rule, shows no clear jump (−0.5407 years, p = 0.3031).

The Wald estimate. The Wald estimate is −0.1985 / 0.4089 = −0.485 points on the three-item scale: among older adults whose referral depended on crossing the threshold, referral lowered the six-month score by about half a point. Restricting the analysis to scores 4 to 7 gives −0.539, a similar value. A percentile bootstrap interval based on 2,000 resamples of older adults runs from −0.927 to −0.047, which is wide because the first stage divides the reduced form by 0.4089, and a fuller analysis would resample clinics to respect clustering.

The two naive comparisons show why the design matters. Referred older adults had a mean six-month score 1.217 points higher than those not referred, because referral selects lonelier people. Among referred older adults, the mean score fell from 6.286 at screening to 5.547 six months later, a drop of 0.739 points that includes regression to the mean, since people selected for high scores tend to score lower when measured again. The regression discontinuity estimate avoids both biases, at the cost of describing only the margin of eligibility.

Two panels: the probability of referral rises slowly from 0.03 to 0.16 for scores 3 to 5 and jumps to about 0.62 at a score of 6; the mean six-month score rises with the screening score and shows a small downward step at 6
The first stage (left) shows a large jump in referral at a score of 6; the reduced form (right) shows a small step down in the six-month score at the same point. Lines are the fitted linear models on each side of the threshold.

Reflection

A provincial home-support program offers subsidized weekly help to adults aged 65 and older whose frailty index, a score from 0 to 1 recorded by a home-care assessor, is 0.25 or higher. Assessors may approve people below 0.25 in exceptional circumstances, and some eligible people decline the subsidy. In a regression discontinuity analysis using local linear regression with a bandwidth of 0.05, the estimated probability of receiving the subsidy is 0.10 just below the threshold and 0.70 just above it, and the estimated mean number of hospital admissions in the following year is 0.42 just below and 0.36 just above. A histogram shows 1,840 assessments with scores from 0.230 to 0.249 and 2,410 with scores from 0.250 to 0.269, whereas the counts in neighbouring bins on both sides change by less than 5 percent from one bin to the next. (a) State whether the design is sharp or fuzzy and calculate the Wald estimate. (b) Explain to whom the estimate applies. (c) Interpret the histogram and describe two further checks you would run.

Model answer

(a) The design is fuzzy, because the probability of receiving the subsidy jumps from 0.10 to 0.70 at the threshold, a jump of 0.60, rather than from 0 to 1. The jump in admissions is 0.36 − 0.42 = −0.06. The Wald estimate is −0.06 / 0.60 = −0.10 admissions per person in the following year.

(b) The estimate applies to compliers near a frailty index of 0.25: older adults who receive the subsidy because their score reaches the threshold and would not receive it otherwise. It does not describe the effect for very frail adults far above the threshold, for people approved by exception regardless of score, or for people who decline.

(c) The histogram shows 2,410 assessments just above the threshold and 1,840 just below, about 31 percent more above (2,410 / 1,840 = 1.31), whereas neighbouring bins differ by less than 5 percent. This heaping suggests that some assessors nudge borderline scores up to 0.25 to secure the subsidy. The people pushed over the threshold may be needier or have stronger advocates, so the comparison at the threshold is no longer clean. I would run a formal density test (for example, the McCrary test or rddensity) and check whether age, prior admissions and living alone are continuous at the threshold. I would also estimate a donut model that excludes scores from 0.24 to 0.26 and examine whether heaping is concentrated among particular assessors.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What distinguishes a fuzzy regression discontinuity design from a sharp one?

In a fuzzy design, some people below the threshold receive the program and some above it do not, so the probability of receipt jumps by less than one. The Cedar Valley rule is fuzzy because clinician judgement allows referral below a score of 6.

Question 2: In the simulated Cedar Valley data, the probability of referral jumps by 0.4089 at a score of 6, and the mean six-month score drops by 0.1985 points. What is the Wald estimate of the effect of referral?

The Wald estimator divides the reduced form by the first stage: −0.1985 / 0.4089 = −0.485. Option b is the reduced form alone, option c multiplies the two jumps (−0.0812), and option d adds them (−0.6074).

Question 3: A histogram of screening scores shows 446 older adults at a score of 5 and 680 at a score of 6, while the counts at other values decline smoothly. What does this pattern most likely indicate?

An excess of observations just above the threshold, against an otherwise smooth decline, is the signature of sorting or manipulation, which density tests such as McCrary's are designed to detect. The size of the first stage (option a) cannot be judged from the distribution of the running variable.

Question 4: The Cedar Valley steering committee asks whether lowering the referral threshold from 6 to 5 is justified. Why is a regression discontinuity estimate well suited to this question?

A regression discontinuity estimate is local to the threshold, and a decision about moving the threshold by one point concerns exactly those people. The estimate does not describe all lonely older adults (option a), which is the main limitation of the design for broader funding decisions.
Section 3 of 5

Natural Experiments and Instrumental Variables

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Natural Experiments and Instrumental Variables

Using events that researchers do not control, and the assumptions that make them informative.

Definition

What makes an experiment natural

Events or interventions not under the control of researchers that divide a population into exposed and unexposed groups.Paraphrased from Craig et al., 2012

John Snow's comparison of two London water companies in 1854 remains the classic example of assignment that was as if random.

Canadian examples

Provinces as natural experiments

Public health insurance

Hanratty (1996) used its staggered introduction across provinces to study infant health.

Quebec child care

Baker, Gruber and Milligan (2008) compared Quebec with other provinces after 1997.

Minimum alcohol prices

Stockwell and colleagues (2012) analyzed price changes in British Columbia and Saskatchewan.

Mincome

Forget (2011) linked health records for Dauphin, Manitoba, decades after the experiment.

Instrumental variables

Four conditions

Relevance

The instrument changes the probability of receiving the program.

Independence

The instrument shares no causes with the outcome.

Exclusion restriction

The instrument affects the outcome only through the program.

Monotonicity

No one receives the program less often because of the instrument.

What the estimate means

Compliers and the local average treatment effect

Wald estimator, binary instrument
\[ \text{LATE} = \frac{E[Y \mid Z=1] - E[Y \mid Z=0]}{P(D=1 \mid Z=1) - P(D=1 \mid Z=0)} \]

With a first-stage F statistic below about 10, the instrument is weak and the estimate is imprecise and fragile.

Worked example

Connector capacity as an instrument

Illustrative Cedar Valley figures, year two
\[ \frac{5.6 - 5.9}{0.80 - 0.35} = \frac{-0.30}{0.45} = -0.67 \]

The main threats are seasonal or clinic-level patterns in capacity, and waiting-list contacts that affect loneliness directly.

Judging the evidence

Strengthening a natural experiment

  • The evaluator documents how exposure was assigned and by whom.
  • Negative control outcomes can reveal hidden confounding.
  • Triangulation compares designs with different sources of bias.
  • Linked administrative data need data access agreements and attention to First Nations data governance.
Carry forward

From evaluation to improvement

  • Natural experiments are only as strong as the evidence about their assignment process.
  • Instrumental variable estimates describe compliers and depend on unverifiable assumptions.
  • Section 4 turns to designs that programs use to improve themselves.

Learning Objectives for this section

  • Define a natural experiment and summarize the Medical Research Council guidance on using natural experiments to evaluate population health interventions.
  • Describe Canadian provincial policies that have been studied as natural experiments, and identify the design and the main threat in each.
  • State the assumptions of an instrumental variable analysis (relevance, independence, the exclusion restriction and monotonicity) and draw them on a causal diagram.
  • Calculate a Wald estimate from a binary instrument and interpret it as a local average treatment effect among compliers.
  • Assess a proposed instrument for a health program and identify the evidence that would strengthen or weaken it.

3.1 What Makes an Experiment Natural

Craig and colleagues (2012), writing the Medical Research Council's guidance on natural experiments, define them as events or interventions that are not under the control of researchers but that divide a population into exposed and unexposed groups whose outcomes can be compared. The term describes how exposure was assigned. The analytical methods are the same ones used in quasi-experimental evaluation: interrupted time series, difference-in-differences, regression discontinuity, synthetic control, matching and instrumental variables. What distinguishes a strong natural experiment is an assignment process that is unrelated, or plausibly unrelated, to the outcome. Dunning (2012) describes the ideal as assignment that is "as if random", and he argues that the credibility of a natural experiment rests mainly on the evidence that this condition holds.

The classic example is John Snow's study of cholera in London in 1854. Two water companies supplied houses on the same streets, and one of them, the Lambeth Company, had moved its intake upstream to a cleaner part of the Thames, while the Southwark and Vauxhall Company still drew water from a stretch contaminated by sewage. Households had generally not chosen their supplier, so the companies' customers were similar in most respects, and Snow (1855) could compare cholera deaths between them. The assignment was set by commercial decisions made years earlier, which is the feature evaluators look for.

The Medical Research Council guidance makes several points that apply directly to program evaluation. Natural experimental studies are most informative when the intervention is expected to have an effect large enough to detect, when good data on exposure and outcomes are available for large populations, and when the process that determines exposure is understood well enough to support a credible comparison. The guidance encourages evaluators to examine the assignment process carefully, to choose comparison groups and analyses that reduce bias from selective exposure, and to test assumptions with more than one method (Craig et al., 2012). A later review summarized the range of methods and their contributions to public health research (Craig et al., 2017), and the updated framework for complex interventions treats natural experimental evaluation as one route to evaluating interventions that researchers do not control (Skivington et al., 2021).

For the Cedar Valley Connector program, the distinction between a planned evaluation and a natural experiment is a matter of degree. The health authority planned the program, but the evaluation team did not control which clinics went first; Lesson 7 established that the 12 wave-one clinics were chosen for readiness. The region-wide emergency department series in Section 1 and the referral threshold in Section 2 are features of the program that the evaluation uses opportunistically, which is the attitude a natural experimental study takes.

3.2 Natural Experiments in Canadian Provinces

Canada's provinces make health and social policy separately and at different times, which creates many natural experiments. The examples below show how evaluators have used provincial differences and the threats that accompany each design.

Saskatchewan introduced public hospital insurance in 1947 and public medical care insurance in 1962, and the other provinces followed under federal cost-sharing legislation over the following years. Hanratty (1996) used the staggered introduction of national health insurance across provinces to estimate its effects on infant health. The design is a difference-in-differences comparison across provinces and years, and its main threat is that provinces adopting insurance earlier may have differed in other trends affecting infant health.

In 1997, Quebec began introducing low-fee child care, while other provinces did not. Baker, Gruber and Milligan (2008) compared families in Quebec with families in the rest of Canada before and after the policy, and reported increases in mothers' employment together with less favourable measures on some child and family outcomes. The design is difference-in-differences, and its main threat is other Quebec-specific changes over the same period.

British Columbia and Saskatchewan set minimum prices for alcoholic beverages and adjusted them at known dates. Stockwell and colleagues (2012) used time-series analyses of these changes and associated increases in minimum prices with reductions in alcohol consumption. The design is an interrupted time series with variation in the size and timing of price changes, and its main threats are co-occurring changes in alcohol policy and in cross-border purchasing.

From 1974 to 1979, the Mincome experiment offered a guaranteed annual income to residents of Dauphin, Manitoba. Decades later, Forget (2011) linked administrative health records to compare Dauphin residents with matched residents of comparison communities and reported fewer hospitalizations during the experiment. The original program was planned as a social experiment, but the health analysis was a natural experimental study, and its main threat is that Dauphin differed from its comparison communities in ways that matching did not capture.

Each of these studies depends on administrative data that cover whole populations over long periods. In British Columbia, Population Data BC provides researchers with access to linked health and social data under data access agreements, which makes provincial policy changes and program rollouts open to similar analysis. Section 2 described one more provincial example, the minimum legal drinking age, which Callaghan and colleagues (2013) studied with a regression discontinuity design.

3.3 The Logic of Instrumental Variables

Notation: in this section Z denotes the instrument and D denotes receipt of the program, which corresponds to the exposure X in the causal diagrams used across the course series. Elsewhere in the series, Z denotes a collider.

The designs in Lesson 7 and in Sections 1 and 2 of this lesson deal with confounding by comparing groups or periods that should be similar. An instrumental variable analysis takes a different route. It looks for a variable, the instrument, that changes the probability of receiving the program but has no other connection to the outcome. Variation in the program that is driven by the instrument is free of the confounding that affects voluntary participation, so the association between the instrument and the outcome, scaled by the association between the instrument and the program, estimates the program's effect.

Instrument (Z) e.g., slot open Program (D) started meetings Outcome (Y) loneliness Unmeasured (U) motivation, frailty No direct path from Z to Y (exclusion restriction) No shared causes of Z and Y (independence) Relevance
An instrument must affect program receipt (relevance), share no causes with the outcome (independence), and affect the outcome only through the program (exclusion); the dashed red paths must be absent. The fourth condition, monotonicity, concerns how individuals respond to the instrument and cannot be drawn on the diagram.

Three core conditions define an instrument (Hernán and Robins, 2006). Relevance requires that the instrument be associated with receipt of the program. Independence, also called exchangeability, requires that the instrument share no common causes with the outcome, as it would if it were randomly assigned. The exclusion restriction requires that the instrument affect the outcome only through the program. Only relevance can be checked directly in the data; the other two rest on knowledge of how the instrument arose, supported by indirect checks. A fourth condition is needed to say whose effect is being estimated, and it is described below.

The Wald estimator for a binary instrument

Effect = (mean outcome when Z = 1 − mean outcome when Z = 0) / (proportion receiving the program when Z = 1 − proportion receiving the program when Z = 0)

The numerator is the effect of the instrument on the outcome (the reduced form), and the denominator is the effect of the instrument on program receipt (the first stage). With covariates or a continuous instrument, the same logic is implemented as two-stage least squares: the first stage regresses program receipt on the instrument and covariates, and the second stage regresses the outcome on the predicted receipt. Standard errors must come from software that accounts for both stages.

Compliers and the local average treatment effect

Imbens and Angrist (1994) and Angrist, Imbens and Rubin (1996) showed what an instrumental variable estimate means when program effects differ between people. With a binary instrument and a binary program, each person belongs to one of four groups, defined by how their program receipt would respond to the instrument. Under the fourth condition, monotonicity, which states that there are no defiers, the Wald estimator equals the average effect among compliers. This quantity is the local average treatment effect. It is the effect for the people whose participation the instrument actually changes, and it may differ from the average effect in the whole population.

CompliersClick to explore
Always-takersClick to explore
Never-takersClick to explore
DefiersClick to explore

Weak instruments create two problems. When the first stage is small, the denominator of the Wald estimator is close to zero, so the estimate becomes very imprecise, and in finite samples it is biased toward the confounded ordinary regression estimate. A small violation of the exclusion restriction is also magnified, because any direct effect of the instrument on the outcome is divided by a small number. A first-stage F statistic below about 10 is a common warning sign (Staiger and Stock, 1997). Hernán and Robins (2006) concluded that instrumental variable methods can be valuable but that their key assumptions are unverifiable, and that an instrumental variable estimate can be more biased than a conventional estimate when the assumptions are even slightly violated.

3.4 Instruments in Health Services Evaluation

Instruments in health services research usually come from features of the care system that steer people toward or away from a treatment for reasons unrelated to their own prognosis. The table summarizes common types and the assumption most likely to fail in each.

InstrumentExampleAssumption most at risk
Random assignment with incomplete uptakeRandomization to a program arm, used in Lesson 6 to estimate a complier average causal effectExclusion, if assignment changes behaviour apart from uptake
Distance to a facilityDifferential distance to hospitals offering cardiac catheterization (McClellan, McNeil and Newhouse, 1994)Independence, because people who live far from facilities differ in income, rurality and health
Provider preferenceA physician's prescribing preference for one drug over another (Brookhart et al., 2006)Independence and exclusion, if preferred providers also differ in other aspects of care
Threshold rulesBeing above an eligibility score in a fuzzy regression discontinuity design (Section 2)Exclusion, if the threshold also triggers other services
Genetic variantsMendelian randomization, in which inherited variants act as instruments for an exposure (Davey Smith and Ebrahim, 2003)Exclusion, if a variant affects the outcome through other pathways

3.5 Worked Example: Connector Capacity as an Instrument

Case: full caseloads in the program's second year

In the second year of the fictional Cedar Valley Connector program, after the wave-two clinics joined, connectors' caseloads were sometimes full. When a referral arrived and no slot was open, the older adult was placed on a waiting list, and many never started. Whether a slot was open depended on when earlier participants completed their twelve weeks, which the evaluator argues is unrelated to the characteristics of the person being referred. Of 1,240 older adults referred in year two, 860 were referred when a slot was open and 380 when caseloads were full. The figures in this example are illustrative.

Among those referred when a slot was open, 80 percent started connector meetings within four weeks, compared with 35 percent of those referred when caseloads were full. The mean six-month three-item UCLA score was 5.6 in the first group and 5.9 in the second. The Wald estimate is (5.6 − 5.9) / (0.80 − 0.35) = −0.30 / 0.45 = −0.67 points. Under the instrumental variable assumptions, starting connector meetings promptly lowered the six-month score by about two thirds of a point among compliers, the older adults who started only because a slot was open.

The evaluator then examines each assumption. Relevance is strong, because the instrument changes the proportion starting by 45 percentage points. Independence is threatened if caseloads fill at particular times or places that also affect loneliness: caseloads may be fuller in winter, when loneliness and illness are more common, or at clinics serving more isolated neighbourhoods. The analysis should therefore compare referrals within the same clinic and season and check whether baseline scores, age and living arrangements are balanced between referrals made when slots were open and when caseloads were full. The exclusion restriction is threatened if waiting-list status affects loneliness through other paths, for example if waitlisted older adults received a welcome call and a list of community groups, or if clinicians referred them to another service. Monotonicity is plausible, because an open slot should never make an older adult less likely to start.

This estimate answers a different question from the regression discontinuity estimate in Section 2. The fuzzy regression discontinuity estimates the effect of referral for older adults at the margin of eligibility, whereas the capacity instrument estimates the effect of starting promptly among older adults who were already referred. The two complier groups differ, and the estimates need not agree. Reporting both, with their assumptions, gives decision-makers a fuller picture than either alone.

3.6 Judging a Natural Experiment

Because the evaluator does not control assignment in a natural experiment, the strength of the evidence depends on how well alternative explanations have been examined. The questions below organize that assessment. Lesson 2 Section 3.5 places this kind of design within the MRC framework for developing and evaluating complex interventions, whose evaluation phase includes natural experimental designs for interventions that cannot be randomized.

How was exposure assigned?v

The evaluator should document who decided which people, places or periods were exposed, on what basis and when. Assignment by readiness, need or political priority is more likely to be related to outcomes than assignment by timing, capacity or administrative accident.

Is there a credible comparison?v

The comparison group or period should be exposed to the same co-occurring events as the exposed group. Pre-period trends, covariate balance and the density of the running variable provide evidence on comparability.

Have falsification tests been run?v

Negative control outcomes, which the intervention should not affect, and negative control exposures, which should not affect the outcome, can reveal confounding (Lipsitch, Tchetgen Tchetgen and Cohen, 2010). For Cedar Valley, emergency visits for fractures could serve as a negative control outcome for the connector program.

Do different designs agree?v

Triangulation compares results from approaches with different sources of bias (Lawlor, Tilling and Davey Smith, 2016). If a controlled interrupted time series, a difference-in-differences comparison of clinics and a regression discontinuity estimate point in the same direction, a common bias is less likely to explain all three.

Are the data and approvals in place?v

Natural experimental studies usually rely on linked administrative data, which require data access agreements, privacy review and, where the work is research, ethics review under TCPS 2. Indigenous data in British Columbia also require attention to First Nations data governance, as Lesson 4 discussed.

Try it: assess a proposed instrument

A colleague proposes using the straight-line distance from each older adult's home to the nearest wave-one clinic as an instrument for program participation, with emergency department visits in the following year as the outcome. Write a short paragraph assessing relevance, independence, the exclusion restriction and monotonicity for this instrument, and name one check that could weaken or strengthen the case for it.

Reflection

An evaluator wants to estimate the effect of attending a community connector program on emergency department visits among adults aged 65 and older in a health region with towns and rural areas. Attendance is voluntary, and attenders differ from non-attenders in health and motivation. The evaluator proposes living within 5 km of a connector office as an instrument. In linked records, 34 percent of older adults living within 5 km attended, compared with 18 percent of those living farther away, and mean emergency department visits in the following year were 0.48 per person within 5 km and 0.52 farther away. The first-stage F statistic is 45. Most residents living farther than 5 km live in rural communities, and the region's two emergency departments are in its two largest towns, where the connector offices are also located. The four instrumental variable conditions are relevance (the instrument is associated with attendance), independence (the instrument shares no causes with the outcome), the exclusion restriction (the instrument affects the outcome only through attendance) and monotonicity (no one attends less because they live closer). (a) Calculate the Wald estimate. (b) Assess each of the four conditions for this instrument. (c) Propose two changes that would make the analysis more credible, and state whom the estimate would describe.

Model answer

(a) The Wald estimate is (0.48 − 0.52) / (0.34 − 0.18) = −0.04 / 0.16 = −0.25 emergency visits per person per year among compliers.

(b) Relevance is satisfied: attendance differs by 16 percentage points and the first-stage F of 45 is well above the usual warning level of 10. Independence is doubtful, because people within 5 km live mainly in towns and differ from rural residents in age, income, health and access to services. The exclusion restriction is also doubtful, because living near a connector office means living near an emergency department, and proximity to an emergency department affects its use directly. Monotonicity is plausible, because living closer should not make anyone less likely to attend, although a few people might avoid a program where neighbours could see them.

(c) First, I would adjust for rurality and for distance to the nearest emergency department, or compare older adults within the same community type, so that the instrument varies among people with similar access to emergency care. Second, I would test the instrument against a negative control outcome, such as emergency visits for fractures, which attending a connector program should not change; an association would signal a violation. The estimate would describe compliers: older adults who attend because they live close and would not attend if they lived farther away, who may be less mobile or lack transportation.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: According to the Medical Research Council guidance (Craig et al., 2012), what defines a natural experiment?

The guidance defines natural experiments by how exposure was assigned: by events or interventions that researchers do not control. Natural experiments often use administrative data (option b), but the data source does not define them.

Question 2: Which instrumental variable assumption can be checked directly in the data?

Relevance, the association between the instrument and program receipt, can be measured with the first stage and its F statistic. Independence and the exclusion restriction depend on how the instrument arose and can only be supported indirectly, for example with balance checks and negative control outcomes.

Question 3: In the Cedar Valley capacity example, 80 percent of older adults referred when a slot was open started meetings within four weeks, compared with 35 percent when caseloads were full, and mean six-month scores were 5.6 and 5.9. What does the Wald estimate of −0.67 describe?

Under monotonicity, the Wald estimate is the local average treatment effect among compliers, the people whose participation the instrument changed. The comparison of attenders with non-attenders (option d) is the confounded comparison that the instrument is meant to avoid.

Question 4: Distance to a facility is often proposed as an instrument. Which assumption is it most likely to violate, and why?

Where people live is related to many determinants of health, so distance may share causes with the outcome. Distance is usually strongly associated with use (option b is wrong), and the timing of measurement (option d) does not bear on the exclusion restriction.
Section 4 of 5

Quality Improvement Designs

⏱ Estimated reading time: 45 minutes
Section 4 of 5

Quality Improvement Designs

Plan-Do-Study-Act cycles, run charts, control charts and the design justification.

Purpose

Quality improvement and evaluation

Improvement

Teams make rapid, data-guided changes to their own local processes.

Evaluation

Evaluators judge the merit, worth or significance of a program.

Research

Researchers seek knowledge that generalizes beyond the setting.

The Model for Improvement asks what we are trying to accomplish, how we will know a change is an improvement, and what change we can make.

Plan-Do-Study-Act

Small tests, linked cycles

  • Each cycle states a prediction before the test begins.
  • Tests start small and grow as confidence builds.
  • Each cycle builds on what the previous one learned.
  • Data are collected over time, together with outcome, process and balancing measures.
Run charts

Four rules for non-random change

Shift

Six or more consecutive points fall on one side of the median.

Trend

Five or more consecutive points all rise or all fall.

Runs

There are too few or too many runs for the number of points.

Astronomical point

One value is obviously different from all the others.

Control charts

Common-cause and special-cause variation

p-chart limits
\[ \bar{p} \pm 3\sqrt{\frac{\bar{p}(1-\bar{p})}{n_i}} \;=\; 0.54 \pm 3\sqrt{\frac{0.54 \times 0.46}{52}} \;=\; 0.333 \text{ to } 0.747 \]

A month at 0.70 lies within the limits; a month at 0.75 would signal special-cause variation.

Reporting

SQUIRE 2.0

  • SQUIRE 2.0 has 18 items that follow the structure of a scientific report.
  • The rationale item links the work to a theory of change.
  • The study of the intervention names the design used to judge whether changes caused the results.
  • Reports describe outcome, process and balancing measures and the context of the work.
Choosing a design

Writing the justification

  • The justification states the question, the causal contrast and how the program was assigned.
  • It names the design and the alternatives rejected, with reasons.
  • It states each assumption in plain language and pairs it with a planned check.
  • It explains what the evaluator will conclude if a check fails.
Carry forward

Bringing the lesson together

  • Every design in Lessons 6 to 8 trades one assumption for another.
  • Combining designs with different weaknesses strengthens a conclusion.
  • The final assessment asks you to apply these ideas to a new program.

Learning Objectives for this section

  • Describe the Model for Improvement and the Plan-Do-Study-Act cycle, and identify the features that make a series of cycles informative.
  • Construct and interpret a run chart using the four run chart rules.
  • Distinguish common-cause from special-cause variation, choose an appropriate statistical process control chart, and calculate control limits for a p-chart.
  • Identify the reporting elements of SQUIRE 2.0.
  • Compare quality improvement designs with the evaluation designs of Lessons 6 to 8, and write a design justification that names its assumptions and the checks for them.

4.1 Quality Improvement and Evaluation

Quality improvement has been defined as systematic, data-guided activities designed to bring about immediate, positive change in the delivery of health care in particular settings (Lynn et al., 2007). Its purpose is local and practical: a team wants its own process to work better, soon. Evaluation, as Lesson 1 defined it, determines the merit, worth or significance of a program, and research seeks knowledge that generalizes beyond the setting. The three overlap in methods, and many programs need all of them. TCPS 2 Article 2.5, discussed in Lesson 1, recognizes that quality improvement and program evaluation activities do not require research ethics board review unless they are conducted for research purposes, although they still raise ethical questions that teams must address.

Programs such as the Cedar Valley Connector program use quality improvement designs to improve their own delivery, and an evaluator needs to understand what such data can show. The designs share the logic of Section 1, because they track a measure over time and look for change after a deliberate intervention, although their standards of evidence differ, as Section 4.6 explains.

4.2 The Model for Improvement and Plan-Do-Study-Act Cycles

The Model for Improvement (Langley et al., 2009) asks three questions: what are we trying to accomplish, how will we know that a change is an improvement, and what change can we make that will result in improvement. The answers become an aim, a set of measures and a set of change ideas. Change ideas are then tested in Plan-Do-Study-Act (PDSA) cycles. The cycle descends from Walter Shewhart's work on industrial quality in the 1920s and 1930s and was developed by W. Edwards Deming, who preferred "Study" to "Check" to emphasize learning from the data. In the Plan stage the team states the change, a prediction and the data it will collect; in Do it carries out the test on a small scale; in Study it compares the results with the prediction; and in Act it decides whether to adopt, adapt or abandon the change.

The Model for Improvement uses three kinds of measures. Outcome measures track the aim, process measures track whether the changes are being carried out, and balancing measures watch for unintended effects elsewhere in the system, such as staff overtime or longer waits for other services.

PDSA cycles are simple to describe and difficult to do well. Taylor and colleagues (2014) reviewed published applications in health care and found that few reports showed the method's key features: a sequence of linked cycles, an explicit prediction for each test, small-scale testing before wider implementation, and the use of data collected over time. Reed and Card (2016) argued that the apparent simplicity of the method hides the skill, resources and organizational support it requires, and that cycles are often used as single projects with little connection to a theory of change. An evaluator reviewing a program's improvement work therefore looks for documented predictions, linked cycles and time-series data, which together distinguish genuine learning from a record of activity.

Case: faster first meetings in Cedar Valley

The fictional Cedar Valley Connector program's coordinator notices that many referred older adults wait more than two weeks for a first meeting, and that some lose interest while they wait. The improvement aim is to raise the percentage of referred older adults who have a first connector meeting within 14 days from about 54 percent to 70 percent within six months. The outcome measure is the weekly percentage of new referrals with a first meeting within 14 days, the process measure is the percentage of new referrals telephoned within two business days, and the balancing measure is connector overtime hours. In the first PDSA cycle, one connector at one clinic telephones every new referral within two business days for two weeks, predicting that 8 of 10 will be reached. Six of nine are reached, and several older adults say they would prefer a text message first. In the second cycle, two connectors offer a text message before the call. Later cycles extend the change to all wave-one clinics from week 11.

4.3 Run Charts

A run chart plots a measure over time in the order the data were collected, with the median as a centre line and the timing of changes annotated on the chart. It is the basic tool for judging whether a process has changed, because it distinguishes patterns that are unlikely to arise by chance from the ordinary ups and downs of a stable process. Perla, Provost and Murray (2011) describe four rules for detecting non-random patterns. A shift is six or more consecutive points all above or all below the median, with points that fall on the median neither breaking nor adding to the run. A trend is five or more consecutive points all going up or all going down. Too few or too many runs, judged against a published table for the number of points, indicate non-random clustering or oscillation. An astronomical data point is a value that is obviously different from all others and that people familiar with the process would recognize as unusual. When a change is tested after a stable baseline, the median is often calculated from the baseline points and extended across the chart, so that new points are judged against the earlier centre.

50% 60% 70% 1 5 10 15 20 Week Baseline median (54%), extended Change spread to all clinics Shift: 10 points above the median
Illustrative run chart for Cedar Valley: weekly percentage of new referrals with a first meeting within 14 days, before (teal) and after (red) the telephone-and-text change was spread to all wave-one clinics.

In the illustrative data, the weekly percentages in the ten baseline weeks were 52, 58, 49, 55, 61, 50, 57, 53, 48 and 56, with a median of 54. After the change was spread from week 11, the percentages were 60, 63, 59, 66, 68, 64, 70, 67, 72 and 69. All ten post-change points lie above the baseline median, which meets the shift rule. No five consecutive points rise or fall, so the trend rule is not met, and no point is astronomical. The shift is a non-random signal that coincides with the change, and the team would reasonably conclude that the process improved. The run chart cannot show, by itself, that the change caused the improvement. A new coordinator, a seasonal fall in referrals or a change in how meetings were recorded could produce the same signal, which is the history threat of Section 1 in a new setting.

4.4 Statistical Process Control Charts

Shewhart (1931) distinguished two kinds of variation. Common-cause variation is the ordinary variation built into a stable process, arising from many small sources that cannot be separated. Special-cause variation comes from specific, identifiable sources that are not part of the usual process. A process that shows only common-cause variation is said to be in statistical control, and its future performance is predictable within limits. A statistical process control chart, or control chart, plots the measure over time with a centre line (usually the mean) and upper and lower control limits, conventionally set three standard deviations from the centre line, with the standard deviation calculated from the distribution that suits the data (Benneyan, Lloyd and Plsek, 2003; Mohammed, Worthington and Woodall, 2008).

Common rule sets for detecting special-cause variation include a single point outside the control limits, eight consecutive points on one side of the centre line, six consecutive points increasing or decreasing, and two of three consecutive points beyond two standard deviations on the same side. Teams should agree on one set of rules in advance, because applying many rules increases false signals. The choice of chart depends on the type of data.

ChartType of dataCedar Valley example
p-chartProportions with varying subgroup sizesMonthly proportion of referrals with a first meeting within 14 days
u-chartRates per unit of exposure with varying denominatorsMonthly emergency visits per 1,000 residents aged 65 and older
c-chartCounts with a constant area of opportunityWeekly number of new referrals at one clinic
Individuals chart (XmR)One continuous value per period, with limits from moving rangesMonthly median days from referral to first meeting

Control limits for a p-chart

Centre line = p̄, the overall proportion in the baseline period. Limits for subgroup i = p̄ ± 3 × √(p̄ × (1 − p̄) / ni), where ni is the number of referrals in that subgroup.

With p̄ = 0.54 and a month with 52 referrals, the standard deviation is √(0.54 × 0.46 / 52) = 0.0691, so the limits are 0.54 ± 0.207, or 0.333 to 0.747. A month in which 75 percent of 52 referrals had a first meeting within 14 days would lie above the upper limit and signal special-cause variation, whereas a month at 70 percent would lie within the limits and would need a run of points to signal a change. Months with fewer referrals have wider limits.

Control charts and interrupted time series use the same kind of data for different purposes: a control chart flags signals quickly for managers to investigate, whereas an interrupted time series estimates the size of an effect against a modelled counterfactual. Both are affected by the issues raised in Section 1. Standard control limits assume independent observations, so positive autocorrelation produces more false signals, and a strong seasonal pattern can create apparent special causes in winter. A u-chart of Cedar Valley's monthly emergency visit rate, for example, would flag winter peaks as special causes unless the seasonal pattern were modelled first.

4.5 Reporting with SQUIRE 2.0

The Standards for QUality Improvement Reporting Excellence, SQUIRE 2.0 (Ogrinc et al., 2016), provide an 18-item guideline for reporting systematic efforts to improve the quality, safety and value of health care. The items follow the structure of a scientific report: title and abstract; an introduction covering the problem, available knowledge, rationale and specific aims; methods covering context, the intervention or interventions, the study of the intervention, measures, analysis and ethical considerations; results; a discussion with a summary, interpretation, limitations and conclusions; and funding.

Four items deserve particular attention from evaluators.

Rationalev

Authors state the informal or formal frameworks, models or theories used to explain the problem and the expected effect of the intervention, which links improvement work to the program theory of Lesson 3.

Contextv

Authors describe the elements of the setting considered important at the outset, such as staffing, leadership support and existing services, so that readers can judge where the results might transfer.

Study of the interventionv

SQUIRE separates the intervention from the approach chosen to assess whether the observed outcomes were due to it. This is where the team names its design, whether a run chart, a control chart, an interrupted time series or a controlled comparison.

Measures and analysisv

Authors report their outcome, process and balancing measures and the methods used to draw inferences, such as the run chart rules applied. A Cedar Valley report on first meetings would also list the co-occurring changes the team considered.

4.6 Quality Improvement Designs and the Evaluation Designs of Lessons 6 to 8

Lessons 6 to 8 have introduced a set of designs that differ in the question they answer, the assumption they rely on and the data they need. The table compares them, using the Cedar Valley program as a common reference.

DesignQuestion answeredKey assumptionCedar Valley use
Cluster or stepped-wedge randomized trial (Lesson 6)Effect of offering the program, averaged over clinicsRandomization was carried out and maintainedHypothetical randomized order of clinic waves
Difference-in-differences (Lesson 7)Effect in program clinics over a periodParallel trends without the programWave-one compared with wave-two clinics in year one
Propensity score methods (Lesson 7)Effect for participants compared with similar non-participantsNo unmeasured confoundingReferred older adults matched to similar older adults at wave-two clinics
Synthetic control (Lesson 7)Effect for one treated unitA weighted combination of donors reproduces the pre-periodRegion-wide rates compared with a weighted set of other regions
Interrupted time series (Lesson 8)Change in level and slope after launchThe pre-launch trend would have continued; a comparison series shares co-occurring eventsMonthly emergency visit rates with Juniper Ridge
Regression discontinuity (Lesson 8)Effect at the eligibility thresholdContinuity at the threshold and no manipulationThe UCLA score of 6 referral rule
Instrumental variable (Lesson 8)Effect among compliersRelevance, independence, exclusion and monotonicityConnector capacity at the time of referral
Run chart or control chart with PDSA cycles (Lesson 8)Whether a local process changed after a tested changeA stable baseline and no co-occurring changeFirst meetings within 14 days

Quality improvement designs place speed and local learning ahead of causal certainty. A run chart with a shift after a tested change is persuasive to the team that made the change, and it is often enough for a local decision, but it leaves history and instrumentation threats largely unaddressed. Improvement teams can strengthen their designs with the methods of this lesson: by annotating co-occurring events, by spreading a change across sites at staggered times so that each site's start serves as a replication (a logic close to the stepped-wedge design of Lesson 6), by adding a comparison site or a balancing measure, and by analyzing a long series with segmented regression. Evaluators, in turn, can use a program's improvement data as evidence of implementation and as process measures within a larger design.

4.7 Choosing and Justifying a Design

No design is best in general. The choice follows from the evaluation question and the decision it informs (Lesson 4), from the way the program was assigned to people or places, from the data that exist or can be collected, from the assumptions that are plausible in the setting, and from ethical and practical constraints. A design justification sets out this reasoning so that readers can judge it. A strong justification states the evaluation question and the causal contrast it requires, describes how the program was or will be assigned, names the chosen design and the alternatives considered with reasons for rejecting them, states the design's assumptions in plain language, specifies in advance the checks for each assumption, explains what the evaluator will conclude if a check fails, and identifies complementary designs or data that address the main remaining threats.

4.8 Worked Example: A Design Justification for Cedar Valley

The following excerpt shows the core of a design justification for the fictional Cedar Valley Connector program, written for the steering committee.

Excerpt: Cedar Valley design justification

The evaluation addresses three questions agreed with the steering committee: whether the program reduces loneliness among referred older adults, whether it reduces emergency department visits among older adults in the region, and whether lowering the referral threshold from a UCLA score of 6 to 5 is justified. Because the health authority chose the 12 wave-one clinics for their readiness, the clinics were not randomized, and comparisons between waves must address selection.

For loneliness, the primary design is a difference-in-differences comparison of the change between screening and follow-up among older adults screened at wave-one and wave-two clinics during the program's first year, combined with the matching on baseline characteristics developed in Lesson 7. Its key assumption is that, without the program, loneliness and screening patterns would have followed parallel trends in the two groups of clinics. Because loneliness screening began only at launch, pre-launch trends in loneliness cannot be observed, so we will examine trends in emergency and primary care visits among screened older adults for the 24 months before launch, from linked records, as an indirect check, and present an event-study plot for those outcomes.

For emergency department visits, the design is a controlled interrupted time series using 48 months before and 24 months after launch, with Juniper Ridge as the comparison region and emergency visits for fractures as a negative control outcome. Its assumptions are that the difference between the regions would have continued along its pre-launch path and that co-occurring events affected both regions. Because the program reaches a small share of older adults, we will report a plausibility check on the size of any region-wide effect and analyze visits among referred older adults through linked records.

For the threshold question, a fuzzy regression discontinuity at a score of 6 will estimate the effect of referral for older adults at the margin of eligibility. We will check the distribution of scores at 5 and 6, the continuity of age and prior emergency visits, placebo thresholds and results in two windows, and we will interpret the estimate cautiously because the score has only seven values. Process improvement data on first meetings within 14 days will be tracked on a p-chart and reported as implementation evidence.

AssumptionPlanned checkIf the check fails
Parallel trends between wave-one and wave-two clinicsPre-launch trends and an event-study plotReport the difference-in-differences estimate as descriptive and give more weight to the other designs
Juniper Ridge shares co-occurring eventsPre-launch trend in the difference series, a list of regional events, and the negative control outcomeSeek a second comparison series, such as residents aged 50 to 64
No manipulation of screening scoresFrequencies at scores 5 and 6, and clinician-level patternsUse a donut analysis that excludes scores of 6, or drop the threshold analysis
Effect size plausible given reachDilution arithmetic using participant countsTreat a region-wide effect as unexplained and investigate co-interventions

What a design justification contains

A design justification states the evaluation questions the design must answer and the causal contrast each requires, describes how the program was or will be assigned, names the chosen design (randomized, quasi-experimental, natural experimental, a quality improvement design, or a combination), and explains why at least one alternative was rejected. It lists each key assumption in plain language with the check that will be run for it, states what the evaluator will conclude if a check fails, and summarizes the assumptions and checks in a short table, as the Cedar Valley excerpt above does.

A sound justification lets the design follow from the evaluation questions, the assignment process and the available data, and it matches each assumption to a specific, feasible check planned in advance. It addresses the main threats to validity, including history and selection, explains how complementary data or designs reduce them, and, where relevant, addresses the number of time points, the bandwidth or the strength of an instrument. It is written clearly enough for a steering committee to act on.

Reflection

A clinic team tested a change in which connectors telephone each referred older adult within two business days. The team recorded the weekly percentage of referrals reached within two business days for eight weeks before the change (41, 47, 44, 39, 50, 45, 42 and 48) and for eight weeks after it began in week 9 (52, 49, 55, 47, 58, 56, 60 and 57). The run chart rules are as follows: a shift is six or more consecutive points all above or all below the median; a trend is five or more consecutive points all increasing or all decreasing; too few or too many runs indicate a non-random pattern; and an astronomical point is an obviously unusual value. SQUIRE 2.0 is an 18-item guideline for reporting quality improvement work, with items that include the rationale, the context, the intervention, the study of the intervention, measures, analysis, ethical considerations and limitations. (a) Calculate the baseline median and apply the shift and trend rules, extending the baseline median across all 16 weeks. (b) State what the team can and cannot conclude. (c) Name two elements that a SQUIRE 2.0 report of this work should contain and explain why each matters here.

Model answer

(a) The sorted baseline values are 39, 41, 42, 44, 45, 47, 48 and 50, so the median is (44 + 45) / 2 = 44.5. Week 8 (48) and all eight post-change weeks (52, 49, 55, 47, 58, 56, 60 and 57) lie above 44.5, giving nine consecutive points above the median, which meets the shift rule. No five consecutive points rise or fall, because the post-change values alternate between rises and falls, so the trend rule is not met, and no value is astronomical.

(b) The team can conclude that the process changed in a non-random way at about the time of the change. It cannot conclude from the run chart alone that the telephone change caused the improvement. The run begins in week 8, one week before the change, which suggests checking whether staff started early or whether another event occurred, such as new staff or a drop in referrals. Annotating co-occurring events, tracking a balancing measure such as connector overtime, and replicating the change at a second clinic would strengthen the conclusion.

(c) The report should describe the study of the intervention, naming the run chart rules and the baseline median used, because readers need to know how the team judged improvement. It should also describe the context, including staffing and referral volumes, because the change may depend on connector capacity and may not transfer to busier clinics.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: On a run chart with a baseline median of 54 percent, which pattern meets the shift rule described by Perla, Provost and Murray (2011)?

A shift is six or more consecutive points on the same side of the median. Five consecutive rising points (option b) meet the trend rule, and a single extreme value (option c) is an astronomical point.

Question 2: For a p-chart with a centre line of 0.54 and a month with 52 referrals, the control limits are 0.333 to 0.747. A month at 0.70 is observed. What is the appropriate interpretation?

A single point within the limits does not signal special-cause variation; a change would need to be shown by a rule such as a run of points on one side of the centre line. One month at 70 percent also cannot show that the aim is sustained (option b).

Question 3: What does SQUIRE 2.0 mean by the study of the intervention?

SQUIRE 2.0 separates the intervention itself (option a) from the study of the intervention, which is the design used to judge whether the intervention caused the observed changes, such as a run chart or an interrupted time series.

Question 4: Which statement best compares a run chart used in PDSA work with a controlled interrupted time series?

Run charts are built for rapid local learning and do not model a counterfactual or include a comparison series, so co-occurring events remain possible explanations. A controlled interrupted time series estimates an effect and uses a comparison series to address history.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson completed the course's sequence on designs for causal claims about programs. Section 1 showed how an interrupted time series uses a long run of observations to project what would have happened without a program, how segmented regression estimates level and slope changes, why seasonality and autocorrelation must be handled, and how a comparison series addresses co-occurring events. The worked example applied these ideas to Cedar Valley's simulated emergency department series and showed that a shared drop in a neighbouring region reduced the estimated effect by about a third.

Section 2 turned to eligibility thresholds and explained how a regression discontinuity design compares people just above and just below a cut-off, why the Cedar Valley referral rule produces a fuzzy design, and how the Wald estimator recovers a local effect for compliers. Section 3 placed these designs within the wider idea of natural experiments, using the Medical Research Council guidance and Canadian provincial examples, and set out the assumptions of instrumental variable analysis. Section 4 described quality improvement designs, from Plan-Do-Study-Act cycles to run charts and control charts, and compared them with every design in Lessons 6 to 8 before showing how to write a design justification.

Key Takeaways from this lesson

  • An interrupted time series uses the pre-intervention level and trend to project a counterfactual, which controls for secular trends and maturation but leaves history, instrumentation and changes in population composition as threats.
  • Segmented regression estimates a level change and a slope change, and the effect at a given month after launch combines both coefficients, so its confidence interval requires their covariance.
  • The expected shape of an effect should be specified in advance from the program theory, because a shape chosen after examining the data can fit random fluctuation.
  • Seasonality can be modelled with calendar-month indicators or Fourier terms, and positive autocorrelation makes ordinary standard errors too small, which Newey-West standard errors or autoregressive models correct.
  • A controlled interrupted time series adds a comparison series that shares co-occurring events, and its key assumption is that the difference between the series would have continued along its pre-launch path.
  • A region-wide outcome can dilute the effect of a program that reaches few people, so evaluators should check whether an estimated effect is plausible given the program's reach.
  • A regression discontinuity design estimates the effect of a program at an eligibility threshold, and in a fuzzy design the Wald estimator divides the jump in the outcome by the jump in the probability of receiving the program.
  • Regression discontinuity analyses need checks for manipulation of the running variable, covariate continuity, placebo thresholds and sensitivity to bandwidth, and discrete running variables require extra caution.
  • An instrumental variable must be relevant, independent of the outcome's causes and related to the outcome only through the program, and under monotonicity it estimates a local average treatment effect among compliers.
  • Run charts and control charts support rapid local learning in quality improvement, but they leave history threats largely unaddressed unless they are combined with comparisons, replication or segmented regression.

Core Concepts Reviewed

Section 1: interrupted time series, segmented regression, level and slope changes, impact models, seasonality and Fourier terms, autocorrelation, the Durbin-Watson test, Newey-West standard errors, the number of time points, dilution and the controlled interrupted time series.

Section 2: the running variable and threshold, continuity, sharp and fuzzy designs, the first stage, the reduced form, the Wald estimator, local linear regression, bandwidth choice, discrete running variables, manipulation and density tests, covariate and placebo checks, and the local interpretation of effects.

Section 3: natural experiments and the Medical Research Council guidance, as-if random assignment, Canadian provincial examples, instrumental variable assumptions, two-stage least squares, compliers and the local average treatment effect, weak instruments and falsification tests.

Section 4: the Model for Improvement, Plan-Do-Study-Act cycles, outcome, process and balancing measures, run chart rules, common-cause and special-cause variation, control charts and p-chart limits, SQUIRE 2.0, and the comparison and justification of evaluation designs.

The final reflection asks you to choose and justify a design for a new program, drawing on all four sections of the lesson.

Reflection

A provincial health ministry introduces free public transit passes for adults aged 65 and older with annual incomes below $30,000 in one health region on 1 April, and plans to extend the program to other regions in two years if it is effective. Monthly emergency department and primary care visit rates are available for every region for the five years before the launch, income is recorded for everyone in tax-linked administrative data, and uptake of the pass is recorded for each eligible person. Uptake is voluntary, and about half of eligible older adults collect a pass. The ministry asks whether the passes reduce health care use among older adults. In about 250 words, choose a design or a combination of designs, justify the choice by reference to how the passes were assigned and the data available, state the key assumption of each design in plain language, and specify one check for each assumption.

Model answer

The passes were assigned by region and launch date, and within the launch region by an income threshold, and uptake is voluntary. These features support two complementary designs.

The first is a controlled interrupted time series of monthly emergency and primary care visit rates among older adults with incomes below $30,000, with 60 months before launch, using the same income group in neighbouring regions as the comparison series. Its key assumption is that, without the passes, the difference between the launch region and the comparison regions would have continued along its pre-launch path, which requires that co-occurring events affect both. I would check the pre-launch trend in the difference series, record provincial and regional events around April, and examine a negative control outcome, such as visits among older adults with incomes well above the threshold in the launch region.

The second is a fuzzy regression discontinuity at the $30,000 income threshold within the launch region, with pass uptake as the treatment. Its key assumptions are that health care use would change smoothly with income at the threshold and that people cannot precisely manipulate reported income. I would check the density of incomes just below and above $30,000, test the continuity of age and prior visits at the threshold, and report estimates across several bandwidths. Because uptake is about half, the estimate would describe compliers near the threshold. Together, the two designs address history and selection with different assumptions, and agreement between them would strengthen the ministry's decision.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Quasi-Experimental Designs II: Time Series, Discontinuities and Natural Experiments (15 Questions)

Question 1: An evaluator has 60 monthly observations of a rate before a program launch and 18 after, and the program theory predicts an effect that builds as enrolment grows. Which impact model should the analysis plan specify?

The impact model should follow the program theory and be fixed in advance (Lopez Bernal et al., 2017). A gradually building effect implies a slope change. Choosing the shape after seeing the data (option c) risks fitting random fluctuation.

Question 2: How do Newey-West standard errors differ from ordinary least squares standard errors in a segmented regression?

Newey-West standard errors leave the ordinary least squares estimates unchanged and correct the variance for autocorrelation and unequal variance. Modelling the residuals as an autoregressive process (option c) describes generalized least squares methods such as Prais-Winsten.

Question 3: In the Section 1 worked example, the single-series effect in month 72 was −2.029 visits per 1,000 and the controlled effect was −1.378. Which conclusion is best supported?

The comparison region dropped at the launch month without the program, so the controlled analysis removes the shared change and gives a smaller estimate. A narrower interval (option b) is no reason to prefer an estimate that ignores a known history threat.

Question 4: Which design from Lesson 7 is most closely related to a controlled interrupted time series?

A controlled interrupted time series compares the change in a program series with the change in a comparison series, as difference-in-differences does, with many periods and modelled trends on each side of the interruption.

Question 5: A subsidy makes adults eligible at age 65. Which threat is specific to using this threshold in a regression discontinuity design?

When several programs share a threshold, the discontinuity reflects all of them, a problem known for age 65. Misreporting of date of birth (option d) is possible in principle but is difficult where age is verified administratively.

Question 6: Why is the confidence interval for a fuzzy regression discontinuity estimate wider than that for the reduced-form jump on which it is based?

The Wald estimator divides the reduced form by the first stage, so with a first stage of 0.4089 the estimate and its uncertainty are about 2.4 times those of the reduced form. A discrete running variable creates other difficulties, but it does not double the variance (option b).

Question 7: A fuzzy regression discontinuity design is a special case of which method?

In a fuzzy design, being above the threshold acts as an instrument for receiving the program, and the Wald estimator is the instrumental variable estimator for a binary instrument.

Question 8: Referred older adults in the simulated data had six-month scores 1.217 points higher than those not referred, while the regression discontinuity estimate was −0.485. What explains the difference?

People are referred because they are lonelier, so a comparison of referred and non-referred adults reflects selection. The regression discontinuity design avoids this by comparing people on either side of the threshold, and it also avoids the regression to the mean that affects a pre-post comparison (option a).

Question 9: Which kind of evidence most strengthens the claim that assignment in a natural experiment was as if random?

Dunning (2012) argues that a natural experiment's credibility rests on evidence about the assignment process. Sample size (option b) affects precision, and a significant outcome difference (option c) says nothing about whether assignment was independent of the outcome.

Question 10: Under monotonicity, whose effect does an instrumental variable estimate describe?

Imbens and Angrist (1994) showed that, when there are no defiers, the instrumental variable estimate is the local average treatment effect among compliers. Always-takers and never-takers do not change their participation in response to the instrument, so they contribute nothing to the estimate.

Question 11: An instrumental variable analysis has a first-stage F statistic of 4. Which problem does this most strongly suggest?

An F statistic well below about 10 signals a weak instrument (Staiger and Stock, 1997), which makes the estimate imprecise, biased toward the confounded estimate and sensitive to small violations of the other assumptions. The F statistic cannot test the exclusion restriction (option a).

Question 12: Which Canadian study used a regression discontinuity design at a provincial legal threshold?

Callaghan and colleagues (2013) compared hospital admissions just below and just above the provincial minimum legal drinking age. The Quebec child care study used difference-in-differences, and the minimum price studies used time-series analyses.

Question 13: In a Plan-Do-Study-Act cycle, what is the purpose of stating a prediction in the Plan stage?

A prediction turns a test into a learning opportunity, because the team compares what happened with what it expected. Taylor and colleagues (2014) found that few published PDSA applications documented predictions.

Question 14: A u-chart of monthly emergency visits per 1,000 older adults flags every January as special-cause variation. What is the most likely explanation?

Standard control limits assume a stable process without seasonality, so a regular winter peak appears as special-cause variation unless it is modelled. A p-chart (option d) is for proportions, and a u-chart is the usual choice for rates with varying denominators.

Question 15: A design justification states that a controlled interrupted time series assumes that the comparison region is similar. How could this statement be improved?

The precise assumption concerns the trajectory of the difference between series, and the justification should pair it with checks such as the pre-launch trend in the difference series and a negative control outcome. Equal average rates (option a) are unnecessary, because the model allows the regions to differ in level.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Quasi-Experimental Designs II: Time Series, Discontinuities and Natural Experiments.

You can now fit and interpret a segmented regression with seasonal terms and autocorrelation-consistent standard errors, add a comparison series, judge whether a time series is sensitive enough for a program's reach, estimate a fuzzy regression discontinuity with the Wald estimator and check its assumptions, assess a natural experiment and a proposed instrument, interpret run charts and control charts, and justify an evaluation design by naming its assumptions and the checks for them.

Lesson 9 turns from whether a program works to how it is implemented. It distinguishes efficacy, effectiveness and implementation, introduces implementation outcomes and frameworks such as CFIR 2.0 and RE-AIM, covers implementation strategies and adaptation, and presents hybrid effectiveness-implementation designs that combine the designs of Lessons 6 to 8 with the study of implementation.

Continue to Lesson 9 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

This glossary defines the terms, tools and people introduced in Lesson 8, grouped by the lesson's main themes.

Core Concepts
Interrupted time series A design that measures an outcome at regular intervals before and after an intervention begins at a known time, and asks whether the series changed by more than its earlier pattern predicts.
Segmented regression A regression model that fits separate lines before and after an interruption and estimates the change in level and the change in slope.
Level change The immediate step in an interrupted time series at the start of an intervention, relative to the projected pre-intervention trend.
Slope change The difference between the post-intervention and pre-intervention slopes of an interrupted time series.
Autocorrelation Correlation between observations of a series at different times, usually positive between adjacent months in health data.
Fourier terms Pairs of sine and cosine functions with a fixed period, used to model a smooth seasonal cycle with few parameters.
Controlled interrupted time series An interrupted time series with a comparison series that shares co-occurring events but not the intervention, used to address history threats.
Dilution The reduction in the measured effect of a program when the outcome is measured for a whole population, most of whom the program does not reach.
Regression discontinuity design A design that estimates a program's effect by comparing people just above and just below a threshold on the score that determines eligibility.
Running variable The score whose value relative to a threshold determines eligibility in a regression discontinuity design; also called the forcing or assignment variable.
Fuzzy regression discontinuity A regression discontinuity design in which the probability of receiving the program jumps at the threshold by less than one.
Bandwidth The window of the running variable on each side of the threshold within which a regression discontinuity analysis uses observations.
Manipulation of the running variable Precise control of scores by participants or staff to fall on a preferred side of the threshold, which threatens a regression discontinuity design.
Natural experiment An event or intervention outside researchers' control that divides a population into exposed and unexposed groups whose outcomes can be compared.
Instrumental variable A variable that changes the probability of receiving a program, shares no causes with the outcome and affects the outcome only through the program.
Exclusion restriction The assumption that an instrument affects the outcome only through its effect on receipt of the program.
Monotonicity The assumption that no one is less likely to receive the program because of the instrument, which rules out defiers.
Local average treatment effect The average effect of a program among compliers, the people whose program receipt is changed by the instrument.
Weak instrument An instrument only weakly associated with program receipt, which produces imprecise estimates that are sensitive to small violations of the other assumptions.
Common-cause and special-cause variation Common-cause variation is the ordinary variation of a stable process; special-cause variation arises from specific, identifiable sources outside the usual process.
Frameworks & Tools
Durbin-Watson test A test for first-order autocorrelation in regression residuals, whose statistic is near 2 when there is none and falls toward 0 with positive autocorrelation.
Newey-West standard errors Standard errors that remain valid under autocorrelation and unequal variance up to a chosen lag, while keeping the ordinary least squares estimates.
Wald estimator The ratio of the effect of an instrument (or threshold) on the outcome to its effect on program receipt.
Local linear regression The standard regression discontinuity estimator, which fits straight lines on each side of the threshold within the bandwidth.
Density test A test for a discontinuity in the distribution of the running variable at the threshold, such as McCrary's test or the rddensity method, used to detect manipulation.
Two-stage least squares An instrumental variable estimator that regresses program receipt on the instrument and then regresses the outcome on the predicted receipt.
Medical Research Council guidance on natural experiments Guidance by Craig and colleagues (2012) on when and how to use natural experiments to evaluate population health interventions.
Model for Improvement A quality improvement framework that asks three questions about aims, measures and changes and tests changes with Plan-Do-Study-Act cycles (Langley et al., 2009).
Plan-Do-Study-Act cycle An iterative method for testing a change on a small scale by planning with a prediction, carrying out the test, studying the results and acting on them.
Run chart A plot of a measure over time with the median as a centre line, interpreted with rules for shifts, trends, runs and astronomical points.
Statistical process control chart A plot of a measure over time with a centre line and control limits, usually three standard deviations from the centre, used to detect special-cause variation.
SQUIRE 2.0 The Standards for QUality Improvement Reporting Excellence, an 18-item guideline for reporting quality improvement work (Ogrinc et al., 2016).
Key People
Donald T. Campbell American social scientist (1916 to 1996) who introduced the regression discontinuity design with Thistlethwaite in 1960 and catalogued quasi-experimental designs, including the time-series design, with Stanley and later with Cook and Shadish.
Anita K. Wagner Health policy researcher at Harvard Medical School who led the widely used 2002 paper on segmented regression analysis of interrupted time series in medication use research.
John Snow English physician (1813 to 1858) whose comparison of cholera deaths among households supplied by two London water companies is often cited as an early natural experiment.
Peter Craig Public health researcher at the University of Glasgow and lead author of the Medical Research Council guidance on using natural experiments to evaluate population health interventions (2012).
Joshua D. Angrist Economist who, with Guido Imbens, developed the interpretation of instrumental variable estimates as local average treatment effects; they shared the 2021 Nobel memorial prize in economic sciences.
Guido W. Imbens Economist who developed the local average treatment effect with Joshua Angrist and wrote influential guidance on regression discontinuity estimation and bandwidth choice.
Walter A. Shewhart Physicist and statistician at Bell Telephone Laboratories (1891 to 1967) who developed the control chart and the distinction between chance and assignable causes of variation.
W. Edwards Deming American statistician (1900 to 1993) who promoted quality management and developed the Plan-Do-Study-Act cycle from Shewhart's earlier work.
No matching entries. Try a different search term.