# Lesson 9: Implementation Science

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are on Lesson nine of Program Planning and Evaluation, which is about implementation science.

**Sarah:** For the last few weeks we have been deep in study designs: randomized trials, difference-in-differences, interrupted time series. All of that was about whether a program works. Where does implementation science fit?

**Kiffer:** It picks up a question those lessons left open. Suppose a program has some evidence behind it. Why does it take hold in some clinics and fade in others, and what can a health authority do about that? Implementation science studies the methods that get evidence-based programs taken up, delivered well and kept going in routine practice.

**Sarah:** And our running case is the Cedar Valley Connector program again.

**Kiffer:** It is. Cedar Valley is a fictional health authority in British Columbia, and the connector program is a fictional social prescribing program. Primary care clinicians refer adults aged sixty-five and older who are lonely or isolated to a community connector, who meets them up to six times over twelve weeks, builds a plan with them, and links them to groups, volunteer roles, transport help and services.

**Sarah:** And this lesson is about the second wave.

**Kiffer:** Right. The first twelve clinics were chosen because they were ready, which we discussed in Lesson seven. Now the other twelve are joining. In this lesson those second-wave clinics are smaller and more stretched. Four are rural, two depend heavily on locum physicians, and several have no spare room for a connector.

**Sarah:** Let's start with Section one, the research-to-practice gap. People often quote a figure of seventeen years. Where does that come from?

**Kiffer:** It comes from an estimate by Balas and Boren in two thousand that it takes about seventeen years for research findings to reach routine clinical practice. Morris, Wooding and Grant later reviewed the studies behind that kind of figure and found that the estimates vary a lot, depending on where you put the start and the end of the lag. So I treat the number as a sign that translation is slow, with real uncertainty about how slow.

**Sarah:** And the lesson also talks about a voltage drop.

**Kiffer:** That is the second problem. Programs tested under controlled conditions often deliver smaller benefits once ordinary staff deliver them to more varied people with fewer resources. Chambers, Glasgow and Stange discussed this in twenty thirteen. They also questioned the idea that a program can be perfected once and then delivered unchanged.

**Sarah:** Social prescribing seems to have the opposite problem, though. It spread before the evidence was strong.

**Kiffer:** That is a fair observation. Connector programs spread across the United Kingdom and Canada faster than the evidence accumulated, and early reviews judged the evidence to be limited. Even so, whether a connector program helps a population depends on implementation: whether clinicians refer, whether people engage, and whether there are community resources to connect people to.

**Sarah:** The section distinguishes efficacy, effectiveness and implementation research. Can you walk through those?

**Kiffer:** Brian Flay drew the first distinction in nineteen eighty-six. An efficacy trial asks whether an intervention does more good than harm under optimum conditions. An effectiveness trial asks the same question under real-world conditions. Lesson six called these explanatory and pragmatic trials. Implementation research asks a third question. It takes the program's effectiveness as given and tests the strategies used to get it adopted and delivered.

**Sarah:** So what is being compared?

**Kiffer:** Usually two ways of supporting delivery, such as a standard launch in second-wave clinics against a launch with a practice facilitator, judged by implementation outcomes like the share of eligible older adults who get referred.

**Sarah:** The reading uses Geoffrey Curran's phrase, the thing. I liked that.

**Kiffer:** It is a very useful teaching device. Curran calls the intervention the thing. Effectiveness research asks whether the thing works. Implementation strategies are the things we do to help people and places do the thing. And implementation outcomes describe how much and how well they do it. For Cedar Valley, the thing is screening, referral and connector support. The strategies are training, changes to the medical record, feedback reports and so on.

**Sarah:** Section one then turns to how to interpret a null result.

**Kiffer:** Yes. If an outcome evaluation finds no effect, there are two broad explanations. Either the program was delivered and its theory of change did not hold, which is intervention failure, or it was never delivered well enough to have a chance, which is implementation failure. Basch and colleagues called the evaluation of a program that was never adequately implemented a Type Three error. Without implementation data you cannot tell the two apart, which is one reason Enola Proctor and colleagues argued for measuring implementation outcomes explicitly.

**Sarah:** Couldn't you just compare the well-implemented clinics with the poorly implemented ones?

**Kiffer:** You can, and it is informative, but you have to remember that the comparison is observational, even inside a randomized trial. Clinics that implement well may also have better staffing or richer community resources, and those things can affect loneliness directly. All the confounding logic from Lesson seven applies.

**Sarah:** Let's get to the implementation outcomes themselves. Proctor and colleagues defined eight.

**Kiffer:** They did, in twenty eleven. First, though, they had a model with three kinds of outcomes. Implementation outcomes are the effects of deliberate efforts to implement a program. Service outcomes describe the quality of services, things like timeliness, equity and safety. Client outcomes describe changes in the people served, such as their symptoms or function. Implementation outcomes are preconditions for the other two.

**Sarah:** And the eight are?

**Kiffer:** Acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration and sustainability. Proctor and colleagues also suggested that the first four matter most early on, when settings are deciding whether to take a program up. Fidelity and cost matter during delivery, and penetration and sustainability matter later.

**Sarah:** Acceptability and appropriateness sound like the same thing to me.

**Kiffer:** They sound close, but they are measured separately for a good reason. Appropriateness is perceived fit: does this program suit our patients and our setting? Acceptability is whether people find it agreeable or satisfactory. A rural clinician might say the connector program is exactly what her lonely patients need, which is high appropriateness, and still dislike the scripted screening questions, which is low acceptability.

**Sarah:** Are there standard measures for these?

**Kiffer:** For three of them, yes. Bryan Weiner and colleagues developed the Acceptability of Intervention Measure, the Intervention Appropriateness Measure and the Feasibility of Intervention Measure. Each has four items on a five-point scale, and they tested them in twenty seventeen.

**Sarah:** Isn't there an overlap with the process evaluation material from Lesson five?

**Kiffer:** A large overlap. Fidelity appears in both, and penetration is close to reach. The difference is the role the measures play. In process evaluation they help explain an outcome result. In implementation research they are the dependent variables, the outcomes a strategy is judged against.

**Sarah:** The section ends with a worked example on penetration. Can you take us through it?

**Kiffer:** Sure. Lesson two estimated that the twelve first-wave clinics have about twenty-three thousand attached patients aged sixty-five and older. The regional survey found that twenty-four point five percent of older adults score six or higher on the three-item loneliness scale from the University of California, Los Angeles, which is the program's referral threshold. That gives about five thousand six hundred and thirty-five eligible people.

**Sarah:** And there were three hundred and twelve referrals in the first six months.

**Kiffer:** So penetration of the referral step is three hundred and twelve over five thousand six hundred and thirty-five, which is about five and a half percent. Two hundred and forty-one people attended a first meeting, which is about four percent of the eligible population. And across the twelve clinics, penetration ranged from roughly one to ten percent.

**Sarah:** That range seems important.

**Kiffer:** It changes how you read the loneliness result. The mean score among people with two measurements fell from seven point one to six point three, but whatever the program does for participants, it reached fewer than one in twenty eligible older adults. Population impact depends on how many people a program reaches as well as how much it helps them.

**Sarah:** The reading is also cautious about that denominator.

**Kiffer:** It should be. Applying a regional survey prevalence to each clinic's panel assumes the panels look like the survey population. A rural clinic with an older panel may have much higher need. A better denominator counts people who were actually screened and scored six or higher, which means adding a screening field to the medical record. The second-wave plan does that.

**Sarah:** On to Section two, frameworks. Implementation science seems to have a framework for everything.

**Kiffer:** It has a great many, which is why Per Nilsen's taxonomy from twenty fifteen is so helpful. He sorted theories, models and frameworks by purpose. Process models describe or guide the steps of moving research into practice. Determinant frameworks, classic theories and implementation theories all aim to explain what influences implementation outcomes. Evaluation frameworks tell you what to measure to judge whether implementation succeeded. An evaluation usually needs one approach for each purpose.

**Sarah:** The reading covers two process frameworks.

**Kiffer:** The first is the Knowledge-to-Action framework from Ian Graham and colleagues at the University of Ottawa, published in two thousand six. The Canadian Institutes of Health Research use it to describe knowledge translation. It has a knowledge creation funnel and an action cycle, with steps such as adapting knowledge to the local context, assessing barriers to its use, and selecting and tailoring interventions. The Cedar Valley second wave sits at the step of assessing barriers.

**Sarah:** And the second?

**Kiffer:** The Exploration, Preparation, Implementation, Sustainment framework, known as EPIS, from Gregory Aarons and colleagues in twenty eleven. It describes those four phases and the outer and inner context in each. Moullin and colleagues later emphasized bridging factors, which link the outer and inner context. For Cedar Valley, the steering committee and the agreement with the First Nations health centre that hosts a connector are bridging factors.

**Sarah:** Then the centrepiece of the section, the Consolidated Framework for Implementation Research. Everyone calls it CFIR.

**Kiffer:** Laura Damschroder and colleagues published the original in two thousand nine, and an updated version, CFIR two point oh, in twenty twenty-two, based on feedback from people who had used it. It has five domains: the innovation, the outer setting, the inner setting, the individuals involved, and the implementation process. Inside those are forty-eight constructs and nineteen subconstructs.

**Sarah:** What changed in the update?

**Kiffer:** Several things. It pays more attention to the people who receive the innovation and adds constructs related to equity. And the individuals domain now lists roles, such as mid-level leaders and innovation deliverers, and describes their characteristics using the capability, opportunity, motivation and behaviour model, known as COM-B, which students met in Lesson two.

**Sarah:** How do you actually use it?

**Kiffer:** You start with three definitions. What exactly is the innovation, what is the inner setting, and what is the outer setting? For Cedar Valley, the innovation is screening, referral and connector support. The inner setting is each clinic. The outer setting is the clinic's community and the health authority. Then you collect data, usually through interviews built from the construct definitions, code the transcripts to constructs, and rate them.

**Sarah:** Rate them how?

**Kiffer:** Damschroder and Lowery, in a study of a weight management program across Veterans Affairs medical centres, rated each construct for each site from minus two to plus two, combining whether it helped or hindered with how strongly. Zero is neutral and an X marks mixed evidence. Then you compare sites that implemented well with sites that did not.

**Sarah:** And the Cedar Valley team did this before the second wave launched.

**Kiffer:** Three months before. They interviewed a manager and a clinician in each of the twelve second-wave clinics, plus the first-wave connectors. The strongest negatives were local conditions in the four rural clinics, meaning no transit and long distances to groups, and relative priority in the two clinics that depend on locums, where getting patients in to see a doctor outweighs everything else. There were also moderate negatives for the complexity of the referral, which in the first wave needed a separate form outside the medical record, for the lack of space in seven clinics, for clinicians' uncertainty about how to ask about loneliness, and for the fact that first-wave clinics never saw their own referral numbers. On the positive side, clinicians valued having somewhere to send lonely patients. That relative advantage was rated plus two. The ratings matter mainly because they force every barrier to be tied to evidence and linked to a decision.

**Sarah:** The section also covers the Theoretical Domains Framework. When would you use that instead?

**Kiffer:** When the problem is a specific behaviour of health professionals. The framework, validated by Cane, O'Connor and Michie in twenty twelve, has fourteen domains drawn from behaviour change theories. Suppose a clinic has supportive managers and plenty of resources, and clinicians still rarely ask older patients about loneliness. That is a behaviour problem, and the framework helps you ask whether it is knowledge, skills, beliefs about consequences, memory during short visits, or uncertainty about whether loneliness is part of their role.

**Sarah:** Now the evaluation framework, RE-AIM.

**Kiffer:** RE-AIM was proposed by Russell Glasgow, Vogt and Boles in nineteen ninety-nine. The letters stand for reach, effectiveness, adoption, implementation and maintenance. The central argument is that a program's public health impact depends on all five. A program that helps people a great deal but reaches few of them, or is taken up in few settings, or fades when funding ends, has limited impact.

**Sarah:** Where does representativeness come in?

**Kiffer:** It is the heart of reach and adoption. You ask how many people and settings took part, what proportion that is, and whether they look like the people and settings who did not. That makes RE-AIM a natural tool for looking at equity.

**Sarah:** Give us the second-wave numbers.

**Kiffer:** The second-wave clinics have about twenty-one thousand attached patients aged sixty-five and older, so about five thousand one hundred and forty-five eligible. Eleven of the twelve clinics adopted, meaning they signed on and made at least one referral within three months. In six months there were two hundred and sixty-three referrals and one hundred and ninety-eight first meetings, which gives a reach of about three point eight percent. In the adopting clinics, forty-one of seventy-four clinicians referred at least once.

**Sarah:** That sounds close to the first wave.

**Kiffer:** It is fairly close. Reach was about four point three percent in the first wave and about three point eight in the second, and clinician adoption was about fifty-nine percent against fifty-five. The representativeness indicators show where the problem lies. Men were thirty-one percent of participants but forty-five percent of attached older patients. The four rural clinics hold thirty percent of the eligible population and contributed about eighteen percent of participants.

**Sarah:** And PRISM helps explain that.

**Kiffer:** PRISM is a model Feldstein and Glasgow built in two thousand eight to add context to RE-AIM. It looks at the intervention from the organization's and the patient's perspective, the characteristics of recipients, the implementation and sustainability infrastructure, and the external environment. For the rural gap, the external environment, distance and transport, is the obvious first place to look.

**Sarah:** Can I ask how all these frameworks fit together? It feels like a lot.

**Kiffer:** A sensible combination for Cedar Valley uses EPIS to describe the phases, CFIR to assess determinants before launch and again at six months, and RE-AIM with PRISM to evaluate the rollout. Proctor's perception measures, acceptability, appropriateness and feasibility, sit inside the implementation dimension. Penetration lines up with reach, and sustainability with setting-level maintenance.

**Sarah:** Section three is about implementation strategies. What counts as a strategy?

**Kiffer:** Proctor, Powell and McMillen defined strategies as the methods or techniques used to improve a program's adoption, implementation and sustainability. Some are discrete, like a reminder in the medical record. Most real efforts combine several, because barriers sit at several levels.

**Sarah:** Where is the line between the program and the strategy?

**Kiffer:** It can be blurry. Connector meetings and the co-developed plan are clearly the program. Training clinicians and sending feedback reports are clearly strategies. Screening is the awkward case, since it is part of the referral pathway, but a prompt reminding clinicians to screen is a strategy. The important thing is to state the boundary in the logic model so that the evaluation knows what it is testing.

**Sarah:** And the vocabulary for strategies comes from ERIC.

**Kiffer:** The Expert Recommendations for Implementing Change project, led by Byron Powell and published in twenty fifteen. It used a modified Delphi process to produce seventy-three discrete strategies, each with a name and definition. Waltz and colleagues then grouped them into nine clusters, from evaluative and iterative strategies like audit and feedback, through interactive assistance like facilitation, to changes in infrastructure like new record templates.

**Sarah:** How did the Cedar Valley team choose strategies?

**Kiffer:** They took the CFIR barriers to a workshop with the steering committee, first-wave champions, second-wave managers and the connectors. Waltz and colleagues have a matching tool that lists the strategies experts recommend for each CFIR barrier, although the experts disagreed a lot, so it is a starting list. The team ended up with a one-step referral template and a screening prompt, quarterly audit and feedback, clinic champions, short practice-based training, community meeting sites, more transport money and remote meetings for rural participants, and a practice facilitator.

**Sarah:** And then they specified the facilitation.

**Kiffer:** Following Proctor, Powell and McMillen, you specify each strategy on seven dimensions: the actor, the action, the action target, the timing, the dose, the implementation outcome it should change, and the justification. For Cedar Valley, a half-time facilitator works with each clinic's champion and manager, from two months before launch to six months after. The dose comes to five and a half hours of contact per clinic. The target outcomes are clinician adoption and penetration. Labels like training or support are impossible to replicate, and specification also makes cost visible. The whole support bundle costs sixty-four thousand dollars, which is about five thousand eight hundred dollars per adopting clinic. That is the implementation cost, separate from the cost of running the program itself, and Lesson ten will use both.

**Sarah:** Then adaptation. I suspect this is where practitioners and researchers argue.

**Kiffer:** They have argued about it for a long time. The older view treated any departure from protocol as a fidelity problem. Hawe, Shiell and Riley proposed a more useful way of thinking in two thousand four: standardize the function of each component, the purpose it serves, and let the form vary with context. Jolles and colleagues later developed this as core functions and forms.

**Sarah:** What are the core functions for Cedar Valley?

**Kiffer:** The lesson names four: a person-centred conversation about what matters to the older adult, a co-developed written plan, active linkage to at least one community resource, and follow-up to deal with barriers. How those happen can vary. A meeting can be at home or by telephone, a plan can be in the participant's language, and linkage can involve a volunteer driver.

**Sarah:** And FRAME records the changes.

**Kiffer:** FRAME is an expanded framework for reporting adaptations and modifications, published by Shannon Wiltsey Stirman and colleagues in twenty nineteen. It records eight things about each change: when it was made, whether it was planned, who decided, what was changed, at what level, what type of change it was, whether it was fidelity-consistent, and why.

**Sarah:** The reading logs three examples.

**Kiffer:** The first is a shift to mainly telephone and video meetings in the rural clinics, introduced after people missed meetings in winter. That was reactive, but it preserves all four core functions, so it is fidelity-consistent. The second is a connector who replaced the first one-to-one meeting with a group session because her caseload was full.

**Sarah:** That one sounds like a problem.

**Kiffer:** It probably is fidelity-inconsistent, because a group session is unlikely to give each person a conversation about what matters to them. The coordinator kept the group session as an optional extra, restored the one-to-one meeting, and spread referrals across connectors. A FRAME log lets you catch a change like that early.

**Sarah:** And the third example is the land-based pathway.

**Kiffer:** Yes. It was co-designed with a First Nation and led by the Nation's health centre and community members, with the program team in a supporting role. It adds land-based and cultural activities as ways of connecting people. It is planned, and it is fidelity-consistent because it delivers the core functions through culturally grounded forms. The documentation and any data about it follow the agreements with the Nation, in line with the principles of ownership, control, access and possession from Lesson four.

**Sarah:** The section closes with scale-up and sustainability.

**Kiffer:** The World Health Organization and ExpandNet distinguish horizontal scale-up, which extends a program to more sites, from vertical scale-up, which embeds it in budgets, policy and structures. The second wave, going from twelve clinics to twenty-four, is horizontal. Moving the connector positions into the health authority's base budget would be vertical.

**Sarah:** And sustainability?

**Kiffer:** Moore and colleagues proposed a definition in twenty seventeen. After a defined period, the program continues to be delivered, or behaviour change is maintained, and the program may evolve or adapt while continuing to produce benefits. That last part matters, because it accepts change as part of sustaining a program, which fits the Dynamic Sustainability Framework of Chambers and colleagues.

**Sarah:** How did Cedar Valley assess it?

**Kiffer:** The steering committee completed the Program Sustainability Assessment Tool from Luke and colleagues, which rates eight domains. Partnerships and program evaluation scored well. Funding stability scored poorly, because the connector positions sit in a time-limited allocation. So the sustainability plan centres on vertical scale-up.

**Sarah:** Section four, effectiveness-implementation hybrid designs. What problem are they solving?

**Kiffer:** Curran and colleagues argued in twenty twelve that the usual sequence, efficacy then effectiveness then implementation, is slow, and that effectiveness evidence often comes from conditions unlike later implementation. Hybrid designs blend the questions so you learn about implementation while effectiveness is still being established.

**Sarah:** And there are three types.

**Kiffer:** A type one study tests a program's effects on health outcomes as its primary aim and gathers information on implementation as a secondary aim. A type two study has co-primary aims, testing the program's effects and an implementation strategy. A type three study tests an implementation strategy as its primary aim, with health outcomes secondary.

**Sarah:** What did the twenty twenty-two update change?

**Kiffer:** Curran and colleagues reviewed a decade of use. They suggested talking about hybrid studies, because a hybrid can use many research designs, including non-randomized ones. They said the type is set by the study's aims, and secondarily by its outcomes, so you should name both the design and the type, as in a parallel cluster randomized hybrid type three study. They also said type three studies should be powered on implementation outcomes, which are usually measured at the site level.

**Sarah:** How do you choose a type?

**Kiffer:** They offered four considerations: how strong the effectiveness evidence is, how much adaptation you expect, how much you know about the determinants of implementation, and whether you are ready to test an implementation strategy. Strong evidence and little adaptation point toward type three. Weak evidence and little knowledge of determinants point toward type one.

**Sarah:** And Cedar Valley fits each type at a different point.

**Kiffer:** The first wave, with evidence only from a two-clinic pilot, fits type one. The primary aim is the loneliness effect, estimated with the comparison-group methods of Lesson seven, and the CFIR interviews are the secondary aim. The second wave fits type two. Since connector capacity is limited and a waiting list is likely anyway, referred older adults could be randomized to start immediately or after twelve weeks, and the strategy bundle is tested against the first-wave penetration of about five and a half percent.

**Sarah:** That strategy test has no comparison group, though.

**Kiffer:** That is its main weakness, and the reading says so. If penetration rises, you cannot be sure the bundle caused it. A version that also randomized clinics to the bundle or a standard launch would be stronger, but with twelve clinics it would have very little power.

**Sarah:** Which brings us to the calculation.

**Kiffer:** Suppose a later provincial spread offers the program to forty clinics in other regions, and they are randomized to a standard launch or an enhanced launch with facilitation and feedback. That is type three. Penetration is a clinic-level outcome, so the number of clinics drives power. If clinic penetration has a standard deviation of about four percentage points and you expect a five-point difference, the usual approximation gives about ten clinics per arm, which rounds up to eleven, or twenty-two in total. Forty clinics leave room for withdrawals, while six per arm, which is all the second wave could offer, falls well short.

**Sarah:** And finally, reporting.

**Kiffer:** The Standards for Reporting Implementation Studies, known as StaRI, were published by Hilary Pinnock and colleagues in twenty seventeen. There are twenty-seven items, and for many of them you report the implementation strategy and the intervention in parallel. You use it alongside design-specific guidelines, such as the trial reporting standards from Lesson six or the quality improvement standards from Lesson eight.

**Sarah:** Let's finish with the worked example, the implementation part of an evaluation plan.

**Kiffer:** The plan selects an implementation framework for the program and specifies the implementation outcomes to be measured. Name one or two frameworks, classify each using Nilsen's taxonomy, and justify the choice with reference to the setting and the evaluation questions. Then specify four to six implementation outcomes, each with a numerator, a denominator, a data source and a time of measurement.

**Sarah:** And if the program could be a hybrid study?

**Kiffer:** Then say which type and why, using the four considerations. I would also like every plan to include at least one outcome that tracks equity or representativeness, the way the Cedar Valley plan tracks rural older adults and men. The worked example at the end of Section four shows one way to put this together.

**Sarah:** Any last advice?

**Kiffer:** Keep the two layers apart. Effectiveness outcomes tell you whether the program works, and implementation outcomes tell you how much and how well it is delivered. A good evaluation plan needs both, and it needs to say which framework explains the implementation results.

**Sarah:** Thanks, Kiffer. Next week is the final lesson, on economic evaluation, reporting and evaluation use.

**Kiffer:** Thanks, Sarah. See you then.
