# Lesson 1: Foundations of Program Planning and Evaluation

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This is the first episode for Health Sciences eight twenty-six, Program Planning and Evaluation, so we are starting at the very beginning.

**Sarah:** Which raises an obvious first question. Most of our listeners have done research methods courses already. Why does a graduate program need a separate course on evaluation?

**Kiffer:** Because evaluation asks a different kind of question. A research methods course teaches you how to find out whether something is true in general. An evaluation course teaches you how to reach a defensible judgement about a particular program, for particular people who have to make a decision about it. The methods overlap a great deal. The purpose and the audience are different, and that changes a surprising amount.

**Sarah:** We'll come back to that. Before we get into definitions, tell me about the case we'll be using all term.

**Kiffer:** It's a fictional program called the Cedar Valley Connector program, run by a fictional health authority in British Columbia. Family doctors and nurse practitioners screen adults aged sixty-five and older for loneliness and social isolation. People who screen positive are referred to a community connector, who meets them up to six times over twelve weeks, works out a plan with them, and links them to community groups, volunteer roles, transportation help and services.

**Sarah:** So it's a social prescribing program.

**Kiffer:** Yes, a community connector model of social prescribing. The program started in twelve of the region's twenty-four primary care clinics, and the other twelve are scheduled to join a year later. It has a budget, seven connector positions, community partners, a steering committee that includes older adults with lived experience of loneliness, and an Indigenous health partnership with local First Nations. All of its numbers are made up for teaching, and I've kept them consistent so they add up.

**Sarah:** Why that program in particular?

**Kiffer:** Because it raises nearly every problem the course deals with: a need to document, a theory about how connection reduces loneliness, interest holders who want different things, outcomes that are hard to measure, a staggered rollout that might allow a comparison, and an Indigenous partnership that changes how the evaluation should be governed.

**Sarah:** All right. Section one. What is evaluation?

**Kiffer:** The definition most evaluators work from comes from Michael Scriven. Evaluation is the systematic determination of the merit, worth or significance of something. The important word is determination. An evaluation reaches a conclusion about how good or valuable something is. A report that describes what a program did and stops there has stopped short of evaluation in Scriven's sense.

**Sarah:** Merit, worth and significance sound like three ways of saying the same thing.

**Kiffer:** They are three different kinds of value. Merit is intrinsic quality, judged against what a program of that kind ought to do. Worth is value to a particular organization or population in a particular setting, so it brings in cost and alternatives. Significance is importance, including importance beyond the local setting.

**Sarah:** Can you put that in Cedar Valley terms?

**Kiffer:** Suppose the connectors are well trained, the plans fit each person, and participants feel less lonely. That's merit. Now suppose a volunteer visiting program produces similar benefits at a third of the cost. Cedar Valley still has merit, and its worth to the health authority is lower, because there's a cheaper route to the same result. And if loneliness is common in the region and little else addresses it, the program may have real significance.

**Sarah:** You also talk about formative and summative evaluation. Where do those terms come from?

**Kiffer:** Also from Scriven, in a paper from nineteen sixty-seven. Formative evaluation is done to improve a program while it is being developed or delivered. Summative evaluation is done to reach an overall judgement, usually to inform a decision about whether to continue, expand, change or end the program. Robert Stake had a nice analogy, which Scriven reported. When the cook tastes the soup, that's formative. When the guests taste the soup, that's summative.

**Sarah:** So the difference is about who's tasting and why.

**Kiffer:** Yes, and the same method can serve either purpose. If connectors use a participant survey to adjust how they run meetings, the survey is formative. If the executive uses the same survey to decide whether the second wave goes ahead, it is serving a summative purpose.

**Sarah:** Let's go back to your opening point. If evaluation and research use the same methods, what really separates them?

**Kiffer:** Sandra Mathison argued that the clearest differences are purpose and audience. Research aims at knowledge that holds beyond the setting where it was produced, and its main audience is other researchers. Evaluation aims to inform decisions about a particular program, and its audience is the people who run it, fund it or are affected by it. So in evaluation the questions are usually negotiated with users, the timing follows decisions and budgets, and values are explicit.

**Sarah:** But surely some studies are both.

**Kiffer:** Many are. A comparison of first-wave and second-wave Cedar Valley clinics would inform the health authority's decision and could also add to the international evidence on social prescribing. The distinction is a matter of emphasis, and it still matters, partly because it affects whether you need research ethics review, which we'll reach in section three.

**Sarah:** You said values are explicit in evaluation. How does an evaluator get from data to a judgement without it just being their opinion?

**Kiffer:** This is the heart of the first section. Scriven described a general logic of evaluation, and Deborah Fournier set it out as four steps. First, you establish the criteria of merit, the dimensions on which the program will be judged. Second, you construct standards, which say how well the program has to do on each criterion to count as excellent, good, adequate or poor. Third, you measure performance and compare it with the standards. Fourth, you synthesize across criteria into an overall judgement.

**Sarah:** And where do the criteria come from?

**Kiffer:** The program's objectives are the obvious source, though they can be too narrow or too easy to reach. Other sources include assessed needs, ethical and legal requirements, professional standards, and the values of the people involved. At Cedar Valley, cultural safety came onto the list because of the Indigenous health partnership, reach came from managers worried about referrals going nowhere, and cost came from the executive.

**Sarah:** And then those go into a rubric.

**Kiffer:** Right. A rubric is a table with the criteria in rows and descriptions of each level of performance in the columns. The most important thing about a rubric is when you write it. If the steering committee agrees it before the data arrive, nobody can quietly move the goalposts once the results are in.

**Sarah:** Walk me through the Cedar Valley example.

**Kiffer:** In the first six months, clinicians referred three hundred and twelve older adults, and two hundred and forty-one came to a first meeting. That's seventy-seven point two percent, which the rubric rates as good. Among participants with loneliness scores at intake and at twelve weeks, the average score fell from seven point one to six point three, on a scale from three to nine. A fall of eight tenths of a point is also rated good.

**Sarah:** And cost?

**Kiffer:** Spending over the six months was four hundred and twenty thousand dollars, or about one thousand seven hundred and forty-three dollars for each person who came to a first meeting. That lands in the adequate band. Cultural safety was assessed through a talking circle with Indigenous participants and a review with the First Nations partners, and the committee rated it good. So overall, good.

**Sarah:** With some caveats, I assume.

**Kiffer:** Two. The cost per person should fall as connectors fill their caseloads, because start-up costs were spread over very few people. And the fall in loneliness scores can't be credited to the program. People are referred only when they score six or higher, and when you select people for high scores, their scores tend to come down on the next measurement even if nothing happens. That's regression to the mean, and we'll spend real time on it in Lesson seven.

**Sarah:** So the rubric gives you a judgement about performance, and it can't tell you about cause.

**Kiffer:** That's right, and it leads to the types of evaluation. A needs assessment asks how big a problem is and who has it. Formative evaluation improves a program during design and early delivery. Process evaluation asks whether the program is delivered as intended. Outcome evaluation asks whether intended changes happened among participants. Impact evaluation, as we use the term, estimates how much of the change the program caused, by comparison with what would have happened without it. Economic evaluation compares costs with consequences.

**Sarah:** And these fit together somehow?

**Kiffer:** Rossi, Lipsey and Henry arrange them as a hierarchy: need first, then the design and theory of the program, then implementation, then outcomes and impact, then cost and efficiency. Each level depends on the ones beneath it. If you run an impact evaluation on a program that was never delivered properly, a null result tells you very little, because you can't tell a bad idea from a badly delivered one.

**Sarah:** Let's move to section two, approaches to evaluation. What do you mean by an approach?

**Kiffer:** An approach is a set of prescriptions about how to do evaluation. It tells you whose questions to answer, what counts as credible evidence, how to make judgements of value, and what role the evaluator should play. Evaluation theorists call these prescriptive models theories, which can confuse people coming from the sciences.

**Sarah:** And the course organizes them on a tree.

**Kiffer:** Marvin Alkin and Christina Christie drew the history of evaluation theory as a tree. Its roots are accountability and systematic social inquiry. The methods branch grew from Donald Campbell's work on experiments and is concerned with credible evidence about effects. The use branch is concerned with whether evaluations get used. The valuing branch grows from Scriven and asks how judgements of value are made and whose values count. The third edition of their book places culturally responsive and Indigenous approaches on that branch.

**Sarah:** Start with the use branch.

**Kiffer:** The use branch starts from a fairly depressing observation, which is that many evaluations are finished and then ignored. Michael Quinn Patton's early research on federal health evaluations found what he called the personal factor: evaluations got used when an identifiable person or group cared about the findings and helped shape the study. So utilization-focused evaluation begins by identifying the primary intended users and designing everything around their intended use.

**Sarah:** Who would that be at Cedar Valley?

**Kiffer:** The steering committee, and the director who has to decide how the second wave rolls out. The evaluator would sit down with them early and ask what they need to know, by when, and what they would do differently depending on the answer.

**Sarah:** Patton also developed developmental evaluation.

**Kiffer:** He did. Developmental evaluation is for innovations in complex, changing conditions, where the program's form is still emerging. The evaluator joins the team, brings data into decisions as they arise, and documents how the innovation changes. At Cedar Valley, the program and one First Nation are co-designing a land-based connection pathway whose final form nobody knows yet. That's a natural fit, done in partnership with the Nation.

**Sarah:** And participatory and empowerment evaluation?

**Kiffer:** Participatory evaluation shares the work with people who aren't professional evaluators. Brad Cousins, at the University of Ottawa, and Elizabeth Whitmore distinguished a practical form, which involves staff and managers to make findings more useful, from a transformative form, which involves community members to shift power. Empowerment evaluation, from David Fetterman, has programs evaluate their own work with the evaluator as a coach. Critics such as Daniel Stufflebeam worried that it weakens independence.

**Sarah:** Now the methods branch.

**Kiffer:** Two approaches there focus on explaining how programs work. The first is theory-driven evaluation, associated with Huey-Tsyh Chen. Chen argued that if you only test whether outcomes changed, you treat the program as a black box. Theory-driven evaluation spells out what the program does, what it should change first, and how those early changes lead to the final outcomes, and then tests each link.

**Sarah:** Why does that matter in practice?

**Kiffer:** Because it tells you how to read a disappointing result. Carol Weiss distinguished implementation failure from theory failure. Suppose Cedar Valley loneliness scores don't budge. If connectors held very few meetings and made very few linkages, that's implementation failure. If meetings and linkages happened as planned and people still felt just as lonely, that's theory failure, and the program idea itself needs rethinking.

**Sarah:** And realist evaluation?

**Kiffer:** Realist evaluation comes from Ray Pawson and Nick Tilley. Their starting point is that programs work differently for different people, so the useful question is what works, for whom, in what circumstances and why. Programs work by triggering mechanisms, which are the ways people respond to the resources a program offers, and mechanisms only fire in certain contexts. The unit of analysis is the context, mechanism and outcome configuration.

**Sarah:** Give me one for Cedar Valley.

**Kiffer:** Take an older woman who has recently lost her husband and has never belonged to a community group. The connector goes with her to her first walking group, which eases her anxiety about a room of strangers, and she keeps going. Now take a rural man who has stopped driving. For him, the active ingredient is transportation help. Both end up less lonely, through different mechanisms, and an average would hide both stories.

**Sarah:** Which brings us to the valuing branch.

**Kiffer:** The valuing branch begins with Scriven's argument that evaluators must make value judgements defensibly, and later theorists asked whose values should define merit. For health evaluation in Canada, two families on this branch matter a great deal: culturally responsive evaluation and Indigenous evaluation.

**Sarah:** What does culturally responsive evaluation ask of an evaluator?

**Kiffer:** It places the culture and context of the community at the centre of every stage, from framing questions to sharing findings. Stafford Hood, Rodney Hopson and Karen Kirkhart trace its roots to African American evaluators and scholars of education in the United States. Kirkhart made an argument I find persuasive: culture is a matter of validity. If you use measures or relationships that participants don't recognize, your conclusions can simply be wrong.

**Sarah:** And Indigenous evaluation?

**Kiffer:** Indigenous evaluation is grounded in Indigenous knowledge systems, values and governance, and is led by or done in partnership with Indigenous communities. Its principles include relationships built on respect and trust, accountability to the community, and self-determination over questions, data and interpretation. Verna Kirkness and Ray Barnhardt's four Rs, respect, relevance, reciprocity and responsibility, are a common starting point, and Lesson four goes much further.

**Sarah:** What would that look like at Cedar Valley?

**Kiffer:** The First Nations partners would help decide which outcomes matter, which might include connection to family, culture and land as well as loneliness scores. They would agree how data about their members are held and shared, and review findings before anything is published. And nobody would assume that a framework from one Nation applies to another.

**Sarah:** So does a real evaluation pick one approach?

**Kiffer:** Usually it combines several. For Cedar Valley I'd frame the evaluation as utilization-focused, built around the second-wave decision, structure the outcome work as theory-driven, use realist analysis to explain why some clinics do better, use developmental evaluation for the land-based pathway, and make the whole thing culturally responsive, with the Indigenous partnership shaping governance.

**Sarah:** Let's go to section three, evaluation in Canada. Where does the story start?

**Kiffer:** With the federal government. The Treasury Board issued its first formal program evaluation policy in nineteen seventy-seven. There have been several since, including a Policy on Evaluation in two thousand and nine. The current one is the Policy on Results, which took effect on July first, twenty sixteen.

**Sarah:** What does it require?

**Kiffer:** It joins performance measurement and evaluation in one framework. Each department keeps a results framework with its core responsibilities, intended results and indicators, and an inventory of its programs. Deputy heads name a head of evaluation and approve a five-year evaluation plan every year. Departments have discretion over most coverage, but ongoing grant and contribution programs averaging five million dollars or more a year must be evaluated at least once every five years.

**Sarah:** Cedar Valley is a provincial program, though. Does any of this apply?

**Kiffer:** Not directly. British Columbia's health authorities aren't governed by the federal policy. It would matter if a federal department funded part of the program. More generally, the structure is a good model: track a small set of indicators all the time, and schedule evaluations periodically to explain what the indicators show. Lesson five picks that up.

**Sarah:** Next, the profession. Is evaluation a regulated profession in Canada?

**Kiffer:** No. Anyone can call themselves an evaluator. The Canadian Evaluation Society, formed in nineteen eighty-one, responded with competencies and a professional designation. The competencies, revised in twenty eighteen, fall into five domains: reflective, technical, situational, management and interpersonal practice. Situational practice, for example, is reading the setting, meaning the program's history, its politics and the communities it serves.

**Sarah:** And the designation?

**Kiffer:** It's called the Credentialed Evaluator designation. To apply, you need a graduate degree, or a graduate certificate or diploma in evaluation, the equivalent of two years of full-time evaluation work within the past ten years, and written examples showing how you meet the competencies. You keep it through continuing professional learning. The designation is voluntary, and some employers and funders ask for it.

**Sarah:** How do we judge whether an evaluation itself is any good?

**Kiffer:** The main reference is the Program Evaluation Standards, from the Joint Committee on Standards for Educational Evaluation. The third edition, from twenty eleven, has thirty standards under five attributes. Utility asks whether the evaluation serves its users. Feasibility asks whether it's practical. Propriety asks whether it's fair, legal and respectful. Accuracy asks whether its conclusions are dependable. Evaluation accountability, which the third edition added, asks whether the evaluation is documented and itself evaluated.

**Sarah:** Can those pull against each other?

**Kiffer:** Often. A bigger sample improves accuracy and hurts feasibility. Reporting full results for small groups improves transparency and can identify participants. The standards make you name those trade-offs.

**Sarah:** Now ethics. This is the part students always ask me about. When does an evaluation need research ethics board review?

**Kiffer:** In Canada the key document is the Tri-Council Policy Statement, in its twenty twenty-two edition. Article two point one defines research as an undertaking intended to extend knowledge through disciplined inquiry or systematic investigation. Article two point five says that quality improvement studies, program evaluation activities and performance reviews don't need research ethics board review when they are used exclusively for assessment, management or improvement purposes.

**Sarah:** Exclusively is doing a lot of work there.

**Kiffer:** It is. If the purpose includes extending knowledge beyond the program, the activity is research. Two more points. If data collected for evaluation are later proposed for research, that secondary use may need review. And the policy says that the choice of method and the intent or ability to publish are not what determine whether something is research.

**Sarah:** So wanting to present at a conference doesn't automatically turn an evaluation into research.

**Kiffer:** That's right, and the same is true of a randomized design. In both cases, the deciding question is the purpose of the activity. The policy also notes that activities outside research ethics board review can still raise ethical issues that deserve independent consideration.

**Sarah:** Where does that independent consideration come from?

**Kiffer:** Often from a screening tool. The best known in Canada is ARECCI, which stands for A Project Ethics Community Consensus Initiative, offered by Alberta Innovates. Its screening tool walks a project lead through the project's primary purpose and a set of risk filters, and suggests the kind of review that fits, which might be a research ethics board, a second opinion review, or an internal review. In British Columbia, Interior Health asks its quality improvement and evaluation teams to use the ARECCI tools.

**Sarah:** And when Indigenous communities are involved?

**Kiffer:** Chapter nine of the Tri-Council Policy Statement applies to research involving First Nations, Inuit and Métis Peoples, and it requires community engagement when research is likely to affect a community's welfare. An evaluation under Article two point five is formally outside that chapter, but its principles are still the right standard. A university study comparing Cedar Valley clinics to publish evidence for other provinces would be research, so it would need review and engagement with the First Nations partners.

**Sarah:** What other ethical problems come up day to day?

**Kiffer:** Consent for routine data that people gave for their care. Confidentiality in small programs, where a table might describe five people. Participants who depend on a service and hesitate to criticize it. And pressure on findings. Surveys of evaluators, going back to Morris and Cohn in nineteen ninety-three, find that pressure to soften findings is among the most common ethical problems. A written agreement at the start about who owns the report is the best protection.

**Sarah:** Section four. You use a framework from the Centers for Disease Control and Prevention to organize the course. Why that one?

**Kiffer:** Because it gives a clear order to the work, and it's widely taught in public health. It's a United States federal document, and its steps are general enough for Canadian settings. The original, from nineteen ninety-nine, has six steps in a cycle: engage interest holders, describe the program, focus the evaluation design, gather credible evidence, justify conclusions, and ensure use and share lessons learned. At the centre are four standards: utility, feasibility, propriety and accuracy.

**Sarah:** And then it was updated in twenty twenty-four.

**Kiffer:** Yes, by Daniel Kidder and colleagues. The update keeps the six-step cycle and makes four main changes. It adds a new first step, assess context, which looks at readiness for evaluation, the interest holders involved, the place where the program operates, and the evaluation capacity available. It renames several steps, so step three becomes focus the evaluation questions and design, step five becomes generate and support conclusions, and step six becomes act on findings.

**Sarah:** And the other two changes?

**Kiffer:** The third is the one the authors call the biggest. It adds three cross-cutting actions that apply at every step: engage collaboratively, advance equity, and learn from and use insights. Engagement used to be step one, and now it runs through the whole cycle. The fourth change replaces the four standards with five federal evaluation standards: relevance and utility, rigour, independence and objectivity, transparency, and ethics.

**Sarah:** I noticed the update also changes some vocabulary.

**Kiffer:** It replaces an older term, built on the word stake, with interest holder. The authors explain that the older term can imply a power imbalance and has a violent connotation for some American Indian and Alaska Native tribes. We use interest holder in this course for the same reasons.

**Sarah:** Run Cedar Valley through the cycle for me.

**Kiffer:** Assessing context means noticing that routine data exist for the first wave but there are no comparison data, that rural transit is poor, and that the evaluation capacity is one half-time analyst. Describing the program means writing an agreed account of what it is and does. Focusing the questions means agreeing with the executive and the steering committee that the evaluation will inform the second-wave decision. Gathering evidence means loneliness scores, attendance, emergency department visits and talking circles with Indigenous participants.

**Sarah:** And the last two steps?

**Kiffer:** Generating and supporting conclusions means applying the rubric agreed in advance and saying how confident we are in each conclusion. Acting on findings means a briefing for the executive before the decision, a plain-language summary for older adults, and a report reviewed with the First Nations partners before release. And the cross-cutting actions apply throughout, so we keep asking whether the partners are shaping each step and whether people facing the most barriers are being served.

**Sarah:** Which brings us to the evaluation plan itself.

**Kiffer:** An evaluation plan sets out how an evaluation will be done, and its parts follow the cycle in the order an evaluator works: a program description, a needs statement and objectives, a logic model, an interest holder map with evaluation questions, an evaluation matrix, decisions about design, an implementation framework, and finally costing and an executive summary for decision-makers. The course teaches those parts in that order, with Cedar Valley as the example.

**Sarah:** Does writing the plan involve collecting any data?

**Kiffer:** No. The plan comes first. Nobody interviews participants or pulls records at that stage. If you describe a program at your own workplace, use only information you're allowed to share, and leave out anything that identifies clients or staff.

**Sarah:** So what goes into the program description?

**Kiffer:** Six elements: its purpose, population, activities, resources, setting and history. A real description also lists its sources, such as program documents and public reports.

**Sarah:** How do you know a program is ready to be described and evaluated?

**Kiffer:** It helps if it's specific, like a single service, with enough documentation to work from, and at a stage where evaluation decisions are still open. A realistic composite based on real models, like Cedar Valley, should say so. For a program led by or serving First Nations, Inuit or Métis communities, rely on what they've made public and describe the program in their own terms.

**Sarah:** And what makes a good description?

**Kiffer:** Specific numbers, like budget lines, eligibility rules and the number of sites. A neutral, descriptive voice. And a clear line between what the program intends and what it has been shown to achieve. The Cedar Valley example says the planners expect fewer emergency department visits, and it says plainly that this hasn't been tested.

**Sarah:** That's a helpful place to end. Anything you want students to hold on to from this first lesson?

**Kiffer:** An evaluation is a reasoned judgement made for particular people. State your criteria and standards up front, reach a judgement, and be honest about what the evidence can and can't show. Everything else in the course gives you better tools for doing that.

**Sarah:** Thanks, Kiffer. Next time, needs assessment and planning models.

**Kiffer:** Thanks, Sarah. See you then.
