# Lesson 10: Economic Evaluation, Reporting and Evaluation Use

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This is the last episode for Program Planning and Evaluation, which feels a little strange to say.

**Sarah:** It does. We have spent nine weeks building an evaluation plan piece by piece. What is left to do in the final lesson?

**Kiffer:** Two things. The first is money: what a program costs and whether it is good value. The second is what happens at the end of an evaluation, when the evidence has to become a judgement and the judgement has to reach the people who make decisions. The lesson closes with a worked example of a complete evaluation proposal for Cedar Valley, which pulls the whole plan together.

**Sarah:** And we are still with Cedar Valley?

**Kiffer:** We are. The Cedar Valley Connector program is our fictional community connector program for adults aged sixty-five and older in British Columbia. Its first wave runs in twelve primary care clinics on a budget of eight hundred and forty thousand dollars a year. For this lesson I added some illustrative figures for its first full year: six hundred and forty referrals, and five hundred older adults who attended a first meeting with a connector.

**Sarah:** Let's start with costs, then. I would have thought the cost of a program is just its budget. Why does that need a whole section?

**Kiffer:** Because the budget records what a program spends, and economists care about what a program uses. The idea underneath is opportunity cost, which is the value of the best alternative use of a resource. Suppose a clinic lends Cedar Valley a meeting room at no charge. No money changes hands, so the budget ignores it. But that room could have hosted a foot-care clinic, so it has a cost to society.

**Sarah:** So the financial cost and the economic cost can be quite different.

**Kiffer:** They can, and the gap tends to be largest for programs that depend on partners. Social prescribing is a good example. A connector links people to walking groups, volunteer roles and seniors' centres that the health system does not pay for. If you report only the budget, the program looks cheaper than it is, and you hide costs that community organizations may struggle to carry as referrals grow.

**Sarah:** Whose costs you count must depend on who is asking, though.

**Kiffer:** Exactly, and that is the idea of perspective. From the program perspective, Cedar Valley costs its budget. From the health system perspective, you add things the health authority pays for elsewhere, such as corporate overhead, the time clinicians spend screening and referring, the clinic space, and any change in the use of emergency departments or primary care. From the societal perspective, you also add volunteer time, participants' time and travel, and costs that community organizations bear.

**Sarah:** Is there a rule about which one to use?

**Kiffer:** In Canada, the economic evaluation guidelines from the Canadian Agency for Drugs and Technologies in Health, which is now called Canada's Drug Agency, use the publicly funded health care payer perspective as the reference case and allow a societal analysis alongside it. The Second Panel on Cost-Effectiveness in Health and Medicine in the United States recommended reporting both, plus an impact inventory that lists effects inside and outside the health sector even when they cannot be valued.

**Sarah:** For something like Cedar Valley, I'd guess the societal view matters a lot.

**Kiffer:** It does, because so much of what the program mobilizes sits in the community. The impact inventory is a good discipline even when you cannot value everything, because it shows readers what your totals leave out.

**Sarah:** How do you actually go about costing a program?

**Kiffer:** Drummond and colleagues describe three steps. You identify the resources, you measure how much of each is used, and you value each quantity with a unit cost. Identification starts from the inputs and activities in the logic model from Lesson three. Measurement uses natural units, such as hours of staff time, numbers of referrals or rides funded. Valuation attaches a price, such as salary plus benefits per hour, or a fee from the provincial Medical Services Plan schedule for physician services.

**Sarah:** What about things without a price, like volunteer time?

**Kiffer:** That needs a shadow price. One common choice is the replacement cost, the wage you would pay someone to do the same task. Another is what the volunteer would otherwise earn. The choice changes the answer, so you state it and test it. For retired participants, the value of their time is genuinely contested, and many analyses simply report the hours.

**Sarah:** The section spends a lot of time on something called time-driven activity-based costing. Can you explain it without the accounting jargon?

**Kiffer:** I'll try. It comes from Robert Kaplan and Steven Anderson, writing in management accounting, and it needs two numbers. The first is what an hour of a staff member's capacity costs. A Cedar Valley connector position costs eighty-five thousand dollars a year. After vacation, statutory holidays, sick leave, training and team meetings, a connector has about one thousand seven hundred hours available for real work. Divide one by the other and you get fifty dollars an hour.

**Sarah:** And the second number?

**Kiffer:** How long each activity takes. The analyst had connectors keep activity logs for four weeks and reviewed case records. On average, a participant needs about half an hour of triage, an hour and a half for the first meeting, three hours of follow-up meetings, two hours of linkage work such as phoning groups and arranging transport, two hours of travel, and an hour and a half of documentation. That adds to ten and a half hours.

**Sarah:** So ten and a half hours at fifty dollars an hour.

**Kiffer:** Five hundred and twenty-five dollars per participant for connector time.

**Sarah:** That seems low. If I divide the budget by five hundred participants I get something much bigger.

**Kiffer:** You get one thousand six hundred and eighty dollars, which is called gross costing. The gap comes from shared costs such as the coordinator, the analyst and training, from the transport fund and partner grants, and from unused capacity. Five hundred participants used only about forty-four percent of the connectors' available hours in year one, and Kaplan and Anderson point out that this method makes unused capacity visible where a single average hides it.

**Sarah:** Which brings us to fixed and variable costs.

**Kiffer:** Right. Most of Cedar Valley's costs are fixed in the short run. The connectors are paid whether they see five hundred people or eight hundred. Variable costs, such as transport vouchers and clinician referral time, rise with each participant. So the marginal cost of one more participant is small while there is spare capacity, and the average cost falls as caseloads grow. If each connector could serve one hundred and twenty people a year, seven connectors could serve eight hundred and forty, and the same budget would work out to one thousand dollars per participant.

**Sarah:** That sounds like a strong argument for waiting until after the first year to judge a program.

**Kiffer:** It is an argument for reporting both. A start-up year can make a program look expensive. But assuming full capacity flatters the program if the referrals never come. So the lesson recommends reporting the observed cost, stating the program's capacity, and showing a steady-state scenario.

**Sarah:** What did Cedar Valley's full economic cost come to?

**Kiffer:** From the health system perspective, we start with the eight hundred and forty thousand dollar budget and add ten percent for overhead, which is eighty-four thousand dollars. We add forty-five dollars of clinician time for each of the six hundred and forty referrals, which is twenty-eight thousand eight hundred. And we add one thousand six hundred dollars of meeting space for each of the twelve clinics, which is nineteen thousand two hundred. The total is nine hundred and seventy-two thousand dollars, or one thousand nine hundred and forty-four dollars per participant.

**Sarah:** About sixteen percent more than the budget.

**Kiffer:** Yes. And if we add two thousand volunteer hours at twenty-five dollars an hour, the societal total is one million and twenty-two thousand dollars. I should stress that those unit costs are assumptions chosen for teaching. A real costing would document where each one came from.

**Sarah:** Let's go there. What makes an economic evaluation an economic evaluation?

**Kiffer:** Drummond and colleagues give two conditions. You look at both costs and consequences, and you compare at least two alternatives. If you only describe the costs of one program, that's a partial evaluation. It can be useful, but it cannot tell you about value.

**Sarah:** And then there are different flavours.

**Kiffer:** There are. Cost-effectiveness analysis measures consequences in one natural unit. Cost-utility analysis uses quality-adjusted life years. Cost-benefit analysis puts everything in money. And cost-consequence analysis lays out several outcomes beside the costs without combining them, which is common in public health because programs affect many things at once.

**Sarah:** Give me the cost-effectiveness version for Cedar Valley.

**Kiffer:** First we need the incremental cost. In our illustrative results, participants had slightly fewer emergency visits and primary care visits than people in comparison clinics, which offsets seventy-six dollars. So the program costs one thousand eight hundred and sixty-eight dollars more per participant than usual care. Now suppose that at twelve weeks, thirty-four percent of participants and twenty-two percent of comparison adults scored below the referral threshold on the loneliness scale. That is twelve more people out of every hundred no longer screening as lonely.

**Sarah:** So for a hundred people, the extra cost is one hundred and eighty-six thousand eight hundred dollars, and you get twelve people out of loneliness.

**Kiffer:** Which works out to about fifteen thousand five hundred and sixty-seven dollars per additional older adult no longer screening as lonely. That's the incremental cost-effectiveness ratio in natural units.

**Sarah:** Is fifteen thousand dollars good value?

**Kiffer:** That's the problem. Nobody else reports results in that unit, so there's nothing to compare it with. And it ignores any effect on health beyond loneliness. That is why health economists developed the quality-adjusted life year.

**Sarah:** Remind listeners what a quality-adjusted life year is.

**Kiffer:** It weights time by a utility value. One is full health, zero is a state equivalent to death, and states judged worse than death can be negative. A year in full health is one quality-adjusted life year. A year at a utility of zero point seven is zero point seven of one. The utilities come from people's preferences, elicited with methods such as the standard gamble and the time trade-off. George Torrance and colleagues at McMaster developed the time trade-off in the early nineteen seventies, and most evaluations now use a scored questionnaire such as the EuroQol five-dimension questionnaire, which has a Canadian value set.

**Sarah:** So how many quality-adjusted life years does Cedar Valley produce?

**Kiffer:** In our illustration, the groups start with the same utility. At twelve weeks the program group is zero point zero three higher. By twelve months the difference has gone. If you draw that, it is a triangle with a base of one year and a height of zero point zero three, and the area is half of that, zero point zero one five quality-adjusted life years per participant.

**Sarah:** That seems tiny.

**Kiffer:** It is small, which is typical for a light-touch program measured on a generic instrument. Divide the incremental cost of one thousand eight hundred and sixty-eight dollars by zero point zero one five and you get about one hundred and twenty-four thousand, five hundred dollars per quality-adjusted life year gained.

**Sarah:** Is that bad?

**Kiffer:** Canada has no official threshold. Canadian studies often compare results with reference values such as fifty thousand or one hundred thousand dollars per quality-adjusted life year. By those standards the base case looks expensive. But hold on before you judge it, because the shape of that triangle was an assumption.

**Sarah:** Because nobody measured anything between twelve weeks and twelve months.

**Kiffer:** Right. If the twelve-week difference persisted to twelve months, the gain would be about zero point zero two six five, and the ratio would fall to about seventy thousand dollars.

**Sarah:** Before we get to uncertainty, walk me through the cost-effectiveness plane.

**Kiffer:** Picture a graph with the difference in health on the horizontal axis and the difference in cost on the vertical axis, with usual care at the centre. A program in the lower right quadrant saves money and improves health, so it dominates. One in the upper left costs more and does less, so it is dominated. Most new programs land in the upper right, where you pay more for more health and the decision depends on how much you are willing to pay. Cedar Valley sits there.

**Sarah:** And the threshold is a line through the middle.

**Kiffer:** A line through the origin whose slope is the willingness to pay. Below the line, the program is cost-effective. A related number, the net monetary benefit, converts the health gain into dollars at that willingness to pay and subtracts the cost. At one hundred thousand dollars per quality-adjusted life year, Cedar Valley's health gain is worth one thousand five hundred dollars, and the cost is one thousand eight hundred and sixty-eight, so the net benefit is minus three hundred and sixty-eight dollars.

**Sarah:** Let's talk about uncertainty, then, because I suspect that is where this gets interesting.

**Kiffer:** It is. There are two main tools. Scenario analysis changes one assumption or a set of assumptions at a time. Probabilistic sensitivity analysis gives every uncertain input a probability distribution, draws values thousands of times and recalculates the result each time.

**Sarah:** What did the scenarios show for Cedar Valley?

**Kiffer:** Removing the health care offsets or taking a societal perspective barely changed the ratio. Two assumptions mattered a lot. If benefits persist, the ratio is about seventy thousand four hundred dollars. If caseloads fill to eight hundred and forty, it is about seventy-three thousand six hundred. With both, it falls to about forty-one thousand six hundred dollars, under the fifty thousand dollar reference value.

**Sarah:** So the program's value depends on things the first year cannot tell us.

**Kiffer:** That is the main message of the section. And the probabilistic analysis adds a second message. In the worked example, five thousand draws put almost all the results in the upper right quadrant, and the probability that the program is cost-effective is about twenty-eight percent at one hundred thousand dollars per quality-adjusted life year and about fifty-one percent at one hundred and twenty-five thousand.

**Sarah:** Why treat persistence as a scenario instead of putting it into the probabilistic model?

**Kiffer:** Because it's a question about the structure of the model, about a period nobody measured. Probabilistic analysis is good at uncertainty in parameter values. For a structural assumption, you show the alternatives side by side so the decision-maker can see that everything hinges on it.

**Sarah:** The section also takes on social return on investment, which I see a lot in community sector reports.

**Kiffer:** Social return on investment is a form of cost-benefit analysis from the social enterprise world. You map outcomes with the people affected, give each outcome a financial proxy, and adjust for deadweight, which is what would have happened anyway, along with attribution and drop-off. The result is a ratio of social value to investment.

**Sarah:** And the cautions?

**Kiffer:** The ratio depends heavily on the proxies and on the deadweight estimate. In our hypothetical example, someone assigns four thousand dollars to each participant whose loneliness score falls by a point, and three hundred participants qualify. That is one point two million dollars, a ratio of about one point four three to one against the budget. Set deadweight at half, which is what the comparison group suggests, and the ratio drops to zero point seven one. Use the full health system cost as the investment and it drops to zero point six two.

**Sarah:** Same program, very different headline.

**Kiffer:** And the proxy itself is a judgement no data can confirm. A systematic review by Banke-Thomas and colleagues found wide variation in how these analyses are done in public health. The strength of the approach is that it asks the people affected what counts. The weakness is that ratios from different studies cannot be compared.

**Sarah:** Last piece of Section two is reporting.

**Kiffer:** The Consolidated Health Economic Evaluation Reporting Standards, known as CHEERS, were updated in twenty twenty-two to a twenty-eight-item checklist covering the perspective, comparator, time horizon, valuation and uncertainty. The new version added items on analysis plans, distributional effects, and how patients and the public were engaged.

**Sarah:** Let's move to Section three, synthesis and judgement. This felt like a return to Lesson one.

**Kiffer:** It is. In Lesson one we introduced evaluative reasoning, moving from criteria to standards to evidence to a judgement. Section three develops the tools. The starting point is the difference between a finding and a conclusion. A finding says seventy-eight point one percent of referred adults attended a first meeting. A conclusion says reach was good against the standard the committee agreed.

**Sarah:** And the lesson leans on Scriven's distinction between merit and worth.

**Kiffer:** Merit is the intrinsic quality of a program, such as how well it works and how well it is delivered. Worth is its value in a particular context, which brings in cost, need and the alternatives. Keeping them apart lets you say something honest about Cedar Valley: it works, and at current caseloads it is not yet clearly good value.

**Sarah:** How do evaluators keep themselves honest when they make those calls?

**Kiffer:** With an evaluative rubric agreed in advance. You set out the criteria, describe what excellent, good, adequate and poor look like for each one, name the evidence you will use, and state how important each criterion is. Jane Davidson's work and the rubric approach described by King and colleagues are good guides. For Cedar Valley, the steering committee, including its older adults with lived experience and its First Nations representatives, agreed five criteria: reach, equity of reach, cultural safety, effectiveness and value for money.

**Sarah:** So the economic result goes into the rubric as one criterion.

**Kiffer:** Yes, following Julian King, who argues that economic results should be rated against agreed standards alongside everything else. A ratio below one hundred thousand dollars counts as good, for example, and a ratio above that with plausible scenarios below it counts as adequate.

**Sarah:** Then how do you combine five ratings into one judgement? Couldn't you just give each a number and add them up?

**Kiffer:** You could, and the lesson shows why that is risky. Suppose effectiveness has a weight of thirty percent, cultural safety and value for money twenty percent each, and reach and equity fifteen percent each. A program rated excellent on everything except a poor rating for cultural safety scores three point four out of four. That reads as better than good, for a program that harmed some of its participants.

**Sarah:** That is unsettling.

**Kiffer:** It's the compensation problem. Strong ratings buy back a harmful one, and the weights give an impression of precision they don't deserve. Scriven and Davidson recommend what they call qualitative weight and sum, where you compare the pattern of ratings within importance categories and explain your reasoning in words. You add hurdles for criteria that cannot be traded off. Cedar Valley's committee made cultural safety a hurdle, so the program cannot be rated better than adequate overall if cultural safety is poor.

**Sarah:** And within a criterion, the evidence comes from several places.

**Kiffer:** It does, and it can relate in three ways, following Jennifer Greene. It can converge, when sources agree. It can be complementary, when sources describe different aspects of the same change. Or it can be dissonant. The Cedar Valley example has all three.

**Sarah:** Tell me about the dissonance.

**Kiffer:** The difference-in-differences estimate was a fall of zero point four points in loneliness, with a confidence interval from zero point one to zero point seven. Interviews agreed for most people. But rural participants told interviewers they were very satisfied with their connectors, and their loneliness scores barely moved.

**Sarah:** Which would be easy to average away.

**Kiffer:** And averaging would hide the most useful finding. The program records explained it. The transport fund ran out in the ninth month, and the groups rural participants were linked to were often a long drive away. The mechanism in the program theory, getting people into community groups, was blocked by context. That is a finding with a clear implication.

**Sarah:** So what was the overall judgement?

**Kiffer:** Reach was good. Equity of reach was adequate, because attendance was much lower among people whose first language is neither English nor French and in rural areas, although there is a plan to close the gap. Cultural safety was good. Effectiveness was good, with moderate confidence. Value for money was adequate, with low confidence. The hurdle was passed, so the committee rated merit as good and worth at year-one caseloads as adequate.

**Sarah:** And they recommended going ahead?

**Kiffer:** With three conditions: fill the caseloads before adding positions, restore and extend transport support, and measure outcomes at twelve months to see whether the benefits last.

**Sarah:** The section ends with a checklist called TIDieR. Why does that belong here?

**Kiffer:** Because a judgement applies to a program as it was actually delivered. Glasziou and colleagues found that published descriptions of interventions were often too thin to replicate. The Template for Intervention Description and Replication, led by Tammy Hoffmann, has twelve items covering the rationale, materials, procedures, who delivered it, how, where, how much, tailoring, modifications and fidelity. For Cedar Valley, the dose item records up to six meetings over twelve weeks, with an average of four meetings and ten and a half hours of connector time.

**Sarah:** Section four. Reporting and use. Where do you start?

**Kiffer:** With the report itself. It should answer the evaluation questions in order, and it should keep three kinds of statement apart. A finding reports evidence. A conclusion judges it against a standard. A recommendation proposes an action. If you keep them distinct, a reader can trace any recommendation back to the evidence behind it.

**Sarah:** And then the executive summary, which is what most decision-makers read.

**Kiffer:** Often the only thing they read. The Canadian Health Services Research Foundation promoted a format called one, three, twenty-five: one page of main messages, a three-page executive summary, and a report of no more than twenty-five pages. Each layer must make sense for a reader who stops there.

**Sarah:** What separates a strong main message from a weak one?

**Kiffer:** A weak one describes the report, such as saying that cost-effectiveness results were sensitive to assumptions. A strong one tells the reader something they can act on. At current caseloads the program costs about one hundred and twenty-five thousand dollars for each year of full health gained, and that would fall below fifty thousand if caseloads fill and benefits last.

**Sarah:** The section also covers charts, which I suspect many of us get wrong.

**Kiffer:** Cleveland and McGill showed that people compare quantities most accurately when they are positions on a common scale, as in a dot plot or bar chart, and much less accurately as angles or areas, as in pie charts or bubbles. Give each chart one message, make the title state it, sort the categories, label points directly, and show the standard you are judging against. For economic results, a busy executive will usually understand a short table of scenarios better than a cost-effectiveness plane.

**Sarah:** What about knowledge translation more broadly?

**Kiffer:** The Canadian Institutes of Health Research describe knowledge translation as an iterative process of synthesis, dissemination, exchange and ethical application of knowledge, and a utilization-focused evaluation with an active steering committee is integrated knowledge translation by design. Lavis and colleagues give five planning questions: what to transfer, to whom, by whom, how, and with what effect.

**Sarah:** What does that look like for Cedar Valley?

**Kiffer:** The executive gets a short briefing before the second-wave budget decision. The steering committee holds a workshop to interpret the draft findings. Older adults get a plain-language summary in large print and at community meetings, presented with the committee's older adult members. Clinicians hear from a clinician champion. And findings about First Nations participants are reviewed and interpreted with the partners before release, under the principles of ownership, control, access and possession.

**Sarah:** Which leads to evaluation use. What does the research say?

**Kiffer:** Carol Weiss showed that research gets used in more ways than direct application. Reviews since then describe instrumental use, when findings directly inform a decision, conceptual use, when they change how people understand a problem, symbolic use, when an evaluation is used to legitimize a decision already made, and process use, which comes from taking part. Cousins and Leithwood found that use depends on relevance, credibility, communication and timeliness, and on the decision setting. Patton adds what he calls the personal factor, someone who genuinely cares about the findings.

**Sarah:** If you had to pick one thing that would most help Cedar Valley's evaluation get used, what would it be?

**Kiffer:** Timing. If the budget decision is made in month ten and the full report arrives in month fourteen, the report cannot inform that decision, however good it is. So the plan includes an interim briefing before the decision.

**Sarah:** Let's finish with the evaluation proposal. What goes into it?

**Kiffer:** The Cedar Valley proposal runs to about six thousand words, not counting references and appendices. It brings together the parts of the plan from the earlier lessons and adds three new pieces from this lesson: a costing and economic evaluation plan that names a perspective and the resources to be measured, a draft rubric with synthesis rules, and a reporting and use plan for at least three audiences.

**Sarah:** And the executive summary.

**Kiffer:** No more than two pages, written for decision-makers, and it does not count toward the six thousand words. Because a proposal is a plan, the summary says what the evaluation will deliver, how and when. The Cedar Valley worked example shows how the team divided the words across eleven parts, from the program description through design, costing, synthesis, ethics and reporting, and it includes an excerpt of their executive summary.

**Sarah:** Any last advice for anyone writing a proposal like this?

**Kiffer:** Read the proposal as a decision-maker would. Check that the need leads to the questions, the questions lead to the design, and the design leads to a judgement that someone could use in time. And be candid about what the evaluation cannot tell them. That candour is part of what makes the rest credible.

**Sarah:** Thanks, Kiffer, and thanks to everyone who has listened across the course.

**Kiffer:** Thank you, Sarah. It has been a pleasure, and I hope these lessons are useful in the evaluations you go on to do.
