Economic Evaluation, Reporting and Evaluation Use
Program Planning & Evaluation
Learning objectives for this lesson:
- Distinguish the financial cost of a program from its economic cost, and explain how the perspective of an analysis determines which costs are counted.
- Calculate a program's cost per participant with gross costing and time-driven activity-based costing, and explain how the use of capacity changes it.
- Distinguish cost-minimization, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis, and calculate the QALYs gained from utility measurements.
- Calculate an incremental cost-effectiveness ratio and net monetary benefit, place the result on the cost-effectiveness plane, and interpret scenario and probabilistic sensitivity analyses.
- Explain the cautions that apply to social return on investment, and identify the reporting items in CHEERS 2022.
- Construct an evaluative rubric and synthesize mixed evidence into judgements of merit and worth with a stated level of confidence.
- Describe a program with the TIDieR checklist so that others can understand, cost and replicate it.
- Plan reports, executive summaries, data visualizations and knowledge translation products for decision-makers and communities, and explain the conditions under which evaluations are used.
- Describe the structure of a complete evaluation proposal, including its work plan, budget and risk register, its costing, synthesis and reporting plans, and an executive summary for decision-makers.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on Rossi, P. H., Lipsey, M. W., & Henry, G. T. (2019). Evaluation: A Systematic Approach (8th ed.). SAGE; and Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
Costing a Program
Learning Objectives for this section
- Distinguish the financial cost of a program from its economic cost, and explain why a program budget usually understates the resources a program consumes.
- Explain how the perspective of an analysis (program, health system or societal) determines which costs are counted.
- Identify, measure and value the resources a program consumes, including in-kind, volunteer and participant resources.
- Compare gross costing, micro-costing and time-driven activity-based costing, and calculate a cost per participant from connector time data.
- Distinguish fixed, variable, average and marginal costs, and explain how the use of capacity changes the cost per participant.
This lesson completes the evaluation plan that the course has built since Lesson 1. Decision-makers ask what a program costs and whether the money would do more good elsewhere. This section teaches the first question, and Section 2 uses the answer to teach economic evaluation. Costing also supports budgets for scale-up and documents implementation cost, one of the implementation outcomes of Proctor and colleagues (2011) taught in Lesson 9.
The running example is the fictional Cedar Valley Connector program, a community connector (social prescribing) program for adults aged 65 and older run by the fictional Cedar Valley Health Authority in British Columbia. Its first wave operates in 12 of the region's 24 primary care clinics on an annual budget of $840,000: seven connector positions ($595,000), a coordinator ($105,000), a half-time analyst ($55,000), a transport fund ($40,000), partner grants ($30,000), and training, travel and data systems ($15,000). For this lesson, assume that in the program's first full year clinicians made 640 referrals and 500 older adults attended a first meeting with a connector (312 referrals and 241 first meetings in the first six months, and 328 referrals and 259 first meetings in the second six months). These year-one figures, and the unit costs introduced below, are illustrative assumptions for this lesson.
1.1 Costs as Opportunity Costs
Economic evaluation rests on the idea of opportunity cost: the cost of using a resource for one purpose is the value of the best alternative use that is given up (Drummond et al., 2015). When the health authority pays a connector, the money could instead have funded home-care hours or a nurse in a long-term care home. When a clinic lends a meeting room to the program, the room could have been used for a foot-care clinic. The budget records the first of these choices, because money changes hands, and misses the second, because nothing is paid. An evaluator who wants to know the full cost of a program therefore distinguishes the financial cost, which is the money spent and recorded in accounts, from the economic cost, which is the value of all resources used, whether or not anyone paid for them.
Three related ideas follow. In-kind contributions are counted at their opportunity cost, sunk costs are left out of forward-looking decisions, and transfer payments are left out of societal costs because they move money without using resources. The distinction between financial and economic cost matters most for programs that rely on partners. Social prescribing programs link people to community groups and volunteer roles that the health system does not pay for, so a costing that reports only the budget makes them appear cheaper than they are and hides costs that community organizations may not be able to absorb as referrals grow.
1.2 The Perspective of the Analysis
The perspective of an economic analysis defines whose costs and consequences are counted. The program perspective counts costs to the organization delivering the program. The health system perspective, called the publicly funded health care payer perspective in Canadian guidance, counts all costs to the public health system, including changes in the use of other services. The societal perspective counts all costs, whoever bears them.
The guidelines for economic evaluation published by CADTH (2017), the organization now called Canada's Drug Agency, specify the publicly funded health care payer perspective for the reference case and allow a broader societal perspective as an additional analysis when important costs fall outside the health system. The Second Panel on Cost-Effectiveness in Health and Medicine recommended that analysts report two reference cases, one from the health care sector perspective and one from the societal perspective, and that they complete an impact inventory listing the consequences of an intervention inside and outside the health sector, even those that are not valued (Sanders et al., 2016). For a program such as Cedar Valley, whose resources and effects extend into the community, the impact inventory makes visible what a health-system analysis leaves out.
The cost of the first wave is its budget, $840,000 a year. This perspective answers the program manager's question about how much money the program needs, and it omits corporate overhead, clinician referral time, donated rooms and effects on other services.
The cost adds corporate overhead, clinician screening and referral time and clinic space, together with any change in emergency department and primary care use, which may offset part of the program cost. This is the reference perspective in Canadian guidance.
The cost also includes volunteer time, participants' and caregivers' time and travel, and costs that community organizations and the First Nations health centre bear beyond what the program pays them. This perspective reveals cost-shifting from the health system to the community.
| Cost item (Cedar Valley first wave) | Program | Health system | Societal |
|---|---|---|---|
| Connector, coordinator and analyst salaries and benefits | Yes | Yes | Yes |
| Transport fund, partner grants, training, travel and data systems | Yes | Yes | Yes |
| Health authority corporate overhead | No | Yes | Yes |
| Clinician time for screening and referral | No | Yes | Yes |
| Clinic meeting rooms provided in kind | No | Yes | Yes |
| Changes in emergency department and primary care visits | No | Yes | Yes |
| Volunteer time in community groups | No | No | Yes |
| Participants' and caregivers' time and travel | No | No | Yes |
| Costs to community organizations beyond partner grants | No | No | Yes |
1.3 Identifying, Measuring and Valuing Resources
Drummond and colleagues (2015) describe costing as three steps. The evaluator first identifies the resources the program uses, then measures the quantity of each resource, and then values each quantity with a unit cost. The total cost is the sum, across resources, of the quantity multiplied by the unit cost.
The costing identity
Total cost = Σ (quantity of resourcei × unit cost of resourcei)
For example, 640 referrals that each take clinician time valued at $45 cost 640 × $45 = $28,800.
Identification starts from the inputs and activities in the program's logic model (Lesson 3). The evaluator walks through each activity, from screening in the clinic to the last follow-up call, and asks what staff time, space, equipment, travel and partner resources it uses. The inventory should be agreed with staff and partners, who know about resources that never appear in the accounts, such as the volunteer who drives participants to their first group meeting.
Measurement records quantities in natural units, such as staff hours, referrals, rides funded and visits to other services, from payroll records, activity logs, case management systems, administrative data (Lesson 5) and participant surveys.
Valuation attaches a unit cost to each quantity. Staff time is valued at salary plus benefits per hour of work. Purchased items are valued at their market price. Physician services are often valued with the provincial fee schedule, which in British Columbia is the Medical Services Plan payment schedule, and hospital and emergency care can be valued with national cost estimates from the Canadian Institute for Health Information. Resources without a market price need a shadow price. Volunteer time can be valued at the wage of a paid worker doing the same task (the replacement cost method) or at what the volunteer would otherwise earn, and the choice changes the result. The value of retired participants' time is contested, and many analyses report their hours without a dollar value.
Start-up costs are incurred once, such as hiring, initial training, building the resource directory and co-designing the land-based pathway. Recurring costs are incurred every year. Report them separately, because continuing a program depends on recurring costs and adopting it elsewhere depends on both.
Equipment that lasts several years, such as laptops or a vehicle, is spread over its useful life as an equivalent annual cost, which accounts for depreciation and for the return the money could have earned.
Corporate services such as human resources, finance and information technology are allocated to the program, often as a percentage of direct costs or by staff numbers. State the allocation method, because it can change the result.
Costs that occur only because a program is being evaluated, such as research interviews or extra questionnaires, are excluded from the program cost and belong in the evaluation budget.
Costs from different years are expressed in the prices of one price year with an index such as the consumer price index from Statistics Canada, and reports state the currency and price year.
1.4 Gross Costing, Micro-Costing and Time-Driven Activity-Based Costing
Costing methods differ in how finely they measure resource use. Gross costing, also called top-down costing, divides a total expenditure by a measure of output. For the first full year of Cedar Valley, $840,000 divided by 500 participants who attended a first meeting gives a gross cost of $1,680 per participant. Gross costing is quick and uses data that already exist, but it treats every participant as identical and cannot explain why one clinic or one participant costs more than another.
Micro-costing, or bottom-up costing, measures each resource used to deliver the program to each participant and values it separately (Frick, 2009). It is more demanding, and it is the method of choice when a program is new and no unit costs exist, when costs are expected to vary across participants or sites, or when the evaluation needs to show which components drive the total. Micro-costing data can come from time-and-motion studies, activity logs kept by staff, case management records, or structured interviews with staff about how long tasks take.
Time-driven activity-based costing is a form of micro-costing developed in management accounting by Kaplan and Anderson (2004) and applied to health care by Kaplan and Porter (2011). It needs two estimates for each type of staff: the capacity cost rate, which is the cost of supplying that staff capacity divided by the practical capacity in hours, and the time each activity in the process requires. The cost of serving a participant is the time each activity takes multiplied by the capacity cost rate, summed along a process map of the participant's path through the program.
Capacity cost rate for a Cedar Valley connector
Cost of one connector position (salary and benefits) = $595,000 ÷ 7 = $85,000 a year
Paid hours = 37.5 hours × 52 weeks = 1,950 hours a year
Practical capacity, after vacation, statutory holidays, sick leave, training and team meetings = 1,700 hours a year (assumed)
Capacity cost rate = $85,000 ÷ 1,700 hours = $50.00 per hour
The Cedar Valley analyst asked connectors to keep an activity log for four weeks and reviewed case records for a sample of closed cases. Participants attended a mean of four meetings (the first plus three follow-ups) of the six allowed. The table and process map summarize mean connector time per participant.
| Activity | Mean connector time (hours) | Cost at $50.00 per hour |
|---|---|---|
| Referral triage and first telephone contact | 0.5 | $25 |
| First meeting and co-development of the connection plan | 1.5 | $75 |
| Follow-up meetings (mean of 3.0 meetings of 1.0 hour) | 3.0 | $150 |
| Linkage work: calls to groups, arranging transport, accompanied first visits | 2.0 | $100 |
| Travel to homes and community venues | 2.0 | $100 |
| Documentation and case review | 1.5 | $75 |
| Total per participant | 10.5 | $525 |
Direct connector time costs $525 per participant, which is less than a third of the gross cost of $1,680. The difference has three sources. The first is the shared costs that the process map does not include: the coordinator, the analyst, training, travel and data systems. The second is the transport fund and the partner grants, which pay for participants' links to the community and lie outside the connector's own work. The third is unused capacity. Seven connectors supply 7 × 1,700 = 11,900 practical hours a year, and 500 participants at 10.5 hours each used 5,250 hours, or 44.1 percent of that capacity. A time study would show how the remaining hours were spent, for example on partner development, clinic liaison, building the resource directory and supporting referred adults who never attended. Kaplan and Anderson (2004) argued that one advantage of time-driven costing is that it makes unused capacity visible, where a gross cost per participant hides it inside a single average.
1.5 Fixed, Variable, Average and Marginal Costs
Costs behave differently as the number of participants changes. Fixed costs, such as connector and coordinator salaries, the analyst and the data system, do not change with the number of participants in the short run. Variable costs, such as transport vouchers and clinician time for referrals, rise with each participant. Step costs are fixed over a range and then jump, as when a full caseload requires an eighth connector. The average cost is the total cost divided by the number of participants, and the marginal cost is the additional cost of serving one more participant.
When a program has spare capacity, the marginal cost of an extra participant is small, because the connectors are already paid. The average cost then falls as the caseload grows, since the fixed costs are spread over more people. Suppose a connector working at full caseload can serve 120 participants a year, which uses 120 × 10.5 = 1,260 hours and leaves 440 of the 1,700 practical hours for partner development, liaison and other work that does not belong to a single participant. Seven connectors could then serve 840 participants a year, and the gross cost per participant at full capacity would be $840,000 ÷ 840 = $1,000.
Costs measured in a start-up year overstate the steady-state cost per participant, while an analysis that assumes full capacity overstates efficiency if referrals never reach that level. A careful costing reports the observed cost, states the program's capacity, and presents the cost at a plausible steady-state caseload as a scenario, which Section 2 does for Cedar Valley.
Suppose that in year two the first-wave budget is unchanged at $840,000 and 588 older adults attend a first meeting. Calculate the gross cost per participant and the share of the 840-participant capacity in use. Then explain what the marginal cost of a 589th participant would mostly consist of, and at what caseload the health authority would face a step cost. (The gross cost is $840,000 ÷ 588 = $1,428.57, and 588 ÷ 840 = 70 percent of capacity. The marginal cost would consist mainly of variable items such as transport support and the clinician's referral time, since the connectors are already paid. A step cost would arise once demand exceeded about 840 participants a year, when an eighth connector would be needed.)
1.6 The Full Economic Cost of the First Wave
The table below brings the elements of this section together for the first full year of the first wave. The unit costs for overhead, clinician time, rooms and volunteer time are assumptions chosen for teaching, and a real costing would document the source of each one. The health authority applies a corporate overhead rate of 10 percent, clinicians spend time valued at $45 on each screening and referral, each of the 12 clinics provides meeting space valued at $1,600 a year, and community group volunteers contribute about 2,000 hours a year to welcoming and accompanying participants, valued at a replacement wage of $25 an hour.
| Resource | Calculation | Annual cost |
|---|---|---|
| Program budget | Seven connectors, coordinator, analyst, transport fund, partner grants, training, travel and data systems | $840,000 |
| Health authority overhead | 10% × $840,000 | $84,000 |
| Clinician screening and referral time | 640 referrals × $45 | $28,800 |
| Clinic meeting rooms provided in kind | 12 clinics × $1,600 | $19,200 |
| Health system total | Per participant: $972,000 ÷ 500 = $1,944 | $972,000 |
| Volunteer time in community groups | 2,000 hours × $25 | $50,000 |
| Societal total (valued items) | Per participant: $1,022,000 ÷ 500 = $2,044 | $1,022,000 |
The full health-system cost is about 16 percent higher than the budget ($972,000 ÷ $840,000 = 1.157). Some societal items were identified and not valued: participants' and caregivers' time and travel, the costs that community organizations bear beyond their partner grants, and any costs that the First Nations health centre bears in hosting a connector beyond the funded position. The evaluation plan proposes to measure participants' time and travel with a short questionnaire and to estimate partner costs in discussion with each partner, including the First Nations health centre, in a way that respects the partner's authority over its own information. Items that cannot be valued are listed in the impact inventory so that readers know what the totals omit.
The executive asks what the second wave will add to the annual budget. This is a question of affordability, which a budget impact analysis answers for a specific budget holder over one to five years (Sullivan et al., 2014), and it is separate from the question of value that Section 2 addresses. The team assumes that the 12 new clinics need seven more connectors ($595,000), a second coordinator ($105,000), and transport, partner grant, training and data funds at first-wave levels ($85,000), while the half-time analyst is shared. The second wave would add $785,000, for an annual budget across 24 clinics of $1,625,000. The team also notes the first-year start-up costs of hiring and training, including the $64,000 implementation support package for the second-wave clinics that Lesson 9 specified, and the new referrals that community organizations in the second-wave areas would receive.
With the cost of the program established from more than one perspective, the evaluation can ask whether the outcomes justify it. Section 2 combines the health-system cost of $1,944 per participant with estimates of the program's effects to calculate an incremental cost-effectiveness ratio and to judge how confident a decision-maker can be in the result.
Reflection
A health authority runs a community paramedicine program in which paramedics visit frail older adults at home. Its annual budget is $650,000: four paramedic positions ($440,000), a supervisor ($120,000), leased vehicles ($60,000), and equipment and supplies ($30,000). Other resources are used but not paid for by the program: dispatch staff in the provincial ambulance service schedule the visits (time valued at $25,000 a year), family physicians spend time valued at $40 per patient preparing care plans, the municipality provides space in a fire hall at no charge (valued at $8,000 a year), and family caregivers spend time at home visits. In the year, 400 patients received at least one visit. Calculate the cost per patient from the program perspective and from the publicly funded health system perspective, state which further items a societal perspective would add, and explain one reason the program-perspective cost could mislead a decision to expand the program to a second region.
From the program perspective, the cost is the budget, $650,000 ÷ 400 = $1,625 per patient. The health system perspective adds the dispatch staff time ($25,000), because the ambulance service is publicly funded, and the physicians' care-planning time (400 × $40 = $16,000). The total is $650,000 + $25,000 + $16,000 = $691,000, or $1,727.50 per patient.
A societal perspective would add the fire hall space ($8,000), because the municipality is outside the health system but the space has an opportunity cost, bringing the valued total to $699,000, or $1,747.50 per patient. It would also add family caregivers' time at visits, and patients' own time, which should at least be listed in an impact inventory even if they are not valued.
The program-perspective figure could mislead an expansion decision because it hides resources that a second region may not have. A new region might have no spare dispatch capacity and no free municipal space, so its budget would need to cover $33,000 of resources that the first region receives in kind. A strong answer might also note that the first-year cost includes start-up effects, or that cost per patient depends on whether the paramedics' caseloads are full.
Minimum 20 characters required.
Question 1: A clinic lends a meeting room to the Cedar Valley program at no charge. How should a health-system costing treat the room?
Question 2: Which perspective would count the time of volunteers in community groups and participants' own travel time?
Question 3: A Cedar Valley connector position costs $85,000 a year and provides 1,700 practical hours. A participant uses 10.5 hours of connector time. What is the connector cost per participant?
Question 4: Why is the marginal cost of one more Cedar Valley participant in year one much lower than the average cost of $1,680?
Economic Evaluation
Learning Objectives for this section
- Distinguish partial from full economic evaluations, and describe cost-minimization, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis.
- Explain how quality-adjusted life years combine length and quality of life, and calculate QALYs gained as the area between utility profiles.
- Calculate an incremental cost-effectiveness ratio and net monetary benefit, and place a result on the cost-effectiveness plane.
- Describe how scenario analyses, probabilistic sensitivity analysis and cost-effectiveness acceptability curves represent uncertainty.
- Explain the cautions that apply to social return on investment analyses, and identify the main reporting items in CHEERS 2022.
Section 1 established that the first wave of the fictional Cedar Valley Connector program costs the health system $1,944 per participant in its first full year. A cost on its own cannot tell a decision-maker whether the program is worth funding. Economic evaluation compares the costs and consequences of a program with those of an alternative, and it asks whether the additional health gained justifies the additional resources used.
2.1 Full and Partial Economic Evaluation
Drummond and colleagues (2015) define a full economic evaluation by two features: it examines both costs and consequences, and it compares two or more alternatives. Studies that lack one of these features are partial evaluations. A cost description of one program, an outcome description of one program, or a comparison of the costs of two programs without their outcomes are all partial evaluations, which can be useful steps toward a full one. Full economic evaluations differ in how they measure consequences.
| Type | How consequences are measured | Summary result | When it is used |
|---|---|---|---|
| Cost-minimization analysis | Consequences shown or assumed to be equivalent | Difference in cost | Rarely, because equivalence must be demonstrated |
| Cost-effectiveness analysis | A single natural unit, such as an older adult no longer lonely | Cost per unit of effect gained | Comparing programs with the same main outcome |
| Cost-utility analysis | Quality-adjusted life years (QALYs) | Cost per QALY gained | Comparing programs across health areas |
| Cost-benefit analysis | Monetary value of consequences | Net benefit or benefit-cost ratio | Comparing health and non-health investments |
| Cost-consequence analysis | Several outcomes reported side by side without aggregation | A table of costs and outcomes | Complex programs with many outcomes |
Cost-consequence analysis is common in public health, where programs affect several outcomes that do not combine easily into one measure. It leaves the weighing of outcomes to the reader, which is transparent but provides no single decision rule. Section 3 shows how an evaluative rubric can structure that weighing.
2.2 The Incremental Cost-Effectiveness Ratio
Economic evaluation compares a program with what would happen without it, which is usually usual care. The key quantity is the incremental cost-effectiveness ratio (ICER), the difference in cost divided by the difference in effect (Weinstein & Stason, 1977).
Incremental cost-effectiveness ratio
ICER = (C1 − C0) ÷ (E1 − E0) = ΔC ÷ ΔE
Here C1 and E1 are the mean cost and effect per person with the program, and C0 and E0 are the mean cost and effect with the comparator. The ICER is the additional cost of each additional unit of effect.
The ratio is incremental because the decision concerns the additional money and the additional effect. An average ratio, such as total cost divided by the number of people whose loneliness improved, attributes every improvement to the program, including improvements that would have happened anyway.
Both increments come from the evaluation design. For this lesson, assume that a comparison-group evaluation of the first wave (Lesson 7) produced the following illustrative results per participant over twelve months. Emergency department visits were 0.12 lower and primary care visits 0.40 lower than in the comparison group. These figures count participants’ own visits; the Lesson 7 exercise that found a rise in recorded primary care visits across all older patients of first-wave clinics measured a different quantity, which may include visits that connectors arranged or recorded. At assumed unit costs of $450 per emergency visit and $55 per primary care visit, these differences offset 0.12 × $450 + 0.40 × $55 = $54 + $22 = $76 of the program cost. The incremental cost is therefore ΔC = $1,944 − $76 = $1,868 per participant.
At twelve weeks, 34 percent of program participants and 22 percent of comparison participants scored below the referral threshold of 6 on the three-item UCLA Loneliness Scale (illustrative values). The incremental effect is 0.34 − 0.22 = 0.12, or 12 additional older adults no longer screening as lonely for every 100 participants. For 100 participants, the incremental cost is 100 × $1,868 = $186,800, so the ICER is $186,800 ÷ 12 = $15,567 per additional older adult no longer screening as lonely.
This result is easy to explain, but a decision-maker cannot tell from it whether $15,567 is good value, because no other program reports its results in the same unit and the measure ignores effects on health beyond loneliness.
2.3 Cost-Utility Analysis and Quality-Adjusted Life Years
Cost-utility analysis addresses the problem of incomparable units by measuring consequences in a common unit, the quality-adjusted life year (QALY). A QALY weights each period of life by a utility value that represents health-related quality of life on a scale on which 1 is full health and 0 is a state equivalent to death. States judged worse than death have negative values. One year in full health is one QALY, and one year at a utility of 0.70 is 0.70 QALYs (Weinstein et al., 2009).
Utilities are elicited from people's preferences between health states. In the standard gamble, a respondent chooses between a certain health state and a gamble between full health and death. In the time trade-off, the respondent chooses between a longer life in the health state and a shorter life in full health. Torrance and colleagues developed the time trade-off at McMaster University in the early 1970s, and Torrance (1986) reviewed the methods. Most evaluations use a generic instrument with a preference-based scoring algorithm, such as the EQ-5D-5L, which has a Canadian value set derived from the time trade-off (Xie et al., 2016), the Health Utilities Index, or the SF-6D.
QALYs gained
QALYs = Σ (utility in each period × length of the period in years)
QALYs gained by a program are the area between the utility profiles of the program and comparison groups over the time horizon. With measurements at a few time points, the area is usually computed with the trapezoid rule, assuming that utility changes in a straight line between measurements.
Suppose Cedar Valley participants complete the EQ-5D-5L at intake, twelve weeks and twelve months. In the illustrative analysis, the groups have the same mean utility at intake, the program group's mean utility is 0.03 higher at twelve weeks, and the difference has returned to zero by twelve months. The area between the profiles is a triangle with a base of one year and a height of 0.03, so the incremental QALYs are ΔE = 0.5 × 1 × 0.03 = 0.015 per participant. If the difference of 0.03 were instead sustained to twelve months, the area would be 0.5 × 0.03 × (12 ÷ 52) + 0.03 × (40 ÷ 52) = 0.0035 + 0.0231 = 0.0265 QALYs.
The base-case ICER is ΔC ÷ ΔE = $1,868 ÷ 0.015 = $124,533 per QALY gained. The figure shows why the assumption about persistence matters: no one measured utility between twelve weeks and twelve months, so the shape of the profile in that interval is a modelling choice that Section 2.5 tests.
QALYs have known limits for programs such as Cedar Valley. Generic instruments such as the EQ-5D-5L describe mobility, self-care, usual activities, pain and anxiety or depression, and they may miss changes in social connection that participants value. Capability measures such as the ICECAP-O, developed for older people (Coast et al., 2008), assess attachment, security, role, enjoyment and control, and some evaluations of social programs report them alongside QALYs. QALYs also raise equity questions, since a QALY is valued equally whoever gains it, and the Cedar Valley evaluation reports results for subgroups so that readers can see how gains and costs are distributed.
2.4 The Cost-Effectiveness Plane and Net Monetary Benefit
The cost-effectiveness plane plots the incremental effect on the horizontal axis and the incremental cost on the vertical axis, with the comparator at the origin (Black, 1990). Its four quadrants classify results.
A line through the origin with a slope equal to the decision-maker's willingness to pay per QALY, written λ (lambda), divides the north-east quadrant. Results below the line are cost-effective at that threshold. Canada has no official threshold. The CADTH (2017) guidelines ask analysts to show results across a range of willingness-to-pay values, and Canadian studies often report results against reference values such as $50,000 and $100,000 per QALY. The figure places the illustrative Cedar Valley results on the plane.
Ratios are awkward to analyze, because the same ICER can arise in the north-east and south-west quadrants and because a ratio becomes unstable when ΔE is near zero. Net monetary benefit avoids these problems by converting health gains into money at the threshold value (Stinnett & Mullahy, 1998).
Net monetary benefit
NMB = λ × ΔE − ΔC
At λ = $50,000, the Cedar Valley base case gives NMB = $50,000 × 0.015 − $1,868 = $750 − $1,868 = −$1,118. At λ = $100,000, NMB = $1,500 − $1,868 = −$368. A program is cost-effective at a threshold when its NMB is positive, which is the same as an ICER below λ in the north-east quadrant.
2.5 Representing Uncertainty
Every input to the Cedar Valley analysis is uncertain. Deterministic sensitivity analysis changes one input, or a set of inputs that form a scenario, and recalculates the result. Probabilistic sensitivity analysis assigns each uncertain parameter a probability distribution, draws a value for every parameter many times, and recalculates the result for each draw, so that the spread of results reflects the joint uncertainty in all parameters (Briggs et al., 2006). The share of draws with a positive NMB at each threshold is plotted as a cost-effectiveness acceptability curve (van Hout et al., 1994). Probabilistic analysis handles uncertainty in parameter values. Uncertainty about structure, such as whether effects persist, is better shown with scenarios, because the answer changes the model itself.
Analyses with time horizons longer than one year also discount future costs and QALYs to present values, and the CADTH (2017) guidelines specify a rate of 1.5 percent a year for the reference case. The Cedar Valley analysis uses a one-year horizon, so it needs no discounting.
Scenario analyses. The base case uses the inputs from Section 1 and this section. The QALY gain of 0.015 is the area between the utility curves, computed with the trapezoid rule from a utility difference that rises to 0.03 at twelve weeks and returns to zero at twelve months. With an incremental cost of $1,868, the ICER is $124,533 per QALY, and the net monetary benefit is −$1,118 at $50,000 per QALY and −$368 at $100,000 per QALY. Five scenarios change one assumption, or a pair of assumptions, and recalculate the result.
| Scenario | Incremental cost | Incremental QALYs | ICER (dollars per QALY) |
|---|---|---|---|
| Base case | $1,868 | 0.0150 | $124,533 |
| Effect persists to 52 weeks | $1,868 | 0.0265 | $70,388 |
| Full caseloads (840 a year) | $1,104 | 0.0150 | $73,630 |
| No health care offsets | $1,944 | 0.0150 | $129,600 |
| Societal perspective | $1,968 | 0.0150 | $131,200 |
| Persistence and full caseloads | $1,104 | 0.0265 | $41,617 |
The scenarios show which assumptions matter. Removing the health care offsets or adding valued volunteer time changes the ICER only modestly ($129,600 and $131,200 per QALY). Persistence of the effect ($70,388) and full caseloads ($73,630, with an incremental cost of $1,104) each bring the ICER below $100,000, and together they bring it to $41,617, below $50,000. The program's value depends mainly on two things the first year cannot show, whether benefits last and whether caseloads fill.
Probabilistic sensitivity analysis. The probabilistic analysis draws 5,000 sets of inputs: the cost per participant from a gamma distribution (mean $1,944, standard deviation $120), the differences in emergency and primary care visits from normal distributions, and the twelve-week utility difference from a normal distribution with mean 0.03 and standard deviation 0.012. Across the draws, the mean incremental cost is $1,865.60 and the mean QALY gain is 0.0151. The table gives the share of draws with a positive net monetary benefit at each willingness to pay, which is the cost-effectiveness acceptability curve.
| Willingness to pay per QALY | Probability cost-effective |
|---|---|
| $25,000 | 0.000 |
| $50,000 | 0.000 |
| $75,000 | 0.062 |
| $100,000 | 0.282 |
| $125,000 | 0.511 |
| $150,000 | 0.665 |
| $200,000 | 0.830 |
Almost all draws (99.4 percent) lie in the north-east quadrant, and 0.6 percent lie in the north-west, where the program is dominated. The probability that the program is cost-effective is 0.000 at $50,000 per QALY, 0.282 at $100,000 and 0.511 at $125,000, close to the base-case ICER. Under base-case assumptions, a decision-maker willing to pay $100,000 per QALY would face about a 28 percent chance that the program is good value, and the scenario analyses show that this chance would rise considerably if effects persist or caseloads fill.
2.6 Cost-Benefit Analysis and Social Return on Investment
Cost-benefit analysis values all consequences in money and reports the net benefit (benefits minus costs) or the benefit-cost ratio. Health effects can be valued by willingness to pay, elicited through contingent valuation or discrete choice experiments, or with a monetary value per QALY. If the Cedar Valley health gain is valued at $100,000 per QALY, the benefits per participant are 0.015 × $100,000 + $76 of averted health care costs = $1,576. Against the gross health-system cost of $1,944, the net benefit is $1,576 − $1,944 = −$368 and the benefit-cost ratio is $1,576 ÷ $1,944 = 0.81. The net benefit equals the NMB at λ = $100,000, which shows that a cost-utility analysis with a threshold is a restricted form of cost-benefit analysis.
Social return on investment (SROI) is a form of cost-benefit analysis developed in the social enterprise sector and set out in a guide first published by the United Kingdom Cabinet Office in 2009 and revised by the SROI Network in 2012 (Nicholls et al., 2012). It maps outcomes with interest holders, assigns financial proxies to each outcome, and adjusts for deadweight (what would have happened anyway), attribution (the share due to others), displacement and drop-off. The result is a ratio of social value to investment. SROI has appeal for community programs because it involves interest holders in deciding which outcomes count. A systematic review of SROI in public health found wide variation in methods and in the quality of the proxies and adjustments used (Banke-Thomas et al., 2015).
Suppose a hypothetical SROI of Cedar Valley assigns a financial proxy of $4,000 to each participant whose loneliness score falls by at least one point, and 300 of the 500 participants meet that criterion. The claimed social value is 300 × $4,000 = $1,200,000, and the ratio to the $840,000 budget is 1.43 to 1. If deadweight is set at 50 percent, because the comparison group suggests that half of the improvement would have happened anyway, the value falls to $600,000 and the ratio to 0.71 to 1. Using the full health-system cost of $972,000 as the investment, the ratio falls further to 0.62 to 1. The proxy itself is a judgement that no data in the evaluation can confirm.
Proxies are often drawn from value banks or from unrelated contexts, and a ratio is only as credible as its least defensible proxy. Reports should give the source of each proxy and test alternatives.
Many SROI analyses estimate deadweight from general statistics or judgement instead of a comparison group, which can inflate the claimed value. The designs of Lessons 6 to 8 provide a stronger basis for the adjustment.
Counting both a reduction in loneliness and the improved wellbeing that follows from it values the same change twice. Outcome maps should separate final outcomes from intermediate ones.
Because methods vary so widely, a ratio from one SROI cannot be compared with a ratio from another, and a ratio above one does not show that a program is a better use of money than alternatives.
2.7 Reporting with CHEERS 2022
The Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) replaced the 2013 statement with a 28-item checklist for reporting any health economic evaluation (Husereau et al., 2022). Items cover the title and abstract, the setting and comparators, the perspective, time horizon and discount rate, the selection, measurement and valuation of outcomes and resources, the currency and price date, the rationale and description of any model, and the characterization of uncertainty. CHEERS 2022 added items asking authors to state whether a health economic analysis plan was prepared, to describe how distributional effects were examined, and to report how patients, the public and others affected by the study were engaged and what difference that engagement made. Distributional cost-effectiveness analysis, which examines how costs and health gains fall across social groups, is one way to address the distributional item (Cookson et al., 2017).
For each of the following CHEERS topics, write one sentence stating what the illustrative Cedar Valley analysis reports or omits: perspective, comparator, time horizon, discount rate, outcome measure and its valuation, currency and price year, characterization of uncertainty, distributional effects, and engagement with older adults and First Nations partners. (The analysis reports a health-system perspective with a societal scenario, usual care in comparison clinics, a one-year horizon with no discounting, EQ-5D-5L utilities with the Canadian value set, scenario and probabilistic analyses, and subgroup results. It has not yet stated a price year, and it should describe how the steering committee's older adults and First Nations representatives shaped the outcomes and the interpretation.)
An ICER, a probability of cost-effectiveness and an SROI ratio are each one line of evidence about a program. Section 3 turns to the task of combining such evidence with evidence on reach, cultural safety and effectiveness into an overall judgement.
Reflection
A health authority is considering a home-based exercise program to prevent falls among adults aged 75 and older. Compared with usual care, the program costs $900 per participant and reduces hospital admissions for falls enough to save $300 per participant over one year. The program gains 0.008 quality-adjusted life years (QALYs) per participant over the same year. A probabilistic sensitivity analysis found that the probability the program is cost-effective is 0.35 at a willingness to pay of $50,000 per QALY and 0.70 at $100,000 per QALY. A community agency has separately published a social return on investment (SROI) analysis claiming $5 of social value for every $1 invested, without a comparison group. Calculate the incremental cost, the incremental cost-effectiveness ratio (ICER) and the net monetary benefit at $50,000 and at $100,000 per QALY, state where the result lies on the cost-effectiveness plane, interpret the probabilistic results, and explain how you would treat the SROI claim in advice to the health authority.
The incremental cost is $900 − $300 = $600 per participant, and the ICER is $600 ÷ 0.008 = $75,000 per QALY gained. The result lies in the north-east quadrant, because the program costs more and gains health. The net monetary benefit is $50,000 × 0.008 − $600 = $400 − $600 = −$200 at $50,000 per QALY, and $800 − $600 = $200 at $100,000 per QALY. The program is therefore cost-effective at $100,000 per QALY and not at $50,000.
The probabilistic results agree with this. At $50,000 per QALY only 35 percent of simulations show a positive net benefit, while at $100,000 the figure is 70 percent, so a decision-maker willing to pay $100,000 would face about a 30 percent chance that the program is poor value. I would present these probabilities together with the scenario analyses that drive them.
I would not treat the SROI ratio as comparable evidence. Without a comparison group, its deadweight adjustment is a judgement, and the ratio depends on financial proxies whose sources should be checked. I would ask for the proxies, the deadweight and attribution assumptions and a test of alternatives, and I would note that a ratio above one does not show that the program is better value than other uses of the money.
Minimum 20 characters required.
Question 1: What distinguishes a full economic evaluation from a partial one?
Question 2: Compared with usual care, a program costs $1,868 more per participant and gains 0.015 QALYs. What is the ICER, and where does the result lie on the cost-effectiveness plane?
Question 3: At a willingness to pay of $100,000 per QALY, a program has an incremental cost of $1,868 and an incremental effect of 0.015 QALYs. What is its net monetary benefit, and what does it imply?
Question 4: Which statement about social return on investment (SROI) is most accurate?
Synthesis and Judgement
Learning Objectives for this section
- Distinguish descriptive findings from evaluative conclusions about the merit and worth of a program.
- Construct an analytic rubric with criteria, performance levels, descriptors, evidence sources and importance for a program.
- Compare qualitative weight-and-sum, numerical weight-and-sum, hurdle and profile approaches to synthesis, and explain the risks of numerical weighting.
- Synthesize convergent, complementary and dissonant evidence into a judgement with a stated level of confidence.
- Describe a program with the TIDieR checklist so that others can understand, cost and replicate it.
Sections 1 and 2 produced a cost and an economic result for the fictional Cedar Valley Connector program. The evaluation has also produced evidence on reach, cultural safety and changes in loneliness. Decision-makers asked a single evaluative question, whether the program is good enough and good value enough to extend to the second wave, and answering it requires a judgement that brings these lines of evidence together. Lesson 1 introduced evaluative reasoning as a movement from criteria to standards to evidence to synthesis. This section develops the tools that make that synthesis explicit.
3.1 From Findings to Evaluative Conclusions
A descriptive finding states what happened: 78.1 percent of referred adults attended a first meeting. An evaluative conclusion states how good that is: reach was good against the standard the steering committee agreed in advance. Davidson (2005) argues that evaluation is distinguished from other applied research by its obligation to draw conclusions of this second kind, and that an evaluation which reports only descriptive findings leaves the hardest part of the work to readers who have less information than the evaluator.
Scriven's distinctions from Lesson 1 shape the conclusions an evaluation can reach. Merit is the intrinsic quality of a program, judged against criteria such as effectiveness and quality of delivery. Worth is its value in a particular context, which brings in cost, need and the alternatives available to the decision-maker. A program can have good merit and modest worth, as when it works well but costs more than an equally useful alternative. Keeping the two separate lets an evaluation report that the Cedar Valley program delivers what it promises while being candid about its value for money at current caseloads.
Synthesis happens at two levels. Within a criterion, several pieces of evidence are combined into a rating, as when a difference-in-differences estimate, participant interviews and connector records together determine the rating for effectiveness. Across criteria, the ratings are combined into an overall judgement. Each level needs a stated rule, and the rules should be agreed with the primary intended users (Lesson 4) before the evidence arrives.
3.2 Building an Evaluative Rubric
An evaluative rubric sets out the criteria in rows and describes what performance at each level looks like for each criterion. King and colleagues (2013) describe rubrics as a method for surfacing the values of interest holders and for making the basis of a judgement transparent, and Oakden (2013) gives a practical account of developing them with program partners. An analytic rubric rates each criterion separately, while a holistic rubric describes overall levels of performance in a single set of descriptors. Analytic rubrics are more common in program evaluation because they show where a program is strong and where it is weak.
Rubrics are developed in steps. The evaluator drafts criteria from the logic model and the evaluation questions, holds a workshop in which interest holders revise the criteria and draft the descriptors, specifies the evidence for each criterion, agrees the importance of each criterion and any hurdles, and tests the rubric on hypothetical results to check that the descriptors discriminate. For Cedar Valley, the steering committee, including its four older adults with lived experience and its two First Nations representatives, built on the six-month rubric of Lesson 1 to agree the year-one rubric below. The value-for-money criterion follows King (2017), who argues that economic results should enter an evaluation as evidence rated against agreed standards, alongside the other criteria.
| Criterion and importance | Excellent | Good | Adequate | Poor |
|---|---|---|---|---|
| Reach: percentage of referred adults who attend a first meeting (very important) | 80 percent or more | 65 to 79 percent | 50 to 64 percent | Below 50 percent |
| Equity of reach across sex, language, rurality and Indigenous identity (very important) | All groups within 5 percentage points | All groups within 10 points | Gaps over 10 points, with an agreed plan to close them | Gaps over 10 points, with no plan |
| Cultural safety, as judged by Indigenous participants and partners (hurdle and very important) | Safe and respectful, with no unresolved concerns | Generally safe, with concerns resolved promptly | Mixed reports, with a plan in place | Disrespect or harm without a response |
| Effectiveness: program-attributable reduction in mean loneliness at twelve weeks (extremely important) | 0.5 points or more, with a confidence interval excluding zero | 0.3 to 0.49 points, with a confidence interval excluding zero | 0.1 to 0.29 points, or a larger estimate with an interval including zero | Below 0.1 points |
| Value for money (very important) | Dominant, or an ICER below $50,000 per QALY with a probability of at least 0.7 of being cost-effective at that value | Base-case ICER below $100,000 per QALY | Base-case ICER above $100,000, with plausible scenarios below $100,000 | Dominated, or an ICER above $100,000 in all plausible scenarios |
The committee agreed two synthesis rules. Cultural safety is a hurdle: the program cannot be rated better than adequate overall if cultural safety is rated poor. The overall judgement of merit cannot be higher than the rating for effectiveness, because the program exists to reduce loneliness. The committee also agreed to report merit and worth separately, with value for money informing worth.
3.3 Approaches to Synthesis Across Criteria
Several approaches exist for combining criterion ratings into an overall judgement, and they can produce different conclusions from the same ratings. Scriven (1991) and Davidson (2005) recommend qualitative approaches for most program evaluations.
Each criterion is assigned an importance category, such as extremely important, very important or important, and each is rated on the rubric. The evaluator then compares the pattern of ratings within each importance category, giving most attention to the most important criteria, and states the overall judgement in words with the reasoning shown. The method keeps the reasoning visible and avoids arithmetic on ordinal ratings (Scriven, 1991; Davidson, 2005).
Each rating is converted to a number (for example, excellent 4, good 3, adequate 2 and poor 1), multiplied by a weight, and summed. With weights of 0.30 for effectiveness, 0.20 each for cultural safety and value for money, and 0.15 each for reach and equity, a hypothetical program rated excellent on every criterion except a poor rating for cultural safety scores 0.30 × 4 + 0.15 × 4 + 0.15 × 4 + 0.20 × 1 + 0.20 × 4 = 3.40, which reads as better than good. The method lets strong ratings compensate for a harmful one, treats ordinal ratings as if they were measurements, and gives an impression of precision that the weights do not support.
A hurdle, or bar, is a minimum level that a program must reach on a criterion before other criteria can raise the overall rating. Cedar Valley's cultural safety hurdle records the committee's view that a program which harms some participants cannot be rescued by good results for others. Hurdles are combined with one of the other approaches.
Some evaluations report a profile of ratings without an overall judgement. A profile is appropriate when audiences legitimately weigh the criteria differently, for example when a ministry, a health authority and a community partner will each make their own decision. The evaluator then explains the implications of the profile for each audience.
3.4 Synthesizing Mixed Evidence
Within a criterion, evidence from different sources can relate in three ways (Greene, 2007). It can converge, when sources point to the same conclusion. It can be complementary, when sources describe different aspects of the same phenomenon, as when a survey measures how much loneliness changed and interviews describe how. It can be dissonant, when sources disagree. Lesson 5 introduced joint displays, which set quantitative and qualitative results side by side for each question and make these relationships visible.
Dissonance deserves investigation. The evaluator asks whether the sources measure the same thing, cover the same people and time, and are equally credible, and then uses the program theory of Lesson 3 to ask whether a mechanism might operate in some contexts and not others. Averaging dissonant evidence into a middle rating hides the most useful information the evaluation has produced.
Every rating should carry a statement of confidence. Confidence depends on the strength of the design, the consistency of evidence across sources, the precision of estimates, how directly the evidence addresses the criterion, and the risk of bias in each source. This reasoning resembles the GRADE approach to the certainty of evidence (Guyatt et al., 2008), applied here to one program. Many evaluations rate confidence as high, moderate or low and give the main reason for each rating.
The following year-one results are illustrative. Reach was 500 ÷ 640 = 78.1 percent, which is good, with high confidence because it comes from complete program records. Equity of reach showed attendance of 64 percent among referred adults whose first language is neither English nor French, compared with 80 percent among English or French speakers, and 71 percent in rural areas compared with 80 percent in urban areas. The gaps exceed 10 points, and the committee has agreed a plan of interpreter-supported first meetings and a rural transport pilot, so the rating is adequate, with high confidence.
Cultural safety was rated good: talking circles and the review with First Nations partners described the program as generally safe, and a concern about meeting locations was resolved within a month. Confidence is moderate, because the talking circles reached a small number of participants.
Effectiveness drew on three sources. The difference-in-differences estimate was a 0.4-point reduction in mean loneliness at twelve weeks (95 percent confidence interval 0.1 to 0.7), which is good. Interviews converged with this result for most participants, who described new routines and group memberships, and connector records complemented it by showing which linkages were made. The evidence was dissonant for rural participants, who were satisfied with their connectors and yet showed little change in loneliness. Program records explained the dissonance: the transport fund was exhausted in the ninth month, and the groups to which rural participants were linked were often a long drive away, so the mechanism of attending community groups was blocked by context. Effectiveness is rated good with moderate confidence, because the design is non-randomized, although trends in emergency and primary care visits were parallel before launch and the sources mostly agree.
Value for money was rated adequate: the base-case ICER is $124,533 per QALY, plausible scenarios range from $41,617 to $131,200, and the probability of cost-effectiveness at $100,000 per QALY is 0.282. Confidence is low, because persistence of the effect beyond twelve weeks has not been measured.
The cultural safety hurdle is passed. The extremely important criterion, effectiveness, is good, and so the committee rated the program's merit as good. Among the very important criteria, reach and cultural safety are good and equity of reach and value for money are adequate, so the committee rated the program's worth at year-one caseloads as adequate. The recommendation to the executive was to proceed with the second wave on three conditions: fill caseloads before adding positions, replenish the transport fund and pilot rural transport, and measure loneliness and utility at twelve months to test whether effects persist.
Revising descriptors once results are known, so that a disappointing result earns a better rating, undermines the credibility of the whole evaluation. If a standard proves unrealistic, the change and its reason should be reported.
An ICER or an effect estimate looks more authoritative than interview evidence, but precision is a different property from relevance. A precise estimate of a narrow outcome should not override strong evidence on a criterion it does not address.
When a criterion cannot be rated because data are missing, the rubric should record it as not rated, with the reason, instead of assigning the lowest level.
A good average can conceal poor results for a subgroup. Ratings for equity criteria, and subgroup results within other criteria, protect against this.
3.5 Describing the Intervention with TIDieR
A judgement applies to a program as it was actually delivered, and readers can use the judgement only if they know what that program was. Glasziou and colleagues (2008) found that descriptions of interventions in published trials and reviews were often too incomplete for the intervention to be replicated. The Template for Intervention Description and Replication (TIDieR) is a 12-item checklist for describing interventions in reports and protocols (Hoffmann et al., 2014), and TIDieR-PHP adapts it for population health and policy interventions (Campbell et al., 2018). A full description also serves the other parts of this lesson: costing depends on the dose and the staff involved, and an adopter needs both the description and the cost to judge whether the program could work in their setting.
| TIDieR item | Cedar Valley Connector program (fictional) |
|---|---|
| 1. Brief name | Cedar Valley Connector program, a community connector (social prescribing) program for older adults. |
| 2. Why | Loneliness and social isolation harm health, and many isolated older adults face barriers to joining community activities that one-to-one support and practical help can reduce. |
| 3. What: materials | Screening card with the three-item UCLA Loneliness Scale, connection plan template, community resource directory, transport fund and partner grants. |
| 4. What: procedures | Clinician screening and referral, triage call, first meeting with a co-developed plan, linkage to groups and services, follow-up meetings and closure. |
| 5. Who provided | Seven community connectors, one hosted by a First Nations health centre, supported by a coordinator and trained in person-centred planning and cultural safety. |
| 6. How | Face to face, one to one, with telephone follow-up. |
| 7. Where | Twelve first-wave primary care clinics, participants' homes, community venues and the First Nations health centre. |
| 8. When and how much | Up to six meetings over twelve weeks; a mean of four meetings and 10.5 hours of connector time per participant in year one. |
| 9. Tailoring | Each plan is co-developed with the participant, and a land-based connection pathway is being co-designed with one First Nation. |
| 10. Modifications | Changes during delivery are documented with the FRAME approach described in Lesson 9. |
| 11. How well: planned | Fidelity is monitored with a checklist of core components and monthly case review. |
| 12. How well: actual | Reported from the process evaluation: reach, meetings attended, linkages made and fidelity to core components. |
A report describes a program as follows: "Participants received peer support sessions from trained volunteers at community centres." Using the 12 TIDieR items, list the information a reader would need in order to cost and replicate the program. (A strong answer notes the missing rationale, the materials and procedures of a session, the training and background of the volunteers, whether sessions were individual or group and in person or remote, the number, length and frequency of sessions, tailoring, modifications, and planned and actual fidelity.)
A judgement of merit and worth, a statement of confidence and a clear description of the program are the core content of an evaluation report. Section 4 asks how that content should be written, shown and shared so that decision-makers and communities can use it.
Reflection
A peer support program for new parents was evaluated against a rubric agreed with its advisory group. Criteria, importance and levels were as follows. Effectiveness (extremely important): reduction in mean Edinburgh Postnatal Depression Scale score compared with a comparison group, rated excellent at 2.0 points or more, good at 1.0 to 1.9, adequate at 0.5 to 0.9 and poor below 0.5. Reach among low-income parents (very important): share of participants with low income, rated excellent at 60 percent or more, good at 50 to 59 percent, adequate at 35 to 49 percent and poor below 35 percent. Safety (hurdle and very important): good if all safeguarding incidents were handled promptly under the protocol, and poor if any was not; a poor rating caps the overall rating at adequate. Value for money (very important): excellent if the ICER is below $50,000 per QALY with a probability of at least 0.7 of being cost-effective at that value, good if below $100,000, adequate if above $100,000 with plausible scenarios below it, and poor otherwise. The group also agreed that overall merit cannot be rated higher than effectiveness. Evidence: a non-randomized comparison with similar baseline scores found a reduction of 1.5 points (95 percent confidence interval 0.6 to 2.4); 45 percent of participants had low income; one safeguarding incident occurred and was handled under the protocol within a week; the ICER was $38,000 per QALY with a probability of 0.75 of being cost-effective at $50,000. In interviews with 30 parents, most described feeling less alone, but parents who joined more than three months after birth described little benefit. Rate each criterion with a level of confidence, state overall judgements of merit and worth, and explain how you would handle the interview evidence from late joiners.
Effectiveness is good (1.5 points, with an interval that excludes zero), with moderate confidence because the comparison is non-randomized although baseline scores were similar and the interviews mostly converge. Reach among low-income parents is adequate (45 percent), with high confidence because it comes from program records. Safety is good, so the hurdle is passed, with moderate confidence because rare incidents may go unreported. Value for money is excellent ($38,000 per QALY with a probability of 0.75 at $50,000), with moderate confidence because it inherits the uncertainty of the effect estimate.
Merit is good, the highest rating the effectiveness rule allows. Worth is good: the program is good value and safe, and its main weakness is reach among the parents who may need it most.
The late-joiner evidence is dissonant with the overall effect, and I would investigate it before combining it with the other evidence. I would check whether late joiners differ in baseline scores or circumstances, and whether the program theory, in which peer support reduces isolation during the early weeks of parenthood, implies that timing matters. If program records show smaller score changes among late joiners, I would report the subgroup pattern and recommend earlier referral, for example at the postnatal discharge visit, together with outreach to low-income parents.
Minimum 20 characters required.
Question 1: In Scriven's terms, what is the difference between the merit and the worth of a program?
Question 2: A rubric converts ratings to numbers and sums weighted scores across criteria. What is the main risk of this approach?
Question 3: Rural Cedar Valley participants were satisfied with their connectors, yet their loneliness scores changed little. What is the best response in the synthesis?
Question 4: What does the TIDieR item "when and how much" ask for in a description of the Cedar Valley program?
Reporting and Use
Learning Objectives for this section
- Plan an evaluation report that separates findings, conclusions and recommendations and answers the evaluation questions in order.
- Write main messages and an executive summary for decision-makers, using the 1:3:25 format as a guide.
- Apply principles of data visualization to present evaluation findings to decision-makers.
- Plan knowledge translation for different audiences, including communities that hold rights over their data.
- Describe types of evaluation use and the conditions under which evaluations are used or misused.
- Describe an evaluation work plan, budget and risk register, and explain how an evaluation proposal assembles the parts of an evaluation plan behind an executive summary for decision-makers.
An evaluation that reaches a sound judgement has done most of its work, and it can still fail if the judgement never reaches the people who need it, arrives after the decision, or is written in a form they cannot use. This section covers the report, the executive summary, the presentation of data, knowledge translation and the conditions under which evaluations are used. It ends with the work plan, budget and risk register of an evaluation and a worked example of a complete evaluation proposal for the fictional Cedar Valley Connector program.
4.1 The Evaluation Report
The full report is the reference document of an evaluation. It records what was evaluated, the questions asked, the methods used, the evidence and its limits, and the judgements reached, in enough detail that a reader can check the reasoning. The Program Evaluation Standards introduced in Lesson 1 apply directly: the accuracy standards call for conclusions that are justified by the evidence and reasoning, the utility standards call for timely and appropriate communication, and the propriety standards call for full and fair disclosure of findings, including unwelcome ones (Yarbrough et al., 2011).
The title page, acknowledgements (including the contributions of participants, partners and the steering committee), main messages and the executive summary. Main messages and the executive summary are often the only parts decision-makers read, so they are written last and with the most care.
A description of the program as delivered, following TIDieR (Section 3), with its logic model and theory of change (Lesson 3) and the context in which it operates.
The purpose of the evaluation, its primary intended users, the evaluation questions in priority order (Lesson 4), and the evaluative criteria and rubric agreed with interest holders.
The design and its justification (Lessons 6 to 8), data sources and indicators (Lesson 5), implementation measures (Lesson 9), the costing and economic methods (Sections 1 and 2), analysis and synthesis methods, and the ethics review and data governance arrangements.
Findings organized by question, in the same order as the questions, so that a reader looking for the answer to one question finds all the relevant evidence in one place.
Evaluative conclusions for each question with a level of confidence, recommendations that follow from the conclusions, and the limitations that affect how far the conclusions can be trusted.
The evaluation matrix, instruments, detailed tables, the technical economic appendix with a completed CHEERS 2022 checklist, and the rubric with the evidence used for each rating.
Three kinds of statement should be kept distinct throughout. A finding reports evidence: attendance was 64 percent among adults whose first language is neither English nor French. A conclusion makes a judgement against a standard: equity of reach was adequate. A recommendation proposes an action that follows from one or more conclusions: offer interpreter-supported first meetings in all clinics before the second wave. Readers can then trace each recommendation back to its evidence. Good recommendations are specific, feasible, and linked to the conclusions that justify them. They name who would act, and they are developed with the primary intended users so that they fit the resources and authority those users have. Patton (2008) advises that recommendations focus on actions within the control of intended users and state the costs, benefits and challenges of carrying them out.
4.2 Main Messages and Executive Summaries
Background: the 1:3:25 format and Lavis's five questions
Two planning tools from the knowledge translation literature recur in this section and the next. The Canadian Health Services Research Foundation, whose work continues in Healthcare Excellence Canada, promoted a reader-friendly format for reports to decision-makers known as 1:3:25: one page of main messages, a three-page executive summary, and a report of no more than 25 pages, with technical detail in appendices. The proportions matter less than the principle that each layer must stand on its own for a reader who goes no further. Lavis and colleagues (2003) proposed five questions for planning the transfer of research knowledge to decision-makers: what should be transferred, to whom, by whom, how, and with what effect. Applied to an evaluation, the answers produce a plan with different products and messengers for different audiences, as the table in Section 4.4 shows.
HSCI 241 Lesson 12, Sections 3.1 and 3.3 (Reporting and Translating Review Findings), applies both tools to the findings of a systematic review and is optional fuller reading.
Main messages are the conclusions and their implications, written as statements a decision-maker could act on. They are distinct from a summary of the contents of the report. The executive summary then gives the context, the questions, the approach in a few sentences, the answer to each question with its confidence, and the recommendations. Writing for decision-makers involves leading with the answer, using plain language and consistent terms, giving numbers with their context (for example, 12 more older adults out of every 100 no longer screening as lonely), and stating uncertainty in words that a non-specialist can interpret.
Evaluation adds two requirements to this general advice. First, the main messages of an evaluation report are its evaluative conclusions from Section 3, each with its level of confidence, so the rubric and the synthesis rules agreed with the steering committee determine what the messages can claim. A message about worth also names the standard behind it, such as the willingness-to-pay value or the rubric's descriptor for good reach, so that the reader can see the basis of the judgement. Second, the executive summary answers the evaluation questions in the order the primary intended users agreed them, so that each user can find the answer to the question they asked. The executive summary of an evaluation proposal, as in Section 4.7, has a different job: it states what the evaluation will deliver, how and when, so that a funder can decide whether to support it.
| Weaker main message | Stronger main message |
|---|---|
| This report presents the results of a mixed-methods evaluation of the Connector program's first year. | The Connector program reduced loneliness modestly among the older adults it served in its first year, and the evaluation team has moderate confidence in this result. |
| Cost-effectiveness results were sensitive to assumptions. | At current caseloads the program costs about $125,000 for each year of full health gained, and the cost would fall below $50,000 if caseloads fill and benefits last. |
| Some groups had lower attendance. | Older adults whose first language is neither English nor French were less likely to attend a first meeting, and interpreter-supported first meetings could close this gap before the second wave. |
4.3 Data Visualization for Decision-Makers
Charts in evaluation reports communicate a message to a reader who has little time. Cleveland and McGill (1984) ranked visual encodings by how accurately people judge the quantities they show, and their experiments confirmed the upper part of the ranking: positions along a common scale are judged most accurately, lengths and angles less accurately, and, in their ranking, areas and colour shading least accurately. Dot plots and bar charts therefore support more accurate comparison than pie charts or bubble charts. Tufte (2001) urged designers to maximize the share of ink that shows data and to remove decoration that carries no information, and Evergreen (2017) adapted these ideas for evaluators with practical guidance on choosing and formatting charts for reports.
Economic results need particular care. The cost-effectiveness plane and the acceptability curve are useful for technical readers and are hard for most decision-makers to read. For an executive audience, a short scenario table and a plain statement of probability work better, for example: "If the health authority is willing to pay $100,000 for each year of full health gained, there is about a 28 percent chance that the program is good value at current caseloads."
4.4 Knowledge Translation
The Canadian Institutes of Health Research (CIHR) define knowledge translation as a dynamic and iterative process of synthesis, dissemination, exchange and ethically sound application of knowledge to improve health and strengthen the health system. CIHR distinguishes integrated knowledge translation, in which knowledge users take part in the work from the start, from end-of-grant knowledge translation, in which findings are shared once the work is complete. A utilization-focused evaluation with an engaged steering committee is integrated knowledge translation by design. Lesson 9's Knowledge-to-Action framework (Graham et al., 2006) describes what happens after knowledge reaches users.
In an evaluation, the steering committee is the usual vehicle for integrated knowledge translation. When knowledge users sit on the committee that chooses the questions, agrees the rubric and interprets draft findings, they shape the products and their timing and carry the findings back to their own organizations, which answers the Lavis question of who should transfer the knowledge. A finding is more likely to be acted on when the messenger is someone the audience trusts, which for clinicians may be a respected peer and for a community may be one of its own members, so committee members often present findings to their own constituencies. The table applies these ideas to Cedar Valley, using the five questions introduced in Section 4.2.
| Audience | Main interest | Product and messenger | Timing |
|---|---|---|---|
| Health authority executive | Whether to fund the second wave, and on what conditions | Two-page briefing and a presentation by the evaluation lead and the program director | Before the second-wave budget decision |
| Steering committee | Interpretation, recommendations and program improvement | Data interpretation workshop on draft findings | Before the report is finalized |
| First Nations partners | Findings about Indigenous participants and the land-based pathway | Joint review and co-interpretation, and products the partners choose, under agreed data governance | Before any release, as agreed |
| Older adults and families | What the program offers and what it achieved | Plain-language summary in large print and at community meetings, presented with the committee's older adult members | After the executive briefing |
| Referring clinicians | Whether referral helps their patients | One-page summary and a short talk at clinic meetings by a clinician champion | Before second-wave clinics begin referring |
| Other health authorities and researchers | Transferability and methods | Full report and a journal article reported with CHEERS 2022 and TIDieR | After the report is released |
The First Nations row reflects commitments made in Lessons 4 and 5. Under the First Nations principles of OCAP® (ownership, control, access and possession), the partners decide how information about their members is interpreted and shared, and findings are reviewed with them before release. Reporting back to communities and participants is an obligation in its own right, separate from its value as a means of influencing decisions.
4.5 Evaluation Use
Research on evaluation use began when evaluators noticed that many evaluations had no visible effect on decisions. Weiss (1979) showed that research is used in many ways besides the direct application of findings, and reviews of evaluation use (Leviton & Hughes, 1981; Cousins & Leithwood, 1986) organized these into types that remain in use today.
Instrumental use is the most visible type, but conceptual use and process use often matter as much over time, because they change how managers and staff think about a program. Cousins and Leithwood (1986) found that use depended on characteristics of the evaluation, such as its relevance, credibility, quality of communication and timeliness, and on characteristics of the decision setting, such as information needs, the political climate, competing information and the receptiveness of users. A later review of the empirical literature by Johnson and colleagues (2009) emphasized the engagement of interest holders throughout the evaluation. Patton (2008) called the presence of an identifiable person or group who cares about the findings the personal factor, and his utilization-focused approach (Lesson 1) builds the evaluation around such primary intended users.
An evaluation that reports after the decision has little chance of instrumental use. The Cedar Valley plan schedules an interim briefing before the second-wave budget decision, with the full report to follow.
Questions chosen with primary intended users (Lesson 4) produce answers they want. Engagement throughout also builds the trust that makes unwelcome findings easier to accept.
Users act on findings they believe. Credibility comes from a defensible design, transparent synthesis with a rubric agreed in advance, independence where it matters, and candour about limitations.
Products matched to each audience, clear main messages and good charts make findings usable by people who will never read the full report.
Budgets, political commitments and competing priorities limit what an evaluation can change. Evaluators who understand the decision context can frame recommendations that are feasible within it.
Use also has an ethical side. Misuse occurs when findings are suppressed, distorted or selectively reported, or when an evaluation is commissioned only to justify a decision already made. Evaluators can reduce the risk by agreeing in advance, in the evaluation contract or terms of reference, how findings will be released, who owns the report, and how disagreements about interpretation will be handled, and by keeping Indigenous partners' data governance rights separate from any power to suppress unfavourable findings about the program.
4.6 Work Plan, Evaluation Budget and Risk Register
An evaluation proposal closes with the practical commitments that let a funder judge whether the evaluation can be delivered on time and within its means: a work plan, a budget and a risk register. They apply to the evaluation itself the program work plan and budget outline described in Lesson 2, Section 4.6 (Needs Assessment and Planning Models).
Background: timelines, roles and budgets
A Gantt chart lays the tasks of a project along a timeline, with a bar for the duration of each task, markers for key dates and arrows for dependencies between tasks. A roles table names one accountable person for each task, and a budget lists each cost line with the calculation behind it (quantity, unit cost and total). HSCI 207 Lesson 6, Sections 4.1 to 4.3 (Choosing an Approach and Managing a Research Project), introduces these tools for a research project and is optional reading.
The evaluation work plan
The work plan lists each evaluation task with the person or role responsible, its start and end dates and the output that marks its completion, and a Gantt chart shows the same tasks on a timeline. Its milestones, the dated checkpoints on which later work depends, include ethics approval, the data access agreement, the end of baseline data collection, the interim briefing and the final report. For Cedar Valley, the plan works backward from the second-wave budget decision: the interim briefing must reach the executive before that decision, so baseline data collection and the first analysis are scheduled to finish in time for it. Tasks that depend on outside bodies, such as research ethics review, review under the First Nations data governance agreement and requests for linked administrative data, often take longest and are started first.
The evaluation budget
The evaluation budget is separate from the program budget. Its largest line is usually staff time for the evaluation lead, analysts, research coordinators and interviewers, each costed as hours or full-time equivalents multiplied by a rate that includes benefits. Other common lines are incentives or honoraria for participants and for community members who advise the evaluation, transcription of interviews and focus groups, data access and linkage fees charged by data stewards, translation and printing of instruments, travel to clinics and community meetings, and the production of knowledge translation products. Each line shows its calculation, so that a reviewer can check it and the team can revise it when a quantity changes.
The risk register
A risk register lists what could stop the evaluation from answering its questions on time. For each risk it records the likelihood, the impact, the mitigation and the owner, the person responsible for watching the risk and acting on it. Likelihood and impact are often rated low, medium or high, so that the team attends first to risks that are both likely and serious. The table shows three illustrative entries for Cedar Valley.
| Risk | Likelihood | Impact | Mitigation | Owner |
|---|---|---|---|---|
| Approval of the request for linked administrative data arrives after the interim briefing is due | Medium | High | Send the request in the first month, and prepare the briefing from survey data if the linked data are late | Evaluation lead |
| Follow-up survey response is low among older adults who did not attend a first meeting | High | Medium | Telephone follow-up by trained interviewers, a modest honorarium, and an analysis of attrition | Research coordinator |
| Second-wave clinics launch earlier than planned, which shortens the comparison period | Low | High | Agree launch dates with the steering committee in advance and record any change | Program director, with the evaluation lead |
4.7 Worked Example: The Cedar Valley Evaluation Proposal
An evaluation proposal brings the parts of an evaluation plan together in one document that a funder can review. The worked example is the 6,000-word proposal that the Cedar Valley evaluation team wrote at the start of the first wave, before any of the illustrative results in Sections 2 and 3 existed. A proposal describes a plan, so its executive summary states what the evaluation will do, why, and how its findings will be used. The table shows how the team allocated the 6,000 words and which lesson covers each part.
| Proposal section | Source in the course | Words |
|---|---|---|
| 1. Program description, need and objectives | Lessons 1 and 2 | 700 |
| 2. Program theory: logic model and theory of change | Lesson 3 | 600 |
| 3. Interest holders, engagement and evaluation questions | Lesson 4 | 600 |
| 4. Evaluation matrix, indicators and data systems | Lesson 5 | 700 |
| 5. Design and its justification | Lessons 6 to 8 | 1,000 |
| 6. Implementation evaluation | Lesson 9 | 500 |
| 7. Costing and economic evaluation | Lesson 10, Sections 1 and 2 | 600 |
| 8. Synthesis: criteria, rubric and confidence | Lesson 10, Section 3 | 400 |
| 9. Ethics, Indigenous data governance and equity | Lessons 1, 4 and 5 | 300 |
| 10. Reporting, knowledge translation and use | Lesson 10, Section 4 | 300 |
| 11. Timeline, evaluation budget and risks | Lesson 10, Section 4.6, and Lesson 2, Section 4.6 | 300 |
| Total | 6,000 |
Main messages. The Cedar Valley Health Authority will decide in one year whether to extend the Connector program to its remaining 12 primary care clinics. This evaluation will tell the executive, before that decision, whether the program reaches the older adults who need it, whether it reduces loneliness, what it costs, and whether it is good value compared with usual care. It will also tell program staff how to improve delivery for groups who are not being reached.
Questions. The steering committee agreed five questions: how equitably the program reaches referred older adults; whether it is delivered as intended and in a culturally safe way; what effect it has on loneliness, social participation, self-rated health and use of emergency and primary care at twelve weeks and twelve months; what it costs and whether it is good value; and which contexts and mechanisms explain differences in results.
Approach. Because the first-wave clinics were chosen for readiness, the evaluation will compare them with the 12 second-wave clinics before those clinics launch, using a difference-in-differences design with checks of trends before launch. An implementation evaluation will track reach, fidelity and adaptations. A costing will combine program accounts with a time study of connector work, and a cost-utility analysis from the health-system perspective, with a societal scenario, will use the EQ-5D-5L at intake, twelve weeks and twelve months. Results will be judged against a rubric agreed in advance with the steering committee, with cultural safety as a hurdle.
Partnership and use. Older adults with lived experience and First Nations representatives on the steering committee will shape the instruments, the interpretation and the reports. Information about First Nations participants will be governed by an agreement with the partners consistent with OCAP®. The executive will receive an interim briefing before the budget decision, and older adults, clinicians and partners will each receive products designed for them.
What an evaluation proposal contains
An evaluation proposal of this kind, about 6,000 words excluding references and appendices, is organized as in the worked example. Beyond the parts developed in earlier lessons, it includes a costing and economic evaluation plan that states the perspective, the resources to be measured and the outcome measure, a draft rubric with synthesis rules, a reporting and use plan for at least three audiences, and the work plan, budget and risk register of Section 4.6. An executive summary of no more than two pages, written for decision-makers, precedes it.
A sound proposal connects the program's need, theory, questions, design and methods in a consistent line of reasoning. It justifies the design and states its assumptions and threats to validity with the checks that will address them, names a perspective for the costing and identifies resources (including in-kind and partner resources), and makes the rubric and synthesis rules explicit, developed or planned with primary intended users. Its plans for ethics, equity, Indigenous data governance, reporting and use are specific to the program's interest holders, and its executive summary can be read on its own and tells a decision-maker what the evaluation will deliver and when.
Reflection
A provincial ministry commissioned an evaluation of a youth mental health drop-in program. The evaluator worked independently and met the ministry once at the start. The 140-page report arrived three months after the ministry had set the next year's budget. It contained 60 tables, every chart was titled "Results", and it had no executive summary. It was sent only to the ministry, so the program's youth advisory council and its staff never saw it. A year later, staff said nothing in the program had changed. Using what is known about the conditions under which evaluations are used, identify three reasons this evaluation had little use. Then redesign the reporting and knowledge translation plan for three audiences (the ministry, program staff, and young people), naming for each a product, a messenger and a timing.
Three conditions for use were missing. Timeliness: the report arrived after the budget decision, so instrumental use was impossible. Engagement and relevance: the evaluator met the ministry once and did not involve staff or young people, so no primary intended users were invested in the findings, which Patton calls the absence of the personal factor. Communication: a 140-page report without an executive summary or message titles is unusable for busy readers, and sending it only to the ministry excluded the people able to change practice.
For the ministry, I would provide a two-page briefing with main messages and recommendations, presented by the evaluator with the program director, at least one month before the budget decision, followed by a report of about 25 pages with technical appendices. For program staff, I would hold a data interpretation workshop on the draft findings, led by the evaluator and a respected program manager, before the report is finalized, so that staff shape the recommendations they will carry out. For young people, I would co-produce a short visual summary and a social media version with the youth advisory council, presented by council members at drop-in sites after the ministry briefing. Throughout, the evaluator would meet these users regularly so that the questions and products reflect their decisions.
Minimum 20 characters required.
Question 1: Which of the following statements is a conclusion, as distinct from a finding or a recommendation?
Question 2: According to Cleveland and McGill (1984), which visual encoding supports the most accurate comparison of quantities?
Question 3: The executive funds the second wave of Cedar Valley on the conditions the evaluation recommended. Which type of use is this?
Question 4: Which plan follows the 1:3:25 format for a report to decision-makers?
Final Assessment
Bringing It All Together
This lesson completed the evaluation plan by asking what a program costs, whether it is good value, how evidence is combined into a judgement, and how that judgement reaches the people who will use it. Section 1 showed that the budget of the fictional Cedar Valley Connector program ($840,000) understates its economic cost, which rises to $972,000 from the health-system perspective and to $1,022,000 when volunteer time is valued, and that time-driven activity-based costing puts the connector time for each participant at $525. It also showed how spare capacity in a start-up year inflates the average cost per participant. Section 2 used these costs in an illustrative cost-utility analysis, which gave an incremental cost of $1,868, a gain of 0.015 QALYs and an ICER of $124,533 per QALY, with scenarios ranging from $41,617 to $131,200 and a probability of 0.282 of being cost-effective at $100,000 per QALY.
Section 3 placed the economic result beside evidence on reach, equity, cultural safety and effectiveness in a rubric agreed in advance, and reached separate judgements of merit (good) and worth at year-one caseloads (adequate), each with a stated level of confidence. Section 4 turned to reports, executive summaries, data visualization, knowledge translation and the conditions for use, and ended with the work plan, budget and risk register of an evaluation and a worked example of a complete evaluation proposal. The lesson's main argument is that economic and other evidence serve decisions best when they are combined openly, with the standards agreed before the results arrive and the findings delivered in time and in forms that each audience can use.
Key Takeaways from this lesson
- Economic evaluation values resources at their opportunity cost, so donated space, clinician time and volunteer time count even when no money changes hands.
- The perspective of an analysis determines whose costs count, and the CADTH reference case uses the publicly funded health care payer perspective, with a societal analysis when important costs fall outside the health system.
- Costing proceeds by identifying, measuring and valuing resources, and time-driven activity-based costing multiplies the time each activity takes by a capacity cost rate.
- Average cost per participant falls as caseloads approach capacity, so start-up costs overstate steady-state costs and full-capacity assumptions overstate efficiency when demand is uncertain.
- A full economic evaluation compares the costs and consequences of two or more alternatives, and cost-utility analysis expresses health gains in QALYs so that programs in different areas can be compared.
- The ICER divides incremental cost by incremental effect, the cost-effectiveness plane classifies results by quadrant, and net monetary benefit converts health gains into money at a stated willingness to pay.
- Scenario analyses show the effect of structural assumptions such as the persistence of benefits, while probabilistic sensitivity analysis and acceptability curves show the effect of uncertainty in parameter values.
- Social return on investment ratios depend heavily on financial proxies and deadweight adjustments, and ratios from different analyses cannot be compared.
- Evaluative rubrics make the reasoning behind a judgement explicit, and synthesis rules, hurdles and confidence statements should be agreed with primary intended users before the evidence arrives.
- Evaluations are used when they answer questions their users care about, arrive before decisions are made, are credible, and are communicated in products designed for each audience, including communities that hold rights over their data.
Core Concepts Reviewed
Section 1: opportunity cost, financial and economic cost, program, health system and societal perspectives, identification, measurement and valuation, gross costing, micro-costing, time-driven activity-based costing, fixed, variable, average and marginal costs, and budget impact.
Section 2: full and partial economic evaluation, cost-effectiveness, cost-utility, cost-benefit and cost-consequence analysis, QALYs and utilities, the ICER, the cost-effectiveness plane, net monetary benefit, scenario and probabilistic sensitivity analysis, acceptability curves, social return on investment and CHEERS 2022.
Section 3: merit and worth, evaluative rubrics, qualitative and numerical weight and sum, hurdles and profiles, convergent, complementary and dissonant evidence, confidence in judgements, and TIDieR.
Section 4: findings, conclusions and recommendations, main messages and the 1:3:25 format, principles of data visualization, integrated knowledge translation, instrumental, conceptual, symbolic and process use, conditions for use, evaluation work plans, budgets and risk registers, and the structure of an evaluation proposal.
The final reflection asks you to bring the lesson together by turning the evidence from an evaluation into main messages, a recommendation and an explanation of an economic result for a decision-maker.
Reflection
A regional health authority must decide whether to renew a community diabetes self-management program for three years. The program ran for two years in 10 communities with an annual budget of $1.2 million and served 1,000 adults a year. The evaluation found the following. The health-system cost, including overhead and physician time, was $1.38 million a year. Compared with similar communities without the program, average HbA1c (a measure of blood glucose control) fell by an additional 0.3 percentage points at 12 months (95 percent confidence interval 0.1 to 0.5). The incremental cost-effectiveness ratio (ICER) was $45,000 per quality-adjusted life year (QALY), with a probability of 0.62 of being cost-effective at $50,000 per QALY. Participation was 40 percent of eligible adults in remote communities and 65 percent in towns. In interviews, participants valued group sessions led by peers, and remote participants said virtual sessions were hard to join because of poor internet service. Write three main messages for the executive summary and one recommendation, and explain in two or three sentences how you would present the economic result to an executive who is not an economist.
Main messages. First, the program improved blood glucose control modestly among the adults it served, with an additional fall in HbA1c of 0.3 percentage points compared with similar communities, and we have moderate confidence that the program caused this improvement. Second, the program costs about $1,380 per participant a year to the health system and is likely to be good value at commonly used reference values, although there remains a meaningful chance that it is not. Third, adults in remote communities took part much less often than adults in towns (40 percent compared with 65 percent), mainly because virtual sessions depend on internet service that many remote households lack.
Recommendation. Renew the program for three years, and use part of the renewal to offer in-person, peer-led sessions in remote communities, with participation monitored against a target agreed with those communities.
Presenting the economic result. I would say that each year of full health gained costs about $45,000, which is below the $50,000 reference value often used in Canada, and that when the analysis accounts for uncertainty there is roughly a six-in-ten chance that the program is good value at that level. I would show a short table of the scenarios that most change the result, instead of a cost-effectiveness plane, and say which assumption the decision depends on most.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: The fictional Cedar Valley program's health-system cost ($972,000) is higher than its budget ($840,000). Which items account for the difference?
Question 2: Why should an economic evaluation of a social prescribing program report a societal perspective alongside the health-system perspective?
Question 3: Compared with usual care, the utility difference rises from 0 at intake to 0.04 at twelve weeks and returns to 0 at fifty-two weeks. What are the QALYs gained over the year?
Question 4: A program costs $600 more per participant than usual care and gains 0.010 QALYs. At a willingness to pay of $50,000 per QALY, which statement is correct?
Question 5: In the Cedar Valley probabilistic sensitivity analysis, the probability of cost-effectiveness at $100,000 per QALY was 0.282. What does this mean?
Question 6: Which assumption in the Cedar Valley scenario analysis brought the ICER below $50,000 per QALY?
Question 7: Why does the Cedar Valley analysis test the persistence of effects with a scenario instead of only through the probabilistic analysis?
Question 8: A cost-utility analysis that uses a threshold of $100,000 per QALY is a restricted form of which type of analysis?
Question 9: The Cedar Valley committee agreed that the program cannot be rated better than adequate overall if cultural safety is rated poor. What is this rule called, and why is it used?
Question 10: At year-one caseloads, Cedar Valley was rated good on effectiveness and adequate on value for money. Which conclusion follows Scriven's distinction between merit and worth?
Question 11: Why does a complete TIDieR description matter for the economic evaluation of Cedar Valley?
Question 12: An evaluation team learns that the second-wave budget decision will be made in month ten, and its full report is due in month fourteen. What would most improve the chance of instrumental use?
Question 13: Which practice best respects the rights of First Nations partners under OCAP® when the Cedar Valley evaluation reports its findings?
Question 14: Which reporting guideline applies to the Cedar Valley cost-utility analysis, and what did its 2022 version add?
Question 15: A consultant reports that Cedar Valley returns $1.43 in social value for every dollar invested. Which question should an evaluator ask first?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
This glossary defines the terms, tools and people introduced in Lesson 10, grouped by the lesson's main themes.