Synthesizing Quantitative Findings
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Explain what synthesis is and build characteristics tables, cross-tabulations and results tables for the studies included in a review.
- Define groups for synthesis by population, intervention, comparator, outcome and design, and write decision rules that select one result per study.
- Describe the three synthesis methods that the Cochrane Handbook accepts when effects are not pooled, and summarize effect estimates with the median, interquartile range and range.
- Describe the nine items of the SWiM reporting guideline and use them to plan and report a synthesis without meta-analysis.
- Explain, with reference to statistical power, why vote counting by statistical significance misleads, and carry out vote counting based on the direction of effect with a sign test.
- Read and construct effect direction plots and harvest plots from tabulated study results.
- Decide whether pooling is justified for a group of studies by asking five questions, and identify what HSCI 230 Lesson 2 teaches about meta-analytic methods.
- Plan the synthesis for a rapid scoping review, including its groups, decision rules, results table or cross-tabulation, SWiM-based methods paragraph and pooling decision.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
Tabulating and Grouping Studies
Learning Objectives for this section
- Explain what synthesis means in a review and why tabulating the included studies is the first step of a quantitative synthesis.
- Distinguish a characteristics table, a cross-tabulation and a results table, and state the question each one answers.
- Define the groups for a synthesis by population, intervention, comparator, outcome and design, and explain why they should be planned in the protocol.
- Write decision rules that select one result per study when a study reports several scales, time points or analyses.
- Build a results table for one comparison and outcome, with effects aligned in one direction and studies ordered by design and risk of bias.
Introduction
In Lesson 8, the Cedar Valley evidence team extracted data, appraised each study's risk of bias and rated the certainty of evidence for loneliness. Those ratings depend on a step that this lesson teaches: synthesis, the process of bringing the results of several studies together to answer the review question. A synthesis says what the studies, taken together, show about a comparison and an outcome, and how consistent they are.
A quantitative synthesis combines numerical results, such as differences in loneliness scores between program and comparison groups; Lesson 10 introduces the synthesis of qualitative findings. Section 1 of this lesson tabulates and groups the studies. Section 2 introduces synthesis without meta-analysis and the SWiM reporting guideline. Section 3 explains why counting studies by statistical significance misleads and presents better displays. Section 4 asks when pooling results in a meta-analysis is justified, a method whose computation HSCI 230 Lesson 2 teaches.
Before launching a community connector (social prescribing) program for older adults, the planning team of the fictional Cedar Valley Health Authority in British Columbia, which serves about 46,000 people aged 65 and older, asked an evidence team (an evidence officer, a university librarian and a student intern) for a rapid scoping review and environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.
The searches and screening produced 42 included studies: 14 randomized trials (4 of them cluster-randomized), 10 non-randomized controlled studies, 9 uncontrolled before-and-after studies, 6 qualitative studies and 3 mixed-methods studies. In Lesson 8, no trial reached low risk of bias overall for loneliness, and GRADE rated the evidence as low certainty for group-based programs and for one-to-one befriending or telephone programs, and very low certainty for community connector programs. All study labels, results and other numbers in this lesson beyond these shared figures are illustrative and belong to the fictional case.
This lesson works with the 36 studies that report quantitative results: the 33 quantitative studies and the quantitative strands of the 3 mixed-methods studies. The 6 qualitative studies and the qualitative strands of the mixed-methods studies go to the qualitative evidence synthesis in Lesson 10. The figure shows the route from the extraction forms to the synthesis.
1.1 Why Synthesis Starts with Tables
A common first draft of a review's results describes the included studies one at a time: "Smith and colleagues found that ... Jones and colleagues found that ...". Readers of such a draft must hold every study in mind and do the synthesis themselves, and the writer has no structure for noticing patterns. A structured table solves both problems. It lays the studies side by side on the same columns, so that similarities, differences and gaps become visible before anyone writes a sentence of interpretation.
The guidance on narrative synthesis by Popay and colleagues (2006) describes four elements: developing a theory of how the intervention works, why and for whom; developing a preliminary synthesis; exploring relationships within and between studies; and assessing the robustness of the synthesis. Tabulation is the main tool of the preliminary synthesis, and the Cochrane Handbook chapter on synthesis without meta-analysis (McKenzie and Brennan, 2019) likewise recommends structured tables. A review usually needs three kinds of table.
A characteristics table describes each included study in one row: its design, country and setting, participants, the program and its comparator, the outcomes measured and the time points. It answers the question "What was studied, in whom, and how?" Every included study appears, and Lesson 12 shows how the table is presented.
| Study | Design | Participants | Program and comparator | Loneliness measure; time point |
|---|---|---|---|---|
| Trial A | Randomized trial | 96 adults aged 65 to 89 living alone, England | Weekly arts group in a seniors' centre for 8 weeks; usual activities | UCLA Loneliness Scale (20 items); 8 weeks |
| Trial L | Randomized trial | 160 adults aged 70 and older, the Netherlands | Weekly telephone befriending call for 6 months; waiting list | De Jong Gierveld scale (6 items); 6 months |
| Study P | Non-randomized controlled study | 420 adults aged 65 and older referred from primary care, Scotland | Community connector referral; usual care in matched practices | Revised UCLA Loneliness Scale (20 items); 6 months |
| Study BA1 | Before-and-after study | 140 adults aged 65 and older, Ontario | Community connector program run by a seniors' services agency; no comparator | Three-item loneliness scale; 12 months |
A cross-tabulation counts studies in the cells formed by two characteristics, most often intervention category by design or intervention category by outcome. It answers the question "Where is the evidence, and where are the gaps?" Section 1.3 gives the Cedar Valley version, and the evidence and gap maps in Lesson 10 extend the idea.
A results table lists, for one comparison and one outcome, each study's result in a consistent form, together with the information a reader needs to weigh it: the number of participants analyzed, the measure and time point, and the study's risk of bias. It answers the question "What did the studies in this group find?" Section 1.5 builds one for group-based programs and loneliness.
1.2 Grouping Studies for Synthesis
A review question is usually broader than any single synthesis. A statement that "community-based interventions reduce loneliness" would combine programs that work in different ways and would help no one plan a service. The review therefore divides its studies into groups for synthesis, each defined by a population, intervention, comparator and outcome. The Cochrane Handbook calls this the "PICO for each synthesis" and treats it as part of planning the review (McKenzie et al., 2019). Design can also define groups, since randomized and non-randomized studies are usually synthesized separately.
Grouping involves a trade-off called lumping and splitting. Lumping places broadly similar programs in one group, giving more studies per synthesis but possibly combining programs that differ in ways that matter. Splitting places studies in narrow groups, giving specific answers from as few as one or two studies each. The right level depends on the decision the review supports. The Cedar Valley planning team must choose a kind of program to fund, so the team grouped by the main component of each program.
The eligibility criteria already restrict the population to adults aged 65 and older. Few studies selected participants by baseline loneliness, so the team treated that characteristic as a possible source of variation (Section 2) and did not make it a separate group.
The extraction form assigned each study to one of six categories by its main component: group activity; one-to-one befriending or telephone contact; community connector or social prescribing; intergenerational; technology-based; and other. A data dictionary rule classed a connector program that also ran a lunch club as a connector program, because referral defined it.
Comparators were usual care, usual activities, no program or a waiting list. The protocol set apart any study comparing two programs, because that comparison answers a different question; no controlled study in the review used such a comparator.
Loneliness is the main outcome. The protocol (Lesson 2) grouped other outcomes into domains: social isolation and participation, well-being, mental health, physical health and function, and health service use. Each domain is a separate synthesis, so that results for loneliness are never mixed with results for, say, depressive symptoms.
Randomized trials, non-randomized controlled studies and uncontrolled before-and-after studies protect against bias to different degrees. The team synthesized each design separately and based conclusions about effects on controlled designs, a choice Section 2 shows how to report.
Groups should be planned in the protocol, because a group defined after seeing the results can be drawn to make a program look better or worse. When plans must change, for example because a category contains no studies, the change is reported with its reason, as the SWiM guideline in Section 2 asks.
1.3 Cross-Tabulating the Cedar Valley Studies
The team's first table counted the 42 included studies by intervention category and design. Five of the six categories contain studies; no study fell into "other".
| Intervention category | Randomized trials | Non-randomized controlled | Before-and-after | Mixed methods | Qualitative | Total |
|---|---|---|---|---|---|---|
| Group activity | 8 | 0 | 3 | 1 | 2 | 14 |
| One-to-one befriending or telephone | 6 | 0 | 1 | 0 | 1 | 8 |
| Community connector or social prescribing | 0 | 5 | 3 | 2 | 2 | 12 |
| Intergenerational | 0 | 2 | 1 | 0 | 1 | 4 |
| Technology-based | 0 | 3 | 1 | 0 | 0 | 4 |
| Total | 14 | 10 | 9 | 3 | 6 | 42 |
The table shows a pattern that matters for the planning team. Community connector programs, the model Cedar Valley intends to launch, have 12 studies, second only to group activities (14), yet none is a randomized trial; all 14 trials evaluated group or one-to-one programs. This explains in advance why Lesson 8 rated the connector evidence as very low certainty. The table also shows that the intergenerational and technology-based categories hold too few controlled studies for a confident synthesis, which the review should state plainly.
Using the table above, answer three questions. (1) How many controlled studies (randomized or non-randomized) evaluated community connector programs? (2) What proportion of the 42 studies evaluated group activities or one-to-one programs? (3) Name one gap the table reveals that a future study in British Columbia could fill. Suggested answers: (1) five, all non-randomized; (2) 22 of 42, or about 52 percent; (3) the review contains no randomized trial of a community connector program, which a staggered introduction of the Cedar Valley program across clinics could begin to address.
1.4 Choosing One Result per Study
Many studies report several results that could fill one cell of a results table: loneliness on two scales, at three time points, by two kinds of analysis. Reviewers who choose among these after seeing them may pick, without meaning to, the result that fits their expectations. The safeguard is a set of decision rules, written before extraction, that select one result per study for each synthesis, as the Cochrane Handbook recommends (McKenzie et al., 2019). The accordion lists the Cedar Valley rules.
Use the first scale available in this order: the UCLA Loneliness Scale (any version), the De Jong Gierveld Loneliness Scale, the three-item loneliness scale, and a single direct question. The order was added to the protocol as a dated amendment in week 7, after the charting form was piloted and before the results fields were charted.
Use the time point closest to the end of the program, and tabulate later follow-up separately, because effects may fade after a program ends.
For randomized trials, use the analysis that includes all randomized participants (the intention-to-treat analysis) when it is reported. For non-randomized studies, use the estimate adjusted for the confounding domains named in the protocol, which Lesson 8 listed as baseline loneliness, depressive symptoms and living alone.
Use estimates that account for clustering. A cluster trial analyzed as if individuals had been randomized has a confidence interval that is too narrow, and the team flags such results.
When a trial compares two programs with one control group, place each arm in its own category and note that the comparisons share control participants, who must not be counted twice in one synthesis.
Record every effect so that the same sign means the same thing. On the loneliness scales used here, a higher score means more loneliness, so a negative difference favours the program. For outcomes where a higher score is better, such as social participation, the team reversed the sign so that a negative difference still favours the program, and noted the reversal.
1.5 Building a Results Table
With groups defined and one result chosen per study, the team built a results table for each synthesis. To compare results from different loneliness scales, it expressed each as a standardized mean difference (SMD): the difference in mean scores between program and comparison groups divided by a standard deviation, which puts every scale in standard deviation units. HSCI 230 Lesson 2 shows the calculation. The table gives the results for the 8 trials of group-based programs.
| Trial | Design | Participants analyzed | Scale | Program length | SMD (95% CI) | Risk of bias (RoB 2) |
|---|---|---|---|---|---|---|
| C | Cluster-randomized | 210 | De Jong Gierveld (6 items) | 16 weeks | −0.18 (−0.51 to 0.15) | Some concerns |
| F | Randomized | 182 | De Jong Gierveld (6 items) | 24 weeks | −0.15 (−0.44 to 0.14) | Some concerns |
| E | Randomized | 148 | UCLA (20 items) | 16 weeks | −0.20 (−0.52 to 0.12) | Some concerns |
| B | Randomized | 120 | Three-item scale | 12 weeks | 0.03 (−0.33 to 0.39) | Some concerns |
| A | Randomized | 96 | UCLA (20 items) | 8 weeks | −0.22 (−0.62 to 0.18) | Some concerns |
| D | Randomized | 64 | Three-item scale | 10 weeks | −0.55 (−1.05 to −0.05) | Some concerns |
| G | Cluster-randomized | 230 | UCLA (20 items) | 12 weeks | −0.02 (−0.34 to 0.30) | High |
| H | Randomized | 186 | Three-item scale | 8 weeks | −0.31 (−0.60 to −0.02) | High |
SMD: standardized mean difference, program minus comparison; values below zero mean less loneliness in the program group. CI: confidence interval, the range of effects reasonably compatible with the trial's data. Trials C and G report estimates adjusted for clustering. The 8 trials analyzed 1,236 participants in total. All results are illustrative.
Three features make this table useful. Every result points in a stated direction, so a reader sees at once that seven of the eight estimates fall below zero. The rows are ordered by risk of bias and then by size, so the two trials at high risk sit together and a reader can check whether they differ from the rest; here one shows almost no difference and the other a small benefit. And the table includes columns, such as program length, that might explain variation (Section 2). A study that reported only "no significant difference" still gets a row, marked "not estimable", so that readers can see how much evidence lacks usable data.
The table does not yet state a conclusion. A reader might be tempted to count the confidence intervals that exclude zero (two of eight) and decide that group programs mostly fail. Section 3 explains why that reading is wrong and what to do instead, and Section 2 first sets out the methods and reporting standard for synthesis without meta-analysis.
Reflection
A student review examines community walking programs for adults aged 65 and older and includes seven studies. Four are randomized trials comparing a walking group with usual activities: two measure loneliness with the UCLA Loneliness Scale, one with the De Jong Gierveld scale, and one with both scales. Two of these four trials report loneliness at 12 weeks (the end of the program) and again at 6 months. A fifth randomized trial compares a walking group with a seated exercise class. The last two studies are before-and-after evaluations of walking groups with no comparison group. The student's draft results table has one row for every loneliness result in every study, records each result as "effective" or "not effective" following the authors' abstracts, and leaves out one of the four usual-activities trials because it reported only "no significant difference". Propose the groups for synthesis, write two decision rules for choosing one result per study, and identify three problems in the draft table, with a correction for each.
Groups. The main synthesis should contain the four trials comparing walking groups with usual activities, for loneliness. The trial with a seated exercise class as comparator answers a different question (walking compared with another program) and should be reported separately. The two before-and-after studies should be tabulated and described, but conclusions about effect should rest on the controlled trials, because without a comparison group a change over time cannot be attributed to the program.
Decision rules. First, when a study reports more than one loneliness scale, use the UCLA Loneliness Scale, then the De Jong Gierveld scale. Second, use the time point closest to the end of the program (12 weeks) in the main synthesis and tabulate the 6-month results separately.
Problems and corrections. (1) One row per result lets the trial with two scales and the trials with two time points appear more than once; the table should have one row per study for each synthesis, chosen by the decision rules. (2) "Effective" and "not effective" repeat the authors' significance-based interpretation; the table should record each effect estimate with its confidence interval, aligned so that the same sign always favours the program. (3) Omitting the trial with no numbers makes the evidence look more complete; it should keep its row, marked "not estimable", and the synthesis should report that one of four trials lacked usable data.
Minimum 20 characters required.
Question 1: What does a cross-tabulation in a review show?
Question 2: Why should decision rules for choosing one result per study be written before extraction begins?
Question 3: What does the community connector row of the Cedar Valley cross-tabulation show?
Question 4: A trial reports loneliness on a scale where a higher score means more loneliness, and social participation on a scale where a higher score means more participation. In a results table where negative values favour the program, what does the Cedar Valley Rule 6 require?
Synthesis Without Meta-Analysis and the SWiM Guideline
Learning Objectives for this section
- Give reasons why a review may synthesize quantitative results without a meta-analysis, and explain why "narrative synthesis" is an incomplete description of a method.
- Describe the three synthesis methods that the Cochrane Handbook accepts when effects are not pooled, and state the question each answers and the data each needs.
- Summarize a set of effect estimates with the median, interquartile range and range, and interpret the summary correctly.
- Describe the nine items of the SWiM reporting guideline and apply them to a synthesis plan.
- Investigate variation in results without meta-analysis by tabulating results against characteristics planned in advance, and write a structured synthesis paragraph.
Introduction
Many reviews in public health do not combine their results in a meta-analysis, because the studies are too few, too varied or too incompletely reported, or because the review maps evidence. Such reviews still synthesize, and many have described what they did as a narrative synthesis. The term has a defined meaning in the guidance of Popay and colleagues (2006), but it has often been used for any synthesis that is not a meta-analysis, including study-by-study descriptions with no stated method.
Campbell and colleagues (2019) examined a sample of Cochrane reviews that synthesized results without meta-analysis and found that most did not describe their synthesis methods or link their conclusions clearly to the data. The same group then led the development of the Synthesis Without Meta-analysis (SWiM) reporting guideline (Campbell et al., 2020), which sets out nine items to report when a review synthesizes quantitative effects without pooling them.
2.1 When a Review Synthesizes Without Meta-Analysis
The Cochrane Handbook chapter on synthesis without meta-analysis (McKenzie and Brennan, 2019) lists the circumstances that lead reviews to other methods. Effect estimates may be incompletely reported, for example as "no significant difference" with no numbers. Studies may use effect measures that cannot be converted to a common metric. The evidence may be at such high risk of bias that a pooled estimate would give misleading precision. And the studies may differ so much that an average effect would have no clear meaning. Section 4 returns to these conditions.
The Cedar Valley review meets several of them. Its protocol (Lesson 2) planned no meta-analysis, because a twelve-week rapid scoping review aims to map evidence. Its community connector studies are non-randomized and mostly at serious risk of confounding, and several categories contain only a handful of controlled studies. Even so, the planning team asked what the evidence says about loneliness, and the review needs a method that answers that question transparently.
2.2 Three Methods That Synthesize Without Pooling
McKenzie and Brennan (2019) describe three methods that are acceptable when effects are not pooled. Each answers a different question and needs different data. The table summarizes them.
| Method | Question it answers | Minimum data from each study | Main limitation |
|---|---|---|---|
| Summarizing effect estimates | What is the range and distribution of the observed effects? | An effect estimate in a common metric | The summary does not weight studies by size or precision and gives no confidence interval for an average effect. |
| Combining P values | Is there evidence of an effect in at least one study? | An exact P value and the direction of effect | The result says nothing about the size of the effect and is rarely meaningful for health decisions. |
| Vote counting based on the direction of effect | Is there any evidence of an effect? | The direction of effect | The count ignores the size and precision of each result, so a trivial benefit and a large one count the same. |
Summarizing effect estimates
When the studies in a group report effects in a common metric, such as the standardized mean difference introduced in Section 1, the review can describe their distribution with three statistics. The median is the middle value when the estimates are put in order. The interquartile range (IQR) is the span of the middle half of the estimates, from the 25th to the 75th percentile. The range runs from the smallest to the largest estimate. These statistics describe the spread of observed results and carry no weighting by study size.
Worked example: group-based programs and loneliness
The eight standardized mean differences from Section 1, in order, are −0.55, −0.31, −0.22, −0.20, −0.18, −0.15, −0.02 and 0.03.
Median = (−0.20 + −0.18) ÷ 2 = −0.19, because with eight values the median is the mean of the fourth and fifth.
Interquartile range = −0.24 to −0.12 (computed with the default quantile method in R and Python). Range = −0.55 to 0.03.
By a widely used rule of thumb (Cohen, 1988), a standardized mean difference of about 0.2 is described as small. The summary statement is that the middle half of the trials found loneliness lower in the program group by between 0.12 and 0.24 standard deviations, a small difference.
For the 6 one-to-one befriending or telephone trials, the 5 with estimable results had a median of −0.20 (IQR −0.24 to −0.16; range −0.28 to −0.12). Trial L reported only that there was no significant difference, so the summary must state that it rests on 5 of the 6 trials.
This method needs a standard metric. A difference of 2 points means something different on the 20-item UCLA scale, scored from 20 to 80, than on the three-item scale, scored from 3 to 9, and standard deviation units put the two on a common footing. For binary outcomes, such as the proportion who often feel lonely, the metric would be a ratio such as the odds ratio or risk ratio, which HSCI 230 Lesson 2 teaches.
Combining P values
Methods such as Fisher's method combine the exact P values from several studies into one test of the hypothesis that the program has no effect in any study. A small combined P value suggests a real effect in at least one study but says nothing about its size. The Cedar Valley team did not use these methods, since the planning team asked about the size and consistency of effects.
Vote counting based on the direction of effect
The third method counts how many studies found an effect in the direction of benefit and how many in the direction of harm, whatever the size or significance of each result. It needs the least data, and Section 3 explains how it differs from the flawed practice of counting studies by statistical significance.
2.3 The SWiM Reporting Guideline
SWiM is a reporting guideline: it lists what a review should tell readers about a synthesis that does not pool effects, so that they can judge whether the conclusions follow. It was developed through a consensus process and is designed to be used alongside PRISMA 2020 (Campbell et al., 2020). Its nine items cover the methods (items 1 to 7), the results (item 8) and the discussion (item 9).
The accordion describes each item and shows how the Cedar Valley review reports it.
What to report. Item 1a asks for a description of the groups used in the synthesis, such as groupings of populations, interventions, outcomes or designs, with the rationale for each. Item 1b asks for any changes to these groups after the protocol, with reasons.
Cedar Valley. Studies were grouped by the main component of the program, by outcome domain and by design, because the planning team must choose a type of program. An amendment made before the synthesis began, at the planning team's request, added standardized effect summaries for trials and GRADE ratings for loneliness. The planned category "other" contained no studies.
What to report. The standard metric used for each outcome, why it was chosen, and any methods used to transform the results as reported into that metric, citing the guidance followed.
Cedar Valley. Loneliness results were expressed as standardized mean differences because the studies used different scales, with standard deviations derived from standard errors or confidence intervals by methods in the Cochrane Handbook where needed. Where no metric could be calculated, the direction of effect was used.
What to report. The methods used to synthesize the effects for each outcome, with a justification for using them instead of meta-analysis.
Cedar Valley. For groups with three or more estimable results, the team summarized standardized mean differences with the median, interquartile range and range. For all groups, it counted studies by direction of effect and applied a sign test, which Section 3 explains. It did not combine P values. Meta-analysis was not planned, for the reasons given in the protocol.
What to report. Any criteria used to select particular studies for the main synthesis or for drawing conclusions, such as design, risk of bias or directness, with a justification.
Cedar Valley. Conclusions about effects rest on the controlled studies. The before-and-after studies and the quantitative strands of the mixed-methods studies are shown in tables and plots but do not support statements about effects, because without a comparison group a change over time cannot be attributed to the program. Risk of bias was used to order studies, never to exclude them.
What to report. The methods used to examine variation in effects across studies when a meta-analysis and its statistical tools for heterogeneity are not used.
Cedar Valley. The team planned to tabulate results by program length, delivery setting and risk of bias, and to describe any pattern as exploratory. Section 2.4 shows the result.
What to report. The methods used to assess the certainty of the synthesis findings.
Cedar Valley. GRADE was applied to loneliness for three comparisons, following the guidance of Murad and colleagues (2017) for rating certainty without a pooled estimate, as Lesson 8 described.
What to report. The tables and graphs used to present effects, and the study characteristics used to order studies in them, with the studies clearly identified.
Cedar Valley. The review uses results tables ordered by design, risk of bias and size, an effect direction plot for community connector programs and a harvest plot of loneliness across the five categories (Section 3).
What to report. For each comparison and outcome, a description of the synthesized findings and their certainty, in language consistent with the question the synthesis addresses, stating which studies contribute.
Cedar Valley. Section 2.5 gives the paragraph for group-based programs and loneliness.
What to report. The limitations of the synthesis methods and of the groupings, and how they affect the conclusions that can be drawn for the review question.
Cedar Valley. Direction counts ignore the size of effects; most groups hold few controlled studies; one befriending trial had no estimable result; and grouping by main component may hide the contribution of other components in multi-component programs.
SWiM allows any defensible method, provided the review reports it against the nine items. A reader of a review that follows SWiM can find which studies contributed to each statement, what was done with their results, and why.
2.4 Investigating Variation Without Meta-Analysis
Results almost always vary across studies, partly by chance and partly, perhaps, because of real differences between programs, populations or methods. Without the statistical tools of meta-analysis, a review can tabulate or plot results against characteristics that might explain the variation, as SWiM item 5 asks. The characteristics should be few, chosen in the protocol and grounded in a reason to expect a difference, because a search through many characteristics after seeing the results will find patterns by chance.
The Cedar Valley team examined program length for the 8 group-based trials, reasoning that longer programs might allow friendships to form. The table divides the trials at 12 weeks.
| Program length | Trials | Standardized mean differences | Median | Range |
|---|---|---|---|---|
| 12 weeks or less | A, B, D, G, H (5) | −0.22, 0.03, −0.55, −0.02, −0.31 | −0.22 | −0.55 to 0.03 |
| More than 12 weeks | C, E, F (3) | −0.18, −0.20, −0.15 | −0.18 | −0.20 to −0.15 |
The medians are similar, and the shorter programs show a wider spread, including both the largest benefit and the two results near zero. With five and three trials, this pattern could easily arise by chance, so the team reported that program length did not explain the variation; tabulations by setting and by risk of bias gave the same conclusion. The unexplained variation is why GRADE rated down for inconsistency in Lesson 8.
2.5 Writing the Synthesis
SWiM item 8 asks that each written statement address a comparison and an outcome, report the findings and their certainty, and name the contributing studies. The tabs contrast a study-by-study draft with a structured synthesis of the same eight trials.
"Trial D found that the group program significantly reduced loneliness. Trial H also found a significant reduction. Trials A, C, E and F found no significant effect. Trial B found no effect, and Trial G found no effect. Overall, the evidence for group programs is mixed."
This draft repeats each study's significance test, never states a size of effect, ignores risk of bias and reaches a conclusion ("mixed") that the data do not support, since seven of the eight estimates point the same way.
"Eight randomized trials (1,236 participants; Trials A to H) compared group-based programs with usual activities or no program. Seven of the eight found lower loneliness in the program group at the end of the program. Standardized mean differences had a median of −0.19 (interquartile range −0.24 to −0.12; range −0.55 to 0.03), a small difference. The estimates varied, and program length, setting and risk of bias did not explain the variation. All trials had some concerns or high risk of bias for this outcome. Group-based programs may reduce loneliness slightly among older adults living in the community (low-certainty evidence)."
This paragraph names the comparison, the studies and the participants, gives the direction and the size of the effects, reports the exploration of variation and the risk of bias, and ends with a GRADE statement in the standard wording from Lesson 8.
A student review reports its synthesis methods in one paragraph: "Because the studies were heterogeneous, a narrative synthesis was conducted. Studies were grouped by intervention type. Results are presented in tables." List which SWiM items the paragraph addresses, which it omits, and one sentence the student could add for each of two omitted items. Suggested answer: the paragraph partly addresses item 1a (groups, without a rationale) and item 7 (tables, without saying how studies are ordered). It omits items 2 to 6 and gives no basis for items 8 and 9. For item 3, the student could add: "We summarized standardized mean differences with the median and interquartile range and counted studies by direction of effect." For item 4, the student could add: "Conclusions about effects were based on controlled studies; uncontrolled studies are described separately."
Reflection
Six randomized trials compared peer-support programs with usual care for loneliness among older adults living in the community, with 640 participants in total. Five reported results that the review team converted to standardized mean differences, where negative values mean less loneliness in the program group: Trial 1, −0.40; Trial 2, −0.10; Trial 3, −0.25; Trial 4, 0.05; Trial 5, −0.30. Trial 6 reported only that the difference was "not significant". The review also includes three uncontrolled before-and-after studies of peer support. GRADE rated the evidence from the trials as low certainty. Using these data: (a) calculate the median and range of the five standardized mean differences and count how many trials favour the program; (b) write one sentence for SWiM item 3 (synthesis methods) and one for SWiM item 4 (criteria for prioritizing results); (c) write a results paragraph that meets SWiM item 8; and (d) state one limitation of this synthesis for SWiM item 9.
(a) In order, the five values are −0.40, −0.30, −0.25, −0.10 and 0.05, so the median is −0.25 and the range is −0.40 to 0.05. Four of the five trials with estimable results favour the program, and Trial 6 has no estimable direction.
(b) Item 3: "Because the trials used different loneliness scales and a meta-analysis was not planned, we summarized standardized mean differences with the median and range and counted trials by the direction of effect." Item 4: "Conclusions about effect were based on the randomized trials; the three uncontrolled before-and-after studies are described separately because they lack a comparison group."
(c) "Six randomized trials (640 participants) compared peer-support programs with usual care. Four of the five trials with estimable results found lower loneliness in the program group, with a median standardized mean difference of −0.25 (range −0.40 to 0.05), a small difference; Trial 6 reported no usable data. Peer-support programs may reduce loneliness among older adults living in the community (low-certainty evidence)."
(d) The median gives every trial equal weight regardless of size, and the direction count treats the estimate of 0.05, which is close to zero, as a full vote against the program, so the synthesis describes the pattern of results without estimating an average effect with a confidence interval.
Minimum 20 characters required.
Question 1: Which question does summarizing effect estimates with the median and interquartile range answer?
Question 2: Eight standardized mean differences are −0.55, −0.31, −0.22, −0.20, −0.18, −0.15, −0.02 and 0.03. What is their median?
Question 3: Which statement describes SWiM item 4?
Question 4: What did the Cedar Valley team conclude after tabulating the group-based trial results by program length?
Beyond Vote Counting: Effect Direction Plots and Harvest Plots
Learning Objectives for this section
- Describe vote counting by statistical significance and explain, with reference to statistical power, why it misleads.
- Explain why a non-significant result is not evidence of no effect, and why the problem grows as the number of small studies grows.
- Carry out vote counting based on the direction of effect, report the proportion of studies favouring the program, and interpret a sign test.
- Read and construct an effect direction plot and a harvest plot from tabulated study results.
- Choose a display suited to the data available in a group of studies.
Introduction
Section 1 ended with a temptation: to see that only two of eight confidence intervals for group-based programs exclude zero and conclude that the programs mostly failed. Lesson 1 showed the same trap with eight simulated trials of one program, of which only three reached significance. This section explains why counting by significance gives wrong answers, presents vote counting based on the direction of effect, and shows two displays designed for synthesis without meta-analysis.
3.1 Vote Counting by Statistical Significance
A result is called statistically significant when its P value falls below a chosen threshold, usually 0.05, or equivalently when its 95 percent confidence interval excludes the value that means no effect (zero for a difference). Vote counting by statistical significance sorts studies into three piles, significant benefit, significant harm and no significant difference, and then reaches a conclusion from the tallies. If most studies are significant and favourable, the program "works"; if most are non-significant, it "does not work"; and if the piles are similar, the evidence is "mixed". The method still appears in published reviews, including one that the Cedar Valley team verified in Lesson 6, and the Cochrane Handbook advises against it (McKenzie and Brennan, 2019).
3.2 Why Counting by Significance Misleads
A non-significant result is not a finding of no effect
Whether a study reaches significance depends on the size of the true effect and on the study's precision, which rises with the number of participants. A small study of an effective program will often have a confidence interval wide enough to include zero. Altman and Bland (1995) summarized the error of reading such a result as proof of no effect in the title of a short paper: "Absence of evidence is not evidence of absence". A non-significant result means that the study could not distinguish the effect from zero. Its confidence interval often includes both no effect and a worthwhile benefit.
The figure plots the Cedar Valley results for group-based programs. It looks like a forest plot, which HSCI 230 Lesson 2 teaches, but it has no pooled estimate at the bottom; SWiM item 7 lists this kind of ordered display as a way to present results without meta-analysis.
Six of the eight intervals include zero, and all six of those also include reductions in loneliness of 0.3 standard deviations or more. The non-significant trials are compatible with a small benefit, and seven of the eight point estimates favour the program. A significance count reads this pattern as "two successes, six failures"; a direction count shows seven of eight trials pointing toward benefit, a pattern compatible with a small benefit measured imprecisely.
Small studies have low power
Statistical power is the probability that a study will produce a statistically significant result when the program truly has an effect of a given size. The table shows the power of a two-group trial to detect a standardized mean difference of 0.2, a small effect of the size the Cedar Valley trials suggest, using a two-sided test at the 0.05 level.
| Participants in the trial (two equal groups) | 50 | 100 | 200 | 400 |
|---|---|---|---|---|
| Power to detect a standardized mean difference of 0.2 | 11% | 17% | 29% | 52% |
Calculated with the normal approximation for a comparison of two means. Most Cedar Valley trials analyzed between 64 and 230 participants.
A trial with 100 participants has about a one-in-six chance of reaching significance even when the program truly reduces loneliness by 0.2 standard deviations, so a significance count across such trials will usually declare an effective program ineffective.
More studies can make the answer worse
One might expect that adding studies would correct the error. Hedges and Olkin (1980) showed the opposite for the majority rule. When each study has power below one half, the proportion of significant results tends toward that power as studies accumulate, so the chance that most studies are significant falls toward zero. The table applies this to trials of 100 participants each, with a true standardized mean difference of 0.2 and power of 17 percent.
| Number of trials | 5 | 10 | 20 | 40 |
|---|---|---|---|---|
| Expected number with a significant result | 0.9 | 1.7 | 3.4 | 6.8 |
| Probability that more than half are significant | 3.8% | 0.3% | 0.01% | under 0.01% |
With 40 such trials, a vote count is almost certain to conclude that a program with a real effect does not work. As the number of trials grows, a majority-rule count becomes more likely to reach the wrong conclusion.
Size, weight and false conflict
Vote counting by significance also discards information. A trial of 64 people counts the same as a trial of 230, and a large benefit counts the same as a trivial one that reached significance in a very large study. The method also manufactures conflict: when some trials are significant and others are not, reviewers call the evidence "mixed" even when the estimates lie close together, as in Figure 3.1.
3.3 Vote Counting Based on the Direction of Effect
The acceptable alternative counts each study by the direction of its effect estimate alone: favouring the program or favouring the comparison, whatever its size or significance (McKenzie and Brennan, 2019). The review reports the number and proportion of studies favouring the program, ideally with a confidence interval for that proportion, and can apply a sign test. The sign test asks how likely a split at least as uneven as the observed one would be if the program had no effect, so that each study were as likely to point one way as the other, like a fair coin.
Worked example: direction counts in the Cedar Valley trials
Group-based programs. Seven of 8 trials favour the program (88 percent; 95 percent confidence interval for the proportion 53 to 98 percent, Wilson method). Two-sided sign test: P = 0.07.
One-to-one befriending or telephone programs. Five of the 5 trials with an estimable direction favour the program, and Trial L, which reported only "no significant difference", has no estimable direction. Two-sided sign test on 5 trials: P = 0.06. No trial in this group reached statistical significance, so a significance count would have recorded zero successes in six.
These P values need care. With only five studies, the sign test cannot produce a two-sided P value below 0.06, because even five out of five has a probability of 2 in 32 under the hypothesis of no effect. Reading such a P value as proof of no effect would repeat the error of Section 3.2, so the team reports the counts, the proportion with its interval and the P value together, and lets the effect summaries from Section 2 describe size.
Direction counting has its own limits, which SWiM item 9 asks reviews to state. It ignores the size of effects, so an estimate of 0.03 counts as a full vote against the program although it is essentially zero. It gives a small study the same vote as a large one. It requires every result to be aligned in direction before counting (Section 1, Rule 6). And it counts studies, so a study that reports five loneliness measures still casts one vote in each synthesis, chosen by the decision rules.
The one-to-one trials reported these standardized mean differences (95 percent confidence intervals): Trial I −0.28 (−0.70 to 0.14); Trial J −0.12 (−0.53 to 0.29); Trial K −0.24 (−0.64 to 0.16); Trial L not estimable; Trial M −0.16 (−0.52 to 0.20); Trial N −0.20 (−0.61 to 0.21). (1) What would a significance count conclude? (2) What does a direction count show? (3) Write one sentence for the review. Suggested answers: (1) none of the six trials is significant, so a significance count would conclude that befriending does not reduce loneliness. (2) All five trials with an estimable result favour the program, with a median standardized mean difference of −0.20, and one trial has no estimable direction. (3) "All five trials with estimable results found lower loneliness in the befriending group (median standardized mean difference −0.20), but each was small and imprecise, and one further trial reported no usable data."
3.4 Effect Direction Plots
Direction counts become harder to follow when studies report several outcomes measured in different ways. The effect direction plot, developed by Thomson and Thomas (2013) for reviews of complex public health interventions, displays the direction of effect for every study and every outcome domain in one grid. Each row is a study and each column an outcome domain. An upward arrow shows an effect favouring the intervention, a downward arrow an effect favouring the comparison, and a sideways arrow a domain in which the study's several outcomes conflict. The size of the arrow shows the study's sample size, and rows are ordered or shaded by design and risk of bias. Hilton Boon and Thomson (2021) revised the plot to follow the 2019 Cochrane Handbook guidance, so that arrows record the direction of effect regardless of statistical significance and each column can be summarized with a sign test.
When a study reports several outcomes within one domain, such as depressive symptoms and anxiety within mental health, the team needs a rule for the domain's arrow. The Cedar Valley team gave a domain a direction when at least 70 percent of its outcomes in the study pointed the same way and marked it as conflicting otherwise, a threshold of the kind Thomson and Thomas (2013) described. It also defined the direction of each domain in advance: for health service use, fewer unplanned visits to primary care or emergency departments counted as benefit.
For loneliness, all five controlled studies favour the program, as do four of the five uncontrolled studies. On a two-sided sign test, five of five gives P = 0.06, the smallest value possible with five studies. Social isolation and well-being point the same way in every controlled study that measured them. Mental health and health service use are mixed: Studies Q and T recorded more unplanned visits among program participants, which the team interpreted with caution, since connectors may link people to care they had been going without. Because the controlled studies are non-randomized and three of five are at serious risk of confounding, the plot supports a statement about consistency of direction, and the GRADE rating from Lesson 8 (very low certainty) governs how strongly the review can state it.
Following SWiM item 4, the uncontrolled studies appear below a labelled divider, visible to readers, and the review's statements about effect rest on the controlled studies above it.
3.5 Harvest Plots
The harvest plot was developed by Ogilvie and colleagues (2008) to synthesize evidence about the differential effects of interventions, such as whether they narrow or widen health inequalities, and is now used for many kinds of synthesis without meta-analysis. It is a matrix. Rows usually represent outcomes, hypotheses or categories of intervention, and columns represent categories of finding. Each study is drawn as a bar in the cell that matches its finding, and the height and shading of the bar encode study characteristics chosen by the review team, such as design and risk of bias, with a label above the bar identifying the study. The plot combines the information of a vote count with a picture of the strength of the evidence behind each vote.
The Cedar Valley team drew one harvest plot for loneliness, covering the 24 controlled studies across the five intervention categories. It placed studies by direction of effect, consistent with Section 3.3, so that the plot shows a direction-based vote count with design and risk of bias added.
Twenty of the 24 controlled studies favour the program, one has an unclear direction and three favour the comparison. The tall bars, the randomized trials, are concentrated in the group and one-to-one rows. The community connector row holds only short bars, and three of its five are red. The intergenerational and technology rows hold so few studies, mostly at serious risk of bias, that no pattern can be read from them. A planning team can see in one figure that the most consistent and best-designed evidence concerns programs other than the one it plans to launch, which is the same message the cross-tabulation in Section 1 and the GRADE ratings in Lesson 8 gave.
3.6 Choosing a Display
The three displays suit different data, and a review often uses more than one. The table compares them. Other displays exist, such as the albatross plot, which plots each study's P value against its sample size (Harrison et al., 2017), and the choice among them should follow the question and the data, as SWiM item 7 asks reviews to explain.
| Display | What it shows | Suited to | Cedar Valley use |
|---|---|---|---|
| Ordered estimate plot without a pooled result | Each study's effect estimate and confidence interval on a common metric | Groups in which the studies report a common or convertible metric | Group-based trials (Figure 3.1) |
| Effect direction plot | Direction of effect for each study in several outcome domains, with sample size | Studies that measure many outcomes in different ways | Community connector programs (Figure 3.2) |
| Harvest plot | Where studies fall by direction across several categories, with design and risk of bias | Comparing categories, outcomes or hypotheses at a glance | Loneliness across five categories (Figure 3.3) |
Each of these methods describes a body of evidence without combining its results into a single estimate. Section 4 asks when a review should go further and pool the results statistically, and what HSCI 230 Lesson 2 adds when it does.
Reflection
A published review of a home-visiting program for isolated older adults included 12 randomized trials, each with between 40 and 120 participants. Three trials found a statistically significant reduction in loneliness and nine found no significant difference. The authors concluded that "most trials found no effect, so the program is ineffective". The review's supplementary table shows that 10 of the 12 point estimates favoured the program and 2 favoured the comparison, and that most confidence intervals included both zero and a reduction of 0.3 standard deviations or more. For context, a two-group trial of 80 participants has about 15 percent power to detect a standardized mean difference of 0.2, and a two-sided sign test for 10 of 12 studies favouring one direction gives P = 0.04. Explain why the authors' conclusion is not justified, state what a vote count based on the direction of effect shows, and describe a figure that would present these trials more informatively, including what its axes or columns would show.
The authors used vote counting by statistical significance, which treats each non-significant trial as evidence of no effect. A non-significant result only means that a trial could not distinguish the effect from zero. These trials were small: a trial of 80 participants has about 15 percent power to detect a small effect, so even if the program works, most such trials will be non-significant. Hedges and Olkin (1980) showed that when power is this low, a majority-rule count becomes less likely to detect a real effect as more trials are added. The supplementary table supports this reading, since most confidence intervals include a worthwhile reduction as well as zero.
A count based on direction shows that 10 of the 12 trials (83 percent) favour the program, and the sign test gives P = 0.04, so a split this uneven would be unlikely if the program had no effect. The count says nothing about size, so the review should also summarize the standardized mean differences with a median and interquartile range.
An ordered estimate plot without a pooled result would show each trial as a row, with its standardized mean difference and confidence interval on a horizontal axis, a vertical line at zero, and rows ordered by risk of bias and sample size. Readers would see that the estimates cluster on the side favouring the program. A harvest plot by outcome domain would be an alternative strong answer.
Minimum 20 characters required.
Question 1: A trial of 100 participants tests a program whose true standardized mean difference is 0.2. About how likely is the trial to reach statistical significance at the 0.05 level?
Question 2: What did Hedges and Olkin (1980) show about majority-rule vote counting when individual studies have low power?
Question 3: In an effect direction plot, what does a sideways arrow mean?
Question 4: In the Cedar Valley harvest plot, what do bar height and bar colour show?
When Pooling Is Justified
Learning Objectives for this section
- Explain what a meta-analysis adds to a synthesis and what problems pooling cannot correct.
- Apply five questions to decide whether pooling the results of a group of studies is justified.
- Correct common misconceptions about heterogeneity, study similarity and the purpose of pooling.
- Justify a decision to pool or not to pool for each comparison in a review, and report the decision transparently.
- Identify what HSCI 230 Lesson 2 teaches about meta-analytic methods and how the tables built in this lesson prepare the data for them.
Introduction
The methods in Sections 2 and 3 describe a body of evidence without combining its results into a single number. A meta-analysis goes further: it combines the effect estimates from several studies into one weighted average, with a confidence interval, using statistical methods. Gene Glass (1976) coined the term to describe "the analysis of analyses", and the method is now standard in reviews of the effects of health interventions. This section does not teach how to compute a meta-analysis. HSCI 230 Lesson 2 does that, covering fixed-effect and random-effects models, forest plots, statistical heterogeneity, publication bias and influential studies. This section teaches the decision that comes first: whether pooling is justified for a given group of studies, and how to report the decision either way.
4.1 What Pooling Adds and What It Cannot Fix
Pooling adds three things. First, it combines the information in several imprecise studies into one more precise estimate, which addresses the problem of small studies that Section 3 described. The eight Cedar Valley trials of group-based programs, none with more than 230 participants, would together give an estimate based on 1,236. Second, it estimates the size of the effect with a confidence interval, which a direction count cannot do. Third, it provides statistical tools for describing how much results vary between studies and for exploring why, as well as for examining signs of publication bias.
Pooling cannot fix three other problems. It cannot remove bias from the included studies: if every trial overstates the benefit because participants who know their group report less loneliness, the pooled estimate will overstate it too, with a narrower confidence interval that makes the error look more certain. It cannot give meaning to an average of things that should not be averaged: a pooled effect of "community programs" that mixes group activities, befriending calls and connector referrals would describe no program anyone could implement. And it cannot recover studies that were never published or found, which is why the searching methods in Lessons 4 and 5 matter. These limits explain why the decision to pool depends on judgement about the studies as well as on whether the numbers are available.
4.2 Five Questions Before Pooling
The Cochrane Handbook treats the decision as a matter of whether a meta-analysis would be meaningful and appropriate for a given synthesis, considering the diversity of the studies, the data available and the risk of bias (Deeks, Higgins and Altman, 2019; McKenzie and Brennan, 2019). Those considerations can be organized as five questions, which the Cedar Valley team asked of each comparison.
The studies should share the population, intervention, comparator and outcome defined for the synthesis (Section 1.2), closely enough that an average effect would answer a question someone is asking. Reviewers call differences in these features clinical diversity; in public health the same idea applies to programs, settings and populations. Some diversity is expected and acceptable, and the question is whether the average would mean something. A pooled effect of group programs for older adults living in the community answers a planning question; a pooled effect of every community intervention does not.
Each study must provide an effect estimate and a measure of its precision (a standard error or confidence interval), or data from which these can be calculated, in a metric that can be converted to a common one. Studies that report only "no significant difference" or only a P value without direction cannot contribute. If many studies lack usable data, a meta-analysis of the rest may be unrepresentative, and a direction count or effect direction plot, which can include more studies, may give a fairer picture.
Randomized and non-randomized studies are subject to different biases, and Cochrane guidance advises against combining them in one meta-analysis; when both are included, they are analyzed separately. Within a design, problems of the unit of analysis must be handled, such as cluster trials analyzed as if individuals had been randomized and trials with several arms sharing one control group (Section 1.4). Non-randomized studies should contribute estimates adjusted for important confounders.
If most studies are at high risk of bias in the same direction, a pooled estimate will be precise and probably wrong. Reviews can restrict the main analysis to studies at lower risk or run a sensitivity analysis that excludes studies at high risk, both of which are planned in the protocol (Lesson 8). A review should not pool studies at critical risk of bias by ROBINS-I, which the tool's developers advise excluding from synthesis.
The protocol should state whether and how effects will be pooled, so that the choice is not made after seeing which option gives the more favourable answer. Two studies are enough to compute a pooled estimate, but with very few studies the amount of variation between them cannot be estimated reliably, which affects random-effects results; HSCI 230 Lesson 2 explains why. With few studies, reviewers should report the pooled result with caution or describe the studies individually.
4.3 Common Misconceptions About Pooling
Students and some review authors hold beliefs about pooling that lead to poor decisions in both directions. The cards describe five of them.
4.4 Worked Example: Should the Cedar Valley Team Pool?
The team applied the five questions to each comparison for loneliness. The table records its reasoning.
| Comparison (controlled studies) | Questions 1 to 5 | Would pooling be justified? |
|---|---|---|
| Group-based programs versus usual activities or no program (8 randomized trials; 1,236 participants) | Similar programs, populations and comparators; standardized mean differences available for all 8, with clustering handled in Trials C and G; all randomized; no trial at low risk of bias, with Trials G and H at high risk, so a sensitivity analysis excluding them would be needed; not planned in the protocol. | Yes in a full systematic review, with a random-effects model and a sensitivity analysis by risk of bias. Not done in this review because the protocol planned none. |
| One-to-one befriending or telephone programs versus usual care or a waiting list (6 randomized trials; 742 participants) | Similar programs; two kinds of comparator, which would need separate analysis because waiting-list comparisons may inflate self-reported benefit; Trial L has no usable data, leaving 5; three trials at high risk of bias; not planned. | Possible in principle, but with 5 usable trials split by comparator, each analysis would rest on very few studies. Synthesis without meta-analysis is more appropriate here. |
| Community connector programs versus usual care (5 non-randomized controlled studies; 1,480 participants) | Programs differ in referral pathways, connector caseloads and duration; adjusted estimates available in 4 of 5, using different sets of confounders; three studies at serious risk of confounding; designs cannot be combined with trials. | No. A pooled estimate would combine differently adjusted, seriously confounded results and would suggest more precision than the evidence holds. |
| Intergenerational (2) and technology-based (3) programs | Diverse programs and outcomes; very few studies; most at serious risk of bias. | No. |
| All 24 controlled studies as one group | Different programs, comparators and designs; the average would not inform the choice among programs. | No. |
The first row needed the most thought. The group-based trials meet the conditions that the Cochrane Handbook sets out, and a full systematic review of these programs would normally pool them. The team considered amending its protocol to add a meta-analysis and decided against it for two reasons. The decision facing the planning team concerns community connector programs, for which pooling is not justified, and a precise estimate for group programs would not change that decision. And an amendment made after the results table had been seen would be open to the criticism that the method was chosen by the result. The review reports that the data for group-based programs would support a meta-analysis and recommends one in any full review. A team with different aims might reasonably decide the other way, provided it reports the change and its reason, as SWiM item 1b and PRISMA 2020 require.
4.5 From Synthesis to Statements for Decision-Makers
The synthesis ends in statements that combine the findings from Sections 2 and 3 with the GRADE ratings from Lesson 8, written in the standard wording of Santesso and colleagues (2020). The table gives the Cedar Valley statements for loneliness that will go into the evidence brief in Lesson 12.
| Comparison | Synthesis findings | Statement and certainty |
|---|---|---|
| Group-based programs | 7 of 8 trials favour the program; median standardized mean difference −0.19 (interquartile range −0.24 to −0.12) | Group-based programs may reduce loneliness slightly among older adults living in the community (low certainty). |
| One-to-one befriending or telephone programs | 5 of 5 trials with estimable results favour the program; median −0.20; one trial with no usable data | One-to-one befriending or telephone programs may reduce loneliness (low certainty). |
| Community connector programs | 5 of 5 controlled studies favour the program; mixed results for mental health and service use | The evidence is very uncertain about the effect of community connector programs on loneliness (very low certainty). |
| Intergenerational and technology-based programs | 3 of 5 controlled studies favour the program; most at serious risk of bias | Too few controlled studies to support a statement; certainty not rated. |
These statements show how far a careful synthesis without meta-analysis can go. The planning team learns which kinds of program have the most consistent evidence, roughly how large their effects appear to be, how confident it can be, and that the evidence for the model it intends to launch is very uncertain. Lesson 8 drew the practical implication: the Cedar Valley program should be launched with an evaluation built in, so that the program adds to the evidence about community connector programs.
4.6 Connecting to Meta-Analysis
Students who take HSCI 230 Lesson 2, before or after this course, will find that the work of this lesson is the preparation a meta-analysis needs. The groups for synthesis define which studies enter each meta-analysis. The decision rules choose one result per study. The results table, with standardized mean differences, confidence intervals and risk-of-bias judgements, is the data set the meta-analysis uses, and the ordered estimate plot in Figure 3.1 becomes a forest plot once a pooled estimate is added. The exploration of variation in Section 2.4 becomes subgroup analysis, and the restriction to studies at lower risk becomes a sensitivity analysis. HSCI 230 Lesson 2 adds the statistical model, the weights and the heterogeneity statistics, and its section on when pooling is inappropriate returns to the questions asked here.
Reflection
A review team is deciding whether to pool results for three groups of studies of programs for older adults. Group A: 7 randomized trials of telephone befriending compared with usual care, all measuring loneliness; standardized mean differences with confidence intervals are available for all 7; 2 trials are at high risk of bias and 5 have some concerns; the protocol planned a random-effects meta-analysis if at least five trials were comparable. Group B: 4 non-randomized studies of community workshop programs for older men; 2 measure loneliness and all 4 measure well-being with different scales; all report only unadjusted estimates, and 3 are at serious risk of confounding. Group C: a team member proposes pooling all 11 studies from Groups A and B into one estimate of the effect of "community programs" on loneliness. The five questions for pooling are: (1) Do the studies address the same question? (2) Are compatible data available? (3) Are the designs compatible? (4) Is risk of bias acceptable or handled? (5) Was pooling planned, with enough studies to estimate variation? Apply the questions to each group, decide whether pooling is justified, and state what the review should do instead where it is not.
Group A. The trials share a population, intervention, comparator and outcome, so an average effect of telephone befriending on loneliness would be meaningful (question 1). Effect estimates with confidence intervals are available for all seven (question 2), and all are randomized (question 3). Risk of bias can be handled with a sensitivity analysis that excludes the two trials at high risk (question 4), and the protocol planned a meta-analysis with a threshold this group meets (question 5). Pooling is justified, using the methods in HSCI 230 Lesson 2, with heterogeneity reported and the sensitivity analysis shown.
Group B. Only two studies measure loneliness, the estimates are unadjusted, three studies are at serious risk of confounding, and the well-being scales differ. Questions 2 and 4 fail, and question 5 fails for loneliness. Pooling is not justified. The review should tabulate the four studies, present an effect direction plot for loneliness and well-being, and report the synthesis with SWiM.
Group C. Pooling all eleven fails questions 1 and 3: it mixes different programs and combines randomized with non-randomized designs, which Cochrane guidance advises against. The average would describe no program anyone could deliver. The review should keep the groups separate and, if readers need an overview, use a harvest plot across the two categories.
Minimum 20 characters required.
Question 1: Which problem does pooling leave uncorrected?
Question 2: Why did the Cedar Valley team judge pooling unjustified for the community connector studies?
Question 3: What does Cochrane guidance advise about randomized trials and non-randomized studies of the same intervention?
Question 4: A review's results show large variation between trials of similar programs. What is the most appropriate response?
Final Assessment
Bringing It All Together
This lesson followed the Cedar Valley evidence team from extracted data to statements about effect. The team first made the evidence visible: a characteristics table described every study, a cross-tabulation showed that all 14 randomized trials concerned group or one-to-one programs while the 12 community connector studies were all non-randomized or uncontrolled, and results tables set out one result per study, chosen by rules written in advance and aligned in direction. Groups for synthesis, defined by population, intervention, comparator, outcome and design, kept unlike comparisons apart.
Because the review did not pool results, it used the methods that the Cochrane Handbook accepts for synthesis without meta-analysis and reported them against the nine SWiM items. Summaries of standardized mean differences described the size and spread of effects, and counts by the direction of effect, with sign tests, replaced the misleading practice of counting studies by statistical significance, which in the Cedar Valley trials would have reported two successes in eight and none in six. Effect direction plots and harvest plots displayed these counts alongside design, sample size and risk of bias.
Finally, five questions about the question, the data, the designs, the risk of bias and the plan determined whether pooling was justified. The group-based trials could support a meta-analysis of the kind HSCI 230 Lesson 2 teaches, while the connector studies could not. Combined with the GRADE ratings from Lesson 8, the synthesis produced clear statements for decision-makers: group and one-to-one programs may reduce loneliness, and the evidence for community connector programs is very uncertain.
Key Takeaways from this lesson
- Synthesis brings the results of several studies together to answer a review question, and a quantitative synthesis begins with structured tables that make the studies comparable.
- A characteristics table describes every study, a cross-tabulation shows where evidence is concentrated and where gaps lie, and a results table sets out one result per study for each comparison and outcome.
- Groups for synthesis are defined by population, intervention, comparator, outcome and design, planned in the protocol, and any later change is reported with its reason.
- Decision rules written before extraction select one result per study when studies report several scales, time points or analyses, and every effect is aligned so that the same sign means the same thing.
- When effects are not pooled, the Cochrane Handbook accepts summarizing effect estimates, combining P values and vote counting based on the direction of effect, each of which answers a different question.
- The SWiM guideline asks reviews to report nine items covering grouping, metrics, synthesis methods, prioritization, heterogeneity, certainty, data presentation, results and limitations.
- Vote counting by statistical significance misleads because small studies have low power, non-significant results are not evidence of no effect, and the error grows as underpowered studies accumulate.
- Vote counting based on the direction of effect, reported with the proportion favouring the program and a sign test, is acceptable but ignores the size and precision of each result.
- Effect direction plots display direction across several outcome domains for each study, and harvest plots display where studies fall across categories with design and risk of bias encoded in each bar.
- Pooling is justified when the studies address the same question, provide compatible data, share a design, have risk of bias that is acceptable or handled, and were planned for pooling in sufficient number; HSCI 230 Lesson 2 teaches the computation.
Core Concepts Reviewed
Section 1: synthesis, characteristics tables, cross-tabulations, results tables, groups for synthesis, lumping and splitting, decision rules for selecting results, alignment of direction and the standardized mean difference.
Section 2: narrative synthesis, reasons for synthesis without meta-analysis, summarizing effect estimates with the median and interquartile range, combining P values, the nine SWiM items and the investigation of variation without meta-analysis.
Section 3: statistical significance, statistical power, the failure of vote counting by significance shown by Hedges and Olkin, vote counting based on direction with a sign test, effect direction plots and harvest plots.
Section 4: what meta-analysis adds and cannot fix, five questions before pooling, the separation of randomized and non-randomized designs, misconceptions about heterogeneity, and the link to HSCI 230 Lesson 2.
The final reflection asks you to explain the lesson's main ideas to a decision-maker who has read a misleading headline about the evidence.
Reflection
The director of a health authority planning a community connector (social prescribing) program for older adults has read a news story headlined "Social prescribing doesn't work: most studies found no significant effect". The director asks you, the review team's intern, for a reply of about 200 words. Your review found the following for loneliness. Group-based programs: 8 randomized trials, 2 statistically significant, 7 of 8 point estimates favouring the program, median standardized mean difference −0.19 (interquartile range −0.24 to −0.12), low certainty. One-to-one befriending or telephone programs: 6 randomized trials, none statistically significant, 5 of the 5 trials with estimable results favouring the program, low certainty. Community connector programs: 5 non-randomized controlled studies, all favouring the program, three at serious risk of confounding, very low certainty. The review did not pool results, because its protocol planned no meta-analysis and pooling was not justified for the connector studies. Write the reply. It should explain what is wrong with the reasoning in the headline, state what the review found for each kind of program and how certain those findings are, explain why the review did not produce a single pooled number, and say what the findings imply for the planned program.
The headline counts studies by statistical significance. Most studies of these programs are small, and a small study will often miss a real but modest effect, so "no significant effect" in a single study usually means that the study could not tell whether the program worked. Our review counted studies by the direction of their results and summarized the size of their effects instead.
Seven of eight trials of group-based programs found less loneliness in the program group, and the typical difference was small (a median of about 0.2 standard deviations); these programs may reduce loneliness slightly, with low certainty. All five befriending or telephone trials with usable results pointed toward benefit, also with low certainty, although none was significant on its own. All five controlled studies of community connector programs also pointed toward benefit, but they were non-randomized and mostly at serious risk of confounding, so the evidence is very uncertain.
We did not combine the results into one number because the connector studies differ too much and are too open to bias for an average to be trustworthy, and our protocol planned no meta-analysis.
The evidence leaves the effect of connector programs uncertain. The planned program should therefore be launched with an evaluation built in, for example by introducing it across clinics in stages, so that Cedar Valley adds to the evidence.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: How many Cedar Valley studies does this lesson's quantitative synthesis cover, and where are the qualitative findings synthesized?
Question 2: A review team places telephone befriending and home visiting in separate groups for synthesis, although each group then holds only two studies. What kind of grouping choice is this?
Question 3: Which SWiM items are reported in the methods section of a review?
Question 4: A review states that "three of ten trials found a significant benefit, so the evidence is inconclusive". What is the main flaw in this reasoning?
Question 5: A sign test on 5 studies that all favour the program gives a two-sided P value of 0.06. What is the best interpretation?
Question 6: Which display best shows the direction of effect in several outcome domains for each of ten studies that used different measures?
Question 7: In the Cedar Valley effect direction plot, Studies Q and T recorded more unplanned health service visits among program participants. How did the team handle these results?
Question 8: Why did the Cedar Valley team place the uncontrolled studies below a labelled divider in its effect direction plot?
Question 9: Trial L reported only "no significant difference" with no numbers. How should a synthesis treat it?
Question 10: What does pooling in a meta-analysis add that a direction count cannot provide?
Question 11: Why did the Cedar Valley team decide against pooling the group-based trials, even though the data would support a meta-analysis?
Question 12: Which GRADE certainty level matches the statement "Group-based programs may reduce loneliness slightly among older adults living in the community"?
Question 13: When a review later carries out a meta-analysis, what role do its groups for synthesis, decision rules and results tables play?
Question 14: Which pattern in the Cedar Valley harvest plot matters most for the planning team?
Question 15: A student argues that because the studies in a group differ somewhat in populations and program length, pooling is never appropriate. What is the best reply?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, tools and people appear in this lesson on synthesizing quantitative findings.