# Lesson 9: Synthesizing Quantitative Findings

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This is the episode for Lesson nine of Health Sciences two forty-one. Last time, the Cedar Valley team extracted its data, appraised its studies and rated the certainty of the evidence. This week we look at the step that sits underneath those ratings, which is synthesis: bringing the results of several studies together to answer the review question.

**Sarah:** Remind listeners where the team is, and what synthesis means for them in practice.

**Kiffer:** The fictional Cedar Valley Health Authority in British Columbia wants to launch a community connector program, sometimes called social prescribing, for older adults. Before it does, the planning team asked a small evidence team for a rapid scoping review and environmental scan. The review asks which community interventions have been evaluated for loneliness or social isolation among adults aged sixty-five and older, and with what outcomes. After screening, the team has forty-two included studies.

**Sarah:** And all of those go into this lesson's synthesis?

**Kiffer:** Thirty-six of them do. Those are the thirty-three quantitative studies plus the quantitative parts of three mixed-methods studies. The six qualitative studies, and the qualitative parts of the mixed-methods studies, go to a separate synthesis in Lesson ten. I should also say that every study label and result we discuss today is illustrative. They belong to the fictional case.

**Sarah:** So where does a synthesis start? I imagine most students would start writing.

**Kiffer:** Most do, and the first draft usually describes the studies one at a time. Smith found this, and Jones found that. The trouble is that the reader has to hold every study in mind and do the synthesis themselves. The better starting point is a set of tables that put the studies side by side on the same columns.

**Sarah:** How many tables are we talking about?

**Kiffer:** Usually three. A characteristics table describes each study in one row: its design, setting, participants, program, comparator and outcomes. A cross-tabulation counts studies by two characteristics at once, such as the kind of program by the design of the study. And a results table lists one result per study for a single comparison and outcome.

**Sarah:** Let's talk about grouping. Why can't the team just synthesize everything at once?

**Kiffer:** Because the review question is broader than any useful answer. The Cedar Valley question covers all community interventions. A sentence saying that community interventions reduce loneliness would mix group activities, befriending phone calls, intergenerational programs and connector referrals, which work in different ways. Nobody could plan a service from that sentence. So the team divides the studies into groups for synthesis. Each group is defined by a population, an intervention, a comparator and an outcome, and often by design as well. The Cochrane Handbook calls this the population, intervention, comparator and outcome for each synthesis, and it treats the groups as part of planning the review, written into the protocol before anyone sees results. That matters, because a group drawn after seeing the results can be drawn to make a program look better or worse.

**Sarah:** I've heard the phrase lumping and splitting. Is that the same thing?

**Kiffer:** It describes the trade-off inside grouping. Lumping puts broadly similar programs in one group. You get more studies and a broad answer, but you might combine programs that differ in ways that matter. Splitting puts studies in narrow groups. You get specific answers, but a group might hold only one or two studies, which is too few to judge consistency.

**Sarah:** So which did Cedar Valley choose?

**Kiffer:** It grouped by the main component of each program, because the planning team has to choose a kind of program to fund. There were five categories with studies: group activities, one-to-one befriending or telephone contact, community connector programs, intergenerational programs and technology-based programs.

**Sarah:** And the cross-tabulation showed something important, I think.

**Kiffer:** It did. Community connector programs, which is the model Cedar Valley wants to launch, have twelve studies, second only to group activities with fourteen, and none of them is a randomized trial. All fourteen randomized trials in the review evaluated group activities or one-to-one programs. So one simple table explains why Lesson eight rated the connector evidence as very low certainty. The evidence starts from weaker designs.

**Sarah:** Now, decision rules. What problem do they solve?

**Kiffer:** Many studies report several results that could fill the same cell. A trial might measure loneliness on two different scales, at the end of the program and again six months later, with two different kinds of analysis. If you choose among those after seeing them, you might pick, without meaning to, the one that fits your expectations. Decision rules, written before extraction, choose one result per study for each synthesis.

**Sarah:** What did the Cedar Valley rules say?

**Kiffer:** There were six. The team ranked the loneliness scales, used the time point closest to the end of the program, and used the analysis of everyone who was randomized. It used estimates that account for clustering or confounding, and never counted a shared control group twice. And it lined up every result so that a negative number always means less loneliness in the program group.

**Sarah:** Let's get to the results table. What does it look like for the group programs?

**Kiffer:** There are eight trials of group-based programs, with one thousand two hundred and thirty-six participants between them. The trials used three different loneliness scales, so the team expressed each result as a standardized mean difference. That is the difference in average scores between the program and comparison groups divided by a standard deviation, which puts every scale into the same units. The calculation is taught in Health Sciences two thirty, Lesson two.

**Sarah:** And what did the table show?

**Kiffer:** Seven of the eight estimates fall below zero, which means less loneliness in the program group. The rows are ordered by risk of bias and then by size, so the two trials at high risk of bias sit together at the bottom. That lets a reader check whether those two look different. Here they don't stand apart: one shows almost no difference and the other a small benefit.

**Sarah:** And a study with no numbers?

**Kiffer:** It still gets a row, marked as not estimable. Leaving it out would make the evidence look more complete than it is.

**Sarah:** Section two is synthesis without meta-analysis. Why would a review choose not to pool?

**Kiffer:** The Cochrane Handbook lists the usual reasons. Results may be incompletely reported, for example in a paper that says only that there was no significant difference. Studies may use effect measures that can't be converted to a common one. The evidence may be at such high risk of bias that a pooled number would look far more precise than it deserves. Or the studies may be so different that an average would mean nothing.

**Sarah:** And Cedar Valley?

**Kiffer:** It meets several of those. Its protocol, back in Lesson two, planned no meta-analysis, because a twelve-week rapid scoping review is mainly about mapping the evidence. The connector studies are non-randomized and mostly at serious risk of confounding. But the planning team still asked what the evidence says about loneliness, so the review needs a transparent way to answer.

**Sarah:** For a long time, people would just call this narrative synthesis.

**Kiffer:** They would, and the term has a real meaning in guidance that Jennie Popay and colleagues wrote in two thousand six. But it was often used as a label for any synthesis that wasn't a meta-analysis, including study-by-study summaries with no method at all. Mhairi Campbell and colleagues looked at a sample of Cochrane reviews in two thousand nineteen and found that most didn't describe their synthesis methods or link their conclusions clearly to the data.

**Sarah:** Which led to SWiM.

**Kiffer:** Yes. The same group led the development of the Synthesis Without Meta-analysis guideline, known as SWiM, published in twenty twenty. But first, the Cochrane Handbook accepts three methods when effects aren't pooled. The first is summarizing the effect estimates. If the studies share a common metric, you describe their distribution with the median, the interquartile range and the range. The second is combining P values, which tests whether there is an effect in at least one study but says nothing about its size. The third is vote counting based on the direction of effect, which counts how many studies favour the program and how many favour the comparison.

**Sarah:** Let's do the numbers for the group programs. What's the median?

**Kiffer:** With eight values, the median is the average of the fourth and fifth in order. Those are minus zero point two zero and minus zero point one eight, so the median is minus zero point one nine. The interquartile range, which covers the middle half of the trials, runs from minus zero point two four to minus zero point one two, and the full range runs from minus zero point five five to plus zero point zero three. In plain words, the middle half of the trials found loneliness lower in the program group by between roughly an eighth and a quarter of a standard deviation. By a common rule of thumb from Jacob Cohen, a value of about zero point two is called a small effect. So it's a small difference.

**Sarah:** Is the median the same thing as a pooled effect?

**Kiffer:** That's a common confusion. The median gives every trial the same weight, whether it had sixty-four participants or two hundred and thirty, and it has no confidence interval. It describes how the results are spread. A meta-analysis estimates a weighted average with a confidence interval, which is a different quantity.

**Sarah:** What about the befriending trials?

**Kiffer:** There were six. Five had results that could be converted, with a median of minus zero point two zero. The sixth, which we call Trial L, reported only that there was no significant difference. Its result can't be estimated, so the summary has to say that it rests on five of the six trials.

**Sarah:** Now SWiM itself. Nine items, I think.

**Kiffer:** Nine items. Items one to seven belong in the methods section. Item one is about how studies were grouped, with a second part for any changes from the protocol. Item two is the standard metric and how results were converted to it. Item three is the synthesis method. Item four is the criteria used to prioritize some results over others. Item five is how variation between studies was investigated. Item six is how certainty was assessed. Item seven is how the data are presented in tables and graphs. Item eight, in the results, asks for the findings and their certainty for each comparison and outcome, naming the studies that contribute. Item nine, in the discussion, asks for the limitations of the synthesis methods and the groupings.

**Sarah:** Does SWiM tell you which method to use?

**Kiffer:** SWiM is a reporting guideline, so it leaves the choice of method to you. You can use any defensible method, as long as you report it against the nine items. The value is that a reader can trace each statement back to the studies and the method behind it.

**Sarah:** Give me a Cedar Valley example of item four. Prioritizing results sounds abstract.

**Kiffer:** The team based its statements about effect on the controlled studies, meaning the randomized trials and the non-randomized studies with a comparison group. The before-and-after studies appear in tables and plots, but they don't support statements about effect, because without a comparison group you can't tell whether a change over time came from the program. Risk of bias was used to order studies, never to exclude them.

**Sarah:** And item five, investigating variation, without the statistics of a meta-analysis?

**Kiffer:** You tabulate results against a small number of characteristics that you chose in advance and had a reason to expect might matter. Cedar Valley compared group programs of twelve weeks or less with longer ones. The shorter programs had a median of minus zero point two two and the longer ones minus zero point one eight. With five trials in one group and three in the other, that's no clear difference, so the team reported that program length didn't explain the variation.

**Sarah:** Which connects back to the certainty ratings.

**Kiffer:** It does. That unexplained variation is why GRADE, the Grading of Recommendations Assessment, Development and Evaluation approach, rated the group evidence down one level for inconsistency in Lesson eight.

**Sarah:** What does a good written synthesis sound like, compared with a weak one?

**Kiffer:** A weak one goes trial by trial. Trial D found a significant reduction, Trial H did too, the other six found no significant effect, so the evidence is mixed. A structured one names the eight trials and their participants, says that seven of eight found lower loneliness in the program group with a small median difference, reports that program length didn't explain the variation and that every trial had some risk of bias, and ends with the certainty: group programs may reduce loneliness slightly, with low certainty.

**Sarah:** Which brings us to section three. Vote counting by statistical significance. What is it, exactly?

**Kiffer:** You sort studies into three piles: significant benefit, significant harm and no significant difference. Then you draw a conclusion from the piles. If most are significant and favourable, the program works. If most are non-significant, it doesn't. If the piles are similar, the evidence is mixed. It's easy, and you still see it in published reviews, although the Cochrane Handbook advises against it. Look at the group trials again. Only two of eight confidence intervals exclude zero, so a significance count says two successes and six failures. Yet seven of the eight estimates favour the program, and every one of the six non-significant intervals includes a worthwhile reduction in loneliness as well as zero.

**Sarah:** So a non-significant result doesn't mean the program failed.

**Kiffer:** It means the study couldn't tell. Douglas Altman and Martin Bland put it in the title of a short paper in nineteen ninety-five: absence of evidence is not evidence of absence.

**Sarah:** Why do so many studies fail to tell?

**Kiffer:** Because of statistical power, which is the probability that a study will find a significant result when an effect truly exists. A trial of one hundred people has only about a seventeen in one hundred chance of detecting a standardized mean difference of zero point two. Even with four hundred people, it's about a coin toss. Most Cedar Valley trials had between sixty-four and two hundred and thirty participants.

**Sarah:** Surely adding more studies fixes that.

**Kiffer:** That's the natural intuition, and it fails for the majority rule. Larry Hedges and Ingram Olkin showed in nineteen eighty that when each study has power below one half, the share of significant results settles near that power as studies accumulate. So the chance that most studies are significant falls toward zero. With ten trials of one hundred people each, the chance that more than half are significant is about three in a thousand.

**Sarah:** That's striking. The more underpowered studies you have, the more likely the count is to be wrong.

**Kiffer:** That's right. The count also treats a trial of sixty-four people like one of two hundred and thirty, and it manufactures conflict, because reviewers call the evidence mixed whenever some results are significant and others aren't.

**Sarah:** So what's the acceptable version?

**Kiffer:** Vote counting based on the direction of effect. You count each study by which way its estimate points, whatever its size or significance. For the group programs, seven of eight favour the program. That's about eighty-eight in one hundred, with a confidence interval for the proportion from about fifty-three to ninety-eight in one hundred.

**Sarah:** And you can test that?

**Kiffer:** With a sign test. It asks how likely a split at least this uneven would be if the program had no effect, so that each study were as likely to point one way as the other, like tossing a fair coin. For seven of eight, the two-sided P value is about zero point zero seven.

**Sarah:** Which is above zero point zero five. Doesn't that put us right back where we started?

**Kiffer:** It's a fair challenge, and the same caution applies: a P value above zero point zero five from eight studies is weak evidence either way. The befriending trials make the point more clearly. All five with usable results favour the program, which gives a P value of about zero point zero six, and that is the smallest value the test can produce with five studies. So the team reports the count, the proportion and the P value together, and relies on the effect summaries to describe size.

**Sarah:** And the befriending trials under a significance count?

**Kiffer:** None of the six was significant. A significance count would have recorded zero successes in six and concluded that befriending doesn't work. The direction count shows five of five pointing toward benefit.

**Sarah:** Does direction counting have its own weaknesses?

**Kiffer:** It does, and SWiM item nine asks you to state them. It ignores size, so an estimate of plus zero point zero three counts as a full vote against the program, although it's essentially zero. It gives a small study the same vote as a large one. It needs every result aligned in direction. And each study casts one vote per synthesis, however many outcomes it measured.

**Sarah:** Let's talk about the plots. What's an effect direction plot?

**Kiffer:** Hilary Thomson and Sian Thomas developed it in two thousand thirteen for reviews of complex public health interventions. Each row is a study and each column is an outcome domain. An upward arrow means the effect favours the program, a downward arrow means it favours the comparison, and a sideways arrow means the study's outcomes in that domain conflict. The size of the arrow shows how many participants the study had.

**Sarah:** And it's been updated since?

**Kiffer:** Michele Hilton Boon and Hilary Thomson revised it in two thousand twenty-one, so that the arrows follow the direction of effect regardless of significance and each column can carry a sign test. That brings the plot in line with the twenty nineteen Cochrane guidance.

**Sarah:** What did the Cedar Valley plot show?

**Kiffer:** The team drew it for the ten community connector studies across five domains: loneliness, social isolation, well-being, mental health and health service use. For loneliness, all five controlled studies favour the program, and four of the five uncontrolled studies do too. Social isolation and well-being point the same way in every controlled study that measured them. Mental health and service use are mixed.

**Sarah:** Mixed how, for service use?

**Kiffer:** Two of the controlled studies recorded more unplanned visits among people in the program. The protocol had defined fewer unplanned visits as benefit, so those got downward arrows. The team interpreted them carefully in the text, because a connector might link people to care they had been going without. The plot records the direction, and the text explains what it might mean.

**Sarah:** And the uncontrolled studies?

**Kiffer:** They sit below a labelled divider. Readers can see them, but the statements about effect rest on the controlled studies above the line, which is the prioritization from SWiM item four. And because those controlled studies are non-randomized and mostly at serious risk of confounding, the certainty stays very low.

**Sarah:** Now the harvest plot. The name always makes me think of farming.

**Kiffer:** David Ogilvie and colleagues introduced it in two thousand eight to show evidence about the differential effects of interventions, such as whether they narrow or widen health inequalities. It's a grid. Each study is drawn as a bar in the cell that matches its finding, and the height and shading of the bar show features the review team chooses, such as design and risk of bias.

**Sarah:** How did Cedar Valley use it?

**Kiffer:** For loneliness across the five categories, using the twenty-four controlled studies. The columns show whether a study favours the program, has an unclear direction or favours the comparison. Tall bars are randomized trials, short bars are non-randomized studies, and red bars are at high or serious risk of bias. Twenty of the twenty-four favour the program, one is unclear and three favour the comparison.

**Sarah:** And the pattern that matters?

**Kiffer:** The tall bars sit in the group and one-to-one rows. The connector row holds only short bars, and three of its five are red. So in one picture the planning team can see that the best-designed evidence concerns programs other than the one it plans to launch. It's the same message as the cross-tabulation and the GRADE ratings, seen from another angle.

**Sarah:** Section four asks when pooling is justified. First, what does pooling add?

**Kiffer:** A meta-analysis combines effect estimates into one weighted average with a confidence interval. Gene Glass coined the term in nineteen seventy-six. Pooling adds precision, because it combines several imprecise estimates into one. It gives you an estimate of the size of the effect, which a direction count can't. And it gives you tools for describing how much results vary and for looking for publication bias. Health Sciences two thirty, Lesson two, teaches all of that.

**Sarah:** And what can't it fix?

**Kiffer:** Three things. It can't remove a bias that the studies share. If every trial overstates benefit because participants who know their group report their own loneliness, the pooled estimate overstates it too, only more precisely. It can't give meaning to an average of very different programs. And it can't recover studies that were never published or found.

**Sarah:** So how does a team decide?

**Kiffer:** I suggest five questions, drawn from the Cochrane Handbook. Do the studies address the same question, closely enough that an average would mean something? Does each study give an effect estimate and its precision? Are randomized and non-randomized studies kept apart? Is the risk of bias acceptable or handled? And was pooling planned, with enough studies to estimate how much results vary? A no at any step points to synthesis without meta-analysis.

**Sarah:** Students sometimes think studies have to be identical before you can pool them.

**Kiffer:** Many do, but some diversity is expected. The random-effects model, which you'll meet in Health Sciences two thirty, assumes that the true effect varies from study to study and estimates its average. The question is whether that average means something for the decision at hand.

**Sarah:** And large variation between results? Does that rule pooling out?

**Kiffer:** It's a reason to investigate where the variation comes from and to report the spread of effects alongside any average. It rules pooling out when the variation is so large, or so clearly due to different interventions, that an average would mislead.

**Sarah:** Let's apply the five questions to Cedar Valley, starting with the group programs.

**Kiffer:** The answers are mostly yes. The programs and comparators are similar, standardized mean differences are available for all eight trials, they're all randomized, and the two trials at high risk of bias could be handled with a sensitivity analysis. A full systematic review would normally pool them.

**Sarah:** But the team didn't.

**Kiffer:** It considered amending the protocol and decided against it, for two reasons. The decision facing the planning team concerns connector programs, and a precise number for group programs wouldn't change that decision. And an amendment made after the results table had been seen could look as if the method had been chosen by the result. So the review reports that the group data would support a meta-analysis and recommends one in any full review.

**Sarah:** Could another team reasonably decide differently?

**Kiffer:** Yes, provided it reports the change and the reason, as SWiM item one and the Preferred Reporting Items for Systematic reviews and Meta-Analyses, known as PRISMA, require. The decision is a judgement, and the review's job is to make the judgement visible.

**Sarah:** And the connector studies?

**Kiffer:** Those fail. The programs differ in how people are referred and how long connectors work with them. The studies adjusted for different sets of confounders, and three of the five are at serious risk of confounding. A pooled estimate would combine differently adjusted, seriously confounded results and suggest far more precision than the evidence holds. They also could never be pooled with the randomized trials of other programs.

**Sarah:** So where does all this leave the planning team?

**Kiffer:** With clear statements that combine the synthesis with the certainty ratings from Lesson eight. Group-based programs may reduce loneliness slightly, with low certainty. One-to-one befriending or telephone programs may reduce loneliness, also with low certainty. And the evidence is very uncertain about the effect of community connector programs on loneliness.

**Sarah:** So for connectors, the honest answer is that nobody knows yet.

**Kiffer:** The effect remains open. That's why Lesson eight recommended launching the program with an evaluation built in, for example by introducing it across clinics in stages, so that Cedar Valley adds to the evidence.

**Sarah:** For students who take Health Sciences two thirty, before or after this course, how does this lesson connect to it?

**Kiffer:** Almost everything here is the preparation a meta-analysis needs. The groups decide which studies go into each analysis, the decision rules choose one result per study, and the results table is the data set. Add a pooled result to the plot of ordered estimates and you have a forest plot.

**Sarah:** Finally, let's put it together. How would a team plan the synthesis for a rapid scoping review?

**Kiffer:** First, define the groups and build a cross-tabulation of the included studies, noting any changes from the protocol. Second, write decision rules for choosing one result per study. Third, if the studies report effects, build a results table for the main outcome and draw an effect direction plot or a harvest plot. If the review only maps evidence, cross-tabulate concepts by outcome instead.

**Sarah:** And then?

**Kiffer:** Write a synthesis methods paragraph that addresses SWiM items one to seven, saying which items don't apply and why. Then use the five questions to explain, in two or three sentences, whether pooling would be justified for the main outcome. Together, these pieces give the review's report its synthesis methods and the structure of its results section.

**Sarah:** Any last advice?

**Kiffer:** When you write the synthesis, describe the size of effects as well as their direction, and make the path from your tables to your statements visible. That way your readers can judge for themselves how far the evidence goes.

**Sarah:** Thanks, Kiffer. Next time is Lesson ten, on scoping reviews, rapid reviews and qualitative evidence synthesis.

**Kiffer:** Thanks, Sarah. See you then.
