HSCI 241 · Lesson 8

Data Extraction, Risk of Bias and Certainty of Evidence

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Design a data extraction form whose domains follow the review question, using the TIDieR checklist for intervention fields and a data dictionary for every field.
  • Pilot an extraction form with two extractors, calculate percent agreement, and revise the form in response to the pattern of disagreement.
  • Distinguish risk of bias from reporting quality, imprecision and applicability, and explain why domain-based judgements have replaced summary quality scores.
  • Describe the structure and judgement scales of RoB 2, ROBINS-I, ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the Mixed Methods Appraisal Tool, and select a tool for each design in a review.
  • Apply a risk-of-bias tool with a second reviewer through calibration, written decision rules, recorded support and consensus, and calculate Cohen's kappa for the judgements.
  • Present risk-of-bias judgements in a traffic-light plot and explain how a review uses them in synthesis.
  • Rate the certainty of a body of evidence with GRADE, from its starting level through the five domains for rating down and the three for rating up, and write a matching plain-language statement.
  • Plan the charting form, the appraisal approach and, where the question concerns effects, a GRADE evidence profile for a rapid scoping review.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 8 · HSCI 241

Data Extraction, Risk of Bias and Certainty of Evidence

This lesson moves from a list of included studies to a judgement about how far the evidence can be trusted.

About three hours, including the final assessment
Running case

The fictional Cedar Valley evidence review

42included studies
14randomized trials
10non-randomized controlled studies
18before-and-after, qualitative and mixed-methods studies

The review asks which community-based interventions have been evaluated for loneliness or social isolation among adults aged 65 and older. All of the review’s numbers are illustrative.

Overview

Four sections

1. Extraction

Design and pilot a form that two people fill in the same way.

2. Tools

Match RoB 2, ROBINS-I, ROBINS-E, the NOS, JBI and the MMAT to designs.

3. Consistency

Use calibration, decision rules and consensus between two reviewers.

4. GRADE

Rate the certainty of a body of evidence for each outcome.

Where this fits

Concepts elsewhere, tools here

HSCI 230

Lessons 7 to 11 teach the sources of bias, and Lesson 2 teaches meta-analysis.

HSCI 241

This lesson summarizes those concepts and teaches the tools that apply them, and Lessons 9 and 12 use the results.

Lesson goals

What you will be able to do

  • Build and pilot a charting form with a data dictionary.
  • Write an appraisal plan that names a tool for each design.
  • Apply a tool with a second reviewer and record the support for each judgement.
  • Rate the certainty of evidence for an outcome with GRADE.
Section 1 of 5

Designing and Piloting a Data Extraction Form

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Designing and Piloting a Data Extraction Form

This section covers the domains of a form, the data dictionary, piloting and the management of extraction.

35 minutes
Purpose

Four jobs of an extraction form

Describe

The form feeds the characteristics table.

Synthesize

It collects categories and results.

Appraise

It records allocation, attrition and measurement.

Audit

It shows where every number came from.

The unit of extraction is the study, and its companion reports are linked under one identifier.

Domains

What the form collects

Study

Source, eligibility check, methods, participants, outcomes, results, funding and conflicts of interest.

Intervention (TIDieR)

What, who, how, where, when and how much, tailoring, modification and fidelity.

TIDieR is a 12-item checklist for describing interventions (Hoffmann et al., 2014).

Data dictionary

Fields that two people fill in the same way

  • Each field holds one fact, with a definition and allowed values.
  • Rules cover unusual cases, such as mixed settings or several time points.
  • Results are copied as reported, with the page or table.
Not reported
Not applicable
Zero
Piloting

The Cedar Valley pilot

85.8%first pilot: 103 of 120 entries agreed
93.1%second pilot: 67 of 72 entries agreed

Disagreements clustered in setting (6), time point (5) and intervention category (4), with 2 copying errors.

Managing extraction

Who extracts, and how errors are caught

Dual independent

Two people extracted the results of the 24 comparative studies.

Single plus checking

One person extracted the other fields and studies, and a second checked every entry.

Single extraction produced more errors than double extraction (Buscemi et al., 2006).

Carry forward

From the form to the tools

The design recorded on the form decides which risk-of-bias tool applies to each study.

Section 2 introduces the tools and the domains of bias they examine.

Learning Objectives for this section

  • Explain what data extraction is, why a review uses a standard form, and how extraction relates to the charting step of a scoping review.
  • List the main domains of an extraction form and use the TIDieR checklist to plan the fields that describe an intervention.
  • Write a data dictionary with definitions, allowed values and instructions so that two people fill in each field the same way.
  • Plan and run a pilot of an extraction form, calculate percent agreement for the pilot, and revise the form in response.
  • Choose between dual independent extraction and single extraction with verification, and handle multiple reports of one study.

Introduction

By the end of Lesson 7, a review team has a list of included studies and a PRISMA flow diagram showing how it got there. The next task is to take the information the review needs out of each study and record it in a form that can be compared, checked and synthesized. This task is called data extraction: the systematic recording of the same set of details from every included study onto a standard form. Much of the work is clerical, but the form decides what the review can say later. A characteristic that nobody extracted cannot be used to group studies, explain differences between them or describe whom the evidence applies to.

The lesson follows the order of the work. Section 1 designs and pilots the extraction form. Section 2 introduces the tools that assess the risk of bias in each study, matched to its design. Section 3 shows how two reviewers apply a tool consistently. Section 4 rates the certainty of the whole body of evidence for an outcome with GRADE. These tools rest on a few concepts of bias, chiefly selection bias, information bias and confounding. Section 2 summarizes them, and HSCI 230 Lessons 7 to 11 teach them in depth for students who take that course.

Case: The Cedar Valley evidence review (fictional)

The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, of whom about 46,000 are aged 65 and older, across Cedar City, several smaller communities and 24 primary care clinics. Before launching a community connector (social prescribing) program for older adults, its planning team asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The evidence team is a health authority evidence officer, a university librarian and a student intern. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes.

The database searches retrieved 2,480 records. After 610 duplicates were removed, the team screened 1,870 titles and abstracts, assessed 142 full texts and included 38 studies. Citation chasing added 4 studies, for 42 included studies in total, and grey literature and website searching added 26 documents. For this lesson, the 42 studies comprise 14 randomized trials (4 of them cluster-randomized), 10 non-randomized controlled studies, 9 uncontrolled before-and-after studies, 6 qualitative studies and 3 mixed-methods studies. These design counts, like the other numbers in this lesson beyond the shared figures above, are illustrative.

In a scoping review, the JBI method calls this step data charting (Peters et al., 2020). Charting tends to be more descriptive and more iterative than extraction for a systematic review of effects, and Lesson 10 returns to it. The design principles in this section apply to both. The Cedar Valley team uses one form for its 42 studies, because the planning team asked about outcomes as well as about the range of programs studied, and a shorter descriptive form for the 26 grey-literature documents, most of which are short program evaluation reports whose methods are too briefly described to classify by design, and which also feed the environmental scan in Lesson 11.

1.1 What an Extraction Form Is For

An extraction form does four jobs. It describes the included studies, which become the characteristics table that readers use to judge whom the evidence covers (Lesson 12). It collects the information the synthesis needs, such as intervention categories and outcome results (Lesson 9). It records the details on which risk-of-bias assessment depends, such as how participants were allocated and how many were lost to follow-up (Sections 2 and 3). And it leaves an audit trail, so that a reader or a later update can see where every number came from.

The unit of extraction is the study. A single study often produces several reports, such as a registry entry, a protocol, a main results paper and a later follow-up paper. These are called companion reports, and the Cochrane Handbook advises collating all reports of a study so that each study is counted once and its information is complete (Li, Higgins and Deeks, 2019). The reverse problem also occurs, when two papers that look like separate studies report overlapping samples. The Cedar Valley form therefore begins with a study identifier and a field that lists every report linked to it.

1.2 What to Collect: The Domains of a Form

Most extraction forms are organized into domains that follow the elements of the review question. The table shows the domains of the Cedar Valley form.

DomainExample fieldsCedar Valley example (illustrative)
SourceStudy identifier, linked reports, extractor, date, contact with authorsCV-017; main paper and registry entry
Eligibility checkConfirmation that population, concept and context criteria are metMean age 74 years; community setting; loneliness measured
MethodsDesign, country, setting, recruitment, dates, length of follow-upCluster-randomized trial in 12 seniors' centres
ParticipantsNumber enrolled and analyzed, age, gender, living arrangement, baseline loneliness240 randomized, 210 analyzed; 61 percent living alone
Intervention and comparatorTIDieR fields: what, who, how, where, how much, tailoringWeekly volunteer-led group walks for 16 weeks, compared with usual activities
OutcomesOutcome, instrument, scoring, time pointsSix-item De Jong Gierveld Loneliness Scale at 16 weeks and 12 months
ResultsNumbers analyzed and statistics exactly as reportedMeans, standard deviations and group sizes at each time point
OtherFunding, conflicts of interest, notes for appraisal, queriesMunicipal grant; no conflicts declared

The intervention fields need the most planning, because community programs for loneliness vary in ways that matter for synthesis. Hoffmann and colleagues (2014) published the Template for Intervention Description and Replication, known as TIDieR, a 12-item checklist for reporting interventions. Reviewers can use its items to structure extraction, which also shows when a study left out something important. The accordion shows how the Cedar Valley team turned TIDieR items into fields.

Why and what: rationale, materials and proceduresv

The form records the program's stated rationale, any materials used and the activities participants take part in. A free-text summary is paired with a closed field for the intervention category so that studies can be grouped later.

Who provided it and howv

The form records the type of provider (paid staff, trained volunteer, peer or health professional), the mode of delivery (in person, telephone or video), and whether the program was delivered to individuals or groups. For connector programs, it records whether the connector was based in a clinic or a community organization.

Where, when and how muchv

The form records the setting, the number, length and frequency of sessions, and the total duration, so that a six-week telephone program can be compared with a year-long group program on the same terms.

Tailoring, modification and fidelityv

The form records whether the program was adapted to individuals, whether it changed during the study, and whether the authors measured how well it was delivered. A program delivered as planned to few participants can explain a null result, so these fields matter for appraisal too.

1.3 Writing Fields That Two People Fill In the Same Way

A field that two people interpret differently produces data that cannot be trusted. The main safeguard is a data dictionary, sometimes called a codebook, which defines every field: its name, a definition, the allowed values, instructions for unusual cases and an example. Several principles make fields reliable. Each field should hold one fact, so "age and gender of participants" becomes two fields. Closed fields with allowed values suit anything that will be used to group or count studies, and free text suits descriptions that readers will need in the authors' own terms. Every field should allow "not reported" and "not applicable" as answers. Extractors should record the page or table where each number was found, and results should be copied exactly as reported, with the type of statistic, the time point and the group sizes; any conversion is done later and documented.

FieldDefinition and allowed valuesInstruction for unusual cases
SettingWhere the program was delivered: community-dwelling; residential care; mixed; not reported.Use "mixed" only when the study includes both settings and does not separate the results.
Intervention categoryThe main component: group activity; one-to-one befriending or telephone contact; community connector or social prescribing; intergenerational; technology-based; other.Assign the category of the component that takes the most contact time, and record any second component in the secondary category field.
Loneliness instrumentThe scale used, with its version and score range.Record the number of items, for example the three-item scale of Hughes et al. (2004) or the six-item scale of De Jong Gierveld and van Tilburg (2006).
Primary time pointThe measurement closest to the end of the intervention.Extract every reported time point, flag the one closest to the end of the program as primary, and flag the longest follow-up up to twelve months as secondary.

Three answers that mean different things

"Not reported" means the study should have said something and did not, such as the number lost to follow-up. "Not applicable" means the field does not apply, such as the number of clusters in an individually randomized trial. "Zero" is a reported value. If all three are left as blank cells, nobody can tell them apart later, and missing information about attrition is itself a reason for concern in appraisal.

1.4 Piloting the Form

Piloting means that two or more people use the draft form on a small, varied set of included studies, compare what they recorded, and revise the form and the data dictionary where they differed. The figure shows where piloting fits in the extraction workflow.

1. Draft the form and data dictionary 2. Pilot two people, varied studies 3. Compare and revise fields, values, rules repeat until agreement is acceptable 4. Extract all studies dual or single plus check 5. Verify and resolve log every discrepancy 6. Lock the dataset for appraisal and synthesis
The extraction workflow. Piloting and revision repeat until two people record the same information with acceptable agreement, and only then does full extraction begin.

The Cedar Valley team piloted its form with the student intern and the evidence officer. Each extracted the same five studies independently: an individually randomized trial, a cluster-randomized trial, a non-randomized controlled study, a qualitative study and a mixed-methods study. The draft form had 24 fields that could be compared directly, giving 5 × 24 = 120 paired entries, and the extractors disagreed on 17 of them.

Percent agreement in the pilot

Percent agreement = (number of entries on which both extractors recorded the same value ÷ total number of paired entries) × 100

First pilot: 103 ÷ 120 × 100 = 85.8 percent. Second pilot, after revision, on three new studies: 67 ÷ 72 × 100 = 93.1 percent.

The pattern of disagreement mattered more than the total. Six of the 17 disagreements concerned setting, because two studies recruited people living at home and in assisted living without separating their results. Five concerned the time point, because studies reported several follow-ups and the extractors chose different ones. Four concerned the intervention category, because one program combined group activities with one-to-one support, and two were copying errors. The team added the "mixed" setting value, the time-point rule and the secondary category field shown in the data dictionary, then piloted the revised form on three new studies (3 × 24 = 72 paired entries), with 5 disagreements. There is no universal threshold for acceptable agreement, so a team should decide in advance what it will accept and should look at which fields disagree, since a high overall percentage can hide one field that is often wrong.

Try it: Apply the category rule

A study describes a program in which a volunteer visits each older adult at home once a week for an hour, and once a month brings the person to a two-hour group lunch. Using the Cedar Valley rule, which category and secondary category would you record? Over a four-week month, the visits provide about four hours of contact and the lunch about two, so the primary category is one-to-one befriending and the secondary category is group activity. Write one sentence for the notes field that would let a later reader check your decision.

1.5 Managing Extraction: Who, How and With What

Errors in extraction are common. Buscemi and colleagues (2006) found that single extraction produced more errors than double extraction, although it took less time. Cochrane guidance therefore recommends that two people extract data independently, particularly outcome data, and resolve differences by discussion. A common compromise is for one person to extract and a second to check every entry against the source, and the Cochrane Rapid Reviews Methods Group recommends this approach for rapid reviews (Garritty et al., 2021). Lesson 10 discusses the trade-offs that rapid-review shortcuts make.

The Cedar Valley protocol, as amended in week 7 (Lesson 10), uses both approaches. For the 24 comparative studies (the 14 randomized trials and the 10 non-randomized controlled studies), two people extract the results fields independently, because errors in effect estimates would change what the brief says about effectiveness. One person extracts the descriptive fields for those studies, and all fields for the other 18 studies, and the other person checks them in full. A discrepancy log records each disagreement, its cause and its resolution, and the team emails authors when key information is missing, recording the request and any reply on the form.

The software matters less than the form. Spreadsheets work for small reviews if the columns are restricted to the allowed values. Screening platforms such as Covidence (Lesson 7) include extraction modules that support dual extraction, and other tools available in 2026 include SRDR+ (the Systematic Review Data Repository Plus, hosted by the United States Agency for Healthcare Research and Quality) and REDCap, a data capture platform that many universities provide. The cards describe errors that the Cedar Valley team watched for.

Extracting from the abstractClick to explore
Mixing up the groupsClick to explore
Confusing measures of spreadClick to explore
Counting a study twiceClick to explore
Recording interpretation as dataClick to explore
Choosing a time point after seeing resultsClick to explore

What Carries Forward

The form now holds, for each of the 42 studies, its design, its methods and the details that bear on bias, such as allocation, blinding, attrition and the outcome measure. The design field decides which appraisal tool applies to each study, and Section 2 introduces those tools.

Reflection

A student team drafted an extraction form for a review of community programs for loneliness among adults aged 65 and older. The draft includes these four fields: (1) Participants (age, gender, number), completed as free text; (2) Setting, completed as free text; (3) Outcome results, completed as free text; and (4) Was the program effective? (yes or no). Two extractors piloted the form on five varied studies. The form had 24 comparable fields, giving 120 paired entries, and the extractors disagreed on 17 of them. Identify the problem with each of the four fields and rewrite each as one or more fields with a definition, allowed values or an instruction. Then calculate the percent agreement in the pilot and state what the team should do before beginning full extraction.

Model answer

Field 1 combines three facts, so extractors will record them in different orders and formats. It should become separate fields: mean or median age (with the statistic named), percentage of women, number randomized or enrolled, and number analyzed, each allowing "not reported". Field 2 as free text will produce many spellings of the same setting; it should be a closed field with the values community-dwelling, residential care, mixed and not reported, with the rule that "mixed" is used only when results are not separated by setting. Field 3 needs several fields: outcome instrument and version, time point, group sizes, type of statistic, the values for each group exactly as reported, and the table or page where they appear, with any conversion done later and documented. Field 4 records an interpretation; it should be replaced by a field quoting the authors' conclusion, separate from the numerical results, so that the review reaches its own judgement.

Agreement was (120 − 17) ÷ 120 = 103 ÷ 120 = 85.8 percent. Before full extraction, the team should list which fields produced the 17 disagreements, revise those definitions and rules in the data dictionary, and pilot the revised form on a few new studies against an agreement level it set in advance.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the Cedar Valley pilot, two extractors compared 120 paired entries and disagreed on 17 of them. What was their percent agreement?

Percent agreement is the number of matching entries divided by all paired entries: (120 − 17) ÷ 120 = 103 ÷ 120 = 85.8 percent. The figure 14.2 percent is the proportion of entries that disagreed, which is the complement of agreement.

Question 2: Why does the Cedar Valley extraction form begin with a study identifier and a list of linked reports?

The unit of extraction is the study. Registry entries, protocols, main papers and follow-up papers about one study are companion reports, and collating them prevents double counting and gives complete information. Risk of bias is assessed for results from the study, using all of its reports together.

Question 3: A study gives no information about how many participants were lost to follow-up. What should the extractor record in the attrition field?

"Not reported" records that information the study should have given is missing, which is itself relevant to appraisal. "Not applicable" is for fields that do not apply to the study, zero is a reported value, and a blank cell cannot be distinguished from either.

Question 4: Buscemi and colleagues (2006) compared single and double data extraction. Which statement reflects their finding and its implication for review teams?

Single extraction produced more errors than double extraction, although it was faster. Cochrane guidance therefore recommends independent extraction by two people, particularly for outcome data, and rapid reviews often use single extraction with a second person checking every entry.
Section 2 of 5

Risk-of-Bias Tools Matched to Study Design

⏱ Estimated reading time: 40 minutes
Section 2 of 5

Risk-of-Bias Tools Matched to Study Design

This section covers RoB 2, ROBINS-I, ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the MMAT.

40 minutes
Definitions

Risk of bias and its neighbours

Risk of bias

Systematic error from design or conduct.

Reporting quality

Completeness of the paper’s description.

Imprecision

Random error shown by wide intervals.

Different quality scales gave different conclusions about the same trials (Jüni et al., 1999).

RoB 2 (Sterne et al., 2019)

Five domains for randomized trials

  • Domain 1 concerns the randomization process.
  • Domain 2 concerns deviations from intended interventions.
  • Domain 3 concerns missing outcome data.
  • Domain 4 concerns measurement of the outcome.
  • Domain 5 concerns selection of the reported result.

Each domain is judged at low risk, some concerns or high risk, and the overall judgement follows explicit rules.

Non-randomized studies

ROBINS-I and ROBINS-E

ROBINS-I

Seven domains judged against a target trial, at low, moderate, serious or critical risk (Sterne et al., 2016).

ROBINS-E

Seven adapted domains for follow-up studies of exposures, with a category for residual confounding (Higgins et al., 2024).

Reviewers list the confounding domains in the protocol before reading the studies.

Other tools

The NOS, the JBI checklists and the MMAT

Newcastle-Ottawa Scale

Up to nine stars for cohort and case-control studies.

JBI checklists

One checklist per design, answered yes, no, unclear or not applicable.

MMAT (2018)

Two screening questions and five criteria in each of five categories.

Cedar Valley

Allocating 42 studies to tools

Two reviewers independently

RoB 2 for 14 trials and ROBINS-I for 10 non-randomized controlled studies.

One reviewer, fully checked

JBI for 9 before-and-after studies and the MMAT for 9 qualitative and mixed-methods studies.

PRISMA-ScR treats appraisal in scoping reviews as optional (Tricco et al., 2018).

Carry forward

From tools to consistent judgements

The tools structure judgement, and reviewers can still disagree.

Section 3 shows how two reviewers reach consistent, documented judgements.

Learning Objectives for this section

  • Distinguish risk of bias from reporting quality, imprecision and applicability, and explain why domain-based tools replaced summary quality scores.
  • Name the five domains of RoB 2 and explain how signalling questions lead to judgements of low risk, some concerns or high risk for a specific result.
  • Describe the seven domains of ROBINS-I, the idea of a target trial, and the judgement levels from low to critical.
  • Describe the purpose and structure of ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the Mixed Methods Appraisal Tool (MMAT).
  • Select an appropriate tool for each design in a review and justify the choice in a protocol.

Introduction

A study can report its methods clearly and enrol thousands of people and still give a misleading answer, because its design or conduct pushed the result away from the truth in a consistent direction. That systematic error is bias, and the risk of bias of a result is the likelihood that features of the study's design or conduct have made it systematically too large or too small. HSCI 230 Lessons 7 to 11 teach the sources of bias: measurement (Lesson 7), selection (Lesson 8), information bias (Lesson 9), design-specific and temporal biases (Lesson 10) and confounding (Lesson 11). The tools in this section turn those concepts into structured questions that a reviewer answers for each included study.

Background: three sources of bias

Confounding is bias caused by a confounder: a common cause of the exposure (here, receiving a program) and the outcome, or a proxy for one, that is not on the causal pathway between them. In a Cedar Valley study that compares adults referred to a connector program with adults in clinics without one, baseline loneliness can influence both who is referred and later loneliness, which is why the team listed it as a confounding domain. ROBINS-I addresses this in its domain on bias due to confounding; in a trial, randomization addresses it, and RoB 2 checks the randomization process in domain 1.

Selection bias arises when entry into or retention in a study depends jointly on the exposure and the outcome, or on their causes. One Cedar Valley trial lost 35 percent of its participants, with more losses in the comparison group; if the loneliest participants were the least likely to return the twelve-week questionnaire, retention would depend on both group and outcome. RoB 2 addresses this in its domain on missing outcome data, and ROBINS-I in its domains on selection of participants and missing data. A sample that is unlike older adults in general raises a different question, external validity.

Information (measurement) bias arises from error in measuring the exposure or the outcome. Measurement error has two separate dimensions: it can be random or systematic, and it can be differential (different in the groups compared) or non-differential. Participants who know they received a befriending program may report their loneliness differently on a self-completed scale, which RoB 2 addresses in its domain on measurement of the outcome.

HSCI 230 Lessons 7 to 11 are optional fuller reading, in particular Lesson 8 on selection, Lesson 9 on information bias and Lesson 11 on confounding.

Risk of bias should be kept separate from three other properties of a study. Reporting quality describes how completely a paper reports its methods, and a poorly reported trial may have been conducted well. Imprecision is random error, shown by a wide confidence interval, and a small unbiased trial is imprecise without being biased. Applicability asks whether the study's population, intervention and setting match the review question. Risk-of-bias tools address bias alone, and GRADE in Section 4 considers imprecision and applicability (which it calls indirectness) across the body of evidence.

Earlier instruments often added points into a quality score. Jüni and colleagues (1999) applied 25 quality scales to the same set of trials and found that conclusions about whether high-quality trials showed different effects depended on which scale was used. Current tools are domain-based: the reviewer judges each type of bias separately, explains each judgement and combines them with explicit rules. The broader term critical appraisal, used particularly by JBI, covers assessments of conduct and relevance as well as bias.

Should a scoping review appraise studies at all?

Scoping reviews usually do not assess risk of bias. PRISMA-ScR lists critical appraisal as an optional item to be reported, with its rationale, if it is done (Tricco et al., 2018), and JBI guidance for scoping reviews generally does not require it (Peters et al., 2020). The Cedar Valley team decided in its protocol to appraise its studies, because the planning team needs to know how far to trust the outcome findings before funding a program, and to report appraisal descriptively without excluding any study because of it.

2.1 Choosing a Tool by Design

The ways in which a randomized trial can go wrong differ from the ways a cohort study or a qualitative study can go wrong, so each design has its own tools. The figure maps common designs to the tools in this section.

Design that produced the result Tool Randomized trial (parallel, cluster, crossover) RoB 2 and its variants Non-randomized study of an intervention ROBINS-I Follow-up (cohort) study of an exposure ROBINS-E Cohort or case-control study (common in existing reviews) Newcastle-Ottawa Scale widely used, with known limits Cross-sectional, prevalence, case series, quasi-experimental, qualitative and others JBI checklist for that design A review that mixes qualitative, quantitative and mixed-methods studies MMAT (one tool, five categories)
Matching designs to appraisal tools. The design of the study as the reviewers judge it, which may differ from the authors' label, decides which tool applies.

Authors' design labels are sometimes wrong, for example when a "pilot trial" allocated participants by alternation, so the extraction form records the design as the reviewers judge it.

2.2 RoB 2 for Randomized Trials

The revised Cochrane risk-of-bias tool for randomized trials, RoB 2 (Sterne et al., 2019), replaced the original Cochrane tool (Higgins et al., 2011), which rated items such as sequence generation and blinding as low, high or unclear risk. RoB 2 assesses a specific result, such as loneliness at twelve weeks, so one trial can receive different judgements for different outcomes. The reviewer states the effect of interest first: the effect of assignment to the intervention (the intention-to-treat effect, which most reviews of public health programs estimate) or the per-protocol effect of adhering to it. Judgements are reached through signalling questions, factual questions answered "yes", "probably yes", "probably no", "no" or "no information". An algorithm maps the answers in each domain to a proposed judgement of low risk of bias, some concerns or high risk of bias, which the reviewer can override with a written reason. The tabs describe the five domains.

Bias arising from the randomization process. This domain asks whether the allocation sequence was random, whether it was concealed until participants were enrolled and assigned, and whether baseline differences suggest a problem. Allocation concealment means that the people enrolling participants could not know or predict the next assignment, so they could not steer particular people into a group. Example: "Was the allocation sequence concealed until participants were enrolled and assigned to interventions?" In one Cedar Valley trial, a coordinator assigned participants from a list she could see in advance, and the groups differed at baseline in the proportion living alone.

Bias due to deviations from intended interventions. For the effect of assignment, this domain asks whether participants and program staff knew the assignments, whether deviations from the intended intervention arose because of the trial context, and whether the analysis kept participants in the groups to which they were randomized. Example: "Was an appropriate analysis used to estimate the effect of assignment to intervention?"

Bias due to missing outcome data. This domain asks whether outcome data were available for all or nearly all randomized participants and, if not, whether missingness could depend on the true value of the outcome. Example: "Could missingness in the outcome depend on its true value?" The loneliest participants may be the least likely to return a follow-up questionnaire. One Cedar Valley trial lost 35 percent of its participants, with more losses in the comparison group.

Bias in measurement of the outcome. This domain asks whether the measurement method was appropriate, whether it could have differed between groups, and whether knowledge of the assigned intervention could have influenced the assessment. Example: "Were outcome assessors aware of the intervention received by study participants?" For self-reported loneliness, the participant is the outcome assessor and usually knows which group he or she is in. Section 3 shows the decision rule the Cedar Valley reviewers wrote for this domain.

Bias in selection of the reported result. This domain asks whether the result was analyzed according to a plan finalized before unblinded outcome data were available, and whether it may have been selected from several scales, time points or analyses. A registry entry or protocol is the usual evidence for this domain.

The overall judgement follows explicit rules. A result is at low risk of bias overall when all five domains are at low risk. It has some concerns when at least one domain raises some concerns and none is at high risk. It is at high risk when at least one domain is at high risk, or when several domains raise some concerns in a way that substantially lowers confidence in the result. A variant for cluster-randomized trials adds a domain on bias arising from the identification or recruitment of participants into clusters, and a variant for crossover trials addresses carry-over and period effects. The Cedar Valley team used the cluster variant for its 4 cluster-randomized trials. Because the developers revise these tools from time to time, a protocol should name the version used.

2.3 ROBINS-I for Non-randomized Studies of Interventions

The Risk Of Bias In Non-randomised Studies of Interventions tool, ROBINS-I (Sterne et al., 2016), assesses studies that compare people who received a program with people who did not, using as a reference a hypothetical randomized trial that would answer the same question without bias. This target trial is a reference point, and it may be hypothetical even when such a trial would be impractical or unethical (Hernán and Robins, 2016). Describing it makes the reviewer specify the population, intervention, comparator and outcome that the study is trying to estimate. The seven domains are arranged by the stage at which bias arises.

Background: the target trial protocol

A target trial is described with the same elements as the protocol of a real randomized trial (Hernán and Robins, 2016): the eligibility criteria; the treatment strategies compared, such as referral to a connector program or usual care; the assignment procedure, which in the target trial is random; time zero, when eligibility is met, a strategy is assigned and follow-up starts; the follow-up period; the outcome, such as loneliness at twelve weeks; the causal contrast, either the effect of assignment or the effect of adhering to a strategy; and the analysis plan. An observational study designed to mimic each element is said to emulate the target trial.

ROBINS-I asks the reviewer to write the target trial for the review question first and then judges each study by how far it departs from it. Without random assignment, the groups may differ in causes of the outcome, so the confounding domain asks how well the study dealt with them. When eligibility, assignment and the start of follow-up do not coincide, the domain on selection of participants applies. When the strategies are defined with information collected after time zero, the domain on classification of interventions applies. The causal contrast decides how deviations from intended interventions are judged, and the last three domains compare the study's follow-up, outcome measurement and reported analysis with the protocol. HSCI 230 Lesson 3, Section 1 (Introduction to Observational Studies), applies the same reasoning to the design of observational studies and is optional reading.

Before the intervention: confounding and selection of participantsv

Bias due to confounding arises when factors that predict the outcome also influence who receives the intervention. Reviewers list the important confounding domains in the protocol before reading the studies, then judge whether each study measured and controlled for them. The Cedar Valley team listed baseline loneliness, depressive symptoms, living alone, mobility or functional limitation, age, and prior social participation. Bias in selection of participants into the study arises when entry into the study or the analysis is related to both intervention and outcome, for example when follow-up begins weeks after people start a program and early dropouts are never counted.

At the intervention: classification of interventionsv

Bias in classification of interventions arises when it is unclear or misrecorded who received the program, particularly when classification is decided with knowledge of the outcome, such as defining "participants" after the results are known as people who attended at least four sessions.

After the intervention starts: deviations, missing data, measurement and reportingv

The last four domains parallel RoB 2. Bias due to deviations from intended interventions includes imbalances in co-interventions, which are other services that may affect the outcome, such as home support. Bias due to missing data, bias in measurement of outcomes and bias in selection of the reported result follow the same logic as in RoB 2.

Each domain and the overall result are judged at low risk (comparable to a well-performed randomized trial), moderate risk (sound for a non-randomized study but not comparable to a well-performed randomized trial), serious risk (some important problems) or critical risk (too problematic to provide useful evidence on the effect), with a further option of no information. The overall judgement is generally the most severe domain judgement. Because confounding can rarely be ruled out without randomization, moderate risk is a good result for a non-randomized study.

2.4 ROBINS-E for Studies of Exposures

Some questions concern exposures that nobody assigns, such as air pollution, shift work or living alone. The Risk Of Bias In Non-randomized Studies of Exposures tool, ROBINS-E (Higgins et al., 2024), was developed for follow-up (cohort) studies of such exposures. It adapts the ROBINS-I domains to bias due to confounding, bias arising from measurement of the exposure, bias in selection of participants into the study or analysis, bias due to post-exposure interventions, bias due to missing data, bias arising from measurement of the outcome, and bias in selection of the reported result. Judgements range from low risk through some concerns and high risk to very high risk, with an additional category of low risk except for concerns about uncontrolled confounding, which recognizes that residual confounding can almost never be excluded in exposure studies. The Cedar Valley question concerns interventions, so the team does not use ROBINS-E.

2.5 The Newcastle-Ottawa Scale

The Newcastle-Ottawa Scale (NOS), developed by Wells and colleagues at the Universities of Newcastle (Australia) and Ottawa, appraises cohort and case-control studies and appears in many published reviews. The cohort version has three categories: selection, with four items (representativeness of the exposed cohort, selection of the non-exposed cohort, ascertainment of exposure, and absence of the outcome at the start); comparability, with one item worth up to two stars for control of the most important factor and any additional factor; and outcome, with three items (assessment of the outcome, adequate length of follow-up and adequacy of follow-up). The case-control version replaces outcome with an exposure category. The maximum is nine stars.

Stang (2010) criticized the scale for unclear items and cut-offs with no stated basis, and Hartling and colleagues (2013) found low agreement between individual reviewers using it. Many reviews convert stars into "good", "fair" or "poor" ratings with thresholds that vary between reviews, which reproduces the problem with summary scores that Jüni and colleagues (1999) described. A review that uses the NOS should report the stars awarded on each item and explain what earns each star for its question. For non-randomized studies of interventions, ROBINS-I is the more structured choice.

2.6 The JBI Checklists

JBI (formerly the Joanna Briggs Institute, at the University of Adelaide) publishes a family of critical appraisal checklists for many designs, including randomized trials, quasi-experimental studies, cohort, case-control and analytical cross-sectional studies, prevalence studies, case series, case reports, qualitative research, economic evaluations, and text and opinion papers. Questions are answered "yes", "no", "unclear" or "not applicable". The analytical cross-sectional checklist has eight questions, including questions on valid measurement and on identifying and dealing with confounders. The qualitative checklist has ten, most of which ask about congruity between the study's philosophical perspective, methodology, methods, analysis and interpretation, along with the researcher's influence and the representation of participants' voices. JBI has revised its checklists for randomized and quasi-experimental studies so that their questions are grouped by the type of bias they address (Barker et al., 2023).

The family covers designs, such as prevalence studies and case series, that the Cochrane tools do not. JBI guidance asks reviewers to decide in advance how appraisal results will be used, for example which answers would lead to exclusion or to reporting with caution (Aromataris and Munn, 2020), and reporting the answer to each question is more informative than a count of "yes" answers. The Cedar Valley team uses the JBI quasi-experimental checklist for its 9 uncontrolled before-and-after studies.

The CASP checklists, from the Critical Appraisal Skills Programme, are a widely used alternative family, with checklists for randomized trials, cohort, case-control and qualitative studies, among others; HSCI 230 and HSCI 841 Lesson 2 use them. Whichever family a review uses, it names the checklist it applies to each design in its protocol.

2.7 The Mixed Methods Appraisal Tool

A review that includes qualitative, quantitative and mixed-methods studies may prefer one tool for all of them. The Mixed Methods Appraisal Tool (MMAT), first developed by Pluye and colleagues (2009) and revised as version 2018 by Hong and colleagues (2018), begins with two screening questions: whether the study has clear research questions, and whether the collected data allow those questions to be addressed. It then offers five design categories, each with five criteria answered "yes", "no" or "can't tell": qualitative studies, quantitative randomized controlled trials, quantitative non-randomized studies, quantitative descriptive studies and mixed-methods studies. A mixed-methods study is appraised against the mixed-methods criteria, which ask about the rationale for combining methods and the integration of components, and against the criteria for its qualitative and quantitative components. The developers discourage an overall score and recommend reporting the rating for each criterion. The Cedar Valley team uses the MMAT for its 6 qualitative and 3 mixed-methods studies. The table summarizes the allocation of all 42 studies.

Design (number of studies)ToolAssessment process
Randomized trials (14, including 4 cluster-randomized)RoB 2, with the cluster variant where needed, for loneliness at the primary time pointTwo reviewers independently, then consensus
Non-randomized controlled studies (10)ROBINS-I, with confounding domains listed in the protocolTwo reviewers independently, then consensus
Uncontrolled before-and-after studies (9)JBI checklist for quasi-experimental studiesOne reviewer, checked in full by a second
Qualitative studies (6)MMAT, qualitative categoryOne reviewer, checked in full by a second
Mixed-methods studies (3)MMAT, mixed-methods and component categoriesOne reviewer, checked in full by a second
Try it: Match the design to the tool

Name a tool for each study. (a) Twenty seniors' centres are randomly allocated to a peer-led walking group or usual activities. (b) A health authority compares loneliness among adults referred to a connector program with similar adults in clinics without it, adjusting for baseline loneliness. (c) A cohort follows adults aged 65 and older for eight years to see whether hearing loss predicts loneliness. (d) Researchers interview 18 participants about a befriending program. Suggested answers: (a) RoB 2 with the cluster variant; (b) ROBINS-I; (c) ROBINS-E; (d) the JBI qualitative checklist, or the MMAT qualitative category in a review that mixes designs.

What Carries Forward

Each of these tools asks for judgement, and two careful reviewers can answer the same question differently. Section 3 describes the procedures that make appraisal consistent.

Reflection

A review of programs for loneliness among older adults includes the following four studies. (a) An individually randomized trial of a weekly telephone befriending program compared with usual care, with loneliness self-reported by participants at twelve weeks. (b) A study comparing 300 adults referred to a community connector program with 300 adults from clinics without the program, adjusting for age and sex only. (c) An eight-year cohort study of whether living alone predicts later loneliness, which the review includes as background on risk factors. (d) A study that combines a participant survey with interviews about a group exercise program, in a review that includes many designs and wants one tool for qualitative, quantitative and mixed-methods studies. For each study, name an appropriate appraisal tool, identify one domain or criterion most likely to raise concern and explain why, and state the judgement scale or response options the tool uses.

Model answer

(a) RoB 2 fits a randomized trial. Domain 4, measurement of the outcome, is the likely concern, because participants report their own loneliness and know whether they received calls. The domain and overall judgements are low risk, some concerns or high risk, reached through signalling questions answered yes, probably yes, probably no, no or no information.

(b) ROBINS-I fits a non-randomized study of an intervention. Confounding is the main concern, because clinicians refer people for reasons related to loneliness, and adjusting only for age and sex leaves baseline loneliness, depressive symptoms and living alone uncontrolled; the study is likely at serious risk. ROBINS-I uses low, moderate, serious and critical risk, with a no-information option.

(c) ROBINS-E fits a follow-up study of an exposure that nobody assigns. Confounding by health and income, and measurement of the exposure over eight years, are likely concerns. Its judgements run from low risk through some concerns and high to very high risk, with a category of low risk except for concerns about uncontrolled confounding. The Newcastle-Ottawa Scale is an alternative if the review must match earlier reviews, reported item by item.

(d) The MMAT fits, using the mixed-methods criteria together with the qualitative and quantitative descriptive categories. Integration of the survey and interview components is a likely concern. Each criterion is answered yes, no or can't tell, and no overall score is calculated.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which list gives the five domains of RoB 2?

RoB 2 has five domains: the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. Option a lists items from the original 2011 Cochrane tool, option b lists ROBINS-I domains, and option d lists Newcastle-Ottawa items.

Question 2: A reviewer is appraising a study that compares older adults referred to a connector program with similar adults who were not referred. Which tool and reference point fit this study best?

A non-randomized comparison of people who did and did not receive a program is a non-randomized study of an intervention, which ROBINS-I assesses against a target trial. RoB 2 requires random allocation, and ROBINS-E is designed for exposures such as living alone that nobody assigns as a program.

Question 3: Which statement about the Newcastle-Ottawa Scale is accurate?

The Newcastle-Ottawa Scale appraises cohort and case-control studies, awarding up to nine stars across selection, comparability and outcome (or exposure for case-control studies). Thresholds for good or poor studies vary between reviews and have no stated basis, which is one of the criticisms of the scale. Option b describes ROBINS-I.

Question 4: How does the 2018 version of the Mixed Methods Appraisal Tool appraise a mixed-methods study?

A mixed-methods study is appraised against the five mixed-methods criteria, which concern the rationale for and integration of the methods, and against the criteria for its qualitative and appropriate quantitative categories. The MMAT developers discourage calculating an overall score.
Section 3 of 5

Applying a Tool Consistently Between Two Reviewers

⏱ Estimated reading time: 35 minutes
Section 3 of 5

Applying a Tool Consistently Between Two Reviewers

This section covers calibration, decision rules, support records, consensus, agreement and presentation.

35 minutes
Preparation

Before the first independent assessment

  • Both reviewers read the full guidance for each tool.
  • The team fixes the effect of interest and lists confounding domains.
  • Reviewers gather the paper, supplements, registry entry and protocol.
  • A calibration exercise on two or three studies surfaces differences early.
Decision rules

The Cedar Valley rule for domain 4

Usual case

Questions 4.3 and 4.4 are answered yes and probably yes, and 4.5 probably no, giving some concerns.

Specific reason for influence

A waiting list or a promise of less loneliness leads to 4.5 probably yes, giving high risk.

Rules are recorded in an appendix or a protocol amendment.

Support records

Trial G, domain 3

230 of 354completed follow-up (Table 2, p. 6)
80 and 44losses in the comparison and intervention groups
Highdomain judgement from the algorithm
Agreement

Comparing two reviewers

81.4%domain judgements agreed (57 of 70)
64.3%overall judgements agreed (9 of 14)
κ = 0.33fair agreement
Cohen’s kappa
\[ \kappa = \frac{p_o - p_e}{1 - p_e} = \frac{0.643 - 0.469}{1 - 0.469} = 0.33 \]
Presentation

Traffic-light plot and its uses

Group-based trials (8)

All 8 have some concerns in domain 4; overall, 6 have some concerns and 2 are at high risk.

Uses in synthesis

Judgements appear beside results, guide stratified tables and feed GRADE.

robvis draws these plots as an R package or web app (McGuinness and Higgins, 2021).

Carry forward

From single studies to a body of evidence

Study-level judgements become one input to a rating for each outcome.

Section 4 rates the certainty of the evidence with GRADE.

Learning Objectives for this section

  • Explain why two trained reviewers can reach different risk-of-bias judgements for the same study, and why reviews use two independent assessors.
  • Prepare for appraisal by reading the tool's guidance, specifying the effect of interest, gathering all reports of each study and running a calibration exercise.
  • Write decision rules that make recurring judgements consistent, and record the support for each answer with a quotation, its location and a reason.
  • Compare two reviewers' judgements, resolve disagreements by consensus or arbitration, and calculate percent agreement and Cohen's kappa for the judgements.
  • Present risk-of-bias results in a traffic-light plot and describe how a review uses them in synthesis without summary scores.

Introduction

The tools in Section 2 structure judgement without removing it. A signalling question such as "Is it likely that assessment of the outcome was influenced by knowledge of intervention received?" asks a reviewer to weigh the outcome, the comparator, the setting and what participants were told. Studies of the tools show that reviewers often disagree. Hartling and colleagues (2013) reported low agreement between individual reviewers using the Newcastle-Ottawa Scale, and Minozzi and colleagues (2020) reported low agreement between reviewers applying RoB 2 and described the difficulties they met in applying it. Much of this disagreement comes from reviewers reading different parts of a report or making different assumptions that can be written down and agreed. The procedures in this section are designed to remove those avoidable sources of disagreement and to make the remaining judgements transparent.

The standard approach is for two reviewers to assess each study independently, compare their answers and reach consensus, with a third person available to arbitrate. The Cedar Valley protocol uses this approach for the 24 comparative studies, which are the studies whose risk of bias most affects what the evidence brief will say about effectiveness, and uses one assessor with full checking by a second for the other 18 studies (Section 2).

3.1 Preparing to Appraise

Preparation prevents most disagreements before they happen. Both reviewers should read the full guidance document for each tool, since the short signalling questions depend on definitions given in the guidance. For RoB 2, the team agrees on the effect of interest before starting, and the Cedar Valley team chose the effect of assignment to the intervention because the planning team wants to know what happens when a program is offered to older adults, including those who attend rarely. For ROBINS-I, the team lists the important confounding domains and co-interventions in the protocol, as Section 2 described. For every study, the reviewers gather all available sources: the main paper, its supplementary files, the trial registry entry, any published protocol or statistical analysis plan, and companion reports identified during extraction. A judgement about selective reporting made without the registry entry is likely to differ from one made with it.

The team then runs a calibration exercise, in which both reviewers independently appraise the same small set of studies, usually two or three per tool, and discuss every difference before appraisal begins. Calibration serves the same purpose as piloting the extraction form in Section 1, and it usually produces the first decision rules. The figure shows the full process.

Prepare guidance, sources Calibrate two or three studies Decision rules written and shared Assess independently Compare answer by answer new rule when needed Consensus third reviewer if needed Record and report
Dual independent assessment. Comparison often reveals a recurring disagreement, which the team settles with a new decision rule that applies to all remaining studies.

3.2 Writing Decision Rules

A decision rule is a written agreement about how the team will answer a signalling question in a situation that recurs across studies. Decision rules leave the tool itself unchanged and record how the team interprets the tool's guidance for its own topic, so that the same situation receives the same answer in every study and so that readers can see the reasoning. The rules belong in an appendix to the review or in a protocol amendment. The table shows the rules the Cedar Valley team wrote during calibration and the first round of comparison.

Tool and domainRecurring situationCedar Valley decision rule
RoB 2, domain 4 (outcome measurement)Loneliness is self-reported, and participants know which group they are in.Answer 4.3 "yes" and 4.4 "probably yes" for every self-reported loneliness outcome. Answer 4.5 "probably no" (which leads to some concerns) unless the report gives a specific reason to expect that knowledge of assignment shaped responses, such as a comparison group told it was on a waiting list or recruitment materials promising less loneliness, in which case answer "probably yes" (which leads to high risk).
RoB 2, domain 3 (missing outcome data)Some participants did not complete the follow-up questionnaire.Treat data as available for nearly all participants when at least 95 percent of randomized participants are analyzed. The RoB 2 guidance explains that what counts as nearly all depends on the outcome and context, so the team states its threshold.
RoB 2, domain 5 (selection of the reported result)No registry entry, protocol or analysis plan can be found.Answer 5.1 "no information". If there is no sign of selection among several scales, time points or analyses, the domain is judged as some concerns.
RoB 2 cluster variant, domain 1bParticipants in cluster trials are recruited after centres are randomized.Record who recruited participants and whether they knew the centre's allocation. Recruitment after randomization by people who knew the allocation leads to at least some concerns.
ROBINS-I, confoundingA study adjusts for some confounding domains listed in the protocol.A study that does not control for baseline loneliness, depressive symptoms and living alone is at serious risk of confounding unless it shows that the groups were similar on the omitted factors.

The first rule came from the most common disagreement in calibration. In two trials whose comparison groups attended a different social activity of similar length, the student intern had answered 4.4 "probably no", reasoning that both groups expected some benefit, and so had judged domain 4 at low risk. The evidence officer had answered "probably yes", because a participant who knows that she is in the "new program" group may still report her loneliness differently. The team agreed that the active comparator is better considered at question 4.5, where it reduces the likelihood of influence without removing the possibility, and wrote the rule accordingly.

3.3 Recording the Support for Each Judgement

Every answer should be supported by a record that another reader could check: a short quotation or summary from the source, its location, and the reviewer's reason when the answer involves judgement. RoB 2 and ROBINS-I provide a free-text box for this purpose beside each signalling question. The support record turns an opinion into an argument, and it makes consensus meetings faster because reviewers can see where their sources differed. The example shows the support record for domain 3 of one Cedar Valley trial, labelled Trial G.

Signalling questionAnswerSupport (source, location and reason)
3.1 Were data for this outcome available for all, or nearly all, randomized participants?NoOf 354 participants recruited in the 12 randomized centres, 230 (65 percent) completed the twelve-week questionnaire: 133 of 177 in the intervention centres and 97 of 177 in the comparison centres (Table 2, page 6).
3.2 Is there evidence that the result was not biased by missing outcome data?NoOnly a complete-case analysis is reported, with no sensitivity analysis for missing data (page 7).
3.3 Could missingness in the outcome depend on its true value?Probably yesReasons for dropout are not reported, and lonelier participants may be less willing to return a questionnaire.
3.4 Is it likely that missingness in the outcome depended on its true value?Probably yesLosses were almost twice as high in the comparison group (80 compared with 44), consistent with disengagement among people offered no program.
Domain judgementHigh riskThe answers follow the path in the RoB 2 algorithm that leads to high risk of bias for this domain.

3.4 Comparing Judgements and Reaching Consensus

After independent assessment, the two reviewers compare their answers question by question and their judgements domain by domain. It helps to classify each disagreement. A disagreement of fact occurs when one reviewer found information that the other missed, such as a registry entry or a table in a supplementary file, and it is resolved by looking at the source together. A disagreement of interpretation occurs when both reviewers saw the same information and answered differently, and it is resolved by discussion with reference to the guidance. If the same interpretive disagreement recurs, the team writes a decision rule and rechecks the studies already assessed. If the reviewers cannot agree, a third person with methodological experience arbitrates, and the record notes that arbitration was used. When the report lacks information that would change a judgement, the team can write to the authors, using the same log as in Section 1.

Many reviews report agreement between the reviewers' independent judgements, before consensus, as an indicator of how difficult the appraisal was. Lesson 7 introduced percent agreement and Cohen's kappa for screening decisions, and the same statistics apply here. In the Cedar Valley review, the two reviewers made 5 domain judgements for each of 14 trials, giving 70 judgements. They disagreed on 13, so they agreed on 57 of 70 (81.4 percent), and 7 of the 13 disagreements were in domain 4. Their independent overall judgements for the 14 trials are cross-tabulated below.

Intern (rows) by evidence officer (columns)Low riskSome concernsHigh riskTotal
Low risk0202
Some concerns0628
High risk0134
Total09514

Agreement on the overall judgements

Observed agreement, po = (0 + 6 + 3) ÷ 14 = 9 ÷ 14 = 0.643, or 64.3 percent.

Agreement expected by chance, pe = [(2 × 0) + (8 × 9) + (4 × 5)] ÷ 142 = 92 ÷ 196 = 0.469.

Cohen's kappa, κ = (po − pe) ÷ (1 − pe) = (0.643 − 0.469) ÷ (1 − 0.469) = 0.174 ÷ 0.531 = 0.33.

On the benchmarks of Landis and Koch (1977), a kappa between 0.21 and 0.40 indicates fair agreement. The two trials that the intern rated at low risk overall were the attention-control trials described in Section 3.2, and the decision rule resolved both.

A kappa of 0.33 is lower than the 81.4 percent agreement on domain judgements might suggest, for two reasons. The overall judgement takes the most severe domain judgement into account, so a single disagreement in any of five domains can change it. And because most trials fell into one category, chance agreement is high, which lowers kappa for a given level of observed agreement. After consensus, the 14 trials were judged at some concerns (9 trials) or high risk of bias (5 trials) for loneliness, and none reached low risk overall. This pattern is common in reviews of psychosocial programs, because participants report their own loneliness and know which program they received, so domain 4 rarely reaches low risk.

3.5 Common Errors in Applying the Tools

Several errors recur in published appraisals, and the Cedar Valley team reviewed them during calibration.

Judging the report instead of the studyClick to explore
Treating a small study as biasedClick to explore
Appraising the study once for all outcomesClick to explore
Adding up the answersClick to explore
Equating lack of blinding with high riskClick to explore
Letting the result shape the judgementClick to explore

3.6 Presenting Risk of Bias and Using It in Synthesis

Risk-of-bias results are usually presented in two figures. A traffic-light plot shows each study as a row and each domain as a column, with a coloured symbol for each judgement. A summary plot shows, for each domain, the proportion of studies at each level of risk. Reviewers can draw these figures by hand, in a spreadsheet, or with robvis, a free tool released as an R package and a web application (McGuinness and Higgins, 2021). The plot below shows the consensus judgements for the 8 Cedar Valley trials of group-based programs, labelled Trial A to Trial H.

D1 D2 D3 D4 D5 Overall Trial ATrial BTrial CTrial D Trial ETrial FTrial GTrial H + + + − + − + + + − + − + + − − + − − + + − + − + + + − + − + + + − − − + + × − + × × − + − + × D1 randomization; D2 deviations from intended interventions; D3 missing outcome data; D4 measurement of the outcome; D5 selection of the reported result. Low risk Some concerns High risk
Traffic-light plot of consensus RoB 2 judgements for loneliness in the 8 illustrative Cedar Valley trials of group-based programs. Every trial has some concerns in domain 4 because loneliness is self-reported by participants who know their group; Trial G is at high risk because of missing outcome data and Trial H because of a problem with randomization.

The review then puts the assessment to use. It should present each study's judgement beside its results, so that readers can see whether the studies at higher risk show different effects. It can stratify or order studies by risk of bias in tables and figures (Lesson 9 shows how), and a review with a meta-analysis can run a sensitivity analysis restricted to studies at lower risk (HSCI 230 Lesson 2). Excluding studies because of their risk of bias is defensible only when the protocol specified the rule in advance, since an exclusion decided after seeing the results can be used to steer the conclusions. Finally, the judgements feed the risk-of-bias domain of GRADE, which Section 4 describes.

Try it: Apply the domain 4 decision rule

A Cedar Valley trial randomizes 140 older adults to a weekly telephone befriending call or to a waiting list, and participants complete the three-item loneliness scale themselves at twelve weeks. The recruitment flyer said that the program "helps people feel less lonely". Using the team's decision rule from Section 3.2, answer signalling questions 4.3, 4.4 and 4.5 and state the domain judgement. Suggested answer: 4.3 is "yes" because the participants are the assessors and know their group; 4.4 is "probably yes" because loneliness is subjective; and 4.5 is "probably yes" because the comparison group knew it was waiting and the flyer set an expectation of benefit, so the domain is at high risk of bias. Write the support record for 4.5 in one or two sentences.

What Carries Forward

The team now has a consensus judgement, with its support, for every comparative study. Those judgements describe individual studies. Section 4 asks a different question: taking all the studies of one intervention and one outcome together, how confident can the team be in what they show? GRADE answers that question for each outcome.

Reflection

Two reviewers independently applied RoB 2 to ten randomized trials of group programs whose outcome was loneliness, self-reported by participants who knew their group. For two trials whose comparison groups attended a different social activity of similar length, Reviewer 1 answered signalling question 4.4 ("Could assessment of the outcome have been influenced by knowledge of intervention received?") as "probably no" and judged domain 4 at low risk; Reviewer 2 answered "probably yes". In RoB 2, a "probably yes" to 4.4 leads to question 4.5 ("Is it likely that assessment of the outcome was influenced by knowledge of intervention received?"), where "probably no" gives some concerns and "probably yes" gives high risk. Their independent overall judgements were: Reviewer 1 rated 2 trials low risk, 6 some concerns and 2 high risk; Reviewer 2 rated 0 low risk, 7 some concerns and 3 high risk. They agreed on 5 trials rated some concerns and 2 rated high risk, and on none rated low risk. (a) Write a decision rule for domain 4 that would resolve the disagreement and apply to future trials. (b) Describe what each reviewer should record as support for question 4.5. (c) Calculate observed agreement and Cohen's kappa, using kappa = (observed agreement − chance agreement) ÷ (1 − chance agreement), where chance agreement is the sum across categories of (Reviewer 1 proportion × Reviewer 2 proportion). (d) Interpret the result.

Model answer

(a) Rule: for self-reported loneliness, answer 4.3 "yes" and 4.4 "probably yes", since participants are the assessors and the outcome is subjective. Consider the comparator at 4.5: answer "probably no" (some concerns) when the comparison group received an active program of similar contact, and "probably yes" (high risk) when it was on a waiting list or the trial promoted an expectation of benefit. Under this rule both disputed trials have some concerns in domain 4.

(b) Each reviewer records the comparator as described, a quotation about what participants were told, the page or table, and a sentence explaining why influence is or is not likely.

(c) Observed agreement = (0 + 5 + 2) ÷ 10 = 0.70. Chance agreement = (0.2 × 0.0) + (0.6 × 0.7) + (0.2 × 0.3) = 0 + 0.42 + 0.06 = 0.48. Kappa = (0.70 − 0.48) ÷ (1 − 0.48) = 0.22 ÷ 0.52 = 0.42.

(d) A kappa of 0.42 indicates moderate agreement on the Landis and Koch benchmarks. It is lower than 70 percent agreement suggests because most trials fall in one category, which raises chance agreement. The disagreements came from one interpretive question, which the rule should remove, and the team should recheck trials already assessed against it.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: One reviewer finds a trial registry entry that the other reviewer missed, and their answers to a domain 5 question differ as a result. How should the team classify and resolve this disagreement?

When one reviewer has information the other missed, the disagreement is one of fact, and the reviewers resolve it by looking at the source together. Decision rules address recurring differences of interpretation, and arbitration is for disagreements that remain after discussion. Judgements are never averaged.

Question 2: Two reviewers' independent overall judgements for 14 trials agree on 9, and chance agreement calculated from their marginal totals is 0.469. What is Cohen's kappa, to two decimal places?

Observed agreement is 9 ÷ 14 = 0.643. Kappa = (0.643 − 0.469) ÷ (1 − 0.469) = 0.174 ÷ 0.531 = 0.33. The value 0.64 is observed agreement, and 0.17 is the numerator before division.

Question 3: A trial paper does not describe how allocation was concealed, and no protocol or registry entry can be found. How should the reviewer answer the RoB 2 signalling question about allocation concealment?

When a report is silent, the answer is "no information", which keeps the judgement about the study's conduct separate from the quality of its reporting. Reviewers should first search the registry, protocol and companion reports and may write to the authors. RoB 2 does ask directly about concealment in domain 1.

Question 4: Why did the Cedar Valley team write a decision rule for domain 4 of RoB 2?

The intern and the evidence officer repeatedly answered questions 4.4 and 4.5 differently for trials with attention-control comparators. A decision rule records how the team will interpret the guidance in that recurring situation. RoB 2 applies to self-reported outcomes, and lack of blinding does not automatically mean high risk.
Section 4 of 5

Rating the Certainty of a Body of Evidence with GRADE

⏱ Estimated reading time: 40 minutes
Section 4 of 5

Rating the Certainty of a Body of Evidence with GRADE

This section covers certainty levels, starting points, rating down and up, and communicating certainty.

40 minutes
Levels

Four levels of certainty

High

⊕⊕⊕⊕

Moderate

⊕⊕⊕◯

Low

⊕⊕◯◯

Very low

⊕◯◯◯

Definitions follow Balshem et al. (2011).

Starting points

Where the rating begins

Randomized trials

The rating starts at high certainty.

Non-randomized studies

The rating starts at low, or at high with ROBINS-I and rating down (Schünemann et al., 2019).

Qualitative findings are assessed with GRADE-CERQual (Lesson 10).

Rating down

Five domains

Risk of bias
Inconsistency
Indirectness
Imprecision
Publication bias

A serious concern lowers certainty by one level, and a very serious concern by two.

Rating up

Three domains (Guyatt et al., 2011)

Large effect

A relative risk above 2 or below 0.5 can justify one level.

Dose-response gradient

More exposure is associated with a larger effect.

Opposing confounding

Plausible confounding would reduce the observed effect.

Worked example

Cedar Valley evidence profile for loneliness

Group-based

8 trials, 1,236 participants. Rated down for risk of bias and inconsistency: low.

One-to-one

6 trials, 742 participants. Rated down for risk of bias and imprecision: low.

Community connector

5 non-randomized studies, 1,480 participants. Rated down for risk of bias: very low.

All results are illustrative and belong to the fictional case.

Communicating certainty

Wording, and what comes next

High

The program reduces loneliness.

Moderate

The program probably reduces loneliness.

Low

The program may reduce loneliness.

Very low

The evidence is very uncertain.

Complete the Section 4 reflection, then the final assessment.

Learning Objectives for this section

  • Explain what the certainty of evidence means in GRADE, and why it is rated for each outcome across a body of studies.
  • State the starting level of certainty for randomized and non-randomized evidence, and describe the alternative start used with ROBINS-I.
  • Describe the five reasons for rating certainty down (risk of bias, inconsistency, indirectness, imprecision and publication bias) and the three reasons for rating it up (a large effect, a dose-response gradient, and plausible confounding that would reduce the effect).
  • Build a GRADE evidence profile for a body of evidence without a pooled estimate, and write an informative statement that matches each level of certainty.
  • Explain how certainty ratings inform an evidence brief, and plan the extraction, appraisal and certainty steps of a rapid scoping review.

Introduction

Risk-of-bias tools describe single studies. A decision-maker asks a different question: taking all the relevant studies together, how confident can we be about the effect of this intervention on this outcome? The Grading of Recommendations Assessment, Development and Evaluation approach, known as GRADE, answers that question with a rating of the certainty of evidence for each outcome. GRADE was developed by an international working group (Guyatt et al., 2008) and is used by Cochrane, the World Health Organization and many guideline developers. Students who take HSCI 230 Lesson 2, before or after this course, meet GRADE there in a short call-out; this section shows how a review team applies it.

Three features of GRADE shape its use. Certainty is rated for a body of evidence, which means all the studies that address one comparison and one outcome, so the same review can have high certainty for one outcome and very low certainty for another. Certainty is rated for outcomes that matter to the people who will use the review, which for the Cedar Valley planning team means loneliness first. And GRADE separates the certainty of evidence from the strength of any recommendation that follows from it. Guideline panels and decision-makers make recommendations after weighing certainty together with benefits, harms, costs, values and feasibility, so a recommendation can be strong even when certainty is low.

4.1 Four Levels of Certainty

GRADE uses four levels. The descriptions below paraphrase those given by Balshem and colleagues (2011), and the symbols are those used in GRADE tables.

LevelSymbolMeaning
High⊕⊕⊕⊕The review team is very confident that the true effect lies close to the estimate.
Moderate⊕⊕⊕◯The team is moderately confident: the true effect is likely to be close to the estimate, but it could be substantially different.
Low⊕⊕◯◯The team's confidence is limited: the true effect may be substantially different from the estimate.
Very low⊕◯◯◯The team has very little confidence: the true effect is likely to be substantially different from the estimate.

GRADE originally called these levels the quality of evidence. The current term, certainty, makes clear that the rating describes confidence in an estimate of effect, which can be low even when every study was carefully conducted.

4.2 Where the Rating Starts

The rating begins from the design of the studies. A body of evidence from randomized trials starts at high certainty, because randomization protects against confounding. A body of evidence from non-randomized studies, including cohort studies, controlled before-and-after studies and other observational designs, starts at low certainty, because confounding and selection can rarely be excluded. The team then considers reasons to rate the evidence down and, less often, reasons to rate it up, as the figure shows.

Randomized trials start here Non-randomized studies start here High ⊕⊕⊕⊕ Moderate ⊕⊕⊕◯ Low ⊕⊕◯◯ Very low ⊕◯◯◯ Rate down (five domains) Risk of biasInconsistencyIndirectnessImprecisionPublication bias Rate up (three domains) Large effectDose-response gradientPlausible confoundingwould reduce the effect Each serious concern moves the rating down one level; a very serious concern moves it two.
The GRADE certainty ladder. The solid arrow shows rating down, which applies to any body of evidence; the dashed arrow shows rating up, which applies mainly to non-randomized evidence with no serious reason to rate down.

A second starting point exists for teams that use ROBINS-I. Because ROBINS-I compares each study with a target trial, the GRADE Working Group has described an approach in which non-randomized evidence assessed with ROBINS-I starts at high certainty and is then rated down for risk of bias by as many levels as the ROBINS-I judgements warrant (Schünemann et al., 2019). Studies at serious risk of bias usually bring the evidence to low certainty or below under this approach, so the two starting points tend to reach similar results. The Cedar Valley protocol states that it uses the conventional low starting point, which is simpler to explain to the planning team.

Two kinds of evidence in the Cedar Valley review sit outside this profile. The 9 uncontrolled before-and-after studies have no comparison group, so they cannot separate the program's effect from change over time, and the team reports them as supporting information. The 6 qualitative studies and the qualitative parts of the 3 mixed-methods studies describe participants' experiences and are assessed with GRADE-CERQual, the companion approach for qualitative evidence that Lesson 10 introduces.

4.3 Five Reasons to Rate Down

For each domain, the team judges whether there is no serious concern, a serious concern (rate down one level) or a very serious concern (rate down two levels). Certainty cannot fall below very low, and the same problem should be counted in only one domain. The accordion describes each domain.

Risk of biasv

This domain draws on the study-level judgements from Sections 2 and 3, considered across the body of evidence. The question is whether the limitations of the studies, weighted by how much each study contributes, lower confidence in the overall result. If most of the information comes from studies with some concerns or high risk of bias, the team rates down. It helps to check whether studies at high risk show different results from the others, since an effect that appears only in studies at high risk warrants more concern.

Inconsistencyv

Inconsistency is unexplained variation in results across studies. Signs include point estimates that differ widely, confidence intervals that barely overlap, and, in a meta-analysis, statistical measures of heterogeneity such as I-squared, which HSCI 230 Lesson 2 teaches. When variation can be explained by a pre-specified characteristic, such as program length or setting, the team can rate the subgroups separately and need not rate down. When studies differ in the size of a benefit but agree on its direction, the concern is smaller than when some show benefit and others show harm.

Indirectnessv

Indirectness arises when the evidence differs from the review question in its population, intervention, comparator or outcome. Examples include trials in residential care when the question concerns older adults living at home, or a measure of social contact used in place of loneliness.

Imprecisionv

Imprecision is random error. The team asks whether the confidence interval, or the range of results across studies, is narrow enough to support a decision. If it includes both an important benefit and no effect, the evidence is imprecise. GRADE also uses the idea of an optimal information size: the number of participants that a single adequately powered trial would need. When the total number of participants in the body of evidence falls short of it, the team usually rates down even if the result is statistically significant.

Publication biasv

Publication bias arises when the studies that reach the review differ systematically from those conducted, usually because studies with favourable results are published more often or more quickly. Suspicion rises when the evidence consists of a few small studies with positive results, when registered trials have not reported, or when a funnel plot is asymmetric. Funnel plots and tests for asymmetry, taught in HSCI 230 Lesson 2, are generally not used with fewer than ten studies. Searching trial registries and grey literature (Lessons 4 and 5) reduces the problem. GRADE describes publication bias as undetected or strongly suspected, and a team usually rates down by one level when it is strongly suspected.

4.4 Three Reasons to Rate Up

GRADE allows the certainty of non-randomized evidence to be rated up in three circumstances (Guyatt et al., 2011). Rating up is usually considered only when no domain has been rated down, since a large effect in studies at serious risk of bias may itself be a product of bias.

A large effectClick to explore
A dose-response gradientClick to explore
Plausible confounding would reduce the effectClick to explore

4.5 Rating Certainty Without a Pooled Estimate

GRADE is often described in terms of a meta-analysis, with a pooled estimate and its confidence interval. Many reviews, including the Cedar Valley review, synthesize results without pooling them, using the methods and the SWiM reporting guideline that Lesson 9 teaches. Murad and colleagues (2017) described how to rate certainty in this situation. The same five domains apply, but the team judges them from the pattern of results across studies: the spread of effects and their direction for inconsistency, the precision of individual studies and the total number of participants for imprecision, and so on. The team explains each judgement in a footnote, since a reader cannot check it against a single number. Rating certainty in a scoping review is unusual, because scoping reviews are designed to map evidence and seldom estimate effects. The Cedar Valley protocol planned no certainty ratings, and a dated amendment in week 8 added them, limited to loneliness for three intervention categories, because the planning team asked directly how confident it could be; the brief labels the ratings as provisional.

4.6 Worked Example: The Cedar Valley Evidence Profile

Case: Rating certainty for loneliness in the Cedar Valley review (fictional)

The team rated certainty for one outcome, loneliness measured with a validated self-report scale at the end of the program, and for three comparisons: group-based programs compared with usual activities or no program (8 randomized trials, 1,236 participants), one-to-one befriending or telephone programs compared with usual care or a waiting list (6 randomized trials, 742 participants), and community connector (social prescribing) programs compared with usual care (5 non-randomized controlled studies, 1,480 participants). The other 5 non-randomized controlled studies evaluated intergenerational and technology-based programs and were too few in each category to rate. All results described here are illustrative and belong to the fictional case.

An evidence profile sets out the judgement for each domain, the resulting certainty and a footnote explaining each decision.

Comparison (studies; participants)Risk of biasInconsistencyIndirectnessImprecisionPublication biasCertainty
Group-based programs (8 randomized trials; 1,236)SeriousaSeriousbNot seriousNot seriousUndetectedc⊕⊕◯◯ Low
One-to-one befriending or telephone (6 randomized trials; 742)SeriousdNot seriousNot seriousSeriouseUndetectedc⊕⊕◯◯ Low
Community connector programs (5 non-randomized studies; 1,480)SeriousfNot seriousNot seriousNot seriousUndetectedc⊕◯◯◯ Very low

a All 8 trials have some concerns (6) or high risk of bias (2) for loneliness, mainly because participants who know their group report their own loneliness (Section 3). b Five trials reported small reductions in loneliness, two reported little or no difference, and one reported a larger reduction; program length and setting did not explain the variation. c Fewer than ten studies per comparison, so funnel plots were not used; registry searches found no completed but unreported trials. d Three of the 6 trials are at high risk of bias and three have some concerns; their results were similar, so the team rated down one level. e The total sample is modest, and most trials' confidence intervals include both no effect and a meaningful reduction. f Three of the 5 studies are at serious risk of confounding by ROBINS-I, because clinicians referred people they judged likely to benefit; the other two are at moderate risk.

Group-based programs start at high certainty as randomized evidence. The team rated down one level for risk of bias and one for inconsistency, giving low certainty. It judged indirectness not serious because all trials enrolled community-dwelling adults aged 65 and older, and imprecision not serious because the trials together enrolled more than a thousand participants. One-to-one programs also start at high, and the team rated down one level for risk of bias and one for imprecision, again giving low certainty. Community connector programs start at low certainty as non-randomized evidence and were rated down one level for risk of bias, giving very low certainty. Two of these studies showed larger reductions in loneliness among people with more connector contacts, and the team considered rating up for a dose-response gradient. It did not rate up, because the risk of bias was serious and people whose loneliness improved early may have chosen to keep meeting their connector, which would produce the same gradient.

4.7 Communicating Certainty

Certainty ratings help decision-makers when the review states them in consistent, plain terms. Santesso and colleagues (2020) proposed standard wording for the statements in reviews and summaries, linking each level of certainty to a verb phrase. The Cedar Valley evidence brief uses this wording, as the table shows.

CertaintyWording patternCedar Valley statement
HighThe intervention reduces the outcome.No Cedar Valley comparison reached high certainty.
ModerateThe intervention probably reduces the outcome.No Cedar Valley comparison reached moderate certainty.
LowThe intervention may reduce the outcome.Group-based programs may reduce loneliness slightly among older adults living in the community. One-to-one befriending or telephone programs may reduce loneliness.
Very lowThe evidence is very uncertain about the effect of the intervention on the outcome.The evidence is very uncertain about the effect of community connector programs on loneliness.

In a full review, these statements appear in a summary of findings table, which presents, for each important outcome, the number of studies and participants, the effect and the certainty. Software such as GRADEpro GDT, available in 2026, helps teams build these tables, and Lesson 12 shows how to present them to decision-makers.

The Cedar Valley ratings have a direct implication for the planning team. The community connector model it intends to launch rests on very uncertain evidence about loneliness, so the true effect could be substantially larger or smaller than existing studies suggest. That rating leaves open whether the program works, and it strengthens the case for launching the program with a built-in evaluation, for example by staggering its introduction across clinics, which the evidence brief in Lesson 12 recommends.

Reflection

A review team rates the certainty of evidence for one comparison: intergenerational programs compared with usual activities, outcome loneliness among older adults living in the community. The evidence is 4 randomized trials with 510 participants. Risk of bias for loneliness: 1 trial at low risk, 2 with some concerns and 1 at high risk, with most participants in the trials that have some concerns or high risk. Results: two trials found moderate reductions in loneliness and two found slight increases, and no planned subgroup explains the difference. Two of the four trials were conducted in residential care homes. The confidence interval of every trial includes no effect. There are fewer than ten trials, and a registry search found no completed but unreported trials. In GRADE, randomized evidence starts at high certainty, each serious concern rates it down one level and each very serious concern two levels, and certainty cannot fall below very low. Rate each of the five domains with a reason, state the final certainty, and write a plain-language statement using the GRADE wording for that level. Then explain what the rating does and does not tell a planning team.

Model answer

Risk of bias: serious, because most of the information comes from trials with some concerns or high risk, mainly from self-reported loneliness in unblinded trials (rate down one level). Inconsistency: serious, because two trials suggest benefit and two suggest slight harm, with no explanation (rate down one). Indirectness: serious, because half the trials enrolled residents of care homes, while the question concerns older adults living in the community (rate down one). Imprecision: serious, because the total sample is modest and every confidence interval includes no effect. Publication bias: undetected, because funnel plots are not informative with four trials and the registry search found no unreported trials.

Starting at high, the first three domains already bring the rating to very low, and imprecision cannot lower it further, so the certainty is very low. Statement: "The evidence is very uncertain about the effect of intergenerational programs on loneliness among older adults." A team could defensibly judge risk of bias not serious, given one low-risk trial, and the rating would still be very low.

The rating tells the planning team that the true effect could be substantially different from what these trials suggest, in either direction. The rating leaves open whether intergenerational programs work, so a team that adopted one would need to evaluate it.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the conventional GRADE approach, at what level does certainty start for a body of evidence from randomized trials and from non-randomized studies?

Randomized evidence starts at high certainty and non-randomized evidence at low certainty. When ROBINS-I is used, GRADE guidance also allows non-randomized evidence to start at high and be rated down for risk of bias, but the number of levels depends on the ROBINS-I judgements.

Question 2: Which list contains only reasons for rating certainty down in GRADE?

The five reasons for rating down are risk of bias, inconsistency, indirectness, imprecision and publication bias. A large effect, a dose-response gradient and plausible confounding that would reduce the effect are the three reasons for rating up.

Question 3: A body of evidence from eight randomized trials is rated down one level for risk of bias and one level for inconsistency, with no other concerns. What is its certainty, and which statement fits it?

Starting at high and rating down two levels gives low certainty. The GRADE wording for low certainty uses "may" (Santesso et al., 2020); "probably" signals moderate certainty, and "very uncertain" signals very low.

Question 4: Why did the Cedar Valley team decline to rate up the community connector evidence for a dose-response gradient?

Rating up is usually considered only when there is no serious reason to rate down, and the connector studies had serious risk of bias. People whose loneliness improved early may have continued meeting their connector, which would produce a gradient without a causal effect of dose. Rating up applies mainly to non-randomized evidence.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson followed the Cedar Valley evidence team from a list of 42 included studies to a statement about how confident the planning team can be in what those studies show. The extraction form, built on the TIDieR items and a data dictionary and piloted until two people recorded the same information, determined what the review could later compare and appraise. The design recorded on that form decided which appraisal tool applied: RoB 2 for the randomized trials, ROBINS-I for the non-randomized controlled studies, the JBI quasi-experimental checklist for the uncontrolled before-and-after studies, and the MMAT for the qualitative and mixed-methods studies, with ROBINS-E and the Newcastle-Ottawa Scale described for the exposure and observational questions where students will meet them.

The tools structure judgement without removing it, which is why calibration, decision rules, recorded support, independent assessment and consensus matter. In the Cedar Valley review, no trial reached low risk of bias overall for loneliness, because participants report their own loneliness and know which program they received. GRADE then moved the question from single studies to bodies of evidence: group-based and one-to-one programs reached low certainty, and community connector programs, the model the health authority intends to launch, reached very low certainty. That rating leaves the program's effect open, which supports launching it with an evaluation built in.

Key Takeaways from this lesson

  • Data extraction records the same details from every included study on a standard form, and the form determines what the review can later compare, appraise and report.
  • A data dictionary with definitions, allowed values and rules for unusual cases, together with a pilot by two extractors, makes extraction consistent and shows which fields need revision.
  • Risk of bias concerns systematic error from a study's design or conduct, and it is assessed separately from reporting quality, imprecision and applicability.
  • RoB 2 assesses a specific result from a randomized trial across five domains, using signalling questions and an algorithm to reach judgements of low risk, some concerns or high risk.
  • ROBINS-I compares a non-randomized study of an intervention with a target trial across seven domains, and ROBINS-E adapts that approach to follow-up studies of exposures.
  • The Newcastle-Ottawa Scale is widely used for cohort and case-control studies, but its summary stars have known problems, and item-level reporting is preferable.
  • The JBI checklists cover many designs that the Cochrane tools do not, and the MMAT offers one tool for reviews that mix qualitative, quantitative and mixed-methods studies.
  • Consistent appraisal depends on calibration, written decision rules, recorded support for each answer, independent assessment by two reviewers and a consensus process.
  • GRADE rates certainty for each outcome across a body of evidence, starting at high for randomized and low for non-randomized evidence, rating down for five reasons and up for three.
  • Certainty ratings should be communicated with standard wording, so that decision-makers can tell what the evidence shows and how much weight it can bear.

Core Concepts Reviewed

Section 1: data extraction and charting, companion reports, the domains of an extraction form, TIDieR, the data dictionary, piloting and percent agreement, and dual extraction compared with single extraction and verification.

Section 2: risk of bias compared with reporting quality and imprecision, the problems with summary scores, RoB 2 and its five domains, ROBINS-I and the target trial, ROBINS-E, the Newcastle-Ottawa Scale, the JBI checklists and the MMAT.

Section 3: preparation and calibration, decision rules, support records, disagreements of fact and of interpretation, consensus and arbitration, Cohen's kappa for appraisal judgements, traffic-light plots and the use of risk of bias in synthesis.

Section 4: certainty of evidence, the four GRADE levels, starting levels, the five domains for rating down and three for rating up, rating without a pooled estimate, evidence profiles and informative statements.

The final reflection asks you to write the extraction, appraisal and certainty methods for a new rapid scoping review, drawing on all four sections.

Reflection

A health authority asks a two-person evidence team for a twelve-week rapid scoping review of walking-group programs for loneliness among adults aged 65 and older. Preliminary screening suggests the included studies will be randomized trials (some cluster-randomized), non-randomized controlled studies, uncontrolled before-and-after studies and qualitative interview studies. The authority wants to know what has been studied and how confident it can be that walking groups reduce loneliness. Write the methods text (about 200 words) for the extraction, appraisal and certainty steps of the protocol. It should state how data will be extracted and checked, whether studies will be appraised and why, which tool will be used for each design, how the two reviewers will make their judgements consistent, how the appraisal will be used in the synthesis, and whether and how GRADE will be applied, including the starting levels. End with the sentence the team would use in the brief if the randomized evidence were rated low certainty for a small reduction in loneliness.

Model answer

Data will be charted on a form with a data dictionary, structured by the TIDieR items for intervention fields and piloted by both reviewers on five varied studies before full charting. One reviewer will chart descriptive fields and the second will check them in full; outcome results from comparative studies will be extracted independently by both. Although appraisal is optional in scoping reviews, studies will be appraised because the authority must judge how far to trust reported outcomes. Randomized trials will be assessed with RoB 2 (with the cluster variant where needed), non-randomized controlled studies with ROBINS-I against the confounding domains listed in this protocol, uncontrolled before-and-after studies with the JBI quasi-experimental checklist, and qualitative studies with the JBI qualitative checklist or the MMAT. Both reviewers will calibrate on two studies per tool, record decision rules and supporting quotations, assess comparative studies independently and resolve disagreements by consensus, with a third person arbitrating. Judgements will be reported by domain in a traffic-light plot, without summary scores, and no study will be excluded on that basis. Certainty for loneliness will be rated with GRADE, starting at high for randomized and low for non-randomized evidence, using guidance for evidence without a pooled estimate. Brief wording: "Walking-group programs may reduce loneliness slightly among older adults."

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Data Extraction, Risk of Bias and Certainty of Evidence (15 Questions)

Question 1: A review team wants to describe each included study's intervention in enough detail to compare programs. Which resource gives the most suitable structure for those extraction fields?

TIDieR (Hoffmann et al., 2014) lists twelve items for describing an intervention, such as what was delivered, by whom, how, where and how much, and reviewers use it to structure extraction fields. PRISMA 2020 concerns the review report, and GRADE and the MMAT concern certainty and appraisal.

Question 2: A Cedar Valley trial reports self-reported loneliness and emergency department visits taken from health records. Why might RoB 2 give these two results different judgements?

RoB 2 assesses the risk of bias in a specific result. Participants who know their group may report loneliness differently, while health-record counts of visits are less open to that influence, so domain 4 can differ between the two results.

Question 3: Which pairing of appraisal tool and study design is appropriate?

JBI publishes a checklist for prevalence studies, which suits a survey estimating how common loneliness is. A cluster-randomized trial needs RoB 2 with its cluster variant, a cohort study of an exposure suits ROBINS-E or the Newcastle-Ottawa Scale, and an interview study needs a qualitative tool.

Question 4: How do risk of bias and imprecision differ?

Bias is systematic error that pushes a result in a consistent direction because of how a study was designed or conducted. Imprecision is random error, reflected in wide confidence intervals, and GRADE assesses it as a separate domain.

Question 5: A scoping review team asks whether it must appraise its included studies. Which answer reflects PRISMA-ScR and JBI guidance for scoping reviews?

PRISMA-ScR lists critical appraisal as an optional item to be reported with its rationale if done (Tricco et al., 2018), and JBI guidance generally does not require it (Peters et al., 2020). The Cedar Valley team chose to appraise because the planning team needs to judge how far to trust outcome findings.

Question 6: In ROBINS-I, which description matches a study judged at moderate risk of bias overall?

Moderate risk means the study is sound for a non-randomized study but cannot be considered comparable to a well-performed randomized trial. Low risk means comparable to such a trial, critical risk means too problematic to provide useful evidence, and a separate option covers missing information.

Question 7: Which finding would justify rating down the Cedar Valley evidence for indirectness?

Indirectness arises when the population, intervention, comparator or outcome of the evidence differs from the review question. Attrition belongs to risk of bias, unexplained variation to inconsistency, and wide confidence intervals to imprecision.

Question 8: After seeing the results, a team decides to exclude all studies at high risk of bias from its synthesis. Why is this a problem?

Excluding studies by risk of bias is defensible only when the protocol set the rule in advance, because a rule chosen after the results are known can be used to produce a preferred conclusion. Sensitivity analyses and stratified presentation are other ways to use the judgements.

Question 9: Which feature distinguishes ROBINS-E from ROBINS-I?

ROBINS-E (Higgins et al., 2024) is for follow-up studies of exposures such as pollution, occupation or living alone, while ROBINS-I is for interventions. ROBINS-E includes a confounding domain and a judgement category that acknowledges residual confounding.

Question 10: In an extraction pilot, overall agreement is 93 percent, but the primary time-point field disagrees in most studies. What should the team do?

A high overall percentage can hide one unreliable field. Revising the definition, for example with a written time-point rule, and piloting again addresses the problem. Letting extractors choose freely invites selection that depends on the results.

Question 11: Which statement about funnel plots in GRADE's publication bias domain matches the lesson?

With fewer than ten studies, funnel plots and tests for asymmetry have little power and are generally not used. Searches of trial registries and grey literature provide other evidence about unreported studies, and symmetry does not prove that none are missing.

Question 12: Why do RoB 2 assessments of community programs for loneliness rarely reach low risk in domain 4?

For a self-reported outcome, the participant is the outcome assessor, and participants in community programs almost always know their assignment, so knowledge of the intervention could influence the assessment. Validated scales exist, and RoB 2 does not rate unblinded trials as high risk automatically.

Question 13: An evidence brief states: "The evidence is very uncertain about the effect of community connector programs on loneliness." What does this wording tell the planning team?

"Very uncertain" is the GRADE wording for very low certainty, which means the true effect is likely to differ substantially from the estimate. This wording leaves open whether the programs work, and it supports launching a program with a built-in evaluation.

Question 14: Why should a review avoid adding Newcastle-Ottawa stars or checklist answers into a single score?

Summary scores weight unrelated items equally, and a study with one serious flaw can still score well. Jüni and colleagues (1999) showed that different quality scales led to different conclusions about the same trials. Item-level or domain-level reporting is preferred.

Question 15: Starting from low certainty, which situation would most plausibly justify rating non-randomized evidence up by one level?

GRADE suggests rating up one level for a large effect, for example a relative risk below 0.5 or above 2, when the evidence is consistent and free of serious threats to validity (Guyatt et al., 2011). Studies at serious risk of bias or with imprecise results are not candidates for rating up.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Data Extraction, Risk of Bias and Certainty of Evidence.

You can now design and pilot an extraction form, choose and apply a risk-of-bias tool for each design in a review, work with a second reviewer to reach consistent and documented judgements, and rate the certainty of a body of evidence with GRADE.

Lesson 9 turns to synthesizing quantitative findings: tabulating and grouping studies, synthesis without meta-analysis and the SWiM reporting guideline, why vote counting by statistical significance misleads, and when pooling is justified.

Continue to Lesson 9 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

These terms, tools and people appear in this lesson on data extraction, risk of bias and certainty of evidence.

Core Concepts
Data extraction The systematic recording of the same set of details from every included study onto a standard form, so that studies can be compared, appraised and synthesized.
Data charting The JBI term for extraction in a scoping review, which tends to be more descriptive and iterative than extraction for a review of effects.
Companion reports Several reports of the same study, such as a registry entry, protocol, main paper and follow-up paper, which are collated so that the study is counted once.
Data dictionary A document that defines each field on an extraction form, with allowed values, instructions for unusual cases and an example.
Risk of bias The likelihood that features of a study's design or conduct have made a result systematically too large or too small.
Critical appraisal The systematic assessment of a study's conduct, trustworthiness and relevance, a broader term that JBI uses for its checklists.
Reporting quality How completely a paper describes its methods and results, which differs from how well the study was conducted.
Signalling questions Factual questions in RoB 2 and ROBINS-I, answered yes, probably yes, probably no, no or no information, that lead to a domain judgement.
Allocation concealment Arrangements that prevent the people enrolling trial participants from knowing or predicting the next assignment.
Effect of assignment The effect of being assigned to an intervention regardless of adherence, also called the intention-to-treat effect; RoB 2 also allows assessment of the per-protocol effect of adhering to it.
Target trial A hypothetical randomized trial that would answer a study's question without bias, described by the elements of a trial protocol (eligibility, treatment strategies, assignment, time zero, follow-up, outcome, causal contrast and analysis plan) and used by ROBINS-I as the reference for judging a non-randomized study.
Confounding domains Factors, listed in a review protocol before reading the studies, that may predict both receipt of an intervention and the outcome.
Decision rule A written agreement about how a review team will answer a signalling question in a situation that recurs across studies.
Calibration exercise Independent appraisal of the same few studies by all reviewers, followed by discussion of every difference, before full appraisal begins.
Certainty of evidence In GRADE, the extent of confidence that an estimate of effect for an outcome is close to the true effect, rated high, moderate, low or very low.
Inconsistency Unexplained variation in results across studies, one of the five GRADE reasons for rating certainty down.
Indirectness Differences between the evidence and the review question in population, intervention, comparator or outcome, a GRADE reason for rating down.
Imprecision Random error in an estimate, shown by wide confidence intervals or a small total sample, a GRADE reason for rating down.
Publication bias Systematic difference between the studies that reach a review and those conducted, usually because favourable results are published more often.
Frameworks & Tools
RoB 2 The revised Cochrane risk-of-bias tool for randomized trials, with five domains and judgements of low risk, some concerns or high risk (Sterne et al., 2019).
ROBINS-I The Risk Of Bias In Non-randomised Studies of Interventions tool, with seven domains judged at low, moderate, serious or critical risk (Sterne et al., 2016).
ROBINS-E The Risk Of Bias In Non-randomized Studies of Exposures tool for follow-up studies of exposures, with seven domains (Higgins et al., 2024).
Newcastle-Ottawa Scale A star-based tool for cohort and case-control studies, with up to nine stars across selection, comparability and outcome or exposure.
JBI critical appraisal checklists A family of design-specific checklists from JBI, answered yes, no, unclear or not applicable.
Mixed Methods Appraisal Tool (MMAT) A tool with two screening questions and five criteria in each of five design categories, for reviews that mix qualitative, quantitative and mixed-methods studies (Hong et al., 2018).
GRADE The Grading of Recommendations Assessment, Development and Evaluation approach for rating the certainty of a body of evidence for each outcome.
TIDieR The Template for Intervention Description and Replication, a 12-item checklist for describing interventions (Hoffmann et al., 2014).
Traffic-light plot A figure showing each study as a row and each risk-of-bias domain as a column, with a coloured symbol for each judgement.
Evidence profile A GRADE table that sets out the judgement for each domain, the resulting certainty and an explanatory footnote for each outcome.
Summary of findings table A table presenting, for each important outcome, the number of studies and participants, the effect and the certainty of evidence.
Key People
Jonathan Sterne University of Bristol epidemiologist and medical statistician who led the development of ROBINS-I and RoB 2.
Julian Higgins University of Bristol professor of evidence synthesis, an editor of the Cochrane Handbook, who led the original Cochrane risk-of-bias tool and the ROBINS-E paper.
Gordon Guyatt McMaster University clinical epidemiologist who coined the term evidence-based medicine and is a founding member of the GRADE Working Group.
Holger Schünemann McMaster University clinical epidemiologist and a co-chair of the GRADE Working Group, lead author of guidance on using ROBINS-I with GRADE.
George Wells University of Ottawa epidemiologist and biostatistician who led the development of the Newcastle-Ottawa Scale.
Pierre Pluye McGill University professor of family medicine who developed the original Mixed Methods Appraisal Tool.
Quan Nha Hong Researcher who led the development of the 2018 version of the Mixed Methods Appraisal Tool.
Tammy Hoffmann Bond University researcher who led the development of the TIDieR checklist for describing interventions.
Peter Jüni Epidemiologist whose 1999 study with colleagues showed that different quality scales led to different conclusions about the same trials.
No matching entries. Try a different search term.