# Lesson 7: Managing Records and Screening Studies

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. Today we are on Lesson Seven, which is about managing records and screening studies.

**Sarah:** So this is the point where the searching is finished and the reading begins.

**Kiffer:** That's right. In the last few lessons, students built database searches, searched the grey literature and the web, and chased citations. This lesson starts the moment those searches finish. A review team now has thousands of records in several files, and it has to get from that pile to a short list of included studies without losing anything along the way.

**Sarah:** I have to admit that "managing records" sounds like the administrative part of the course.

**Kiffer:** Some of it is administrative, and it is also part of the method. A team that loses a batch of records during import has narrowed its own search without meaning to, and a team that merges two different reports by mistake has thrown away evidence before anyone read it.

**Sarah:** Let's set up the running case first. We are back with Cedar Valley.

**Kiffer:** We are. The Cedar Valley Health Authority is fictional, and it serves about two hundred and ten thousand people in British Columbia. Before it launches a community connector program, which is a form of social prescribing, its planning team has asked an evidence officer, a university librarian and a student intern for a rapid scoping review and an environmental scan. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged sixty-five and older, and with what outcomes.

**Sarah:** How many records are we talking about?

**Kiffer:** The librarian searched five databases: MEDLINE, Embase, CINAHL, PsycINFO and Web of Science. Together they returned two thousand four hundred and eighty records. All of the numbers about Cedar Valley are illustrative, but they are consistent across the lesson.

**Sarah:** Before we get into software, the first section starts with vocabulary. Records, reports and studies.

**Kiffer:** Yes. These definitions come from the PRISMA twenty twenty statement. PRISMA stands for Preferred Reporting Items for Systematic reviews and Meta-Analyses. In its terms, a record is the title or abstract of a report as a database indexes it. A report is a document that tells you about a study, such as a journal article, a conference abstract, a preprint or a government report. And a study is the investigation itself, for example a trial of a befriending program.

**Sarah:** So one study can have several reports.

**Kiffer:** Exactly. A befriending trial might present early results as a conference abstract and then publish a full journal article two years later. That is one study and two reports. And each of those reports might be indexed in three different databases, so you could end up with six records for one study.

**Sarah:** And de-duplication deals with which of those?

**Kiffer:** De-duplication works on records. Its job is to make sure each report enters screening once. So the six records become two. Linking the two reports to the one study happens later, at full-text screening, when a reviewer can read enough to be sure that the abstract and the article really describe the same trial.

**Sarah:** Why wait? It seems tidier to merge them at the start.

**Kiffer:** Because at the start you only have titles and abstracts, and judging whether two reports describe the same study from a title is risky. You might discard a conference abstract that contains results the later article left out. Keeping both and linking them later costs very little.

**Sarah:** Let's talk about the software. What does a reference manager actually do in a review?

**Kiffer:** In a review, a reference manager becomes the team's central store of records. It holds each record as structured fields, such as authors, title, year and abstract, organizes records into folders, finds duplicates, and exports records to other tools. The lesson uses Zotero, which is free and has shared group libraries for teams. Most teams make one folder, which Zotero calls a collection, for each database export, so that the count from each source stays visible.

**Sarah:** How do the records get from a database into Zotero?

**Kiffer:** Usually through a file format called R I S, which is a plain text format where every field sits on its own line with a two-letter tag. T I marks the title, for example. The team exports every record from each final search, with abstracts, and imports each file into its own collection. Then comes the step that people skip, which is comparing the number exported with the number imported.

**Sarah:** Did that matter in Cedar Valley?

**Kiffer:** It did. Embase returned eight hundred and sixteen records, and the first import brought in only five hundred. The librarian had exported them in two batches, and the second file had been saved in a different folder. Because the intern kept an import log, the shortfall was obvious, and the missing batch went in before anyone started screening. Without the log, more than three hundred records would have vanished, and nobody would have known.

**Sarah:** Now the duplicates themselves. Why are there so many?

**Kiffer:** Mainly because databases overlap. A gerontology article in a well-known journal is likely to be indexed in MEDLINE, Embase, CINAHL and PsycINFO, and it may turn up in Web of Science as well. So the same report can arrive four or five times.

**Sarah:** The lesson makes a point about two different kinds of mistake here.

**Kiffer:** Yes. The first kind is a false-positive merge, where the software decides that two records are the same when they actually describe different reports. One of them disappears before anyone screens it, which reduces recall. The second kind is a missed duplicate, where two copies of the same report stay in the set. That wastes some screening time, and nothing is lost. So teams set automated matching to be cautious and send uncertain pairs to a person.

**Sarah:** How does the software decide that two records match?

**Kiffer:** It compares fields. The strongest single field is the digital object identifier, or DOI. Older records and conference abstracts often lack one. Titles differ between databases in punctuation, spelling and capital letters, and author names can appear in full in one database and as initials in another. So good matching cleans up the fields first and then compares several together, such as title, first author, year and journal.

**Sarah:** Walk us through what the Cedar Valley intern did.

**Kiffer:** She worked in three passes and logged each one. Zotero's automated check, with each suggested group reviewed before merging, removed five hundred and seventy-one duplicates. A manual pass, sorting by title and then by first author, found twenty-seven more, mostly records without a DOI whose titles differed in punctuation or spelling. The screening platform flagged another twelve on import. That makes six hundred and ten duplicates and leaves one thousand eight hundred and seventy records. She also found nine look-alike pairs that were different reports, and she kept both records in each pair with a note for the full-text stage.

**Sarah:** What are the other look-alikes students should watch for?

**Kiffer:** A preprint and its published version, a trial protocol and its results paper, and correction notices are all separate reports that resemble the original. Series titles, such as part one and part two, can fool automated matching completely.

**Sarah:** The section ends with the audit trail.

**Kiffer:** An audit trail is a written record that would let someone else repeat the steps and get the same numbers. For this stage it holds the original export files, a log of counts for every source and every pass, and a copy of the library from before anything was merged. Those counts become the first boxes in the flow diagram.

**Sarah:** Let's move to Section Two, which is about screening platforms. First, what do we mean by screening?

**Kiffer:** Screening is deciding, record by record and then report by report, which items meet the review's eligibility criteria. And with nearly two thousand records, the Cedar Valley team needs somewhere for two people to make independent decisions, where disagreements are found automatically, where reasons are recorded, and where the numbers can be counted at the end.

**Sarah:** Could they do that in a spreadsheet?

**Kiffer:** For a few hundred records, a spreadsheet can work. But it makes it hard for two people to decide without seeing each other's choices, it is easy to overwrite rows when you sort them, and you end up counting everything by hand. Dedicated screening platforms were built to solve those problems.

**Sarah:** What do they have in common?

**Kiffer:** They import records and check for duplicates on the way in. They show one record at a time, with chosen keywords highlighted. They let each reviewer decide include, exclude or maybe without seeing the other reviewer's decision, which is called blinding. They list the conflicts, meaning the records on which reviewers disagree. And they store reasons and notes and let you export everything.

**Sarah:** You added a warning about keyword highlighting.

**Kiffer:** I did. Highlighting is useful, because it draws your eye to words that often signal inclusion or exclusion. But a word can appear in a sentence that says the opposite of what the highlight suggests. In the lesson's example, the word residential is highlighted in red, but the abstract says the program reached people at home, unlike residential care. If you screen by colour, you will make mistakes. You still have to read the record.

**Sarah:** The two platforms in the lesson are Covidence and Rayyan. Start with Covidence.

**Kiffer:** Covidence is a web-based subscription service, often available through institutional licences. It organizes a review as a fixed sequence of stages, from import through title and abstract screening and full-text review to extraction. By default it asks for two votes on every record and keeps them hidden from each other. At full text, you cannot exclude a report without choosing a reason from a list the team writes in advance. And it keeps a running count that it can draw as a PRISMA flow diagram.

**Sarah:** And Rayyan?

**Kiffer:** Rayyan was developed at the Qatar Computing Research Institute, and it was described in a paper by Ouzzani and colleagues in twenty sixteen. It has a free plan and paid plans. It puts all the records in one flexible screening space, where reviewers mark include, exclude or maybe and can add reasons and labels. Its blind mode hides each reviewer's decisions until the team switches it off, and then it shows the conflicts.

**Sarah:** So which is better?

**Kiffer:** It depends on the team. Covidence imposes structure, which helps a team with a first-time screener and makes the counts straightforward. Rayyan is more flexible, which suits small teams and student projects, but the team has to supply the structure itself, for example by deciding that every full-text exclusion needs a reason. Cedar Valley used Covidence through the librarian's university licence, and a student team without a licence could do the same work in Rayyan's free plan.

**Sarah:** Then the section turns to setting up a screening project.

**Kiffer:** And the main message is to set everything up before anyone screens a single record, because changing the rules halfway through makes earlier decisions inconsistent with later ones. You create the project and add the team. You set two independent reviewers per record. You import the de-duplicated file and check the count against your log. You write the screening guide into the project, set the exclusion reasons in a fixed order, add keyword highlights, and plan a pilot.

**Sarah:** Tell us about the screening guide.

**Kiffer:** A screening guide turns eligibility criteria into questions a screener can answer from a title, an abstract or a full text. The Cedar Valley guide follows the review's Population, Concept and Context question from Lesson Two. Its five questions ask whether participants are adults aged sixty-five and older living in the community, whether the intervention is delivered in the community, whether loneliness or social isolation was an outcome, whether the study was done in a high-income country, and whether the report is an eligible type of source.

**Sarah:** And each question comes with a rule for borderline cases.

**Kiffer:** That's the important part, because borderline cases cause most disagreements. For the population question, the rule says to include a study if at least eighty percent of participants are sixty-five or older, if results are reported separately for that age group, or, when the ages are not reported in detail, if the mean or median age is sixty-five or older. For the setting question, the rule says that telephone and online delivery count as community-based.

**Sarah:** Why do the exclusion reasons need to be in a particular order?

**Kiffer:** Because many reports fail more than one criterion. Take a commentary about a hospital program for adults aged fifty and older. It fails the population question, the setting question and the publication type question. Under an ordered list, it is recorded once, with the first reason that applies, which is population outside criteria. Without that rule, the counts in your flow diagram would depend on who happened to screen which report.

**Sarah:** Section Three is the screening itself. Two stages.

**Kiffer:** Yes. At title and abstract screening, reviewers remove only what clearly fails a criterion, and everything else moves forward, including records whose abstracts are silent and anything marked maybe. Reasons usually are not recorded at this stage. At full-text screening, reviewers read the whole report and apply every criterion strictly, giving each exclusion exactly one reason from the ordered list.

**Sarah:** Why does the first stage have to be so generous?

**Kiffer:** Because a relevant study excluded on its abstract never reaches anyone who could see that it belonged. The full-text stage can correct a generous decision. Nothing later corrects a strict one.

**Sarah:** What happened at full text in Cedar Valley?

**Kiffer:** One thousand seven hundred and twenty-five records were excluded at title and abstract, so one hundred and forty-five reports were sought. Three could not be obtained, a dissertation and two articles, and they are counted as reports not retrieved. That left one hundred and forty-two reports to assess. One hundred and two were excluded, and forty reports were included.

**Sarah:** And those forty reports describe thirty-eight studies.

**Kiffer:** Right, because two studies each had a conference abstract and a journal article. This is where the linking we talked about earlier actually happens.

**Sarah:** What were the most common reasons for exclusion?

**Kiffer:** Population outside the criteria was the most common, at thirty-two, often studies of adults aged fifty and older with no separate results for the older group. Then no loneliness or social isolation outcome, at twenty-seven, such as an exercise program that measured only falls. Twenty-one were not community-based, nine were not from high-income countries, and thirteen were the wrong kind of publication, such as a protocol with no results yet.

**Sarah:** Before all of this, though, the team ran a pilot.

**Kiffer:** They did. In a pilot, both screeners work through the same random sample of records independently, and then they sit down and compare. The evidence officer and the intern each screened the same two hundred records. They agreed on one hundred and eighty-six and disagreed on fourteen.

**Sarah:** And the disagreements told them something.

**Kiffer:** They told them a lot. Nine of the fourteen involved samples of adults aged sixty and older with no separate results for those sixty-five and older. Three involved telephone befriending programs that one screener had not counted as community-based. And two were simple misreadings. So the team added the rule about mean or median age and the statement that telephone and online programs count. Those rules came directly from the pilot.

**Sarah:** After the pilot, did they split the work?

**Kiffer:** No, both of them screened every remaining record independently, at both stages. That is dual independent screening. Disagreements were resolved by discussion, and anything they could not settle went to the librarian as a third reviewer.

**Sarah:** Is dual screening really necessary? It doubles the work.

**Kiffer:** It does. Careful people make errors when they read hundreds of abstracts, and two people rarely make the same error on the same record. A methodological systematic review by Waffenschmidt and colleagues in twenty nineteen found that single screening missed more eligible studies than double screening, and a randomized trial by Gartlehner and colleagues in twenty twenty reached the same conclusion for abstract screening. The Cochrane Handbook recommends that at least two people decide independently whether each study is eligible.

**Sarah:** But rapid reviews take shortcuts.

**Kiffer:** Some do, for example by having one person screen most records while a second checks a sample or checks all the exclusions. Those shortcuts save time at a known cost, and Lesson Ten looks at when they are acceptable. Cedar Valley is a rapid review, and the team still chose full dual screening, because the number of records was manageable and the planning team was going to use the review to choose a program model.

**Sarah:** Now the statistics. How do you know whether two screeners agree?

**Kiffer:** You start with a two-by-two table. One screener's decisions go across the top and the other's down the side. In the Cedar Valley pilot, both included twenty-two records. The intern included eight that the officer excluded. The officer included six that the intern excluded. And both excluded one hundred and sixty-four.

**Sarah:** And percent agreement is the easy one.

**Kiffer:** Percent agreement is the share of records with matching decisions. Twenty-two plus one hundred and sixty-four is one hundred and eighty-six, out of two hundred, which is zero point nine three, or ninety-three percent. That sounds very good, and that is the problem. Most records in screening are obvious exclusions. Two screeners will agree on most of them whatever their skill, so a high percent agreement can hide poor agreement on the records that actually matter.

**Sarah:** Which is where Cohen's kappa comes in.

**Kiffer:** Yes. Jacob Cohen introduced kappa in nineteen sixty. The idea is to ask how much the two screeners would have agreed by chance, if each made decisions at their own overall rate but without regard to what the records said. The intern included thirty of two hundred records, which is fifteen percent, and the officer included twenty-eight, which is fourteen percent. If their decisions were unrelated, they would both include about two percent of records and both exclude about seventy-three percent. Add those together and chance alone gives agreement of about zero point seven five.

**Sarah:** So most of that ninety-three percent would have happened anyway.

**Kiffer:** Most of it. Kappa takes the agreement they achieved beyond chance, which is zero point nine three minus zero point seven five two, or zero point one seven eight. Then it divides that by the most agreement that was possible beyond chance, which is one minus zero point seven five two, or zero point two four eight. Zero point one seven eight divided by zero point two four eight is about zero point seven two.

**Sarah:** And how should a student read zero point seven two?

**Kiffer:** The most widely cited labels come from Landis and Koch in nineteen seventy-seven, and they would call that substantial. But they proposed those labels as conventions for discussion, and McHugh, writing in twenty twelve, argued they are too generous for health research. The more useful reading comes from the table. Thirty-six records were included by at least one screener, and they agreed on only twenty-two of those. That is the agreement that matters at this stage, and it is why the team revised its guide.

**Sarah:** The lesson also talks about a paradox.

**Kiffer:** Feinstein and Cicchetti described it in nineteen ninety as high agreement with low kappa. When very few records are relevant, chance agreement is already close to one, so even a handful of disagreements pulls kappa down sharply. The practice example in the lesson has two screeners agreeing on more than ninety-five percent of records, with a kappa of only about zero point two eight.

**Sarah:** Is that a flaw in kappa or a real problem with the screening?

**Kiffer:** In that example it is a real problem. Of the eleven records that either screener wanted to include, they agreed on only two. Percent agreement hides that, and kappa reveals it. So my advice is to report the two-by-two table or the number of conflicts alongside kappa, and to look closely at how the screeners did on the includes.

**Sarah:** One more caution about kappa, I think.

**Kiffer:** Yes. Kappa measures consistency between people. It does not measure accuracy. Two screeners who misread a criterion in the same way will agree perfectly. What keeps the decisions faithful to the criteria is the guide, the pilot and the discussion of conflicts.

**Sarah:** Let's move to Section Four, the flow diagram.

**Kiffer:** The PRISMA twenty twenty statement, published by Page and colleagues in twenty twenty-one, is a twenty-seven item checklist for reporting systematic reviews. Item sixteen a asks authors to describe the results of the search and selection process, from the records identified to the studies included, ideally with a flow diagram. Item sixteen b asks them to cite studies that might look eligible but were excluded, and to explain why.

**Sarah:** What is a reader looking for in the diagram?

**Kiffer:** Practical answers about how large the search was, how many records were duplicates, how many survived each stage, whether any reports could not be obtained, why reports were excluded at full text, and how many studies and reports the review included. Those answers help a reader judge whether the search was broad enough and whether the criteria were applied as described.

**Sarah:** The twenty twenty diagram is different from the older one.

**Kiffer:** The original PRISMA diagram, from Moher and colleagues in two thousand nine, had four phases. The twenty twenty version counts records, reports and studies consistently, separates duplicates from records removed by automation tools, adds a box for reports that could not be retrieved, and places study registers beside databases. For reviews that searched beyond databases, it adds a second column for other methods, such as websites, organizations and citation searching.

**Sarah:** And there is more than one template.

**Kiffer:** Four. There are templates for new reviews and updated reviews, and each comes with or without the column for other methods. Cedar Valley is a new review that also searched websites, contacted organizations and chased citations, so it uses the template for new reviews with other sources.

**Sarah:** Walk us down the left-hand column.

**Kiffer:** Records identified from databases: two thousand four hundred and eighty, broken down by database. Records removed before screening: six hundred and ten duplicates, with zero removed by automation tools and zero for other reasons. Records screened: one thousand eight hundred and seventy. Records excluded: one thousand seven hundred and twenty-five. Reports sought: one hundred and forty-five. Not retrieved: three. Reports assessed: one hundred and forty-two. Reports excluded, with the five reasons and their counts: one hundred and two.

**Sarah:** And the right-hand column?

**Kiffer:** The grey literature searches produced five hundred and fifty-seven records after duplicates were removed, and sixty-one reports were sought from them. Organizations the team contacted sent seven documents. Citation chasing produced one thousand one hundred and sixty new records, and twenty-three reports were sought from them. That is ninety-one reports sought. Two could not be obtained, eighty-nine were assessed, and fifty-nine were excluded with reasons. Thirty were included: twenty-six grey-literature documents, such as program evaluation reports and agency reports, and four reports of four more studies from citation chasing.

**Sarah:** So what does the final box say?

**Kiffer:** Forty-two studies included, which is thirty-eight from the databases and four from citation chasing. Forty-four reports of those studies, which is forty plus four. And twenty-six grey-literature documents, shown separately. The team decided at the protocol stage to chart grey-literature documents apart from research studies, so it adapted the final box and explained the adaptation in its methods. PRISMA presents the diagram as a template that authors can adapt.

**Sarah:** And this is a scoping review, which has its own reporting guideline.

**Kiffer:** It does. The PRISMA extension for scoping reviews, developed by Tricco and colleagues in twenty eighteen, asks for the same selection information and uses the phrase sources of evidence, because a scoping review may include documents that are not research studies. That fits the Cedar Valley diagram well. The environmental scan is reported separately, and Lesson Eleven covers that.

**Sarah:** How should a team check its diagram?

**Kiffer:** Someone other than the person who drew it should check every number against the log. Each box should equal the box above it minus the exclusions beside it, the reasons should add up to their total, and the number of studies can never be larger than the number of included reports. Numbers from a drawing tool or a screening platform still have to match the team's own log.

**Sarah:** What errors do you see most often?

**Kiffer:** The most common is numbers that do not add up, usually because the diagram was drawn from memory. Others are final boxes that mix studies with reports, full-text exclusions with no reasons, unretrieved reports folded into the exclusions, and diagrams that disagree with the abstract after a revision.

**Sarah:** Let's finish by putting the whole workflow together. What does a team do once its searches are final?

**Kiffer:** First, the team exports the results of each final search from each database, imports each file into its own Zotero collection, and checks the counts against its search log. It de-duplicates in an automated pass and a manual pass, records what each pass removed, and imports the records into Rayyan or Covidence, checking the count again.

**Sarah:** And then the pilot.

**Kiffer:** Yes. Two screeners pilot-screen a random sample of fifty records independently, using a screening guide built from the review's eligibility criteria. They build their own two-by-two table, calculate percent agreement and Cohen's kappa, talk through every disagreement, and revise the guide.

**Sarah:** And after the pilot?

**Kiffer:** Then they screen the remaining titles and abstracts, assess the full texts, and draft a PRISMA twenty twenty flow diagram with the numbers they have so far. Records removed before screening for any reason other than duplication or an automation tool belong in the box for records removed for other reasons, with a note explaining why. And documents found through grey-literature and artificial intelligence assisted searching belong in the column for other methods. Keep the log as you go, and the diagram will almost draw itself.

**Sarah:** Thanks, Kiffer. Next time, Lesson Eight takes the included studies into data extraction and risk of bias.

**Kiffer:** Thanks, Sarah. See you then.
