# Lesson 6: Artificial Intelligence in Evidence Retrieval and Synthesis

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are in Lesson six of Finding and Synthesizing Health Evidence, and the topic is artificial intelligence in evidence retrieval and synthesis.

**Sarah:** I suspect this is the lesson a lot of students have been waiting for. Many of them already use these tools.

**Kiffer:** They do, and so do I. These tools help with parts of a review, and some of what they produce is wrong in ways that are hard to see. If you understand how a tool produces its answer, you can decide where it helps and how to check it.

**Sarah:** Let's set the scene with the running case. Remind us about Cedar Valley.

**Kiffer:** The Cedar Valley evidence review is a fictional project. The planning team at the fictional Cedar Valley Health Authority in British Columbia is about to launch a community connector program, a form of social prescribing, for older adults. They ask a small evidence team, made up of an evidence officer, a university librarian and a student intern, for a rapid scoping review and an environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged sixty-five and older, and with what outcomes.

**Sarah:** And by this lesson they have already searched.

**Kiffer:** They have searched five databases and retrieved two thousand four hundred and eighty records, and after removing six hundred and ten duplicates, one thousand eight hundred and seventy remain. Now the intern suggests asking a conversational assistant for the key studies. The librarian agrees on two conditions: every reference is checked against a primary source and recorded in a log, and the test is documented so that a reader can see what was done. The evidence officer adds a third, which is that nothing from the environmental scan goes into any tool.

**Sarah:** Those conditions sound cautious. Before we get to why, what is a large language model actually doing when it answers me?

**Kiffer:** It works with tokens, which are short pieces of text such as a word, part of a word or a punctuation mark. In training, the model reads an enormous amount of text and learns to predict the next token from the ones before it, and it stores what it learns as billions of numbers called parameters. When you ask a question, it calculates a probability for every possible next token, picks one, adds it to the text and calculates again. A setting often called temperature controls how much chance enters each pick.

**Sarah:** So it is never looking anything up.

**Kiffer:** A model on its own has no catalogue of articles to consult. What it knows about the literature is spread through its parameters as patterns, so when it writes a reference, it is producing text that resembles references it has seen.

**Sarah:** The lesson lists four consequences of that design.

**Kiffer:** The first is the one I just described: references come from patterns. The second is the knowledge cutoff, since without a search tool the model knows nothing published after its training text ends. The third is variability, because chance enters each choice, so a repeated prompt can give a different list of studies. The fourth is that fluency tells you nothing about accuracy. The model writes confidently whether or not it is right.

**Sarah:** That last one is uncomfortable, because confidence is exactly what persuades us.

**Kiffer:** It is. Bender and colleagues argued in twenty twenty-one that these models assemble plausible language without grounding in meaning. You don't need to settle that debate. You need to recognise that the output is shaped by the training text and by the prompt. If you ask for studies showing that befriending reduces loneliness, you invite a list of supportive studies. A neutral prompt that asks what has been evaluated, and with what results, gives the model less to agree with. HSCI eight forty-one, Lesson twelve, applies the same cautions to using these models for qualitative coding.

**Sarah:** A lot of tools now search the web before they answer. Does that fix the problem?

**Kiffer:** It changes it. The approach is called retrieval-augmented generation, a term Lewis and colleagues introduced in twenty twenty. A retriever searches an index of documents, the passages it ranks highest go into the model's context window, which is the text the model can consider at once, and the model writes an answer that cites them.

**Sarah:** So what does retrieval change?

**Kiffer:** Because each citation points to a document the retriever actually returned, the tool is much less likely to invent a reference outright. That is a real improvement. Other errors remain. The model can attach a claim to a source that does not support it, for example by turning a cautious conclusion into a confident one, and the retriever can bring back weak or off-topic documents, which the model will summarize just as fluently as strong ones.

**Sarah:** I think there is a connection to Lesson three here.

**Kiffer:** A direct one. Lesson three contrasted Boolean retrieval in a database, which returns every record that matches, with relevance ranking in a web search engine. Retrieval-augmented tools usually use semantic search. They turn the question and the documents into embeddings, which are lists of numbers representing meaning, and return the closest matches. That can find documents that use different words from your question. It is still relevance ranking, though, so the tool returns a top set and its recall, the share of all relevant records it finds, is low and unknown.

**Sarah:** Let's get to the part everyone talks about, which is fabricated citations. Why do they happen?

**Kiffer:** References are some of the most patterned text that exists: surnames and initials, a year, a title, a journal, a volume, pages and a digital object identifier, all in a predictable order. A model trained on millions of reference lists learns that shape very well. Ask it for sources without a retrieval step, and it produces text in that shape from patterns linked to the topic.

**Sarah:** So the parts can be real even when the whole is invented.

**Kiffer:** Exactly. The authors may be real researchers and the journal may be real. The title reads like a typical title. The combination simply matches nothing that was ever published.

**Sarah:** Is there evidence on how often this happens?

**Kiffer:** Walters and Wilder asked ChatGPT in twenty twenty-three to write short literature reviews on forty-two topics and checked all six hundred and thirty-six references. GPT stands for generative pre-trained transformer. With the version called GPT three point five, fifty-five percent of the references were fabricated, and with GPT four, eighteen percent. Chelli and colleagues found that chatbots given the criteria of published systematic reviews returned only a small fraction of the included studies.

**Sarah:** Those numbers come from older models, though.

**Kiffer:** They do, and I want to be careful about that. Newer tools may do better. What the studies show is that fabrication fell with a newer model without disappearing, and I know of no published evaluation showing that any tool has eliminated it. So every reference still needs checking.

**Sarah:** The lesson describes several types of citation error.

**Kiffer:** There is the fabricated reference, to a work that does not exist. There is the conflated reference, which merges parts of two real works, such as one article's title with another article's authors and year. There are real works with wrong details, real works attached to claims they don't make, and real works that don't fit the review or have been retracted.

**Sarah:** What makes them so hard to catch?

**Kiffer:** They look exactly like correct references, and they arrive mixed in with correct ones. If you recognise four studies on a list, you tend to trust the other four. The only reliable check is to look each source up in an independent system and read what it says.

**Sarah:** Section two sorts tools into categories instead of ranking products. Why?

**Kiffer:** Because products change faster than a degree program. If you understand the category, you can predict a new tool's strengths and weaknesses. I name products only as examples available in twenty twenty-six, and naming them is no endorsement. Two questions sort the tools. Where does the answer come from, whether training data, a search of an index or the team's own records? And does the tool produce written text, or a ranking or classification of records?

**Sarah:** Give us the four categories.

**Kiffer:** First, conversational assistants, such as ChatGPT, Claude, Gemini and Microsoft Copilot, many of which can now run a web search. Second, artificial intelligence search engines and research assistants, such as Perplexity for the web, and Elicit, Consensus and Scopus AI for scholarly literature. Third, citation-context tools, such as scite and Semantic Scholar, which look at the sentences in which later papers cite a study. Fourth, screening tools, such as ASReview, and the machine-learning features in Rayyan, EPPI-Reviewer and DistillerSR.

**Sarah:** What is each one good for?

**Kiffer:** Assistants are good with words. They can suggest synonyms for a concept table, which the librarian then tests, explain a method, or help edit text the team wrote. Search engines can find a handful of relevant papers quickly, supply seed studies for citation chasing and check the database search. Citation-context tools show how a study was received, and screening tools order records so that relevant ones appear early.

**Sarah:** And what can't they do?

**Kiffer:** An assistant without search can't be trusted to name sources. A search engine returns a ranked top set chosen by criteria you usually can't see, so its recall is low and its results can't be repeated. A citation-context label describes a citing sentence and tells you nothing about the quality of the study being cited. A screening tool only predicts, and a human still decides on every record.

**Sarah:** Here is the question students will ask. If a research assistant finds twenty great papers in ten minutes, why spend weeks on a database search?

**Kiffer:** Because a review promises its readers a documented, reproducible attempt to find all of the eligible evidence. The database search keeps that promise, because it returns every matching record, the strategy can be published, and another team can run it again. An AI search engine was designed to give a busy person a good answer quickly, so it ranks, selects and summarises, and each step costs recall and reproducibility. These tools can add to the search and check it, and the database search stays the backbone.

**Sarah:** And if an AI tool does find something useful?

**Kiffer:** It goes through the usual screening. In the flow diagram from the twenty twenty Preferred Reporting Items for Systematic reviews and Meta-Analyses, usually called PRISMA, it enters as a record identified by other methods, the same route as websites and citation searching.

**Sarah:** The Cedar Valley team tested three tools. What happened?

**Kiffer:** The intern's test is the verification exercise in section three, so I'll save that. The librarian entered the review's question, framed by population, concept and context, into a scholarly research assistant and exported the twenty sources it returned. Sixteen were already among the database records. The other four were outside the review's scope: two studied adults under sixty-five, one was a commentary and one evaluated a hospital discharge program.

**Sarah:** So the tool found nothing new.

**Kiffer:** Nothing new and eligible. The team's conclusion was measured. The result offers some reassurance about the database search, but a top set of twenty cannot show what the tool missed. In the third test, the evidence officer used a citation-context tool to see how later papers cited a social prescribing review by Bickerdike and colleagues. It helped her understand the debate, and when she read the citing sentences herself, two of the tool's labels did not match them.

**Sarah:** Section two ends with questions for evaluating a tool you've never seen.

**Kiffer:** There are six, and the tool's own documentation should answer them. What does it search, and what does its index cover? How does it decide what to show? Has it been independently evaluated for a task like yours? Can you save the prompt, version, date and full output? What happens to information you enter? And who provides it, and what are their interests? If the documentation can't answer one of them, that gap is useful information too.

**Sarah:** Let's move to section three, on verification. The lesson opens with a principle.

**Kiffer:** The principle is that the authors of a review are responsible for its content, whatever tools they used. In twenty twenty-five, Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence issued a joint position statement on artificial intelligence in evidence synthesis that says this directly, and adds that these tools should be used with human oversight. The Committee on Publication Ethics said in twenty twenty-three that these tools can't be authors, because they can't take responsibility for the work.

**Sarah:** And responsibility, in practice, means checking.

**Kiffer:** It means a verification log with four checks for every suggested source. First, existence: search the exact title in a database, Google Scholar and the library's discovery tool, and put any digital object identifier into the resolver at doi dot org or look it up in Crossref, the agency that registers most of them. Second, details: compare the authors, year, journal, volume and pages with the publisher's record. Third, the claim: read the abstract and the relevant part of the full text. Fourth, fit: check the eligibility criteria and check for a retraction, which Crossref now makes easier through the Retraction Watch database.

**Sarah:** Now the exercise. The intern asked an assistant for eight sources.

**Kiffer:** And I want to be clear that the output in the lesson was invented for teaching. It imitates what an assistant without web search can produce, it mixes real publications with fabricated and mis-attributed entries, and the answers say which are which.

**Sarah:** Walk us through what the team found.

**Kiffer:** Three entries were correct as given. They were a systematic review by Dickens and colleagues from twenty eleven, an integrative review by Gardiner and colleagues, and a consensus report from the National Academies of Sciences, Engineering, and Medicine, which the team recorded as a report, a form of grey literature, since the prompt had asked for peer-reviewed sources.

**Sarah:** And the problems?

**Kiffer:** One entry was a conflated reference. It gave the title of a twenty fifteen article by Holt-Lunstad and colleagues with the authors, year and journal of a different twenty ten article by the same lead author. A review by Cattan and colleagues existed but had the wrong journal, year and volume. And the Bickerdike review of social prescribing had correct details and a summary claiming consistent evidence that link workers reduce loneliness in older adults.

**Sarah:** Which it didn't find.

**Kiffer:** It didn't. The review rated all fifteen of its included evaluations at high risk of bias and concluded that the evidence was insufficient to judge success or value for money. It also had no age restriction. The tool turned a cautious conclusion into a confident one.

**Sarah:** And the last two entries?

**Kiffer:** They couldn't be found anywhere. One claimed to be a cluster randomized trial of community connectors for rural Canadian seniors, in a real journal, the Canadian Journal on Aging, with a digital object identifier that returns a not-found message. The other claimed to be a meta-analysis of link worker programs for adults over sixty-five.

**Sarah:** That second one sounds like exactly what the planning team wants to hear.

**Kiffer:** That is why the lesson calls it the most tempting entry. A fabricated source that fits the reader's question closely is the one most likely to survive an unchecked reading. The team removed both and kept them in the log so that the removal is documented.

**Sarah:** So what is the tally?

**Kiffer:** Six of the eight sources exist, only three were correct in every detail and claim, and none was an eligible primary study. The tool did name several well-cited reviews that are useful for citation chasing. A student who pasted the list straight in, though, would have cited two articles that do not exist, attributed a confident finding to a review that judged the evidence insufficient, and given wrong details for two more.

**Sarah:** Section three also covers reproducibility. Can an AI search be repeated?

**Kiffer:** Only approximately. The model samples its output, providers update and retire models, the index changes, and small changes in wording change the results. So the goal is a record full enough that a reader can see what was done. That is the AI-assisted search log, which records the tool and provider, the version, the date and the person, the settings, the exact prompt, the full output saved as a file, and how the output was verified and screened.

**Sarah:** Then privacy.

**Kiffer:** Whatever you enter into a tool leaves your control to some degree. Inputs may be stored, reviewed or used to train future models, and may be held in another country. Public bodies in British Columbia, including health authorities and universities, fall under the Freedom of Information and Protection of Privacy Act and have policies on approved tools. For Cedar Valley, the survey responses from eleven programs and the notes from seven key informant interviews in the environmental scan stay out of consumer tools.

**Sarah:** What about confidential documents and copyright?

**Kiffer:** Unpublished manuscripts, grant applications and internal drafts belong to their authors or organizations, and many journals and funders forbid reviewers from entering them into these tools. The United States National Institutes of Health, for example, prohibits its peer reviewers from using them to analyze or critique grant applications. On copyright, library licences may restrict uploading full texts, so you ask the librarian first.

**Sarah:** And finally, disclosure.

**Kiffer:** A disclosure statement tells readers which tools you used, for what, at which stages, and how you checked them. The emerging guidance is the RAISE recommendations, short for responsible use of artificial intelligence in evidence synthesis, which were still being revised in twenty twenty-five. The joint position statement that endorses them asks authors to report fully any use that makes or suggests judgements, such as eligibility, risk of bias, extraction or synthesis, giving each tool's name, version and dates, its purpose and stages, the justification, any related interests, and its limitations.

**Sarah:** Do I have to disclose a grammar checker?

**Kiffer:** Editing limited to spelling and grammar generally doesn't need to be reported, though journal policies vary. The lesson gives a template, a completed Cedar Valley example and a short version for course work.

**Sarah:** Let's go to section four, on active-learning screening. Why is screening the stage with the most evidence?

**Kiffer:** Because it is usually the most time-consuming stage of a review. In Cedar Valley, two people independently screening one thousand eight hundred and seventy records means three thousand seven hundred and forty decisions before anyone reads a full text.

**Sarah:** The lesson distinguishes three ways a tool can help.

**Kiffer:** In prioritized screening, the model sets the order and every record is still screened. As a second screener, the model takes the place of one of two human reviewers. In truncation, the team stops before every record has been read. O'Mara-Eves and colleagues reviewed the evidence on text mining in twenty fifteen and concluded that prioritization was safe and ready for live reviews, that a second screener could be used with caution, and that automatic elimination was promising and not yet fully proven.

**Sarah:** How does active learning actually work?

**Kiffer:** The reviewer starts by labelling a few records as relevant or irrelevant. The tool turns each title and abstract into numerical features, and a classifier, such as naive Bayes or logistic regression, learns which features predict relevance. It scores every unscreened record and shows you the one it ranks highest. You decide, it retrains, and the cycle repeats.

**Sarah:** So the human still makes every decision.

**Kiffer:** Every one. The model changes only the order. ASReview, an open-source tool from Utrecht University described by van de Schoot and colleagues, implements this cycle, and its simulation mode replays a completed review to show how early the relevant records would have appeared.

**Sarah:** And the lesson uses that simulation mode with Cedar Valley.

**Kiffer:** As a look ahead. Lesson seven describes how the team screens all the records in duplicate. One hundred and forty-five pass title and abstract screening, and one hundred and forty-two of their full texts are assessed. This lesson imagines the intern replaying those decisions, counting the one hundred and forty-two assessed records as the relevant ones, with illustrative figures built for teaching. Eighty-eight of the one hundred and forty-two relevant records turn up in the first ten percent of records. Ninety-five percent recall is reached after five hundred and forty records, under a third of the total. The last relevant record doesn't appear until record one thousand three hundred and ninety-five.

**Sarah:** And there's a measure called work saved over sampling.

**Kiffer:** Cohen and colleagues introduced it in twenty oh six. It compares the screening needed to reach a target recall with what random order would need. In random order you would screen about ninety-five percent of the records to find ninety-five percent of the relevant ones. Here the tool got there after about twenty-nine percent, so the work saved at ninety-five percent recall is about sixty-six percent.

**Sarah:** That sounds like a big saving. What's the catch?

**Kiffer:** Those numbers are calculated after the fact, when you already know there are one hundred and forty-two relevant records. In a live review nobody knows the total, so nobody knows the recall. If you screen every record in the tool's order, you lose nothing. The time saving comes only from stopping early, and the decision to stop is where the risk lies.

**Sarah:** How do teams decide when to stop?

**Kiffer:** With a stopping rule set in advance. A common heuristic stops after a run of consecutive irrelevant records, say one hundred. In the simulation, that rule would have stopped at record seven hundred and ten, saving about six in every ten records, with one hundred and thirty-eight of the one hundred and forty-two found.

**Sarah:** So it misses four.

**Kiffer:** It misses four. The lesson asks you to suppose that one was a qualitative evaluation of a men's shed program whose abstract never mentioned loneliness or isolation, and that it became an included study. Records that are hard to find tend to use unexpected language or have weak abstracts, and a scoping review of a varied public health literature needs exactly those records.

**Sarah:** Are there better stopping rules?

**Kiffer:** There are more careful ones. Boetje and van de Schoot proposed the SAFE procedure in twenty twenty-four, a deliberately conservative combination of heuristics. Callaghan and Müller-Hansen proposed a statistical approach in twenty twenty that screens a random sample of the remaining records and tests whether a recall target has been reached with stated confidence. Callaghan and colleagues argued in twenty twenty-four that stopping criteria still need to be well evaluated and transparently reported.

**Sarah:** What does the evidence on performance say overall?

**Kiffer:** Most of the studies that O'Mara-Eves and colleagues reviewed suggested savings of thirty to seventy percent, sometimes at ninety-five percent recall. The evidence comes mostly from simulations on completed reviews, and performance varies with the topic, the share of relevant records, the quality of the abstracts and the size of the dataset.

**Sarah:** Simulations treat the human decisions as the truth. Are human screeners that reliable?

**Kiffer:** Not entirely. Gartlehner and colleagues found in a randomized trial that single-reviewer abstract screening missed thirteen percent of relevant studies. The joint position statement uses that finding to illustrate a trade-off: in a rapid review that would otherwise use one human screener, a tool acting as a second reviewer might catch studies the human would miss.

**Sarah:** And large language models doing the screening directly?

**Kiffer:** That evidence is newer and less settled. Khraisha and colleagues tested GPT four on screening and extraction across several types of literature and languages. Its accuracy matched humans on some tasks, and performance fell across all stages once they adjusted for chance agreement and for the imbalance between included and excluded records. They concluded that substantial caution was warranted.

**Sarah:** The lesson ends with a decision table. What is it for?

**Kiffer:** It brings the lesson together. For each stage of a review, it sorts uses of artificial intelligence into three columns: generally acceptable with human checking, acceptable only with added safeguards, and not acceptable. It applies three principles, which are human oversight of every judgement, verification of every output, and disclosure of any use that makes or suggests judgements. It is a teaching tool for this course, and stricter rules from journals, funders or institutions take precedence.

**Sarah:** Give us a few rows.

**Kiffer:** At search development, suggesting synonyms is generally acceptable, translating a search between databases needs the librarian to check every line, and running an unreviewed search string written by a tool is not acceptable. At title and abstract screening, ordering records while every record is screened is generally acceptable, a model as second screener or an early stop under a pre-specified rule needs safeguards, and letting a tool exclude records with no human check is not acceptable.

**Sarah:** What did the Cedar Valley team decide in the end?

**Kiffer:** They will use an assistant for synonyms, tested by the librarian, and they ran the two supplementary searches with every source verified. They will screen all the records in duplicate and use active learning only in simulation, to plan future updates. No tool will make eligibility, charting or synthesis decisions, and nothing from the scan goes into any tool. The final report will include a disclosure statement, with the logs in an appendix.

**Sarah:** Let's finish with the practical side. Where should a team start if it wants to try these tools?

**Kiffer:** Start small. Using a tool that your institution's policy allows, run one assisted search on the review question and record it in a search log. Verify every source the tool suggests with the four steps, recording each check in a verification log. Mark a copy of the decision table to show where you will use these tools in the rest of the review and the safeguard at each stage. Then write a short disclosure statement using the template, and enter no personal or confidential information into any tool.

**Sarah:** And next week?

**Kiffer:** Lesson seven, on managing records and screening studies. The Cedar Valley team imports its records into a reference manager, removes the duplicates, and screens titles and abstracts in a screening platform with two independent reviewers.

**Sarah:** That's it for this week. Thanks, Kiffer.

**Kiffer:** Thanks, Sarah. See you next week.
