Artificial Intelligence in Evidence Retrieval and Synthesis
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Describe how a large language model produces text by predicting one token after another, and explain why this process can generate fabricated or mis-attributed citations.
- Explain how retrieval-augmented generation works and why it reduces some citation errors while leaving others in place.
- Distinguish conversational assistants, AI search engines and research assistants, citation-context tools and AI-assisted screening tools by where their answers come from and by what each can and cannot do in a review.
- Evaluate an unfamiliar AI research tool by asking about its index, its ranking method, its evaluation, its documentation and its handling of data.
- Verify AI-suggested citations with a four-step procedure covering existence, bibliographic details, claim and fit, and record the results in a verification log.
- Document an AI-assisted search, identify privacy, confidentiality and copyright risks, and write a disclosure statement that follows emerging guidance such as the RAISE recommendations.
- Explain how active-learning screening works and interpret recall, work saved over sampling and stopping rules from a screening simulation.
- Judge when AI assistance is acceptable at each stage of a review, and prepare an AI-assisted search log, a verification log and a disclosure statement for a review.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
How Large Language Models and Retrieval-Augmented Tools Generate Answers
Learning Objectives for this section
- Describe in plain terms how a large language model produces text by predicting one token after another from patterns learned in training.
- Explain how retrieval-augmented generation adds a search step to a language model and why that step reduces some errors while leaving others in place.
- Explain why fabricated and mis-attributed citations occur, and distinguish the main types of citation error.
- Relate the behaviour of AI tools to the retrieval ideas from Lesson 3, including relevance ranking, recall and reproducibility.
1.1 Why a Review Team Needs to Understand AI Tools
By the middle of a review, a team has a question, a protocol, a database search and a plan for grey literature, and it has probably been told that an artificial intelligence (AI) tool could do much of the remaining work. Some of these tools help with parts of a review, and some of their outputs are wrong in ways that are hard to see. A reviewer who understands how the tools produce their answers can decide where they help and how to check them.
The running example in this course is the Cedar Valley evidence review, a fictional project. Before launching a community connector (social prescribing) program for older adults, the planning team of the fictional Cedar Valley Health Authority in British Columbia asked its small evidence team, made up of an evidence officer, a university librarian and a student intern, for a rapid scoping review and environmental scan within twelve weeks. The review asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged 65 and older, and with what outcomes. By this lesson the team has searched MEDLINE, Embase, CINAHL, PsycINFO and Web of Science (Lesson 4), retrieving 2,480 records, of which 1,870 remain after 610 duplicates are removed, and it has planned its grey-literature and web searching (Lesson 5).
The intern has used a conversational AI assistant for coursework and suggests that the team ask it for the key studies on loneliness interventions for older adults. The librarian agrees to a test on two conditions. Every reference the tool produces will be checked against a primary source and recorded in a verification log, and the test will be documented in enough detail that a reader of the final report can see what was done. The evidence officer adds a third condition: nothing from the environmental scan, such as interview notes, will be entered into any tool. Section 3 works through the intern's results, and this section explains why the librarian's conditions are needed.
Artificial intelligence is a broad label for computer systems that perform tasks usually associated with human judgement. Machine learning is the part of AI in which a system learns patterns from examples, as the screening tools in Section 4 do. A large language model is a generative AI system that produces text. HSCI 841 Lesson 12 complements this lesson by examining the use of large language models to code qualitative data.
1.2 How a Large Language Model Produces Text
Tokens and next-token prediction
A large language model works with tokens, which are short pieces of text such as a whole word, part of a word or a punctuation mark. A long or unusual word, such as a surname or a technical term, is usually split into several tokens. In training, the model reads a very large collection of text and learns to predict the next token in a passage from the tokens before it, storing what it learns as billions of numerical settings called parameters. Most current models use the transformer architecture introduced by Vaswani and colleagues (2017), which lets the model weigh every earlier token when it predicts the next one.
Developers then train the model further on examples of instructions and on human ratings of its answers, so that it follows requests and holds a conversation. When it answers, it repeats one step many times: it calculates a probability for every possible next token, selects one, adds it to the text and calculates again. A setting often called temperature controls how much randomness enters the selection, which is one reason the same prompt can produce different answers on different occasions.
What next-token prediction means for a reviewer
Four consequences follow for evidence work. First, a language model on its own has no catalogue of sources. What it knows about the literature is spread across its parameters as patterns, so when it writes a reference it produces text that resembles references it has seen. Second, the model's information stops at its knowledge cutoff, so recent studies are missing unless the tool can search. Third, because selection involves chance, a repeated prompt can give a different list of studies, which makes the output difficult to reproduce. Fourth, fluency and confidence in the text carry no information about accuracy, because the model produces fluent text whether or not the content is correct.
Bender and colleagues (2021) described large language models as systems that assemble plausible sequences of language without grounding in meaning, and argued that this creates risks when people treat the output as reliable. Reviewers need only recognize that a model's output is shaped by its training text and by the wording of the prompt. A prompt that asks for "studies showing that befriending reduces loneliness" invites a list of supportive studies, whereas a neutral prompt that asks what has been evaluated and with what results gives the model less to agree with.
1.3 Retrieval-Augmented Generation
Many AI research tools add a search step to the language model. Lewis and colleagues (2020) introduced the term retrieval-augmented generation for systems that combine a retriever, which finds relevant passages in a collection of documents, with a generator, which writes an answer using those passages. In a typical tool, the user's question is turned into a search, the retriever returns the passages it ranks highest from its index (the web, or a database of scholarly abstracts), those passages are placed in the model's context window, and the model writes an answer with links to the passages it drew on. Many conversational assistants available in 2026 can also run a web search during a conversation, which turns them into retrieval-augmented systems for that response.
What retrieval changes and what it leaves in place
Because each citation points to a document the retriever actually returned, a retrieval-augmented tool is much less likely to invent a reference outright. Other errors remain. The model can attach a claim to a source that does not support it, for example by turning a cautious conclusion into a confident one or by combining findings from two passages into a statement neither makes. The retriever can return weak, outdated or off-topic documents, and the model will summarize them with the same fluency as strong ones.
That last limit connects directly to Lesson 3. A bibliographic database running a Boolean search returns every record that matches the search, which is why reviews depend on it for high recall, the proportion of all relevant records that a search finds. A retrieval-augmented tool usually converts the question into an embedding, a list of numbers that represents its meaning, and returns the documents whose embeddings are closest. This approach, often called semantic search, can find relevant documents that use different words from the question, which is useful. It is still relevance ranking: the tool returns a top set of perhaps ten to a few dozen sources and has no notion of retrieving every eligible study. Its index also limits what it can find: many scholarly AI tools index abstracts from open sources and hold little of the content of Embase or CINAHL and little Canadian grey literature.
1.4 Why Fabricated and Mis-Attributed Citations Occur
References are among the most patterned text that exists. A reference list follows a predictable shape: surnames and initials, a year, a title with a colon, a journal name, a volume, an issue, a page range and a digital object identifier (DOI). A model trained on millions of reference lists learns that shape very well. Asked for sources without a retrieval step, it produces text in that shape from patterns associated with the topic. The surnames may belong to real researchers, the journal may be real and the title may read like a typical title, while the combination corresponds to no published work. Ji and colleagues (2023), in a survey of the problem across text-generation systems, describe this kind of output as hallucination: content that is fluent but unsupported by the source or by the facts.
Studies that tested chatbots in 2023 documented the problem. Walters and Wilder (2023) asked ChatGPT to write short literature reviews on 42 topics and checked the 636 references in the 84 reviews it produced. They found that 55 percent of the references produced by GPT-3.5 and 18 percent of those produced by GPT-4 were fabricated, and that 43 percent and 24 percent of the real references, respectively, contained substantive errors. Chelli and colleagues (2024) gave ChatGPT and Bard (later renamed Gemini) the inclusion criteria of eleven published systematic reviews on shoulder rotator cuff conditions and compared the references the tools returned with the references in the reviews. They classified 39.6 percent of the GPT-3.5 references, 28.6 percent of the GPT-4 references and 91.4 percent of the Bard references as hallucinated, and the GPT models found fewer than one in seven of the studies the human reviews had included. These figures describe the model versions tested at that time, and newer tools may perform differently. They show that fabrication fell with a newer model without disappearing.
| Type of citation error | What it looks like | How a reviewer detects it |
|---|---|---|
| Fabricated reference | A complete, plausible reference to an article, report or book that does not exist. The DOI, if given, fails to resolve or leads to a different work. | No match for the exact title in a bibliographic database, Google Scholar or Crossref, and no match for the DOI. |
| Conflated reference | Parts of two or more real works merged into one reference, such as the authors and year of one article with the title of another. | The title search finds a real article, but its authors, year or journal differ from those given. |
| Real work, wrong details | A real article with an incorrect year, journal, volume, page range or DOI. | The publisher's record or the database record shows different bibliographic details. |
| Real work, unsupported claim | A correct reference attached to a statement that the source does not make, or that overstates what it found. | Reading the abstract and the relevant part of the full text shows a different finding or conclusion. |
| Real work, wrong fit | An accurate reference and summary for a source that does not meet the review's eligibility criteria, or that has been retracted. | Comparing the source with the review's question and criteria, and checking for a retraction notice. |
Two features of these errors make them hard to see. They look exactly like correct references, so a reader cannot tell them apart by inspection, and they are often mixed in with real references, so a list that contains several familiar studies earns trust it has not earned. The only reliable check is to look each source up in an independent system and to read what it says, which is the procedure Section 3 sets out.
Each statement describes what a reviewer found when checking a reference that an AI tool suggested. Name the type of error from the table above before opening the answers below. (a) The title matches a real 2015 article, but the reference gives the authors and year of a different 2010 article by the same lead author. (b) The article exists and the details are correct, but the tool's summary says it found a large reduction in loneliness, and the abstract reports that most included evaluations were at high risk of bias and that effectiveness could not be judged. (c) No database, search engine or DOI registry has any record of the title, and the DOI returns an error.
Statement (a) describes a conflated reference, because elements of two real works have been merged. Statement (b) describes a real work with an unsupported claim, because the reference is correct but the summary misstates the source. Statement (c) describes a fabricated reference, because nothing matching it can be found anywhere and the identifier does not resolve. Section 3 meets all three of these errors in the Cedar Valley verification exercise.
A DOI is a string of characters with a predictable format, beginning with 10 and a publisher prefix, and a language model can produce one as easily as it produces a title. A DOI shows that a reference is real only when it resolves at doi.org to the same work that the reference describes.
Walters and Wilder (2023) found far fewer fabricated references from GPT-4 than from GPT-3.5, which shows that newer models can improve. Their GPT-4 results still included fabricated references and errors in real ones, and no published evaluation that this course is aware of shows that any tool has eliminated them. Every reference therefore still needs checking, whichever tool produced it.
A retrieval-augmented tool links its claims to retrieved documents, which reduces fabrication. The claims can still misstate the documents, and the documents can still be weak or irrelevant. Reviewers check the claim against the source in the same way for these tools as for any other.
1.5 From Mechanism to Practice
The mechanism explains the rules the rest of the lesson applies. A generated reference is a lead to be checked. Retrieval-augmented tools rank and summarize a small top set, so they can supplement a documented database search without replacing it (Section 2). Outputs vary with the prompt, the model version and the date, so every use has to be recorded (Section 3). Machine learning can rank records well without making inclusion decisions, which is why ordering records for screening is its most established use in reviews (Section 4).
The Cedar Valley team's working rule
After discussing how these tools work, the Cedar Valley team adopts a working rule for the rest of the review. Any source suggested by an AI tool is treated as a lead. It enters the review only after a team member has found it in an independent system, confirmed its details, read enough of it to confirm the claim, and recorded these checks in the verification log.
Reflection
A colleague working on a review of school-based physical activity programs asks a conversational AI assistant, with its web search turned off, for ten key studies. The assistant returns ten complete references in APA style, each with a digital object identifier (DOI) and a one-sentence summary, and several of the authors are well-known researchers in the field. The colleague wants to paste the list straight into the reference list. A second colleague suggests switching to a retrieval-augmented tool, which first searches an index of documents and then writes an answer that cites the passages it retrieved, and says that this will make every claim correct. Write a reply to both colleagues that (a) explains how a large language model without a search step produces a reference and why a DOI and familiar author names do not show that a reference is real, (b) explains what a retrieval-augmented tool changes, (c) names two kinds of error that can remain when retrieval is used, and (d) ends with one practical rule for the review.
A language model without search has no catalogue of articles to consult. It produces a reference one token at a time, following the patterns of the millions of references in its training text, so the result has the right shape (authors, year, title, journal and DOI) whether or not the work exists. Familiar author names show only that those researchers are associated with the topic in the training text, and a DOI is a predictable string that the model can generate as easily as a title. A DOI shows that a reference is real only when it resolves to the same work at doi.org.
A retrieval-augmented tool first retrieves real documents and then writes an answer citing them, so it is much less likely to invent a source. Two errors remain. The tool can attach a claim to a real source that does not support it, for example by turning a cautious conclusion into a confident one, and it can retrieve weak or off-topic documents and summarize them as fluently as strong ones. Because it returns a ranked top set, it also misses relevant studies.
The practical rule is that every AI-suggested source is a lead: it enters the review only after someone has found it in an independent database, checked its details, read enough of it to confirm the claim and recorded those checks in a verification log.
Minimum 20 characters required.
Question 1: When a large language model without a search tool writes a reference list, what is it doing?
Question 2: Which error is a retrieval-augmented tool still likely to make?
Question 3: A reference gives the title of a 2015 article but the authors, year and journal of a 2010 article by the same lead author. What type of error is this?
Question 4: What did Walters and Wilder (2023) find when they compared references produced by GPT-3.5 and GPT-4?
Categories of AI Research Tools and What Each Can and Cannot Do
Learning Objectives for this section
- Distinguish four categories of AI research tools: conversational assistants, AI search engines and research assistants, citation-context tools, and AI-assisted screening and extraction tools.
- Explain, for each category, where its answers come from and which review tasks it can support.
- Identify the limitations of each category that matter for a systematic or scoping review, including coverage, recall, reproducibility and transparency.
- Evaluate an unfamiliar AI tool with a short set of questions about its index, its method, its evaluation and its handling of data.
2.1 Categories Last Longer Than Products
AI research tools change quickly. Products are launched, renamed, merged, given new features and withdrawn within the span of a single degree program, and a guide that ranked individual products would be out of date before the term ended. This section therefore organizes tools into four categories by how they work. Product names appear only as examples available in 2026, and their presence in the text is no endorsement. A student who understands the categories can place a new tool into one of them, predict its strengths and weaknesses, and ask the right questions about it.
Two questions sort most tools. The first asks where the answer comes from: from patterns the model learned in training, from a search of an index of documents, or from the review team's own records and decisions. The second asks what the tool produces: written text, such as an answer or a summary, or a ranking or classification of records. Figure 6.2 places the four categories on these two dimensions.
2.2 The Four Categories
The tabs below describe each category in turn: how it works, examples available in 2026, the review tasks it can support and the limits that matter in a review.
How they work. A conversational assistant is a general-purpose large language model with a chat interface. Without a search tool, it answers from patterns learned in training, as Section 1 described. Many assistants can now run a web search during a conversation, and several offer extended research modes that run a series of searches and compile a report with links. Examples available in 2026 include ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google) and Microsoft Copilot.
What they can support. Assistants are good at work with words. They can suggest synonyms, spelling variants and related terms for a concept table, which the librarian then tests against the databases (Lesson 4). They can explain an unfamiliar method in plain language, draft a plain-language summary from text the team has written, check a protocol section for clarity, or suggest how a search might be translated from one database's syntax to another, which a librarian must then check line by line.
What they cannot do. An assistant without search cannot be trusted to name sources, because it generates references from patterns. An assistant with search returns a small, ranked set of web pages chosen by criteria the user cannot see. Neither produces a search that another team could repeat, and neither has access to subscription databases such as Embase unless an institution has connected them.
How they work. AI search engines and research assistants are retrieval-augmented systems built for finding information. General tools search the web, and scholarly tools search large collections of article records and abstracts. They typically return a written answer with citations, a list of sources, or a table summarizing each source. Examples available in 2026 include Perplexity, which searches the web, and Elicit, Consensus and Scopus AI, which search scholarly literature; Scopus AI draws on the Scopus database.
What they can support. These tools can find a handful of relevant papers quickly, which helps when a team is scoping a topic or deciding whether a review already exists. They can supply seed studies for citation chasing (Lesson 5), and they can serve as a check on a database search: if a tool finds an eligible study that the database search missed, the librarian can ask why and revise the search.
What they cannot do. They return a ranked top set, so their recall is low by design and unknown in any particular case. Their index may hold only abstracts, may favour open-access and English-language work, and may include little grey literature. The ranking method is usually undisclosed, results change as the index and model change, and the written summaries can misstate the sources they cite.
How they work. Citation-context tools analyze the sentences in which later papers cite an earlier one. Some use machine learning to classify each citing statement, for example as supporting, contrasting or simply mentioning the cited work. Examples available in 2026 include scite, whose Smart Citations classify citing statements in this way, and Semantic Scholar, which shows the sentences in which a paper is cited and flags citations it judges influential. The citation-network tools introduced in Lesson 5, such as Connected Papers, ResearchRabbit and Litmaps, recommend related papers from patterns of citation and sit close to this category.
What they can support. They show quickly how a study has been received, whether later work questioned it, and which later papers built on it. This helps a reviewer interpret an influential study and can surface related work for citation chasing.
What they cannot do. A classification describes what a citing sentence says about a paper. It says nothing about the quality of the cited study, and an automated classifier can misread a sentence. Coverage depends on the full texts the tool can access, so citation counts and classifications are incomplete.
How they work. Screening tools use machine learning on the review team's own records. In active learning, a model learns from the reviewer's decisions on a few records and then shows the records it predicts are most likely to be relevant, retraining as each decision is made. Classifiers trained on large labelled collections can identify a particular kind of record, such as reports of randomized trials. Large language models are also being tested for screening and for extracting data from full texts. Examples available in 2026 include ASReview, an open-source active-learning tool, and the machine-learning features of screening platforms such as Rayyan, EPPI-Reviewer and DistillerSR.
What they can support. Ordering records so that relevant ones appear early is the most established use of machine learning in reviews. It lets a team begin full-text work sooner and supports decisions about when screening can safely end, which Section 4 examines.
What they cannot do. The model only predicts. A human reviewer still decides on every record the model shows, and a decision to stop screening before every record has been seen carries a risk of missing studies that the team must estimate and report. Extraction by a large language model produces values that each need checking against the full text.
The categories side by side
| Category | Where answers come from | Review tasks it can support | Main limits in a review |
|---|---|---|---|
| Conversational assistants | Patterns learned in training, plus web search when that feature is on | Synonyms for a concept table, plain-language explanations, drafting and editing the team's own text | Fabricated references without search, no reproducible search, no access to subscription databases |
| AI search engines and research assistants | A ranked search of a web or scholarly index, summarized by a language model | Scoping a topic, finding seed studies, checking a database search for missed studies | Low and unknown recall, partial coverage, undisclosed ranking, summaries that need checking |
| Citation-context tools | Citing sentences in later papers, classified by machine learning | Interpreting how a study was received and finding related work | Incomplete coverage, misclassified sentences, no judgement of study quality |
| AI-assisted screening and extraction | The team's own records and screening decisions | Ordering records for screening, flagging study designs, drafting extraction for checking | Uncertain stopping point, performance that varies by topic, extraction errors |
2.3 Why an AI Search Cannot Replace the Database Search
The difference between these tools and the search built in Lesson 4 is a difference of purpose. A systematic or scoping review promises its readers that the team has made a documented, reproducible attempt to find all of the eligible evidence. The database search keeps that promise because it is exhaustive within each database, because its full strategy can be published, and because another team can run it again and obtain essentially the same records from the same databases on the same date. An AI search engine is designed to give a busy user a good answer quickly, so it ranks, selects and summarizes. Each of those steps reduces recall and reproducibility. For this reason the AI tools in this lesson can add to a review's search and can check it, and the database search remains the backbone of the review.
The same reasoning explains where AI-found studies belong in reporting. PRISMA 2020 separates records identified from databases and registers from records identified by other methods, such as websites, organizations and citation searching. A study that an AI search engine finds, and that the team then confirms and screens, enters the review through the other-methods route, and the search that found it is documented in the same way as a web search (Lesson 5 and Section 3).
The team runs three short tests during one week, recording each in its search log. First, the intern asks a conversational assistant, with web search turned off, for key studies on community-based interventions for loneliness in older adults; Section 3 verifies the eight references it produces. Second, the librarian enters the review's PCC question (population, concept and context) into a scholarly AI research assistant and exports the 20 sources it returns. After checking them against the team's reference library, the librarian finds that 16 were already among the 1,870 database records and that the other 4 are outside the review's scope: two studied adults younger than 65, one was a commentary, and one evaluated a hospital discharge program. Third, the evidence officer uses a citation-context tool to see how later papers cited Bickerdike and colleagues' (2017) systematic review of social prescribing, and reads several of the citing sentences to confirm the tool's labels.
The team draws a measured conclusion. The AI research assistant found no eligible study that the database search had missed, which offers some reassurance about the search. It cannot show what the tool itself missed, because a ranked top set of 20 says nothing about the hundreds of other relevant records in the index. The citation-context tool helped the team understand the debate about the evidence for social prescribing, and two of the labels it assigned did not match the citing sentences when the evidence officer read them.
2.4 Evaluating a Tool You Have Not Seen Before
A new tool will appear during this term, and another will appear next year. The questions below turn the categories into an evaluation that a student can complete in a few minutes using the tool's own documentation. They echo the expectations that the 2025 joint position statement on AI in evidence synthesis, discussed in Section 3, sets for tool developers, which include publishing how a tool works, making evaluations of it available and stating its limitations and potential biases.
A tool can only return what is in its index. The documentation should state the source of its records (for example the open web, a named scholarly database or a set of open abstracts), the date range, and whether it holds full texts or abstracts only. A tool that indexes open abstracts in English will underrepresent some journals, languages and grey literature.
Relevance ranking, semantic search, filters for study design, and limits on the number of results all shape the output. If the documentation does not explain the ranking, the team should treat the results as an unrepresentative sample of the relevant literature.
An evaluation published by independent methodologists for a task like the team's, such as identifying intervention studies in health services research, carries more weight than a developer's general claims. A team that finds no evaluation can pilot the tool on a small set of studies it already knows are eligible.
The team needs to save the prompt, settings, date, version and full output, and to export the sources in a format a reference manager can import. A tool that allows none of this cannot be used in a way that meets the reporting expectations in Section 3.
The terms of use should state whether inputs are stored, whether they are used to train future models, and where the data are held. These terms decide whether anything beyond published material can be entered, a question Section 3 takes up for privacy and copyright.
Funding, ownership and commercial relationships with publishers can affect what a tool indexes and promotes. The joint position statement asks review authors to report financial and non-financial interests related to the AI tools they use.
Choose one AI tool that you have used or that your library offers. Using its documentation, place it in one of the four categories, state where its answers come from, and answer the six questions above in one or two sentences each. Note any question that the documentation does not answer, because a missing answer is itself useful information about whether the tool belongs in a review. Keep your notes, since an AI-assisted search log (Section 3.4) asks for the same details.
A note on product names
The examples named in this section were available in 2026. Products in this field are renamed, merged, given new features and withdrawn often. When a product named here has changed, the category and the questions still apply, and the tool's current documentation is the source to consult.
Reflection
A health authority manager proposes to replace the database search for a scoping review of community transportation programs for older adults with an AI research assistant. The tool returned 25 relevant-looking papers with summaries in ten minutes, whereas the librarian's Boolean search of MEDLINE, Embase, CINAHL, PsycINFO and Web of Science retrieved 3,100 records that will take weeks to screen. The tool's documentation states that it searches a large collection of open scholarly abstracts, returns its 25 highest-ranked results by relevance, and does not describe how it ranks them. Recall is the proportion of all relevant records that a search finds. Write a reply to the manager that (a) explains the difference between exhaustive Boolean retrieval and relevance ranking in terms of recall and reproducibility, (b) identifies two limits of this tool's index or method that matter for this review, and (c) proposes a role for the tool that the review could report transparently.
A Boolean search in a bibliographic database returns every record that matches the search, which is why it can achieve high recall, and because the full strategy can be published, another team can run it again in the same databases and obtain essentially the same records. The AI research assistant ranks documents by predicted relevance and returns its top 25. Its recall for this question is therefore low by design and unknown, because the 25 results say nothing about the relevant records it did not show. Its ranking is undisclosed and its index and model change over time, so the search cannot be repeated.
Two limits of this tool matter here. Its index holds open scholarly abstracts, so it is likely to underrepresent subscription content from Embase and CINAHL and to hold little grey literature, such as program evaluations published by health authorities and transit agencies, which a scoping review of community programs needs. It also gives no information about how it ranks results, so the team cannot judge which kinds of studies it favours.
An appropriate role is as a supplementary search. The librarian can run the tool once, save the prompt, date, version and full output in an AI-assisted search log, check each source against the database results, and screen any new source through the usual process. New studies found this way are reported as records identified by other methods, and a new eligible study is also a signal to revisit the database search.
Minimum 20 characters required.
Question 1: Why does an AI search engine have low and unknown recall for a review question?
Question 2: Which task is best suited to a citation-context tool such as scite?
Question 3: Of 20 sources from an AI research assistant, 16 were already in the Cedar Valley database results and the other 4 were out of scope. What is the most defensible conclusion?
Question 4: If a tool's documentation does not explain how it ranks results, how should a review team treat those results?
Verifying AI Output: Reproducibility, Privacy and Disclosure
Learning Objectives for this section
- Apply a four-step verification procedure (existence, bibliographic details, claim and fit) to AI-suggested citations and record the results in a verification log.
- Explain why AI-assisted searches are difficult to reproduce, and document an AI-assisted search so that a reader can see exactly what was done.
- Identify the privacy, confidentiality and copyright risks of entering material into AI tools.
- Write a disclosure statement for AI use that follows emerging guidance, including the RAISE recommendations and the 2025 joint position statement on AI in evidence synthesis.
3.1 The Reviewer Remains Responsible
Every use of AI in a review rests on one principle: the authors of the review are responsible for its content, whatever tools they used. The 2025 position statement on AI use in evidence synthesis issued jointly by Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (Flemyng et al., 2025) states that evidence synthesists remain ultimately responsible for their work, including the decision to use AI, and that AI should be used with human oversight. The Committee on Publication Ethics (COPE) stated in 2023 that AI tools cannot be authors because they cannot take responsibility for the work. In practice, every source an AI tool suggests is checked against a primary source, and every use is recorded and reported.
3.2 A Four-Step Verification Procedure
A verification log is a table in which a reviewer records, for each AI-suggested source, the checks performed, the result and the action taken. The procedure below fills one row of the log for each source and takes a few minutes per reference.
Step 1, existence. Search for the exact title, in quotation marks, in a bibliographic database such as MEDLINE, in Google Scholar and in the library's discovery tool, and enter any DOI at doi.org or look it up in Crossref, the agency that registers most journal DOIs. A source that none of these systems can find is treated as fabricated. Step 2, details. Compare the authors, year, title, journal, volume, issue, pages and DOI with the publisher's record. Many articles appear online a year or more before they are assigned to an issue, so a difference between the two dates is common. Step 3, claim. Read the abstract and the relevant part of the full text, and decide whether the source says what the AI output says it says. Step 4, fit. Compare the source with the eligibility criteria, and check for a retraction or correction on the publisher's page or in the Retraction Watch database, which Crossref now makes openly available.
3.3 Worked Verification Exercise: The Cedar Valley Citation Check
This exercise uses AI output invented for teaching, imitating what a conversational assistant without web search can produce. It mixes real publications with fabricated and mis-attributed entries, which the answers identify. Do not cite the fabricated entries.
Invented AI output for teaching (this is not output from any real tool)
Prompt entered by the intern: "List eight peer-reviewed sources on community-based interventions to reduce loneliness or social isolation among adults aged 65 and older. Give a full APA reference and a one-sentence summary of the main finding for each."
Invented response:
- Dickens, A. P., Richards, S. H., Greaves, C. J., & Campbell, J. L. (2011). Interventions targeting social isolation in older people: A systematic review. BMC Public Health, 11, 647. https://doi.org/10.1186/1471-2458-11-647. Summary given: group-based and participatory interventions were more often associated with improved outcomes than one-to-one and non-participatory interventions.
- Holt-Lunstad, J., Smith, T. B., & Layton, J. B. (2010). Loneliness and social isolation as risk factors for mortality: A meta-analytic review. PLoS Medicine, 7(7), e1000316. Summary given: loneliness and social isolation are associated with an increased risk of death.
- Thompson, R., Nguyen, L., & Bartlett, K. (2021). Community connector programs and loneliness among rural Canadian seniors: A cluster randomized trial. Canadian Journal on Aging, 40(3), 412–425. https://doi.org/10.1017/S0714980821000187. Summary given: rural seniors referred to community connectors reported lower loneliness at six months than those receiving usual care.
- Gardiner, C., Geldenhuys, G., & Gott, M. (2018). Interventions to reduce social isolation and loneliness among older people: An integrative review. Health & Social Care in the Community, 26(2), 147–157. https://doi.org/10.1111/hsc.12367. Summary given: adaptability, a community development approach and productive engagement were features of the most effective interventions, although the quality of the evidence was generally weak.
- Bickerdike, L., Booth, A., Wilson, P. M., Farley, K., & Wright, K. (2017). Social prescribing: Less rhetoric and more reality. A systematic review of the evidence. BMJ Open, 7(4), e013384. https://doi.org/10.1136/bmjopen-2016-013384. Summary given: the review found consistent evidence that social prescribing through link workers reduces loneliness among older adults in primary care.
- Cattan, M., White, M., Bond, J., & Learmouth, A. (2008). Preventing social isolation and loneliness among older people: A systematic review of health promotion interventions. Journal of Aging and Health, 20(1), 41–67. Summary given: group activities with an educational or support component were the most effective interventions, and the effectiveness of home visiting and befriending remained unclear.
- Okafor, M., Lindqvist, E., & Harland, J. (2023). Social prescribing link workers and loneliness in adults over 65: A systematic review and meta-analysis. Journal of Applied Gerontology, 42(9), 1874–1889. Summary given: link worker programs produced a moderate reduction in loneliness, with larger effects in longer programs.
- National Academies of Sciences, Engineering, and Medicine. (2020). Social isolation and loneliness in older adults: Opportunities for the health care system. The National Academies Press. https://doi.org/10.17226/25663. Summary given: a consensus report on how the health care system can identify and respond to social isolation and loneliness in older adults.
Choose at least four of the eight entries, including entries 2, 5 and 7, and apply the four steps using PubMed, Google Scholar, doi.org or Crossref, and your library's discovery tool. Record your result for each step, then compare your results with the team's results below.
Exists: yes; the title, the DOI and PubMed lead to the same article. Details: correct. Claim: supported. The abstract reports that 79 percent of group-based and 55 percent of one-to-one interventions reported at least one improved outcome, and over 80 percent of participatory compared with 44 percent of non-participatory interventions. These figures come from vote counting across studies at medium to high risk of bias, a method whose limits Lesson 9 explains. Fit: a systematic review, used for citation chasing. Verdict: verified as given.
Exists: the title search finds a real article, but it is a 2015 article by Holt-Lunstad, Smith, Baker, Harris and Stephenson in Perspectives on Psychological Science, 10(2), 227–237. Details: the authors, year, journal, volume and article number given belong to a different real article, Holt-Lunstad, Smith and Layton (2010), "Social relationships and mortality risk: A meta-analytic review," PLoS Medicine, 7(7), e1000316. Claim: both articles concern social relationships, loneliness or isolation as risk factors for mortality, so the summary fits the topic; the team cites whichever article it reads. Fit: neither evaluates an intervention, so both are background. Verdict: conflated reference, corrected.
Exists: no. No record matching the title or authors appears in MEDLINE, Google Scholar or Crossref, and the DOI returns a "DOI not found" message at doi.org. The Canadian Journal on Aging is a real journal, which helps the entry look credible. Verdict: not found, treated as fabricated, removed and recorded in the log. This entry was invented for this exercise.
Exists: yes. Details: correct for the 2018 issue, 26(2), 147–157; the article appeared online in 2016, which is why some databases list that year. Claim: supported. The abstract reports that adaptability, a community development approach and productive engagement were associated with the most effective interventions, and that the evidence was generally weak. Fit: an integrative review, used for citation chasing. Verdict: verified as given.
Exists: yes. Details: correct. Claim: unsupported. The review included 15 evaluations of United Kingdom programs that referred primary care patients, without an age restriction, to a link worker, and rated all 15 at high risk of bias. The authors noted that most evaluations presented positive conclusions despite clear methodological shortcomings, and concluded that the evidence was insufficient to judge either success or value for money. The AI summary turned a cautious conclusion into a confident one and added an older-adult focus. Verdict: real source with an unsupported claim; the team rewrites the summary from the abstract and uses the review for citation chasing.
Exists: yes, under the same title and authors. Details: wrong. The article was published in Ageing & Society, 25(1), 41–67, in 2005, and the journal, year and volume given are incorrect. Claim: supported; the abstract reports that nine of the ten effective interventions were group activities with an educational or support input, and that the effectiveness of home visiting and befriending remained unclear. Verdict: real source with wrong details, corrected and used for citation chasing.
Exists: no. No record matching the title or the authors appears in any of the systems searched, and the entry has no DOI, which is unusual for a 2023 journal article. Verdict: not found, treated as fabricated and removed. This entry was invented for this exercise. It is the most tempting entry because it appears to answer the planning team's question directly, and a fabricated source that fits the reader's question closely is the kind most likely to survive an unverified reading.
Exists: yes; the DOI resolves to the report. Details: correct. Claim: the summary restates the report's stated scope, which the title confirms. Fit: the prompt asked for peer-reviewed sources, and this is a consensus study report, so the team records it as a report, a form of grey literature, and uses it as background. Verdict: verified, with the source type corrected.
The team's verification log
| No. | Source as given | Exists | Details | Claim | Verdict and action |
|---|---|---|---|---|---|
| 1 | Dickens et al., 2011 | Yes | Correct | Supported | Verified; citation chasing |
| 2 | Holt-Lunstad et al., 2010 | Yes | Two articles merged | Fits topic | Conflated; corrected; background |
| 3 | Thompson et al., 2021 | No | Not applicable | Not applicable | Fabricated; removed |
| 4 | Gardiner et al., 2018 | Yes | Correct | Supported | Verified; citation chasing |
| 5 | Bickerdike et al., 2017 | Yes | Correct | Unsupported | Summary rewritten; citation chasing |
| 6 | Cattan et al., 2008 | Yes | Journal, year, volume wrong | Supported | Corrected to 2005; citation chasing |
| 7 | Okafor et al., 2023 | No | Not applicable | Not applicable | Fabricated; removed |
| 8 | National Academies, 2020 | Yes | Correct | Supported | Verified; recorded as report |
Six of the eight sources exist, only three were correct in every detail and claim as given, and none is an eligible primary study for the scoping review. The tool did name several well-cited reviews that are useful for citation chasing. A reviewer who copied the list unchecked would have cited two articles that do not exist, attributed a confident finding to a review that judged the evidence insufficient, and given wrong details for two more.
3.4 Reproducibility and the AI-Assisted Search Log
Lesson 3 showed that web search engines personalize and re-rank results. AI tools add further variation: the model samples its output, providers update and retire models, the index behind a retrieval tool changes, and small changes in a prompt's wording can change the results. Because an AI-assisted search cannot be repeated exactly, the goal is a record full enough that a reader can see what was done, judge its influence and repeat it approximately.
| Field in the AI-assisted search log | What to record | Cedar Valley entry for the second test |
|---|---|---|
| Tool and provider | Product name, provider, and type of account or licence | A scholarly AI research assistant, recorded by name; university licence |
| Model or version | The model or version shown, or a note that none was shown | The version shown on the settings page |
| Date and person | The date of the search and the team member who ran it | The date of the test; the librarian |
| Settings | Any filters, modes, date limits or source options selected | Default settings, with results limited to journal articles |
| Prompt | The exact wording entered, including any follow-up prompts | The review's PCC question as written in the protocol |
| Output | The full output saved as a file, with its file name | Exported list of 20 sources and a saved copy of the answer page |
| Handling | How the output was verified, de-duplicated and screened | 16 duplicates of database records; 4 outside the review's scope; none added |
This log sits with the web search documentation from Lesson 5 and the PRISMA-S log from Lesson 4. PRISMA 2020 also asks authors to report automation tools used in selecting studies and collecting data, which covers the tools in Section 4.
3.5 Privacy, Confidentiality and Copyright
Whatever is entered into an AI tool leaves the team's control to some degree. Depending on the provider's terms and the account, inputs may be stored, reviewed or used to train future models, and may be held in another country. Public bodies in British Columbia, including health authorities and universities, are subject to the Freedom of Information and Protection of Privacy Act and have policies on which AI tools staff and students may use. The team checks both institutions' policies before entering anything other than published material.
Interview notes, survey responses and any other information about identifiable people stay out of consumer AI tools. For the Cedar Valley environmental scan, this covers the responses from the 11 programs that complete the survey and the notes from the 7 key informant interviews (Lesson 11). Data that seem anonymous can still identify a person or a small organization.
Unpublished manuscripts, grant applications, internal health authority documents and draft reports belong to their authors or organizations. Many journals and funders prohibit peer reviewers from entering confidential material into generative AI tools; the United States National Institutes of Health, for example, prohibits its peer reviewers from using them to analyze or critique grant applications.
Published articles are protected by copyright, and the library's licences with publishers set terms for subscribed content, which in some cases restrict uploading full texts into AI tools. Before uploading full texts, the team asks the librarian whether the licence allows it.
3.6 Disclosure
A disclosure statement tells readers which AI tools were used, for what purpose, at which stages, and how their output was checked. The RAISE recommendations (Responsible use of AI in evidence SynthEsis), developed through work involving Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence, are emerging guidance on responsible AI use in evidence synthesis that was still being revised in 2025. The joint position statement that endorses them (Flemyng et al., 2025) asks authors to report fully any AI use that makes or suggests judgements, such as eligibility, risk of bias, extraction or synthesis decisions, giving each tool's name, version and dates of use, its purpose and the stages affected, the justification for using it, any related interests, and its limitations. Editing limited to spelling and grammar generally need not be reported, although journal policies vary. The template below covers these elements.
Copy the template and replace each bracketed field.
We used [tool name, provider and version or model] between [dates] to [purpose] at the [review stage or stages] stage. [The tool was used with the following settings: settings.] The prompts and full outputs are provided in [appendix or repository]. We chose this tool because [justification, including any evaluation of the tool or the team's own pilot]. Every output was [verification method, for example checked against the primary source by one reviewer and confirmed by a second], and [number] of [number] AI-suggested sources were retained after verification. The tool did not [list the judgements it did not make, such as decisions to include or exclude studies]. The authors declare [interests related to the tool, or no such interests]. Limitations of this use include [limitations and their likely effect on the review]. The authors take full responsibility for the content of this review.
During week five of the review we used a general-purpose conversational assistant, with web search turned off, and a scholarly AI research assistant, each recorded by name and version with the prompts and full outputs in Appendix C, to identify sources for citation chasing and to check the sensitivity of the database search. These tools supplemented a peer-reviewed database search. The student intern checked every AI-suggested source for existence, bibliographic accuracy and support for the stated finding, and the librarian confirmed the checks. Six of the eight sources suggested by the conversational assistant existed; two could not be found and were removed, three required corrections to their details or summaries, and none was an eligible primary study. None of the 20 sources from the research assistant was a new eligible study. No AI tool made screening, eligibility, charting or synthesis decisions, and no environmental scan data were entered into any AI tool. The authors declare no interests related to these tools. These searches cannot be reproduced exactly because AI outputs vary over time. The authors take full responsibility for the content of this review.
I used [tool and version] on [date] to [purpose]. My prompts and the full output are in [appendix]. I checked every source it suggested against [databases or systems] and recorded the results in my verification log; [number] of [number] sources were verified and kept. I made all screening, eligibility and synthesis decisions myself, and I did not enter any personal or confidential information into the tool.
Where the disclosure goes
Place the statement where the journal or funder specifies, usually the methods section, and put the prompts, outputs and verification log in an appendix or repository.
Reflection
You are checking AI-suggested references with a four-step procedure: (1) existence: can the exact title be found in a bibliographic database, Google Scholar or Crossref, and does the DOI resolve at doi.org; (2) details: do the authors, year, journal, volume and pages match the publisher's record; (3) claim: does the source say what the AI summary says; and (4) fit: does the source meet the review's eligibility criteria and has it been retracted. Verdicts are: verified as given; real with wrong details (correct it); real with an unsupported claim (rewrite the summary); conflated (corrected); or not found (treat as fabricated and remove). One AI-suggested reference describes a 2019 randomized trial of volunteer befriending for older adults by two named authors, with the summary "befriending halved loneliness scores over twelve months." Suppose that your checks find that the exact title appears in no database, that the DOI returns "DOI not found", and that Google Scholar shows a 2019 trial of volunteer befriending with a similar title, by different authors in a different journal, which reported a small reduction in loneliness that was not statistically significant. Write the verification-log entry for the reference (existence, details, claim, verdict and action), explain whether the similar trial can replace it and on what conditions, and write a two-sentence disclosure statement for this use of the AI tool.
Log entry. Existence: not found; the title is absent from MEDLINE, Google Scholar and Crossref, and the DOI does not resolve. Details and claim: not applicable, because there is no source to compare. Verdict: not found, treated as fabricated. Action: removed from the reference list, with the entry kept in the log so that the removal is documented.
The similar trial. The real trial is a different source, by different authors in a different journal, so it cannot be substituted as a correction of the AI reference. It can enter the review only as a new lead found while verifying: I would retrieve it, confirm its details from the publisher's record, read it to describe its finding accurately (a small, non-significant reduction in loneliness, which contradicts the AI summary), check it against the eligibility criteria, and screen it through the usual process. If it is included, it is reported as a record identified by other methods.
Disclosure. I used [tool name and version] on [date] to suggest sources on volunteer befriending, and the prompt and full output are in Appendix B. I checked every suggested source against MEDLINE, Google Scholar and Crossref, removed one reference that could not be found, and made all eligibility decisions myself.
Minimum 20 characters required.
Question 1: In the four-step verification procedure, what is checked in step 3?
Question 2: In the Cedar Valley exercise, entry 5 (Bickerdike and colleagues, 2017) existed and had correct details. Why was it flagged?
Question 3: Why can an AI-assisted search not be reproduced exactly?
Question 4: Under the 2025 joint position statement on AI in evidence synthesis, which use of AI must be fully reported?
Active-Learning Screening and the Evidence on Its Performance
Learning Objectives for this section
- Explain how active learning orders records during title and abstract screening, and how prioritized screening differs from using a model as a second screener or as a basis for stopping early.
- Interpret recall, the proportion of records screened and work saved over sampling from a screening simulation.
- Describe the stopping problem and compare heuristic and statistical stopping rules.
- Summarize what published evaluations show about active learning and large language model screening, and the conditions that limit those findings.
- Use a decision table to judge when AI assistance is acceptable at each stage of a review, and apply it to the Cedar Valley review.
4.1 The Screening Workload
Title and abstract screening is usually the most time-consuming stage of a review. In the fictional Cedar Valley review, 1,870 records remain after de-duplication, and Lesson 7 describes how the evidence officer and the intern screen each one independently, which amounts to 3,740 decisions before any full text is read. Machine learning has been applied to this stage for longer than to any other, and its use here has the most evidence behind it.
Screening tools can use machine learning in three ways, and the distinction matters more than the choice of product. In prioritized screening, the tool changes the order in which records are shown, and the team still screens every record. With a model acting as a second screener, the model's prediction takes the place of one of two human reviewers. In screening truncation, the team stops screening before it has seen every record, because the model's ranking suggests that few relevant records remain. O'Mara-Eves and colleagues (2015), in a systematic review of text mining for identifying studies, concluded that using text mining to prioritize the order of screening should be considered safe and ready for use in live reviews, that its use as a second screener could be considered with caution, and that using it to eliminate studies automatically was promising but not yet fully proven.
4.2 How Active Learning Works
Active learning is a form of machine learning in which the model chooses which unlabelled record a human should label next, learns from that label, and chooses again. In screening, the reviewer first labels a few records as relevant or irrelevant. The tool converts each title and abstract into numerical features, such as word counts weighted by how distinctive the words are, or embeddings of the kind described in Section 1. A classifier, such as naive Bayes or logistic regression, learns from the labelled records which features predict relevance, and then scores every unscreened record. The tool shows the reviewer the record with the highest score. The reviewer decides, the model retrains on the larger set of decisions, and the cycle repeats.
ASReview, described by van de Schoot and colleagues (2021), is an open-source tool developed at Utrecht University that implements this cycle and lets the user choose among several feature extractors and classifiers. Its simulation mode replays a completed review whose decisions are known, showing how early the relevant records would have appeared. Screening platforms such as Rayyan, EPPI-Reviewer and DistillerSR offer comparable machine-learning features within their own workflows. A different kind of tool, the classifier trained once on a large labelled collection, is used for specific tasks. Thomas and colleagues (2021), for example, developed and evaluated a classifier that identifies reports of randomized controlled trials for Cochrane Reviews.
4.3 Measuring Performance: A Worked Cedar Valley Simulation
Three measures describe how well a screening tool performs. Recall, also called sensitivity, is the proportion of all relevant records that have been found at a given point. The proportion screened is the share of all records the reviewer has read at that point. Work saved over sampling, introduced by Cohen and colleagues (2006), compares the work needed to reach a target recall with the work needed to reach the same recall by screening in random order.
Work saved over sampling at 95 percent recall
WSS@95 = (N − n95) ÷ N − 0.05
Here N is the total number of records and n95 is the number screened when 95 percent of the relevant records have been found. In random order, a reviewer must screen about 95 percent of the records to find 95 percent of the relevant ones, so subtracting 0.05 removes the saving that random order would give.
Lesson 7 describes how the team screens all 1,870 records in duplicate and sends 145 forward for full-text retrieval, of which 142 full texts are obtained and assessed. To learn what active learning might offer when the review is updated, the intern replays those completed decisions in ASReview's simulation mode, treating the 142 records whose full texts were assessed as the relevant records. The figures that follow are illustrative and were constructed for teaching. In the simulation, 88 of the 142 relevant records appear within the first 187 records screened, which is the first 10 percent. The 135th relevant record, which takes recall past 95 percent (135 ÷ 142 = 95.1 percent), appears at record 540. The last relevant record appears at record 1,395.
| Point in the simulation | Records screened | Proportion screened | Relevant found | Recall |
|---|---|---|---|---|
| First 10 percent of records | 187 | 10.0% | 88 | 62.0% |
| 95 percent recall reached | 540 | 28.9% | 135 | 95.1% |
| Stopping heuristic triggered | 710 | 38.0% | 138 | 97.2% |
| Last relevant record found | 1,395 | 74.6% | 142 | 100% |
| All records screened | 1,870 | 100% | 142 | 100% |
The work saved over sampling at 95 percent recall is (1,870 − 540) ÷ 1,870 − 0.05 = 0.711 − 0.05 = 0.661, or about 66 percent. In words, prioritized screening reached 95 percent recall after the reviewer had read 540 records, whereas random order would have required about 1,777 records (95 percent of 1,870) to reach the same recall on average. Figure 6.5 shows the same simulation as a curve.
4.4 The Stopping Problem
The simulation measures are calculated after the fact, when the team knows that 142 records are relevant. In a live review nobody knows the total, so nobody knows the recall at any moment. Prioritized screening that continues to the last record is safe, and the time saving from stopping early is where the risk lies. A stopping rule is a pre-specified criterion for ending screening before every record has been read.
Heuristic stopping rules are simple to apply. The most common stops after a fixed number of consecutive irrelevant records, such as 50 or 100, on the reasoning that a long run without a relevant record means few remain. Boetje and van de Schoot (2024) proposed the SAFE procedure, a practical and deliberately conservative combination of stopping heuristics for active-learning screening in tools such as ASReview. Statistical stopping rules take a different approach. Callaghan and Müller-Hansen (2020) proposed drawing a random sample from the records that remain unscreened and using the number of relevant records in the sample to test whether a recall target, such as 95 percent, has been reached with a stated level of confidence. Callaghan and colleagues (2024) later argued that stopping criteria for computer-assisted screening need to be well evaluated and transparently reported before reviews rely on them.
In the Cedar Valley simulation, a heuristic of 100 consecutive irrelevant records would have stopped screening at record 710, after the 138th relevant record had appeared at record 610. Stopping there would have saved reading 1,160 records, or 62.0 percent of the total, and would have missed 4 of the 142 relevant records (recall 97.2 percent). Suppose that one of the four missed records was a qualitative evaluation of a men's shed program whose abstract never used the words loneliness or isolation, and that it was later among the included studies. A missed record can matter even at high recall. Hard-to-find records usually describe the concept in unexpected language or have weak abstracts, and a scoping review of a varied public health literature most needs to catch them.
4.5 What the Evidence Shows
The evidence on active learning comes mainly from simulations on completed reviews. O'Mara-Eves and colleagues (2015) found that most of the studies in their systematic review suggested workload savings of between 30 and 70 percent, sometimes accompanied by the loss of 5 percent of relevant studies, which is 95 percent recall. They noted that studies rarely replicated each other, and that more evaluation was needed outside highly technical and clinical areas. Van de Schoot and colleagues (2021) reported simulation studies with ASReview in which relevant records were found after screening a fraction of the total. Hamel and colleagues (2021) evaluated active machine learning across ten completed systematic reviews and developed a seven-step framework for teams adopting it, which ends with guidance on truncating screening.
Several conditions limit how far these findings transfer to a new review. Performance varies with the topic, the share of records that are relevant, the clarity of the abstracts and the size of the dataset. Simulations treat the original human decisions as the truth, although human screeners also miss studies; Gartlehner and colleagues (2020) found in a randomized trial that single-reviewer abstract screening missed 13 percent of relevant studies. The joint position statement on AI in evidence synthesis (Flemyng et al., 2025) cites that finding as an example of the trade-offs a team should weigh, since in a rapid review an AI tool used as a second reviewer may catch studies that a single human reviewer would miss.
Evaluations of large language models for screening are newer and less settled. Khraisha and colleagues (2024) tested GPT-4 on title and abstract screening, full-text screening and data extraction across several literature types and languages. GPT-4's accuracy matched human performance on some tasks, but performance fell across all stages once the authors adjusted for chance agreement and for the imbalance between included and excluded records. The authors concluded that substantial caution was warranted, while noting that for certain tasks under specific conditions the model approached human performance.
A reader can test a claim that a tool saves a stated share of screening work by asking whether the result came from a simulation or a live review, at what recall the saving was measured, how similar the evaluation datasets were to the reader's topic, whether the stopping point was chosen in advance, and whether the evaluators had an interest in the tool.
PRISMA 2020 asks authors to describe any automation tools used in selecting studies. A full description names the tool and version, the model settings, the starting labels, whether the tool prioritized, acted as a second screener or supported stopping early, the stopping rule and when it was specified, the number of records left unscreened, and any estimate of recall.
4.6 When AI Assistance Is Acceptable: A Decision Table
The table below brings the lesson together. It is a teaching tool for this course that applies three principles found in the joint position statement and in the earlier sections: human oversight of every judgement, verification of every output against primary sources, and full disclosure of any AI use that makes or suggests judgements. Journals, funders, the university and the health authority may set stricter rules, and those rules take precedence.
| Review stage | Generally acceptable, with human checking | Acceptable only with added safeguards | Not acceptable |
|---|---|---|---|
| Question, scope and protocol (Lessons 1 and 2) | Explaining question frameworks; checking the clarity of the team's own draft | Suggesting options for scope, with the team and interest holders deciding | Letting a tool set the question or the eligibility criteria |
| Search development (Lessons 3 and 4) | Suggesting synonyms and spelling variants for the concept table | Translating a search between databases, with the librarian checking every line and a PRESS peer review | Running an AI-generated search string without librarian review and testing |
| Searching (Lessons 4 and 5) | Suggesting organizations and websites to search | AI search engines as a supplementary, documented search, with every source verified | Replacing the database search with an AI search; citing unverified AI-suggested sources |
| Citation chasing (Lesson 5) | Citation-network and citation-context tools to find related work, with records screened as usual | Relying on a tool's recommendations as the only form of citation chasing, with the limits reported | Treating a "supporting" citation label as evidence of a study's quality |
| Title and abstract screening (Lesson 7) | Active learning to order records while every record is screened | A model as second screener, or stopping early under a pre-specified, evaluated stopping rule, with recall estimated and reported | Excluding records by a tool alone, with no human check and no estimate of recall |
| Full-text screening (Lesson 7) | Tools that find full texts or highlight passages | Model-suggested decisions, with a human deciding on every record | A model making final eligibility decisions |
| Data extraction and charting (Lessons 8 and 10) | Formatting a charting table the team designed | Model-drafted values, with every value checked against the full text | Unchecked model output entering the charting table |
| Appraisal and certainty (Lesson 8) | Explaining the items of an appraisal tool | Model suggestions as one input, with two human reviewers deciding | Reporting model judgements as the reviewers' own |
| Synthesis and reporting (Lessons 9 and 12) | Editing the grammar and clarity of the team's text | A model-drafted plain-language summary, checked line by line against the findings and disclosed | Model-written findings, or text containing unverified references or claims |
| Environmental scan (Lesson 11) | Suggesting survey wording for the team to revise | Transcription or analysis only with tools approved for personal information and with participants' consent | Entering interview notes or survey responses into consumer AI tools |
The team records its decisions in the protocol's amendment log. A conversational assistant may suggest synonyms, which the librarian tests in each database. It has run the two supplementary AI searches described in Sections 2 and 3, with every source verified. It will screen all 1,870 records in duplicate, as Lesson 7 describes, and will use active learning only in a simulation to plan future updates. No AI tool will make eligibility, charting or synthesis decisions, and no information from the 18 programs in the environmental scan will be entered into any AI tool. The final report will include the disclosure statement from Section 3, and the AI-assisted search log and verification log will appear in an appendix. Lesson 10 returns to these choices when it considers the shortcuts that rapid reviews take.
Putting the lesson into practice
A review team that plans to use AI assistance usually works through four steps. Using a tool its institution permits, it runs one AI-assisted search on its question and records the tool, version, date, settings, exact prompt and full output in an AI-assisted search log. It verifies every source the tool suggests with the four-step procedure from Section 3 and records each check in a verification log, noting which sources exist, which needed corrections and which, if any, are eligible for the review. It marks, in a copy of the decision table, the stages at which it plans to use AI assistance and the safeguard it will apply at each one. And it writes a disclosure statement from the template in Section 3.6, entering no personal or confidential information into any tool.
Reflection
A team replays a completed review in an active-learning tool's simulation mode. The review had 2,000 records after de-duplication, of which 80 were relevant at the title and abstract stage. In the simulation, 95 percent recall (76 of 80 relevant records) was reached after 600 records had been screened. A heuristic stopping rule of 100 consecutive irrelevant records would have stopped screening at record 750, when 77 of the 80 relevant records had been found. The last relevant record appeared at record 1,420. Recall is the proportion of all relevant records found so far. Work saved over sampling at 95 percent recall is WSS@95 = (N − n95) ÷ N − 0.05, where N is the total number of records and n95 is the number screened when 95 percent recall is reached. (a) Calculate the recall and the proportion of records screened at the stopping point, and WSS@95. (b) Explain why these figures would not be available during a live review. (c) Recommend whether a rapid review for a health authority should stop screening at the heuristic's stopping point, and state what the review would need to report.
(a) At the stopping point, recall is 77 ÷ 80 = 96.25 percent, and the proportion screened is 750 ÷ 2,000 = 37.5 percent. WSS@95 = (2,000 − 600) ÷ 2,000 − 0.05 = 0.70 − 0.05 = 0.65, so prioritized screening saved about 65 percent of the work that random order would need to reach the same recall.
(b) Recall and WSS depend on knowing the total number of relevant records, which is known in a simulation only because every record was screened. In a live review the total is unknown, so the team cannot tell how many relevant records remain when the heuristic triggers. The last relevant record here appeared at record 1,420, long after the stopping point.
(c) For a rapid review, stopping at the heuristic's point can be defensible if the stopping rule was specified in the protocol before screening began, if the topic resembles those on which the method has been evaluated, and if the team accepts a small risk of missing studies. I would prefer to add a statistical check, such as screening a random sample of the remaining records, to estimate recall. The review would need to report the tool and version, the starting labels, the stopping rule and when it was chosen, the number of records left unscreened, any estimate of recall, and the possibility that relevant studies were missed. If the review must find every eligible study, the team should screen all records and use the tool only to order them.
Minimum 20 characters required.
Question 1: What distinguishes prioritized screening from screening truncation?
Question 2: In the Cedar Valley simulation, 95 percent recall was reached after 540 of 1,870 records. What is the work saved over sampling at 95 percent recall?
Question 3: What does a statistical stopping rule such as the one proposed by Callaghan and Müller-Hansen (2020) do?
Question 4: Which conclusion is consistent with Khraisha and colleagues' (2024) evaluation of GPT-4 for screening and extraction?
Final Assessment
Bringing It All Together
This lesson examined how artificial intelligence tools produce answers and what that means for a review team. Section 1 showed that a large language model writes text, including references, by predicting one token after another from patterns learned in training, which explains why fabricated and conflated references occur. Retrieval-augmented generation ties citations to retrieved documents and reduces fabrication, but its claims can still misstate their sources, and its relevance ranking returns a top set with unknown recall. Section 2 sorted AI research tools into conversational assistants, AI search engines and research assistants, citation-context tools and AI-assisted screening tools, and set out questions for evaluating any new tool.
Section 3 turned these ideas into practice. The fictional Cedar Valley team verified eight AI-suggested references with a four-step procedure and found that six existed, three were correct as given and two had been fabricated. The section introduced the AI-assisted search log, the privacy, confidentiality and copyright limits on what may be entered into a tool, and a disclosure statement built on the RAISE recommendations and the 2025 joint position statement. Section 4 explained active-learning screening, the measures used to evaluate it, the stopping problem and the evidence on performance, and ended with a decision table for judging AI use at each stage of a review.
Key Takeaways from this lesson
- A large language model generates references from learned patterns, so a reference produced without retrieval is a lead to be checked and never evidence that a source exists.
- A DOI and familiar author names do not show that a reference is real; a DOI counts only when it resolves to the same work at doi.org.
- Retrieval-augmented tools reduce fabricated references, but their summaries can misstate sources and their relevance ranking returns a top set with low and unknown recall.
- Conversational assistants, AI search engines, citation-context tools and AI-assisted screening tools differ in where their answers come from, and each supports different review tasks.
- AI search tools can supplement and check a documented database search, and studies they find enter the review as records identified by other methods.
- Every AI-suggested source passes four checks (existence, bibliographic details, claim and fit), and each check is recorded in a verification log.
- An AI-assisted search cannot be repeated exactly, so the tool, version, date, settings, prompt and full output are recorded in a search log.
- Personal information, confidential documents and licensed full texts stay out of AI tools unless institutional policy and licence terms allow them.
- Disclosure follows emerging guidance such as the RAISE recommendations, and reports each tool's name, version, dates, purpose, stages, justification, interests and limitations.
- Active learning safely orders records for screening, while stopping early saves more time at the cost of a recall that is unknown unless it is estimated.
Core Concepts Reviewed
Section 1: tokens, next-token prediction, parameters, knowledge cutoff, context window, temperature, retrieval-augmented generation, embeddings and semantic search, hallucination, and fabricated, conflated and mis-attributed references.
Section 2: conversational assistants, AI search engines and research assistants, citation-context tools, AI-assisted screening and extraction, relevance ranking and recall, and questions for evaluating an unfamiliar tool.
Section 3: reviewer responsibility, the four-step verification procedure, the verification log, the AI-assisted search log, privacy, confidentiality and copyright, the RAISE recommendations, the 2025 joint position statement and the disclosure statement.
Section 4: prioritized screening, the model as second screener, screening truncation, active learning, recall, work saved over sampling, heuristic and statistical stopping rules, evidence on performance and the decision table for AI use.
The final reflection asks you to write an AI-use plan and disclosure statement for a new rapid review by the Cedar Valley evidence team.
Reflection
The fictional Cedar Valley Health Authority asks its three-person evidence team (an evidence officer, a university librarian and a student intern) for a rapid review, due in eight weeks, of community-based falls prevention programs for adults aged 65 and older. The librarian expects the database search to yield about 3,000 records after de-duplication. The team has access to a conversational AI assistant, a scholarly AI research assistant and an active-learning screening tool. The health authority's policy forbids entering personal information into consumer AI tools, and the team also plans to interview staff from five falls prevention programs. Write an AI-use plan of 250 to 400 words that states, for (1) searching, (2) title and abstract screening, (3) data extraction and (4) the program interviews, whether and how AI will be used and what safeguard applies at each stage; describes how AI-suggested sources will be verified, using the four steps of existence, bibliographic details, claim and fit; and ends with a disclosure statement of three or four sentences that names the tools, purposes and stages, the verification method, and the judgements the tools did not make.
Searching. The librarian will build and peer-review the database search. The conversational assistant may suggest synonyms, which the librarian will test in each database. The research assistant will be run once as a supplementary search, with the prompt, version, date, settings and full output saved in an AI-assisted search log; new sources will be screened as records identified by other methods.
Verification. The intern will check every AI-suggested source for existence, bibliographic details, claim and fit, record each check in a verification log for the librarian to confirm, and remove any source that cannot be found.
Screening. One reviewer will screen every record in the order set by the active-learning tool, which will act as a second screener. Any early stop will follow a stopping rule set in the protocol, with recall estimated from a random sample of the remaining records.
Extraction and interviews. Two reviewers will extract data by hand, and no interview notes will be entered into any AI tool.
Disclosure. We used [assistant, version] to suggest search terms, [research assistant, version] for one supplementary search and [screening tool, version] to prioritize and second-screen titles and abstracts between [dates]. Prompts, outputs and verification logs are in Appendix C. No AI tool made eligibility, extraction or synthesis decisions, and the stopping rule and estimated recall are reported in the methods. The authors take full responsibility for the review.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: The intern asks a conversational assistant with web search turned off for studies published last month. What is the most likely problem?
Question 2: Which pairing of a tool category with a review task is most appropriate?
Question 3: An AI-suggested reference has a DOI that returns "DOI not found", and no database or search engine has the title. What belongs in the verification log?
Question 4: How does the semantic search used in many retrieval-augmented tools differ from Boolean retrieval in a bibliographic database?
Question 5: In the Cedar Valley log, entry 6 (Cattan and colleagues) had the right title and authors but the wrong journal, year and volume. What is the correct action?
Question 6: Which element belongs in a disclosure statement under the 2025 joint position statement?
Question 7: Which material could the Cedar Valley team enter into a consumer AI tool, following the lesson's guidance?
Question 8: Why does the formula for work saved over sampling at 95 percent recall subtract 0.05?
Question 9: A team stops screening after 100 consecutive irrelevant records. Which description of this rule is accurate?
Question 10: According to the lesson's decision table, which use of AI at title and abstract screening is generally acceptable with human checking?
Question 11: Why was entry 7 (Okafor and colleagues, 2023) described as the most tempting entry in the Cedar Valley exercise?
Question 12: Which statement about active learning in a tool such as ASReview is accurate?
Question 13: What did O'Mara-Eves and colleagues (2015) conclude about text mining for screening?
Question 14: The 2025 joint position statement cites the finding that single-reviewer abstract screening missed 13 percent of relevant studies. How does it use this finding?
Question 15: A student writes: "The AI tool cited twelve papers on my topic, so my search is done." Which response best applies the lesson?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
Definitions of the terms, tools, frameworks and people introduced in this lesson.