How Databases and Search Engines Find Information
Finding & Synthesizing Health Evidence
Learning objectives for this lesson:
- Describe how bibliographic databases index records, and distinguish the subject headings added by a database from the fields written by authors.
- Explain how the controlled vocabularies MeSH, Emtree and CINAHL Subject Headings work, including entry terms, hierarchies, explosion and major topics.
- Explain why a sensitive search combines controlled vocabulary with free-text terms for every concept.
- Apply the Boolean operators AND, OR and NOT to sets of records, and explain how an inverted index supports exhaustive retrieval.
- Explain in general terms how relevance ranking orders results in Google and Google Scholar, and distinguish what is publicly documented from what is undisclosed.
- Calculate recall, precision and the number needed to read, and estimate recall with a benchmark set of known studies.
- Explain why systematic and scoping reviews favour recall, and identify when a team might knowingly accept lower recall.
- Identify the sources of variation that limit the reproducibility of web searches, and record a web search so that it is transparent.
- Describe the groundwork a review team prepares before building a full search: subject headings and free-text terms for each concept, a benchmark set of known studies and a documented test web search.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.
Indexing, Controlled Vocabularies and Free Text
Learning Objectives for this section
- Describe what a bibliographic database record contains and distinguish the fields written by authors from the fields added by the database.
- Explain how a controlled vocabulary works, including preferred terms, entry terms, the hierarchy of broader and narrower headings, explosion and major topics.
- Compare MeSH, Emtree and CINAHL Subject Headings and identify the database that each one indexes.
- Explain why a sensitive search combines controlled vocabulary with free-text terms for each concept.
1.1 Why the Same Question Returns Different Results
A student who types the same words into PubMed, Embase, CINAHL, Google and Google Scholar will receive five different sets of results. The counts differ, the first records differ, and some studies appear in one tool and are missing from another. These differences arise from how each tool stores information and how it decides which records match a query. A reviewer who understands those mechanisms can build a search that finds what it should, explain the search to others, and judge which tools can support a systematic search and which cannot.
This section explains indexing and contrasts searching subject terms with searching the words that authors wrote. Section 2 compares Boolean retrieval in databases with relevance ranking in Google and Google Scholar, Section 3 introduces recall and precision, and Section 4 examines the reproducibility of web search. Writing and translating a full search string belongs to Lesson 4, and grey literature and targeted web searching belong to Lesson 5.
The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, including about 46,000 adults aged 65 and older. Before launching a community connector (social prescribing) program, its planning team has asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The team is an evidence officer, a university librarian and a student intern. In Lesson 2 the team wrote its question in PCC form: the population is adults aged 65 and older, the concept is community-based interventions evaluated for loneliness or social isolation, and the context is any community setting. In this lesson the librarian explains to the intern how the databases they will search find information.
1.2 What a Database Record Contains
A bibliographic database is an organized collection of records, each describing one published item such as a journal article. MEDLINE, the database of the United States National Library of Medicine (NLM), is the best-known example in the health sciences, and PubMed is the free interface through which most people search it. Embase, CINAHL, APA PsycInfo and Web of Science are other major databases that the Cedar Valley team will search.
Each record contains two kinds of fields. The first kind comes from the authors and the publisher: the title, the abstract, the author names, the journal, and any keywords the authors chose. The second kind is added by the database: the subject headings that describe what the article is about, the publication type (for example, randomized controlled trial or review), the language, and other descriptive codes. Adding this second kind of information is called indexing. A search can look in either kind of field, and the results differ depending on which kind it uses.
The example shows why indexing matters. The authors described their participants as "older adults" and "seniors" and their outcome as being "lonely". Another team writing about the same topic might say "elderly people" and "social disconnection". Indexing replaces this variety with a single standard heading for each concept, so that every article about loneliness in people aged 65 and older carries the same headings regardless of the words its authors chose.
1.3 How a Controlled Vocabulary Works
A controlled vocabulary is a fixed list of approved terms used to describe the subject of each record. When the list is organized with relationships among the terms, it is called a thesaurus. Health databases use thesauri with five main features, and each one affects how a search behaves.
First, each concept has one preferred term, also called a subject heading or descriptor, such as Loneliness. Second, each preferred term has entry terms, which are synonyms and variant spellings that point to it. In MeSH, the entry terms for Social Isolation include "Social Exclusion", and the entry terms for Aged include "Elderly". Third, the headings are arranged in a hierarchy of broader and narrower terms, which MeSH calls its tree structure. Fourth, each heading has a scope note that defines how indexers use it; the scope note for Aged, for example, defines it as a person 65 through 79 years of age and directs indexers to the narrower heading "Aged, 80 and over" for older people. Fifth, many databases allow subheadings (MeSH calls them qualifiers) that narrow a heading to an aspect such as psychology, therapy or epidemiology.
The hierarchy makes possible an operation called explode. Exploding a heading retrieves records indexed with that heading and with every narrower heading beneath it. Indexers are instructed to assign the most specific heading that fits an article, so an article about loneliness is indexed with Loneliness rather than with the broader Social Isolation. A search for Social Isolation that is not exploded will therefore miss articles indexed only with its narrower headings. A related option, major topic, restricts a search to records in which the heading describes a main point of the article rather than a minor one. Exploding increases the number of records retrieved, and restricting to major topic decreases it; Section 3 explains how these choices trade off.
1.4 Three Thesauri: MeSH, Emtree and CINAHL Subject Headings
Each major health database has its own thesaurus, and the headings, hierarchies and search syntax differ from one database to the next. A search built with MeSH headings therefore has to be translated, heading by heading, before it can run in Embase or CINAHL. Lesson 4 teaches that translation. The cards below summarize the thesauri the Cedar Valley team will meet.
Because each thesaurus is maintained separately, the same concept can carry a different heading, a different position in the hierarchy, or no heading at all, depending on the database. The tabs below show how a heading for the loneliness concept is written in four common interfaces. The aim at this stage is to recognize that each interface writes headings in its own way; Lesson 4 returns to the details.
In PubMed, a MeSH heading is written with the field tag [Mesh], as in "Social Isolation"[Mesh]. PubMed explodes MeSH headings by default, so this search includes Loneliness and the other narrower headings. Adding [Mesh:NoExp] turns explosion off, and the tag [Majr] restricts the search to records in which the heading is a major topic.
On the Ovid platform, a heading ends with a forward slash, as in Loneliness/. Explosion is requested explicitly by placing exp in front of the heading, as in exp Social Isolation/. Published systematic reviews often report MEDLINE searches in this form.
On Elsevier's Embase.com platform, an Emtree term is written in single quotation marks followed by a slash and an option, as in 'loneliness'/exp for an exploded search or 'loneliness'/mj for a search restricted to records in which the term is a major focus.
In CINAHL on the EBSCOhost platform, a heading is written with the code MH, as in (MH "Loneliness"). A plus sign inside the quotation marks explodes the heading, as in (MH "Social Isolation+"), and the code MM in place of MH restricts the search to major concepts.
1.5 Free-Text Searching and Why It Is Still Needed
A free-text search (also called a keyword or text-word search) looks for words in the fields written by the authors, usually the title, the abstract and the author keywords. It finds a record only when the record contains the exact words searched, or their variants when the searcher uses truncation, which Lesson 4 teaches. A free-text search for "loneliness" would miss the illustrative record shown in Section 1.2, whose title says "lonely", unless the searcher also included "lonely" or used truncation.
Free text is still needed because indexing has gaps, and each gap causes a subject-heading search to miss relevant records. The items below describe the main gaps, and Web of Science and web search engines, which have no thesaurus, rely on free text entirely.
A thesaurus adds a heading only after a concept appears often enough in the literature. MeSH added Social Prescribing as a descriptor in 2025, although studies of social prescribing had been published for many years before that. Records indexed before a new heading exists were indexed with other headings and generally do not carry the new one, so a search using only the new heading would miss most of the earlier studies. Free-text terms such as "social prescribing", "community connector" and "link worker" find those older records.
PubMed contains records that are not indexed for MEDLINE, including some articles supplied by publishers and articles available through PubMed Central from journals outside MEDLINE. These records have no MeSH headings at all, so only a free-text search can find them. Other databases have similar groups of records that are unindexed or only partly indexed.
Whether indexing is done by people, by software or by both, two similar articles can receive different headings. An article about a walking group might be indexed under Exercise and Social Support without any loneliness heading, even though loneliness was its primary outcome. Free-text terms catch articles whose headings missed the concept the reviewer is interested in.
Many intervention names, such as "befriending", "men's sheds" or "intergenerational programs", have no heading of their own. They are indexed under broader headings that also cover many unrelated studies. Searching the specific phrase in free text is often the only practical way to find these interventions.
PubMed also performs a step called automatic term mapping. When a user types words without field tags, PubMed tries to match them to MeSH headings, journal names and author names, and then combines the matched heading with a search of the words in all fields. Because an untagged search therefore does more than was typed, systematic searchers write each heading and free-text term with explicit field tags, and they check PubMed's "Details" display to see how a query was translated.
| Feature | Controlled vocabulary search | Free-text search |
|---|---|---|
| What it searches | Subject headings assigned by the database | Words in the title, abstract and keyword fields |
| Handles synonyms and spelling variants | Yes, through entry terms that point to one heading | Only for the variants the searcher lists or truncates |
| Includes narrower concepts | Yes, when the heading is exploded | Only when the narrower terms are listed |
| Finds new or unindexed records | No, because these records lack headings | Yes, provided the words appear in the record |
| Finds concepts without a heading | No, the concept falls under broader headings | Yes, by searching the specific phrase |
| Transfers between databases | No, each database has its own thesaurus | Mostly, with changes to syntax and field tags |
1.6 Combining the Two Approaches
Because each approach covers the other's gaps, the standard practice in systematic searching is to search every main concept with both its subject headings and a set of free-text terms, joined with OR, and then to join the concepts with AND. Chapter 4 of the Cochrane Handbook for Systematic Reviews of Interventions (Lefebvre et al., 2023) recommends this approach for searches that aim to be comprehensive. Section 2 explains exactly what OR and AND do to sets of records, and Lesson 4 shows how to assemble the full strategy in a concept table.
The pattern for one concept
Concept block for loneliness in MEDLINE (written in words): the exploded heading Social Isolation, OR the heading Loneliness, OR the free-text words loneliness, lonely, social isolation, socially isolated and social exclusion in the title or abstract.
Each concept in the PCC question gets a block like this, and the blocks are then combined: population block AND concept block (interventions) AND outcome block.
The two approaches also help each other. Librarians often read the headings and abstracts of a few known relevant articles to find headings and wording to add, which is one reason the Cedar Valley librarian asks the intern to collect known relevant studies (Section 3).
The intern types "social prescribing loneliness older adults" into PubMed and likes the first page. The librarian opens the MeSH record for Social Prescribing and points out that the heading was introduced in 2025. A study of a link worker program published in 2019 would have been indexed with headings such as Social Support, Community Networks and Loneliness, and would not carry the Social Prescribing heading. The librarian then opens the MeSH tree for Social Isolation and shows that Loneliness sits beneath it, so an exploded Social Isolation search also finds studies indexed only with Loneliness. The intern leaves with a draft list of headings and free-text terms for each concept, which becomes a concept table in Lesson 4.
Open the MeSH Browser on the National Library of Medicine website and look up the heading Aged. Record (a) its scope note, (b) at least one entry term, (c) the narrower headings that sit beneath it in the tree, and (d) the year it was introduced. Then decide whether a search for adults aged 65 and older should use the exploded heading or the unexploded one, and write one sentence explaining your choice. Repeat the exercise for the heading Loneliness, one of the Cedar Valley concepts.
Common errors at this stage
First searches often contain three errors: searching a heading without checking whether it was exploded, relying on a single heading for a concept that has no heading or only a recent one, and copying MeSH headings into CINAHL or Embase without checking each heading in that database's own thesaurus.
Reflection
A health authority team is preparing a MEDLINE search, run in PubMed, on social prescribing programs for adults aged 65 and older. The librarian gives the team four facts from the MeSH Browser. First, the heading Social Prescribing was introduced in 2025 and sits beneath the heading Social Support. Second, the heading Social Isolation has four narrower headings: Loneliness, Ostracism, Social Alienation and Social Deprivation. Third, the heading Aged is defined as a person 65 through 79 years of age, and it has two narrower headings, Aged, 80 and over, and Frail Elderly. Fourth, PubMed explodes headings by default.
(a) Explain why a search that uses only the heading Social Prescribing would miss relevant studies, and list at least four free-text terms you would add for the intervention concept. (b) Explain what the exploded heading Social Isolation retrieves and why this matters for the loneliness concept. (c) State whether you would explode the heading Aged, with a reason based on its definition.
(a) Because Social Prescribing became a MeSH heading only in 2025, records indexed before then were described with other headings, such as Social Support or Community Networks, and generally do not carry it. A search on the new heading alone would therefore find mainly recent records and miss most of the earlier evaluations. I would add free-text terms in the title and abstract fields for the ways authors describe these programs: social prescribing, social prescription, community connector, link worker, community navigator and community referral. I would keep the heading as well, because it will index future records consistently.
(b) The exploded heading Social Isolation retrieves records indexed with Social Isolation itself and with any of its four narrower headings. This matters because indexers assign the most specific heading available, so a study whose main outcome is loneliness is likely to be indexed with Loneliness rather than with Social Isolation. Exploding the broader heading captures those records in one step; I would still add Loneliness explicitly and add free-text terms such as lonely and social isolation.
(c) I would explode Aged. Its scope note covers ages 65 to 79 and sends indexers to Aged, 80 and over for older people, so an unexploded search would miss studies indexed only with the narrower heading, which is a large and relevant group for this question. Frail Elderly would also be included, which is appropriate.
Minimum 20 characters required.
Question 1: In a MEDLINE record, which of the following fields is added by the database rather than written by the authors?
Question 2: A librarian searches PubMed for "Social Isolation"[Mesh]. Loneliness is a narrower heading beneath Social Isolation. Which records does this search retrieve?
Question 3: Why would a MEDLINE search that uses only the MeSH heading Social Prescribing miss many relevant studies?
Question 4: Which statement best describes the relationship between MeSH, Emtree and CINAHL Subject Headings?
Boolean Retrieval and Relevance Ranking
Learning Objectives for this section
- Describe how an inverted index allows a database to find matching records quickly.
- Apply the Boolean operators AND, OR and NOT to small sets of records and predict the result of a query.
- Explain in general terms how relevance ranking orders results, using term frequency, inverse document frequency and link analysis, and identify which parts of commercial ranking are not disclosed.
- Compare bibliographic databases, Google and Google Scholar on exhaustiveness, transparency and suitability for systematic searching.
2.1 The Inverted Index
Section 1 described the fields in a database record. This section explains how a search system uses those fields to answer a query. A database with tens of millions of records cannot read every record each time a user searches. Instead, it prepares in advance a structure called an inverted index. For every word and every subject heading in the collection, the inverted index stores a list of the records that contain it, called a postings list. The idea resembles the index at the back of a textbook, which lists for each term the pages where it appears. Manning, Raghavan and Schütze (2008) give a full account of the inverted index in their textbook Introduction to Information Retrieval, and the description here follows theirs in simplified form.
When a user enters a query, the system looks up each term in the inverted index and works with the postings lists rather than with the records themselves. A search for "loneliness" therefore does not find "lonely", because the two words have separate entries in the index, unless the searcher asks for both or uses truncation. A search on a subject heading consults the entries for headings, which are separate from the entries for words in the title and abstract.
2.2 Boolean Retrieval in Bibliographic Databases
Boolean retrieval treats a query as a logical statement and returns every record for which the statement is true. It takes its name from the English mathematician George Boole, whose 1854 book The Laws of Thought set out an algebra of logical operations. Three operators do almost all the work in health searching. AND returns records that are in both sets (the intersection), so it narrows a search. OR returns records that are in either set or both (the union), so it broadens a search. NOT returns records in the first set that are absent from the second (the difference), so it removes records.
Boolean retrieval in a bibliographic database has three properties that matter for evidence synthesis. It is exhaustive: the system returns every record that satisfies the logic, and the count it reports is the exact size of that set. It is transparent: anyone who reads the search strategy can work out why each record was or was not retrieved. It is repeatable: the same strategy run in the same database on the same date returns the same set, so another team can check the search and a later team can update it. These properties are the reason systematic reviews rely on bibliographic databases for their main searches, and they are the reason PRISMA-S (Rethlefsen et al., 2021) asks reviewers to report every strategy in full.
Boolean logic also has well-known weaknesses. A record either matches or it does not, with no partial credit, so a record that says "lonely" fails a query that asks for "loneliness". The searcher must anticipate every relevant word in advance. NOT is risky, because it removes any record that mentions the excluded term, including relevant ones; a search for older adults NOT dementia would drop a trial of a befriending program that mentions dementia once in passing. Databases also differ in the order in which they process operators when parentheses are missing, so careful searchers always use parentheses to group terms. Lesson 4 teaches the full syntax.
| Record | Title (illustrative) | Indexed title words |
|---|---|---|
| 1 | Loneliness among older adults: a group intervention trial | loneliness, older adults, intervention |
| 2 | Social isolation and mortality in older adults | social isolation, older adults |
| 3 | A befriending program for lonely older adults | lonely, older adults |
| 4 | Loneliness in university students | loneliness |
| 5 | Social prescribing for older adults: a pilot evaluation | older adults |
| 6 | Telephone calls to reduce loneliness in care homes | loneliness |
| 7 | Social isolation in rural communities | social isolation |
| 8 | Exercise for older adults with arthritis | older adults |
Using the toy collection in the table and the postings lists in the inverted-index figure, write down the records each query returns: (1) loneliness AND older adults; (2) (loneliness OR lonely OR social isolation) AND older adults; (3) older adults NOT loneliness. Then decide which of the three queries finds the most of the records you would want for the Cedar Valley question, and which relevant record none of them can find from title words alone.
Query 1 returns record 1 only, because 1 is the only number in both the loneliness list (1, 4, 6) and the older adults list (1, 2, 3, 5, 8). Query 2 first forms the union of the three loneliness-related lists, which is records 1, 2, 3, 4, 6 and 7, and then intersects it with the older adults list, giving records 1, 2 and 3. Query 3 removes records 1, 4 and 6 from the older adults list, giving records 2, 3, 5 and 8; it discards record 1, the most relevant record in the collection, which shows why NOT is dangerous.
Query 2 finds the most relevant records. Record 5, on social prescribing for older adults, may well measure loneliness, but its title does not mention it, so none of the three queries can find it from title words. A search of its abstract or its subject headings, or a fourth concept for the intervention, would be needed.
Sorting a Boolean result set: PubMed Best Match
Many databases now offer a relevance sort, and since 2020 PubMed's default display order has been "Best Match", a machine-learning ranking described by Fiorini et al. (2018). Sorting changes only the order in which the retrieved records are displayed. The set of records, and the count PubMed reports, are the same whichever sort is chosen. A systematic search that screens every retrieved record is therefore unaffected by the sort order, whereas a search that looks at only the first page is strongly affected by it.
2.3 Relevance Ranking
Relevance ranking (also called ranked retrieval) works differently. Instead of deciding whether each record matches, the system gives each document a score for how well it appears to answer the query and presents the documents in order of that score. Documents that contain only some of the query words can still be returned, lower in the list. Users of a ranked system look at the first page or two and stop, so the ranking, rather than the logic of the query, decides what the user sees.
Ranking methods developed through decades of published research in information retrieval. Gerard Salton and colleagues at Cornell University developed the vector space model, which represents queries and documents as weighted lists of terms and ranks documents by their similarity to the query. Two weighting ideas from that research remain central. Term frequency gives more weight to a document that uses a query word more often. Inverse document frequency, proposed by Karen Spärck Jones (1972), gives more weight to words that are rare in the collection, because a rare word says more about what a document is about than a common one. Probabilistic models developed by Stephen Robertson and colleagues led to the widely used BM25 formula (Robertson and Zaragoza, 2009), which adds limits on the effect of repeated words and adjusts for document length.
A worked illustration of term weighting
Inverse document frequency for a term = log10(N ÷ df), where N is the number of documents in the collection and df is the number of documents that contain the term.
Suppose a collection holds 1,000,000 documents. The word "adults" appears in 200,000 of them, so its weight is log10(1,000,000 ÷ 200,000) = log10(5) = 0.699. The word "loneliness" appears in 2,000 of them, so its weight is log10(500) = 2.699, about 3.9 times the weight of "adults".
Document A uses "loneliness" 6 times and "adults" 2 times; a simple score of frequency multiplied by weight gives 6 × 2.699 + 2 × 0.699 = 17.59. Document B uses "loneliness" once and "adults" 10 times, giving 2.699 + 10 × 0.699 = 9.69. Document A ranks above Document B. Real systems dampen repeated words and adjust for length, but the principle is the same.
The same principle applies to the toy collection. In a ranked system, a query for loneliness older adults would return record 1 first, because it contains both words. Records 4 and 6, which contain only "loneliness", would come next, because "loneliness" appears in three of the eight records and is rarer than "older adults", which appears in five. Records 2, 3, 5 and 8 would follow. Record 7 contains neither word and would not appear. The ranked system returns seven records in order, where the Boolean query in the exercise returned exactly one.
Web search engines add information that ordinary text collections lack: the links between pages. Sergey Brin and Lawrence Page (1998) described PageRank, which treats a link from one page to another as a kind of endorsement and gives a page a higher score when many well-linked pages point to it. PageRank was part of the original design of Google. Link analysis is now one signal among many, and the details of how it is currently used are not public.
Google's public documentation describes its ranking systems only in general terms. It states that the systems consider the meaning of the query, the relevance of page content, signals of quality such as expertise and trustworthiness, the usability of pages, and the user's context and settings, including location, search history and language. It also states that the weight of each factor varies with the type of query. Google does not publish the signals in full, their weights or their interactions, and it changes its systems continually. Any more specific claim about how Google ranks a given page should be treated with caution.
Google Scholar applies ranking to scholarly documents. Its own help pages state that it aims to rank documents in the way researchers do, weighing the full text of each document, where it was published, who wrote it, and how often and how recently it has been cited in other scholarly literature. An early independent analysis by Beel and Gipp (2009) concluded that citation counts carried heavy weight in the ranking. One consequence is that older, highly cited articles tend to appear near the top, while new studies, reports and theses may sit far down the list even when they answer the question directly.
Bibliographic databases increasingly offer ranked displays, and some offer searching by meaning rather than by exact words, which Lesson 6 discusses. In a Boolean database, a relevance sort reorders a set that the query has already defined, as with PubMed Best Match. In a fully ranked system, the ranking also decides which documents the user will realistically see, because the list is long and the user stops reading.
2.4 Comparing the Two Models for Evidence Synthesis
Relevance ranking is well suited to finding a good answer quickly, which is what most web users want. A systematic or scoping review has a different goal: it aims to find all the eligible studies, or as close to all as resources allow, and to show others exactly how they were found. The table compares the three kinds of tool the Cedar Valley team will use against that goal.
| Feature | Bibliographic database (for example MEDLINE) | Google Scholar | |
|---|---|---|---|
| How matches are decided | Boolean logic, with an optional relevance sort | Undisclosed ranking with many signals | Undisclosed ranking that weights citations |
| Size of the result set | Exact count of every matching record | An estimated count | An estimated count |
| Can every result be viewed | Yes | No, the list ends long before the estimate | No, about 1,000 results at most |
| Bulk export of results | Yes, to a reference manager | No | No official bulk export |
| Controlled vocabulary | Yes (MeSH, Emtree, CINAHL headings) | No | No |
| Truncation and nested Boolean logic | Fully supported | Limited | No truncation and limited grouping |
| Coverage | Documented journal lists | The open web, undocumented | Broad scholarly coverage, undocumented |
| Role in a systematic search | Principal search | Supplementary, for grey literature | Supplementary, for grey literature and checking |
Published evaluations support the roles in the last row. Haddaway et al. (2015) found that Google Scholar shows at most the first 1,000 results of any search and recommended using it as a supplementary source, particularly for grey literature, with a defined number of results screened. Gusenbauer and Haddaway (2020) tested 28 academic search systems against criteria for systematic searching, including Boolean functionality, reproducibility and bulk export, and concluded that Google Scholar was unsuitable as a principal search system for systematic reviews, although useful as a supplementary one. Bramer et al. (2017) showed that no single database retrieved all the studies included in a large set of reviews, which is why the Cedar Valley team plans to search five databases rather than one.
The intern runs a quick test. In PubMed, a short tagged search for the loneliness concept and the older-adult concept returns a fixed set of a few hundred records, every one of which can be exported to the team's reference manager. In Google Scholar, similar words produce an estimate of tens of thousands of results, ordered with highly cited reviews at the top. The intern can page through only the first 1,000 and cannot export them in bulk. The librarian explains that the two numbers measure different things: the PubMed figure is the exact size of a defined set, and the Google Scholar figure is an estimate attached to a ranked list that no one can read in full. The team decides that the five bibliographic databases will provide the main search, and that Google Scholar will be used in Lesson 5 as a supplementary source with a fixed screening limit.
Reflection
A student is preparing a scoping review on community programs that reduce loneliness in older adults. In PubMed, a search that combines a loneliness concept and an older-adult concept with AND returns 412 records, and the count stays at 412 whether the results are sorted by Best Match or by most recent. In Google Scholar, a similar set of words produces the message "About 18,600 results", with highly cited reviews on the first page. Google Scholar displays at most about 1,000 results for a query and has no official bulk export. The student proposes to replace the PubMed search with the first 50 Google Scholar results, arguing that they look more relevant.
(a) Explain why the two numbers, 412 and about 18,600, differ in meaning. (b) Explain why sorting by Best Match did not change the PubMed count. (c) Give two reasons, based on how each tool retrieves records, why the proposal is unsuitable for the main search of a scoping review, and describe a suitable role for Google Scholar.
(a) The PubMed figure of 412 is the exact size of a set defined by Boolean logic: every record that satisfies the query, all of which can be viewed and exported. The Google Scholar figure is an estimate attached to a ranked list. It is not the size of a set that anyone can see, because only about the first 1,000 results are displayed.
(b) Best Match is a sort order. PubMed first retrieves the Boolean set and then decides only the order in which the 412 records are displayed, so the set and the count stay the same under any sort.
(c) First, the first 50 Google Scholar results are chosen by an undisclosed ranking that gives weight to citation counts, so recent studies, reports and evaluations from small programs are likely to sit lower in the list. Looking good at the top of the list is a sign of precision and says nothing about recall, which is what a scoping review needs. Second, the Google Scholar search cannot be fully reported or rerun: results cannot be exported in full, truncation is not supported, and the ranking changes between users and over time, so readers could not check the search or update it. A suitable role for Google Scholar is supplementary: after the database searches, the student could run a few documented queries, screen a fixed number of results set in advance, such as the first 100, and record the date, the query and the settings.
Minimum 20 characters required.
Question 1: In a toy collection, loneliness appears in records 1, 4 and 6, and older adults appears in records 1, 2, 3, 5 and 8. Which records does the query loneliness OR older adults return?
Question 2: What does sorting a PubMed result set by Best Match rather than by publication date change?
Question 3: In relevance ranking, under which condition does inverse document frequency give a query word more weight?
Question 4: Which statement about Google Scholar is supported by published evaluations and by Google Scholar's own documentation?
Recall and Precision: Why Reviews Favour Recall
Learning Objectives for this section
- Define recall, precision and the number needed to read, and calculate each from the results of a search.
- Explain why recall and precision usually move in opposite directions as a search is broadened or narrowed.
- Describe how recall is estimated in practice with a benchmark set of known relevant studies.
- Explain why systematic and scoping reviews favour recall, and identify the conditions under which a review team might accept lower recall.
3.1 Two Questions About Any Search
Sections 1 and 2 explained how databases find records. This section asks how well a particular search performs. Two questions capture most of what matters. The first is how many of the relevant records the search found. The second is how much of what the search returned is relevant. The measures that answer these questions, recall and precision, became the standard pair through the Cranfield experiments that Cyril Cleverdon led in England in the late 1950s and 1960s, which compared indexing systems by testing them against a set of documents whose relevance had been judged in advance. The same two measures are still used to evaluate search engines, search filters and the AI-assisted screening tools that Lesson 6 describes.
Any record in a database falls into one of four groups, depending on whether it is relevant to the review question and whether the search retrieved it. The figure shows these groups for a search of a single database.
The formulas
Recall = relevant records retrieved ÷ all relevant records in the database = a ÷ (a + c)
Precision = relevant records retrieved ÷ all records retrieved = a ÷ (a + b)
Number needed to read (NNR) = all records retrieved ÷ relevant records retrieved = 1 ÷ precision
Recall is the same quantity that a diagnostic test calls sensitivity: the proportion of true cases that the test detects. Precision resembles the positive predictive value of a test: the proportion of positive results that are true cases. Bachmann et al. (2002) proposed the number needed to read as a more intuitive form of precision, by analogy with the number needed to treat. It states how many records a reviewer must read, on average, to find one relevant record. The fourth group, d, is very large in any bibliographic database, so a measure built on it, such as specificity, is close to 100 percent for almost every search and says little about performance. For that reason, search evaluation uses recall and precision rather than sensitivity and specificity as a pair.
3.2 A Worked Example from the Cedar Valley Review
The following example is a teaching device. It assumes that the true number of relevant records is known, which never happens in practice; Section 3.4 explains how recall is estimated when it is unknown. Suppose that MEDLINE contains exactly 30 studies that meet the Cedar Valley eligibility criteria. The librarian drafts two searches. Search A is narrow: it looks for the words loneliness, older adults and intervention in titles only. Search B is broad: it combines the subject headings from Section 1 with free-text synonyms in titles and abstracts for each of the three concepts. All numbers are illustrative.
| Quantity | Search A (narrow) | Search B (broad) |
|---|---|---|
| Records retrieved (a + b) | 150 | 750 |
| Relevant records retrieved (a) | 18 | 27 |
| Relevant records missed (c) | 30 − 18 = 12 | 30 − 27 = 3 |
| Recall, a ÷ (a + c) | 18 ÷ 30 = 0.60 (60%) | 27 ÷ 30 = 0.90 (90%) |
| Precision, a ÷ (a + b) | 18 ÷ 150 = 0.12 (12%) | 27 ÷ 750 = 0.036 (3.6%) |
| Number needed to read | 150 ÷ 18 = 8.3 | 750 ÷ 27 = 27.8 |
| Screening time at 30 seconds per record, one reviewer | 75 minutes | 375 minutes (6.25 hours) |
| Screening time with two independent reviewers | 2.5 hours in total | 12.5 hours in total |
The comparison shows the trade-off in concrete terms. Search B costs five times as many records to screen and about ten more hours of reviewer time with dual screening. In exchange, it finds 9 more eligible studies and misses 3 instead of 12. Each additional study found by Search B costs about 67 extra records to screen (600 extra records divided by 9 extra studies). Search A misses 40 percent of the eligible studies, which would leave the evidence brief for Cedar Valley built on little more than half of the relevant evidence in MEDLINE.
The example also shows why specificity is uninformative. If MEDLINE holds about 30 million records, both searches leave out more than 99.99 percent of the irrelevant ones, so specificity cannot distinguish a good search from a poor one. Recall and precision, by contrast, differ clearly between the two searches.
3.3 Why Recall and Precision Move in Opposite Directions
The curves in the figure have a typical shape. The first terms a librarian adds find many relevant records, so recall climbs quickly. Each further synonym or heading finds fewer new relevant records and many more irrelevant ones, because rarer wordings are also used in unrelated contexts. Recall therefore levels off while precision keeps falling. No search reaches 100 percent recall at a workable precision in most health topics, and the practical question becomes how far along the curve the team can afford to go.
Most decisions in search design move a search along this curve. The table summarizes the usual direction of each effect. These are tendencies, and an individual change can behave differently in a particular database.
| Search choice | Usual effect on recall | Usual effect on precision |
|---|---|---|
| Add synonyms or variant spellings with OR | Increases | Decreases |
| Explode a subject heading | Increases | Decreases |
| Combine subject headings with free text | Increases | Decreases |
| Add another concept with AND | Decreases | Increases |
| Search titles only rather than titles and abstracts | Decreases | Increases |
| Restrict a heading to major topic | Decreases | Increases |
| Exclude a term with NOT | Decreases, sometimes sharply | Increases |
| Search additional databases | Increases across the review | Decreases, and adds duplicates |
Lesson 4 introduces validated search filters, such as the filters for randomized trials in the Cochrane Handbook, which are published in sensitivity-maximizing and precision-maximizing versions. Those two labels refer directly to the recall and precision trade-off described here.
3.4 Estimating Recall When the Total Is Unknown
In a real review, nobody knows how many relevant records a database contains, so recall cannot be calculated directly. Precision is easier: once the team has screened the retrieved records, precision is the number judged eligible divided by the number screened. Recall has to be estimated, and two methods are common.
The first method uses a benchmark set (also called a validation set or test set) of known relevant studies, collected independently of the search being tested. Sources include the included studies of earlier reviews, studies suggested by content experts, and studies the team already knows. The librarian checks that each benchmark study is present in the database and then runs the draft search. The estimated recall is the proportion of benchmark studies that the search retrieves. A missed benchmark study is also informative, because its title, abstract and headings show which terms the search lacks. The second method, relative recall (Sampson et al., 2006), pools all the eligible studies found by every search method in the review, including other databases and citation chasing, and treats that pool as the reference standard for each individual search.
The intern collects 20 eligible studies from the reference lists of two earlier reviews on interventions for loneliness in older adults and confirms that all 20 are indexed in MEDLINE. Draft Search A retrieves 12 of the 20, an estimated recall of 60 percent. Draft Search B retrieves 18 of the 20, an estimated recall of 90 percent. The librarian reads the two studies that Search B missed. One describes its intervention only as "befriending" and its outcome only as "social participation"; the other was indexed under broader headings without Loneliness, and its abstract uses the phrase "social disconnection". The team adds those terms to the free-text lists. Because adjusting a search to fit a known set can make it look better than it is, the librarian keeps ten further known studies aside as a final check, the test set that Lesson 4 describes.
A librarian tests a search for a different review against a benchmark set of 25 known eligible studies, all present in the database. The search retrieves 1,200 records, including 23 of the 25 benchmark studies. After screening all 1,200 records, the team judges 40 of them eligible. Calculate (a) the estimated recall, (b) the precision, (c) the number needed to read, and (d) the screening time for one reviewer at 30 seconds per record. Then state one change that would raise the estimated recall and its likely effect on precision.
(a) Estimated recall is 23 ÷ 25 = 0.92, or 92 percent. (b) Precision is 40 ÷ 1,200 = 0.033, or about 3.3 percent. (c) The number needed to read is 1,200 ÷ 40 = 30, so the team reads 30 records for each eligible one. (d) At 30 seconds per record, one reviewer needs 1,200 × 0.5 = 600 minutes, or 10 hours. To raise estimated recall, the librarian should read the two missed benchmark studies and add the terms or headings they use with OR; this would retrieve more records and would usually lower precision.
3.5 Why Reviews Favour Recall
Chapter 4 of the Cochrane Handbook (Lefebvre et al., 2023) advises that searches for systematic reviews should aim for high sensitivity, that is, high recall, while accepting that precision will be low. The reasons can be stated in terms of costs and of bias, and the cards below set them out.
Precision still matters, because reviewer time is finite. The Cedar Valley team has twelve weeks, and screening thousands of records with two reviewers competes with charting, appraisal, the environmental scan and writing. Its plan therefore aims for high recall in the main database searches, and it manages the workload in other ways: removing duplicates before screening (Lesson 7), using a screening platform, piloting eligibility criteria so that decisions are quick, and possibly using active-learning screening tools (Lesson 6). The team will consider lowering recall only after these options, and it will report any limit it sets.
A common misunderstanding
Students sometimes judge a search by whether its first page looks relevant. That is a judgement about precision at the top of a ranked list, which is how web search engines are designed to perform. A systematic search is judged by recall, which cannot be seen on any single page of results. A search whose first page is full of relevant records can still miss a large share of the eligible studies, and a broad search with a mixed first page can be the better search for a review.
Reflection
Two draft searches for a review are tested in MEDLINE against a benchmark set of 25 known eligible studies, all of which are present in MEDLINE. Search X retrieves 2,000 records, includes 24 of the 25 benchmark studies, and on screening yields 60 eligible studies. Search Y retrieves 600 records, includes 19 of the 25 benchmark studies, and on screening yields 48 eligible studies. The team has two reviewers who each screen every record independently at about 30 seconds per record, and it must deliver a scoping review within twelve weeks. Recall is relevant records retrieved divided by all relevant records; precision is relevant records retrieved divided by all records retrieved; the number needed to read is records retrieved divided by relevant records retrieved.
(a) Calculate the estimated recall, the precision and the number needed to read for each search, and the total dual-screening time for each. (b) Recommend one search and justify your choice with reference to why reviews favour recall. (c) Describe one way to improve the search you did not choose, or to reduce the workload of the one you chose, without simply accepting lower recall.
(a) Search X has an estimated recall of 24 ÷ 25 = 96 percent, a precision of 60 ÷ 2,000 = 3 percent, and a number needed to read of 2,000 ÷ 60 = 33.3. Dual screening takes 2,000 records × 0.5 minutes × 2 reviewers = 2,000 minutes, about 33.3 hours. Search Y has an estimated recall of 19 ÷ 25 = 76 percent, a precision of 48 ÷ 600 = 8 percent, and a number needed to read of 600 ÷ 48 = 12.5. Dual screening takes 600 × 0.5 × 2 = 600 minutes, or 10 hours.
(b) I would choose Search X. Search Y misses about one in four eligible studies, and those it misses may differ systematically from those it finds, for example by using different terms or coming from another discipline, which could bias the scoping review's account of the evidence. The extra 23 hours of screening are a known and manageable cost over twelve weeks, whereas the missed studies would be invisible.
(c) To make Search X manageable, the team could remove duplicates across databases before screening, pilot the eligibility criteria on 50 records so that decisions are quick, and use a screening platform. Alternatively, to improve Search Y, the librarian could read the six benchmark studies it missed, identify the headings and free-text terms they use, and add those terms with OR, then retest against benchmark studies held back for that purpose.
Minimum 20 characters required.
Question 1: A search retrieves 400 records, of which 20 are relevant. The database contains 25 relevant records in total. What are the recall and precision of the search?
Question 2: A search has a precision of 5 percent. What is its number needed to read?
Question 3: Why do systematic and scoping reviews usually favour recall over precision?
Question 4: A librarian has 20 known eligible studies, all present in MEDLINE, and a draft search retrieves 17 of them. What does this tell the team?
Personalization, Ranking Opacity and the Reproducibility of Web Search
Learning Objectives for this section
- Define reproducibility and transparency as they apply to a literature search, and explain why the two differ for web searches.
- Identify the main sources of variation in web search results, including personalization, location, language, index changes and updates to ranking systems.
- Compare the reproducibility of bibliographic database searches, Google searches and Google Scholar searches.
- Describe the minimum information a reviewer should record about a web search so that readers can judge it.
4.1 What Reproducibility Means for a Search
A search is reproducible when another person, following the documented method, can run it again and obtain the same results, or can explain any differences. Reproducibility matters in evidence synthesis for three reasons. It allows peer reviewers and readers to check that a review searched as it claims. It allows the same team, or a different one, to update the review later by rerunning the search for newer records. It also allows a reader to judge the risk that relevant studies were missed. PRISMA-S, the reporting guideline for literature searches developed by Rethlefsen et al. (2021), exists largely to make searches reproducible, and Lesson 4 teaches it in detail.
A related idea is transparency: the extent to which a report shows exactly what was done, even when the result could not be obtained again. A search can be fully transparent without being reproducible. This distinction is central to web searching, because, as this section shows, a reviewer can report a Google search completely and still be unable to guarantee that anyone else will see the same results.
Bibliographic database searches come close to reproducibility, although even they fall short of it. Databases add new records every day, sometimes including older articles added later. Thesauri change each year, as the Social Prescribing heading did in 2025. Platforms change their search engines and syntax. A rerun of a MEDLINE search a year later, limited to the original date range, will usually return a very similar set, and the differences can be traced to these known causes. Web search engines behave very differently, for the reasons set out below.
4.2 Personalization and Context
Personalization means that a search engine adjusts results to the person searching. Google's public documentation states that its systems use the user's location, search history and settings, that the language of the query shapes the language of the results, and that pages a user has visited often may be moved higher in that user's results. It also states that users can see whether personal information affected a result and can turn personalization off. Hannak et al. (2013) measured personalization in Google web search by comparing the results that real accounts and controlled test accounts received for the same queries. They found measurable personalization, with being signed in to an account and geographic location as the main sources of difference.
The term "filter bubble", popularized by Eli Pariser (2011), describes the concern that personalization shows people mainly what fits their past behaviour. Google states that its systems are not designed to create filter bubbles. Studies that have measured personalization in web search have generally found it present but modest in average size. For evidence synthesis, the size of the effect matters less than its existence: any personalization means that two reviewers running the same query can see different results, and neither can tell from the results page alone which records were affected.
Some sources of variation remain even when personalization is turned off. Results depend on the approximate location that the search engine infers from the network connection, which a private browsing window does not hide. They depend on the country version of the search engine, on the interface language, and on settings such as filters for explicit content. They can also depend on the device, because page usability is one of the factors Google describes. A reviewer can control some of these factors and record the rest, but cannot remove them.
4.3 Ranking Opacity and Change Over Time
Ranking opacity means that the method used to order results is not disclosed in enough detail for an outsider to predict or check it. Section 2 showed that Google publishes only general categories of ranking factors and that Google Scholar describes its ranking in a single paragraph. Commercial search engines have reasons to keep the details private, including competition and the need to resist manipulation of rankings by website owners. For a reviewer, the consequence is that no one can explain why a given page appeared in position three rather than position thirty, or whether it would appear at all if the search were run again.
Opacity combines with change over time. Search engines revise their ranking systems continually, and the web index is rebuilt as pages change. Google Scholar's documented use of citation counts in ranking has a further consequence: as articles gain citations, their positions change, so the order of results for the same query drifts from month to month even if no new articles are published. Because Google Scholar shows at most about the first 1,000 results, articles near that boundary can move in and out of view.
Other features of the results page make the problem harder. The counts that Google and Google Scholar display ("about" so many results) are estimates that can change between pages of the same search and should not be reported as the size of a result set. Results pages may include advertisements labelled as sponsored, highlighted answer boxes drawn from a single page, and, since 2024, AI-generated summaries above the results for some queries. Each of these elements is produced by its own undisclosed process, and each can differ between users and days.
| Question about reproducibility | Bibliographic database | Google Scholar | |
|---|---|---|---|
| Can the exact query be recorded and reported? | Yes | Yes | Yes |
| Can the full set of results be saved? | Yes, by export | Only the pages actually viewed | Only the pages viewed, up to about 1,000 results |
| Will a rerun return the same records? | Largely, within the original date limits | Unlikely | Unlikely |
| Can a reader see why a record was retrieved? | Yes, from the strategy | No | No |
| Is the searcher's context a factor? | No | Yes, location, language, account and settings | Possibly; language settings can change results, and other personalization is undocumented |
| Is the collection searched documented? | Yes, with published journal lists | No | No |
Gusenbauer and Haddaway (2020) included reproducibility among the criteria they used to judge academic search systems, and they judged that Google Scholar did not meet their requirements for reproducible searching. This finding, together with the cap on viewable results and the lack of bulk export, is why reviews use Google Scholar and Google as supplementary sources. Supplementary searching is still valuable. Much of the information the Cedar Valley environmental scan needs, such as program descriptions, funding announcements and evaluation reports from community organisations, exists only on the open web and will never appear in MEDLINE.
On the same afternoon, the evidence officer in Cedar City and the intern working from campus in Burnaby type the same query into Google: social prescribing older adults British Columbia. The officer is signed out; the intern is signed in to an account that has been used for weeks of searching on social prescribing. Their first pages share only three of ten results, and the shared results appear in different positions. The officer's page includes local news and a seniors' centre page, while the intern's page includes a review article and a national charity report. When the intern repeats the search two weeks later, one program page on the first list has moved to a new address and a new news story has appeared. All details here are illustrative. The team concludes that it cannot make its web searches reproducible, and it decides instead to make them transparent: it will control the settings it can, record the rest, and save copies of everything it screens.
4.4 Making Web Searches Transparent
Because a web search cannot be reproduced exactly, the reviewer's aim shifts to recording enough for a reader to understand what was searched and to judge what might have been missed. PRISMA-S asks reviewers to report web searches and other supplementary methods alongside database searches, and Lesson 5 provides a full documentation template for grey literature and website searching. The items below give the minimum record that this lesson's mechanisms imply.
Note the search engine and its country version, the date of the search, and the query exactly as typed, including quotation marks and operators. A query that is reported only in summary cannot be repeated even approximately.
Search while signed out, in a private browsing window, with the interface language and region set deliberately, and record these settings and the approximate location of the searcher. These steps reduce personalization from the account and browser history, and the record allows a reader to see what was left uncontrolled.
Decide in advance how many results to screen, for example the first 100 results of each query, and report the rule. Haddaway et al. (2015) recommended a defined limit of this kind for Google Scholar, because no one can screen an estimated count of tens of thousands. A rule set in advance prevents the reviewer from stopping when results happen to look poor or continuing when they happen to look good.
Save the results pages screened, as files or screenshots, and save copies of included documents, because web pages change or disappear. Record each included document's address and the date it was accessed. These copies allow the team, and anyone checking its work, to see the results as they were.
If a results count is reported, describe it as the estimate the search engine displayed, and report separately the number of results actually screened. The number screened is the figure that belongs in the PRISMA flow diagram, which Lesson 7 covers.
With a friend or colleague, agree on one query about a health topic that interests you. At the same time, each of you runs it in Google, one signed in and one signed out in a private window, and each records the first ten results. Count how many results appear on both lists, and note how many of the shared results appear in the same position. Then write two sentences explaining which differences you could have controlled and which you could only record.
4.5 From Search Engines to AI Tools
The mechanisms in this lesson carry forward to the artificial intelligence tools that Lesson 6 examines. AI search engines and research assistants usually begin with a retrieval step that resembles relevance ranking, so they inherit the opacity and the dependence on an undisclosed index described here. They then add a step that generates text from what was retrieved, and that step can produce different wording, different sources and, at times, citations to works that do not exist, even when the same question is asked twice. The same distinction between reproducibility and transparency applies: a reviewer who uses these tools must record what was asked, when, with which tool and version, and must verify each source against the original publication.
Summary of the lesson's argument
Bibliographic databases describe records with controlled vocabularies and retrieve them with Boolean logic, which makes their searches exhaustive, transparent and close to reproducible. Web search engines rank documents with undisclosed methods that depend on the searcher, the place and the day, which makes them useful for finding grey literature and unsuitable as the principal search for a review. Recall and precision give a common language for judging any of these searches, and reviews favour recall because missed studies are invisible and can bias the findings.
Groundwork before the full search
Before it builds a full search, a review team usually prepares three records. First, for each main concept in its PCC question, it looks up the MeSH heading in the MeSH Browser and the matching CINAHL Subject Heading, and records each heading's scope note, at least two entry terms, its narrower headings, whether it will be exploded, and the year it was introduced, together with a list of free-text synonyms for each concept. Second, it assembles a benchmark set of three to five studies already known to meet the eligibility criteria, notes where each one was found, and confirms that each is in PubMed. Third, it runs one test search in Google Scholar while signed out, records the date, the exact query, the settings, the estimated count and the first five results, and reruns the search a few days later to note any differences. Lesson 4 uses these records to build and test the full search.
Reflection
Two members of a review team in British Columbia run the Google query community connector program seniors loneliness BC on the same afternoon. One is signed in on a laptop in Vancouver. The other is signed out, using a private browsing window on a phone in Kamloops. Their first ten results share four pages, and the shared pages appear in different positions. A week later, one of the shared pages, a program description, returns an error because it has moved. Google's public documentation states that its ranking considers the meaning of the query, the relevance and quality of content, the usability of pages, and the user's context and settings, including location, search history and language, and it does not publish the weights of these factors. The team must report its web searching in the review.
(a) Identify three distinct sources of variation that could explain the differences, and for each state whether the team could have controlled it or could only record it. (b) Write a documentation entry for one of these searches that would allow a reader to judge it. (c) Explain in two or three sentences why this search can be transparent even though it cannot be reproduced.
(a) The first source is personalization from the signed-in account, including search history; the team could have controlled this by having both members search while signed out in private windows. The second is location: Vancouver and Kamloops may receive different local results for a query that mentions British Columbia, and the team could only record the approximate location, because a private window does not hide it. The third is change over time in the web index and in Google's ranking systems, shown by the program page that moved; the team could only record the date and save copies of what it screened. Device differences between laptop and phone are a further possible source that could be recorded.
(b) Search engine: Google, Canadian version (google.ca). Date: 14 October 2026, 2:10 pm. Query, exactly as typed: community connector program seniors loneliness BC. Settings: signed out, private window, interface language English, region Canada, SafeSearch default. Searcher location: Kamloops, British Columbia, mobile phone. Stopping rule: first 50 results, set in advance. Displayed estimate recorded separately from the 50 results screened. Results pages saved as PDF files; addresses and access dates recorded for the 3 documents retained.
(c) Transparency means that a reader can see exactly what was searched, when, how and with what limits. Because the ranking and the index change in ways the team cannot see, nobody can guarantee the same results on a rerun, but the record allows a reader to judge what the search could and could not have found.
Minimum 20 characters required.
Question 1: Which source of variation in Google results remains when a reviewer searches while signed out in a private browsing window?
Question 2: What is the difference between a reproducible search and a transparent search?
Question 3: Why does the order of Google Scholar results for the same query tend to change from month to month even when no new articles are added?
Question 4: A reviewer reports only: "Google Scholar search, about 24,300 results." What is the main problem with this report?
Final Assessment
Bringing It All Together
This lesson has explained why the same question returns different results in different search tools. Bibliographic databases describe each record with subject headings drawn from a controlled vocabulary, such as MeSH for MEDLINE, Emtree for Embase and CINAHL Subject Headings for CINAHL. Headings gather varied wording under one term and allow a search to include narrower concepts through explosion, but they have gaps: new concepts such as social prescribing receive headings late, some records are never indexed, and indexing varies. A sensitive search therefore combines headings with free-text terms for every concept.
Databases retrieve records with Boolean logic applied to an inverted index, which makes their searches exhaustive, transparent and close to reproducible. Google and Google Scholar rank documents with methods that are documented only in general terms, display only part of their results, and change with the searcher, the place and the day. Recall and precision provide a common language for judging any search. Reviews favour recall because an irrelevant record costs seconds to exclude, while a missed study is invisible and may bias the findings.
For the fictional Cedar Valley review, these principles produce a plan with five databases searched with both headings and free text, a benchmark set to estimate recall, and web searching used as a documented supplement. Lesson 4 turns that plan into a full search string.
Key Takeaways from this lesson
- A bibliographic database record contains fields written by authors and fields added by the database during indexing, and a search can look in either kind.
- A controlled vocabulary assigns one preferred heading to each concept, with entry terms, a hierarchy of broader and narrower headings, and scope notes that define its use.
- MeSH, Emtree and CINAHL Subject Headings index different databases, and their headings and hierarchies differ, so each heading must be checked when a search is translated.
- Exploding a heading retrieves records indexed with it and with every narrower heading, which matters because indexers assign the most specific heading available.
- Free-text terms are needed alongside headings because new concepts receive headings late, some records are never indexed, and indexing varies between records.
- Boolean retrieval returns every record that satisfies the logic of a query, which makes it exhaustive, transparent and repeatable, while relevance sorting only reorders that set.
- Relevance ranking orders documents by an estimated match to the query, and Google and Google Scholar describe their ranking factors only in general terms.
- Recall is the proportion of relevant records a search finds, precision is the proportion of retrieved records that are relevant, and the number needed to read equals one divided by precision.
- Reviews favour recall because missed studies are invisible and can bias findings, and recall is estimated in practice with a benchmark set of known studies.
- Web searches cannot be fully reproduced because of personalization, location, index changes and undisclosed ranking updates, so reviewers aim to make them transparent instead.
Core Concepts Reviewed
Section 1: indexing, controlled vocabularies and thesauri, MeSH, Emtree and CINAHL Subject Headings, entry terms, explosion, major topics, free-text searching and automatic term mapping.
Section 2: the inverted index, Boolean operators and exhaustive retrieval, relevance ranking with term frequency and inverse document frequency, link analysis, and the limits of Google and Google Scholar for systematic searching.
Section 3: recall, precision and the number needed to read, the trade-off between recall and precision, benchmark sets and relative recall, and the reasons reviews favour recall.
Section 4: reproducibility and transparency, personalization and context, ranking opacity, change over time in web indexes and rankings, and the minimum record for a web search.
The final reflection asks you to explain the lesson's main ideas to a decision-maker in the Cedar Valley case.
Reflection
You are the student intern on the fictional Cedar Valley evidence team, which is preparing a rapid scoping review on community-based interventions for loneliness and social isolation in adults aged 65 and older. The planning director, who is not a researcher, asks why the team needs about five weeks for searching and screening when Google finds information in seconds. Use these facts. The team plans to search MEDLINE, Embase, CINAHL, APA PsycInfo and Web of Science, using subject headings (such as MeSH in MEDLINE) together with free-text words. In a pilot, a broad MEDLINE search retrieved 750 records and found an estimated 90 percent of known eligible studies, while a narrow search retrieved 150 records and found about 60 percent. Google and Google Scholar rank results with methods that are not published, show results that vary by user, place and day, and cannot be exported in full; Google Scholar displays at most about 1,000 results.
Write a reply of 200 to 300 words that explains (a) how bibliographic databases and web search engines find information differently, (b) what recall and precision mean and why the team favours recall, and (c) what role web searching will play and how it will be documented.
Thank you for the question. Google and the research databases we use work in different ways. Databases such as MEDLINE describe each article with standard subject headings, so a study that says "lonely seniors" and one that says "isolated older adults" can both be found under the same headings. They return every record that matches our search logic, give an exact count, and let us save the whole set. Google and Google Scholar rank pages using methods they do not publish, and the order changes with the person searching, their location and the day. Only part of the list can be viewed, and none of it can be exported in full.
We judge a search by two measures. Recall is the share of all relevant studies that the search finds; precision is the share of what it returns that is relevant. In our pilot, a narrow search gave us 150 records to read but found only about 60 percent of the studies we know are relevant. A broad search gave us 750 records and found about 90 percent. We favour recall because a missed study is invisible: we would never know it was absent, and if missed studies differ from the ones we find, the brief could mislead your planning. Reading extra records takes time but carries no such risk.
Web searching still matters, because many program descriptions and local evaluations exist only online. We will use Google and Google Scholar as supplementary sources, screen a fixed number of results set in advance, and record the date, exact wording and settings of each search, saving copies of what we screen, so that you and others can see exactly what we did.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: A team's MEDLINE search uses only free-text terms in titles and abstracts. Which relevant record is it most likely to miss?
Question 2: Which pairing correctly matches each thesaurus to the database it indexes?
Question 3: Search A in the Cedar Valley example looked for its words in titles only. If the librarian extends it to titles and abstracts, which change is most likely?
Question 4: Which feature makes Boolean retrieval in a bibliographic database suitable for the main search of a systematic review?
Question 5: A search retrieves 2,000 records, of which 50 are eligible. What are its precision and number needed to read?
Question 6: Why does an exploded MeSH heading usually increase recall?
Question 7: According to Google's public documentation, which statement about Google's ranking is accurate?
Question 8: The Cedar Valley intern checks a draft MEDLINE search against 20 known eligible studies and finds that 2 are missed. What is the most useful next step?
Question 9: Why is specificity rarely used to evaluate a search of a bibliographic database?
Question 10: A rapid review team decides to search two databases instead of five. How should this decision be handled?
Question 11: Which pair of actions best improves the transparency of a Google search for grey literature?
Question 12: Why can a PubMed search that uses only MeSH headings miss records that a free-text search finds?
Question 13: In a toy collection, lonely appears in record 3, loneliness in records 1, 4 and 6, and older adults in records 1, 2, 3, 5 and 8. What does (loneliness OR lonely) AND older adults return?
Question 14: Google Scholar states that its ranking weighs the full text, the source, the authors and citations. What consequence might this have for a scoping review of a new type of program?
Question 15: Which combination of methods fits the Cedar Valley team's aim of high recall within its twelve-week timeline?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, tools and people appear in this lesson on how databases and search engines find information.