HSCI 241 · Lesson 3

How Databases and Search Engines Find Information

Finding & Synthesizing Health Evidence

Learning objectives for this lesson:

  • Describe how bibliographic databases index records, and distinguish the subject headings added by a database from the fields written by authors.
  • Explain how the controlled vocabularies MeSH, Emtree and CINAHL Subject Headings work, including entry terms, hierarchies, explosion and major topics.
  • Explain why a sensitive search combines controlled vocabulary with free-text terms for every concept.
  • Apply the Boolean operators AND, OR and NOT to sets of records, and explain how an inverted index supports exhaustive retrieval.
  • Explain in general terms how relevance ranking orders results in Google and Google Scholar, and distinguish what is publicly documented from what is undisclosed.
  • Calculate recall, precision and the number needed to read, and estimate recall with a benchmark set of known studies.
  • Explain why systematic and scoping reviews favour recall, and identify when a team might knowingly accept lower recall.
  • Identify the sources of variation that limit the reproducibility of web searches, and record a web search so that it is transparent.
  • Describe the groundwork a review team prepares before building a full search: subject headings and free-text terms for each concept, a benchmark set of known studies and a documented test web search.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University, drawing on the Cochrane Handbook for Systematic Reviews of Interventions and the JBI Manual for Evidence Synthesis.

Lesson 3 · HSCI 241

How Databases and Search Engines Find Information

This short walkthrough orients you before you work through the lesson at your own pace.

Finding & Synthesizing Health Evidence
Why this lesson

Two ways of finding information

Bibliographic databases

They describe records with subject headings and return every record that matches the logic of a search.

Web search engines

They rank pages with undisclosed methods and show the top of a long list that no one can read in full.

The running case

The fictional Cedar Valley evidence review

210,000residents are served by the health authority.
46,000residents are aged 65 and older.
12weeks are available to deliver the evidence brief.

The evidence team is an evidence officer, a university librarian and a student intern.

The plan

Four sections

1. Indexing

Controlled vocabularies such as MeSH work alongside free text.

2. Retrieval

Boolean logic is compared with relevance ranking.

3. Recall

Recall and precision are used to judge a search.

4. Reproducibility

Web searches vary by person, place and day.

Before the full search

Groundwork for a search

  • A list of subject headings and free-text terms for each concept in the question.
  • A benchmark set of three to five known eligible studies.
  • One documented test search in Google Scholar.

These records are the starting point for the database search in Lesson 4 and the web search plan in Lesson 5.

Section 1 of 5

Indexing, Controlled Vocabularies and Free Text

⏱ Estimated reading time: 35 minutes
Section 1 of 5

Indexing, Controlled Vocabularies and Free Text

Databases describe each record with standard subject terms, and searches combine those terms with the words authors wrote.

About 35 minutes
Inside a record

Two kinds of fields

Written by the authors

The title, abstract and author keywords are searched with free-text terms.

Added by the database

Subject headings and the publication type are searched with controlled vocabulary.

The authors wrote "lonely older adults", and the indexing assigned the headings Loneliness and Aged.

How a thesaurus works

Preferred terms, entry terms and explosion

Social Isolation Loneliness Ostracism SocialAlienation SocialDeprivation Exploding Social Isolation retrieves all five headings.

Indexers assign the most specific heading, so an unexploded broad heading misses records indexed beneath it.

Three thesauri

Each database has its own vocabulary

MeSH

The National Library of Medicine uses it to index MEDLINE, and it holds more than 30,000 descriptors.

Emtree

Elsevier uses it to index Embase, with detailed terms for drugs and devices.

CINAHL Subject Headings

They index CINAHL and add nursing and allied health terms to a MeSH-like structure.

APA PsycInfo uses the APA Thesaurus, and Web of Science has no subject-heading thesaurus.

Why free text is still needed

The gaps in indexing

Late headings

Social Prescribing became a MeSH heading only in 2025.

Unindexed records

Some PubMed records are never indexed for MEDLINE.

Varied indexing

Similar articles can receive different headings.

No heading at all

Specific intervention names such as befriending lack their own heading.

Cedar Valley (fictional)

The intern's first lesson in indexing

  • The intern's first PubMed search looked good on the first page.
  • The librarian showed that Social Prescribing entered MeSH in 2025.
  • An exploded Social Isolation search also finds records indexed with Loneliness.
  • The intern left with headings and free-text terms for each concept.
Carry forward

From vocabulary to retrieval

A sensitive search uses subject headings and free text for every concept.

Section 2 explains how databases and search engines turn those terms into results.

Learning Objectives for this section

  • Describe what a bibliographic database record contains and distinguish the fields written by authors from the fields added by the database.
  • Explain how a controlled vocabulary works, including preferred terms, entry terms, the hierarchy of broader and narrower headings, explosion and major topics.
  • Compare MeSH, Emtree and CINAHL Subject Headings and identify the database that each one indexes.
  • Explain why a sensitive search combines controlled vocabulary with free-text terms for each concept.

1.1 Why the Same Question Returns Different Results

A student who types the same words into PubMed, Embase, CINAHL, Google and Google Scholar will receive five different sets of results. The counts differ, the first records differ, and some studies appear in one tool and are missing from another. These differences arise from how each tool stores information and how it decides which records match a query. A reviewer who understands those mechanisms can build a search that finds what it should, explain the search to others, and judge which tools can support a systematic search and which cannot.

This section explains indexing and contrasts searching subject terms with searching the words that authors wrote. Section 2 compares Boolean retrieval in databases with relevance ranking in Google and Google Scholar, Section 3 introduces recall and precision, and Section 4 examines the reproducibility of web search. Writing and translating a full search string belongs to Lesson 4, and grey literature and targeted web searching belong to Lesson 5.

Case: the Cedar Valley evidence review (fictional)

The fictional Cedar Valley Health Authority in British Columbia serves about 210,000 residents, including about 46,000 adults aged 65 and older. Before launching a community connector (social prescribing) program, its planning team has asked a small evidence team for a rapid scoping review and environmental scan within twelve weeks. The team is an evidence officer, a university librarian and a student intern. In Lesson 2 the team wrote its question in PCC form: the population is adults aged 65 and older, the concept is community-based interventions evaluated for loneliness or social isolation, and the context is any community setting. In this lesson the librarian explains to the intern how the databases they will search find information.

1.2 What a Database Record Contains

A bibliographic database is an organized collection of records, each describing one published item such as a journal article. MEDLINE, the database of the United States National Library of Medicine (NLM), is the best-known example in the health sciences, and PubMed is the free interface through which most people search it. Embase, CINAHL, APA PsycInfo and Web of Science are other major databases that the Cedar Valley team will search.

Each record contains two kinds of fields. The first kind comes from the authors and the publisher: the title, the abstract, the author names, the journal, and any keywords the authors chose. The second kind is added by the database: the subject headings that describe what the article is about, the publication type (for example, randomized controlled trial or review), the language, and other descriptive codes. Adding this second kind of information is called indexing. A search can look in either kind of field, and the results differ depending on which kind it uses.

One record in a bibliographic database (illustrative) Fields written by the authors Title A volunteer telephone befriending program for lonely older adults Abstract Seniors living alone received weekly calls from trained volunteers ... Author keywords befriending; isolation; ageing Fields added by the database Subject headings (MeSH) Loneliness Aged Social Support Telephone Volunteers Publication type Randomized Controlled Trial A free-text search looks here A subject-heading search looks here
An illustrative record. The authors wrote "lonely", "seniors" and "older adults", while the indexing assigned the standard headings Loneliness and Aged, so a search on either kind of field can find the record by a different route.

The example shows why indexing matters. The authors described their participants as "older adults" and "seniors" and their outcome as being "lonely". Another team writing about the same topic might say "elderly people" and "social disconnection". Indexing replaces this variety with a single standard heading for each concept, so that every article about loneliness in people aged 65 and older carries the same headings regardless of the words its authors chose.

1.3 How a Controlled Vocabulary Works

A controlled vocabulary is a fixed list of approved terms used to describe the subject of each record. When the list is organized with relationships among the terms, it is called a thesaurus. Health databases use thesauri with five main features, and each one affects how a search behaves.

First, each concept has one preferred term, also called a subject heading or descriptor, such as Loneliness. Second, each preferred term has entry terms, which are synonyms and variant spellings that point to it. In MeSH, the entry terms for Social Isolation include "Social Exclusion", and the entry terms for Aged include "Elderly". Third, the headings are arranged in a hierarchy of broader and narrower terms, which MeSH calls its tree structure. Fourth, each heading has a scope note that defines how indexers use it; the scope note for Aged, for example, defines it as a person 65 through 79 years of age and directs indexers to the narrower heading "Aged, 80 and over" for older people. Fifth, many databases allow subheadings (MeSH calls them qualifiers) that narrow a heading to an aspect such as psychology, therapy or epidemiology.

The hierarchy makes possible an operation called explode. Exploding a heading retrieves records indexed with that heading and with every narrower heading beneath it. Indexers are instructed to assign the most specific heading that fits an article, so an article about loneliness is indexed with Loneliness rather than with the broader Social Isolation. A search for Social Isolation that is not exploded will therefore miss articles indexed only with its narrower headings. A related option, major topic, restricts a search to records in which the heading describes a main point of the article rather than a minor one. Exploding increases the number of records retrieved, and restricting to major topic decreases it; Section 3 explains how these choices trade off.

Sociological Factors Social Isolation Loneliness Ostracism Social Alienation Social Deprivation Exploding Social Isolation retrieves all five headings inside this boundary A newer heading Social Support Social Prescribing Added to MeSH in 2025. Older records were indexed without it.
Part of the MeSH hierarchy relevant to the Cedar Valley question, as listed by the National Library of Medicine. Loneliness also appears in a second branch under Emotions, which is not shown.

1.4 Three Thesauri: MeSH, Emtree and CINAHL Subject Headings

Each major health database has its own thesaurus, and the headings, hierarchies and search syntax differ from one database to the next. A search built with MeSH headings therefore has to be translated, heading by heading, before it can run in Embase or CINAHL. Lesson 4 teaches that translation. The cards below summarize the thesauri the Cedar Valley team will meet.

MeSH (Medical Subject Headings)Click to explore
EmtreeClick to explore
CINAHL Subject HeadingsClick to explore
APA Thesaurus and Web of ScienceClick to explore

Because each thesaurus is maintained separately, the same concept can carry a different heading, a different position in the hierarchy, or no heading at all, depending on the database. The tabs below show how a heading for the loneliness concept is written in four common interfaces. The aim at this stage is to recognize that each interface writes headings in its own way; Lesson 4 returns to the details.

In PubMed, a MeSH heading is written with the field tag [Mesh], as in "Social Isolation"[Mesh]. PubMed explodes MeSH headings by default, so this search includes Loneliness and the other narrower headings. Adding [Mesh:NoExp] turns explosion off, and the tag [Majr] restricts the search to records in which the heading is a major topic.

On the Ovid platform, a heading ends with a forward slash, as in Loneliness/. Explosion is requested explicitly by placing exp in front of the heading, as in exp Social Isolation/. Published systematic reviews often report MEDLINE searches in this form.

On Elsevier's Embase.com platform, an Emtree term is written in single quotation marks followed by a slash and an option, as in 'loneliness'/exp for an exploded search or 'loneliness'/mj for a search restricted to records in which the term is a major focus.

In CINAHL on the EBSCOhost platform, a heading is written with the code MH, as in (MH "Loneliness"). A plus sign inside the quotation marks explodes the heading, as in (MH "Social Isolation+"), and the code MM in place of MH restricts the search to major concepts.

1.5 Free-Text Searching and Why It Is Still Needed

A free-text search (also called a keyword or text-word search) looks for words in the fields written by the authors, usually the title, the abstract and the author keywords. It finds a record only when the record contains the exact words searched, or their variants when the searcher uses truncation, which Lesson 4 teaches. A free-text search for "loneliness" would miss the illustrative record shown in Section 1.2, whose title says "lonely", unless the searcher also included "lonely" or used truncation.

Free text is still needed because indexing has gaps, and each gap causes a subject-heading search to miss relevant records. The items below describe the main gaps, and Web of Science and web search engines, which have no thesaurus, rely on free text entirely.

New concepts arrive before their headingsv

A thesaurus adds a heading only after a concept appears often enough in the literature. MeSH added Social Prescribing as a descriptor in 2025, although studies of social prescribing had been published for many years before that. Records indexed before a new heading exists were indexed with other headings and generally do not carry the new one, so a search using only the new heading would miss most of the earlier studies. Free-text terms such as "social prescribing", "community connector" and "link worker" find those older records.

Some records are never indexedv

PubMed contains records that are not indexed for MEDLINE, including some articles supplied by publishers and articles available through PubMed Central from journals outside MEDLINE. These records have no MeSH headings at all, so only a free-text search can find them. Other databases have similar groups of records that are unindexed or only partly indexed.

Indexing varies between recordsv

Whether indexing is done by people, by software or by both, two similar articles can receive different headings. An article about a walking group might be indexed under Exercise and Social Support without any loneliness heading, even though loneliness was its primary outcome. Free-text terms catch articles whose headings missed the concept the reviewer is interested in.

The concept may be too specific for any headingv

Many intervention names, such as "befriending", "men's sheds" or "intergenerational programs", have no heading of their own. They are indexed under broader headings that also cover many unrelated studies. Searching the specific phrase in free text is often the only practical way to find these interventions.

PubMed also performs a step called automatic term mapping. When a user types words without field tags, PubMed tries to match them to MeSH headings, journal names and author names, and then combines the matched heading with a search of the words in all fields. Because an untagged search therefore does more than was typed, systematic searchers write each heading and free-text term with explicit field tags, and they check PubMed's "Details" display to see how a query was translated.

FeatureControlled vocabulary searchFree-text search
What it searchesSubject headings assigned by the databaseWords in the title, abstract and keyword fields
Handles synonyms and spelling variantsYes, through entry terms that point to one headingOnly for the variants the searcher lists or truncates
Includes narrower conceptsYes, when the heading is explodedOnly when the narrower terms are listed
Finds new or unindexed recordsNo, because these records lack headingsYes, provided the words appear in the record
Finds concepts without a headingNo, the concept falls under broader headingsYes, by searching the specific phrase
Transfers between databasesNo, each database has its own thesaurusMostly, with changes to syntax and field tags

1.6 Combining the Two Approaches

Because each approach covers the other's gaps, the standard practice in systematic searching is to search every main concept with both its subject headings and a set of free-text terms, joined with OR, and then to join the concepts with AND. Chapter 4 of the Cochrane Handbook for Systematic Reviews of Interventions (Lefebvre et al., 2023) recommends this approach for searches that aim to be comprehensive. Section 2 explains exactly what OR and AND do to sets of records, and Lesson 4 shows how to assemble the full strategy in a concept table.

The pattern for one concept

Concept block for loneliness in MEDLINE (written in words): the exploded heading Social Isolation, OR the heading Loneliness, OR the free-text words loneliness, lonely, social isolation, socially isolated and social exclusion in the title or abstract.

Each concept in the PCC question gets a block like this, and the blocks are then combined: population block AND concept block (interventions) AND outcome block.

The two approaches also help each other. Librarians often read the headings and abstracts of a few known relevant articles to find headings and wording to add, which is one reason the Cedar Valley librarian asks the intern to collect known relevant studies (Section 3).

Case: the intern's first lesson in indexing

The intern types "social prescribing loneliness older adults" into PubMed and likes the first page. The librarian opens the MeSH record for Social Prescribing and points out that the heading was introduced in 2025. A study of a link worker program published in 2019 would have been indexed with headings such as Social Support, Community Networks and Loneliness, and would not carry the Social Prescribing heading. The librarian then opens the MeSH tree for Social Isolation and shows that Loneliness sits beneath it, so an exploded Social Isolation search also finds studies indexed only with Loneliness. The intern leaves with a draft list of headings and free-text terms for each concept, which becomes a concept table in Lesson 4.

Try it: read a MeSH record

Open the MeSH Browser on the National Library of Medicine website and look up the heading Aged. Record (a) its scope note, (b) at least one entry term, (c) the narrower headings that sit beneath it in the tree, and (d) the year it was introduced. Then decide whether a search for adults aged 65 and older should use the exploded heading or the unexploded one, and write one sentence explaining your choice. Repeat the exercise for the heading Loneliness, one of the Cedar Valley concepts.

Common errors at this stage

First searches often contain three errors: searching a heading without checking whether it was exploded, relying on a single heading for a concept that has no heading or only a recent one, and copying MeSH headings into CINAHL or Embase without checking each heading in that database's own thesaurus.

Reflection

A health authority team is preparing a MEDLINE search, run in PubMed, on social prescribing programs for adults aged 65 and older. The librarian gives the team four facts from the MeSH Browser. First, the heading Social Prescribing was introduced in 2025 and sits beneath the heading Social Support. Second, the heading Social Isolation has four narrower headings: Loneliness, Ostracism, Social Alienation and Social Deprivation. Third, the heading Aged is defined as a person 65 through 79 years of age, and it has two narrower headings, Aged, 80 and over, and Frail Elderly. Fourth, PubMed explodes headings by default.

(a) Explain why a search that uses only the heading Social Prescribing would miss relevant studies, and list at least four free-text terms you would add for the intervention concept. (b) Explain what the exploded heading Social Isolation retrieves and why this matters for the loneliness concept. (c) State whether you would explode the heading Aged, with a reason based on its definition.

Model answer

(a) Because Social Prescribing became a MeSH heading only in 2025, records indexed before then were described with other headings, such as Social Support or Community Networks, and generally do not carry it. A search on the new heading alone would therefore find mainly recent records and miss most of the earlier evaluations. I would add free-text terms in the title and abstract fields for the ways authors describe these programs: social prescribing, social prescription, community connector, link worker, community navigator and community referral. I would keep the heading as well, because it will index future records consistently.

(b) The exploded heading Social Isolation retrieves records indexed with Social Isolation itself and with any of its four narrower headings. This matters because indexers assign the most specific heading available, so a study whose main outcome is loneliness is likely to be indexed with Loneliness rather than with Social Isolation. Exploding the broader heading captures those records in one step; I would still add Loneliness explicitly and add free-text terms such as lonely and social isolation.

(c) I would explode Aged. Its scope note covers ages 65 to 79 and sends indexers to Aged, 80 and over for older people, so an unexploded search would miss studies indexed only with the narrower heading, which is a large and relevant group for this question. Frail Elderly would also be included, which is appropriate.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a MEDLINE record, which of the following fields is added by the database rather than written by the authors?

Subject headings such as MeSH terms are assigned during indexing by the database. The title, abstract and author keywords come from the authors and publisher, so they are searched with free-text terms.

Question 2: A librarian searches PubMed for "Social Isolation"[Mesh]. Loneliness is a narrower heading beneath Social Isolation. Which records does this search retrieve?

PubMed explodes MeSH headings by default, so the search includes Loneliness and the other narrower headings. Ovid, by contrast, explodes only when exp is added, which is why the first option is a common confusion. A major-topic restriction requires the [Majr] tag.

Question 3: Why would a MEDLINE search that uses only the MeSH heading Social Prescribing miss many relevant studies?

Social Prescribing became a MeSH descriptor in 2025. Studies indexed before then were described with other headings, so free-text terms such as social prescribing and link worker are needed to find them. MeSH is applied to all study designs, and the heading is a genuine MeSH descriptor.

Question 4: Which statement best describes the relationship between MeSH, Emtree and CINAHL Subject Headings?

MeSH indexes MEDLINE, Emtree indexes Embase and CINAHL Subject Headings index CINAHL. All three are controlled vocabularies, but their headings and hierarchies differ, so each heading must be looked up in each database when a search is translated (Lesson 4).
Section 2 of 5

Boolean Retrieval and Relevance Ranking

⏱ Estimated reading time: 35 minutes
Section 2 of 5

Boolean Retrieval and Relevance Ranking

Databases return exact sets of records, while search engines return ranked lists.

About 35 minutes
The inverted index

Lists of records for each term

TermRecords containing it
loneliness1, 4, 6
lonely3
social isolation2, 7
older adults1, 2, 3, 5, 8

A query is answered by comparing these lists rather than by reading the records.

Boolean operators

AND, OR and NOT act on sets

AND

It keeps records in both sets, so the search narrows.

OR

It keeps records in either set, so the search broadens.

NOT

It removes records in the second set and can discard relevant ones.

(loneliness OR lonely OR social isolation) AND older adults returns records 1, 2 and 3.

Why reviews rely on it

Three properties of Boolean retrieval

Exhaustive

Every matching record is returned, and the count is exact.

Transparent

Readers can see from the strategy why each record was retrieved.

Repeatable

The same strategy, database and date return the same set.

PubMed Best Match changes the display order and leaves the set unchanged.

Relevance ranking

Scoring and ordering documents

Term weighting

Frequent words in a document and rare words in the collection raise a document's score.

Web signals

Link analysis and many other signals are used, and Google does not publish their weights.

Inverse document frequency
\[ \text{idf} = \log_{10}\left(\frac{N}{df}\right) \]
Google Scholar

A useful supplement with firm limits

1,000results is about the most that Google Scholar will display for a query.
  • Its ranking weighs full text, source, authors and citations.
  • It offers no official bulk export and no truncation.
  • Evaluations judge it unsuitable as a principal search system.
Carry forward

From retrieval to evaluation

The main search of a review runs in bibliographic databases, and ranked web tools play a supporting role.

Section 3 introduces recall and precision for judging any search.

Learning Objectives for this section

  • Describe how an inverted index allows a database to find matching records quickly.
  • Apply the Boolean operators AND, OR and NOT to small sets of records and predict the result of a query.
  • Explain in general terms how relevance ranking orders results, using term frequency, inverse document frequency and link analysis, and identify which parts of commercial ranking are not disclosed.
  • Compare bibliographic databases, Google and Google Scholar on exhaustiveness, transparency and suitability for systematic searching.

2.1 The Inverted Index

Section 1 described the fields in a database record. This section explains how a search system uses those fields to answer a query. A database with tens of millions of records cannot read every record each time a user searches. Instead, it prepares in advance a structure called an inverted index. For every word and every subject heading in the collection, the inverted index stores a list of the records that contain it, called a postings list. The idea resembles the index at the back of a textbook, which lists for each term the pages where it appears. Manning, Raghavan and Schütze (2008) give a full account of the inverted index in their textbook Introduction to Information Retrieval, and the description here follows theirs in simplified form.

When a user enters a query, the system looks up each term in the inverted index and works with the postings lists rather than with the records themselves. A search for "loneliness" therefore does not find "lonely", because the two words have separate entries in the index, unless the searcher asks for both or uses truncation. A search on a subject heading consults the entries for headings, which are separate from the entries for words in the title and abstract.

Term Postings list (record numbers) loneliness 1 4 6 lonely 3 social isolation 2 7 older adults 1 2 3 5 8 intervention 1 Query: loneliness AND older adults The system compares the two dark rows. Only record 1 appears in both lists, so the result set is record 1. Record 3 says "lonely" and sits in a different row, so this query does not find it. The toy collection has eight records and uses title words only.
An inverted index for the eight-record toy collection used in the exercise below. Each search works on these lists of record numbers rather than on the records themselves.

2.2 Boolean Retrieval in Bibliographic Databases

Boolean retrieval treats a query as a logical statement and returns every record for which the statement is true. It takes its name from the English mathematician George Boole, whose 1854 book The Laws of Thought set out an algebra of logical operations. Three operators do almost all the work in health searching. AND returns records that are in both sets (the intersection), so it narrows a search. OR returns records that are in either set or both (the union), so it broadens a search. NOT returns records in the first set that are absent from the second (the difference), so it removes records.

A AND B A B It returns records in both sets. The search narrows. A OR B A B It returns records in either set. The search broadens. A NOT B A B It returns records in A only. It can remove good records. A = records with a loneliness term; B = records with an older-adult term.
The three Boolean operators shown as operations on two sets of records. Shaded areas are the records each query returns.

Boolean retrieval in a bibliographic database has three properties that matter for evidence synthesis. It is exhaustive: the system returns every record that satisfies the logic, and the count it reports is the exact size of that set. It is transparent: anyone who reads the search strategy can work out why each record was or was not retrieved. It is repeatable: the same strategy run in the same database on the same date returns the same set, so another team can check the search and a later team can update it. These properties are the reason systematic reviews rely on bibliographic databases for their main searches, and they are the reason PRISMA-S (Rethlefsen et al., 2021) asks reviewers to report every strategy in full.

Boolean logic also has well-known weaknesses. A record either matches or it does not, with no partial credit, so a record that says "lonely" fails a query that asks for "loneliness". The searcher must anticipate every relevant word in advance. NOT is risky, because it removes any record that mentions the excluded term, including relevant ones; a search for older adults NOT dementia would drop a trial of a befriending program that mentions dementia once in passing. Databases also differ in the order in which they process operators when parentheses are missing, so careful searchers always use parentheses to group terms. Lesson 4 teaches the full syntax.

RecordTitle (illustrative)Indexed title words
1Loneliness among older adults: a group intervention trialloneliness, older adults, intervention
2Social isolation and mortality in older adultssocial isolation, older adults
3A befriending program for lonely older adultslonely, older adults
4Loneliness in university studentsloneliness
5Social prescribing for older adults: a pilot evaluationolder adults
6Telephone calls to reduce loneliness in care homesloneliness
7Social isolation in rural communitiessocial isolation
8Exercise for older adults with arthritisolder adults
Try it: run three Boolean queries by hand

Using the toy collection in the table and the postings lists in the inverted-index figure, write down the records each query returns: (1) loneliness AND older adults; (2) (loneliness OR lonely OR social isolation) AND older adults; (3) older adults NOT loneliness. Then decide which of the three queries finds the most of the records you would want for the Cedar Valley question, and which relevant record none of them can find from title words alone.

Answers to the three queriesv

Query 1 returns record 1 only, because 1 is the only number in both the loneliness list (1, 4, 6) and the older adults list (1, 2, 3, 5, 8). Query 2 first forms the union of the three loneliness-related lists, which is records 1, 2, 3, 4, 6 and 7, and then intersects it with the older adults list, giving records 1, 2 and 3. Query 3 removes records 1, 4 and 6 from the older adults list, giving records 2, 3, 5 and 8; it discards record 1, the most relevant record in the collection, which shows why NOT is dangerous.

Query 2 finds the most relevant records. Record 5, on social prescribing for older adults, may well measure loneliness, but its title does not mention it, so none of the three queries can find it from title words. A search of its abstract or its subject headings, or a fourth concept for the intervention, would be needed.

Sorting a Boolean result set: PubMed Best Match

Many databases now offer a relevance sort, and since 2020 PubMed's default display order has been "Best Match", a machine-learning ranking described by Fiorini et al. (2018). Sorting changes only the order in which the retrieved records are displayed. The set of records, and the count PubMed reports, are the same whichever sort is chosen. A systematic search that screens every retrieved record is therefore unaffected by the sort order, whereas a search that looks at only the first page is strongly affected by it.

2.3 Relevance Ranking

Relevance ranking (also called ranked retrieval) works differently. Instead of deciding whether each record matches, the system gives each document a score for how well it appears to answer the query and presents the documents in order of that score. Documents that contain only some of the query words can still be returned, lower in the list. Users of a ranked system look at the first page or two and stop, so the ranking, rather than the logic of the query, decides what the user sees.

Ranking methods developed through decades of published research in information retrieval. Gerard Salton and colleagues at Cornell University developed the vector space model, which represents queries and documents as weighted lists of terms and ranks documents by their similarity to the query. Two weighting ideas from that research remain central. Term frequency gives more weight to a document that uses a query word more often. Inverse document frequency, proposed by Karen Spärck Jones (1972), gives more weight to words that are rare in the collection, because a rare word says more about what a document is about than a common one. Probabilistic models developed by Stephen Robertson and colleagues led to the widely used BM25 formula (Robertson and Zaragoza, 2009), which adds limits on the effect of repeated words and adjusts for document length.

A worked illustration of term weighting

Inverse document frequency for a term = log10(N ÷ df), where N is the number of documents in the collection and df is the number of documents that contain the term.

Suppose a collection holds 1,000,000 documents. The word "adults" appears in 200,000 of them, so its weight is log10(1,000,000 ÷ 200,000) = log10(5) = 0.699. The word "loneliness" appears in 2,000 of them, so its weight is log10(500) = 2.699, about 3.9 times the weight of "adults".

Document A uses "loneliness" 6 times and "adults" 2 times; a simple score of frequency multiplied by weight gives 6 × 2.699 + 2 × 0.699 = 17.59. Document B uses "loneliness" once and "adults" 10 times, giving 2.699 + 10 × 0.699 = 9.69. Document A ranks above Document B. Real systems dampen repeated words and adjust for length, but the principle is the same.

The same principle applies to the toy collection. In a ranked system, a query for loneliness older adults would return record 1 first, because it contains both words. Records 4 and 6, which contain only "loneliness", would come next, because "loneliness" appears in three of the eight records and is rarer than "older adults", which appears in five. Records 2, 3, 5 and 8 would follow. Record 7 contains neither word and would not appear. The ranked system returns seven records in order, where the Boolean query in the exercise returned exactly one.

Web search engines add information that ordinary text collections lack: the links between pages. Sergey Brin and Lawrence Page (1998) described PageRank, which treats a link from one page to another as a kind of endorsement and gives a page a higher score when many well-linked pages point to it. PageRank was part of the original design of Google. Link analysis is now one signal among many, and the details of how it is currently used are not public.

Google's public documentation describes its ranking systems only in general terms. It states that the systems consider the meaning of the query, the relevance of page content, signals of quality such as expertise and trustworthiness, the usability of pages, and the user's context and settings, including location, search history and language. It also states that the weight of each factor varies with the type of query. Google does not publish the signals in full, their weights or their interactions, and it changes its systems continually. Any more specific claim about how Google ranks a given page should be treated with caution.

Google Scholar applies ranking to scholarly documents. Its own help pages state that it aims to rank documents in the way researchers do, weighing the full text of each document, where it was published, who wrote it, and how often and how recently it has been cited in other scholarly literature. An early independent analysis by Beel and Gipp (2009) concluded that citation counts carried heavy weight in the ranking. One consequence is that older, highly cited articles tend to appear near the top, while new studies, reports and theses may sit far down the list even when they answer the question directly.

Bibliographic databases increasingly offer ranked displays, and some offer searching by meaning rather than by exact words, which Lesson 6 discusses. In a Boolean database, a relevance sort reorders a set that the query has already defined, as with PubMed Best Match. In a fully ranked system, the ranking also decides which documents the user will realistically see, because the list is long and the user stops reading.

2.4 Comparing the Two Models for Evidence Synthesis

Relevance ranking is well suited to finding a good answer quickly, which is what most web users want. A systematic or scoping review has a different goal: it aims to find all the eligible studies, or as close to all as resources allow, and to show others exactly how they were found. The table compares the three kinds of tool the Cedar Valley team will use against that goal.

FeatureBibliographic database (for example MEDLINE)GoogleGoogle Scholar
How matches are decidedBoolean logic, with an optional relevance sortUndisclosed ranking with many signalsUndisclosed ranking that weights citations
Size of the result setExact count of every matching recordAn estimated countAn estimated count
Can every result be viewedYesNo, the list ends long before the estimateNo, about 1,000 results at most
Bulk export of resultsYes, to a reference managerNoNo official bulk export
Controlled vocabularyYes (MeSH, Emtree, CINAHL headings)NoNo
Truncation and nested Boolean logicFully supportedLimitedNo truncation and limited grouping
CoverageDocumented journal listsThe open web, undocumentedBroad scholarly coverage, undocumented
Role in a systematic searchPrincipal searchSupplementary, for grey literatureSupplementary, for grey literature and checking

Published evaluations support the roles in the last row. Haddaway et al. (2015) found that Google Scholar shows at most the first 1,000 results of any search and recommended using it as a supplementary source, particularly for grey literature, with a defined number of results screened. Gusenbauer and Haddaway (2020) tested 28 academic search systems against criteria for systematic searching, including Boolean functionality, reproducibility and bulk export, and concluded that Google Scholar was unsuitable as a principal search system for systematic reviews, although useful as a supplementary one. Bramer et al. (2017) showed that no single database retrieved all the studies included in a large set of reviews, which is why the Cedar Valley team plans to search five databases rather than one.

Strength of Boolean retrievalClick to explore
Weakness of Boolean retrievalClick to explore
Strength of relevance rankingClick to explore
Weakness of relevance rankingClick to explore
Case: two counts for the same idea

The intern runs a quick test. In PubMed, a short tagged search for the loneliness concept and the older-adult concept returns a fixed set of a few hundred records, every one of which can be exported to the team's reference manager. In Google Scholar, similar words produce an estimate of tens of thousands of results, ordered with highly cited reviews at the top. The intern can page through only the first 1,000 and cannot export them in bulk. The librarian explains that the two numbers measure different things: the PubMed figure is the exact size of a defined set, and the Google Scholar figure is an estimate attached to a ranked list that no one can read in full. The team decides that the five bibliographic databases will provide the main search, and that Google Scholar will be used in Lesson 5 as a supplementary source with a fixed screening limit.

Reflection

A student is preparing a scoping review on community programs that reduce loneliness in older adults. In PubMed, a search that combines a loneliness concept and an older-adult concept with AND returns 412 records, and the count stays at 412 whether the results are sorted by Best Match or by most recent. In Google Scholar, a similar set of words produces the message "About 18,600 results", with highly cited reviews on the first page. Google Scholar displays at most about 1,000 results for a query and has no official bulk export. The student proposes to replace the PubMed search with the first 50 Google Scholar results, arguing that they look more relevant.

(a) Explain why the two numbers, 412 and about 18,600, differ in meaning. (b) Explain why sorting by Best Match did not change the PubMed count. (c) Give two reasons, based on how each tool retrieves records, why the proposal is unsuitable for the main search of a scoping review, and describe a suitable role for Google Scholar.

Model answer

(a) The PubMed figure of 412 is the exact size of a set defined by Boolean logic: every record that satisfies the query, all of which can be viewed and exported. The Google Scholar figure is an estimate attached to a ranked list. It is not the size of a set that anyone can see, because only about the first 1,000 results are displayed.

(b) Best Match is a sort order. PubMed first retrieves the Boolean set and then decides only the order in which the 412 records are displayed, so the set and the count stay the same under any sort.

(c) First, the first 50 Google Scholar results are chosen by an undisclosed ranking that gives weight to citation counts, so recent studies, reports and evaluations from small programs are likely to sit lower in the list. Looking good at the top of the list is a sign of precision and says nothing about recall, which is what a scoping review needs. Second, the Google Scholar search cannot be fully reported or rerun: results cannot be exported in full, truncation is not supported, and the ranking changes between users and over time, so readers could not check the search or update it. A suitable role for Google Scholar is supplementary: after the database searches, the student could run a few documented queries, screen a fixed number of results set in advance, such as the first 100, and record the date, the query and the settings.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In a toy collection, loneliness appears in records 1, 4 and 6, and older adults appears in records 1, 2, 3, 5 and 8. Which records does the query loneliness OR older adults return?

OR returns the union of the two postings lists: every record that contains either term or both. Record 1 alone would be the result of AND, and the other two options describe the results of NOT in each direction.

Question 2: What does sorting a PubMed result set by Best Match rather than by publication date change?

Best Match is a relevance sort applied to the set the query has already retrieved. The set, the count and the explosion of headings are unchanged; only the display order differs.

Question 3: In relevance ranking, under which condition does inverse document frequency give a query word more weight?

Inverse document frequency, proposed by Karen Spärck Jones, weights rare words more heavily because they say more about what a document is about. Repetition within a document is term frequency, and links between pages belong to link analysis such as PageRank.

Question 4: Which statement about Google Scholar is supported by published evaluations and by Google Scholar's own documentation?

Google Scholar's documentation says it weighs citations among other factors, and evaluations such as Haddaway et al. (2015) and Gusenbauer and Haddaway (2020) report the display cap and conclude that it is suitable only as a supplementary source. It has no MeSH indexing and does not disclose its formula.
Section 3 of 5

Recall and Precision: Why Reviews Favour Recall

⏱ Estimated reading time: 35 minutes
Section 3 of 5

Recall and Precision: Why Reviews Favour Recall

Two measures describe how well a search performs, and reviews weight one of them heavily.

About 35 minutes
Definitions

Recall, precision and the number needed to read

Recall and precision
\[ \text{Recall} = \frac{\text{relevant retrieved}}{\text{all relevant}} \qquad \text{Precision} = \frac{\text{relevant retrieved}}{\text{all retrieved}} \]
Number needed to read
\[ \text{NNR} = \frac{1}{\text{Precision}} \]
Worked example (illustrative)

Two draft MEDLINE searches, 30 eligible studies

MeasureSearch ASearch B
Retrieved150750
Eligible found1827
Recall60%90%
Precision12%3.6%
Number needed to read8.327.8
Dual screening time2.5 hours12.5 hours
The trade-off

Recall and precision move in opposite directions

These choices raise recall

Adding synonyms with OR, exploding headings and adding free text raise recall and lower precision.

These choices raise precision

Adding a concept with AND, searching titles only and restricting to major topics raise precision and lower recall.

Each extra study found by Search B cost about 67 extra records to screen.

Estimating recall

Testing against a benchmark set

12 of 20benchmark studies were found by Search A, an estimated recall of 60%.
18 of 20benchmark studies were found by Search B, an estimated recall of 90%.

The studies a search misses show which terms it lacks.

Why reviews favour recall

The costs of the two errors differ

Unequal costs

An irrelevant record costs seconds, while a missed study is invisible.

Risk of bias

Missed studies can differ systematically from those found.

Readers rely on it

Decision-makers assume the review is close to comprehensive.

Carry forward

Judge a search by what it finds

A relevant first page shows precision, and it says nothing about recall.

Section 4 asks whether web searches can be reproduced.

Learning Objectives for this section

  • Define recall, precision and the number needed to read, and calculate each from the results of a search.
  • Explain why recall and precision usually move in opposite directions as a search is broadened or narrowed.
  • Describe how recall is estimated in practice with a benchmark set of known relevant studies.
  • Explain why systematic and scoping reviews favour recall, and identify the conditions under which a review team might accept lower recall.

3.1 Two Questions About Any Search

Sections 1 and 2 explained how databases find records. This section asks how well a particular search performs. Two questions capture most of what matters. The first is how many of the relevant records the search found. The second is how much of what the search returned is relevant. The measures that answer these questions, recall and precision, became the standard pair through the Cranfield experiments that Cyril Cleverdon led in England in the late 1950s and 1960s, which compared indexing systems by testing them against a set of documents whose relevance had been judged in advance. The same two measures are still used to evaluate search engines, search filters and the AI-assisted screening tools that Lesson 6 describes.

Any record in a database falls into one of four groups, depending on whether it is relevant to the review question and whether the search retrieved it. The figure shows these groups for a search of a single database.

All records in the database b retrieved but not relevant a relevant and retrieved c relevant but missed d neither Teal circle: retrieved Red circle: relevant
The four groups of records for any search. Recall compares a with all relevant records (a plus c), and precision compares a with all retrieved records (a plus b).

The formulas

Recall = relevant records retrieved ÷ all relevant records in the database = a ÷ (a + c)

Precision = relevant records retrieved ÷ all records retrieved = a ÷ (a + b)

Number needed to read (NNR) = all records retrieved ÷ relevant records retrieved = 1 ÷ precision

Recall is the same quantity that a diagnostic test calls sensitivity: the proportion of true cases that the test detects. Precision resembles the positive predictive value of a test: the proportion of positive results that are true cases. Bachmann et al. (2002) proposed the number needed to read as a more intuitive form of precision, by analogy with the number needed to treat. It states how many records a reviewer must read, on average, to find one relevant record. The fourth group, d, is very large in any bibliographic database, so a measure built on it, such as specificity, is close to 100 percent for almost every search and says little about performance. For that reason, search evaluation uses recall and precision rather than sensitivity and specificity as a pair.

3.2 A Worked Example from the Cedar Valley Review

The following example is a teaching device. It assumes that the true number of relevant records is known, which never happens in practice; Section 3.4 explains how recall is estimated when it is unknown. Suppose that MEDLINE contains exactly 30 studies that meet the Cedar Valley eligibility criteria. The librarian drafts two searches. Search A is narrow: it looks for the words loneliness, older adults and intervention in titles only. Search B is broad: it combines the subject headings from Section 1 with free-text synonyms in titles and abstracts for each of the three concepts. All numbers are illustrative.

QuantitySearch A (narrow)Search B (broad)
Records retrieved (a + b)150750
Relevant records retrieved (a)1827
Relevant records missed (c)30 − 18 = 1230 − 27 = 3
Recall, a ÷ (a + c)18 ÷ 30 = 0.60 (60%)27 ÷ 30 = 0.90 (90%)
Precision, a ÷ (a + b)18 ÷ 150 = 0.12 (12%)27 ÷ 750 = 0.036 (3.6%)
Number needed to read150 ÷ 18 = 8.3750 ÷ 27 = 27.8
Screening time at 30 seconds per record, one reviewer75 minutes375 minutes (6.25 hours)
Screening time with two independent reviewers2.5 hours in total12.5 hours in total

The comparison shows the trade-off in concrete terms. Search B costs five times as many records to screen and about ten more hours of reviewer time with dual screening. In exchange, it finds 9 more eligible studies and misses 3 instead of 12. Each additional study found by Search B costs about 67 extra records to screen (600 extra records divided by 9 extra studies). Search A misses 40 percent of the eligible studies, which would leave the evidence brief for Cedar Valley built on little more than half of the relevant evidence in MEDLINE.

The example also shows why specificity is uninformative. If MEDLINE holds about 30 million records, both searches leave out more than 99.99 percent of the irrelevant ones, so specificity cannot distinguish a good search from a poor one. Recall and precision, by contrast, differ clearly between the two searches.

Recall (0 to 100%) 0 50 100 A: 60% B: 90% 0 500 1,000 1,500 Precision (0 to 25%) 0 10 20 A: 12% B: 3.6% 0 500 1,000 1,500 Records retrieved as the search is broadened (illustrative curves)
As a search is broadened, recall rises quickly at first and then levels off, while precision falls. The curves are illustrative and are drawn to agree with the Search A and Search B figures in the table.

3.3 Why Recall and Precision Move in Opposite Directions

The curves in the figure have a typical shape. The first terms a librarian adds find many relevant records, so recall climbs quickly. Each further synonym or heading finds fewer new relevant records and many more irrelevant ones, because rarer wordings are also used in unrelated contexts. Recall therefore levels off while precision keeps falling. No search reaches 100 percent recall at a workable precision in most health topics, and the practical question becomes how far along the curve the team can afford to go.

Most decisions in search design move a search along this curve. The table summarizes the usual direction of each effect. These are tendencies, and an individual change can behave differently in a particular database.

Search choiceUsual effect on recallUsual effect on precision
Add synonyms or variant spellings with ORIncreasesDecreases
Explode a subject headingIncreasesDecreases
Combine subject headings with free textIncreasesDecreases
Add another concept with ANDDecreasesIncreases
Search titles only rather than titles and abstractsDecreasesIncreases
Restrict a heading to major topicDecreasesIncreases
Exclude a term with NOTDecreases, sometimes sharplyIncreases
Search additional databasesIncreases across the reviewDecreases, and adds duplicates

Lesson 4 introduces validated search filters, such as the filters for randomized trials in the Cochrane Handbook, which are published in sensitivity-maximizing and precision-maximizing versions. Those two labels refer directly to the recall and precision trade-off described here.

3.4 Estimating Recall When the Total Is Unknown

In a real review, nobody knows how many relevant records a database contains, so recall cannot be calculated directly. Precision is easier: once the team has screened the retrieved records, precision is the number judged eligible divided by the number screened. Recall has to be estimated, and two methods are common.

The first method uses a benchmark set (also called a validation set or test set) of known relevant studies, collected independently of the search being tested. Sources include the included studies of earlier reviews, studies suggested by content experts, and studies the team already knows. The librarian checks that each benchmark study is present in the database and then runs the draft search. The estimated recall is the proportion of benchmark studies that the search retrieves. A missed benchmark study is also informative, because its title, abstract and headings show which terms the search lacks. The second method, relative recall (Sampson et al., 2006), pools all the eligible studies found by every search method in the review, including other databases and citation chasing, and treats that pool as the reference standard for each individual search.

Case: testing the Cedar Valley draft against a benchmark set

The intern collects 20 eligible studies from the reference lists of two earlier reviews on interventions for loneliness in older adults and confirms that all 20 are indexed in MEDLINE. Draft Search A retrieves 12 of the 20, an estimated recall of 60 percent. Draft Search B retrieves 18 of the 20, an estimated recall of 90 percent. The librarian reads the two studies that Search B missed. One describes its intervention only as "befriending" and its outcome only as "social participation"; the other was indexed under broader headings without Loneliness, and its abstract uses the phrase "social disconnection". The team adds those terms to the free-text lists. Because adjusting a search to fit a known set can make it look better than it is, the librarian keeps ten further known studies aside as a final check, the test set that Lesson 4 describes.

Try it: calculate recall, precision and the number needed to read

A librarian tests a search for a different review against a benchmark set of 25 known eligible studies, all present in the database. The search retrieves 1,200 records, including 23 of the 25 benchmark studies. After screening all 1,200 records, the team judges 40 of them eligible. Calculate (a) the estimated recall, (b) the precision, (c) the number needed to read, and (d) the screening time for one reviewer at 30 seconds per record. Then state one change that would raise the estimated recall and its likely effect on precision.

Worked answerv

(a) Estimated recall is 23 ÷ 25 = 0.92, or 92 percent. (b) Precision is 40 ÷ 1,200 = 0.033, or about 3.3 percent. (c) The number needed to read is 1,200 ÷ 40 = 30, so the team reads 30 records for each eligible one. (d) At 30 seconds per record, one reviewer needs 1,200 × 0.5 = 600 minutes, or 10 hours. To raise estimated recall, the librarian should read the two missed benchmark studies and add the terms or headings they use with OR; this would retrieve more records and would usually lower precision.

3.5 Why Reviews Favour Recall

Chapter 4 of the Cochrane Handbook (Lefebvre et al., 2023) advises that searches for systematic reviews should aim for high sensitivity, that is, high recall, while accepting that precision will be low. The reasons can be stated in terms of costs and of bias, and the cards below set them out.

Unequal costs of errorClick to explore
Missed studies can bias findingsClick to explore
Readers rely on comprehensivenessClick to explore
When lower recall is acceptedClick to explore

Precision still matters, because reviewer time is finite. The Cedar Valley team has twelve weeks, and screening thousands of records with two reviewers competes with charting, appraisal, the environmental scan and writing. Its plan therefore aims for high recall in the main database searches, and it manages the workload in other ways: removing duplicates before screening (Lesson 7), using a screening platform, piloting eligibility criteria so that decisions are quick, and possibly using active-learning screening tools (Lesson 6). The team will consider lowering recall only after these options, and it will report any limit it sets.

A common misunderstanding

Students sometimes judge a search by whether its first page looks relevant. That is a judgement about precision at the top of a ranked list, which is how web search engines are designed to perform. A systematic search is judged by recall, which cannot be seen on any single page of results. A search whose first page is full of relevant records can still miss a large share of the eligible studies, and a broad search with a mixed first page can be the better search for a review.

Reflection

Two draft searches for a review are tested in MEDLINE against a benchmark set of 25 known eligible studies, all of which are present in MEDLINE. Search X retrieves 2,000 records, includes 24 of the 25 benchmark studies, and on screening yields 60 eligible studies. Search Y retrieves 600 records, includes 19 of the 25 benchmark studies, and on screening yields 48 eligible studies. The team has two reviewers who each screen every record independently at about 30 seconds per record, and it must deliver a scoping review within twelve weeks. Recall is relevant records retrieved divided by all relevant records; precision is relevant records retrieved divided by all records retrieved; the number needed to read is records retrieved divided by relevant records retrieved.

(a) Calculate the estimated recall, the precision and the number needed to read for each search, and the total dual-screening time for each. (b) Recommend one search and justify your choice with reference to why reviews favour recall. (c) Describe one way to improve the search you did not choose, or to reduce the workload of the one you chose, without simply accepting lower recall.

Model answer

(a) Search X has an estimated recall of 24 ÷ 25 = 96 percent, a precision of 60 ÷ 2,000 = 3 percent, and a number needed to read of 2,000 ÷ 60 = 33.3. Dual screening takes 2,000 records × 0.5 minutes × 2 reviewers = 2,000 minutes, about 33.3 hours. Search Y has an estimated recall of 19 ÷ 25 = 76 percent, a precision of 48 ÷ 600 = 8 percent, and a number needed to read of 600 ÷ 48 = 12.5. Dual screening takes 600 × 0.5 × 2 = 600 minutes, or 10 hours.

(b) I would choose Search X. Search Y misses about one in four eligible studies, and those it misses may differ systematically from those it finds, for example by using different terms or coming from another discipline, which could bias the scoping review's account of the evidence. The extra 23 hours of screening are a known and manageable cost over twelve weeks, whereas the missed studies would be invisible.

(c) To make Search X manageable, the team could remove duplicates across databases before screening, pilot the eligibility criteria on 50 records so that decisions are quick, and use a screening platform. Alternatively, to improve Search Y, the librarian could read the six benchmark studies it missed, identify the headings and free-text terms they use, and add those terms with OR, then retest against benchmark studies held back for that purpose.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A search retrieves 400 records, of which 20 are relevant. The database contains 25 relevant records in total. What are the recall and precision of the search?

Recall is 20 relevant retrieved divided by 25 relevant in total, which is 80 percent. Precision is 20 relevant retrieved divided by 400 retrieved, which is 5 percent. The second option reverses the two measures.

Question 2: A search has a precision of 5 percent. What is its number needed to read?

The number needed to read is 1 divided by precision: 1 ÷ 0.05 = 20. On average, a reviewer reads 20 records to find one relevant record. The figure of 95 confuses the proportion of irrelevant records with the number needed to read.

Question 3: Why do systematic and scoping reviews usually favour recall over precision?

The costs of the two errors are unequal. An irrelevant record costs seconds at screening, while a missed study leaves no trace and may differ systematically from the studies found. Precision can be calculated after screening, and no search guarantees complete recall, which is why reviews also use citation chasing.

Question 4: A librarian has 20 known eligible studies, all present in MEDLINE, and a draft search retrieves 17 of them. What does this tell the team?

Seventeen of 20 benchmark studies gives an estimated recall of 85 percent for this search in MEDLINE. The benchmark set says nothing about precision or the number needed to read, which require screening the retrieved records, and the estimate applies to MEDLINE only.
Section 4 of 5

Personalization, Ranking Opacity and the Reproducibility of Web Search

⏱ Estimated reading time: 35 minutes
Section 4 of 5

Personalization, Ranking Opacity and the Reproducibility of Web Search

Web search results change with the person, the place and the day.

About 35 minutes
Two standards

Reproducibility and transparency

Reproducibility

A rerun of the search returns the same results, or the differences can be explained.

Transparency

The report shows exactly what was searched, when, how and with what limits.

A web search can be made transparent even when it cannot be reproduced.

Personalization and context

Who, where and in what language

Account and search history Inferred location Language and country version Device and settings

A private window removes browser history, and the search engine can still infer location.

Ranking opacity and change

Why the same query drifts over time

  • Ranking systems are described only in general terms and are updated continually.
  • The web index changes as pages are added, moved and deleted.
  • Google Scholar positions shift as articles gain citations.
  • The displayed results count is an estimate rather than the size of a set.
Cedar Valley (fictional)

The same query on two screens

Evidence officer

The officer searched signed out in Cedar City and saw local news first.

Student intern

The intern searched signed in on campus and saw a review article first.

The two first pages shared only three of ten results.

Making it transparent

The minimum record for a web search

  • Record the tool, the date and the exact query.
  • Search signed out in a private window, and record the settings and location.
  • Set a stopping rule before you begin screening.
  • Save the results pages and any documents you keep.
  • Report the displayed count as an estimate, and give the number screened.
Next steps

Reflection, knowledge check and final assessment

Databases carry the main search, and web tools serve as documented supplements.

Complete the reflection and knowledge check below, then the final assessment. Lesson 4 builds the full database search from the groundwork described in this lesson.

Learning Objectives for this section

  • Define reproducibility and transparency as they apply to a literature search, and explain why the two differ for web searches.
  • Identify the main sources of variation in web search results, including personalization, location, language, index changes and updates to ranking systems.
  • Compare the reproducibility of bibliographic database searches, Google searches and Google Scholar searches.
  • Describe the minimum information a reviewer should record about a web search so that readers can judge it.

4.1 What Reproducibility Means for a Search

A search is reproducible when another person, following the documented method, can run it again and obtain the same results, or can explain any differences. Reproducibility matters in evidence synthesis for three reasons. It allows peer reviewers and readers to check that a review searched as it claims. It allows the same team, or a different one, to update the review later by rerunning the search for newer records. It also allows a reader to judge the risk that relevant studies were missed. PRISMA-S, the reporting guideline for literature searches developed by Rethlefsen et al. (2021), exists largely to make searches reproducible, and Lesson 4 teaches it in detail.

A related idea is transparency: the extent to which a report shows exactly what was done, even when the result could not be obtained again. A search can be fully transparent without being reproducible. This distinction is central to web searching, because, as this section shows, a reviewer can report a Google search completely and still be unable to guarantee that anyone else will see the same results.

Bibliographic database searches come close to reproducibility, although even they fall short of it. Databases add new records every day, sometimes including older articles added later. Thesauri change each year, as the Social Prescribing heading did in 2025. Platforms change their search engines and syntax. A rerun of a MEDLINE search a year later, limited to the original date range, will usually return a very similar set, and the differences can be traced to these known causes. Web search engines behave very differently, for the reasons set out below.

4.2 Personalization and Context

Personalization means that a search engine adjusts results to the person searching. Google's public documentation states that its systems use the user's location, search history and settings, that the language of the query shapes the language of the results, and that pages a user has visited often may be moved higher in that user's results. It also states that users can see whether personal information affected a result and can turn personalization off. Hannak et al. (2013) measured personalization in Google web search by comparing the results that real accounts and controlled test accounts received for the same queries. They found measurable personalization, with being signed in to an account and geographic location as the main sources of difference.

The term "filter bubble", popularized by Eli Pariser (2011), describes the concern that personalization shows people mainly what fits their past behaviour. Google states that its systems are not designed to create filter bubbles. Studies that have measured personalization in web search have generally found it present but modest in average size. For evidence synthesis, the size of the effect matters less than its existence: any personalization means that two reviewers running the same query can see different results, and neither can tell from the results page alone which records were affected.

Some sources of variation remain even when personalization is turned off. Results depend on the approximate location that the search engine infers from the network connection, which a private browsing window does not hide. They depend on the country version of the search engine, on the interface language, and on settings such as filters for explicit content. They can also depend on the device, because page usability is one of the factors Google describes. A reviewer can control some of these factors and record the rest, but cannot remove them.

Who is searchingClick to explore
Where and in what languageClick to explore
When the search is runClick to explore
How the ranking works that dayClick to explore
Identical query Evidence officer Cedar City, signed out Student intern Burnaby, signed in Ranking system undisclosed and updated continually Web index on that day Officer's first results 1. Local news story 2. Health authority page 3. Seniors' centre page Intern's first results 1. Review article 2. National charity report 3. Health authority page
An illustrative example of how the same query can produce different first results for two people. The ranking system and the state of the web index are outside the reviewer's view, while the searcher's context is partly within the reviewer's control.

4.3 Ranking Opacity and Change Over Time

Ranking opacity means that the method used to order results is not disclosed in enough detail for an outsider to predict or check it. Section 2 showed that Google publishes only general categories of ranking factors and that Google Scholar describes its ranking in a single paragraph. Commercial search engines have reasons to keep the details private, including competition and the need to resist manipulation of rankings by website owners. For a reviewer, the consequence is that no one can explain why a given page appeared in position three rather than position thirty, or whether it would appear at all if the search were run again.

Opacity combines with change over time. Search engines revise their ranking systems continually, and the web index is rebuilt as pages change. Google Scholar's documented use of citation counts in ranking has a further consequence: as articles gain citations, their positions change, so the order of results for the same query drifts from month to month even if no new articles are published. Because Google Scholar shows at most about the first 1,000 results, articles near that boundary can move in and out of view.

Other features of the results page make the problem harder. The counts that Google and Google Scholar display ("about" so many results) are estimates that can change between pages of the same search and should not be reported as the size of a result set. Results pages may include advertisements labelled as sponsored, highlighted answer boxes drawn from a single page, and, since 2024, AI-generated summaries above the results for some queries. Each of these elements is produced by its own undisclosed process, and each can differ between users and days.

Question about reproducibilityBibliographic databaseGoogleGoogle Scholar
Can the exact query be recorded and reported?YesYesYes
Can the full set of results be saved?Yes, by exportOnly the pages actually viewedOnly the pages viewed, up to about 1,000 results
Will a rerun return the same records?Largely, within the original date limitsUnlikelyUnlikely
Can a reader see why a record was retrieved?Yes, from the strategyNoNo
Is the searcher's context a factor?NoYes, location, language, account and settingsPossibly; language settings can change results, and other personalization is undocumented
Is the collection searched documented?Yes, with published journal listsNoNo

Gusenbauer and Haddaway (2020) included reproducibility among the criteria they used to judge academic search systems, and they judged that Google Scholar did not meet their requirements for reproducible searching. This finding, together with the cap on viewable results and the lack of bulk export, is why reviews use Google Scholar and Google as supplementary sources. Supplementary searching is still valuable. Much of the information the Cedar Valley environmental scan needs, such as program descriptions, funding announcements and evaluation reports from community organisations, exists only on the open web and will never appear in MEDLINE.

Case: the same query on two screens

On the same afternoon, the evidence officer in Cedar City and the intern working from campus in Burnaby type the same query into Google: social prescribing older adults British Columbia. The officer is signed out; the intern is signed in to an account that has been used for weeks of searching on social prescribing. Their first pages share only three of ten results, and the shared results appear in different positions. The officer's page includes local news and a seniors' centre page, while the intern's page includes a review article and a national charity report. When the intern repeats the search two weeks later, one program page on the first list has moved to a new address and a new news story has appeared. All details here are illustrative. The team concludes that it cannot make its web searches reproducible, and it decides instead to make them transparent: it will control the settings it can, record the rest, and save copies of everything it screens.

4.4 Making Web Searches Transparent

Because a web search cannot be reproduced exactly, the reviewer's aim shifts to recording enough for a reader to understand what was searched and to judge what might have been missed. PRISMA-S asks reviewers to report web searches and other supplementary methods alongside database searches, and Lesson 5 provides a full documentation template for grey literature and website searching. The items below give the minimum record that this lesson's mechanisms imply.

Record the tool, the date and the exact queryv

Note the search engine and its country version, the date of the search, and the query exactly as typed, including quotation marks and operators. A query that is reported only in summary cannot be repeated even approximately.

Control and record the searcher's contextv

Search while signed out, in a private browsing window, with the interface language and region set deliberately, and record these settings and the approximate location of the searcher. These steps reduce personalization from the account and browser history, and the record allows a reader to see what was left uncontrolled.

Set a stopping rule before searchingv

Decide in advance how many results to screen, for example the first 100 results of each query, and report the rule. Haddaway et al. (2015) recommended a defined limit of this kind for Google Scholar, because no one can screen an estimated count of tens of thousands. A rule set in advance prevents the reviewer from stopping when results happen to look poor or continuing when they happen to look good.

Save what was screenedv

Save the results pages screened, as files or screenshots, and save copies of included documents, because web pages change or disappear. Record each included document's address and the date it was accessed. These copies allow the team, and anyone checking its work, to see the results as they were.

Report the estimate as an estimatev

If a results count is reported, describe it as the estimate the search engine displayed, and report separately the number of results actually screened. The number screened is the figure that belongs in the PRISMA flow diagram, which Lesson 7 covers.

Try it: compare two screens

With a friend or colleague, agree on one query about a health topic that interests you. At the same time, each of you runs it in Google, one signed in and one signed out in a private window, and each records the first ten results. Count how many results appear on both lists, and note how many of the shared results appear in the same position. Then write two sentences explaining which differences you could have controlled and which you could only record.

4.5 From Search Engines to AI Tools

The mechanisms in this lesson carry forward to the artificial intelligence tools that Lesson 6 examines. AI search engines and research assistants usually begin with a retrieval step that resembles relevance ranking, so they inherit the opacity and the dependence on an undisclosed index described here. They then add a step that generates text from what was retrieved, and that step can produce different wording, different sources and, at times, citations to works that do not exist, even when the same question is asked twice. The same distinction between reproducibility and transparency applies: a reviewer who uses these tools must record what was asked, when, with which tool and version, and must verify each source against the original publication.

Summary of the lesson's argument

Bibliographic databases describe records with controlled vocabularies and retrieve them with Boolean logic, which makes their searches exhaustive, transparent and close to reproducible. Web search engines rank documents with undisclosed methods that depend on the searcher, the place and the day, which makes them useful for finding grey literature and unsuitable as the principal search for a review. Recall and precision give a common language for judging any of these searches, and reviews favour recall because missed studies are invisible and can bias the findings.

Groundwork before the full search

Before it builds a full search, a review team usually prepares three records. First, for each main concept in its PCC question, it looks up the MeSH heading in the MeSH Browser and the matching CINAHL Subject Heading, and records each heading's scope note, at least two entry terms, its narrower headings, whether it will be exploded, and the year it was introduced, together with a list of free-text synonyms for each concept. Second, it assembles a benchmark set of three to five studies already known to meet the eligibility criteria, notes where each one was found, and confirms that each is in PubMed. Third, it runs one test search in Google Scholar while signed out, records the date, the exact query, the settings, the estimated count and the first five results, and reruns the search a few days later to note any differences. Lesson 4 uses these records to build and test the full search.

Reflection

Two members of a review team in British Columbia run the Google query community connector program seniors loneliness BC on the same afternoon. One is signed in on a laptop in Vancouver. The other is signed out, using a private browsing window on a phone in Kamloops. Their first ten results share four pages, and the shared pages appear in different positions. A week later, one of the shared pages, a program description, returns an error because it has moved. Google's public documentation states that its ranking considers the meaning of the query, the relevance and quality of content, the usability of pages, and the user's context and settings, including location, search history and language, and it does not publish the weights of these factors. The team must report its web searching in the review.

(a) Identify three distinct sources of variation that could explain the differences, and for each state whether the team could have controlled it or could only record it. (b) Write a documentation entry for one of these searches that would allow a reader to judge it. (c) Explain in two or three sentences why this search can be transparent even though it cannot be reproduced.

Model answer

(a) The first source is personalization from the signed-in account, including search history; the team could have controlled this by having both members search while signed out in private windows. The second is location: Vancouver and Kamloops may receive different local results for a query that mentions British Columbia, and the team could only record the approximate location, because a private window does not hide it. The third is change over time in the web index and in Google's ranking systems, shown by the program page that moved; the team could only record the date and save copies of what it screened. Device differences between laptop and phone are a further possible source that could be recorded.

(b) Search engine: Google, Canadian version (google.ca). Date: 14 October 2026, 2:10 pm. Query, exactly as typed: community connector program seniors loneliness BC. Settings: signed out, private window, interface language English, region Canada, SafeSearch default. Searcher location: Kamloops, British Columbia, mobile phone. Stopping rule: first 50 results, set in advance. Displayed estimate recorded separately from the 50 results screened. Results pages saved as PDF files; addresses and access dates recorded for the 3 documents retained.

(c) Transparency means that a reader can see exactly what was searched, when, how and with what limits. Because the ranking and the index change in ways the team cannot see, nobody can guarantee the same results on a rerun, but the record allows a reader to judge what the search could and could not have found.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which source of variation in Google results remains when a reviewer searches while signed out in a private browsing window?

Signing out removes the account's history and saved preferences, and a private window prevents the browser from using stored history. The search engine can still infer an approximate location from the network connection, so location remains a source of variation that must be recorded.

Question 2: What is the difference between a reproducible search and a transparent search?

Reproducibility concerns whether a rerun yields the same results; transparency concerns whether the report shows exactly what was done. Web searches can be made transparent but cannot be made fully reproducible, while database searches can approach both.

Question 3: Why does the order of Google Scholar results for the same query tend to change from month to month even when no new articles are added?

Google Scholar's documentation states that ranking considers how often and how recently a document has been cited. As citations accumulate, positions shift. There is no evidence of deliberate randomization, and Google Scholar does not remove older articles or sort by date by default.

Question 4: A reviewer reports only: "Google Scholar search, about 24,300 results." What is the main problem with this report?

The displayed count is an estimate, and no one can screen it in full. A transparent report gives the date, the exact query, the settings and the number of results actually screened, and that number belongs in the flow diagram in place of the estimate.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson has explained why the same question returns different results in different search tools. Bibliographic databases describe each record with subject headings drawn from a controlled vocabulary, such as MeSH for MEDLINE, Emtree for Embase and CINAHL Subject Headings for CINAHL. Headings gather varied wording under one term and allow a search to include narrower concepts through explosion, but they have gaps: new concepts such as social prescribing receive headings late, some records are never indexed, and indexing varies. A sensitive search therefore combines headings with free-text terms for every concept.

Databases retrieve records with Boolean logic applied to an inverted index, which makes their searches exhaustive, transparent and close to reproducible. Google and Google Scholar rank documents with methods that are documented only in general terms, display only part of their results, and change with the searcher, the place and the day. Recall and precision provide a common language for judging any search. Reviews favour recall because an irrelevant record costs seconds to exclude, while a missed study is invisible and may bias the findings.

For the fictional Cedar Valley review, these principles produce a plan with five databases searched with both headings and free text, a benchmark set to estimate recall, and web searching used as a documented supplement. Lesson 4 turns that plan into a full search string.

Key Takeaways from this lesson

  • A bibliographic database record contains fields written by authors and fields added by the database during indexing, and a search can look in either kind.
  • A controlled vocabulary assigns one preferred heading to each concept, with entry terms, a hierarchy of broader and narrower headings, and scope notes that define its use.
  • MeSH, Emtree and CINAHL Subject Headings index different databases, and their headings and hierarchies differ, so each heading must be checked when a search is translated.
  • Exploding a heading retrieves records indexed with it and with every narrower heading, which matters because indexers assign the most specific heading available.
  • Free-text terms are needed alongside headings because new concepts receive headings late, some records are never indexed, and indexing varies between records.
  • Boolean retrieval returns every record that satisfies the logic of a query, which makes it exhaustive, transparent and repeatable, while relevance sorting only reorders that set.
  • Relevance ranking orders documents by an estimated match to the query, and Google and Google Scholar describe their ranking factors only in general terms.
  • Recall is the proportion of relevant records a search finds, precision is the proportion of retrieved records that are relevant, and the number needed to read equals one divided by precision.
  • Reviews favour recall because missed studies are invisible and can bias findings, and recall is estimated in practice with a benchmark set of known studies.
  • Web searches cannot be fully reproduced because of personalization, location, index changes and undisclosed ranking updates, so reviewers aim to make them transparent instead.

Core Concepts Reviewed

Section 1: indexing, controlled vocabularies and thesauri, MeSH, Emtree and CINAHL Subject Headings, entry terms, explosion, major topics, free-text searching and automatic term mapping.

Section 2: the inverted index, Boolean operators and exhaustive retrieval, relevance ranking with term frequency and inverse document frequency, link analysis, and the limits of Google and Google Scholar for systematic searching.

Section 3: recall, precision and the number needed to read, the trade-off between recall and precision, benchmark sets and relative recall, and the reasons reviews favour recall.

Section 4: reproducibility and transparency, personalization and context, ranking opacity, change over time in web indexes and rankings, and the minimum record for a web search.

The final reflection asks you to explain the lesson's main ideas to a decision-maker in the Cedar Valley case.

Reflection

You are the student intern on the fictional Cedar Valley evidence team, which is preparing a rapid scoping review on community-based interventions for loneliness and social isolation in adults aged 65 and older. The planning director, who is not a researcher, asks why the team needs about five weeks for searching and screening when Google finds information in seconds. Use these facts. The team plans to search MEDLINE, Embase, CINAHL, APA PsycInfo and Web of Science, using subject headings (such as MeSH in MEDLINE) together with free-text words. In a pilot, a broad MEDLINE search retrieved 750 records and found an estimated 90 percent of known eligible studies, while a narrow search retrieved 150 records and found about 60 percent. Google and Google Scholar rank results with methods that are not published, show results that vary by user, place and day, and cannot be exported in full; Google Scholar displays at most about 1,000 results.

Write a reply of 200 to 300 words that explains (a) how bibliographic databases and web search engines find information differently, (b) what recall and precision mean and why the team favours recall, and (c) what role web searching will play and how it will be documented.

Model answer

Thank you for the question. Google and the research databases we use work in different ways. Databases such as MEDLINE describe each article with standard subject headings, so a study that says "lonely seniors" and one that says "isolated older adults" can both be found under the same headings. They return every record that matches our search logic, give an exact count, and let us save the whole set. Google and Google Scholar rank pages using methods they do not publish, and the order changes with the person searching, their location and the day. Only part of the list can be viewed, and none of it can be exported in full.

We judge a search by two measures. Recall is the share of all relevant studies that the search finds; precision is the share of what it returns that is relevant. In our pilot, a narrow search gave us 150 records to read but found only about 60 percent of the studies we know are relevant. A broad search gave us 750 records and found about 90 percent. We favour recall because a missed study is invisible: we would never know it was absent, and if missed studies differ from the ones we find, the brief could mislead your planning. Reading extra records takes time but carries no such risk.

Web searching still matters, because many program descriptions and local evaluations exist only online. We will use Google and Google Scholar as supplementary sources, screen a fixed number of results set in advance, and record the date, exact wording and settings of each search, saving copies of what we screen, so that you and others can see exactly what we did.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: How Databases and Search Engines Find Information (15 Questions)

Question 1: A team's MEDLINE search uses only free-text terms in titles and abstracts. Which relevant record is it most likely to miss?

A free-text search finds only the words the team listed. A record that uses an unlisted synonym is missed, even though a subject-heading search would have found it through indexing. The other three records contain the team's words and would be found.

Question 2: Which pairing correctly matches each thesaurus to the database it indexes?

MeSH is the National Library of Medicine's thesaurus for MEDLINE, and Emtree is Elsevier's thesaurus for Embase. CINAHL has its own subject headings, APA PsycInfo uses the APA Thesaurus, and Web of Science has no subject-heading thesaurus of this kind.

Question 3: Search A in the Cedar Valley example looked for its words in titles only. If the librarian extends it to titles and abstracts, which change is most likely?

Searching more fields retrieves more records, including relevant ones whose titles lack the words, so recall rises. Many of the additional records mention the words in passing, so precision usually falls. This is the trade-off shown by the recall and precision curves.

Question 4: Which feature makes Boolean retrieval in a bibliographic database suitable for the main search of a systematic review?

Exhaustive, transparent and repeatable retrieval is what a review's main search needs. Ranking on the first page and tolerance of partial matches describe relevance ranking, and adjustment to location and history describes web search personalization.

Question 5: A search retrieves 2,000 records, of which 50 are eligible. What are its precision and number needed to read?

Precision is 50 ÷ 2,000 = 0.025, or 2.5 percent. The number needed to read is 2,000 ÷ 50 = 40. The other options misplace the decimal point or invert the ratio.

Question 6: Why does an exploded MeSH heading usually increase recall?

Explosion follows the hierarchy downward. Because indexers assign the most specific heading available, an unexploded broad heading misses records indexed only with its narrower headings. Restricting to major topic lowers recall, and headings do not transfer between databases.

Question 7: According to Google's public documentation, which statement about Google's ranking is accurate?

Google describes factors including the meaning of the query, relevance, quality, usability and the user's context, and says their weight varies with the query. It does not publish the signals in full or their weights. PageRank was part of the original design and is now one signal among many.

Question 8: The Cedar Valley intern checks a draft MEDLINE search against 20 known eligible studies and finds that 2 are missed. What is the most useful next step?

Missed benchmark studies show which words or headings the search lacks. Removing a concept with AND would increase rather than decrease the records retrieved, a major-topic restriction would lower recall further, and Google Scholar cannot serve as the main search.

Question 9: Why is specificity rarely used to evaluate a search of a bibliographic database?

A database holds millions of irrelevant records, and any search leaves out almost all of them, so specificity cannot separate good searches from poor ones. Specificity can be defined for retrieval, but it differs from precision and is uninformative here.

Question 10: A rapid review team decides to search two databases instead of five. How should this decision be handled?

Rapid reviews may accept lower recall for speed, but the shortcut must be reported so that readers can judge its effect (Lesson 10). Precision does not restore recall, and Google Scholar is unsuitable as a principal search system.

Question 11: Which pair of actions best improves the transparency of a Google search for grey literature?

Transparency requires a record of what was searched, when and how, and copies of what was seen. Being signed in increases personalization, the estimated count is not the number screened, and Google supports neither subject headings nor bulk export.

Question 12: Why can a PubMed search that uses only MeSH headings miss records that a free-text search finds?

PubMed includes records that are not indexed for MEDLINE, such as some publisher-supplied records and PubMed Central articles from journals outside MEDLINE. These records have no headings, so only free text can find them. Headings are applied to all designs and are not removed with age.

Question 13: In a toy collection, lonely appears in record 3, loneliness in records 1, 4 and 6, and older adults in records 1, 2, 3, 5 and 8. What does (loneliness OR lonely) AND older adults return?

The OR step gives records 1, 3, 4 and 6. Intersecting that set with the older adults list (1, 2, 3, 5, 8) leaves records 1 and 3. Record 1 alone is the result without lonely, which shows why synonyms raise recall.

Question 14: Google Scholar states that its ranking weighs the full text, the source, the authors and citations. What consequence might this have for a scoping review of a new type of program?

Documents that have had little time to gather citations, including new studies, reports and theses, tend to rank lower than older, highly cited articles. Citations are one factor among several, and no ranking guarantees that all relevant documents appear early.

Question 15: Which combination of methods fits the Cedar Valley team's aim of high recall within its twelve-week timeline?

High recall requires several bibliographic databases searched with both controlled vocabulary and free text. Web searching adds grey literature as a documented supplement with a stopping rule. The other options sacrifice recall or rely on tools that cannot serve as the principal search.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: How Databases and Search Engines Find Information.

You can now explain how bibliographic databases index and retrieve records, read a thesaurus entry and decide whether to explode a heading, predict the results of a Boolean query, describe in accurate general terms how web search engines rank results, calculate recall, precision and the number needed to read, and record a web search so that readers can judge it.

Lesson 4 puts these ideas to work. You will build a concept table, write a full search string with Boolean operators, truncation, phrase searching and field tags, translate it across MEDLINE, Embase, CINAHL, APA PsycInfo and Web of Science, test it against your benchmark set, and prepare it for peer review with PRESS and reporting with PRISMA-S.

Continue to Lesson 4 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

These terms, tools and people appear in this lesson on how databases and search engines find information.

Core Concepts
Bibliographic database An organized collection of records, each describing one published item with fields such as title, abstract, authors, journal and subject headings, as in MEDLINE, Embase and CINAHL.
Indexing The process of describing each record with standard subject terms from a controlled vocabulary, carried out by people, by software or by both.
Controlled vocabulary A fixed list of approved terms used to describe the subject of records; when the terms are organized with relationships among them, the list is called a thesaurus.
Subject heading The preferred term for a concept in a controlled vocabulary, also called a descriptor, such as the MeSH heading Loneliness.
Entry term A synonym or variant spelling that points to a preferred heading, such as Elderly for the MeSH heading Aged.
Explode A search option that retrieves records indexed with a heading and with every narrower heading beneath it in the hierarchy.
Major topic A restriction that retrieves only records in which a heading describes a main point of the article; it raises precision and lowers recall.
Free-text search A search for words in the fields written by authors, usually the title, abstract and author keywords; also called a keyword or text-word search.
Inverted index A structure that lists, for each word or heading in a collection, the records that contain it, so that a query can be answered by comparing lists.
Boolean retrieval Retrieval that treats a query as a logical statement with AND, OR and NOT and returns every record for which the statement is true.
Relevance ranking Retrieval that scores documents by how well they appear to match a query and presents them in order of that score, including partial matches.
Term frequency In ranking, the number of times a query word appears in a document; more occurrences give a higher score, usually with a limit on the effect of repetition.
Inverse document frequency In ranking, a weight that is higher for words that appear in few documents in the collection, calculated from the logarithm of the number of documents divided by the number containing the word.
Recall The proportion of all relevant records in a source that a search retrieves; equivalent to sensitivity in diagnostic testing.
Precision The proportion of retrieved records that are relevant; similar to the positive predictive value of a diagnostic test.
Number needed to read The average number of retrieved records a reviewer must read to find one relevant record, equal to one divided by precision.
Benchmark set A collection of known relevant studies, gathered independently of the search being tested, used to estimate recall; also called a validation set or test set.
Relative recall An estimate of recall that treats the pooled relevant studies found by all search methods in a review as the reference standard for each individual search.
Personalization The adjustment of search results to the person searching, using information such as location, search history and settings.
Ranking opacity The condition in which the method used to order search results is not disclosed in enough detail for an outsider to predict or check it.
Reproducibility The property of a search that allows another person following the documented method to obtain the same results, or to explain any differences.
Transparency The extent to which a report shows exactly what was searched, when, how and with what limits, whether or not the search can be repeated with the same results.
Frameworks & Tools
MeSH (Medical Subject Headings) The thesaurus of the United States National Library of Medicine, used to index MEDLINE, with more than 30,000 descriptors arranged in tree structures and revised every year.
Emtree Elsevier's thesaurus for indexing Embase, larger than MeSH and particularly detailed for drugs, medical devices and pharmacology.
CINAHL Subject Headings The thesaurus used to index CINAHL, based on the structure of MeSH with additional headings for nursing and allied health.
APA Thesaurus of Psychological Index Terms The American Psychological Association's controlled vocabulary for indexing APA PsycInfo.
Automatic term mapping PubMed's process of matching untagged search words to MeSH headings, journals and authors and combining them with an all-fields search.
PubMed Best Match PubMed's default relevance sort, based on a machine-learning ranking model, which changes the display order of retrieved records without changing the set.
PageRank A link-analysis method described by Brin and Page in 1998 that scores a web page higher when many well-linked pages link to it.
BM25 A widely used ranking formula from the probabilistic relevance framework that combines term frequency, inverse document frequency and document length.
PRISMA-S An extension of the PRISMA statement for reporting literature searches in systematic reviews, published by Rethlefsen and colleagues in 2021.
Key People
George Boole English mathematician (1815 to 1864) whose 1854 book The Laws of Thought set out the algebra of logic on which Boolean searching is based.
Cyril Cleverdon British librarian who led the Cranfield experiments in the late 1950s and 1960s, which established test collections and the use of recall and precision to evaluate retrieval.
Gerard Salton Computer scientist at Cornell University who developed the vector space model and the SMART retrieval system, foundations of relevance ranking.
Karen Spärck Jones British computer scientist at the University of Cambridge who proposed inverse document frequency in 1972.
Stephen Robertson British information scientist whose work on probabilistic retrieval models led to the BM25 ranking formula.
Sergey Brin and Lawrence Page Founders of Google who, as Stanford graduate students, described the prototype search engine and the PageRank method in 1998.
Eli Pariser Author who popularized the term filter bubble in his 2011 book on personalization in online services.
Melissa Rethlefsen Medical librarian and lead author of PRISMA-S, the reporting guideline for literature searches in systematic reviews.
No matching entries. Try a different search term.