# Lesson 3: How Databases and Search Engines Find Information

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are on Lesson three of Health Sciences two forty-one, which is about how databases and search engines find information.

**Sarah:** That sounds like a technical lesson for a course about evidence. Why does a student who wants to do a review need to know how search tools work on the inside?

**Kiffer:** Because the same words typed into PubMed, Embase, Google and Google Scholar give four different sets of results, and a reviewer has to decide which of those sets to trust. If you know how each tool stores and finds information, you can build a search that finds what it should and explain to a reader why you searched where you did.

**Sarah:** Let's set up the case first, since it runs through the whole course.

**Kiffer:** The case is the fictional Cedar Valley Health Authority in British Columbia. It serves about two hundred and ten thousand people, about forty-six thousand of them aged sixty-five or older. The planning team wants to launch a community connector program, which is a form of social prescribing, and it has asked a small evidence team for a rapid scoping review and an environmental scan within twelve weeks. The evidence team is an evidence officer, a university librarian and a student intern, and this week the librarian explains to the intern how the databases they will search actually work.

**Sarah:** So start where the lesson starts. What is indexing?

**Kiffer:** Indexing is the process of describing each record in a database with standard subject terms. A record in a bibliographic database, such as MEDLINE, has two kinds of fields. Some come from the authors: the title, the abstract and any keywords they chose. Others are added by the database, such as subject headings, the publication type and the language. Adding those is indexing.

**Sarah:** Why would the database add its own words when the authors have already written an abstract?

**Kiffer:** Because authors describe the same thing in many different ways. In the reading's example, the authors wrote lonely, seniors and older adults. Indexing replaces that variety with one standard heading per concept, so every article about loneliness in people over sixty-five carries the same headings, whatever words the authors chose.

**Sarah:** And the list of standard headings is the controlled vocabulary.

**Kiffer:** Right. A controlled vocabulary is a fixed list of approved terms, and when the terms are organized with relationships among them, we call it a thesaurus. The best known in health is the Medical Subject Headings, usually shortened to MeSH, which the United States National Library of Medicine uses to index MEDLINE. It has more than thirty thousand headings and is revised every year.

**Sarah:** Who actually assigns the headings? I picture someone reading every article.

**Kiffer:** For decades that was largely how it worked. Since April of twenty twenty-two, the library has indexed all MEDLINE journals with an automated system, and human indexers review selected sets of citations for quality.

**Sarah:** The reading lists several features of a thesaurus. Walk me through them.

**Kiffer:** Each concept has one preferred term, which is the heading. Each heading has entry terms, which are synonyms that point to it, so Elderly is an entry term for the heading Aged. The headings sit in a hierarchy of broader and narrower terms. Each heading also has a scope note that defines how it is used.

**Sarah:** The hierarchy is where explode comes in. What happens if I search a broad heading without exploding it?

**Kiffer:** Exploding a heading retrieves records indexed with that heading and with every heading beneath it. In MeSH, Social Isolation has four narrower headings: Loneliness, Ostracism, Social Alienation and Social Deprivation. Indexers use the most specific heading that fits, so a loneliness trial gets the heading Loneliness. If you search Social Isolation without exploding it, you miss that trial.

**Sarah:** That seems like an easy mistake to make.

**Kiffer:** It is, partly because the platforms behave differently. PubMed explodes headings by default. Ovid, which many libraries use for MEDLINE, explodes a heading only when you type the letters e x p in front of it.

**Sarah:** What about the other databases?

**Kiffer:** Each has its own thesaurus. Embase uses Emtree, which Elsevier maintains, and it is especially detailed for drugs and devices. CINAHL, the Cumulative Index to Nursing and Allied Health Literature, uses the CINAHL Subject Headings, which follow the structure of MeSH and add headings for nursing and allied health. Web of Science has no subject-heading thesaurus of this kind.

**Sarah:** So a MeSH heading can't simply be copied into Embase.

**Kiffer:** No. The headings, the hierarchies and the syntax all differ, so every heading has to be looked up again in each database. That translation is a large part of Lesson four.

**Sarah:** If subject headings solve the wording problem, why does anyone bother with free text?

**Kiffer:** Because indexing has gaps. The clearest example in our case is social prescribing. MeSH added Social Prescribing as a heading in twenty twenty-five, even though studies of social prescribing had been published for years. Older studies were indexed with other headings, such as Social Support, and generally don't carry the new one. To find them you search the words authors used, such as social prescribing, link worker and community connector.

**Sarah:** What are the other gaps?

**Kiffer:** Some records in PubMed are never indexed for MEDLINE, so they have no headings. Indexing varies between similar articles, so a walking group study might be indexed under Exercise and Social Support without any loneliness heading. And many intervention names, like befriending or men's sheds, have no heading of their own. So the standard practice, which the Cochrane Handbook describes, is to search each concept with its headings and a set of free-text terms joined by OR, and then to join the concepts with AND.

**Sarah:** You mentioned that PubMed does something automatic with the words you type.

**Kiffer:** It's called automatic term mapping. If you type words without field tags, PubMed tries to match them to MeSH headings, journal names and authors, and it combines the matched heading with a search of the words in all fields. Because an untagged search does more than you typed, systematic searchers tag every term explicitly.

**Sarah:** How did the intern take all this?

**Kiffer:** The intern had typed social prescribing loneliness older adults into PubMed and liked the first page. The librarian opened the MeSH record for Social Prescribing, showed the year it was introduced, and pointed out that a link worker study from twenty nineteen would carry headings like Social Support and Loneliness instead. The intern left with a draft list of headings and free-text words for each concept.

**Sarah:** Let me summarize the first section. A database record has fields from the authors and fields added by indexing. Controlled vocabularies gather synonyms under one heading and let you include narrower concepts by exploding. Each database has its own thesaurus, and free text covers the gaps, especially for new concepts like social prescribing.

**Kiffer:** That's it exactly.

**Sarah:** Section two is about what happens when you press search. Start with the database.

**Kiffer:** A database with tens of millions of records can't read each one every time someone searches, so it builds an inverted index in advance. For every word and every heading, it keeps a list of the records that contain it, called a postings list. When you search, the system works with those lists of numbers, which is why it's fast.

**Sarah:** And presumably why the word lonely isn't found by a search for loneliness.

**Kiffer:** Right. They're separate entries with separate lists. Unless you ask for both, or use truncation, which Lesson four covers, the system has no way to know they're related.

**Sarah:** The reading has a toy collection of eight records. Can you talk through it without the table in front of us?

**Kiffer:** Imagine that loneliness appears in records one, four and six, and older adults appears in records one, two, three, five and eight. Lonely appears only in record three. If you search loneliness AND older adults, the system looks for numbers on both lists, and only record one is on both. AND narrows, because it keeps the intersection.

**Sarah:** And OR broadens.

**Kiffer:** OR keeps anything on either list. If you search loneliness OR lonely OR social isolation, and then combine that with older adults, you get records one, two and three. Record three, which says lonely, is now found.

**Sarah:** And NOT?

**Kiffer:** NOT removes anything on the second list, and it's the risky one. In the toy collection, older adults NOT loneliness throws away record one, which is the most relevant record in the collection.

**Sarah:** Why does this kind of logic matter so much for reviews?

**Kiffer:** Boolean retrieval, named after the mathematician George Boole, returns every record for which your query is true. That gives database searching three properties. It's exhaustive, because every matching record comes back and the count is exact. It's transparent, because anyone reading your strategy can see why a record was or wasn't retrieved. And it's repeatable, because the same strategy in the same database on the same date returns the same set.

**Sarah:** PubMed shows results sorted by Best Match now, though. Isn't that ranking?

**Kiffer:** It's a sort order. PubMed retrieves the Boolean set first and then uses a machine-learning model to decide the order in which those records appear. The set and the count stay the same whichever sort you choose. If you screen every record, the sort doesn't matter. If you only look at the first page, it matters a great deal.

**Sarah:** So what's different about Google?

**Kiffer:** Google uses relevance ranking. Instead of deciding whether each document matches, it gives each one a score and shows them in order. Documents that match only some of your words can still appear, lower down.

**Sarah:** How is the score worked out?

**Kiffer:** Gerard Salton and his colleagues developed the vector space model, and two ideas from that line of work are still central. Term frequency gives more weight to a document that uses your word more often. Inverse document frequency, which Karen Spärck Jones proposed in nineteen seventy-two, gives more weight to words that are rare in the collection, because a rare word says more about what a document is about.

**Sarah:** Can you put numbers on that?

**Kiffer:** In the reading's example, out of a million documents, the word adults appears in two hundred thousand and the word loneliness in two thousand. Loneliness ends up with roughly four times the weight. So a document that mentions loneliness six times and adults twice outranks one that mentions loneliness once and adults ten times.

**Sarah:** And web search adds links. Is PageRank still how Google works?

**Kiffer:** Sergey Brin and Lawrence Page described PageRank in nineteen ninety-eight. It treats a link from one page to another as a kind of endorsement, and it was part of Google's original design. Beyond that I want to be careful, because many confident claims about Google are guesses. Google's public documentation says its systems consider the meaning of the query, the relevance of content, signals of quality, the usability of pages, and the user's context and settings, and that the weight of each factor varies by query. It doesn't publish the signals in full or their weights, and it changes the systems continually.

**Sarah:** What about Google Scholar?

**Kiffer:** Google Scholar's own help pages say it tries to rank documents the way researchers would, weighing the full text, where a document was published, who wrote it, and how often and how recently it has been cited. An early analysis by Beel and Gipp concluded that citation counts carry heavy weight. So a new evaluation report may sit far down the list even if it answers your question directly.

**Sarah:** So could a reviewer use Google Scholar as the main search?

**Kiffer:** The published evaluations say no. Haddaway and colleagues in twenty fifteen pointed out that Google Scholar shows at most the first thousand results of any search. It doesn't support truncation, its support for grouping terms with parentheses is limited, and it has no official bulk export. Gusenbauer and Haddaway in twenty twenty tested twenty-eight academic search systems and concluded that Google Scholar was unsuitable as a principal search system for systematic reviews, although useful as a supplement.

**Sarah:** And the intern tried this for real.

**Kiffer:** A short tagged PubMed search gave a fixed, exportable set of a few hundred records, while similar words in Google Scholar produced an estimate of tens of thousands of results. The librarian explained that the first is the exact size of a defined set, and the second is an estimate attached to a list that no one can read in full.

**Sarah:** Let me try to summarize. Databases use an inverted index and Boolean logic, so they return an exact set that can be checked and rerun. Search engines score and rank, which is convenient but hides most of the list and the method. So the main search runs in databases, and Google Scholar is a supplement.

**Kiffer:** Yes. And Bramer and colleagues showed in twenty seventeen that no single database retrieved all the studies included in a large set of reviews, which is why the Cedar Valley team plans to search five.

**Sarah:** That leads into Section three. How do you know whether a search is any good?

**Kiffer:** You ask two questions. How many of the relevant records did it find, and how much of what it returned is relevant? The first is recall and the second is precision. They became the standard pair through the Cranfield experiments that Cyril Cleverdon led in England in the late nineteen fifties and the nineteen sixties. Recall is the number of relevant records you retrieved divided by all the relevant records in the database. Precision is the number of relevant records you retrieved divided by everything you retrieved.

**Sarah:** Those sound like sensitivity and positive predictive value.

**Kiffer:** They are close relatives. Recall is the same idea as sensitivity in a diagnostic test, and precision resembles positive predictive value. There's also a third measure, the number needed to read, which Bachmann and colleagues proposed. It's one divided by precision, and it tells you how many records you read, on average, to find one relevant record.

**Sarah:** Why not use specificity as well?

**Kiffer:** Because a database contains millions of irrelevant records and every search leaves almost all of them out. Specificity comes out above ninety-nine point nine nine percent for good and bad searches alike, so it can't tell them apart.

**Sarah:** Let's do the worked example.

**Kiffer:** It's a teaching device, because it pretends we know the true number of relevant studies. Suppose MEDLINE contains exactly thirty studies that meet the Cedar Valley criteria. Search A is narrow: it looks for loneliness, older adults and intervention in titles only. It retrieves one hundred and fifty records and finds eighteen eligible studies. Search B is broad: it combines subject headings with free-text synonyms in titles and abstracts. It retrieves seven hundred and fifty records and finds twenty-seven.

**Sarah:** So Search A has a recall of eighteen out of thirty, which is sixty percent.

**Kiffer:** And a precision of eighteen out of one hundred and fifty, which is twelve percent, so the number needed to read is about eight. Search B has a recall of ninety percent and a precision of three point six percent, so the number needed to read is about twenty-eight.

**Sarah:** Those precision figures look terrible.

**Kiffer:** They look poor to someone used to Google's first page, and they're normal for a review search. What matters is the cost. At thirty seconds a record, with two reviewers screening independently, Search A takes two and a half hours in total and Search B takes twelve and a half. That's about ten extra hours to find nine more studies. And Search A misses forty percent of the eligible studies, so an evidence brief built on it would rest on little more than half the evidence in MEDLINE.

**Sarah:** Is it always a trade like that?

**Kiffer:** Almost always. As you broaden a search, recall climbs quickly at first and then levels off, while precision keeps falling. Adding synonyms with OR, exploding headings and adding free text push recall up and precision down. Adding a concept with AND, searching titles only or restricting to major topics does the opposite.

**Sarah:** Here's my pushback. In a real review nobody knows there are thirty eligible studies. So how do you ever calculate recall?

**Kiffer:** You estimate it. The usual method is a benchmark set, which is a group of known eligible studies collected independently of the search, for example from earlier reviews or from experts.

**Sarah:** What did the Cedar Valley team do?

**Kiffer:** The intern collected twenty eligible studies from two earlier reviews and confirmed that all twenty were in MEDLINE. Search A caught twelve, an estimated recall of sixty percent, and Search B caught eighteen, or ninety percent. The librarian then read the two that Search B missed. One described its intervention only as befriending and its outcome as social participation. The other had been indexed under broader headings without Loneliness, and its abstract used the phrase social disconnection. Those terms went into the free-text lists.

**Sarah:** Isn't there a danger of tuning the search to the studies you already know?

**Kiffer:** There is, and it's a good point. A search adjusted to fit a known set can look better than it really is, so the librarian kept ten more known studies aside as a final check, the test set in Lesson four. There's also a second method, called relative recall, from Sampson and colleagues, which pools every eligible study found by every method in the review and uses that pool as the reference standard.

**Sarah:** So why do reviews favour recall, when it costs so much time?

**Kiffer:** Because the two kinds of error don't cost the same. An irrelevant record costs a few seconds at screening, and then it's gone. A missed study costs nothing visible at the time, because nobody knows it's missing, and you can't recover it later unless another method happens to find it.

**Sarah:** There's also a bias argument in the reading.

**Kiffer:** Yes. The studies a narrow search misses can differ systematically from the ones it finds. They may come from another discipline or report small or null results. If they differ in their findings, the review is biased. Chapter four of the Cochrane Handbook advises that review searches aim for high sensitivity and accept that precision will be low.

**Sarah:** But the Cedar Valley team only has twelve weeks.

**Kiffer:** Right, and that constraint is real. So the team keeps recall high in the database searches and manages the workload in other ways. It removes duplicates before screening, pilots the eligibility criteria so decisions are quick, uses a screening platform, and may use active-learning tools, which Lesson six covers. Rapid reviews sometimes accept lower recall, for example by searching fewer databases, and when they do, they report it. That's Lesson ten.

**Sarah:** One more thing. Students tend to judge a search by its first page.

**Kiffer:** That's a judgement about precision at the top of a ranked list, and it tells you nothing about recall. A search with an excellent first page can still miss a large share of the eligible studies.

**Sarah:** Section four is about reproducibility. What does that mean for a search?

**Kiffer:** A search is reproducible when another person, following your documented method, can run it again and get the same results, or explain the differences. It matters because readers need to check the search and because someone will want to update the review later.

**Sarah:** And the lesson separates that from transparency.

**Kiffer:** Transparency is about the report. A transparent search is one where the report shows exactly what was searched, when, how and with what limits, so you can have full transparency without reproducibility. The literature search extension of the Preferred Reporting Items for Systematic reviews and Meta-Analyses, which everyone calls PRISMA-S, was published by Rethlefsen and colleagues in twenty twenty-one, and it aims at both.

**Sarah:** Are database searches fully reproducible?

**Kiffer:** They come close, though they fall short of perfect. Databases keep adding records, thesauri change every year, as the social prescribing heading showed, and platforms change. Still, a rerun with the original date limits gives a very similar set, and the differences have known causes.

**Sarah:** And web searches?

**Kiffer:** Web searches are a different matter, and the first reason is personalization. Google's own documentation says its systems use your location, your search history and your settings, that the language of your query shapes the language of the results. Hannak and colleagues measured personalization in twenty thirteen by comparing real accounts with controlled test accounts. They found measurable personalization, with being signed in and geographic location as the main sources of difference.

**Sarah:** How big is the effect?

**Kiffer:** Studies have generally found it present and modest on average, which is smaller than the phrase filter bubble, popularized by Eli Pariser, might suggest. For evidence synthesis, though, any personalization means two reviewers can see different results for the same query, and neither can tell from the page which results were affected. The existence of the effect matters more than its average size.

**Sarah:** Can't you simply sign out and use a private window?

**Kiffer:** That helps with the account and the browser history. It doesn't hide your approximate location, which the search engine infers from your network connection, and the country version, the interface language and some settings still matter. You can control some of these and record the rest, and you can't remove them entirely.

**Sarah:** The second reason is ranking opacity.

**Kiffer:** Ranking opacity means the method for ordering results isn't disclosed in enough detail for an outsider to predict or check it. The consequence for a reviewer is that nobody can explain why a page sits in position three, or whether it would appear at all on a rerun. And the systems change continually, as does the web index, as pages are added, moved and deleted.

**Sarah:** Does Google Scholar drift in the same way?

**Kiffer:** It has its own version of the problem. Because it ranks partly on citation counts, the order of results drifts as articles collect citations, even if nothing new is published. And since it shows only about the first thousand results, articles near that edge drift in and out of view.

**Sarah:** What about the number at the top, the about so many results?

**Kiffer:** That's an estimate, and it can even change between pages of the same search, so it should never be reported as the size of a result set. The page may also carry advertisements and, since twenty twenty-four, summaries generated by artificial intelligence, each produced by its own undisclosed process.

**Sarah:** Tell me what happened in Cedar Valley.

**Kiffer:** The evidence officer in Cedar City and the intern in Burnaby ran the same Google query on the same afternoon. The officer was signed out, and the intern was signed in to an account used for weeks of searching on social prescribing. Their first pages shared only three of ten results, in different positions. When the intern reran the search two weeks later, one program page had moved and a new news story had appeared. So the team concluded that it couldn't make its web searches reproducible, and it decided to make them transparent instead.

**Sarah:** What does that look like in practice?

**Kiffer:** Record the search engine and its country version, the date, and the query exactly as typed. Search signed out in a private window, set the language and region deliberately, and record those settings and your approximate location. Decide before you start how many results you'll screen, such as the first hundred. Save the results pages and copies of anything you keep, because pages disappear. Report the displayed count as an estimate, and report separately the number you actually screened, which is the figure that goes into the flow diagram in Lesson seven.

**Sarah:** Given all of these problems, is web searching worth doing?

**Kiffer:** It is. Much of what the Cedar Valley scan needs, such as program descriptions, funding announcements and evaluation reports from community organizations, exists only on the open web and will never appear in MEDLINE. The aim is to use web searching as a documented supplement, and Lesson five gives a full template for documenting it.

**Sarah:** And this connects to the lesson on artificial intelligence.

**Kiffer:** Directly. AI search engines and research assistants usually start with a retrieval step that resembles relevance ranking, so they inherit the same opacity. Then they generate text, which can differ each time and sometimes cites works that don't exist. Lesson six takes that up, and the same logic applies there: record what you asked, when, and with which tool, and check every source against the original.

**Sarah:** Let me pull the whole lesson together. Databases describe records with controlled vocabularies and find them with Boolean logic, so their searches are exhaustive, transparent and close to reproducible. Search engines rank with methods they don't publish, and the results change with the person, the place and the day. Recall and precision let you judge any search, and reviews favour recall because missed studies are invisible and can bias the findings.

**Kiffer:** That's a good summary of the whole lesson.

**Sarah:** So before anyone builds a full search, what groundwork does a team prepare?

**Kiffer:** Three things, which feed the database search and the grey literature and web search plan. First, for each main concept in the question, the team looks up the MeSH heading and the matching CINAHL Subject Heading. It records the scope note, entry terms and narrower headings, decides whether to explode each one, and adds free-text synonyms.

**Sarah:** And the other two?

**Kiffer:** Second, the team assembles a benchmark set of three to five studies it already knows meet its eligibility criteria, and confirms that each is in PubMed. Third, it runs one test search in Google Scholar while signed out, records the date, the exact query, the settings, the estimated count and the first five results, and reruns it a few days later to see what changed. Lesson four uses all of that to build the full search and test it against the benchmark set.

**Sarah:** Thanks, Kiffer.

**Kiffer:** Thanks, Sarah. See everyone next week.
