# Lesson 4: Building and Documenting a Database Search

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are talking about the database search, which is the step in a review that students tend to underestimate.

**Sarah:** Why underestimate? Most of us have typed words into a library database before.

**Kiffer:** Most people have, and that is part of the problem. A review search is a planned piece of method. Another team should be able to read it, judge whether it was likely to find the relevant studies, and run it again. A search typed on the fly can't meet that standard, because nobody, including the person who typed it, can say exactly what it did. If a search misses a group of relevant studies, it still returns hundreds of records, so nothing in the results tells you what is missing.

**Sarah:** So the lesson is about building habits that protect against that.

**Kiffer:** Four habits: plan the search, write it carefully, have someone check it, and report it fully.

**Sarah:** Let's remind listeners of our running case.

**Kiffer:** The Cedar Valley Health Authority is a fictional health authority in British Columbia. Its planning team wants to launch a community connector program, which is a form of social prescribing, for older adults. A small evidence team, made up of an evidence officer, a university librarian and a student intern, is doing a rapid scoping review. The question asks which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged sixty-five and older, and with what outcomes. The team wrote it in lesson two using the population, concept and context framework, and this week the intern and the librarian turn it into a search.

**Sarah:** So where does that start? With the databases?

**Kiffer:** It starts with deciding which parts of the question become search concepts. A search concept is one idea that a relevant record has to address, searched as a group of interchangeable terms. Concepts are joined with AND, which means a record only comes back if it has at least one term from every concept. Each added concept therefore removes the studies whose titles and abstracts happen to describe that idea in words you didn't search. That's why most review searches use two to four concepts and leave the rest of the question to the screeners.

**Sarah:** How did Cedar Valley decide which parts to search?

**Kiffer:** The test I teach is to ask, for each part of the question, whether the authors of a relevant study would almost always describe it in the title, the abstract or the subject headings. Older adults and loneliness pass that test easily. Community-based interventions are described more loosely, but without that concept the search would pull in thousands of studies that measure loneliness without testing anything. So the team searched three concepts: loneliness or social isolation, older adults, and community-based interventions.

**Sarah:** What about the context? The protocol says community settings in high-income countries.

**Kiffer:** That fails the test. A paper about a seniors' centre program may never use the word community, and plenty of abstracts don't name the country. If you search for setting, you lose those studies. So the screeners apply the context when they read the records.

**Sarah:** And words like evaluation or effectiveness? The question is about evaluated interventions.

**Kiffer:** That's a tempting fourth concept, and the team decided against it. Many evaluations describe themselves only by their design or by what they measured. A before-and-after study of a walking group might never use the word evaluation in its abstract. Adding those words would quietly remove studies the review wants.

**Sarah:** Once you know the concepts, what's the concept table?

**Kiffer:** It's a grid. For each concept you list the subject headings, the free-text terms and notes on your decisions. It puts the whole plan on one page before any syntax is written, so a colleague can check it, and it becomes the starting point when you translate the search for other databases.

**Sarah:** Remind us what subject headings and free-text terms are.

**Kiffer:** Subject headings come from a controlled vocabulary, a fixed list of terms that indexers assign to records. In MEDLINE that vocabulary is called Medical Subject Headings, or MeSH. Free-text terms are words you search in the title, abstract and author keywords. Lesson three explained why both exist. Headings gather studies that use different words for the same idea, and free text finds records whatever headings they were given.

**Sarah:** Give me an example from the Cedar Valley table.

**Kiffer:** For older adults, the heading is Aged, exploded, and we will come back to what exploded means. The free-text terms include older adults, older people, older persons, seniors, elderly, geriatric, later life and old age. One word the team deliberately left out was aged. It works well as a heading, but abstracts use it to report the ages of every group, such as children aged five to twelve or adults aged eighteen to forty. Searching it in the abstract would pull in a huge number of irrelevant records.

**Sarah:** Where do the synonyms come from? I imagine students just brainstorm.

**Kiffer:** Brainstorming is a start, and it isn't enough. The best source is the records of studies you already know are relevant. Look at their titles, abstracts and author keywords, and look at the subject headings the indexers gave them. The thesaurus itself helps, because every heading lists entry terms, which are synonyms that point to it. The MeSH record for Aged lists elderly as an entry term, for instance. Other sources are the search strategies published with earlier reviews on similar questions and the people who run programs, who know the everyday names. The Cedar Valley planners pointed out that local programs use names like friendly visiting and seniors' peer support.

**Sarah:** Are there tools that help?

**Kiffer:** There are free tools, such as the Yale MeSH Analyzer and PubMed PubReMiner, that summarize the headings and frequent words across a set of records. After that you check spelling variants, plurals, hyphens and older terms. Elderly is a good example of an older term. A team might choose to write about older adults in its own report, for good reasons, but the search still has to include elderly, because a large share of the literature on this population used that word.

**Sarah:** You mentioned some words cause trouble because they mean other things.

**Kiffer:** Isolation is the obvious one. It also describes laboratory isolation of organisms and isolation in infection control. Connector is another, since it names parts in engineering. The answer is to keep the word and tie it to a context word, so isolation only counts when it is near words for people or social life. We'll see how to do that in section two.

**Sarah:** Now tell us what exploded means, as you promised.

**Kiffer:** Headings sit in hierarchies called trees, with broader headings above narrower ones. To explode a heading is to search it along with everything beneath it. In MEDLINE, Aged has two narrower headings, Aged, eighty and over, and Frail Elderly. Explode Aged and you get all three. Search Aged without exploding and you only get records indexed with Aged itself.

**Sarah:** Is there ever a reason not to explode?

**Kiffer:** Sometimes there is, if a narrower heading is clearly outside the question. For a review, the default is to explode headings that have useful narrower terms. The opposite move is major-topic restriction, where you only accept records in which the heading is marked as a main topic. That raises precision and lowers recall, and reviews favour recall, so it's rarely used.

**Sarah:** Why not rely on headings alone, if they are so well organized?

**Kiffer:** Headings are assigned after a record arrives, and some records are never indexed at all, so the newest studies are often the ones without headings. New ideas are a second problem. Social prescribing is a relatively recent term, and vocabularies change every year, so the librarian checks the current MeSH browser for a matching heading. Whatever the librarian finds, the search keeps free-text terms for it, because older studies describe the same activity in other words. The rule is headings and free text together, combined with OR, inside every concept.

**Sarah:** How do you know whether the draft is any good?

**Kiffer:** You test it against a test set, which is a group of studies you already know are relevant. Cedar Valley had ten. Nine were indexed in MEDLINE, and the first draft found eight of them.

**Sarah:** Eight of nine sounds pretty good.

**Kiffer:** It sounds good, and the ninth tells you something useful. The missed study described its outcome only as social connectedness, and it was indexed under headings the draft didn't include. So the intern added a line for social connection and connectedness, and the revised search found all nine. That one study revealed a whole family of studies the search would have missed.

**Sarah:** Did they test the third concept too?

**Kiffer:** They did. They ran the loneliness and older adults concepts on their own. That also found all nine known studies, along with many more records that measured loneliness without testing any intervention. Since the third concept cost nothing in the test set and removed a lot of noise, they kept it.

**Sarah:** Is there a catch with test sets?

**Kiffer:** There is. A test set tells you whether the search finds studies like the ones you already know. It can't tell you about studies described in ways nobody anticipated. And if you used those studies to choose your terms, the search will find them almost by construction. So it's one check among several, and peer review is another.

**Sarah:** Let's move to section two, where the table becomes an actual search.

**Kiffer:** The Cedar Valley librarian uses the Ovid platform for MEDLINE. Ovid keeps a numbered search history, so each line is either a set of terms or a combination of earlier lines. The section walks through the operators and symbols one at a time and then reads the whole strategy.

**Sarah:** Start with Boolean operators.

**Kiffer:** OR retrieves records with any of its terms, so it widens a search and belongs inside a concept. AND retrieves records with a term from every set, so it narrows a search and joins concepts. NOT removes records, and it needs care, because it removes a record if the excluded word appears anywhere in the searched fields. Imagine the intern wanted to drop studies of children and added a line removing every record that mentions children. That line would also remove intergenerational programs, where older adults spend time with children, and those are exactly the kind of intervention the review is looking for.

**Sarah:** But the final strategy does use NOT somewhere.

**Kiffer:** It uses it once, in a tested form. The line removes records indexed with animal headings that are not also indexed with the heading Humans, so a study that involves both people and animals stays in. That line is widely used and appears in Cochrane guidance on searching.

**Sarah:** You also talk about parentheses.

**Kiffer:** When one line mixes operators, the database decides the order, and platforms decide differently. If you type lonely or isolated and seniors, the database might read it as lonely, or isolated seniors, which brings back loneliness at every age. Parentheses settle it. Better still, build the search in numbered lines, one line per heading or group of terms, then a line that combines each concept with OR, and a final line that joins the concepts with AND. People call this the building-block approach. Every combination can be checked, and you can see how many records each line adds.

**Sarah:** Let's take truncation next.

**Kiffer:** Truncation finds every word that begins with a stem. In Ovid, the dollar sign does that, and the asterisk works too. Volunteer with a dollar sign finds volunteer, volunteers, volunteered and volunteering. You can also add a number to cap the extra characters, so club with a dollar sign and the number one finds club and clubs and stops.

**Sarah:** I assume the trap is a stem that's too short.

**Kiffer:** It is, and the team's favourite example is art. The group-activity line wanted art programs, but art with a dollar sign also finds artery, arterial, arthritis and article. So the strategy searches art or arts. Class has the same problem, since it brings in classification, so the line searches class or classes.

**Sarah:** How do wildcards work?

**Kiffer:** A wildcard stands for a character inside a word. In Ovid, the number sign stands for exactly one character, so woman and women can be found together. The question mark stands for one character or none, which handles spelling differences. Neighbourhood with a question mark after the o finds both the Canadian and the American spelling, and the strategy uses exactly that.

**Sarah:** How does Ovid handle phrases?

**Kiffer:** Ovid treats words typed side by side as a phrase, so later life is searched as a phrase. Phrases are precise, but rigid. Social isolation as a phrase misses isolated socially, and it misses social and emotional isolation. That is where proximity comes in. In Ovid, the adjacency operator followed by a number finds two terms within that many words of each other, in either order. The number counts the search words themselves, so adjacency three allows up to two words in between.

**Sarah:** Give me the Cedar Valley example.

**Kiffer:** Social within three words of isolation finds social isolation, isolated socially, and social and emotional isolation. It does not find social activities reduced their isolation, because three words sit between the terms.

**Sarah:** How do you pick the number?

**Kiffer:** It's a judgement about recall and precision. A small number behaves almost like a phrase. A big number starts pairing words from different clauses. Two to five is common in reviews, and you check by looking at what each setting adds.

**Sarah:** Field tags are the last piece.

**Kiffer:** A field tag tells Ovid where to look. A slash after a term searches it as a subject heading, and the word exp in front explodes it. For free text, the team searches the title, abstract and author keyword fields together. Ovid also has a multipurpose field, which is what you get if you type a word with no tag, and a text word field. Both cover different sets of fields in different databases, so the librarian avoids them in a review strategy. Listing the fields explicitly means a reader, a peer reviewer or a translator can see exactly what each line searched.

**Sarah:** Walk us through the final MEDLINE strategy.

**Kiffer:** It's thirty numbered lines in four blocks. Lines one to seven are loneliness and isolation. There are two headings, the truncated stem for lonely, an adjacency line for social or emotional isolation and related ideas, a line that only accepts isolated or isolation near words for people, and the line for social connectedness that the test set prompted. Line seven combines them.

**Sarah:** Then come the older adults.

**Kiffer:** Lines eight to twelve cover them. There is the exploded heading Aged, an adjacency line that catches older adults, older Canadian adults and adults who are older, a line of single words like seniors and gerontological, and a line of phrases such as later life and old age.

**Sarah:** And the interventions block is the big one.

**Kiffer:** It runs from line thirteen to line twenty-six. It has six headings, such as Social Support, Volunteers, Peer Group and Intergenerational Relations, and then seven free-text lines for social prescribing, connectors and link workers, befriending and friendly visiting, peer and volunteer support, intergenerational programs, group activities, and community programs. Line twenty-seven joins the three concepts with AND, lines twenty-eight and twenty-nine remove animal-only records, and line thirty limits to publications from two thousand ten onward, as the protocol requires. In the fictional search, it returned seven hundred and forty-two records.

**Sarah:** Is there a language limit?

**Kiffer:** There is none. The protocol makes English and French reports eligible and asks the team to list relevant reports in other languages, and the team can only list them if the search finds them.

**Sarah:** In section three the team moves beyond MEDLINE. Why not stop there?

**Kiffer:** No single database covers all the relevant literature. Loneliness among older adults is studied in nursing, psychology, gerontology, social work and public health. So Cedar Valley searches MEDLINE, Embase, the Cumulative Index to Nursing and Allied Health Literature, which everyone calls CINAHL, the American Psychological Association's PsycInfo, and the Web of Science Core Collection. For Cochrane intervention reviews, the handbook treats the Cochrane trials register, MEDLINE and Embase as the minimum.

**Sarah:** And rewriting the MEDLINE search for each of those is translation.

**Kiffer:** That is search translation. The key idea is to separate the database from the platform. The database supplies the records and the thesaurus, and the platform supplies the search software and its syntax. MEDLINE can be searched on Ovid or PubMed, and Embase on Ovid or on Elsevier's own platform. Moving from Ovid MEDLINE to Ovid Embase changes the headings and keeps most of the syntax. Moving to CINAHL on EBSCOhost changes both.

**Sarah:** Where do people go wrong?

**Kiffer:** Most errors happen in three places. The first is wildcards. Ovid and EBSCOhost give opposite meanings to the question mark and the number sign. In Ovid the question mark is optional, and in EBSCOhost the number sign is optional, so a wildcard copied across without thinking searches something else. The second is that Web of Science treats the dollar sign as one character or none. If you paste the Ovid stem for lonely with its dollar sign into Web of Science, you miss the word loneliness entirely, so you have to change it to an asterisk.

**Sarah:** That's an easy one to miss. What is the third?

**Kiffer:** The third is proximity counting. Ovid counts the search words themselves, while EBSCOhost and Web of Science count only the words in between. So adjacency three in Ovid matches N two in EBSCOhost and near two in Web of Science. If you copy the number straight across, every proximity line becomes slightly broader than you intended.

**Sarah:** What did the CINAHL version look like?

**Kiffer:** It is longer, because CINAHL wants the title and abstract fields searched separately, so each free-text line is written twice inside one search line. Headings are written in a different format, with a plus sign to explode, so Aged plus is the exploded heading.

**Sarah:** What happened to the headings themselves?

**Kiffer:** That's where translation takes the most care, because each thesaurus is separate. Loneliness and Social Isolation keep their names in all three. The MeSH heading Volunteers becomes Volunteer Workers in CINAHL and the singular volunteer in Emtree, which is Embase's thesaurus. Social Support becomes Support, Psychosocial in CINAHL. For some MeSH headings the team didn't find a close match, so it relied on the free-text lines and wrote the decision down.

**Sarah:** Are there tools that do this for you?

**Kiffer:** There are tools for syntax. The Polyglot Search Translator converts a strategy between common platforms, and a randomized trial found it cut the time searchers spent on translation. It leaves the headings to you, though. After translating, the librarian runs each line, compares the counts with the MEDLINE lines, looks into anything that returns nothing or far too much, and repeats the test-set check in every database.

**Sarah:** Embase has a lot of conference abstracts. Did the team remove them?

**Kiffer:** Some teams do, with a publication-type line. Cedar Valley kept them. The protocol excludes conference abstracts, but it asks screeners to look for a full report of any relevant abstract first, and they can't do that if the search throws the abstracts away.

**Sarah:** Let's turn to search filters.

**Kiffer:** A search filter is a ready-made block of lines that retrieves a type of record, most often a study design. A validated filter has been tested against a set of records whose status is known, so its sensitivity and precision are reported. The Cochrane filter for randomized trials in MEDLINE is the best-known example.

**Sarah:** Did Cedar Valley use one?

**Kiffer:** It did not. A design filter suits a review limited to one design, and a scoping review that charts every kind of evaluation would lose most of its studies to a trial filter.

**Sarah:** What about limits more generally?

**Kiffer:** Every limit removes records, so every limit needs a reason that comes from the protocol, and the reason gets reported. The date limit follows the protocol. Language limits can bias a review toward English-language studies. And limits that depend on indexing, like the age-group limits in Ovid, miss every unindexed record. That's why the population is searched with headings and free text instead.

**Sarah:** In section four, someone else looks at the search before it runs for real.

**Kiffer:** Right. Studies of the search strategies published with systematic reviews have found errors of many kinds, including missing synonyms, unexploded headings, misspellings and wrong line numbers. Margaret Sampson and Jessie McGowan described these in two thousand six. The frustrating thing is that none of these errors shows up in the results.

**Sarah:** So peer review is the safety check.

**Kiffer:** It's the same logic as peer reviewing a protocol, applied to the step everything else depends on. The guideline is called PRESS, which stands for Peer Review of Electronic Search Strategies. The current version is the twenty fifteen guideline statement, and it organizes the review into six elements.

**Sarah:** What are they?

**Kiffer:** First, does the search translate the research question properly, with the right number of concepts? Second, are the Boolean and proximity operators used correctly? Third, are the subject headings relevant, complete and exploded where they should be? Fourth, do the free-text terms include the synonyms and variants, with sensible truncation and field tags? Fifth, are there errors in spelling, syntax or line numbers? And sixth, are the limits and filters justified? For each one, the reviewer records no revisions, suggested revisions or required revisions.

**Sarah:** The lesson shows the intern's first draft. How bad was it?

**Kiffer:** It was a normal first draft, which means it had several problems. It added a fourth concept of evaluation words. It used NOT to remove records mentioning children. It searched Aged without exploding it. It truncated art too short. It misspelled loneliness. One combination line joined lines eight and nine and skipped line ten. And it applied an English-language limit and an age limit that the protocol didn't support.

**Sarah:** And the reviewer caught all of that?

**Kiffer:** The second librarian did, and rated the draft as needing required revisions overall. The intern revised the strategy, wrote a short response to every comment, and sent both back. The reviewer confirmed the changes, and the team reran the test-set check. The result is the strategy we walked through in section two.

**Sarah:** What makes a peer review comment useful?

**Kiffer:** It names the line number, it says what the problem would do to the results, and it proposes a fix. It separates what must change from what is a matter of judgement. And it asks before overruling a decision the searcher made on purpose, such as leaving the setting out of the search. That's the structure to use whenever you review someone else's search.

**Sarah:** Let's turn to reporting. What is PRISMA-S?

**Kiffer:** It's a sixteen-item extension of the Preferred Reporting Items for Systematic reviews and Meta-Analyses statement, written specifically for literature searches. Melissa Rethlefsen and colleagues published it in twenty twenty-one. The items fall into four groups: information sources and methods, search strategies, peer review, and managing records.

**Sarah:** Which ones matter most for the database searches?

**Kiffer:** You name each database with its platform, and you give the full strategy for every database exactly as run, usually in an appendix. You describe and justify every limit, cite any filter, say whether you reused an earlier strategy, give the date of each search, describe the peer review, and report the number of records from each source. Lesson five handles the items about websites, registries and citation chasing, and lesson seven handles de-duplication.

**Sarah:** How do you remember all of that after the fact?

**Kiffer:** You don't, which is why the team keeps a search log while searching. For every search it records the database, the platform, the coverage label the platform shows, the date, who ran it, the full strategy with the count for each line, the limits, and the name of the exported file. Ovid and EBSCOhost will export the search history with the counts, and saving that on the day is the simplest way to keep the strategy exactly as run.

**Sarah:** What did Cedar Valley end up reporting?

**Kiffer:** All five databases were searched on the tenth of February, twenty twenty-six. MEDLINE returned seven hundred and forty-two records, Embase eight hundred and sixteen, CINAHL three hundred and eighty-eight, PsycInfo two hundred and ninety-six, and Web of Science two hundred and thirty-eight. That's two thousand four hundred and eighty in total, and in lesson seven the team removes six hundred and ten duplicates, which leaves one thousand eight hundred and seventy records to screen.

**Sarah:** What about keeping a search current?

**Kiffer:** Cochrane's standards expect searches to be rerun within twelve months before a review is published. Cedar Valley's brief is due in twelve weeks, so no update is planned in that time, but every strategy is saved on its platform so the searches can be rerun quickly if the review is published later.

**Sarah:** Let's finish by putting it together. What are the steps, from start to finish?

**Kiffer:** Start from the review question and build a concept table with two to four concepts, with headings and free-text terms for each, and a note of anything you'll handle at screening. Write the MEDLINE strategy in Ovid syntax as numbered lines, and check it against three to five studies you already know are relevant. Then translate it into one more database, such as CINAHL or PsycInfo, and write down your heading decisions.

**Sarah:** And then the peer review?

**Kiffer:** Right. A second searcher reviews the MEDLINE strategy using the six PRESS elements, and you revise the search and respond to each comment. And start a search log that records the database, platform, coverage, date, limits and number of records for every search you run.

**Sarah:** That's a lot of work.

**Kiffer:** It is, and it pays off at every later stage, because screening, charting and synthesis all depend on what this search finds.

**Sarah:** Thanks, Kiffer. Next time we turn to grey literature and citation chasing.

**Kiffer:** Thanks, Sarah. See you then.
