# Lesson 5: Grey Literature, Web Searching and Citation Chasing

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are on Lesson five, which is about grey literature, web searching and citation chasing.

**Sarah:** Last week was all about database searching: building a search string, translating it across databases and having it peer reviewed. Why isn't that enough?

**Kiffer:** Because bibliographic databases index journal articles, and a lot of the evidence on health programs never appears in a journal. It sits in reports on government websites, in theses held by university libraries, in conference abstracts and in records of trials that were registered and then never published. A review that searches only databases can miss that material, and the missing material is not a random sample of the evidence.

**Sarah:** We'll come back to that last point. First, remind us about the running case.

**Kiffer:** The running case is the Cedar Valley evidence review, which is fictional. The Cedar Valley Health Authority in British Columbia serves about two hundred and ten thousand people, and about forty-six thousand of them are sixty-five or older. Before it launches a community connector program, which is a form of social prescribing, its planning team has asked a small evidence team for a rapid scoping review and an environmental scan within twelve weeks. There is an evidence officer from the health authority, a university librarian who works with the team, and a student intern. The review question is which community-based interventions have been evaluated for reducing loneliness or social isolation among adults aged sixty-five and older, and with what outcomes. By this point their database searches have been run, and they retrieved two thousand four hundred and eighty records.

**Sarah:** Let's start with section one. What exactly is grey literature?

**Kiffer:** The most widely cited definition comes from an international conference on grey literature held in Luxembourg in nineteen ninety-seven and extended in New York in two thousand four. In plain terms, grey literature is material produced by government, academia, business and industry whose publication is not controlled by commercial publishers, because publishing is not the producer's main activity.

**Sarah:** So it's defined by who publishes it.

**Kiffer:** Exactly. The definition says nothing about quality. A grey document might be a careful evaluation with a good sample and clear methods. It might also be a two-page promotional summary with no methods at all. Students sometimes assume grey means weak, and that assumption misleads them.

**Sarah:** Why does a review need to go looking for it? Couldn't you argue that if something wasn't good enough for a journal, it's not good enough for a review?

**Kiffer:** That argument is common, and the evidence runs against it. The first reason to search grey literature is publication bias. Studies with favourable or statistically significant results are more likely to be submitted and accepted. Hopewell and colleagues reviewed methodological studies for the Cochrane Collaboration in two thousand seven and found that trials published in journals tended to show larger intervention effects than trials found only in the grey literature.

**Sarah:** So if you leave out the grey trials, the intervention looks better than it is.

**Kiffer:** That's the risk. A journal-only review can overestimate effects. Searching grey literature is one of the main ways to reduce that bias at the search stage. Statistical checks for publication bias, like funnel plots, belong to the epidemiology course's meta-analysis lesson.

**Sarah:** What are the other reasons?

**Kiffer:** The second is time-lag bias. Favourable results tend to be published sooner, and null results can take years or never appear. Abstracts, preprints and registry records show work that is recent or still under way. The third reason is practice-based evidence, and for Cedar Valley it may be the most important. Many community programs are evaluated by the organizations that run them or by their funders, and those evaluations end up as reports on websites.

**Sarah:** Like a seniors' centre evaluating its own telephone check-in service.

**Kiffer:** Exactly that. Those reports describe precisely the kinds of programs the planning team is considering, and almost none of them are in MEDLINE. The fourth reason is context. Reports tell you who runs programs and how they're funded, which decision makers care about, and which also feeds the environmental scan in Lesson eleven.

**Sarah:** Is there a downside?

**Kiffer:** Cost. Mahood, Van Eerd and Irvin described grey literature searching for a systematic review as time-consuming and hard to document, and the documents vary enormously in format and quality. A team with twelve weeks cannot search everywhere. So the answer is to plan, focus on the sources most likely to matter, limit how much of each one you examine, and write everything down.

**Sarah:** Let's go through the kinds of grey literature. Which ones matter most for health reviews?

**Kiffer:** Government and agency reports come first. For Cedar Valley, a good example is the report on the social isolation of seniors that Canada's National Seniors Council published in twenty fourteen. Then program evaluation reports, written by the delivering organization, an external evaluator or a funder. Then theses and dissertations, which often contain full methods and null results that later journal articles shorten or drop. Conference abstracts, which are short and preliminary but show recent work. Preprints, which are full manuscripts posted before peer review, on servers like medRxiv for health and PsyArXiv for psychology. Trial registry records. Guidelines and policy briefs, such as the World Health Organization's advocacy brief on social isolation and loneliness among older people. And statistical reports from bodies like Statistics Canada, which mostly give you context.

**Sarah:** You spend extra time on trial registries in the lesson. Why?

**Kiffer:** Because they can reveal studies that would otherwise be invisible. Since two thousand five, the International Committee of Medical Journal Editors has required trials to be registered at or before the enrolment of the first participant as a condition of consideration for publication in the journals that follow its recommendations. So a registry record exists for many trials whose results were never published.

**Sarah:** Give me a Cedar Valley example.

**Kiffer:** Suppose the intern finds a record on ClinicalTrials dot gov for a completed trial of telephone befriending for older adults, with no linked publication. That's a signal. The team should check whether results were posted to the registry, search again for a publication, and contact the investigators. Excluding the trial because there's no article would make publication bias worse.

**Sarah:** Are social programs registered as reliably as drug trials?

**Kiffer:** Registration may be less consistent for social and community interventions, so a registry search for Cedar Valley will be incomplete. It's still worth doing, because the cost is low and an unpublished trial is a valuable find.

**Sarah:** Where does a Canadian team actually look?

**Kiffer:** The search is organized around organizations. At the federal level there's the Public Health Agency of Canada, whose publications are on Canada dot c a, the National Seniors Council, Employment and Social Development Canada, which runs the New Horizons for Seniors Program, Statistics Canada and the Canadian Institute for Health Information. Provincially, there's the British Columbia Ministry of Health, the Office of the Seniors Advocate, the British Columbia Centre for Disease Control and the regional health authorities. United Way British Columbia, which manages the provincially funded Better at Home program, and local non-profits and seniors' centres. Nationally, the Canadian Institute for Social Prescribing and the National Institute on Ageing. For theses, Theses Canada, which is a program of Library and Archives Canada with Canadian universities, along with university repositories like Summit at Simon Fraser University and cIRcle at the University of British Columbia.

**Sarah:** Why both Theses Canada and the repositories?

**Kiffer:** Because deposit practices vary by university, so Theses Canada won't hold every recent thesis. Checking the repositories directly closes that gap. I'd also mention the Grey Matters checklist, which came from the Canadian Agency for Drugs and Technologies in Health. That agency became Canada's Drug Agency in twenty twenty-four, and the checklist is now published there.

**Sarah:** Once the team has a pile of grey documents, how does it decide which ones to trust?

**Kiffer:** The lesson uses a checklist developed by Jess Tyndall at Flinders University, known by its letters, A A C O D S. They stand for authority, accuracy, coverage, objectivity, date and significance. Authority asks who produced the document and with what expertise. Accuracy asks whether methods and data are stated. Coverage asks whether the scope and limits are clear. Objectivity asks whether the producer's interests could have shaped the findings. Date asks whether the document is dated and current, and significance asks whether it's relevant to the question and adds something. So a report with no named authors, no date and funding from the program it evaluates raises concerns about authority, date and objectivity.

**Sarah:** Does passing that checklist mean the study is good?

**Kiffer:** No, and that's an important distinction. The checklist is a credibility screen. A credible report can still have a weak design, such as a pre-post survey with no comparison group. The risk-of-bias tools in Lesson eight judge design.

**Sarah:** Let's move to section two. Once you know what you're looking for, how do you actually search the web for it?

**Kiffer:** I start with a case study by Katelyn Godin and colleagues, published in two thousand fifteen. They searched for Canadian guidelines on school breakfast programs using four strategies: grey literature databases, customized Google search engines, the websites of targeted organizations, and emails to content experts. Fifteen publications were included. Targeted website searching found fourteen of them. The content experts identified nine. The grey literature databases found only one. And four publications were found by only one strategy, so dropping any strategy would have lost something.

**Sarah:** That's striking for the databases.

**Kiffer:** It is, although the yields will differ by topic. For a question about community programs, like ours, the case study suggests that organizational websites and contacts with the people who run programs are likely to be productive.

**Sarah:** What does targeted website searching involve in practice?

**Kiffer:** First you build a reasoned list of organizations, the ones that fund, deliver, evaluate or set policy for the programs in question. For Cedar Valley, the librarian proposes candidates, the intern and the evidence officer add others, and each site goes into the plan with a short reason. Then each site is searched in up to three ways. You browse the pages where documents are listed, like Publications or Reports, and scan every relevant item. You use the site's own search box with simple terms, knowing that most site search boxes handle Boolean logic and phrases poorly. And you run a Google search restricted to that site, which often turns up portable document format reports that the site's own search misses.

**Sarah:** That brings us to the operators. Walk me through them.

**Kiffer:** Quotation marks match an exact phrase, like community connector. The word OR, in capital letters, offers alternatives. A minus sign placed directly in front of a word excludes it, which is handy for removing job postings. Then there are three operators the lesson concentrates on. The site operator limits results to a domain, such as the British Columbia government's domain, or to a whole top-level domain like dot c a. The file type operator limits results to one format, and for reports that is usually portable document format. The in title operator requires a word to appear in the page title. Each of these is typed as the operator name, a colon and the term, with no space after the colon.

**Sarah:** What happens if you do put a space there?

**Kiffer:** Google reads the operator name as an ordinary word, and your restriction quietly disappears. That's the most common error I see. The second is writing or in lower case, which Google treats as an ordinary word, so your alternatives aren't combined.

**Sarah:** Students coming from Lesson four will want to use truncation.

**Kiffer:** Right, and Google has no truncation. The asterisk in Google stands for a whole word inside a quoted phrase. So typing the stem isolat with an asterisk doesn't give you isolated and isolation. You either write the forms you need or rely on Google's automatic matching of some variants.

**Sarah:** Can't you just paste the database string from last week into Google?

**Kiffer:** No. A long Boolean string with field tags and nested parentheses exceeds Google's limit of thirty-two words and uses syntax Google ignores. The string has to be rewritten as several short queries, each combining one or two concepts with operators.

**Sarah:** Give me one of the Cedar Valley queries.

**Kiffer:** One query looks for the word loneliness or the phrase social isolation, plus seniors and evaluation, restricted to the British Columbia government domain and to portable document format files. Another covers the regional health authorities by joining their domains with OR. And there's a federal query with a wrinkle, because Government of Canada content is split between the Canada dot c a domain and older g c dot c a domains, such as the one Statistics Canada uses.

**Sarah:** So a federal query that only names one domain will miss part of the government.

**Kiffer:** Exactly, so the team includes both. There's also a query for the World Health Organization's site, which uses the phrase older people, because that's the organization's own vocabulary. Matching each organization's language matters on the web the way controlled vocabulary matters in a database.

**Sarah:** Where does Google Scholar fit?

**Kiffer:** Google Scholar indexes journal articles along with theses, preprints and reports from academic and organizational sites. It supports phrases, OR, the minus sign, a title operator and an author operator. But it ranks results by an undisclosed method, displays at most one thousand results for any query, and can't export a full result set in one step.

**Sarah:** Is there evidence on how useful it is for grey literature?

**Kiffer:** Neal Haddaway and colleagues tested it in two thousand fifteen using case studies of environmental science reviews. They found a moderate amount of grey literature in the results, much of it well beyond the first pages, and Google Scholar missed important studies in five of six case studies. They recommended that title searches for grey literature focus on the first two hundred to three hundred results, and that Google Scholar never be used alone. A later comparison by Gusenbauer and Haddaway concluded it was inadequate as the principal search system for a systematic review.

**Sarah:** So it's a supplement.

**Kiffer:** Yes. The Cedar Valley team uses one Google Scholar query and screens the first two hundred results. Which brings us to stopping rules. A stopping rule is a decision, made before searching, about how many ranked results you'll examine. For Cedar Valley, it's the first one hundred Google results for each query, or all the results if fewer are shown.

**Sarah:** Why does it have to be decided in advance?

**Kiffer:** Two reasons. It keeps the effort consistent from query to query, and it lets the team report exactly how much of each ranking was examined. Deciding while you look lets what you see, or fatigue, set the stopping point.

**Sarah:** Let's turn to section three, citation chasing. What is it?

**Kiffer:** Every article cites earlier work and is later cited by newer work. Citation chasing uses those links to find studies that keyword searches miss. If an included study cites an evaluation of a befriending program, that evaluation is probably relevant, whatever words happen to be in its title.

**Sarah:** I've heard it called snowballing.

**Kiffer:** It has had a lot of names: snowballing, pearl growing, reference list checking, citation tracking. In twenty twenty-four, Julian Hirt and colleagues published the TARCiS statement, which stands for Terminology, Application and Reporting of Citation Searching. It came out of a Delphi consensus study with international methods experts and has ten recommendations.

**Sarah:** What are the key terms?

**Kiffer:** Backward citation searching retrieves and screens the references that your seed references cite. Forward citation searching retrieves and screens the references that cite your seeds. Seed references are the known relevant documents you start from, and the statement recommends using all the records included after full-text screening. Repeating the process with newly found studies as seeds is called iterative citation searching.

**Sarah:** When is it worth the effort?

**Kiffer:** When a topic is hard to search with words, and loneliness interventions are a very good example. The same kind of program might be called social prescribing, community referral, a link worker service, befriending, friendly visiting or a telephone check-in service. Studies on the same topic tend to cite each other even when they describe themselves differently.

**Sarah:** Is there evidence that it finds much?

**Kiffer:** The best-known study is by Trisha Greenhalgh and Richard Peacock in two thousand five. They audited the sources for a review of complex evidence on the diffusion of innovations in health service organizations. Fewer than a third of the sources came from the planned database and hand searches, and about half came from snowballing.

**Sarah:** That sounds like databases barely matter.

**Kiffer:** I'd push back on that reading. That review covered a scattered literature from many disciplines, so its proportions are unusually high. A Cochrane methodology review by Horsley and colleagues found limited but supportive evidence for checking reference lists. And the statement is clear that citation searching shouldn't replace extensive database searching in a review that aims to find all relevant studies, because it can only find documents linked to the seeds.

**Sarah:** Are there biases built into citation chasing?

**Kiffer:** Two. It reinforces whatever bias the seed set has, so if the included studies come mostly from one country, their citations tend to lead back to the same networks. And forward chasing favours older seeds, because recent studies haven't had time to be cited.

**Sarah:** Walk me through backward searching.

**Kiffer:** You assemble the seed set and record it. Then you retrieve the cited references, ideally from a citation index such as Web of Science or Lens dot org, because that gives you titles and abstracts to screen. A reference list read by hand usually gives only titles. Then you remove duplicates within the set and remove anything already screened from the main searches, and screen what's left with the same process as everything else.

**Sarah:** You also say backward searching is good for finding grey literature.

**Kiffer:** It is. Authors of included studies often cite program reports, government documents and theses, and those citations point to documents no index covers. So the Cedar Valley team also reads the reference lists of the grey documents it includes, by hand.

**Sarah:** And forward searching?

**Kiffer:** Forward searching retrieves the documents that cite each seed, using Web of Science, Scopus, Lens dot org or Google Scholar. They cover different material, so the statement suggests considering two indexes. Because citation counts grow over time, you record the date you searched.

**Sarah:** The lesson also covers citation network tools. What are they doing under the hood?

**Kiffer:** Two old ideas from information science. Co-citation, described by Henry Small in nineteen seventy-three, links two documents that a later document cites together. Bibliographic coupling, described by Kessler in nineteen sixty-three, links two documents that cite the same earlier work. Both rest on the citation index that Eugene Garfield proposed in nineteen fifty-five, which became the Science Citation Index. They fall into two groups. A tool like citationchaser, which Haddaway and colleagues released in twenty twenty-two, takes a list of seeds, retrieves their complete cited and citing lists from Lens dot org, and exports them for screening. That fits the recommendations well. Tools like Connected Papers, Litmaps or ResearchRabbit draw maps of similar papers chosen by an algorithm.

**Sarah:** Are the map tools useful?

**Kiffer:** For exploration, yes. They help you see clusters of related work and spot terms or authors to add to a search. But they show a selection, not a complete list, and their data and ranking methods can change, so they can't replace a complete backward and forward search.

**Sarah:** How does Cedar Valley use all this?

**Kiffer:** Twice. Before screening, the librarian does a quick sensitivity check. The team found three earlier reviews of interventions for loneliness and social isolation in older people, by Dickens and colleagues, by Gardiner and colleagues, and by Fakoya and colleagues. The intern checks whether the eligible studies those reviews cite are among the database records. If, say, two were missing and both described friendly visiting, the librarian would add that term and rerun the search. After screening, the team uses its thirty-eight included studies as seeds. Backward and forward searching together retrieve two thousand four hundred and forty-eight records. After removing duplicates and anything already screened, one thousand one hundred and sixty new records are screened. Twenty-three full texts are assessed, and four studies are included, which brings the review to forty-two.

**Sarah:** Do they run a second round?

**Kiffer:** They decide not to, because of the twelve-week deadline and because the first round added few studies relative to the number screened. The key is that they report that decision and the reason. A team with more time might reasonably have continued.

**Sarah:** Which brings us to section four, documentation. The lesson calls it the weakest part of most grey literature searches. Is that fair?

**Kiffer:** The evidence suggests it is. Simon Briscoe looked at systematic reviews from the United Kingdom's Health Technology Assessment programme published between two thousand four and two thousand thirteen. Of three hundred reviews, one hundred and eight reported searching the web. Most of those gave only the names of the websites or search engines, and only six met the highest standard of reporting.

**Sarah:** So "we searched Google" is the norm.

**Kiffer:** More or less. And from the searcher's side, Mahood and colleagues and Godin and colleagues describe sources that can't export results and web pages that move or vanish. Documentation is hard, which is why it needs a plan and a log.

**Sarah:** The lesson separates transparent from reproducible. What's the difference?

**Kiffer:** A search is transparent when a reader can see exactly what was done: the sources, queries, settings, dates, how far down each list the searcher went, and what was kept. A search is reproducible when someone repeating those steps gets the same results. Lesson three explained why web searches rarely achieve that, because rankings are personalized and pages change daily.

**Sarah:** So what's the realistic goal?

**Kiffer:** Full transparency, and as much reproducibility as the sources allow. You record enough for someone to rerun each search, and you keep copies of everything you saved, so a reader can check the evidence you actually used even after the websites change.

**Sarah:** What goes in the search plan?

**Kiffer:** It's written before searching and belongs in the protocol. It lists source groups and specific sources with a reason for each, states the method for each source, gives the queries and any limits, sets a stopping rule for every ranked source, and names who will search and check each one and when. Writing it first fixes the stopping rules before anyone sees the results.

**Sarah:** And the web search log?

**Kiffer:** One row for every search, filled in at the time. It records an identifier, the date and searcher, the source and its web address, the settings, the exact query or the pages browsed, the number of results screened and the items saved. Settings matter, so for Google the team notes the Google domain, the private browsing window, that it was signed out, the apparent location and any filters.

**Sarah:** Give me one row from Cedar Valley.

**Kiffer:** Search W four, run on the eighteenth of February, twenty twenty-six, on the Canadian Google domain, in a private window while signed out. The query required loneliness in the title, added seniors, program and evaluation, restricted results to dot c a sites and portable document format files, and excluded jobs. The intern screened the first one hundred results and saved eleven items, numbered W four dash one to W four dash eleven.

**Sarah:** And do you log searches that find nothing?

**Kiffer:** Always. The Cedar Valley log includes a Statistics Canada site search that saved nothing. A reader needs to know the source was covered.

**Sarah:** What about preserving documents?

**Kiffer:** Web addresses stop working, which people call link rot. So every saved item gets an identifier tied to its search, it goes into the reference manager with its address and access date, web pages get an archived copy, for example through the Wayback Machine's Save Page Now function, and the first page of results for each query is saved as a screenshot.

**Sarah:** How do the Cedar Valley numbers come together?

**Kiffer:** The twelve web searches saved sixty-three items. Adding theses, trial registry records, preprints and conference abstracts gives seven hundred and seventy-one records and items. Removing two hundred and fourteen duplicates leaves five hundred and fifty-seven for screening. In the completed review, grey literature and website searching contributed twenty-six documents, and citation searching added the four studies we discussed.

**Sarah:** And reporting?

**Kiffer:** Reporting follows PRISMA-S, the literature search extension of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses. It has items for study registries, online resources and browsing, citation searching, contacts, full search strategies and dates. In the flow diagram from the twenty twenty version of the statement, registries are counted with databases, and websites, organizations and citation searching are counted in a separate column for other methods. Lesson seven builds that diagram.

**Sarah:** Let's finish with the practical side. What does a team come away with from all this?

**Kiffer:** Three records. First, a grey literature search plan for a British Columbia question that covers at least four source groups, with at least five organizational websites, and gives each source a reason, a method, a stopping rule, a person and a time. Second, at least three Google queries that use the site, file type and in title operators correctly, each run and recorded in a web search log with every field, with each item saved under its log identifier. Third, a small backward and forward citation search from two relevant articles found in the pilot search, run with one citation index, recording the index, the date and the counts.

**Sarah:** And Lesson six adds to that.

**Kiffer:** Yes. Lesson six adds an artificial intelligence assisted search and a verification log, and it explains why those tools sometimes produce citations that don't exist. The habits from this week, exact queries, dates, stopping rules and saved copies, are exactly what you'll need there.

**Sarah:** Thanks, Kiffer. That's it for this week's Office Hours.

**Kiffer:** Thanks, Sarah, and thanks to everyone listening. See you next week.
