# Lesson 5: Screening and Diagnostic Tests

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson five, Screening and Diagnostic Tests.

**Sarah:** It's meant for after you've worked through the lesson. We'll bring in some perspective and some critique, spend time on the questions students usually find thorny in this material, and work through a few extra examples, including some that are harder than the ones on the lesson page.

**Kiffer:** In a few places we'll ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. What does a positive result from a good test actually tell you? Are sensitivity and specificity really fixed properties of a test? What can go wrong when you combine a test result with what you already know, or with a second test? And how can screening make survival look better without anyone living longer? We'll finish with screening for depression, where Kiffer and I land in different places.

**Kiffer:** Let's start with the first one.

**Sarah:** I'd like to use a real questionnaire for this, because it connects back to Lesson three. The PHQ-9 is a nine-item questionnaire used to screen for depression. In 2019, a team led by researchers at McGill University and the Jewish General Hospital in Montreal published a large pooled analysis of it in the BMJ. In the twenty-nine studies that compared it with a semi-structured diagnostic interview, a score of ten or more had a sensitivity of about eighty-eight percent and a specificity of about eighty-five percent.

**Kiffer:** And a set of questionnaire items is a test in exactly the sense this lesson uses, so the same arithmetic applies. Let's put it in a clinic. Suppose, as a made-up figure, that one in ten adults attending a clinic has major depression, and everyone completes the PHQ-9.

**Sarah:** So here's a question for everyone listening. Of all the people who score ten or more, roughly what share actually has major depression? Take a few seconds.

*(Pause)*

**Kiffer:** Let's count it in a group of one thousand. One hundred have major depression, and eighty-eight percent of them screen positive, which is eighty-eight people. Nine hundred don't, and fifteen percent of them screen positive anyway, which is one hundred and thirty-five people. So two hundred and twenty-three people screen positive, and eighty-eight of them have depression. Eighty-eight divided by two hundred and twenty-three is about 0.39.

**Sarah:** So it's just under forty percent. Most people who screen positive don't have major depression, even though the questionnaire picks up almost nine in ten people who do. I think that surprises people because eighty-eight percent sounds like the answer to the question they're asking.

**Kiffer:** It does, and it's one of the most common errors with this material. Eighty-eight percent is the probability of a positive screen given depression. The person holding the result wants the probability of depression given a positive screen. Those are different conditional probabilities, and the second one depends on how common depression is among the people tested.

**Sarah:** Can I ask where the cutoff of ten comes from?

**Kiffer:** In that analysis, ten was the cutoff that gave the largest value of sensitivity plus specificity. That's the Youden index from the lesson, sensitivity plus specificity minus one. On an ROC curve, it picks the point that sits farthest above the diagonal chance line. You'll sometimes hear it described as the point closest to the top-left corner. That's a separate rule. The two pick the same cutpoint when the curve is symmetric, and they can pick different cutpoints when it's lopsided.

**Sarah:** And maximizing sensitivity plus specificity gives the two equal weight. That's a choice somebody made.

**Kiffer:** It is. And equal weight on sensitivity and specificity means unequal weight on the two kinds of error. In a clinic where one in ten people has depression, there are nine people without it for each person with it, so the rule treats one missed case as costing as much as nine false alarms. If missing a case were worse than that, you'd lower the cutoff and accept a lower specificity. Either way it's a value judgement, and it's worth saying so when you report a cutpoint.

**Sarah:** Question two. Are sensitivity and specificity really fixed properties of a test? We usually describe them as the portable numbers, the ones that travel with the test from one setting to another.

**Kiffer:** They travel much better than predictive values, but they can still change. Here's a clear example. In 2022, a Cochrane review of rapid antigen tests for COVID-19 found an average sensitivity of about seventy-three percent in people with symptoms and about fifty-five percent in people without symptoms. Among people with symptoms, it was about eighty-one percent in the first week after symptoms began and about fifty-four percent in the second week. Specificity was above ninety-nine percent in both groups.

**Sarah:** It's the same kind of test in both groups, so most of the difference must come from the people being tested.

**Kiffer:** Mostly from how much virus the swab picks up. The review found that sensitivity fell steadily as the viral load fell, and viral loads tend to be highest early in a symptomatic illness. This is often called the spectrum effect. Sensitivity depends on what the cases look like, and specificity depends on what the non-cases look like.

**Sarah:** So a test evaluated in people with obvious, advanced disease will look more sensitive than it turns out to be in a screening program, where most cases are early and mild. And on the specificity side, I'd guess a hospital is a harder place for a test, because the people without the disease often have other illnesses that could produce a false positive.

**Kiffer:** That's the idea. There's evidence that it happens across many tests. A 2013 study in the Canadian Medical Association Journal looked at twenty-three meta-analyses of diagnostic tests. In eight of them, sensitivity or specificity changed significantly with the prevalence in the included studies, and overall, specificity tended to be lower where prevalence was higher. The authors suggested that differences in the kinds of patients studied were a likely explanation.

**Sarah:** So before you borrow a sensitivity from a published study, you check whether the people in that study look like the people you plan to test.

**Kiffer:** Yes, and you check who received the reference standard. Here's a made-up example. A study evaluates a new screening test in one thousand people, and only those who test positive go on to a biopsy. Suppose fifty people truly have the disease and the test picks up forty of them.

**Sarah:** Then the ten people it misses test negative and never get a biopsy, so the study never learns that they have the disease. Every confirmed case is a test positive, so the sensitivity comes out at one hundred percent, when the true value is forty out of fifty, or eighty percent.

**Kiffer:** That's verification bias. The QUADAS-2 tool mentioned in the lesson asks directly whether all patients received a reference standard. One way around the problem is to follow the people who tested negative and count how many are diagnosed with the disease over the following months.

**Sarah:** Question three. What can go wrong when you combine a test result with what you already know, or with a second test?

**Kiffer:** Let's take the first half with a real test, the rapid strep test used for children with sore throats.

**Sarah:** A 2016 Cochrane review of that test found a summary sensitivity of 85.6 percent and a specificity of 95.4 percent, compared with throat culture. In the studies where every child had both tests, the median share of children with strep was about thirty percent. So the likelihood ratio for a positive result is 0.856 divided by one minus 0.954. One minus 0.954 is 0.046, and 0.856 divided by 0.046 is about 18.6. For a negative result, it's one minus 0.856, which is 0.144, divided by 0.954. That's about 0.151.

**Kiffer:** Good. Now a child comes in, and based on the symptoms, the doctor puts the chance of strep at thirty percent. The rapid test is positive. What's the probability that the child has strep? Let's give everyone a few seconds to try it first.

*(Pause)*

**Sarah:** Okay. The pre-test probability is thirty percent and the likelihood ratio is 18.6, so 0.3 times 18.6 is 5.58. That's a probability of five hundred and fifty-eight percent, which can't be right.

**Kiffer:** That's a good catch. What went wrong?

**Sarah:** The likelihood ratio multiplies the odds, and I multiplied the probability. So first I turn thirty percent into odds. That's 0.3 divided by 0.7, which is about 0.429. Times 18.6 gives post-test odds of about 7.98. Then the probability is 7.98 divided by 8.98, which is about 0.89.

**Kiffer:** So about eighty-nine percent. Now do the negative result.

**Sarah:** 0.429 times 0.151 is about 0.065. Then 0.065 divided by 1.065 is about 0.061, so about six percent.

**Kiffer:** Notice something about that. If you'd made the same mistake with the negative result, thirty percent times 0.151 gives about 4.5 percent. That's wrong, but it looks believable, so no sanity check would catch it. The shortcut only comes close when the pre-test probability is small, because then the odds and the probability are nearly the same. At thirty percent they aren't. I'd convert to odds every time, whatever the numbers look like.

**Sarah:** And six percent is the leftover chance after a negative result. Isn't this where SnNout is supposed to apply?

**Kiffer:** SnNout works when the sensitivity is very high and the pre-test probability is modest. Here the sensitivity is about eighty-six percent, which isn't high enough for that, so a negative result still leaves about one child in sixteen with strep. That's part of the reason the Infectious Diseases Society of America's 2012 guideline recommends backing up a negative rapid test with a throat culture in children and adolescents, while saying it isn't routinely necessary in adults.

**Sarah:** Now the second half. What goes wrong with a second test?

**Kiffer:** Here's the harder example, with made-up numbers. We screen ten thousand people for a disease with a prevalence of one percent. Test A has a sensitivity of ninety percent and a specificity of ninety-five percent. Everyone who tests positive gets test B, which has a sensitivity of ninety percent and a specificity of ninety-eight percent. A person counts as positive only if both tests are positive.

**Sarah:** That's testing in series. With the formulas from the lesson, the combined sensitivity is 0.9 times 0.9, which is 0.81, and the combined specificity is one minus the product of the two false positive rates. That's 0.05 times 0.02, which is 0.001, and one minus 0.001 is 0.999.

**Kiffer:** Let's count it. One hundred people have the disease. Test A catches ninety of them, and test B confirms eighty-one. Nine thousand nine hundred people don't have the disease. Test A wrongly flags five percent of them, which is four hundred and ninety-five people, and test B wrongly flags two percent of those, which is about ten.

**Sarah:** So about ninety-one people end up positive, and eighty-one of them have the disease. That's a positive predictive value of about eighty-nine percent.

**Kiffer:** That two percent is an assumption hiding inside the formula. It says test B is just as likely to be fooled by someone test A got wrong as by any other healthy person. Statisticians call this conditional independence. Now suppose some healthy people carry a related infection that fools both tests. Will the positive predictive value be higher or lower than eighty-nine percent? Take a few seconds.

*(Pause)*

**Sarah:** It'll be lower. If the two tests tend to make the same mistakes, the second test clears fewer of the first test's false positives.

**Kiffer:** Right. Let's put a number on it. Suppose that among the four hundred and ninety-five people test A wrongly flagged, test B is wrong twenty percent of the time, because many of them have that related infection. Test B can still have a specificity of ninety-eight percent across all healthy people. What changes is how it performs on the people test A got wrong.

**Sarah:** Then twenty percent of four hundred and ninety-five is ninety-nine false positives. Add the eighty-one true positives and that's one hundred and eighty positives, so the positive predictive value is eighty-one divided by one hundred and eighty, which is 0.45.

**Kiffer:** So it falls from about eighty-nine percent to forty-five percent, with ten times as many false positives. The combined false positive rate is now ninety-nine out of nine thousand nine hundred healthy people, which is one percent, so the combined specificity is ninety-nine percent. The formula promised 99.9 percent.

**Sarah:** Does this happen with real tests? I'd expect it whenever two tests detect the same substance, because anything that fools one could easily fool the other.

**Kiffer:** That's one way it happens. Between 2020 and 2023, the World Health Organization supported fourteen countries to check HIV rapid tests for their national testing services, including whether pairs of tests gave false positives on the same specimens. In more than half of those countries, at least one pair did, and most of the countries changed the tests they use as a result. The same assumption sits behind multiplying two likelihood ratios one after the other, so it's worth asking whether two tests can be fooled by the same thing.

**Sarah:** Question four. How can screening make survival look better without anyone living longer?

**Kiffer:** Let's do it with one made-up person. She has a cancer that will cause her death at age seventy-four, whatever we do. Without screening, she develops symptoms and is diagnosed at seventy-one. With screening, the cancer is found at sixty-seven. Has screening helped her, and what happens to her five-year survival? Take a few seconds.

*(Pause)*

**Sarah:** Screening hasn't helped her, because she dies at seventy-four either way. Without screening she lives three years after diagnosis, so she doesn't count as a five-year survivor. With screening she lives seven years after diagnosis, so she does.

**Kiffer:** That's lead-time bias. An earlier diagnosis starts the survival clock sooner, and the survival figure improves even though her death comes at the same age.

**Sarah:** And length bias adds to it. A test given every year or two is more likely to find slow-growing tumours, because they spend longer in the stage where screening can detect them. The fast-growing ones tend to show up through symptoms between screens.

**Kiffer:** Right, so screen-detected cancers have a better outlook on average, partly because of which cancers screening tends to find.

**Sarah:** And overdiagnosis is the extreme case, where the disease that's found would never have caused any harm.

**Kiffer:** A well-known real example is thyroid cancer in South Korea. Thyroid ultrasound was offered as an add-on to a national cancer screening program that began in 1999. By 2011, the rate of thyroid cancer diagnosis was fifteen times the 1993 rate, while deaths from thyroid cancer stayed about the same.

**Sarah:** If the rate of diagnosis rose fifteen-fold and the death rate didn't move, most of the extra diagnoses were probably cancers that would never have caused harm.

**Kiffer:** That was the conclusion of the researchers who described it in the New England Journal of Medicine in 2014. In 2015, a Korean expert committee recommended against ultrasound screening for thyroid cancer in healthy people.

**Sarah:** So what's the right way to judge a screening program?

**Kiffer:** By deaths from the disease among everyone offered screening, compared with a similar group who weren't offered it, ideally in a randomized trial. That rate counts deaths in the whole group, so an earlier date of diagnosis doesn't change it, and harmless cancers found by screening don't lower it. Survival among screen-detected cases is the figure I'd treat with the most suspicion.

**Sarah:** Okay. Now the part where we disagree, which is screening for depression.

**Kiffer:** That brings back the PHQ-9.

**Sarah:** I lean toward routine screening in primary care. In 2023, the United States Preventive Services Task Force recommended screening all adults for depression, including pregnant and postpartum people and older adults. It concluded with moderate certainty that screening has a moderate net benefit. And without a questionnaire, doctors miss a lot. A 2009 meta-analysis in the Lancet found that general practitioners correctly identified depression in only about forty-seven percent of the patients who had it.

**Kiffer:** Those are real points. The Canadian Task Force on Preventive Health Care reached a different conclusion in 2013. It recommended that clinicians not routinely screen adults for depression, whether they were at average risk or in groups at higher risk. Its review found no high-quality evidence that screening for depression was effective, and it rated these as weak recommendations based on very low quality evidence.

**Sarah:** But a shortage of good trials doesn't show that screening fails. The questionnaire takes a few minutes, and a positive score leads to a conversation with the doctor. I think the cost of missing half the people with depression is larger than the cost of some extra conversations.

**Kiffer:** That's a fair argument, and it's the strongest one for screening. My concern is the arithmetic from question one. If one in ten patients has major depression, about six in ten positive screens are false positives, and the Canadian Task Force was concerned about false positive diagnoses leading to unnecessary treatment. There's also the third of Wilson and Jungner's principles, that facilities for diagnosis and treatment should be available. If people who screen positive then wait months for care, screening adds to the queue without helping much.

**Sarah:** Then the answer is to fund the follow-up care and keep screening.

**Kiffer:** I'd support screening where that care is in place. Where it isn't, I think the money does more good spent on the care first. And I'd want any new screening program evaluated, so that we learn whether the people offered screening actually do better.

**Sarah:** I'm still more in favour of screening than Kiffer is, but we agree on the practical rule. Before backing a screening program, ask whether people offered screening end up better off than people who aren't, and whether there's care ready for everyone who screens positive. And a positive screen leads to a clinical assessment before anyone is given a diagnosis.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, sensitivity is the probability of a positive result given disease, and the person holding the result wants the reverse. To get it you also need the specificity and the prevalence, and where the disease is uncommon, many positives will be false.

**Kiffer:** Second, every rule for combining information carries an assumption. Likelihood ratios multiply odds, so convert probabilities to odds first. The series formula and chained likelihood ratios both assume the tests don't share their mistakes.

**Sarah:** And third, judge a test by the people it was evaluated in, and judge a screening program by deaths among everyone offered screening. Survival among screen-detected cases can improve through lead time, length bias and overdiagnosis alone.

**Kiffer:** If you'd like more practice, rework the series testing example from this episode with your own numbers. Change how often test B is fooled by test A's false positives, and watch what happens to the positive predictive value.

**Sarah:** Next time, it's Lesson six, Measures of Association, where the two by two table comes back to compare exposed and unexposed groups with risk ratios, rate ratios and odds ratios.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson six.
