# Lesson 8: Time-to-Event Data

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson eight, Time-to-Event Data.

**Sarah:** It's for after you've worked through the lesson. We'll add some perspective on where these methods came from and how they're used, and some critique. We'll also work through a few of the questions students tend to find thorny, and do some extra worked examples, including a couple that are harder than the ones on the lesson page.

**Kiffer:** At a few points we'll ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the four questions. When should the clock start, and how can starting it at the wrong moment create an effect out of nothing? What exactly does a Kaplan-Meier curve tell you, and where does it run short of information? Why does one minus Kaplan-Meier overstate the risk when people can die of something else? And what does a hazard ratio mean in terms of risk? We'll finish with whether the hazard ratio deserves to be the headline result, because Kiffer and I see that a little differently.

**Kiffer:** Let's start with the clock.

**Sarah:** Question one. When should the clock start?

**Kiffer:** Here's a real case, and the researchers on both sides of it worked in Canada. In 2001, two researchers at the University of Toronto published a study in the Annals of Internal Medicine comparing actors and actresses who had won an Academy Award with other, less recognized performers. They reported that the winners' life expectancy was about 3.9 years longer.

**Sarah:** That's a big difference. I can think of possible explanations, like status, money, or people taking better care of their health once they're famous.

**Kiffer:** Those were the kinds of explanations people offered. Look at how the groups were formed, though. In the headline comparison, a performer counted as a winner from birth. Someone who won at sixty had to survive to sixty to win, so every year before the win was a year in which that winner, by definition, could not have died.

**Sarah:** So that's immortal time. The winners were credited with decades they were guaranteed to survive.

**Kiffer:** Exactly. In 2006, a team at McGill led by James Hanley published a reanalysis that counted each performer as a winner only from the year of the win. The advantage shrank to about one year, and it was no longer statistically significant.

**Sarah:** And this is the time zero rule from the lesson. A person should become eligible, have their exposure determined and start follow-up at the same moment. At birth, nobody knows who's going to win.

**Kiffer:** Right. Let's put numbers on it with a made-up example so you can see how the bias works. One thousand people are discharged from hospital after a heart attack and followed for up to two years for death. Three hundred of them start a cardiac rehabilitation program, on average four months after discharge. These numbers are hypothetical, and we've built them so that the program has no effect on death at all.

**Sarah:** So any difference we find comes entirely from the analysis.

**Kiffer:** Yes. The three hundred program users had twenty deaths in the four thousand person-months after they started. The seven hundred who never started had sixty-six deaths over twelve thousand person-months. A naive analysis calls the users exposed from the day of discharge, so it adds their waiting time to the exposed group. That's three hundred people times four months, or twelve hundred person-months, with no deaths in them, because you had to be alive to start the program.

**Sarah:** Then the exposed rate is twenty divided by five thousand two hundred, which is about 0.38 per hundred person-months. The unexposed rate is sixty-six divided by twelve thousand, which is 0.55 per hundred. The rate ratio is about 0.7, so the program seems to cut the death rate by thirty percent.

**Kiffer:** Here's a question for everyone listening. Where should those twelve hundred waiting months go, and what's the rate ratio once they're in the right place? Take a few seconds.

*(Pause)*

**Kiffer:** They belong to the unexposed group, because during those months the users hadn't started the program yet. The exposed rate becomes twenty divided by four thousand, which is 0.5 per hundred person-months. The unexposed group now has sixty-six deaths over twelve thousand plus twelve hundred, or thirteen thousand two hundred person-months, and that's also 0.5 per hundred. The rate ratio is one, and the effect is gone.

**Sarah:** And the bias moved both rates. Adding death-free months lowered the exposed rate, and taking those months away raised the unexposed rate.

**Kiffer:** That's the part people tend to miss. The fix is to treat program use as an exposure that changes over time, so each person's months before they start count as unexposed time, or to define the exposure at time zero.

**Sarah:** Question two. What exactly does a Kaplan-Meier curve tell you?

**Kiffer:** A bit of history first. Edward Kaplan and Paul Meier worked on the problem separately, and each submitted a paper to the Journal of the American Statistical Association. The editor persuaded them to merge the two. By Kaplan's own account, it took about four years of correspondence to reconcile their approaches, and the 1958 paper that resulted became one of the most cited papers in statistics.

**Sarah:** I'd like to try a calculation, because I think the bookkeeping is where people slip.

**Kiffer:** Here's a made-up dataset. Twelve university athletes return to play after an ankle sprain and are followed for up to twelve months for a first re-injury. One is re-injured at month two. One graduates and leaves at month three. At month four, two are re-injured and one transfers to another school. One is re-injured at month seven. One quits the sport at month eight. Two are re-injured at month ten. The last three finish twelve months without a re-injury.

**Sarah:** Okay. At month two, twelve are at risk and one is re-injured, so the estimate is eleven twelfths, about 0.917. The athlete who leaves at month three doesn't change the estimate. At month four there are two re-injuries, and the athlete who transferred at month four has left, so nine are at risk. Nine minus two, over nine, is seven ninths, and 0.917 times seven ninths is about 0.713.

**Kiffer:** Check the risk set at month four. Who counts?

**Sarah:** Everyone whose recorded time is month four or later. That's the two who were re-injured, the one who transferred, and the seven with later times. So it's ten. I took the athlete who transferred out too early.

**Kiffer:** Right. When an event and a censoring share a time, the convention is that the event comes first. That athlete was seen injury-free at month four, so they were at risk when those two injuries happened.

**Sarah:** So the step at month four is eight tenths, and 0.917 times eight tenths is about 0.733. At month seven, seven are at risk and one is re-injured, so I multiply by six sevenths and get about 0.629. The athlete who quits at month eight leaves. At month ten, five are at risk and two are re-injured, so I multiply by three fifths and get about 0.377. It stays there to month twelve.

**Kiffer:** Here's the next question. What's the median time to re-injury for these twelve athletes? Take a few seconds.

*(Pause)*

**Sarah:** The tempting answer is month seven, because 0.629 is the last value above one half, or something between seven and ten if you draw a line between the steps. The median is the first time the curve falls to 0.50 or below. At month seven it's still at 0.629. It first falls below one half at month ten, so the median is ten months.

**Kiffer:** Now be careful about which number you report. The 0.377 is the probability of getting through twelve months without a re-injury. The twelve-month risk of re-injury is one minus that, about 0.623, or sixty-two percent. Mixing up the survival and the risk is one of the most common slips in reading these curves, and it's especially easy to make when you're describing a curve out loud.

**Sarah:** Let me sanity-check the risk. Six of the twelve were re-injured, so the risk can't be below one half. Three were censored early, so it can't be above nine twelfths, which is seventy-five percent. Sixty-two percent sits between the two.

**Kiffer:** Good. Now some critique. The drop at month ten rests on five athletes. If one of the three who finished had been re-injured, the curve would have dropped again, from about 0.38 to about 0.25. That's why the numbers at risk under a curve matter, and it's also why two small groups can have curves that look very different and still be compatible with chance. The information comes from the events.

**Sarah:** There's also the athlete who quit at month eight. If people quit because the ankle felt unstable, that censoring tells us something about their risk.

**Kiffer:** That would be informative censoring, and it would make the curve too optimistic. The follow-up times alone can't tell you whether it's happening, so I'd want the reason for every departure recorded from the start.

**Sarah:** Question three. Why does one minus Kaplan-Meier overstate the risk when people can die of something else?

**Kiffer:** This one combines the life table with competing events, and it's harder than the examples on the lesson page. Made-up numbers again. Two hundred adults aged eighty and over are followed for two years for a first diagnosis of dementia through linked health records, so nobody is lost except by dying. In the first year, twenty are diagnosed with dementia and thirty die without it. The second year is the same, twenty diagnosed and thirty deaths.

**Sarah:** So forty of the two hundred were diagnosed. That's twenty percent, and because everyone was followed until diagnosis, death or two years, twenty percent is simply what happened.

**Kiffer:** Exactly. Now suppose an analyst treats the deaths as censoring and builds an actuarial life table. A Kaplan-Meier analysis that censored the deaths would behave the same way. Before we compute it, will one minus the life-table survival come out above twenty percent, below it, or right on it? Take a few seconds.

*(Pause)*

**Sarah:** Above. Treating the deaths as censoring assumes that the people who died had the same future risk of dementia as the people who survived. So the method imagines that some of them would have gone on to be diagnosed.

**Kiffer:** Let's see by how much. Can you run the first year?

**Sarah:** Two hundred enter, and the thirty deaths are treated as withdrawals, so the effective number at risk is two hundred minus half of thirty, which is one hundred eighty-five. Twenty divided by one hundred eighty-five is about 0.108, so the probability of getting through the year without dementia is about 0.892.

**Kiffer:** And the second year?

**Sarah:** One hundred fifty enter, which is two hundred minus thirty minus twenty. The effective number at risk is one hundred fifty minus fifteen, or one hundred thirty-five. Twenty divided by one hundred thirty-five is about 0.148, so the conditional survival is about 0.852. Then 0.892 times 0.852 is about 0.760.

**Kiffer:** So one minus the life-table survival is about 0.24.

**Sarah:** Twenty-four percent, against the twenty percent who were actually diagnosed. That's about one fifth higher than what happened.

**Kiffer:** And the gap grows as deaths become more common. Even if the people who died would have had the same risk of dementia as the survivors, the twenty-four percent estimates the risk in a hypothetical world where nobody could die of anything else. The twenty percent is the proportion of real people who developed dementia, which is the number you'd want for planning memory clinics or long-term care beds.

**Sarah:** Is the hypothetical number ever the one you want?

**Kiffer:** Sometimes, yes. Cancer survival is a Canadian example. The Canadian Cancer Statistics reports give net survival, which the Canadian Cancer Society defines as the probability of surviving cancer in the absence of other causes of death. It's the preferred measure for comparing cancer survival between populations, because it adjusts for differences in their background risk of dying. That's a deliberate choice of the hypothetical world, made for comparison.

**Sarah:** So decide which question you're asking before you choose the method.

**Kiffer:** Yes. If you want the proportion of people who will actually have the event, use methods built for competing risks, which estimate the cumulative incidence directly. Here, because nobody was lost for any other reason, that's just forty divided by two hundred. Once ordinary loss to follow-up is added, you need those methods to get it.

**Sarah:** Question four. What does a hazard ratio mean in terms of risk?

**Kiffer:** Take a made-up trial of a home-visiting program for people with heart failure. Over three years, it reports a hazard ratio of 0.5 for a first hospital admission.

**Sarah:** And a headline might say the program halves the risk of being admitted.

**Kiffer:** That's the common misreading. The hazard ratio says that, among people who haven't been admitted yet, the admission rate in the program group is half the rate in the control group at any point in the three years, assuming the hazards are proportional. To get risks, you need a survival curve. Suppose sixty-five percent of the control group are still free of admission at three years. Using the relationship from the lesson, the program group's survival is the control group's survival raised to the power of the hazard ratio. What's the three-year risk ratio? Take a few seconds.

*(Pause)*

**Sarah:** The program group's survival is 0.65 raised to the power 0.5, which is the square root of 0.65, about 0.806. So the three-year risks are thirty-five percent in the control group and about nineteen percent with the program. Dividing 0.194 by 0.35 gives a risk ratio of about 0.55. That's a drop in risk of about forty-five percent, a little short of half.

**Kiffer:** And the risk difference is about sixteen fewer people admitted per hundred over three years, which is the number a patient or a health authority can use. The risk ratio sits closer to one than the hazard ratio, and the gap widens as the event becomes more common, because risks can't go above one.

**Sarah:** Okay. Now the part where we disagree. I'd like to see the hazard ratio lose its place as the headline result. It compares rates among the people who are still event-free, which is hard to explain. It's routinely read as a risk ratio, as we just heard. And it's a complete summary only if the hazards stay proportional.

**Kiffer:** Those are fair points. What would you put in its place?

**Sarah:** Absolute risks at prespecified times, and the restricted mean survival time. In 2013, Royston and Parmar argued that the hazard ratio can't be recommended as a general measure of treatment effect in randomized trials, and they proposed the restricted mean as a practical alternative. And Hernán's 2010 paper, The Hazards of Hazard Ratios, makes a deeper point. Even in a randomized trial, the people still at risk later on are a selected group, because the most susceptible people in the higher-risk group have already had the event.

**Kiffer:** I agree that the selection point is a strong one. Where I'd push back is on demoting the hazard ratio. When the hazards are roughly proportional, one number summarizes the whole follow-up, and the Cox model can adjust for several confounders at once, which matters a great deal in observational studies. Under that same condition, a hazard ratio can also be compared across studies with different lengths of follow-up, where a two-year risk and a five-year risk can't be compared directly.

**Sarah:** The restricted mean is easier to explain, though. Over twenty-four months, the exposed group had, say, six fewer event-free months on average. Most people can follow that.

**Kiffer:** It is easier, and I like it as a companion to the hazard ratio. It has a choice built into it, too. The restricted mean depends on the horizon you pick, so a team that picks the horizon after seeing the curves can make a difference look bigger or smaller. It has to be fixed in the protocol, like the intervals in a life table.

**Sarah:** The same is true of the time you choose for an absolute risk. So we both want those choices made in advance.

**Kiffer:** We do, yes. And I think a hazard ratio should always be reported with the curves and with absolute risks at stated times.

**Sarah:** I'm more skeptical of the hazard ratio than Kiffer is, but we agree on the practical rule. Show the Kaplan-Meier curves with their numbers at risk, report absolute risks or the restricted mean at times fixed in the protocol, and put any hazard ratio alongside those, with a check of proportional hazards.

**Kiffer:** Let's pull it together with three things to take away.

**Sarah:** First, time zero sets up everything else. When eligibility, exposure and the start of follow-up don't line up, event-free time gets credited to the wrong group, as it did for the Oscar winners and in the rehabilitation example.

**Kiffer:** Second, know which quantity you're reading. A Kaplan-Meier curve shows the probability of staying event-free, the risk is one minus that, the tail is only as good as its numbers at risk, and when competing events are common, one minus Kaplan-Meier describes a hypothetical world.

**Sarah:** And third, a hazard ratio is a ratio of rates among people still at risk. Before you explain one to anyone, translate it into absolute risks at a stated time.

**Kiffer:** If you'd like more practice, rework the dementia life table from this episode. Change the number of deaths in each year, and watch how far one minus the life-table survival drifts from the proportion who were actually diagnosed.

**Sarah:** That brings us to the end of the course. Lesson eight is the last lesson, and this is the last Office Hours episode for this course. Thanks for working through the material with us.

**Kiffer:** Take care, everyone.

**Sarah:** And all the best with whatever you study next.
