HSCI 341, Lesson 8

Time-to-Event
Data

Fundamental Epidemiological Concepts and Approaches

Learning objectives for this lesson:

  • Define time-to-event data and specify time zero, the event and the end of follow-up for a study.
  • Distinguish right, left and interval censoring from late entry, and explain the difference between informative and non-informative censoring.
  • Explain why treating censored participants as event-free, or dropping them, misstates the risk of the event, and recognize immortal time and lead-time bias as errors in the choice of time zero.
  • Compute an actuarial life table by hand and state the assumptions of the actuarial method.
  • Compute a Kaplan-Meier estimate by hand and read survival at a chosen time, the median survival time, censoring marks and numbers at risk from a Kaplan-Meier curve.
  • Explain the logic of the log-rank test as a comparison of observed and expected events.
  • Interpret the hazard as an event rate and the hazard ratio as a ratio of rates, state the proportional hazards idea in plain language, and explain why a hazard ratio differs from a risk ratio.
  • Write the time-to-event subsection of a study protocol, including censoring rules, an analysis plan and the number of events needed.

Expected completion time: about two and a half hours. This lesson follows Lesson 7: Designing Against Bias and is the last lesson of HSCI 341.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.

Lesson 8 · HSCI 341

Time-to-Event Data

Censoring, life tables, Kaplan-Meier estimation and the hazard ratio, taught for people who design studies.

About two and a half hours across four sections and a final assessment
Why this lesson

Estimating incidence when follow-up is incomplete

Lesson 4 gave us

Incidence risk needs complete follow-up, and the incidence rate summarizes the whole period in one number.

This lesson adds

Survival methods estimate how the probability of remaining event-free changes over time while using every participant's follow-up.

Running case (hypothetical)

The Campus Connection Follow-up

600 participants

All are aged 18 to 29 and screen negative at baseline.

24 months

Each participant completes an online PHQ-9 every month.

200 and 400

Participants with high and low baseline loneliness are compared.

The event is the first monthly PHQ-9 score of 10 or more.

Lesson map

Four sections

1. Design and censoring

Time zero, the event, follow-up, and the types of censoring are defined.

2. Actuarial life table

Cumulative survival is built from fixed intervals by hand.

3. Kaplan-Meier

The estimate is computed at each event time and the curve is read.

4. Comparing groups

The log-rank test, the hazard ratio and the protocol are covered.

How to work through it

Using this lesson

  • Each section pairs worked examples with calculators that show every step.
  • Each reflection reveals a model answer after you save your own response.
  • Each knowledge check must be passed before the next section opens.
Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, methods and people you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Designing a Time-to-Event Study
Time-to-event data Data in which each participant contributes a length of time from a defined starting point to a defined event, together with a status recording whether the event was observed. Also called survival data.
Time zero The moment at which the clock starts for each participant. It should be the moment at which eligibility is met, exposure is assigned and follow-up begins (Hernán et al., 2016).
Study time scale Time measured from each participant's own time zero, so that every participant starts at zero regardless of the calendar date of enrolment.
Event The outcome that stops the clock, defined by a criterion, an ascertainment method and schedule, and a dating rule that are the same in every exposure group.
Status variable The variable recording how a participant's follow-up ended, usually coded 1 for an observed event and 0 for censoring.
Risk set The participants who are still under observation and event-free just before a given time, and who could therefore have the event at that time.
Censoring and Incomplete Follow-Up
Right censoring Incomplete follow-up in which the event, if it occurs, would occur after the participant's last observation. It arises from loss to follow-up, withdrawal, moving away and the administrative end of a study.
Administrative censoring Censoring of participants who are still event-free when the study reaches its planned end. It is usually non-informative.
Left censoring Incomplete information in which the event is known to have occurred before observation began, at an unknown time.
Interval censoring Incomplete information in which the event is known only to have occurred between two observation times, as when events are detected at scheduled visits.
Late entry (left truncation) Entry of a participant into the risk set some time after their time zero. People who had the event before they could enter are never seen, so late entrants are counted at risk only from entry.
Non-informative censoring The assumption that participants who are censored at a given time have the same future risk of the event as participants in the same group who remain under observation.
Informative censoring Censoring that is related to a participant's prognosis, so that those who leave have a different future risk from those who remain. It biases survival estimates.
Competing event An event, such as death from another cause, that makes the event of interest impossible. Treating it as censoring overstates the cumulative incidence of the event of interest (Austin et al., 2016).
Immortal time bias Bias that arises when membership of an exposure group depends on remaining event-free for a period after time zero, so that the period is wrongly credited to the exposure (Suissa, 2008). Taught in HSCI 230 Lesson 10.
Lead-time bias Apparent extension of survival from diagnosis that arises when one group's condition is diagnosed earlier, for example by screening, without any delay in the event. Taught in HSCI 230 Lesson 10.
Estimating Survival
Actuarial life table A method that divides follow-up into fixed intervals, estimates the conditional probability of surviving each interval and multiplies these probabilities to obtain cumulative survival (Cutler & Ederer, 1958).
Effective number at risk In an actuarial life table, the number entering an interval minus half the withdrawals during it, reflecting the assumption that withdrawals were at risk for half the interval on average.
Conditional probability of survival The probability of getting through an interval or past an event time without the event, among those who were event-free at its start.
Survivor function The probability of remaining event-free beyond time t, written S(t). One minus the survivor function is the cumulative risk of the event by time t.
Kaplan-Meier estimator The product-limit estimator of the survivor function, which multiplies the conditional probabilities of getting past each observed event time and keeps censored participants in the risk set until they leave (Kaplan & Meier, 1958).
Median survival time The earliest time at which the estimated survival falls to 0.50 or below. If the curve never reaches 0.50 during follow-up, the median is reported as not reached.
Numbers-at-risk table A table beneath a Kaplan-Meier plot giving the number of participants still at risk at regular times, which shows how much information supports each part of the curve.
Greenwood's formula The standard formula for the variance of a Kaplan-Meier estimate, used to compute confidence intervals that widen as the number at risk falls (Greenwood, 1926).
Restricted mean survival time The area under a survival curve up to a chosen time, interpreted as the average event-free time within that horizon. Differences in it do not depend on proportional hazards (Royston & Parmar, 2013).
Comparing Groups
Log-rank test A test of the null hypothesis that two or more groups have the same survival, based on summing observed minus expected events across event times and referring the result to a chi-squared distribution.
Hazard The rate at which events occur among participants who are still event-free at a given time, per unit of time. It can rise, fall or stay constant over follow-up.
Hazard ratio The hazard in one group divided by the hazard in another at the same time since time zero. Under constant hazards it equals the incidence rate ratio.
Proportional hazards The assumption that the ratio of the hazards in two groups is the same at every time during follow-up, although each hazard may change over time.
Cox proportional hazards model A regression model that estimates adjusted hazard ratios while leaving the shape of the baseline hazard unspecified (Cox, 1972). HSCI 341 interprets the hazard ratio from a Cox model without fitting one.
Events-based sample size A sample size calculation that first finds the number of events needed to detect a hazard ratio (Schoenfeld, 1983) and then divides by the proportion expected to have the event.
Key People
John Graunt London haberdasher whose 1662 analysis of the Bills of Mortality included an early life table estimating how many of 100 people would survive to successive ages.
Major Greenwood British epidemiologist and statistician whose 1926 report on the natural duration of cancer gave the variance formula still used for survival estimates.
Sidney J. Cutler and Fred Ederer Statisticians at the United States National Cancer Institute who set out the actuarial life-table method for clinical follow-up studies in 1958.
Edward L. Kaplan and Paul Meier Statisticians who published the product-limit estimator of survival from incomplete observations in 1958, now known as the Kaplan-Meier estimator.
Nathan Mantel American biostatistician whose 1966 paper on evaluating survival data is one of the origins of the log-rank test.
Richard Peto British epidemiologist and statistician who, with Julian Peto, developed rank tests for comparing survival curves in 1972 and later led large collaborative trial overviews.
David R. Cox British statistician who introduced the proportional hazards regression model in 1972, the most widely used model for time-to-event data.
David Schoenfeld American biostatistician who derived the number of events needed to detect a hazard ratio with a given power in 1983.
No matching entries. Try a different search term.
Section 1

Time-to-Event Questions and Study Design

⏱ Estimated reading time: 35 minutes

Section 1 of 4

Time-to-Event Questions and Study Design

Time zero, the event, follow-up, censoring and the errors that start at time zero.

The outcome

A time and a status for every participant

Time

The interval from time zero to the event, to the last event-free contact, or to the end of the study.

Status

The status is 1 if the event ended follow-up and 0 if the participant was censored.

A censored time of 8 months means that the event time is known only to exceed 8 months.

Three design decisions

Time zero, the event and the end of follow-up

Time zero

Eligibility, exposure assignment and the start of follow-up coincide.

The event

The criterion, ascertainment schedule and dating rule are the same in every group.

End of follow-up

Follow-up ends at the event, the last event-free contact or the end of the study.

Types of incomplete follow-up

Right, left and interval censoring, and late entry

Right censoring

Loss to follow-up, withdrawal, moving away and the end of the study all produce it.

Left censoring

The event happened before observation began, at an unknown time.

Interval censoring

The event is known only to fall between two assessments.

Late entry

A participant is counted at risk only from the date of entry.

The key assumption

Informative and non-informative censoring

Non-informative

Censored participants have the same future risk as those who remain under observation.

Informative

Censoring is related to prognosis, so survival estimates are biased.

  • Retention procedures keep losses small.
  • Recorded reasons and baseline comparisons show who left.
  • A worst-case sensitivity analysis tests whether the conclusions survive.
Why shortcuts fail

Ten participants, three censored early

Two shortcuts and the limits set by the data
\[ \frac{E}{N} = \frac{5}{10} = 0.50 \qquad \frac{E}{N-C} = \frac{5}{7} = 0.71 \qquad \frac{E}{N} \le \text{risk} \le \frac{E+C}{N} = 0.80 \]

The Kaplan-Meier estimate of the 24-month risk for these participants is 0.64, which lies within the limits.

Errors that start at time zero

Immortal time and lead time

Immortal time bias

Exposure defined after time zero credits event-free waiting time to the exposure.

Lead-time bias

Earlier diagnosis starts the clock earlier without delaying the event.

HSCI 230 Lesson 10 treats both biases in detail.

Carry forward

Into the life table

  • Each participant contributes a time and a status defined by the protocol.
  • Censoring is expected, and the methods assume that it is non-informative.
  • The life table uses censored follow-up up to the point of censoring and assumes nothing after it.

Introduction and Overview

Many questions in public health research concern when an event happens as well as whether it happens. A trial of a smoking cessation program asks how long participants remain abstinent, a cohort study of young adults asks how long people who start without depressive symptoms remain free of them, and a tuberculosis program asks how long after diagnosis people complete treatment. In each case the outcome for every participant is a length of time from a defined starting point to a defined event. Data of this kind are called time-to-event data, and the methods for analyzing them are called survival analysis. The name reflects the origin of these methods in studies of mortality, although the event can be any well-defined change of state.

Time-to-event analysis extends Lesson 4. That lesson defined incidence risk as the proportion of an at-risk group that develops an outcome over a defined period, and incidence rate as new cases divided by person-time. Both measures make assumptions about follow-up. The risk requires that everyone is followed for the whole period, and the rate treats a person-month early in follow-up as equivalent to a person-month late in follow-up. Real cohorts rarely meet the first condition, because some participants leave before the period ends, and the second is often implausible. The methods in this lesson estimate the probability of remaining event-free over time while using all of the follow-up that each participant provides. Lesson 6 compared groups with the incidence rate ratio, and Section 4 of this lesson presents the hazard ratio as the time-to-event counterpart of that measure.

The lesson is conceptual and is written for people who design studies. This section sets out the design decisions that define a time-to-event outcome and the forms of incomplete follow-up called censoring. Section 2 computes an actuarial life table by hand, Section 3 computes a Kaplan-Meier estimate by hand and shows how to read a Kaplan-Meier curve, and Section 4 compares groups, explains the hazard ratio and the proportional hazards idea, and lists what a protocol with a time-to-event outcome must specify. These decisions belong in the methods of a study protocol, the written plan for a study as HSCI 207 Lesson 6 Section 3.1 defines it, and in a funding proposal they are summarized in the analysis plan.

Learning Objectives for this section

  • Explain what makes an outcome time-to-event data and why both a time and an event status are recorded for each participant.
  • Specify time zero, the event and the end of follow-up for a study, and convert follow-up records into a time and a status for each participant.
  • Distinguish right, left and interval censoring from late entry, and identify the sources of right censoring in a study description.
  • Explain the difference between informative and non-informative censoring and describe design steps that make informative censoring less likely and easier to detect.
  • Show why treating censored participants as event-free, or dropping them from the analysis, misstates the risk of the event.
  • Recognize immortal time bias and lead-time bias as consequences of the choice of time zero.

The section begins with the two parts of a time-to-event outcome and the kinds of question that need them, and then introduces the running case that the whole lesson uses.

Questions About Whether and When

A time-to-event outcome has two parts for each participant. The first is a length of time, measured from a starting point that is defined in the same way for everyone. The second is a status that records how that time ended: either the event was observed, or observation stopped before any event was seen. Analysts usually code the status as 1 for an observed event and 0 for a participant whose observation stopped without the event, who is described as censored. Neither part is sufficient alone. Two participants who were each followed for eight months carry very different information if one had the event in month eight and the other moved away in month eight without having it.

Two features distinguish these data from other continuous outcomes. Follow-up times cannot be negative and are usually right-skewed, because a few participants remain event-free for a long time. More fundamentally, the event time of a censored participant is known only to exceed the time at which observation stopped. A participant who was followed for eight months without the event and then moved away tells the investigator that their event time is longer than eight months, and nothing more. Survival methods use this partial information. Methods that ignore it, or that treat the censoring time as if it were an event time, give biased answers, as the last part of this section shows.

The question being asked also matters. A question about the risk of the event by a fixed time (what proportion of participants screen positive for depressive symptoms within 24 months) can be answered with the methods of this lesson even when follow-up is incomplete. A question about timing (how quickly positive screens accumulate, and whether they accumulate faster in one group) can only be answered with them, because it concerns the whole course of follow-up.

Box 8.1 describes the hypothetical study that serves as the running case for this lesson. It extends the Campus Connection Study from Lesson 7 into a follow-up of young adults who screen negative for depressive symptoms at baseline, and its event is a first positive screen.

Box 8.1: Case: The Campus Connection Follow-up (hypothetical)

The Campus Connection Study, the running example of HSCI 341 Lesson 7, is a cross-sectional survey of adults aged 18 to 29 in British Columbia on loneliness and depressive symptoms. This lesson imagines a follow-up extension of that study, and all of its numbers are hypothetical teaching data. Participants whose baseline score on the Patient Health Questionnaire-9 (PHQ-9) is below 10, the threshold commonly used for at least moderate depressive symptoms (Kroenke et al., 2001), are invited to complete a brief online PHQ-9 each month for up to 24 months. The event is the first monthly score of 10 or more, dated to the month in which it occurred. The exposure is high versus low loneliness at baseline. Six hundred participants enrol: 200 report high loneliness and 400 report low loneliness.

The running case fixes the setting, and the next part turns to the three design decisions that define its time-to-event outcome.

Three Design Decisions: Time Zero, the Event and the End of Follow-Up

Before any data are collected, a time-to-event protocol must define when the clock starts, what stops it, and how observation ends for participants who do not have the event. These three decisions determine what the time and status variables mean, and errors in any of them cannot be repaired at the analysis stage.

Box 8.2 recalls the thought experiment from HSCI 230 Lesson 3 that later work calls the target trial, which guides the first of these decisions, the choice of time zero.

Box 8.2: Recall: The Target Trial

HSCI 230 Lesson 3, Section 1 (A Unified Approach to Study Design), described the thought experiment that Hernán (2005) recommended: specify an observational study as if it were a randomized trial, naming the study group and its selection, the assignment to exposure, the procedures for follow-up and the detection of the outcome. Later work by Hernán and colleagues calls this hypothetical trial the target trial, and its protocol is the template for the observational study. In a trial, eligibility is confirmed, the treatment is assigned and follow-up begins at the same moment, so no participant contributes follow-up before being eligible or before the exposure group is known. An observational study that starts the clock at a moment with the same alignment avoids the time-related biases described at the end of this section.

Retrieval question. In a randomized trial, which event marks the moment at which eligibility, assignment and the start of follow-up coincide?

Show the answer▼

Randomization. Eligibility is confirmed just before it, the treatment group is determined by it, and follow-up for the outcome starts from it, so it serves as time zero.

Time zero

Time zero is the moment at which the clock starts for each participant. Common choices are the date of enrolment, the date of randomization in a trial, the date of diagnosis in a clinical cohort, and the date of a first prescription in a study of a medication. A good choice satisfies three conditions at the same moment: the participant meets the eligibility criteria, the participant's exposure group is determined, and follow-up for the event begins. Hernán et al. (2016) describe this alignment as part of specifying the target trial that an observational study tries to emulate, and they show that several avoidable biases arise when the three conditions are met at different times. In the Campus Connection Follow-up, time zero is the date on which the participant completes the baseline survey, because that is when eligibility (a PHQ-9 score below 10) is confirmed, when loneliness is measured, and when monthly follow-up begins.

Enrolment is usually spread over weeks or months, so time zero falls on different calendar dates for different participants. Survival analysis therefore measures time on a study time scale, defined as the time since each person's own time zero, and every participant starts at zero on that scale. Some cohort studies of chronic disease use age as the time scale instead, so that participants are compared with others of the same age. Figure 8.1 shows ten participants from the high-loneliness group of the Campus Connection Follow-up on the study time scale.

The event

The event must be defined precisely enough that two investigators reviewing the same record would agree on whether and when it occurred. The definition states the criterion (a PHQ-9 score of 10 or more), the method of ascertainment (the monthly online questionnaire) and the dating rule (the event is assigned to the month of the first qualifying score). The ascertainment schedule should be the same in every exposure group. If participants with high loneliness were contacted more often than others, their events would be detected sooner, and the comparison between groups would reflect the schedule as well as any effect of loneliness. Most analyses use the first occurrence of the event, so a participant leaves the group at risk once the event has happened. A composite event, such as hospitalization or death, counts whichever component occurs first, and each component needs its own definition.

Follow-up and how it ends

Each participant's follow-up ends at the earliest of three dates: the date of the event, the last date on which the participant was known to be event-free before leaving the study, and the administrative end of the study. The time variable is the interval from time zero to that date, and the status variable records whether the event ended it. Table 8.1 converts the follow-up records of ten participants from the high-loneliness group into times and statuses. Sections 2 to 4 use these ten participants, together with ten from the low-loneliness group, whenever a short dataset keeps the arithmetic manageable.

Table 8.1. Follow-up records of ten high-loneliness participants in the hypothetical Campus Connection Follow-up, reduced to a time and a status.

ParticipantWhat happenedTime (months)Status
1Screened positive in month 221 (event)
2Withdrew consent after the month 3 questionnaire30 (censored)
3Screened positive in month 551 (event)
4Screened positive in month 551 (event)
5Stopped responding after the month 8 questionnaire80 (censored)
6Screened positive in month 10101 (event)
7Moved overseas after the month 12 questionnaire120 (censored)
8Screened positive in month 15151 (event)
9Completed 24 months without a positive screen240 (censored)
10Completed 24 months without a positive screen240 (censored)
Event (status 1) Censored (status 0) How it ended P1 Event P2 Withdrew P3 Event P4 Event P5 Lost P6 Event P7 Moved away P8 Event P9 Study ended P10 Study ended 0 6 12 18 24 Months since time zero (baseline survey)
Figure 8.1. Ten high-loneliness participants in the hypothetical Campus Connection Follow-up on the study time scale. Each line runs from time zero to the end of that participant’s follow-up; red circles are events and vertical bars are censoring times. The dashed line marks the administrative end of the study at 24 months.

The protocol should also state what counts as loss to follow-up, because participants in an online cohort rarely announce that they are leaving. A common rule is that a participant who misses a stated number of consecutive questionnaires, for example three, is censored at the date of the last completed questionnaire. Writing the rule in advance prevents the analyst from choosing censoring dates after seeing the data, and it makes the rule the same for every exposure group.

Time zero, the event and the rules for ending follow-up together reduce each participant's record to a time and a status, as Table 8.1 does for ten participants. Five of those participants, numbers 2, 5, 7, 9 and 10, have a status of 0, and the next part of the section describes the forms that this incomplete information can take.

Censoring and Its Types

Censoring occurs when a participant's event time is only partly known. The form that standard survival methods handle is right censoring, in which the event, if it occurs at all, would occur after the participant's last observation. Right censoring has several sources, and a protocol should anticipate each of them. Click each card to see how the source arises in the Campus Connection Follow-up.

Loss to follow-upClick to learn more
Withdrawal of consentClick to learn more
Administrative censoringClick to learn more
Leaving the study populationClick to learn more
Competing eventsClick to learn more

Two other forms of censoring arise less often in prospective cohorts. Left censoring occurs when the event happened before observation began at an unknown time. In a study of the age at which adolescents first use cannabis, a participant who reports at the first interview that they have already used it is left-censored, because the event is known only to have occurred before that interview. The Campus Connection Follow-up avoids left censoring by excluding participants who screen positive at baseline. Interval censoring occurs when the event is known only to have occurred between two observation times. Every event in the Campus Connection Follow-up is interval-censored within a month, because a score of 10 or more in month 5 means that symptoms reached the threshold at some point after the month 4 questionnaire. Treating the month of detection as the event time is acceptable when the gap between assessments is short relative to the length of follow-up. When assessments are a year apart, methods designed for interval-censored data are needed.

Late entry, also called left truncation, differs from censoring. It arises when participants join the group at risk some time after their time zero, for example when time zero is the date of diagnosis but people are recruited from a clinic months or years later. Anyone who had the event before they could be recruited is never seen. A late entrant is therefore counted as at risk only from the date of entry, and an analysis that counts the earlier time as observed time overstates early survival. Figure 8.2 places the four situations on one time axis.

A. Right censoring Event-free to month 14, then lost B. Left censoring Event before the first visit C. Interval censoring Found at month 12, began after 6 D. Late entry Followed from month 5 onward event time unknown, later event somewhere here first visit: event already happened event in this window not observed entry 0 6 12 18 24 Months since time zero
Figure 8.2. Four kinds of incomplete information on one time axis. Right censoring (A) is the form handled by all of the methods in this lesson. Left censoring (B) and interval censoring with wide gaps between assessments (C) need specialized methods. A late entrant (D) is counted as at risk only from the date of entry.

Right censoring, the first situation in Figure 8.2, is the form of incomplete information that every method in the rest of the lesson handles. Those methods also make an assumption about why right-censored participants leave the study, and the next part examines that assumption.

Informative and Non-Informative Censoring

Every method in this lesson rests on the assumption that censoring is non-informative. The assumption states that, at any time during follow-up, participants who are censored have the same future risk of the event as participants in the same group who remain under observation. Censoring then removes people from the risk set without changing the risk of the people who remain, and the estimate for the remaining participants applies to the whole group.

Administrative censoring usually satisfies the assumption, because the closing date of a study has nothing to do with any participant's health. Loss to follow-up and withdrawal may not. In the Campus Connection Follow-up, a participant whose mood is worsening may stop answering monthly questionnaires before their score reaches 10. If participants like this are censored, the people who remain in the risk set are healthier than the group as a whole, fewer events are observed than actually occur, and the estimated probability of remaining event-free is too high. Censoring that depends on the participant's prognosis in this way is called informative censoring. If it is more common in one exposure group than in the other, it also biases the comparison between groups.

The assumption cannot be tested from the observed data, because the event times of censored participants are unknown. Design decisions make the assumption more plausible and its failure easier to assess. A protocol can include retention procedures that keep losses small (reminders by more than one channel, short questionnaires and modest incentives for completion), a record of the reason for every withdrawal, a plan to compare the baseline characteristics of participants who were censored early with those who were not, and a sensitivity analysis that recomputes the estimates under extreme assumptions, such as the assumption that every participant who was lost to follow-up had the event at the time of censoring. If the conclusions hold under those assumptions, informative censoring is an unlikely explanation for them.

Box 8.3 describes competing events, which the methods in this lesson treat as censoring even though a person who has had one can no longer have the event of interest.

Box 8.3: Competing events

A competing event is an event that prevents the event of interest from occurring. In a study of time to a first diagnosis of dementia among older adults, death from another cause is a competing event. The methods in this lesson treat a competing event as censoring, which assumes that the person could still have the event of interest later. Because that assumption is false for a person who has died, one minus the Kaplan-Meier estimate overstates the proportion of people who actually develop the event of interest, and the overstatement grows as competing events become more common. Methods that estimate the cumulative incidence in the presence of competing risks are described by Austin et al. (2016). Death is rare among adults aged 18 to 29, so it has little effect in the Campus Connection Follow-up, but the protocol should still name any competing events and state how they will be handled.

Non-informative censoring is an assumption that design can make more plausible and that the observed data cannot test. The next part shows what happens when the ten participants in Table 8.1 are summarized with ordinary methods that ignore censoring altogether.

Why Ordinary Summaries Fail

Consider the ten participants in Table 8.1 and the question of what proportion screen positive within 24 months. Five participants had the event (E = 5), three were censored before 24 months (C = 3), and two completed follow-up without the event. Two shortcuts are common, and both give the wrong answer.

The first shortcut treats every censored participant as event-free and divides the events by everyone enrolled, giving 5 ÷ 10 = 0.50. This understates the risk, because participants 2, 5 and 7 might have screened positive after they left the study, and the shortcut assumes that none did. The second shortcut drops the censored participants and divides the events by the participants who remain, giving 5 ÷ 7 = 0.71. This tends to overstate the risk, because the people removed are those known to have been event-free for part of the period, while every participant who had an event is kept. The data themselves only rule out values below 0.50 and above 0.80, which is the risk if all three censored participants had screened positive immediately after leaving. Section 3 shows that the Kaplan-Meier estimate of the 24-month risk for these ten participants is 0.64. It lies within those limits, and it uses the event-free months of each censored participant without inventing events for them.

Equation 8.1 sets out the two shortcuts and the limits that the data place on the risk. Calculator 8.1 computes the two shortcuts and the upper limit from the number enrolled, the number with an observed event and the number censored before the end of follow-up, and its presets load the ten participants in Table 8.1 and the 200 high-loneliness participants.

Two shortcuts and the limits set by the data
Shortcut 1 = E / N    Shortcut 2 = E / (N − C)    E / N ≤ risk ≤ (E + C) / NEq 8.1
The first shortcut divides the number of events by the number enrolled. The second shortcut divides the events by the number enrolled minus the number censored before the end of follow-up. The true risk lies between the first shortcut and the value obtained by counting every early-censored participant as an event.

Averaging the observed times is also misleading. The mean of the ten follow-up times in Table 8.1 is 10.8 months, but it mixes event times with censoring times, and three of the censoring times are shorter than the event times they would eventually have become. The incidence rate from Lesson 4 does use all of the follow-up correctly: the ten participants contributed 2 + 3 + 5 + 5 + 8 + 10 + 12 + 15 + 24 + 24 = 108 person-months and had 5 events, a rate of 5 ÷ 108 = 0.046 per person-month, or 4.6 per 100 person-months. The rate summarizes the whole period in one number, however, and it assumes that the event rate stayed the same throughout. It cannot show whether positive screens clustered in the first months of follow-up or spread evenly across it. The life table in Section 2 and the Kaplan-Meier estimator in Section 3 describe how the probability of remaining event-free changes over time, and Section 4 returns to the rate as the basis of the hazard ratio.

Before those methods are introduced, the last part of this section returns to the choice of time zero.

Design Errors That Start at Time Zero

Two biases follow directly from the choice of time zero. HSCI 230 Lesson 10, Design-Specific and Temporal Biases, teaches both in detail. They are summarized here because each is prevented by a decision written into the protocol.

The accordion below describes immortal time bias and lead-time bias in turn, and its third panel states the feature that the two biases share.

Immortal time bias▼

Immortal time bias arises when membership of the exposed group depends on something that happens after time zero. Suppose the Campus Connection team classified participants as users of a campus peer-support program if they joined it before any positive screen, and started everyone's clock at the baseline survey. A participant who joined in month 9 must have remained event-free for months 0 to 8 in order to be classified as a user. Those months are immortal with respect to the event, and counting them as exposed person-time makes the program look protective even if it has no effect (Suissa, 2008). The protocol prevents the bias by defining exposure at time zero, or by treating program use as a time-varying exposure so that each participant's months before joining are counted as unexposed time.

Lead-time bias▼

Lead-time bias arises when time zero is the date of diagnosis and one group's condition is detected earlier, for example by screening. Survival measured from diagnosis is then longer in the screened group even if screening does not delay death at all, because the clock started earlier. The protocol prevents the bias by measuring time from a starting point that does not depend on when the condition was detected, such as the date of randomization to screening or no screening, or by comparing mortality in the whole screened and unscreened populations.

The common thread▼

In both biases the clock starts at a moment when eligibility, exposure assignment and the start of follow-up do not coincide. Checking that the three conditions are met at the same moment for every participant, as Hernán et al. (2016) recommend, is the single most useful check on a time-to-event protocol.

This section has defined a time-to-event outcome through time zero, the event and the end of follow-up, described the forms of censoring and the non-informative censoring assumption, shown why ordinary summaries misstate the risk when follow-up is incomplete, and traced immortal time bias and lead-time bias to the choice of time zero. The exercise below applies these decisions to a study of readmission after a first heart attack, and the reflection, key takeaways and knowledge check that follow review the section. Section 2 then builds the actuarial life table, which uses the follow-up of every participant to describe how the probability of remaining event-free changes over time.

Try it: Define a time-to-event outcome

A health authority plans to follow adults discharged from hospital after a first heart attack to study time to readmission for any cardiac cause. Enrolment runs for one year and follow-up ends on a fixed date one year after the last enrolment. Readmissions are identified from provincial hospital records. Write down the time zero, the event, the sources of right censoring you would expect, and one source of censoring that could be informative.

One possible answer. Time zero is the date of discharge from the first heart attack admission, when eligibility and follow-up begin together. The event is the first readmission with a cardiac diagnosis, dated to the admission date. Right censoring arises from the administrative end of follow-up (with longer potential follow-up for early enrolees), from moving out of the province, and from death from a non-cardiac cause, which is better treated as a competing event. Moving out of the province to live with family because of declining health could be informative, since those people may have a higher risk of readmission.

Reflection

A research team plans a cohort study of time to first relapse among 300 adults who have completed a residential treatment program for alcohol use disorder. Enrolment will run for 12 months, and every participant will be followed until a fixed study end date 18 months after the last enrolment, so potential follow-up ranges from 18 to 30 months. Relapse will be identified at follow-up interviews held every three months, and participants who miss two consecutive interviews will be treated as lost to follow-up. Some participants are expected to move to other cities for work or family reasons. For this exercise, time zero is the moment the clock starts for a participant; a participant is right-censored when follow-up ends before the event is observed; and censoring is informative when censored participants have a different future risk of the event from participants who remain under observation. (a) Propose a time zero and justify it. (b) List the sources of right censoring you expect. (c) Identify the source most likely to be informative and the direction in which it would bias the estimated probability of remaining relapse-free. (d) Propose two design steps that would reduce or help detect this problem. (e) Explain what kind of censoring the three-monthly interviews introduce.

Model answer(a) Time zero should be the date of discharge from the residential program, because at that moment each participant becomes eligible, has completed the same exposure (the program) and starts to be at risk of relapse in the community. A later date, such as the first follow-up interview, would exclude relapses in the first months and could introduce immortal time. (b) Right censoring will come from the administrative end of the study (with 18 to 30 months of potential follow-up depending on enrolment date), from loss to follow-up after two missed interviews, from withdrawal of consent, and from moving to another city if interviews cannot continue there. Deaths should be recorded as competing events. (c) Loss to follow-up is the most likely to be informative, because people who have started drinking again may stop attending interviews before a relapse is recorded. Censoring them removes high-risk people from the risk set, so the estimated probability of remaining relapse-free will be too high. (d) The protocol could use retention procedures such as several contact methods and brief telephone interviews, and it could compare the baseline characteristics of participants lost to follow-up with those retained and run a worst-case sensitivity analysis that treats each loss as a relapse at the time of censoring. (e) Relapse is known only to have occurred between two interviews, so every event is interval-censored within a three-month window. Dating events to the interview at which they are detected is a reasonable approximation when the window is short relative to follow-up.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • A time-to-event outcome is a follow-up time and an event status for each participant, defined by time zero, the event and the end of follow-up.
  • Time zero should be the moment at which eligibility, exposure assignment and the start of follow-up coincide; misalignment produces immortal time and lead-time bias.
  • Right censoring comes from loss to follow-up, withdrawal, moving away and the administrative end of the study, and it differs from left censoring, interval censoring and late entry.
  • Survival methods assume non-informative censoring, which design can make more plausible through retention procedures, recorded reasons, comparisons and sensitivity analyses.
  • Treating censored participants as event-free understates the risk, and dropping them tends to overstate it.
Knowledge Check: this section

1. A participant in a cohort study was followed for 14 months without the event and then moved to another province, where follow-up could not continue. How should this participant's follow-up be recorded?

The participant was observed event-free for 14 months and then left, so the time is 14 months and the status is 0 (right-censored). Status 1 would record an event that was never observed, removal would discard 14 event-free months, and assigning the planned follow-up time would invent months of observation.

2. According to the target trial principle described by Hernán and colleagues, time zero should be the moment at which:

Time zero should be the moment at which eligibility, exposure assignment and the start of follow-up coincide for each participant. When they occur at different times, biases such as immortal time bias can arise. The first enrolment date is a calendar date for the study, and measuring exposure without confirming eligibility does not align the three conditions.

3. Which situation is most likely to produce informative censoring?

Censoring is informative when the people who are censored have a different future risk from those who remain. Participants whose symptoms are worsening are likely to be at higher risk of the event, so their departure leaves a healthier risk set. Administrative closure, staggered entry and random data loss are unrelated to prognosis and are usually non-informative.

4. Of ten participants, five had the event, three were censored before the end of follow-up and two completed follow-up event-free. An analyst drops the three censored participants and reports a risk of 5/7 = 0.71. Which statement is correct?

Dropping censored participants removes only people known to have been event-free for part of the period, while every participant with an event is kept, so the estimate tends to be too high. Treating censored participants as event-free (5/10 = 0.50) would understate the risk. The Kaplan-Meier estimate for these data is 0.64, which lies between the two.

5. A study classifies participants as users of a peer-support program if they joined it before any event, and starts every participant's follow-up at enrolment. Which bias does this design create?

To be classified as a user, a participant must remain event-free until joining, so the time between enrolment and joining is immortal and is wrongly credited to the program. The program appears protective even if it has no effect. Defining exposure at time zero, or treating program use as time-varying, prevents the bias. Lead-time bias concerns earlier diagnosis, which is a different problem.
← Lesson 7

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 2

The Actuarial Life Table

⏱ Estimated reading time: 30 minutes

Section 2 of 4

The Actuarial Life Table

Cumulative survival as a product of conditional probabilities over fixed intervals.

The logic

Surviving one interval at a time

Cumulative survival
\[ \color{#0B7B6B}{S_j} = \color{#1D4ED8}{p_1} \times \color{#1D4ED8}{p_2} \times \cdots \times \color{#1D4ED8}{p_j} = S_{j-1} \times \color{#1D4ED8}{p_j} \]
Sj survival to the end of interval jpj conditional probability of surviving interval j
One row

The columns of the table

One interval
\[ \color{#047857}{r_j} = \color{#1D4ED8}{l_j} - \frac{\color{#6D28D9}{w_j}}{2} \qquad \color{#BE185D}{q_j} = \frac{\color{#C2410C}{d_j}}{\color{#047857}{r_j}} \qquad p_j = 1 - \color{#BE185D}{q_j} \]
l number enteringw withdrawalsd eventsr effective number at riskq probability of the event
Worked example

High-loneliness group, six-month intervals

MonthslwdrS
0 to 620012141940.928
6 to 1217410121690.862
12 to 181521491450.808
18 to 24129871250.763
The half-interval rule

Three ways to count withdrawals

0.070
14 / 200: withdrawals counted as fully at risk
0.072
14 / 194: the actuarial half-interval rule
0.074
14 / 188: withdrawals removed entirely
Assumptions

What the actuarial method assumes

  • Withdrawals have the same risk as those who remain.
  • Withdrawals occur, on average, at the midpoint of the interval.
  • Survival at a given time since time zero does not depend on the date of enrolment.
  • Interval widths are fixed in the protocol before the data are seen.
Carry forward

From fixed intervals to event times

  • The life table reports survival only at interval boundaries.
  • Recomputing the conditional probability at every event time removes the need to choose intervals.
  • Censored participants are then counted fully at risk until the moment they leave.

Introduction and Overview

The life table is the oldest method for describing survival. John Graunt's analysis of the London Bills of Mortality in 1662 estimated how many of 100 people born would still be alive at successive ages, and the same arithmetic later became the basis for pricing life insurance and annuities. Two kinds of life table are in use. A period life table, of the kind used to compute life expectancy in Lesson 4, applies the age-specific death rates of one calendar period to a hypothetical population. A cohort or clinical life table, the subject of this section, follows a real group of people from time zero and estimates the proportion who remain event-free at the end of successive intervals of follow-up. Cutler and Ederer (1958) set out the method for clinical follow-up studies, and it is often called the actuarial method.

The actuarial life table answers the question that defeated the shortcuts in Section 1. It uses the follow-up of censored participants up to the point at which they leave, and it never assumes anything about what happened to them afterwards. It does so by breaking follow-up into intervals and estimating, for each interval, the probability of getting through it without the event among those who started it event-free. Multiplying these conditional probabilities together gives the probability of remaining event-free to the end of any interval. The same logic, with intervals that end at each event time, becomes the Kaplan-Meier estimator in Section 3.

Learning Objectives for this section

  • Explain why cumulative survival is the product of conditional probabilities of surviving each interval.
  • Define the columns of an actuarial life table: the number entering, the withdrawals, the events, the effective number at risk, the conditional probabilities of the event and of survival, and the cumulative survival.
  • Compute an actuarial life table by hand from counts of events and withdrawals in each interval.
  • Explain the half-interval assumption for withdrawals and show how the estimate changes if withdrawals are counted as fully at risk or not at risk.
  • State the assumptions of the actuarial method and the trade-off involved in choosing the width of the intervals.

The section first explains why survival over several intervals can be written as a product of conditional probabilities. It then defines the columns of the table, works through a small example and the running case, sets out the assumptions behind the method, and closes with the step from grouped intervals to exact event times.

The Logic: Surviving One Interval at a Time

Suppose follow-up is divided into intervals of six months. To remain event-free for 18 months, a participant must get through the first interval without the event, then get through the second interval given that they reached it event-free, then get through the third interval given that they reached it event-free. Probability theory states that the probability of a sequence of events is the product of each event's probability conditional on the events before it. The probability of remaining event-free to the end of interval j is therefore the product of the conditional probabilities of surviving intervals 1 to j.

The value of this decomposition lies in the denominators. Each conditional probability is estimated only from the participants who were present and event-free at the start of that interval. A participant who was censored in month 8 contributes fully to the first interval, contributes partly to the second, and is simply absent from the third and fourth. Nobody has to guess what happened to that participant after month 8, and their eight event-free months are not wasted. Figure 8.3 shows the chain of conditional probabilities for the Campus Connection Follow-up life table computed later in this section.

0 to 6 months 200 enter p = 0.928 S = 0.928 6 to 12 months 174 enter p = 0.929 S = 0.862 12 to 18 months 152 enter p = 0.938 S = 0.808 18 to 24 months 129 enter p = 0.944 S = 0.763 Each link uses only those event-free at its start; S is the running product of the p values S at 24 months = 0.928 × 0.929 × 0.938 × 0.944 = 0.763
Figure 8.3. The actuarial life table as a chain. Each box gives the number entering a six-month interval, the conditional probability of surviving it (p) and the cumulative survival to its end (S), for the 200 high-loneliness participants in the hypothetical Campus Connection Follow-up.

Each link of the chain in Figure 8.3 is computed from counts recorded for its interval. The next part sets out those counts, and the quantities derived from them, as the columns of an actuarial life table.

The Columns of an Actuarial Life Table

An actuarial life table has one row per interval and seven columns. The subscript j numbers the intervals, so l1 is the number entering the first interval and S3 is the cumulative survival to the end of the third interval. Table 8.2 defines each column and shows how it is obtained.

Table 8.2. Columns of an actuarial life table, with the symbol for each quantity and how it is obtained.

SymbolQuantityHow it is obtained
ljNumber entering the intervalParticipants event-free and under observation at the start of interval j; for later intervals, lj = lj−1 − wj−1 − dj−1
wjWithdrawalsParticipants censored during the interval, for any reason
djEventsParticipants who had the event during the interval
rjEffective number at riskrj = lj − wj / 2
qjConditional probability of the eventqj = dj / rj
pjConditional probability of surviving the intervalpj = 1 − qj
SjCumulative survival to the end of the intervalSj = p1 × p2 × … × pj = Sj−1 × pj

The effective number at risk is the only column that needs explanation. Withdrawals leave at various points during the interval, so they were at risk for part of it. The actuarial method assumes that, on average, they were at risk for half of it, and it counts each withdrawal as half a person in the denominator. Equation 8.2 gives the formulas for one interval, and Calculator 8.2 applies them. The calculator starts from the first interval of the Campus Connection Follow-up, in which 200 participants entered, 12 withdrew and 14 screened positive.

One interval of an actuarial life table
r = l − w / 2     q = d / r     p = 1 − qEq 8.2
The effective number at risk is the number entering the interval minus half the withdrawals. The conditional probability of the event is the number of events divided by the effective number at risk, and the conditional probability of surviving the interval is one minus that probability.

Equation 8.2 deals with one interval at a time. A full table chains the intervals together, because the number entering each interval depends on the withdrawals and events in the interval before it, and the next part builds such a table for a small cohort followed in yearly intervals.

Worked Example: Three Yearly Intervals

The first worked example is deliberately small. One hundred people are followed in yearly intervals. In the first year, 10 withdraw and 8 have the event; in the second year, 6 withdraw and 10 have the event; and in the third year, 12 withdraw and 6 have the event. Each row's number entering is the previous row's number entering minus its withdrawals and events, so 100 − 10 − 8 = 82 people enter the second year and 82 − 6 − 10 = 66 enter the third.

Worked Example 8.1 applies Equation 8.2 to each of the three years and multiplies the conditional probabilities of surviving them to obtain the cumulative survival.

Worked Example 8.1: A three-interval actuarial life table

IntervalljwjdjrjqjpjSj
0 to 1 year100108950.0840.9160.916
1 to 2 years82610790.1270.8730.916 × 0.873 = 0.800
2 to 3 years66126600.1000.9000.800 × 0.900 = 0.720

The estimated probability of remaining event-free through three years is 0.72, so the estimated three-year risk of the event is 1 − 0.72 = 0.28. Each value in the last column is computed from unrounded values; multiplying the rounded values shown gives the same result to three decimal places.

Worked Example 8.1 shows every step of the table on round numbers. The next part applies the same steps to the 200 high-loneliness participants of the running case.

The Running Case: 200 Participants Over 24 Months

The 200 high-loneliness participants in the Campus Connection Follow-up were followed in four six-month intervals. The counts in Worked Example 8.2 record, for each interval, how many participants screened positive and how many were censored before 24 months for any of the reasons listed in Section 1. The 114 participants who were still event-free and under observation at 24 months were censored administratively at the end of the study, and they do not appear as withdrawals because they completed the final interval.

Worked Example 8.2: Actuarial life table, high-loneliness group (hypothetical)

MonthsljwjdjrjqjpjSj
0 to 620012141940.0720.9280.928
6 to 1217410121690.0710.9290.862
12 to 181521491450.0620.9380.808
18 to 24129871250.0560.9440.763

The estimated probability of remaining free of a positive screen for 24 months is 0.763, so the estimated 24-month risk is 0.237. Across the four intervals, 42 participants screened positive and 44 were censored before 24 months. The first shortcut from Section 1 gives 42 ÷ 200 = 0.21 and the second gives 42 ÷ 156 = 0.27; the life-table estimate lies between them.

Equation 8.3 states the rule behind the last column of Worked Examples 8.1 and 8.2: the cumulative survival to the end of an interval is the product of the conditional probabilities of surviving every interval up to and including that one.

Calculator 8.3 computes a full life table with up to five intervals. It opens with the running case; the second preset loads the three-interval example. Changing the number of events or withdrawals in one interval changes the number entering every later interval, which shows how each row depends on the rows before it. Figure 8.4, below the calculator, plots the four estimates for the running case.

Cumulative survival in an actuarial life table
Sj = p1 × p2 × … × pj   with   pj = 1 − dj / (lj − wj / 2)Eq 8.3
The cumulative survival to the end of interval j is the product of the conditional probabilities of surviving each interval up to and including that one. Each conditional probability is one minus the events divided by the number entering the interval less half the withdrawals.
0.50 0.60 0.70 0.80 0.90 1.00 0 6 12 18 24 Months since time zero Probability event-free 0.928 0.862 0.808 0.763
Figure 8.4. Actuarial survival estimates at the interval boundaries for the 200 high-loneliness participants. The vertical axis starts at 0.50. Actuarial estimates are usually drawn as points joined by lines, because the table says nothing about survival within an interval.

The running case gives an estimated 24-month risk of 0.237 in the high-loneliness group. That estimate depends on how the 44 participants censored before 24 months were counted in the denominators, and the next part examines the rule that determined their contribution.

Why Withdrawals Are Counted as Half a Person

The half-interval assumption is a compromise between two extremes, and comparing the three versions for the first interval of the running case shows what is at stake. If the 12 withdrawals were counted as at risk for the whole interval, the denominator would be 200 and the conditional probability of the event would be 14 ÷ 200 = 0.070. This treats the withdrawals as if they had been observed event-free to month 6, which they were not, and it understates the probability of the event. If the withdrawals were removed entirely, the denominator would be 188 and the probability would be 14 ÷ 188 = 0.074. This discards the months that withdrawals spent at risk before they left, and it overstates the probability. The actuarial value, 14 ÷ 194 = 0.072, lies between them. The difference between the three versions grows with the number of withdrawals relative to the number entering and with the width of the interval.

The half-interval rule is one of three assumptions on which the actuarial method rests. The accordion below sets out each assumption and then the related choice of interval width.

Assumption 1: censoring is non-informative▼

The withdrawals in each interval are assumed to have the same risk of the event as the participants who remained. If participants whose mood was worsening were more likely to stop responding, the denominators would contain too many people who would have had the event, the conditional probabilities of the event would be too low, and cumulative survival would be overstated. This is the same assumption that underlies every method in this lesson.

Assumption 2: withdrawals occur, on average, at the midpoint▼

Counting each withdrawal as half a person assumes that withdrawals are spread evenly through the interval. If most withdrawals happened in the first week of each interval, the half-interval rule would overstate their time at risk. Narrower intervals make this assumption matter less.

Assumption 3: survival does not depend on when a participant enrolled▼

Participants who enrolled early contribute to later intervals, while participants who enrolled late are censored administratively before reaching them. Combining them assumes that the probability of the event at a given time since enrolment did not change over the calendar period of enrolment. A change in campus mental health services during recruitment, for example, could violate this assumption.

The choice of interval width▼

Wide intervals give stable conditional probabilities but hide changes in risk within each interval and make the midpoint assumption less accurate. Narrow intervals show the timing of events in more detail but contain few events each, so their conditional probabilities are imprecise. Intervals should be fixed in the protocol before the data are seen, usually at clinically or practically meaningful points such as six-month assessment waves. The Kaplan-Meier estimator in Section 3 avoids the choice entirely by letting each interval end at an event time.

Box 8.4 describes the kinds of data for which the actuarial method remains in use despite these assumptions.

Box 8.4: When the actuarial life table is still the right tool

The actuarial method remains in use when event and censoring times are known only by interval. Cancer registries that record survival by year since diagnosis, surveillance systems that report counts by quarter, and cohorts with fixed assessment waves all produce data in this form. It is also convenient for very large cohorts, because a single table summarizes the data. When exact or near-exact times are available for each participant, the Kaplan-Meier estimator uses them more fully.

The half-interval assumption and the choice of interval width both follow from grouping follow-up into fixed intervals. The final part of this section considers what changes when each interval is allowed to end at an observed event time.

From Grouped Intervals to Exact Event Times

The life table shows the probability of remaining event-free only at the interval boundaries: 0.928 at 6 months, 0.862 at 12 months, 0.808 at 18 months and 0.763 at 24 months. Between the boundaries it says nothing, and an analyst who chose three-month intervals would obtain somewhat different values. If the conditional probabilities are recomputed every time an event occurs, two things change. The intervals no longer need to be chosen, and withdrawals no longer need the half-interval assumption, because each censored participant is counted as fully at risk at every event time before their censoring time and as absent after it. That is the Kaplan-Meier estimator.

This section has built the actuarial life table from conditional probabilities, applied it to a small cohort and to the running case, and set out the assumptions on which it rests. The exercise below works through one row of a life table by hand, and the reflection, key takeaways and knowledge check that follow review the section. Section 3 then develops the Kaplan-Meier estimator in full.

Try it: Complete one row by hand

In a four-interval life table, 150 people enter the third interval, 10 withdraw during it and 14 have the event. Cumulative survival at the end of the second interval is 0.850. Compute r3, q3, p3, S3 and the number entering the fourth interval, then check the answers with the single-interval calculator, Calculator 8.2.

Answer. r3 = 150 − 10 / 2 = 145; q3 = 14 / 145 = 0.097; p3 = 0.903; S3 = 0.850 × 0.903 = 0.768; and 150 − 10 − 14 = 126 people enter the fourth interval.

Reflection

A community clinic follows 150 adults who have completed a smoking cessation program, in yearly intervals, to estimate the probability of remaining free of relapse to daily smoking. In year 1, 10 participants withdraw and 12 relapse; in year 2, 8 withdraw and 9 relapse; in year 3, 14 withdraw and 6 relapse. The actuarial method uses the number entering each interval (l), the effective number at risk r = l − w/2 (where w is the withdrawals), the conditional probability of relapse q = d/r (where d is the relapses), the conditional probability of remaining relapse-free p = 1 − q, and cumulative survival S, the product of the p values up to that interval. The number entering each later interval is the previous number entering minus that interval's withdrawals and relapses. (a) Compute the life table and the estimated three-year probability of remaining relapse-free. (b) Recompute the three-year estimate with withdrawals counted as fully at risk (r = l) and explain why it differs. (c) State one assumption of the actuarial method that could fail in this clinic and the direction of bias if it did.

Model answer(a) Year 1: l = 150, r = 150 − 5 = 145, q = 12/145 = 0.083, p = 0.917, S = 0.917. Year 2: l = 150 − 10 − 12 = 128, r = 124, q = 9/124 = 0.073, p = 0.927, S = 0.917 × 0.927 = 0.851. Year 3: l = 128 − 8 − 9 = 111, r = 104, q = 6/104 = 0.058, p = 0.942, S = 0.851 × 0.942 = 0.802. The estimated probability of remaining relapse-free for three years is about 0.80, so the three-year risk of relapse is about 0.20. (b) Counting withdrawals as fully at risk gives q values of 12/150 = 0.080, 9/128 = 0.070 and 6/111 = 0.054, and three-year survival of 0.809 (a risk of 0.191). The estimate is higher because the withdrawals are treated as if they had been observed relapse-free to the end of the year in which they left, which overstates the time at risk in the denominators. (c) The method assumes non-informative censoring. If people who started smoking again were more likely to stop attending the clinic, the withdrawals would include many unrecorded relapses, the relapse probabilities would be too low, and the three-year survival would be overstated.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • Cumulative survival is the product of the conditional probabilities of surviving each interval among those who started it event-free.
  • Each row of an actuarial life table uses the number entering, the withdrawals and the events, with the effective number at risk equal to the number entering minus half the withdrawals.
  • The half-interval rule lies between counting withdrawals as fully at risk, which understates the probability of the event, and removing them, which overstates it.
  • The actuarial method assumes non-informative censoring, withdrawals at the midpoint on average and stable survival across enrolment dates, and its intervals should be fixed in the protocol.
Knowledge Check: this section

1. In an actuarial life table, 120 participants enter an interval, 16 withdraw during it and 9 have the event. What is the conditional probability of the event in the interval?

The effective number at risk is 120 − 16/2 = 112, so the conditional probability is 9/112 = 0.080. Using 120 treats the withdrawals as fully at risk, using 104 removes them entirely, and using 95 subtracts both the withdrawals and the events, which confuses the denominator with the number entering the next interval.

2. Why is the cumulative survival in a life table computed as a product of conditional probabilities?

To be event-free at the end of interval j, a participant must survive interval 1, then survive interval 2 given survival of interval 1, and so on, and the probability of such a sequence is the product of the conditional probabilities. Multiplication does not correct for informative censoring, and the risk set changes from interval to interval.

3. The cumulative survival at the end of interval 2 is 0.86, and the conditional probability of surviving interval 3 is 0.94. What is the cumulative survival at the end of interval 3?

Cumulative survival is the previous cumulative survival multiplied by the conditional probability for the new interval: 0.86 × 0.94 = 0.81. The value 0.80 comes from subtracting the conditional probability of the event (0.06) from 0.86, which is incorrect, and 0.94 is the conditional probability for interval 3 alone.

4. What does the actuarial method assume when it counts each withdrawal as half a person in the denominator?

Withdrawals leave at various points during the interval, and the method assumes that, on average, they were at risk for half of it. The assumption concerns time at risk and says nothing about whether withdrawals would have had the event. If withdrawals all occurred at the start, they should not be counted at all.

5. Which change to a study's life-table plan makes the estimates depend less on the half-interval assumption?

With narrower intervals, the uncertainty about when within an interval a withdrawal occurred covers less time, so the half-interval assumption matters less. Wider intervals have the opposite effect. Removing withdrawals or counting them as fully at risk replaces the assumption with a more extreme one that biases the conditional probabilities.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 3

The Kaplan-Meier Estimator and Reading Survival Curves

⏱ Estimated reading time: 35 minutes

Section 3 of 4

The Kaplan-Meier Estimator and Reading Survival Curves

The product-limit estimate, step by step, and how to read the curve it produces.

The product-limit idea

Kaplan-Meier estimator

Estimated survival at time t
\[ \color{#0B7B6B}{\hat{S}(t)} = \prod_{t_j \le t} \frac{\color{#1D4ED8}{n_j} - \color{#C2410C}{d_j}}{\color{#1D4ED8}{n_j}} \]
Ŝ(t) estimated survivalnj at risk just before tjdj events at tj
Worked example

Ten participants, step by step

MonthAt riskEventsCensoredŜ(t)
210100.900
39010.900
58200.675
86010.675
105100.540
124010.540
153100.360
Conventions

Ties, censoring and the end of the curve

  • Tied events at the same time are handled together in one step.
  • When an event and a censoring share a time, the censored participant counts at risk at that time.
  • The estimate is defined only up to the longest follow-up time.
Reading the curve

Four things to read

Survival at a time

The curve is at 0.540 at 12 months, so the 12-month risk is 0.460.

Median survival

The curve first falls to 0.50 or below at month 15.

Censoring marks

Vertical ticks show censoring times and never move the curve.

Numbers at risk

Only 2 participants remain at risk from month 15, so the tail is imprecise.

Two groups

High versus low loneliness, ten participants each

Survival at 12 months

The estimate is 0.54 with high loneliness and 0.76 with low loneliness.

Median survival

The median is 15 months with high loneliness and is not reached with low loneliness.

Restricted mean to 24 months

Participants were event-free for 14.0 and 19.9 months on average.

Carry forward

From describing to comparing

  • The Kaplan-Meier curve drops only at event times and is read with its numbers at risk.
  • The median is the first time the curve reaches 0.50 or below.
  • Comparing two curves formally requires the log-rank test and the hazard ratio.

Introduction and Overview

Edward Kaplan and Paul Meier published the product-limit estimator in 1958 (Kaplan & Meier, 1958), and it has become the standard way to describe survival in clinical and epidemiological research. The estimator applies the logic of the life table with one change. Instead of fixed intervals chosen in advance, each interval ends at an observed event time. At every event time the estimator computes the conditional probability of getting past that time among the participants still at risk, and it multiplies these conditional probabilities together. Censored participants are counted as fully at risk at every event time before they leave and are absent after it, so the half-interval assumption of the actuarial method is no longer needed.

This section computes a Kaplan-Meier estimate by hand for the ten high-loneliness participants introduced in Section 1, using a calculator that shows each step. It then explains how to read a published Kaplan-Meier curve: the survival probability at a chosen time, the median survival time, the censoring marks and the table of numbers at risk that should appear beneath every curve.

Learning Objectives for this section

  • Explain how the Kaplan-Meier estimator extends the actuarial life table by ending each interval at an observed event time.
  • Compute a Kaplan-Meier estimate by hand, including tied event times and censoring between event times.
  • Read the estimated survival probability and cumulative risk at a chosen time from a Kaplan-Meier curve.
  • Find the median survival time on a Kaplan-Meier curve and explain what "not reached" means.
  • Interpret censoring marks and the numbers-at-risk table, and explain why the right-hand end of a curve is the least precise part.
  • State the assumptions of the Kaplan-Meier estimator and list what a report of a Kaplan-Meier analysis should include.

The section begins with the product-limit idea in symbols.

The Product-Limit Idea

Order the distinct event times as t1 < t2 < t3 and so on. At each event time tj, let nj be the number of participants at risk just before that time, meaning those who have neither had the event nor been censored before tj, and let dj be the number who have the event at tj. The conditional probability of getting past tj is (nj − dj) / nj. The Kaplan-Meier estimate of survival at time t, written Ŝ(t), is the product of these conditional probabilities over every event time up to and including t. The hat on S marks a quantity estimated from data.

Each step needs only the estimate at the previous event time, the number at risk and the number of events. Equation 8.4 writes one such step in symbols. Calculator 8.4 performs a single step, starting from the third event time in Worked Example 8.3: the estimate just before month 5 is 0.900, eight participants are at risk, and two screen positive.

One step of the Kaplan-Meier estimator
Ŝ(tj) = Ŝ(tj−1) × (nj − dj) / njEq 8.4
The estimated survival at an event time is the estimated survival at the previous event time multiplied by the number at risk just before this time minus the number of events at this time, divided by the number at risk.

Repeating this step at each event time in turn produces the whole estimate. The next part does so for the ten high-loneliness participants whose records appear in Table 8.1.

Worked Example: Ten Participants, Step by Step

Table 8.3 applies the estimator to the ten high-loneliness participants from Table 8.1. The rows are the distinct follow-up times in order. A row with an event changes the estimate; a row with only censoring leaves the estimate unchanged and reduces the number at risk at the next event time. Worked Example 8.3 sets out this calculation and interprets the resulting 24-month estimate.

Worked Example 8.3: Kaplan-Meier estimate for ten high-loneliness participants

Table 8.3. Kaplan-Meier estimate for the ten high-loneliness participants in Table 8.1, with one row for each distinct follow-up time.

MonthAt risk just before (n)Events (d)Censored(n − d) / nŜ(t)
210109/10 = 0.9000.900
3901no event0.900
58206/8 = 0.7500.900 × 0.750 = 0.675
8601no event0.675
105104/5 = 0.8000.675 × 0.800 = 0.540
12401no event0.540
153102/3 = 0.6670.540 × 0.667 = 0.360
24202no event0.360

The estimated probability of remaining free of a positive screen for 24 months is 0.360, so the estimated 24-month risk is 1 − 0.360 = 0.640, the value quoted in Section 1. Participant 2, censored at month 3, counts in the risk set at month 2 and is absent from month 5 onward. The two events at month 5 are tied and are handled in a single step with d = 2.

Three conventions are worth stating explicitly, because published analyses and software follow them. Tied events at the same time are handled together in one step. When an event and a censoring are recorded at the same time, the event is assumed to come first, so the censored participant is still counted at risk at that time. The estimate is defined only up to the longest follow-up time; beyond month 24 the data say nothing about these participants.

Equation 8.5 gathers the steps of Table 8.3 into a single expression: the estimate at time t is the product of the conditional probabilities at every event time up to and including t.

The step-by-step Calculator 8.5 takes up to ten participants, each with a follow-up time and a status (1 for an event, 0 for censored). It sorts the times, forms the risk set at each time, and shows the working for all time points at once or for one time point at a time, together with the curve drawn so far. Try changing participant 5 from censored to an event at month 8 and watch every later step change. Then change participant 2's time from 3 to 20 months while keeping the status at 0 and notice that a censored participant who stays longer raises the number at risk at every later event time.

Kaplan-Meier estimator
Ŝ(t) = ∏tj ≤ t (nj − dj) / njEq 8.5
The estimated survival at time t is the product, over every event time up to and including t, of the number at risk just before that time minus the number of events at that time, divided by the number at risk. The symbol ∏ means "multiply together".

Worked Example 8.3 gives the estimate as a sequence of numbers. Published analyses present it as a curve, and the next part explains how to read one.

Reading a Kaplan-Meier Curve

A Kaplan-Meier curve plots the estimate against time as a step function. It starts at 1 at time zero, stays flat between event times, and drops at each event time by an amount that depends on the number of events and the number at risk. Figure 8.5 shows the curve for the ten participants in Table 8.3.

0.00 0.25 0.50 0.75 1.00 0 6 12 18 24 Months since time zero Probability event-free median = 15 months 0.900 0.675 0.540 0.360 Number at risk 10 6 4 2 2
Figure 8.5. Kaplan-Meier curve for the ten high-loneliness participants in Table 8.3. The curve drops only at event times; vertical ticks mark censoring times. The dashed red line reads the median from 0.50 on the vertical axis. The numbers at risk beneath the plot count participants still under observation and event-free at each time.

Four features of Figure 8.5 carry most of the information. The four tabs below explain how to read the survival at a chosen time, the median survival time, the censoring marks and the numbers at risk.

Reading up from a time on the horizontal axis to the curve and across to the vertical axis gives the estimated probability of remaining event-free at that time. At 12 months the curve is at 0.540, so the estimated 12-month risk of a positive screen is 1 − 0.540 = 0.460. Because the curve is right-continuous, the value exactly at an event time is the value after the drop: Ŝ(10) = 0.540 describes the probability of remaining event-free beyond month 10. Some reports plot the cumulative risk, 1 − Ŝ(t), which rises from 0 as the survival curve falls from 1; the two plots carry the same information.

The median survival time is the earliest time at which the estimated survival falls to 0.50 or below, so that half of the group is estimated to have had the event. Reading across from 0.50 on the vertical axis to the curve and down to the horizontal axis gives a median of 15 months in Figure 8.5, because the curve first falls to or below 0.50 at month 15. When the curve never falls to 0.50 during follow-up, the median is reported as not reached. The median is preferred to the mean as a summary of survival time because the mean cannot be computed without knowing the event times of participants who were censored at the end of follow-up.

The small vertical ticks on the curve mark the times at which participants were censored. They do not change the estimate. A cluster of ticks early in follow-up indicates heavy early loss, which should prompt a question about whether censoring was informative. Many ticks at a single late time usually reflect administrative censoring at the end of the study.

The table beneath the plot gives the number of participants still at risk at regular times. It shows how much information supports each part of the curve. In Figure 8.5 only two participants remain at risk from month 15 onward, so the last part of the curve rests on very few people, and a single additional event would have dropped it from 0.360 to 0.180. Readers should treat the right-hand end of any curve with caution when the numbers at risk are small, and a curve published without a numbers-at-risk table is harder to interpret.

Uncertainty in the estimate is usually shown as a 95 percent confidence band around the curve. The standard error is computed with Greenwood's formula (Greenwood, 1926), which adds a contribution from each event time. The practical point for study design is that the band widens as the number at risk falls, so precision late in follow-up depends on retaining participants and on enrolling enough of them in the first place.

These reading rules apply to any Kaplan-Meier curve. The next part uses them to set the high-loneliness curve beside a curve for ten low-loneliness participants.

Comparing Two Curves by Eye

Figure 8.6 adds ten participants from the low-loneliness group to the ten from the high-loneliness group. In the low-loneliness group, participant 1 was censored at month 4, participants 2, 4 and 6 screened positive at months 6, 12 and 20, participants 3 and 5 were censored at months 9 and 18, and the remaining four completed 24 months without a positive screen.

0.00 0.25 0.50 0.75 1.00 0 6 12 18 24 Months since time zero Probability event-free Low loneliness (B) High loneliness (A) 0.610 0.360 Number at risk B 10 9 7 6 4 A 10 6 4 2 2
Figure 8.6. Kaplan-Meier curves for ten high-loneliness (A) and ten low-loneliness (B) participants in the hypothetical Campus Connection Follow-up, with censoring marks and numbers at risk. The median is 15 months in group A and is not reached in group B.

The curves can be compared at a fixed time, by their medians, or by the area beneath them. At 12 months the estimated probability of remaining event-free is 0.540 in the high-loneliness group and 0.762 in the low-loneliness group. The median is 15 months in the high-loneliness group and is not reached in the low-loneliness group, whose curve ends at 0.610. The area under each curve up to 24 months is the restricted mean survival time, the average number of event-free months within the first 24 (Royston & Parmar, 2013). It is 14.0 months for the high-loneliness participants and 19.9 months for the low-loneliness participants, a difference of about 5.9 event-free months. Each of these comparisons describes the two samples. Whether the difference is larger than chance would produce with ten participants per group is the question for the log-rank test in Section 4.

Before that comparison, the next part returns to the high-loneliness group alone and sets its Kaplan-Meier estimate beside the two shortcuts from Section 1.

Kaplan-Meier Compared with the Shortcuts

For the same ten high-loneliness participants, the three approaches give the following estimates of the probability of remaining event-free for 24 months. Table 8.4 sets out the three estimates and the assumption that each makes about the three participants censored early.

Table 8.4. Estimated probability of remaining event-free for 24 months among the ten high-loneliness participants under the two shortcuts and the Kaplan-Meier estimator, with the assumption each approach makes about the three early-censored participants.

ApproachEstimated 24-month survivalWhat it assumes about the three early-censored participants
Treat censored participants as event-free5/10 event-free = 0.500None of them would ever have screened positive
Drop censored participants2/7 event-free = 0.286Their event-free months carry no information
Kaplan-Meier0.360After censoring, each had the same risk as those who remained

The Kaplan-Meier estimate lies between the two shortcuts, and it is the only one of the three whose assumption is stated in terms that a design can make plausible. Efron (1967) showed that the estimator can be computed by redistributing each censored participant's share of the probability equally among the participants who remain at risk after the censoring time, which is a concrete way to see the non-informative censoring assumption at work.

The Kaplan-Meier estimate rests on assumptions of its own, and the final part of this section states them and lists what a report of a Kaplan-Meier analysis should include.

What the Estimator Assumes and What a Report Should Include

The Kaplan-Meier estimator makes no assumption about the shape of the survival curve, which is why it is called non-parametric. It does make three assumptions about the data. Censoring must be non-informative, as discussed in Section 1. Participants who enrolled early and late must have the same survival prospects at the same time since time zero, as for the life table. Event times must be recorded precisely enough that the order of events and censorings is correct, which is why monthly ascertainment works well in the Campus Connection Follow-up and annual ascertainment would not.

Box 8.5 lists the elements that a report of a Kaplan-Meier analysis should contain.

Box 8.5: Reporting a Kaplan-Meier analysis

A report of a Kaplan-Meier analysis should state the time zero, the event definition and the censoring rules; give the number of participants and the number of events in each group; present the curve with censoring marks and a numbers-at-risk table; report survival at one or two prespecified times with 95 percent confidence intervals; and give the median survival time with its confidence interval when it is reached. Bland and Altman (1998) and Clark et al. (2003) give accessible accounts of the method and its presentation.

This section has computed the Kaplan-Meier estimate by hand, shown how to read and compare its curves, and stated its assumptions and the content of a report. The reflection, key takeaways and knowledge check that follow review the section, and Section 4 then turns from describing survival in each group to formal comparisons between groups.

Reflection

Eight participants in a pilot cohort are followed for up to 12 months. Their follow-up times in months and statuses (1 = event, 0 = censored) are: 1 (1), 3 (0), 4 (1), 4 (1), 6 (0), 7 (1), 9 (0) and 12 (0). The Kaplan-Meier estimator multiplies, at each event time, the conditional probability (n − d)/n, where n is the number at risk just before that time (participants with a follow-up time equal to or later than it) and d is the number of events at that time; censored participants leave the risk set after their censoring time without changing the estimate. The median survival time is the earliest time at which the estimate falls to 0.50 or below. (a) Compute the Kaplan-Meier estimate at each event time. (b) Give the median survival time. (c) Give the number at risk at 6 months. (d) Explain why the estimate after month 7 should be interpreted with caution and what a numbers-at-risk table would show a reader.

Model answer(a) Month 1: 8 at risk and 1 event, so the estimate is 7/8 = 0.875. The participant censored at month 3 leaves the risk set. Month 4: 6 at risk and 2 events, so the estimate is 0.875 × 4/6 = 0.583. The participant censored at month 6 leaves. Month 7: 3 at risk and 1 event, so the estimate is 0.583 × 2/3 = 0.389. No further events occur, so the estimate stays at 0.389 until follow-up ends at month 12. (b) The estimate is 0.583 after month 4 and first falls to 0.50 or below at month 7, so the median survival time is 7 months. (c) Four participants have follow-up times of 6 months or longer (6, 7, 9 and 12), so 4 are at risk at 6 months. (d) After month 7 only two participants remain at risk, so a single further event would halve the estimate. The curve beyond month 7 therefore rests on very little information and its confidence interval would be wide. A numbers-at-risk table beneath the curve would show the reader how few participants support each part of it, which is the main protection against over-reading the tail.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • The Kaplan-Meier estimator multiplies (n − d) / n at each event time, so censored participants count fully until they leave and never cause a drop.
  • Tied events are handled in one step, a participant censored at an event time counts as at risk at that time, and the estimate ends at the longest follow-up.
  • The median survival time is the first time the curve falls to 0.50 or below, and it is reported as not reached if the curve stays above 0.50.
  • Censoring marks and the numbers-at-risk table show where a curve is well supported; the right-hand end, with few participants at risk, is the least precise part.
Knowledge Check: this section

1. At an event time, 12 participants are at risk and 3 have the event. The Kaplan-Meier estimate at the previous event time was 0.80. What is the new estimate?

The conditional probability of getting past this time is (12 − 3)/12 = 0.75, and the new estimate is 0.80 × 0.75 = 0.60. The value 0.55 comes from subtracting 3/12 from 0.80, and 0.75 is the conditional probability alone.

2. How does a participant who is censored at month 8 affect a Kaplan-Meier estimate?

A censored participant is counted in the risk set at every event time up to the censoring time and then leaves it. Censoring does not change the estimate at the time it occurs; it reduces the number at risk at later event times. Dropping the participant from earlier risk sets would waste eight event-free months.

3. A Kaplan-Meier curve falls to 0.52 at month 18 and to 0.47 at month 22, and follow-up ends at month 30 with the curve at 0.41. What is the median survival time?

The median is the earliest time at which the estimated survival falls to 0.50 or below. The curve is above 0.50 at month 18 (0.52) and first falls below it at month 22 (0.47), so the median is 22 months. The median is not interpolated between event times, and it is reached because the curve falls below 0.50 during follow-up.

4. Why should the right-hand end of a Kaplan-Meier curve be interpreted with caution?

As the number at risk falls, each event removes a larger share of the remaining probability and the confidence interval widens, so the tail rests on little information. Censoring marks at the end usually reflect administrative censoring, and the estimator makes no assumption about the shape of the hazard.

5. Which element of a published Kaplan-Meier figure lets readers judge how much information supports each part of the curve?

The numbers-at-risk table shows how many participants remain at risk at regular times, which tells readers where the curve is well supported and where it rests on few people. A p-value compares groups and says nothing about precision at particular times, and the mean follow-up time and overall event proportion do not show how information changes over time.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 4

Comparing Groups: The Log-Rank Test and the Hazard Ratio

⏱ Estimated reading time: 40 minutes

Section 4 of 4

Comparing Groups: The Log-Rank Test and the Hazard Ratio

Observed and expected events, hazards as rates, proportional hazards and the protocol.

The log-rank test

Observed versus expected events

Expected events in group A
\[ \color{#0B7B6B}{E_A} = \sum_j \color{#C2410C}{d_j} \times \frac{\color{#1D4ED8}{n_{Aj}}}{n_j} \qquad O_A = 5,\; E_A = 3.24,\; \chi^2_1 = 1.67,\; p = 0.20 \]

Eight events give the test little power to detect even a large difference.

The hazard

An event rate among those still at risk

MonthsEventsPerson-monthsPer 100 person-months
0 to 6141,1221.25
6 to 12129781.23
12 to 1898431.07
18 to 2477290.96
The hazard ratio

A ratio of rates

Hazard ratio under constant hazards
\[ \color{#0B7B6B}{HR} = \frac{\color{#C2410C}{42/3672}}{\color{#1D4ED8}{40/8800}} = 2.52 \qquad 95\%\ \text{CI } 1.63 \text{ to } 3.88 \]

The Cox model (Cox, 1972) estimates adjusted hazard ratios and assumes only that the ratio of the hazards is constant.

Proportional hazards

The ratio stays the same over time

Plausible

The curves separate steadily and do not cross.

Doubtful

The curves cross, or separate early and then converge.

Survival at fixed times and restricted mean survival time do not depend on the assumption.

Hazard ratio and risk ratio

Translating a hazard ratio into risks

Survival under proportional hazards
\[ \color{#0B7B6B}{S_1(t)} = \color{#1D4ED8}{S_0(t)}^{\color{#C2410C}{HR}} \qquad 0.90^{2.5} = 0.768 \;\Rightarrow\; RR = \frac{0.232}{0.100} = 2.32 \]

With a common event, 0.502.5 = 0.177 and the risk ratio is 1.65.

The protocol

What a time-to-event protocol must specify

  • The protocol states time zero, eligibility, exposure, the event and its ascertainment.
  • It sets out follow-up, censoring rules and the plan for informative censoring.
  • It describes the estimators, the test, the effect measure and the proportional hazards check.
  • It sizes the study in events first: 103 events, or 687 participants if 15% have the event.
Next steps

Reflection, knowledge check and final assessment

  • The section reflection asks you to interpret a hazard ratio and translate it into risks.
  • The knowledge check covers the log-rank test, the hazard ratio and proportional hazards.
  • The final assessment asks you to write the time-to-event subsection of a new protocol.

Introduction and Overview

Sections 2 and 3 described survival in one group at a time. Most time-to-event studies compare groups: participants with and without an exposure, or participants randomized to an intervention or to control. This section introduces the two tools used for that comparison. The log-rank test asks whether two survival curves differ by more than chance would produce. The hazard ratio measures how much faster events occur in one group than in the other, and it is the time-to-event counterpart of the incidence rate ratio from Lesson 6. The section then explains the proportional hazards idea, which most hazard ratio analyses assume, and closes with the decisions that a protocol with a time-to-event outcome must state before data collection begins.

Learning Objectives for this section

  • Explain the logic of the log-rank test as a comparison of observed and expected events at each event time.
  • Define the hazard as the event rate among participants still at risk and estimate it from life-table intervals.
  • Compute a hazard ratio as a ratio of incidence rates with its 95 percent confidence interval, and interpret it in words.
  • State the proportional hazards assumption in plain language, recognize patterns that suggest it does not hold, and name alternatives to a single hazard ratio.
  • Explain why a hazard ratio differs from a risk ratio when the event is common.
  • List what a protocol with a time-to-event outcome must specify, including the number of events needed for adequate power.

The section begins with the log-rank test.

Comparing Two Curves: The Log-Rank Test

Figure 8.6 in Section 3 showed that the high-loneliness participants had a lower Kaplan-Meier curve than the low-loneliness participants. The log-rank test asks whether a difference of that size could plausibly arise if the two groups had the same survival at every time. Its logic builds on the risk set. At each event time, the participants at risk in both groups are pooled, and the events at that time are shared out between the groups in proportion to their numbers at risk. If the null hypothesis is true, a group that makes up 40 percent of the pooled risk set should, on average, account for 40 percent of the events at that time. The expected number of events in group A at event time j is therefore the total number of events at that time multiplied by group A's share of the risk set.

Equation 8.6 adds these expected numbers over the event times and gives the test statistic that compares them with the observed number of events in group A.

Expected events in group A under the null hypothesis
EA = ∑j dj × nAj / nj     χ² = (OA − EA)² / VEq 8.6
The expected number of events in group A adds up, over every event time, the total events at that time multiplied by the number at risk in group A divided by the total number at risk. The test statistic compares the observed events in group A with this expected number, scaled by its variance V, and is referred to a chi-squared distribution with one degree of freedom.

Table 8.5 applies this logic to the twenty participants in Figure 8.6. Group A is the high-loneliness group. At month 12, one low-loneliness participant screened positive and high-loneliness participant 7 was censored; following the convention from Section 3, participant 7 is counted in the risk set at month 12. Worked Example 8.4 sets out this calculation and the resulting log-rank statistic.

Worked Example 8.4: Observed and expected events at each event time

Table 8.5. Observed and expected events in the high-loneliness group (A) at each event time, for the twenty participants in Figure 8.6.

MonthAt risk, AEvents, AAt risk, BEvents, BTotal eventsExpected in A
210110011 × 10/20 = 0.500
5829022 × 8/17 = 0.941
6609111 × 6/15 = 0.400
10517011 × 5/12 = 0.417
12407111 × 4/11 = 0.364
15316011 × 3/9 = 0.333
20205111 × 2/7 = 0.286
TotalOA = 5OB = 38EA = 3.24

The high-loneliness group had 5 events where 3.24 were expected under the null hypothesis, and the low-loneliness group had 3 where 8 − 3.24 = 4.76 were expected. The variance of OA − EA, summed over the event times, is 1.86, so the log-rank statistic is (5 − 3.24)² / 1.86 = 1.67 with one degree of freedom, giving p = 0.20. A simpler approximation that is often taught, (OA − EA)² / EA + (OB − EB)² / EB, gives 1.61 and p = 0.21.

The curves in Figure 8.6 look quite different, yet the test does not reject the null hypothesis. With ten participants per group and eight events in total, a large difference in survival is compatible with chance. This is a property of the design and says nothing in favour of equal survival. The information in a time-to-event comparison comes from the events, and eight events provide very little of it; the last part of this section shows how to plan for enough events. The ratio of observed to expected events also gives a rough estimate of the hazard ratio: (5 ÷ 3.24) ÷ (3 ÷ 4.76) = 2.45.

Calculator 8.6 performs the observed-and-expected arithmetic for up to eight event times. It opens with Table 8.5. Moving one event from group A to group B at the same time shows how sensitive a comparison with few events is.

The test is usually attributed to Mantel (1966) and to Peto and Peto (1972), and Bland and Altman (2004) give a short account for applied readers. Three of its properties matter for interpretation. It gives every event time the same weight, so late differences, where few participants remain at risk, count as much as early ones. It is most powerful when one group has a consistently higher hazard throughout follow-up, the situation described by the proportional hazards assumption below, and it can miss a real difference when curves cross. It extends directly to more than two groups and to comparisons within strata of a confounder, the stratified log-rank test, in the same way that the Mantel-Haenszel approach in Lesson 7 extends a single two-by-two table.

The log-rank test addresses whether two curves differ by more than chance would produce. A measure of how much faster events occur in one group needs the hazard, which the next part defines.

The Hazard: An Event Rate Among Those Still at Risk

The hazard at time t is the rate at which events occur among participants who are still event-free at t, expressed per unit of time. It is the incidence rate from Lesson 4 computed over a very short interval around t. The survival probability can only stay level or fall as time passes, but the hazard can rise, fall or stay constant. The hazard of death after major surgery is highest in the days after the operation and then falls; the hazard of death from most chronic diseases rises with age; and the hazard of a sudden accidental event may be roughly constant.

The life table from Section 2 gives an estimate of the hazard within each interval. Dividing the events in an interval by the person-time at risk in that interval gives the average hazard over the interval. With the half-interval assumption applied to both withdrawals and events, the person-time in an interval of width h is approximately h × (rj − dj / 2).

Table 8.6 applies this approximation to the four six-month intervals of the high-loneliness life table in Worked Example 8.2.

Table 8.6. Estimated hazard of a first positive screen in each six-month interval for the 200 high-loneliness participants in the hypothetical Campus Connection Follow-up.

MonthsEventsPerson-months at riskHazard per person-monthPer 100 person-months
0 to 6146 × (194 − 7) = 1,1220.01251.25
6 to 12126 × (169 − 6) = 9780.01231.23
12 to 1896 × (145 − 4.5) = 8430.01071.07
18 to 2476 × (125 − 3.5) = 7290.00960.96
0 to 24423,6720.01141.14

In this hypothetical cohort the hazard of a first positive screen declined slowly over follow-up, from about 1.25 to about 0.96 per 100 person-months. The overall rate in the last row, 42 events over 3,672 person-months, is the average of the interval hazards weighted by person-time.

The overall hazard of 1.14 per 100 person-months summarizes the high-loneliness group. The next part compares it with the hazard in the low-loneliness group.

The Hazard Ratio as a Ratio of Rates

The hazard ratio compares the hazard in one group with the hazard in another at the same time since time zero. If both hazards are constant over follow-up, the hazard ratio equals the incidence rate ratio from Lesson 6, and it can be computed from the events and person-time in each group. The 400 low-loneliness participants in the Campus Connection Follow-up had 40 first positive screens over 8,800 person-months, a rate of 0.45 per 100 person-months, compared with 1.14 per 100 person-months in the high-loneliness group.

Equation 8.7 gives the hazard ratio under constant hazards with its confidence interval, and Calculator 8.7 computes both from the events and person-time in each group, with a preset for the Campus Connection cohort.

Hazard ratio under constant hazards
HR = (a1 / T1) / (a0 / T0)    95% CI = exp[ln(HR) ± 1.96 × √(1/a1 + 1/a0)]Eq 8.7
The hazard ratio is the event rate in the exposed group (events divided by person-time) divided by the event rate in the unexposed group. Its confidence interval is computed on the log scale, with a standard error that depends only on the number of events in the exposed group and the number of events in the unexposed group.

The hazard ratio is (42 ÷ 3,672) ÷ (40 ÷ 8,800) = 2.52, with a 95 percent confidence interval from 1.63 to 3.88. In words, at any point during the 24 months, participants with high loneliness who had not yet screened positive did so at about two and a half times the rate of participants with low loneliness. The standard error depends only on the numbers of events, 42 and 40, which is another way of seeing that events carry the information in a time-to-event comparison.

In practice, hazard ratios are usually estimated with the Cox proportional hazards model (Cox, 1972). The Cox model does not assume that the hazards are constant. It allows the baseline hazard to take any shape over time and assumes only that the ratio of the hazards stays the same. Like logistic regression for the odds ratio, it can adjust a hazard ratio for confounders, and it is the time-to-event analogue of the regression adjustment described in Lesson 7. HSCI 341 interprets the hazard ratio from a Cox model without fitting one. For protocol writing, the point is that a hazard ratio from a Cox model has the same interpretation as the rate ratio above, provided its central assumption holds.

A single hazard ratio, whether computed from rates or from a Cox model, is a complete summary only if the ratio of the hazards stays the same over follow-up. The next part examines that assumption.

Proportional Hazards in Plain Language

The proportional hazards assumption states that the ratio of the hazards in the two groups is the same at every time during follow-up. Each hazard may rise or fall, but they rise and fall together. If participants with high loneliness screen positive at two and a half times the rate of participants with low loneliness in month 2, the assumption says they also do so at two and a half times the rate in month 20. A single hazard ratio is then a complete summary of how the groups differ over time.

Figure 8.7 shows schematic hazards and survival curves for a situation in which the assumption holds and for one in which it fails.

Proportional hazards 0.00 0.05 0.10 0 12 24 Hazard 0.0 0.5 1.0 0 12 24 Months Survival Hazards not proportional 0.00 0.05 0.10 0 12 24 Hazard 0.0 0.5 1.0 0 12 24 Months Survival Unexposed Exposed
Figure 8.7. Schematic hazards (top) and survival curves (bottom). On the left, the exposed hazard is 2.5 times the unexposed hazard at every time, although both fall, and the survival curves separate steadily. On the right, the exposed hazard is higher early and lower later, the hazard ratio changes over time, and the survival curves cross.

The assumption fails when the effect of the exposure changes over follow-up. A surgical treatment compared with medical management may carry a higher hazard of death in the weeks after the operation and a lower hazard afterwards, so the hazard ratio is above 1 early and below 1 later, and a single hazard ratio averages these into a number that describes neither period. Figure 8.7 contrasts the two situations. Kaplan-Meier curves that cross, or that separate early and then converge, suggest that the hazards are not proportional. More formal checks compare the log of the cumulative hazard in each group over time or test whether the hazard ratio drifts with time.

When the hazards are not proportional, the protocol can provide alternatives. The analyst can report separate hazard ratios for prespecified periods, such as the first six months and the remainder of follow-up; report the difference in survival at prespecified times; or report the difference in restricted mean survival time, which Royston and Parmar (2013) recommend as a summary that does not depend on proportional hazards. Hernán (2010) adds a further caution. Period-specific hazard ratios compare the participants who remain at risk at each time, and in later periods those participants are a selected group, because the most susceptible members of the higher-risk group have already had the event. Survival curves and risks at fixed times are less vulnerable to this problem, and he concludes that survival curves are more informative than hazard ratios and should be used more widely in observational studies.

A hazard ratio is not a risk ratio

Under proportional hazards, the survival curve of the exposed group is the survival curve of the unexposed group raised to the power of the hazard ratio: S1(t) = S0(t)HR. This relationship shows how far the hazard ratio can differ from the risk ratio at a fixed time. If 90 percent of low-loneliness participants remain event-free at 24 months and the hazard ratio is 2.5, then 0.902.5 = 0.768 of high-loneliness participants remain event-free. The 24-month risks are 0.100 and 0.232, a risk ratio of 2.32. If the event were common, with only half of the unexposed group event-free at 24 months, the same hazard ratio of 2.5 would give risks of 0.500 and 0.823, a risk ratio of 1.65. The hazard ratio and the risk ratio are close when the event is rare and diverge as it becomes common, in the same way that the odds ratio and the risk ratio diverged in Lesson 6.

Equation 8.8 states the relationship between the two survival curves. Calculator 8.8 applies it to translate a hazard ratio into the probability of remaining event-free, the cumulative risk, the risk ratio and the risk difference at a chosen time, and its presets include the uncommon and common events described above.

Survival under proportional hazards
S1(t) = S0(t)HREq 8.8
Under proportional hazards, the probability of remaining event-free in the exposed group at time t equals the probability of remaining event-free in the unexposed group at the same time raised to the power of the hazard ratio.

The check of the proportional hazards assumption, the summary to report if it fails and the choice of effect measure are decisions for the analysis plan. The final part of this section gathers them, with the other decisions in this lesson, into the content of a protocol.

What a Protocol With a Time-to-Event Outcome Must Specify

The decisions in this lesson must be written into the protocol before data collection begins. Most of them cannot be corrected later, and stating them in advance prevents analytic choices from being shaped by the results. Table 8.7 lists the elements with the Campus Connection Follow-up as the example.

Table 8.7. Elements that a protocol with a time-to-event outcome must specify, with the Campus Connection Follow-up as the example.

ElementWhat the protocol statesCampus Connection Follow-up
Time zeroThe moment at which eligibility, exposure assignment and follow-up coincideDate of completing the baseline survey
Eligibility at time zeroCriteria that make every participant event-free and at riskAged 18 to 29 and a baseline PHQ-9 score below 10
ExposureMeasured at or before time zero, or handled as time-varyingBaseline loneliness, classified as high or low at a prespecified cut-off
EventCriterion, ascertainment method, schedule and dating rule, the same in every groupFirst monthly PHQ-9 score of 10 or more, dated to that month
Follow-upMaximum duration and assessment scheduleMonthly questionnaires for 24 months
Censoring rulesDefinitions of loss, withdrawal, administrative end and competing eventsCensored at the last completed questionnaire after three consecutive misses, at withdrawal, or at 24 months; deaths recorded as a competing event and, because they are very rare at these ages, censored in the main analysis
Informative censoringRetention procedures, records of reasons, comparisons and sensitivity analysesReminders by e-mail and text, a reason recorded for each withdrawal, baseline comparison of censored and retained participants, and a worst-case analysis
AnalysisEstimators, tests, effect measure, assumption checks and alternativesKaplan-Meier curves with numbers at risk; survival at 12 and 24 months; log-rank test; hazard ratio from a Cox model adjusted for prespecified confounders; proportional hazards check, with restricted mean survival time as the prespecified alternative
Sample sizeNumber of events required, then participantsComputed with Equation 8.9 below

Planning for enough events

The log-rank example showed that a time-to-event comparison with few events has little power, whatever the number of participants. Schoenfeld (1983) derived the number of events needed to detect a given hazard ratio with a two-sided test. The number of participants follows by dividing the required events by the proportion of participants expected to have the event during follow-up.

Equation 8.9 gives both quantities. Calculator 8.9 evaluates them and adds an allowance for loss to follow-up, with a preset for the Campus Connection protocol.

Events needed to detect a hazard ratio
D = (z1−α/2 + z1−β)² / [p(1 − p) × (ln HR)²]     N = D / πEq 8.9
The number of events required is the squared sum of the standard normal values for the significance level and the power, divided by the product of the proportion of participants who are exposed, one minus that proportion, and the squared natural logarithm of the hazard ratio to be detected. Dividing by the proportion expected to have the event gives the number of participants.

For a new cohort designed to detect a hazard ratio of 1.8 with a two-sided significance level of 0.05 and 80 percent power, with one third of participants in the high-loneliness group, the formula gives (1.960 + 0.842)² ÷ [(1/3) × (2/3) × (ln 1.8)²] = 102.2, which rounds up to 103 events. If about 15 percent of participants are expected to screen positive within 24 months, the study needs 103 ÷ 0.15 = 687 participants (rounded up) before any allowance for loss to follow-up. The protocol targets a hazard ratio of 1.8, smaller than the 2.52 observed in the hypothetical pilot, because pilot estimates from small samples are imprecise and tend to be too large when a pilot is continued only if its results look promising, and because the target should be the smallest effect that would matter for campus mental health planning.

Worked Example 8.5 shows how the elements of Table 8.7 and this sample size calculation are written as the time-to-event subsection of the Campus Connection Follow-up protocol.

Worked Example 8.5: The time-to-event subsection of the Campus Connection Follow-up protocol

Outcome and follow-up. The primary outcome is time from completion of the baseline survey (time zero) to the first monthly PHQ-9 score of 10 or more. Eligible participants are adults aged 18 to 29 with a baseline PHQ-9 score below 10. Participants will receive an online PHQ-9 at the same point in each month for 24 months, with identical schedules and reminders in both loneliness groups, and an event will be dated to the month of the first qualifying score.

Censoring. Participants who miss three consecutive questionnaires will be censored at the date of their last completed questionnaire. Participants who withdraw will be censored at the date of withdrawal, and those who remain event-free will be censored administratively at 24 months. Deaths will be recorded as competing events and, because they are expected to be very rare at ages 18 to 29, censored in the main analysis. The reason for each withdrawal will be recorded, the baseline characteristics of participants censored before 24 months will be compared with those of participants followed to an event or to 24 months, and the hazard ratio will be recomputed under the assumption that every participant lost to follow-up screened positive at the time of censoring.

Analysis. Kaplan-Meier curves will be presented for each loneliness group with censoring marks and numbers at risk at 0, 6, 12, 18 and 24 months, together with the estimated probability of remaining event-free at 12 and 24 months and 95 percent confidence intervals. The groups will be compared with the log-rank test, and the hazard ratio with its 95 percent confidence interval will be estimated with a Cox model adjusted for the confounders identified in the study's causal diagram. If the proportional hazards assumption is not supported, the difference in restricted mean survival time to 24 months will be reported as the primary comparison.

Sample size. To detect a hazard ratio of 1.8 with 80 percent power at a two-sided significance level of 0.05, with one third of participants in the high-loneliness group, 103 events are required. Assuming that 15 percent of participants screen positive within 24 months, 687 participants are needed, increased to 859 to allow for 20 percent loss to follow-up.

The inflation for loss to follow-up in the last paragraph of Worked Example 8.5 divides 687 by 0.80 and rounds up, the same approach used to inflate a survey sample for expected non-response. Losses reduce the number of observed events, so a protocol that expects heavy loss should either plan a larger cohort or a longer follow-up.

This section has compared groups with the log-rank test and the hazard ratio, set out the proportional hazards assumption and its alternatives, and gathered the design decisions of the whole lesson into the content of a protocol. The note below describes how the lesson applies to writing the time-to-event subsection of a protocol, and the reflection, key takeaways and knowledge check that follow review the section before the final review of the lesson.

Applying this lesson: writing a time-to-event subsection

A protocol that follows participants over time contains a time-to-event subsection that covers each row of Table 8.7, as Worked Example 8.5 does. It states time zero and confirms that eligibility, exposure assignment and the start of follow-up coincide at that moment; defines the event, its ascertainment and its dating rule; lists every expected source of censoring and how each will be recorded; and reports the number of events and participants needed, which Calculator 8.9 computes. For a cross-sectional study, a useful check on the design is to work out what time zero, the event and the censoring rules would be if a follow-up phase were added, and which source of censoring would most threaten the non-informative censoring assumption. Informative censoring and loss to follow-up also belong in the protocol's threats to validity and mitigation subsection, described in Lesson 7 Section 4.

Reflection

A randomized trial compares a peer-support intervention with usual care for preventing a first episode of moderate depressive symptoms among first-year university students over 18 months. The report gives a hazard ratio of 0.65 (95 percent confidence interval 0.48 to 0.88). The Kaplan-Meier curves separate steadily over follow-up and do not cross, the median is not reached in either group, and the estimated probabilities of remaining free of moderate symptoms at 18 months are 0.78 in the intervention group and 0.68 in the usual-care group. For this exercise, the hazard is the event rate among participants still event-free; the proportional hazards assumption states that the ratio of the hazards is the same at every time; under that assumption the intervention group's survival equals the usual-care survival raised to the power of the hazard ratio; and the cumulative risk at a time is one minus the survival at that time. (a) Interpret the hazard ratio in one or two sentences. (b) Explain whether the description of the curves supports the proportional hazards assumption. (c) Check that the reported survival probabilities are consistent with the hazard ratio, then compute the 18-month risk ratio and risk difference and explain why the risk ratio differs from the hazard ratio. (d) Name two elements the trial protocol must have specified for the time-to-event outcome.

Model answer(a) At any time during the 18 months, students in the peer-support group who had not yet developed moderate symptoms did so at about 65 percent of the rate in the usual-care group, a 35 percent lower hazard, and the confidence interval of 0.48 to 0.88 excludes 1. (b) Curves that separate steadily and do not cross are consistent with proportional hazards, although a formal check would be needed to confirm it. Crossing curves, or curves that separate early and then run parallel, would suggest that the hazard ratio changes over time. (c) Under proportional hazards the intervention survival should be 0.68 raised to the power 0.65, which is 0.78, matching the report. The 18-month risks are 1 − 0.78 = 0.22 and 1 − 0.68 = 0.32, so the risk ratio is 0.22/0.32 = 0.69 and the risk difference is −0.10, or 10 fewer students with moderate symptoms per 100. The risk ratio is closer to 1 than the hazard ratio because risks accumulate over time and are capped at 1, while the hazard ratio compares rates among those still at risk. (d) The protocol must have specified time zero (randomization) and the event definition with its ascertainment schedule, which must be the same in both arms; it should also have stated the censoring rules and the planned check of proportional hazards.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • The log-rank test compares observed events with the events expected if both groups had the same survival, and with few events it has little power.
  • The hazard is the event rate among those still at risk, and under constant hazards the hazard ratio equals the incidence rate ratio.
  • Proportional hazards means that the ratio of the hazards is the same throughout follow-up; crossing curves suggest otherwise, and fixed-time survival or restricted mean survival time are alternatives.
  • A hazard ratio is closer to the risk ratio when the event is rare and further from it when the event is common.
  • A time-to-event protocol states time zero, eligibility, the event, follow-up, censoring rules, the informative-censoring plan, the analysis and the number of events needed.
Knowledge Check: this section

1. In the log-rank test, how is the expected number of events in group A at a single event time calculated?

Under the null hypothesis that both groups have the same survival, the events at each time are shared in proportion to the numbers at risk, so the expected number in group A is the total events multiplied by nA/n. Splitting events in half ignores unequal risk sets, and using group A's own rate reproduces the observed count.

2. Group A has 30 events over 2,000 person-years and group B has 20 events over 4,000 person-years. Assuming constant hazards, what is the hazard ratio comparing A with B?

The rates are 30/2,000 = 0.015 and 20/4,000 = 0.005 per person-year, so the hazard ratio is 0.015/0.005 = 3.0. The value 1.5 is the ratio of event counts, which ignores the different amounts of person-time, and 0.33 compares B with A.

3. Which statement expresses the proportional hazards assumption?

Proportional hazards means that the ratio of the two hazards stays the same over follow-up, while each hazard may rise or fall. Constant hazards in each group is a stronger assumption that the Cox model does not require, and equal proportions with the event would describe no difference between the groups.

4. Under proportional hazards, 60 percent of unexposed participants remain event-free at 24 months and the hazard ratio is 2.0. What is the 24-month risk ratio?

The exposed survival is 0.602.0 = 0.36, so the risks are 0.64 and 0.40 and the risk ratio is 0.64/0.40 = 1.6. Because the event is common, the risk ratio lies closer to 1 than the hazard ratio. The risk ratio equals the hazard ratio only approximately, and only when the event is rare.

5. A team is planning a cohort study to detect a hazard ratio of 1.5. Which quantity most directly determines the study's power?

The information in a time-to-event comparison comes from the events, and Schoenfeld's formula gives the number of events needed for a chosen hazard ratio, significance level and power. Enrolling many participants helps only if enough of them have the event during follow-up.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 5

Final Review & Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson treated time-to-event analysis as a way of estimating incidence when follow-up is incomplete. Section 1 set out the design decisions that give a time-to-event outcome its meaning (time zero, the event and the end of follow-up) and the forms of censoring that make follow-up incomplete. It showed that the shortcuts of treating censored participants as event-free or dropping them misstate the risk, and that the methods of the lesson depend on censoring being non-informative. Sections 2 and 3 estimated the probability of remaining event-free over time, first with the actuarial life table and then with the Kaplan-Meier estimator, and showed how to read survival at a time, the median, censoring marks and numbers at risk from a curve.

Section 4 compared groups. The log-rank test compares observed with expected events, the hazard ratio compares event rates among participants still at risk, and the proportional hazards assumption states that this ratio is the same throughout follow-up. Because the information in these comparisons comes from events, the number of events needed is the starting point of a sample size calculation. Each of these decisions belongs in the protocol, written before any data are collected.

Key Takeaways from this lesson

  • A time-to-event outcome records a follow-up time and an event status for each participant, measured from a time zero at which eligibility, exposure assignment and follow-up coincide.
  • Right censoring arises from loss to follow-up, withdrawal, moving away and the administrative end of a study; left censoring, interval censoring and late entry are distinct situations that need their own handling.
  • All of the methods in this lesson assume non-informative censoring, which cannot be tested directly but can be made more plausible and checked through retention procedures, recorded reasons, comparisons and sensitivity analyses.
  • Treating censored participants as event-free understates the risk, and dropping them tends to overstate it; the true risk lies between the first shortcut and the value obtained by counting every early censoring as an event.
  • The actuarial life table multiplies conditional probabilities of surviving fixed intervals and counts each withdrawal as half a person at risk.
  • The Kaplan-Meier estimator multiplies conditional probabilities at each observed event time, keeps censored participants in the risk set until they leave, and gives a step curve whose median and tail must be read with the numbers at risk.
  • The log-rank test compares the observed events in each group with the events expected if both groups shared the same survival, and its power depends on the number of events.
  • The hazard is the event rate among those still at risk, and under constant hazards the hazard ratio equals the incidence rate ratio.
  • Proportional hazards means that the ratio of the hazards is the same throughout follow-up; crossing curves suggest otherwise, and differences in survival at fixed times or in restricted mean survival time are alternatives.
  • A protocol with a time-to-event outcome must specify time zero, eligibility, the event and its ascertainment, follow-up, censoring rules, the plan for informative censoring, the analysis and the number of events needed.

Core Concepts Reviewed

Section 1: time-to-event data, time zero, the event, follow-up, right, left and interval censoring, late entry, informative and non-informative censoring, competing events, and immortal time and lead-time bias.

Section 2: the actuarial life table, conditional probabilities of the event and of survival, the effective number at risk and the half-interval assumption, and cumulative survival as a running product.

Section 3: the Kaplan-Meier product-limit estimator, risk sets at event times, ties and censoring conventions, survival at a time, median survival, censoring marks, numbers at risk and restricted mean survival time.

Section 4: the log-rank test, the hazard, the hazard ratio as a ratio of rates, proportional hazards, the relationship between hazard ratios and risk ratios, protocol elements and the number of events needed.

The final reflection asks you to bring these decisions together for a new study.

Reflection

A public health unit plans to study whether adults aged 65 and older who live alone reach a first fall-related emergency department visit sooner than those who live with others. It can recruit 1,200 adults through community centres over six months and follow each for two years through linked provincial emergency department records, with an annual telephone interview to update contact details and living arrangements. About one third of participants are expected to live alone. Earlier work suggests that about 12 percent will have a fall-related visit within two years, and the team considers a hazard ratio of 1.5 the smallest difference that would change its programs. The number of events needed to detect a hazard ratio with a two-sided significance level of 0.05 and 80 percent power is D = (1.960 + 0.842)² / [p(1 − p) × (ln HR)²], where p is the proportion exposed, and the number of participants is D divided by the proportion expected to have the event. Time zero is the moment the clock starts; right censoring occurs when follow-up ends before the event; censoring is non-informative when censored participants have the same future risk as those who remain; and a competing event is one that makes the event of interest impossible. Write the time-to-event subsection of this study's protocol. It should (a) define time zero, eligibility and the event; (b) list the sources of censoring, including any competing event; (c) identify the main threat of informative censoring and how the design will address it; (d) outline the analysis, including how proportional hazards will be checked and what will be reported if the assumption does not hold; and (e) compute the number of events needed and judge whether 1,200 participants are enough.

Model answer(a) Time zero is the date of the baseline interview at the community centre, when eligibility (age 65 or older, living in the community in the province and able to complete the interview), living arrangement and the start of record linkage coincide. The event is the first emergency department visit with a fall-related diagnosis code in the linked records, dated to the visit. (b) Participants are censored administratively at two years, at the date they leave the province (identified from the annual interview or loss of health-plan coverage) and at withdrawal. Death from any cause is a competing event and will be analyzed as such, because a person who has died cannot have a later fall-related visit. (c) Linked records make loss to follow-up for the event itself small, so the main threat is moving into long-term care or out of the province because of declining health, which removes people at high risk of falls. The protocol will record the reason for every departure, compare baseline frailty and mobility in censored and retained participants, and run a sensitivity analysis that treats each such departure as an event. (d) Kaplan-Meier curves by living arrangement with censoring marks and numbers at risk, the cumulative incidence of a fall-related visit at one and two years estimated with death as a competing event (Austin et al., 2016), a log-rank test, and a hazard ratio from a Cox model adjusted for prespecified confounders (age, sex, frailty and prior falls). Proportional hazards will be checked by inspecting the curves and testing whether the hazard ratio changes over time; if the assumption fails, the difference in restricted mean event-free time to two years will be the primary comparison. (e) D = 7.849 / [(1/3)(2/3)(ln 1.5)²] = 7.849 / (0.2222 × 0.1644) = 214.8, so 215 events are needed. With 12 percent having the event, 1,200 participants yield about 144 events, which is not enough; about 215 / 0.12 = 1,792 participants would be needed before allowing for loss, or the team would need longer follow-up or a higher-risk population.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Complete all 15 questions below with 100% accuracy to finish this lesson. You must also pass every section knowledge check and complete every reflection, including the final reflection above, before submitting.

Final Assessment: Time-to-Event Data

1. Which two pieces of information does every time-to-event analysis need for each participant?

Each participant contributes a time from time zero to the end of follow-up and a status recording whether that time ended with the event (1) or with censoring (0). Event dates alone cannot represent censored participants, and counts and person-time totals summarize groups.

2. A study of time to hospital readmission uses the date of discharge as time zero. A participant moves out of the province 40 days after discharge without being readmitted. In a Kaplan-Meier analysis, this participant:

The participant is right-censored at day 40: they are part of the risk set at every event time up to day 40 and then leave it without causing a drop. Treating the move as a readmission invents an event, and counting the participant after day 40 assumes observation that did not happen.

3. Which design feature makes informative censoring easiest to detect?

Informative censoring cannot be tested directly, but recording why each participant left and comparing the baseline characteristics of censored and retained participants shows whether those who left differed. Excluding lost participants makes the problem worse, and the choice of time zero or interval width does not address it.

4. In an actuarial life table, 60 participants enter an interval, 12 withdraw and 9 have the event. The cumulative survival at the end of the previous interval was 0.90. What is the cumulative survival at the end of this interval?

The effective number at risk is 60 − 6 = 54, so q = 9/54 = 0.167 and p = 0.833, and the cumulative survival is 0.90 × 0.833 = 0.75. Ignoring withdrawals gives 0.77, removing them gives 0.73, and 0.83 is the conditional probability alone.

5. What is the main difference between the Kaplan-Meier estimator and the actuarial life table?

Both methods multiply conditional probabilities of surviving successive intervals. Kaplan-Meier lets each interval end at an observed event time, so censored participants are counted fully at risk until they leave and no half-interval assumption is needed. Both assume non-informative censoring, and neither estimates a hazard ratio.

6. Ten participants are at risk at month 3, when 2 have the event. One participant is censored at month 4. At month 6, 7 participants are at risk and 1 has the event. What is the Kaplan-Meier estimate at month 6?

The estimate after month 3 is 8/10 = 0.800. At month 6 the conditional probability is 6/7, so the estimate is 0.800 × 6/7 = 0.686. Forgetting the censoring gives 0.800 × 7/8 = 0.700, treating the censoring as an event gives 0.600, and dropping the censored participant from the start gives 0.667.

7. A published Kaplan-Meier curve for one group ends at 0.58 after 36 months of follow-up. What can be said about the median survival time?

The median is the earliest time at which the curve falls to 0.50 or below. A curve that ends at 0.58 never reaches 0.50, so the median is reported as not reached. Censoring does not prevent a median from being estimated when the curve does fall to 0.50.

8. A log-rank test comparing two groups of ten participants, with eight events in total, gives a chi-squared statistic of 1.67 and p = 0.20. Which interpretation is most appropriate?

A non-significant result means that a difference of the observed size is compatible with chance. With only eight events, the test has little power, so the result is weak evidence in either direction. It does not show that survival is equal or that the hazard ratio is exactly 1, and the p-value says nothing about proportional hazards.

9. A trial reports a hazard ratio of 0.70 for an intervention compared with control. Which interpretation is correct?

A hazard ratio compares event rates among participants who are still at risk, so a value of 0.70 means a 30 percent lower rate at any time, under proportional hazards. The cumulative risk ratio at the end of follow-up is closer to 1 than 0.70 when the event is common, and the hazard ratio is not a proportion of participants or a delay.

10. Which pattern in two Kaplan-Meier curves most strongly suggests that the hazards are not proportional?

Crossing curves mean that the group with the higher hazard early in follow-up has the lower hazard later, so the hazard ratio changes over time. Curves that separate steadily are consistent with proportional hazards, and a shared shape or different amounts of censoring do not by themselves indicate a changing ratio.

11. Why does a hazard ratio differ from the risk ratio at 24 months when the event is common?

Under proportional hazards, S1(t) = S0(t)HR. As risks accumulate they approach their ceiling of 1, so the ratio of cumulative risks moves toward 1 while the ratio of rates among those still at risk does not. The hazard ratio is a ratio of rates that depends directly on person-time.

12. A screening program reports longer survival from diagnosis among screen-detected cancers than among cancers diagnosed after symptoms appeared. What is the most likely design explanation?

Screening moves the date of diagnosis earlier. If time zero is the date of diagnosis, screen-detected cases gain survival time even if death occurs at the same moment, which is lead-time bias. HSCI 230 Lesson 10 treats it in detail; a protocol avoids it by starting the clock at a point that does not depend on when the disease was detected.

13. Using Schoenfeld's formula with a two-sided significance level of 0.05, 80 percent power and equal-sized groups, about how many events are needed to detect a hazard ratio of 2.0?

D = (1.960 + 0.842)² / [0.5 × 0.5 × (ln 2.0)²] = 7.849 / (0.25 × 0.4805) = 65.3, which rounds up to 66 events. The number 248 is roughly what a hazard ratio of 0.7 requires, which shows how quickly the required events rise as the hazard ratio approaches 1.

14. In a study of time to a first diagnosis of dementia, deaths from other causes are treated as censoring in a Kaplan-Meier analysis. What is the consequence?

Treating death as censoring assumes that the people who died could still develop dementia later, which is false. One minus the Kaplan-Meier estimate therefore overstates the cumulative incidence of dementia, and the overstatement grows with the number of deaths. Competing-risks methods estimate the cumulative incidence correctly.

15. Which statement belongs in the time-to-event subsection of a study protocol?

A protocol states censoring rules before data collection so that they are the same for every group and cannot be shaped by the results. Deciding rules or intervals after seeing the data invites biased choices, and excluding participants lost to follow-up discards their event-free time and tends to overstate the risk.

✦ Complete the final reflection above before submitting