HSCI 341, Lesson 6

Measures of
Association

Fundamental Epidemiological Concepts and Approaches

Learning objectives for this lesson:

  • Calculate and interpret the risk ratio, incidence rate ratio, and odds ratio
  • Compute risk difference, attributable fraction (exposed), and population attributable measures
  • Understand when to use each measure of association
  • Correctly distinguish between strength of association and statistical significance
  • Understand the basis for hypothesis tests and confidence intervals

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Ratio Measures of Association
Risk Ratio (Relative Risk, RR) The ratio of the cumulative incidence (risk) in the exposed group to that in the unexposed group: RR = R₁ / R₀. Used in cohort studies and RCTs over a defined follow-up period.
Incidence Rate Ratio (IRR) The ratio of incidence rates (events per person-time) in the exposed and unexposed groups. Appropriate when follow-up time varies across individuals.
Odds Ratio (OR) The ratio of the odds of disease in the exposed group to the odds in the unexposed group (Cornfield, 1951). The standard measure in case-control studies and the natural output of logistic regression.
Prevalence Ratio (PR) The ratio of prevalence in the exposed group to prevalence in the unexposed group. Preferred over the prevalence odds ratio when the outcome is not rare in cross-sectional analyses.
Hazard Ratio (HR) The ratio of instantaneous event rates (hazards) between groups, typically from Cox proportional hazards regression. Approximates the rate ratio under proportional hazards.
Rare-Disease Assumption When disease prevalence is low (commonly < 5%), the odds ratio approximates the risk ratio. Important for interpreting case-control results.
Difference Measures of Association
Risk Difference (RD, Attributable Risk) The absolute difference in risk between exposed and unexposed groups: RD = R₁ − R₀. Captures the public-health impact in absolute terms.
Rate Difference The absolute difference in incidence rates between exposed and unexposed groups, in events per person-time.
Attributable Fraction in the Exposed (AFe) Among the exposed, the proportion of disease attributable to the exposure: AFe = (R₁ − R₀) / R₁ = (RR − 1) / RR.
Population Attributable Risk (PAR) The excess risk in the total population attributable to the exposure: PAR = Rpop − R₀. Reflects both effect size and exposure prevalence.
Population Attributable Fraction (AFp) The proportion of disease in the total population attributable to the exposure (Northridge, 1995). Useful for prioritising public-health interventions.
Number Needed to Treat (NNT) 1 / |risk difference|: the average number of patients who must receive a beneficial intervention for one to avoid the bad outcome (Laupacis, Sackett, & Roberts, 1988).
Number Needed to Harm (NNH) 1 / |risk difference| for harmful exposures: the average number exposed for one additional case of harm.
Inference & Interpretation
Null Hypothesis (H₀) A statement of no effect or no difference (e.g., RR = 1, RD = 0). The hypothesis tested by significance tests.
p-Value The probability of observing data as extreme or more extreme than the observed, assuming the null hypothesis is true. Not the probability that the null is true.
Confidence Interval A range of values consistent with the data at a chosen confidence level. For a ratio measure, an interval that excludes 1 implies statistical significance at the corresponding level.
Strength of Association vs. Statistical Significance The size of an effect (e.g., RR) is conceptually distinct from how confident we are that it differs from chance (p-value, CI width). Large samples can yield significant findings for trivial effects, and vice versa.
Effect Modification (Interaction) When the magnitude of an exposure-outcome association differs across levels of a third variable. Reported as stratum-specific estimates rather than adjusted away. Additive and multiplicative scales can give different interaction conclusions; reporting both is recommended (Knol & VanderWeele, 2012; VanderWeele & Knol, 2014).
No matching entries. Try a different search term.
Section 1

Introduction & Ratio Measures of Association

⏱ Estimated reading time: 15 minutes

Lesson 6 · HSCI 341

Comparison Is the Core Operation

A single frequency tells you how common a disease is. A measure of association tells you how much the exposure changes it.

Section 1 of 4

Ratio Measures of Association

Risk ratio, incidence rate ratio, and odds ratio: formulas, relationships, and worked examples.

The core idea

Why ratios, and why not just P-values

Measure of association

How strongly is the exposure linked to disease? Independent of sample size. The number that answers the biological or causal question.

P-value

How likely are these data if there were truly no association? Sensitive to sample size. Does not describe effect magnitude.

Null value for ratio measures: 1. Values > 1 indicate increased risk; values < 1 indicate protection.

Three formulas

RR, IR, and OR from the 2x2 table

Risk Ratio (RR)
\[ \color{#0B7B6B}{RR} = \frac{\color{#C2410C}{a_1/(a_1+b_1)}}{\color{#1D4ED8}{a_0/(a_0+b_0)}} = \frac{\color{#C2410C}{R_{E+}}}{\color{#1D4ED8}{R_{E-}}} \]
RR risk ratioRE+ risk in the exposedRE− risk in the unexposed
Incidence Rate Ratio (IR)
\[ \color{#0B7B6B}{IR} = \frac{\color{#C2410C}{a_1 / T_1}}{\color{#1D4ED8}{a_0 / T_0}} \]
IR incidence rate ratioa1/T1 rate in the exposeda0/T0 rate in the unexposed
Odds Ratio (OR)
\[ \color{#0B7B6B}{OR} = \frac{\color{#C2410C}{a_1 / b_1}}{\color{#1D4ED8}{a_0 / b_0}} = \frac{\color{#C2410C}{a_1} \color{#6D28D9}{b_0}}{\color{#1D4ED8}{a_0} \color{#BE185D}{b_1}} \]
OR odds ratioa1 exposed casesb1 exposed non-casesa0 unexposed casesb0 unexposed non-cases
Worked example 1

Brazil water cistern study

Exposure: household water cistern
Outcome: diarrhea

RR
\[ RR = \frac{194/1782}{303/1617} = \frac{0.109}{0.187} = 0.58 \]
OR
\[ OR = \frac{194 \times 1314}{303 \times 1588} = 0.53 \]

Interpretation

Both measures are below 1: cistern access is protective. The RR of 0.58 means the risk of diarrhea is 42% lower in the cistern group. OR is further from the null than RR, as the general relationship predicts.

Worked example 2

Migraine incidence rates

Exposure: female sex (vs. male)
Outcome: migraine onset, ages 30-40

Incidence Rate Ratio
\[ IR = \frac{131/250}{44/236} = \frac{0.524}{0.186} = 2.81 \]

Interpretation

The migraine rate is 2.81 times higher in females than males aged 30-40. Person-months in the denominator account for variable follow-up time across participants.

The OR's special role

Why OR is the case-control measure

OR is symmetric: the odds of disease given exposure equal the odds of exposure given disease. Case-control studies sample on outcome, so only OR can be computed directly.

Cornfield approximation (rare disease)
\[ \text{When } \color{#C2410C}{p(D+)} < 0.05:\quad \color{#0B7B6B}{OR} \approx \color{#1D4ED8}{RR} \]
p(D+) disease frequencyOR odds ratioRR risk ratio

As disease prevalence rises, OR diverges from RR and moves further from the null. IR sits between RR and OR. This ordering holds on both sides of one.

Carry forward

Three anchors from this section

  • RR and IR come from cohort studies where the population at risk is observed directly.
  • OR is the case-control measure; it approximates RR when disease is rare (under 5%).
  • Ordering from the null: IR further than RR, OR furthest of all, on both sides of 1.

Introduction and Overview

An earlier lesson gave us measures of disease frequency: prevalence, incidence, risk, and rate. An earlier lesson took the same probabilistic vocabulary and applied it at the level of a single test. This lesson brings these strands together: it combines the disease-frequency vocabulary with the 2×2 contingency logic to produce measures of association, the quantitative comparison between exposed and unexposed groups that is the central output of analytic epidemiology. The four content sections build up from the three ratio measures (this section: risk ratio, rate ratio, odds ratio), through difference measures and exposed-group attributable fractions (a later section), to population-level attributable measures and how each measure relates to study design (a later section), and finally to the hypothesis-testing and confidence-interval machinery that turns each of these point estimates into a defensible inference (a later section).

Learning Objectives

  • Explain why measures of association are used in epidemiology.
  • Set up incidence risk and incidence rate data in 2×2 tables.
  • Calculate and interpret the risk ratio (RR), incidence rate ratio (IR), and odds ratio (OR).
  • Describe the relationships among RR, IR, and OR.

The section begins with the reason for measuring association, sets out the 2×2 tables from which the ratio measures are computed, defines the risk ratio, the incidence rate ratio and the odds ratio, and closes with the relationships among the three.

Why Measure Association?

Measures of association assess the magnitude of the relationship between an exposure (a potential cause) and a disease. Unlike measures of statistical significance, which are heavily dependent on sample size, measures of association indicate the strength of the effect, that is, how much more (or less) likely disease is in exposed compared to non-exposed groups (Tripepi et al., 2007).

Box 6.1 sets out the difference between the strength of an association and its statistical significance, a distinction to which Section 4 returns when it introduces hypothesis tests and confidence intervals.

Box 6.1: Strength vs. Significance

A measure of association tells you how strongly an exposure is linked to disease. A P-value tells you how likely the observed data would be under the null hypothesis of no association. A strong association can be non-significant (small sample), and a weak association can be highly significant (large sample). Always report both.

A measure of association compares the frequency of disease in exposed and non-exposed groups, so the data must first be arranged in a form that makes this comparison possible. The next part sets out that arrangement.

Data Layout

Depending on study design, disease frequency can be expressed as incidence risk, incidence rate, prevalence, or odds. The formulae later in this lesson are written in terms of the cell labels of Table 6.1, for risk data, and Table 6.2, for rate data. For risk data, the standard 2×2 table is:

Table 6.1. Layout of the 2×2 table for incidence risk data, with exposure status in the columns and disease status in the rows.

ExposedNon-exposedTotal
Diseaseda1a0m1
Non-diseasedb1b0m0
Totaln1n0n

For rate data, the denominator is person-time at risk rather than the number of individuals:

Table 6.2. Layout of the table for incidence rate data, giving the number of cases and the person-time at risk in each exposure group.

ExposedNon-exposedTotal
Number of casesa1a0m1
Person-time at riskt1t0t

A quick word on odds, since the third measure below is built from them. The risk of an outcome is the number of cases divided by everyone at risk, meaning cases plus non-cases. The odds of the same outcome is the number of cases divided by the non-cases alone. If 20 of 100 people get sick, the risk is 20/100 = 0.20, but the odds are 20/80 = 0.25. Risk and odds stay close while an outcome is uncommon and pull apart as it becomes common, which is the same fact that lets the odds ratio stand in for the risk ratio only when disease is rare.

With the two layouts in place and the difference between risk and odds defined, the next part introduces the three ratio measures that are computed from these tables.

Three Ratio Measures of Association

The risk ratio, the incidence rate ratio and the odds ratio each divide a measure of disease frequency in the exposed group by the same measure in the non-exposed group. Box 6.2 recalls how HSCI 230 defined the three measures and poses a short retrieval question on them.

Box 6.2: Recall: Risk, Rate and Odds Ratios

HSCI 230 Lesson 5, Sections 1 and 2 (Introduction and Cohort Study Design; Risk-Based and Rate-Based Designs), defined the risk ratio as the risk of disease in the exposed divided by the risk in the unexposed, and the incidence rate ratio as the same comparison of incidence rates, each a count of new cases over person-time. HSCI 230 Lesson 4, Section 1 (Introduction and the Study Base), defined the odds ratio as the cross-product of the two-by-two table, the measure that a case-control study estimates. For all three, a value of 1 means no association, values above 1 mean more disease in the exposed and values below 1 mean less. This section adds which designs supply each measure and how the three relate to each other, and Section 4 adds their confidence intervals.

Retrieval question. In a cohort, 30 of 200 exposed and 10 of 200 unexposed people develop the disease. What are the risk ratio and the odds ratio?

Show the answer▼

RR = (30/200) ÷ (10/200) = 0.15 ÷ 0.05 = 3.0, and OR = (30 × 190) ÷ (10 × 170) = 5,700 ÷ 1,700 = 3.35. The odds ratio lies further from 1, as the relationships later in this section explain.

The three measures differ in the data they require, and so in the study designs that can supply them. The flip cards below give the formula for each measure and the designs from which it can be computed.

Click each card for what each measure requires of the study design:

Risk Ratio (RR)Click to learn more
Incidence Rate Ratio (IR)Click to learn more
Odds Ratio (OR)Click to learn more

Two worked examples show the measures in use. Worked Example 6.1 applies the risk ratio and the odds ratio to a study from Brazil that compared diarrhea among those with and without a water cistern, using the risk layout of Table 6.1.

Worked Example 6.1: Brazil Water Cistern Study

Diarrhea & Water Cistern Presence
Water CisternNo CisternTotal
Diarrhea Present194303497
Diarrhea Absent1,5881,3142,902
Total1,7821,6173,399
  • RR = (194/1782) / (303/1617) = 0.109 / 0.187 = 0.58
  • OR = (194 × 1314) / (303 × 1588) = 0.53

Both measures indicate that having a water cistern is protective against diarrhea (values < 1). The RR of 0.58 means the risk is 42% lower among those with cisterns.

Worked Example 6.2 turns to rate data in the layout of Table 6.2 and computes the incidence rate ratio for migraine in females and males aged 30 to 40.

Worked Example 6.2: Migraine Incidence Rates

Gender and Migraine (Ages 30–40)
FemaleMaleTotal
Cases of migraine13144175
Person-months250236486

IR = (131/250) / (44/236) = 0.524 / 0.186 = 2.81

The rate of migraine is 2.81 times higher in females than males aged 30–40.

The Three Formulae, with Calculators

Each formula below is paired with a calculator that starts from the worked examples above. Entering the cistern data into both the risk ratio and the odds ratio calculators shows how far the two measures separate when the outcome is common.

Equation 6.1 defines the risk ratio, and Calculator 6.1 computes it, starting from the cistern data of Worked Example 6.1. Equation 6.2 defines the odds ratio, which Calculator 6.2 computes from the same data, and Equation 6.3 defines the incidence rate ratio, which Calculator 6.3 computes from the migraine data of Worked Example 6.2.

Risk ratio
RR = (a1/n1) / (a0/n0)Eq 6.1
The risk ratio is the risk in the exposed group divided by the risk in the unexposed group.
Odds ratio
OR = (a1/b1) / (a0/b0) = (a1 × b0) / (a0 × b1)Eq 6.2
The odds ratio is the odds of disease in the exposed group divided by the odds of disease in the unexposed group, which simplifies to the cross-product of the table.
Incidence rate ratio
IR = (a1/t1) / (a0/t0)Eq 6.3
The incidence rate ratio is the incidence rate in the exposed group divided by the incidence rate in the unexposed group, each a count of cases over person-time.

The risk ratio and the odds ratio give different values for the same data, as the cistern example shows with an RR of 0.58 and an OR of 0.53. The next part explains how the three measures are related and when one of them can stand in for another.

Relationships Among RR, IR, and OR

In general, IR values are further from the null (1) than RR values, and OR values are even further away (Cornfield, 1951; Knol et al., 2008). This can be visualised on a number line, as in Figure 6.1:

1 (null) 0 ∞ OR IR RR RR IR OR

Figure 6.1. General relationships among RR, IR, and OR. OR is always furthest from the null value of 1.

Under some conditions the measures come close to one another, and the accordion below explains when the OR approximates the RR, when the RR approximates the IR, and when the OR from a case-control study estimates the IR.

When is OR ≈ RR?▼

When the disease is rare (prevalence or incidence risk < 5%), OR approximates RR, the classic Cornfield (1951) approximation, which showed how case-control studies such as the landmark study of smoking and lung cancer by Doll and Hill (1950) could be interpreted in terms of relative risk. This is because when a1 is small relative to n1, the denominator of the odds (b1) is approximately equal to n1, and similarly for the non-exposed group. In the cistern example, the overall risk was 14.6%, so OR (0.53) was more extreme than RR (0.58); when outcomes are not rare, treating OR as RR can substantially overstate the effect (Knol et al., 2008).

When is RR ≈ IR?▼

RR and IR will be close to each other if the exposure has a negligible impact on the total time at risk in the study population. This occurs when the disease is rare or when IR is close to the null value (IR ≈ 1).

OR as an Estimator of IR▼

OR is a good estimator of IR under certain conditions in case-control studies. If controls are selected using cumulative or risk-based sampling (all non-cases after cases have occurred), then OR estimates IR only if the disease is rare. If controls are selected using density sampling (a control selected from non-cases each time a case occurs), then OR is a direct estimate of IR regardless of disease rarity.

Interactive 6.1 shows the rare-disease condition at work: it computes the RR and the OR from one 2×2 table and plots how far the OR moves from the RR as the outcome becomes more common, so that the effect of outcome prevalence on the gap between the two measures can be seen directly.

⚖ Interactive 6.1: Risk Ratio vs. Odds Ratio

This tool calculates the risk ratio (RR) and the odds ratio (OR) from the same 2×2 table and plots how far apart they are as the outcome becomes more common. It is meant to show you why the OR can stand in for the RR only when the outcome is rare. As you move the outcome prevalence shifter to the right, notice that the RR stays about the same (the flat RR line on the chart) while the OR moves away from it: above the RR for a harmful exposure and below it for a protective one. You can also click any cell of the table to type your own counts, or start from one of the presets.
2×2 table (click cells to edit)
Y+Y−Total
E+4060100
E−2080100
Scales the table to give this overall prevalence (preserving RR).
Presets:
RR vs. OR as outcome prevalence climbs
Risk in E+
–
Risk in E−
–
Outcome prevalence
–
Risk Ratio
–
Odds Ratio
–
OR / RR
–
Try the Common outcome preset: a true RR of 2.0 produces an OR around 3.0+. Reporting the OR as if it were a "risk ratio" overstates the harm by 50%. The "rare disease assumption" is what justifies treating OR as RR, and it should not be assumed without checking.

This section has explained why association is measured, set out the layouts for risk and rate data, defined the risk ratio, the incidence rate ratio and the odds ratio, and shown that the OR lies furthest from the null and approximates the RR only when the outcome is rare. The Key Takeaways below summarise these points, the knowledge check tests them, and Section 2 turns from ratio measures to difference measures.

Key Takeaways

  • Measures of association quantify the strength of the exposure-disease relationship, unlike P-values which reflect sample size.
  • RR compares risks, IR compares incidence rates, and OR compares odds between exposed and non-exposed groups.
  • OR is the only measure that can be computed from case-control studies due to its symmetry property.
  • When disease is rare (<5%), OR ≈ RR. IR values are further from the null than RR, and OR values further still.
Knowledge Check: this section

1. The odds ratio (OR) is the only ratio measure of association applicable to case-control studies because:

The OR exhibits symmetry: (a1×b0)/(a0×b1) is the same regardless of whether you view it as odds of disease or odds of exposure. In case-control studies, the investigator sets the number of cases and controls, making RR incalculable, but OR remains valid.

2. A risk ratio of 0.58 for diarrhea in a cistern study indicates:

An RR of 0.58 means the risk in the exposed group is 58% of the risk in the non-exposed group, which is a 42% reduction (1 − 0.58 = 0.42). Since RR < 1, the exposure (cistern) is protective.

3. Under what condition does OR best approximate RR?

When disease is rare, the number of cases (a) is small relative to the total (n), so odds and risk become approximately equal. Under density sampling, OR estimates IR regardless of rarity.

✦ Pass the knowledge check with 100% to continue

Section 2

Measures of Effect in the Exposed Group

⏱ Estimated reading time: 15 minutes

Section 2 of 4

Measures of Effect in the Exposed Group

Risk difference, attributable fraction in the exposed, and vaccine efficacy as a special case.

Risk difference

RD: the absolute increase in risk

Risk Difference (RD) / Attributable Risk
\[ \color{#0B7B6B}{RD} = \color{#C2410C}{R_{E+}} - \color{#1D4ED8}{R_{E-}} = \frac{\color{#C2410C}{a_1}}{a_1+b_1} - \frac{\color{#1D4ED8}{a_0}}{a_0+b_0} \]
RD risk differenceRE+ risk in the exposedRE− risk in the unexposed
Incidence Rate Difference (ID)
\[ \color{#0B7B6B}{ID} = \frac{\color{#C2410C}{a_1}}{\color{#C2410C}{T_1}} - \frac{\color{#1D4ED8}{a_0}}{\color{#1D4ED8}{T_0}} \]
ID rate differencea1/T1 rate in the exposeda0/T0 rate in the unexposed

Null value: 0. RD > 0 means excess risk in the exposed; RD < 0 means a protective effect.

Worked example

Smoking and low birth weight

Smokers: 40/351 = 0.114
Non-smokers: 331/4,649 = 0.071

RD
\[ RD = 0.114 - 0.071 = 0.043 \]

Per 100 women who smoked, approximately 4.3 additional low-birth-weight babies attributable to smoking.

Why RD matters

A risk ratio of 1.60 sounds substantial. An RD of 4.3 per 100 tells you the actual scale of the problem when planning an intervention.

Attributable fraction

AFe: proportion of exposed cases due to exposure

Attributable Fraction in the Exposed
\[ \color{#0B7B6B}{AF_e} = \frac{\color{#1D4ED8}{RR} - 1}{\color{#1D4ED8}{RR}} = \frac{\color{#C2410C}{RD}}{\color{#6D28D9}{R_{E+}}} \]
AFe attributable fraction in the exposedRR risk ratioRD risk differenceRE+ risk in the exposed

Ranges from 0 (when RR = 1, no excess) to 1 (when all disease in the exposed is due to the exposure).

From the smoking example:

\[ AF_e = \frac{1.60 - 1}{1.60} = \frac{0.60}{1.60} \approx 0.375 \]

37.5% of low-birth-weight cases among smoking women were attributable to their smoking.

Special case

Vaccine efficacy as AFe

Unvaccinated (exposure+): 20% develop disease
Vaccinated (exposure-): 5% develop disease

Vaccine Efficacy
\[ \color{#0B7B6B}{VE} = AF_e = \frac{\color{#C2410C}{RD}}{\color{#6D28D9}{R_{E+}}} = \frac{0.15}{0.20} = 0.75 \]
VE vaccine efficacyRD risk difference (unvaccinated vs vaccinated)RE+ risk in the unvaccinated

Plain language

The vaccine prevented 75% of the cases that would have occurred in vaccinated individuals had they remained unvaccinated.

Carry forward

Relative vs. absolute measures

Ratio measures (section 1)

Relative strength of association. Tend to be stable across different baseline risks. Used in clinical risk communication.

Difference measures (section 2)

Absolute excess risk and proportional causal share. Needed for resource allocation and policy prioritisation.

A later section takes these ideas and asks: what is the impact on the whole population, not just on those who are exposed?

Introduction and Overview

An earlier section covered the three ratio measures (RR, IR, OR), that is, how many times more likely disease is in the exposed group compared to the unexposed. This section turns to the parallel set of difference measures, which answer a different question: not how many times more, but how many extra cases occur because of the exposure. Difference measures lead naturally to attributable fractions in the exposed, which quantify how much disease in the exposed group can be attributed to the exposure itself.

Learning Objectives

  • Distinguish between “ratio” (relative) and “difference” (absolute) measures of association.
  • Calculate and interpret the risk difference (RD) and incidence rate difference (ID).
  • Calculate and interpret the attributable fraction in the exposed (AFe).
  • Explain the concept of vaccine efficacy as a special case of AFe.

The section first contrasts ratio and difference measures, then defines the risk difference and the incidence rate difference, and then turns to the attributable fraction in the exposed and its use as a measure of vaccine efficacy.

Ratio vs. Difference Measures

The ratio measures from an earlier section (RR, IR, OR) tell us the relative strength of association, but they do not indicate the absolute number of cases attributable to the exposure. Difference (absolute effect) measures address this gap by computing how many additional cases occur because of the exposure; the choice between ratio and difference measures is itself a substantive scientific decision rather than a statistical convenience (Greenland & Pearce, 2015; Tripepi et al., 2007).

It helps to name the two scales this contrast lives on. Ratio measures work on a multiplicative scale: a risk ratio of 2 says the exposed risk is the baseline risk multiplied by two. Difference measures work on an additive scale: a risk difference of 0.04 says the exposed risk is the baseline plus four percentage points. The same association can look striking on one scale and slight on the other, so the choice of scale is a real part of the scientific question rather than a matter of presentation.

Box 6.3 explains why both kinds of measure are needed when the public health impact of an exposure is assessed.

Box 6.3: Why Both Matter

Even when an exposure is very strongly associated with disease (high RR), if the exposure is rare in a population, it may contribute very few cases. Conversely, a relatively weak risk factor (modest RR) that is common can be responsible for many cases. Difference measures capture this “public health impact.”

The next part defines the first of the difference measures, the risk difference, together with its counterpart for rate data, the incidence rate difference.

Risk Difference (RD): Attributable Risk

The risk difference (RD), also called attributable risk (Walter, 1976), is simply the risk in the exposed group minus the risk in the non-exposed group:

Risk difference
RD = p(D+|E+) − p(D+|E−) = (a1/n1) − (a0/n0) Eq 6.4
The risk difference is the risk in the exposed group minus the risk in the unexposed group.

Similarly, the incidence rate difference (ID) is the difference between two incidence rates:

Incidence rate difference
ID = (a1/t1) − (a0/t0) Eq 6.5
The rate difference is the rate in the exposed minus the rate in the unexposed, each a count of cases over person-time.

The three tabs below interpret and apply these two formulae. The first tab explains how the sign of a difference measure is read. In the second tab, Worked Example 6.3 applies Equation 6.4 to a cohort of 5,000 women followed through pregnancy, with smoking as the exposure and low birth weight as the outcome. In the third tab, Worked Example 6.4 applies Equation 6.5 to the migraine data of Worked Example 6.2. Calculator 6.4 and Calculator 6.5, which follow the tabs, reproduce the two worked examples so that the effect of changing the counts on each difference can be seen.

Interpretation of Difference Measures

  • RD or ID < 0 → Exposure is protective
  • RD or ID = 0 → No effect of exposure
  • RD or ID > 0 → Exposure is positively associated with disease

RD indicates the increase (or decrease) in the probability of disease in the exposed group, beyond the baseline risk. It tells you: “For every X exposed individuals, how many additional cases occur because of the exposure?”

Worked Example 6.3: Smoking & Low Birth Weight

From a cohort of 5,000 women followed through pregnancy:

SmokerNon-smokerTotal
Low birth weight40331371
Normal birth weight3114,3184,629
Total3514,6495,000
  • Risk in exposed: RE+ = 40/351 = 0.114
  • Risk in non-exposed: RE− = 331/4649 = 0.071
  • RD = 0.114 − 0.071 = 0.043

For every 100 women who smoked, approximately 4.3 had a low-birth-weight baby due to the fact that they smoked (assuming causal relationship).

Worked Example 6.4: Migraine Incidence Rates

Using the migraine data from the previous section, 131 cases occurred among women over 250 person-months and 44 cases among men over 236 person-months.

  • Rate in women: 131/250 = 0.524 cases per person-month.
  • Rate in men: 44/236 = 0.186 cases per person-month.
  • ID = 0.524 − 0.186 = 0.338 cases per person-month.

Women aged 30 to 40 experienced about 34 more cases of migraine per 100 person-months of observation than men of the same age. The rate ratio of 2.81 from the previous section describes the same data on a relative scale.

Worked Example 6.5 takes the reciprocal of the risk difference from Worked Example 6.3 to obtain the number needed to harm, and applies the same calculation to a protective intervention to obtain the number needed to treat.

Worked Example 6.5: Number Needed to Harm and Number Needed to Treat

The reciprocal of a risk difference gives the number of people who must be exposed for one additional case to occur. In the smoking example, RD = 0.043, so the number needed to harm is 1 ÷ 0.043 = 23.3: about 23 women who smoke for each additional low-birth-weight baby, assuming a causal relationship.

For a protective intervention the same calculation gives the number needed to treat. If, for example, an intervention lowered the risk of an outcome from 0.20 to 0.15, RD = −0.05 and NNT = 1 ÷ 0.05 = 20 people treated for one outcome avoided.

Both numbers apply to the follow-up period of the study from which the risks come, here the pregnancy, so each is quoted with its time frame.

The risk difference and its reciprocal describe the excess risk in absolute terms. The next part expresses the same excess as a proportion of the risk in the exposed group, which gives the attributable fraction in the exposed.

Attributable Fraction in the Exposed (AFe)

The AFe (also called the attributable fraction among the exposed) expresses the proportion of disease in exposed individuals that is due to the exposure, assuming the relationship is causal (Walter, 1976). It can be viewed as the proportion of disease in the exposed group that would be avoided if the exposure were removed.

Equation 6.6 gives the AFe in three forms: from the risk difference and the risk in the exposed, from the risk ratio, and approximately from the odds ratio.

Attributable fraction in the exposed
AFe = RD / p(D+|E+) = (RR − 1) / RR ≈ (OR − 1) / OR Eq 6.6
The attributable fraction in the exposed is the risk difference over the risk in the exposed; equivalently it can be written from the risk ratio, or approximated from the odds ratio.

AFe ranges from 0 (where risk is equal, RR = 1) to 1 (where all disease in the exposed group is due to the exposure, RR = ∞). In case-control studies, AFe can be approximated by substituting OR for RR.

Worked Example 6.6 applies Equation 6.6 to the smoking cohort of Worked Example 6.3, and Calculator 6.6 reproduces the calculation from the two risks, from the RR or from the OR.

Worked Example 6.6: AFe for Smoking

From the smoking example above:

  • RR = 0.114 / 0.071 = 1.60
  • AFe = (1.60 − 1) / 1.60 = 0.60 / 1.60 = 0.375 (37.5%)

Among women who smoked, 37.5% of the low-birth-weight cases were attributable to smoking. Alternatively: 0.043 / 0.114 = 0.377 ≈ 37.7% (slight rounding difference).

Vaccine Efficacy

Vaccine efficacy is a special form of AFe, where “not vaccinated” is the exposure (factor positive) and “vaccinated” is the comparison group. Worked Example 6.7 works through a simple case, Equation 6.7 then states the general formula, and Calculator 6.7 reproduces the example. If 20% of unvaccinated individuals develop disease versus 5% of vaccinated individuals:

Worked Example 6.7: Vaccine Efficacy Calculation

  • RD = 0.20 − 0.05 = 0.15
  • AFe = 0.15 / 0.20 = 0.75 (75%)

The vaccine has prevented 75% of the cases of disease that would have occurred in the vaccinated group if the vaccine had not been used.

Written as a formula, with Ru the risk in the unvaccinated group and Rv the risk in the vaccinated group:

Vaccine efficacy
VE = (Ru − Rv) / Ru = 1 − Rv / RuEq 6.7
Vaccine efficacy is the risk in the unvaccinated group minus the risk in the vaccinated group, divided by the risk in the unvaccinated group; equivalently, it is one minus the risk ratio comparing vaccinated with unvaccinated people.

Box 6.4 distinguishes the AFe, which measures the excess fraction of cases in the exposed group, from the etiologic fraction, for which the AFe provides a lower bound.

Box 6.4: AFe vs. Etiologic Fraction

The etiologic fraction is the proportion of cases in the exposed group for which exposure was a component of the sufficient cause (Rothman, 1976). While AFe measures the excess fraction, the etiologic fraction can be higher because exposure may contribute to cases even when the baseline risk would have produced them eventually. In general, AFe provides a lower bound for the etiologic fraction (Greenland & Robins, 1988).

This section has defined the risk difference and the incidence rate difference, the numbers needed to harm and to treat, and the attributable fraction in the exposed, with vaccine efficacy as a special case. The reflection that follows asks for these measures to be computed for a single exposure, and the Key Takeaways and knowledge check consolidate the section. Section 3 moves from the exposed group to the whole population.

Reflection

In a cohort study, a new environmental pollutant is found to have an RR of 3.0 for respiratory disease. The risk of respiratory disease in the non-exposed population is 2%. Calculate the risk in the exposed (RR × risk in the non-exposed), the risk difference (RD = risk in the exposed − risk in the non-exposed), and the attributable fraction in the exposed, AFe = (RR − 1) / RR. If 1,000 people are exposed, how many additional cases would you expect due to the exposure? Discuss why RD and AFe give different but complementary perspectives.

Model answerRD = baseline × (RR−1) = 0.02 × 2 = 0.04 (4 per 100 exposed). AFe = (RR−1)/RR = 2/3 ≈ 0.667, so 67% of disease in exposed people is attributable to the exposure. With 1,000 exposed individuals, expected cases at baseline = 1,000 × 0.02 = 20; with exposure = 1,000 × 0.06 = 60; additional cases due to exposure = 40. RD gives the public-health-relevant absolute number (40 extra cases per 1000 exposed); AFe gives the within-exposed fraction (67% of exposed cases would not have occurred without exposure). They are complementary: RD scales to population impact; AFe addresses individual-level questions like the legal standard "but for the exposure, would this person have gotten sick?"

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • RD (attributable risk) measures the absolute increase in risk due to exposure; the null value is 0.
  • AFe = (RR − 1)/RR gives the proportion of disease in the exposed that is due to the exposure.
  • Vaccine efficacy is a special case of AFe where the “exposure” is being unvaccinated.
  • AFe provides a lower bound for the etiologic fraction.
Knowledge Check: this section

1. If the risk of disease is 12% in the exposed group and 4% in the non-exposed group, the risk difference (RD) is:

RD = 0.12 − 0.04 = 0.08 or 8%. This is the absolute increase in risk attributable to the exposure. (The RR would be 3.0, which is a ratio measure.)

2. A vaccine efficacy of 75% means:

Vaccine efficacy = AFe = (risk in unvaccinated − risk in vaccinated) / risk in unvaccinated. A value of 75% means 75% of cases were prevented by vaccination.

3. AFe = (RR − 1)/RR. If RR = 2.5, what is AFe?

AFe = (2.5 − 1) / 2.5 = 1.5 / 2.5 = 0.60. This means 60% of disease among the exposed is attributable to the exposure.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 3

Population-Level Measures & Study Design

⏱ Estimated reading time: 12 minutes

Section 3 of 4

Population-Level Measures & Study Design

PAR and AFp: scaling attributable risk to the whole population, and matching measures to study designs.

The two population measures

PAR and AFp

Population Attributable Risk (PAR)
\[ \color{#0B7B6B}{PAR} = \color{#C2410C}{p(D+)_{\text{pop}}} - \color{#1D4ED8}{p(D+)_{E-}} \]
PAR population attributable riskp(D+)pop risk in the whole populationp(D+)E− risk in the unexposed
Population Attributable Fraction (Levin, 1953)
\[ \color{#0B7B6B}{AF_p} = \frac{\color{#C2410C}{PAR}}{\color{#1D4ED8}{p(D+)_{\text{pop}}}} = \frac{\color{#6D28D9}{p_e}(\color{#BE185D}{RR}-1)}{\color{#6D28D9}{p_e}(\color{#BE185D}{RR}-1)+1} \]
AFp population attributable fractionPAR population attributable riskp(D+)pop total population riskpe exposure prevalenceRR risk ratio

where \(p_e\) is the prevalence of exposure. Both association strength and exposure prevalence determine the result.

Worked example

Same cohort, population perspective

Overall risk: 371/5,000 = 0.074
Risk in non-smokers: 331/4,649 = 0.071
Exposure prevalence: 351/5,000 = 7%

\[ PAR = 0.074 - 0.071 = 0.003 \]
\[ AF_p = 0.003/0.074 = 4.1\% \]

The key insight

AFe = 37.5% but AFp = 4.1%. A strong association with a rare exposure produces a small population fraction.

A critical contrast

Strong and rare vs. weak and common

High RR, rare exposure

Example: intravenous drug use and HIV. Very strong association, but small AFp because the exposure is uncommon in the full population.

Modest RR, common exposure

Example: poor diet and chronic disease. Weaker association, but large AFp because most of the population is exposed.

Both the strength of association and the prevalence of exposure drive public-health impact.

Study design constraints

Which measures each design delivers

Design Direct measures Constraints Cohort RR, IR, OR, RD, ID,AFe, PAR, AFp All available Case-control OR directly RR, PAR need external data Cross-sectional Prevalence ratio (PR), OR No incidence
Carry forward

Two things into the next section

  • Population impact depends on exposure prevalence, not only on association strength. A common weak factor can outweigh a rare strong one.
  • Study design constrains measures: cohort gives everything; case-control gives OR directly; cross-sectional gives prevalence ratios.

A later section adds the last piece: standard errors and confidence intervals to quantify how precise these estimates really are.

Introduction and Overview

Earlier sections worked at the level of the exposed group. This section zooms out: even if an exposure powerfully causes disease in the exposed, its public-health importance also depends on how common it is in the population. The population attributable fraction (AFp) captures this combination, and the section closes by mapping each measure of association onto the study design that produces it, tying back to earlier lessons.

Learning Objectives

  • Calculate and interpret the population attributable risk (PAR) and population attributable fraction (AFp).
  • Explain how the prevalence of exposure affects population-level measures.
  • Identify which measures of association can be computed from each study design.

The section first defines the population attributable risk and the population attributable fraction, then shows how the fraction is estimated when confounding calls for an adjusted risk ratio, and ends with a summary of the measures that each study design can supply.

From the Exposed Group to the Entire Population

While RD and AFe describe the effect of exposure among exposed individuals, public health decisions often require understanding the impact of an exposure on the entire population. Two key population-level measures address this:

Population Attributable Risk (PAR)

The PAR is the increase in overall population risk attributable to the exposure. It reflects both the strength of the association and the frequency of the exposure in the population, an idea originally developed by Levin (1953) for lung cancer and reviewed by Northridge (1995) as a link between causal inference and public-health action.

Equation 6.8 gives the PAR in three equivalent forms, the last of which shows that it equals the risk difference multiplied by the proportion of the population that is exposed.

Population attributable risk
PAR = p(D+) − p(D+|E−) = (m1/n) − (a0/n0) = RD × p(E+) Eq 6.8
The population attributable risk is the disease risk in the whole population minus the risk in the unexposed.

Population Attributable Fraction (AFp)

The population attributable fraction is the proportion of all disease in the population that is attributable to the exposure. Box 6.5 recalls how HSCI 341 Lesson 1 introduced it, and Equation 6.9 then sets out two ways of computing it.

Box 6.5: Recall: The Population Attributable Fraction

HSCI 341 Lesson 1, Section 3 (Seeking Causes and Models of Causation), introduced the population attributable fraction with Levin's formula (1953), AFp = pe(RR − 1) ÷ [pe(RR − 1) + 1], where pe is the proportion of the population exposed. It is the proportion of disease in the whole population that is attributable to the exposure, and that would be avoided if the exposure were removed, assuming causation and no confounding. HSCI 230 Lesson 5, Section 3 (The Exposure), showed a published estimate of it for alcohol and cancer in the EPIC cohort.

Retrieval question. If 20% of a population is exposed and the risk ratio is 3, what is the population attributable fraction?

Show the answer▼

AFp = 0.20 × 2 ÷ (0.20 × 2 + 1) = 0.40 ÷ 1.40 = 0.29, so about 29% of cases would be attributable to the exposure.

This section adds a second route to the same quantity: AFp is the population attributable risk divided by the total population risk, which links it to Equation 6.8.

Population attributable fraction
AFp = PAR / p(D+) = p(E+)(RR − 1) / [p(E+)(RR − 1) + 1] Eq 6.9
The population attributable fraction is the population attributable risk as a share of the total population risk; it rises with both the exposure prevalence and the risk ratio.

The second form of Equation 6.9 contains both the exposure prevalence p(E+) and the risk ratio, and Box 6.6 explains why the prevalence of exposure shapes the population impact of a risk factor.

Box 6.6: Why Exposure Prevalence Matters

A strong risk factor (high RR) that is rare in the population will have a small AFp. A weaker risk factor (modest RR) that is common may have a large AFp. For example, intravenous drug use has a very high RR for HIV, but if it is rare in the population, eliminating it would prevent few total cases. A modestly elevated risk factor like poor diet, affecting millions, may account for more total cases.

Worked Example 6.8 returns to the smoking and low birth weight cohort of Worked Example 6.3, computes the PAR with Equation 6.8, and divides it by the population risk, as in the first form of Equation 6.9. Calculator 6.8 and Calculator 6.9, which follow the example, reproduce the two results, and Calculator 6.9 also offers presets that compare a common, weak exposure with a rare, strong one.

Worked Example 6.8: Smoking & Low Birth Weight (Population Level)

From the cohort of 5,000 women (351 smokers, 4,649 non-smokers):

  • Overall risk: p(D+) = 371/5000 = 0.074
  • Risk in non-exposed: 331/4649 = 0.071
  • PAR = 0.074 − 0.071 = 0.003
  • AFp = 0.003 / 0.074 = 0.041 (4.1%)

Only 4.1% of all low-birth-weight babies in the population were attributable to smoking. The low AFp is because very few women (351/5000 = 7%) smoked during the 2nd trimester, despite the relatively strong association (RR = 1.60).

Worked Example 6.8 used the unadjusted risks of a single cohort, and the AFp it gives assumes that the association is not confounded. The next part gives the form of the AFp that is used with a confounding-adjusted risk ratio.

Confounding and AFp

If confounding is present, adjusted estimates of RR should be used. The AFp can then be estimated using:

Population attributable fraction (adjusted)
AFp = pd × (aRR − 1) / aRR Eq 6.10
The population attributable fraction can be estimated from the proportion of cases exposed and the confounding-adjusted risk ratio.

where pd is the proportion of cases exposed to the risk factor, and aRR is the adjusted risk ratio. For multiple exposure categories, a summation formula is used.

Worked Example 6.9 applies Equation 6.10 to a hypothetical cohort study of physical inactivity and type 2 diabetes, and Calculator 6.10 reproduces the example.

Worked Example 6.9: AFp with an Adjusted Risk Ratio

Suppose a cohort study of type 2 diabetes finds that 45% of the people who developed diabetes were physically inactive at baseline (pd = 0.45), and that the risk ratio for inactivity, adjusted for age and body mass index, is aRR = 1.6 (hypothetical values).

AFp = 0.45 × (1.6 − 1) / 1.6 = 0.45 × 0.375 = 0.169, or about 17%.

About 17% of the diabetes cases in this population would be attributable to physical inactivity, assuming a causal relationship and that the adjustment has removed the confounding. The factor (aRR − 1)/aRR is the attributable fraction among the exposed, and multiplying it by the proportion of cases who were exposed scales it to the whole population.

The measures in this section and the two before it differ in the data they require. The next part sets out which of them each study design can supply.

Study Design and Measures of Association

Not all measures can be computed from all study designs. Table 6.3 summarises which measures are available:

Table 6.3. Measures of association that can be computed from cross-sectional, cohort and case-control studies.

MeasureCross-sectionalCohortCase-control
RR✓✓
IR✓
OR✓✓✓
RD✓✓
AFe✓✓✓b
PAR✓✓a
AFp✓✓a✓c

a Requires independent estimate of p(D+) or p(E+). b Estimated using OR. c Requires OR and independent estimate of p(E+|D+).

This section has extended the measures of Section 2 from the exposed group to the whole population, shown how a confounding-adjusted risk ratio enters the AFp, and matched each measure to the designs that can supply it. The reflection that follows applies Equation 6.9 to two risk factors that differ in strength and prevalence, and the Key Takeaways and knowledge check consolidate the section. Section 4 turns to the precision of these estimates.

Reflection

Consider two risk factors for a disease: Factor A has RR = 5.0 and affects 2% of the population. Factor B has RR = 1.5 and affects 40% of the population. Calculate AFp for each factor using the formula AFp = p(E+)(RR − 1) / [p(E+)(RR − 1) + 1]. Which factor would you prioritise in a public health intervention, and why?

Model answerFactor A: AFp = 0.02(4)/(0.02(4)+1) = 0.08/1.08 = 0.074 (7.4%). Factor B: AFp = 0.40(0.5)/(0.40(0.5)+1) = 0.20/1.20 = 0.167 (16.7%). Despite a much smaller per-person RR, Factor B prevents over twice as many population cases because it is so much more common. Prioritise B for a population-level public-health intervention, because the population attributable fraction is what determines burden averted. A nuance: B's lower per-person effect may mean the intervention is harder to deliver, less attractive to individuals (low perceived risk), and politically tougher; A's larger per-person effect may justify a targeted (high-risk) intervention even though its population impact is smaller. The best portfolio often combines both: high-risk strategies for A and population-wide strategies for B.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • PAR = overall population risk − risk in unexposed; it reflects both strength and prevalence of exposure.
  • AFp = PAR / p(D+); it gives the proportion of all disease in the population attributable to the exposure.
  • A common risk factor with modest RR can have a larger AFp than a rare factor with high RR.
  • Different study designs support different measures: only OR is available from case-control studies.
Knowledge Check: this section

1. A risk factor has RR = 4.0 but affects only 1% of the population. The AFp is:

AFp = p(E+)(RR−1) / [p(E+)(RR−1)+1] = 0.01×3 / (0.01×3+1) = 0.03/1.03 = 0.029 or about 2.9%. Despite the strong association, the low prevalence of exposure limits the population impact.

2. Which measure cannot be computed directly from a case-control study?

RD requires actual disease risks in the exposed and non-exposed groups. In case-control studies, these risks cannot be computed because the investigator determines the ratio of cases to controls.

3. PAR differs from RD in that:

PAR = p(D+) − p(D+|E−). It is the overall population-level risk increase attributable to the exposure, incorporating both the strength of association and the prevalence of exposure. RD only measures the difference between exposed and non-exposed groups.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 4

Hypothesis Testing & Confidence Intervals

⏱ Estimated reading time: 15 minutes

Section 4 of 4

Hypothesis Testing & Confidence Intervals

Standard errors, P-values, confidence intervals, and a guide to choosing the right statistical test.

Standard error

Precision of a point estimate

Difference measures (RD, ID)

Variance computed directly from cell counts.

Var(RD)
\[ \frac{\color{#C2410C}{R_{E+}}(1-\color{#C2410C}{R_{E+}})}{\color{#C2410C}{n_1}} + \frac{\color{#1D4ED8}{R_{E-}}(1-\color{#1D4ED8}{R_{E-}})}{\color{#1D4ED8}{n_0}} \]
RE+, n1 risk and size, exposed groupRE−, n0 risk and size, unexposed group

Ratio measures (RR, IR, OR)

Variance on the log scale using Taylor series approximations.

Var(ln OR)
\[ \frac{1}{\color{#C2410C}{a_1}} + \frac{1}{\color{#BE185D}{b_1}} + \frac{1}{\color{#1D4ED8}{a_0}} + \frac{1}{\color{#6D28D9}{b_0}} \]
a1 exposed casesb1 exposed non-casesa0 unexposed casesb0 unexposed non-cases
Confidence intervals

Uncertainty around the point estimate

CI for difference measures (symmetric)
\[ \color{#0B7B6B}{\hat{\theta}} \pm \color{#1D4ED8}{z_{\alpha/2}} \cdot \color{#C2410C}{SE(\hat{\theta})} \]
θ̂ point estimatezα/2 critical valueSE standard error
CI for ratio measures (log scale, then exponentiate)
\[ \exp\!\left[\color{#0B7B6B}{\ln\hat{\theta}} \pm \color{#1D4ED8}{z_{\alpha/2}} \cdot \color{#C2410C}{SE(\ln\hat{\theta})}\right] \]
ln θ̂ estimate on the log scalezα/2 critical valueSE standard error of the log estimate

Significance criterion: the CI for OR, RR, or IR must exclude 1; the CI for RD or ID must exclude 0.

From the lesson's examples

Computed confidence intervals

Measure Point estimate 95% CI RD (smoking)0.043(0.009, 0.077) RR (smoking)1.60(1.174, 2.182) OR (smoking)1.68(1.154, 2.387) ID (migraine)0.338(0.232, 0.443) IR (migraine)2.81(1.983, 4.050)
Forest plot of the three ratio-measure confidence intervals from the lesson examples, on a log axis, all sitting to the right of the null value of 1.
The three ratio measures plotted on a log axis. Each interval lies entirely to the right of the null (1), so each is significant at the 5% level; the differing interval widths show how much more the data constrain some estimates than others.
Four test statistics

Tests for 2x2 tables and regression

Pearson chi-squared

For categorical outcomes with expected cell counts all above 5. Compares observed to expected counts under independence.

Fisher's exact

Small expected counts (below 5). Exact P-value, no large-sample assumption. Often reported with the OR and its exact CI.

Wald test

Estimate divided by its standard error. Fast, but less accurate in small samples.

Likelihood ratio test

Generally preferred in regression settings; more accurate than Wald in small samples.

Choosing the right test

A decision framework

Situation Parametric Non-parametric Continuous, 2 indep. groupsTwo-sample t-testMann-Whitney U Continuous, 2 paired groupsPaired t-testWilcoxon signed-rank Continuous, 3+ groupsOne-way ANOVAKruskal-Wallis Categorical, independentPearson chi-sq / Fisher-- Categorical, pairedMcNemar's test-- Two continuous variablesPearson rSpearman rho
Carry forward

Point estimates need uncertainty bands

  • Standard error quantifies precision; computed differently for difference vs. ratio measures.
  • Confidence intervals show the range of plausible effects, not just a binary signal.
  • P-values are not the only verdict: a P-value of 0.049 and one of 0.051 carry almost identical information.

The final assessment is just below. Take the synthesis materials slowly before attempting the fifteen questions.

Introduction and Overview

Earlier sections produced point estimates of association: single numbers like RR = 2.5 or AFp = 30%. Those numbers are useless without a quantification of how uncertain they are. This section closes the lesson by introducing the standard error, hypothesis tests, and confidence intervals, the same statistical-inference machinery you previewed in an earlier course, now applied directly to the measures you just learned to compute.

Learning Objectives

  • Explain the concepts of standard error, null hypothesis, and P-value.
  • Describe the four common test statistics for evaluating associations.
  • Interpret confidence intervals for measures of association.
  • Distinguish between statistical significance and the strength of association.

The section first defines the standard error of each measure, then uses it in hypothesis tests and confidence intervals, and ends with a guide to choosing a statistical test for other kinds of outcome and comparison.

Standard Error

The standard error (SE) provides a measure of the precision of a point estimate, that is, how much uncertainty exists in the estimate. This part gives the variance of the risk difference (Equation 6.11) and of the log risk ratio and the log odds ratio (Equation 6.12 and Equation 6.13), each with a worked example on the smoking and low birth weight cohort and a calculator. For difference measures (RD, ID), the variance can be computed directly:

Variance of the risk difference
var(RD) = [(a1/n1)(1 − a1/n1)] / n1 + [(a0/n0)(1 − a0/n0)] / n0 Eq 6.11
The variance of the risk difference adds a contribution from the exposed group and one from the unexposed group, each computed straight from the cell counts.

Worked Example 6.10 applies Equation 6.11 to the smoking cohort, and Calculator 6.11 reproduces the calculation so that the contributions of the two groups to the variance can be compared.

Worked Example 6.10: Variance and Standard Error of the RD

For the smoking and low birth weight cohort, a1/n1 = 40/351 = 0.1140 and a0/n0 = 331/4,649 = 0.0712.

  • Exposed contribution: (0.1140 × 0.8860) / 351 = 0.000288.
  • Unexposed contribution: (0.0712 × 0.9288) / 4,649 = 0.0000142.
  • var(RD) = 0.000288 + 0.0000142 = 0.000302, so SE(RD) = √0.000302 = 0.0174.

Almost all of the variance comes from the exposed group, because it contains only 351 women. The estimate would become more precise mainly through a larger exposed group.

For ratio measures (RR, IR, OR), the variance is computed on the log scale using Taylor series approximations:

The first tab gives the variance of ln(RR) (Equation 6.12), which Worked Example 6.11 applies to the smoking cohort and Calculator 6.12 reproduces. The second tab gives the variance of ln(OR) (Equation 6.13), which Worked Example 6.12 applies to the same cohort and Calculator 6.13 reproduces.

Variance of ln(RR)
var(ln RR) = 1/a1 − 1/n1 + 1/a0 − 1/n0 Eq 6.12
The variance of the log risk ratio is built from the diseased counts and group totals of the exposed group and the unexposed group; the variance is calculated for ln(RR) because its sampling distribution is approximately symmetric, whereas that of RR is skewed.

Worked Example 6.11: var(ln RR)

For the smoking cohort (a1 = 40, n1 = 351, a0 = 331, n0 = 4,649):

var(ln RR) = 1/40 − 1/351 + 1/331 − 1/4,649 = 0.025000 − 0.002849 + 0.003021 − 0.000215 = 0.024957, so SE(ln RR) = √0.024957 = 0.158.

The term 1/40 is by far the largest, because the exposed group has only 40 cases.

Variance of ln(OR)
var(ln OR) = 1/a1 + 1/a0 + 1/b1 + 1/b0 Eq 6.13
The variance of the log odds ratio is the sum of the reciprocals of all four cells of the table, two from the exposed column and two from the unexposed column.

Worked Example 6.12: var(ln OR)

For the smoking cohort (a1 = 40, a0 = 331, b1 = 311, b0 = 4,318):

var(ln OR) = 1/40 + 1/331 + 1/311 + 1/4,318 = 0.025000 + 0.003021 + 0.003215 + 0.000232 = 0.031468, so SE(ln OR) = √0.031468 = 0.177.

The variance of ln(OR) is larger than that of ln(RR) for the same data, so the odds ratio is estimated slightly less precisely.

The standard errors computed in this part are used in the Wald statistic and in the confidence intervals later in this section. The next part turns to hypothesis testing.

Hypothesis Testing

A hypothesis test asks how likely the observed data would be if there were no association between exposure and disease. Box 6.7 recalls how HSCI 230 defined the p-value and separated statistical significance from practical significance, and the paragraphs that follow it state the null hypotheses for the measures in this lesson.

Box 6.7: Recall: P-values and Practical Significance

The glossary of HSCI 230 Lesson 11 defined the p-value as the probability, assuming the null hypothesis is true, of observing data at least as extreme as those obtained. It says nothing directly about the probability that the null hypothesis is true or about the size of an effect. HSCI 230 Lesson 12, Section 2 (Stepwise Critical Appraisal), separated statistical significance from practical, clinical or public health significance: a very large study can detect, with a small p-value, an effect too small to change any decision.

Retrieval question. A study of 200,000 people finds a risk ratio of 1.03 with p = 0.001. Is the association statistically significant, and is it practically important?

Show the answer▼

It is statistically significant at the 5% level, because p is below 0.05. A 3% relative increase is small, so its practical importance depends on how common the outcome and the exposure are; the small p-value reflects the precision of a very large study, and the risk ratio with its confidence interval carries the practical message.

For a measure of association, the null hypothesis states that there is no association:

  • For difference measures (RD, ID): H0: θ = 0
  • For ratio measures (RR, IR, OR): H0: θ = 1

An alternative hypothesis can be 1-tailed or 2-tailed. In general, 2-tailed hypotheses are preferred because 1-tailed hypotheses are harder to justify.

Box 6.8 describes what is lost when a P-value is reduced to a verdict of significant or non-significant.

Box 6.8: Limitations of P-values

P-values are often dichotomised into “significant” or “non-significant” at α = 0.05, but this entails a huge loss of information (Wasserstein & Lazar, 2016; Greenland et al., 2016). A P-value of 0.049 and 0.051 lead to different conclusions despite being virtually identical. Always report the actual P-value and a confidence interval, which conveys both significance and precision.

Test Statistics

Four test statistics are in common use for measures of association: the Pearson χ², exact tests, the Wald statistic and the likelihood ratio test. The flip cards below describe each statistic and the conditions under which it applies.

Click each card to explore:

Pearson χ²Click to explore
Exact TestsClick to explore
Wald StatisticClick to explore
Likelihood Ratio TestClick to explore

The Pearson χ² and the Wald statistic can both be calculated by hand. The tabs below give each formula with a worked example and a calculator.

The first tab gives the Pearson χ² statistic (Equation 6.14). Worked Example 6.13 applies it to the smoking and low birth weight table, and Calculator 6.14 reproduces the example. The second tab gives the Wald statistic (Equation 6.15), which Worked Example 6.14 applies to the risk difference and the risk ratio of the same cohort, using the standard errors from Worked Example 6.10 and Worked Example 6.11, and which Calculator 6.15 reproduces.

Pearson χ² statistic
χ² = Σ (O − E)² / E,  with E = (row total × column total) / grand totalEq 6.14
The Pearson χ² statistic sums, over the four cells, the squared difference between the observed count and the expected count, divided by the expected count; each expected count is the row total times the column total divided by the grand total.

Worked Example 6.13: Pearson χ²

For the smoking and low birth weight table, the expected number of low-birth-weight babies among smokers, if smoking and birth weight were unrelated, is (371 × 351) / 5,000 = 26.0442, compared with 40 observed. The four expected counts are 26.0442 (a1), 344.9558 (a0), 324.9558 (b1), and 4,304.0442 (b0).

χ² = (40 − 26.0442)²/26.0442 + (331 − 344.9558)²/344.9558 + (311 − 324.9558)²/324.9558 + (4,318 − 4,304.0442)²/4,304.0442 = 7.478 + 0.565 + 0.599 + 0.045 = 8.69.

With 1 degree of freedom, χ² = 8.69 corresponds to P = 0.003, so an association this strong would be unlikely if smoking and low birth weight were unrelated.

Wald statistic
ZWald = (θ − θ0) / SE(θ)Eq 6.15
The Wald statistic is the distance between the point estimate and its null value, measured in standard errors; for ratio measures, θ and θ0 are replaced by their natural logarithms and the standard error is that of ln θ.

Worked Example 6.14: Wald Statistic

For the risk difference in the smoking cohort, RD = 0.0428 and SE(RD) = 0.0174. Under the null hypothesis θ0 = 0, Z = (0.0428 − 0) / 0.0174 = 2.46, which corresponds to a two-sided P-value of 0.014.

For the risk ratio, the statistic is computed on the log scale with θ0 = 1: Z = (ln 1.6006 − ln 1) / 0.158 = 0.4704 / 0.158 = 2.98, with a two-sided P-value of 0.003.

Box 6.8 noted that a confidence interval conveys both significance and precision. The next part defines the confidence interval and shows how it is computed for the difference and ratio measures of this lesson.

Confidence Intervals

A confidence interval (CI) reflects the level of uncertainty in a point estimate. A 95% CI means that if the study were repeated many times under identical conditions, 95% of the computed CIs would contain the true parameter value. This is a property of the procedure, not a probability statement about the parameter (Greenland et al., 2016).

Computing CIs

The method of computing a confidence interval depends on whether the measure is a difference or a ratio. Equation 6.16 gives the interval for a difference measure and Equation 6.17 the interval for a ratio measure, and each is followed by a worked example on the smoking cohort and a calculator.

For difference measures, the CI is computed directly:

Confidence interval, difference measures
θ ± Zα × √var(θ) Eq 6.16
A symmetric interval places the point estimate at the centre, then steps out by a critical value times the standard error.

Worked Example 6.15 applies Equation 6.16 to the risk difference, and Calculator 6.16 reproduces it.

Worked Example 6.15: CI for the Risk Difference

For the smoking cohort, RD = 0.0428 and var(RD) = 0.000302, so SE(RD) = √0.000302 = 0.0174.

0.0428 ± 1.96 × 0.0174 = 0.0428 ± 0.0341, giving a 95% CI from 0.0087 to 0.0769 (0.009 to 0.077 after rounding, as in Table 6.4).

The interval excludes 0. It indicates that smoking is compatible with anywhere from about 1 to 8 additional low-birth-weight babies per 100 women who smoke.

For ratio measures, the CI is computed on the log scale and then exponentiated:

Confidence interval, ratio measures
θ × exp(± Zα × √var(ln θ)) Eq 6.17
For a ratio measure the interval is built on the log scale, then converted back by exponentiating, so it ends up asymmetric around the point estimate.

The CI is symmetrical about lnθ but not about θ itself, which is why confidence intervals for ratio measures appear asymmetric.

Worked Example 6.16 applies Equation 6.17 to the risk ratio and compares the result with the interval for the odds ratio, and Calculator 6.17 reproduces both intervals.

Worked Example 6.16: CI for the Risk Ratio

For the smoking cohort, RR = 1.6006 and var(ln RR) = 0.024957, so SE(ln RR) = 0.15798 and the margin on the log scale is 1.96 × 0.15798 = 0.30964.

  • Lower limit = 1.6006 × e−0.30964 = 1.6006 × 0.73371 = 1.174.
  • Upper limit = 1.6006 × e0.30964 = 1.6006 × 1.36293 = 2.182.

The interval runs from about 0.43 below the estimate to about 0.58 above it, the asymmetry described above. Applying the same formula to the odds ratio (OR = 1.678, var(ln OR) = 0.031468) gives 1.185 to 2.375. Table 6.4 lists 1.154 to 2.387 for the OR because the textbook computed that interval by a different method, and small differences of this kind between interval methods are expected.

Interpreting CIs

A computed interval can be read in two ways: as a test of the null value and as a range of plausible effect sizes. Box 6.9 sets out both readings for ratio and difference measures, and Box 6.10 applies the first of them to the intervals for the smoking and migraine examples, which Table 6.4 lists.

Box 6.9: Interpreting CIs for Measures of Association

  • For RR, IR, OR: if the 95% CI includes 1, the association is not statistically significant at α = 0.05.
  • For RD, ID: if the 95% CI includes 0, the association is not statistically significant.

However, this “surrogate significance test” is an under-use of the CI. The CI also shows the range of plausible effect sizes, which is far more informative than a binary significant/non-significant classification.

Box 6.10: Example CIs from the Textbook

Table 6.4. Point estimates and 95% confidence intervals for the smoking and migraine examples, as reported in the textbook.

MeasurePoint Estimate95% CI
RD (smoking)0.043(0.009, 0.077)
RR (smoking)1.601(1.174, 2.182)
OR (smoking)1.678(1.154, 2.387)
ID (migraine)0.338(0.232, 0.443)
IR (migraine)2.811(1.983, 4.050)

None of the CIs for ratio measures include 1, and none for difference measures include 0, confirming statistical significance for all associations.

The next part turns from the measures of association in this lesson to the choice of a statistical test for other kinds of outcome and comparison.

Choosing the Right Statistical Test

The Pearson χ², Fisher’s exact, Wald, and likelihood ratio tests above are the workhorses for 2×2 tables and the regression-based measures of association you will meet in a later course. But epidemiological analyses often require comparing means, proportions, or whole distributions across groups, sometimes paired, sometimes not, sometimes badly skewed. The right test depends on three structural questions about your data:

  • Outcome type: continuous (means / medians) or categorical (counts / proportions)?
  • Group structure: one group, two groups, or three or more? Independent or paired/matched?
  • Distributional assumptions: can you defend approximate normality (or appeal to a large-sample CLT argument), or do you need a non-parametric (rank-based) alternative?

The matrix in Table 6.5 answers those three questions together. Find the row whose outcome type and comparison structure match your data, then read across to the test. The parametric column assumes approximate normality (or a large enough sample to appeal to the central limit theorem), and the rank-based column is the fallback when that assumption cannot be defended, when the sample is small, or when the outcome is ordinal and the spacing between its categories is not meaningful.

Table 6.5. Choosing a statistical test by outcome type, comparison structure and distributional assumptions.

Outcome variableComparisonParametric / standard testRank-based or exact alternativeWhat to report alongside the p-value
ContinuousOne group against a hypothesised valueOne-sample t-testWilcoxon signed-rankMean (or median) difference with 95% CI
ContinuousTwo independent groupsTwo-sample t-test (usually the Welch version)Mann–Whitney U (Wilcoxon rank-sum)Difference in means with 95% CI
ContinuousTwo paired observationsPaired t-testWilcoxon signed-rank on the differencesMean within-pair change with 95% CI
ContinuousThree or more independent groupsOne-way ANOVAKruskal–WallisGroup means, then pairwise contrasts with a multiplicity adjustment
ContinuousOrdered exposure categories (dose, quartile, stage)Linear contrast within ANOVA, or regression on the category scoreJonckheere–Terpstra trend testChange in mean per category step with 95% CI
OrdinalTwo independent groupsNot recommended; the numeric codes do not carry equal spacingMann–Whitney UMedian and IQR by group, or the proportional-odds OR from an ordinal model
OrdinalThree or more independent groupsNot recommended, as aboveKruskal–WallisMedian and IQR by group
OrdinalTwo paired observationsNot recommended, as aboveWilcoxon signed-rankMedian of the within-pair differences
BinaryTwo independent groupsPearson χ²Fisher’s exact test when any expected count is below 5Risk difference, risk ratio, or odds ratio with 95% CI
BinaryPaired or matchedMcNemar’s χ² on the discordant pairsExact binomial on the discordant pairsMatched odds ratio b/c with 95% CI
BinaryOrdered exposure categoriesCochran–Armitage test for trendExact trend testOdds ratio per category step with 95% CI
Nominal, 3+ categoriesAny categorical exposurePearson χ² on the r × c tableFisher–Freeman–Halton exact testCramér’s V, or category-specific odds ratios against a reference
Two continuousAssociation between themPearson r (linear)Spearman ρ (monotonic)Coefficient with 95% CI
Ordinal with ordinal or continuousAssociation between themNot recommended, as aboveSpearman ρ, or Kendall’s τb when ties are commonCoefficient with 95% CI

Box 6.11 lists three judgements that remain with the analyst after Table 6.5 has been consulted.

Box 6.11: Three Things the Matrix Does Not Decide for You

  • The measurement scale is a claim the analyst makes about the data, and storing a variable as a number does not settle it. A 1-to-5 satisfaction item stored as a number is ordinal, and averaging it assumes the step from 1 to 2 is the same size as the step from 4 to 5.
  • A non-significant test is not evidence that no association exists. Every row of the matrix names an effect estimate for this reason, because the estimate and its confidence interval answer the question the study actually asked.
  • Rank-based tests answer a different question from their parametric counterparts. The Mann–Whitney U compares whole distributions, and it can be read as a comparison of medians only when the two distributions have a similar shape.

Comparison Table: When Used, How Calculated, How Interpreted

Table 6.6 summarises the tests most commonly reported alongside measures of association.

Table 6.6. When each test is used, how it is calculated and how it is interpreted.

TestWhen to useHow calculatedHow to interpret
One-sample t-test Compare a single sample mean to a known/hypothesised value μ0; continuous, approximately normal (or n large). t = (x̄ − μ0) / (s/√n); df = n−1. If p < α (or 95% CI for the mean excludes μ0), the population mean differs from μ0.
Two-sample (independent) t-test Compare means of 2 independent groups; continuous, approximately normal. Welch’s version does not assume equal variances. t = (x̄1 − x̄2) / SEdiff; df = n1+n2−2 (Student) or Welch–Satterthwaite df (Welch). p < α → the two group means differ. Always report the mean difference and its 95% CI.
Paired t-test Two related observations on the same unit (before/after, twin pairs, matched cases-controls); continuous differences approximately normal. Compute within-pair differences di; t = d̄ / (sd/√n); df = n−1. p < α → mean within-pair change ≠ 0. Reduces between-subject variability, and is usually more powerful than treating data as unpaired.
One-way ANOVA Compare means across 3+ independent groups; continuous, approximately normal, roughly equal variances. F = MSbetween / MSwithin; df1 = k−1, df2 = N−k. Significant F → at least one group mean differs. Follow with post-hoc pairwise comparisons (Tukey HSD, Bonferroni) to see which.
Pearson χ² Test independence of two categorical variables (any r×c table). Assumption: all expected counts > 1 and ≥ 80% > 5. χ² = Σ(O−E)²/E; df = (r−1)(c−1). Expected = (row total × column total) / N. p < α → row and column variables are associated. The test signals whether there is association; report a measure of association (OR, RR) for strength.
Fisher’s exact test 2×2 (or larger) tables with small expected counts where the χ² approximation is suspect. Hypergeometric: enumerates every table with the same margins, sums probabilities of tables as extreme or more extreme than observed. Exact p-value, with no large-sample assumption. With modern computing, fine to use even when χ² would also be valid.
McNemar’s test Paired/matched binary outcomes (before/after on the same person; matched case-control on exposure; agreement of two diagnostic tests). Look only at discordant pairs b and c: χ² = (b−c)² / (b+c); df = 1. Concordant pairs ignored. p < α → the discordant pairs are unbalanced, so a real change/effect exists. The matched OR is b/c.
Wilcoxon signed-rank Non-parametric alternative to one-sample / paired t-test. Use when differences are skewed, ordinal, or have outliers. Rank |di|, attach signs, sum positive (or negative) ranks; compare to its null distribution (or large-sample z). p < α → the median difference is non-zero (under symmetry). Robust to outliers.
Mann–Whitney U (Wilcoxon rank-sum) Non-parametric alternative to two-sample t-test. Two independent groups, continuous or ordinal. Pool all observations, rank them, sum ranks in one group; U = R1 − n1(n1+1)/2. p < α → the two distributions differ. If shapes are similar, this is a test of medians; otherwise it tests stochastic dominance.
Kruskal–Wallis Non-parametric alternative to one-way ANOVA. 3+ independent groups; continuous or ordinal. H from rank sums; approximately χ²k−1 under H0. p < α → at least one group’s distribution differs. Follow with pairwise rank-sum tests (Dunn’s test or pairwise Wilcoxon with adjustment).
Pearson correlation (r) Linear association between two continuous variables; assumes approximate bivariate normality, sensitive to outliers. r = Σ(x−x̄)(y−ȳ) / √[Σ(x−x̄)² Σ(y−ȳ)²]; tested via t = r√[(n−2)/(1−r²)]. Range −1 to +1: sign = direction, magnitude = strength of linear association. Always inspect a scatterplot first.
Spearman ρ Monotonic (not necessarily linear) association between two ordinal or non-normally distributed continuous variables. Pearson r applied to the ranks of x and y. Same −1 to +1 interpretation, but for monotonic association. Robust to outliers and non-linearity.

Table 6.6 recommends pairwise comparisons after a significant one-way ANOVA or Kruskal–Wallis test, and Box 6.12 recalls how the significance threshold is adjusted for such multiple comparisons.

Box 6.12: Recall: Multiple Comparisons

HSCI 230 Lesson 5, Section 7 (Randomised Controlled Trials: Conduct, Analysis, Vaccine Trials and Reporting), introduced the Bonferroni correction for multiple comparisons: when k comparisons are tested, each is tested at α ÷ k, so that the chance of at least one false positive across all of them stays at or below α. After a significant one-way ANOVA the comparisons are usually all pairs of groups. The Holm procedure is a less conservative alternative to Bonferroni: it orders the p-values from smallest to largest, tests them in turn against α ÷ k, α ÷ (k − 1) and so on, and stops at the first one that is not significant. Tukey's honestly significant difference is the usual method for all pairwise comparisons of means after ANOVA. A single primary outcome and comparison, named before the data are analyzed, avoids most multiplicity problems.

Retrieval question. A study makes four comparisons with an overall α of 0.05. What threshold does the Bonferroni correction set for each comparison?

Show the answer▼

0.05 ÷ 4 = 0.0125.

Box 6.13 closes the part with a rule that applies to every test in Table 6.5 and Table 6.6: each test is reported together with an effect estimate and its confidence interval.

Box 6.13: Pair every test with an effect estimate

None of the tests above are themselves measures of effect size. A statistically significant chi-square confirms that some association exists, but the magnitude must come from the OR, RR, mean difference, or correlation coefficient. Likewise, a non-significant t-test in a small study is not evidence of no effect. Always pair every test with the corresponding effect estimate and its 95% CI.

This section has defined the standard error of each measure, set out the test statistics in common use for measures of association, used the standard error in the Wald statistic and in confidence intervals, and shown how a test is chosen for other kinds of outcome and comparison. The reflection that follows compares two studies by strength of association, statistical significance and precision, and the Key Takeaways and knowledge check consolidate the section.

Reflection

A study reports an OR of 1.45 with a 95% CI of (0.92, 2.28). A second study reports an OR of 1.15 with a 95% CI of (1.02, 1.30). A confidence interval that includes 1.0 indicates that the association is not statistically significant at the 5% level, and a narrower interval indicates a more precise estimate. Compare these two findings in terms of: (a) strength of association, (b) statistical significance, and (c) precision. Which finding might be more concerning from a public health perspective, and why?

Model answerStrength: Study 1 reports a larger point estimate (OR 1.45 vs. 1.15), suggesting a stronger per-person association if both are correct. Statistical significance: Study 1's CI (0.92, 2.28) includes 1, so it is non-significant; Study 2's CI (1.02, 1.30) excludes 1, so it is significant. Precision: Study 2 is much more precise, with CI width 0.28 vs. Study 1's 1.36, reflecting larger sample size and/or better measurement. Public-health concern: Study 2 is more concerning for action because the effect, though smaller, is reliably non-zero and applies (likely) to a much larger population; small effects in big populations produce many cases. Study 1's effect, if real, is larger per person but the data don't yet exclude no effect. The decision depends on the absolute baseline risk and population size, but precision and population scope usually win over magnitude alone.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • Standard errors quantify precision; they are computed differently for difference vs. ratio measures.
  • Hypothesis testing uses a null hypothesis (no effect) and a test statistic to generate a P-value.
  • Four common test statistics for measures of association: Pearson χ², exact tests, Wald, and likelihood ratio tests.
  • Beyond 2×2 tables, choose tests by outcome type, group structure, and distribution: t-tests / ANOVA for normal continuous data, χ² / Fisher / McNemar for categorical, and Wilcoxon / Mann–Whitney / Kruskal–Wallis / Spearman as non-parametric alternatives.
  • Confidence intervals are more informative than P-values: they show the range of plausible effect sizes.
  • For ratio measures, the CI containing 1 (or 0 for differences) indicates non-significance at the corresponding α level.
Knowledge Check: this section

1. A 95% confidence interval for an odds ratio is (1.2, 3.8). This means:

Since the CI does not include 1 (the null value for ratio measures), the association is statistically significant. The interval 1.2 to 3.8 represents the range of plausible values for the true OR. Note: the CI is a property of the procedure, not a probability statement about the parameter.

2. Why are confidence intervals for ratio measures (like OR) asymmetric around the point estimate?

The variance of ratio measures is computed on the log scale, where the CI is symmetric around ln(θ). When exponentiated back to the original scale, the CI becomes asymmetric around θ.

3. Which test statistic is generally considered superior in regression settings?

Likelihood ratio tests are generally superior to Wald tests, especially in regression settings. They compare the likelihood of the data under the estimated parameters versus the null hypothesis parameters.

✦ Pass the knowledge check with 100% and complete the reflection to continue

Section 5

Final Review & Assessment

⏱ Estimated time: 20 minutes

Bringing It All Together

This lesson translated the disease-frequency measures of an earlier lesson into measures of association between an exposure and an outcome. You worked through the three ratio measures (RR, IR, OR), the difference measures of effect in the exposed (RD, AFe), the population-level extensions (PAR, AFp), and the inferential machinery (standard errors, hypothesis tests, confidence intervals) that surrounds every estimate.

Lesson 7, Designing Against Bias, uses these measures to examine validity and confounding in study protocols, and Lesson 8, Time-to-Event Data, applies them to survival curves and hazard ratios. The measures you computed here will reappear in each of those lessons as the outputs that designs deliver and that systematic reviews eventually pool. For deeper treatment of these measures and their statistical foundations, see Greenland and Pearce (2015) and the historical landmark papers by Cornfield (1951) and Doll and Hill (1950).

Key Takeaways from this lesson

  • Ratio measures (RR, IR, OR) compare disease frequency between exposed and unexposed; RR and IR come from cohort studies, OR is the natural measure for case-control. When disease is rare, OR ≈ RR.
  • Risk difference (RD) and the attributable fraction in the exposed (AFe) express absolute and proportional impact in the exposed group; vaccine efficacy is a special case of AFe.
  • Population attributable risk (PAR) and AFp depend on both the strength of association and the prevalence of exposure: a common weak risk factor can dwarf a rare strong one in population terms.
  • The study design dictates the measure: cohort designs deliver risks/rates and ratios; case-control designs deliver odds ratios; cross-sectional designs deliver prevalence ratios.
  • Every measure of association needs a standard error and confidence interval; for ratio measures, CIs are constructed on the log scale and are therefore asymmetric.
  • The right hypothesis test follows the data: t-tests/ANOVA for normal continuous outcomes, χ² / Fisher / McNemar for categorical, and Wilcoxon / Mann–Whitney / Kruskal–Wallis / Spearman as non-parametric analogues.

Reflection

A colleague presents findings from a case-control study showing OR = 2.3 (95% CI: 1.1, 4.8) for the association between a workplace chemical exposure and bladder cancer. She concludes the chemical “causes 2.3 times the risk of bladder cancer.” In a case-control study participants are selected on outcome status, so absolute risks and risk ratios cannot be estimated directly; the odds ratio approximates the risk ratio only under the rare disease assumption (the outcome is uncommon, roughly under 5%, in the source population); and a confidence interval that excludes 1.0 indicates statistical significance at the 5% level but says nothing about bias or confounding. Evaluate this statement. What can and cannot be concluded from this study? Discuss the roles of strength of association, statistical significance, the rare disease assumption, and the distinction between OR and RR.

Model answerThe colleague has conflated three things: strength of association (OR 2.3 is moderate-to-strong, consistent with causation if confirmed); statistical significance (CI excludes 1, so chance is unlikely the sole explanation); and risk. The OR is not a risk ratio: case-control sampling fixes the row totals on outcome status, so absolute risks cannot be estimated. "2.3 times the risk" is correct only under the rare-disease approximation (OR ≈ RR when outcome prevalence < 5%); for bladder cancer in occupational cohorts that is plausible but not given. Further, association is not causation: the OR could be inflated by recall bias (cases recall workplace exposures more thoroughly), selection bias (hospital cases vs. community controls), or residual confounding (smoking, age, occupational co-exposures). What can be concluded: there is a statistically significant association of moderate strength in this case-control study; the rare-disease approximation makes the OR a reasonable proxy for the RR; the association is consistent with causation but not sufficient on its own. What is needed: a prospective cohort with biomarker exposure assessment, biological-mechanism evidence, and dose-response data.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Complete all 15 questions below with 100% accuracy to finish this lesson. You must also complete the reflection above before submitting.

Final Assessment: Measures of Association

1. Measures of association differ from measures of statistical significance in that they:

Measures of association (RR, OR, etc.) assess the strength of the relationship between exposure and disease. Statistical significance (P-values) is heavily influenced by sample size, not the magnitude of the effect.

2. In a cohort study, 150 of 2,000 exposed individuals and 75 of 2,000 non-exposed individuals develop the disease. What is the risk ratio?

RR = (150/2000) / (75/2000) = 0.075 / 0.0375 = 2.00. The risk of disease is twice as high in the exposed group.

3. The odds ratio can be calculated from a case-control study because:

OR = (a1×b0)/(a0×b1) is the same whether viewed as the ratio of disease odds or the ratio of exposure odds. Since case-control studies sample based on disease status, only the OR can be validly computed.

4. RD = 0.043 in the smoking and low-birth-weight example means:

RD is the absolute difference in risk: 0.114 − 0.071 = 0.043. This means 4.3 additional cases per 100 exposed women, above what would be expected based on the baseline risk.

5. If RR = 4.0, the attributable fraction in the exposed (AFe) is:

AFe = (RR−1)/RR = (4−1)/4 = 3/4 = 0.75. So 75% of disease in the exposed group is attributable to the exposure.

6. Vaccine efficacy of 80% indicates:

Vaccine efficacy = AFe = (risk unvaccinated − risk vaccinated)/risk unvaccinated. A value of 80% means the vaccine prevented 80% of expected cases.

7. A common risk factor with a modest RR may have a larger AFp than a rare risk factor with a high RR because:

AFp = p(E+)(RR−1)/[p(E+)(RR−1)+1]. Both the RR and the prevalence of exposure (p(E+)) contribute. A high p(E+) with a modest RR can produce a larger AFp than a low p(E+) with a high RR.

8. PAR is best described as:

PAR = p(D+) − p(D+|E−). It reflects the increase in disease risk in the entire population that is attributable to the exposure, incorporating both the strength of association and how common the exposure is.

9. The null value for the risk ratio (RR) is:

For ratio measures (RR, IR, OR), the null value is 1, meaning the risk (or rate or odds) is the same in both groups. For difference measures (RD, ID), the null value is 0.

10. A 95% CI for RR of (0.85, 1.32) suggests:

Since the 95% CI includes the null value of 1, we cannot reject H0 at the 0.05 level. The range of plausible values spans from a modest protective effect (0.85) to a modest risk increase (1.32).

11. Confidence intervals for OR are asymmetric around the point estimate because:

The CI is symmetric on the ln(OR) scale. When exponentiated back to the OR scale, the interval becomes asymmetric because the exponential function is non-linear.

12. In the formula var(ln OR) = 1/a1 + 1/a0 + 1/b1 + 1/b0, increasing all cell counts will:

Since the variance is a sum of reciprocals of cell counts, larger cell counts produce smaller reciprocals, reducing the overall variance. This leads to a narrower (more precise) confidence interval.

13. Which of the following measures CANNOT be estimated from a case-control study unless external data on disease incidence are available?

Because the investigator fixes the numbers of cases and controls, a case-control study cannot estimate the incidence rate in either exposure group unless external data on disease incidence are available. The OR can be computed directly, and when controls are selected by density sampling it also estimates the incidence rate ratio. AFe and AFp can be approximated using OR with appropriate external data.

14. The Pearson χ² test is most appropriate when:

The Pearson χ² has an approximate χ² distribution provided all expected cell values are >1 and at least 80% have expected values >5; in a 2×2 table, 80% of four cells means all four cells. For small samples, exact tests (like Fisher’s) are preferred.

15. A study finds RR = 1.8 (P = 0.40). Which interpretation is most appropriate?

An RR of 1.8 suggests a meaningful increase in risk. However, the P-value of 0.40 indicates the result is not statistically significant, likely due to insufficient sample size. A non-significant P-value does not prove the null hypothesis; it means we lack sufficient evidence to reject it. The CI would be more informative here.

✦ Complete the final reflection above before submitting