Chapter 1: A Definition of Causal Effect

Published

Last modified: 2026-10-09 12:11:16 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

You already reason about causes and effects every day, and you already know that seeing two things happen together does not mean that one caused the other. Someone who could not tell the difference would not last long: they would copy whatever the people who were later rewarded happened to do, however dangerous.

This chapter therefore does not try to teach new causal intuitions. Its job is to introduce the mathematical notation that formalizes the intuition you already have, so that causal concepts can be defined precisely. The rest of the book uses this notation throughout.

This content is based on Hernán and Robins (2020, chap. 1, pp. 3-12).

1 1.1 Individual Causal Effects (pp. 3-4)


Two vignettes:

  • Zeus receives a heart transplant on January 1 and dies five days later. Suppose we somehow knew that, had he not received the transplant, he would have been alive five days later. Then the transplant caused his death.
  • Hera also receives a heart transplant on January 1 and is alive five days later. Suppose we knew that she would also have been alive without the transplant. Then the transplant had no causal effect on her five-day survival.

In each case we compare (usually only in our heads) the outcome when an action \(A\) is taken with the outcome when it is withheld. If the two differ, \(A\) has a causal effect (causative or preventive) on the outcome.

Epidemiologists, statisticians, economists, and other social scientists call the action \(A\) an intervention, an exposure, a policy, or a treatment (Hernán and Robins 2020, 3).

1.1 Notation

  • Treatment \(A\): 1 if treated, 0 if untreated.
  • Outcome \(Y\): 1 if death, 0 if survival.
  • \(Y^{a=1}\) (“\(Y\) under treatment \(a = 1\)”): the outcome that would have been observed under treatment value \(a = 1\).
  • \(Y^{a=0}\): the outcome that would have been observed under treatment value \(a = 0\).

Zeus has \(Y^{a=1} = 1\) and \(Y^{a=0} = 0\); Hera has \(Y^{a=1} = 0\) and \(Y^{a=0} = 0\).

Capital letters denote random variables, that is, variables that may take different values for different individuals; lower-case letters denote particular values of a random variable. \(Y^{a=1}\) and \(Y^{a=0}\) are also random variables. We sometimes write \(Y_i^a\) for the outcome of individual \(i\); when \(i\) refers to a specific individual such as Zeus, \(Y_i^a\) is not random, because the book assumes for now that individual counterfactual outcomes are deterministic (Section 1.4 relaxes this) (Hernán and Robins 2020, 3–4).

1.2 Definition of an Individual Causal Effect

Definition 1 (Individual causal effect) Treatment \(A\) has a causal effect on an individual’s outcome \(Y\) if \(Y^{a=1} \neq Y^{a=0}\) for that individual.

The transplant has a causal effect on Zeus (\(Y^{a=1} = 1 \neq 0 = Y^{a=0}\)) but not on Hera (\(Y^{a=1} = 0 = Y^{a=0}\)).

The individual causal effect itself is the contrast \(Y_i^{a=1} - Y_i^{a=0}\): it is 1 for Zeus and 0 for Hera, and an effect exists exactly when this difference is not 0 (Hernán and Robins 2020, 4, margin note).

\(Y^{a=1}\) and \(Y^{a=0}\) are called potential outcomes or counterfactual outcomes.

“Potential outcomes” emphasizes that, depending on the treatment received, either outcome could potentially be observed. “Counterfactual outcomes” emphasizes that these outcomes describe situations that may not actually happen (counter to the fact) (Hernán and Robins 2020, 4).

1.3 Consistency

For each individual, the counterfactual outcome that corresponds to the treatment actually received is factual. Zeus was treated (\(A = 1\)), so his counterfactual outcome under treatment, \(Y^{a=1} = 1\), equals his observed outcome \(Y = 1\).

Definition 2 (Consistency) An individual with observed treatment \(A = a\) has observed outcome equal to the counterfactual outcome under \(a\): \[ \text{if } A_i = a, \text{ then } Y_i^a = Y_i^A = Y_i, \] written compactly as \(Y = Y^A\), where \(Y^A\) is the counterfactual \(Y^a\) evaluated at the individual’s observed treatment value.

1.4 Individual Effects Are Not Identified

Only one counterfactual outcome is observed per individual (the one for the treatment actually received); the others are missing. Because of this missing data, individual causal effects cannot be identified, that is, they cannot be expressed as a function of the observed data.

The book points to crossover experiments (Fine Point 2.1, in Chapter 2) as a possible exception, under strong assumptions (Hernán and Robins 2020, 4).

2 1.2 Average Causal Effects (pp. 4-7)


An individual causal effect needs three things: an outcome, the two actions \(a = 1\) and \(a = 0\) to compare, and the individual. An average causal effect replaces the individual with a well-defined population of individuals.

Take Zeus’s extended family (20 people) as the population. Table 1.1 lists both counterfactual outcomes for each member.

Table 1.1: Counterfactual outcomes for the 20 members of Zeus’s family (Hernán and Robins 2020, 5)

Name \(Y^{a=0}\) \(Y^{a=1}\)
Rheia 0 1
Kronos 1 0
Demeter 0 0
Hades 0 0
Hestia 0 0
Poseidon 1 0
Hera 0 0
Zeus 0 1
Artemis 1 1
Apollo 1 0
Leto 0 1
Ares 1 1
Athena 1 1
Hephaestus 0 1
Aphrodite 0 1
Polyphemus 0 1
Persephone 1 1
Hermes 1 0
Hebe 1 0
Dionysus 1 0

From Table 1.1:

  • 10 of 20 would have died had everyone been treated: \(\Pr[Y^{a=1} = 1] = 10/20 = 0.5\).
  • 10 of 20 would have died had no one been treated: \(\Pr[Y^{a=0} = 1] = 10/20 = 0.5\).

Counting the deaths and dividing by 20 is the same as averaging the 0/1 counterfactual outcomes over the 20 individuals: for a dichotomous outcome, the risk equals the average. You can check this by averaging the \(Y^{a=1}\) column of Table 1.1.

2.1 Definition of Average Causal Effect

Definition 3 (Average causal effect) An average causal effect of treatment \(A\) on outcome \(Y\) is present if \[ \Pr[Y^{a=1} = 1] \neq \Pr[Y^{a=0} = 1] \] in the population of interest. Because the risk of a dichotomous outcome equals its mean, the definition can be written as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), which also applies to nondichotomous outcomes.

In Zeus’s family both counterfactual risks are 0.5: whether all or none receive a transplant, half would die. The null hypothesis of no average causal effect is true.

On the difference scale, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), which for a binary outcome equals \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\); in Zeus’s family it is \(0.5 - 0.5 = 0\) (Hernán and Robins 2020, 6, margin note).

With more than two possible actions, the contrast of interest must be specified. “The causal effect of aspirin” means nothing until we say, for example, “taking, while alive, 150 mg of aspirin by mouth (or nasogastric tube if need be) daily for 5 years” versus “not taking aspirin.” That effect is well defined even if counterfactual outcomes under other interventions (e.g., absorbing 500 mg through the skin daily) are not (Hernán and Robins 2020, 6).

2.2 Null Average Effect, Non-Null Individual Effects

Absence of an average causal effect does not imply absence of individual effects. In Table 1.1, 12 individuals have \(Y^{a=1} \neq Y^{a=0}\):

  • 6 were harmed, \(Y^{a=1} - Y^{a=0} = 1\) (Rheia, Zeus, Leto, Hephaestus, Aphrodite, Polyphemus);
  • 6 were helped, \(Y^{a=1} - Y^{a=0} = -1\) (Kronos, Poseidon, Apollo, Hermes, Hebe, Dionysus).

The two groups being the same size is not an accident: a difference of averages equals the average of the differences, \[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{}, \] so a null average effect means the individual effects average to zero.

In Table 1.1 the right-hand side is \(\frac{6 \times 1 + 6 \times (-1) + 8 \times 0}{20} = \frac{0}{20} = 0\), matching \(0.5 - 0.5 = 0\) on the left.

Definition 4 (Sharp causal null hypothesis) The sharp causal null hypothesis holds when there is no causal effect for any individual in the population: \(Y^{a=1} = Y^{a=0}\) for all individuals.

The sharp causal null implies the null hypothesis of no average effect, but (as Table 1.1 shows) not conversely.

Even though individual effects are not identified, average causal effects sometimes are (Chapters 2 and 3). From here on, the book says “causal effect” for “average causal effect” and calls the null hypothesis of no average effect the causal null hypothesis (Hernán and Robins 2020, 6).

NoteFine Point 1.1: Interference

The definition of \(Y^a\) assumes that an individual’s counterfactual outcome does not depend on other individuals’ treatments. If Hera’s getting a new heart upset Zeus so much that he would not survive his own transplant (although he would have survived it had Hera not been transplanted), Hera’s treatment would interfere with Zeus’s outcome. Interference is common with contagious agents and educational programs.

Under interference, \(Y_i^a\) is not well defined; one must speak of, e.g., “the effect of transplant on Zeus when Hera does not get a new heart,” and the effect may differ for every allocation of hearts. Cox (1958) called the assumption of no interference “no interaction between units”; it is part of Rubin’s (1980) stable-unit-treatment-value assumption (SUTVA). The book assumes no interference unless stated otherwise (Hernán and Robins 2020, 5).

NoteTechnical Point 1.1: Causal Effects in the Population

\(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) is the mean counterfactual outcome had everyone received \(a\):

  • discrete outcomes: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_y y \, p_{Y^a}(y)\) with \(p_{Y^a}(y) = \Pr[Y^a = y]\); for dichotomous outcomes \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \Pr[Y^a = 1]\);
  • continuous outcomes: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \int y f_{Y^a}(y)\, dy\);
  • both: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \int y \, dF_{Y^a}(y)\), with \(F_{Y^a}\) the cdf of \(Y^a\).

There is a non-null average causal effect if \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a'}\right]\mathclose{}\) for any two values \(a\) and \(a'\).

A population causal effect can also contrast other functionals (median, variance, hazard, cdf) of the marginal distributions of the counterfactual outcomes. In Table 1.1, \(Y^{a=1}\) and \(Y^{a=0}\) have the same distribution (10 deaths out of 20), so the population causal effect on any functional is zero, e.g. \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{} = 0\).

Unlike the mean, a difference in variances is not in general the variance of the individual effects. Let \(D = Y^{a=1} - Y^{a=0}\), which is \(-1\) for 6 individuals, \(1\) for 6, and \(0\) for 8. Then \[\begin{align} \operatorname{E}\mathopen{}\left[D\right]\mathclose{} &= \frac{6 \times (-1) + 6 \times 1 + 8 \times 0}{20} = 0, \\ \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} &= \frac{6 \times 1 + 6 \times 1 + 8 \times 0}{20} = 0.6, \\ \operatorname{Var}\mathopen{}\left(D\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[D\right]\mathclose{}\right)\mathclose{}^2 = 0.6 - 0 = 0.6 > 0. \end{align}\] A randomized trial identifies \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{}\) but not \(\operatorname{Var}\mathopen{}\left(Y^{a=1} - Y^{a=0}\right)\mathclose{}\), because the covariance of \(Y^{a=1}\) and \(Y^{a=0}\) is never observed. The same holds for any nonlinear functional (Hernán and Robins 2020, 6).

3 1.3 Measures of Causal Effect (p. 7)


In Zeus’s family the causal null holds because both counterfactual risks equal 0.5. The causal null can be represented in equivalent ways:

  1. \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1] = 0\) (here \(0.5 - 0.5 = 0\))

  2. \(\dfrac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]} = 1\) (here \(0.5 / 0.5 = 1\))

  3. \(\dfrac{\Pr[Y^{a=1} = 1] / \Pr[Y^{a=1} = 0]}{\Pr[Y^{a=0} = 1] / \Pr[Y^{a=0} = 0]} = 1\)

The left-hand sides are the causal risk difference, causal risk ratio, and causal odds ratio.

When the causal null does not hold (say, smoking and lung cancer), these are not 0, 1, and 1; they quantify the same causal effect on different scales. Because they measure the causal effect, they are called effect measures.

The causal risk difference is the average of the individual effects \(Y^{a=1} - Y^{a=0}\) on the difference scale. The causal risk ratio is not the average of individual effects \(Y^{a=1}/Y^{a=0}\) on the ratio scale: it measures the causal effect in the population but is not an average of any individual causal effects (Hernán and Robins 2020, 7).

3.1 Which Measure?

Each measure serves a purpose. Suppose 3 in a million would develop the outcome if treated and 1 in a million if untreated:

  • causal risk ratio \(= \dfrac{3/10^6}{1/10^6} = 3\): treatment triples the risk (multiplicative scale);
  • causal risk difference \(= \dfrac{3}{10^6} - \dfrac{1}{10^6} = 0.000002\): the absolute number of cases attributable to treatment (additive scale).

The choice of scale depends on the goal of the inference.

NoteFine Point 1.2: Number Needed to Treat

In a population of 100 million, suppose 20 million would die within five years if treated and 30 million if untreated. Equivalent summaries:

  • causal risk difference \(= 0.2 - 0.3 = -0.1\);
  • treating all 100 million yields 10 million fewer deaths than treating none;
  • one must treat 100 million to save 10 million lives;
  • on average, one must treat 10 patients to save 1 life.

The number needed to treat (NNT) is the average number of individuals who must receive \(a = 1\) to reduce the number of cases by one. For treatments with a negative causal risk difference, \[ \text{NNT} = \frac{-1}{\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]} = \frac{-1}{-0.1} = 10 . \] For harmful treatments (positive risk difference) the symmetric quantity is the number needed to harm. The NNT was introduced by Laupacis, Sackett, and Roberts (1988); like the risk difference, it applies only to the population and time interval on which it is based (Hernán and Robins 2020, 8).

4 1.4 Random Variability (pp. 7-9)


Two liberties so far: the immortal Zeus cannot actually die, and real populations are much larger than 20. In practice investigators collect data on a sample of the population of interest, so population risks can only be estimated.

4.1 First Source of Random Error: Sampling Variability

View the 20 individuals of Table 1.1 as a random sample from a near-infinite super-population.

  • Sample proportion: \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^{a=0} = 1\right]\mathclose{} = 10/20 = 0.50\).
  • Super-population probability \(\Pr[Y^{a=0} = 1]\) need not equal it: it could be 0.57, with the sample giving 0.50 because of sampling variability.

The “hat” marks an estimator: \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^a = 1\right]\mathclose{}\) estimates \(\Pr[Y^a = 1]\).

Definition 5 (Consistent estimator) An estimator \(\hat\theta\) of \(\theta\) is consistent if, with probability approaching 1, \(\hat\theta- \theta\) approaches zero as the sample size increases towards infinity.

\(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^a = 1\right]\mathclose{}\) is consistent for \(\Pr[Y^a = 1]\) because sampling error is random and obeys the law of large numbers.

Caution: “consistency” of an estimator has a different meaning from “consistency” of counterfactual outcomes (Definition 2) (Hernán and Robins 2020, 9).

Because super-population probabilities can only be consistently estimated, never computed, we cannot conclude with certainty that there is or is not a causal effect; a statistical procedure is needed to evaluate the evidence about the causal null hypothesis \(\Pr[Y^{a=1} = 1] = \Pr[Y^{a=0} = 1]\) (Chapter 10).

4.2 Second Source of Random Error: Nondeterministic Counterfactuals

So far counterfactual outcomes are deterministic: Zeus has a 100% chance of dying if treated and 0% if untreated.

Alternatively, Zeus could have a 90% chance of dying if treated and 10% if untreated. Then his counterfactual outcomes are stochastic (nondeterministic), and the values in Table 1.1 are realizations of “random flips of mortality coins.” These probabilities would likely vary across individuals.

Quantum mechanics, unlike classical mechanics, holds that outcomes are inherently nondeterministic: if Zeus’s quantum mechanical probability of dying is 90%, no amount of data about Zeus removes the uncertainty about whether he will die if treated (Hernán and Robins 2020, 9).

4.3 Simplifying Assumptions Until Chapter 10

Random error comes from sampling variability, nondeterministic counterfactuals, or both. Until Chapter 10, the book assumes:

  • counterfactual outcomes are deterministic; and
  • we have data on every individual in a very large super-population, e.g. the 20 individuals stand for 20 billion, with 1 billion identical to Zeus, 1 billion identical to Hera, and so on.

Chapter 10 shows that estimates and confidence intervals for causal effects in the super-population are the same whether the world is stochastic (quantum) or deterministic (classical) at the individual level, whereas confidence intervals for the average causal effect in the actual study sample differ between the two cases. Super-population effects are usually the effects of substantive interest (Hernán and Robins 2020, 9).

NoteTechnical Point 1.2: Nondeterministic Counterfactuals

For nondeterministic counterfactuals, \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_y y \, p_{Y^a}(y)\), where \(p_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[Q_{Y^a}(\cdot)\right]\mathclose{}\) and \(Q_{Y^a}(y)\) is an individual’s random probability of outcome \(y\) under \(a\) (in the text, \(Q_{Y^{a=1}}(1) = 0.9\) for Zeus).

More generally, each individual has a distribution \(\Theta_{Y^a}(\cdot)\) of \(Y^a\), a random cdf. Then \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^a \mid \Theta_{Y^a}(\cdot)\right]\mathclose{}\right]\mathclose{} && \text{(law of total expectation)} \\ &= \operatorname{E}\mathopen{}\left[\int y \, d\Theta_{Y^a}(y)\right]\mathclose{} && \text{(mean of the individual distribution)} \\ &= \int y \, d\,\operatorname{E}\mathopen{}\left[\Theta_{Y^a}(y)\right]\mathclose{} && \text{(exchange expectation and integral)} \\ &= \int y \, dF_{Y^a}(y), && \text{with } F_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[\Theta_{Y^a}(\cdot)\right]\mathclose{}. \end{align}\]

For binary nondeterministic outcomes, the causal risk ratio \(\operatorname{E}\mathopen{}\left[Q_{Y^{a=1}}(1)\right]\mathclose{} / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\) equals the weighted average \(\operatorname{E}\mathopen{}\left[W \, Q_{Y^{a=1}}(1) / Q_{Y^{a=0}}(1)\right]\mathclose{}\) of the individual ratio-scale effects, with weights \(W = Q_{Y^{a=0}}(1) / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\), provided \(Q_{Y^{a=0}}(1)\) is never 0 (Hernán and Robins 2020, 10).

5 1.5 Causation versus Association (pp. 9-12)


Real data do not look like Table 1.1: we observe only one of each individual’s counterfactual outcomes, the one for the treatment actually received. We observe the treatment \(A\) and the outcome \(Y\), as in Table 1.2.

Table 1.2: Observed treatment and outcome in Zeus’s family (Hernán and Robins 2020, 9)

Name \(A\) \(Y\)
Rheia 0 0
Kronos 0 1
Demeter 0 0
Hades 0 0
Hestia 1 0
Poseidon 1 0
Hera 1 0
Zeus 1 1
Artemis 0 1
Apollo 0 1
Leto 0 0
Ares 1 1
Athena 1 1
Hephaestus 1 1
Aphrodite 1 1
Polyphemus 1 1
Persephone 1 1
Hermes 1 0
Hebe 1 0
Dionysus 1 0

The conditional probability \(\Pr[Y = 1 \mid A = a]\) is the proportion who developed the outcome among those who happened to receive treatment \(a\):

  • 7 of the 13 treated died: \(\Pr[Y = 1 \mid A = 1] = 7/13\);
  • 3 of the 7 untreated died: \(\Pr[Y = 1 \mid A = 0] = 3/7\).

5.1 Independence and Association

Definition 6 (Independence) Treatment \(A\) and outcome \(Y\) are independent (\(A\) is not associated with \(Y\), or \(A\) does not predict \(Y\)) when \(\Pr[Y = 1 \mid A = 1] = \Pr[Y = 1 \mid A = 0]\). Independence is written \(Y \perp\!\!\!\perp A\), or equivalently \(A \perp\!\!\!\perp Y\).

Equivalent statements of independence:

  1. \(\Pr[Y = 1 \mid A = 1] - \Pr[Y = 1 \mid A = 0] = 0\)

  2. \(\dfrac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = 1\)

  3. \(\dfrac{\Pr[Y = 1 \mid A = 1] / \Pr[Y = 0 \mid A = 1]}{\Pr[Y = 1 \mid A = 0] / \Pr[Y = 0 \mid A = 0]} = 1\)

The left-hand sides are the associational risk difference, risk ratio, and odds ratio, collectively association measures. \(A\) and \(Y\) are associated (dependent) when \(\Pr[Y = 1 \mid A = 1] \neq \Pr[Y = 1 \mid A = 0]\).

Dawid (1979) introduced the symbol \(\perp\!\!\!\perp\) for independence. For a continuous outcome, mean independence is \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\); for dichotomous outcomes, independence and mean independence coincide. So association can be written \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) for both kinds of outcome. For binary \(A\), \(Y\) and \(A\) are not associated if and only if they are uncorrelated (Hernán and Robins 2020, 10–11).

5.2 Association Measures in Zeus’s Family

Treatment and outcome are associated because \(7/13 \neq 3/7\): \[\begin{align} \text{risk difference} &= \frac{7}{13} - \frac{3}{7} = \frac{49 - 39}{91} = \frac{10}{91} \approx 0.11, \\ \text{risk ratio} &= \frac{7/13}{3/7} = \frac{7 \times 7}{13 \times 3} = \frac{49}{39} \approx 1.26, \\ \text{odds ratio} &= \frac{(7/13)/(6/13)}{(3/7)/(4/7)} = \frac{7/6}{3/4} = \frac{28}{18} \approx 1.56 . \end{align}\]

Association measures are also subject to random variability; until Chapter 10 the book ignores this by assuming the population of Table 1.2 is extremely large (Hernán and Robins 2020, 11).

5.3 Causation Versus Association

In the same population of 20:

  • no causation: risk if all treated \(= 10/20\) equals risk if all untreated \(= 10/20\);
  • association: risk in the 13 treated \(= 7/13\) is greater than risk in the 7 untreated \(= 3/7\).

Figure 1.1 in the book draws the population as a diamond divided into a white area (the treated) and a smaller grey area (the untreated). Causation contrasts the whole diamond all white (everyone treated) with the whole diamond all grey (everyone untreated). Association contrasts the white area with the grey area of the original diamond (Hernán and Robins 2020, 11).

Causation Association
Question “What would the risk be if everybody had been treated / untreated?” “What is the risk in the treated / the untreated?”
World counterfactual actual
Risk marginal \(\Pr[Y^a = 1]\), in the whole population conditional \(\Pr[Y = 1 \mid A = a]\), in the subset with \(A = a\)
Comparison same population, two treatment values two disjoint subsets defined by actual treatment

These different definitions explain the adage “association is not causation.”

The book often writes the redundant “causal effect” to avoid confusion with the common use of “effect” to mean association.

The discrepancy in Zeus’s family would not be surprising if the transplant recipients were, on average, sicker than those who did not receive one. Chapter 7 calls this discrepancy confounding.

Why the distinction matters: suppose the causal risk ratio of 5-year mortality for aspirin vs. no aspirin is 0.5, but the associational risk ratio is 1.5 because people at high cardiovascular risk are preferentially prescribed aspirin. A physician who withholds aspirin because the treated die more often is acting on association, not causation, and “will be sued for malpractice” (Hernán and Robins 2020, 12).

5.4 The Question for the Rest of the Book

Causal inference needs data like the hypothetical Table 1.1, but real data look like Table 1.2. Under which conditions can real-world data be used for causal inference? Chapter 2 gives one answer: conduct a randomized experiment.

6 Summary


  1. Individual causal effect: the contrast \(Y^{a=1} - Y^{a=0}\) for an individual, present when \(Y^{a=1} \neq Y^{a=0}\); not identified, because only one counterfactual outcome is observed (consistency: \(Y = Y^A\)).
  2. Average causal effect: the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{}\) in a population, present when \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\); a null average effect does not rule out individual effects (the sharp null is stronger).
  3. Effect measures: causal risk difference, risk ratio, and odds ratio quantify one effect on different scales; for a beneficial treatment, the NNT is the reciprocal of the absolute risk difference.
  4. Random variability: from sampling and possibly from nondeterministic counterfactuals; ignored until Chapter 10.
  5. Causation vs. association: causation compares the whole population under two treatment values; association compares two subsets defined by the treatment actually received.

7 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top