You already reason about causes and effects every day, and you already know that seeing two things happen together does not mean that one caused the other. Someone who could not tell the difference would not last long: they would copy whatever the people who were later rewarded happened to do, however dangerous.
This chapter therefore does not try to teach new causal intuitions. Its job is to introduce the mathematical notation that formalizes the intuition you already have, so that causal concepts can be defined precisely. The rest of the book uses this notation throughout.
Two vignettes:
In each case we compare (usually only in our heads) the outcome when an action \(A\) is taken with the outcome when it is withheld. If the two differ, \(A\) has a causal effect (causative or preventive) on the outcome.
Zeus has \(Y^{a=1} = 1\) and \(Y^{a=0} = 0\); Hera has \(Y^{a=1} = 0\) and \(Y^{a=0} = 0\).
Definition 1 (Individual causal effect) Treatment \(A\) has a causal effect on an individual’s outcome \(Y\) if \(Y^{a=1} \neq Y^{a=0}\) for that individual.
The transplant has a causal effect on Zeus (\(Y^{a=1} = 1 \neq 0 = Y^{a=0}\)) but not on Hera (\(Y^{a=1} = 0 = Y^{a=0}\)).
The individual causal effect itself is the contrast \(Y_i^{a=1} - Y_i^{a=0}\): it is 1 for Zeus and 0 for Hera, and an effect exists exactly when this difference is not 0 (Hernán and Robins 2020, 4, margin note).
\(Y^{a=1}\) and \(Y^{a=0}\) are called potential outcomes or counterfactual outcomes.
For each individual, the counterfactual outcome that corresponds to the treatment actually received is factual. Zeus was treated (\(A = 1\)), so his counterfactual outcome under treatment, \(Y^{a=1} = 1\), equals his observed outcome \(Y = 1\).
Definition 2 (Consistency) An individual with observed treatment \(A = a\) has observed outcome equal to the counterfactual outcome under \(a\): \[ \text{if } A_i = a, \text{ then } Y_i^a = Y_i^A = Y_i, \] written compactly as \(Y = Y^A\), where \(Y^A\) is the counterfactual \(Y^a\) evaluated at the individual’s observed treatment value.
Only one counterfactual outcome is observed per individual (the one for the treatment actually received); the others are missing. Because of this missing data, individual causal effects cannot be identified, that is, they cannot be expressed as a function of the observed data.
An individual causal effect needs three things: an outcome, the two actions \(a = 1\) and \(a = 0\) to compare, and the individual. An average causal effect replaces the individual with a well-defined population of individuals.
Take Zeus’s extended family (20 people) as the population. Table 1.1 lists both counterfactual outcomes for each member.
Table 1.1: Counterfactual outcomes for the 20 members of Zeus’s family (Hernán and Robins 2020, 5)
| Name | \(Y^{a=0}\) | \(Y^{a=1}\) |
|---|---|---|
| Rheia | 0 | 1 |
| Kronos | 1 | 0 |
| Demeter | 0 | 0 |
| Hades | 0 | 0 |
| Hestia | 0 | 0 |
| Poseidon | 1 | 0 |
| Hera | 0 | 0 |
| Zeus | 0 | 1 |
| Artemis | 1 | 1 |
| Apollo | 1 | 0 |
| Leto | 0 | 1 |
| Ares | 1 | 1 |
| Athena | 1 | 1 |
| Hephaestus | 0 | 1 |
| Aphrodite | 0 | 1 |
| Polyphemus | 0 | 1 |
| Persephone | 1 | 1 |
| Hermes | 1 | 0 |
| Hebe | 1 | 0 |
| Dionysus | 1 | 0 |
From Table 1.1:
Definition 3 (Average causal effect) An average causal effect of treatment \(A\) on outcome \(Y\) is present if \[ \Pr[Y^{a=1} = 1] \neq \Pr[Y^{a=0} = 1] \] in the population of interest. Because the risk of a dichotomous outcome equals its mean, the definition can be written as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), which also applies to nondichotomous outcomes.
In Zeus’s family both counterfactual risks are 0.5: whether all or none receive a transplant, half would die. The null hypothesis of no average causal effect is true.
On the difference scale, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), which for a binary outcome equals \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\); in Zeus’s family it is \(0.5 - 0.5 = 0\) (Hernán and Robins 2020, 6, margin note).
Absence of an average causal effect does not imply absence of individual effects. In Table 1.1, 12 individuals have \(Y^{a=1} \neq Y^{a=0}\):
The two groups being the same size is not an accident: a difference of averages equals the average of the differences, \[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{}, \] so a null average effect means the individual effects average to zero.
Definition 4 (Sharp causal null hypothesis) The sharp causal null hypothesis holds when there is no causal effect for any individual in the population: \(Y^{a=1} = Y^{a=0}\) for all individuals.
The sharp causal null implies the null hypothesis of no average effect, but (as Table 1.1 shows) not conversely.
Fine Point 1.1: Interference
The definition of \(Y^a\) assumes that an individual’s counterfactual outcome does not depend on other individuals’ treatments. If Hera’s getting a new heart upset Zeus so much that he would not survive his own transplant (although he would have survived it had Hera not been transplanted), Hera’s treatment would interfere with Zeus’s outcome. Interference is common with contagious agents and educational programs.
Under interference, \(Y_i^a\) is not well defined; one must speak of, e.g., “the effect of transplant on Zeus when Hera does not get a new heart,” and the effect may differ for every allocation of hearts. Cox (1958) called the assumption of no interference “no interaction between units”; it is part of Rubin’s (1980) stable-unit-treatment-value assumption (SUTVA). The book assumes no interference unless stated otherwise (Hernán and Robins 2020, 5).
Technical Point 1.1: Causal Effects in the Population
\(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) is the mean counterfactual outcome had everyone received \(a\):
There is a non-null average causal effect if \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a'}\right]\mathclose{}\) for any two values \(a\) and \(a'\).
A population causal effect can also contrast other functionals (median, variance, hazard, cdf) of the marginal distributions of the counterfactual outcomes. In Table 1.1, \(Y^{a=1}\) and \(Y^{a=0}\) have the same distribution (10 deaths out of 20), so the population causal effect on any functional is zero, e.g. \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{} = 0\).
Unlike the mean, a difference in variances is not in general the variance of the individual effects. Let \(D = Y^{a=1} - Y^{a=0}\), which is \(-1\) for 6 individuals, \(1\) for 6, and \(0\) for 8. Then \[\begin{align} \operatorname{E}\mathopen{}\left[D\right]\mathclose{} &= \frac{6 \times (-1) + 6 \times 1 + 8 \times 0}{20} = 0, \\ \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} &= \frac{6 \times 1 + 6 \times 1 + 8 \times 0}{20} = 0.6, \\ \operatorname{Var}\mathopen{}\left(D\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[D\right]\mathclose{}\right)\mathclose{}^2 = 0.6 - 0 = 0.6 > 0. \end{align}\] A randomized trial identifies \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{}\) but not \(\operatorname{Var}\mathopen{}\left(Y^{a=1} - Y^{a=0}\right)\mathclose{}\), because the covariance of \(Y^{a=1}\) and \(Y^{a=0}\) is never observed. The same holds for any nonlinear functional (Hernán and Robins 2020, 6).
In Zeus’s family the causal null holds because both counterfactual risks equal 0.5. The causal null can be represented in equivalent ways:
\(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1] = 0\) (here \(0.5 - 0.5 = 0\))
\(\dfrac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]} = 1\) (here \(0.5 / 0.5 = 1\))
\(\dfrac{\Pr[Y^{a=1} = 1] / \Pr[Y^{a=1} = 0]}{\Pr[Y^{a=0} = 1] / \Pr[Y^{a=0} = 0]} = 1\)
The left-hand sides are the causal risk difference, causal risk ratio, and causal odds ratio.
When the causal null does not hold (say, smoking and lung cancer), these are not 0, 1, and 1; they quantify the same causal effect on different scales. Because they measure the causal effect, they are called effect measures.
Each measure serves a purpose. Suppose 3 in a million would develop the outcome if treated and 1 in a million if untreated:
The choice of scale depends on the goal of the inference.
Fine Point 1.2: Number Needed to Treat
In a population of 100 million, suppose 20 million would die within five years if treated and 30 million if untreated. Equivalent summaries:
The number needed to treat (NNT) is the average number of individuals who must receive \(a = 1\) to reduce the number of cases by one. For treatments with a negative causal risk difference, \[ \text{NNT} = \frac{-1}{\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]} = \frac{-1}{-0.1} = 10 . \] For harmful treatments (positive risk difference) the symmetric quantity is the number needed to harm. The NNT was introduced by Laupacis, Sackett, and Roberts (1988); like the risk difference, it applies only to the population and time interval on which it is based (Hernán and Robins 2020, 8).
Two liberties so far: the immortal Zeus cannot actually die, and real populations are much larger than 20. In practice investigators collect data on a sample of the population of interest, so population risks can only be estimated.
View the 20 individuals of Table 1.1 as a random sample from a near-infinite super-population.
The “hat” marks an estimator: \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^a = 1\right]\mathclose{}\) estimates \(\Pr[Y^a = 1]\).
Definition 5 (Consistent estimator) An estimator \(\hat\theta\) of \(\theta\) is consistent if, with probability approaching 1, \(\hat\theta- \theta\) approaches zero as the sample size increases towards infinity.
\(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^a = 1\right]\mathclose{}\) is consistent for \(\Pr[Y^a = 1]\) because sampling error is random and obeys the law of large numbers.
So far counterfactual outcomes are deterministic: Zeus has a 100% chance of dying if treated and 0% if untreated.
Alternatively, Zeus could have a 90% chance of dying if treated and 10% if untreated. Then his counterfactual outcomes are stochastic (nondeterministic), and the values in Table 1.1 are realizations of “random flips of mortality coins.” These probabilities would likely vary across individuals.
Random error comes from sampling variability, nondeterministic counterfactuals, or both. Until Chapter 10, the book assumes:
Technical Point 1.2: Nondeterministic Counterfactuals
For nondeterministic counterfactuals, \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_y y \, p_{Y^a}(y)\), where \(p_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[Q_{Y^a}(\cdot)\right]\mathclose{}\) and \(Q_{Y^a}(y)\) is an individual’s random probability of outcome \(y\) under \(a\) (in the text, \(Q_{Y^{a=1}}(1) = 0.9\) for Zeus).
More generally, each individual has a distribution \(\Theta_{Y^a}(\cdot)\) of \(Y^a\), a random cdf. Then \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^a \mid \Theta_{Y^a}(\cdot)\right]\mathclose{}\right]\mathclose{} && \text{(law of total expectation)} \\ &= \operatorname{E}\mathopen{}\left[\int y \, d\Theta_{Y^a}(y)\right]\mathclose{} && \text{(mean of the individual distribution)} \\ &= \int y \, d\,\operatorname{E}\mathopen{}\left[\Theta_{Y^a}(y)\right]\mathclose{} && \text{(exchange expectation and integral)} \\ &= \int y \, dF_{Y^a}(y), && \text{with } F_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[\Theta_{Y^a}(\cdot)\right]\mathclose{}. \end{align}\]
For binary nondeterministic outcomes, the causal risk ratio \(\operatorname{E}\mathopen{}\left[Q_{Y^{a=1}}(1)\right]\mathclose{} / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\) equals the weighted average \(\operatorname{E}\mathopen{}\left[W \, Q_{Y^{a=1}}(1) / Q_{Y^{a=0}}(1)\right]\mathclose{}\) of the individual ratio-scale effects, with weights \(W = Q_{Y^{a=0}}(1) / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\), provided \(Q_{Y^{a=0}}(1)\) is never 0 (Hernán and Robins 2020, 10).
Real data do not look like Table 1.1: we observe only one of each individual’s counterfactual outcomes, the one for the treatment actually received. We observe the treatment \(A\) and the outcome \(Y\), as in Table 1.2.
Table 1.2: Observed treatment and outcome in Zeus’s family (Hernán and Robins 2020, 9)
| Name | \(A\) | \(Y\) |
|---|---|---|
| Rheia | 0 | 0 |
| Kronos | 0 | 1 |
| Demeter | 0 | 0 |
| Hades | 0 | 0 |
| Hestia | 1 | 0 |
| Poseidon | 1 | 0 |
| Hera | 1 | 0 |
| Zeus | 1 | 1 |
| Artemis | 0 | 1 |
| Apollo | 0 | 1 |
| Leto | 0 | 0 |
| Ares | 1 | 1 |
| Athena | 1 | 1 |
| Hephaestus | 1 | 1 |
| Aphrodite | 1 | 1 |
| Polyphemus | 1 | 1 |
| Persephone | 1 | 1 |
| Hermes | 1 | 0 |
| Hebe | 1 | 0 |
| Dionysus | 1 | 0 |
The conditional probability \(\Pr[Y = 1 \mid A = a]\) is the proportion who developed the outcome among those who happened to receive treatment \(a\):
Definition 6 (Independence) Treatment \(A\) and outcome \(Y\) are independent (\(A\) is not associated with \(Y\), or \(A\) does not predict \(Y\)) when \(\Pr[Y = 1 \mid A = 1] = \Pr[Y = 1 \mid A = 0]\). Independence is written \(Y \perp\!\!\!\perp A\), or equivalently \(A \perp\!\!\!\perp Y\).
Equivalent statements of independence:
\(\Pr[Y = 1 \mid A = 1] - \Pr[Y = 1 \mid A = 0] = 0\)
\(\dfrac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = 1\)
\(\dfrac{\Pr[Y = 1 \mid A = 1] / \Pr[Y = 0 \mid A = 1]}{\Pr[Y = 1 \mid A = 0] / \Pr[Y = 0 \mid A = 0]} = 1\)
The left-hand sides are the associational risk difference, risk ratio, and odds ratio, collectively association measures. \(A\) and \(Y\) are associated (dependent) when \(\Pr[Y = 1 \mid A = 1] \neq \Pr[Y = 1 \mid A = 0]\).
Treatment and outcome are associated because \(7/13 \neq 3/7\): \[\begin{align} \text{risk difference} &= \frac{7}{13} - \frac{3}{7} = \frac{49 - 39}{91} = \frac{10}{91} \approx 0.11, \\ \text{risk ratio} &= \frac{7/13}{3/7} = \frac{7 \times 7}{13 \times 3} = \frac{49}{39} \approx 1.26, \\ \text{odds ratio} &= \frac{(7/13)/(6/13)}{(3/7)/(4/7)} = \frac{7/6}{3/4} = \frac{28}{18} \approx 1.56 . \end{align}\]
In the same population of 20:
| Causation | Association | |
|---|---|---|
| Question | “What would the risk be if everybody had been treated / untreated?” | “What is the risk in the treated / the untreated?” |
| World | counterfactual | actual |
| Risk | marginal \(\Pr[Y^a = 1]\), in the whole population | conditional \(\Pr[Y = 1 \mid A = a]\), in the subset with \(A = a\) |
| Comparison | same population, two treatment values | two disjoint subsets defined by actual treatment |
These different definitions explain the adage “association is not causation.”
Causal inference needs data like the hypothetical Table 1.1, but real data look like Table 1.2. Under which conditions can real-world data be used for causal inference? Chapter 2 gives one answer: conduct a randomized experiment.