Chapter 1: A Definition of Causal Effect

You already reason about causes and effects every day, and you already know that seeing two things happen together does not mean that one caused the other. Someone who could not tell the difference would not last long: they would copy whatever the people who were later rewarded happened to do, however dangerous.

This chapter therefore does not try to teach new causal intuitions. Its job is to introduce the mathematical notation that formalizes the intuition you already have, so that causal concepts can be defined precisely. The rest of the book uses this notation throughout.

Check Each Symbol Against Your Intuition

As each piece of notation appears, restate it in words and check that it says what your intuition already says. If a formula and your intuition disagree, settle the disagreement before reading on: later chapters build on these definitions.

1 1.1 Individual Causal Effects (pp. 3-4)

Example 1 (Two heart transplants)  

  • Zeus receives a heart transplant on January 1 and dies five days later. Suppose we somehow knew that, had he not received the transplant, he would have been alive five days later. Then the transplant caused his death.
  • Hera also receives a heart transplant on January 1 and is alive five days later. Suppose we knew that she would also have been alive without the transplant. Then the transplant had no causal effect on her five-day survival.

In each case we compare (usually only in our heads) the outcome when an action \(A\) is taken with the outcome when it is withheld. If the two differ, \(A\) has a causal effect (causative or preventive) on the outcome.

Notation

Definition 1 (Treatment, outcome, and counterfactual outcomes) Let \(A\) be a dichotomous treatment (1: treated, 0: untreated) and \(Y\) a dichotomous outcome (1: death, 0: survival). For each treatment value \(a\), the counterfactual outcome \(Y^a\) (read “\(Y\) under treatment \(a\)”) is the outcome that would have been observed had the individual received treatment value \(a\). With a dichotomous treatment, each individual has two counterfactual outcomes, \(Y^{a=1}\) and \(Y^{a=0}\). They are also called potential outcomes. Indexing \(Y^a\) by the individual’s own treatment value alone assumes that the treatments of other individuals do not affect that individual’s outcome. For now, each counterfactual outcome is a fixed value for each individual (Hernán and Robins 2020, 3–4).

Example 2 (Counterfactual outcomes of Zeus and Hera) In Example 1, Zeus died when treated and would have survived untreated, so \(Y^{a=1} = 1\) and \(Y^{a=0} = 0\) for him. Hera survived when treated and would also have survived untreated, so \(Y^{a=1} = 0\) and \(Y^{a=0} = 0\) for her.

Writing \(Y^a\) for an individual takes for granted that the individual’s own treatment value is all that matters.

Definition 2 (No interference) There is no interference between individuals when each individual’s counterfactual outcome \(Y^a\) (Definition 1) depends only on that individual’s own treatment value \(a\), and not on the treatment values of anyone else in the population (Hernán and Robins 2020, 5).

Example 3 (Interference between Zeus and Hera) Suppose Zeus would survive his own transplant if Hera did not get a new heart, but Hera’s transplant would upset him so much that he would die after his own. Then Zeus’s outcome under \(a = 1\) depends on Hera’s treatment, so there is interference, and “Zeus’s \(Y^{a=1}\)” has no single value until Hera’s treatment is also fixed.

Definition of an Individual Causal Effect

Definition 3 (Individual causal effect) Let \(Y^{a=1}\) and \(Y^{a=0}\) be an individual’s counterfactual outcomes under the two values of a dichotomous treatment \(A\) (Definition 1). Treatment \(A\) has a causal effect on that individual’s outcome \(Y\) if \(Y^{a=1} \neq Y^{a=0}\). The individual causal effect of individual \(i\) is the contrast \(Y_i^{a=1} - Y_i^{a=0}\), which is nonzero exactly when such an effect exists (Hernán and Robins 2020, 4).

Example 4 (Individual causal effects of Zeus and Hera) With the counterfactual outcomes of Example 2, the transplant has a causal effect on Zeus, because \(Y^{a=1} = 1 \neq 0 = Y^{a=0}\), and his individual causal effect is \(1 - 0 = 1\). It has no causal effect on Hera, because \(Y^{a=1} = 0 = Y^{a=0}\), and her individual causal effect is \(0 - 0 = 0\).

Consistency

For each individual, the counterfactual outcome that corresponds to the treatment actually received is factual.

Definition 4 (Consistency) Let \(A\) be an individual’s observed treatment, \(Y\) the observed outcome, and \(Y^a\) the counterfactual outcomes (Definition 1). Consistency holds if an individual with observed treatment \(A = a\) has observed outcome equal to the counterfactual outcome under \(a\): \[ \text{if } A_i = a, \text{ then } Y_i^a = Y_i^{A_i} = Y_i, \tag{1}\] written compactly as \(Y = Y^A\), where \(Y^A\) is the counterfactual \(Y^a\) evaluated at the individual’s observed treatment value (Hernán and Robins 2020, 4).

Example 5 (Consistency for Zeus and Hera) Zeus and Hera were both treated (\(A = 1\)). Under consistency, each one’s observed outcome is their \(Y^{a=1}\): Zeus’s observed outcome is \(Y = 1 = Y^{a=1}\), and Hera’s is \(Y = 0 = Y^{a=1}\). Consistency says nothing about \(Y^{a=0}\) for either of them, because neither was untreated.

Individual Effects Are Not Identified

Only one counterfactual outcome is observed per individual (the one for the treatment actually received); the others are missing.

Definition 5 (Identified) A quantity is identified if it can be expressed as a function of the observed data. Equivalently, any two states of the world that produce the same observed data give the quantity the same value (Hernán and Robins 2020, 4).

Example 6 (An identified quantity) Zeus’s counterfactual outcome under the treatment he received is identified under consistency (Definition 4): whatever the rest of the world looks like, it equals his observed outcome, \(Y^{a=1} = Y = 1\).

Proposition 1 (Individual causal effects are not identified) Suppose that, for each individual, only the treatment \(A\) and the outcome \(Y\) are observed, that \(A\) and \(Y\) are dichotomous, that consistency (Definition 4) holds, and that nothing else is assumed about the counterfactual outcome under the treatment value the individual did not receive. Then the individual causal effect \(Y^{a=1} - Y^{a=0}\) (Definition 3) is not identified (Definition 5).

Proof. Take Zeus, with observed \(A = 1\) and \(Y = 1\). By consistency, \(Y^{a=1} = 1\). Two states of the world agree with these observed data: one with \(Y^{a=0} = 0\), where Zeus’s individual causal effect is \(1 - 0 = 1\), and one with \(Y^{a=0} = 1\), where it is \(1 - 1 = 0\). The same observed data give two different values, so the individual causal effect is not a function of the observed data. The same argument applies to any individual, with the roles of \(a = 1\) and \(a = 0\) swapped for an untreated one.

2 1.2 Average Causal Effects (pp. 4-7)

An individual causal effect needs three things: an outcome, the two actions \(a = 1\) and \(a = 0\) to compare, and the individual. An average causal effect replaces the individual with a well-defined population of individuals.

Take Zeus’s extended family (20 people) as the population. Table 1.1 lists both counterfactual outcomes for each member.

Table 1.1: Counterfactual outcomes for the 20 members of Zeus’s family (Hernán and Robins 2020, 5)

Name \(Y^{a=0}\) \(Y^{a=1}\)
Rheia 0 1
Kronos 1 0
Demeter 0 0
Hades 0 0
Hestia 0 0
Poseidon 1 0
Hera 0 0
Zeus 0 1
Artemis 1 1
Apollo 1 0
Leto 0 1
Ares 1 1
Athena 1 1
Hephaestus 0 1
Aphrodite 0 1
Polyphemus 0 1
Persephone 1 1
Hermes 1 0
Hebe 1 0
Dionysus 1 0

Definition 6 (Counterfactual risk) For a dichotomous outcome \(Y\) and a treatment value \(a\), the counterfactual risk \(\Pr[Y^a = 1]\) is the proportion of the population who would develop the outcome had everyone in the population received treatment value \(a\). Because \(Y^a\) takes only the values 0 and 1, this proportion equals the mean \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) (Hernán and Robins 2020, 5).

Example 7 (Counterfactual risks in Zeus’s family) From Table 1.1:

  • 10 of 20 would have died had everyone been treated: \(\Pr[Y^{a=1} = 1] = 10/20 = 0.5\).
  • 10 of 20 would have died had no one been treated: \(\Pr[Y^{a=0} = 1] = 10/20 = 0.5\).

Averaging the 0/1 entries of the \(Y^{a=1}\) column gives the same \(10/20\), as Definition 6 says it must.

Definition of Average Causal Effect

Definition 7 (Average causal effect) Let \(A\) be a dichotomous treatment and \(Y\) an outcome, with counterfactual outcomes \(Y^{a=1}\) and \(Y^{a=0}\) (Definition 1), in a well-defined population of interest. An average causal effect of \(A\) on \(Y\) is present in that population if \[ \Pr[Y^{a=1} = 1] \neq \Pr[Y^{a=0} = 1] \tag{2}\] for a dichotomous outcome, or, for any outcome with finite means, \[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}. \tag{3}\] When the two sides are equal, the null hypothesis of no average causal effect is true. On the difference scale, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), which for a dichotomous outcome equals \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\) (Hernán and Robins 2020, 5–6).

Example 8 (No average causal effect in Zeus’s family) In Example 7 both counterfactual risks are 0.5: whether all or none receive a transplant, half would die. The null hypothesis of no average causal effect is true, and the average causal effect on the difference scale is \(0.5 - 0.5 = 0\).

Name Both Treatment Values

With more than two possible actions, the contrast of interest must be specified. “The causal effect of aspirin” is undefined until both arms are spelled out, with dose, route, frequency, and duration: for example, a daily low-dose tablet for five years versus no aspirin. Such an effect can be well defined even when the counterfactual outcomes under other aspirin interventions, at another dose or by another route, are not (Hernán and Robins 2020, 6).

Null Average Effect, Non-Null Individual Effects

Absence of an average causal effect does not imply absence of individual effects.

Example 9 (Individual effects that cancel out) In Table 1.1, 12 individuals have \(Y^{a=1} \neq Y^{a=0}\):

  • 6 were harmed, \(Y^{a=1} - Y^{a=0} = 1\) (Rheia, Zeus, Leto, Hephaestus, Aphrodite, Polyphemus);
  • 6 were helped, \(Y^{a=1} - Y^{a=0} = -1\) (Kronos, Poseidon, Apollo, Hermes, Hebe, Dionysus).

The other 8 have \(Y^{a=1} - Y^{a=0} = 0\). The average of the 20 individual causal effects is \(\frac{6 \times 1 + 6 \times (-1) + 8 \times 0}{20} = 0\), the same as the average causal effect \(0.5 - 0.5 = 0\) of Example 8.

The two groups being the same size is not an accident.

Proposition 2 (The average causal effect is the average of the individual effects) In any population in which \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) are finite, \[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{}. \tag{4}\] So the average causal effect on the difference scale (Definition 7) is zero exactly when the individual causal effects (Definition 3) average to zero (Hernán and Robins 2020, 6).

Proof. Expectation is linear: the mean of a difference is the difference of the means. In a finite population of \(n\) individuals, for example, \(\frac{1}{n}\sum_i Y_i^{a=1} - \frac{1}{n}\sum_i Y_i^{a=0} = \frac{1}{n}\sum_i \mathopen{}\left(Y_i^{a=1} - Y_i^{a=0}\right)\mathclose{}\).

Definition 8 (Sharp causal null hypothesis) Let \(A\) be a dichotomous treatment with counterfactual outcomes \(Y^{a=1}\) and \(Y^{a=0}\) (Definition 1). The sharp causal null hypothesis holds in a population when there is no individual causal effect (Definition 3) for anyone in it: \(Y^{a=1} = Y^{a=0}\) for all individuals (Hernán and Robins 2020, 6).

Example 10 (Where the sharp null does and does not hold) In the whole of Zeus’s family the sharp causal null is false: Zeus, for one, has \(Y^{a=1} = 1 \neq 0 = Y^{a=0}\). Among the 8 family members with no individual effect (Demeter, Hades, Hestia, and Hera, with \(Y^{a=1} = Y^{a=0} = 0\), and Artemis, Ares, Athena, and Persephone, with \(Y^{a=1} = Y^{a=0} = 1\)), it holds.

Proposition 3 (The sharp null implies the null of no average effect) Let \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) be finite. If the sharp causal null hypothesis (Definition 8) holds in a population, then the null hypothesis of no average causal effect (Definition 7) holds in that population. The converse is false.

Proof. Under the sharp null, \(Y^{a=1} - Y^{a=0} = 0\) for every individual, so \(\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{} = 0\), and by Proposition 2 \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\). For the converse, Zeus’s family is a counterexample: the null of no average effect holds (Example 8), but the sharp null does not (Example 10).

Terminology for the Rest of the Book

Individual causal effects are not identified (Proposition 1), but average causal effects sometimes are (Chapters 2 and 3). From here on, “causal effect” means average causal effect, and the null hypothesis of no average causal effect is called the causal null hypothesis (Hernán and Robins 2020, 6).

Fine Point 1.1: Interference

Definition 1 assumes no interference (Definition 2), and Example 3 shows what goes wrong without it. Interference is common with contagious agents and educational programs, where an individual’s outcome is influenced by social interaction with other members of the population.

Under interference, \(Y_i^a\) is not well defined; one must speak of, e.g., “the effect of transplant on Zeus when Hera does not get a new heart,” and the effect may differ for every allocation of hearts. Cox (1958) called the assumption of no interference “no interaction between units”; it is part of Rubin’s (1980) stable-unit-treatment-value assumption (SUTVA). The book assumes no interference unless stated otherwise (Hernán and Robins 2020, 5).

Technical Point 1.1: Causal Effects in the Population

\(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) is the mean counterfactual outcome had everyone received \(a\):

  • discrete outcomes: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_y y \, p_{Y^a}(y)\) with \(p_{Y^a}(y) = \Pr[Y^a = y]\); for dichotomous outcomes \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \Pr[Y^a = 1]\);
  • continuous outcomes: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \int y f_{Y^a}(y)\, dy\);
  • both: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \int y \, dF_{Y^a}(y)\), with \(F_{Y^a}\) the cdf of \(Y^a\).

There is a non-null average causal effect if \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a'}\right]\mathclose{}\) for any two values \(a\) and \(a'\).

A population causal effect can also contrast other functionals (median, variance, hazard, cdf) of the marginal distributions of the counterfactual outcomes. In Table 1.1, \(Y^{a=1}\) and \(Y^{a=0}\) have the same distribution (10 deaths out of 20), so the population causal effect on any functional is zero, e.g. \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{} = 0\).

Unlike the mean, a difference in variances is not in general the variance of the individual effects. Let \(D = Y^{a=1} - Y^{a=0}\), which is \(-1\) for 6 individuals, \(1\) for 6, and \(0\) for 8. Then \[\begin{align} \operatorname{E}\mathopen{}\left[D\right]\mathclose{} &= \frac{6 \times (-1) + 6 \times 1 + 8 \times 0}{20} = 0, \\ \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} &= \frac{6 \times 1 + 6 \times 1 + 8 \times 0}{20} = 0.6, \\ \operatorname{Var}\mathopen{}\left(D\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[D^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[D\right]\mathclose{}\right)\mathclose{}^2 = 0.6 - 0 = 0.6 > 0. \end{align}\] A randomized trial identifies \(\operatorname{Var}\mathopen{}\left(Y^{a=1}\right)\mathclose{} - \operatorname{Var}\mathopen{}\left(Y^{a=0}\right)\mathclose{}\) but not \(\operatorname{Var}\mathopen{}\left(Y^{a=1} - Y^{a=0}\right)\mathclose{}\), because the covariance of \(Y^{a=1}\) and \(Y^{a=0}\) is never observed. The same holds for any nonlinear functional (Hernán and Robins 2020, 6).

3 1.3 Measures of Causal Effect (p. 7)

In Zeus’s family the causal null holds because both counterfactual risks equal 0.5. The causal null can be represented in equivalent ways, each on its own scale.

Definition 9 (Causal risk difference, risk ratio, and odds ratio) For a dichotomous treatment and a dichotomous outcome in a population, with counterfactual risks \(\Pr[Y^{a=1} = 1]\) and \(\Pr[Y^{a=0} = 1]\) (Definition 6):

  • the causal risk difference is \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\);
  • the causal risk ratio is \(\dfrac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]}\), defined when \(\Pr[Y^{a=0} = 1] > 0\);
  • the causal odds ratio is \(\dfrac{\Pr[Y^{a=1} = 1] / \Pr[Y^{a=1} = 0]}{\Pr[Y^{a=0} = 1] / \Pr[Y^{a=0} = 0]}\), defined when both counterfactual risks lie strictly between 0 and 1.

Because they measure the causal effect, these are called effect measures (Hernán and Robins 2020, 7).

Example 11 (Effect measures in Zeus’s family) With both counterfactual risks equal to 0.5 (Example 7):

  • causal risk difference: \(0.5 - 0.5 = 0\);
  • causal risk ratio: \(0.5 / 0.5 = 1\);
  • causal odds ratio: \(\dfrac{0.5 / 0.5}{0.5 / 0.5} = 1\).

Proposition 4 (The causal null on each scale) For a dichotomous treatment and a dichotomous outcome, whenever the measure in question is defined (Definition 9), each of the following statements is equivalent to the causal null hypothesis \(\Pr[Y^{a=1} = 1] = \Pr[Y^{a=0} = 1]\):

  1. the causal risk difference equals 0;

  2. the causal risk ratio equals 1;

  3. the causal odds ratio equals 1.

Proof. Statements (i) and (ii) restate the equality of the two risks. For statement (iii), the odds \(p / (1 - p)\) is strictly increasing in \(p\) on \((0, 1)\), so two risks in \((0, 1)\) have equal odds exactly when they are equal.

When the causal null does not hold (say, smoking and lung cancer), these are not 0, 1, and 1; they quantify the same causal effect on different scales.

Remark 1 (The causal risk ratio is not an average of individual ratios). By Proposition 2, the causal risk difference is the average of the individual effects \(Y^{a=1} - Y^{a=0}\) on the difference scale. The causal risk ratio has no such reading: it measures the causal effect in the population, but, with the deterministic counterfactual outcomes used so far, it is not the average of the individual ratios \(Y^{a=1}/Y^{a=0}\) (Hernán and Robins 2020, 7, margin note). In Zeus’s family, for instance, the individual ratio is \(1/0\) for Zeus and \(0/0\) for Hera, so it is not even defined for them. (Section 1.4 shows that with nondeterministic counterfactual outcomes the causal risk ratio can be written as a weighted average of individual ratio-scale effects.)

Which Measure?

Example 12 (A rare outcome on two scales) Suppose 3 in a million would develop the outcome if treated and 1 in a million if untreated:

  • causal risk ratio \(= \dfrac{3/10^6}{1/10^6} = 3\);
  • causal risk difference \(= \dfrac{3}{10^6} - \dfrac{1}{10^6} = 0.000002\).

Choose the Scale for the Question

Each measure answers a different question. Use the causal risk ratio (multiplicative scale) to say how many times treatment multiplies the risk: in Example 12, it triples it. Use the causal risk difference (additive scale) to count the cases attributable to treatment: in Example 12, 2 per million treated. The choice of scale depends on the goal of the inference (Hernán and Robins 2020, 7).

Definition 10 (Number needed to treat) For a dichotomous treatment and a dichotomous outcome in a population, suppose the treatment reduces the number of cases, that is, the causal risk difference (Definition 9) is negative. The number needed to treat (NNT) is how many individuals, on average, must be treated (\(a = 1\) rather than \(a = 0\)) to avoid one case. It is the reciprocal of the absolute causal risk difference: \[ \text{NNT} = \frac{1}{\mathopen{}\left|\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\right|\mathclose{}}. \tag{5}\] For a treatment with a positive causal risk difference, the reciprocal of the risk difference is the number needed to harm. Like the risk difference, the NNT is specific to the population and the follow-up period it was computed for (Hernán and Robins 2020, 8).

Fine Point 1.2: Number Needed to Treat

In a population of 100 million, suppose 20 million would die within five years if treated and 30 million if untreated. Equivalent summaries:

  • causal risk difference \(= 0.2 - 0.3 = -0.1\);
  • treating all 100 million yields 10 million fewer deaths than treating none;
  • one must treat 100 million to save 10 million lives;
  • on average, one must treat 10 patients to save 1 life.

The last line is the NNT (Definition 10): \(\text{NNT} = 1 / \mathopen{}\left|-0.1\right|\mathclose{}\), which is 10. The NNT was introduced by Laupacis, Sackett, and Roberts (1988) (Hernán and Robins 2020, 8).

4 1.4 Random Variability (pp. 7-9)

Two liberties so far: the immortal Zeus cannot actually die, and real populations are much larger than 20. In practice investigators collect data on a sample of the population of interest, so population risks can only be estimated.

First Source of Random Error: Sampling Variability

Definition 11 (Super-population and sampling variability) A super-population is a population so large that it can be treated as infinite, from which the study individuals are viewed as a random sample. Sampling variability is the random difference between a proportion computed in the sample and the corresponding proportion in the super-population (Hernán and Robins 2020, 8).

Example 13 (Zeus’s family as a sample) View the 20 individuals of Table 1.1 as a random sample from a near-infinite super-population (all immortals, say). The proportion of the sample who would have died untreated is \(10/20 = 0.50\). The super-population probability \(\Pr[Y^{a=0} = 1]\) need not equal it: it could be 0.57, with the sample giving 0.50 because of sampling variability.

Definition 12 (Estimator) An estimator is a rule that turns sample data into a value used as a guess for a super-population quantity. A hat marks it: \(\hat\theta\) is an estimator of \(\theta\) (Hernán and Robins 2020, 8).

Example 14 (The sample proportion as an estimator) When each sampled individual’s \(Y^a\) is known, as in Table 1.1, the sample proportion \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^a = 1\right]\mathclose{}\) is an estimator of the super-population risk \(\Pr[Y^a = 1]\). In Example 13 it gives the estimate \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^{a=0} = 1\right]\mathclose{} = 0.50\) of \(\Pr[Y^{a=0} = 1]\).

Definition 13 (Consistent estimator) An estimator \(\hat\theta\) of \(\theta\) (Definition 12) is consistent if, with probability approaching 1, \(\hat\theta- \theta\) approaches zero as the sample size increases towards infinity: for every \(\epsilon > 0\), \(\Pr\mathopen{}\left[\mathopen{}\left|\hat\theta- \theta\right|\mathclose{} > \epsilon\right]\mathclose{} \to 0\) as the sample size \(n \to \infty\), written \(\hat\theta\overset{\scriptstyle p}{\longrightarrow}\theta\) (Hernán and Robins 2020, 8, margin note).

Example 15 (The sample proportion is consistent) If the sample in Example 13 is a simple random sample from the super-population, the counterfactual outcomes \(Y^{a=0}\) of its members are independent draws with mean \(\Pr[Y^{a=0} = 1]\). By the law of large numbers, \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y^{a=0} = 1\right]\mathclose{} \overset{\scriptstyle p}{\longrightarrow}\Pr[Y^{a=0} = 1]\): with 20 individuals the estimate can be off by 0.07, but the chance of an error that large shrinks toward 0 as the sample grows.

Two Meanings of Consistency

“Consistency” of an estimator (Definition 13) has nothing to do with “consistency” of counterfactual outcomes (Definition 4) (Hernán and Robins 2020, 9, margin note). The first is a property of a statistical procedure as the sample grows; the second links observed outcomes to counterfactual ones.

Second Source of Random Error: Nondeterministic Counterfactuals

Definition 14 (Deterministic and nondeterministic counterfactuals) For a dichotomous outcome, an individual’s counterfactual outcomes are deterministic if each \(Y^a\) has a fixed value for that individual, so the individual’s probability of the outcome under each treatment value is 0 or 1. They are nondeterministic (stochastic) if, for some treatment value \(a\), the individual’s probability of \(Y^a = 1\) lies strictly between 0 and 1 (Hernán and Robins 2020, 9).

Example 16 (Zeus’s mortality coins) So far Zeus’s counterfactual outcomes have been deterministic: a 100% chance of dying if treated and 0% if untreated. Alternatively, Zeus could have a 90% chance of dying if treated and 10% if untreated. His counterfactual outcomes would then be nondeterministic, and the values in Table 1.1 would be realizations of “random flips of mortality coins.” These probabilities would likely vary across individuals.

Simplifying Assumptions Until Chapter 10

Random error comes from sampling variability, nondeterministic counterfactuals, or both.

Assumed Until Chapter 10

To set random error aside, the book assumes until Chapter 10 that:

  • counterfactual outcomes are deterministic (Definition 14);
  • we have data on every individual in a very large super-population (Definition 11).

The second assumption amounts to letting the 20 individuals stand for 20 billion, with 1 billion identical to Zeus, 1 billion identical to Hera, and so on.

Technical Point 1.2: Nondeterministic Counterfactuals

For nondeterministic counterfactuals, \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_y y \, p_{Y^a}(y)\), where \(p_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[Q_{Y^a}(\cdot)\right]\mathclose{}\) and \(Q_{Y^a}(y)\) is an individual’s random probability of outcome \(y\) under \(a\) (in the text, \(Q_{Y^{a=1}}(1) = 0.9\) for Zeus).

More generally, each individual has a distribution \(\Theta_{Y^a}(\cdot)\) of \(Y^a\), a random cdf. Then \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^a \mid \Theta_{Y^a}(\cdot)\right]\mathclose{}\right]\mathclose{} && \text{(law of total expectation)} \\ &= \operatorname{E}\mathopen{}\left[\int y \, d\Theta_{Y^a}(y)\right]\mathclose{} && \text{(mean of the individual distribution)} \\ &= \int y \, d\,\operatorname{E}\mathopen{}\left[\Theta_{Y^a}(y)\right]\mathclose{} && \text{(exchange expectation and integral)} \\ &= \int y \, dF_{Y^a}(y), && \text{with } F_{Y^a}(\cdot) = \operatorname{E}\mathopen{}\left[\Theta_{Y^a}(\cdot)\right]\mathclose{}. \end{align}\]

For binary nondeterministic outcomes, the causal risk ratio \(\operatorname{E}\mathopen{}\left[Q_{Y^{a=1}}(1)\right]\mathclose{} / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\) equals the weighted average \(\operatorname{E}\mathopen{}\left[W \, Q_{Y^{a=1}}(1) / Q_{Y^{a=0}}(1)\right]\mathclose{}\) of the individual ratio-scale effects, with weights \(W = Q_{Y^{a=0}}(1) / \operatorname{E}\mathopen{}\left[Q_{Y^{a=0}}(1)\right]\mathclose{}\), provided \(Q_{Y^{a=0}}(1)\) is never 0 (Hernán and Robins 2020, 10).

5 1.5 Causation versus Association (pp. 9-12)

Real data do not look like Table 1.1: we observe only one of each individual’s counterfactual outcomes, the one for the treatment actually received. We observe the treatment \(A\) and the outcome \(Y\), as in Table 1.2.

Table 1.2: Observed treatment and outcome in Zeus’s family (Hernán and Robins 2020, 9)

Name \(A\) \(Y\)
Rheia 0 0
Kronos 0 1
Demeter 0 0
Hades 0 0
Hestia 1 0
Poseidon 1 0
Hera 1 0
Zeus 1 1
Artemis 0 1
Apollo 0 1
Leto 0 0
Ares 1 1
Athena 1 1
Hephaestus 1 1
Aphrodite 1 1
Polyphemus 1 1
Persephone 1 1
Hermes 1 0
Hebe 1 0
Dionysus 1 0

Definition 15 (Risk among those who received a treatment value) For a dichotomous outcome \(Y\) and observed treatment \(A\), the conditional probability \(\Pr[Y = 1 \mid A = a]\) is the proportion who developed the outcome among those in the population who happened to receive treatment value \(a\) (Hernán and Robins 2020, 10).

Example 17 (Risks in the treated and the untreated in Zeus’s family) From Table 1.2:

  • 7 of the 13 treated died: \(\Pr[Y = 1 \mid A = 1] = 7/13\);
  • 3 of the 7 untreated died: \(\Pr[Y = 1 \mid A = 0] = 3/7\).

Independence and Association

Definition 16 (Independence) For a dichotomous treatment \(A\) and a dichotomous outcome \(Y\), \(A\) and \(Y\) are independent (\(A\) is not associated with \(Y\), or \(A\) does not predict \(Y\)) when \(\Pr[Y = 1 \mid A = 1] = \Pr[Y = 1 \mid A = 0]\) (Definition 15). Independence is written \(Y \perp\!\!\!\perp A\), or equivalently \(A \perp\!\!\!\perp Y\). \(A\) and \(Y\) are associated (dependent) when \(\Pr[Y = 1 \mid A = 1] \neq \Pr[Y = 1 \mid A = 0]\) (Hernán and Robins 2020, 10).

Example 18 (Treatment and outcome are associated in Zeus’s family) In Example 17, \(\Pr[Y = 1 \mid A = 1] = 7/13 \approx 0.54\) and \(\Pr[Y = 1 \mid A = 0] = 3/7 \approx 0.43\). These differ, so \(A\) and \(Y\) are associated: \(A \perp\!\!\!\perp Y\) does not hold.

Definition 17 (Associational risk difference, risk ratio, and odds ratio) For a dichotomous treatment \(A\) and a dichotomous outcome \(Y\), with conditional risks \(\Pr[Y = 1 \mid A = a]\) (Definition 15):

  • the associational risk difference is \(\Pr[Y = 1 \mid A = 1] - \Pr[Y = 1 \mid A = 0]\);
  • the associational risk ratio is \(\dfrac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]}\), defined when \(\Pr[Y = 1 \mid A = 0] > 0\);
  • the associational odds ratio is \(\dfrac{\Pr[Y = 1 \mid A = 1] / \Pr[Y = 0 \mid A = 1]}{\Pr[Y = 1 \mid A = 0] / \Pr[Y = 0 \mid A = 0]}\), defined when both conditional risks lie strictly between 0 and 1.

Together they are called association measures. Whenever the measure in question is defined, \(A \perp\!\!\!\perp Y\) (Definition 16) holds if and only if the associational risk difference is 0, if and only if the associational risk ratio is 1, and if and only if the associational odds ratio is 1, by the argument of Proposition 4.

Example 19 (Association measures in Zeus’s family) With the risks of Example 17: \[\begin{align} \text{risk difference} &= \frac{7}{13} - \frac{3}{7} = \frac{49 - 39}{91} = \frac{10}{91} \approx 0.11, \\ \text{risk ratio} &= \frac{7/13}{3/7} = \frac{7 \times 7}{13 \times 3} = \frac{49}{39} \approx 1.26, \\ \text{odds ratio} &= \frac{(7/13)/(6/13)}{(3/7)/(4/7)} = \frac{7/6}{3/4} = \frac{28}{18} \approx 1.56 . \end{align}\]

Causation Versus Association

Example 20 (No causation, yet association) In the same population of 20:

  • no causation: risk if all treated \(= 10/20\) equals risk if all untreated \(= 10/20\) (Example 8);
  • association: risk in the 13 treated \(= 7/13\) is greater than risk in the 7 untreated \(= 3/7\) (Example 18).

Remark 2 (Two different contrasts). Causation (Definition 7) and association (Definition 16) compare different things (Hernán and Robins 2020, 11–12):

Causation Association
Question “What would the risk be if everybody had been treated / untreated?” “What is the risk in the treated / the untreated?”
World counterfactual actual
Risk marginal \(\Pr[Y^a = 1]\), in the whole population conditional \(\Pr[Y = 1 \mid A = a]\), in the subset with \(A = a\)
Comparison same population, two treatment values two disjoint subsets defined by actual treatment

Association Is Not Causation

The differences listed in Remark 2 are why “association is not causation”: an association measure (Definition 17) can differ from the corresponding effect measure (Definition 9), as it does in Example 20. The book often writes the redundant “causal effect” to keep “effect” from being read as association (Hernán and Robins 2020, 12).

Example 21 (Acting on association instead of causation) Suppose the causal risk ratio of 5-year mortality for aspirin versus no aspirin is 0.5, but the associational risk ratio is 1.5 because people at high cardiovascular risk are preferentially prescribed aspirin. A physician who stops prescribing aspirin because its users die more often is acting on the association, and withholds a treatment that would halve the risk (Hernán and Robins 2020, 12).

The Question for the Rest of the Book

Under Which Conditions Can Real-World Data Be Used for Causal Inference?

Causal inference needs data like the hypothetical Table 1.1, but real data look like Table 1.2. The rest of the book asks when, and how, data like Table 1.2 can answer causal questions. Chapter 2 gives one answer: conduct a randomized experiment.

6 Summary

  1. Individual causal effect: the contrast \(Y^{a=1} - Y^{a=0}\) for an individual, present when \(Y^{a=1} \neq Y^{a=0}\); not identified, because only one counterfactual outcome is observed (consistency: \(Y = Y^A\)).
  2. Average causal effect: the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{}\) in a population, present when \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\); a null average effect does not rule out individual effects (the sharp null is stronger).
  3. Effect measures: causal risk difference, risk ratio, and odds ratio quantify one effect on different scales; for a beneficial treatment, the NNT is the reciprocal of the absolute risk difference.
  4. Random variability: from sampling and possibly from nondeterministic counterfactuals; ignored until Chapter 10.
  5. Causation vs. association: causation compares the whole population under two treatment values; association compares two subsets defined by the treatment actually received.

7 References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.