Does your looking up at the sky make other pedestrians look up too? The question has the parts of any causal question: an action (your looking up), an outcome (others looking up), and a population (say, residents of Madrid in 2019). A natural design: flip a coin each time someone approaches, look up on heads and straight ahead on tails, repeat a few thousand times, and compare the proportions of pedestrians who look up within 10 seconds. Suppose 55% looked up when you did and 1% when you did not.
That design is a randomized experiment: an experiment because the investigator carries out the action, and randomized because a random device (the coin) decides whom to act on. A nonrandomized experiment, e.g. looking up only when a man approaches, would be much less convincing: critics could say that men and women differ in how often they look up, so the comparison is between “noncomparable” groups. This chapter explains why randomization supports convincing causal inferences.
In a real study we never know both of Zeus’s potential outcomes, only his observed outcome under the treatment he received. Table 2.1 adds the counterfactual columns to the observed data of Table 1.2; “?” marks the counterfactual outcome that is missing.
Table 2.1: Observed data with counterfactual outcomes, Zeus’s family (Hernán and Robins 2020, 14)
| Name | \(A\) | \(Y\) | \(Y^0\) | \(Y^1\) |
|---|---|---|---|---|
| Rheia | 0 | 0 | 0 | ? |
| Kronos | 0 | 1 | 1 | ? |
| Demeter | 0 | 0 | 0 | ? |
| Hades | 0 | 0 | 0 | ? |
| Hestia | 1 | 0 | ? | 0 |
| Poseidon | 1 | 0 | ? | 0 |
| Hera | 1 | 0 | ? | 0 |
| Zeus | 1 | 1 | ? | 1 |
| Artemis | 0 | 1 | 1 | ? |
| Apollo | 0 | 1 | 1 | ? |
| Leto | 0 | 0 | 0 | ? |
| Ares | 1 | 1 | ? | 1 |
| Athena | 1 | 1 | ? | 1 |
| Hephaestus | 1 | 1 | ? | 1 |
| Aphrodite | 1 | 1 | ? | 1 |
| Polyphemus | 1 | 1 | ? | 1 |
| Persephone | 1 | 1 | ? | 1 |
| Hermes | 1 | 0 | ? | 0 |
| Hebe | 1 | 0 | ? | 0 |
| Dionysus | 1 | 0 | ? | 0 |
With half the counterfactual outcomes missing, Table 2.1 by itself only supports association measures. Randomized experiments also have missing counterfactuals, but randomization ensures the values are missing by chance, so effect measures can be consistently estimated despite the missing data.
Take the near-infinite population drawn as a diamond in Figure 1.1. For each individual, flip a coin with probability of heads below 50%:
Five days later: \(\Pr[Y = 1 \mid A = 1] = 0.3\) and \(\Pr[Y = 1 \mid A = 0] = 0.6\), so the associational risk ratio is \(0.3/0.6 = 0.5\) and the associational risk difference is \(0.3 - 0.6 = -0.3\).
Suppose the assistants mistakenly treated the grey group instead of the white group. Our conclusions would not change: the risk in the treated (now grey) would still be expected to be 0.3 and in the untreated 0.6. When group membership is randomized, which group got the treatment is irrelevant. The groups are exchangeable.
Definition 1 (Exchangeability) The treated and the untreated are exchangeable if the risk under each potential treatment value \(a\) is the same in both groups: \[ \Pr[Y^a = 1 \mid A = 1] = \Pr[Y^a = 1 \mid A = 0] \quad \text{for } a = 0 \text{ and } a = 1. \] Equivalently, the counterfactual outcome is independent of the actual treatment: \(Y^a \perp\!\!\!\perp A\) for all \(a\).
If a conditional risk is the same in every treatment group, it equals the marginal risk: \[ \Pr[Y^a = 1 \mid A = 1] = \Pr[Y^a = 1 \mid A = 0] = \Pr[Y^a = 1]. \] So the actual treatment \(A\) does not predict the counterfactual outcome \(Y^a\). Randomization is valued because it is expected to produce exchangeability.
The risk under treatment in the white (treated) group is not counterfactual at all, because that group was treated. Hence \[\begin{align} \Pr[Y^{a=1} = 1] &= \Pr[Y^{a=1} = 1 \mid A = 1] && \text{(exchangeability)} \\ &= \Pr[Y = 1 \mid A = 1] && \text{(consistency: } Y = Y^{a=1} \text{ when } A = 1\text{)} \\ &= 0.3 . \end{align}\] The same steps with \(a = 0\) give \(\Pr[Y^{a=0} = 1] = \Pr[Y = 1 \mid A = 0] = 0.6\). The causal risk ratio is \(0.3/0.6 = 0.5\) and the causal risk difference is \(0.3 - 0.6 = -0.3\), the same as the associational measures.
Technical Point 2.1: Full Exchangeability and Mean Exchangeability
Let \(\mathcal{A} = \{a, a', a'', \ldots\}\) be the set of treatment values in the population and \(Y^{\mathcal{A}} = \{Y^a, Y^{a'}, Y^{a''}, \ldots\}\) the set of all counterfactual outcomes. Randomization makes \(Y^{\mathcal{A}} \perp\!\!\!\perp A\), full exchangeability (for a dichotomous treatment, \(\{Y^{a=1}, Y^{a=0}\} \perp\!\!\!\perp A\)). Full exchangeability implies, but is not implied by, \(Y^a \perp\!\!\!\perp A\) for each \(a\).
Mean exchangeability is \(\operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) for all \(a, a'\). For a dichotomous outcome it is the same as exchangeability; for a continuous outcome, exchangeability implies mean exchangeability but not conversely (e.g., the variance of \(Y^a\) may depend on treatment).
Mean exchangeability is all that is needed for \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = a\right]\mathclose{}\): \[\begin{align} \operatorname{E}\mathopen{}\left[Y \mid A = a\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^a \mid A = a\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} && \text{(mean exchangeability)}. \end{align}\] (Hernán and Robins 2020, 15)
In a randomized experiment with exchangeability and a causal effect, \(Y \perp\!\!\!\perp A\) fails: treatment is associated with the observed outcome.
Check \(Y^a \perp\!\!\!\perp A\) for \(a = 0\), pretending we can see Table 1.1. Among the 13 treated, \(Y^{a=0} = 1\) for Poseidon, Ares, Athena, Persephone, Hermes, Hebe, and Dionysus; among the 7 untreated, for Kronos, Artemis, and Apollo: \[ \Pr[Y^{a=0} = 1 \mid A = 1] = \frac{7}{13} \approx 0.54 \;>\; \frac{3}{7} \approx 0.43 = \Pr[Y^{a=0} = 1 \mid A = 0]. \] The treated have a worse prognosis than the untreated: exchangeability does not hold (it also fails for \(a = 1\): \(7/13\) in the treated vs. \(3/7\) in the untreated, the latter from Rheia, Artemis, and Leto).
Fine Point 2.1: Crossover Experiments
Suppose Zeus calls a lightning strike (\(a = 1\)) yesterday and his blood pressure rises; today he refrains (\(a = 0\)) and it does not. Having seen “both” outcomes, can we conclude that lightning-bolt use affects his blood pressure? Only under strong assumptions.
In a crossover experiment, individual \(i\) receives treatment \(A_{it}\) in periods \(t = 0, 1\). Let \(Y_{i1}^{a_0, a_1}\) be the outcome at \(t = 1\) under \(a_0\) at \(t = 0\) and \(a_1\) at \(t = 1\), and \(Y_{i0}^{a_0}\) the outcome at \(t = 0\). The individual effect \(Y_{it}^{a_t = 1} - Y_{it}^{a_t = 0}\) is identified if:
If \(A_{i1} = 1\) and \(A_{i0} = 0\), then \[\begin{align} Y_{i1} - Y_{i0} &= Y_{i1}^{a_1 = 1} - Y_{i0}^{a_0 = 0} && \text{(consistency and (i))} \\ &= \mathopen{}\left(Y_{i1}^{a_1 = 1} - Y_{i1}^{a_1 = 0}\right)\mathclose{} + \mathopen{}\left(Y_{i1}^{a_1 = 0} - Y_{i0}^{a_0 = 0}\right)\mathclose{} && \text{(add and subtract } Y_{i1}^{a_1 = 0}\text{)} \\ &= \alpha_i + (\beta_i - \beta_i) && \text{(by (ii) and (iii))} \\ &= \alpha_i . \end{align}\] Symmetrically, if \(A_{i1} = 0\) and \(A_{i0} = 1\), then \(Y_{i0} - Y_{i1} = \alpha_i\).
Condition (i) requires an outcome with abrupt onset that fully resolves by the next period, so crossover designs cannot study an irreversible action (heart transplant) on an irreversible outcome (death) (Hernán and Robins 2020, 16).
Table 2.2 adds a prognostic factor \(L\) (1 if in critical condition, 0 otherwise), measured before treatment was assigned.
Table 2.2: Heart transplant study with prognostic factor \(L\) (Hernán and Robins 2020, 17)
| Name | \(L\) | \(A\) | \(Y\) |
|---|---|---|---|
| Rheia | 0 | 0 | 0 |
| Kronos | 0 | 0 | 1 |
| Demeter | 0 | 0 | 0 |
| Hades | 0 | 0 | 0 |
| Hestia | 0 | 1 | 0 |
| Poseidon | 0 | 1 | 0 |
| Hera | 0 | 1 | 0 |
| Zeus | 0 | 1 | 1 |
| Artemis | 1 | 0 | 1 |
| Apollo | 1 | 0 | 1 |
| Leto | 1 | 0 | 0 |
| Ares | 1 | 1 | 1 |
| Athena | 1 | 1 | 1 |
| Hephaestus | 1 | 1 | 1 |
| Aphrodite | 1 | 1 | 1 |
| Polyphemus | 1 | 1 | 1 |
| Persephone | 1 | 1 | 1 |
| Hermes | 1 | 1 | 0 |
| Hebe | 1 | 1 | 0 |
| Dionysus | 1 | 1 | 0 |
Two mutually exclusive designs could have produced these data:
Design 1 is a marginally randomized experiment (one unconditional randomization probability for all). Design 2 is a conditionally randomized experiment (randomization probabilities depend on \(L\)).
A marginally randomized experiment is expected to give exchangeability, \(Y^a \perp\!\!\!\perp A\). A conditionally randomized one generally does not, because each treatment group may have a different share of individuals with a bad prognosis.
In Table 2.2:
The treated would have had a higher risk than the untreated had they remained untreated: \(A\) predicts \(Y^{a=0}\), so \(Y^a \perp\!\!\!\perp A\) fails. Since the study was randomized, it must have been randomized conditional on \(L\) (design 2).
A conditionally randomized experiment is two marginally randomized experiments, one in \(L = 1\) and one in \(L = 0\). Within each stratum, the treated and the untreated are exchangeable: \[ \Pr[Y^a = 1 \mid A = 1, L = 1] = \Pr[Y^a = 1 \mid A = 0, L = 1], \quad \text{i.e. } Y^a \perp\!\!\!\perp A \mid L = 1, \] and likewise \(Y^a \perp\!\!\!\perp A \mid L = 0\).
Definition 2 (Conditional exchangeability) \(Y^a \perp\!\!\!\perp A \mid L\) for all \(a\): within every level \(l\) of \(L\), the counterfactual outcome is independent of the actual treatment, \(Y^a \perp\!\!\!\perp A \mid L = l\).
If Table 2.2 had come from a marginally randomized experiment, the causal risk ratio would simply be the associational one: \[ \frac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = \frac{7/13}{3/7} = \frac{49}{39} \approx 1.26 . \] Under conditional randomization we have two options:
Fine Point 2.2: Risk Periods
A risk is the proportion developing the outcome during a specified period (e.g., 5-day mortality); the book often states the period once and then omits it.
Why the period matters: in a randomized trial of antibiotics for elderly people with plague, one investigator reports a causal risk ratio of 0.05 and another a ratio of 1. Both are right: the first used 1-year risks, the second 100-year risks, which are 1 under either treatment. A treatment “having a causal effect on mortality” means death is delayed, not prevented (Hernán and Robins 2020, 19).
In our conditionally randomized study, hearts were assigned with probability 50% to the 8 individuals with \(L = 0\) and 75% to the 12 with \(L = 1\).
From Table 2.2:
| \(A = 1\) | \(A = 0\) | |
|---|---|---|
| \(L = 0\) (8 individuals) | \(\Pr[Y = 1 \mid L = 0, A = 1] = 1/4\) | \(\Pr[Y = 1 \mid L = 0, A = 0] = 1/4\) |
| \(L = 1\) (12 individuals) | \(\Pr[Y = 1 \mid L = 1, A = 1] = 6/9 = 2/3\) | \(\Pr[Y = 1 \mid L = 1, A = 0] = 2/3\) |
Because of conditional exchangeability, each observed stratum risk equals the corresponding counterfactual risk, e.g. \(\Pr[Y = 1 \mid L = 0, A = 1] = \Pr[Y^{a=1} = 1 \mid L = 0]\).
The risk if all 20 had been treated is a weighted average of \(1/4\) (in \(L = 0\)) and \(2/3\) (in \(L = 1\)), with weights \(\Pr[L = 0] = 8/20 = 0.4\) and \(\Pr[L = 1] = 12/20 = 0.6\): \[\begin{align} \Pr[Y^{a=1} = 1] &= \tfrac{1}{4} \times 0.4 + \tfrac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5, \\ \Pr[Y^{a=0} = 1] &= \tfrac{1}{4} \times 0.4 + \tfrac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5 . \end{align}\] The causal risk ratio is \(0.5/0.5 = 1\).
\[\begin{align} \Pr[Y^a = 1] &= \sum_l \Pr[Y^a = 1 \mid L = l] \Pr[L = l] && \text{(law of total probability)} \\ &= \sum_l \Pr[Y = 1 \mid L = l, A = a] \Pr[L = l] && \text{(conditional exchangeability, consistency)} \end{align}\] The left side is an unobserved counterfactual risk; the right side uses only observed quantities (\(L\), \(A\), \(Y\)).
Definition 3 (Identification) A counterfactual quantity is identified (identifiable) if it can be expressed as a function of the distribution of the observed data; otherwise it is unidentified.
Definition 4 (Standardization) The standardized mean of \(Y\) for treatment level \(a\), using the population as the standard, is \[ \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \times \Pr[L = l]. \] Under conditional exchangeability it equals the counterfactual mean \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\), i.e. the mean that would have been observed had everyone in the population received \(a\).
So the causal risk ratio can be computed by standardization as \[ \frac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]} = \frac{\sum_l \Pr[Y = 1 \mid L = l, A = 1] \Pr[L = l]}{\sum_l \Pr[Y = 1 \mid L = l, A = 0] \Pr[L = l]} . \]
The same causal risk ratio can be computed by inverse probability (IP) weighting.
Display Table 2.2 as a tree, from left to right:
So \(\Pr[Y^{a=0} = 1] = (2 + 8)/(8 + 12) = 10/20 = 0.5\). This relies on exchangeability within each stratum: the treated, had they been untreated, would have had the same risk as the untreated in their stratum.
So \(\Pr[Y^{a=1} = 1] = (2 + 8)/20 = 0.5\), and the causal risk ratio is \(0.5/0.5 = 1\), as with standardization.
The two simulated trees (Figure 2.2: everybody untreated; everybody treated) pooled together form a pseudo-population twice the size of the original, in which every individual appears once treated and once untreated (Figure 2.3).
In the pseudo-population, \(L\) is independent of \(A\), so under \(Y^a \perp\!\!\!\perp A \mid L\) in the original population the treated and untreated are unconditionally exchangeable. Hence the associational risk ratio in the pseudo-population equals the causal risk ratio in both the pseudo-population and the original population.
Each individual is weighted by the inverse of the probability of receiving the treatment they actually received:
| Group | \(\Pr[A = a \mid L]\) | Weight | Count | Pseudo-population size |
|---|---|---|---|---|
| \(L = 0\), \(A = 0\) | 0.5 | \(1/0.5 = 2\) | 4 | 8 |
| \(L = 0\), \(A = 1\) | 0.5 | \(1/0.5 = 2\) | 4 | 8 |
| \(L = 1\), \(A = 0\) | 0.25 | \(1/0.25 = 4\) | 3 | 12 |
| \(L = 1\), \(A = 1\) | 0.75 | \(1/0.75 \approx 1.33\) | 9 | 12 |
Definition 5 (IP weights) \[ W^A = \frac{1}{f[A \mid L]}, \] where \(f[a \mid l] = \Pr[A = a \mid L = l]\) for discrete \(A\) and \(L\). A treated individual with \(L = l\) gets \(1/\Pr[A = 1 \mid L = l]\); an untreated one with \(L = l'\) gets \(1/\Pr[A = 0 \mid L = l']\).
Technical Point 2.2: Formal Definition of IP Weights
The weight’s denominator is the conditional density of \(A\) given \(L\) evaluated at the individual’s own values, i.e. at the random arguments \(A\) and \(L\): \(f[A \mid L]\). This notation is needed because \(\Pr[A = A \mid L = L]\) is tautologically 1. In a conditionally randomized experiment, \(f[a \mid l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\).
The mean outcome in the pseudo-population equals the IP weighted mean in the population. With \(I(A = a) = 1\) if \(A = a\) and 0 otherwise: \[\begin{align} \operatorname{E}_{ps}\mathopen{}\left[Y \mid A = a\right]\mathclose{} &= \frac{\operatorname{E}_{ps}\mathopen{}\left[Y\, I(A = a)\right]\mathclose{}}{\operatorname{E}_{ps}\mathopen{}\left[I(A = a)\right]\mathclose{}} && \text{(laws of probability)} \\ &= \frac{\operatorname{E}\mathopen{}\left[W^A Y\, I(A = a)\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[I(A = a)\, W^A\right]\mathclose{}} && \text{(definition of } \text{E}_{ps}\text{)} \\ &= \frac{\operatorname{E}\mathopen{}\left[Y\, I(A = a) / \Pr[A = a \mid L]\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[I(A = a) / \Pr[A = a \mid L]\right]\mathclose{}} && \text{(} I(A = a)/f[A \mid L] = I(A = a)/f[a \mid L]\text{)} \\ &= \operatorname{E}\mathopen{}\left[\frac{Y\, I(A = a)}{\Pr[A = a \mid L]}\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[I(A = a)/\Pr[A = a \mid L] \mid L\right]\mathclose{} = 1\text{)}. \end{align}\] (Hernán and Robins 2020, 23)
Both gave a causal risk ratio of 1 here, and this is no coincidence (Technical Point 2.3). Both build a counterfactual tree in which everyone receives \(a\), from different pieces:
Because both simulate what would have happened had \(L\) not been used to decide the probability of treatment, we say these methods adjust for \(L\). “Controlling for” \(L\) is a slight abuse of language: this analytic control is quite different from the physical control of a randomized experiment.
Technical Point 2.3: Equivalence of IP Weighting and Standardization
Let \(A\) be discrete with finitely many values and assume positivity: \(f[a \mid l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\) (guaranteed in conditionally randomized experiments). Define
They are equal (no counterfactuals needed): \[\begin{align} \operatorname{E}\mathopen{}\left[\frac{I(A = a)\, Y}{f[A \mid L]}\right]\mathclose{} &= \sum_l \frac{1}{f[a \mid l]} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} f[a \mid l] \Pr[L = l] && \text{(only } A = a \text{ contributes)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(cancel } f[a \mid l]\text{)}. \end{align}\] For continuous \(L\), replace the sum by an integral.
Under conditional exchangeability, both equal \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\). Standardized mean: \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(conditional exchangeability, positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(consistency)}. \end{align}\] IP weighted mean: \[\begin{align} \operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[A \mid L]} Y\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} Y^a\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} Y^a \;\middle|\; L\right]\mathclose{}\right]\mathclose{} && \text{(law of total expectation; } f[a \mid L] > 0\text{)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} \;\middle|\; L\right]\mathclose{} \operatorname{E}\mathopen{}\left[Y^a \mid L\right]\mathclose{}\right]\mathclose{} && \text{(conditional exchangeability)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^a \mid L\right]\mathclose{}\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[I(A = a)/f[a \mid L] \mid L\right]\mathclose{} = 1\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} . \end{align}\]
With a continuous treatment the IP weighted mean is no longer equal to the standardized mean and is biased for \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) even under exchangeability: with \(f\) a density, \(\operatorname{E}\mathopen{}\left[I(A = a)/f(a \mid l) \mid L = l\right]\mathclose{} = 0\), not 1; with \(f\) a probability, \(f(a \mid L = l) = 0\) with probability 1 and positivity fails. Section 12.4 generalizes IP weighting to continuous treatments; Technical Point 3.1 shows the results fail without positivity even for discrete \(A\) (Hernán and Robins 2020, 25).
An ideal randomized experiment plus standardization or IP weighting gives average causal effects. But randomized experiments are often unethical, impractical, or untimely. For the heart transplant study:
Frequently an observational study is the least bad option (Chapter 3).