In Chapter 2, randomization gave us exchangeability by design. Many scientific questions, however, are answered with observational studies, in which the investigators observe and record the relevant data but do not assign treatment. Much of what we know (evolution, plate tectonics, the fact that hot coffee can burn) comes from observation rather than experiment.
This chapter asks: under what conditions can an observational study support a valid causal inference?
In an ideal randomized experiment, an associational risk ratio of, say, 0.7 is expected to equal the causal risk ratio, because randomization makes the treated and untreated exchangeable.
In an observational study of heart transplant where sicker patients were more likely to be transplanted, an associational risk ratio of 1.1 may be a compromise between:
An observational study can be conceptualized as a conditionally randomized experiment if:
These are the identifiability conditions (or identifiability assumptions).
Table 3.1 contains the same data as Table 2.2: \(L\) is critical condition at baseline (1: yes), \(A\) is heart transplant, and \(Y\) is death.
| Name | \(L\) | \(A\) | \(Y\) | Name | \(L\) | \(A\) | \(Y\) |
|---|---|---|---|---|---|---|---|
| Rheia | 0 | 0 | 0 | Leto | 1 | 0 | 0 |
| Kronos | 0 | 0 | 1 | Ares | 1 | 1 | 1 |
| Demeter | 0 | 0 | 0 | Athena | 1 | 1 | 1 |
| Hades | 0 | 0 | 0 | Hephaestus | 1 | 1 | 1 |
| Hestia | 0 | 1 | 0 | Aphrodite | 1 | 1 | 1 |
| Poseidon | 0 | 1 | 0 | Polyphemus | 1 | 1 | 1 |
| Hera | 0 | 1 | 0 | Persephone | 1 | 1 | 1 |
| Zeus | 0 | 1 | 1 | Hermes | 1 | 1 | 0 |
| Artemis | 1 | 0 | 1 | Hebe | 1 | 1 | 0 |
| Apollo | 1 | 0 | 1 | Dionysus | 1 | 1 | 0 |
If these data came from an observational study in which the three identifiability conditions held, we would compute the same causal risk ratio as in Chapter 2: 1.
Example 1 (Standardized risks in Table 3.1) Stratum-specific risks from Table 3.1 (8 individuals with \(L = 0\), 12 with \(L = 1\)):
with \(\Pr[L = 0] = 8/20 = 0.4\) and \(\Pr[L = 1] = 12/20 = 0.6\). Standardizing,
\[ \begin{aligned} \Pr[Y^{a=1} = 1] &= \Pr[Y = 1 \mid A = 1, L = 0]\Pr[L = 0] + \Pr[Y = 1 \mid A = 1, L = 1]\Pr[L = 1] \\ &= (1/4)(0.4) + (2/3)(0.6) \\ &= 0.1 + 0.4 = 0.5, \end{aligned} \]
\[ \begin{aligned} \Pr[Y^{a=0} = 1] &= \Pr[Y = 1 \mid A = 0, L = 0]\Pr[L = 0] + \Pr[Y = 1 \mid A = 0, L = 1]\Pr[L = 1] \\ &= (1/4)(0.4) + (2/3)(0.6) \\ &= 0.1 + 0.4 = 0.5, \end{aligned} \]
so the causal risk ratio is \(0.5 / 0.5 = 1\).
Definition 1 (Identifiability) An average causal effect is (nonparametrically) identifiable under a set of assumptions if those assumptions imply that the distribution of the observed data is compatible with a single value of the effect measure. It is nonidentifiable if the observed data distribution is compatible with several values of the effect measure (Hernán and Robins 2020, Fine Point 3.1, p. 29).
Fine Point 3.1: Identifiability of causal effects
If Table 3.1 came from a conditionally randomized experiment, \(Y^a \perp\!\!\!\perp A \mid L\) holds by design and the causal risk ratio of 1 is identified with no further assumptions. If it came from an observational study, the risk ratio equals 1 only if we add the identifying assumption \(Y^a \perp\!\!\!\perp A \mid L\), which is external to the data. Without it, the same data are compatible with a causal risk ratio
When any of the three conditions fails, the analogy with a conditionally randomized experiment breaks down. Other approaches rely on different identifiability conditions; for example, instrumental variable methods (Chapter 16) assume that a predictor of treatment behaves as if randomly assigned conditional on measured covariates.
In a marginally randomized experiment, the treated and untreated are exchangeable (\(Y^a \perp\!\!\!\perp A\)) because randomization balances the independent predictors of the outcome between groups.
An independent predictor of the outcome is a covariate associated with \(Y\) within levels of treatment; for dichotomous outcomes these are often called risk factors.
In Table 3.1, 69% of the treated (\(9/13\)) but only 43% of the untreated (\(3/7\)) were in critical condition (\(L = 1\)), so marginal exchangeability does not hold.
Definition 2 (Conditional exchangeability) The treated and untreated are conditionally exchangeable given \(L\) when
\[Y^a \perp\!\!\!\perp A \mid L \quad \text{for all } a.\]
Equivalently, for a dichotomous outcome, \(\Pr[Y^a = 1 \mid A = 1, L = l] = \Pr[Y^a = 1 \mid A = 0, L = l]\) for all \(a\) and \(l\).
Conditional exchangeability holds in a conditionally randomized experiment because, within levels of \(L\), all other outcome predictors are equally distributed between the treated and the untreated.
Heart transplants are allocated by HLA compatibility, so HLA genes are unequally distributed between treatment groups. But HLA genes do not predict mortality given \(L\) and \(A\), so treatment is still effectively random within levels of \(L\), and HLA need not be included in the analysis.
Suppose, unknown to the investigators, doctors prefer to transplant hearts into nonsmokers. Let \(U\) denote smoking (an unmeasured variable). Within the stratum \(L = 1\), smokers (\(U = 1\)) are less likely to be treated, so smoking is less common among the treated than the untreated, and \(Y^a \perp\!\!\!\perp A \mid L\) fails.
Remark 1 (Conditional exchangeability cannot be verified). Conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) fails whenever there is an unmeasured independent predictor \(U\) of the outcome such that the probability of treatment depends on \(U\) within strata of \(L\). Even when it holds, it cannot be checked: verifying it would require comparing \(\Pr[Y^a = 1 \mid A = a, L = l]\) with \(\Pr[Y^a = 1 \mid A \neq a, L = l]\), but \(Y^a\) is unknown for individuals with \(A \neq a\).
An experiment that assigned everyone to \(A = 1\) (or everyone to \(A = 0\)) could not estimate the average causal effect. Treatment must be assigned so that each treatment level has a positive probability.
Definition 3 (Positivity) \[\Pr[A = a \mid L = l] > 0 \quad \text{for all values } l \text{ with } \Pr[L = l] \neq 0 \text{ in the population of interest,}\]
for every treatment value \(a\) involved in the causal contrast.
Example 2 (A positivity violation) If doctors always transplanted a heart to individuals in critical condition, then \(\Pr[A = 0 \mid L = 1] = 0\) (book Figure 3.1). The data would contain no untreated individuals with \(L = 1\) who could stand in for what would have happened to the treated with \(L = 1\) had they been untreated.
Unlike exchangeability, positivity can sometimes be checked empirically (Chapter 12). In Table 3.1 there are treated and untreated individuals in both levels of \(L\).
Technical Point 3.1: Positivity for standardization and IP weighting
The standardized mean \(\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l]\) is defined only if \(\operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{}\) is defined for every \(l\) with \(\Pr[L = l] \neq 0\), i.e., only under positivity; otherwise it is undefined.
For IP weighting, \(\operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[a \mid L]}\right]\mathclose{}\) is undefined without positivity (it involves \(0/0\)). The version \(\operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[A \mid L]}\right]\mathclose{}\) is always defined, because \(f[A \mid L]\) is never 0, but it no longer equals the counterfactual mean. Let \(Q(a) \stackrel{\text{def}}{=}\{l : \Pr[A = a \mid L = l] > 0\}\). Then
\[ \operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[A \mid L]}\right]\mathclose{} = \Pr[L \in Q(a)] \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L \in Q(a)\right]\mathclose{} \Pr[L = l \mid L \in Q(a)], \]
which under exchangeability equals \(\operatorname{E}\mathopen{}\left[Y^a \mid L \in Q(a)\right]\mathclose{} \Pr[L \in Q(a)]\). For binary \(A\) without positivity, \(Q(0) \neq Q(1)\), so the IP weighted contrast compares two different groups and has no causal interpretation even under exchangeability. Under positivity \(Q(0) = Q(1)\) and the contrast is the average causal effect if exchangeability holds.
Definition 4 (Consistency) The observed outcome of every treated individual equals her outcome had she received treatment, and the observed outcome of every untreated individual equals her outcome had she remained untreated:
\[Y = Y^A,\]
where \(Y^A\) is the counterfactual \(Y^a\) evaluated at the individual’s actual treatment \(A\).
The book takes the counterfactuals \(Y^a\) as primitives; the observed \(Y\) is a function of them. For binary \(A\):
\[ Y^A = A\,Y^{a=1} + (1 - A)\,Y^{a=0}, \]
which equals \(Y^{a=1}\) when \(A = 1\) (the second term vanishes) and \(Y^{a=0}\) when \(A = 0\) (the first term vanishes).
An individual’s causal effect \(Y^{a=1} - Y^{a=0}\) can be well-defined only if her \(Y^a\) is well-defined for both \(a = 1\) and \(a = 0\). In turn, the average causal effect \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\) in a given population can be well-defined only if every member of that population has a well-defined individual causal effect.
How do we know the counterfactuals are well-defined? A natural sufficient condition: when \(a\) is itself a well-defined intervention, \(Y^a\) is well-defined, as the outcome that would have been seen had that intervention been carried out.
Example 3 (Two heart transplant trials) Two ideal randomized trials of heart transplant (\(a = 1\)) versus medical therapy (\(a = 0\)) in the same population:
In both trials \(a = 1\) is well-defined by the protocol, so \(Y^{a=1}\) is well-defined. But the two counterfactuals, and hence the two causal effects, will likely differ.
Even Trial 1 did not specify the surgeon’s training. Because surgical experience affects post-transplant mortality, \(\Pr[Y^{a=1} = 1]\) depends on the (unknown) mix of surgeons. Had the protocol specified training, some other detail would still have been left unspecified or open to interpretation.
In experiments, the protocol makes the interventions explicit. In observational studies, investigators who talk about “the effect of heart transplant” without defining the intervention may mean different things, so \(Y^a\) is ill-defined until the experts agree on one intervention at a time.
An intervention is sufficiently well-defined when, for all practical purposes, no meaningful vagueness remains for \(Y^a\).
How do we know no meaningful vagueness remains? We don’t. It is a matter of agreement among experts, based on the substantive knowledge available at a particular time, and experts who agree today may judge differently once new knowledge arrives.
Fine Point 3.3: Possible worlds
Philosophers (Stalnaker 1968, Lewis 1973) defined \(Y^a\) as the value of \(Y\) in the closest possible world in which the individual received \(a\). Since the closest world to the actual world is itself, \(Y^a = Y\) when \(A = a\), so consistency always holds under this definition. When \(A \neq a\), the closest possible world is vague. Robins and Greenland (2000) argued that well-defined interventions should replace closest possible worlds, because in observational studies counterfactuals are vague to the degree that the hypothetical interventions are not made precise.
Biological states (blood pressure, LDL cholesterol, body weight) and social factors (socioeconomic status) can be changed only by intervening on their causes.
Definition 5 (Effective no-direct-effect (ENDE) intervention) One way to read the claim that blood pressure \(A\) affects this particular individual’s stroke outcome \(Y\) is as a belief that some intervention \(X\) exists (possibly not yet discovered) that can change this individual’s value of \(A\) to \(a\) and that affects her outcome \(Y\) only through \(A\), with no (direct) effect of its own. Such an \(X\) is an effective no-direct-effect intervention for \(a\), written ENDE(\(a\)). Its existence is logically equivalent to the statement that the counterfactuals \(Y^a\) and the joint counterfactuals \(Y^{x,a}\) are both well-defined and equal.
Because an average causal effect requires a well-defined individual effect for every member of the population, the average effect of a non-intervention \(A\), comparing \(a\) with \(a'\), is well-defined if every individual in the population has both an ENDE(\(a\)) and an ENDE(\(a'\)) intervention.
Considering small changes in \(A\) makes ENDE interventions more credible. For a change in blood pressure of \(\Delta\) mm Hg with \(\Delta\) near 0:
So \(Y^{\Delta}\) and \(\operatorname{E}\mathopen{}\left[Y^{\Delta} - Y^{\Delta = 0}\right]\mathclose{}\) are (essentially) well-defined for all individuals, where \(Y^{\Delta}\) is \(Y^a\) evaluated at \(a = A - \Delta\) and \(A\) is the individual’s blood pressure before the intervention begins.
Fine Point 3.4: An interventionist approach to causal inference
The book’s framework, sometimes called “interventionist”, rests on well-defined counterfactual outcomes, which are in turn defined through well-defined interventions. It supplies the formal language for precise discussion of causal inference about interventions and for interpreting the numbers that data analyses produce. Many published papers, however, attach causal interpretations to numerical quantities for questions that involve no recognizable intervention. The book therefore extends the interventionist approach to non-interventions by spelling out when quantitative causal inference about them is meaningful. The authors do not claim that this is the only valid philosophy of causality, but they know of no alternative framework that produces practically interpretable estimates of causal effects (Hernán and Robins 2020, Fine Point 3.4, p. 39).
The second component of consistency is the “\(=\)” in \(Y^a = Y\) for individuals with \(A = a\).
Example 4 (A well-defined intervention the data cannot link to) Investigators define \(a = 1\) as heart transplant with specified preoperative procedures, anesthesia, surgical technique, postoperative care, and immunosuppressive therapy. The data contain only an indicator \(B\) of whether a person had a transplant. For an individual with \(B = 1\), the well-defined \(Y^{a=1}\) need not equal the observed \(Y\).
The book uses a different letter (\(B\)) for the recorded variable precisely because the same letter is reserved for cases in which \(A = a\) corresponds to the well-defined intervention \(a\).
Fine Point 3.5: Attributable fraction
The excess fraction, one version of the attributable fraction, is the share of observed cases that would not have occurred had everyone been untreated: \((\Pr[Y = 1] - \Pr[Y^{a=0} = 1]) / \Pr[Y = 1]\).
Example 5 (The ambrosia dinner) All 20 individuals attended a dinner: 10 ate ambrosia (\(A = 1\)), 10 ate nectar (\(A = 0\)). The next day 7 of the 10 ambrosia eaters and 1 of the 10 nectar drinkers were sick. Assuming exchangeability,
The excess fraction is
\[ \frac{\Pr[Y = 1] - \Pr[Y^{a=0} = 1]}{\Pr[Y = 1]} = \frac{0.4 - 0.1}{0.4} = \frac{0.3}{0.4} = 0.75. \]
Of the 8 observed cases, only \(0.1 \times 20 = 2\) would have occurred had everyone received \(a = 0\), so 75% of the cases are attributable to ambrosia.
Assuming the three identifiability conditions amounts to viewing an observational analysis as an attempt to emulate a hypothetical randomized experiment.
Definition 6 (Target trial) The target trial (or target experiment) is the hypothetical randomized experiment that would quantify the causal effect of interest. For each causal effect, we may (i) specify the protocol of the target trial that we would like to, but cannot, conduct, and (ii) describe how the observational data would be used to emulate it.
If the emulation were successful, the observational study and the target trial (had it been conducted) would give the same results.
Remark 2 (Summary of the target trial for interventions and non-interventions). If \(A\) is an intervention, the target trial randomizes eligible individuals to values \(a\). If \(A\) is a non-intervention and we believe ENDE interventions \(X\) exist, the target trial randomizes eligible individuals to an \(X\) that sets \(A\) to \(a\), and the effect of \(A\) is equated with the effect of \(X\). Without that belief, \(Y^a\) and the effect of \(A\) are not well-defined.
Fine Point 3.6: When causal inference for non-interventions goes unquestioned
In a trial with assignment \(Z\) and received treatment \(A\), the intention-to-treat effect of \(Z\) is well-defined because \(Z\) is an intervention (Hernán and Robins 2020, Fine Point 3.6, p. 43). With perfect adherence and no direct effect of \(Z\), \(Z = A\) and \(Z\) is an ENDE intervention for \(A\). Without perfect adherence, \(Z\) is not an ENDE intervention for everyone, because it does not succeed in setting \(A\) to 1 or 0 for every individual (even if \(Z\) has no direct effect on \(Y\)). The effect of \(A\) in the whole population is then well-defined only if some ENDE intervention \(X\) on \(A\) exists (e.g., adding aspirin to food without the participants’ knowledge). Most experts find this so self-evident that “the effect of received treatment” is discussed without mentioning \(X\).
Whether or not such an \(X\) exists, \(Z\) is effective for one subset: individuals who would have \(A = 1\) under \(Z = 1\) and \(A = 0\) under \(Z = 0\). For them, \(Z\) is an ENDE intervention, so the effect of \(A\) in that subset is well-defined and equals the effect of \(Z\). Chapter 16 calls this subset the compliers; its members cannot be identified individually, but the effect of \(A\) among them can be identified under additional assumptions.
If \(Z\) has direct effects on \(Y\) not through \(A\), the effect of \(Z\) differs from that of \(X\) even under full adherence.
Some authors treat “the causal effect of \(A\) on \(Y\)” as well-defined even where many experts would not grant that any ENDE intervention \(X\) on \(A\) exists (Pearl 2009; Schwartz et al. 2016; Glymour and Spiegelman 2016) (Hernán and Robins 2020, 42).
Chapter 16 considers alternative identifying conditions for emulating a target trial.
Example 6 (Body mass index and death) Comparing the risk of death between people with body mass index (BMI) 25 versus 30 at the start of follow-up implies a target trial of an instantaneous, possibly very large, weight change, for which no ENDE intervention is known. A modified analysis could emulate a more reasonable target trial, e.g., one assigning some individuals to a 5% reduction in BMI every year, starting at age 40, for as long as their BMI stays over 25 (a sustained strategy of the kind studied in Part III).
Technical Point 3.2: Recursive substitution
For chronologically ordered variables \(L, A, M, Y\) with well-defined interventions on \(L, A, M\), the one-step-ahead counterfactuals are \(L, A^l, M^{l,a}, Y^{l,a,m}\). All other factuals and well-defined counterfactuals are functions of them via recursive substitution, e.g., \(A = A^L\), \(M^a = M^{L,a}\), \(Y^a = Y^{L, a, M^a}\), and \(Y = Y^{L, A, M}\). Applied to the two transplant trials, recursive substitution shows why Trial 2’s outcome distribution is harder to transport: it changes whenever the distribution of the natural values \(A_1^{a_0=1}, A_2^{a_0=1}, A_3^{a_0=1}\) differs between populations. The argument applies equally to observational and randomized studies.
The identifiability conditions under which an observational study can be analyzed like a conditionally randomized experiment:
Specifying the target trial makes the causal question, and what the data must supply to emulate it, explicit. For non-intervention variables, the target trial randomizes an ENDE intervention \(X\).