Suppose an investigator observes that when one pedestrian looks up at the sky, a second pedestrian tends to look up too. She also notices that pedestrians look up when they hear a thunderous noise overhead, so she cannot tell whether the second pedestrian looked up because of the first one or because of the noise: the effect of one person’s looking up is confounded by the noise. In a randomized experiment treatment is assigned by a coin flip, but in an observational study treatment may be determined by factors that also affect the outcome. The effects of those factors then become entangled with the effect of treatment. This chapter defines confounding structurally, links it to exchangeability, contrasts the structural definition with the traditional definition of a confounder, introduces single-world intervention graphs (SWIGs), and reviews methods to adjust for confounding.
Confounding is the bias due to common causes of treatment and outcome.
In Figure 7.1 there are two sources of association between \(A\) and \(Y\):
The second path is a backdoor path.
Definition 1 (Backdoor Path) In a causal DAG, a backdoor path is a noncausal path between treatment and outcome that remains even if all arrows pointing from treatment to other variables (the descendants of treatment) are removed. That is, the path has an arrow pointing into treatment (Hernán and Robins 2020, 91).
If \(L\) did not exist, all the association between \(A\) and \(Y\) would be causal, and the associational risk ratio \(\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]\) would equal the causal risk ratio \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\). The common cause \(L\) adds a second source of association, which is confounding for the effect of \(A\) on \(Y\).
When marginal exchangeability \(Y^a \perp\!\!\!\perp A\) holds, as in a marginally randomized experiment, the average causal effect is identified without adjustment:
\[\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\]
When only conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) holds, as in a conditionally randomized experiment, the average causal effect is identified by adjusting for \(L\) via standardization or IP weighting, and the conditional effects \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L = l\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid L = l\right]\mathclose{}\) are identified by stratification.
Theorem 1 (Standardization Under Conditional Exchangeability) Under conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), positivity, and consistency,
\[\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \Pr[L = l]\]
Proof. \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l, A = a\right]\mathclose{} \Pr[L = l] && \text{(conditional exchangeability; positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \Pr[L = l] && \text{(consistency)} \end{align}\]
Positivity, \(\Pr[A = a \mid L = l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\), ensures that the conditional mean in the second line is defined.
This hypothetical example is not from the book. Twenty individuals receive a drug \(A\); \(L = 1\) indicates a high-risk group; \(Y = 1\) indicates death.
| \(L\) | \(A\) | \(n\) | Deaths | \(\Pr[Y = 1 \mid A, L]\) |
|---|---|---|---|---|
| 1 | 1 | 8 | 4 | 0.50 |
| 1 | 0 | 2 | 1 | 0.50 |
| 0 | 1 | 2 | 0 | 0.00 |
| 0 | 0 | 8 | 0 | 0.00 |
Crude risks: \(\Pr[Y = 1 \mid A = 1] = (4 + 0) / (8 + 2) = 0.40\) and \(\Pr[Y = 1 \mid A = 0] = (1 + 0) / (2 + 8) = 0.10\).
With \(\Pr[L = 1] = 10/20 = 0.50\), standardization (Theorem 1) gives
\[\begin{align} \Pr[Y^{a=1} = 1] &= 0.50 \times 0.50 + 0.00 \times 0.50 = 0.25 \\ \Pr[Y^{a=0} = 1] &= 0.50 \times 0.50 + 0.00 \times 0.50 = 0.25 \end{align}\]
so the standardized risk difference is \(0\), while the crude risk difference is \(0.40 - 0.10 = 0.30\).
If we know the true causal DAG, how can we tell whether some set \(L\) yields conditional exchangeability? The book gives two approaches: the backdoor criterion (Pearl 1995) and the transformation of the DAG into a SWIG (Section 7.5).
Definition 2 (Backdoor Criterion) A set of covariates \(L\) satisfies the backdoor criterion if
Under faithfulness (and the causal model discussed in Technical Point 7.1), conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) holds if and only if \(L\) satisfies the backdoor criterion. Checking every subset of measured non-descendants of \(A\) therefore tells us whether any of them achieves conditional exchangeability.
Technical Point 7.1: Does Conditional Exchangeability Imply the Backdoor Criterion?
That \(L\) satisfies the backdoor criterion always implies conditional exchangeability given \(L\), even without faithfulness. The converse (given faithfulness) holds under an FFRCISTG model (Technical Point 6.3). Under an NPSEM-IE model, conditional exchangeability can hold without the backdoor criterion, for example in the DAG with nodes \(A, L, Y\) and arrows \(A \rightarrow L\), \(A \rightarrow Y\).
The difference arises because the NPSEM-IE assumes cross-world independencies between counterfactuals, which no randomized experiment could ever verify; for that reason Robins did not assume them in the FFRCISTG model. The book assumes an FFRCISTG model and faithfulness unless stated otherwise (Hernán and Robins 2020, 94).
Definition 3 (Sufficient Set for Confounding Adjustment) A set \(L\) of measured non-descendants of \(A\) is a sufficient set for confounding adjustment when conditioning on \(L\) blocks all backdoor paths, that is, when the treated and the untreated are exchangeable within levels of \(L\).
Applying the backdoor criterion to Figures 7.1-7.3:
In all three there is confounding, but no unmeasured confounding given \(L\).
In Figure 7.4 \(A\) and \(Y\) share no common cause, so there is no confounding: the backdoor path \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\) is blocked by the collider \(L\), and \(\Pr[Y^a = 1] = \Pr[Y = 1 \mid A = a]\).
Example 1 (Physical Activity and Cervical Cancer) Let \(A\) be physical activity, \(Y\) cervical cancer, \(U_1\) a pre-cancer lesion, \(L\) a Pap smear (a diagnostic test for pre-cancer), and \(U_2\) a health-conscious personality that leads to more physical activity and more doctor visits. Under Figure 7.4 the effect of \(A\) on \(Y\) is unconfounded, and there is no need to adjust for \(L\) to compute \(\Pr[Y^{a=1} = 1]\) or \(\Pr[Y^{a=0} = 1]\) (Hernán and Robins 2020, 95).
Adjusting for \(L\) in Figure 7.4 would create bias, because conditioning on the collider \(L\) opens the path \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\). Here unconditional exchangeability holds but conditional exchangeability given \(L\) does not: the average causal effect is identified, but the effects within levels of \(L\) generally are not. The book calls the resulting bias selection bias (Chapter 8); it is also known as M-bias.
Figure 7.5 adds the arrow \(L \rightarrow A\), which creates the open backdoor path \(A \leftarrow L \leftarrow U_1 \rightarrow Y\): there is confounding.
There is neither unconditional exchangeability nor conditional exchangeability given \(L\). The definition of collider is path-specific: \(L\) is a collider on the second path but not on the first.
A solution is to measure either a variable \(L_1\) between \(U_1\) and \(A\) or \(Y\), which yields conditional exchangeability given \(L_1\), or a variable \(L_2\) between \(U_2\) and \(A\) or \(L\), which yields conditional exchangeability given \(\{L_2, L\}\). Figure 7.6 adds \(L_1\) between \(U_1\) and \(Y\) and \(L_2\) between \(U_2\) and \(A\).
Fine Point 7.1: The Strength and Direction of Confounding Bias
Suppose a study of heart transplant \(A\) on death \(Y\) finds a risk ratio of 0.6, and a critic suspects confounding by smoking \(L\). Smokers are less likely to receive a transplant and more likely to die. Because the transplant group has fewer smokers, it would have lower mortality even under the null, so adjusting for smoking moves the estimate upward (the book’s illustration: from 0.6 to 0.7). Failing to adjust exaggerates the apparent benefit.
Signed causal diagrams. With dichotomous \(L\), \(A\), \(Y\) in Figure 7.1, put a \(+\) on the arrow \(L \rightarrow A\) if \(L\) has a positive average causal effect on \(A\), otherwise a \(-\); likewise for \(L \rightarrow Y\).
In the smoking example, \(L \rightarrow A\) is \(-\) and \(L \rightarrow Y\) is \(+\), so the confounding is negative, consistent with the unadjusted 0.6 lying below the adjusted 0.7. This simple rule may fail in more complex diagrams or with non-dichotomous variables (VanderWeele, Hernán, and Robins 2008).
Magnitude. A large confounding bias requires a strong confounder-treatment association and a strong confounder-outcome association (conditional on treatment); for discrete confounders it also depends on the confounder’s prevalence. For unknown confounders, sensitivity analyses can organize educated guesses about the size of the bias (Hernán and Robins 2020, 96).
Suppose data on \(L\), \(A\), and \(Y\) suffice to identify the causal effect, as in Figures 7.1-7.4.
Definition 4 (Confounder (Given Data on L, A, Y)) \(L\) is a confounder if data on \(A\) and \(Y\) alone do not suffice for identification, that is, if there is conditional exchangeability given \(L\) but not unconditional exchangeability (structural confounding). \(L\) is a non-confounder if data on \(A\) and \(Y\) alone suffice, that is, if there is unconditional exchangeability (Hernán and Robins 2020, 96).
The causal diagrams of this section show two structural sources of lack of exchangeability through open backdoor paths:
An alternative definition calls confounding any “bias due to an open backdoor path between \(A\) and \(Y\)”, equivalently “any systematic bias that would be eliminated by randomized assignment of \(A\)”. It differs only in labeling the bias from conditioning on \(L\) in Figure 7.4 as confounding.
Fine Point 7.2: Identification of Conditional and Unconditional Effects
Which effects can be identified depends on which variables are measured. In Figure 7.6:
The structural approach needs prior knowledge of the causal DAG, including all shared causes (measured or not) of \(A\) and \(Y\); the backdoor criterion then says what to adjust for. The traditional approach instead labels as confounders the variables meeting mostly associational conditions, mandates adjusting for them, and declares confounding when adjusted and unadjusted estimates differ.
Definition 5 (Traditional Definition of Confounder) A variable is a confounder under the traditional approach if it
Replacing condition 2 by the structural condition “it is a cause of the outcome” fixes Figure 7.4, but then \(L\) in Figure 7.2 would no longer count as a confounder, although it must be adjusted for (Technical Point 7.2).
The structural approach first identifies the sources of confounding (the common causes of treatment and outcome) and then a sufficient adjustment set. Whether a variable belongs to a sufficient set depends on the other variables in it. In Figures 7.2 and 7.3, \(L\) is needed only because \(U\) is unmeasured; given \(U\), \(L\) would not be a confounder. Given a causal DAG, confounding is an absolute concept, whereas confounder is a relative one (Hernán and Robins 2020, 100).
The structural approach has two advantages:
Fine Point 7.3: Surrogate Confounders
In Figure 7.8 the unmeasured \(U\) (e.g., socioeconomic status) confounds the effect of physical activity \(A\) on cardiovascular disease \(Y\), and the measured \(L\) (e.g., income) is a proxy for \(U\). \(L\) is not on a backdoor path, but adjusting for it may remove some of the confounding by \(U\); if \(L\) were perfectly correlated with \(U\), conditioning on \(L\) would be the same as conditioning on \(U\). If \(L\) is a binary, nondifferentially misclassified version of \(U\), conditioning on \(L\) partially blocks \(A \leftarrow U \rightarrow Y\) under some weak conditions (Greenland 1980; Ogburn and VanderWeele 2012). So one typically prefers to adjust for \(L\).
Variables that can reduce, but never completely eliminate, confounding bias without lying on a backdoor path are surrogate confounders (Hernán and Robins 2020, 100). One strategy is to measure and adjust for as many surrogate confounders as possible (see Chapter 18).
Technical Point 7.2: Fixing the Traditional Definition of Confounder
The traditional definition relies on two incorrect statistical criteria (conditions 1 and 2) and one incorrect causal criterion (condition 3). To fix it:
If both conditions hold, \(U\) is a non-confounder given data on \(L\). These conditions were proposed by Robins (1997, Theorem 4.3) and discussed by Greenland, Pearl, and Robins (1999), who used them to show there is no confounding in Figure 7.4 (Hernán and Robins 2020, 101).
The equivalence between exchangeability and the backdoor criterion seems “rather magical” because counterfactuals do not appear on causal diagrams. Single-world intervention graphs (SWIGs) put the counterfactual variables on the graph, so that exchangeability can be read off directly by d-separation.
Definition 6 (Single-World Intervention Graph (SWIG)) A SWIG depicts the variables and causal relations that would be observed in a hypothetical world in which all individuals received treatment level \(a\), a counterfactual world created by a single intervention. It is obtained from a causal DAG as follows:
Example 2 (SWIGs for Figures 7.2 and 7.4)
On the SWIG, \(Y^a\) is d-separated from \(A\) given \(L\) if and only if \(L\) is a non-descendant of \(A\) that blocks all backdoor paths from \(A\) to \(Y\). This is the simple SWIG-based argument for the equivalence stated in Section 7.2.
Without randomization, causal inference relies on the uncheckable assumption that the measured \(L\) is a sufficient set for confounding adjustment. Under that assumption, methods that adjust for \(L\) fall into two categories:
Fine Point 7.4: Confounders Cannot Be Descendants of Treatment, but Can Be in the Future of Treatment
In Figure 7.11, \(L\) is a descendant of \(A\) that blocks all backdoor paths, and the effect of \(A\) on \(Y\) is entirely through \(L\). Conditioning on \(L\) opens no collider path, but it blocks the causal pathway, so adjusting for \(L\) does not remove bias. Hence conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) must fail. On the SWIG (Figure 7.12) \(L\) is replaced by the counterfactual \(L^a\), and we can read off \(Y^a \perp\!\!\!\perp A \mid L^a\) but not \(Y^a \perp\!\!\!\perp A \mid L\), since \(L\) is not on the graph. (Under an FFRCISTG model, an independence that cannot be read off the SWIG cannot be assumed to hold.)
The problem is that \(L\) is a descendant of \(A\), not that \(L\) occurs after \(A\). Without the arrow \(A \rightarrow L\), \(L\) would be a non-descendant that blocks all backdoor paths, and adjusting for it would remove all bias even if \(L\) were measured after \(A\). What matters is the topology of the causal diagram, not the time order of the nodes (Hernán and Robins 2020, 103).
Some methods handle confounding without conditional exchangeability:
They replace conditional exchangeability with other assumptions that are just as unverifiable, so the choice of method depends on which unverifiable assumptions are more plausible in a given setting.
Conditional exchangeability may be unrealistic, but expert knowledge about the causal structure helps get close to it:
A critic who says only “your observational study may be confounded” makes a logical, not a scientific, statement: it is true of every observational study. A scientific criticism names a source, such as “confounding due to cigarette smoking, a common cause through which a backdoor path may remain open”. That gives a testable challenge: adjust for smoking or, if smoking was not measured, conduct a sensitivity analysis (Hernán and Robins 2020, 104–5).
Technical Point 7.3: Difference-in-Differences and Negative Outcome Controls
Suppose unmeasured \(U\) (e.g., history of heart disease) confounds the effect of aspirin \(A\) on blood pressure \(Y\), and we also measured the outcome just before treatment, \(C\), a negative outcome control (Figure 7.13). \(A\) cannot cause \(C\), so \(\operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\) measures additive confounding for the effect of \(A\) on \(C\).
Under additive equi-confounding, \(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{} = \operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\), the effect in the treated is identified:
\[\begin{align} \operatorname{E}\mathopen{}\left[Y^1 - Y^0 \mid A = 1\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^1 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} && \text{(linearity)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(add and subtract } \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(consistency)} \\ &= \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\right)\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(equi-confounding)} \end{align}\]
The association of \(A\) with \(Y\) (effect plus confounding) minus the confounding measured through \(C\). The arrow \(C \rightarrow Y\) is not necessary for \(C\) to be a negative outcome control. This is difference-in-differences (Card 1990; Meyer 1995; Angrist and Krueger 1999). It is a somewhat restrictive use of negative outcome controls: it needs \(Y\) and \(C\) on the same scale and additive equi-confounding; Sofer et al. (2016) describe more general methods.
With both a negative outcome control \(C\) and a negative treatment control \(Z\), the effect can be identified nonparametrically despite unmeasured \(U\) under additional assumptions; for discrete \(U\), \(C\), \(Z\) with \(C\) and \(Z\) having at least as many levels as \(U\), identification holds quite generally (Miao et al. 2018). This is proximal causal inference (Cui et al. 2024); Figure 7.15 is an example (Hernán and Robins 2020, 106).
Technical Point 7.4: The Front Door Criterion
In Figure 7.14, unmeasured \(U\) confounds \(A\) and \(Y\), and \(M\) fully mediates the effect of \(A\) on \(Y\) and shares no unmeasured causes with \(A\) or \(Y\). Standardization and IP weighting cannot be used because \(U\) is unavailable, but Pearl (1995) showed that \(\Pr[Y^a = 1]\) is identified by the front door formula
\[\Pr[Y^a = 1] = \sum_m \Pr[M = m \mid A = a] \sum_{a'} \Pr[Y = 1 \mid M = m, A = a'] \Pr[A = a']\]
(Pearl’s “backdoor formula” is what the book calls standardization or the point-treatment g-formula.)
Proof.
This proof requires well-defined counterfactuals \(Y^m\); Technical Points 21.11 and 21.12 give proofs without that condition (Hernán and Robins 2020, 107).