Chapter 7: Confounding

Published

Last modified: 2026-10-09 13:46:40 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Suppose an investigator observes that when one pedestrian looks up at the sky, a second pedestrian tends to look up too. She also notices that pedestrians look up when they hear a thunderous noise overhead, so she cannot tell whether the second pedestrian looked up because of the first one or because of the noise: the effect of one person’s looking up is confounded by the noise. In a randomized experiment treatment is assigned by a coin flip, but in an observational study treatment may be determined by factors that also affect the outcome. The effects of those factors then become entangled with the effect of treatment. This chapter defines confounding structurally, links it to exchangeability, contrasts the structural definition with the traditional definition of a confounder, introduces single-world intervention graphs (SWIGs), and reviews methods to adjust for confounding.

This chapter is based on Hernán and Robins (2020, chap. 7, pp. 91-107).

The book describes confounding as “just a form of lack of exchangeability between the treated and the untreated” (Hernán and Robins 2020, 91). In the presence of confounding, “association is not causation” holds no matter how large the study population is.

1 7.1 The Structure of Confounding (pp. 91-93)


Confounding is the bias due to common causes of treatment and outcome.

Figure 7.1 (same as Figure 6.1). L is a common cause of treatment A and outcome Y.

Figure 7.1 (same as Figure 6.1). L is a common cause of treatment A and outcome Y.

In Figure 7.1 there are two sources of association between \(A\) and \(Y\):

  • the causal path \(A \rightarrow Y\);
  • the path \(A \leftarrow L \rightarrow Y\) through the common cause \(L\).

The second path is a backdoor path.

Definition 1 (Backdoor Path) In a causal DAG, a backdoor path is a noncausal path between treatment and outcome that remains even if all arrows pointing from treatment to other variables (the descendants of treatment) are removed. That is, the path has an arrow pointing into treatment (Hernán and Robins 2020, 91).

If \(L\) did not exist, all the association between \(A\) and \(Y\) would be causal, and the associational risk ratio \(\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]\) would equal the causal risk ratio \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\). The common cause \(L\) adds a second source of association, which is confounding for the effect of \(A\) on \(Y\).

1.1 Examples of Confounding


Left: Figure 7.2, in which L causes A, and L and Y share the unmeasured cause U. Right: Figure 7.3, in which L causes Y, and L and A share the unmeasured cause U. Unmeasured nodes are gray.

Left: Figure 7.2, in which L causes A, and L and Y share the unmeasured cause U. Right: Figure 7.3, in which L causes Y, and L and A share the unmeasured cause U. Unmeasured nodes are gray.

  • Occupational factors (Figure 7.1): physical fitness \(L\) causes both working as a firefighter \(A\) and lower mortality \(Y\) (healthy worker bias).
  • Clinical decisions (Figure 7.1 or 7.2): heart disease \(L\) is an indication for aspirin \(A\) and a risk factor for stroke \(Y\), either directly or through unmeasured atherosclerosis \(U\) (confounding by indication or channeling).
  • Lifestyle (Figure 7.3): personality and social factors \(U\) lead to both lack of exercise \(A\) and smoking \(L\), which affects death \(Y\).
  • Genetic factors (Figure 7.3): a DNA sequence \(L\) that affects trait \(Y\) is more frequent among carriers of sequence \(A\) (linkage disequilibrium or population stratification).
  • Social factors (Figure 7.1): disability at age 55 \(L\) affects income at age 65 \(A\) and disability at age 75 \(Y\).
  • Environmental exposures (Figure 7.3): weather \(U\) affects levels of particulate matter \(A\) and of other pollutants \(L\) that cause coronary heart disease \(Y\).

Terminology in the examples:

  • “Channeling” is often reserved for bias created by patient-specific risk factors \(L\) that encourage doctors to prescribe a particular drug \(A\) within a class of drugs.
  • When subclinical disease \(U\) causes both lack of exercise \(A\) and clinical disease \(Y\), and \(L\) is unknown, this form of confounding is often called reverse causation.
  • “Population stratification” is often reserved for bias from studying a mixture of individuals of different ethnic groups, so \(U\) can stand for ethnicity.
  • Some authors replace an unmeasured common cause \(U\) and its two arrows with a bidirected edge between the measured variables that \(U\) causes.

A historical note (Hernán and Robins 2020, 92): early statistical descriptions of confounding were given by Yule (1903) for discrete variables and by Pearson et al. (1899) for continuous variables, who called these associations “fictitious”, “illusory”, “apparent”, or “spurious”. But associations due to common causes are real associations; they simply cannot be interpreted as treatment effects.

Common structure: in every example the bias is due to a cause (\(L\) or \(U\)) shared by treatment and outcome, which opens a backdoor path between \(A\) and \(Y\). The book reserves the term confounding for this structure and uses other names for biases with other structures. Throughout the chapter it assumes positivity and consistency, perfectly measured nodes, no selection nodes, and no random variability (Chapters 8, 9, and 10 relax these assumptions).

2 7.2 Confounding and Exchangeability (pp. 93-95)


When marginal exchangeability \(Y^a \perp\!\!\!\perp A\) holds, as in a marginally randomized experiment, the average causal effect is identified without adjustment:

\[\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\]

When only conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) holds, as in a conditionally randomized experiment, the average causal effect is identified by adjusting for \(L\) via standardization or IP weighting, and the conditional effects \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L = l\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid L = l\right]\mathclose{}\) are identified by stratification.

Theorem 1 (Standardization Under Conditional Exchangeability) Under conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), positivity, and consistency,

\[\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \Pr[L = l]\]

Proof. \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l, A = a\right]\mathclose{} \Pr[L = l] && \text{(conditional exchangeability; positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \Pr[L = l] && \text{(consistency)} \end{align}\]

Positivity, \(\Pr[A = a \mid L = l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\), ensures that the conditional mean in the second line is defined.

2.1 Supplement: A Numerical Example


This hypothetical example is not from the book. Twenty individuals receive a drug \(A\); \(L = 1\) indicates a high-risk group; \(Y = 1\) indicates death.

\(L\) \(A\) \(n\) Deaths \(\Pr[Y = 1 \mid A, L]\)
1 1 8 4 0.50
1 0 2 1 0.50
0 1 2 0 0.00
0 0 8 0 0.00

Crude risks: \(\Pr[Y = 1 \mid A = 1] = (4 + 0) / (8 + 2) = 0.40\) and \(\Pr[Y = 1 \mid A = 0] = (1 + 0) / (2 + 8) = 0.10\).

With \(\Pr[L = 1] = 10/20 = 0.50\), standardization (Theorem 1) gives

\[\begin{align} \Pr[Y^{a=1} = 1] &= 0.50 \times 0.50 + 0.00 \times 0.50 = 0.25 \\ \Pr[Y^{a=0} = 1] &= 0.50 \times 0.50 + 0.00 \times 0.50 = 0.25 \end{align}\]

so the standardized risk difference is \(0\), while the crude risk difference is \(0.40 - 0.10 = 0.30\).

The high-risk group (\(L = 1\)) is treated more often (\(8/10\) versus \(2/10\)) and dies more often (\(0.50\) versus \(0.00\)). Within each level of \(L\) the treated and untreated have the same risk, so the whole crude risk difference of \(0.30\) is confounding, assuming \(L\) is the only common cause of \(A\) and \(Y\).

2.2 The Backdoor Criterion


If we know the true causal DAG, how can we tell whether some set \(L\) yields conditional exchangeability? The book gives two approaches: the backdoor criterion (Pearl 1995) and the transformation of the DAG into a SWIG (Section 7.5).

Definition 2 (Backdoor Criterion) A set of covariates \(L\) satisfies the backdoor criterion if

  1. all backdoor paths between \(A\) and \(Y\) are blocked by conditioning on \(L\), and
  2. \(L\) contains no variables that are descendants of treatment \(A\).

Under faithfulness (and the causal model discussed in Technical Point 7.1), conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) holds if and only if \(L\) satisfies the backdoor criterion. Checking every subset of measured non-descendants of \(A\) therefore tells us whether any of them achieves conditional exchangeability.

NoteTechnical Point 7.1: Does Conditional Exchangeability Imply the Backdoor Criterion?

That \(L\) satisfies the backdoor criterion always implies conditional exchangeability given \(L\), even without faithfulness. The converse (given faithfulness) holds under an FFRCISTG model (Technical Point 6.3). Under an NPSEM-IE model, conditional exchangeability can hold without the backdoor criterion, for example in the DAG with nodes \(A, L, Y\) and arrows \(A \rightarrow L\), \(A \rightarrow Y\).

The difference arises because the NPSEM-IE assumes cross-world independencies between counterfactuals, which no randomized experiment could ever verify; for that reason Robins did not assume them in the FFRCISTG model. The book assumes an FFRCISTG model and faithfulness unless stated otherwise (Hernán and Robins 2020, 94).

2.3 Two Settings in Which the Backdoor Criterion Holds


  1. No common causes of treatment and outcome (e.g., Figure 6.2): there are no backdoor paths, the empty set satisfies the criterion, and there is no confounding. This is a marginally randomized experiment; marginal exchangeability is equivalent to no common causes of treatment and outcome.
  2. Common causes, but a set \(L\) of measured non-descendants of \(A\) blocks all backdoor paths (e.g., Figure 7.1): there is confounding, but no unmeasured confounding. This is a conditionally randomized experiment, with an arrow \(L \rightarrow A\) by design.

Definition 3 (Sufficient Set for Confounding Adjustment) A set \(L\) of measured non-descendants of \(A\) is a sufficient set for confounding adjustment when conditioning on \(L\) blocks all backdoor paths, that is, when the treated and the untreated are exchangeable within levels of \(L\).

The heart transplant study (Chapter 2) is an example of the second setting. The treated had a higher frequency of severe heart disease \(L\), a common cause of \(A\) and \(Y\), so had they remained untreated their risk of death would have been higher than that of the untreated. They are not marginally exchangeable, but they are conditionally exchangeable given \(L\). The second setting is also what investigators hope for in observational studies that measure many variables \(L\).

Magnitude and direction: the backdoor criterion says nothing about the size or direction of confounding. Some unblocked backdoor paths may be weak, and strong backdoor paths may induce biases in opposite directions that partly cancel. Unmeasured confounding is not “all or nothing”, so investigators should consider its expected direction and magnitude (Fine Point 7.1). The book refers to Greenland and Robins (1986, 2009) for a detailed discussion of the relations between confounding and exchangeability.

3 7.3 Confounding and the Backdoor Criterion (pp. 95-98)


Applying the backdoor criterion to Figures 7.1-7.3:

  • Figure 7.1: the backdoor path \(A \leftarrow L \rightarrow Y\) is blocked by conditioning on \(L\).
  • Figure 7.2: the backdoor path \(A \leftarrow L \leftarrow U \rightarrow Y\) could be blocked by \(U\), which is unmeasured, but it is also blocked by \(L\).
  • Figure 7.3: the backdoor path \(A \leftarrow U \rightarrow L \rightarrow Y\) is also blocked by \(L\).

In all three there is confounding, but no unmeasured confounding given \(L\).

3.1 M-bias: Figure 7.4


Left: Figure 7.4, in which L is a collider on the path A <- U2 -> L <- U1 -> Y. Right: Figure 7.5, which adds the arrow L -> A.

Left: Figure 7.4, in which L is a collider on the path A <- U2 -> L <- U1 -> Y. Right: Figure 7.5, which adds the arrow L -> A.

In Figure 7.4 \(A\) and \(Y\) share no common cause, so there is no confounding: the backdoor path \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\) is blocked by the collider \(L\), and \(\Pr[Y^a = 1] = \Pr[Y = 1 \mid A = a]\).

Example 1 (Physical Activity and Cervical Cancer) Let \(A\) be physical activity, \(Y\) cervical cancer, \(U_1\) a pre-cancer lesion, \(L\) a Pap smear (a diagnostic test for pre-cancer), and \(U_2\) a health-conscious personality that leads to more physical activity and more doctor visits. Under Figure 7.4 the effect of \(A\) on \(Y\) is unconfounded, and there is no need to adjust for \(L\) to compute \(\Pr[Y^{a=1} = 1]\) or \(\Pr[Y^{a=0} = 1]\) (Hernán and Robins 2020, 95).

Adjusting for \(L\) in Figure 7.4 would create bias, because conditioning on the collider \(L\) opens the path \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\). Here unconditional exchangeability holds but conditional exchangeability given \(L\) does not: the average causal effect is identified, but the effects within levels of \(L\) generally are not. The book calls the resulting bias selection bias (Chapter 8); it is also known as M-bias.

Why “M-bias”: the structure was described by Greenland, Pearl, and Robins (1999) and named M-bias by Greenland (2003), because \(U_2\), \(L\), \(U_1\) resemble an M lying on its side (Hernán and Robins 2020, 97). That unconditional effects can be identified while conditional effects are not was shown non-graphically by Greenland and Robins (1986).

If \(U_1\) caused \(U_2\), \(U_2\) caused \(U_1\), or an unmeasured \(U_3\) caused both, \(A\) and \(Y\) would have a common cause, and there would be neither unconditional nor conditional exchangeability given \(L\).

3.2 Intractable Bias: Figure 7.5


Figure 7.5 adds the arrow \(L \rightarrow A\), which creates the open backdoor path \(A \leftarrow L \leftarrow U_1 \rightarrow Y\): there is confounding.

  • Conditioning on \(L\) blocks \(A \leftarrow L \leftarrow U_1 \rightarrow Y\),
  • but opens \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\), on which \(L\) is a collider.

There is neither unconditional exchangeability nor conditional exchangeability given \(L\). The definition of collider is path-specific: \(L\) is a collider on the second path but not on the first.

A solution is to measure either a variable \(L_1\) between \(U_1\) and \(A\) or \(Y\), which yields conditional exchangeability given \(L_1\), or a variable \(L_2\) between \(U_2\) and \(A\) or \(L\), which yields conditional exchangeability given \(\{L_2, L\}\). Figure 7.6 adds \(L_1\) between \(U_1\) and \(Y\) and \(L_2\) between \(U_2\) and \(A\).

NoteFine Point 7.1: The Strength and Direction of Confounding Bias

Suppose a study of heart transplant \(A\) on death \(Y\) finds a risk ratio of 0.6, and a critic suspects confounding by smoking \(L\). Smokers are less likely to receive a transplant and more likely to die. Because the transplant group has fewer smokers, it would have lower mortality even under the null, so adjusting for smoking moves the estimate upward (the book’s illustration: from 0.6 to 0.7). Failing to adjust exaggerates the apparent benefit.

Signed causal diagrams. With dichotomous \(L\), \(A\), \(Y\) in Figure 7.1, put a \(+\) on the arrow \(L \rightarrow A\) if \(L\) has a positive average causal effect on \(A\), otherwise a \(-\); likewise for \(L \rightarrow Y\).

  • Both signs equal: positive confounding; the unadjusted estimate is biased upward.
  • Signs differ: negative confounding; the unadjusted estimate is biased downward.

In the smoking example, \(L \rightarrow A\) is \(-\) and \(L \rightarrow Y\) is \(+\), so the confounding is negative, consistent with the unadjusted 0.6 lying below the adjusted 0.7. This simple rule may fail in more complex diagrams or with non-dichotomous variables (VanderWeele, Hernán, and Robins 2008).

Magnitude. A large confounding bias requires a strong confounder-treatment association and a strong confounder-outcome association (conditional on treatment); for discrete confounders it also depends on the confounder’s prevalence. For unknown confounders, sensitivity analyses can organize educated guesses about the size of the bias (Hernán and Robins 2020, 96).

3.3 Confounders and Non-confounders


Suppose data on \(L\), \(A\), and \(Y\) suffice to identify the causal effect, as in Figures 7.1-7.4.

Definition 4 (Confounder (Given Data on L, A, Y)) \(L\) is a confounder if data on \(A\) and \(Y\) alone do not suffice for identification, that is, if there is conditional exchangeability given \(L\) but not unconditional exchangeability (structural confounding). \(L\) is a non-confounder if data on \(A\) and \(Y\) alone suffice, that is, if there is unconditional exchangeability (Hernán and Robins 2020, 96).

  • In Figures 7.1-7.3, \(L\) is a confounder: \(\Pr[Y^a = 1]\) is identified by the standardized risk \(\sum_l \Pr[Y = 1 \mid A = a, L = l] \Pr[L = l]\). In Figures 7.2 and 7.3, \(L\) is not a common cause of \(A\) and \(Y\), yet it is a confounder because it is needed to block the backdoor path through \(U\).
  • In Figure 7.4, \(L\) is a non-confounder, and \(\Pr[Y^a = 1] = \Pr[Y = 1 \mid A = a]\). Standardizing by \(L\) would be biased.

An informal definition for Figures 7.1 to 7.4 is “a confounder is any variable that can be used to adjust for confounding.” This definition is not circular, because confounding was defined first, just as “a musician is a person who plays music” is not circular once music has been defined (Hernán and Robins 2020, 96).

3.4 Two Definitions of Confounding


The causal diagrams of this section show two structural sources of lack of exchangeability through open backdoor paths:

  • common causes of treatment and outcome, which the book calls confounding;
  • conditioning on a common effect, which the book calls selection bias.

An alternative definition calls confounding any “bias due to an open backdoor path between \(A\) and \(Y\)”, equivalently “any systematic bias that would be eliminated by randomized assignment of \(A\)”. It differs only in labeling the bias from conditioning on \(L\) in Figure 7.4 as confounding.

Why randomization would eliminate the Figure 7.4 bias: random assignment of \(A\) rules out an unmeasured common cause \(U_2\) of \(A\) and \(L\), so conditioning on \(L\) would no longer open a backdoor path.

A difference between the two definitions: under the structural definition, whether confounding exists is a fact about the population, independent of the analysis. Under the “eliminated by randomization” definition it depends on the analysis: in Figure 7.4 there is no confounding if we do not adjust for \(L\), but there is if we do. The choice is a matter of taste with no practical implications, because identifiability depends only on whether conditional or unconditional exchangeability holds (Hernán and Robins 2020, 98).

NoteFine Point 7.2: Identification of Conditional and Unconditional Effects

Which effects can be identified depends on which variables are measured. In Figure 7.6:

  • Measuring only \(L_2\): no exchangeability given \(L_2\); no causal effects are identified.
  • Measuring \(L_2\) and \(L\): conditional exchangeability given \(\{L_2, L\}\) (but not given either alone). Identified are:
    • effects within joint strata of \(L\) and \(L_2\), via \(\operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L_2 = l_2\right]\mathclose{}\);
    • the unconditional effect, via \(\sum_{l, l_2} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L_2 = l_2\right]\mathclose{} \Pr[L = l, L_2 = l_2]\);
    • effects within strata of \(L\), via \(\sum_{l_2} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L_2 = l_2\right]\mathclose{} \Pr[L_2 = l_2 \mid L = l]\);
    • effects within strata of \(L_2\), via \(\sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L_2 = l_2\right]\mathclose{} \Pr[L = l \mid L_2 = l_2]\).
  • Measuring only \(L_1\): effects within strata of \(L_1\) and the unconditional effect.
  • Measuring \(L_1\) and \(L\): additionally, effects within joint strata of \(L_1\) and \(L\), and within strata of \(L\).
  • Measuring \(L\), \(L_1\), and \(L_2\): additionally, effects within joint strata of all three (Hernán and Robins 2020, 98).

4 7.4 Confounding and Confounders (pp. 98-101)


The structural approach needs prior knowledge of the causal DAG, including all shared causes (measured or not) of \(A\) and \(Y\); the backdoor criterion then says what to adjust for. The traditional approach instead labels as confounders the variables meeting mostly associational conditions, mandates adjusting for them, and declares confounding when adjusted and unadjusted estimates differ.

Definition 5 (Traditional Definition of Confounder) A variable is a confounder under the traditional approach if it

  1. is associated with the treatment,
  2. is associated with the outcome conditional on the treatment (often replaced by “in the untreated”), and
  3. does not lie on a causal pathway between treatment and outcome.

4.1 Where the Traditional Approach Agrees and Disagrees


  • Figures 7.1-7.3: \(L\) meets all three conditions, matching the backdoor criterion.
  • Figure 7.4: \(L\) meets all three conditions (it shares \(U_2\) with \(A\) and \(U_1\) with \(Y\), and is not on the causal pathway), so the traditional approach says to adjust for it, although there is no confounding and adjusting causes selection bias.
  • Figure 7.7 is a second example in which the traditional approach leads to harmful adjustment for \(L\).

Replacing condition 2 by the structural condition “it is a cause of the outcome” fixes Figure 7.4, but then \(L\) in Figure 7.2 would no longer count as a confounder, although it must be adjusted for (Technical Point 7.2).

Associational criteria are insufficient to characterize confounding. A definition of confounder that relies almost exclusively on statistical considerations can advise adjusting for a “confounder” even when structural confounding does not exist.

Change in estimate is not a criterion for confounding. Adjusted and unadjusted estimates can differ for reasons other than confounding, including selection bias from adjusting for non-confounders (Chapter 8) and the noncollapsibility of some effect measures (Fine Point 4.3). Attempts to define confounding by change in estimates were abandoned long ago because of these problems (Hernán and Robins 2020, 100).

Technically, investigators do not need full structural knowledge; they only need to know a set of variables that guarantees conditional exchangeability.

4.2 Confounding Is Absolute; Confounder Is Relative


The structural approach first identifies the sources of confounding (the common causes of treatment and outcome) and then a sufficient adjustment set. Whether a variable belongs to a sufficient set depends on the other variables in it. In Figures 7.2 and 7.3, \(L\) is needed only because \(U\) is unmeasured; given \(U\), \(L\) would not be a confounder. Given a causal DAG, confounding is an absolute concept, whereas confounder is a relative one (Hernán and Robins 2020, 100).

The structural approach has two advantages:

  1. it prevents inconsistencies between beliefs and actions (if you believe Figure 7.4, you will not adjust for \(L\), whatever a non-structural definition says);
  2. it makes the researchers’ assumptions explicit, so others can criticize them.

VanderWeele and Shpitser (2013) also proposed a formal definition of confounder. No approach guarantees that the researchers’ DAG is correct, so a chosen adjustment set may still fail to remove confounding or may introduce selection bias.

Figure 7.8. L is a surrogate (proxy) for the unmeasured common cause U of A and Y.

Figure 7.8. L is a surrogate (proxy) for the unmeasured common cause U of A and Y.

NoteFine Point 7.3: Surrogate Confounders

In Figure 7.8 the unmeasured \(U\) (e.g., socioeconomic status) confounds the effect of physical activity \(A\) on cardiovascular disease \(Y\), and the measured \(L\) (e.g., income) is a proxy for \(U\). \(L\) is not on a backdoor path, but adjusting for it may remove some of the confounding by \(U\); if \(L\) were perfectly correlated with \(U\), conditioning on \(L\) would be the same as conditioning on \(U\). If \(L\) is a binary, nondifferentially misclassified version of \(U\), conditioning on \(L\) partially blocks \(A \leftarrow U \rightarrow Y\) under some weak conditions (Greenland 1980; Ogburn and VanderWeele 2012). So one typically prefers to adjust for \(L\).

Variables that can reduce, but never completely eliminate, confounding bias without lying on a backdoor path are surrogate confounders (Hernán and Robins 2020, 100). One strategy is to measure and adjust for as many surrogate confounders as possible (see Chapter 18).

NoteTechnical Point 7.2: Fixing the Traditional Definition of Confounder

The traditional definition relies on two incorrect statistical criteria (conditions 1 and 2) and one incorrect causal criterion (condition 3). To fix it:

  1. Replace condition 3 by: there exist variables \(L\) and \(U\) such that \(Y^a \perp\!\!\!\perp A \mid L, U\). This implies that \(L\) is not on a causal pathway from \(A\) to \(Y\) and that \(\operatorname{E}\mathopen{}\left[Y^a \mid L = l, U = u\right]\mathclose{}\) is identified by \(\operatorname{E}\mathopen{}\left[Y \mid L = l, U = u, A = a\right]\mathclose{}\).
  2. Replace conditions 1 and 2 by: \(U\) can be split into disjoint subsets \(U_1\) and \(U_2\) (\(U = U_1 \cup U_2\), \(U_1 \cap U_2 = \emptyset\)) such that
    1. \(U_1\) and \(A\) are not associated within strata of \(L\), and
    2. \(U_2\) and \(Y\) are not associated within joint strata of \(A\), \(L\), and \(U_1\).

If both conditions hold, \(U\) is a non-confounder given data on \(L\). These conditions were proposed by Robins (1997, Theorem 4.3) and discussed by Greenland, Pearl, and Robins (1999), who used them to show there is no confounding in Figure 7.4 (Hernán and Robins 2020, 101).

5 7.5 Single-World Intervention Graphs (pp. 101-102)


The equivalence between exchangeability and the backdoor criterion seems “rather magical” because counterfactuals do not appear on causal diagrams. Single-world intervention graphs (SWIGs) put the counterfactual variables on the graph, so that exchangeability can be read off directly by d-separation.

Definition 6 (Single-World Intervention Graph (SWIG)) A SWIG depicts the variables and causal relations that would be observed in a hypothetical world in which all individuals received treatment level \(a\), a counterfactual world created by a single intervention. It is obtained from a causal DAG as follows:

  • Split the treatment node into a left side \(A\) (the natural value of treatment, the value that would have been observed without intervention), which keeps all arrows into \(A\), and a right side \(a\) (the intervention value), which inherits all arrows out of \(A\). There is no arrow from \(A\) to \(a\), because \(a\) is a constant.
  • Replace each descendant of treatment by its counterfactual, e.g., \(Y\) by \(Y^a\). Non-descendants of \(A\) keep their factual labels, because treatment does not affect them.

Example 2 (SWIGs for Figures 7.2 and 7.4)  

  • Figure 7.9 is the SWIG of Figure 7.2: \(U \rightarrow L \rightarrow A\), \(U \rightarrow Y^a\), \(a \rightarrow Y^a\). Every path between \(Y^a\) and \(A\) is blocked by conditioning on \(L\), so \(Y^a \perp\!\!\!\perp A \mid L\).
  • Figure 7.10 is the SWIG of Figure 7.4: \(U_2 \rightarrow A\), \(U_2 \rightarrow L \leftarrow U_1 \rightarrow Y^a\), \(a \rightarrow Y^a\). Without conditioning, the path through the collider \(L\) is blocked, so \(Y^a \perp\!\!\!\perp A\). Conditioning on \(L\) opens \(Y^a \leftarrow U_1 \rightarrow L \leftarrow U_2 \rightarrow A\), so \(Y^a \perp\!\!\!\perp A \mid L\) fails.

On the SWIG, \(Y^a\) is d-separated from \(A\) given \(L\) if and only if \(L\) is a non-descendant of \(A\) that blocks all backdoor paths from \(A\) to \(Y\). This is the simple SWIG-based argument for the equivalence stated in Section 7.2.

Origins: Richardson and Robins (2013) showed that SWIGs overcome some shortcomings of the twin causal diagrams previously proposed by Balke and Pearl (1994) (Hernán and Robins 2020, 101). Under an FFRCISTG model, d-separation on the SWIG also implies statistical independence.

Is the natural value measurable? The book assumes the natural value \(A\) is well defined even though it is generally not observed under intervention \(a\). It notes experiments suggesting that electroencephalogram recordings can detect a choice up to 1/2 second before people become conscious of it, which in principle would allow measuring \(A\) and still intervening.

Arrows from \(a\): in the single-intervention world \(a\) is a constant and cannot affect other variables; SWIGs nonetheless draw arrows from \(a\) to keep track of the variables directly affected by \(A\) in the original DAG.

6 7.6 Confounding Adjustment (pp. 102-107)


Without randomization, causal inference relies on the uncheckable assumption that the measured \(L\) is a sufficient set for confounding adjustment. Under that assumption, methods that adjust for \(L\) fall into two categories:

  • G-methods (“g” for generalized): standardization, IP weighting, and g-estimation. They estimate the causal effect in the whole population or in any subset. The heart transplant study used standardization (Section 2.4) and IP weighting (Section 2.5); Part II covers the parametric g-formula, IP weighting of marginal structural models, and g-estimation of structural nested models.
  • Conventional stratification-based methods: stratification (including restriction) and matching. They estimate the association between \(A\) and \(Y\) within subsets defined by \(L\) (Sections 4.4 and 4.5); Part II covers their model-based extension, outcome regression.

“Deleting” versus conditioning. Standardization and IP weighting simulate the \(A\)-\(Y\) association in the population if the backdoor paths through \(L\) did not exist; IP weighting, for example, creates a pseudo-population in which \(A\) is independent of \(L\), “deleting” the arrow \(L \rightarrow A\). Stratification does not delete that arrow but computes the effect in a subset, represented by a selection box. Part III explains why deleting the arrow is advantageous with time-varying treatments, why g-estimation is the only generally valid stratification-based method, and why conventional stratification-based methods can cause selection bias with time-varying confounders (Chapter 20).

A common variation of stratification and matching replaces \(L\) by the estimated probability of treatment \(\Pr[A = 1 \mid L]\), the propensity score (Rosenbaum and Rubin 1983; Chapter 15).

NoteFine Point 7.4: Confounders Cannot Be Descendants of Treatment, but Can Be in the Future of Treatment

In Figure 7.11, \(L\) is a descendant of \(A\) that blocks all backdoor paths, and the effect of \(A\) on \(Y\) is entirely through \(L\). Conditioning on \(L\) opens no collider path, but it blocks the causal pathway, so adjusting for \(L\) does not remove bias. Hence conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) must fail. On the SWIG (Figure 7.12) \(L\) is replaced by the counterfactual \(L^a\), and we can read off \(Y^a \perp\!\!\!\perp A \mid L^a\) but not \(Y^a \perp\!\!\!\perp A \mid L\), since \(L\) is not on the graph. (Under an FFRCISTG model, an independence that cannot be read off the SWIG cannot be assumed to hold.)

The problem is that \(L\) is a descendant of \(A\), not that \(L\) occurs after \(A\). Without the arrow \(A \rightarrow L\), \(L\) would be a non-descendant that blocks all backdoor paths, and adjusting for it would remove all bias even if \(L\) were measured after \(A\). What matters is the topology of the causal diagram, not the time order of the nodes (Hernán and Robins 2020, 103).

6.1 Methods That Do Not Require Conditional Exchangeability


Some methods handle confounding without conditional exchangeability:

  • difference-in-differences and negative outcome controls (Technical Point 7.3);
  • proximal inference (Technical Point 7.3);
  • the front door criterion (Technical Point 7.4);
  • instrumental variable estimation (Chapter 16).

They replace conditional exchangeability with other assumptions that are just as unverifiable, so the choice of method depends on which unverifiable assumptions are more plausible in a given setting.

6.2 Expert Knowledge and the Critic


Conditional exchangeability may be unrealistic, but expert knowledge about the causal structure helps get close to it:

  • measure many non-descendants \(L\) of treatment in the hope of blocking all backdoor paths;
  • avoid adjusting for variables affected by treatment or by the outcome;
  • when several causal structures are plausible, analyze under each and state the assumptions each requires.

A critic who says only “your observational study may be confounded” makes a logical, not a scientific, statement: it is true of every observational study. A scientific criticism names a source, such as “confounding due to cigarette smoking, a common cause through which a backdoor path may remain open”. That gives a testable challenge: adjust for smoking or, if smoking was not measured, conduct a sensitivity analysis (Hernán and Robins 2020, 104–5).

Hernán et al. (2002) give a practical example of using expert knowledge of the causal structure to evaluate confounding. One can never be certain that the causal structures considered include the true one; this uncertainty is unavoidable with observational data.

Valid causal inference from observational data also requires the absence of selection and measurement biases, which, unlike confounding, can arise in randomized experiments too. Chapter 8 turns to selection bias.

Left: Figure 7.13, with negative outcome control C (the pre-treatment outcome). Right: Figure 7.14, in which M fully mediates the effect of A on Y. U is unmeasured.

Left: Figure 7.13, with negative outcome control C (the pre-treatment outcome). Right: Figure 7.14, in which M fully mediates the effect of A on Y. U is unmeasured.

NoteTechnical Point 7.3: Difference-in-Differences and Negative Outcome Controls

Suppose unmeasured \(U\) (e.g., history of heart disease) confounds the effect of aspirin \(A\) on blood pressure \(Y\), and we also measured the outcome just before treatment, \(C\), a negative outcome control (Figure 7.13). \(A\) cannot cause \(C\), so \(\operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\) measures additive confounding for the effect of \(A\) on \(C\).

Under additive equi-confounding, \(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{} = \operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\), the effect in the treated is identified:

\[\begin{align} \operatorname{E}\mathopen{}\left[Y^1 - Y^0 \mid A = 1\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^1 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} && \text{(linearity)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(add and subtract } \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y^0 \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^0 \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(consistency)} \\ &= \mathopen{}\left(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\right)\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[C \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[C \mid A = 0\right]\mathclose{}\right)\mathclose{} && \text{(equi-confounding)} \end{align}\]

The association of \(A\) with \(Y\) (effect plus confounding) minus the confounding measured through \(C\). The arrow \(C \rightarrow Y\) is not necessary for \(C\) to be a negative outcome control. This is difference-in-differences (Card 1990; Meyer 1995; Angrist and Krueger 1999). It is a somewhat restrictive use of negative outcome controls: it needs \(Y\) and \(C\) on the same scale and additive equi-confounding; Sofer et al. (2016) describe more general methods.

With both a negative outcome control \(C\) and a negative treatment control \(Z\), the effect can be identified nonparametrically despite unmeasured \(U\) under additional assumptions; for discrete \(U\), \(C\), \(Z\) with \(C\) and \(Z\) having at least as many levels as \(U\), identification holds quite generally (Miao et al. 2018). This is proximal causal inference (Cui et al. 2024); Figure 7.15 is an example (Hernán and Robins 2020, 106).

NoteTechnical Point 7.4: The Front Door Criterion

In Figure 7.14, unmeasured \(U\) confounds \(A\) and \(Y\), and \(M\) fully mediates the effect of \(A\) on \(Y\) and shares no unmeasured causes with \(A\) or \(Y\). Standardization and IP weighting cannot be used because \(U\) is unavailable, but Pearl (1995) showed that \(\Pr[Y^a = 1]\) is identified by the front door formula

\[\Pr[Y^a = 1] = \sum_m \Pr[M = m \mid A = a] \sum_{a'} \Pr[Y = 1 \mid M = m, A = a'] \Pr[A = a']\]

(Pearl’s “backdoor formula” is what the book calls standardization or the point-treatment g-formula.)

Proof.

  1. By the law of total probability, \(\Pr[Y^a = 1] = \sum_m \Pr[M^a = m] \Pr[Y^a = 1 \mid M^a = m]\).
  2. \(\Pr[M^a = m] = \Pr[M = m \mid A = a]\), because there is no confounding for the effect of \(A\) on \(M\) (\(M^a \perp\!\!\!\perp A\)).
  3. \(\Pr[Y^a = 1 \mid M^a = m] = \Pr[Y^m = 1]\), because
    1. \(Y^a = Y^m\) when \(M^a = m\) (\(A\) affects \(Y\) only through \(M\)), and
    2. \(Y^m \perp\!\!\!\perp M^a\) by d-separation on the SWIG for the joint intervention setting \(M\) to \(m\) and \(A\) to \(a\).
  4. \(\Pr[Y^m = 1] = \sum_{a'} \Pr[Y = 1 \mid M = m, A = a'] \Pr[A = a']\), by conditional exchangeability \(Y^m \perp\!\!\!\perp M \mid A\) on the SWIG intervening on \(M\) alone (standardization over \(A\)).
  5. Substituting steps 2-4 into step 1 gives the front door formula.

This proof requires well-defined counterfactuals \(Y^m\); Technical Points 21.11 and 21.12 give proofs without that condition (Hernán and Robins 2020, 107).

7 Summary


  1. Confounding is the bias due to common causes of treatment and outcome, which create open backdoor paths.
  2. No confounding is equivalent to marginal exchangeability; with confounding, a sufficient set \(L\) of measured non-descendants of \(A\) gives conditional exchangeability, and standardization or IP weighting identifies the average causal effect.
  3. Under faithfulness and an FFRCISTG model, \(Y^a \perp\!\!\!\perp A \mid L\) holds if and only if \(L\) satisfies the backdoor criterion.
  4. Adjusting for a collider such as \(L\) in Figure 7.4 (M-bias) creates selection bias; in Figure 7.5 no adjustment set based on \(L\) alone works.
  5. The traditional associational definition of confounder can recommend harmful adjustment; confounding is absolute, but “confounder” is relative to the other adjustment variables.
  6. SWIGs split the treatment node and show exchangeability directly via d-separation.
  7. G-methods and conventional stratification-based methods both require conditional exchangeability; difference-in-differences, proximal inference, the front door criterion, and instrumental variables rely on other unverifiable assumptions.

Looking ahead:

  • Chapter 8: selection bias, the other structural source of lack of exchangeability through open paths.
  • Chapter 9: measurement bias.
  • Chapters 12-15: IP weighting, standardization, g-estimation, and outcome regression in practice.

8 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top