Causal inference needs expert knowledge and untestable assumptions about the causal network linking treatment, outcome, and other variables. In the simple settings of Chapters 1-5 that network could stay implicit; in realistic settings we must state what we know and what we assume.
This chapter introduces causal diagrams, a graphical tool for encoding qualitative causal assumptions. The book’s advice: draw your assumptions before your conclusions.
Definition 1 (Causal Markov Assumption) Conditional on its direct causes (parents), any variable on a causal DAG is independent of any variable for which it is not a cause (its non-descendants) (Hernán and Robins 2020, 76–77).
A consequence: all common causes of any pair of variables on the graph must also be on the graph, even if unmeasured.
Technical Point 6.1: Causal DAGs, Formally
For a DAG \(G\) with nodes \(V = (V_1, \ldots, V_M)\), ordered so that \(V_m\) is not an ancestor of \(V_j\) when \(m > j\):
The distribution of \(V\) is Markov with respect to \(G\) if each \(V_j\) is independent of its non-descendants given its parents, which is equivalent to the Markov factorization
\[ f(v) = \prod_{j=1}^{M} f(v_j \mid pa_j). \]
Example 1 (Markov Factorization Implies an Independence) Take the DAG \(A \leftarrow L \to Y\) (Figure 3). The parents are: none for \(L\); \(L\) for \(A\); \(L\) for \(Y\). Start from the chain rule, which holds for any distribution, and apply the Markov factorization:
\[ \begin{aligned} f(l, a, y) &= f(l)\, f(a \mid l)\, f(y \mid l, a) && \text{(chain rule)} \\ &= f(l)\, f(a \mid l)\, f(y \mid l) && \text{(the only parent of } Y \text{ is } L\text{)}. \end{aligned} \]
Dividing both sides by \(f(l)\):
\[ \begin{aligned} f(a, y \mid l) &= \frac{f(l, a, y)}{f(l)} && \text{(definition of conditional density)} \\ &= f(a \mid l)\, f(y \mid l) && \text{(substitute the factorization)}, \end{aligned} \]
so \(A \perp\!\!\!\perp Y \mid L\): the missing arrow from \(A\) to \(Y\) implies a conditional independence.
Each causal DAG has an underlying counterfactual model (Technical Points 6.2 and 6.3). That model justifies the intuitive graphical rules in this chapter, but conventional diagrams do not show the counterfactual variables.
Single World Intervention Graphs (SWIGs) (Richardson and Robins 2013) put counterfactuals on the graph; they are introduced in Chapter 7.
Technical Point 6.2: Nonparametric Structural Equation Models
A nonparametric structural equation model (NPSEM) for a DAG with ordered nodes \(V_1, \ldots, V_M\) assumes unobserved errors \(\epsilon_m\) and unknown deterministic functions \(f_m\) such that
Only the parents of \(V_m\) have a direct effect on it. Every variable can be intervened on, and all factual and counterfactual values are obtained recursively.
Example 2 (The NPSEM for Book Figure 6.1) The structural equations are \(L = f_L(\epsilon_L)\), \(A = f_A(L, \epsilon_A)\), \(Y = f_Y(L, A, \epsilon_Y)\). Setting \(A\) to \(a\) leaves \(L\) unchanged, because \(L\) is not a descendant of \(A\), so the counterfactual outcome is obtained by substituting \(a\) for \(A\): \(Y^a = f_Y(L, a, \epsilon_Y)\).
Technical Point 6.3: Which Independences?
An FCISTG alone does not imply the causal Markov assumption; independence assumptions must be added.
| Model | Assumption | Fig. 6.2 implies |
|---|---|---|
| NPSEM-IE (Pearl) | all errors \(\epsilon_m\) mutually independent | full exchangeability \((Y^{a=0}, Y^{a=1}) \perp\!\!\!\perp A\) |
| FFRCISTG (Robins 1986) | one-step-ahead counterfactuals jointly independent | marginal exchangeability \(Y^a \perp\!\!\!\perp A\) for each \(a\) |
Both imply the causal Markov assumption. An NPSEM-IE is an FFRCISTG, not vice versa. Unless stated otherwise, the book’s causal DAGs represent an FFRCISTG.
Three examples, all with no conditioning:
| Figure | Structure | Example | \(A\) and \(Y\) |
|---|---|---|---|
| Figure 2 | \(A \to Y\) | aspirin prevents heart disease (randomized) | associated |
| Figure 3 | \(A \leftarrow L \to Y\) | lighter, smoking, lung cancer | associated (no effect of \(A\)) |
| Figure 4 | \(A \to L \leftarrow Y\) | haplotype, heart disease, smoking | independent |
Definition 2 (Path; Causal Path) A path between \(R\) and \(S\) is a route connecting them along a sequence of edges that visits no variable more than once. A path is causal if all its arrows point in the same direction; otherwise it is noncausal (Hernán and Robins 2020, 78, margin note).
Picture paths as pipes through which association flows. Association is symmetric, so it flows regardless of the direction of the arrows.
Definition 3 (Collider) A variable on a path is a collider on that path if two arrowheads on the path collide at it, as \(L\) does on \(A \to L \leftarrow Y\).
Summary: two variables are marginally associated if one causes the other or if they share a common cause; otherwise they are marginally independent.
A square box around a node means we condition on it (e.g., restrict to one of its levels).
| Structure | Conditioning on | Path | \(A\), \(Y\) given the conditioning variable |
|---|---|---|---|
| \(A \to B \to Y\) (Fig. 6.5) | mediator \(B\) | blocked | independent: \(A \perp\!\!\!\perp Y \mid B\) |
| \(A \leftarrow L \to Y\) (Fig. 6.6) | common cause \(L\) | blocked | independent: \(A \perp\!\!\!\perp Y \mid L\) |
| \(A \to L \leftarrow Y\) (Fig. 6.7) | collider \(L\) | opened | associated |
In Figure 4, restrict to people with heart disease (\(L=1\), book Figure 6.7). Among them, someone without the haplotype is more likely to have the other cause, smoking: \(A\) and \(Y\) are inversely associated given \(L=1\).
In the extreme, if \(A\) and \(Y\) were the only causes of \(L\), then among those with \(L=1\) the absence of one would perfectly predict the presence of the other.
Conditioning on \(C\), a variable affected by the collider \(L\), also opens \(A \to L \leftarrow Y\). The path stays blocked only if we condition on neither \(L\) nor \(C\).
Two variables can be associated because
A fourth, nonstructural source is chance (random variability), which shrinks as the study population grows. Until Chapter 10 the book assumes a very large population, so all associations discussed are structural.
Fine Point 6.1: D-Separation
A path is blocked if and only if
Otherwise it is open. Two variables are d-separated if all paths between them are blocked; otherwise they are d-connected. Two sets are d-separated if each variable in one is d-separated from every variable in the other.
Example 3 (D-Separation in Book Figures 6.1 and 6.4)
Fine Point 6.2: Faithfulness
An arrow \(A \to Y\) means \(A\) affects \(Y\) for at least one individual, yet the average causal effect and the association can still both be null, when effects in different individuals cancel exactly. Such exact cancellation is rare, so the book assumes faithfulness: lack of d-separation can almost always be equated with a non-zero association.
Example 4 (Unfaithfulness from Cancelling Effects) In the book’s Table 4.1, transplant increases the risk of death in women and decreases it in men, each half of the population, and the effects cancel exactly: \(\Pr[Y^{a=1}=1] = \Pr[Y^{a=0}=1]\). Figure 2 is still the correct diagram, because \(A\) affects every individual’s \(Y\), but the expected association is absent: the distribution is not faithful to the DAG (Hernán and Robins 2020, Fine Point 6.2, p. 83).
Standardization and IP weighting (Chapter 2) can also be derived from causal graph theory, as part of what is sometimes called the do-calculus (reviewed in Pearl 2009). So using counterfactuals in Chapters 1-5 privileged a notation, not an approach.
Whatever the notation, causal inference via standardization or IP weighting requires exchangeability, positivity, and consistency. This section covers positivity and consistency; exchangeability is translated into graphs in Section 6.5 and Chapters 7-8.
In the book’s diagrams, positivity is implicit unless stated and consistency is embedded in the notation.
Diagrams read as NPSEMs with independent errors seem to give all variables equal status, which can mislead when some nodes correspond to ill-defined interventions.
When there are several ways to intervene on \(A\) and some of them affect \(Y\) directly, “the effect of \(A\) on \(Y\)” is unclear: losing weight by caloric restriction, by exercise, or by genetic manipulation would lead to different mortality. Being explicit about the intervention is a step toward a well-defined effect, relevant data, and the right adjustment variables.
Fine Point 6.3: Discovery of Causal Structure
Discovery is learning parts of the causal structure from data. It sometimes works if we assume faithfulness, but often it does not: a strong association between \(B\) and \(C\) fits \(B \to C\), \(C \to B\), a shared unmeasured cause, a conditioned-on common effect, and combinations.
Example 5 (Learning \(Z \to A \to Y\) (Hernán and Robins 2020, Fine Point 6.3, p. 85)) Suppose \(Z\) precedes \(A\), which precedes \(Y\), and with infinite data all three are marginally associated and the only conditional independence is \(Z \perp\!\!\!\perp Y \mid A\). Under faithfulness the only consistent causal DAGs are \(Z \to A \to Y\), possibly with a common cause of \(Z\) and \(A\) in addition to, or instead of, \(Z \to A\). Then there is no unmeasured common cause of \(A\) and \(Y\), and the average causal effect is identified by \(\operatorname{E}\mathopen{}\left[Y \mid A=1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A=0\right]\mathclose{}\).
Definition 4 (Systematic Bias) There is systematic bias when the data are insufficient to identify the causal effect even with an infinite sample size. Informally: any structural association between treatment and outcome that does not arise from the causal effect of treatment on outcome (Hernán and Robins 2020, 85).
Lack of exchangeability causes bias even when treatment has no effect: bias under the null.
Example 6 (Table 3.1) In the book’s observational study of Table 3.1 the causal risk ratio was 1, whereas the associational risk ratio was 1.26 (Hernán and Robins 2020, 86).
Any structure that causes bias under the null also causes bias under the alternative (a non-null effect); the converse is false.
A third source is measurement bias (information bias), from mismeasured treatment, outcome, or covariates; some types also cause bias under the null (Chapter 9). Random variability is a separate, nonstructural source (Chapter 10).
Causal diagrams are good at locating sources of bias, but less helpful for illustrating effect modification.
Transplant \(A\) is randomized, so association is causation; investigators stratify by quality of care \(V\) and find additive effect modification.
A surrogate effect modifier is associated with the causal effect modifier \(V\) but does not affect \(Y\). The association can come from any structure:
| Book figure | Surrogate | Structure linking it to \(V\) |
|---|---|---|
| 6.14 | cost of treatment \(S\) | cause and effect: \(V \to S\) |
| 6.15 | passport nationality \(P\) | common cause: residence \(U\) affects \(V\) and \(P\) |
| 6.16 | bottled mineral water \(W\) | conditioning on a common effect: \(V\) and \(W\) affect cost \(S\), analysis restricted to \(S=0\) |
Causal diagrams are agnostic about interaction between two treatments \(A\) and \(E\). They can encode it if augmented with nodes for sufficient-component causes (Chapter 5), with deterministic arrows from the treatments to those nodes; the book develops such diagrams in Chapter 8.
Fine Point 6.4: Evidence That a State Has Well-Defined Counterfactuals
Consider drug \(Z\), systolic blood pressure \(A\) (a state, not directly manipulable), and stroke \(Y\). If (i) \(Z\) is associated with \(A\) and \(Y\), (ii) \(A\) and \(Y\) are associated, and (iii) \(Z \perp\!\!\!\perp Y \mid A\), then under faithfulness the only causal DAG has \(A \to Y\), no \(Z \to Y\), and no unmeasured common cause of \(A\) and \(Y\) or of \(Z\) and \(Y\) (the discovery argument of Fine Point 6.3).