Chapter 6: Graphical Representation of Causal Effects

Published

Last modified: 2026-10-09 10:19:15 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~43% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Causal inference needs expert knowledge and untestable assumptions about the causal network linking treatment, outcome, and other variables. In the simple settings of Chapters 1-5 that network could stay implicit; in realistic settings we must state what we know and what we assume.

This chapter introduces causal diagrams, a graphical tool for encoding qualitative causal assumptions.

TipDraw Your Assumptions Before Your Conclusions

Put the causal structure you are willing to assume into a diagram before you analyze the data. A diagram makes the assumptions explicit, which clarifies conceptual problems and helps investigators communicate (Hernán and Robins 2020, 75).

This chapter is based on Hernán and Robins (2020, chap. 6, pp. 75-89). Chapters 6-9 of the book use causal diagrams to conceptualize problems.

1 6.1 Causal Diagrams (pp. 75-77)


Figure 1: Book Figure 6.1: disease severity \(L\), heart transplant \(A\), death \(Y\).

Definition 1 (Directed Acyclic Graph) A directed acyclic graph (DAG) is a set of nodes joined by arrows (directed edges) in which no node can be reached from itself by following arrows in their direction.

Example 1 (Book Figure 6.1 Is a DAG) Figure 1 has nodes \(L\), \(A\), \(Y\) and arrows \(L \to A\), \(L \to Y\), \(A \to Y\). The diagrams in this chapter put earlier variables to the left, so every arrow points rightward and following arrows never returns to a node already visited. Adding an arrow \(Y \to L\) would create the cycle \(L \to A \to Y \to L\), in which \(L\) would be a cause of itself, and the graph would no longer be a DAG.


Definition 2 (Parents, Descendants, and Ancestors) In a DAG:

  • the parents of a node \(V\) are the nodes with an arrow pointing directly into \(V\);
  • a node \(W\) is a descendant of \(V\), and \(V\) is an ancestor of \(W\), if \(W\) can be reached from \(V\) by following one or more arrows in their direction;
  • the non-descendants of \(V\) are the nodes other than \(V\) that are not descendants of \(V\).

Example 2 (Parents and Descendants in Book Figure 6.1) In Figure 1, \(L\) has no parents, the only parent of \(A\) is \(L\), and the parents of \(Y\) are \(L\) and \(A\). \(A\) and \(Y\) are descendants of \(L\), \(Y\) is the only descendant of \(A\), and the only non-descendant of \(A\) is its parent \(L\).


Definition 3 (Causal DAG) A causal DAG is a DAG (Definition 1) whose nodes are random variables and in which:

  1. an arrow \(V \to W\) means that \(V\) has a direct causal effect on \(W\) for at least one individual in the population, where “direct” means not mediated by the other variables on the graph;
  2. the absence of an arrow from \(V\) to \(W\) means that \(V\) has no direct causal effect on \(W\) for any individual;
  3. every common cause of any pair of variables on the graph, measured or not, is also on the graph;
  4. every variable is a cause of its descendants (Definition 2).

These conditions combine the book’s description on p. 75 with its formal definition in Technical Point 6.1 (Hernán and Robins 2020, 75–76).

Example 3 (Reading Book Figure 6.1 as a Causal DAG) In Figure 1, \(L\) is disease severity, \(A\) is heart transplant, and \(Y\) is death. The arrow \(L \to A\) says that severity affects whether at least one individual receives a transplant, and the arrow \(A \to Y\) says that transplant directly affects at least one individual’s death. Drawing the graph as a causal DAG also asserts that \(L\) is the only common cause of any two of these variables: for instance, no unmeasured factor affects both transplant and death.

WarningWhat an Arrow Does Not Say

A standard causal diagram does not say whether an effect is harmful or protective. When a variable has two causes, as \(Y\) does in Figure 1, the diagram does not say how they interact (Hernán and Robins 2020, 75).

The modern theory of causal diagrams arose in computer science and artificial intelligence; see Pearl (2009) and Spirtes, Glymour, and Scheines (2000).


1.1 Directed Acyclic Graphs and the Causal Markov Assumption

Definition 4 (Markov Property) The joint distribution of the variables on a DAG \(G\) is Markov with respect to \(G\) if, conditional on its parents, each variable is independent of its non-descendants (Definition 2).

Example 4 (The Markov Property for a Chain) Take the DAG \(V_1 \to V_2 \to V_3\). \(V_1\) has no non-descendants, and the only non-descendant of \(V_2\) is its parent \(V_1\), so these two variables add no requirement. The non-descendants of \(V_3\) are \(V_1\) and its parent \(V_2\), so the Markov property requires exactly \(V_3 \perp\!\!\!\perp V_1 \mid V_2\).


Definition 5 (Causal Markov Assumption) The causal Markov assumption is that the joint distribution of the variables on a causal DAG (Definition 3) is Markov with respect to that DAG (Definition 4): conditional on its direct causes (its parents), each variable is independent of every variable that it does not cause (its non-descendants) (Hernán and Robins 2020, 76–77).

Example 5 (Lighters and Lung Cancer) Suppose smoking \(L\) causes both carrying a lighter \(A\) and lung cancer \(Y\), and carrying a lighter does not cause lung cancer, so the causal DAG is \(A \leftarrow L \to Y\). \(Y\) is a non-descendant of \(A\), whose only parent is \(L\), so the causal Markov assumption gives \(A \perp\!\!\!\perp Y \mid L\): among people with the same smoking status, carrying a lighter tells us nothing more about lung cancer.

Remark 1 (Why Every Common Cause Must Be on the Graph). Condition 3 of Definition 3 is needed for the causal Markov assumption to be credible. In Example 5, leave smoking \(L\) off the graph. \(A\) and \(Y\) then have no parents and no arrow between them, so the Markov property would require \(A \perp\!\!\!\perp Y\). But carrying a lighter and lung cancer share the cause smoking, so they are generically (for all but special parameter values) associated, and the distribution is then not Markov with respect to the graph without \(L\).


Proposition 1 (Markov Factorization) Let \(G\) be a DAG with nodes \(V = (V_1, \ldots, V_M)\), numbered so that no node is an ancestor of a node with a smaller number (such a numbering exists because \(G\) has no cycles). Suppose \(V\) is discrete with probability mass function \(f\), and \(f(v) > 0\) for every \(v\) in the product of the variables’ ranges. Write \(PA_j\) for the parents of \(V_j\) and \(pa_j\) for their values in \(v\). Then the distribution of \(V\) is Markov with respect to \(G\) (Definition 4) if and only if, for every \(v\),

\[ f(v) = \prod_{j=1}^{M} f(v_j \mid pa_j). \tag{1}\]

Proof. Suppose first that the distribution is Markov with respect to \(G\). By the chain rule, \(f(v) = \prod_{j=1}^{M} f(v_j \mid v_1, \ldots, v_{j-1})\). A descendant of \(V_j\) has \(V_j\) as an ancestor, so it has a larger number than \(V_j\); hence \(V_1, \ldots, V_{j-1}\) are all non-descendants of \(V_j\). They include every parent of \(V_j\), because a parent is an ancestor and so has a smaller number. The Markov property then gives \(f(v_j \mid v_1, \ldots, v_{j-1}) = f(v_j \mid pa_j)\), which is Equation 1.

Conversely, suppose Equation 1 holds, and fix \(j\). Let \(D\) be the set of descendants of \(V_j\) and \(N\) its set of non-descendants. No variable in \(N\) has a parent in \(D\) or equal to \(V_j\), since it would then be a descendant of \(V_j\). Sum Equation 1 over the values of the variables in \(D\), one at a time, starting with the one with the largest number. Each variable summed out at that point has all its children in \(D\) with larger numbers, already summed out, so it appears only in its own factor, which sums to 1. What remains is

\[ f(v_j, n) = f(v_j \mid pa_j) \prod_{k : V_k \in N} f(v_k \mid pa_k), \]

where \(n\) is the value of \(N\) in \(v\) and the product does not involve \(v_j\). Summing over \(v_j\) gives \(f(n) = \prod_{k : V_k \in N} f(v_k \mid pa_k)\), so \(f(v_j \mid n) = f(v_j \mid pa_j)\). The parents of \(V_j\) are in \(N\), so this conditional distribution depends on \(n\) only through \(pa_j\): \(V_j\) is independent of its non-descendants given its parents.

Example 6 (Markov Factorization Implies an Independence) Take the lighter DAG \(A \leftarrow L \to Y\) of Example 5. The parents are: none for \(L\); \(L\) for \(A\); \(L\) for \(Y\). Start from the chain rule, which holds for any distribution, and apply Proposition 1:

\[ \begin{aligned} f(l, a, y) &= f(l)\, f(a \mid l)\, f(y \mid l, a) && \text{(chain rule)} \\ &= f(l)\, f(a \mid l)\, f(y \mid l) && \text{(the only parent of } Y \text{ is } L\text{)}. \end{aligned} \]

Dividing both sides by \(f(l)\):

\[ \begin{aligned} f(a, y \mid l) &= \frac{f(l, a, y)}{f(l)} && \text{(definition of conditional probability)} \\ &= f(a \mid l)\, f(y \mid l) && \text{(substitute the factorization)}, \end{aligned} \]

so \(A \perp\!\!\!\perp Y \mid L\): the missing arrow from \(A\) to \(Y\) implies a conditional independence.

NoteTechnical Point 6.1: Causal DAGs, Formally

The book’s Technical Point 6.1 states formally the definitions above: DAGs, parents, descendants, and ancestors (Definition 1, Definition 2), causal DAGs (Definition 3), the Markov property and its factorization (Definition 4, Proposition 1), and the causal Markov assumption (Definition 5). A causal DAG on its own makes no claim about the distribution of the data; the causal Markov assumption is what links the two (Hernán and Robins 2020, Technical Point 6.1, p. 76).


1.2 Randomized Experiments and Observational Studies

Figure 2: Book Figure 6.2: a marginally randomized experiment.

Example 7 (Randomized Experiments as Causal DAGs)  

  • If transplant is randomized with a probability that depends on severity \(L\), then \(L\) is a common cause of \(A\) and \(Y\) and must appear (condition 3 of Definition 3): Figure 1 depicts this conditionally randomized experiment.
  • If everyone has the same probability of transplant, \(L\) is not a common cause and can be omitted: Figure 2 depicts this marginally randomized experiment.
  • Figure 1 can also depict an observational study in which the only parent of \(A\) is \(L\), so that no other cause of \(Y\) affects \(A\).

Chapter 7 shows that being willing to draw Figure 1 for an observational study is the graphical translation of conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) for all \(a\) (Hernán and Robins 2020, 77).


1.3 Graphs and Counterfactuals

Each causal DAG has an underlying counterfactual model (Technical Points 6.2 and 6.3 below). That model justifies the intuitive graphical rules in this chapter, but conventional diagrams do not show the counterfactual variables.

Single World Intervention Graphs (SWIGs) (Richardson and Robins 2013) put counterfactuals on the graph; they are introduced in Chapter 7.

Causal diagrams encode both causation (our subject-matter knowledge) and the associations that the causal structure implies. That simultaneous representation is what makes them attractive (Hernán and Robins 2020, 77). The book’s treatment is informal, aiming at conceptual insight rather than rigor.


NoteTechnical Point 6.2: Nonparametric Structural Equation Models

In the book, every causal DAG stands for a counterfactual model. The basic one is defined next. Robins (1986) introduced it as a finest causally interpreted structural tree graph (FCISTG) “as detailed as the data”, and Pearl (2009) showed how to represent it with a DAG. For exposition the book assumes that every variable can be intervened on, although its statistical methods do not require this (Hernán and Robins 2020, Technical Point 6.2, p. ;77).

Definition 6 (Nonparametric Structural Equation Model) Let \(G\) be a DAG with nodes \(V_1, \ldots, V_M\), numbered as in Proposition 1. A nonparametric structural equation model (NPSEM) represented by \(G\) assumes that there are unobserved random variables (errors) \(\epsilon_1, \ldots, \epsilon_M\) and unknown deterministic functions \(f_1, \ldots, f_M\) such that, for each \(m\) and each value \(pa_m\) of the parents of \(V_m\), the one-step-ahead counterfactual \(V_m^{pa_m}\), the value \(V_m\) would take if its parents were set to \(pa_m\), equals \(f_m(pa_m, \epsilon_m)\). For a node \(V_m\) with no parents, such as \(V_1\), this reads \(V_m = f_m(\epsilon_m)\).

Counterfactuals are assumed to exist for interventions on every variable, and every factual and counterfactual value is obtained by applying these equations recursively in the numbered order.

Example 8 (The NPSEM for Book Figure 6.1) For Figure 1 the structural equations (Definition 6) are \(L = f_L(\epsilon_L)\), \(A = f_A(L, \epsilon_A)\), \(Y = f_Y(L, A, \epsilon_Y)\). Setting \(A\) to \(a\) leaves \(L\) unchanged, because \(L\) is not a descendant of \(A\), so the counterfactual outcome is obtained by substituting \(a\) for \(A\): \(Y^a = f_Y(L, a, \epsilon_Y)\).


NoteTechnical Point 6.3: Which Independences?

An FCISTG alone does not imply the causal Markov assumption (Definition 5): assumptions about the joint distribution of the errors, or of the counterfactuals, must be added. Two such models are defined next (Hernán and Robins 2020, Technical Point 6.3, p. 78).

Definition 7 (NPSEM with Independent Errors) An NPSEM-IE is an NPSEM (Definition 6) whose errors \(\epsilon_1, \ldots, \epsilon_M\) are mutually independent. This is the model that Pearl usually assumes.

Definition 8 (FFRCISTG) A finest fully randomized causally interpreted structured tree graph (FFRCISTG) is an FCISTG, that is, a model as in Definition 6, in which, for every value \(v = (v_1, \ldots, v_M)\) of the nodes, the one-step-ahead counterfactuals \(V_1^{pa_1}, \ldots, V_M^{pa_M}\), with each \(pa_m\) read off from \(v\), are mutually independent (Robins 1986).

Example 9 (Two Models for Book Figure 6.2) For Figure 2 the structural equations (Definition 6) are \(A = f_A(\epsilon_A)\) and \(Y^a = f_Y(a, \epsilon_Y)\).

  • As an NPSEM-IE (Definition 7), \(\epsilon_A \perp\!\!\!\perp\epsilon_Y\). Then \(A\), a function of \(\epsilon_A\), is independent of \((Y^{a=0}, Y^{a=1})\), a function of \(\epsilon_Y\): full exchangeability \((Y^{a=0}, Y^{a=1}) \perp\!\!\!\perp A\) holds.
  • As an FFRCISTG (Definition 8), the one-step-ahead counterfactuals for \(v = (a, y)\) are \(A\) and \(Y^a\), so the model gives \(Y^a \perp\!\!\!\perp A\) for each \(a\) separately: marginal exchangeability, which says nothing about \(A\) and the pair \((Y^{a=0}, Y^{a=1})\) jointly.

Proposition 2 (Every NPSEM-IE Is an FFRCISTG, Not Conversely) Every NPSEM-IE (Definition 7) is an FFRCISTG (Definition 8). The converse fails: some models that satisfy the FFRCISTG condition are not NPSEM-IE models.

Proof. Fix \(v\). In an NPSEM-IE, each one-step-ahead counterfactual \(V_m^{pa_m} = f_m(pa_m, \epsilon_m)\), with \(pa_m\) fixed by \(v\), is a function of \(\epsilon_m\) alone. Functions of mutually independent random variables are mutually independent, so these counterfactuals are mutually independent, as an FFRCISTG requires.

For the converse, take Figure 2 with \(Y^{a=0}\) and \(Y^{a=1}\) independent fair coin flips, and let \(A = 1\) if exactly one of them equals 1 and \(A = 0\) otherwise. Setting \(\epsilon_Y = (Y^{a=0}, Y^{a=1})\) and \(\epsilon_A = A\) gives an NPSEM. For each \(a\), whatever the value of \(Y^a\), the other, independent fair coin makes \(A\) equal 1 with probability \(1/2\), so \(A \perp\!\!\!\perp Y^a\) and the model is an FFRCISTG. But \(A\) is a function of \((Y^{a=0}, Y^{a=1})\) and is not constant, so full exchangeability fails. By Example 9, no NPSEM-IE for Figure 2 gives this joint distribution of \((A, Y^{a=0}, Y^{a=1})\).

Proposition 3 (Counterfactual Models Imply the Causal Markov Assumption) If the counterfactuals for a causal DAG \(G\) follow an FFRCISTG (Definition 8), and so in particular if they follow an NPSEM-IE (Proposition 2), then the joint distribution of the factual variables is Markov with respect to \(G\) (Definition 4). Robins (1986) proved this (Hernán and Robins 2020, Technical Point 6.3, p. 78).

NoteThe Default Counterfactual Model

Unless stated otherwise, a causal DAG in the book, and in these notes, represents an FFRCISTG (Definition 8) (Hernán and Robins 2020, Technical Point 6.3, p. 78).

Robins and Richardson (2010) showed that Robins’s original formulation is equivalent to the joint-independence assumption for a positive distribution, and that an NPSEM-IE makes many more independence assumptions than an FFRCISTG.

2 6.2 Causal Diagrams and Marginal Independence (pp. 77-79)


Three examples, all with no conditioning:

Figure 3: Book Figure 6.3: smoking \(L\) causes carrying a lighter \(A\) and lung cancer \(Y\).
Figure 4: Book Figure 6.4: haplotype \(A\) and smoking \(Y\) both cause heart disease \(L\), a collider.
Table 1: Marginal association between \(A\) and \(Y\) in three causal DAGs. “Associated” holds generically (for all but special parameter values).
Figure Structure Example \(A\) and \(Y\)
Figure 2 \(A \to Y\) aspirin prevents heart disease (randomized) associated
Figure 3 \(A \leftarrow L \to Y\) lighter, smoking, lung cancer associated (no effect of \(A\))
Figure 4 \(A \to L \leftarrow Y\) haplotype, heart disease, smoking independent

2.1 Paths and the Flow of Association

Definition 9 (Path; Causal Path) A path between \(R\) and \(S\) is a route connecting them along a sequence of edges that visits no variable more than once. A path is causal if all its arrows point in the same direction; otherwise it is noncausal (Hernán and Robins 2020, 78, margin note).

Example 10 (Paths in Book Figure 6.1) In Figure 1 there are two paths between \(A\) and \(Y\): \(A \to Y\), which is causal, and \(A \leftarrow L \to Y\), which is noncausal because its two arrows point in different directions.

Picture paths as pipes through which association flows. Association is symmetric, so it flows regardless of the direction of the arrows.


Definition 10 (Collider) A variable on a path is a collider on that path if two arrowheads on the path collide at it, as \(L\) does on \(A \to L \leftarrow Y\).

Example 11 (Being a Collider Depends on the Path) In Figure 1, \(Y\) is a collider on the path \(A \to Y \leftarrow L\) between \(A\) and \(L\). On the path \(L \to A \to Y\) between \(L\) and \(Y\), by contrast, \(A\) is not a collider: one arrow on the path enters it and the other leaves it.


2.2 The Three Examples

Example 12 (Three Marginal Associations)  

  • Figure 2: causation implies association; in an ideal randomized experiment \(\Pr[Y^{a=1}=1] \neq \Pr[Y^{a=0}=1]\) if and only if \(\Pr[Y=1 \mid A=1] \neq \Pr[Y=1 \mid A=0]\).
  • Figure 3: learning that Hera carries a lighter makes it likelier she smokes, hence likelier she gets lung cancer; association flows through the common cause \(L\), though \(A\) has no effect on \(Y\).
  • Figure 4: learning that Apollo lacks the haplotype says nothing about his smoking; the collider \(L\) blocks the flow of association, so \(A \perp\!\!\!\perp Y\).

Remark 2 (Marginal Association in a Causal DAG). Two variables on a causal DAG are generically (for all but special parameter values) marginally associated if one causes the other or if they share a common cause. Otherwise, under the causal Markov assumption (Definition 5), they are marginally independent.

In Figure 3, an investigator who concludes from the association that carrying a lighter causes lung cancer makes a mistake: information about \(A\) improves prediction of \(Y\) without \(A\) affecting \(Y\) (Hernán and Robins 2020, 79). In Figure 4, that both \(A\) and \(Y\) cause heart disease \(L\) is irrelevant to the marginal association between \(A\) and \(Y\).

3 6.3 Causal Diagrams and Conditional Independence (pp. 80-81)


NoteNotation: A Box Means Conditioning

A square box around a node means that the analysis conditions on that variable, for example by restricting to one of its levels.

Figure 5: Book Figure 6.5: aspirin \(A\) lowers platelet aggregation \(B\), which affects heart disease \(Y\) (the square node \(B\) marks conditioning).

Example 13 (Conditioning in Three Structures)  

Structure Conditioning on Path \(A\), \(Y\) given the conditioning variable
\(A \to B \to Y\) (Fig. 6.5) mediator \(B\) blocked independent: \(A \perp\!\!\!\perp Y \mid B\)
\(A \leftarrow L \to Y\) (Fig. 6.6) common cause \(L\) blocked independent: \(A \perp\!\!\!\perp Y \mid L\)
\(A \to L \leftarrow Y\) (Fig. 6.7) collider \(L\) opened generically associated

The two independences follow from the causal Markov assumption (Definition 5). The association in the collider row holds generically (for all but special parameter values), in at least one level of \(L\).

  • Mediator (Hernán and Robins 2020, 80): among those with low platelet aggregation (\(B=0\)), knowing whether someone took aspirin adds no information about heart disease, because aspirin acts only through \(B\): \(\Pr[Y=1 \mid A=1, B=b] = \Pr[Y=1 \mid A=0, B=b]\) for all \(b\).
  • Common cause: among nonsmokers (\(L=0\)), carrying a lighter no longer predicts lung cancer. Blocking the flow of association through common causes is the graph-based justification for stratification as a way to achieve exchangeability.
  • Because complete diagrams (all possible arrows present) imply no conditional independences, it is often said that the information about associations is in the missing arrows.

3.1 Conditioning on a Collider Opens the Path

Example 14 (Haplotype and Smoking Among People with Heart Disease) In Figure 4, restrict to people with heart disease (\(L=1\), book Figure 6.7). Intuitively, among them someone without the haplotype is more likely to have the other cause, smoking, which would make \(A\) and \(Y\) inversely associated given \(L=1\). In the extreme case where \(L = 1\) exactly when \(A = 1\) or \(Y = 1\), among those with \(L=1\) the absence of one cause guarantees the presence of the other.

WarningThe Diagram Does Not Fix the Sign

The diagram alone implies only that, generically (for all but special parameter values), \(A\) and \(Y\) are associated in at least one level of \(L\). The sign of that association, and whether it appears in the level \(L = 1\) in particular, depend on how \(A\) and \(Y\) act together on \(L\) (Chapter 8).

Intuition from the book (Hernán and Robins 2020, 81): whether two causes are associated cannot depend on a future event (their common effect), but two causes of an effect generally become associated once we stratify on that effect. Chapter 8 develops associations due to conditioning on common effects.


3.2 Conditioning on a Descendant of a Collider

Figure 6: Book Figure 6.8: diuretic use \(C\) follows a diagnosis of heart disease \(L\) (the square node \(C\) marks conditioning).

Example 15 (Conditioning on Diuretic Use) In Figure 6, diuretic use \(C\) is affected by the collider heart disease \(L\). Conditioning on \(C\) also opens the path \(A \to L \leftarrow Y\), so haplotype and smoking are generically (for all but special parameter values) associated in at least one level of \(C\). The path stays blocked when we condition on neither \(L\) nor \(C\).


3.3 Three Structural Sources of Association

Remark 3 (Three Structural Sources of Association). Two variables can be associated because

  1. one causes the other;
  2. they share a common cause;
  3. they share a common effect and the analysis is restricted to a level of that effect (or of its descendants).

A fourth, nonstructural source is chance (random variability), which shrinks as the study population grows. Until Chapter 10 the book assumes a very large population, so all associations discussed are structural.

The heuristic arguments of this section are formalized by d-separation (Pearl 1995); Fine Point 6.1 lists the rules and Fine Point 6.2 introduces faithfulness (Hernán and Robins 2020, 81).

4 6.4 Positivity and Consistency in Causal Diagrams (pp. 81-85)


NoteFine Point 6.1: D-Separation

The graphical rules of Sections 6.2 and 6.3, for a path (Definition 9) between two variables, are:

  1. with nothing conditioned on, a path is blocked exactly when it contains a collider (Definition 10);
  2. a path that contains a conditioned-on non-collider is blocked;
  3. a conditioned-on collider does not block a path;
  4. a collider with a conditioned-on descendant does not block a path.

The two definitions below summarize these rules (Hernán and Robins 2020, Fine Point 6.1, p. 82).

Definition 11 (Blocked and Open Paths) Given a set \(\mathcal{C}\) of conditioned-on variables, a path is blocked if and only if

  • it contains a non-collider that is in \(\mathcal{C}\), or
  • it contains a collider that is not in \(\mathcal{C}\) and has no descendant in \(\mathcal{C}\).

Otherwise the path is open.

Definition 12 (D-Separation) Two variables are d-separated given a set \(\mathcal{C}\) of other variables (\(\mathcal{C}\) may be empty) if every path between them is blocked given \(\mathcal{C}\) (Definition ;11); otherwise they are d-connected given \(\mathcal{C}\). Two disjoint sets of variables are d-separated given \(\mathcal{C}\) if each variable in one is d-separated from every variable in the other.

Example 16 (D-Separation in Book Figures 6.1, 6.4, and 6.8)  

  • Figure 1, \(A\) and \(L\), given nothing: the path \(L \to A\) is open, so they are d-connected, even though the other path \(A \to Y \leftarrow L\) is blocked by the collider \(Y\).
  • Figure 4, \(A\) and \(Y\), given nothing: the only path is blocked by the collider \(L\), so they are d-separated.
  • Figure 6, \(A\) and \(Y\), given \(\{C\}\): the collider \(L\) has the descendant \(C\) in the conditioning set, so the only path is open and they are d-connected.

4.1 From D-Separation to Independence

Theorem 1 (D-Separation Implies Independence) Let \(\mathcal{A}\), \(\mathcal{B}\), \(\mathcal{C}\) be disjoint sets of variables on a causal DAG, and suppose the causal Markov assumption (Definition 5) holds. If \(\mathcal{A}\) and \(\mathcal{B}\) are d-separated given \(\mathcal{C}\) (Definition 12), then \(\mathcal{A} \perp\!\!\!\perp\mathcal{B} \mid \mathcal{C}\). Pearl (1988) proved this (Hernán and Robins 2020, Fine Point 6.1, p. 82).

Example 17 (Reading an Independence off Book Figure 6.5) In Figure 5, the only path between \(A\) and \(Y\) is \(A \to B \to Y\), and it contains the conditioned-on non-collider \(B\). So \(A\) and \(Y\) are d-separated given \(\{B\}\), and Theorem 1 gives \(A \perp\!\!\!\perp Y \mid B\), the conclusion reached informally in Section 6.3.


Definition 13 (Faithfulness) The joint distribution of the variables on a causal DAG is faithful to the DAG if, for all disjoint sets \(\mathcal{A}\), \(\mathcal{B}\), \(\mathcal{C}\) of its variables (\(\mathcal{C}\) possibly empty), \(\mathcal{A} \perp\!\!\!\perp\mathcal{B} \mid \mathcal{C}\) implies that \(\mathcal{A}\) and \(\mathcal{B}\) are d-separated given \(\mathcal{C}\): the converse of Theorem 1. The faithfulness assumption is that this holds (Hernán and Robins 2020, Fine Point 6.2, p. 83).

Example 18 (Faithfulness and a Conditioned-On Collider) In Figure 4 with \(L\) conditioned on (book Figure 6.7), \(A\) and \(Y\) are d-connected given \(\{L\}\). Under faithfulness, then, \(A\) and \(Y\) are not independent given \(L\): they are associated in at least one level of \(L\).

“d-” stands for directional. An equivalent set of graphical rules, moralization, was developed by Lauritzen et al. (1990) (Hernán and Robins 2020, Fine Point 6.1, p. 82).


NoteFine Point 6.2: Faithfulness

An arrow \(A \to Y\) means \(A\) affects \(Y\) for at least one individual, yet the average causal effect and the association can still both be null, when effects in different individuals cancel exactly. Such exact cancellation is rare, so the book assumes faithfulness (Definition ;13): lack of d-separation can almost always be equated with a non-zero association.

Example 19 (Unfaithfulness from Cancelling Effects) In the book’s Table 4.1, transplant increases the risk of death in women and decreases it in men, each half of the population, and the effects cancel exactly: \(\Pr[Y^{a=1}=1] = \Pr[Y^{a=0}=1]\). Figure 2 is still the correct diagram, because \(A\) affects every individual’s \(Y\), but the expected association is absent: the distribution is not faithful to the DAG (Definition 13) (Hernán and Robins 2020, Fine Point 6.2, p. 83).

Example 20 (Matching Builds In a Cancellation) Take a matched study (book Section 4.5, Figure 6.9). The matching factor \(L\) affects treatment \(A\), and whether an individual enters the matched sample, \(S = 1\), depends on both \(A\) and \(L\). Conditioning on \(S\) leaves two open paths between \(L\) and \(A\), the arrow \(L \to A\) and \(L \to S \leftarrow A\), so d-separation does not predict independence of \(L\) and \(A\) given \(\{S\}\). But matching is designed to give the treated and the untreated in the sample the same distribution of \(L\), so among those with \(S = 1\) the two variables are independent: the two paths carry associations that cancel exactly. Faithfulness (Definition 13) rules out such exact cancellations, so matching produces an unfaithful distribution on purpose (Hernán and Robins 2020, Fine Point 6.2, p. 83).

Faithfulness may also fail when variables are linked by deterministic arrows: two variables can then be independent even though some paths between them are open.


4.2 Graphs, Do-Calculus, and Notation

Standardization and IP weighting (Chapter 2) can also be derived from causal graph theory, as part of what is sometimes called the do-calculus (reviewed in Pearl 2009). So using counterfactuals in Chapters 1-5 privileged a notation, not an approach.

Whatever the notation, causal inference via standardization or IP weighting requires exchangeability, positivity, and consistency. This section covers positivity and consistency; exchangeability is translated into graphs in Section 6.5 and Chapters 7-8.


4.3 Positivity and Consistency in the Graph

Remark 4 (Where Positivity and Consistency Sit in a Graph).

  • Positivity concerns arrows into treatment nodes. Causal graphs generally cannot encode its violations, except special cases, e.g., \(A\) a deterministic function of a pretreatment \(L\) (drawn as a bold \(L \to A\) arrow).
  • Consistency (well-defined counterfactuals) concerns arrows out of treatment nodes: \(A \to Y\) must correspond to a possibly hypothetical but relatively unambiguous intervention (or, if \(A\) is a state, to well-defined interventions that affect \(Y\) only through \(A\); Fine Point 6.4).

In the book’s diagrams, positivity is implicit unless stated and consistency is embedded in the notation.

Treatment nodes thus have a special status (Hernán and Robins 2020, 83–84). Some authors draw it explicitly with decision nodes; influence diagrams are causal diagrams augmented with decision nodes (Dawid 2000, 2002). The book omits them, because it is always explicit about the interventions on \(A\), but gives treatment nodes a distinct status in SWIGs later. The causal trees of Chapter 2 and the sufficient-cause “pies” of Chapter 5 also distinguished treatments from other variables.


4.4 Treatment Nodes Need Well-Defined Interventions

WarningNot Every Node Can Be a Treatment

Diagrams read as NPSEMs with independent errors (Definition 7) seem to give all variables equal status, which can mislead when some nodes correspond to ill-defined interventions.

  • A node for “obesity” may be acceptable as an outcome \(Y\) or a covariate \(L\).
  • It is generally not acceptable as a treatment \(A\) (Chapter 3).

Pearl (2018, 2019) has proposed a concept of causation based on variables that “listen to others”, which still assumes well-defined counterfactuals for every variable (Hernán and Robins 2020, 84, margin note).


Figure 7: Book Figure 6.10, as described in the text: caloric intake \(Z\), exercise \(L\), and genetic traits \(U\) affect weight loss \(A\); \(L\) and \(U\) also affect mortality \(Y\) through other pathways.

Example 21 (Many Ways to Lose Weight) In Figure 7 there are several ways to intervene on weight loss \(A\), and some of them, through exercise \(L\) or genetic traits \(U\), also affect mortality \(Y\) directly. Losing weight by caloric restriction, by exercise, or by genetic manipulation would then lead to different mortality, so “the effect of \(A\) on \(Y\)” is unclear.

TipName the Intervention

Being explicit about the intervention is a step toward a well-defined effect, relevant data, and the right adjustment variables.


Definition 14 (Causal Discovery) Causal discovery is learning parts of the causal structure, such as which arrows are absent from the causal DAG, by analyzing data rather than from expert knowledge (Hernán and Robins 2020, Fine Point 6.3, p. 85).

Example 22 (Learning \(Z \to A \to Y\)) Suppose \(Z\) precedes \(A\), which precedes \(Y\), and assume faithfulness (Definition 13). With infinite data, suppose all three variables are marginally associated and the only conditional independence is \(Z \perp\!\!\!\perp Y \mid A\). Then the only causal DAGs compatible with the data are \(Z \to A \to Y\), possibly with a common cause of \(Z\) and \(A\) in addition to, or instead of, \(Z \to A\). So there is no unmeasured common cause of \(A\) and \(Y\), and, if \(Y^a\) is well defined and consistency and positivity hold, the average causal effect is identified by \(\operatorname{E}\mathopen{}\left[Y \mid A=1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A=0\right]\mathclose{}\) (Hernán and Robins 2020, Fine Point 6.3, p. 85).

NoteFine Point 6.3: Discovery of Causal Structure

Discovery (Definition 14) sometimes works if we assume faithfulness, so that independences in the data imply missing arrows, but often it does not: a strong association between \(B\) and \(C\) fits \(B \to C\), \(C \to B\), a shared unmeasured cause, a conditioned-on common effect, and combinations of these.

Why the learned DAG in Example 22: a direct arrow \(Z \to Y\), a common cause of \(Z\) and \(Y\), or an unmeasured common cause of \(A\) and \(Y\) would each make \(Z\) and \(Y\) dependent given \(A\) (assuming faithfulness); the marginal association of \(A\) and \(Y\) then requires an arrow \(A \to Y\). If an unmeasured common cause of \(A\) and \(Y\) existed, no conditional independence would be found, and discovery could not tell whether \(A\) causes \(Y\). Approaches to discovery are described by Spirtes et al. (2000) and Peters et al. (2017); finite samples are discussed in the book’s Technical Point 10.7.

5 6.5 A Structural Classification of Bias (pp. 85-87)


Definition 15 (Systematic Bias) There is systematic bias when the data are insufficient to identify the causal effect even with an infinite sample size. Informally: any structural association between treatment and outcome that does not arise from the causal effect of treatment on outcome (Hernán and Robins 2020, 85).

Example 23 (Lighters and Lung Cancer Again) In Figure 3, carrying a lighter \(A\) has no effect on lung cancer \(Y\), yet the two are generically (for all but special parameter values) associated through smoking \(L\). That association is structural: it does not shrink as the sample grows. So even infinite data on \(A\) and \(Y\) alone cannot reveal that the effect is null: an analysis that ignores \(L\) has systematic bias.

With systematic bias no estimator can be consistent (see Chapter 1 for consistent estimators). In this chapter “bias” means systematic bias, because the sample size is assumed infinite.


Definition 16 (Unconditional and Conditional Bias) For a dichotomous treatment \(A\), a dichotomous outcome \(Y\), and a covariate \(L\):

  • there is unconditional bias when \(\Pr[Y^{a=1}=1] - \Pr[Y^{a=0}=1] \neq \Pr[Y=1 \mid A=1] - \Pr[Y=1 \mid A=0]\);
  • there is conditional bias given \(L\) when \(\Pr[Y^{a=1}=1 \mid L=l] - \Pr[Y^{a=0}=1 \mid L=l] \neq \Pr[Y=1 \mid L=l, A=1] - \Pr[Y=1 \mid L=l, A=0]\) for at least one \(l\).

Under consistency and positivity, unconditional exchangeability \(Y^a \perp\!\!\!\perp A\) rules out unconditional bias, and conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) rules out conditional bias. When exchangeability fails, bias is generically present. Absence of unconditional bias means the association measure in the population is a consistent estimate of the corresponding effect measure.

Example 24 (Unconditional but No Conditional Bias in Table 3.1) In the book’s Table 3.1, assume, as in Chapter 3, consistency, positivity, and conditional exchangeability given \(L\). Within each level of \(L\) the observed risks are equal in the treated and the untreated (1/4 versus 1/4 when \(L=0\), 2/3 versus 2/3 when \(L=1\)), so the conditional causal risk differences are 0 and there is no conditional bias. The causal risk difference is also 0, but the associational risk difference is \(7/13 - 3/7 \approx 0.11\): there is unconditional bias.


5.1 Bias Under the Null

Definition 17 (Bias Under the Null) There is bias under the null when treatment has no causal effect on the outcome and yet, in the population, treatment and outcome are associated. Lack of exchangeability causes bias under the null.

Example 25 (Bias Under the Null in Table 3.1) In the book’s observational study of Table 3.1 (Example ;24), the causal risk ratio was 1, whereas the associational risk ratio was 1.26 (Hernán and Robins 2020, 86).

Remark 5 (Under the Null and Under the Alternative). Any causal structure that causes bias under the null also causes bias under the alternative, that is, when treatment does have a non-null effect; the converse is false.

For example, conditioning on some variables may cause selection bias under the alternative but not under the null (Greenland 1977; Hernán 2017; see also Chapter 18) (Hernán and Robins 2020, 86, margin note).


5.2 Two Structures That Break Exchangeability

Remark 6 (Sources of Systematic Bias). Lack of exchangeability has two causal structures:

  1. common causes of treatment and outcome: what many epidemiologists call confounding (Chapter 7);
  2. conditioning on common effects: what many epidemiologists call selection bias under the null (Chapter 8).

A third source of systematic bias is measurement bias (information bias), from mismeasured treatment, outcome, or covariates; some types also cause bias under the null (Chapter 9). Random variability is a separate, nonstructural source of error (Chapter 10).

All three biases can arise in observational studies and in randomized experiments. Earlier chapters used idealized experiments (no loss to follow-up, full adherence, blinded assignment), but real experiments rarely look like that; the remaining chapters of Part I examine the boundary between experimenting and observing (Hernán and Robins 2020, 86–87).

6 6.6 The Structure of Effect Modification (pp. 87-89)


Causal diagrams are good at locating sources of bias, but less helpful for illustrating effect modification.

Figure 8: Book Figure 6.12: randomized transplant \(A\); quality of care \(V\) affects death \(Y\).

Example 26 (Transplant and Quality of Care) In Figure 8, transplant \(A\) is randomized, so association is causation. Investigators stratify by quality of care \(V\) and find effect modification by \(V\) on the additive scale.


6.1 Two Caveats

WarningWhat the Diagram Does Not Say About Effect Modification
  1. Figure 8 would be valid without \(V\), since \(V\) is not a common cause of \(A\) and \(Y\); \(V\) appears only because the question refers to it. Variables on the path from \(V\) to \(Y\), such as therapy complications \(N\) (book Figure 6.13), could also be effect modifiers.
  2. Figure 8 does not show whether, or how, \(V\) modifies the effect of \(A\). It cannot distinguish:
    • effects in the same direction in both strata of \(V\);
    • effects in opposite directions (qualitative effect modification);
    • an effect in one stratum only (e.g., \(A\) kills only those with \(V=0\)).

6.2 Surrogate Effect Modifiers

Definition 18 (Causal and Surrogate Effect Modifiers) A causal effect modifier is a variable that modifies the effect of treatment on the outcome (Chapter 4) and itself affects the outcome. A surrogate effect modifier is a variable that is associated with a causal effect modifier but does not itself affect the outcome (Hernán and Robins 2020, 88).

Example 27 (Three Surrogates for Quality of Care) In the study of Example 26, an analysis stratified by a surrogate of quality of care \(V\), but not by \(V\) itself, will generally detect effect modification by the surrogate. The association between surrogate and \(V\) can come from any structure:

Book figure Surrogate Structure linking it to \(V\)
6.14 cost of treatment \(S\) cause and effect: \(V \to S\)
6.15 passport nationality \(P\) common cause: residence \(U\) affects \(V\) and \(P\)
6.16 bottled mineral water \(W\) conditioning on a common effect: \(V\) and \(W\) affect cost \(S\), analysis restricted to \(S=0\)
WarningA Surrogate Is Not a Lever

Reading surrogate effect modification causally misleads. “Cost modifies the effect of transplant” could suggest raising prices without raising the quality of care (Hernán and Robins 2020, 88).

Causal and surrogate effect modifiers are often indistinguishable in practice, so “effect modification” covers both (Section 4.2); some prefer the neutral term “heterogeneity of causal effects”.

Why the mineral-water example works (book margin note, p. 88): a hospital can keep its total cost low while paying for bottled water only by economizing elsewhere, including on the care that reduces deaths. Among low-cost hospitals, then, bottled water tends to go with lower quality of care. VanderWeele and Robins (2007b) give a finer classification of effect modification via causal diagrams.


6.3 Interaction in Causal Diagrams

Remark 7 (Diagrams and Interaction). Causal diagrams are agnostic about interaction between two treatments \(A\) and \(E\). They can encode it if augmented with nodes for sufficient-component causes (Chapter 5), with deterministic arrows from the treatments to those nodes; the book develops such diagrams in Chapter 8.


NoteFine Point 6.4: Evidence That a State Has Well-Defined Counterfactuals

Consider drug \(Z\), systolic blood pressure \(A\) (a state, not directly manipulable), and stroke \(Y\). If (i) \(Z\) is associated with \(A\) and \(Y\), (ii) \(A\) and \(Y\) are associated, and (iii) \(Z \perp\!\!\!\perp Y \mid A\), then under faithfulness the only causal DAG has \(A \to Y\), no \(Z \to Y\), and no unmeasured common cause of \(A\) and \(Y\) or of \(Z\) and \(Y\) (the discovery argument of Example ;22).

  • If \(Y^a\) is well defined and consistency and positivity hold, \(Y^a \perp\!\!\!\perp A\) holds and \(\operatorname{E}\mathopen{}\left[Y \mid A=1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A=0\right]\mathclose{}\) is the average causal effect.
  • An investigator who doubted that \(Y^a\) was well defined might take (i)-(iii) as empirical evidence that it is.

In practice \(A\) and \(Y\) usually share unmeasured causes. Then \(A\) is a collider between \(Z\) and those causes, conditioning on it links \(Z\) to \(Y\), and (iii) fails, so the data say nothing about whether \(Y^a\) is well defined (Hernán and Robins 2020, Fine Point 6.4, p. 89).

The book gets around this with a dose-finding trial that randomizes two things: which drug an individual gets (\(Z\)), and which of two blood-pressure reductions (\(a_1\) or \(a_2\)) the dose \(D\) is titrated to reach and hold. Blood pressure \(A\) is then fixed by the randomized target, so no unmeasured cause can point into it (book Figure 6.18). If, in addition, \(Y \perp\!\!\!\perp D \mid A\) and \(\operatorname{E}\mathopen{}\left[Y \mid A=a_1\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid A=a_2\right]\mathclose{}\), then, assuming faithfulness and that \(Y^{d,a}\) and \(Y^a\) are well defined, the dose affects \(Y\) only through \(A\), and \(\operatorname{E}\mathopen{}\left[Y \mid A=a_1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A=a_2\right]\mathclose{}\) identifies \(\operatorname{E}\mathopen{}\left[Y^{a_1} - Y^{a_2}\right]\mathclose{}\), whatever unmeasured causes \(A\) and \(Y\) would otherwise share. Without titration (book Figure 6.17), \(Z\) and \(Y\) stay associated given \(A\), and a direct effect of \(Z\) on \(Y\) cannot be excluded.

7 Summary


  • A causal DAG encodes qualitative causal knowledge: an arrow means a direct effect for at least one individual; a missing arrow means none for anyone; all common causes of variables on the graph must be on the graph (needed for the causal Markov assumption).
  • Each causal DAG represents a counterfactual model (by default an FFRCISTG).
  • Association flows along open paths: marginally, causal chains and common causes transmit association, colliders block it; conditioning on a non-collider blocks a path, conditioning on a collider or its descendant opens it (d-separation).
  • Faithfulness, assumed throughout the book, lets us equate d-connection with association.
  • Graphs rarely show positivity violations; consistency requires treatment nodes with well-defined interventions.
  • Systematic bias from lack of exchangeability has two structures: common causes (confounding) and conditioning on common effects (selection bias); measurement bias is a third type.
  • Causal diagrams do not show whether or how a variable modifies an effect, and cannot separate causal from surrogate effect modifiers.

Looking ahead: Chapter 7 (confounding), Chapter 8 (selection bias), Chapter 9 (measurement bias), Chapter 10 (random variability).

The diagrams in this chapter are drawn with the R packages dagitty and ggdag.

8 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top