Chapter 23: Causal Mediation

Published

Last modified: 2026-10-09 10:17:06 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~2% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Part III asked how outcomes would change under sustained treatment strategies; it did not ask how treatment produces its effect. Causal mediation studies the causal pathways through which a treatment \(A\) affects an outcome \(Y\), in particular the pathways that run through an intermediate variable, the mediator \(M\).

Mediation can be viewed as a special case of causal inference with time-varying treatments: instead of one treatment taking values at several times, we have two different variables, the treatment and the mediator, at different times. This chapter develops a framework for mediation based on hypothetical interventions that can be mapped into a target trial. Unlike approaches built on pure direct and total indirect effects, this interventionist framework allows the causal estimates to be checked empirically, at least in principle.

This chapter is based on Hernán and Robins (2020, chap. 23, pp. 323-331).

Key message: the standard decomposition of a total effect into a pure direct effect and a total indirect effect relies on cross-world counterfactuals whose identifying assumptions no experiment can check. The interventionist alternative reframes the mediation question as a question about separately intervenable components of treatment. When the identifying assumptions hold, the resulting g-formula is the same mediation formula, but now it estimates the effect of an intervention that could be carried out, and its assumptions can be refuted by a future randomized trial.

1 23.1 Mediation Analysis Under Attack (pp. 323-325)


Example 1 (The smoking cessation trial) Smokers are randomized to quit smoking (\(A = 0\)) or to keep smoking (\(A = 1\)), with perfect adherence. The trial finds that cessation lowers the 1-year risk of myocardial infarction \(Y\), i.e., \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} > \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\).

The investigators then ask whether the benefit operates through reduced hypertension \(M\), measured at 6 months (assume no one has the outcome during the first 6 months). The book’s Figure 23.1 is the causal diagram \(A \to M \to Y\) with an additional arrow \(A \to Y\), assuming interventions on \(M\) are sufficiently well defined.

Decomposing the total effect of \(A\) on \(Y\) into a part through \(M\) (indirect) and a part not through \(M\) (direct) is a causal mediation analysis. Chapter 22 formalized direct effects (Technical Points 22.1 and 22.2) but not indirect effects.

Hernán and Robins (2020, 323) attributes the pure direct effect and the total indirect effect to Robins and Greenland (1992). Pearl (2001) called the same quantities the natural direct effect and the natural indirect effect.


Definition 1 (Pure direct effect) The pure direct effect of \(A\) on \(Y\) not through \(M\) is the average causal effect of \(A\) on \(Y\) if each individual’s mediator had been set to the value \(M^{a=0}\) it would have taken under \(a = 0\):

\[\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{}.\]

Definition 2 (Total indirect effect) The total indirect effect of \(A\) on \(Y\) through \(M\) is

\[\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}.\]

In the smoking example, \(M^{a=0}\) is known for those who actually quit but unknown for those who kept smoking. \(Y^{a=1, M^{a=0}}\) is a cross-world quantity: it is indexed by two treatment values, \(a = 1\) and \(a = 0\), that cannot occur simultaneously for the same individual in the same world. Both the pure direct effect and the total indirect effect are cross-world quantities.


Theorem 1 (The decomposition of the total effect) The pure direct effect and the total indirect effect sum to the total effect:

\[\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}.\]

Proof. Adding the two contrasts, the cross-world term cancels:

\[ \begin{aligned} &\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} \\ &\quad = \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{} && \text{(the } \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} \text{ terms cancel)} \\ &\quad = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} && \text{(consistency: } Y^{a, M^{a}} = Y^{a} \text{)}. \end{aligned} \]


1.1 Identification: the mediation formula

Figure 23.1 has no unmeasured common causes of \(M\) and \(Y\). If such common causes existed, no direct effect could be identified: neither controlled direct effects nor the pure direct and total indirect effects. Identifying the mediation effects requires identifying the cross-world mean \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\), which the mediation formula does:

\[\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} = \sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr[M = m \mid A = 0].\]

Theorem 2 (The mediation formula (Technical Point 23.1)) Under the causal diagram in Figure 23.1, assuming exchangeability and consistency for \(A\) and \(M\) and the cross-world independence \(Y^{a=1, m} \perp\!\!\!\perp M^{a=0}\),

\[\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} = \sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr[M = m \mid A = 0].\]

Proof. \[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} &= \sum_m \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}} \mid M^{a=0} = m\right]\mathclose{} \Pr[M^{a=0} = m] && \text{(law of total expectation)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y^{a=1, m} \mid M^{a=0} = m\right]\mathclose{} \Pr[M^{a=0} = m] && \text{(on } \{M^{a=0} = m\}, \ Y^{a=1, M^{a=0}} = Y^{a=1, m} \text{)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y^{a=1, m}\right]\mathclose{} \Pr[M^{a=0} = m] && \text{(cross-world independence } Y^{a=1, m} \perp\!\!\!\perp M^{a=0} \text{)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y^{a=1, m} \mid A = 1, M = m\right]\mathclose{} \Pr[M^{a=0} = m \mid A = 0] && \text{(exchangeability for } A \text{ and for } M \text{)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr[M = m \mid A = 0] && \text{(consistency)}. \end{aligned} \]

The fourth line uses randomization of \(A\) (so \(Y^{a=1, m}\) and \(M^{a=0}\) are independent of \(A\)) and the absence of unmeasured common causes of \(M\) and \(Y\) (so \(Y^{a=1, m}\) is independent of \(M\) given \(A = 1\)). The book compresses the last two lines into one step “by exchangeability and consistency” (Hernán and Robins 2020, 324).


NoteTechnical Point 23.1: Proof of the mediation formula

The proof of Theorem 2 makes three moves. It averages over the value the mediator would take without treatment; it drops the conditioning on that value by cross-world independence; and it swaps the remaining counterfactual quantities for observed ones by exchangeability and consistency. Only the middle move is controversial: the cross-world independence \(Y^{a=1, m} \perp\!\!\!\perp M^{a=0}\) holds if Figure 23.1 is read as an NPSEM-IE, but not if it is read as an FFRCISTG model (Hernán and Robins 2020, 324).


1.2 Why the mediation formula is under attack

  • The mediation formula identifies a cross-world quantity whose value cannot be confirmed by any experiment that intervenes on \(A\) and \(M\), not even in principle.
  • Its proof needs the cross-world independence \(Y^{a=1, m} \perp\!\!\!\perp M^{a=0}\). \(Y^{a=1, m}\) and \(M^{a=0}\) can never be observed together for the same individual, so no trial randomizing \(A\) and \(M\), singly or jointly, can check their independence.
  • That independence holds under an NPSEM-IE but not under an FFRCISTG model, the counterfactual model used throughout the book (Technical Point 6.2). This is one reason the book prefers the FFRCISTG: it wants methods whose results are, in principle, verifiable.
  • Skeptical policy makers may regard the pure direct effect as having no policy relevance, because it corresponds to no intervention.

Under an FFRCISTG model, the pure direct and total indirect effects are not point identified, but sharp bounds exist (Robins and Richardson 2010, cited in Hernán and Robins (2020, 325)).

In very unusual settings, a crossover trial could identify the cross-world quantity (Fine Points 2.1 and 3.2) (Hernán and Robins 2020, 324).

2 23.2 A Defense of Mediation Analysis (pp. 325-327)


The trial investigators, as advocates of the NPSEM-IE, defend the pure direct effect with a policy story (Pearl 2001 gave a similar argument):

  • Nicotine-free cigarettes will become available in a year, and policy makers want to learn their benefits as soon as possible.
  • Suppose strong experimental evidence shows
    1. nicotine affects heart disease \(Y\) only through hypertension \(M\), and
    2. the non-nicotine components of cigarettes have no effect on hypertension \(M\).
  • Then a smoker of nicotine-free cigarettes would have the hypertension status she would have had without cigarettes, so \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) is the risk of heart disease if all smokers switched to nicotine-free cigarettes, which the mediation formula computes from the existing trial.

The book observes two oddities in this defense (Hernán and Robins 2020, 325): the story is about an intervention on the nicotine content of cigarettes, which makes no reference to the mediator \(M\) at all; and assumptions (i) and (ii) concern direct effects of variables that do not even appear in Figure 23.1.


2.1 Separable components of treatment

The story decomposes treatment \(A\) into two separable components:

  • \(N\): nicotine exposure, which affects \(M\) but not \(Y\) directly;
  • \(O\): exposure to the other, non-nicotine components, which affects \(Y\) but not \(M\).

Each component can, in principle, be intervened on separately. For example, \(Y^{n=0, o=1}\) is the outcome under an intervention that removes only the nicotine from cigarettes.

The book’s Figure 23.2 is the FFRCISTG causal DAG for this story: bold (deterministic) arrows \(A \to N\) and \(A \to O\), and arrows \(N \to M\), \(M \to Y\), and \(O \to Y\). The arrows from \(A\) are deterministic because in the trial either \(A = N = O = 1\) (kept smoking regular cigarettes) or \(A = N = O = 0\) (quit).

  • No arrow \(N \to Y\) encodes assumption (i).
  • No arrow \(O \to M\) encodes assumption (ii).

Figure 23.3 is the corresponding SWIG. If \(O\) causes \(M\) for no individual, we may write \(M^{n, o}\) as \(M^{n}\).

Assumptions (i) and (ii) are “no controlled direct effect” assumptions. For example, (i) says that \(N\) has no direct effect on \(Y\) when the value of \(M\) is set (Hernán and Robins 2020, 326).


Theorem 3 (When the mediation formula is the g-formula (Technical Point 23.2)) Under the FFRCISTG represented by Figure 23.2, with assumptions (i) and (ii),

\[\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} = \sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr(M = m \mid A = 0),\]

which is the mediation formula.

Proof. Exchangeability holds for \(N\) and \(O\) in Figure 23.3 (the SWIG of Figure 23.2), so if \(N\) and \(O\) were observed, the g-formula would give

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} &= \sum_m \operatorname{E}\mathopen{}\left[Y \mid N = 0, O = 1, M = m\right]\mathclose{} \Pr(M = m \mid N = 0, O = 1) && \text{(g-formula)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y \mid O = 1, M = m\right]\mathclose{} \Pr(M = m \mid N = 0) && \text{(} N \text{ is not a parent of } Y\text{; } O \text{ is not a parent of } M \text{)} \\ &= \sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr(M = m \mid A = 0) && \text{(in the data, } O = 1 \iff A = 1 \text{ and } N = 0 \iff A = 0 \text{)}. \end{aligned} \]

No one in the trial has \((N = 0, O = 1)\), so positivity fails. But positivity is sufficient, not necessary: given exchangeability and consistency, identification by the g-formula only requires that the g-formula be a function of the observed data distribution, which the last line shows it is.


NoteTechnical Point 23.2: When the mediation formula is the g-formula

The separable components \(N\) and \(O\) are exchangeable (Figure 23.3), so the g-formula would identify \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\), except that nobody in the trial has \((N = 0, O = 1)\), so positivity fails. Positivity, however, is sufficient rather than necessary: given exchangeability and consistency, the g-formula identifies the mean whenever it can be written as a function of the observed data. The deterministic arrows \(A \to N\) and \(A \to O\), together with assumptions (i) and (ii) of no direct effect, make that function exactly the mediation formula (Theorem 3).

The book calls this derivation “somewhat heuristic” because determinism between \(A\), \(N\), and \(O\) creates null sets; Robins et al. (2022) give a rigorous proof using the SWIG Markov property of Technical Point 21.12. In the same way, under the expanded diagram of Figure 23.7 the g-formula equals the front door formula (Technical Point 23.3) (Hernán and Robins 2020, 326).


2.2 Lessons from the separable-components story

  • For an NPSEM-IE advocate, \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) was already identified; the story only shows that this parameter, and hence the pure direct and total indirect effects, encodes \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\), a parameter of public health interest.
  • For an FFRCISTG user, the story gives \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) an interventional interpretation as \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\), and makes both identifiable by the g-formula (the mediation formula).
  • If absent arrows mean absent direct effects for every individual, then \(Y^{a=1, M^{a=0}} = Y^{n=0, o=1}\) for every individual.

Proof (Why \(Y^{n=0, o=1} = Y^{a=1, M^{a=0}}\) for every individual). \[ \begin{aligned} Y^{n=0, o=1} &= Y^{o=1, M^{n=0}} && \text{(no } N \to Y \text{: } N \text{ affects } Y \text{ only through } M \text{; no } O \to M \text{)} \\ &= Y^{a=1, M^{n=0}} && \text{(with } N \text{ irrelevant for } Y \text{, setting } o = 1 \text{ acts on } Y \text{ like setting } a = 1 \text{)} \\ &= Y^{a=1, M^{a=0}} && \text{(with } O \text{ irrelevant for } M \text{, setting } n = 0 \text{ acts on } M \text{ like setting } a = 0 \text{)}. \end{aligned} \]


2.3 Separable effects

Definition 3 (Separable effect of the nicotine component) The separable direct effect of the component \(N\) on \(Y\) is

\[\operatorname{E}\mathopen{}\left[Y^{n=1, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{},\]

a controlled direct effect for the separable component \(N\).

It equals the total indirect effect:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{n=1, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} && \text{(} n = o = 1 \text{ is the same intervention as } a = 1 \text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} && \text{(} Y^{n=0, o=1} = Y^{a=1, M^{a=0}} \text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} && \text{(consistency: } Y^{a=1} = Y^{a=1, M^{a=1}} \text{)}. \end{aligned} \]

Likewise the separable effect of \(O\) with nicotine removed equals the pure direct effect:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=0}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=0}\right]\mathclose{} && \text{(} Y^{n=0, o=1} = Y^{a=1, M^{a=0}} \text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} && \text{(} n = o = 0 \text{ is the same intervention as } a = 0 \text{)}. \end{aligned} \]

A note on p. 327. The book (Hernán and Robins 2020, 327) writes the pure direct effect \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) as “analogously” equal to \(\operatorname{E}\mathopen{}\left[Y^{n=1, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=1, o=0}\right]\mathclose{}\). The same argument as in the proof above gives \[ \begin{aligned} Y^{n=1, o=0} &= Y^{o=0, M^{n=1}} && \text{(no } N \to Y \text{ arrow, and no } O \to M \text{ arrow)} \\ &= Y^{a=0, M^{a=1}} && \text{(} o = 0 \text{ acts on } Y \text{ like } a = 0 \text{, and } n = 1 \text{ acts on } M \text{ like } a = 1 \text{)}. \end{aligned} \] So that contrast equals \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=1}}\right]\mathclose{}\) (the effect of \(O\) with nicotine present), which differs from the pure direct effect in general. For a counterexample, let \(M^n = n\) and \(Y = O \cdot M\). Then the book’s contrast is \(\operatorname{E}\mathopen{}\left[Y^{n=1, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=1, o=0}\right]\mathclose{} = 1 - 0 = 1\), but the pure direct effect is \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = 0 - 0 = 0\). Also, the book’s pairing does not sum with the total indirect effect to the total effect, whereas the pairing in these notes does: \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=0}\right]\mathclose{}\) is the pure direct effect, and it adds to the total indirect effect \(\operatorname{E}\mathopen{}\left[Y^{n=1, o=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) to give \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\).

3 23.3 Empirically Verifiable Mediation (pp. 327-329)


The interventional reading of \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) as \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) is valid only if the separable-components story is correct and Figure 23.3 represents an FFRCISTG. Its advantage is that the story can be refuted by a randomized trial.

Example 2 (A three-arm trial of nicotine-free cigarettes) Once nicotine-free cigarettes exist, randomize smokers to:

  1. smoking cessation (\(A = N = O = 0\));
  2. continued smoking of standard cigarettes (\(A = N = O = 1\));
  3. continued smoking of nicotine-free cigarettes (\(N = 0\), \(O = 1\)).

Without temporal trends, arms 1 and 2 should reproduce the original trial’s arm means (assume samples large enough to ignore sampling variability). By randomization, \(\operatorname{E}\mathopen{}\left[Y \mid N = 0, O = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) in the new trial. If the story is correct, \(\operatorname{E}\mathopen{}\left[Y \mid N = 0, O = 1\right]\mathclose{}\) equals the mediation formula from the original trial.


3.1 If the prediction fails

If \(\operatorname{E}\mathopen{}\left[Y \mid N = 0, O = 1\right]\mathclose{}\) differs from the mediation formula, at least one assumption is false:

    1. no direct effect of nicotine on \(Y\);
    1. no direct effect of the non-nicotine components on \(M\);
    1. no unmeasured common cause \(U\) of \(M\) and \(Y\) in Figure 23.2.

The new trial’s data help locate the failure:

  • If \(\operatorname{E}\mathopen{}\left[M \mid N = 0, O = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[M \mid N = 0, O = 0\right]\mathclose{} \neq 0\), then \(O\) and \(M\) are associated given \(N = 0\),
    1. is refuted, and an arrow \(O \to M\) must be added.
  • If \(\operatorname{E}\mathopen{}\left[Y \mid N = 1, O = 1, M = m\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid N = 0, O = 1, M = m\right]\mathclose{} \neq 0\) for some \(m\), then either (i) is false (add \(N \to Y\)) or an unmeasured common cause \(U\) of \(M\) and \(Y\) exists (add \(U\)), or both.
NoteFine Point 23.1: Empirical falsification of the assumptions for separable effects

Telling (i) apart from (iii) needs a further trial, say with 8 arms, that also intervenes on \(M\):

  • if \(N\) affects \(Y\) only through \(M\), \(N\) is independent of \(Y\) given \(M\) and \(O\) in the eight-arm trial;
  • if \(M\) and \(Y\) share no unmeasured common cause, the distribution of \(Y\) given \(M\) within levels of \(N\) and \(O\) is the same in the three-arm and the eight-arm trials.

How can Figure 23.2 omit a common cause of \(M\) and \(Y\) when Figure 23.1 is an FFRCISTG with no such common cause? Common causes may be active only under interventions that give \(N\) and \(O\) different values; the FFRCISTG for Figure 23.1 only considers \(A = N = O = 1\) and \(A = N = O = 0\). If such a \(U\) is the only reason for the discrepancy, \(Y^{n=0, o=1} = Y^{a=1, M^{a=0}}\) still holds for every individual, but \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) is no longer identified by the mediation formula. Robins et al. (2022) describe a realistic hypothetical study of treatment for river blindness with exactly this structure (Hernán and Robins 2020, 328).


3.2 How different analysts react

Suppose the new trial reproduces the original arms but its nicotine-free arm differs from the mediation formula.

  • Everyone agrees the separable-components story was wrong, and that \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) is estimated by the third arm, not by the mediation formula.
  • NPSEM-IE advocates may still believe \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) equals the mediation formula, but can no longer justify the pure direct effect as the effect of switching to nicotine-free cigarettes.
  • FFRCISTG users may lose interest in the mediation formula, and instead value what the three-arm trial taught them about nicotine-free cigarettes and about which of assumptions (i)-(iii) failed.

3.3 Can treatment always be decomposed?

Keep assuming Figure 23.1 is an FFRCISTG and that the nicotine-free arm differs from the mediation formula. Could \(A\) always be split into other, possibly unknown, components \(N'\) and \(O'\) for which Figure 23.2 is an FFRCISTG and (i) and (ii) hold, so that \(\operatorname{E}\mathopen{}\left[Y^{n'=0, o'=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) = the mediation formula?

No. Otherwise \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) would always be point identified by the mediation formula, which it is not; this would contradict the sharp bounds of Robins and Richardson (2010) (Hernán and Robins 2020, 329).

4 23.4 An Interventionist Theory of Mediation (pp. 329-331)


The chapter’s interventionist theory of mediation reframes the mediation question as a question about the effects of interventions on substantively meaningful, separable components (\(N\) and \(O\)) of \(A\). It can stand on its own, without any reference to cross-world (nested) counterfactuals.

If \(N\) and \(O\) are separable components of \(A\), then in a future six-arm trial with arms

  1. \(a = 1\);
  2. \(a = 0\);
  3. \(n = 1, o = 1\);
  4. \(n = 0, o = 0\);
  5. \(n = 0, o = 1\);
  6. \(n = 1, o = 0\),

two statements hold:

  • the mean of \(Y\) in arm 1 equals that in arm 3, and the mean in arm 2 equals that in arm 4;
  • \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) is the mean outcome in arm 5 and \(\operatorname{E}\mathopen{}\left[Y^{n=1, o=0}\right]\mathclose{}\) the mean outcome in arm 6, whether or not \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) is identified from the currently observed data.

The theory was presented by Robins and Richardson (2010) and extended by Robins, Richardson, and Shpitser (2022). Extensions cover survival analysis (Didelez 2019; Aalen et al. 2020), competing risks (Stensrud et al. 2020, 2021), and interference (Shpitser et al. 2021). Stensrud et al. (2021) stressed the equalities of arm means across the six-arm trial (Hernán and Robins 2020, 329).


4.1 Identification from current data

With data on \(A\), \(M\), and \(Y\) only, the separable effects of \(N\) and \(O\) are identified if

  • there is no unmeasured common cause of \(M\) and \(Y\), and
  • \(O\) has no direct effect on \(M\) and \(N\) has no direct effect on \(Y\).

In particular, \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\) then equals the mediation formula (Theorem 3).


4.2 When interventions on the mediator are not well defined

Often interventions on the putative mediator \(M\) are not well defined, so counterfactuals like \(Y^{a, m}\) are not meaningful. Then neither the pure direct effect nor the controlled direct effects based on \(M\) (Technical Point 22.1) exist. The interventionist effects still exist, as long as meaningful separable components \(N\) and \(O\) can be intervened on.

NoteFine Point 23.2: Separable effects with a surrogate mediator

If interventions on \(M\) are not well defined, the arrow \(M \to Y\) in Figure 23.1 is not causal (Section 9.5): \(M\) is a surrogate for an unknown true mediator \(H\) (Figure 23.4: \(A \to H\), \(H \to M\), \(H \to Y\), \(A \to Y\)). Adding separable components \(N\) and \(O\) gives Figure 23.5, in which, unlike Figure 23.2, \(N\) and \(Y\) are not d-separated given \(M\) and \(O\).

Suppose the three-arm trial nevertheless shows \(N \perp\!\!\!\perp Y \mid M, O\). Four explanations are possible:

  1. Figure 23.1, not 23.4, is the true diagram, and \(M\) is the true mediator;
  2. \(M\) is a one-to-one deterministic function of \(H\);
  3. a non-deterministic faithfulness violation in Figure 23.4;
  4. the predicted dependence exists but the trial was too small to detect it.

Causal-discovery advocates would tend to choose (a) if the trial was large (Technical Point 10.7) (Hernán and Robins 2020, 330).


4.3 Summary of the interventionist approach

  • Treatment \(A\) is hypothesized to decompose into multiple substantively meaningful separable components, each contributing to the overall effect and each, in principle, independently intervenable.
  • The assumptions that identify the separable effects are, in principle, verifiable in future randomized trials that intervene on the components.
  • Well-defined interventions on the purported mediator need not exist.
  • Specific separable components improve communication with subject-matter experts.
  • When the identifying assumptions hold, the identifying g-formula equals the mediation formula, but the effects it identifies refer to interventions on the separable components.

The framework accommodates more than two components, including components that vary over time.


NoteTechnical Point 23.3: Path-specific effects and the front door formula

Take the front-door setting of Fine Point 9.5 (BMI \(L\), drug \(A\), outcome \(Y\), unmeasured \(H\) confounding \(L\) and \(Y\)), modified to include an edge \(L \to Y\) (Figure 23.6). The total effect of \(L\) on \(Y\) is then not identified, because the \(L \to Y\) path cannot be separated from confounding by \(H\). Is the effect of \(L\) along \(L \to A \to Y\) identified?

Expand the graph (Figure 23.7): \(N\) is the BMI reported to the physician who prescribes \(A\), and \(O\) is the BMI used for referral to physical therapy and diet counseling; in the data \(L = N = O\). Figure 23.8 is the SWIG for an (unethical) intervention that reports a value \(n\) to the physician while the true \(L = O\) drives referral. Because \(L \equiv O\) even in the intervened world, \(O\) can be dropped. Since \(Y^{n} \perp\!\!\!\perp N \mid L\) in Figure 23.8, the g-formula gives

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{n}\right]\mathclose{} &= \sum_{l, a} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] \Pr[A = a \mid N = n] && \text{(g-formula)} \\ &= \sum_{a} \mathopen{}\left\{\sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l]\right\}\mathclose{} \Pr[A = a \mid N = n] && \text{(regroup the sum)} \\ &= \sum_{a} \mathopen{}\left\{\sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l]\right\}\mathclose{} \Pr[A = a \mid L = n] && \text{(} L \equiv N \text{ in the data)}, \end{aligned} \]

which is the front door formula. This derivation is heuristic because of the null sets created by determinism; a rigorous proof combines determinism with the approach of Technical Point 21.12. \(N\) and \(O\) need not be separable components of \(L\); what matters is that the substantive story implies an expanded graph with an intervention variable \(N\) deterministically related to \(L\) in the actual world (Stensrud et al. 2023; Wen et al. 2023 obtained essentially equivalent results). Fulcher et al. (2020) had shown that the front door formula identifies the cross-world quantity \(\operatorname{E}\mathopen{}\left[Y^{L, A^{l=n}}\right]\mathclose{}\) under the NPSEM-IE for Figure 23.6 (Hernán and Robins 2020, 331).


4.4 Practical caveats

  • The chapter treats theory, not the practice of estimating pure direct effects and separable direct effects.
  • Without trials that actually intervene on the mediator (controlled direct effects) or on treatment components (separable direct effects), these analyses rest on observational data and exchangeability: valid mediation analyses must adjust for confounders of both the treatment-outcome and the mediator-outcome relationships.
  • The chapter’s simple diagram is a teaching device, not a realistic study. In practice, causal mediation analyses are observational and rely on more heroic assumptions than non-mediation analyses.

For separable effects, further issues arise when a measured common cause of \(M\) and \(Y\) is a child of \(N\) and \(O\) (Robins and Richardson 2010; Robins et al. 2022) (Hernán and Robins 2020, 331).

5 Summary


  • The pure direct effect \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{}\) and the total indirect effect \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) sum to the total effect, but both are cross-world quantities.
  • The mediation formula \(\sum_m \operatorname{E}\mathopen{}\left[Y \mid A = 1, M = m\right]\mathclose{} \Pr[M = m \mid A = 0]\) identifies \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{}\) only under a cross-world independence that holds under an NPSEM-IE, not an FFRCISTG, and that no experiment can check.
  • Decomposing \(A\) into separable components \(N\) and \(O\) gives the mediation formula an interventional meaning, \(\operatorname{E}\mathopen{}\left[Y^{n=0, o=1}\right]\mathclose{}\), identified by the g-formula under no-direct-effect assumptions and no unmeasured \(M\)-\(Y\) confounding.
  • Those assumptions can be refuted by a future trial that intervenes on the components.
  • The interventionist theory does not require well-defined interventions on \(M\).

6 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top