Chapter 19: Time-Varying Treatments

Published

Last modified: 2026-10-09 13:46:40 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Parts I and II of the book considered time-fixed treatments, whose value is determined once, at the start of follow-up. Many causal questions instead involve treatments whose value can change over time for the same individual: medical treatments, lifestyle habits, employment or marital status, occupational exposures. Part III extends the framework to these time-varying treatments, and this chapter introduces the terminology and concepts it needs.

This chapter is based on Hernán and Robins (2020, chap. 19, pp. 255-266).

The authors warn that this is one of the most technical chapters of the book, because simplifying the concepts further would cost too much rigor. The chapter defines causal effects for time-varying treatments, introduces treatment strategies and sequentially randomized experiments, states the identifiability conditions (sequential exchangeability, positivity, consistency), and uses SWIGs to show that some strategies can be identified when others cannot. How to estimate these effects is the subject of Chapters 20 and 21.

1 19.1 The Causal Effect of Time-Varying Treatments (pp. 255-256)


For a time-fixed treatment \(A\) (1: treated, 0: untreated) given at time zero and an outcome \(Y\) measured 60 months later, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\). It need not mention when treatment occurs, because everybody’s treatment is determined at the same time.

A time-varying treatment needs time to be made explicit.

Definition 1 (Time-indexed treatment) Let \(K\) be the last time at which treatment can be given. For each time \(k = 0, 1, \ldots, K\), \(A_k\) denotes the value of a dichotomous treatment at time \(k\).

Example 1 (HIV and antiretroviral therapy) In a 5-year follow-up study of individuals infected with HIV, time \(k\) is counted in months, so \(K = 59\). \(A_k = 1\) if the individual receives antiretroviral therapy in month \(k\) and \(A_k = 0\) otherwise. Nobody was treated before the study started: \(A_{-1} = 0\) for all individuals. The outcome \(Y\) measures health status (higher is better) at the end of follow-up, time \(K + 1 = 60\).

Remark 1 (Simplifying assumptions). For simplicity the book provisionally assumes no losses to follow-up, no deaths, and perfectly measured variables. Time is indexed from zero (the first time of possible treatment is \(k = 0\)) for compatibility with much of the published literature. Although the example has an outcome measured at a fixed time, the concepts carry over to time-varying and failure-time outcomes (see Technical Point 21.8 of the book).


Definition 2 (Treatment history) An overbar denotes history: \(\bar{A}_k = (A_0, A_1, \ldots, A_k)\) is the treatment history from time 0 through time \(k\). The whole history through \(K\), \(\bar{A}_K\), is often written \(\bar{A}\) without a subscript. Lower-case letters denote realizations: \(a_k\) is a value of \(A_k\).

Example 2 (Treatment histories)  

  • An individual treated in every month has \(\bar{A} = (1, 1, \ldots, 1) = \bar{1}\).
  • An individual never treated has \(\bar{A} = (0, 0, \ldots, 0) = \bar{0}\).
  • Most individuals are treated during only part of follow-up, so their histories mix 1s and 0s and have no compact symbol.

Remark 2 (The effect of a time-varying treatment is not unique). A contrast at a single time,

\[\operatorname{E}\mathopen{}\left[Y^{a_k = 1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_k = 0}\right]\mathclose{},\]

quantifies the effect of treatment at time \(k\) only, not the effect of treatment at all times between 0 and \(K\). The average causal effect of a time-varying treatment must instead contrast mean counterfactual outcomes under two rules, each of which specifies treatment at every time from \(k = 0\) to \(k = K\). Because there are many such rules, the average causal effect of a time-varying treatment is not uniquely defined.

2 19.2 Treatment Strategies (pp. 256-257)


Definition 3 (Treatment strategy) A treatment strategy (also called a plan, policy, protocol, or regime) is a rule to assign treatment at each time \(k\) of follow-up.

The rules contrasted in Remark 2 are treatment strategies.

Example 3 (Always treat versus never treat) The strategies “always treat”, \(\bar{a} = \bar{1}\), and “never treat”, \(\bar{a} = \bar{0}\), define one average causal effect of \(\bar{A}\) on \(Y\):

\[\operatorname{E}\mathopen{}\left[Y^{\bar{a} = \bar{1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\bar{a} = \bar{0}}\right]\mathclose{}.\]

A general counterfactual theory for comparing treatment strategies was first articulated by Robins (1986, 1987, 1997a).


Many other contrasts are possible.

Example 4 (Other contrasts of fixed treatment sequences) For example, \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\bar{a}'}\right]\mathclose{}\) could compare

  • “treat every other month”, \(\bar{a} = (1, 0, 1, 0, \ldots)\), with
  • “treat in all months except the first”, \(\bar{a}' = (0, 1, 1, 1, \ldots)\).

Proposition 1 (Number of fixed treatment sequences) With a dichotomous treatment \(a_k \in \{0, 1\}\) at each time \(k = 0, 1, \ldots, K\), there are \(2^{K+1}\) fixed treatment sequences \(\bar{a} = (a_0, a_1, \ldots, a_K) \in \{0, 1\}^{K+1}\).

Proof. The vector \((a_0, \ldots, a_K)\) has \(K + 1\) components, each taking one of two values, so there are \(2^{K+1}\) such vectors.

Fixed treatment sequences do not exhaust all strategies.

The book writes “at least \(2^K\)”, a lower bound, and uses \(2^K\) again when counting fixed sequences in Technical Point 19.3; these notes write \(2^{K+1}\) throughout. The point is only that the number grows exponentially with the length of follow-up.


2.1 Dynamic Strategies

Example 5 (Starting treatment when the CD4 count drops) Let \(L_k\) be the CD4 cell count (cells per microliter) measured at month \(k\), coded \(L_k = 1\) when low (bad prognosis) and \(L_k = 0\) otherwise; everybody starts with a high count, \(L_0 = 0\). Consider the strategy

do not treat while \(L_k = 0\); start treatment when \(L_k = 1\) and treat continuously after that.

This strategy cannot be written as a fixed vector \(\bar{a} = (a_0, a_1, \ldots, a_K)\) that gives every individual the same \(a_k\) at time \(k\): at each time, who is treated depends on each individual’s evolving \(L_k\).

Definition 4 (Dynamic and static strategies) A dynamic treatment strategy is a rule in which the treatment \(a_k\) at time \(k\) depends on the evolution of an individual’s time-varying covariates \(\bar{L}_k\). Strategies \(\bar{a}\) in which treatment does not depend on covariates are non-dynamic or static treatment strategies.


NoteFine Point 19.1: Deterministic and Random Treatment Strategies
  • A deterministic dynamic strategy is a rule \(g = [g_0(\bar{a}_{-1}, l_0), \ldots, g_K(\bar{a}_{K-1}, \bar{l}_K)]\), where \(g_k(\bar{a}_{k-1}, \bar{l}_k)\) is the treatment assigned at \(k\) to an individual with past history \((\bar{a}_{k-1}, \bar{l}_k)\). Example: \(g_k(\bar{a}_{k-1}, \bar{l}_k) = 1\) if the CD4 count was low at or before \(k\), and 0 otherwise.
  • A deterministic static strategy is a rule \(g = [g_0(\bar{a}_{-1}), \ldots, g_K(\bar{a}_{K-1})]\) that does not depend on \(\bar{l}_k\).
  • Random strategies assign a probability of treatment rather than a value. They can be static (“each month, independently, treat with probability 0.3”) or dynamic (“each month, independently, treat with probability 0.3 if the CD4 count is low; do not treat if it is high”).

The strategy \(g\) that maximizes \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) (when higher outcomes are better) is the optimal treatment strategy. For a drug, it will almost always be dynamic, since treatment must stop when toxicity develops. No random strategy can be preferred to the optimal deterministic one, but random strategies (randomized trials) remain scientifically necessary, because before the trial nobody knows which deterministic strategy is optimal. Unless noted otherwise, \(g\) denotes a deterministic strategy.

The book cites Young et al. (2014) for a taxonomy of treatment strategies.


2.2 Causal Effects Require Specified Strategies

Remark 3 (Causal effects are contrasts of specified strategies). The average causal effect of a time-varying treatment is well defined only once the strategies being compared are specified, e.g.

  • \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\bar{a}'}\right]\mathclose{}\): “always treat” versus “never treat”;
  • \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{g}\right]\mathclose{}\): “always treat” versus “treat only after the CD4 count is low”.

\(g\) denotes any strategy, static or dynamic; for a static strategy we sometimes write \(Y^{g = \bar{a}}\) rather than \(Y^g\) or \(Y^{\bar{a}}\). Even with only two options (treat or not) at each time, there are as many causal effects as there are pairs of strategies.

NoteTechnical Point 19.1: On the Definition of Dynamic Strategies

Every dynamic strategy \(g\) that depends on past treatment and covariates has a counterpart \(g'\) that depends only on past covariates, defined recursively by \(g'_0(l_0) = g_0(\bar{a}_{-1} = 0, l_0)\) and \(g'_k(\bar{l}_k) = g_k\mathopen{}\left(g'_{k-1}(\bar{l}_{k-1}), \bar{l}_k\right)\mathclose{}\). By consistency, an individual has the same treatment, covariate, and outcome history whether following \(g\) or \(g'\) from time zero, so \(Y^g = Y^{g'}\). In the observed data, an individual has followed \(g\) through time \(t\) if and only if they have followed \(g'\) through \(t\). If treatment under \(g\) already does not depend on past treatment, \(g\) and \(g'\) are identical.

In the recursion, \(g'_{k-1}(\bar{l}_{k-1})\) stands for the whole treatment history \(\bar{a}_{k-1}\) that \(g'\) would have produced through \(k-1\) given \(\bar{l}_{k-1}\).

3 19.3 Sequentially Randomized Experiments (pp. 257-259)


The book uses three causal diagrams with the first two times, \(k = 0\) and \(k = 1\). In each, \(A_k\) is treatment, \(L_k\) the measured variables, \(U_k\) unmeasured common causes of at least two variables, and \(Y\) the outcome. In the HIV example, CD4 count \(L_k\) is a consequence of unmeasured immune damage \(U_k\), which also lowers health status \(Y\).

Table 1: Which variables have arrows into treatment \(A_k\) in the book’s Figures 19.1-19.3.
Figure Arrows from \(\bar{A}_{k-1}\) into \(A_k\) Arrows from \(\bar{L}_k\) into \(A_k\) Arrows from \(\bar{U}_k\) into \(A_k\)
19.1 yes (for \(k \geq 1\)) no no
19.2 yes (for \(k \geq 1\)) yes no
19.3 yes (for \(k \geq 1\)) yes yes

All participants are assumed to adhere to the assigned treatment. A causal graph must include all common causes of any two variables on it, which is why the \(U_k\) appear even though they are not measured.


3.1 Figure 19.1: Treatment Depends Only on Past Treatment

Example 6 (Assignment depending only on past treatment) Example: each month, treatment is assigned with probability 0.5 to those untreated the previous month (\(A_{k-1} = 0\)) and with probability 1 to those treated the previous month (\(A_{k-1} = 1\)).

  • For static strategies (Definition 4), treatment assignment that depends only on past treatment is the time-varying generalization of no confounding by measured or unmeasured variables.
  • So under Figure 19.1 (where treatment is randomized at each time), association is causation for static strategies: the counterfactual mean under a static strategy \(\bar{a}\) is the mean outcome among those who followed \(\bar{a}\) (provided some individuals followed \(\bar{a}\), and under consistency).
WarningThe Equality Does Not Carry Over to Dynamic Strategies

The equality \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\) does not in general extend to dynamic strategies. Under a dynamic strategy \(g\) that depends on \(L\), generically, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) equals the mean outcome among those who followed \(g\) only if the probability of \(A_k = 1\) is exactly 0.5 at every time \(k\) at which treatment under \(g\) depends on \(L_k\). Otherwise, identifying \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) requires g-methods applied to data on \(L\), \(A\), and \(Y\), under either Figure 19.1 or Figure 19.2.


3.2 Figure 19.2: Treatment Also Depends on Measured Covariates

Example 7 (Assignment depending also on CD4 count) Example: each month, treatment is assigned with probability

  • 0.4 to untreated individuals with high CD4 count,
  • 0.8 to untreated individuals with low CD4 count,
  • 0.5 to previously treated individuals (\(A_{k-1} = 1\)), whatever their CD4 count.

There is confounding by measured, but not unmeasured, variables.

The book labels these strata “high CD4 (\(A_{k-1} = 0, L_k = 1\))” and “low CD4 (\(A_{k-1} = 0, L_k = 0\))”, which reverses the coding of Section 19.2 (\(L_k = 1\) means low CD4). We describe the strata by CD4 level only; under the Section 19.2 coding, high CD4 is \(L_k = 0\) and low CD4 is \(L_k = 1\).


Definition 5 (Sequentially randomized experiment) An experiment in which treatment is randomly assigned to each individual at each time \(k\), with assignment probabilities that depend only on the measured history \((\bar{A}_{k-1}, \bar{L}_k)\), is a sequentially randomized experiment.

Remark 4 (Which diagrams can be sequentially randomized).

  • Figures 19.1 and 19.2 (Table 1) can represent sequentially randomized experiments.
  • Figure 19.3 cannot: there, \(A_k\) depends partly on unmeasured \(U\), which investigators cannot use to assign treatment.
  • A sequentially randomized experiment has no arrows from unmeasured prognostic factors \(U\) into any \(A_k\), however many time points \(k = 0, 1, \ldots, K\) there are.

3.3 Observational Studies

Treatment decisions usually depend on prognostic factors, so observational studies are typically represented by Figure 19.2 or 19.3, not 19.1.

  • If clinicians’ decisions at each month depend on \((\bar{A}_{k-1}, \bar{L}_k)\) but on no unmeasured \(\bar{U}_k\), the study is represented by Figure 19.2. It then differs from a sequentially randomized experiment (Definition 5) only in that the assignment probabilities are unknown (though estimable from the data).
WarningThe Data Cannot Rule Out Unmeasured Confounding

The data cannot tell us whether Figure 19.2 or Figure 19.3 is correct. Figure 19.3 has unmeasured confounding.

Sequentially randomized experiments (Definition 5) are rarely run in practice, but the concept clarifies the conditions needed for valid estimation of the effects of time-varying treatments.

4 19.4 Sequential Exchangeability (pp. 259-260)


For a time-fixed treatment, valid inference typically requires conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), which holds in conditionally randomized experiments and in observational studies where treatment depends on measured \(L\) and, given \(L\), on no unmeasured common cause of treatment and outcome.

For a time-varying treatment we need conditional exchangeability at each time, given the covariate history \(\bar{L}_k\).


Definition 6 (Sequential exchangeability for \(Y^g\)) \[Y^g \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = g(\bar{A}_{k-2}, \bar{L}_{k-1}), \bar{L}_k \quad \text{for all strategies } g \text{ and } k = 0, 1, \ldots, K.\]

That is, at each time \(k\), the treated and untreated are exchangeable for \(Y^g\) given the covariate history \(\bar{L}_k\) through time \(k\) and any observed treatment history compatible with \(g\).

Example 8 (Sequential exchangeability with two times) With two times (\(K = 1\)), sequential exchangeability for \(Y^g\) (Definition 6) says that, for all strategies \(g\),

\[Y^g \perp\!\!\!\perp A_0 \mid L_0 \quad\text{and}\quad Y^g \perp\!\!\!\perp A_1 \mid A_0 = g(L_0), L_0, L_1.\]

The book drops the word “conditional” and says simply sequential exchangeability.

For individuals whose treatment history \([A_0 = g(L_0), A_1 = g(A_0, L_0, L_1)]\) is compatible with strategy \(g\) through the end of follow-up, consistency gives \(Y^g = Y\), which also equals the counterfactual outcome under the static strategy \((a_0, a_1) = (A_0, A_1)\).


Remark 5 (When sequential exchangeability holds).

  • A sequentially randomized experiment (Definition 5) implies sequential exchangeability for \(Y^g\) (Definition 6).
  • More generally, it holds in any causal graph that, like Figure 19.2, has no arrows from unmeasured \(U\) into treatment.
  • So under Figure 19.2, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is identified for all strategies \(g\), provided that the treatment level assigned by \(g\) has positive probability given each history compatible with \(g\) that occurs with positive probability, and under consistency.
  • Under Figure 19.3, sequential exchangeability for \(Y^g\) does not hold (generically, that is, under faithfulness, which rules out exact cancellations of effects). The book states that \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is then identified for none of the strategies (Hernán and Robins 2020, sec. 19.4).
  • Under other diagrams, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) may be identified for some strategies but not others.

“This form of sequential exchangeability” signals that there are others, some weaker and some stronger.

Whenever the book talks about identification, the identifying formula is the g-formula (Chapter 21). In rare cases, not relevant here, effects can be identified by formulas related to but different from the g-formula (e.g., the front door formula of Technical Point 7.3).


4.1 Unconditional Sequential Exchangeability

Definition 7 (Unconditional sequential exchangeability) Unconditional sequential exchangeability holds for a static strategy \(\bar{a}\) if

\[Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \qquad k = 0, 1, \ldots, K.\]

Under Figure 19.1 (no arrows from \(\bar{L}_k\) or \(\bar{U}_k\) into \(A_k\); Table 1), unconditional sequential exchangeability holds for every static strategy \(\bar{a}\).

Proposition 2 (Association is causation under unconditional sequential exchangeability) Let \(\bar{a}\) be a static strategy. Suppose unconditional sequential exchangeability (Definition 7) holds for \(\bar{a}\), positivity holds along \(\bar{a}\) (\(\Pr[A_k = a_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}] > 0\) for \(k = 0, 1, \ldots, K\)), and consistency holds (\(Y^{\bar{a}} = Y\) whenever \(\bar{A} = \bar{a}\)). Suppose also that \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{\bar{a}}\right|\mathclose{}\right]\mathclose{} < \infty\). Then association is causation: \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\).

Proof. The positivity conditions give \(\Pr[\bar{A}_k = \bar{a}_k] = \prod_{j=0}^{k} \Pr[A_j = a_j \mid \bar{A}_{j-1} = \bar{a}_{j-1}] > 0\) for every \(k\), so each conditional expectation below is well defined (for \(k = 0\), conditioning on the empty history \(\bar{A}_{-1} = \bar{a}_{-1}\) is no conditioning).

One-step identity: for each \(k = 0, 1, \ldots, K\),

\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}, A_k = a_k\right]\mathclose{} && \text{(exchangeability at } k\text{; positivity at } k\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_k = \bar{a}_k\right]\mathclose{} && \text{(since } \bar{A}_k = (\bar{A}_{k-1}, A_k)\text{)} \end{align}\]

Applying the one-step identity for \(k = 0\), then \(k = 1\), and so on through \(k = K\):

\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_0 = \bar{a}_0\right]\mathclose{} && \text{(one-step identity, } k = 0\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_1 = \bar{a}_1\right]\mathclose{} && \text{(one-step identity, } k = 1\text{)} \\ &\;\;\vdots \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_K = \bar{a}_K\right]\mathclose{} && \text{(one-step identity, } k = K\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{} && \text{(consistency: } Y^{\bar{a}} = Y \text{ when } \bar{A} = \bar{a}\text{)} \end{align}\]


4.2 A Diagram With Partial Identification

Example 9 (An unrecorded clinic visit) Figure 19.4 includes an unmeasured \(W_0\) that causes both \(A_0\) and the measured CD4 count \(L_1\), e.g. an unrecorded scheduled clinic visit at time 0; \(U_1\) is the true (unknown) CD4 count, which \(L_1\) measures with error.

  • \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{}\) is still identified for every static strategy \(\bar{a}\) (Definition 4), provided that each treatment level in \(\bar{a}\) has positive probability given each history compatible with \(\bar{a}\) that occurs with positive probability, and under consistency.
  • The book states that \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is not identified for any dynamic strategy \(g\) (Definition 4) whose assignment depends on \(L_1\) (Hernán and Robins 2020, sec. 19.4).

4.3 Positivity and Consistency

Remark 6 (Three identifiability conditions). Identification also needs sequential versions of positivity and consistency. Both are expected to hold in a sequentially randomized experiment. Under all three conditions, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is identified by methods that adjust appropriately for \((\bar{A}_{k-1}, \bar{L}_k)\): the g-formula (standardization), IP weighting, and g-estimation.

NoteTechnical Point 19.2: Positivity and Consistency for Time-Varying Treatments

Sequential positivity:

\[\text{if } f_{\bar{A}_{k-1}, \bar{L}_k}(\bar{a}_{k-1}, \bar{l}_k) \neq 0, \text{ then } f_{A_k \mid \bar{A}_{k-1}, \bar{L}_k}(a_k \mid \bar{a}_{k-1}, \bar{l}_k) > 0 \text{ for all } \bar{a}_k, \bar{l}_k.\]

It holds in a sequentially randomized experiment (Definition 5) if the randomization probabilities are never 0 or 1, whatever the past history. For a particular strategy \(g\), it need hold only for treatment histories compatible with \(g\), i.e., \(a_k = g(\bar{a}_{k-1}, \bar{l}_k)\) for each \(k\).

Sequential consistency:

\[Y^{\bar{a}} = Y^{\bar{a}^*} \text{ if } \bar{a}^* = \bar{a}; \quad Y^{\bar{a}} = Y \text{ if } \bar{A} = \bar{a}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k^{\bar{a}^*} \text{ if } \bar{a}^*_{k-1} = \bar{a}_{k-1}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k \text{ if } \bar{A}_{k-1} = \bar{a}_{k-1},\]

where \(\bar{L}_k^{\bar{a}}\) is the counterfactual covariate history through \(k\) under \(\bar{a}\). Weaker conditions suffice for identification (“if \(\bar{A} = \bar{a}\) then \(Y^{\bar{a}} = Y\)” for static strategies; “if \(A_k = g_k(\bar{A}_{k-1}, \bar{L}_k)\) at each \(k\) then \(Y^g = Y\)” for dynamic ones), but the book always accepts the stronger sequential version. If “treat in month \(k\)” and “do not treat in month \(k\)” are sufficiently well defined at all \(k\), so are all static and dynamic strategies built from them.

5 19.5 Identifiability Under Some but Not All Treatment Strategies (pp. 261-264)


Remark 7 (A generalized backdoor criterion). The backdoor criterion of Chapter 7 generalizes to time-varying treatments. For static strategies, a sufficient condition for identification (given sequential positivity and consistency) is that, for every \(k\), conditioning on \((\bar{A}_{k-1}, \bar{L}_k)\) blocks every backdoor path between \(A_k\) and \(Y\) that does not pass through a later treatment.

But this generalized criterion works on causal DAGs, which contain no counterfactual outcomes, so it does not show the connection to sequential exchangeability directly. SWIGs (Chapter 7) do, and are especially useful here.

Pearl and Robins (1995) proposed a generalized backdoor criterion for static strategies; Robins (1997a) extended it to dynamic strategies.


5.1 Two Simplified Diagrams

Example 10 (The diagrams of Figures 19.5 and 19.6) Figures 19.5 and 19.6 simplify Figures 19.2 and 19.4: they drop \(U_0\), \(L_0\), the arrow \(A_0 \to U_1\), and the arrow \(L_1 \to Y\). Write \(\text{pa}(V)\) for the set of parents of a node \(V\). The treatment nodes have these complete parent sets:

  • Figure 19.5: \(\text{pa}(A_0) = \emptyset\) and \(\text{pa}(A_1) = \{A_0, L_1\}\).
  • Figure 19.6: \(\text{pa}(A_0) = \{W_0\}\) and \(\text{pa}(A_1) = \{A_0, L_1\}\), where \(W_0\) is unmeasured.

In both diagrams, no arrow points from any \(U\) or \(W\) into \(A_1\). The covariate \(L_1\) has parents including \(A_0\) and \(U_1\) (and also \(W_0\) in Figure 19.6, so \(W_0\) is a shared cause of \(A_0\) and \(L_1\)). The outcome \(Y\) has parents including \(U_1\) and \(A_1\), but not \(L_1\).

For \(L_1\) and \(Y\) we list only some parents; the book’s figures give the complete diagrams.


5.2 Building the SWIG for a Static Strategy

Algorithm 1 (Building a SWIG for a static strategy) Given a causal DAG and a static strategy \(\bar{a} = (a_0, \ldots, a_K)\):

  1. Split each treatment node \(A_k\) into two halves. The right half is the value \(a_k\) set by the intervention; the left half is the treatment that would be observed after intervening on all earlier treatments, \(A_k^{\bar{a}_{k-1}}\) (for \(k = 0\), simply \(A_0\)).
  2. Rewire the arrows: arrows into a treatment point into its left half; arrows out of a treatment leave from its right half.
  3. Relabel every descendant of a treatment node by its counterfactual, with a superscript listing the interventions on the treatments that precede it: a covariate \(L_k\) becomes \(L_k^{\bar{a}_{k-1}}\), and the outcome \(Y\) becomes \(Y^{\bar{a}}\).

Example 11 (Figure 19.7: the SWIG for Figure 19.5) Applying alg. 1 to Figure 19.5 (Example 10) under the static strategy \((a_0, a_1)\), e.g. “always treat” \((1, 1)\) or “never treat” \((0, 0)\), gives Figure 19.7:

  • \(A_0\) splits into \(A_0\) (left half) and \(a_0\) (right half);
  • \(A_1\) splits into \(A_1^{a_0}\) (left half) and \(a_1\) (right half);
  • the covariate becomes \(L_1^{a_0}\) and the outcome becomes \(Y^{a_0, a_1}\);
  • \(U_1\), which no treatment affects, keeps its label.

\(L_1\) is written \(L_1^{a_0}\), not \(L_1^{a_0, a_1}\), because a later intervention on \(A_1\) cannot affect the earlier \(L_1\). Unlike the DAG, the SWIG contains the counterfactual outcome, so exchangeability can be checked by d-separation.


5.3 Static Sequential Exchangeability

Example 12 (Reading exchangeability off the SWIG of Figure 19.7) In Figure 19.7, d-separation gives, for every \((a_0, a_1)\),

\[Y^{a_0, a_1} \perp\!\!\!\perp A_0 \quad\text{and}\quad Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}.\]

The path \(A_1^{a_0} \leftarrow a_0 \to L_1^{a_0} \leftarrow U_1 \to Y^{a_0, a_1}\) looks open, but it is blocked: \(a_0\) is a constant in the counterfactual world, and a path through a constant intervention node is always blocked, because a constant is implicitly conditioned on.

Proposition 3 (From a SWIG independence to an observed-data independence) Fix a static strategy \((a_0, a_1)\) with \(\Pr[A_0 = a_0] > 0\). If \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}\) and consistency holds (\(L_1^{a_0} = L_1\) and \(A_1^{a_0} = A_1\) whenever \(A_0 = a_0\)), then \(Y^{a_0, a_1} \perp\!\!\!\perp A_1 \mid A_0 = a_0, L_1\).

Proof (From the SWIG statement to an observed-data statement).

  1. \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}\) implies, in particular, \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0 = a_0, L_1^{a_0}\) (restrict to those who received \(A_0 = a_0\)).
  2. Among those with \(A_0 = a_0\), consistency gives \(L_1^{a_0} = L_1\) and \(A_1^{a_0} = A_1\).
  3. Substituting: \(Y^{a_0, a_1} \perp\!\!\!\perp A_1 \mid A_0 = a_0, L_1\).

Definition 8 (Static sequential exchangeability) Static sequential exchangeability holds for a given static strategy \(\bar{a}\) if

\[Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \bar{L}_k \quad \text{for } k = 0, 1, \ldots, K.\]

Remark 8 (Static sequential exchangeability is weaker).

  • Static sequential exchangeability (Definition 8) is weaker than sequential exchangeability for \(Y^g\) (Definition 6): it concerns only counterfactual outcomes indexed by static strategies \(g = \bar{a}\).
  • When it holds for every static \(\bar{a}\), it suffices, given sequential positivity and consistency, to identify \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{}\) for every static strategy.
  • It holds in Figure 19.6 (Example 10) too (check d-separation on its SWIG, Figure 19.8), so under Figure 19.6 (and given sequential positivity and consistency) every static strategy is identified.

Example 13 (Two variations on Figure 19.5)  

  • Without the arrow \(L_1 \to A_1\), the SWIG would lack \(L_1^{a_0} \to A_1^{a_0}\), so unconditional sequential exchangeability (Definition 7) would hold and every static strategy could be identified without data on \(L_1\) (given positivity and consistency).
  • With an added arrow \(U_1 \to A_1\), the SWIG would have \(U_1 \to A_1^{a_0}\), no form of sequential exchangeability would hold (generically, that is, under faithfulness), and the book states that no strategy would be identified (Hernán and Robins 2020, sec. 19.5).

NoteTechnical Point 19.3: The Many Forms of Sequential Exchangeability

For a sequentially randomized experiment with times \(k = 0, 1, \ldots, K\) (a longer version of Figure 19.7), the SWIG gives

\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k^{\bar{a}_{k-1}} \mid \bar{A}_{k-1}^{\bar{a}_{k-2}}, \bar{L}_k^{\bar{a}_{k-1}},\]

where \(\underline{L}_{k+1}^{\bar{a}}\) is the counterfactual covariate history from \(k + 1\) to the end of follow-up. Restricting to \(\bar{A}_{k-1}^{\bar{a}_{k-2}} = \bar{a}_{k-1}\) and applying consistency gives

\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \bar{L}_k.\]

When this holds for all \(\bar{a}\), the book calls it sequential exchangeability; these notes call it joint sequential exchangeability, to keep it distinct from Definition 6, which it implies. Although it mentions only static strategies, it is equivalent to

\[\mathopen{}\left(Y^{g}, \underline{L}_{k+1}^{g}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = g(\bar{A}_{k-2}, \bar{L}_{k-1}), \bar{L}_k \quad \text{for all } g,\]

and so, with positivity and consistency, it identifies the outcome and covariate distributions under every static and dynamic strategy (Robins 1986). This rests on the joint independence of \((Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}})\) and \(A_k\); for dynamic strategies, the two separate independences of \(Y^{\bar{a}}\) and of \(\underline{L}_{k+1}^{\bar{a}}\) from \(A_k\) do not suffice.

A sequentially randomized trial implies still stronger conditions, which cannot be read from SWIGs and are not needed for identification, e.g.

  • \(\mathopen{}\left\{\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{}; \text{all } \bar{a}\right\}\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1}, \bar{L}_k\);
  • full sequential exchangeability: \(\mathopen{}\left(Y^{\bar{\mathcal{A}}}, \bar{L}^{\bar{\mathcal{A}}}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1}, \bar{L}_k\), where \(\bar{\mathcal{A}}\) is the set of all \(2^{K+1}\) static strategies, and \(Y^{\bar{\mathcal{A}}}\), \(\bar{L}^{\bar{\mathcal{A}}}\) the corresponding sets of counterfactuals (by analogy with Technical Point 2.1).

We write \(\underline{L}_{k+1}\) (an underbar for “from \(k + 1\) onward”) and \(\bar{\mathcal{A}}\) for the set of static strategies; the book’s typography for these symbols differs slightly. In the dynamic version we write the conditioning history as \(g(\bar{A}_{k-2}, \bar{L}_{k-1})\), matching Definition 6.


5.4 SWIGs for Dynamic Strategies

Figure 19.9 represents Figure 19.5 under a dynamic strategy \(g = [g_0, g_1(L_1)]\): \(A_0\) is set to a fixed \(g_0\), and \(A_1\) to \(g_1(L_1^g)\), which depends on the \(L_1^g\) observed after setting \(A_0 = g_0\).

Example 14 (“Treat at time 1 only if CD4 is low”) Strategy: do not treat at time 0; at time 1 treat only if the CD4 count is low (\(L_1^g = 1\)). So \(g_0 = 0\) for everybody, \(g_1(L_1^g) = 1\) when \(L_1^g = 1\), and \(g_1(L_1^g) = 0\) when \(L_1^g = 0\).

Remark 9 (The SWIG of a dynamic strategy).

  • The SWIG gains an arrow \(L_1^g \to g_1(L_1^g)\) that is not in the original DAG; it exists only in the counterfactual world of this strategy. It is drawn differently but treated like any other arrow for d-separation.
  • The outcome node is \(Y^g\).

5.5 Figure 19.9 Versus Figure 19.10

Example 15 (Dynamic strategies under Figures 19.5 and 19.6)  

  • Figure 19.9 (from Figure 19.5, Example 10): d-separation gives \(Y^g \perp\!\!\!\perp A_0\) and \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, L_1^g\) for every \(g\). Restricting the second to those with \(A_0 = g_0\) and applying consistency (\(L_1^g = L_1\) and \(A_1^g = A_1\) whenever \(A_0 = g_0\)) gives \(Y^g \perp\!\!\!\perp A_1 \mid A_0 = g_0, L_1\). Sequential exchangeability for \(Y^g\) holds, so (given sequential positivity and consistency) all strategies, static and dynamic, are identified.
  • Figure 19.10 (from Figure 19.6): because of the open path \[A_0 \leftarrow W_0 \to L_1^g \to g_1(L_1^g) \to Y^g,\] neither \(Y^g \perp\!\!\!\perp A_0\) nor sequential exchangeability for \(Y^g\) holds (generically, that is, under faithfulness). The book states that \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is then not identified for the dynamic strategy (Hernán and Robins 2020, sec. 19.5).

Another way to see the Figure 19.10 problem: the distribution of \(Y^g\) depends on that of \(g_1(L_1^g)\), hence on that of \(L_1^g\), and the open path \(A_0 \leftarrow W_0 \to L_1^g\) means that \(L_1^g \perp\!\!\!\perp A_0\) generically (under faithfulness) fails, so the distribution of \(L_1^g\) cannot be read off the data among those with \(A_0 = a_0\); the book concludes that it is not identified (Hernán and Robins 2020, sec. 19.5). Under a static strategy, no arrow from \(L_1^{a_0}\) into the treatment node feeds \(Y\), which is why static strategies remain identified under Figure 19.6.


NoteFine Point 19.2: Arrows From Intervention Nodes in SWIGs

SWIGs include arrows from intervention nodes such as \(a\) into later variables, even though in a single intervention world \(a\) is a constant that cannot affect anything.

  • The arrows record which variables \(A\) directly affects in the original DAG.
  • Without them, the SWIGs of Figure 7.14 and of Figure 7.14 plus an arrow \(A \to Y\) would be identical, and we could not tell that the effect is identified (by the front door formula) in the first DAG but not the second.

So when applying d-separation to a SWIG, every path through an intervention node is blocked, even though the node is not written in the conditioning event.

The same holds for deterministic strategies that depend on a baseline confounder \(L_0\) (replace \(g_0\) by \(g_0(L_0)\) in Figure 19.9): given \(L_0\), \(g_0(L_0)\) is a constant, so paths through it are blocked, and \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, L_0, L_1^g\) becomes, after setting \(A_0 = g(L_0)\) and using consistency, \(Y^g \perp\!\!\!\perp A_1 \mid A_0 = g(L_0), L_0, L_1\).

For a random strategy that draws \(A_0^{+,g}\) from a distribution that may depend on \(L_0\), \(A_0^{+,g}\) appears on the SWIG and must be conditioned on explicitly. The book reports that Richardson and Robins (2013) showed that \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, A_0^{+,g}, L_0, L_1^g\) is a necessary condition for identification by the g-formula for such a strategy (Hernán and Robins 2020, Fine Point 19.2). They also treated strategies that depend on the natural value of treatment \(A_t^g\) (Robins et al. 2004), recently called “modified treatment policies” (Diaz et al. 2021).


5.6 One Last Example

Example 16 (Adding an arrow from \(L_1\) to \(Y\)) Figure 19.11 is Figure 19.6 (Example 10) plus an arrow \(L_1 \to Y\) (Figure 19.12 is its SWIG under a static strategy). Here neither sequential exchangeability for \(Y^g\) (Definition 6) nor static sequential exchangeability for \(Y^{\bar{a}}\) (Definition 8) holds (generically, that is, under faithfulness). The book states that then no strategy’s effect is identified (Hernán and Robins 2020, sec. 19.5).

Table 2: Identification of static strategies and of dynamic strategies whose assignment depends on \(L_1\), under Figures 19.5, 19.6, and 19.11. “Not identified” is the book’s conclusion (Hernán and Robins 2020, sec. 19.5); the exchangeability conditions behind it fail generically, that is, under faithfulness.
Diagram Static strategies Dynamic strategies depending on \(L_1\)
Figure 19.5 identified identified
Figure 19.6 identified not identified
Figure 19.11 not identified not identified

Why Figure 19.11 fails even for static strategies: in its SWIG, \(L_1^{a_0}\) is now a cause of \(Y^{a_0, a_1}\), so the path \(A_0 \leftarrow W_0 \to L_1^{a_0} \to Y^{a_0, a_1}\) is open, and \(Y^{a_0, a_1} \perp\!\!\!\perp A_0\) generically (under faithfulness) fails; conditioning on \(L_1\) to block it is not allowed at time 0, because \(L_1\) comes after \(A_0\). This explanation is ours; the book states only the conclusion.

6 19.6 Time-Varying Confounding and Time-Varying Confounders (pp. 265-266)


WarningSequential Exchangeability Is Never Guaranteed

No form of sequential exchangeability is guaranteed in an observational study. Approximating it requires expert knowledge to decide which time-varying variables \(\bar{L}_k\) to measure (in HIV: CD4 count, viral load, symptoms). Whether the measured covariates suffice can never be known with certainty, but causal diagrams such as Figures 19.1-19.4 organize our beliefs.

These diagrams were drawn without selection (e.g., censoring) so as to focus on confounding.


6.1 \(L_1\) as a Confounder for \(A_1\)

Example 17 (\(L_1\) as a confounder for the effect of \(A_1\)) In Figure 19.5 (Example 10), consider the effect of \(A_1\) alone on \(Y\).

  • \(A_1\) and \(Y\) share the cause \(U_1\): the backdoor path \(A_1 \leftarrow L_1 \leftarrow U_1 \to Y\) is open, so there is confounding.
  • \(U_1\) is unmeasured, but conditioning on the measured \(L_1\) blocks the path.
  • So if \(L_1\) is measured, there is no unmeasured confounding for the effect of \(A_1\), and we call \(L_1\) a confounder for that effect, even though the actual common cause is \(U_1\) (see Section 7.3) and even though there is no arrow \(L_1 \to Y\).
WarningConditioning on a Collider Opens Another Path

Conditioning on \(L_1\), a collider, opens a second backdoor path, \(A_1 \leftarrow A_0 \to L_1 \leftarrow U_1 \to Y\). That path can be blocked by also conditioning on prior treatment \(A_0\), if it is available.


Now consider the long diagram over all times \(k = 0, 1, 2, \ldots\), in which each \(L_k\) affects later treatments \(A_k, A_{k+1}, \ldots\) and shares unmeasured causes \(U_k\) with \(Y\). To estimate the effects of strategies defined by interventions on \(A_0, A_1, A_2, \ldots\), we must condition at each \(k\) on the covariate history \(\bar{L}_k\), together with \(\bar{A}_{k-1}\), to block the backdoor paths between \(A_k\) and \(Y\).

Definition 9 (Time-varying confounders) Covariates \(\bar{L}_k\) that, together with \(\bar{A}_{k-1}\), are needed at each time \(k\) to block the backdoor paths between \(A_k\) and \(Y\) are time-varying confounders for the effect of \(\bar{A}\) on \(Y\).

  • No unmeasured confounding for \(\bar{A}\) requires that \(\bar{L}_k\) be measured for everyone.
  • Time-varying confounders are sometimes called time-dependent confounders.
WarningMeasuring the Confounders Is Not Enough
  • Whether all confounders are measured cannot be checked from the data: as Section 19.3 noted, the data cannot distinguish Figure 19.2 from Figure 19.3.
  • Even with all confounders measured and modeled correctly, most adjustment methods can give biased comparisons of strategies; Chapter 20 explains why g-methods are needed.

NoteFine Point 19.3: A Definition of Time-Varying Confounding

Assume no selection bias.

  • There is confounding for effects involving \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{}\) if \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\).
  • The confounding is solely time-fixed (wholly due to baseline covariates) if \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid L_0\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}, L_0\right]\mathclose{}\), as when the only arrows into \(A_1\) in Figure 19.2 come from \(A_0\) and \(L_0\).
  • If the identifiability conditions hold but \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid L_0\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}, L_0\right]\mathclose{}\), there is time-varying confounding.
  • If the identifiability conditions fail, as they generically do in Figure 19.3 (under faithfulness), there is unmeasured confounding.

A sufficient condition for no time-varying confounding is unconditional sequential exchangeability (Definition 7), \(Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}\), as in Figure 19.1. That diagram can in fact be pruned to \(A_0\), \(A_1\), and \(Y\): \(L_1\) is not a common cause of two nodes, so drop it; then \(L_0\) and \(U_1\) are no longer common causes, so drop them; then \(U_0\) is not either.

Remark 10 (Treatment-confounder feedback is not yet required). Note what the chapter does not yet require: a time-varying confounder here is simply a covariate needed to block backdoor paths at some time \(k\). Whether that covariate is also affected by earlier treatment (treatment-confounder feedback) is the subject of Chapter 20, and it is that feedback that makes conventional adjustment fail.

7 Summary


  • The average causal effect of a time-varying treatment is a contrast of \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) under two treatment strategies; it is not uniquely defined until the strategies are specified.
  • Strategies are static (\(\bar{a}\), not depending on covariates) or dynamic (\(g\), depending on the evolving \(\bar{L}_k\)) (Definition 4), and deterministic or random.
  • In a sequentially randomized experiment (Definition 5), treatment at each \(k\) is randomized given \((\bar{A}_{k-1}, \bar{L}_k)\); such randomization implies sequential exchangeability (Definition 6).
  • Identification of all strategies is guaranteed by sequential exchangeability, sequential positivity, and consistency.
  • Static sequential exchangeability (Definition 8) is weaker and identifies only static strategies; SWIGs show when a diagram (e.g., Figure 19.6) supports static but not dynamic strategies.
  • Time-varying confounders are covariates \(\bar{L}_k\) needed at each time to block backdoor paths; Chapter 20 explains why conventional methods may fail to adjust for them when there is treatment-confounder feedback.

8 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top