Chapter 19: Time-Varying Treatments
Parts I and II of the book considered time-fixed treatments, whose value is determined once, at the start of follow-up. Many causal questions instead involve treatments whose value can change over time for the same individual: medical treatments, lifestyle habits, employment or marital status, occupational exposures. Part III extends the framework to these time-varying treatments, and this chapter introduces the terminology and concepts it needs.
This chapter is based on Hernán and Robins (2020, chap. 19, pp. 255-266).
The authors warn that this is one of the most technical chapters of the book, because simplifying the concepts further would cost too much rigor. The chapter defines causal effects for time-varying treatments, introduces treatment strategies and sequentially randomized experiments, states the identifiability conditions (sequential exchangeability, positivity, consistency), and uses SWIGs to show that some strategies can be identified when others cannot. How to estimate these effects is the subject of Chapters 20 and 21.
1 19.1 The Causal Effect of Time-Varying Treatments (pp. 255-256)
For a time-fixed treatment \(A\) (1: treated, 0: untreated) given at time zero and an outcome \(Y\) measured 60 months later, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\). It need not mention when treatment occurs, because everybody’s treatment is determined at the same time.
A time-varying treatment needs time to be made explicit.
2 19.2 Treatment Strategies (pp. 256-257)
The rules contrasted in Remark 2 are treatment strategies.
A general counterfactual theory for comparing treatment strategies was first articulated by Robins (1986, 1987, 1997a).
Many other contrasts are possible.
Proof. The vector \((a_0, \ldots, a_K)\) has \(K + 1\) components, each taking one of two values, so there are \(2^{K+1}\) such vectors.
Fixed treatment sequences do not exhaust all strategies.
The book writes “at least \(2^K\)”, a lower bound, and uses \(2^K\) again when counting fixed sequences in Technical Point 19.3; these notes write \(2^{K+1}\) throughout. The point is only that the number grows exponentially with the length of follow-up.
2.1 Dynamic Strategies
- A deterministic dynamic strategy is a rule \(g = [g_0(\bar{a}_{-1}, l_0), \ldots, g_K(\bar{a}_{K-1}, \bar{l}_K)]\), where \(g_k(\bar{a}_{k-1}, \bar{l}_k)\) is the treatment assigned at \(k\) to an individual with past history \((\bar{a}_{k-1}, \bar{l}_k)\). Example: \(g_k(\bar{a}_{k-1}, \bar{l}_k) = 1\) if the CD4 count was low at or before \(k\), and 0 otherwise.
- A deterministic static strategy is a rule \(g = [g_0(\bar{a}_{-1}), \ldots, g_K(\bar{a}_{K-1})]\) that does not depend on \(\bar{l}_k\).
- Random strategies assign a probability of treatment rather than a value. They can be static (“each month, independently, treat with probability 0.3”) or dynamic (“each month, independently, treat with probability 0.3 if the CD4 count is low; do not treat if it is high”).
The strategy \(g\) that maximizes \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) (when higher outcomes are better) is the optimal treatment strategy. For a drug, it will almost always be dynamic, since treatment must stop when toxicity develops. No random strategy can be preferred to the optimal deterministic one, but random strategies (randomized trials) remain scientifically necessary, because before the trial nobody knows which deterministic strategy is optimal. Unless noted otherwise, \(g\) denotes a deterministic strategy.
The book cites Young et al. (2014) for a taxonomy of treatment strategies.
2.2 Causal Effects Require Specified Strategies
Every dynamic strategy \(g\) that depends on past treatment and covariates has a counterpart \(g'\) that depends only on past covariates, defined recursively by \(g'_0(l_0) = g_0(\bar{a}_{-1} = 0, l_0)\) and \(g'_k(\bar{l}_k) = g_k\mathopen{}\left(g'_{k-1}(\bar{l}_{k-1}), \bar{l}_k\right)\mathclose{}\). By consistency, an individual has the same treatment, covariate, and outcome history whether following \(g\) or \(g'\) from time zero, so \(Y^g = Y^{g'}\). In the observed data, an individual has followed \(g\) through time \(t\) if and only if they have followed \(g'\) through \(t\). If treatment under \(g\) already does not depend on past treatment, \(g\) and \(g'\) are identical.
In the recursion, \(g'_{k-1}(\bar{l}_{k-1})\) stands for the whole treatment history \(\bar{a}_{k-1}\) that \(g'\) would have produced through \(k-1\) given \(\bar{l}_{k-1}\).
3 19.3 Sequentially Randomized Experiments (pp. 257-259)
The book uses three causal diagrams with the first two times, \(k = 0\) and \(k = 1\). In each, \(A_k\) is treatment, \(L_k\) the measured variables, \(U_k\) unmeasured common causes of at least two variables, and \(Y\) the outcome. In the HIV example, CD4 count \(L_k\) is a consequence of unmeasured immune damage \(U_k\), which also lowers health status \(Y\).
| Figure | Arrows from \(\bar{A}_{k-1}\) into \(A_k\) | Arrows from \(\bar{L}_k\) into \(A_k\) | Arrows from \(\bar{U}_k\) into \(A_k\) |
|---|---|---|---|
| 19.1 | yes (for \(k \geq 1\)) | no | no |
| 19.2 | yes (for \(k \geq 1\)) | yes | no |
| 19.3 | yes (for \(k \geq 1\)) | yes | yes |
All participants are assumed to adhere to the assigned treatment. A causal graph must include all common causes of any two variables on it, which is why the \(U_k\) appear even though they are not measured.
3.1 Figure 19.1: Treatment Depends Only on Past Treatment
- For static strategies (Definition 4), treatment assignment that depends only on past treatment is the time-varying generalization of no confounding by measured or unmeasured variables.
- So under Figure 19.1 (where treatment is randomized at each time), association is causation for static strategies: the counterfactual mean under a static strategy \(\bar{a}\) is the mean outcome among those who followed \(\bar{a}\) (provided some individuals followed \(\bar{a}\), and under consistency).
The equality \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\) does not in general extend to dynamic strategies. Under a dynamic strategy \(g\) that depends on \(L\), generically, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) equals the mean outcome among those who followed \(g\) only if the probability of \(A_k = 1\) is exactly 0.5 at every time \(k\) at which treatment under \(g\) depends on \(L_k\). Otherwise, identifying \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) requires g-methods applied to data on \(L\), \(A\), and \(Y\), under either Figure 19.1 or Figure 19.2.
3.2 Figure 19.2: Treatment Also Depends on Measured Covariates
The book labels these strata “high CD4 (\(A_{k-1} = 0, L_k = 1\))” and “low CD4 (\(A_{k-1} = 0, L_k = 0\))”, which reverses the coding of Section 19.2 (\(L_k = 1\) means low CD4). We describe the strata by CD4 level only; under the Section 19.2 coding, high CD4 is \(L_k = 0\) and low CD4 is \(L_k = 1\).
3.3 Observational Studies
Treatment decisions usually depend on prognostic factors, so observational studies are typically represented by Figure 19.2 or 19.3, not 19.1.
- If clinicians’ decisions at each month depend on \((\bar{A}_{k-1}, \bar{L}_k)\) but on no unmeasured \(\bar{U}_k\), the study is represented by Figure 19.2. It then differs from a sequentially randomized experiment (Definition 5) only in that the assignment probabilities are unknown (though estimable from the data).
The data cannot tell us whether Figure 19.2 or Figure 19.3 is correct. Figure 19.3 has unmeasured confounding.
Sequentially randomized experiments (Definition 5) are rarely run in practice, but the concept clarifies the conditions needed for valid estimation of the effects of time-varying treatments.
4 19.4 Sequential Exchangeability (pp. 259-260)
For a time-fixed treatment, valid inference typically requires conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), which holds in conditionally randomized experiments and in observational studies where treatment depends on measured \(L\) and, given \(L\), on no unmeasured common cause of treatment and outcome.
For a time-varying treatment we need conditional exchangeability at each time, given the covariate history \(\bar{L}_k\).
The book drops the word “conditional” and says simply sequential exchangeability.
For individuals whose treatment history \([A_0 = g(L_0), A_1 = g(A_0, L_0, L_1)]\) is compatible with strategy \(g\) through the end of follow-up, consistency gives \(Y^g = Y\), which also equals the counterfactual outcome under the static strategy \((a_0, a_1) = (A_0, A_1)\).
“This form of sequential exchangeability” signals that there are others, some weaker and some stronger.
Whenever the book talks about identification, the identifying formula is the g-formula (Chapter 21). In rare cases, not relevant here, effects can be identified by formulas related to but different from the g-formula (e.g., the front door formula of Technical Point 7.3).
4.1 Unconditional Sequential Exchangeability
Under Figure 19.1 (no arrows from \(\bar{L}_k\) or \(\bar{U}_k\) into \(A_k\); Table 1), unconditional sequential exchangeability holds for every static strategy \(\bar{a}\).
Proof. The positivity conditions give \(\Pr[\bar{A}_k = \bar{a}_k] = \prod_{j=0}^{k} \Pr[A_j = a_j \mid \bar{A}_{j-1} = \bar{a}_{j-1}] > 0\) for every \(k\), so each conditional expectation below is well defined (for \(k = 0\), conditioning on the empty history \(\bar{A}_{-1} = \bar{a}_{-1}\) is no conditioning).
One-step identity: for each \(k = 0, 1, \ldots, K\),
\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}, A_k = a_k\right]\mathclose{} && \text{(exchangeability at } k\text{; positivity at } k\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_k = \bar{a}_k\right]\mathclose{} && \text{(since } \bar{A}_k = (\bar{A}_{k-1}, A_k)\text{)} \end{align}\]
Applying the one-step identity for \(k = 0\), then \(k = 1\), and so on through \(k = K\):
\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_0 = \bar{a}_0\right]\mathclose{} && \text{(one-step identity, } k = 0\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_1 = \bar{a}_1\right]\mathclose{} && \text{(one-step identity, } k = 1\text{)} \\ &\;\;\vdots \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_K = \bar{a}_K\right]\mathclose{} && \text{(one-step identity, } k = K\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{} && \text{(consistency: } Y^{\bar{a}} = Y \text{ when } \bar{A} = \bar{a}\text{)} \end{align}\]
4.2 A Diagram With Partial Identification
4.3 Positivity and Consistency
Sequential positivity:
\[\text{if } f_{\bar{A}_{k-1}, \bar{L}_k}(\bar{a}_{k-1}, \bar{l}_k) \neq 0, \text{ then } f_{A_k \mid \bar{A}_{k-1}, \bar{L}_k}(a_k \mid \bar{a}_{k-1}, \bar{l}_k) > 0 \text{ for all } \bar{a}_k, \bar{l}_k.\]
It holds in a sequentially randomized experiment (Definition 5) if the randomization probabilities are never 0 or 1, whatever the past history. For a particular strategy \(g\), it need hold only for treatment histories compatible with \(g\), i.e., \(a_k = g(\bar{a}_{k-1}, \bar{l}_k)\) for each \(k\).
Sequential consistency:
\[Y^{\bar{a}} = Y^{\bar{a}^*} \text{ if } \bar{a}^* = \bar{a}; \quad Y^{\bar{a}} = Y \text{ if } \bar{A} = \bar{a}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k^{\bar{a}^*} \text{ if } \bar{a}^*_{k-1} = \bar{a}_{k-1}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k \text{ if } \bar{A}_{k-1} = \bar{a}_{k-1},\]
where \(\bar{L}_k^{\bar{a}}\) is the counterfactual covariate history through \(k\) under \(\bar{a}\). Weaker conditions suffice for identification (“if \(\bar{A} = \bar{a}\) then \(Y^{\bar{a}} = Y\)” for static strategies; “if \(A_k = g_k(\bar{A}_{k-1}, \bar{L}_k)\) at each \(k\) then \(Y^g = Y\)” for dynamic ones), but the book always accepts the stronger sequential version. If “treat in month \(k\)” and “do not treat in month \(k\)” are sufficiently well defined at all \(k\), so are all static and dynamic strategies built from them.
5 19.5 Identifiability Under Some but Not All Treatment Strategies (pp. 261-264)
But this generalized criterion works on causal DAGs, which contain no counterfactual outcomes, so it does not show the connection to sequential exchangeability directly. SWIGs (Chapter 7) do, and are especially useful here.
Pearl and Robins (1995) proposed a generalized backdoor criterion for static strategies; Robins (1997a) extended it to dynamic strategies.
5.1 Two Simplified Diagrams
For \(L_1\) and \(Y\) we list only some parents; the book’s figures give the complete diagrams.
5.2 Building the SWIG for a Static Strategy
\(L_1\) is written \(L_1^{a_0}\), not \(L_1^{a_0, a_1}\), because a later intervention on \(A_1\) cannot affect the earlier \(L_1\). Unlike the DAG, the SWIG contains the counterfactual outcome, so exchangeability can be checked by d-separation.
5.3 Static Sequential Exchangeability
Proof (From the SWIG statement to an observed-data statement).
- \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}\) implies, in particular, \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0 = a_0, L_1^{a_0}\) (restrict to those who received \(A_0 = a_0\)).
- Among those with \(A_0 = a_0\), consistency gives \(L_1^{a_0} = L_1\) and \(A_1^{a_0} = A_1\).
- Substituting: \(Y^{a_0, a_1} \perp\!\!\!\perp A_1 \mid A_0 = a_0, L_1\).
For a sequentially randomized experiment with times \(k = 0, 1, \ldots, K\) (a longer version of Figure 19.7), the SWIG gives
\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k^{\bar{a}_{k-1}} \mid \bar{A}_{k-1}^{\bar{a}_{k-2}}, \bar{L}_k^{\bar{a}_{k-1}},\]
where \(\underline{L}_{k+1}^{\bar{a}}\) is the counterfactual covariate history from \(k + 1\) to the end of follow-up. Restricting to \(\bar{A}_{k-1}^{\bar{a}_{k-2}} = \bar{a}_{k-1}\) and applying consistency gives
\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \bar{L}_k.\]
When this holds for all \(\bar{a}\), the book calls it sequential exchangeability; these notes call it joint sequential exchangeability, to keep it distinct from Definition 6, which it implies. Although it mentions only static strategies, it is equivalent to
\[\mathopen{}\left(Y^{g}, \underline{L}_{k+1}^{g}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = g(\bar{A}_{k-2}, \bar{L}_{k-1}), \bar{L}_k \quad \text{for all } g,\]
and so, with positivity and consistency, it identifies the outcome and covariate distributions under every static and dynamic strategy (Robins 1986). This rests on the joint independence of \((Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}})\) and \(A_k\); for dynamic strategies, the two separate independences of \(Y^{\bar{a}}\) and of \(\underline{L}_{k+1}^{\bar{a}}\) from \(A_k\) do not suffice.
A sequentially randomized trial implies still stronger conditions, which cannot be read from SWIGs and are not needed for identification, e.g.
- \(\mathopen{}\left\{\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{}; \text{all } \bar{a}\right\}\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1}, \bar{L}_k\);
- full sequential exchangeability: \(\mathopen{}\left(Y^{\bar{\mathcal{A}}}, \bar{L}^{\bar{\mathcal{A}}}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1}, \bar{L}_k\), where \(\bar{\mathcal{A}}\) is the set of all \(2^{K+1}\) static strategies, and \(Y^{\bar{\mathcal{A}}}\), \(\bar{L}^{\bar{\mathcal{A}}}\) the corresponding sets of counterfactuals (by analogy with Technical Point 2.1).
We write \(\underline{L}_{k+1}\) (an underbar for “from \(k + 1\) onward”) and \(\bar{\mathcal{A}}\) for the set of static strategies; the book’s typography for these symbols differs slightly. In the dynamic version we write the conditioning history as \(g(\bar{A}_{k-2}, \bar{L}_{k-1})\), matching Definition 6.
5.4 SWIGs for Dynamic Strategies
Figure 19.9 represents Figure 19.5 under a dynamic strategy \(g = [g_0, g_1(L_1)]\): \(A_0\) is set to a fixed \(g_0\), and \(A_1\) to \(g_1(L_1^g)\), which depends on the \(L_1^g\) observed after setting \(A_0 = g_0\).
5.5 Figure 19.9 Versus Figure 19.10
Another way to see the Figure 19.10 problem: the distribution of \(Y^g\) depends on that of \(g_1(L_1^g)\), hence on that of \(L_1^g\), and the open path \(A_0 \leftarrow W_0 \to L_1^g\) means that \(L_1^g \perp\!\!\!\perp A_0\) generically (under faithfulness) fails, so the distribution of \(L_1^g\) cannot be read off the data among those with \(A_0 = a_0\); the book concludes that it is not identified (Hernán and Robins 2020, sec. 19.5). Under a static strategy, no arrow from \(L_1^{a_0}\) into the treatment node feeds \(Y\), which is why static strategies remain identified under Figure 19.6.
SWIGs include arrows from intervention nodes such as \(a\) into later variables, even though in a single intervention world \(a\) is a constant that cannot affect anything.
- The arrows record which variables \(A\) directly affects in the original DAG.
- Without them, the SWIGs of Figure 7.14 and of Figure 7.14 plus an arrow \(A \to Y\) would be identical, and we could not tell that the effect is identified (by the front door formula) in the first DAG but not the second.
So when applying d-separation to a SWIG, every path through an intervention node is blocked, even though the node is not written in the conditioning event.
The same holds for deterministic strategies that depend on a baseline confounder \(L_0\) (replace \(g_0\) by \(g_0(L_0)\) in Figure 19.9): given \(L_0\), \(g_0(L_0)\) is a constant, so paths through it are blocked, and \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, L_0, L_1^g\) becomes, after setting \(A_0 = g(L_0)\) and using consistency, \(Y^g \perp\!\!\!\perp A_1 \mid A_0 = g(L_0), L_0, L_1\).
For a random strategy that draws \(A_0^{+,g}\) from a distribution that may depend on \(L_0\), \(A_0^{+,g}\) appears on the SWIG and must be conditioned on explicitly. The book reports that Richardson and Robins (2013) showed that \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, A_0^{+,g}, L_0, L_1^g\) is a necessary condition for identification by the g-formula for such a strategy (Hernán and Robins 2020, Fine Point 19.2). They also treated strategies that depend on the natural value of treatment \(A_t^g\) (Robins et al. 2004), recently called “modified treatment policies” (Diaz et al. 2021).
5.6 One Last Example
| Diagram | Static strategies | Dynamic strategies depending on \(L_1\) |
|---|---|---|
| Figure 19.5 | identified | identified |
| Figure 19.6 | identified | not identified |
| Figure 19.11 | not identified | not identified |
Why Figure 19.11 fails even for static strategies: in its SWIG, \(L_1^{a_0}\) is now a cause of \(Y^{a_0, a_1}\), so the path \(A_0 \leftarrow W_0 \to L_1^{a_0} \to Y^{a_0, a_1}\) is open, and \(Y^{a_0, a_1} \perp\!\!\!\perp A_0\) generically (under faithfulness) fails; conditioning on \(L_1\) to block it is not allowed at time 0, because \(L_1\) comes after \(A_0\). This explanation is ours; the book states only the conclusion.
6 19.6 Time-Varying Confounding and Time-Varying Confounders (pp. 265-266)
No form of sequential exchangeability is guaranteed in an observational study. Approximating it requires expert knowledge to decide which time-varying variables \(\bar{L}_k\) to measure (in HIV: CD4 count, viral load, symptoms). Whether the measured covariates suffice can never be known with certainty, but causal diagrams such as Figures 19.1-19.4 organize our beliefs.
These diagrams were drawn without selection (e.g., censoring) so as to focus on confounding.
6.1 \(L_1\) as a Confounder for \(A_1\)
Conditioning on \(L_1\), a collider, opens a second backdoor path, \(A_1 \leftarrow A_0 \to L_1 \leftarrow U_1 \to Y\). That path can be blocked by also conditioning on prior treatment \(A_0\), if it is available.
Now consider the long diagram over all times \(k = 0, 1, 2, \ldots\), in which each \(L_k\) affects later treatments \(A_k, A_{k+1}, \ldots\) and shares unmeasured causes \(U_k\) with \(Y\). To estimate the effects of strategies defined by interventions on \(A_0, A_1, A_2, \ldots\), we must condition at each \(k\) on the covariate history \(\bar{L}_k\), together with \(\bar{A}_{k-1}\), to block the backdoor paths between \(A_k\) and \(Y\).
- No unmeasured confounding for \(\bar{A}\) requires that \(\bar{L}_k\) be measured for everyone.
- Time-varying confounders are sometimes called time-dependent confounders.
- Whether all confounders are measured cannot be checked from the data: as Section 19.3 noted, the data cannot distinguish Figure 19.2 from Figure 19.3.
- Even with all confounders measured and modeled correctly, most adjustment methods can give biased comparisons of strategies; Chapter 20 explains why g-methods are needed.
Assume no selection bias.
- There is confounding for effects involving \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{}\) if \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\).
- The confounding is solely time-fixed (wholly due to baseline covariates) if \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid L_0\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}, L_0\right]\mathclose{}\), as when the only arrows into \(A_1\) in Figure 19.2 come from \(A_0\) and \(L_0\).
- If the identifiability conditions hold but \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid L_0\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}, L_0\right]\mathclose{}\), there is time-varying confounding.
- If the identifiability conditions fail, as they generically do in Figure 19.3 (under faithfulness), there is unmeasured confounding.
A sufficient condition for no time-varying confounding is unconditional sequential exchangeability (Definition 7), \(Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}\), as in Figure 19.1. That diagram can in fact be pruned to \(A_0\), \(A_1\), and \(Y\): \(L_1\) is not a common cause of two nodes, so drop it; then \(L_0\) and \(U_1\) are no longer common causes, so drop them; then \(U_0\) is not either.
7 Summary
- The average causal effect of a time-varying treatment is a contrast of \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) under two treatment strategies; it is not uniquely defined until the strategies are specified.
- Strategies are static (\(\bar{a}\), not depending on covariates) or dynamic (\(g\), depending on the evolving \(\bar{L}_k\)) (Definition 4), and deterministic or random.
- In a sequentially randomized experiment (Definition 5), treatment at each \(k\) is randomized given \((\bar{A}_{k-1}, \bar{L}_k)\); such randomization implies sequential exchangeability (Definition 6).
- Identification of all strategies is guaranteed by sequential exchangeability, sequential positivity, and consistency.
- Static sequential exchangeability (Definition 8) is weaker and identifies only static strategies; SWIGs show when a diagram (e.g., Figure 19.6) supports static but not dynamic strategies.
- Time-varying confounders are covariates \(\bar{L}_k\) needed at each time to block backdoor paths; Chapter 20 explains why conventional methods may fail to adjust for them when there is treatment-confounder feedback.