Parts I and II of the book considered time-fixed treatments, whose value is determined once, at the start of follow-up. Many causal questions instead involve treatments whose value can change over time for the same individual: medical treatments, lifestyle habits, employment or marital status, occupational exposures. Part III extends the framework to these time-varying treatments, and this chapter introduces the terminology and concepts it needs.
For a time-fixed treatment \(A\) (1: treated, 0: untreated) given at time zero and an outcome \(Y\) measured 60 months later, the average causal effect is the contrast \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\). It need not mention when treatment occurs, because everybody’s treatment is determined at the same time.
A time-varying treatment needs time to be made explicit.
Definition 1 (Time-indexed treatment) Let \(K\) be the last time at which treatment can be given. For each time \(k = 0, 1, \ldots, K\), \(A_k\) denotes the value of a dichotomous treatment at time \(k\).
Example 1 (HIV and antiretroviral therapy) In a 5-year follow-up study of individuals infected with HIV, time \(k\) is counted in months, so \(K = 59\). \(A_k = 1\) if the individual receives antiretroviral therapy in month \(k\) and \(A_k = 0\) otherwise. Nobody was treated before the study started: \(A_{-1} = 0\) for all individuals. The outcome \(Y\) measures health status (higher is better) at the end of follow-up, time \(K + 1 = 60\).
Definition 2 (Treatment history) An overbar denotes history: \(\bar{A}_k = (A_0, A_1, \ldots, A_k)\) is the treatment history from time 0 through time \(k\). The whole history through \(K\), \(\bar{A}_K\), is often written \(\bar{A}\) without a subscript. Lower-case letters denote realizations: \(a_k\) is a value of \(A_k\).
Example 2 (Treatment histories)
Remark 2 (The effect of a time-varying treatment is not unique). A contrast at a single time,
\[\operatorname{E}\mathopen{}\left[Y^{a_k = 1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_k = 0}\right]\mathclose{},\]
quantifies the effect of treatment at time \(k\) only, not the effect of treatment at all times between 0 and \(K\). The average causal effect of a time-varying treatment must instead contrast mean counterfactual outcomes under two rules, each of which specifies treatment at every time from \(k = 0\) to \(k = K\). Because there are many such rules, the average causal effect of a time-varying treatment is not uniquely defined.
Definition 3 (Treatment strategy) A treatment strategy (also called a plan, policy, protocol, or regime) is a rule to assign treatment at each time \(k\) of follow-up.
The rules contrasted in Remark 2 are treatment strategies.
Example 3 (Always treat versus never treat) The strategies “always treat”, \(\bar{a} = \bar{1}\), and “never treat”, \(\bar{a} = \bar{0}\), define one average causal effect of \(\bar{A}\) on \(Y\):
\[\operatorname{E}\mathopen{}\left[Y^{\bar{a} = \bar{1}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\bar{a} = \bar{0}}\right]\mathclose{}.\]
Many other contrasts are possible.
Example 4 (Other contrasts of fixed treatment sequences) For example, \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\bar{a}'}\right]\mathclose{}\) could compare
Proposition 1 (Number of fixed treatment sequences) With a dichotomous treatment \(a_k \in \{0, 1\}\) at each time \(k = 0, 1, \ldots, K\), there are \(2^{K+1}\) fixed treatment sequences \(\bar{a} = (a_0, a_1, \ldots, a_K) \in \{0, 1\}^{K+1}\).
Proof. The vector \((a_0, \ldots, a_K)\) has \(K + 1\) components, each taking one of two values, so there are \(2^{K+1}\) such vectors.
Fixed treatment sequences do not exhaust all strategies.
Example 5 (Starting treatment when the CD4 count drops) Let \(L_k\) be the CD4 cell count (cells per microliter) measured at month \(k\), coded \(L_k = 1\) when low (bad prognosis) and \(L_k = 0\) otherwise; everybody starts with a high count, \(L_0 = 0\). Consider the strategy
do not treat while \(L_k = 0\); start treatment when \(L_k = 1\) and treat continuously after that.
This strategy cannot be written as a fixed vector \(\bar{a} = (a_0, a_1, \ldots, a_K)\) that gives every individual the same \(a_k\) at time \(k\): at each time, who is treated depends on each individual’s evolving \(L_k\).
Definition 4 (Dynamic and static strategies) A dynamic treatment strategy is a rule in which the treatment \(a_k\) at time \(k\) depends on the evolution of an individual’s time-varying covariates \(\bar{L}_k\). Strategies \(\bar{a}\) in which treatment does not depend on covariates are non-dynamic or static treatment strategies.
Fine Point 19.1: Deterministic and Random Treatment Strategies
The strategy \(g\) that maximizes \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) (when higher outcomes are better) is the optimal treatment strategy. For a drug, it will almost always be dynamic, since treatment must stop when toxicity develops. No random strategy can be preferred to the optimal deterministic one, but random strategies (randomized trials) remain scientifically necessary, because before the trial nobody knows which deterministic strategy is optimal. Unless noted otherwise, \(g\) denotes a deterministic strategy.
Remark 3 (Causal effects are contrasts of specified strategies). The average causal effect of a time-varying treatment is well defined only once the strategies being compared are specified, e.g.
\(g\) denotes any strategy, static or dynamic; for a static strategy we sometimes write \(Y^{g = \bar{a}}\) rather than \(Y^g\) or \(Y^{\bar{a}}\). Even with only two options (treat or not) at each time, there are as many causal effects as there are pairs of strategies.
Technical Point 19.1: On the Definition of Dynamic Strategies
Every dynamic strategy \(g\) that depends on past treatment and covariates has a counterpart \(g'\) that depends only on past covariates, defined recursively by \(g'_0(l_0) = g_0(\bar{a}_{-1} = 0, l_0)\) and \(g'_k(\bar{l}_k) = g_k\mathopen{}\left(g'_{k-1}(\bar{l}_{k-1}), \bar{l}_k\right)\mathclose{}\). By consistency, an individual has the same treatment, covariate, and outcome history whether following \(g\) or \(g'\) from time zero, so \(Y^g = Y^{g'}\). In the observed data, an individual has followed \(g\) through time \(t\) if and only if they have followed \(g'\) through \(t\). If treatment under \(g\) already does not depend on past treatment, \(g\) and \(g'\) are identical.
The book uses three causal diagrams with the first two times, \(k = 0\) and \(k = 1\). In each, \(A_k\) is treatment, \(L_k\) the measured variables, \(U_k\) unmeasured common causes of at least two variables, and \(Y\) the outcome. In the HIV example, CD4 count \(L_k\) is a consequence of unmeasured immune damage \(U_k\), which also lowers health status \(Y\).
| Figure | Arrows from \(\bar{A}_{k-1}\) into \(A_k\) | Arrows from \(\bar{L}_k\) into \(A_k\) | Arrows from \(\bar{U}_k\) into \(A_k\) |
|---|---|---|---|
| 19.1 | yes (for \(k \geq 1\)) | no | no |
| 19.2 | yes (for \(k \geq 1\)) | yes | no |
| 19.3 | yes (for \(k \geq 1\)) | yes | yes |
Example 6 (Assignment depending only on past treatment) Example: each month, treatment is assigned with probability 0.5 to those untreated the previous month (\(A_{k-1} = 0\)) and with probability 1 to those treated the previous month (\(A_{k-1} = 1\)).
Example 7 (Assignment depending also on CD4 count) Example: each month, treatment is assigned with probability
There is confounding by measured, but not unmeasured, variables.
Definition 5 (Sequentially randomized experiment) An experiment in which treatment is randomly assigned to each individual at each time \(k\), with assignment probabilities that depend only on the measured history \((\bar{A}_{k-1}, \bar{L}_k)\), is a sequentially randomized experiment.
Remark 4 (Which diagrams can be sequentially randomized).
Treatment decisions usually depend on prognostic factors, so observational studies are typically represented by Figure 19.2 or 19.3, not 19.1.
The Data Cannot Rule Out Unmeasured Confounding
The data cannot tell us whether Figure 19.2 or Figure 19.3 is correct. Figure 19.3 has unmeasured confounding.
For a time-fixed treatment, valid inference typically requires conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), which holds in conditionally randomized experiments and in observational studies where treatment depends on measured \(L\) and, given \(L\), on no unmeasured common cause of treatment and outcome.
For a time-varying treatment we need conditional exchangeability at each time, given the covariate history \(\bar{L}_k\).
Definition 6 (Sequential exchangeability for \(Y^g\)) \[Y^g \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = g(\bar{A}_{k-2}, \bar{L}_{k-1}), \bar{L}_k \quad \text{for all strategies } g \text{ and } k = 0, 1, \ldots, K.\]
That is, at each time \(k\), the treated and untreated are exchangeable for \(Y^g\) given the covariate history \(\bar{L}_k\) through time \(k\) and any observed treatment history compatible with \(g\).
Example 8 (Sequential exchangeability with two times) With two times (\(K = 1\)), sequential exchangeability for \(Y^g\) (Definition 6) says that, for all strategies \(g\),
\[Y^g \perp\!\!\!\perp A_0 \mid L_0 \quad\text{and}\quad Y^g \perp\!\!\!\perp A_1 \mid A_0 = g(L_0), L_0, L_1.\]
Remark 5 (When sequential exchangeability holds).
Definition 7 (Unconditional sequential exchangeability) Unconditional sequential exchangeability holds for a static strategy \(\bar{a}\) if
\[Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \qquad k = 0, 1, \ldots, K.\]
Under Figure 19.1 (no arrows from \(\bar{L}_k\) or \(\bar{U}_k\) into \(A_k\); Table 1), unconditional sequential exchangeability holds for every static strategy \(\bar{a}\).
Proposition 2 (Association is causation under unconditional sequential exchangeability) Let \(\bar{a}\) be a static strategy. Suppose unconditional sequential exchangeability (Definition 7) holds for \(\bar{a}\), positivity holds along \(\bar{a}\) (\(\Pr[A_k = a_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}] > 0\) for \(k = 0, 1, \ldots, K\)), and consistency holds (\(Y^{\bar{a}} = Y\) whenever \(\bar{A} = \bar{a}\)). Suppose also that \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{\bar{a}}\right|\mathclose{}\right]\mathclose{} < \infty\). Then association is causation: \(\operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{}\).
Proof. The positivity conditions give \(\Pr[\bar{A}_k = \bar{a}_k] = \prod_{j=0}^{k} \Pr[A_j = a_j \mid \bar{A}_{j-1} = \bar{a}_{j-1}] > 0\) for every \(k\), so each conditional expectation below is well defined (for \(k = 0\), conditioning on the empty history \(\bar{A}_{-1} = \bar{a}_{-1}\) is no conditioning).
One-step identity: for each \(k = 0, 1, \ldots, K\),
\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_{k-1} = \bar{a}_{k-1}, A_k = a_k\right]\mathclose{} && \text{(exchangeability at } k\text{; positivity at } k\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_k = \bar{a}_k\right]\mathclose{} && \text{(since } \bar{A}_k = (\bar{A}_{k-1}, A_k)\text{)} \end{align}\]
Applying the one-step identity for \(k = 0\), then \(k = 1\), and so on through \(k = K\):
\[\begin{align} \operatorname{E}\mathopen{}\left[Y^{\bar{a}}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_0 = \bar{a}_0\right]\mathclose{} && \text{(one-step identity, } k = 0\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_1 = \bar{a}_1\right]\mathclose{} && \text{(one-step identity, } k = 1\text{)} \\ &\;\;\vdots \\ &= \operatorname{E}\mathopen{}\left[Y^{\bar{a}} \mid \bar{A}_K = \bar{a}_K\right]\mathclose{} && \text{(one-step identity, } k = K\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid \bar{A} = \bar{a}\right]\mathclose{} && \text{(consistency: } Y^{\bar{a}} = Y \text{ when } \bar{A} = \bar{a}\text{)} \end{align}\]
Example 9 (An unrecorded clinic visit) Figure 19.4 includes an unmeasured \(W_0\) that causes both \(A_0\) and the measured CD4 count \(L_1\), e.g. an unrecorded scheduled clinic visit at time 0; \(U_1\) is the true (unknown) CD4 count, which \(L_1\) measures with error.
Remark 6 (Three identifiability conditions). Identification also needs sequential versions of positivity and consistency. Both are expected to hold in a sequentially randomized experiment. Under all three conditions, \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) is identified by methods that adjust appropriately for \((\bar{A}_{k-1}, \bar{L}_k)\): the g-formula (standardization), IP weighting, and g-estimation.
Technical Point 19.2: Positivity and Consistency for Time-Varying Treatments
Sequential positivity:
\[\text{if } f_{\bar{A}_{k-1}, \bar{L}_k}(\bar{a}_{k-1}, \bar{l}_k) \neq 0, \text{ then } f_{A_k \mid \bar{A}_{k-1}, \bar{L}_k}(a_k \mid \bar{a}_{k-1}, \bar{l}_k) > 0 \text{ for all } \bar{a}_k, \bar{l}_k.\]
It holds in a sequentially randomized experiment (Definition 5) if the randomization probabilities are never 0 or 1, whatever the past history. For a particular strategy \(g\), it need hold only for treatment histories compatible with \(g\), i.e., \(a_k = g(\bar{a}_{k-1}, \bar{l}_k)\) for each \(k\).
Sequential consistency:
\[Y^{\bar{a}} = Y^{\bar{a}^*} \text{ if } \bar{a}^* = \bar{a}; \quad Y^{\bar{a}} = Y \text{ if } \bar{A} = \bar{a}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k^{\bar{a}^*} \text{ if } \bar{a}^*_{k-1} = \bar{a}_{k-1}; \quad \bar{L}_k^{\bar{a}} = \bar{L}_k \text{ if } \bar{A}_{k-1} = \bar{a}_{k-1},\]
where \(\bar{L}_k^{\bar{a}}\) is the counterfactual covariate history through \(k\) under \(\bar{a}\). Weaker conditions suffice for identification (“if \(\bar{A} = \bar{a}\) then \(Y^{\bar{a}} = Y\)” for static strategies; “if \(A_k = g_k(\bar{A}_{k-1}, \bar{L}_k)\) at each \(k\) then \(Y^g = Y\)” for dynamic ones), but the book always accepts the stronger sequential version. If “treat in month \(k\)” and “do not treat in month \(k\)” are sufficiently well defined at all \(k\), so are all static and dynamic strategies built from them.
Remark 7 (A generalized backdoor criterion). The backdoor criterion of Chapter 7 generalizes to time-varying treatments. For static strategies, a sufficient condition for identification (given sequential positivity and consistency) is that, for every \(k\), conditioning on \((\bar{A}_{k-1}, \bar{L}_k)\) blocks every backdoor path between \(A_k\) and \(Y\) that does not pass through a later treatment.
But this generalized criterion works on causal DAGs, which contain no counterfactual outcomes, so it does not show the connection to sequential exchangeability directly. SWIGs (Chapter 7) do, and are especially useful here.
Example 10 (The diagrams of Figures 19.5 and 19.6) Figures 19.5 and 19.6 simplify Figures 19.2 and 19.4: they drop \(U_0\), \(L_0\), the arrow \(A_0 \to U_1\), and the arrow \(L_1 \to Y\). Write \(\text{pa}(V)\) for the set of parents of a node \(V\). The treatment nodes have these complete parent sets:
In both diagrams, no arrow points from any \(U\) or \(W\) into \(A_1\). The covariate \(L_1\) has parents including \(A_0\) and \(U_1\) (and also \(W_0\) in Figure 19.6, so \(W_0\) is a shared cause of \(A_0\) and \(L_1\)). The outcome \(Y\) has parents including \(U_1\) and \(A_1\), but not \(L_1\).
Algorithm 1 (Building a SWIG for a static strategy) Given a causal DAG and a static strategy \(\bar{a} = (a_0, \ldots, a_K)\):
Example 11 (Figure 19.7: the SWIG for Figure 19.5) Applying alg. 1 to Figure 19.5 (Example 10) under the static strategy \((a_0, a_1)\), e.g. “always treat” \((1, 1)\) or “never treat” \((0, 0)\), gives Figure 19.7:
\(L_1\) is written \(L_1^{a_0}\), not \(L_1^{a_0, a_1}\), because a later intervention on \(A_1\) cannot affect the earlier \(L_1\). Unlike the DAG, the SWIG contains the counterfactual outcome, so exchangeability can be checked by d-separation.
Example 12 (Reading exchangeability off the SWIG of Figure 19.7) In Figure 19.7, d-separation gives, for every \((a_0, a_1)\),
\[Y^{a_0, a_1} \perp\!\!\!\perp A_0 \quad\text{and}\quad Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}.\]
The path \(A_1^{a_0} \leftarrow a_0 \to L_1^{a_0} \leftarrow U_1 \to Y^{a_0, a_1}\) looks open, but it is blocked: \(a_0\) is a constant in the counterfactual world, and a path through a constant intervention node is always blocked, because a constant is implicitly conditioned on.
Proposition 3 (From a SWIG independence to an observed-data independence) Fix a static strategy \((a_0, a_1)\) with \(\Pr[A_0 = a_0] > 0\). If \(Y^{a_0, a_1} \perp\!\!\!\perp A_1^{a_0} \mid A_0, L_1^{a_0}\) and consistency holds (\(L_1^{a_0} = L_1\) and \(A_1^{a_0} = A_1\) whenever \(A_0 = a_0\)), then \(Y^{a_0, a_1} \perp\!\!\!\perp A_1 \mid A_0 = a_0, L_1\).
Proof (From the SWIG statement to an observed-data statement).
Definition 8 (Static sequential exchangeability) Static sequential exchangeability holds for a given static strategy \(\bar{a}\) if
\[Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \bar{L}_k \quad \text{for } k = 0, 1, \ldots, K.\]
Remark 8 (Static sequential exchangeability is weaker).
Technical Point 19.3: The Many Forms of Sequential Exchangeability
For a sequentially randomized experiment with times \(k = 0, 1, \ldots, K\) (a longer version of Figure 19.7), the SWIG gives
\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k^{\bar{a}_{k-1}} \mid \bar{A}_{k-1}^{\bar{a}_{k-2}}, \bar{L}_k^{\bar{a}_{k-1}},\]
where \(\underline{L}_{k+1}^{\bar{a}}\) is the counterfactual covariate history from \(k + 1\) to the end of follow-up. Restricting to \(\bar{A}_{k-1}^{\bar{a}_{k-2}} = \bar{a}_{k-1}\) and applying consistency gives
\[\mathopen{}\left(Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}, \bar{L}_k.\]
When this holds for all \(\bar{a}\), the book calls it sequential exchangeability; these notes call it joint sequential exchangeability, to keep it distinct from Definition 6, which it implies. Although it mentions only static strategies, it is equivalent to
\[\mathopen{}\left(Y^{g}, \underline{L}_{k+1}^{g}\right)\mathclose{} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = g(\bar{A}_{k-2}, \bar{L}_{k-1}), \bar{L}_k \quad \text{for all } g,\]
and so, with positivity and consistency, it identifies the outcome and covariate distributions under every static and dynamic strategy (Robins 1986). This rests on the joint independence of \((Y^{\bar{a}}, \underline{L}_{k+1}^{\bar{a}})\) and \(A_k\); for dynamic strategies, the two separate independences of \(Y^{\bar{a}}\) and of \(\underline{L}_{k+1}^{\bar{a}}\) from \(A_k\) do not suffice.
A sequentially randomized trial implies still stronger conditions, which cannot be read from SWIGs and are not needed for identification, e.g.
Figure 19.9 represents Figure 19.5 under a dynamic strategy \(g = [g_0, g_1(L_1)]\): \(A_0\) is set to a fixed \(g_0\), and \(A_1\) to \(g_1(L_1^g)\), which depends on the \(L_1^g\) observed after setting \(A_0 = g_0\).
Example 14 (“Treat at time 1 only if CD4 is low”) Strategy: do not treat at time 0; at time 1 treat only if the CD4 count is low (\(L_1^g = 1\)). So \(g_0 = 0\) for everybody, \(g_1(L_1^g) = 1\) when \(L_1^g = 1\), and \(g_1(L_1^g) = 0\) when \(L_1^g = 0\).
Remark 9 (The SWIG of a dynamic strategy).
Example 15 (Dynamic strategies under Figures 19.5 and 19.6)
Fine Point 19.2: Arrows From Intervention Nodes in SWIGs
SWIGs include arrows from intervention nodes such as \(a\) into later variables, even though in a single intervention world \(a\) is a constant that cannot affect anything.
So when applying d-separation to a SWIG, every path through an intervention node is blocked, even though the node is not written in the conditioning event.
The same holds for deterministic strategies that depend on a baseline confounder \(L_0\) (replace \(g_0\) by \(g_0(L_0)\) in Figure 19.9): given \(L_0\), \(g_0(L_0)\) is a constant, so paths through it are blocked, and \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, L_0, L_1^g\) becomes, after setting \(A_0 = g(L_0)\) and using consistency, \(Y^g \perp\!\!\!\perp A_1 \mid A_0 = g(L_0), L_0, L_1\).
For a random strategy that draws \(A_0^{+,g}\) from a distribution that may depend on \(L_0\), \(A_0^{+,g}\) appears on the SWIG and must be conditioned on explicitly. The book reports that Richardson and Robins (2013) showed that \(Y^g \perp\!\!\!\perp A_1^g \mid A_0, A_0^{+,g}, L_0, L_1^g\) is a necessary condition for identification by the g-formula for such a strategy (Hernán and Robins 2020, Fine Point 19.2). They also treated strategies that depend on the natural value of treatment \(A_t^g\) (Robins et al. 2004), recently called “modified treatment policies” (Diaz et al. 2021).
Example 16 (Adding an arrow from \(L_1\) to \(Y\)) Figure 19.11 is Figure 19.6 (Example 10) plus an arrow \(L_1 \to Y\) (Figure 19.12 is its SWIG under a static strategy). Here neither sequential exchangeability for \(Y^g\) (Definition 6) nor static sequential exchangeability for \(Y^{\bar{a}}\) (Definition 8) holds (generically, that is, under faithfulness). The book states that then no strategy’s effect is identified (Hernán and Robins 2020, sec. 19.5).
| Diagram | Static strategies | Dynamic strategies depending on \(L_1\) |
|---|---|---|
| Figure 19.5 | identified | identified |
| Figure 19.6 | identified | not identified |
| Figure 19.11 | not identified | not identified |
Sequential Exchangeability Is Never Guaranteed
No form of sequential exchangeability is guaranteed in an observational study. Approximating it requires expert knowledge to decide which time-varying variables \(\bar{L}_k\) to measure (in HIV: CD4 count, viral load, symptoms). Whether the measured covariates suffice can never be known with certainty, but causal diagrams such as Figures 19.1-19.4 organize our beliefs.
These diagrams were drawn without selection (e.g., censoring) so as to focus on confounding.
Example 17 (\(L_1\) as a confounder for the effect of \(A_1\)) In Figure 19.5 (Example 10), consider the effect of \(A_1\) alone on \(Y\).
Now consider the long diagram over all times \(k = 0, 1, 2, \ldots\), in which each \(L_k\) affects later treatments \(A_k, A_{k+1}, \ldots\) and shares unmeasured causes \(U_k\) with \(Y\). To estimate the effects of strategies defined by interventions on \(A_0, A_1, A_2, \ldots\), we must condition at each \(k\) on the covariate history \(\bar{L}_k\), together with \(\bar{A}_{k-1}\), to block the backdoor paths between \(A_k\) and \(Y\).
Definition 9 (Time-varying confounders) Covariates \(\bar{L}_k\) that, together with \(\bar{A}_{k-1}\), are needed at each time \(k\) to block the backdoor paths between \(A_k\) and \(Y\) are time-varying confounders for the effect of \(\bar{A}\) on \(Y\).
Measuring the Confounders Is Not Enough
Fine Point 19.3: A Definition of Time-Varying Confounding
Assume no selection bias.
A sufficient condition for no time-varying confounding is unconditional sequential exchangeability (Definition 7), \(Y^{\bar{a}} \perp\!\!\!\perp A_k \mid \bar{A}_{k-1} = \bar{a}_{k-1}\), as in Figure 19.1. That diagram can in fact be pruned to \(A_0\), \(A_1\), and \(Y\): \(L_1\) is not a common cause of two nodes, so drop it; then \(L_0\) and \(U_1\) are no longer common causes, so drop them; then \(U_0\) is not either.