Chapter 19 identified sequential exchangeability as a key condition for identifying the causal effects of time-varying treatments. Suppose a study satisfies the strongest form of sequential exchangeability, so the measured time-varying confounders suffice to estimate the effect of any treatment strategy. Which confounding adjustment method should we then use? The answer exposes a central problem for time-varying treatments: a time-varying confounder can drive later treatment while itself being affected by earlier treatment, or while sharing with earlier treatment a cause outside that treatment’s measured past, through a path that conditioning on that measured past does not block. When that happens, traditional adjustment methods can be biased even though all the information needed for valid estimation is available. This chapter describes the structure of the problem and explains why traditional methods fail.
Definition 1 (Treatment-Confounder Feedback) Consider a time-varying treatment \(A_k\) and a time-varying confounder \(L_k\), measured at times \(k = 0, 1, \ldots\). For \(j \ge 0\), write \((\bar{L}_j, \bar{A}_{j-1})\) for the measured past of treatment \(A_j\). There is treatment-confounder feedback when, for some \(k \ge 1\), the time-varying confounder affects treatment (\(L_k \to A_k\)) and, for some earlier treatment \(A_j\) with \(j < k\), either
The book’s one-line description is the effect form (a directed path \(A_j \to \cdots \to L_k\) for some \(j < k\)): “the confounder affects the treatment and the treatment affects the confounder” (Hernán and Robins 2020, 267). The book gives the same name to the case in which \(L_1\) shares an unmeasured cause \(W_0\) with \(A_0\) (its Figures 20.4 and 20.6), because conditioning on \(L_1\) produces the same bias in that case as in the effect form; the observed data also cannot tell the two cases apart (Hernán and Robins 2020, 269 and 272). The second clause of Definition 1 covers that case only. It excludes a common cause in the measured past, such as an earlier confounder \(L_{k-1}\) that affects both \(A_{k-1}\) and \(L_k\); otherwise almost every persistent confounder would count as feedback.
Example 1 (The Sequentially Randomized HIV Trial) Return to the sequentially randomized HIV trial of Chapter 19:
The time-varying covariates \(L_k\) are time-varying confounders: CD4 count \(L_k\) affects treatment \(A_k\). In addition, treatment \(A_{k-1}\) raises future CD4 count \(L_k\). So the trial has treatment-confounder feedback (Definition 1).
Remark 1 (Time-Varying Confounding Without Feedback). Time-varying confounding can occur without treatment-confounder feedback (Definition 1). In the book’s Figure 20.2, the arrows from treatment to later \(L\) and \(U\) variables are deleted: \(L_k\) still confounds the effect of \(A_k\), but no earlier treatment \(A_j\) (\(j < k\)) affects it. Nor does \(L_k\) share with any earlier \(A_j\) a common cause outside the measured past of \(A_j\) in the sense of Definition 1: the diagram is a sequentially randomized trial, so the unmeasured \(U\) variables have no arrows into treatment, and every path from a common cause to \(A_j\) passes through the measured confounders \(\bar{L}_j\), which block it. There is time-varying confounding but no feedback.
Fine Point 20.1: Representing Feedback Cycles with Acyclic Graphs
A feedback loop between treatment and confounder can be drawn on an acyclic graph by unrolling it in time: \(A_{k-1} \to L_k \to A_k \to L_{k+1} \to \cdots\). This representation requires discrete time: treatment and covariates may change during each interval \([k, k+1)\), without specifying when within the interval. In practice the intervals can be made as short as the data’s granularity requires (months in the HIV example, where patients see their doctors at most monthly), and time is usually recorded in discrete units anyway (Hernán and Robins 2020, Fine Point 20.1, p. 268).
Example 2 (The Simplest Feedback Diagram) The smallest diagram that shows treatment-confounder feedback in a two-time-point sequentially randomized trial is the book’s Figure 20.3. Its arrows are
\[A_0 \to L_1, \qquad L_1 \to A_1, \qquad U_1 \to L_1, \qquad U_1 \to Y.\]
It makes four simplifications relative to Figure 20.1:
Example 3 (Feedback Through a Shared Cause) The same problem arises when the time-varying confounder is not affected by prior treatment but shares an unmeasured cause \(W_0\) with it (the book’s Figure 20.4, a subset of Figure 19.4). Its arrows are
\[W_0 \to A_0, \qquad W_0 \to L_1, \qquad L_1 \to A_1, \qquad U_1 \to L_1, \qquad U_1 \to Y.\]
Figure 20.4 represents an observational study. The path \(A_0 \leftarrow W_0 \to L_1\) runs through a cause outside the measured past of \(A_0\), so this is the second clause of Definition 1. The book refers to both Figures 20.3 and 20.4 (and 19.2 and 19.4) as examples of treatment-confounder feedback (Definition 1).
Remark 3 (The Claim of This Chapter). With time-varying confounders and treatment-confounder feedback, traditional methods (stratification, matching, and outcome regression whose treatment coefficients are read as the effect) cannot correctly adjust for the confounders, even when the data suffice for sequential exchangeability. G-methods adjust correctly even in the presence of feedback.
Example 4 (A Hypothetical Two-Time-Point Trial) A hypothetical sequentially randomized trial with 32,000 individuals with HIV and two time points \(k = 0, 1\) (full adherence, no loss to follow-up):
The data are in Table 1.
| \(N\) | \(A_0\) | \(L_1\) | \(A_1\) | Mean \(Y\) |
|---|---|---|---|---|
| 2400 | 0 | 0 | 0 | 84 |
| 1600 | 0 | 0 | 1 | 84 |
| 2400 | 0 | 1 | 0 | 52 |
| 9600 | 0 | 1 | 1 | 52 |
| 4800 | 1 | 0 | 0 | 76 |
| 3200 | 1 | 0 | 1 | 76 |
| 1600 | 1 | 1 | 0 | 44 |
| 6400 | 1 | 1 | 1 | 44 |
Remark 4 (Identifiability Holds by Design). The trial is built so that all three identifiability conditions hold (Hernán and Robins 2020, 269):
Had the trial run past month 1, with each month’s treatment assigned using that month’s CD4 count, each of those later CD4 measurements would be a time-varying confounder as well. Random variability is ignored throughout.
Example 6 (No Effect of \(A_1\)) In every stratum of \((A_0, L_1)\), the mean outcome is the same for \(A_1 = 1\) and \(A_1 = 0\) (rows 1 vs 2, 3 vs 4, 5 vs 6, 7 vs 8). For example,
\[\operatorname{E}\mathopen{}\left[Y \mid A_0 = 0, L_1 = 0, A_1 = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0, L_1 = 0, A_1 = 0\right]\mathclose{} = 84 - 84 = 0,\]
and because the identifiability conditions hold this equals \(\operatorname{E}\mathopen{}\left[Y^{a_1=1} \mid A_0 = 0, L_1 = 0\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_1=0} \mid A_0 = 0, L_1 = 0\right]\mathclose{}\).
Example 7 (No Effect of \(A_0\)) Because \(A_0\) is randomized (so exchangeability holds for \(A_0\) without adjustment), and positivity and consistency also hold (Remark 4), \(\operatorname{E}\mathopen{}\left[Y \mid A_0 = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0\right]\mathclose{}\) estimates \(\operatorname{E}\mathopen{}\left[Y^{a_0=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_0=0}\right]\mathclose{}\). Averaging rows 1-4 (16,000 individuals with \(A_0 = 0\)):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0\right]\mathclose{} &= \frac{2400 \times 84 + 1600 \times 84 + 2400 \times 52 + 9600 \times 52}{16000} \\ &= \frac{4000 \times 84 + 12000 \times 52}{16000} = \frac{336000 + 624000}{16000} = 60. \end{aligned} \]
Averaging rows 5-8 (16,000 individuals with \(A_0 = 1\)):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y \mid A_0 = 1\right]\mathclose{} &= \frac{4800 \times 76 + 3200 \times 76 + 1600 \times 44 + 6400 \times 44}{16000} \\ &= \frac{8000 \times 76 + 8000 \times 44}{16000} = \frac{608000 + 352000}{16000} = 60. \end{aligned} \]
So the average causal effect of \(A_0\) is \(60 - 60 = 0\).
Technical Point 20.1: G-Null Test
The g-null theorem (Robins 1986) ties the sharp null to two conditional independencies in the observed data. With two time points, the independencies are
\[ Y \perp\!\!\!\perp A_0 \mid L_0 \quad \text{and} \quad Y \perp\!\!\!\perp A_1 \mid A_0, L_0, L_1. \]
Assume sequential randomization for every strategy \(g\). The theorem says that both independencies hold exactly when the distribution of \(Y^g\) does not depend on \(g\) and coincides with the distribution of the observed \(Y\), so that in particular \(\operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\) for every \(g\) (Hernán and Robins 2020, Technical Point 20.1, p. 270).
To see why the sharp null leads to the independencies, note that if no strategy changes anyone’s outcome, then \(Y^g = Y\) for every \(g\). Sequential exchangeability for the counterfactuals \(Y^g\) then becomes a statement about \(Y\) itself. That step relies on every pair \((a_0, l_0)\) being produced by some strategy with \(a_0 = g(l_0)\).
Each independency has a causal reading:
When sequential exchangeability holds, checking these independencies in the data therefore tests the sharp null. That check is the g-null test (Robins 1986).
The effects of \(A_0\) and \(A_1\) are each null. Each null taken alone does not establish the joint null for the strategy \((a_0, a_1)\). The g-null theorem of Technical Point 20.1 does not apply directly, because its premise is conditional independence and the examples check only conditional means. The mean joint null says that \(\operatorname{E}\mathopen{}\left[Y^{a_0, a_1}\right]\mathclose{}\) is the same for every static strategy. The equal conditional means of Example 6 and Example 7 are enough for it, as computing the mean outcome under each static strategy with the g-formula shows. That computation uses all three identifiability conditions (Remark 4): sequential exchangeability (unconditionally for the randomized \(A_0\), and given \((A_0, L_1)\) for \(A_1\)), positivity, and consistency. Under these conditions,
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a_0, a_1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a_0, a_1} \mid A_0 = a_0\right]\mathclose{} && \text{(} A_0 \text{ randomized)} \\ &= \sum_{l} \operatorname{E}\mathopen{}\left[Y^{a_0, a_1} \mid A_0 = a_0, L_1 = l\right]\mathclose{} \Pr[L_1 = l \mid A_0 = a_0] && \text{(law of total expectation)} \\ &= \sum_{l} \operatorname{E}\mathopen{}\left[Y^{a_0, a_1} \mid A_0 = a_0, L_1 = l, A_1 = a_1\right]\mathclose{} \Pr[L_1 = l \mid A_0 = a_0] && \text{(sequential exchangeability)} \\ &= \sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A_0 = a_0, L_1 = l, A_1 = a_1\right]\mathclose{} \Pr[L_1 = l \mid A_0 = a_0] && \text{(consistency)}. \end{aligned} \]
From Table 1, \(\Pr[L_1 = 0 \mid A_0 = 0] = 4000/16000 = 0.25\) and \(\Pr[L_1 = 0 \mid A_0 = 1] = 8000/16000 = 0.5\), and the cell means do not depend on \(a_1\), so for either value of \(a_1\)
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a_0=0, a_1}\right]\mathclose{} &= 84 \times 0.25 + 52 \times 0.75 = 21 + 39 = 60, \\ \operatorname{E}\mathopen{}\left[Y^{a_0=1, a_1}\right]\mathclose{} &= 76 \times 0.5 + 44 \times 0.5 = 38 + 22 = 60. \end{aligned} \]
So \(\operatorname{E}\mathopen{}\left[Y^{a_0, a_1}\right]\mathclose{} = 60\) for every \((a_0, a_1)\), and in particular \(\operatorname{E}\mathopen{}\left[Y^{a_0=1, a_1=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_0=0, a_1=0}\right]\mathclose{} = 0\). Do conventional analyses recover this?
Example 8 (Analysis 1: No Adjustment for \(L_1\)) Compare the 9600 individuals treated at both times (rows 6 and 8) with the 4800 untreated at both times (rows 1 and 3):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y \mid A_0 = 1, A_1 = 1\right]\mathclose{} &= \frac{3200 \times 76 + 6400 \times 44}{9600} = \frac{243200 + 281600}{9600} \approx 54.7, \\ \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0, A_1 = 0\right]\mathclose{} &= \frac{2400 \times 84 + 2400 \times 52}{4800} = \frac{201600 + 124800}{4800} = 68.0. \end{aligned} \]
The difference \(54.7 - 68 = -13.3\) is non-null, because \(\operatorname{E}\mathopen{}\left[Y \mid A_0 = a_0, A_1 = a_1\right]\mathclose{}\) is not a valid estimator of \(\operatorname{E}\mathopen{}\left[Y^{a_0, a_1}\right]\mathclose{}\): adjustment for the confounder \(L_1\) is needed.
Example 9 (Analysis 2: Stratification on \(L_1\)) Within \(L_1 = 0\), compare treated at both times (row 6) with untreated at both times (row 1):
\[\operatorname{E}\mathopen{}\left[Y \mid A_0 = 1, L_1 = 0, A_1 = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0, L_1 = 0, A_1 = 0\right]\mathclose{} = 76 - 84 = -8.\]
Within \(L_1 = 1\) (rows 8 and 3): \(44 - 52 = -8\).
The analysis adjusted for the confounder also gives the wrong answer.
Remark 6 (The Problem Is the Adjustment Method). All three identifiability conditions hold in Table 1; the unmeasured \(U_1\) (immunosuppression level) is not needed because we have data on \(L_1\). Yet neither analysis gave the correct answer. The problem is the adjustment method: stratification cannot handle treatment-confounder feedback (Definition 1).
Remark 7 (Stratification Opens a Collider Path). Stratifying on \(L_1\) means estimating the treatment-outcome association separately in the subsets \(L_1 = 0\) (high CD4) and \(L_1 = 1\) (low CD4). But \(L_1\) is affected by prior treatment \(A_0\), so it is a collider on the path
\[A_0 \to L_1 \leftarrow U_1 \to Y\]
(the book’s Figure 20.5: Figure 20.3 with \(L_1\) conditioned on). Conditioning on \(L_1\) opens this path and generically (that is, under faithfulness) induces a noncausal association between \(A_0\) and \(U_1\), and hence between \(A_0\) and \(Y\), within levels of \(L_1\).
Remark 8 (Trading Confounding for Selection Bias). Stratification eliminates confounding for \(A_1\) at the cost of introducing selection bias for \(A_0\). The associational differences
\[\operatorname{E}\mathopen{}\left[Y \mid A_0 = 1, L_1 = l, A_1 = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A_0 = 0, L_1 = l, A_1 = 0\right]\mathclose{}\]
may be non-zero even when treatment has no effect on anyone’s outcome at any time. The net bias depends on the relative sizes of the confounding removed and the selection bias created.
Remark 9 (The Same Bias Under a Shared Cause). The same bias arises when the confounder shares an unmeasured cause \(W_0\) with prior treatment (an observational study, the book’s Figure 20.6): conditioning on the collider \(L_1\) opens
\[A_0 \leftarrow W_0 \to L_1 \leftarrow U_1 \to Y.\]
Because conditioning on \(L_1\) creates the same bias in Figure 20.4 as in Figure 20.3, both settings are called treatment-confounder feedback; the observed data also cannot tell them apart.
The Bias Is Not Limited to the Null or to Two Time Points
Fine Point 20.2: Confounders on the Causal Pathway
Conditioning on a confounder \(L_1\) affected by prior treatment can create selection bias even when \(L_1\) is not on a causal pathway from treatment to outcome (no such pathway exists in Figures 20.5 and 20.6).
In Figure 20.7, by contrast, \(L_1\) is a confounder for \(A_1\) and lies on the causal pathway \(A_0 \to L_1 \to Y\). Even if \(U_1\) were not a common cause of \(L_1\) and \(Y\) (no selection bias), the \(A\)-\(Y\) associations within strata of \(L_1\) would estimate only the direct effect of \(A_0\) not through \(L_1\), not the overall effect of \(\bar{A}\) on \(Y\).
The common rule that variables on a causal pathway cannot be confounders is inaccurate for time-varying treatments: a confounder for later treatment \(A_1\) can lie on a pathway from earlier treatment \(A_0\) to \(Y\). Whether adjusting for it induces bias depends on the method: stratification does; g-methods do not (Hernán and Robins 2020, Fine Point 20.2, p. 272).
Could parametric outcome regression succeed where nonparametric stratification failed? The question matters most with high-dimensional data, where a simple stratified analysis is impossible.
Example 11 (Counting Static Strategies)
Remark 10 (Many Strategies Require Modeling). As argued since Chapter 11, many possible strategies require modeling: a dose-response function for the effect of treatment history \(\bar{a}\) on the mean outcome. For example, assume the effect is linear in cumulative treatment, so every strategy with exactly three months of treatment has the same effect, whenever those three months occur. The price is a new threat to validity: misspecification of the dose-response model.
Conventional Outcome Regression Does Not Remove the Bias
Paying the price of a dose-response model buys no protection if we read the treatment coefficients of an outcome regression as the effect while the model conditions on the time-varying confounder. That conventional use of regression is a stratification-based method, so it cannot remove the bias of stratification under treatment-confounder feedback (Definition 1). The g-formula (Chapter 21) can fit the same kind of outcome model, but it averages the model’s predictions over the distribution of each time-varying confounder given past treatment and past confounders (in this chapter’s two-time-point example, which has no \(L_0\): \(L_1\) given \(A_0\)) instead of reading the effect off the coefficients.
Example 12 (A Cumulative-Treatment Regression) For data generated under Figure 20.5, define \(\text{cum}(\bar{A}) = A_0 + A_1 \in \{0, 1, 2\}\). “Always treat” is \(\text{cum}(\bar{a}) = 2\) and “never treat” is \(\text{cum}(\bar{a}) = 0\); the target is \(\operatorname{E}\mathopen{}\left[Y^{\text{cum}(\bar{a})=2}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{\text{cum}(\bar{a})=0}\right]\mathclose{}\), whose true value is 0.
Fit the outcome regression model
\[\operatorname{E}\mathopen{}\left[Y \mid \bar{A}, L_1\right]\mathclose{} = \theta_0 + \theta_1 \, \text{cum}(\bar{A}) + \theta_2 L_1.\]
Within either level \(l\) of \(L_1\),
\[ \begin{aligned} &\operatorname{E}\mathopen{}\left[Y \mid \text{cum}(\bar{A}) = 2, L_1 = l\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid \text{cum}(\bar{A}) = 0, L_1 = l\right]\mathclose{} \\ &\quad = (\theta_0 + 2\theta_1 + \theta_2 l) - (\theta_0 + 0 \cdot \theta_1 + \theta_2 l) = 2\theta_1 . \end{aligned} \]
Do Not Read \(2\theta_1\) as the Causal Effect
It is tempting to read \(2\theta_1\) as the effect of “always treat” versus “never treat” within levels of \(L_1\). But conditioning on \(L_1\) generically induces an association between \(A_0\) (a component of \(\text{cum}(\bar{A})\)) and \(Y\), so, generically, \(\theta_1 \neq 0\) even though the true effect is zero (Hernán and Robins 2020, 274). The bias comes from conditioning on the collider \(L_1\), not from the form of the model or from the sample size: it would persist with unlimited data.
Remark 11 (Past Treatment Opens New Backdoor Paths). So far the diagrams had no arrow \(A_0 \to A_1\). Now suppose doctors use past treatment history \(\bar{A}_{k-1}\) when deciding on treatment \(A_k\). Adding \(A_0 \to A_1\) to Figures 20.3 and 20.4 gives the book’s Figures 20.8 and 20.9.
Under treatment-confounder feedback (Definition 1), adjusting for \(L_1\) alone no longer closes every backdoor path from \(A_1\) to \(Y\). \(L_1\) is a collider on each of the following paths, so conditioning on it opens them:
\[ \begin{aligned} &A_1 \leftarrow A_0 \to L_1 \leftarrow U_1 \to Y && \text{(Figure 20.8)}, \\ &A_1 \leftarrow A_0 \leftarrow W_0 \to L_1 \leftarrow U_1 \to Y && \text{(Figure 20.9)}. \end{aligned} \]
And whenever past treatment affects the outcome (Figure 20.10), \(A_0\) is a confounder for the effect of \(A_1\), feedback or not.
Remark 12 (Sequential Exchangeability Conditions on Treatment History). Sequential exchangeability at time \(k\) generally requires conditioning on the treatment history \(\bar{A}_{k-1}\) as well as on the covariates; conditioning only on \(L\) is not enough. That is why every sequential exchangeability statement in Chapters 19 and 20 conditions on treatment history.
Definition 2 (Short-Term Effect) In a two-time-point treatment \(\bar{A} = (A_0, A_1)\), the short-term effect of \(A_1\) is the contrast \[\operatorname{E}\mathopen{}\left[Y^{a_1=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a_1=0}\right]\mathclose{},\] in which only \(A_1\) is set by intervention and \(A_0\) is not intervened on, so it keeps whatever value it would naturally take. That is, it is the effect of \(A_1\) when \(A_1\) is treated as a time-fixed treatment.
Suppose the target is the short-term effect of \(A_1\) (Definition 2), treating \(A_1\) as a time-fixed treatment.
Ignoring Past Treatment Biases the Short-Term Effect
Without adjustment for \(A_0\):
So \(\operatorname{E}\mathopen{}\left[Y \mid A_1 = 1, L_1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A_1 = 0, L_1\right]\mathclose{}\) need not be zero even if \(A_1\) has no effect on anyone’s outcome (Figures 20.8-20.10).
Mismeasured Past Treatment
The need to adjust for past treatment matters when past treatment is mismeasured. As in Section 9.3, adjusting for a mismeasured confounder can leave bias in either direction.
Example 14 (Self-Reported Treatment History) Suppose HIV investigators lack medical records and ascertain prior treatment by questionnaire, so they observe a mismeasured \(A_0^*\) rather than \(A_0\). Adding the arrow \(A_0 \to A_0^*\) to Figures 20.8-20.10 shows the problem. \(A_0^*\) is only a noisy child of \(A_0\), so conditioning on it leaves open the backdoor paths from \(A_1\) to \(Y\) on which \(A_0\) lies. Investigators would generically (under faithfulness) find an association between \(A_1\) and \(Y\) even after adjusting for \(A_0^*\) and \(L_1\), although \(A_1\) has no effect on \(Y\).