---
title: "Chapter 20: Treatment-Confounder Feedback"
format:
html: default
revealjs:
output-file: 20-treatment-confounder-feedback-slides.html
pdf:
output-file: 20-treatment-confounder-feedback-handout.pdf
docx:
output-file: 20-treatment-confounder-feedback.docx
preview-changed: true
---
{{< include ../latex-macros/macros.qmd >}}
Chapter 19 identified sequential exchangeability as a key condition
for identifying the causal effects of time-varying treatments.
Suppose a study satisfies the strongest form of sequential exchangeability,
so the measured time-varying confounders suffice to estimate the effect of any treatment strategy.
Which confounding adjustment method should we then use?
The answer exposes a central problem for time-varying treatments:
a time-varying confounder can drive later treatment
while itself being affected by earlier treatment,
or while sharing with earlier treatment a cause outside that treatment's measured past,
through a path that conditioning on that measured past does not block.
When that happens, traditional adjustment methods can be biased
even though all the information needed for valid estimation is available.
This chapter describes the structure of the problem
and explains why traditional methods fail.
::: {.notes}
This chapter is based on @hernan2020causal [Chapter 20, pp. 267-275].
**Key insight**: when a time-varying confounder $L_1$ that affects later treatment $A_1$
is also affected by earlier treatment $A_0$,
or shares with $A_0$ a cause outside the measured past of $A_0$
through a path that conditioning on that measured past does not block,
stratifying on $L_1$ removes confounding for the later treatment $A_1$
but opens a collider path that creates selection bias for $A_0$.
Not adjusting for $L_1$ leaves $A_1$ confounded.
Either way the estimate of a treatment-strategy effect is biased;
g-methods (Chapter 21) avoid both problems.
:::
## 20.1 The Elements of Treatment-Confounder Feedback (pp. 267-269)
---
::: {#def-treatment-confounder-feedback}
## Treatment-Confounder Feedback
Consider a time-varying treatment $A_k$ and a time-varying confounder $L_k$, measured at times $k = 0, 1, \ldots$.
For $j \ge 0$, write $(\bar{L}_j, \bar{A}_{j-1})$ for the measured past of treatment $A_j$.
There is **treatment-confounder feedback** when, for some $k \ge 1$,
the time-varying confounder affects treatment ($L_k \to A_k$)
*and*, for some earlier treatment $A_j$ with $j < k$, either
- $L_k$ is affected by $A_j$, or
- $L_k$ and $A_j$ share a common cause outside the measured past of $A_j$,
through a path between them (such as $A_j \leftarrow W \to L_k$ with $W$ unmeasured)
that conditioning on that measured past does not block.
:::
The book's one-line description is the effect form (a directed path $A_j \to \cdots \to L_k$ for some $j < k$):
"the confounder affects the treatment and the treatment affects the confounder"
[@hernan2020causal, p. 267].
The book gives the same name to the case in which $L_1$ shares an unmeasured cause $W_0$ with $A_0$ (its Figures 20.4 and 20.6),
because conditioning on $L_1$ produces the same bias in that case as in the effect form;
the observed data also cannot tell the two cases apart [@hernan2020causal, pp. 269 and 272].
The second clause of @def-treatment-confounder-feedback covers that case only.
It excludes a common cause in the measured past,
such as an earlier confounder $L_{k-1}$ that affects both $A_{k-1}$ and $L_k$;
otherwise almost every persistent confounder would count as feedback.
::: {#exm-hiv-trial-time-varying}
## The Sequentially Randomized HIV Trial
Return to the sequentially randomized HIV trial of Chapter 19:
- treatment $A_k$ (1: treated, 0: untreated) at each month $k = 0, 1, \ldots, K$;
- covariates $L_k$ (CD4 cell count) at each month;
- an outcome $Y$ measuring health status at month $K + 1$.
The time-varying covariates $L_k$ are time-varying confounders:
CD4 count $L_k$ affects treatment $A_k$.
In addition, treatment $A_{k-1}$ raises future CD4 count $L_k$.
So the trial has treatment-confounder feedback (@def-treatment-confounder-feedback).
:::
::: {.notes}
The book's Figure 20.1 (the same diagram as Figure 19.2) shows the first two months of the trial.
As in Chapter 19, the example has no censoring, so we can focus on confounding.
:::
---
### Time-Varying Confounding Without Feedback
::: {#rem-confounding-without-feedback}
## Time-Varying Confounding Without Feedback
Time-varying confounding can occur *without* treatment-confounder feedback (@def-treatment-confounder-feedback).
In the book's Figure 20.2, the arrows from treatment to later $L$ and $U$ variables are deleted:
$L_k$ still confounds the effect of $A_k$,
but no earlier treatment $A_j$ ($j < k$) affects it.
Nor does $L_k$ share with any earlier $A_j$ a common cause outside the measured past of $A_j$
in the sense of @def-treatment-confounder-feedback:
the diagram is a sequentially randomized trial, so the unmeasured $U$ variables have no arrows into treatment,
and every path from a common cause to $A_j$ passes through the measured confounders $\bar{L}_j$, which block it.
There is time-varying confounding but no feedback.
:::
::: {.callout-note title="Fine Point 20.1: Representing Feedback Cycles with Acyclic Graphs"}
A feedback loop between treatment and confounder can be drawn on an acyclic graph
by unrolling it in time:
$A_{k-1} \to L_k \to A_k \to L_{k+1} \to \cdots$.
This representation requires discrete time:
treatment and covariates may change during each interval $[k, k+1)$,
without specifying when within the interval.
In practice the intervals can be made as short as the data's granularity requires
(months in the HIV example, where patients see their doctors at most monthly),
and time is usually recorded in discrete units anyway [@hernan2020causal, Fine Point 20.1, p. 268].
:::
---
### The Simplest Case: Figure 20.3
::: {#exm-minimal-feedback-diagram}
## The Simplest Feedback Diagram
The smallest diagram that shows treatment-confounder feedback in a two-time-point
sequentially randomized trial is the book's Figure 20.3.
Its arrows are
$$A_0 \to L_1, \qquad L_1 \to A_1, \qquad U_1 \to L_1, \qquad U_1 \to Y.$$
It makes four simplifications relative to Figure 20.1:
- no baseline node $L_0$: $A_0$ is marginally randomized and $A_1$ is randomized conditional on $L_1$;
- no unmeasured $U_0$;
- no arrow $A_0 \to A_1$: treatment at time 1 is assigned using $L_1$ only;
- no arrows from $A_0$, $L_1$, or $A_1$ into $Y$: the [sharp null](01-introduction.qmd#def-sharp-null) holds.
:::
::: {.notes}
None of these simplifications affects the arguments of the chapter;
a larger diagram would add no conceptual insight and would only be harder to read.
::: {#rem-sharp-null-figure-20-3}
## The True Effect Under Figure 20.3 Is Zero
Because Figure 20.3 has no forward-directed path from $A_0$ or $A_1$ to $Y$,
the true average causal effect of "always treat" versus "never treat" is zero:
$$\E{Y^{a_0=1, a_1=1}} - \E{Y^{a_0=0, a_1=0}} = 0.$$
There are no arrows from the unmeasured $U_1$ into the treatments,
so Figure 20.3 can represent a sequentially randomized trial
(or an observational study with no unmeasured confounding),
and Chapter 19's results say the observed data on $(A_0, L_1, A_1, Y)$
should identify this effect as 0.
:::
:::
---
### Feedback Through Shared Causes: Figure 20.4
::: {#exm-shared-cause-feedback}
## Feedback Through a Shared Cause
The same problem arises when the time-varying confounder is not affected by prior treatment
but **shares an unmeasured cause** $W_0$ with it (the book's Figure 20.4, a subset of Figure 19.4).
Its arrows are
$$W_0 \to A_0, \qquad W_0 \to L_1, \qquad L_1 \to A_1, \qquad U_1 \to L_1, \qquad U_1 \to Y.$$
Figure 20.4 represents an observational study.
The path $A_0 \leftarrow W_0 \to L_1$ runs through a cause outside the measured past of $A_0$,
so this is the second clause of @def-treatment-confounder-feedback.
The book refers to both Figures 20.3 and 20.4 (and 19.2 and 19.4)
as examples of treatment-confounder feedback (@def-treatment-confounder-feedback).
:::
::: {#rem-traditional-vs-g}
## The Claim of This Chapter
With time-varying confounders and treatment-confounder feedback,
traditional methods (stratification, matching,
and outcome regression whose treatment coefficients are read as the effect)
cannot correctly adjust for the confounders,
even when the data suffice for sequential exchangeability.
G-methods adjust correctly even in the presence of feedback.
:::
## 20.2 The Bias of Traditional Methods (pp. 269-271)
---
::: {#exm-hypothetical-hiv-trial}
## A Hypothetical Two-Time-Point Trial
A hypothetical sequentially randomized trial with 32,000 individuals with HIV
and two time points $k = 0, 1$ (full adherence, no loss to follow-up):
- $A_0 = 1$ is randomly assigned at baseline with probability 0.5;
- $A_1$ is randomly assigned at month 1 with probability 0.4 if $L_1 = 0$ (high CD4)
and 0.8 if $L_1 = 1$ (low CD4);
- the outcome $Y$, measured at the end of follow-up, is a health score (higher is better).
The data are in @tbl-hiv-feedback-trial.
:::
::: {#tbl-hiv-feedback-trial}
| $N$ | $A_0$ | $L_1$ | $A_1$ | Mean $Y$ |
|-----:|:---:|:---:|:---:|-----:|
| 2400 | 0 | 0 | 0 | 84 |
| 1600 | 0 | 0 | 1 | 84 |
| 2400 | 0 | 1 | 0 | 52 |
| 9600 | 0 | 1 | 1 | 52 |
| 4800 | 1 | 0 | 0 | 76 |
| 3200 | 1 | 0 | 1 | 76 |
| 1600 | 1 | 1 | 0 | 44 |
| 6400 | 1 | 1 | 1 | 44 |
Each row is a combination of $(A_0, L_1, A_1)$,
with its number of individuals $N$ and $\E{Y \mid A_0, L_1, A_1}$.
Source: @hernan2020causal [Table 20.1, p. 269].
:::
::: {.notes}
::: {#exm-design-probabilities}
## Reading the Design Back From the Data
The design probabilities can be read back from @tbl-hiv-feedback-trial.
Within $L_1 = 0$: $1600 / (2400 + 1600) = 0.4$ and $3200 / (4800 + 3200) = 0.4$.
Within $L_1 = 1$: $9600 / (2400 + 9600) = 0.8$ and $6400 / (1600 + 6400) = 0.8$.
Also, $16{,}000$ of the $32{,}000$ individuals have $A_0 = 1$, matching probability 0.5.
:::
:::
---
::: {#rem-identifiability-by-design}
## Identifiability Holds by Design
The trial is built so that all three identifiability conditions hold
[@hernan2020causal, p. 269]:
- **sequential exchangeability**:
randomization leaves $A_0$ with no confounders,
and $A_1$ is assigned using $L_1$ alone,
so exchangeability for $A_1$ holds once we condition on $L_1$;
- **positivity**:
each of the eight rows of the table contains individuals;
- **consistency**:
each treatment history $\bar{a} = (a_0, a_1)$ is a well-defined intervention,
and a person whose observed history is $\bar{a}$ has observed outcome $Y = Y^{\bar{a}}$;
in this ideal trial the recorded history is the treatment actually received.
Had the trial run past month 1, with each month's treatment assigned using that month's CD4 count,
each of those later CD4 measurements would be a time-varying confounder as well.
Random variability is ignored throughout.
:::
---
### Each Treatment, Taken Separately, Has No Effect
::: {#exm-no-effect-a1}
## No Effect of $A_1$
In every stratum of $(A_0, L_1)$, the mean outcome is the same for $A_1 = 1$ and $A_1 = 0$
(rows 1 vs 2, 3 vs 4, 5 vs 6, 7 vs 8).
For example,
$$\E{Y \mid A_0 = 0, L_1 = 0, A_1 = 1} - \E{Y \mid A_0 = 0, L_1 = 0, A_1 = 0} = 84 - 84 = 0,$$
and because the identifiability conditions hold this equals
$\E{Y^{a_1=1} \mid A_0 = 0, L_1 = 0} - \E{Y^{a_1=0} \mid A_0 = 0, L_1 = 0}$.
:::
---
::: {#exm-no-effect-a0}
## No Effect of $A_0$
Because $A_0$ is randomized (so exchangeability holds for $A_0$ without adjustment),
and positivity and consistency also hold (@rem-identifiability-by-design),
$\E{Y \mid A_0 = 1} - \E{Y \mid A_0 = 0}$
estimates $\E{Y^{a_0=1}} - \E{Y^{a_0=0}}$.
Averaging rows 1-4 (16,000 individuals with $A_0 = 0$):
$$
\begin{aligned}
\E{Y \mid A_0 = 0}
&= \frac{2400 \times 84 + 1600 \times 84 + 2400 \times 52 + 9600 \times 52}{16000} \\
&= \frac{4000 \times 84 + 12000 \times 52}{16000}
= \frac{336000 + 624000}{16000} = 60.
\end{aligned}
$$
Averaging rows 5-8 (16,000 individuals with $A_0 = 1$):
$$
\begin{aligned}
\E{Y \mid A_0 = 1}
&= \frac{4800 \times 76 + 3200 \times 76 + 1600 \times 44 + 6400 \times 44}{16000} \\
&= \frac{8000 \times 76 + 8000 \times 44}{16000}
= \frac{608000 + 352000}{16000} = 60.
\end{aligned}
$$
So the average causal effect of $A_0$ is $60 - 60 = 0$.
:::
::: {.callout-note title="Technical Point 20.1: G-Null Test"}
The **g-null theorem** [@robins1986new] ties the sharp null
to two conditional independencies in the observed data.
With two time points, the independencies are
$$
Y \ind A_0 \mid L_0
\quad \text{and} \quad
Y \ind A_1 \mid A_0, L_0, L_1.
$$
Assume sequential randomization for every strategy $g$.
The theorem says that both independencies hold
exactly when the distribution of $Y^g$ does not depend on $g$
and coincides with the distribution of the observed $Y$,
so that in particular $\E{Y^g} = \E{Y}$ for every $g$
[@hernan2020causal, Technical Point 20.1, p. 270].
To see why the sharp null leads to the independencies,
note that if no strategy changes anyone's outcome,
then $Y^g = Y$ for every $g$.
Sequential exchangeability for the counterfactuals $Y^g$
then becomes a statement about $Y$ itself.
That step relies on every pair $(a_0, l_0)$
being produced by some strategy with $a_0 = g(l_0)$.
Each independency has a causal reading:
- the first says $A_0$ has no effect within any stratum of $L_0$;
- the second says $A_1$ has no effect within any stratum of $(A_0, L_0, L_1)$,
which in Figure 20.3 (where there is no $L_0$) means within strata of $(A_0, L_1)$.
When sequential exchangeability holds,
checking these independencies in the data therefore tests the sharp null.
That check is the **g-null test** [@robins1986new].
:::
---
### Analyzing the Joint Treatment
The effects of $A_0$ and $A_1$ are each null.
Each null taken alone does not establish the joint null for the strategy $(a_0, a_1)$.
The g-null theorem of Technical Point 20.1 does not apply directly,
because its premise is conditional independence
and the examples check only conditional means.
The mean joint null says that $\E{Y^{a_0, a_1}}$ is the same for every static strategy.
The equal conditional means of @exm-no-effect-a1 and @exm-no-effect-a0 are enough for it,
as computing the mean outcome under each static strategy with the g-formula shows.
That computation uses all three identifiability conditions (@rem-identifiability-by-design):
sequential exchangeability (unconditionally for the randomized $A_0$, and given $(A_0, L_1)$ for $A_1$),
positivity, and consistency.
Under these conditions,
$$
\begin{aligned}
\E{Y^{a_0, a_1}}
&= \E{Y^{a_0, a_1} \mid A_0 = a_0} && \text{(} A_0 \text{ randomized)} \\
&= \sum_{l} \E{Y^{a_0, a_1} \mid A_0 = a_0, L_1 = l} \Pr[L_1 = l \mid A_0 = a_0] && \text{(law of total expectation)} \\
&= \sum_{l} \E{Y^{a_0, a_1} \mid A_0 = a_0, L_1 = l, A_1 = a_1} \Pr[L_1 = l \mid A_0 = a_0] && \text{(sequential exchangeability)} \\
&= \sum_{l} \E{Y \mid A_0 = a_0, L_1 = l, A_1 = a_1} \Pr[L_1 = l \mid A_0 = a_0] && \text{(consistency)}.
\end{aligned}
$$
From @tbl-hiv-feedback-trial, $\Pr[L_1 = 0 \mid A_0 = 0] = 4000/16000 = 0.25$ and $\Pr[L_1 = 0 \mid A_0 = 1] = 8000/16000 = 0.5$,
and the cell means do not depend on $a_1$, so for either value of $a_1$
$$
\begin{aligned}
\E{Y^{a_0=0, a_1}} &= 84 \times 0.25 + 52 \times 0.75 = 21 + 39 = 60, \\
\E{Y^{a_0=1, a_1}} &= 76 \times 0.5 + 44 \times 0.5 = 38 + 22 = 60.
\end{aligned}
$$
So $\E{Y^{a_0, a_1}} = 60$ for every $(a_0, a_1)$, and in particular
$\E{Y^{a_0=1, a_1=1}} - \E{Y^{a_0=0, a_1=0}} = 0$.
Do conventional analyses recover this?
::: {#exm-analysis-no-adjustment}
## Analysis 1: No Adjustment for $L_1$
Compare the 9600 individuals treated at both times (rows 6 and 8)
with the 4800 untreated at both times (rows 1 and 3):
$$
\begin{aligned}
\E{Y \mid A_0 = 1, A_1 = 1} &= \frac{3200 \times 76 + 6400 \times 44}{9600}
= \frac{243200 + 281600}{9600} \approx 54.7, \\
\E{Y \mid A_0 = 0, A_1 = 0} &= \frac{2400 \times 84 + 2400 \times 52}{4800}
= \frac{201600 + 124800}{4800} = 68.0.
\end{aligned}
$$
The difference $54.7 - 68 = -13.3$ is non-null,
because $\E{Y \mid A_0 = a_0, A_1 = a_1}$ is not a valid estimator of $\E{Y^{a_0, a_1}}$:
adjustment for the confounder $L_1$ is needed.
:::
---
::: {#exm-analysis-stratification}
## Analysis 2: Stratification on $L_1$
Within $L_1 = 0$, compare treated at both times (row 6) with untreated at both times (row 1):
$$\E{Y \mid A_0 = 1, L_1 = 0, A_1 = 1} - \E{Y \mid A_0 = 0, L_1 = 0, A_1 = 0} = 76 - 84 = -8.$$
Within $L_1 = 1$ (rows 8 and 3): $44 - 52 = -8$.
The analysis adjusted for the confounder also gives the wrong answer.
:::
::: {.notes}
::: {#rem-no-stratum-weighting}
## No Weighting of the Strata Recovers the Null
Because the stratum-specific difference is $-8$ in *both* strata of $L_1$,
no weighted average of the stratum-specific differences can equal the correct value 0.
This estimate reflects the bias of traditional methods
when there is treatment-confounder feedback (@def-treatment-confounder-feedback) [@hernan2020causal, p. 271].
:::
::: {#exm-gformula-preview}
## Preview of Chapter 21: The G-Formula
The computation that gave $\E{Y^{a_0, a_1}} = 60$ for every strategy is the g-formula:
it standardizes over the distribution of $L_1$ *given* $A_0$
rather than conditioning on $L_1$, as the stratified analysis does.
Applied to "always treat" versus "never treat", it gives $60 - 60 = 0$, as it should.
Chapter 21 develops this computation.
:::
:::
## 20.3 Why Traditional Methods Fail (pp. 271-273)
---
::: {#rem-problem-is-method}
## The Problem Is the Adjustment Method
All three identifiability conditions hold in @tbl-hiv-feedback-trial;
the unmeasured $U_1$ (immunosuppression level) is not needed
because we have data on $L_1$.
Yet neither analysis gave the correct answer.
The problem is the **adjustment method**:
stratification cannot handle treatment-confounder feedback (@def-treatment-confounder-feedback).
:::
---
### Stratification Opens a Collider Path
::: {#rem-stratification-collider}
## Stratification Opens a Collider Path
Stratifying on $L_1$ means estimating the treatment-outcome association
separately in the subsets $L_1 = 0$ (high CD4) and $L_1 = 1$ (low CD4).
But $L_1$ is affected by prior treatment $A_0$,
so it is a **collider** on the path
$$A_0 \to L_1 \leftarrow U_1 \to Y$$
(the book's Figure 20.5: Figure 20.3 with $L_1$ conditioned on).
Conditioning on $L_1$ opens this path
and generically (that is, under faithfulness) induces a noncausal association between $A_0$ and $U_1$,
and hence between $A_0$ and $Y$, within levels of $L_1$.
:::
::: {.notes}
Intuition: among those with low CD4 count ($L_1 = 1$),
being on treatment ($A_0 = 1$) marks severe immunosuppression (high $U_1$),
since treatment would otherwise have raised their CD4.
Among those with high CD4 count ($L_1 = 0$),
being off treatment ($A_0 = 0$) marks milder immunosuppression (low $U_1$).
:::
::: {#rem-confounding-for-selection-bias}
## Trading Confounding for Selection Bias
Stratification eliminates confounding for $A_1$
at the cost of introducing selection bias for $A_0$.
The associational differences
$$\E{Y \mid A_0 = 1, L_1 = l, A_1 = 1} - \E{Y \mid A_0 = 0, L_1 = l, A_1 = 0}$$
may be non-zero even when treatment has no effect on anyone's outcome at any time.
The net bias depends on the relative sizes of the confounding removed
and the selection bias created.
:::
---
### Feedback Through Shared Causes
::: {#rem-shared-cause-collider}
## The Same Bias Under a Shared Cause
The same bias arises when the confounder shares an unmeasured cause $W_0$ with prior treatment
(an observational study, the book's Figure 20.6):
conditioning on the collider $L_1$ opens
$$A_0 \leftarrow W_0 \to L_1 \leftarrow U_1 \to Y.$$
Because conditioning on $L_1$ creates the same bias in Figure 20.4 as in Figure 20.3,
both settings are called treatment-confounder feedback;
the observed data also cannot tell them apart.
:::
---
### Beyond the Null and Beyond Two Time Points
::: {.callout-note title="The Bias Is Not Limited to the Null or to Two Time Points"}
- The bias is not limited to the sharp null.
With a non-null effect (the book's Figure 20.7, where $A_0 \to L_1 \to Y$),
adding arrows into $Y$ from $A_0$, $A_1$, or $L_1$ leaves the collider path through $L_1$ in place,
so conditioning on $L_1$ still generically (that is, under faithfulness) associates $A_0$ with $Y$
in a way that has no causal interpretation.
- With many time points and variables, the bias still arises;
confounders affected by prior treatment at multiple times
increase the possibility of a large bias.
- Valid estimation of a treatment-strategy effect requires estimating
the joint effect of all treatment components $A_k$ simultaneously and without bias,
which stratification may be unable to do even with data on all time-varying confounders.
:::
::: {.callout-note title="Fine Point 20.2: Confounders on the Causal Pathway"}
Conditioning on a confounder $L_1$ affected by prior treatment can create selection bias
even when $L_1$ is not on a causal pathway from treatment to outcome
(no such pathway exists in Figures 20.5 and 20.6).
In Figure 20.7, by contrast, $L_1$ is a confounder for $A_1$
*and* lies on the causal pathway $A_0 \to L_1 \to Y$.
Even if $U_1$ were not a common cause of $L_1$ and $Y$ (no selection bias),
the $A$-$Y$ associations within strata of $L_1$ would estimate only the direct effect of $A_0$
not through $L_1$, not the overall effect of $\bar{A}$ on $Y$.
The common rule that variables on a causal pathway cannot be confounders
is inaccurate for time-varying treatments:
a confounder for later treatment $A_1$ can lie on a pathway from earlier treatment $A_0$ to $Y$.
Whether adjusting for it induces bias depends on the method:
stratification does; g-methods do not
[@hernan2020causal, Fine Point 20.2, p. 272].
:::
## 20.4 Why Traditional Methods Cannot Be Fixed (pp. 273-274)
---
Could parametric **outcome regression** succeed where nonparametric stratification failed?
The question matters most with high-dimensional data,
where a simple stratified analysis is impossible.
::: {#exm-counting-strategies}
## Counting Static Strategies
- With two time points there are $2^2 = 4$ static strategies $\bar{a}$.
- With about 100 time points (not unusual in practice),
there are about $2^{100}$ static strategies,
far more than the sample size of any study,
and even more once dynamic strategies are considered.
:::
::: {.notes}
The number of combinations of the *data* is larger still,
because several confounders $L_k$ are measured at each time point.
:::
---
### Modeling Is Needed, but Does Not Help
::: {#rem-modeling-cumulative-treatment}
## Many Strategies Require Modeling
As argued since Chapter 11, many possible strategies require **modeling**:
a dose-response function for the effect of treatment history $\bar{a}$ on the mean outcome.
For example, assume the effect is linear in **cumulative treatment**,
so every strategy with exactly three months of treatment has the same effect,
whenever those three months occur.
The price is a new threat to validity: misspecification of the dose-response model.
:::
::: {.callout-warning title="Conventional Outcome Regression Does Not Remove the Bias"}
Paying the price of a dose-response model buys no protection
if we read the treatment coefficients of an outcome regression as the effect
while the model conditions on the time-varying confounder.
That conventional use of regression **is** a stratification-based method,
so it cannot remove the bias of stratification under treatment-confounder feedback (@def-treatment-confounder-feedback).
The g-formula (Chapter 21) can fit the same kind of outcome model,
but it averages the model's predictions over the distribution of each time-varying confounder
given past treatment and past confounders
(in this chapter's two-time-point example, which has no $L_0$: $L_1$ given $A_0$)
instead of reading the effect off the coefficients.
:::
---
### Example: A Cumulative-Treatment Regression
::: {#exm-cumulative-treatment-regression}
## A Cumulative-Treatment Regression
For data generated under Figure 20.5, define $\text{cum}(\bar{A}) = A_0 + A_1 \in \{0, 1, 2\}$.
"Always treat" is $\text{cum}(\bar{a}) = 2$ and "never treat" is $\text{cum}(\bar{a}) = 0$;
the target is $\E{Y^{\text{cum}(\bar{a})=2}} - \E{Y^{\text{cum}(\bar{a})=0}}$, whose true value is 0.
Fit the outcome regression model
$$\E{Y \mid \bar{A}, L_1} = \theta_0 + \theta_1 \, \text{cum}(\bar{A}) + \theta_2 L_1.$$
Within either level $l$ of $L_1$,
$$
\begin{aligned}
&\E{Y \mid \text{cum}(\bar{A}) = 2, L_1 = l} - \E{Y \mid \text{cum}(\bar{A}) = 0, L_1 = l} \\
&\quad = (\theta_0 + 2\theta_1 + \theta_2 l) - (\theta_0 + 0 \cdot \theta_1 + \theta_2 l)
= 2\theta_1 .
\end{aligned}
$$
:::
::: {.callout-warning title="Do Not Read $2\theta_1$ as the Causal Effect"}
It is tempting to read $2\theta_1$ as the effect of "always treat" versus "never treat" within levels of $L_1$.
But conditioning on $L_1$ generically induces an association between $A_0$ (a component of $\text{cum}(\bar{A})$) and $Y$,
so, generically, $\theta_1 \neq 0$ even though the true effect is zero [@hernan2020causal, p. 274].
The bias comes from conditioning on the collider $L_1$,
not from the form of the model or from the sample size:
it would persist with unlimited data.
:::
::: {.notes}
A similar argument applies to **matching**.
G-methods are needed to adjust appropriately for time-varying confounders
in the presence of treatment-confounder feedback [@hernan2020causal, p. 274].
::: {#exm-correct-outcome-model}
## A Correctly Specified Outcome Model for the Trial Data
The book states that $\theta_1$ is non-zero
"even if the true causal effect is zero and the regression model for $\E{Y \mid \bar{A}, L_1}$ is correct"
[@hernan2020causal, p. 274].
Under Figure 20.5, the cumulative-treatment model itself generically cannot be the correct one.
$Y$ is independent of $A_1$ given $(A_0, L_1)$,
so a correct model for $\E{Y \mid A_0, A_1, L_1}$ has no $A_1$ term,
whereas $\text{cum}(\bar{A})$ forces $A_0$ and $A_1$ to share the coefficient $\theta_1$.
The data in @tbl-hiv-feedback-trial illustrate this.
Their cell means depend on $(A_0, L_1)$ only, and they are fit exactly by
$$\E{Y \mid A_0, A_1, L_1} = 84 - 8 A_0 + 0 \cdot A_1 - 32 L_1,$$
which reproduces all four distinct cell means:
- $A_0 = 0, L_1 = 0$: $84$;
- $A_0 = 0, L_1 = 1$: $84 - 32 = 52$;
- $A_0 = 1, L_1 = 0$: $84 - 8 = 76$;
- $A_0 = 1, L_1 = 1$: $84 - 8 - 32 = 44$.
The conclusion is unchanged:
even this correctly specified model gives $A_0$ a non-zero coefficient ($-8$)
although the true effect of $A_0$ is zero,
because the model conditions on the collider $L_1$.
:::
:::
## 20.5 Adjusting for Past Treatment (pp. 274-275)
---
::: {#rem-past-treatment-backdoor}
## Past Treatment Opens New Backdoor Paths
So far the diagrams had no arrow $A_0 \to A_1$.
Now suppose doctors use past treatment history $\bar{A}_{k-1}$
when deciding on treatment $A_k$.
Adding $A_0 \to A_1$ to Figures 20.3 and 20.4 gives the book's Figures 20.8 and 20.9.
Under treatment-confounder feedback (@def-treatment-confounder-feedback),
adjusting for $L_1$ alone no longer closes every backdoor path from $A_1$ to $Y$.
$L_1$ is a collider on each of the following paths,
so conditioning on it opens them:
$$
\begin{aligned}
&A_1 \leftarrow A_0 \to L_1 \leftarrow U_1 \to Y && \text{(Figure 20.8)}, \\
&A_1 \leftarrow A_0 \leftarrow W_0 \to L_1 \leftarrow U_1 \to Y && \text{(Figure 20.9)}.
\end{aligned}
$$
And whenever past treatment affects the outcome (Figure 20.10),
$A_0$ is a confounder for the effect of $A_1$, feedback or not.
:::
::: {#rem-condition-on-treatment-history}
## Sequential Exchangeability Conditions on Treatment History
Sequential exchangeability at time $k$ generally requires conditioning
on the treatment history $\bar{A}_{k-1}$ as well as on the covariates;
conditioning only on $L$ is not enough.
That is why every sequential exchangeability statement in Chapters 19 and 20
conditions on treatment history.
:::
---
### Past Treatment and Time-Fixed Treatments
::: {#def-short-term-effect}
## Short-Term Effect
In a two-time-point treatment $\bar{A} = (A_0, A_1)$,
the **short-term effect** of $A_1$ is the contrast
$$\E{Y^{a_1=1}} - \E{Y^{a_1=0}},$$
in which only $A_1$ is set by intervention
and $A_0$ is not intervened on, so it keeps whatever value it would naturally take.
That is, it is the effect of $A_1$ when $A_1$ is treated as a time-fixed treatment.
:::
Suppose the target is the short-term effect of $A_1$ (@def-short-term-effect),
treating $A_1$ as a time-fixed treatment.
::: {.callout-warning title="Ignoring Past Treatment Biases the Short-Term Effect"}
Without adjustment for $A_0$:
- with treatment-confounder feedback, there is generically (under faithfulness) selection bias;
- if $A_0$ affects $Y$ directly, there is generically confounding.
So $\E{Y \mid A_1 = 1, L_1} - \E{Y \mid A_1 = 0, L_1}$ need not be zero
even if $A_1$ has no effect on anyone's outcome (Figures 20.8-20.10).
:::
::: {.notes}
::: {.callout-tip title="Adjust for Treatment History or Restrict to New Users"}
In practice, this bias tends to show up
when current users ($A_1 = 1$) of a time-fixed treatment are compared with nonusers ($A_1 = 0$).
Either of two remedies avoids it:
- adjust for the prior treatment history;
- keep only individuals who share one particular treatment history.
**New-user designs** take the second route,
keeping only people who have never used the treatment before.
The restriction is a stand-in for adjustment:
an analysis that already adjusts correctly for past treatment
gains nothing against this bias by also discarding prevalent users.
:::
:::
---
### Mismeasured Past Treatment
::: {.callout-warning title="Mismeasured Past Treatment"}
The need to adjust for past treatment matters when past treatment is **mismeasured**.
As in [Section 9.3](09-measurement-bias.qmd#mismeasured-confounders-and-colliders-pp.-128-130), adjusting for a mismeasured confounder can leave bias in either direction.
:::
::: {#exm-mismeasured-a0}
## Self-Reported Treatment History
Suppose HIV investigators lack medical records and ascertain prior treatment by questionnaire,
so they observe a mismeasured $A_0^*$ rather than $A_0$.
Adding the arrow $A_0 \to A_0^*$ to Figures 20.8-20.10 shows the problem.
$A_0^*$ is only a noisy child of $A_0$,
so conditioning on it leaves open the backdoor paths from $A_1$ to $Y$ on which $A_0$ lies.
Investigators would generically (under faithfulness) find an association between $A_1$ and $Y$
even after adjusting for $A_0^*$ and $L_1$,
although $A_1$ has no effect on $Y$.
:::
::: {.notes}
::: {.callout-warning title="Measurement Error Can Bias Even Under the Null"}
With a time-varying treatment, errors in recording treatment **can bias the estimate when treatment has no effect**,
even if the errors are independent and non-differential,
although such errors are often assumed to be harmless under the null.
Earlier treatment confounds later treatment whether or not it affects the outcome,
and an error-prone record of earlier treatment removes that confounding only in part.
When treatment does have an effect, the same partial adjustment can overstate it.
@robins1987addendum showed that random errors in measured treatment can push estimates away from the null
[@hernan2020causal, p. 275].
:::
:::
## Summary
---
- **Treatment-confounder feedback** (@def-treatment-confounder-feedback):
a time-varying confounder $L_k$ affects later treatment $A_k$
and, for some earlier treatment $A_j$ ($j < k$), either is affected by $A_j$
or shares with $A_j$ a cause outside the measured past of $A_j$,
through a path that conditioning on that measured past does not block.
- In @tbl-hiv-feedback-trial the true effect of "always treat" vs. "never treat" is 0,
yet the unadjusted analysis gives $-13.3$ and stratification on $L_1$ gives $-8$.
- Stratification on $L_1$ removes confounding for $A_1$
but opens the collider path $A_0 \to L_1 \leftarrow U_1 \to Y$,
creating selection bias for $A_0$.
- Conventional outcome regression (reading the effect off treatment coefficients in a model that conditions on $L_1$)
and matching are stratification-based too,
so they are biased even when the regression model is correct.
- Sequential exchangeability generally requires conditioning on past treatment;
mismeasured past treatment can bias estimates even under the null.
- G-methods (Chapter 21) handle treatment-confounder feedback correctly.
## References
---
::: {#refs}
:::