Every adjustment method in this book (stratification and outcome regression, standardization and the parametric g-formula, IP weighting, g-estimation) relies on the same condition: the adjustment set \(L\) must be enough to make the treated and the untreated conditionally exchangeable, i.e., to block all backdoor paths between \(A\) and \(Y\) without opening other biasing paths. This chapter asks how to choose \(L\), and how to estimate effects when the candidate covariates \(X\) are high-dimensional.
Example 1 (Prediction Is Not Intervention) Investigators use outcome regression to identify patients at high risk of heart failure, a classification (prediction) task. A prior hospitalization may be a useful predictor, yet nobody would propose to stop admitting people to hospitals to prevent heart failure. Identifying patients with bad prognosis is different from identifying the best action to prevent or treat a disease (Hernán and Robins 2020, 241–42).
In a predictive model all covariates have the same status: there is no treatment \(A\) and no adjustment set \(L\).
A causal analysis requires a thoughtful selection of confounders. Automatic variable selection may work for prediction but not, in general, for causal inference: it may select adjustment variables that introduce bias (Hernán and Robins 2020, 242).
Fine Point 18.1: Variable Selection Procedures for Regression Models
For prediction with very many candidate predictors (perhaps more than individuals) (Hernán and Robins 2020, 243):
Fine Point 18.2: Overfitting and Cross-Validation
Overfitting: a model selected to fit the data as well as possible also fits random variation, so it predicts well for the individuals used to fit it and poorly for new ones. The same holds for random forests, neural networks, and other machine learning algorithms.
Imagine unlimited computing power and a quasi-infinite number of individuals, with \(A\), \(Y\), and a moderate number of discrete variables \(X\), some of which may be confounders.
Colliders (Figure 18.1, same as Figure 7.7): the effect of \(A\) on \(Y\) is null and \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) estimates it without bias. The g-formula contrast adjusted for \(L\),
\[ \textstyle\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = 1, L = l\right]\mathclose{} \Pr(L = l) - \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = 0, L = l\right]\mathclose{} \Pr(L = l), \]
is biased, because \(L\) is associated with \(Y\) given \(A\) and with \(A\), so \(\Pr(L = l) \neq \Pr(L = l \mid A)\). This is selection bias under the null. The same holds for a descendant of a collider (Figure 18.2).
Noncolliders affected by treatment (Figure 18.3): adjusting for \(L\) biases a non-null effect, but under the null \(A\) no longer affects \(L\) and the adjusted contrast is unbiased. This is selection bias under the alternative (off the null; Section 6.5).
Mediators (Figure 18.4, \(A \rightarrow L \rightarrow Y\) and \(A \rightarrow Y\)): adjusting for \(L\) or its descendants blocks the part of the effect through \(L\), sometimes called overadjustment for mediators when the total effect is of interest (Hernán and Robins 2020, 243–45).
All the variables above are affected by treatment, but avoiding every variable measured after \(A\) can exclude useful adjustment variables (Fine Point 7.4).
So whether to adjust for \(L\) cannot be decided by any automated procedure based only on statistical associations; for example, data alone cannot distinguish a collider from a confounder. The book cites Hernán et al. (2002) as an example of applying expert knowledge to adjustment.
Even with temporal order \(L, A, Y\) and a sample so large that only bias matters, adjusting for all pre-treatment covariates does not minimize bias, for two reasons (Hernán and Robins 2020, 245).
Reason 1: colliders. In Figure 18.6 (same as Figure 7.4), \(L\) shares a cause \(U_2\) with \(A\) and a cause \(U_1\) with \(Y\), so \(L\) is a collider on \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\). Adjusting for \(L\) opens that path: M-bias.
Reason 2: bias amplification. In Figure 18.7 (same as Figure 16.1), \(Z \rightarrow A \rightarrow Y\) and an unmeasured \(U\) confounds \(A\) and \(Y\). \(Z\) is an instrument, not on any backdoor path, so adjusting for \(Z\) cannot remove the confounding by \(U\). But it can amplify it: the \(Z\)-adjusted \(A\)-\(Y\) association may be further from the effect than the unadjusted one. This is called Z-bias (Hernán and Robins 2020, 246).
Even with no computational constraints and a quasi-infinite sample, adjusting for all available variables is not advisable. Ideally the adjustment set excludes every variable that may introduce (colliders, their descendants, mediators, other variables affected by treatment) or amplify (instruments) bias. Because these variables cannot be identified by purely statistical algorithms, expert knowledge must guide variable selection (Hernán and Robins 2020, 246).
dagittyOnce subject-matter knowledge is written down as a causal diagram, software can list adjustment sets that satisfy the backdoor criterion. For the M-bias diagram of Figure 18.6:
From here on, assume \(X\) includes all confounders \(L\) and no variable that would induce or amplify bias, and that positivity holds. The problem is now estimation when \(X\) is very high-dimensional or has many continuous variables. Since good estimates of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) give a good estimate of their difference, focus on \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\).
Definition 1 (The Two Nuisance Functions) \[ b(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y \mid A = 1, X\right]\mathclose{}, \qquad \pi(X) \stackrel{\text{def}}{=}\Pr[A = 1 \mid X]. \]
The plug-in g-formula (standardization) estimates \(b(X)\); IP weighting estimates \(\pi(X)\) (Hernán and Robins 2020, 246–47).
Options (Hernán and Robins 2020, 247):
With large samples and many covariates, machine learning usually predicts conditional expectations better than traditional parametric models.
But prediction is not enough: to get valid 95% Wald confidence intervals (covering the true parameter at least 95% of the time), machine learning must be combined with doubly robust estimators using sample splitting and cross-fitting (Section 18.4).
For \(\psi = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Hernán and Robins 2020, 247):
The full-sample augmented IP weighted (AIPW) estimator of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Technical Point 13.2) is
\[ \frac{1}{n} \sum_{i=1}^n\mathopen{}\left\{\hat b(X_i) + \frac{A_i}{\hat\pi(X_i)} \mathopen{}\left[Y_i - \hat b(X_i)\right]\mathclose{}\right\}\mathclose{}, \]
with \(\hat b\) and \(\hat\pi\) fit by machine learning on all \(n\) individuals.
Sample splitting’s cost: the standard error corresponds to a sample of size \(n/2\), so intervals are wider. Cross-fitting recovers the lost efficiency (Hernán and Robins 2020, 248–49):
Technical Point 18.1: Augmented IP Weighted Split-Sample and Cross-Fit Estimator
For \(\psi = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Hernán and Robins 2020, 249):
A version with better finite-sample behavior uses \(M > 2\) equal-sized random subsamples: compute \(\hat\psi^{(m)}\) using subsample \(m\) for estimation and the other \(M - 1\) for training, and set \(\hat\psi_{\text{cross-fit}} = \frac{1}{M} \sum_{m=1}^{M} \hat\psi^{(m)}\).
Technical Point 18.2: Statistical Properties of Split-Sample and Cross-Fit Estimators
Conditional on the training data \(Tr\), \(\hat b\) and \(\hat\pi\) are fixed functions, so \(\hat\psi\) is an average of independent and identically distributed terms and, by the central limit theorem, is asymptotically normal given \(Tr\) with standard error proportional to \(n^{-1/2}\).
If (i) \(\hat b\) and \(\hat\pi\) are consistent in mean square and (ii) the conditional bias divided by the standard error converges to 0 in probability, then \(\hat\psi\) and \(\hat\psi_{\text{cross-fit}}\) are asymptotically normal and unbiased, and their Wald intervals (with bootstrap standard errors) are valid and calibrated. \(\hat\psi_{\text{cross-fit}}\) is then semiparametric efficient, with standard error smaller than that of \(\hat\psi\) by a factor of \(1/\sqrt{2}\) (Hernán and Robins 2020, 250).
Proposition 1 (Conditional Bias of the Split-Sample AIPW Estimator) Assume conditional exchangeability given \(X\) and positivity, so \(\psi = \operatorname{E}\mathopen{}\left[b(X)\right]\mathclose{}\). Then
\[ \operatorname{E}\mathopen{}\left[\hat\psi - \psi \mid Tr\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\pi(X) \mathopen{}\left(\frac{1}{\pi(X)} - \frac{1}{\hat\pi(X)}\right)\mathclose{} \mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mid Tr\right]\mathclose{} . \]
Proof. Given \(Tr\), each term of \(\hat\psi\) has the same mean, so \(\operatorname{E}\mathopen{}\left[\hat\psi \mid Tr\right]\mathclose{}\) is the mean of one term. Conditioning on \(X\) first:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\frac{A}{\hat\pi(X)} \mathopen{}\left[Y - \hat b(X)\right]\mathclose{} \mid X, Tr\right]\mathclose{} &= \frac{\Pr[A = 1 \mid X]}{\hat\pi(X)} \operatorname{E}\mathopen{}\left[Y - \hat b(X) \mid A = 1, X, Tr\right]\mathclose{} && \text{(only } A = 1 \text{ contributes)} \\ &= \frac{\pi(X)}{\hat\pi(X)} \mathopen{}\left(b(X) - \hat b(X)\right)\mathclose{}. && (\text{estimation sample independent of } Tr) \end{aligned} \]
So
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\hat\psi - \psi \mid Tr\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\hat b(X) + \frac{\pi(X)}{\hat\pi(X)} \mathopen{}\left(b(X) - \hat b(X)\right)\mathclose{} - b(X) \mid Tr\right]\mathclose{} && (\psi = \operatorname{E}\mathopen{}\left[b(X)\right]\mathclose{}) \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mathopen{}\left(1 - \frac{\pi(X)}{\hat\pi(X)}\right)\mathclose{} \mid Tr\right]\mathclose{} && \text{(collect terms)} \\ &= \operatorname{E}\mathopen{}\left[\pi(X) \mathopen{}\left(\frac{1}{\pi(X)} - \frac{1}{\hat\pi(X)}\right)\mathclose{} \mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mid Tr\right]\mathclose{} . && \text{(factor out } \pi(X)) \end{aligned} \]
Doubly robust machine learning refutes the belief that any data-adaptive selection of adjustment variables must give incorrect confidence intervals. But it does not solve every problem, for at least three reasons (Hernán and Robins 2020, 250):
General, prescriptive guidelines for variable selection may not be possible while so much research is ongoing. As in Section 13.5, the book’s advice is to run multiple sensitivity analyses: apply several analytic methods and compare the estimates. If they agree, we gain confidence; if they disagree, our job is to understand why (Hernán and Robins 2020, 251).