Chapter 18: Variable Selection and High-Dimensional Data

Published

Last modified: 2026-10-09 13:21:38 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Every adjustment method in this book (stratification and outcome regression, standardization and the parametric g-formula, IP weighting, g-estimation) relies on the same condition: the adjustment set \(L\) must be enough to make the treated and the untreated conditionally exchangeable, i.e., to block all backdoor paths between \(A\) and \(Y\) without opening other biasing paths. This chapter asks how to choose \(L\), and how to estimate effects when the candidate covariates \(X\) are high-dimensional.

This chapter is based on Hernán and Robins (2020, chap. 18, pp. 241-251).

Key theme: variable selection for causal inference is not variable selection for prediction. Some variables induce or amplify bias when adjusted for, and no purely statistical algorithm can find them; subject-matter knowledge is needed. Given a good set of candidate covariates, machine learning can estimate the nuisance functions, but valid confidence intervals need doubly robust estimators combined with sample splitting and cross-fitting.

1 18.1 The Different Goals of Variable Selection (pp. 241-242)


1.1 Prediction Does Not Need Confounding Adjustment

  • Causal analysis: if confounders \(L\) may explain part of the \(A\)-\(Y\) association, adjustment is required before the association can be read as an effect.
  • Predictive analysis: no adjustment for confounding is needed, because there is no treatment whose effect could be confounded. To predict weight gain we simply include whatever predicts it (e.g., smoking cessation, baseline weight, income).

Confounding is a causal concept; it does not apply when the estimand is an association (Hernán and Robins 2020, 241). The distinction between predictive and causal models was discussed in Section 15.5.

Even an outcome model that includes all confounders of the effect of \(A\) on \(Y\) does not give the coefficients of those confounders a causal interpretation, because nothing adjusts for the confounders of the confounders (Hernán and Robins 2020, 241).


Example 1 (Prediction Is Not Intervention) Investigators use outcome regression to identify patients at high risk of heart failure, a classification (prediction) task. A prior hospitalization may be a useful predictor, yet nobody would propose to stop admitting people to hospitals to prevent heart failure. Identifying patients with bad prognosis is different from identifying the best action to prevent or treat a disease (Hernán and Robins 2020, 241–42).

In a predictive model all covariates have the same status: there is no treatment \(A\) and no adjustment set \(L\).


1.2 Choosing Predictors: Tuning Parameters and Black Boxes

  • Most prediction algorithms have tuning parameters chosen to optimize predictive accuracy, e.g., the regularization parameter of the lasso or ridge regression (degree of shrinkage), or the depth and width of a neural network.
  • They are usually chosen by cross-validation (Fine Point 18.2), which is also used to choose between algorithms, since no single algorithm predicts best in all data sets.
  • Black-box algorithms (e.g., deep neural nets): one view is that anything that improves prediction is fair game; another is that interpretability matters (Hernán and Robins 2020, 242).

Arguments for interpretable algorithms (Hernán and Robins 2020, 242):

  • Physicians may not feel comfortable acting on predictions they cannot explain to themselves and to the patient, especially for unusual combinations of symptoms.
  • Black-box algorithms may perform poorly in new settings because they may rely on local, often noncausal, features of the training setting.
  • They can bake in predictive but socially discriminatory practices that were in place when the training data were collected.

1.3 Causal Analysis Needs Thoughtful Selection

A causal analysis requires a thoughtful selection of confounders. Automatic variable selection may work for prediction but not, in general, for causal inference: it may select adjustment variables that introduce bias (Hernán and Robins 2020, 242).


NoteFine Point 18.1: Variable Selection Procedures for Regression Models

For prediction with very many candidate predictors (perhaps more than individuals) (Hernán and Robins 2020, 243):

  • Best subset selection: fit every model with a given number of variables and pick the best by a pre-specified criterion (e.g., Akaike’s information criterion). Computationally infeasible for many variables, and in a finite data set not guaranteed to give the smallest prediction error.
  • Forward selection, backward elimination, and stepwise selection (a combination): efficient, but do not explore all subsets.
  • Shrinkage: add a penalty that pulls parameter estimates (other than the intercept) toward zero, lowering variance and stabilizing predictions. Ridge regression shrinks; the lasso (“least absolute shrinkage and selection operator”) can set some parameters exactly to zero, so it also selects variables.

NoteFine Point 18.2: Overfitting and Cross-Validation

Overfitting: a model selected to fit the data as well as possible also fits random variation, so it predicts well for the individuals used to fit it and poorly for new ones. The same holds for random forests, neural networks, and other machine learning algorithms.

  • Training/validation split: with sample size \(n\), fit on \(n - v\) individuals and evaluate predictions on the other \(v\); with the lasso, the validation performance can guide the degree of shrinkage.
  • Cross-validation: repeat the split many times and average the predictive accuracy over the validation samples, so the algorithm effectively uses more individuals.
  • Leave-\(v\)-out examines all partitions with a validation set of size \(v\); this can be infeasible, so one can take \(v = 1\) or evaluate only some partitions.
  • \(k\)-fold: split into \(k\) equal subsamples, use each as the validation sample with the other \(k - 1\) as training; \(k = 10\) is common (Hernán and Robins 2020, 244).

Deep neural networks with many layers can have thousands to billions of parameters and often fit the training data exactly, yet with massive training data (e.g., speech recognition, images) they often seem nearly immune to overfitting and still predict new individuals well. Explaining this is an active research area (Hernán and Robins 2020, 244).

2 18.2 Variables That Induce or Amplify Bias (pp. 242-246)


2.1 An Idealized Setting

Imagine unlimited computing power and a quasi-infinite number of individuals, with \(A\), \(Y\), and a moderate number of discrete variables \(X\), some of which may be confounders.

  • For prediction (under squared-error loss), the best predictor of \(Y\) is the mean of \(Y\) in each joint stratum of \(A\) and \(X\).
  • For the average causal effect \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), we could adjust for all of \(X\). Is there any reason to adjust for only a subset? Yes: adjusting for some variables induces bias (Hernán and Robins 2020, 242–43).

Collapsibility reminder: an adjusted association measure may differ from the unadjusted one even without confounding (Fine Point 4.3) (Hernán and Robins 2020, 243).


2.2 Post-Treatment Variables That Induce Bias

  • Colliders (Figure 18.1, same as Figure 7.7): the effect of \(A\) on \(Y\) is null and \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) estimates it without bias. The g-formula contrast adjusted for \(L\),

    \[ \textstyle\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = 1, L = l\right]\mathclose{} \Pr(L = l) - \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = 0, L = l\right]\mathclose{} \Pr(L = l), \]

    is biased, because \(L\) is associated with \(Y\) given \(A\) and with \(A\), so \(\Pr(L = l) \neq \Pr(L = l \mid A)\). This is selection bias under the null. The same holds for a descendant of a collider (Figure 18.2).

  • Noncolliders affected by treatment (Figure 18.3): adjusting for \(L\) biases a non-null effect, but under the null \(A\) no longer affects \(L\) and the adjusted contrast is unbiased. This is selection bias under the alternative (off the null; Section 6.5).

  • Mediators (Figure 18.4, \(A \rightarrow L \rightarrow Y\) and \(A \rightarrow Y\)): adjusting for \(L\) or its descendants blocks the part of the effect through \(L\), sometimes called overadjustment for mediators when the total effect is of interest (Hernán and Robins 2020, 243–45).

In Figure 18.4, the \(L\)-adjusted \(A\)-\(Y\) association is a biased estimator of the total effect of \(A\) on \(Y\) but an unbiased estimator of the direct effect not mediated through \(L\) (Hernán and Robins 2020, 245). Chapter 8 has more structures with colliders and their descendants.


2.3 “Never Adjust for Post-Treatment Variables” Is Not the Rule

All the variables above are affected by treatment, but avoiding every variable measured after \(A\) can exclude useful adjustment variables (Fine Point 7.4).

  • In Figure 18.5, \(L\) occurs after \(A\) but can block a backdoor path: the \(L\)-adjusted association is unbiased and the unadjusted one is biased.
  • Causal graphs do not care about temporal order: when \(A\) does not affect \(L\), the correct analysis is the same whether \(L\) comes before or after \(A\).
  • But the data cannot tell whether \(A\) affects \(L\): with temporal order \(A, L, Y\), any joint distribution of \((A, L, Y)\) with no independencies is compatible with several causal graphs (Hernán and Robins 2020, 245).

So whether to adjust for \(L\) cannot be decided by any automated procedure based only on statistical associations; for example, data alone cannot distinguish a collider from a confounder. The book cites Hernán et al. (2002) as an example of applying expert knowledge to adjustment.


2.4 Pre-Treatment Variables That Induce Bias: M-Bias

Even with temporal order \(L, A, Y\) and a sample so large that only bias matters, adjusting for all pre-treatment covariates does not minimize bias, for two reasons (Hernán and Robins 2020, 245).

Reason 1: colliders. In Figure 18.6 (same as Figure 7.4), \(L\) shares a cause \(U_2\) with \(A\) and a cause \(U_1\) with \(Y\), so \(L\) is a collider on \(A \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\). Adjusting for \(L\) opens that path: M-bias.

The data cannot distinguish confounders from colliders, so external information must decide. If there were also an arrow \(L \rightarrow A\) in Figure 18.6, \(L\) would be both a confounder and a collider, and the average causal effect would not be identified whether or not we adjust for \(L\) (Hernán and Robins 2020, 245).


2.5 Pre-Treatment Variables That Amplify Bias: Z-Bias

Reason 2: bias amplification. In Figure 18.7 (same as Figure 16.1), \(Z \rightarrow A \rightarrow Y\) and an unmeasured \(U\) confounds \(A\) and \(Y\). \(Z\) is an instrument, not on any backdoor path, so adjusting for \(Z\) cannot remove the confounding by \(U\). But it can amplify it: the \(Z\)-adjusted \(A\)-\(Y\) association may be further from the effect than the unadjusted one. This is called Z-bias (Hernán and Robins 2020, 246).

Amplification is guaranteed when all equations of the structural equation model for the diagram are linear (Wooldridge 2010, Pearl 2011), and may also occur in more realistic settings (Ding et al. 2017). It is not guaranteed in general: adjusting for \(Z\) could also reduce the bias, and usually we cannot know which (Hernán and Robins 2020, 246).


2.6 Summary of Section 18.2

Even with no computational constraints and a quasi-infinite sample, adjusting for all available variables is not advisable. Ideally the adjustment set excludes every variable that may introduce (colliders, their descendants, mediators, other variables affected by treatment) or amplify (instruments) bias. Because these variables cannot be identified by purely statistical algorithms, expert knowledge must guide variable selection (Hernán and Robins 2020, 246).


2.7 Supplement: Encoding Expert Knowledge with dagitty

Once subject-matter knowledge is written down as a causal diagram, software can list adjustment sets that satisfy the backdoor criterion. For the M-bias diagram of Figure 18.6:

library(dagitty)
m_bias <- dagitty("dag {
  U1 -> L
  U1 -> Y
  U2 -> L
  U2 -> A
  A -> Y
}")
adjustmentSets(m_bias, exposure = "A", outcome = "Y")

Here the only minimal adjustment set is the empty set: no adjustment is needed, and adjusting for \(L\) alone would open the M path. The output is only as good as the diagram, which is exactly the subject-matter knowledge Section 18.2 says cannot come from the data. This supplement is not in the book.

3 18.3 Causal Inference and Machine Learning (pp. 246-247)


From here on, assume \(X\) includes all confounders \(L\) and no variable that would induce or amplify bias, and that positivity holds. The problem is now estimation when \(X\) is very high-dimensional or has many continuous variables. Since good estimates of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) give a good estimate of their difference, focus on \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\).

Definition 1 (The Two Nuisance Functions) \[ b(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y \mid A = 1, X\right]\mathclose{}, \qquad \pi(X) \stackrel{\text{def}}{=}\Pr[A = 1 \mid X]. \]

The plug-in g-formula (standardization) estimates \(b(X)\); IP weighting estimates \(\pi(X)\) (Hernán and Robins 2020, 246–47).


3.1 Why Traditional Parametric Models Fail Here

  • With high-dimensional \(X\), low-dimensional parametric models for \(b\) and \(\pi\) are certain to be misspecified, so \(\hat b\) and \(\hat\pi\) are inconsistent.
  • Richly parameterized generalized linear models use a linear predictor \(\theta^\top s(x) = \sum_j \theta_j s_j(x)\), where \(s(X)\) is a high-dimensional vector of transformations of \(X\) (e.g., cubic splines and their cross-products), e.g., \(\hat b(x) = \operatorname{expit}\mathopen{}\left\{\hat\theta^\top s(x)\right\}\mathclose{}\) and \(\hat\pi(x) = \operatorname{expit}\mathopen{}\left\{\hat\theta_\pi^\top s(x)\right\}\mathclose{}\).
  • Even when consistent, their finite-sample errors are much larger than those of a correct low-dimensional model, and if the dimension of \(s(X)\) exceeds \(n\) the fit fails to converge (Hernán and Robins 2020, 247).

Remember that some variables in \(X\) may not be confounders; we would not need to adjust for them if we knew which they were (Hernán and Robins 2020, 247).


3.2 Machine Learning for the Nuisance Functions

Options (Hernán and Robins 2020, 247):

  • the parametric model with a lasso or ridge penalty (Fine Point 18.1);
  • a variable selection algorithm such as stepwise selection;
  • other machine learning algorithms: tree-based (e.g., random forests) or neural networks (deep learning), which effectively fit thousands or millions of parameters.

With large samples and many covariates, machine learning usually predicts conditional expectations better than traditional parametric models.

But prediction is not enough: to get valid 95% Wald confidence intervals (covering the true parameter at least 95% of the time), machine learning must be combined with doubly robust estimators using sample splitting and cross-fitting (Section 18.4).

4 18.4 Doubly Robust Machine Learning Estimators (pp. 247-250)


4.1 What a Valid Wald Interval Requires

For \(\psi = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Hernán and Robins 2020, 247):

  • The standard error of most estimators \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) scales as \(1/\sqrt{n}\), so the bias must be much smaller than \(1/\sqrt{n}\).
  • The estimator must also be asymptotically normal (usually easier than small bias).

Undercoverage is worse when there is confounding in the superpopulation, because the Wald interval is then centered on a biased estimator (Chapter 10) (Hernán and Robins 2020, 247).


4.2 Why Double Robustness Helps

  • The bias of a doubly robust estimator depends on the product of the errors in \(1/\hat\pi(x)\) and \(\hat b(x)\), so it is below \(1/\sqrt{n}\) if both errors are much smaller than \(n^{-1/4}\) (since \(n^{-1/4} \times n^{-1/4} = n^{-1/2}\)). Machine learning can often achieve that when \(\pi\) and \(b\) are smooth (many derivatives) or sparse (depend on few, unknown components of \(X\)). This is the second-order bias property (Technical Point 13.2).
  • The IP weighted and plug-in g-formula estimators can have bias as large as the error in \(1/\hat\pi\) or in \(\hat b\) alone, and with high-dimensional \(X\) those errors must exceed \(1/\sqrt{n}\), so neither can generally center a valid Wald interval (Hernán and Robins 2020, 247–48).

4.3 Sample Splitting

  1. Randomly divide the \(n\) individuals into an estimation sample and a training sample, each of size \(n/2\).
  2. Apply machine learning to the training sample to estimate \(\hat b(x)\) and \(\hat\pi(x)\).
  3. Compute the doubly robust estimator in the estimation sample, using \(\hat b\) and \(\hat\pi\) from the training sample (Hernán and Robins 2020, 248).

The training sample is also called the nuisance sample, since it estimates the nuisance regressions (Fine Point 15.1). Sample splitting and cross-fitting are not new; combining them with machine learning has only recently been emphasized (Hernán and Robins 2020, 248).


4.4 Why Split? The Full-Sample AIPW Estimator

The full-sample augmented IP weighted (AIPW) estimator of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Technical Point 13.2) is

\[ \frac{1}{n} \sum_{i=1}^n\mathopen{}\left\{\hat b(X_i) + \frac{A_i}{\hat\pi(X_i)} \mathopen{}\left[Y_i - \hat b(X_i)\right]\mathclose{}\right\}\mathclose{}, \]

with \(\hat b\) and \(\hat\pi\) fit by machine learning on all \(n\) individuals.

  • Then \(\hat b\) and \(\hat\pi\) are correlated with the estimator itself, which can affect its bias, variance, and asymptotic normality in unpredictable ways, by an unknown amount.
  • In the split-sample version, the estimation sample is independent of \(\hat b\) and \(\hat\pi\), so under weak conditions (Technical Point 18.2) the estimator is asymptotically normal with standard error of order \(1/\sqrt{n}\) and only the product bias. If the product bias is below \(1/\sqrt{n}\), Wald intervals are valid (Hernán and Robins 2020, 248).

4.5 Cross-Fitting

Sample splitting’s cost: the standard error corresponds to a sample of size \(n/2\), so intervals are wider. Cross-fitting recovers the lost efficiency (Hernán and Robins 2020, 248–49):

  1. Swap the roles: the former estimation half becomes the training sample and vice versa, and recompute the doubly robust estimate.
  2. Average the two estimates to estimate the effect in the whole study population.
  3. Get a 95% confidence interval by bootstrapping (estimate \(\pm 1.96\) bootstrap standard errors, or the 2.5 and 97.5 percentiles of the bootstrap estimates).

NoteTechnical Point 18.1: Augmented IP Weighted Split-Sample and Cross-Fit Estimator

For \(\psi = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) (Hernán and Robins 2020, 249):

  1. Randomly split the \(n\) subjects into an estimation sample of size \(q\) and a training sample of size \(n_{tr} = n - q\), with \(q/n \approx 1/2\).
  2. From the training sample, estimate \(b(x) = \operatorname{E}\mathopen{}\left[Y \mid A = 1, X = x\right]\mathclose{}\) and \(\pi(x) = \Pr[A = 1 \mid X = x]\) by machine learning.
  3. In the estimation sample, compute \[ \hat\psi = \frac{1}{q} \sum_{i=1}^{q} \mathopen{}\left\{\hat b(X_i) + \frac{A_i}{\hat\pi(X_i)} \mathopen{}\left[Y_i - \hat b(X_i)\right]\mathclose{}\right\}\mathclose{}. \]
  4. Compute \(\hat\psi_{\text{cross-fit}} = (\hat\psi + \tilde\psi)/2\), where \(\tilde\psi\) is \(\hat\psi\) with the training and estimation samples swapped.

A version with better finite-sample behavior uses \(M > 2\) equal-sized random subsamples: compute \(\hat\psi^{(m)}\) using subsample \(m\) for estimation and the other \(M - 1\) for training, and set \(\hat\psi_{\text{cross-fit}} = \frac{1}{M} \sum_{m=1}^{M} \hat\psi^{(m)}\).

Open problems: detecting whether the bias of split-sample doubly robust estimators is as large as the standard error, and reducing it without redoing the machine learning. Lin et al. (2020) built estimators based on higher-order influence functions with smaller bias than doubly robust cross-fit estimators and without much larger variance (Hernán and Robins 2020, 249).


NoteTechnical Point 18.2: Statistical Properties of Split-Sample and Cross-Fit Estimators

Conditional on the training data \(Tr\), \(\hat b\) and \(\hat\pi\) are fixed functions, so \(\hat\psi\) is an average of independent and identically distributed terms and, by the central limit theorem, is asymptotically normal given \(Tr\) with standard error proportional to \(n^{-1/2}\).

If (i) \(\hat b\) and \(\hat\pi\) are consistent in mean square and (ii) the conditional bias divided by the standard error converges to 0 in probability, then \(\hat\psi\) and \(\hat\psi_{\text{cross-fit}}\) are asymptotically normal and unbiased, and their Wald intervals (with bootstrap standard errors) are valid and calibrated. \(\hat\psi_{\text{cross-fit}}\) is then semiparametric efficient, with standard error smaller than that of \(\hat\psi\) by a factor of \(1/\sqrt{2}\) (Hernán and Robins 2020, 250).


Proposition 1 (Conditional Bias of the Split-Sample AIPW Estimator) Assume conditional exchangeability given \(X\) and positivity, so \(\psi = \operatorname{E}\mathopen{}\left[b(X)\right]\mathclose{}\). Then

\[ \operatorname{E}\mathopen{}\left[\hat\psi - \psi \mid Tr\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\pi(X) \mathopen{}\left(\frac{1}{\pi(X)} - \frac{1}{\hat\pi(X)}\right)\mathclose{} \mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mid Tr\right]\mathclose{} . \]

Proof. Given \(Tr\), each term of \(\hat\psi\) has the same mean, so \(\operatorname{E}\mathopen{}\left[\hat\psi \mid Tr\right]\mathclose{}\) is the mean of one term. Conditioning on \(X\) first:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\frac{A}{\hat\pi(X)} \mathopen{}\left[Y - \hat b(X)\right]\mathclose{} \mid X, Tr\right]\mathclose{} &= \frac{\Pr[A = 1 \mid X]}{\hat\pi(X)} \operatorname{E}\mathopen{}\left[Y - \hat b(X) \mid A = 1, X, Tr\right]\mathclose{} && \text{(only } A = 1 \text{ contributes)} \\ &= \frac{\pi(X)}{\hat\pi(X)} \mathopen{}\left(b(X) - \hat b(X)\right)\mathclose{}. && (\text{estimation sample independent of } Tr) \end{aligned} \]

So

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\hat\psi - \psi \mid Tr\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\hat b(X) + \frac{\pi(X)}{\hat\pi(X)} \mathopen{}\left(b(X) - \hat b(X)\right)\mathclose{} - b(X) \mid Tr\right]\mathclose{} && (\psi = \operatorname{E}\mathopen{}\left[b(X)\right]\mathclose{}) \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mathopen{}\left(1 - \frac{\pi(X)}{\hat\pi(X)}\right)\mathclose{} \mid Tr\right]\mathclose{} && \text{(collect terms)} \\ &= \operatorname{E}\mathopen{}\left[\pi(X) \mathopen{}\left(\frac{1}{\pi(X)} - \frac{1}{\hat\pi(X)}\right)\mathclose{} \mathopen{}\left(\hat b(X) - b(X)\right)\mathclose{} \mid Tr\right]\mathclose{} . && \text{(factor out } \pi(X)) \end{aligned} \]

The bias is a product of the two errors, the second-order property used in this section. (The book writes the same product with both factors negated, \(\operatorname{E}\mathopen{}\left[\pi(X) \mathopen{}\left(\frac{1}{\hat\pi(X)} - \frac{1}{\pi(X)}\right)\mathclose{} \mathopen{}\left(b(X) - \hat b(X)\right)\mathclose{} \mid Tr\right]\mathclose{}\).)

Rates: if \(1/\hat\pi(x) - 1/\pi(x)\) converges at rate \(n^{-\alpha}\) and \(\hat b(x) - b(x)\) at rate \(n^{-\epsilon}\), the conditional bias is \(o(n^{-1/2})\) when \(\alpha + \epsilon > 1/2\). So \(\hat b\) may converge more slowly than \(n^{-1/4}\) if \(\hat\pi\) converges fast enough, and vice versa (Hernán and Robins 2020, 250).

The efficient standard error of \(\hat\psi_{\text{cross-fit}}\) (scaled by \(n^{1/2}\)) is \(\mathopen{}\left\{\operatorname{Var}\mathopen{}\left(b(X) + \frac{A}{\pi(X)} \mathopen{}\left[Y - b(X)\right]\mathclose{}\right)\mathclose{}\right\}\mathclose{}^{1/2}\).

5 18.5 Variable Selection Is a Difficult Problem (pp. 250-251)


Doubly robust machine learning refutes the belief that any data-adaptive selection of adjustment variables must give incorrect confidence intervals. But it does not solve every problem, for at least three reasons (Hernán and Robins 2020, 250):

  1. Knowledge: subject-matter knowledge may be insufficient to include all important confounders or to exclude bias-inducing or bias-amplifying variables, so small bias is not guaranteed.
  2. Implementation: doubly robust estimators have been hard (and, with machine learning, computationally expensive) to implement for high-dimensional time-varying treatments, especially in survival analysis. Most published longitudinal analyses use singly robust estimators, the ones emphasized in Part III, although these methods are becoming routine in some fields.
  3. Variance: the estimator can reach the variance we would have if \(b(X)\) and \(\pi(X)\) were known, but that variance may still be too large to be useful.

5.1 The Variance Problem

  • If some covariates in \(X\) are strongly associated with \(A\), \(\pi(X)\) can be near 0 or 1 for some individuals, giving a very wide (but often correct) 95% confidence interval, even with small product bias.
  • Discarding the covariates that cause the “problem” after seeing the data changes the game: dropping covariates that are associated with treatment but not much with the outcome makes it impossible to guarantee that the interval is valid.
  • The tension between including all potential confounders (less bias) and excluding some (less variance) is hard to resolve (Hernán and Robins 2020, 251).

The book poses a puzzle: if the interval is invalid when we use the data to discard, say, 5 variables that make the variance too large, why would it be valid had we received a data set that never included those 5 variables? Since some potential confounders are always unrecorded, how should we interpret confidence intervals from any observational analysis (Hernán and Robins 2020, 251)?


5.2 Practical Advice

General, prescriptive guidelines for variable selection may not be possible while so much research is ongoing. As in Section 13.5, the book’s advice is to run multiple sensitivity analyses: apply several analytic methods and compare the estimates. If they agree, we gain confidence; if they disagree, our job is to understand why (Hernán and Robins 2020, 251).

6 Summary


  • Prediction needs no confounding adjustment and may use any algorithm that predicts well; causal inference needs a carefully chosen adjustment set.
  • Adjusting for some variables induces bias (colliders and their descendants, noncolliders affected by treatment, mediators, M-bias structures) or amplifies it (instruments; Z-bias). Data alone cannot find them, so expert knowledge must guide selection.
  • Temporal order alone is not the rule: a post-treatment variable can be a valid adjustment variable, and a pre-treatment variable can be a collider.
  • With high-dimensional \(X\), estimate \(b(X)\) and \(\pi(X)\) with machine learning inside a doubly robust estimator, with sample splitting and cross-fitting, to get valid Wald intervals.
  • Valid intervals may still be very wide; dropping covariates after seeing the data threatens validity. Run multiple sensitivity analyses.
Back to top

References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.