This chapter estimates the same causal effect as Chapter 12 (smoking cessation on weight gain in NHEFS) by standardization, now combined with outcome models: the parametric g-formula. IP weighting models the treatment (and censoring); standardization models the outcome. Both rely on the same identifiability conditions but on different modeling assumptions.
Remark 1 (NHEFS: Setting and Target Effect). Setting (same as Chapter 12):
Causal effect of interest:
\[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\]
the difference in mean weight gain had everybody been treated and uncensored versus untreated and uncensored.
Definition 1 (Standardized mean) The standardized mean outcome in the uncensored who received treatment level \(a\) is
\[\sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \times \Pr[L = l],\]
the average of the stratum-specific means, weighted by the prevalence of each stratum \(l\) in the study population. When some components of \(L\) are continuous, \(\Pr[L = l]\) is replaced by the density \(f_L(l)\) and the sum becomes an integral, \(\int \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \, dF_L(l)\), where \(F_L\) is the joint cdf of \(L\).
Theorem 1 (Standardization identifies the counterfactual mean) Let \(L\) be discrete, and fix a treatment level \(a\). Assume
Then the standardized mean of Definition 1 equals \(\operatorname{E}\mathopen{}\left[Y^{a,c=0}\right]\mathclose{}\).
Proof. For discrete \(L\) (Technical Point 2.3 of the book, with \(A = a\) replaced by \((A = a, C = 0)\)):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a,c=0}\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] && \text{(exchangeability for } (A, C) \text{ given } L\text{; positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] && \text{(consistency)} \end{aligned} \]
Positivity is needed in the second line so that the conditional mean given \((A = a, C = 0, L = l)\) is defined for every \(l\) with \(\Pr[L = l] > 0\).
The same argument, with sums replaced by integrals, covers continuous components of \(L\). So, under these conditions, any estimator that converges to the standardized mean with \(a = 1\) also converges to \(\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{}\); likewise with \(a = 0\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\).
Remark 2 (Standardizing to the treated). The average causal effect in the treated can also be estimated by standardization (book’s Technical Point 4.1): replace \(\Pr[L = l]\) by \(\Pr[L = l \mid A = 1]\) in Definition 1.
Fine Point 13.1: Structural Positivity
Remark 3 (Why we model the outcome). Problem: nonparametric estimation of \(\operatorname{E}\mathopen{}\left[Y \mid A = 1, C = 0, L = l\right]\mathclose{}\) (as in Section 2.3) is out of the question. Only 403 uncensored treated individuals are spread across possibly millions of strata of the 9 confounders. Same for the untreated.
Solution: fit a parametric model for the mean outcome.
Example 1 (NHEFS outcome model) A linear regression model for \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, L\right]\mathclose{}\) (among the uncensored) with covariates:
The model assumes each covariate contributes additively to the mean, except that the contribution of \(A\) varies linearly with smoking intensity.
Remark 4 (Other outcome and treatment types). The same approach applies with other outcome models:
We do not need to estimate \(\Pr[L = l]\).
Proposition 1 (The standardized mean is an iterated expectation) For discrete \(L\),
\[\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] = \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L\right]\mathclose{}\right]\mathclose{}\]
Proof. The outer expectation is over the marginal distribution of \(L\). Writing \(g(l) = \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{}\), we have \(\operatorname{E}\mathopen{}\left[g(L)\right]\mathclose{} = \sum_l g(l) \Pr[L = l]\) by the definition of expectation.
By Proposition 1, the standardized mean can be estimated by averaging the model predictions over the empirical distribution of \(L\):
\[\frac{1}{n} \sum_{i=1}^n\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid A = a, C = 0, L = L_i\right]\mathclose{}\]
where \(n\) is the number of individuals in the study.
Definition 2 (Standardization by expanding the dataset)
Example 2 (Standardization of the Table 2.2 data) The 20-row Table 2.2 becomes 60 rows. With a saturated model (main effects of \(A\) and \(L\) plus \(A \times L\)), the predictions equal the stratum-specific risks: \(\Pr[Y = 1 \mid A = a, L = 0] = 1/4\) and \(\Pr[Y = 1 \mid A = a, L = 1] = 2/3\) for both \(a = 0\) and \(a = 1\). Since 60% of rows have \(L = 1\) and 40% have \(L = 0\), each block’s average prediction is
\[\frac{1}{4} \times 0.4 + \frac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5,\]
exactly the standardized risks of Section 2.3 (0.5 under both \(a = 0\) and \(a = 1\)).
Example 3 (Parametric g-formula in NHEFS) Applying the four steps with the outcome model of Example 1 (book’s Program 13.3):
With a nonparametric bootstrap (Program 13.4), the 95% confidence interval is 2.6 to 4.5 kg.
Technical Point 13.1: Bootstrapping
The nonparametric bootstrap estimates standard errors when analytic or large-sample formulas are impractical:
Remark 5 (IP weighting and standardization model different things).
Definition 3 (Parametric g-formula) IP weighting and standardization each estimate the g-formula (Robins 1986). Standardization is a plug-in g-formula estimator: it substitutes an estimate for the conditional mean outcome that appears in the g-formula. When that estimate comes from a parametric model, the method is the parametric g-formula.
Remark 6 (What the parametric g-formula needs to model).
Definition 4 (Doubly robust estimator) An estimator of a counterfactual mean such as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) that combines a treatment model and an outcome model is doubly robust if, under the identifying conditions (conditional exchangeability and positivity given \(L\), and consistency), it is consistent for that counterfactual mean when either model is correctly specified, without our knowing which. (With censoring, replace \(Y^{a}\) by \(Y^{a,c=0}\) and \(A\) by the joint treatment \((A, C)\) throughout.)
An estimator of a causal contrast, such as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), built from doubly robust estimators (Definition 4) of each counterfactual mean is then consistent for that contrast, provided the identifying conditions hold for every treatment level in the contrast (here both \(a = 1\) and \(a = 0\)) and, for each counterfactual mean, at least one of its two models is correctly specified.
Use Both Methods, and Doubly Robust Ones
Fine Point 13.2: A Doubly Robust Plug-In Estimator
For dichotomous \(A\) and no censoring (Bang and Robins 2005):
Definition 5 (Augmented IP weighted (AIPW) estimator) Let \(A\) be dichotomous, with no censoring, and define
Let \(\hat b(L)\) be the predicted mean outcome under \(A = 1\) from a fitted model for \(\operatorname{E}\mathopen{}\left[Y \mid A, L\right]\mathclose{}\) (the outcome model), and let \(\hat\pi(L)\) be the predicted probability of \(A = 1\) from a fitted model for \(\Pr[A = 1 \mid L]\) (the treatment model). The AIPW estimator of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) is
\[ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a=1}\right]\mathclose{}_{DR} = \frac{1}{n} \sum_{i=1}^n\mathopen{}\left[\hat b(L_i) + \frac{A_i}{\hat\pi(L_i)} \mathopen{}\left\{Y_i - \hat b(L_i)\right\}\mathclose{}\right]\mathclose{} \]
Technical Point 13.2: The Augmented IP Weighted Estimator
For dichotomous \(A\) and no censoring, with \(b\), \(\pi\), \(\hat b\) and \(\hat\pi\) as in Definition 5, under the identifiability conditions,
\[\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\frac{AY}{\pi(L)}\right]\mathclose{}.\]
Theorem 2 (Double robustness of the AIPW functional) Let \(A\) be dichotomous, with no censoring, and let \(b\) and \(\pi\) be as in Definition 5. Assume
Let \(b^*\) and \(\pi^*\) be any functions of \(L\) with \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|b^*(L)\right|\mathclose{}\right]\mathclose{} < \infty\) and \(\pi^*(l) \geq \epsilon\) for some \(\epsilon > 0\) and every \(l\) in the support of \(L\). Then
\[ \operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\pi(L) \mathopen{}\left(\frac{1}{\pi^*(L)} - \frac{1}{\pi(L)}\right)\mathclose{} \mathopen{}\left(b(L) - b^*(L)\right)\mathclose{}\right]\mathclose{}, \]
which is zero if \(\pi^* = \pi\) or \(b^* = b\).
Proof. First condition on \(L\):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\frac{A\{Y - b^*(L)\}}{\pi^*(L)} \,\middle|\, L\right]\mathclose{} &= \frac{\Pr[A = 1 \mid L] \, \operatorname{E}\mathopen{}\left[Y - b^*(L) \mid A = 1, L\right]\mathclose{}}{\pi^*(L)} && \text{(only } A = 1 \text{ contributes)} \\ &= \frac{\pi(L) \{b(L) - b^*(L)\}}{\pi^*(L)} && \text{(definitions of } \pi, b\text{)} \end{aligned} \]
Next, identify the counterfactual mean:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{}\right]\mathclose{} && \text{(iterated expectation)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^{a=1} \mid A = 1, L\right]\mathclose{}\right]\mathclose{} && \text{(exchangeability; positivity, so the conditioning is defined)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid A = 1, L\right]\mathclose{}\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(definition of } b\text{)} \end{aligned} \]
Each of the three terms \(b^*(L)\), \(A\{Y - b^*(L)\}/\pi^*(L)\) and \(b(L)\) is integrable: \(b^*(L)\) by assumption; \(A\{Y - b^*(L)\}/\pi^*(L)\) because its absolute value is at most \(\mathopen{}\left(\mathopen{}\left|Y\right|\mathclose{} + \mathopen{}\left|b^*(L)\right|\mathclose{}\right)\mathclose{}/\epsilon\); and \(b(L)\) because the steps of the identification display hold conditionally on \(L\), giving \(b(L) = \operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{}\), so \(\mathopen{}\left|b(L)\right|\mathclose{} \leq \operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a=1}\right|\mathclose{} \mid L\right]\mathclose{}\), whose expectation \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a=1}\right|\mathclose{}\right]\mathclose{}\) is finite. Then:
\[ \begin{aligned} &\operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(identification display)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L)\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(linearity; each term integrable)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L)\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\frac{\pi(L)\{b(L) - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(iterated expectation and the first display)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L) - b(L) + \frac{\pi(L)}{\pi^*(L)}\{b(L) - b^*(L)\}\right]\mathclose{} && \text{(linearity)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\frac{\pi(L)}{\pi^*(L)} - 1\right)\mathclose{} \{b(L) - b^*(L)\}\right]\mathclose{} && \text{(factor out } b(L) - b^*(L)\text{)} \\ &= \operatorname{E}\mathopen{}\left[\pi(L) \mathopen{}\left(\frac{1}{\pi^*(L)} - \frac{1}{\pi(L)}\right)\mathclose{} \{b(L) - b^*(L)\}\right]\mathclose{} && \text{(factor out } \pi(L)\text{)} \end{aligned} \]
If \(\pi^* = \pi\), the factor \(1/\pi^*(L) - 1/\pi(L)\) is 0; if \(b^* = b\), the factor \(b(L) - b^*(L)\) is 0. Either way the expectation is zero.
Remark 7 (From the functional to the estimator). Suppose \(\hat b\) and \(\hat\pi\) are predictions from parametric models whose parameter estimates converge in probability, so that \(\hat b\) and \(\hat\pi\) settle down to fixed functions \(b^*\) and \(\pi^*\), with \(\pi^*\) bounded away from 0. Then, under the usual regularity conditions, the AIPW estimator of Definition 5 converges in probability to the functional in Theorem 2, evaluated at these limits. A correctly specified outcome model gives \(b^* = b\), and a correctly specified treatment model gives \(\pi^* = \pi\); either one makes the difference in Theorem 2 vanish. So the AIPW estimator is consistent for \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) whenever at least one of the two models is correct: it is doubly robust in the sense of Definition 4.
Technical Point 13.3: The Augmented IP Weighted Estimator and the Doubly Robust Plug-In Estimator
This technical point shows that the doubly robust plug-in estimator of Fine Point 13.2 is the AIPW estimator in which the outcome model is fit so that the augmentation term sums to zero in every sample (a TMLE). Using two clever covariates, \(A/\hat\pi(L)\) and \((1-A)/\{1 - \hat\pi(L)\}\), also makes each plug-in counterfactual mean doubly robust.
Example 4 (NHEFS: First Estimates from Real Data) Our first estimates from real data, by IP weighting and by the parametric g-formula:
Agreement across methods (and, in the next chapters, g-estimation, outcome regression, and propensity scores) is reassuring because each relies on different modeling assumptions. But observational estimates remain open to serious criticism.
Remark 8 (Conditions for a causal interpretation).
The Conditions Are Not Testable
Key concepts introduced: