Chapter 13: Standardization and the Parametric G-Formula

This chapter estimates the same causal effect as Chapter 12 (smoking cessation on weight gain in NHEFS) by standardization, now combined with outcome models: the parametric g-formula. IP weighting models the treatment (and censoring); standardization models the outcome. Both rely on the same identifiability conditions but on different modeling assumptions.

1 13.1 Standardization as an Alternative to IP Weighting (pp. 175-177)

Remark 1 (NHEFS: Setting and Target Effect). Setting (same as Chapter 12):

  • NHEFS: 1629 cigarette smokers aged 25-74 with a baseline visit and a follow-up visit about 10 years later
  • 1566 had weight measured at follow-up, i.e., are uncensored (\(C = 0\))
  • Treatment \(A\): smoking cessation (1: yes, 0: no); outcome \(Y\): weight gain (kg)

Causal effect of interest:

\[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\]

the difference in mean weight gain had everybody been treated and uncensored versus untreated and uncensored.

The Standardized Mean

Definition 1 (Standardized mean) The standardized mean outcome in the uncensored who received treatment level \(a\) is

\[\sum_{l} \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \times \Pr[L = l],\]

the average of the stratum-specific means, weighted by the prevalence of each stratum \(l\) in the study population. When some components of \(L\) are continuous, \(\Pr[L = l]\) is replaced by the density \(f_L(l)\) and the sum becomes an integral, \(\int \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \, dF_L(l)\), where \(F_L\) is the joint cdf of \(L\).

Theorem 1 (Standardization identifies the counterfactual mean) Let \(L\) be discrete, and fix a treatment level \(a\). Assume

  • conditional exchangeability for the joint treatment \((A, C)\): \(Y^{a,c=0} \perp\!\!\!\perp(A, C) \mid L\);
  • positivity: \(\Pr[A = a, C = 0 \mid L = l] > 0\) for every \(l\) with \(\Pr[L = l] > 0\);
  • consistency: \(Y^{a,c=0} = Y\) for every individual with \(A = a\) and \(C = 0\).

Then the standardized mean of Definition 1 equals \(\operatorname{E}\mathopen{}\left[Y^{a,c=0}\right]\mathclose{}\).

Proof. For discrete \(L\) (Technical Point 2.3 of the book, with \(A = a\) replaced by \((A = a, C = 0)\)):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a,c=0}\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] && \text{(exchangeability for } (A, C) \text{ given } L\text{; positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] && \text{(consistency)} \end{aligned} \]

Positivity is needed in the second line so that the conditional mean given \((A = a, C = 0, L = l)\) is defined for every \(l\) with \(\Pr[L = l] > 0\).

The same argument, with sums replaced by integrals, covers continuous components of \(L\). So, under these conditions, any estimator that converges to the standardized mean with \(a = 1\) also converges to \(\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{}\); likewise with \(a = 0\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\).

Supplement: Standardizing to the Treated

Remark 2 (Standardizing to the treated). The average causal effect in the treated can also be estimated by standardization (book’s Technical Point 4.1): replace \(\Pr[L = l]\) by \(\Pr[L = l \mid A = 1]\) in Definition 1.

Fine Point 13.1: Structural Positivity

  • Positivity is needed for standardization too: if \(\Pr[A = a, C = 0 \mid L = l] = 0\) while \(\Pr[L = l] \neq 0\), then \(\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{}\) is undefined.
  • With a parametric outcome model, one can ignore structural nonpositivity by extrapolating over the empty strata, at the cost of bias (95% confidence intervals then cover the truth less than 95% of the time).
  • Near positivity violations, standardization typically has a smaller standard error than IP weighting, but differences in bias may outweigh differences in standard error.

2 13.2 Estimating the Mean Outcome via Modeling (pp. 177-178)

Remark 3 (Why we model the outcome). Problem: nonparametric estimation of \(\operatorname{E}\mathopen{}\left[Y \mid A = 1, C = 0, L = l\right]\mathclose{}\) (as in Section 2.3) is out of the question. Only 403 uncensored treated individuals are spread across possibly millions of strata of the 9 confounders. Same for the untreated.

Solution: fit a parametric model for the mean outcome.

Example 1 (NHEFS outcome model) A linear regression model for \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, L\right]\mathclose{}\) (among the uncensored) with covariates:

  • treatment \(A\) and all 9 confounders in \(L\)
  • linear and quadratic terms for age, weight, smoking intensity, and smoking duration
  • a product term \(A \times\) smoking intensity

The model assumes each covariate contributes additively to the mean, except that the contribution of \(A\) varies linearly with smoking intensity.

Supplement: Other Outcome and Treatment Types

Remark 4 (Other outcome and treatment types). The same approach applies with other outcome models:

  • Dichotomous \(Y\): fit, e.g., a logistic model for \(\Pr[Y = 1 \mid A, C = 0, L]\); the standardized means are then standardized risks, from which a risk difference, risk ratio, or odds ratio can be computed.
  • Nondichotomous \(A\): standardize the predicted means at each treatment level \(a\) of interest.

3 13.3 Standardizing the Mean Outcome to the Confounder Distribution (pp. 178-179)

We do not need to estimate \(\Pr[L = l]\).

Proposition 1 (The standardized mean is an iterated expectation) For discrete \(L\),

\[\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] = \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L\right]\mathclose{}\right]\mathclose{}\]

Proof. The outer expectation is over the marginal distribution of \(L\). Writing \(g(l) = \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{}\), we have \(\operatorname{E}\mathopen{}\left[g(L)\right]\mathclose{} = \sum_l g(l) \Pr[L = l]\) by the definition of expectation.

By Proposition 1, the standardized mean can be estimated by averaging the model predictions over the empirical distribution of \(L\):

\[\frac{1}{n} \sum_{i=1}^n\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid A = a, C = 0, L = L_i\right]\mathclose{}\]

where \(n\) is the number of individuals in the study.

The Four-Step Algorithm

Definition 2 (Standardization by expanding the dataset)  

  1. Expansion: make three copies of the data. Block 1 is the original data. In block 2 set \(A = 0\) for everyone; in block 3 set \(A = 1\) for everyone. Set \(Y\) to missing in blocks 2 and 3.
  2. Outcome modeling: fit the outcome model on the 3-block dataset. Only block 1 contributes, because \(Y\) is missing elsewhere.
  3. Prediction: use the fitted model to predict \(Y\) for every row of blocks 2 and 3.
  4. Standardization by averaging: the mean prediction in block 2 estimates the standardized mean under \(A = 0\); in block 3, under \(A = 1\).

Example: Table 2.2 (No Censoring, One Binary \(L\))

Example 2 (Standardization of the Table 2.2 data) The 20-row Table 2.2 becomes 60 rows. With a saturated model (main effects of \(A\) and \(L\) plus \(A \times L\)), the predictions equal the stratum-specific risks: \(\Pr[Y = 1 \mid A = a, L = 0] = 1/4\) and \(\Pr[Y = 1 \mid A = a, L = 1] = 2/3\) for both \(a = 0\) and \(a = 1\). Since 60% of rows have \(L = 1\) and 40% have \(L = 0\), each block’s average prediction is

\[\frac{1}{4} \times 0.4 + \frac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5,\]

exactly the standardized risks of Section 2.3 (0.5 under both \(a = 0\) and \(a = 1\)).

Example: NHEFS

Example 3 (Parametric g-formula in NHEFS) Applying the four steps with the outcome model of Example 1 (book’s Program 13.3):

  • standardized mean in the untreated (block 2): 1.66 kg
  • standardized mean in the treated (block 3): 5.18 kg
  • estimated effect: \(5.18 - 1.66 = 3.52 \approx 3.5\) kg

With a nonparametric bootstrap (Program 13.4), the 95% confidence interval is 2.6 to 4.5 kg.

Technical Point 13.1: Bootstrapping

The nonparametric bootstrap estimates standard errors when analytic or large-sample formulas are impractical:

  1. Draw a sample of size 1629 with replacement from the 1629 individuals (a bootstrap sample).
  2. Compute the effect estimate in the bootstrap sample with the same method (here, standardization).
  3. Repeat many times (the book uses 1000).
  4. The standard deviation of the bootstrap estimates estimates the standard error; the 95% CI is the estimate \(\pm 1.96 \times\) that standard error.

4 13.4 IP Weighting or Standardization? (pp. 179-180)

Remark 5 (IP weighting and standardization model different things).

  • Without models, the IP weighted and standardized means are exactly equal (Technical Point 2.3).
  • With models, they generally differ, because they model different things:
    • IP weighting: \(\Pr[A = a, C = 0 \mid L]\) (in Chapter 12, logistic models for \(\Pr[A = a \mid L]\) and \(\Pr[C = 0 \mid A = a, L]\))
    • Standardization: \(\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L\right]\mathclose{}\) (here, a linear model)
  • Some misspecification is inescapable, and it will usually bias the two estimates differently.

The (Parametric) G-Formula

Definition 3 (Parametric g-formula) IP weighting and standardization each estimate the g-formula (Robins 1986). Standardization is a plug-in g-formula estimator: it substitutes an estimate for the conditional mean outcome that appears in the g-formula. When that estimate comes from a parametric model, the method is the parametric g-formula.

Remark 6 (What the parametric g-formula needs to model).

  • For an average causal effect we only need the conditional mean outcome; for a standardized distribution (pdf) we would need the conditional distribution of \(Y\) given \(A\) and \(L\).
  • Without time-varying confounders, the parametric g-formula does not require modeling the distribution of \(L\).

Use Both, and Use Doubly Robust Methods

Definition 4 (Doubly robust estimator) An estimator of a counterfactual mean such as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) that combines a treatment model and an outcome model is doubly robust if, under the identifying conditions (conditional exchangeability and positivity given \(L\), and consistency), it is consistent for that counterfactual mean when either model is correctly specified, without our knowing which. (With censoring, replace \(Y^{a}\) by \(Y^{a,c=0}\) and \(A\) by the joint treatment \((A, C)\) throughout.)

An estimator of a causal contrast, such as \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\), built from doubly robust estimators (Definition 4) of each counterfactual mean is then consistent for that contrast, provided the identifying conditions hold for every treatment level in the contrast (here both \(a = 1\) and \(a = 0\)) and, for each counterfactual mean, at least one of its two models is correctly specified.

Use Both Methods, and Doubly Robust Ones

  • When both IP weighting and the parametric g-formula can be used, use both.
  • Whenever possible, use doubly robust estimators (Definition 4).
  • Both methods can also target a subset of the population, e.g., estimating standardized means separately in men and women to study effect modification by sex.

Fine Point 13.2: A Doubly Robust Plug-In Estimator

For dichotomous \(A\) and no censoring (Bang and Robins 2005):

  1. Estimate the IP weights \(W^A = 1/f(A \mid L)\) as in Chapter 12.
  2. Fit an outcome model (a GLM with canonical link) for \(\operatorname{E}\mathopen{}\left[Y \mid A, L, R\right]\mathclose{}\) that adds the covariate \(R = W^A\) if \(A = 1\) and \(R = -W^A\) if \(A = 0\).
  3. Predict for everyone with \(A\) set to 1 and \(R\) recomputed at that value, \(R = 1/f(1 \mid L)\), and average; repeat with \(A\) set to 0 and \(R = -1/f(0 \mid L)\).
  4. The difference of the two averages is a doubly robust plug-in estimate of the average causal effect.

Definition 5 (Augmented IP weighted (AIPW) estimator) Let \(A\) be dichotomous, with no censoring, and define

  • the outcome regression \(b(L) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y \mid A = 1, L\right]\mathclose{}\);
  • the propensity score \(\pi(L) \stackrel{\text{def}}{=}\Pr[A = 1 \mid L]\).

Let \(\hat b(L)\) be the predicted mean outcome under \(A = 1\) from a fitted model for \(\operatorname{E}\mathopen{}\left[Y \mid A, L\right]\mathclose{}\) (the outcome model), and let \(\hat\pi(L)\) be the predicted probability of \(A = 1\) from a fitted model for \(\Pr[A = 1 \mid L]\) (the treatment model). The AIPW estimator of \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) is

\[ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a=1}\right]\mathclose{}_{DR} = \frac{1}{n} \sum_{i=1}^n\mathopen{}\left[\hat b(L_i) + \frac{A_i}{\hat\pi(L_i)} \mathopen{}\left\{Y_i - \hat b(L_i)\right\}\mathclose{}\right]\mathclose{} \]

Technical Point 13.2: The Augmented IP Weighted Estimator

For dichotomous \(A\) and no censoring, with \(b\), \(\pi\), \(\hat b\) and \(\hat\pi\) as in Definition 5, under the identifiability conditions,

\[\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\frac{AY}{\pi(L)}\right]\mathclose{}.\]

  • Plug-in g-formula estimator: \(\frac{1}{n}\sum_i \hat b(L_i)\)
  • Horvitz-Thompson IP weighted estimator: \(\frac{1}{n}\sum_i A_i Y_i / \hat\pi(L_i)\)

Why the AIPW Estimator Is Doubly Robust

Theorem 2 (Double robustness of the AIPW functional) Let \(A\) be dichotomous, with no censoring, and let \(b\) and \(\pi\) be as in Definition 5. Assume

  • conditional exchangeability: \(Y^{a=1} \perp\!\!\!\perp A \mid L\);
  • positivity: \(\pi(l) > 0\) for every \(l\) in the support of \(L\);
  • consistency: \(Y^{a=1} = Y\) for every individual with \(A = 1\);
  • \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a=1}\right|\mathclose{}\right]\mathclose{} < \infty\).

Let \(b^*\) and \(\pi^*\) be any functions of \(L\) with \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|b^*(L)\right|\mathclose{}\right]\mathclose{} < \infty\) and \(\pi^*(l) \geq \epsilon\) for some \(\epsilon > 0\) and every \(l\) in the support of \(L\). Then

\[ \operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[\pi(L) \mathopen{}\left(\frac{1}{\pi^*(L)} - \frac{1}{\pi(L)}\right)\mathclose{} \mathopen{}\left(b(L) - b^*(L)\right)\mathclose{}\right]\mathclose{}, \]

which is zero if \(\pi^* = \pi\) or \(b^* = b\).

Proof. First condition on \(L\):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\frac{A\{Y - b^*(L)\}}{\pi^*(L)} \,\middle|\, L\right]\mathclose{} &= \frac{\Pr[A = 1 \mid L] \, \operatorname{E}\mathopen{}\left[Y - b^*(L) \mid A = 1, L\right]\mathclose{}}{\pi^*(L)} && \text{(only } A = 1 \text{ contributes)} \\ &= \frac{\pi(L) \{b(L) - b^*(L)\}}{\pi^*(L)} && \text{(definitions of } \pi, b\text{)} \end{aligned} \]

Next, identify the counterfactual mean:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{}\right]\mathclose{} && \text{(iterated expectation)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^{a=1} \mid A = 1, L\right]\mathclose{}\right]\mathclose{} && \text{(exchangeability; positivity, so the conditioning is defined)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid A = 1, L\right]\mathclose{}\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(definition of } b\text{)} \end{aligned} \]

Each of the three terms \(b^*(L)\), \(A\{Y - b^*(L)\}/\pi^*(L)\) and \(b(L)\) is integrable: \(b^*(L)\) by assumption; \(A\{Y - b^*(L)\}/\pi^*(L)\) because its absolute value is at most \(\mathopen{}\left(\mathopen{}\left|Y\right|\mathclose{} + \mathopen{}\left|b^*(L)\right|\mathclose{}\right)\mathclose{}/\epsilon\); and \(b(L)\) because the steps of the identification display hold conditionally on \(L\), giving \(b(L) = \operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{}\), so \(\mathopen{}\left|b(L)\right|\mathclose{} \leq \operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a=1}\right|\mathclose{} \mid L\right]\mathclose{}\), whose expectation \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a=1}\right|\mathclose{}\right]\mathclose{}\) is finite. Then:

\[ \begin{aligned} &\operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[b^*(L) + \frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(identification display)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L)\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\frac{A\{Y - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(linearity; each term integrable)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L)\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\frac{\pi(L)\{b(L) - b^*(L)\}}{\pi^*(L)}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[b(L)\right]\mathclose{} && \text{(iterated expectation and the first display)} \\ &= \operatorname{E}\mathopen{}\left[b^*(L) - b(L) + \frac{\pi(L)}{\pi^*(L)}\{b(L) - b^*(L)\}\right]\mathclose{} && \text{(linearity)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\frac{\pi(L)}{\pi^*(L)} - 1\right)\mathclose{} \{b(L) - b^*(L)\}\right]\mathclose{} && \text{(factor out } b(L) - b^*(L)\text{)} \\ &= \operatorname{E}\mathopen{}\left[\pi(L) \mathopen{}\left(\frac{1}{\pi^*(L)} - \frac{1}{\pi(L)}\right)\mathclose{} \{b(L) - b^*(L)\}\right]\mathclose{} && \text{(factor out } \pi(L)\text{)} \end{aligned} \]

If \(\pi^* = \pi\), the factor \(1/\pi^*(L) - 1/\pi(L)\) is 0; if \(b^* = b\), the factor \(b(L) - b^*(L)\) is 0. Either way the expectation is zero.

Remark 7 (From the functional to the estimator). Suppose \(\hat b\) and \(\hat\pi\) are predictions from parametric models whose parameter estimates converge in probability, so that \(\hat b\) and \(\hat\pi\) settle down to fixed functions \(b^*\) and \(\pi^*\), with \(\pi^*\) bounded away from 0. Then, under the usual regularity conditions, the AIPW estimator of Definition 5 converges in probability to the functional in Theorem 2, evaluated at these limits. A correctly specified outcome model gives \(b^* = b\), and a correctly specified treatment model gives \(\pi^* = \pi\); either one makes the difference in Theorem 2 vanish. So the AIPW estimator is consistent for \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\) whenever at least one of the two models is correct: it is doubly robust in the sense of Definition 4.

Technical Point 13.3: The Augmented IP Weighted Estimator and the Doubly Robust Plug-In Estimator

This technical point shows that the doubly robust plug-in estimator of Fine Point 13.2 is the AIPW estimator in which the outcome model is fit so that the augmentation term sums to zero in every sample (a TMLE). Using two clever covariates, \(A/\hat\pi(L)\) and \((1-A)/\{1 - \hat\pi(L)\}\), also makes each plug-in counterfactual mean doubly robust.

5 13.5 How Seriously Do We Take Our Estimates? (pp. 181-185)

Example 4 (NHEFS: First Estimates from Real Data) Our first estimates from real data, by IP weighting and by the parametric g-formula:

  • mean weight gain: 5.2 kg if everybody had quit, 1.7 kg if nobody had quit
  • effect: 3.5 kg by both methods (95% CI: 2.5, 4.5 for IP weighting; 2.6, 4.5 for the g-formula)

Agreement across methods (and, in the next chapters, g-estimation, outcome regression, and propensity scores) is reassuring because each relies on different modeling assumptions. But observational estimates remain open to serious criticism.

Three Groups of Conditions

Remark 8 (Conditions for a causal interpretation).

  1. Identifiability: exchangeability, positivity, and consistency, so that the observational study emulates the target trial.
  2. No measurement error in \(A\), \(Y\), or \(L\) (Chapter 9).
  3. No model misspecification (Chapter 11).

Sensitivity Analysis and Skepticism

The Conditions Are Not Testable

  • None of these conditions is empirically testable; we assume they hold approximately, based on expert knowledge.
  • Expert knowledge is incomplete, so analyses should explore sensitivity to the assumptions.

6 Summary

Key concepts introduced:

  1. Standardized mean: \(\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} \Pr[L = l] = \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L\right]\mathclose{}\right]\mathclose{}\)
  2. Parametric g-formula: estimate the conditional mean with a parametric model and average predictions over the empirical distribution of \(L\) (the four-step algorithm)
  3. NHEFS: standardized means 5.18 kg (quit) and 1.66 kg (did not quit); effect 3.5 kg, essentially the same as IP weighting
  4. Bootstrap: resample individuals with replacement to estimate the standard error
  5. IP weighting vs standardization: equal without models; with models they rely on different assumptions, so compare them
  6. Doubly robust estimators (Definition 4): consistent if either the treatment or the outcome model is correct
  7. Conditions for causal interpretation: identifiability, no measurement error, no model misspecification
Bang, Heejung, and James M Robins. 2005. “Doubly Robust Estimation in Missing Data and Causal Inference Models.” Biometrics 61: 962–73.
Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Robins, James M. 1986. “A New Approach to Causal Inference in Mortality Studies with a Sustained Exposure Period – Application to Control of the Healthy Worker Survivor Effect.” Mathematical Modelling 7: 1393–512.