This chapter uses inverse probability (IP) weighting to estimate the average causal effect of smoking cessation on weight gain from the observational NHEFS data. Chapter 2 introduced IP weighting as a nonparametric method; here IP weights are estimated with models, which lets us handle many covariates and nondichotomous treatments, and the weighted analysis is interpreted as fitting a marginal structural model.
Part II is organized around one question: on average, how much does quitting smoking change body weight?
Example 1 (NHEFS: Does Quitting Smoking Cause Weight Gain?) Population: 1566 cigarette smokers aged 25-74 years who had an NHEFS baseline visit (1971-75) and a follow-up visit about 10 years later (1982).
Treatment: \(A = 1\) if the individual reported quitting smoking before the follow-up visit, \(A = 0\) otherwise.
Outcome: \(Y\) = body weight at follow-up minus body weight at baseline (kg).
Causal estimand (average causal effect on the additive scale): \[\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\]
Example 2 (NHEFS: Unadjusted Comparison) Most people gained weight, but quitters gained more on average:
The associational difference \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) need not equal the causal difference if quitters and non-quitters differ in characteristics that affect weight gain.
| Mean baseline characteristic | Quitters (\(A=1\)) | Non-quitters (\(A=0\)) |
|---|---|---|
| Age, years | 46.2 | 42.8 |
| Men, % | 54.6 | 46.6 |
| White, % | 91.1 | 85.4 |
| University, % | 15.4 | 9.9 |
| Weight, kg | 72.4 | 70.3 |
| Cigarettes/day | 18.6 | 21.2 |
| Years smoking | 26.0 | 24.1 |
| Little exercise, % | 40.7 | 37.9 |
| Inactive life, % | 11.2 | 8.9 |
Example 3 (NHEFS: Measured Confounders) We assume that these 9 baseline variables, collected in the vector \(L\), suffice to adjust for confounding:
Assumption: conditional exchangeability given \(L\): \[Y^a \perp\!\!\!\perp A \mid L\]
Definition 1 (Inverse Probability Weights) The individual-specific IP weights for treatment \(A\) are
\[W^A \stackrel{\text{def}}{=}\frac{1}{f(A \mid L)}\]
i.e., the inverse of the conditional probability of receiving the treatment level the individual actually received.
For a dichotomous treatment:
So only \(\Pr[A = 1 \mid L]\) needs to be estimated.
IP weighting reweights individuals so that \(L\) no longer predicts \(A\). Chapter 2 did this by hand (IP weights in Chapter 2); the next definition makes the reweighted population precise for a general weight.
Definition 2 (Pseudo-Population) Let each of the \(n\) individuals of a study population carry a weight \(W \geq 0\) that is a function of the observed variables, with \(0 < \operatorname{E}\mathopen{}\left[W\right]\mathclose{} < \infty\). The pseudo-population created by \(W\) is the population in which each individual is replicated \(W\) times (a non-integer number of copies is allowed). Its size is \(\sum_{i=1}^n W_i\). Its distribution is defined by its expectations: for any random variable \(h\) that is a function of the observed variables and the counterfactual outcomes, with \(\operatorname{E}\mathopen{}\left[W \mathopen{}\left|h\right|\mathclose{}\right]\mathclose{} < \infty\), \[\operatorname{E}_{ps}\mathopen{}\left[h\right]\mathclose{} \stackrel{\text{def}}{=}\frac{\operatorname{E}\mathopen{}\left[W h\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[W\right]\mathclose{}}.\] Probabilities are expectations of indicators, \(\Pr_{ps}[E] \stackrel{\text{def}}{=}\operatorname{E}_{ps}\mathopen{}\left[\mathbb{1}\mathopen{}\left(E\right)\mathclose{}\right]\mathclose{}\) for any event \(E\), and conditional means are \(\operatorname{E}_{ps}\mathopen{}\left[h \mid A = a\right]\mathclose{} \stackrel{\text{def}}{=}\operatorname{E}_{ps}\mathopen{}\left[h \mathbb{1}\mathopen{}\left(A = a\right)\mathclose{}\right]\mathclose{} / \Pr_{ps}[A = a]\) whenever \(\Pr_{ps}[A = a] > 0\).
If the \(n\) individuals are drawn independently from the population, so that each \(W_i\) has the distribution of \(W\), the expected size of the pseudo-population is \(\operatorname{E}\mathopen{}\left[\sum_{i=1}^n W_i\right]\mathclose{} = \operatorname{E}\mathopen{}\left[W\right]\mathclose{} \cdot n\).
Different weights create different pseudo-populations; Proposition 1 is about the IP weights \(W = W^A\) of Definition 1.
Proposition 1 (Properties of the IP Weighted Pseudo-Population) Suppose that
Then, with \(k\) the number of treatment levels:
Here and in the proof, every sum over \(l\) runs over the values \(l\) with \(\Pr[L = l] > 0\). These three properties need no exchangeability. If, in addition, conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) holds in the actual population for every \(a\), and consistency holds (\(Y = Y^a\) for every individual with \(A = a\)), then:
By Proposition 1, \(\operatorname{E}\mathopen{}\left[W^A\right]\mathclose{} = k\), so for a dichotomous treatment the expected size of this pseudo-population is \(2n\), twice the study population.
Remark 1 (Nonparametric Weights Are Infeasible Here). In Section 2.4, \(\Pr[A = 1 \mid L]\) was estimated nonparametrically, by the proportion treated within each stratum of \(L\). That is impossible here:
This is the curse of dimensionality introduced in Chapter 10, so we resort to a model.
Algorithm 1 (Estimating IP Weights with a Treatment Model)
Fit a logistic regression model for \(\Pr[A = 1 \mid L]\) including all 9 confounders:
For each individual, compute
Set \(\hat{W}^A = 1 / \hat{f}(A \mid L)\).
Fit the (saturated) linear mean model \[\operatorname{E}\mathopen{}\left[Y \mid A\right]\mathclose{} = \theta_0 + \theta_1 A\] by weighted least squares with weights \(\hat{W}^A\).
Definition 3 (Weighted Least Squares) Given weights \(\hat{W}_i\), the weighted least squares estimates \(\hat\theta_0, \hat\theta_1\) of the model \(\operatorname{E}\mathopen{}\left[Y \mid A\right]\mathclose{} = \theta_0 + \theta_1 A\) are the values that minimize \[\sum_i \hat{W}_i \mathopen{}\left[Y_i - (\theta_0 + \theta_1 A_i)\right]\mathclose{}^2.\]
Example 4 (NHEFS: Nonstabilized IP Weighting)
Compare with the unadjusted difference of 2.5 kg.
Remark 2 (Confidence Intervals Must Account for the Weighting). Three options (Hernán and Robins 2020, chap. 12, p. 166):
The robust-variance intervals are valid but conservative: their actual coverage of the super-population parameter exceeds the nominal 95%. All confidence intervals for IP weighted estimates in this chapter are of this conservative kind.
Technical Point 12.1: Horvitz-Thompson and Hajek Estimators
Two sample estimators of the IP weighted mean \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) differ only in how they normalize the weights.
Definition 4 (Horvitz-Thompson and Hajek Estimators) With \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[\cdot\right]\mathclose{}\) the sample average:
For binary \(A\), the IP weighted least squares estimate \(\hat\theta_0 + \hat\theta_1 a\) is the Hajek estimator.
The goal of IP weighting is a pseudo-population (Definition 2) in which \(L\) does not predict \(A\). The weights \(1/f(A \mid L)\) are only one way to get there.
Remark 3 (Other Weights Give the Same Pseudo-Population Contrast).
Proposition 2 (Rescaling the Weights Does Not Change the Estimate) For a constant \(p > 0\), the Hajek (weighted least squares) mean for treatment level \(a\) computed with weights \(p/f(A \mid L)\) equals the one computed with weights \(1/f(A \mid L)\), so the difference between treatment levels is unchanged too.
Proof. \[ \begin{aligned} \frac{\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[p \, \mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} Y / f(A \mid L)\right]\mathclose{}}{\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[p \, \mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} / f(A \mid L)\right]\mathclose{}} &= \frac{p \, \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[\mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} Y / f(A \mid L)\right]\mathclose{}}{p \, \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[\mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} / f(A \mid L)\right]\mathclose{}} && \text{(constants factor out of averages)}\\ &= \frac{\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[\mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} Y / f(A \mid L)\right]\mathclose{}}{\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[\mathbb{1}\mathopen{}\left(A = a\right)\mathclose{} / f(A \mid L)\right]\mathclose{}} && \text{(cancel } p\text{)} \end{aligned} \] The same holds for every treatment level \(a\), so the difference between levels is unchanged.
The key requirement is only that treatment probability in the pseudo-population not depend on \(L\); it may differ between people. A common choice gives the treated probability \(\Pr[A = 1]\) and the untreated \(\Pr[A = 0]\), as in the original population.
Definition 5 (Stabilized IP Weights) \[SW^A \stackrel{\text{def}}{=}\frac{f(A)}{f(A \mid L)}\]
For dichotomous \(A\):
The weights \(W^A = 1/f(A \mid L)\) are called nonstabilized weights, and \(f(A)\) is the stabilizing factor.
The stabilized weights \(SW^A\) create their own pseudo-population (Definition 2), different from the one created by \(W^A\).
Algorithm 2 (Estimating Stabilized IP Weights)
Example 5 (NHEFS: Stabilized IP Weighting)
Remark 4 (When Stabilization Helps). If both kinds of weights give the same estimate, why stabilize?
Fine Point 12.2: Checking Positivity
In the NHEFS, there are 4 white women aged 66, and none of them quit smoking: positivity is empirically violated in that stratum.
Definition 6 (Marginal Structural Model) A marginal structural mean model is a model for the marginal mean of a counterfactual outcome, such as
\[\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \beta_0 + \beta_1 a\]
Because the outcome \(Y^a\) is counterfactual (generally unobserved), the model cannot be fit directly to the data of any real-world study.
Proposition 3 (MSM Treatment Parameters Are Average Causal Effects) In the marginal structural model \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \beta_0 + \beta_1 a\) of Definition 6, \[\beta_1 = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}.\]
Proof. \[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} &= (\beta_0 + \beta_1 \cdot 1) - (\beta_0 + \beta_1 \cdot 0) && \text{(evaluate the model at } a = 1, 0\text{)}\\ &= \beta_1 && \text{(cancel } \beta_0\text{)} \end{aligned} \]
So Sections 12.2-12.3 were already estimating \(\beta_1\) of a marginal structural model.
Algorithm 3 (Estimating Marginal Structural Model Parameters by IP Weighting)
New treatment: \(A\) = cigarettes per day in 1982 minus cigarettes per day at baseline (e.g., \(-25\) or \(+40\)), restricted to the 1162 individuals whose baseline smoking was at most 25 cigarettes per day.
Goal: \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a'}\right]\mathclose{}\) for any \(a, a'\). A saturated model is impractical with dozens of treatment values, so we posit a nonsaturated marginal structural model (a parabolic dose-response curve):
Definition 7 (Nonsaturated Marginal Structural Model for a Continuous Treatment) \[\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \beta_0 + \beta_1 a + \beta_2 a^2\]
where \(\beta_0 = \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) is the mean weight gain with no change in smoking intensity.
Example 6 (Effect of 20 More Cigarettes per Day) Under the model of Definition 7, \[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=20}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} &= (\beta_0 + 20\beta_1 + 20^2 \beta_2) - \beta_0 && \text{(evaluate the model at } a = 20, 0\text{)}\\ &= 20\beta_1 + 400\beta_2 && \text{(cancel } \beta_0\text{; } 20^2 = 400\text{)} \end{aligned} \]
To estimate \(\beta_1, \beta_2\): estimate \(SW^A\), then fit \(\operatorname{E}\mathopen{}\left[Y \mid A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\) to the pseudo-population.
For continuous \(A\), \(f(A \mid L)\) is a probability density function, which is generally hard to estimate when \(L\) is high-dimensional.
Example 7 (NHEFS: Weights for Change in Smoking Intensity)
Example 8 (NHEFS: Change in Smoking Intensity) IP weighted estimates: \(\hat \beta_0 = 2.005\), \(\hat \beta_1 = -0.109\), \(\hat \beta_2 = 0.003\).
| Counterfactual mean | Estimate (95% CI) |
|---|---|
| \(\operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) (no change in intensity) | 2.0 kg (1.4, 2.6) |
| \(\operatorname{E}\mathopen{}\left[Y^{a=20}\right]\mathclose{}\) (+20 cigarettes/day) | 0.9 kg (\(-1.7\), 3.5) |
| \(\operatorname{E}\mathopen{}\left[Y^{a=20}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) | \(0.9 - 2.0 = -1.1\) kg |
Definition 8 (Marginal Structural Logistic Model) For a dichotomous outcome \(D\) (1: yes, 0: no), a marginal structural logistic model is a model for the counterfactual risk on the logit scale, such as \[\operatorname{logit}\Pr[D^a = 1] = \alpha_0 + \alpha_1 a.\]
For a dichotomous treatment, \(\operatorname{exp}\mathopen{}\left\{\alpha_1\right\}\mathclose{}\) in Definition 8 is the causal odds ratio comparing \(a = 1\) with \(a = 0\). \(\operatorname{exp}\mathopen{}\left\{\alpha_1\right\}\mathclose{}\) is estimated by \(\operatorname{exp}\mathopen{}\left\{\hat\theta_1\right\}\mathclose{}\) from fitting \(\operatorname{logit}\Pr[D = 1 \mid A] = \theta_0 + \theta_1 A\) to the IP weighted pseudo-population (Definition 2).
Example 9 (NHEFS: Quitting Smoking and Death by 1992) With \(A\) = quitting smoking and \(D\) = death by 1992, the IP weighted estimate of the causal odds ratio was \(\operatorname{exp}\mathopen{}\left\{\hat\theta_1\right\}\mathclose{} = 1.0\) (95% CI: 0.8, 1.4).
Marginal structural models for the population average causal effect include no covariates. To assess effect modification by a covariate \(V\) (which need not be a confounder), add it to the model.
Definition 9 (Marginal Structural Model with an Effect Modifier) \[\operatorname{E}\mathopen{}\left[Y^a \mid V\right]\mathclose{} = \beta_0 + \beta_1 a + \beta_2 V a + \beta_3 V\]
Additive effect modification by \(V\) is present if \(\beta_2 \neq 0\). The causal effect within level \(V = v\) is \[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a=1} \mid V = v\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid V = v\right]\mathclose{} &= (\beta_0 + \beta_1 \cdot 1 + \beta_2 v \cdot 1 + \beta_3 v) - (\beta_0 + \beta_1 \cdot 0 + \beta_2 v \cdot 0 + \beta_3 v) && \text{(model evaluated at } a = 1 \text{ and } a = 0 \text{, with } V = v\text{)}\\ &= (\beta_0 + \beta_1 + \beta_2 v + \beta_3 v) - (\beta_0 + \beta_3 v) && \text{(multiply out the factors 1 and 0)}\\ &= \beta_1 + \beta_2 v && \text{(cancel } \beta_0 \text{ and } \beta_3 v\text{)}. \end{aligned} \]
Algorithm 4 (Fitting a Marginal Structural Model with an Effect Modifier)
Definition 10 (Stabilized IP Weights Given \(V\)) \[SW^A(V) \stackrel{\text{def}}{=}\frac{f[A \mid V]}{f[A \mid L]}\]
The stabilized weights \(SW^A(V)\) are estimated like \(SW^A\), but with \(V\) added to the numerator’s logistic model.
Either \(SW^A\) or \(SW^A(V)\) (Definition 10) can be used to fit the model.
Example 10 (NHEFS: Does the Effect of Quitting Vary by Sex?) Let \(V\) = sex (0: male, 1: female). The 95% confidence interval for \(\hat\theta_2\) (the product-term coefficient) was \((-2.2, 1.9)\): no strong evidence of additive effect modification by sex.
Remark 5 (Choosing the Covariates of a Marginal Structural Model).
Definition 11 (Faux Marginal Structural Model) A faux marginal structural model is a marginal structural model that includes all of \(L\) as covariates. The book coins the name half-jokingly.
For a faux marginal structural model (Definition 11), \(SW^A(L) = f[A \mid L] / f[A \mid L] = 1\), so no weighting is done: the model is fit as the unweighted outcome regression that fully adjusts for \(L\) (Chapter 15).
The analysis so far used the 1566 individuals with a 1982 weight measurement. Another 63 eligible individuals were excluded because their 1982 weight was unknown.
Definition 12 (Censoring Indicator) The censoring indicator \(C\) equals 1 if an individual’s outcome is unmeasured (censored) and 0 if it is measured.
The analyses of Sections 12.2-12.4 actually fit \[\operatorname{E}\mathopen{}\left[Y \mid A, C = 0\right]\mathclose{} = \theta_0 + \theta_1 A\] among individuals with \(C = 0\).
Example 11 (NHEFS: Censoring as a Possible Source of Selection Bias) Restricting to uncensored individuals is expected to induce bias, even under the null, when \(C\) is a collider on a pathway between \(A\) and \(Y\), or a descendant of one (Figures 8.3-8.6). The NHEFS data are consistent with that structure:
Definition 13 (Causal Effect in the Absence of Censoring) The average causal effect of \(A\) had nobody been censored is \[\operatorname{E}\mathopen{}\left[Y^{a=1, c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, c=0}\right]\mathclose{},\] a joint effect of \(A\) and \(C\) (Chapter 8).
The superscript \(c = 0\) makes explicit what many investigators mean by “the causal effect of \(A\)” even when they omit it.
Definition 14 (IP Weights for Censoring) \[ W^C \stackrel{\text{def}}{=} \begin{cases} \dfrac{1}{\Pr[C = 0 \mid L, A]} & \text{for uncensored individuals} \\ 0 & \text{for censored individuals} \end{cases} \]
\[W^{A,C} \stackrel{\text{def}}{=}W^A \times W^C = \frac{1}{f(A, C = 0 \mid L)}\]
using the factorization \(f(A, C = 0 \mid L) = f(A \mid L) \times \Pr[C = 0 \mid L, A]\).
Definition 15 (Identifiability Conditions for the Joint Treatment \((A, C)\)) For each treatment level \(a\) compared in the marginal structural model, the identifiability conditions for the joint treatment \((A, C)\) are:
Definition 16 (Stabilized IP Weights for Censoring) \[ SW^C \stackrel{\text{def}}{=} \begin{cases} \dfrac{\Pr[C = 0 \mid A]}{\Pr[C = 0 \mid L, A]} & \text{for uncensored individuals} \\ 0 & \text{for censored individuals} \end{cases} \]
\[SW^{A,C} \stackrel{\text{def}}{=}SW^A \times SW^C\]
| Nonstabilized \(W^C = 1/\Pr[C = 0 \mid L, A]\) | Stabilized \(SW^C = \Pr[C = 0 \mid A] / \Pr[C = 0 \mid L, A]\) | |
|---|---|---|
| Pseudo-population size | original population before censoring (about \(1566 + 63 = 1629\)) | original population after censoring (about 1566) |
| Arrows removed | from both \(L\) and \(A\) into \(C\) | from \(L\) into \(C\) |
| Result | no selection, hence no selection bias | selection at random with respect to \(L\): selection but no selection bias |
With either, fit the weighted model \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0\right]\mathclose{} = \theta_0 + \theta_1 A\) to estimate the marginal structural model \(\operatorname{E}\mathopen{}\left[Y^{a, c=0}\right]\mathclose{} = \beta_0 + \beta_1 a\); the stabilized version uses \(SW^{A,C}\) (Definition 16).
Example 12 (NHEFS: Adjusting for Censoring)
This is almost the same as the estimate with \(SW^A\) alone (3.4 kg), so either censoring introduces no selection bias here, or the measured covariates are not enough to remove it.
Key concepts introduced:
NHEFS results: unadjusted 2.5 kg; IP weighted 3.4 kg (95% CI: 2.4, 4.5); also adjusting for censoring 3.5 kg (95% CI: 2.5, 4.5).
Cautions