Chapters 12 and 13 estimated the average causal effect of smoking cessation on weight gain with IP weighting and standardization. This chapter describes the third g-method: g-estimation, applied to the same NHEFS data. Models whose parameters are estimated by g-estimation are called structural nested models.
Setting (as in Chapters 12-13): 1566 cigarette smokers aged 25-74 in NHEFS.
Chapters 12-13 targeted the population average causal effect \[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\] the mean weight difference had everybody been treated and uncensored vs. untreated and uncensored.
This chapter instead targets the effect within each stratum of \(L\): \[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0} \mid L\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0} \mid L\right]\mathclose{}\]
Conditional exchangeability says that, within levels of \(L\), the treated and untreated would have the same outcome distribution under the same treatment level: \[Y^a \perp\!\!\!\perp A \mid L \quad \text{for } a = 0, 1\]
For \(Y^{a=0}\): knowing \(Y^{a=0}\) does not help distinguish quitters from non-quitters with the same \(L\). For a dichotomous \(A\) this is equivalent to \[\Pr[A = 1 \mid Y^{a=0}, L] = \Pr[A = 1 \mid L]\]
\[\operatorname{logit}\Pr[A = 1 \mid Y^{a=0}, L] = \alpha_0 + \alpha_1 Y^{a=0} + \alpha_2 L\]
Question: If we could fit it, and exchangeability held and the model were correct, what would \(\hat\alpha_1\) be?
Answer: its expected value is zero, because \(Y^{a=0}\) does not predict \(A\) given \(L\). This is half of g-estimation; the other half is the structural model.
Target: \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid L\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid L\right]\mathclose{}\) (the difference of means equals the mean of differences). Ignore censoring for now.
No effect modification by \(L\): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid L\right]\mathclose{} = \beta_1 a\]
Effect modification by \(L\) (e.g., larger effect in heavy smokers): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid L\right]\mathclose{} = \beta_1 a + \beta_2 a L\]
Under \(Y^a \perp\!\!\!\perp A \mid L\), the treated and untreated are the same type of people within levels of \(L\), so the conditional effect is the same in both groups and we can write:
Definition 1 (Structural Nested Mean Model) \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a L\]
The parameters \(\beta_1\) and \(\beta_2\) (a vector), estimated by g-estimation, quantify the average causal effect of \(A\) on \(Y\) within levels of \(A\) and \(L\).
Fine Point 14.1: Marginal vs. Structural Nested Models
Consider the marginal structural model for a continuous covariate \(V\) (a component of \(L\)): \[\operatorname{E}\mathopen{}\left[Y^a \mid V\right]\mathclose{} = \beta_0 + \beta_1 a + \beta_2 a V + \beta_3 V\]
If only the effect is of interest, leave \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid V\right]\mathclose{}\) unspecified: \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid V\right]\mathclose{} = \beta_1 a + \beta_2 a V\]
This semiparametric marginal structural mean model is more robust when \(V\) is continuous or high-dimensional, because a misspecified model \(\beta_0 + \beta_3 V\) can bias \((\hat \beta_1, \hat \beta_2)\) through correlated estimates.
With censoring, the target is \(\operatorname{E}\mathopen{}\left[Y^{a=1,c=0} - Y^{a=0,c=0} \mid A, L\right]\mathclose{}\), which requires adjusting for both confounding and selection bias.
The superscript \(c = 0\) is omitted hereafter.
Technical Point 14.1: Multiplicative Structural Nested Mean Models
For positive outcomes, a multiplicative model is often preferred: \[\log\mathopen{}\left(\frac{\operatorname{E}\mathopen{}\left[Y^a \mid A = a, L\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A = a, L\right]\mathclose{}}\right)\mathclose{} = \beta_1 a + \beta_2 a L\] It can be g-estimated with \(H(\psi^\dagger) = Y \operatorname{exp}\mathopen{}\left\{-\psi_1^\dagger a - \psi_2^\dagger a L\right\}\mathclose{}\).
Rank all individuals by their observed \(Y\): individual 23522 first (48.5 kg), 6928 second (47.5 kg), …, 23321 last (\(-41.3\) kg). Now imagine ranking everyone by \(Y^{a=1}\) and, separately, by \(Y^{a=0}\).
Definition 2 (Rank Preservation) Rank preservation holds if the ranking of individuals by \(Y^{a=1}\) is identical to their ranking by \(Y^{a=0}\).
Example 1 (Rank-Preserving Structural Model)
Goal: estimate \(\beta_1\) in \(\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a\) (no effect modification by \(L\)).
Temporarily assume the rank-preserving model \(Y_i^a - Y_i^{a=0} = \psi_1 a\) for all \(i\), so \(\psi_1 = \beta_1\). Rewrite it as \[Y^{a=0} = Y^a - \psi_1 a\]
By consistency, \(Y = Y^A\). Replacing \(a\) by each individual’s observed \(A\): \[ \begin{aligned} Y^{a=0} &= Y^A - \psi_1 A && \text{(set } a = A\text{)} \\ &= Y - \psi_1 A && \text{(consistency: } Y^A = Y\text{)} \end{aligned} \]
If we knew \(\psi_1\) we could compute \(Y^{a=0}\) for everyone. We don’t.
A friend says \(\psi_1\) is one of \(\psi^\dagger \in \{-20, 0, 10\}\). For each candidate compute \[H(\psi^\dagger) = Y - \psi^\dagger A\]
Important
G-estimation does not test whether conditional exchangeability holds; it assumes it (Hernán and Robins 2020, 194).
Example 2 (G-Estimate of the Effect of Smoking Cessation (Program 14.2))
The g-estimation algorithm for \(\psi_1\) gives a consistent estimate of \(\beta_1\) in the mean model whenever the mean model is correctly specified, regardless of whether rank preservation holds.
Fine Point 14.2: Sensitivity Analysis for Unmeasured Confounding
G-estimation uses the fact that \(\alpha_1 = 0\) under conditional exchangeability given \(L\).
Without the product term \(\beta_2 a L\), the model assumes no effect modification by \(L\).
Suppose the effect depends on baseline smoking intensity \(V\): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a V\]
Example 3 (Two-Parameter Model in NHEFS (Program 14.3)) G-estimates: \(\hat \beta_1 = 2.86\) and \(\hat \beta_2 = 0.03\). Confidence intervals are most easily obtained by bootstrapping.
For dichotomous \(A\), the unsaturated model \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + a \sum_{j=1}^{p} \beta_{2j} L_j\] has \(p + 1\) parameters. The average causal effect in the study population is then estimated by averaging the stratum-specific effect over the \(n\) individuals: \[ \begin{aligned} \frac{1}{n} \sum_{i=1}^n\mathopen{}\left(\beta_1 + \sum_{j=1}^{p} \beta_{2j} L_{ij}\right)\mathclose{} &= \frac{1}{n} \sum_{i=1}^n\beta_1 + \frac{1}{n} \sum_{i=1}^n\sum_{j=1}^{p} \beta_{2j} L_{ij} \\ &= \beta_1 + \frac{1}{n} \sum_{i=1}^n\sum_{j=1}^{p} \beta_{2j} L_{ij} \end{aligned} \]
Technical Point 14.2: G-Estimation of Structural Nested Mean Models
General model: \(\operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid A, L\right]\mathclose{} = A \gamma(L; \beta)\), with \(\gamma(L; \beta^\dagger)\) known and \(\gamma(L; 0) = 0\).
Estimating equation. Choose \(\beta^\dagger\) minimizing the association between \(H(\beta^\dagger) = Y - A\gamma(L; \beta^\dagger)\) and \(A\) given \(L\). Basing this on a score test means solving \[\sum_{i=1}^nI[C_i = 0] \, W_i^C \, H_i(\beta^\dagger) \mathopen{}\left(A_i - \operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{}\right)\mathclose{} q(L_i) = 0\] where \(q(L_i)\) is a user-chosen vector function of the same dimension as \(\beta\), and \(W_i^C\) and \(\operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{} = \Pr[A = 1 \mid L_i]\) are replaced by estimates. The choice of \(q\) affects efficiency (CI length), not consistency; Robins (1994) derived the choice that minimizes CI length.
Closed form. If \(\gamma(L; \beta^\dagger) = \beta^{\dagger T} d(L)\) for a known \(d(L)\), choose \(q(L) = d(L)\) and write \(R_i = I[C_i = 0] W_i^C (A_i - \operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{})\). Then \[ \begin{aligned} 0 &= \sum_i R_i \mathopen{}\left(Y_i - A_i\, d(L_i)^T \beta\right)\mathclose{} d(L_i) && \text{(substitute } H_i\text{; } \beta^T d = d^T \beta\text{)} \\ \sum_i R_i A_i\, d(L_i)\, d(L_i)^T \beta &= \sum_i R_i Y_i\, d(L_i) && \text{(move the } \beta \text{ term to the left)} \\ \hat \beta&= \mathopen{}\left(\sum_i R_i A_i\, d(L_i)\, d(L_i)^T\right)\mathclose{}^{-1} \sum_i R_i Y_i\, d(L_i) && \text{(solve for } \beta\text{)} \end{aligned} \]
Mean vs. distributional independence. Structural nested mean models only impose \(\operatorname{E}\mathopen{}\left[H(\beta_1) \mid A, L\right]\mathclose{} = \operatorname{E}\mathopen{}\left[H(\beta_1) \mid L\right]\mathclose{}\), so nonlinear functions of \(H_i(\beta^\dagger)\) (e.g., its cube) cannot be used in the estimating equation; they can be for structural nested distribution models, which impose \(H(\beta_1) \perp\!\!\!\perp A \mid L\).
Double robustness. The estimator above is consistent only if the models for \(\operatorname{E}\mathopen{}\left[A \mid L\right]\mathclose{}\) and \(\Pr[C = 1 \mid A, L]\) are both correct. Replacing \(H(\beta^\dagger)\) by \(H(\beta^\dagger) - \operatorname{E}\mathopen{}\left[H(\beta^\dagger) \mid L\right]\mathclose{}\), with the latter estimated by an unweighted linear model among the uncensored, gives an estimator that is consistent if either (i) the model for \(\operatorname{E}\mathopen{}\left[H(\beta^\dagger) \mid L\right]\mathclose{}\) or (ii) both the treatment and censoring models are correct: a doubly robust estimator (Hernán and Robins 2020, 198).