Chapter 14: G-Estimation of Structural Nested Models

Chapters 12 and 13 estimated the average causal effect of smoking cessation on weight gain with IP weighting and standardization. This chapter describes the third g-method: g-estimation, applied to the same NHEFS data. Models whose parameters are estimated by g-estimation are called structural nested models.

1 14.1 The Causal Question Revisited (pp. 187-188)

Setting (as in Chapters 12-13): 1566 cigarette smokers aged 25-74 in NHEFS.

  • Treatment \(A\): quitting smoking (\(A = 1\)) vs. not (\(A = 0\))
  • Outcome \(Y\): weight gain (kg)
  • Covariates \(L\): sex, age, race, education, intensity and duration of smoking, physical activity in daily life, recreational exercise, and weight
  • Assumption: conditional exchangeability given \(L\)

Chapters 12-13 targeted the population average causal effect \[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\] the mean weight difference had everybody been treated and uncensored vs. untreated and uncensored.

Effects Within Strata of \(L\)

This chapter instead targets the effect within each stratum of \(L\): \[\operatorname{E}\mathopen{}\left[Y^{a=1,c=0} \mid L\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0} \mid L\right]\mathclose{}\]

  • With extremely large data, one could estimate the effect separately in every stratum \(L = l\) (Chapter 4), e.g., {non-quitter, female, white, age 26, college dropout, 15 cigarettes/day, 12 years of smoking, moderate exercise, very active, weight 112 kg}
  • Alternatively, add all of \(L\) and all \(A \times L\) product terms to the marginal structural model (Section 12.5): then \(SW^A(L) = 1\) and the unweighted outcome regression, if correctly specified, adjusts for all confounding by \(L\) (Chapter 15)
  • G-estimation of a structural nested model is a third route

2 14.2 Exchangeability Revisited (pp. 188-189)

Conditional exchangeability says that, within levels of \(L\), the treated and untreated would have the same outcome distribution under the same treatment level: \[Y^a \perp\!\!\!\perp A \mid L \quad \text{for } a = 0, 1\]

For \(Y^{a=0}\): knowing \(Y^{a=0}\) does not help distinguish quitters from non-quitters with the same \(L\). For a dichotomous \(A\) this is equivalent to \[\Pr[A = 1 \mid Y^{a=0}, L] = \Pr[A = 1 \mid L]\]

A Logistic Model Including the Counterfactual

\[\operatorname{logit}\Pr[A = 1 \mid Y^{a=0}, L] = \alpha_0 + \alpha_1 Y^{a=0} + \alpha_2 L\]

  • \(\alpha_2 L = \sum_{j=1}^{p} \alpha_{2j} L_j\) for \(L = (L_1, \ldots, L_p)\)
  • Same as the model for the IP weight denominator in Chapter 12, plus \(Y^{a=0}\) as a covariate
  • We cannot fit it, since \(Y^{a=0}\) is unknown for the treated

Question: If we could fit it, and exchangeability held and the model were correct, what would \(\hat\alpha_1\) be?

Answer: its expected value is zero, because \(Y^{a=0}\) does not predict \(A\) given \(L\). This is half of g-estimation; the other half is the structural model.

3 14.3 Structural Nested Mean Models (pp. 189-191)

Target: \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid L\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid L\right]\mathclose{}\) (the difference of means equals the mean of differences). Ignore censoring for now.

No effect modification by \(L\): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid L\right]\mathclose{} = \beta_1 a\]

  • No intercept \(\beta_0\) and no main effect of \(L\): these cancel in the difference

Effect modification by \(L\) (e.g., larger effect in heavy smokers): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid L\right]\mathclose{} = \beta_1 a + \beta_2 a L\]

Conditioning on Treatment

Under \(Y^a \perp\!\!\!\perp A \mid L\), the treated and untreated are the same type of people within levels of \(L\), so the conditional effect is the same in both groups and we can write:

Definition 1 (Structural Nested Mean Model) \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a L\]

The parameters \(\beta_1\) and \(\beta_2\) (a vector), estimated by g-estimation, quantify the average causal effect of \(A\) on \(Y\) within levels of \(A\) and \(L\).

Why “Semiparametric”?

  • The parametric g-formula (Chapter 13) also models the outcome conditional on \(A\) and \(L\)
  • But a structural nested model is agnostic about the intercept and the main effect of \(L\): there is no \(\beta_0\) and no \(\beta_3 L\)
  • Leaving these unspecified means fewer assumptions and possibly more robustness to model misspecification than the parametric g-formula

Fine Point 14.1: Marginal vs. Structural Nested Models

Consider the marginal structural model for a continuous covariate \(V\) (a component of \(L\)): \[\operatorname{E}\mathopen{}\left[Y^a \mid V\right]\mathclose{} = \beta_0 + \beta_1 a + \beta_2 a V + \beta_3 V\]

  • \(\beta_1 + \beta_2 v = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid V = v\right]\mathclose{}\), the effect in individuals with \(V = v\)
  • \(\beta_0 + \beta_3 v = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid V = v\right]\mathclose{}\)

If only the effect is of interest, leave \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid V\right]\mathclose{}\) unspecified: \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid V\right]\mathclose{} = \beta_1 a + \beta_2 a V\]

This semiparametric marginal structural mean model is more robust when \(V\) is continuous or high-dimensional, because a misspecified model \(\beta_0 + \beta_3 V\) can bias \((\hat \beta_1, \hat \beta_2)\) through correlated estimates.

  • Conditional on a strict subset \(V\) of \(L\), it equals a structural nested model for a “blip” of treatment given \(V\); estimate it by g-estimation with \(V\) in place of \(L\), weighting each individual by \(SW^A(V) = f(A \mid V)/f(A \mid L)\)
  • Conditional on all of \(L\) (\(SW^A = 1\)), the “faux semiparametric MSM” \(\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid L\right]\mathclose{} = \beta_1 a + \beta_2 a L\) is, under conditional exchangeability, the structural nested mean model of this chapter

Censoring

With censoring, the target is \(\operatorname{E}\mathopen{}\left[Y^{a=1,c=0} - Y^{a=0,c=0} \mid A, L\right]\mathclose{}\), which requires adjusting for both confounding and selection bias.

  • G-estimation adjusts for confounding only, not selection bias
  • So first create a pseudo-population with nonstabilized IP weights for censoring, \(W^C = 1/\Pr[C = 0 \mid L, A]\) (from Chapter 12), then g-estimate in that pseudo-population
  • In the pseudo-population, the model is a structural model in the absence of censoring: \[\operatorname{E}\mathopen{}\left[Y^{a,c=0} - Y^{a=0,c=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a L\]

The superscript \(c = 0\) is omitted hereafter.

Technical Point 14.1: Multiplicative Structural Nested Mean Models

For positive outcomes, a multiplicative model is often preferred: \[\log\mathopen{}\left(\frac{\operatorname{E}\mathopen{}\left[Y^a \mid A = a, L\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A = a, L\right]\mathclose{}}\right)\mathclose{} = \beta_1 a + \beta_2 a L\] It can be g-estimated with \(H(\psi^\dagger) = Y \operatorname{exp}\mathopen{}\left\{-\psi_1^\dagger a - \psi_2^\dagger a L\right\}\mathclose{}\).

  • For binary outcomes, this model was originally usable only for rare outcomes (to avoid predicted probabilities above 1); Richardson, Robins, and Wang (2017) removed that restriction, and the approach was extended to time-varying treatments with doubly robust g-estimation (Wang et al. 2022)
  • The older alternative, a structural nested logistic model (\(\operatorname{logit}\) contrast of \(Y^a\) vs. \(Y^{a=0}\)), is not collapsible (Fine Point 4.3) and does not generalize easily to time-varying treatments (Hernán and Robins 2020, 191)

4 14.4 Rank Preservation (pp. 191-193)

Rank all individuals by their observed \(Y\): individual 23522 first (48.5 kg), 6928 second (47.5 kg), …, 23321 last (\(-41.3\) kg). Now imagine ranking everyone by \(Y^{a=1}\) and, separately, by \(Y^{a=0}\).

Definition 2 (Rank Preservation) Rank preservation holds if the ranking of individuals by \(Y^{a=1}\) is identical to their ranking by \(Y^{a=0}\).

  • Additive rank preservation: the effect of \(A\) on \(Y\) is exactly the same, on the additive scale, for every individual
  • Conditional additive rank preservation: the effect is exactly the same for all individuals with the same value of \(L\)

Example 1 (Rank-Preserving Structural Model)  

  • If quitting increases everyone’s weight by exactly 3 kg, additive rank preservation holds
  • The sharp null hypothesis (Chapter 1) is a special case
  • An (additive conditional) rank-preserving structural model: \[Y_i^a - Y_i^{a=0} = \psi_1 a + \psi_2 a L_i \quad \text{for all individuals } i\] so every individual with \(L = l\) has \(Y_i^{a=1} = Y_i^{a=0} + \psi_1 + \psi_2 l\)

Rank Preservation Is Implausible, and Not Needed

  • Individual effects of quitting surely vary even within levels of \(L\): some gain a lot, some little, some may lose weight
  • None of the methods in the book require rank preservation: the MSM of Chapter 12 (estimate 3.5 kg, 95% CI 2.5 to 4.5) and the structural nested mean model of Section 14.3 are models for average effects
  • The rank-preserving model is used only to introduce g-estimation, because g-estimation is easier to understand for it and the procedure is the same for rank-preserving and non-rank-preserving models

5 14.5 G-Estimation (pp. 193-196)

Goal: estimate \(\beta_1\) in \(\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a\) (no effect modification by \(L\)).

Temporarily assume the rank-preserving model \(Y_i^a - Y_i^{a=0} = \psi_1 a\) for all \(i\), so \(\psi_1 = \beta_1\). Rewrite it as \[Y^{a=0} = Y^a - \psi_1 a\]

Linking the Model to the Observed Data

By consistency, \(Y = Y^A\). Replacing \(a\) by each individual’s observed \(A\): \[ \begin{aligned} Y^{a=0} &= Y^A - \psi_1 A && \text{(set } a = A\text{)} \\ &= Y - \psi_1 A && \text{(consistency: } Y^A = Y\text{)} \end{aligned} \]

If we knew \(\psi_1\) we could compute \(Y^{a=0}\) for everyone. We don’t.

The Guessing Game

A friend says \(\psi_1\) is one of \(\psi^\dagger \in \{-20, 0, 10\}\). For each candidate compute \[H(\psi^\dagger) = Y - \psi^\dagger A\]

  • \(H(\psi^\dagger) = Y^{a=0}\) only when \(\psi^\dagger = \psi_1\)
  • Under conditional exchangeability, \(Y^{a=0}\) does not predict \(A\) given \(L\) (Section 14.2)
  • So fit, for each candidate, \[\operatorname{logit}\Pr[A = 1 \mid H(\psi^\dagger), L] = \alpha_0 + \alpha_1 H(\psi^\dagger) + \alpha_2 L\]
  • The candidate with \(\alpha_1 = 0\) is \(Y^{a=0}\); e.g., if \(H(10)\) is unassociated with \(A\) given \(L\), then \(\hat\psi_1 = 10\)

Important

G-estimation does not test whether conditional exchangeability holds; it assumes it (Hernán and Robins 2020, 194).

NHEFS Results

Example 2 (G-Estimate of the Effect of Smoking Cessation (Program 14.2))  

  • Computed 31 candidates \(H(2.0), H(2.1), \ldots, H(5.0)\) and fit 31 logistic models (the Chapter 12 IP-weight denominator model plus \(H(\psi^\dagger)\))
  • \(\hat\alpha_1\) was closest to zero for \(H(3.4)\) and \(H(3.5)\); a finer search found \(\hat\alpha_1 \approx 0\) at \(H(3.446)\)
  • G-estimate: \(\hat\psi_1 = 3.4\) kg
  • 95% CI: 2.5 to 4.5 kg

Confidence Intervals by Test Inversion

  • For each \(\psi^\dagger\), test \(\alpha_1 = 0\) (Wald test); at \(\psi^\dagger = 3.446\) the P-value was 0.998
  • The P-value exceeded 0.05 for \(\psi^\dagger\) between about 2.5 and 4.5
  • The 95% CI is the set of \(\psi^\dagger\) with P-value \(> 0.05\): inversion of the test

Back to Non-Rank-Preserving Models

The g-estimation algorithm for \(\psi_1\) gives a consistent estimate of \(\beta_1\) in the mean model whenever the mean model is correctly specified, regardless of whether rank preservation holds.

  • It does not need \(H(\beta_1) = Y^{a=0}\) for every individual
  • It only needs \(H(\beta_1)\) and \(Y^{a=0}\) to have the same conditional mean given \(L\)

Fine Point 14.2: Sensitivity Analysis for Unmeasured Confounding

G-estimation uses the fact that \(\alpha_1 = 0\) under conditional exchangeability given \(L\).

  • Suppose quitting is less likely when one’s spouse smokes, and spouse’s smoking is associated with determinants of weight gain not in \(L\): unmeasured confounding, so \(\alpha_1 \neq 0\)
  • G-estimation does not actually require \(\alpha_1 = 0\), only that the magnitude of nonexchangeability (the value of \(\alpha_1\)) be known: if \(\alpha_1 = 0.1\), search for the \(\psi^\dagger\) at which \(\hat\alpha_1 = 0.1\)
  • Repeating the analysis over a range of \(\alpha_1\) values and plotting the estimates gives a sensitivity analysis for unmeasured confounding
  • A practical difficulty: how much confounding is \(\alpha_1 = 0.1\)? See Robins, Rotnitzky, and Scharfstein (1999) for technical details

6 14.6 Structural Nested Models with Two or More Parameters (pp. 196-197)

Without the product term \(\beta_2 a L\), the model assumes no effect modification by \(L\).

  • If some components \(V\) of \(L\) modify the effect and the term \(\beta_2 a V\) is omitted, the structural nested model is misspecified and generally biased
  • Contrast: the saturated MSM \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \beta_0 + \beta_1 a\) is not misspecified when it omits \(\beta_2 a V\) and \(\beta_3 V\), because it targets the population effect
  • Structural nested models, by definition, target effects within levels of \(L\)

Two Parameters

Suppose the effect depends on baseline smoking intensity \(V\): \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a V\]

  • Candidate: \(\beta^\dagger = (\beta_1^\dagger, \beta_2^\dagger)\) and \(H(\beta^\dagger) = Y - \beta_1^\dagger A - \beta_2^\dagger A V\)
  • Fit the IP-weighted logistic model \[\operatorname{logit}\Pr[A = 1 \mid H(\beta^\dagger), L] = \alpha_0 + \alpha_1 H(\beta^\dagger) + \alpha_2 H(\beta^\dagger) V + \alpha_3 L\]
  • Find \((\beta_1^\dagger, \beta_2^\dagger)\) making both \(\alpha_1\) and \(\alpha_2\) zero: a two-dimensional search

No Search Needed in General

  • For linear structural nested mean models the estimator has closed form (Technical Point 14.2)
  • For nonlinear models, g-estimation solves an estimating equation, so derivative-based methods (e.g., Newton-Raphson) apply
  • Some structural nested survival models need a search, because their estimating equation is not differentiable (Chapter 17)

Example 3 (Two-Parameter Model in NHEFS (Program 14.3)) G-estimates: \(\hat \beta_1 = 2.86\) and \(\hat \beta_2 = 0.03\). Confidence intervals are most easily obtained by bootstrapping.

Effect Modification by All of \(L\)

For dichotomous \(A\), the unsaturated model \[\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + a \sum_{j=1}^{p} \beta_{2j} L_j\] has \(p + 1\) parameters. The average causal effect in the study population is then estimated by averaging the stratum-specific effect over the \(n\) individuals: \[ \begin{aligned} \frac{1}{n} \sum_{i=1}^n\mathopen{}\left(\beta_1 + \sum_{j=1}^{p} \beta_{2j} L_{ij}\right)\mathclose{} &= \frac{1}{n} \sum_{i=1}^n\beta_1 + \frac{1}{n} \sum_{i=1}^n\sum_{j=1}^{p} \beta_{2j} L_{ij} \\ &= \beta_1 + \frac{1}{n} \sum_{i=1}^n\sum_{j=1}^{p} \beta_{2j} L_{ij} \end{aligned} \]

Technical Point 14.2: G-Estimation of Structural Nested Mean Models

General model: \(\operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid A, L\right]\mathclose{} = A \gamma(L; \beta)\), with \(\gamma(L; \beta^\dagger)\) known and \(\gamma(L; 0) = 0\).

Estimating equation. Choose \(\beta^\dagger\) minimizing the association between \(H(\beta^\dagger) = Y - A\gamma(L; \beta^\dagger)\) and \(A\) given \(L\). Basing this on a score test means solving \[\sum_{i=1}^nI[C_i = 0] \, W_i^C \, H_i(\beta^\dagger) \mathopen{}\left(A_i - \operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{}\right)\mathclose{} q(L_i) = 0\] where \(q(L_i)\) is a user-chosen vector function of the same dimension as \(\beta\), and \(W_i^C\) and \(\operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{} = \Pr[A = 1 \mid L_i]\) are replaced by estimates. The choice of \(q\) affects efficiency (CI length), not consistency; Robins (1994) derived the choice that minimizes CI length.

Closed form. If \(\gamma(L; \beta^\dagger) = \beta^{\dagger T} d(L)\) for a known \(d(L)\), choose \(q(L) = d(L)\) and write \(R_i = I[C_i = 0] W_i^C (A_i - \operatorname{E}\mathopen{}\left[A \mid L_i\right]\mathclose{})\). Then \[ \begin{aligned} 0 &= \sum_i R_i \mathopen{}\left(Y_i - A_i\, d(L_i)^T \beta\right)\mathclose{} d(L_i) && \text{(substitute } H_i\text{; } \beta^T d = d^T \beta\text{)} \\ \sum_i R_i A_i\, d(L_i)\, d(L_i)^T \beta &= \sum_i R_i Y_i\, d(L_i) && \text{(move the } \beta \text{ term to the left)} \\ \hat \beta&= \mathopen{}\left(\sum_i R_i A_i\, d(L_i)\, d(L_i)^T\right)\mathclose{}^{-1} \sum_i R_i Y_i\, d(L_i) && \text{(solve for } \beta\text{)} \end{aligned} \]

Mean vs. distributional independence. Structural nested mean models only impose \(\operatorname{E}\mathopen{}\left[H(\beta_1) \mid A, L\right]\mathclose{} = \operatorname{E}\mathopen{}\left[H(\beta_1) \mid L\right]\mathclose{}\), so nonlinear functions of \(H_i(\beta^\dagger)\) (e.g., its cube) cannot be used in the estimating equation; they can be for structural nested distribution models, which impose \(H(\beta_1) \perp\!\!\!\perp A \mid L\).

Double robustness. The estimator above is consistent only if the models for \(\operatorname{E}\mathopen{}\left[A \mid L\right]\mathclose{}\) and \(\Pr[C = 1 \mid A, L]\) are both correct. Replacing \(H(\beta^\dagger)\) by \(H(\beta^\dagger) - \operatorname{E}\mathopen{}\left[H(\beta^\dagger) \mid L\right]\mathclose{}\), with the latter estimated by an unweighted linear model among the uncensored, gives an estimator that is consistent if either (i) the model for \(\operatorname{E}\mathopen{}\left[H(\beta^\dagger) \mid L\right]\mathclose{}\) or (ii) both the treatment and censoring models are correct: a doubly robust estimator (Hernán and Robins 2020, 198).

7 Summary

  1. Structural nested mean models \(\operatorname{E}\mathopen{}\left[Y^a - Y^{a=0} \mid A = a, L\right]\mathclose{} = \beta_1 a + \beta_2 a L\) model the effect within levels of \(L\) and leave \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid L\right]\mathclose{}\) unspecified (semiparametric).
  2. Exchangeability can be written as: \(Y^{a=0}\) does not predict \(A\) given \(L\) (\(\alpha_1 = 0\)).
  3. G-estimation finds the parameter value whose candidate counterfactual \(H(\psi^\dagger)\) is unassociated with \(A\) given \(L\); NHEFS: 3.4 kg (95% CI 2.5, 4.5).
  4. Rank preservation is implausible and not required: it only simplifies the explanation.
  5. Censoring is handled by IP weights; g-estimation itself adjusts only for confounding.
  6. Models with effect modification need product terms; linear models have a closed-form g-estimator, which can be made doubly robust.

8 References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.