This chapter covers outcome regression and propensity score methods (stratification, standardization, and matching on the propensity score): the most commonly used parametric methods for causal inference. They work well for treatments fixed at a single time, but they are not designed for the complexities of time-varying treatments, which is why the book presents them only after the g-methods.
Structural nested models (Chapter 14) contain parameters for \(A\) and for \(A \times L\) product terms only, so g-estimation is agnostic about the \(L\)-\(Y\) relation. If instead we are willing to specify the \(L\)-\(Y\) relation within levels of \(A\), we can use a structural model that also has parameters for \(L\).
Definition 1 (Structural Model with Parameters for \(L\)) The structural model with parameters for \(L\) states that, for each \(a \in \{0, 1\}\),
\[ \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid L\right]\mathclose{} = \beta_0 + \beta_1 a + a {\beta_2}^{\top} L + {\beta_3}^{\top} L, \]
where \(\beta_0\) and \(\beta_1\) are scalar parameters and \(\beta_2\) and \(\beta_3\) are vector parameters.
Proposition 1 (Stratum-Specific Effect) Under the structural model of Definition 1, the average causal effect in stratum \(L = l\) is \(\beta_1 + {\beta_2}^{\top} l\).
Proof. Evaluate the model at \(a = 1\) and \(a = 0\) and subtract:
\[ \begin{aligned} &\operatorname{E}\mathopen{}\left[Y^{a=1,c=0} \mid L = l\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0} \mid L = l\right]\mathclose{} \\ &\quad = (\beta_0 + \beta_1 \cdot 1 + 1 \cdot {\beta_2}^{\top} l + {\beta_3}^{\top} l) - (\beta_0 + \beta_1 \cdot 0 + 0 \cdot {\beta_2}^{\top} l + {\beta_3}^{\top} l) \\ &\quad = (\beta_0 + \beta_1 + {\beta_2}^{\top} l + {\beta_3}^{\top} l) - (\beta_0 + {\beta_3}^{\top} l) \\ &\quad = \beta_1 + {\beta_2}^{\top} l. \end{aligned} \]
So the stratum-specific effects depend only on \((\beta_1, \beta_2)\); the counterfactual means under no treatment depend on \((\beta_0, \beta_3)\).
Proposition 2 (Outcome Regression Identifies the Structural Model) Suppose that \(L\) is discrete (it takes finitely or countably many values), and that, for each \(a \in \{0, 1\}\), \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a,c=0}\right|\mathclose{}\right]\mathclose{} < \infty\). Suppose also that \(L\) suffices to adjust for confounding and selection bias, that is, conditional exchangeability \(Y^{a,c=0} \perp\!\!\!\perp(A, C) \mid L\) holds for \(a \in \{0, 1\}\); that positivity \(\Pr[A = a, C = 0 \mid L = l] > 0\) holds for \(a \in \{0, 1\}\) and every value \(l\) with \(\Pr[L = l] > 0\); and that consistency (well-defined interventions) holds: \(Y = Y^{a,c=0}\) whenever \(A = a\) and \(C = 0\).
The conditional counterfactual mean is identified:
\[ \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid L = l\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{}, \]
for \(a \in \{0, 1\}\) and every value \(l\) with \(\Pr[L = l] > 0\).
Suppose further that the structural model of Definition 1 holds. Let \(X \stackrel{\text{def}}{=}{(1, A, A {L}^{\top}, {L}^{\top})}^{\top}\) be the regressor vector, and stack the structural parameters as \(\beta \stackrel{\text{def}}{=}{(\beta_0, \beta_1, {\beta_2}^{\top}, {\beta_3}^{\top})}^{\top}\), and suppose that \(\operatorname{E}\mathopen{}\left[Y^2 \mid C = 0\right]\mathclose{} < \infty\), that \(\operatorname{E}\mathopen{}\left[\mathopen{}\left\lVert X\right\rVert^2\mathclose{} \mid C = 0\right]\mathclose{} < \infty\), and that the second-moment matrix \(\operatorname{E}\mathopen{}\left[X {X}^{\top} \mid C = 0\right]\mathclose{}\) is invertible. Then the outcome regression model
\[ \operatorname{E}\mathopen{}\left[Y \mid A, C = 0, L\right]\mathclose{} = \alpha_0 + \alpha_1 A + A {\alpha_2}^{\top} L + {\alpha_3}^{\top} L \]
with stacked coefficients \(\alpha \stackrel{\text{def}}{=}{(\alpha_0, \alpha_1, {\alpha_2}^{\top}, {\alpha_3}^{\top})}^{\top}\) is correctly specified, and its population least-squares coefficients satisfy \(\alpha = \beta\).
Proof. Part 1:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid L = l\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid A = a, C = 0, L = l\right]\mathclose{} && \text{(conditional exchangeability; positivity makes the conditioning event non-null)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} && \text{(consistency)}. \end{aligned} \]
Part 2: by part 1 and the structural model, \(\operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, L = l\right]\mathclose{} = \beta_0 + \beta_1 a + a {\beta_2}^{\top} l + {\beta_3}^{\top} l\) for every \((a, l)\) with positive probability among the uncensored, so the linear model with \(\alpha = \beta\) equals the conditional mean and is correctly specified; in vector form, \(\operatorname{E}\mathopen{}\left[Y \mid A, L, C = 0\right]\mathclose{} = {X}^{\top} \beta\). The moment conditions make every expectation in the normal equations finite. The population least-squares coefficients \(\alpha\) are the solutions of the normal equations \(\operatorname{E}\mathopen{}\left[X (Y - {X}^{\top} \alpha) \mid C = 0\right]\mathclose{} = 0\). The vector \(\beta\) solves them:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[X (Y - {X}^{\top} \beta) \mid C = 0\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[X (Y - {X}^{\top} \beta) \mid A, L, C = 0\right]\mathclose{} \mid C = 0\right]\mathclose{} && \text{(iterated expectations)} \\ &= \operatorname{E}\mathopen{}\left[X (\operatorname{E}\mathopen{}\left[Y \mid A, L, C = 0\right]\mathclose{} - {X}^{\top} \beta) \mid C = 0\right]\mathclose{} && \text{(} X \text{ is a function of } (A, L)\text{)} \\ &= \operatorname{E}\mathopen{}\left[X ({X}^{\top} \beta - {X}^{\top} \beta) \mid C = 0\right]\mathclose{} && \text{(correct specification)} \\ &= 0. \end{aligned} \]
Because \(\operatorname{E}\mathopen{}\left[X {X}^{\top} \mid C = 0\right]\mathclose{}\) is invertible, the normal equations have a unique solution, so \(\alpha = \beta\).
So the structural parameters can be estimated by ordinary least squares, fitting the outcome regression model of Proposition 2.
Example 1 (Smoking Cessation and Weight Gain (NHEFS)) Using the outcome model of Section 13.2, which has a single product term (between \(A\) and smoking intensity), the book reports \(\hat \beta_1 = 2.6\) and \(\hat \beta_2 = 0.05\) (Hernán and Robins 2020, 200).
Definition 2 (Nuisance Parameter) A nuisance parameter is a parameter that a model must contain but that is not itself the target of inference.
Fine Point 15.1: Nuisance Parameters
Two methods for the same effect rely on different nuisance parameters (Definition 2):
Definition 3 (Propensity Score) For a dichotomous treatment \(A\), the propensity score is the conditional probability of treatment given the covariates:
\[ \pi(L) \stackrel{\text{def}}{=}\Pr[A = 1 \mid L]. \]
Example 2 (Propensity Score in a Randomized Trial) In an ideal randomized trial assigning half the individuals to \(A = 1\), \(\pi(L) = 0.5\) for everyone, whatever \(L\) is.
Remark 2 (The Propensity Score Must Usually Be Estimated). In an observational study, \(\pi(L)\) is unknown and must be estimated, e.g., with the same logistic model used for IP weighting and g-estimation in Chapters 12 and 14.
Example 3 (Estimated Propensity Scores in NHEFS) From the logistic model for quitting smoking given \(L\) (Hernán and Robins 2020, 201):
Figure 15.1 of the book shows the two distributions; they overlap over most of the range.
Definition 4 (Balancing Score and Prognostic Score (Technical Point 15.1))
Proposition 3 (The Propensity Score Balances \(L\)) Let \(A\) be dichotomous and \(\pi(L)\) its propensity score (Definition 3). Then \(\pi(L)\) is a balancing score (Definition 4):
\[ A \perp\!\!\!\perp L \mid \pi(L). \]
Proof. The book argues graphically (Technical Point 15.1); the same result follows from iterated expectations. Since \(\pi(L)\) is a function of \(L\), conditioning on \((L, \pi(L))\) is the same as conditioning on \(L\):
\[ \Pr[A = 1 \mid L, \pi(L)] = \Pr[A = 1 \mid L] = \pi(L). \]
Conditioning on \(\pi(L)\) alone, by the law of iterated expectations:
\[ \begin{aligned} \Pr[A = 1 \mid \pi(L)] &= \operatorname{E}\mathopen{}\left[\Pr[A = 1 \mid L] \mid \pi(L)\right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[\pi(L) \mid \pi(L)\right]\mathclose{} \\ &= \pi(L). \end{aligned} \]
The two probabilities are equal, so \(A\) does not depend on \(L\) once \(\pi(L)\) is fixed: \(A \perp\!\!\!\perp L \mid \pi(L)\).
In words: individuals with the same \(\pi(L)\) may differ in \(L\) (e.g., smoking intensity and exercise), but among all individuals with a given value of \(\pi(L)\) in the super-population, the distribution of \(L\) is the same in the treated and the untreated. A function \(b(L)\) is a balancing score exactly when \(\pi(L)\) can be written as a function of \(b(L)\) (Rosenbaum and Rubin 1983, Theorem 2), so the propensity score is the coarsest balancing score.
Theorem 1 (Exchangeability and Positivity Given the Propensity Score) Let \(A\) be dichotomous and \(\pi(L)\) its propensity score (Definition 3).
Proof (Proof of part 1). \[ \begin{aligned} \Pr[A = 1 \mid Y^a, \pi(L)] &= \operatorname{E}\mathopen{}\left[\Pr[A = 1 \mid Y^a, L] \mid Y^a, \pi(L)\right]\mathclose{} && \text{(iterated expectations; } \pi(L) \text{ is a function of } L\text{)} \\ &= \operatorname{E}\mathopen{}\left[\Pr[A = 1 \mid L] \mid Y^a, \pi(L)\right]\mathclose{} && (Y^a \perp\!\!\!\perp A \mid L) \\ &= \operatorname{E}\mathopen{}\left[\pi(L) \mid Y^a, \pi(L)\right]\mathclose{} && \text{(definition of } \pi(L)\text{)} \\ &= \pi(L), \end{aligned} \]
which does not depend on \(Y^a\); hence \(Y^a \perp\!\!\!\perp A \mid \pi(L)\).
Proof (Proof of part 2). The proof of Proposition 3 showed \(\Pr[A = 1 \mid \pi(L)] = \pi(L)\), so \(\Pr[A = 1 \mid \pi(L) = s] = s\) and \(\Pr[A = 0 \mid \pi(L) = s] = 1 - s\) for every \(s\) with \(\Pr[\pi(L) = s] > 0\). Because \(L\) is discrete, \(\Pr[\pi(L) = s] = \sum_{l : \pi(l) = s} \Pr[L = l]\), so the values \(s\) with \(\Pr[\pi(L) = s] > 0\) are exactly the values \(\pi(l)\) with \(\Pr[L = l] > 0\). Positivity within levels of \(\pi(L)\) therefore holds if and only if \(0 < \pi(l) < 1\) for every \(l\) with \(\Pr[L = l] > 0\). Positivity within levels of \(L\) requires \(\Pr[A = 1 \mid L = l] = \pi(l)\) and \(\Pr[A = 0 \mid L = l] = 1 - \pi(l)\) to be positive for every \(l\) with \(\Pr[L = l] > 0\), which is the same condition.
Algorithm 1 (Adjusting for the Propensity Score) Suppose exchangeability and positivity hold given \(\pi(L)\), consistency holds, and the model for \(\Pr[A = 1 \mid L]\) is correctly specified; for outcome regression on, or standardization over, \(\hat\pi(L)\), suppose also that the model for the relation between \(\pi(L)\) and \(Y\) is correctly specified. Then causal effects can be estimated as follows.
Proposition 4 (Effects Within Levels of the Propensity Score) Let \(A\) be dichotomous and \(\pi(L)\) its propensity score. Suppose that \(L\) is discrete (so \(\pi(L)\) is discrete too), and that, for each \(a \in \{0, 1\}\), \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y^{a,c=0}\right|\mathclose{}\right]\mathclose{} < \infty\). Suppose also conditional exchangeability given the propensity score, \(Y^{a,c=0} \perp\!\!\!\perp(A, C) \mid \pi(L)\) for \(a \in \{0, 1\}\); positivity, \(\Pr[A = a, C = 0 \mid \pi(L) = s] > 0\) for \(a \in \{0, 1\}\) and every \(s\) with \(\Pr[\pi(L) = s] > 0\); and consistency, \(Y = Y^{a,c=0}\) whenever \(A = a\) and \(C = 0\). Then, for each value \(s\) of the propensity score with \(\Pr[\pi(L) = s] > 0\),
\[ \operatorname{E}\mathopen{}\left[Y^{a=1,c=0} \mid \pi(L) = s\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0} \mid \pi(L) = s\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 1, C = 0, \pi(L) = s\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0, C = 0, \pi(L) = s\right]\mathclose{}. \]
Proof. For each \(a \in \{0, 1\}\),
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid \pi(L) = s\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a,c=0} \mid A = a, C = 0, \pi(L) = s\right]\mathclose{} && \text{(exchangeability given } \pi(L)\text{; positivity)} \\ &= \operatorname{E}\mathopen{}\left[Y \mid A = a, C = 0, \pi(L) = s\right]\mathclose{} && \text{(consistency)}. \end{aligned} \]
Subtracting the \(a = 0\) line from the \(a = 1\) line gives the result.
For treatment alone, without censoring, Theorem 1 shows that exchangeability and positivity given \(L\) carry over to \(\pi(L)\). When no one is censored, the joint \((A, C)\) conditions of Proposition 4 reduce to these treatment-only conditions, so they hold whenever exchangeability and positivity given \(L\) hold (with \(L\) discrete, as in Proposition 4). With censoring, \(\pi(L)\) alone need not suffice; one would instead condition on a score \(b(L)\) with \((A, C) \perp\!\!\!\perp L \mid b(L)\), for example the vector of joint probabilities \(\Pr[A = a, C = c \mid L]\) over all \((a, c)\).
Few Individuals Share a Propensity Score Value
In practice \(\pi(L)\) is generally continuous, so almost no two individuals share a value \(s\) (e.g., only one NHEFS individual had an estimated \(\pi(L)\) of 0.6563).
Example 4 (Decile Stratification in NHEFS) Classify individuals into 10 strata (deciles) of the estimated \(\pi(L)\), about 162 individuals each, and estimate the effect within each decile (Hernán and Robins 2020, 203):
Residual Imbalance Within Propensity Strata
Within a decile, the distribution of the continuous \(\pi(L)\) may still differ between treated and untreated (e.g., higher average \(\pi(L)\) in the treated), so the groups may not be exchangeable within that decile.
Use the Propensity Score as a Continuous Covariate
Remedy: use the estimated \(\pi(L)\) as a continuous covariate in the outcome model \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, \pi(L)\right]\mathclose{}\).
Example 5 (Continuous Propensity Score in NHEFS) In NHEFS (linear in \(\pi(L)\)): 3.6 kg (95% CI: 2.7, 4.5).
The outcome model on \(\pi(L)\) estimates effects within levels \(s\) of \(\pi(L)\).
Algorithm 2 (Standardization over the Propensity Score) For the population average causal effect \(\operatorname{E}\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\), standardize the fitted conditional means over the distribution of an estimated adjustment score \(b(L)\): the Chapter 13 procedure with \(b(L)\) in place of \(L\). When no one is censored, \(b(L) = \pi(L)\), which suffices for adjustment when exchangeability and positivity hold given \(L\) (Theorem 1). With censoring, \(b(L)\) is a joint score with \((A, C) \perp\!\!\!\perp L \mid b(L)\), as discussed after Proposition 4, such as the vector of joint probabilities \(\Pr[A = a, C = c \mid L]\) over all \((a, c)\). This joint score suffices when, for each \(a \in \{0, 1\}\), \(Y^{a,c=0} \perp\!\!\!\perp(A, C) \mid L\) and \(\Pr[A = a, C = 0 \mid L] > 0\): because \((A, C) \perp\!\!\!\perp L \mid b(L)\), these conditions carry over from \(L\) to \(b(L)\), by the argument Theorem 1 gives for \(A\) alone. In either case, consistency must hold, and the treatment, censoring (if any), and outcome models must be correctly specified.
Estimate the score \(\hat b(L_i)\) for each individual \(i = 1, \ldots, n\), censored or not. Without censoring, \(\hat b(L_i) = \hat\pi(L_i)\). With censoring, combine a treatment model and a censoring model, \(\hat{\Pr}[A = a, C = c \mid L] = \hat{\Pr}[A = a \mid L] \, \hat{\Pr}[C = c \mid L, A = a]\) for each \((a, c)\).
Fit an outcome model for \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, b(L)\right]\mathclose{}\) once, among the uncensored, with \(\hat b(L)\) as the covariate(s); write \(\hat m(a, s)\) for its fitted value at \(A = a\) and score value \(s\).
For each \(a \in \{0, 1\}\), predict each individual’s outcome at \(A = a\) and their own \(\hat b(L_i)\), and average the predictions over all \(n\) individuals:
\[ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a,c=0}\right]\mathclose{} = \frac{1}{n} \sum_{i=1}^n\hat m(a, \hat b(L_i)). \]
Estimate the average causal effect by the difference \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a=1,c=0}\right]\mathclose{} - \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y^{a=0,c=0}\right]\mathclose{}\).
Example 6 (Standardized Estimate in NHEFS) Standardizing over the estimated propensity score (the outcome model may include an \(A \times \pi(L)\) product term) gives 3.6 kg (95% CI: 2.7, 4.6) (Hernán and Robins 2020, 203).
Matching on \(\pi(L)\) is analogous to matching on a single continuous variable (Chapter 4).
Definition 5 (Propensity-Matched Population) A propensity-matched population is a subset of the study population, possibly reweighted, in which the distribution of \(\pi(L)\) is the same among the treated and the untreated, or approximately the same when matches need only fall within a caliper.
Fine Point 15.2: Effect Modification and the Propensity Score
Exact matches on a continuous \(\pi(L)\) are rare, so we match on a close value, e.g., untreated individuals within \(\pm 0.05\) of the treated individual’s estimated \(\pi(L)\) (e.g., a treated individual with estimated \(\pi(L) = 0.6563\) could be matched to an untreated individual with 0.6579).
Choosing How Close a Match Must Be
Bias-variance trade-off:
Example 7 (Matching and Positivity in NHEFS)
Matching Ignores the Type of Nonpositivity
Matching does not distinguish random from structural nonpositivity.
Remark 4 (Target Population of a Matched Analysis).
Propensity-Defined Populations Are Hard to Describe
Restricting to, say, estimated \(\pi(L) < 0.67\) yields a population that is hard to describe: “individuals do not come with a propensity score tattooed on their forehead” (Hernán and Robins 2020, 205), so transportability to other populations is hard to judge.
Example 8 (Restricting on Age and Smoking Duration in NHEFS) In NHEFS, the 2 treated individuals with estimated \(\pi(L) > 0.67\) were the only participants over age 50 who had smoked for fewer than 10 years. Excluding them gives an estimate for
which is a far more natural target population than “estimated \(\pi(L) < 0.67\)” (Hernán and Robins 2020, 205).
Part II used two kinds of models for causal inference.
Definition 6 (Propensity Models and Structural Models)
Remark 5 (Which Parameters Have a Causal Interpretation). We fit a propensity model only to compute the score, so its coefficients are nuisance parameters (Definition 2): they describe the association between \(L\) and \(A\), not the effect of \(A\). The treatment parameters of a structural model do have a causal interpretation.
Remark 6 (Marginal and Nested Structural Models).
Remark 7 (Outcome Regression for Prediction). Outcome regression is also widely used for prediction:
There, no variable is “treatment” and the parameters need not have any causal interpretation.
Prediction Tools Can Harm Causal Models
Forward selection, backward elimination, stepwise selection, and machine learning are built to improve prediction. Applied to propensity or structural models, they can add superfluous or harmful covariates, causing
A Propensity Model Need Not Predict Treatment Well
It is tempting to judge a propensity model by how accurately it predicts \(A\). It need not; it only needs to include the \(L\) that guarantee exchangeability.
Example 10 (Hospitals Aceso and Panacea) Everyone attends hospital Aceso, where 99% receive \(A = 1\), or hospital Panacea, where 99% receive \(A = 0\). Hospital affects the outcome only through \(A\), so it is not needed for exchangeability (Hernán and Robins 2020, 207).
Good Predictors Can Be Bad Adjustment Variables
Remark 8 (Requirements Shared by Propensity and Structural Models). Both propensity and structural models require a correctly specified functional form for the covariates (flexible terms such as cubic splines help), plus exchangeability, positivity, and well-defined interventions. Chapter 16 presents a method that does not require exchangeability of treatment.
| Method | Model used | Additional modeling requirement |
|---|---|---|
| Outcome regression (faux MSM) | \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, L\right]\mathclose{}\) | Correct outcome model (nuisance parameters \(\beta_0, \beta_3\); Definition 2) |
| G-estimation (SNM) | \(\Pr[A = 1 \mid L]\); with censoring, also \(\Pr[C = 0 \mid L, A]\) for the censoring weights | Correct treatment model and correct structural nested model; with censoring, also a correct censoring model |
| Propensity stratification / outcome regression on \(\pi(L)\) | \(\pi(L)\) and \(\operatorname{E}\mathopen{}\left[Y \mid A, C = 0, \pi(L)\right]\mathclose{}\); with censoring, the joint score \(b(L)\) in place of \(\pi(L)\) | Correct propensity model and correct score-\(Y\) relation; with censoring, also a correct censoring model |
| Propensity standardization | as above | as above; targets the population effect |
| Propensity matching | \(\pi(L)\); with censoring, also \(\Pr[C = 0 \mid L, A]\) | Correct propensity model; with censoring, also a correct censoring model; estimate refers to the matched population |
Example 11 (NHEFS Estimates Compared) NHEFS estimates of the effect of smoking cessation on weight gain (kg):