Part I of the book was mostly conceptual, with calculations simple enough to do by hand. Part II requires computers to fit regression models such as linear and logistic models. This chapter contrasts the nonparametric estimators used in Part I with the parametric (model-based) estimators used in Part II, reviews smoothing, and briefly introduces the bias-variance trade-off behind every modeling decision. The need for models arises whether the goal is causal inference or, say, prediction, so the chapter sets causal considerations aside until Chapter 12.
Consider 16 individuals infected with HIV, randomly sampled from a large, possibly hypothetical super-population (the target population). Unlike in Part I, they are not viewed as representatives of a billion identical individuals.
Estimand: the population parameter \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\).
\(A \in \{0, 1\}\), with 8 individuals in each group.
\(A \in \{1, 2, 3, 4\}\) (none, low dose, medium dose, high dose), with 4 individuals per level.
With fewer individuals per category, each sample average is still exactly unbiased, but it is less likely to be close to the population mean: the 95% confidence intervals are wider than in Figure 11.1.
\(A\) is a dose in mg/day taking integer values from 0 to 100.
Question: How can we estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)?
Suppose we knew that \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) changes linearly with \(A\): it starts at some value \(\theta_0\) when \(A = 0\) and changes by \(\theta_1\) units per unit of \(A\):
\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\]
Example 1 (Linear Model for the Dose Data (Figure 11.4)) The OLS estimates are \(\hat\theta_0 = 24.55\) and \(\hat\theta_1 = 2.14\) (Program 11.2). The predicted mean at \(A = 90\) is
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 \\ &= 24.55 + 90 \times 2.14 \\ &= 216.9 \end{aligned} \]
Wald 95% confidence intervals, assuming homoscedastic residuals: \((-21.2, 70.3)\) for \(\theta_0\), \((1.28, 2.99)\) for \(\theta_1\), and \((172.1, 261.6)\) for \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).
A model is “an a priori restriction on the joint distribution of the data” (Hernán and Robins 2020, 155).
Return to the dichotomous treatment of Figure 11.1 and fit the same linear model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \theta_0 + 0 \times \theta_1 = \theta_0 \\ \operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \theta_0 + 1 \times \theta_1 = \theta_0 + \theta_1 \end{aligned} \]
OLS gives \(\hat\theta_0 = 67.5\) and \(\hat\theta_1 = 78.75\), so
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \hat\theta_0 = 67.5 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \hat\theta_0 + \hat\theta_1 = 67.5 + 78.75 = 146.25 \end{aligned} \]
These are exactly the sample averages from Section 11.1, as expected.
Rewrite the model as \(\operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} + \theta_1\). Since \(\theta_1\) can be any number, this statement is always true: the “model” imposes no restriction on the two means.
Definition 1 (Saturated Model) A conditional mean “model” is saturated when it imposes no restrictions on the distribution of the data. This occurs whenever the number of parameters equals the number of unknown conditional means in the population (“the same number of unknowns on both sides of the equal sign”).
Example 2 (Saturated vs. Parsimonious)
A model with few parameters used to estimate many population quantities is parsimonious.
Definition 2 (Nonparametric Estimator) A nonparametric estimator of the conditional mean function produces estimates from the data without any a priori restrictions on that function.
Fine Point 11.1: Fisher Consistency
The book’s nonparametric estimators coincide with what statistics calls Fisher consistent estimators (Hernán and Robins 2020, 157): estimators that, computed on the entire population instead of a sample, return the true value of the population parameter.
In \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\), the single number \(\theta_1\) forces the change in mean outcome per unit of dose to be the same over the whole range of \(A\).
But the effect of one more unit could be larger at low doses and smaller at high doses. Linear models can be made more flexible:
\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\]
Example 3 (Quadratic Model for the Dose Data (Figure 11.5)) OLS estimates (Program 11.3): \(\hat\theta_0 = -7.41\), \(\hat\theta_1 = 4.11\), \(\hat\theta_2 = -0.02\). The predicted mean at \(A = 90\) is
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 + 90 \times 90 \, \hat\theta_2 \\ &= 197.1 \end{aligned} \]
with Wald 95% confidence interval \((142.8, 251.5)\) under homoscedasticity.
Smoothing results from estimating \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) by borrowing information from individuals with \(A \neq a\). All parametric estimators smooth to some degree.
The two models give different estimates of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\):
| Model | Estimate | 95% CI |
|---|---|---|
| \(\theta_0 + \theta_1 A\) | 216.9 | (172.1, 261.6) |
| \(\theta_0 + \theta_1 A + \theta_2 A^2\) | 197.1 | (142.8, 251.5) |
Which is closer to the population mean?
Fine Point 11.2: Model Dimensionality and Frequentist vs. Bayesian Intervals
Technical Point 11.1: A Taxonomy of Commonly Used Models
Conditional mean models (generalized linear models) have a linear predictor \(\sum_{i=0}^{p} \theta_i X_i\) (with \(X_0 = 1\)) and a link function \(g\{\cdot\}\):
\[g\{\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\} = \sum_{i=0}^{p} \theta_i X_i\]
For these canonical links, \(\theta\) can be estimated by maximum likelihood under a normal, Poisson, or logistic working model, respectively; the estimates are consistent as long as the model for \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\) is correct. GEE models for repeated measures are another example (Hernán and Robins 2020, 161).
A conditional mean model leaves the rest of the distribution of \(Y \mid X\) and the distribution of \(X\) unrestricted, so with continuous \(X\) or \(Y\) it is a semiparametric model for the joint distribution of \((X, Y)\).
Relaxing the parametric form:
Kernel regression borrows information only from values of \(X\) near \(x\), so it is “nonparametric” in a different sense from this chapter’s nonparametric estimators, which used only individuals with \(X\) exactly equal to \(x\).