Part I of the book was mostly conceptual, with calculations simple enough to do by hand. Part II needs a computer to fit regression models, for example linear and logistic regression. This chapter contrasts the nonparametric estimators used in Part I with the parametric (model-based) estimators used in Part II, reviews smoothing, and briefly introduces the bias-variance trade-off behind every modeling decision. The need for models arises whether the goal is causal inference or, say, prediction, so the chapter sets causal considerations aside until Chapter 12.
Example 1 (Sixteen Individuals with HIV) Consider 16 individuals infected with HIV, drawn at random from a large (possibly hypothetical) target population, called the super-population (see Chapter 10). Unlike in Part I, they are not viewed as representatives of a billion identical individuals.
The book uses a separate 16-person data set for each of Figures 11.1-11.3, one for each way of coding the treatment in the examples below.
Estimand: the population parameter \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), the mean outcome among individuals in the super-population with treatment level \(a\).
As in Chapter 10, an estimator \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{}\) is a function of the data used to estimate \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), and an estimate is the number obtained by applying the estimator to a particular data set. Informally, a consistent estimator is one whose estimates get closer to the population value as the sample size grows.
Example 2 (A Consistent and an Inconsistent Estimator) In Example 1, two candidate estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) are:
We require consistency, so we use the sample average.
Example 3 (Two Treatment Levels) In the setting of Example 1, let \(A \in \{0, 1\}\), with 8 individuals in each group.
Remark 1 (Association, Not Yet Causation). Under exchangeability of the treated and the untreated, together with positivity and consistency, the difference \(146.25 - 67.50 = 78.75\) in Example 3 could be read as an estimate of the average causal effect of \(A\) on \(Y\) in the target population. This chapter, however, is not about causal inference: the goal is only to motivate models for estimating population quantities such as \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), whether or not they have a causal interpretation.
Example 4 (Four Treatment Levels) In the setting of Example 1, with a different set of 16 individuals, let \(A \in \{1, 2, 3, 4\}\) (none, low dose, medium dose, high dose), with 4 individuals per level.
Example 5 (Dose from 0 to 100 mg/day) In the setting of Example 1, with yet another set of 16 individuals, let \(A\) be a dose in mg/day taking integer values from 0 to 100.
Question: How can we estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)?
Remark 2 (Continuous Treatments). If \(A\) were truly continuous, almost every treatment level would have nobody in the sample, so the sample average would be undefined almost everywhere: think of a continuous variable as a categorical one with uncountably many categories.
Data Cannot Always Speak for Themselves
Example 5 shows that letting the data “speak for themselves” does not always yield a meaningful estimate. Then a model has to fill the gap the data leave: an a priori restriction, made precise in Definition 4 below.
Suppose we knew that \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) changes linearly with \(A\): it starts at some value \(\theta_0\) when \(A = 0\) and changes by \(\theta_1\) units per unit of \(A\):
\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\]
Definition 1 (Linear Mean Model; Parametric Conditional Mean Model) Let \(Y\) be an outcome and \(A\) a treatment.
Remark 3 (Why Not “Dose-Response Curve”?). Some authors call the functional form the “dose-response curve”. The book avoids that term because it hints that the dose causes the response, and confounding can make that hint wrong.
Definition 2 (Residual, Least Squares Estimate, Predicted Value) Fix a parametric conditional mean model and a data set of \(n\) individuals.
Algorithm 1 (Estimating a Conditional Mean with a Linear Model) To estimate \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) under the linear mean model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):
Example 6 (Linear Model for the Dose Data (Figure 11.4)) Apply alg. 1 to the dose data of Example 5. The OLS estimates are \(\hat\theta_0 = 24.55\) and \(\hat\theta_1 = 2.14\) (Program 11.2). The predicted mean at \(A = 90\) is
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 \\ &\approx 216.9 \end{aligned} \]
This uses the unrounded estimates; plugging in the rounded ones gives \(24.55 + 90 \times 2.14 = 217.15\).
Wald 95% confidence intervals, assuming homoscedastic residuals (residual variance that does not depend on \(A\)): \((-21.2, 70.3)\) for \(\theta_0\), \((1.28, 2.99)\) for \(\theta_1\), and \((172.1, 261.6)\) for \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).
Definition 3 (Borrowing Information) An estimator of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) borrows information if it uses data from individuals whose treatment value is not \(a\).
In Example 6, every one of the 16 individuals affects the fitted line, so the estimate of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) borrows information from individuals with doses other than 90.
Definition 4 (Model as an A Priori Restriction) A model is “an a priori restriction on the joint distribution of the data” (Hernán and Robins 2020, 155). A model is correctly specified if its restrictions hold for the true distribution in the target population, and misspecified otherwise.
Example 7 (What the Linear Model Assumes) The linear mean model of Definition 1 restricts \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) to be a straight line, so, for example, \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) must lie between \(\operatorname{E}\mathopen{}\left[Y|A = 80\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\). Through these a priori restrictions, the model adds information to compensate for the lack of information in the data of Example 5.
No Free Lunch
Parametric estimators let us estimate quantities that would otherwise be out of reach (e.g., \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) when nobody received \(A = 90\)), but the inferences are correct only if the model is correctly specified (Definition 4).
Example 8 (Linear Model for a Dichotomous Treatment) Return to the dichotomous treatment of Example 3 and fit the same linear model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \theta_0 + 0 \times \theta_1 = \theta_0 \\ \operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \theta_0 + 1 \times \theta_1 = \theta_0 + \theta_1 \end{aligned} \]
OLS gives \(\hat\theta_0 = 67.5\) and \(\hat\theta_1 = 78.75\), so
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \hat\theta_0 = 67.5 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \hat\theta_0 + \hat\theta_1 = 67.5 + 78.75 = 146.25 \end{aligned} \]
These are exactly the sample averages from Example 3.
The match in Example 8 is no coincidence:
Proposition 1 (OLS with a Dichotomous Treatment Reproduces the Sample Averages) Let \(A_i \in \{0, 1\}\) and \(Y_i\), \(i = 1, \ldots, n\), be data in which each treatment level occurs at least once, and let \(\bar{Y}_a\) be the sample average of \(Y\) among individuals with \(A_i = a\). Then the OLS estimates for the linear mean model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) are unique and satisfy
\[ \hat\theta_0 = \bar{Y}_0, \qquad \hat\theta_0 + \hat\theta_1 = \bar{Y}_1. \]
Proof. Write \(\mu_0 = \theta_0\) and \(\mu_1 = \theta_0 + \theta_1\); the map \((\theta_0, \theta_1) \mapsto (\mu_0, \mu_1)\) is one-to-one. Because \(A_i\) is 0 or 1, the sum of squared residuals splits into two sums:
\[ \sum_{i=1}^n(Y_i - \theta_0 - \theta_1 A_i)^2 = \sum_{i: A_i = 0} (Y_i - \mu_0)^2 + \sum_{i: A_i = 1} (Y_i - \mu_1)^2. \]
The first sum involves only \(\mu_0\) and the second only \(\mu_1\), so each can be minimized separately. A sum \(\sum_i (Y_i - \mu)^2\) over a nonempty set of individuals is a strictly convex quadratic in \(\mu\), uniquely minimized at the average of those \(Y_i\). So the minimizing values are \(\hat{\mu}_0 = \bar{Y}_0\) and \(\hat{\mu}_1 = \bar{Y}_1\), and mapping back gives \(\hat\theta_0 = \bar{Y}_0\) and \(\hat\theta_0 + \hat\theta_1 = \bar{Y}_1\).
Rewrite the model as \(\operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} + \theta_1\). Since \(\theta_1\) can be any number, this statement is always true: the “model” imposes no restriction on the two means.
Definition 5 (Saturated Model) A conditional mean “model” is saturated when it imposes no restrictions on the distribution of the data. For the models in this chapter, this happens when the parameters can reproduce any values of the unknown conditional means, typically because there are as many parameters as unknown conditional means in the population (“the same number of unknowns on both sides of the equal sign” (Hernán and Robins 2020, 156)).
Definition 6 (Parsimonious Model) A model with few parameters that is used to estimate many population quantities is parsimonious.
Example 9 (Saturated vs. Parsimonious)
Remark 4 (A Saturated Model Is Not Really a Model). Strictly, a saturated “model” is not a model under Definition 4, because, like the sample average, it leaves the data to speak for themselves. Because saturated models look like models, they are still ordinarily called models. Part I was titled “Causal inference without models” because it only used saturated models.
A saturated model is one way to avoid restrictions altogether.
Definition 7 (Nonparametric Estimator) A nonparametric estimator of the conditional mean function computes its estimates from the data alone, imposing no a priori restriction on the shape of that function.
Example 10 (Nonparametric Estimators in Part I)
Sometimes No Nonparametric Estimator Exists
If \(A\) has 101 levels and nobody has \(A = 90\), as in Example 5, no nonparametric estimator of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) exists. In contrast, most methods in Part II use estimators that are parametric for some part of how the data are distributed.
Definition 8 (Identifiability Assumptions vs. Modeling Assumptions)
Fine Point 11.1: Fisher Consistency
The book’s notion of a nonparametric estimator (Definition 7) matches a standard statistical concept, defined next (Hernán and Robins 2020, 157).
Definition 9 (Fisher Consistent Estimator) An estimator of a population quantity is Fisher consistent if, when computed on the entire population instead of a sample, it returns the true value of that population quantity.
Remark 5 (Fisher Consistency and Nonparametric Estimation).
In \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\), the single number \(\theta_1\) forces the change in mean outcome per unit of dose to be the same over the whole range of \(A\).
But the difference in mean outcome per unit of dose could be larger at low doses and smaller at high doses. Linear models can be made more flexible:
\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\]
Definition 10 (Linear Model (Linear in the Parameters)) A conditional mean model is linear when it has the form \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \sum_{j} \theta_j f_j(X)\) for known functions \(f_j\) of the covariates \(X\), even if the \(f_j\) are nonlinear (powers, logarithms).
Example 11 (A Parabola Is Still a Linear Model) The model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\) is a linear model in the sense of Definition 10, even though, whenever \(\theta_2 \neq 0\), \((\theta_0, \theta_1, \theta_2)\) define a parabola rather than a straight line. \(\theta_1\) is the parameter for the linear term \(A\) and \(\theta_2\) the parameter for the quadratic term \(A^2\).
Two Meanings of Linear
“Linear” can mean “a straight line in the covariate” or “linear in the parameters” (Definition 10). The quadratic model of Example 11 is linear in the second sense but not the first.
Example 12 (Quadratic Model for the Dose Data (Figure 11.5)) Fit the quadratic model of Example 11 to the dose data of Example 5 by OLS. The estimates (Program 11.3) are \(\hat\theta_0 = -7.41\), \(\hat\theta_1 = 4.11\), \(\hat\theta_2 = -0.02\). The predicted mean at \(A = 90\) is
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 + 90 \times 90 \, \hat\theta_2 \\ &\approx 197.1 \end{aligned} \]
using the unrounded estimates, with Wald 95% confidence interval \((142.8, 251.5)\) under homoscedasticity.
Remark 6 (Curves Get Bumpier with More Parameters). Consider adding cubic, quartic, …, up to 15th-degree terms to the model of Example 12 for the 16 individuals of Example 5.
Definition 11 (Smoothing) Smoothing is estimating \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) by borrowing information (Definition 3) from individuals with \(A \neq a\).
Every parametric estimator based on a non-saturated model smooths to some degree.
Example 13 (Degrees of Smoothing for the Dose Data) For the dose data of Example 5:
Remark 7 (Smoothing with Many Covariates). With many covariates, the fitted “curves” are hyperdimensional surfaces, but the principle is the same: fewer parameters give a smoother prediction (response) surface. Models for binary outcomes, such as the logistic models of Technical Point 11.1 below (Example 15), smooth in the same way.
The two models give different estimates of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) (Table 1):
| Model | Estimate | 95% CI |
|---|---|---|
| \(\theta_0 + \theta_1 A\) | 216.9 | (172.1, 261.6) |
| \(\theta_0 + \theta_1 A + \theta_2 A^2\) | 197.1 | (142.8, 251.5) |
Which is closer to the population mean?
Proposition 2 (A Larger Model Is Correctly Specified Whenever a Smaller One Is) Let \(\mathcal{M}_1\) and \(\mathcal{M}_2\) be conditional mean models for \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) such that every mean function allowed by \(\mathcal{M}_1\) is also allowed by \(\mathcal{M}_2\) (that is, \(\mathcal{M}_1\) is nested in \(\mathcal{M}_2\)). If \(\mathcal{M}_1\) is correctly specified (Definition 4), then so is \(\mathcal{M}_2\).
Proof. \(\mathcal{M}_1\) is correctly specified, so the true function \(a \mapsto \operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) is one that \(\mathcal{M}_1\) allows. By assumption \(\mathcal{M}_2\) allows it too, so \(\mathcal{M}_2\) is correctly specified.
Example 14 (Linear vs. Quadratic Model) The straight-line model is the quadratic model with \(\theta_2 = 0\), so Proposition 2 applies with \(\mathcal{M}_1\) the straight-line and \(\mathcal{M}_2\) the quadratic model.
Remark 8 (More Parameters, More Protection Against Bias). When one model is nested in another, as in Proposition 2, the larger model imposes fewer restrictions, so by Proposition 2 it is correctly specified in every case where the smaller one is. The book’s heuristic is that, in this sense, less restrictive models give more protection against bias from model misspecification.
Definition 12 (Calibrated Confidence Interval) A nominal 95% confidence interval for an estimand is calibrated if, over repeated samples, it covers the estimand 95% of the time.
Remark 9 (The Price of Flexibility). Among nested models, the less smooth one usually has larger variance. In Table 1, the 95% CI from the 3-parameter model is 108.7 units wide, against 89.5 for the 2-parameter model.
A Narrow Interval Around a Biased Estimate
When the 2-parameter estimate is biased, its nominal 95% CI is not calibrated (Definition 12): it does not cover \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) 95% of the time. A narrower interval is no comfort if it is centered in the wrong place.
Definition 13 (Bias-Variance Trade-Off) The bias-variance trade-off is the choice investigators face between more protection against bias from model misspecification (e.g., more parameters) and the resulting cost in variance.
Remark 10 (How Models Are Chosen in Practice). Formal procedures exist to navigate the bias-variance trade-off (Definition 13), but in practice models are often chosen by tradition, by how interpretable the parameters are, and by what software is available.
Remark 11 (The Book’s Working Convention). The book usually assumes parametric models are correctly specified. This assumption is unrealistic, but it lets the book focus on problems specific to causal analysis; misspecification can affect any data analysis.
Check Sensitivity to Model Specification
Question your models: rerun the analysis under alternative specifications compatible with expert knowledge, and see how much the estimates move.
Fine Point 11.2: Model Dimensionality and Frequentist vs. Bayesian Intervals
Frequentist inference defines probability as frequency; Bayesian inference defines it as degree of belief. Chapter 10 covered frequentist confidence intervals; the boxes that follow define their Bayesian counterpart and compare the two (Hernán and Robins 2020, 159). Partly because it requires specifying a prior, Bayesian inference is used less often than frequentist inference.
Definition 14 (Bayesian Credible Interval) A Bayesian 95% credible interval is an interval such that, given the observed data and a prior distribution for all unknown parameters, the estimand lies in it with probability 95%.
Remark 12 (When Credible Intervals Are Confidence Intervals).
Technical Point 11.1: A Taxonomy of Commonly Used Models
The linear models of this chapter belong to a larger family of conditional mean models (Definition 1). The boxes that follow define that family, give its common link functions, and describe two ways to relax its parametric form (Hernán and Robins 2020, 161).
Definition 15 (Conditional Mean Model with a Link Function) Let \(Y\) be an outcome, \(X = (X_0, X_1, \ldots, X_p)\) covariates with \(X_0 = 1\), and \(\theta = (\theta_0, \ldots, \theta_p)\) unknown parameters. A conditional mean model with link function \(g\{\cdot\}\), a known invertible function, restricts \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\) through a linear predictor \(\sum_{i=0}^{p} \theta_i X_i\):
\[g\{\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\} = \sum_{i=0}^{p} \theta_i X_i\]
Example 15 (Identity, Log, and Logit Links) In Definition 15:
Remark 13 (Estimation with Canonical Links). The identity, log, and logit links of Example 15 are the canonical links of the normal, Poisson, and Bernoulli (logistic) distributions. With these links, \(\theta\) can be estimated by maximum likelihood under the matching distribution, used only as a working model. With independent, identically distributed observations and standard regularity conditions, the estimates are consistent for \(\theta\) as long as the model for \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\) is correct, even if \(Y \mid X\) does not follow that distribution. GEE models for repeated measures are another example of a conditional mean model (Hernán and Robins 2020, 161).
Definition 16 (Semiparametric Model) A model for a joint distribution is semiparametric if it describes some parts of the distribution with a finite number of parameters and leaves other parts unrestricted, where the unrestricted parts cannot be described by finitely many parameters.
Remark 14 (A Conditional Mean Model Is Semiparametric). A conditional mean model with a link function (Definition 15) restricts only \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\); it leaves the rest of the distribution of \(Y \mid X\) and the distribution of \(X\) unrestricted. If \(X\) or \(Y\) is continuous, those unrestricted parts cannot be described by finitely many parameters, so the model is semiparametric (Definition 16) for the joint distribution of \((X, Y)\).
Definition 17 (Kernel Regression) Let \((X_i, Y_i)\), \(i = 1, \ldots, n\), be the data and \(x\) a covariate value. Let \(w_h(z)\) be a kernel: a positive function that is maximal at \(z = 0\) and decreases to 0 as \(|z|\) grows, at a rate set by the bandwidth \(h\). Kernel regression estimates \(\operatorname{E}\mathopen{}\left[Y|X = x\right]\mathclose{}\) by
\[\sum_{i=1}^nw_h(x - X_i) Y_i \bigg/ \sum_{i=1}^nw_h(x - X_i).\]
Definition 18 (Generalized Additive Model) A generalized additive model replaces the linear predictor \(\sum_i \theta_i X_i\) of Definition 15 by a sum \(\sum_i f_i(X_i)\) of unknown smooth (e.g., differentiable) functions \(f_i\) of one covariate each. It can be fitted by backfitting: update each \(f_i\) in turn by smoothing (e.g., by kernel regression) the part of the outcome the other terms leave unexplained, and repeat until the fit stops changing.
Remark 15 (Two Senses of Nonparametric). Kernel regression (Definition 17) borrows information only from values of \(X\) near \(x\), with “near” set by the bandwidth. So it is “nonparametric” in a different sense from Definition 7, whose estimators used only individuals with \(X\) exactly equal to \(x\).