Chapter 11: Why Model?

Published

Last modified: 2026-10-09 13:21:38 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~47% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Part I of the book was mostly conceptual, with calculations simple enough to do by hand. Part II needs a computer to fit regression models, for example linear and logistic regression. This chapter contrasts the nonparametric estimators used in Part I with the parametric (model-based) estimators used in Part II, reviews smoothing, and briefly introduces the bias-variance trade-off behind every modeling decision. The need for models arises whether the goal is causal inference or, say, prediction, so the chapter sets causal considerations aside until Chapter 12.

This chapter is based on Hernán and Robins (2020, chap. 11, pp. 153-161).

Important context: Most examples in Part II use real data that can be downloaded from the book’s website, together with code in R, SAS, Stata, and Python. The book assumes a working knowledge of linear and logistic regression; the “code” margin notes point to the program (e.g., Program 11.1) that reproduces each analysis.

1 11.1 Data Cannot Speak for Themselves (pp. 153-155)


1.1 Example: HIV Treatment and CD4 Count

Example 1 (Sixteen Individuals with HIV) Consider 16 individuals infected with HIV, drawn at random from a large (possibly hypothetical) target population, called the super-population (see Chapter 10). Unlike in Part I, they are not viewed as representatives of a billion identical individuals.

  • Treatment \(A\): antiretroviral therapy, assigned at baseline and maintained
  • Outcome \(Y\): CD4 cell count (cells/mm³) at the end of the study

The book uses a separate 16-person data set for each of Figures 11.1-11.3, one for each way of coding the treatment in the examples below.

Estimand: the population parameter \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), the mean outcome among individuals in the super-population with treatment level \(a\).

As in Chapter 10, an estimator \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{}\) is a function of the data used to estimate \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), and an estimate is the number obtained by applying the estimator to a particular data set. Informally, a consistent estimator is one whose estimates get closer to the population value as the sample size grows.

Example 2 (A Consistent and an Inconsistent Estimator) In Example 1, two candidate estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) are:

  1. the sample average of \(Y\) among individuals with \(A = a\), which is consistent;
  2. the value of \(Y\) for the first individual in the data set with \(A = a\), which is not consistent, because it uses one observation however large the sample is.

We require consistency, so we use the sample average.

1.2 Dichotomous Treatment (Figure 11.1)

Example 3 (Two Treatment Levels) In the setting of Example 1, let \(A \in \{0, 1\}\), with 8 individuals in each group.

  • Estimated mean in the treated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} = 146.25\)
  • Estimated mean in the untreated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} = 67.50\)

Remark 1 (Association, Not Yet Causation). Under exchangeability of the treated and the untreated, together with positivity and consistency, the difference \(146.25 - 67.50 = 78.75\) in Example 3 could be read as an estimate of the average causal effect of \(A\) on \(Y\) in the target population. This chapter, however, is not about causal inference: the goal is only to motivate models for estimating population quantities such as \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), whether or not they have a causal interpretation.


1.3 Polytomous Treatment (Figure 11.2)

Example 4 (Four Treatment Levels) In the setting of Example 1, with a different set of 16 individuals, let \(A \in \{1, 2, 3, 4\}\) (none, low dose, medium dose, high dose), with 4 individuals per level.

  • Sample averages: 70.0, 80.0, 117.5, and 195.0 for \(A = 1, 2, 3, 4\)
WarningUnbiased Does Not Mean Close

With fewer individuals per category, each sample average in Example 4 is still an exactly unbiased estimator of its population mean, but it is less likely to be close to that mean: its 95% confidence interval will tend to be wider than in Example 3.


1.4 Treatment Dose (Figure 11.3)

Example 5 (Dose from 0 to 100 mg/day) In the setting of Example 1, with yet another set of 16 individuals, let \(A\) be a dose in mg/day taking integer values from 0 to 100.

  • There are 101 possible treatment values but only 16 individuals
  • Nobody in the study received, for example, \(A = 90\)
  • The treatment-specific sample average is undefined for every dose that nobody received

Question: How can we estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)?

Remark 2 (Continuous Treatments). If \(A\) were truly continuous, almost every treatment level would have nobody in the sample, so the sample average would be undefined almost everywhere: think of a continuous variable as a categorical one with uncountably many categories.

ImportantData Cannot Always Speak for Themselves

Example 5 shows that letting the data “speak for themselves” does not always yield a meaningful estimate. Then a model has to fill the gap the data leave: an a priori restriction, made precise in Definition 4 below.

2 11.2 Parametric Estimators of the Conditional Mean (pp. 155-156)


Suppose we knew that \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) changes linearly with \(A\): it starts at some value \(\theta_0\) when \(A = 0\) and changes by \(\theta_1\) units per unit of \(A\):

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\]

Definition 1 (Linear Mean Model; Parametric Conditional Mean Model) Let \(Y\) be an outcome and \(A\) a treatment.

  • A conditional mean model is a restriction on the shape of the conditional mean function \(a \mapsto \operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\); the restricted shape is called the functional form.
  • The linear mean model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) restricts that function to a straight line (“linear” also has a broader meaning, given in Definition 10 below), with parameters \(\theta_0\) (intercept) and \(\theta_1\) (slope).
  • A conditional mean model that describes the conditional mean function with a finite number of parameters is a parametric conditional mean model.
  • A parametric estimator is an estimator based on a parametric conditional mean model, such as the predicted value from a fitted linear mean model.

Remark 3 (Why Not “Dose-Response Curve”?). Some authors call the functional form the “dose-response curve”. The book avoids that term because it hints that the dose causes the response, and confounding can make that hint wrong.

2.1 Estimation by Ordinary Least Squares

Definition 2 (Residual, Least Squares Estimate, Predicted Value) Fix a parametric conditional mean model and a data set of \(n\) individuals.

  • For a candidate value of the parameters, the residual of individual \(i\) is \(Y_i\) minus the model’s mean at \(A_i\): the vertical distance from the data point to the curve with those parameter values.
  • The ordinary least squares (OLS) estimates are the parameter values that minimize the sum of the \(n\) squared residuals.
  • The predicted value at \(A = a\) is the model’s mean at \(a\), evaluated at the estimated parameters.

Algorithm 1 (Estimating a Conditional Mean with a Linear Model) To estimate \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) under the linear mean model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):

  1. Estimate \(\theta_0\) and \(\theta_1\) by OLS (Definition 2): among all candidate lines, choose the one minimizing \(\sum_{i=1}^n(Y_i - \theta_0 - \theta_1 A_i)^2\).
  2. Report the predicted value \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{} = \hat\theta_0 + \hat\theta_1 a\).

Example 6 (Linear Model for the Dose Data (Figure 11.4)) Apply alg. 1 to the dose data of Example 5. The OLS estimates are \(\hat\theta_0 = 24.55\) and \(\hat\theta_1 = 2.14\) (Program 11.2). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 \\ &;\approx 216.9 \end{aligned} \]

This uses the unrounded estimates; plugging in the rounded ones gives \(24.55 + 90 \times 2.14 = 217.15\).

Wald 95% confidence intervals, assuming homoscedastic residuals (residual variance that does not depend on \(A\)): \((-21.2, 70.3)\) for \(\theta_0\), \((1.28, 2.99)\) for \(\theta_1\), and \((172.1, 261.6)\) for \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).

Definition 3 (Borrowing Information) An estimator of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) borrows information if it uses data from individuals whose treatment value is not \(a\).

In Example 6, every one of the 16 individuals affects the fitted line, so the estimate of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) borrows information from individuals with doses other than 90.

2.2 What Is a Model?

Definition 4 (Model as an A Priori Restriction) A model is “an a priori restriction on the joint distribution of the data” (Hernán and Robins 2020, 155). A model is correctly specified if its restrictions hold for the true distribution in the target population, and misspecified otherwise.

Example 7 (What the Linear Model Assumes) The linear mean model of Definition 1 restricts \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) to be a straight line, so, for example, \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) must lie between \(\operatorname{E}\mathopen{}\left[Y|A = 80\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\). Through these a priori restrictions, the model adds information to compensate for the lack of information in the data of Example 5.

WarningNo Free Lunch

Parametric estimators let us estimate quantities that would otherwise be out of reach (e.g., \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) when nobody received \(A = 90\)), but the inferences are correct only if the model is correctly specified (Definition 4).

Much of the remainder of the book uses model-based causal inference, which relies on (approximately) no model misspecification. Because a parametric model is hardly ever exactly right, we should expect at least some misspecification; nonparametric estimators can partially address it.

3 11.3 Nonparametric Estimators of the Conditional Mean (pp. 156-157)


Example 8 (Linear Model for a Dichotomous Treatment) Return to the dichotomous treatment of Example 3 and fit the same linear model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \theta_0 + 0 \times \theta_1 = \theta_0 \\ \operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \theta_0 + 1 \times \theta_1 = \theta_0 + \theta_1 \end{aligned} \]

OLS gives \(\hat\theta_0 = 67.5\) and \(\hat\theta_1 = 78.75\), so

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \hat\theta_0 = 67.5 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \hat\theta_0 + \hat\theta_1 = 67.5 + 78.75 = 146.25 \end{aligned} \]

These are exactly the sample averages from Example 3.

The match in Example 8 is no coincidence:

Proposition 1 (OLS with a Dichotomous Treatment Reproduces the Sample Averages) Let \(A_i \in \{0, 1\}\) and \(Y_i\), \(i = 1, \ldots, n\), be data in which each treatment level occurs at least once, and let \(\bar{Y}_a\) be the sample average of \(Y\) among individuals with \(A_i = a\). Then the OLS estimates for the linear mean model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) are unique and satisfy

\[ \hat\theta_0 = \bar{Y}_0, \qquad \hat\theta_0 + \hat\theta_1 = \bar{Y}_1. \]

Proof. Write \(\mu_0 = \theta_0\) and \(\mu_1 = \theta_0 + \theta_1\); the map \((\theta_0, \theta_1) \mapsto (\mu_0, \mu_1)\) is one-to-one. Because \(A_i\) is 0 or 1, the sum of squared residuals splits into two sums:

\[ \sum_{i=1}^n(Y_i - \theta_0 - \theta_1 A_i)^2 = \sum_{i: A_i = 0} (Y_i - \mu_0)^2 + \sum_{i: A_i = 1} (Y_i - \mu_1)^2. \]

The first sum involves only \(\mu_0\) and the second only \(\mu_1\), so each can be minimized separately. A sum \(\sum_i (Y_i - \mu)^2\) over a nonempty set of individuals is a strictly convex quadratic in \(\mu\), uniquely minimized at the average of those \(Y_i\). So the minimizing values are \(\hat{\mu}_0 = \bar{Y}_0\) and \(\hat{\mu}_1 = \bar{Y}_1\), and mapping back gives \(\hat\theta_0 = \bar{Y}_0\) and \(\hat\theta_0 + \hat\theta_1 = \bar{Y}_1\).

3.1 Saturated Models

Rewrite the model as \(\operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} + \theta_1\). Since \(\theta_1\) can be any number, this statement is always true: the “model” imposes no restriction on the two means.

Definition 5 (Saturated Model) A conditional mean “model” is saturated when it imposes no restrictions on the distribution of the data. For the models in this chapter, this happens when the parameters can reproduce any values of the unknown conditional means, typically because there are as many parameters as unknown conditional means in the population (“the same number of unknowns on both sides of the equal sign” (Hernán and Robins 2020, 156)).

Definition 6 (Parsimonious Model) A model with few parameters that is used to estimate many population quantities is parsimonious.

Example 9 (Saturated vs. Parsimonious)  

  • \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) with dichotomous \(A\): 2 parameters, 2 unknown means. Saturated.
  • The same model with \(A \in \{0, 1, \ldots, 100\}\): 2 parameters, 101 unknown means \(\operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{}, \ldots, \operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\). Parsimonious: its estimates of the 101 means can all be unbiased only if those means lie on a straight line.

Remark 4 (A Saturated Model Is Not Really a Model). Strictly, a saturated “model” is not a model under Definition 4, because, like the sample average, it leaves the data to speak for themselves. Because saturated models look like models, they are still ordinarily called models. Part I was titled “Causal inference without models” because it only used saturated models.

3.2 Nonparametric Estimators

A saturated model is one way to avoid restrictions altogether.

Definition 7 (Nonparametric Estimator) A nonparametric estimator of the conditional mean function computes its estimates from the data alone, imposing no a priori restriction on the shape of that function.

Example 10 (Nonparametric Estimators in Part I)  

  • For a dichotomous treatment, the sample average (equivalently, by Proposition 1, the saturated model) is a nonparametric estimator of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\).
  • Standardization, IP weighting, stratification, and matching in Part I were based on nonparametric estimators under a saturated model.
WarningSometimes No Nonparametric Estimator Exists

If \(A\) has 101 levels and nobody has \(A = 90\), as in Example 5, no nonparametric estimator of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) exists. In contrast, most methods in Part II use estimators that are parametric for some part of how the data are distributed.

Definition 8 (Identifiability Assumptions vs. Modeling Assumptions)  

  • Identifiability assumptions are those we would need even with unlimited data to compute a parameter; formally, they make the parameter equal to exactly one function of the distribution of the observed data.
  • Modeling assumptions are the extra assumptions we need to estimate the parameter from a finite sample.

The book makes this distinction in a margin note on p. 157.

NoteFine Point 11.1: Fisher Consistency

The book’s notion of a nonparametric estimator (Definition 7) matches a standard statistical concept, defined next (Hernán and Robins 2020, 157).

Definition 9 (Fisher Consistent Estimator) An estimator of a population quantity is Fisher consistent if, when computed on the entire population instead of a sample, it returns the true value of that population quantity.

Remark 5 (Fisher Consistency and Nonparametric Estimation).

  • A Fisher consistent estimator imposes no model restrictions, but for many population quantities none exists.
  • When they exist, Fisher consistent estimators are the nonparametric maximum likelihood estimators under a saturated model.
  • In the wider statistical literature, “nonparametric” is sometimes also applied to estimators that impose only weak restrictions without being Fisher consistent, such as kernel regression (Definition 17, in Technical Point 11.1 below).

4 11.4 Smoothing (pp. 157-159)


In \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\), the single number \(\theta_1\) forces the change in mean outcome per unit of dose to be the same over the whole range of \(A\).

But the difference in mean outcome per unit of dose could be larger at low doses and smaller at high doses. Linear models can be made more flexible:

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\]

Definition 10 (Linear Model (Linear in the Parameters)) A conditional mean model is linear when it has the form \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \sum_{j} \theta_j f_j(X)\) for known functions \(f_j\) of the covariates \(X\), even if the \(f_j\) are nonlinear (powers, logarithms).

Example 11 (A Parabola Is Still a Linear Model) The model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\) is a linear model in the sense of Definition 10, even though, whenever \(\theta_2 \neq 0\), \((\theta_0, \theta_1, \theta_2)\) define a parabola rather than a straight line. \(\theta_1\) is the parameter for the linear term \(A\) and \(\theta_2\) the parameter for the quadratic term \(A^2\).

WarningTwo Meanings of Linear

“Linear” can mean “a straight line in the covariate” or “linear in the parameters” (Definition 10). The quadratic model of Example 11 is linear in the second sense but not the first.

Example 12 (Quadratic Model for the Dose Data (Figure 11.5)) Fit the quadratic model of Example 11 to the dose data of Example 5 by OLS. The estimates (Program 11.3) are \(\hat\theta_0 = -7.41\), \(\hat\theta_1 = 4.11\), \(\hat\theta_2 = -0.02\). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 + 90 \times 90 \, \hat\theta_2 \\ &\approx 197.1 \end{aligned} \]

using the unrounded estimates, with Wald 95% confidence interval \((142.8, 251.5)\) under homoscedasticity.

The estimate 197.1 uses the unrounded coefficients. Plugging in the rounded values gives \(-7.41 + 369.9 - 162 = 200.49\); the discrepancy comes from rounding \(\hat\theta_2\), which is multiplied by \(8100\).

4.1 More Parameters, Less Smoothness

Remark 6 (Curves Get Bumpier with More Parameters). Consider adding cubic, quartic, …, up to 15th-degree terms to the model of Example 12 for the 16 individuals of Example 5.

  • A 15th-degree polynomial has 16 parameters, as many as data points.
  • In general, more parameters means more inflection points: the curve becomes more “wiggly”.
  • Among these polynomial models, the 2-parameter straight line is the smoothest.
  • The least smooth has one parameter per data point. If the 16 doses are distinct, it interpolates the data: every data point lies on the estimated curve.

4.2 Smoothing as Borrowing Information

Definition 11 (Smoothing) Smoothing is estimating \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) by borrowing information (Definition 3) from individuals with \(A \neq a\).

Every parametric estimator based on a non-saturated model smooths to some degree.

Example 13 (Degrees of Smoothing for the Dose Data) For the dose data of Example 5:

  • The 2-parameter model borrows information from all 16 individuals to estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).
  • With one parameter per individual, the model borrows no information at the observed values of \(A\) (but does borrow, by interpolation, at unobserved values).
  • Intermediate smoothing: use an intermediate number of parameters, or restrict which individuals contribute. For example, fit the straight line only to individuals with doses in a window around 90, such as 80 to 100; the wider the window, the more smoothing.

Remark 7 (Smoothing with Many Covariates). With many covariates, the fitted “curves” are hyperdimensional surfaces, but the principle is the same: fewer parameters give a smoother prediction (response) surface. Models for binary outcomes, such as the logistic models of Technical Point 11.1 below (Example 15), smooth in the same way.

5 11.5 The Bias-Variance Trade-Off (pp. 159-160)


The two models give different estimates of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) (Table 1):

Table 1: Estimates of the mean CD4 count at dose 90 under the linear (Example 6) and quadratic (Example 12) models.
Model Estimate 95% CI
\(\theta_0 + \theta_1 A\) 216.9 (172.1, 261.6)
\(\theta_0 + \theta_1 A + \theta_2 A^2\) 197.1 (142.8, 251.5)

Which is closer to the population mean?

5.1 Bias

Proposition 2 (A Larger Model Is Correctly Specified Whenever a Smaller One Is) Let \(\mathcal{M}_1\) and \(\mathcal{M}_2\) be conditional mean models for \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) such that every mean function allowed by \(\mathcal{M}_1\) is also allowed by \(\mathcal{M}_2\) (that is, \(\mathcal{M}_1\) is nested in \(\mathcal{M}_2\)). If \(\mathcal{M}_1\) is correctly specified (Definition 4), then so is \(\mathcal{M}_2\).

Proof. \(\mathcal{M}_1\) is correctly specified, so the true function \(a \mapsto \operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) is one that \(\mathcal{M}_1\) allows. By assumption \(\mathcal{M}_2\) allows it too, so \(\mathcal{M}_2\) is correctly specified.

Example 14 (Linear vs. Quadratic Model) The straight-line model is the quadratic model with \(\theta_2 = 0\), so Proposition 2 applies with \(\mathcal{M}_1\) the straight-line and \(\mathcal{M}_2\) the quadratic model.

  • If the true relation is a straight line, both models are correctly specified.
  • If the true relation is a parabola (\(\theta_2 \neq 0\)), only the quadratic model is correctly specified, and the straight-line model’s estimate in Table 1 is in general biased.

Remark 8 (More Parameters, More Protection Against Bias). When one model is nested in another, as in Proposition 2, the larger model imposes fewer restrictions, so by Proposition 2 it is correctly specified in every case where the smaller one is. The book’s heuristic is that, in this sense, less restrictive models give more protection against bias from model misspecification.


5.2 Variance

Definition 12 (Calibrated Confidence Interval) A nominal 95% confidence interval for an estimand is calibrated if, over repeated samples, it covers the estimand 95% of the time.

Remark 9 (The Price of Flexibility). Among nested models, the less smooth one usually has larger variance. In Table 1, the 95% CI from the 3-parameter model is 108.7 units wide, against 89.5 for the 2-parameter model.

WarningA Narrow Interval Around a Biased Estimate

When the 2-parameter estimate is biased, its nominal 95% CI is not calibrated (Definition ;12): it does not cover \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) 95% of the time. A narrower interval is no comfort if it is centered in the wrong place.

Definition 13 (Bias-Variance Trade-Off) The bias-variance trade-off is the choice investigators face between more protection against bias from model misspecification (e.g., more parameters) and the resulting cost in variance.

Remark 10 (How Models Are Chosen in Practice). Formal procedures exist to navigate the bias-variance trade-off (Definition 13), but in practice models are often chosen by tradition, by how interpretable the parameters are, and by what software is available.

Remark 11 (The Book’s Working Convention). The book usually assumes parametric models are correctly specified. This assumption is unrealistic, but it lets the book focus on problems specific to causal analysis; misspecification can affect any data analysis.

TipCheck Sensitivity to Model Specification

Question your models: rerun the analysis under alternative specifications compatible with expert knowledge, and see how much the estimates move.


NoteFine Point 11.2: Model Dimensionality and Frequentist vs. Bayesian Intervals

Frequentist inference defines probability as frequency; Bayesian inference defines it as degree of belief. Chapter 10 covered frequentist confidence intervals; the boxes that follow define their Bayesian counterpart and compare the two (Hernán and Robins 2020, 159). Partly because it requires specifying a prior, Bayesian inference is used less often than frequentist inference.

Definition 14 (Bayesian Credible Interval) A Bayesian 95% credible interval is an interval such that, given the observed data and a prior distribution for all unknown parameters, the estimand lies in it with probability 95%.

Remark 12 (When Credible Intervals Are Confidence Intervals).

  • In low-dimensional parametric models with large samples, 95% credible intervals are approximately 95% confidence intervals, because the data swamp any reasonable prior.
  • In high-dimensional or nonparametric models, a 95% credible interval can cover the estimand far less often than 95% of the time. If the prior gives the true parameter values little weight, the interval is pulled toward values the prior favors.

NoteTechnical Point 11.1: A Taxonomy of Commonly Used Models

The linear models of this chapter belong to a larger family of conditional mean models (Definition 1). The boxes that follow define that family, give its common link functions, and describe two ways to relax its parametric form (Hernán and Robins 2020, 161).

Definition 16 (Semiparametric Model) A model for a joint distribution is semiparametric if it describes some parts of the distribution with a finite number of parameters and leaves other parts unrestricted, where the unrestricted parts cannot be described by finitely many parameters.

Remark 14 (A Conditional Mean Model Is Semiparametric). A conditional mean model with a link function (Definition 15) restricts only \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\); it leaves the rest of the distribution of \(Y \mid X\) and the distribution of \(X\) unrestricted. If \(X\) or \(Y\) is continuous, those unrestricted parts cannot be described by finitely many parameters, so the model is semiparametric (Definition 16) for the joint distribution of \((X, Y)\).

Definition 17 (Kernel Regression) Let \((X_i, Y_i)\), \(i = 1, \ldots, n\), be the data and \(x\) a covariate value. Let \(w_h(z)\) be a kernel: a positive function that is maximal at \(z = 0\) and decreases to 0 as \(|z|\) grows, at a rate set by the bandwidth \(h\). Kernel regression estimates \(\operatorname{E}\mathopen{}\left[Y|X = x\right]\mathclose{}\) by

\[\sum_{i=1}^nw_h(x - X_i) Y_i \bigg/ \sum_{i=1}^nw_h(x - X_i).\]

Definition 18 (Generalized Additive Model) A generalized additive model replaces the linear predictor \(\sum_i \theta_i X_i\) of Definition 15 by a sum \(\sum_i f_i(X_i)\) of unknown smooth (e.g., differentiable) functions \(f_i\) of one covariate each. It can be fitted by backfitting: update each \(f_i\) in turn by smoothing (e.g., by kernel regression) the part of the outcome the other terms leave unexplained, and repeat until the fit stops changing.

Remark 15 (Two Senses of Nonparametric). Kernel regression (Definition 17) borrows information only from values of \(X\) near \(x\), with “near” set by the bandwidth. So it is “nonparametric” in a different sense from Definition 7, whose estimators used only individuals with \(X\) exactly equal to \(x\).

6 Summary


  1. Data cannot always speak for themselves: with sparse data, nonparametric (saturated) estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) can be imprecise or undefined.
  2. Parametric models that are not saturated are a priori restrictions on the distribution of the data; they let us estimate otherwise inestimable quantities by borrowing information, but only correctly if the model is correctly specified.
  3. Saturated models impose no restrictions; Part I’s methods were nonparametric estimators under saturated models.
  4. Smoothing: fewer parameters, smoother fit, more borrowing of information.
  5. Bias-variance trade-off: more flexible models protect against misspecification bias at the cost of variance.

Looking ahead: The following chapters use models for causal inference:

  • Chapter 12: IP weighting and marginal structural models
  • Chapter 13: Standardization and the parametric g-formula
  • Chapter 14: G-estimation of structural nested models
  • Chapter 15: Outcome regression and propensity scores
  • Chapter 16: Instrumental variable estimation

7 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top