Chapter 11: Why Model?

Published

Last modified: 2026-10-09 12:11:16 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Part I of the book was mostly conceptual, with calculations simple enough to do by hand. Part II requires computers to fit regression models such as linear and logistic models. This chapter contrasts the nonparametric estimators used in Part I with the parametric (model-based) estimators used in Part II, reviews smoothing, and briefly introduces the bias-variance trade-off behind every modeling decision. The need for models arises whether the goal is causal inference or, say, prediction, so the chapter sets causal considerations aside until Chapter 12.

This chapter is based on Hernán and Robins (2020, chap. 11, pp. 153-161).

Important context: Most examples in Part II use real data that can be downloaded from the book’s website, together with code in R, SAS, Stata, and Python. The book assumes a working knowledge of linear and logistic regression; the “code” margin notes point to the program (e.g., Program 11.1) that reproduces each analysis.

1 11.1 Data Cannot Speak for Themselves (pp. 153-155)


1.1 Example: HIV Treatment and CD4 Count

Consider 16 individuals infected with HIV, randomly sampled from a large, possibly hypothetical super-population (the target population). Unlike in Part I, they are not viewed as representatives of a billion identical individuals.

  • Treatment \(A\): antiretroviral therapy, assigned at baseline and maintained
  • Outcome \(Y\): CD4 cell count (cells/mm³) at the end of the study

Estimand: the population parameter \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\).

As in Chapter 10, an estimator \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{}\) is a function of the data used to estimate \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), and an estimate is the number obtained by applying the estimator to a particular data set.

Informally, a consistent estimator is one for which “the larger the sample size, the closer the estimate to the population value.”

Two candidate estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\):

  1. the sample average of \(Y\) among individuals with \(A = a\) (consistent);
  2. the value of \(Y\) for the first individual in the data set with \(A = a\) (not consistent).

We require consistency, so we use the sample average.

1.2 Dichotomous Treatment (Figure 11.1)

\(A \in \{0, 1\}\), with 8 individuals in each group.

  • Estimated mean in the treated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} = 146.25\)
  • Estimated mean in the untreated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} = 67.50\)

Under exchangeability of the treated and the untreated, the difference \(146.25 - 67.50 = 78.75\) could be interpreted as an estimate of the average causal effect of \(A\) on \(Y\) in the target population. This chapter, however, is not about causal inference: the goal is only to motivate models for estimating population quantities such as \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\), whether or not they have a causal interpretation.


1.3 Polytomous Treatment (Figure 11.2)

\(A \in \{1, 2, 3, 4\}\) (none, low dose, medium dose, high dose), with 4 individuals per level.

  • Sample averages: 70.0, 80.0, 117.5, and 195.0 for \(A = 1, 2, 3, 4\)

With fewer individuals per category, each sample average is still exactly unbiased, but it is less likely to be close to the population mean: the 95% confidence intervals are wider than in Figure 11.1.


1.4 Treatment Dose (Figure 11.3)

\(A\) is a dose in mg/day taking integer values from 0 to 100.

  • There are 101 possible treatment values but only 16 individuals
  • Nobody in the study received, for example, \(A = 90\)
  • The treatment-specific sample average is undefined for every dose that nobody received

Question: How can we estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)?

If \(A\) were truly continuous, the sample average would be undefined for nearly all treatment levels: a continuous variable can be viewed as a categorical variable with uncountably many categories.

The lesson is that we cannot always let the data “speak for themselves” to obtain a meaningful estimate. We often need to supplement the data with a model.

2 11.2 Parametric Estimators of the Conditional Mean (pp. 155-156)


Suppose we knew that \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) changes linearly with \(A\): it starts at some value \(\theta_0\) when \(A = 0\) and changes by \(\theta_1\) units per unit of \(A\):

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\]

  • This restriction on the shape of \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) is a linear mean model
  • \(\theta_0\) (intercept) and \(\theta_1\) (slope) are the parameters of the model
  • A model that describes the conditional mean with a finite number of parameters is a parametric conditional mean model

The restriction on the shape of the relation is known as the functional form. Some authors call it the “dose-response curve”; the book avoids that term because it suggests that the dose causally affects the response, which could be false in the presence of confounding.

2.1 Estimation by Ordinary Least Squares

  1. Estimate \(\theta_0\) and \(\theta_1\) by ordinary least squares (OLS): among all candidate lines, choose the one minimizing the sum of the 16 squared residuals (vertical distances from each point to the line).
  2. Use the estimates to compute the predicted value \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{} = \hat\theta_0 + \hat\theta_1 a\) for any \(a\).

Example 1 (Linear Model for the Dose Data (Figure 11.4)) The OLS estimates are \(\hat\theta_0 = 24.55\) and \(\hat\theta_1 = 2.14\) (Program 11.2). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 \\ &= 24.55 + 90 \times 2.14 \\ &= 216.9 \end{aligned} \]

Wald 95% confidence intervals, assuming homoscedastic residuals: \((-21.2, 70.3)\) for \(\theta_0\), \((1.28, 2.99)\) for \(\theta_1\), and \((172.1, 261.6)\) for \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).

Multiplying out the rounded estimates gives \(24.55 + 192.6 = 217.15\); the book’s 216.9 is consistent with the unrounded estimates.

Because OLS uses all data points to find the best line, \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) is estimated by borrowing information from individuals whose treatment value is not 90.

2.2 What Is a Model?

A model is “an a priori restriction on the joint distribution of the data” (Hernán and Robins 2020, 155).

  • The linear model restricts \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) to be a straight line, so, for example, \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) must lie between \(\operatorname{E}\mathopen{}\left[Y|A = 80\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\)
  • Through its a priori restrictions, the model adds information to compensate for the lack of information in the data

No free lunch: parametric estimators let us estimate quantities that cannot be estimated otherwise (e.g., \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) when nobody received \(A = 90\)), but the inferences are correct only if the model’s restrictions are correct, i.e., if the model is correctly specified.

Much of the remainder of the book uses model-based causal inference, which relies on (approximately) no model misspecification. Since parametric models are rarely if ever perfectly specified, some misspecification is almost always expected; nonparametric estimators can partially address it.

3 11.3 Nonparametric Estimators of the Conditional Mean (pp. 156-157)


Return to the dichotomous treatment of Figure 11.1 and fit the same linear model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \theta_0 + 0 \times \theta_1 = \theta_0 \\ \operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \theta_0 + 1 \times \theta_1 = \theta_0 + \theta_1 \end{aligned} \]

OLS gives \(\hat\theta_0 = 67.5\) and \(\hat\theta_1 = 78.75\), so

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \hat\theta_0 = 67.5 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \hat\theta_0 + \hat\theta_1 = 67.5 + 78.75 = 146.25 \end{aligned} \]

These are exactly the sample averages from Section 11.1, as expected.

3.1 Saturated Models

Rewrite the model as \(\operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} + \theta_1\). Since \(\theta_1\) can be any number, this statement is always true: the “model” imposes no restriction on the two means.

Definition 1 (Saturated Model) A conditional mean “model” is saturated when it imposes no restrictions on the distribution of the data. This occurs whenever the number of parameters equals the number of unknown conditional means in the population (“the same number of unknowns on both sides of the equal sign”).

Example 2 (Saturated vs. Parsimonious)  

  • \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) with dichotomous \(A\): 2 parameters, 2 unknown means. Saturated.
  • The same model with \(A \in \{0, 1, \ldots, 100\}\): 2 parameters, 101 unknown means \(\operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{}, \ldots, \operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\). It is unbiased only if all 101 means happen to lie on a straight line.

A model with few parameters used to estimate many population quantities is parsimonious.

Strictly, a saturated “model” is not a model under the book’s definition, because it lets the data speak for themselves just as the sample average does. Because saturated models look like models, they are still ordinarily called models. Part I was titled “Causal inference without models” because it only used saturated models.

3.2 Nonparametric Estimators

Definition 2 (Nonparametric Estimator) A nonparametric estimator of the conditional mean function produces estimates from the data without any a priori restrictions on that function.

  • Example: the sample average (equivalently, the saturated model) for a dichotomous treatment
  • If \(A\) has 101 levels and nobody has \(A = 90\), no nonparametric estimator of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) exists
  • Standardization, IP weighting, stratification, and matching in Part I were nonparametric estimators under a saturated model
  • Most methods in Part II use estimators that are parametric for some part of the distribution of the data

Identifiability vs. modeling assumptions (margin note, p. 157):

  • Identifiability assumptions are needed to compute the parameter even with an infinite amount of data; formally, they make the parameter a unique function of the joint distribution of the observed data.
  • Modeling assumptions are the additional assumptions needed to estimate the parameter because we do not have an infinite amount of data.
NoteFine Point 11.1: Fisher Consistency

The book’s nonparametric estimators coincide with what statistics calls Fisher consistent estimators (Hernán and Robins 2020, 157): estimators that, computed on the entire population instead of a sample, return the true value of the population parameter.

  • A Fisher consistent estimator lacks any model restrictions, but may not exist for many population quantities.
  • When they exist, Fisher consistent estimators are the nonparametric maximum likelihood estimators under a saturated model.
  • In the wider statistical literature, “nonparametric” sometimes also refers to estimators that are not Fisher consistent but impose only weak restrictions, such as kernel regression (see Technical Point 11.1 in Section 5).

4 11.4 Smoothing (pp. 157-159)


In \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\), the single number \(\theta_1\) forces the change in mean outcome per unit of dose to be the same over the whole range of \(A\).

But the effect of one more unit could be larger at low doses and smaller at high doses. Linear models can be made more flexible:

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\]

Two meanings of “linear”: a model is linear when the mean is a linear combination of parameters and functions of the covariates, even if those functions are nonlinear (powers, logarithms). So the model with \(A^2\) is still a linear model, even though, whenever \(\theta_2 \neq 0\), \((\theta_0, \theta_1, \theta_2)\) define a parabola rather than a straight line. \(\theta_1\) is the parameter for the linear term \(A\) and \(\theta_2\) the parameter for the quadratic term \(A^2\).

Example 3 (Quadratic Model for the Dose Data (Figure 11.5)) OLS estimates (Program 11.3): \(\hat\theta_0 = -7.41\), \(\hat\theta_1 = 4.11\), \(\hat\theta_2 = -0.02\). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 + 90 \times 90 \, \hat\theta_2 \\ &= 197.1 \end{aligned} \]

with Wald 95% confidence interval \((142.8, 251.5)\) under homoscedasticity.

The estimate 197.1 uses the unrounded coefficients. Plugging in the rounded values gives \(-7.41 + 369.9 - 162 = 200.49\); the discrepancy comes from rounding \(\hat\theta_2\), which is multiplied by \(8100\).

4.1 More Parameters, Less Smoothness

  • Adding cubic, quartic, …, up to a 15th-degree term gives 16 parameters, as many as data points
  • In general, more parameters means more inflection points: the curve becomes more “wiggly”
  • The 2-parameter straight line is the smoothest model
  • A model with as many parameters as data points is the least smooth: it interpolates the data, so every data point lies on the estimated curve

4.2 Smoothing as Borrowing Information

Smoothing results from estimating \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) by borrowing information from individuals with \(A \neq a\). All parametric estimators smooth to some degree.

  • The 2-parameter model borrows information from all 16 individuals to estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)
  • A model with as many parameters as individuals borrows no information at the observed values of \(A\) (but does borrow, by interpolation, at unobserved values)
  • Intermediate smoothing: use an intermediate number of parameters, or restrict which individuals contribute (e.g., fit the straight line only to individuals with doses 80 to 100; the wider the window, the more smoothing)

With many covariates, the fitted “curves” are hyperdimensional surfaces, but the principle is the same: the fewer parameters in the model, the smoother the prediction (response) surface.

The same reasoning applies to models for dichotomous outcomes, such as logistic models (Technical Point 11.1 in Section 5).

5 11.5 The Bias-Variance Trade-Off (pp. 159-160)


The two models give different estimates of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\):

Model Estimate 95% CI
\(\theta_0 + \theta_1 A\) 216.9 (172.1, 261.6)
\(\theta_0 + \theta_1 A + \theta_2 A^2\) 197.1 (142.8, 251.5)

Which is closer to the population mean?

5.1 Bias

  • If the true relation is curvilinear, the 2-parameter model is misspecified, so its estimate is biased
  • If the true relation is a straight line, both models are correctly specified (the 3-parameter model with \(\theta_2 = 0\))
  • So the 3-parameter model is less likely to be biased: the more parameters, the fewer restrictions, and the more protection against bias from model misspecification

5.2 Variance

  • Less smooth models have larger variance: here, the 95% CI from the 3-parameter model is much wider
  • But when the 2-parameter estimate is biased, its nominal 95% CI is not calibrated: it does not cover \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) 95% of the time

The bias-variance trade-off: investigators must decide whether added protection against bias (e.g., more parameters) is worth the cost in variance. Formal procedures exist, but in practice models are often chosen by tradition, interpretability of the parameters, and software availability.

The book’s working convention: it usually assumes parametric models are correctly specified. This assumption is unrealistic, but it lets the book focus on problems specific to causal analysis; misspecification can affect any data analysis. Careful investigators question their models and run alternative analyses under model specifications compatible with expert knowledge, to assess the sensitivity of their estimates to model specification.


NoteFine Point 11.2: Model Dimensionality and Frequentist vs. Bayesian Intervals
  • Frequentist inference defines probability as frequency; Bayesian inference defines it as degree of belief.
  • A Bayesian 95% credible interval means that, given the observed data, there is a 95% probability that the estimand lies in the interval. Partly because it requires specifying a prior, Bayesian inference is less commonly used.
  • In low-dimensional parametric models with large samples, 95% credible intervals are also 95% confidence intervals: the data swamp reasonable priors.
  • In high-dimensional or nonparametric models, a 95% credible interval may trap the estimand much less than 95% of the time: if the true parameter values are unlikely under the prior, the interval is pulled toward values the prior favors.

NoteTechnical Point 11.1: A Taxonomy of Commonly Used Models

Conditional mean models (generalized linear models) have a linear predictor \(\sum_{i=0}^{p} \theta_i X_i\) (with \(X_0 = 1\)) and a link function \(g\{\cdot\}\):

\[g\{\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\} = \sum_{i=0}^{p} \theta_i X_i\]

  • Identity link: the linear models of this chapter
  • Log link (strictly positive outcomes, e.g., counts): \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \operatorname{exp}\mathopen{}\left\{\sum_{i=0}^{p} \theta_i X_i\right\}\mathclose{} > 0\)
  • Logit link (dichotomous outcomes; logistic regression): \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \operatorname{expit}\mathopen{}\left\{\sum_{i=0}^{p} \theta_i X_i\right\}\mathclose{} \in (0, 1)\)

For these canonical links, \(\theta\) can be estimated by maximum likelihood under a normal, Poisson, or logistic working model, respectively; the estimates are consistent as long as the model for \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\) is correct. GEE models for repeated measures are another example (Hernán and Robins 2020, 161).

A conditional mean model leaves the rest of the distribution of \(Y \mid X\) and the distribution of \(X\) unrestricted, so with continuous \(X\) or \(Y\) it is a semiparametric model for the joint distribution of \((X, Y)\).

Relaxing the parametric form:

  • Kernel regression estimates \(\operatorname{E}\mathopen{}\left[Y|X = x\right]\mathclose{}\) by \(\sum_{i=1}^nw_h(x - X_i) Y_i \big/ \sum_{i=1}^nw_h(x - X_i)\), where the kernel \(w_h(z)\) is positive, maximal at \(z = 0\), and decreases to 0 as \(|z|\) grows, at a rate set by the bandwidth \(h\)
  • Generalized additive models replace \(\sum_i \theta_i X_i\) by a sum of smooth functions \(\sum_i f_i(X_i)\), fitted by backfitting

Kernel regression borrows information only from values of \(X\) near \(x\), so it is “nonparametric” in a different sense from this chapter’s nonparametric estimators, which used only individuals with \(X\) exactly equal to \(x\).

6 Summary


  1. Data cannot always speak for themselves: with sparse data, nonparametric (saturated) estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) can be imprecise or undefined.
  2. Parametric models are a priori restrictions on the distribution of the data; they let us estimate otherwise inestimable quantities by borrowing information, but only correctly if the model is correctly specified.
  3. Saturated models impose no restrictions; Part I’s methods were nonparametric estimators under saturated models.
  4. Smoothing: fewer parameters, smoother fit, more borrowing of information.
  5. Bias-variance trade-off: more flexible models protect against misspecification bias at the cost of variance.

Looking ahead: The following chapters use models for causal inference:

  • Chapter 12: IP weighting and marginal structural models
  • Chapter 13: Standardization and the parametric g-formula
  • Chapter 14: G-estimation of structural nested models
  • Chapter 15: Outcome regression and propensity scores
  • Chapter 16: Instrumental variable estimation

7 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top