Chapter 11: Why Model?

Part I of the book was mostly conceptual, with calculations simple enough to do by hand. Part II requires computers to fit regression models such as linear and logistic models. This chapter contrasts the nonparametric estimators used in Part I with the parametric (model-based) estimators used in Part II, reviews smoothing, and briefly introduces the bias-variance trade-off behind every modeling decision. The need for models arises whether the goal is causal inference or, say, prediction, so the chapter sets causal considerations aside until Chapter 12.

1 11.1 Data Cannot Speak for Themselves (pp. 153-155)

Example: HIV Treatment and CD4 Count

Consider 16 individuals infected with HIV, randomly sampled from a large, possibly hypothetical super-population (the target population). Unlike in Part I, they are not viewed as representatives of a billion identical individuals.

  • Treatment \(A\): antiretroviral therapy, assigned at baseline and maintained
  • Outcome \(Y\): CD4 cell count (cells/mm³) at the end of the study

Estimand: the population parameter \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\).

Dichotomous Treatment (Figure 11.1)

\(A \in \{0, 1\}\), with 8 individuals in each group.

  • Estimated mean in the treated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} = 146.25\)
  • Estimated mean in the untreated: \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} = 67.50\)

Polytomous Treatment (Figure 11.2)

\(A \in \{1, 2, 3, 4\}\) (none, low dose, medium dose, high dose), with 4 individuals per level.

  • Sample averages: 70.0, 80.0, 117.5, and 195.0 for \(A = 1, 2, 3, 4\)

With fewer individuals per category, each sample average is still exactly unbiased, but it is less likely to be close to the population mean: the 95% confidence intervals are wider than in Figure 11.1.

Treatment Dose (Figure 11.3)

\(A\) is a dose in mg/day taking integer values from 0 to 100.

  • There are 101 possible treatment values but only 16 individuals
  • Nobody in the study received, for example, \(A = 90\)
  • The treatment-specific sample average is undefined for every dose that nobody received

Question: How can we estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)?

2 11.2 Parametric Estimators of the Conditional Mean (pp. 155-156)

Suppose we knew that \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) changes linearly with \(A\): it starts at some value \(\theta_0\) when \(A = 0\) and changes by \(\theta_1\) units per unit of \(A\):

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\]

  • This restriction on the shape of \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) is a linear mean model
  • \(\theta_0\) (intercept) and \(\theta_1\) (slope) are the parameters of the model
  • A model that describes the conditional mean with a finite number of parameters is a parametric conditional mean model

Estimation by Ordinary Least Squares

  1. Estimate \(\theta_0\) and \(\theta_1\) by ordinary least squares (OLS): among all candidate lines, choose the one minimizing the sum of the 16 squared residuals (vertical distances from each point to the line).
  2. Use the estimates to compute the predicted value \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = a\right]\mathclose{} = \hat\theta_0 + \hat\theta_1 a\) for any \(a\).

Example 1 (Linear Model for the Dose Data (Figure 11.4)) The OLS estimates are \(\hat\theta_0 = 24.55\) and \(\hat\theta_1 = 2.14\) (Program 11.2). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 \\ &= 24.55 + 90 \times 2.14 \\ &= 216.9 \end{aligned} \]

Wald 95% confidence intervals, assuming homoscedastic residuals: \((-21.2, 70.3)\) for \(\theta_0\), \((1.28, 2.99)\) for \(\theta_1\), and \((172.1, 261.6)\) for \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\).

What Is a Model?

A model is “an a priori restriction on the joint distribution of the data” (Hernán and Robins 2020, 155).

  • The linear model restricts \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{}\) to be a straight line, so, for example, \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) must lie between \(\operatorname{E}\mathopen{}\left[Y|A = 80\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\)
  • Through its a priori restrictions, the model adds information to compensate for the lack of information in the data

3 11.3 Nonparametric Estimators of the Conditional Mean (pp. 156-157)

Return to the dichotomous treatment of Figure 11.1 and fit the same linear model \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \theta_0 + 0 \times \theta_1 = \theta_0 \\ \operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \theta_0 + 1 \times \theta_1 = \theta_0 + \theta_1 \end{aligned} \]

OLS gives \(\hat\theta_0 = 67.5\) and \(\hat\theta_1 = 78.75\), so

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 0\right]\mathclose{} &= \hat\theta_0 = 67.5 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 1\right]\mathclose{} &= \hat\theta_0 + \hat\theta_1 = 67.5 + 78.75 = 146.25 \end{aligned} \]

These are exactly the sample averages from Section 11.1, as expected.

Saturated Models

Rewrite the model as \(\operatorname{E}\mathopen{}\left[Y|A = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{} + \theta_1\). Since \(\theta_1\) can be any number, this statement is always true: the “model” imposes no restriction on the two means.

Definition 1 (Saturated Model) A conditional mean “model” is saturated when it imposes no restrictions on the distribution of the data. This occurs whenever the number of parameters equals the number of unknown conditional means in the population (“the same number of unknowns on both sides of the equal sign”).

Example 2 (Saturated vs. Parsimonious)  

  • \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\) with dichotomous \(A\): 2 parameters, 2 unknown means. Saturated.
  • The same model with \(A \in \{0, 1, \ldots, 100\}\): 2 parameters, 101 unknown means \(\operatorname{E}\mathopen{}\left[Y|A = 0\right]\mathclose{}, \ldots, \operatorname{E}\mathopen{}\left[Y|A = 100\right]\mathclose{}\). It is unbiased only if all 101 means happen to lie on a straight line.

A model with few parameters used to estimate many population quantities is parsimonious.

Nonparametric Estimators

Definition 2 (Nonparametric Estimator) A nonparametric estimator of the conditional mean function produces estimates from the data without any a priori restrictions on that function.

  • Example: the sample average (equivalently, the saturated model) for a dichotomous treatment
  • If \(A\) has 101 levels and nobody has \(A = 90\), no nonparametric estimator of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) exists
  • Standardization, IP weighting, stratification, and matching in Part I were nonparametric estimators under a saturated model
  • Most methods in Part II use estimators that are parametric for some part of the distribution of the data

Fine Point 11.1: Fisher Consistency

The book’s nonparametric estimators coincide with what statistics calls Fisher consistent estimators (Hernán and Robins 2020, 157): estimators that, computed on the entire population instead of a sample, return the true value of the population parameter.

  • A Fisher consistent estimator lacks any model restrictions, but may not exist for many population quantities.
  • When they exist, Fisher consistent estimators are the nonparametric maximum likelihood estimators under a saturated model.
  • In the wider statistical literature, “nonparametric” sometimes also refers to estimators that are not Fisher consistent but impose only weak restrictions, such as kernel regression (see Technical Point 11.1 in Section 5).

4 11.4 Smoothing (pp. 157-159)

In \(\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A\), the single number \(\theta_1\) forces the change in mean outcome per unit of dose to be the same over the whole range of \(A\).

But the effect of one more unit could be larger at low doses and smaller at high doses. Linear models can be made more flexible:

\[\operatorname{E}\mathopen{}\left[Y|A\right]\mathclose{} = \theta_0 + \theta_1 A + \theta_2 A^2\]

Example 3 (Quadratic Model for the Dose Data (Figure 11.5)) OLS estimates (Program 11.3): \(\hat\theta_0 = -7.41\), \(\hat\theta_1 = 4.11\), \(\hat\theta_2 = -0.02\). The predicted mean at \(A = 90\) is

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y|A = 90\right]\mathclose{} &= \hat\theta_0 + 90 \hat\theta_1 + 90 \times 90 \, \hat\theta_2 \\ &= 197.1 \end{aligned} \]

with Wald 95% confidence interval \((142.8, 251.5)\) under homoscedasticity.

More Parameters, Less Smoothness

  • Adding cubic, quartic, …, up to a 15th-degree term gives 16 parameters, as many as data points
  • In general, more parameters means more inflection points: the curve becomes more “wiggly”
  • The 2-parameter straight line is the smoothest model
  • A model with as many parameters as data points is the least smooth: it interpolates the data, so every data point lies on the estimated curve

Smoothing as Borrowing Information

Smoothing results from estimating \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) by borrowing information from individuals with \(A \neq a\). All parametric estimators smooth to some degree.

  • The 2-parameter model borrows information from all 16 individuals to estimate \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\)
  • A model with as many parameters as individuals borrows no information at the observed values of \(A\) (but does borrow, by interpolation, at unobserved values)
  • Intermediate smoothing: use an intermediate number of parameters, or restrict which individuals contribute (e.g., fit the straight line only to individuals with doses 80 to 100; the wider the window, the more smoothing)

5 11.5 The Bias-Variance Trade-Off (pp. 159-160)

The two models give different estimates of \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\):

Model Estimate 95% CI
\(\theta_0 + \theta_1 A\) 216.9 (172.1, 261.6)
\(\theta_0 + \theta_1 A + \theta_2 A^2\) 197.1 (142.8, 251.5)

Which is closer to the population mean?

Bias

  • If the true relation is curvilinear, the 2-parameter model is misspecified, so its estimate is biased
  • If the true relation is a straight line, both models are correctly specified (the 3-parameter model with \(\theta_2 = 0\))
  • So the 3-parameter model is less likely to be biased: the more parameters, the fewer restrictions, and the more protection against bias from model misspecification

Variance

  • Less smooth models have larger variance: here, the 95% CI from the 3-parameter model is much wider
  • But when the 2-parameter estimate is biased, its nominal 95% CI is not calibrated: it does not cover \(\operatorname{E}\mathopen{}\left[Y|A = 90\right]\mathclose{}\) 95% of the time

Fine Point 11.2: Model Dimensionality and Frequentist vs. Bayesian Intervals

  • Frequentist inference defines probability as frequency; Bayesian inference defines it as degree of belief.
  • A Bayesian 95% credible interval means that, given the observed data, there is a 95% probability that the estimand lies in the interval. Partly because it requires specifying a prior, Bayesian inference is less commonly used.
  • In low-dimensional parametric models with large samples, 95% credible intervals are also 95% confidence intervals: the data swamp reasonable priors.
  • In high-dimensional or nonparametric models, a 95% credible interval may trap the estimand much less than 95% of the time: if the true parameter values are unlikely under the prior, the interval is pulled toward values the prior favors.

Technical Point 11.1: A Taxonomy of Commonly Used Models

Conditional mean models (generalized linear models) have a linear predictor \(\sum_{i=0}^{p} \theta_i X_i\) (with \(X_0 = 1\)) and a link function \(g\{\cdot\}\):

\[g\{\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\} = \sum_{i=0}^{p} \theta_i X_i\]

  • Identity link: the linear models of this chapter
  • Log link (strictly positive outcomes, e.g., counts): \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \operatorname{exp}\mathopen{}\left\{\sum_{i=0}^{p} \theta_i X_i\right\}\mathclose{} > 0\)
  • Logit link (dichotomous outcomes; logistic regression): \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{} = \operatorname{expit}\mathopen{}\left\{\sum_{i=0}^{p} \theta_i X_i\right\}\mathclose{} \in (0, 1)\)

For these canonical links, \(\theta\) can be estimated by maximum likelihood under a normal, Poisson, or logistic working model, respectively; the estimates are consistent as long as the model for \(\operatorname{E}\mathopen{}\left[Y|X\right]\mathclose{}\) is correct. GEE models for repeated measures are another example (Hernán and Robins 2020, 161).

A conditional mean model leaves the rest of the distribution of \(Y \mid X\) and the distribution of \(X\) unrestricted, so with continuous \(X\) or \(Y\) it is a semiparametric model for the joint distribution of \((X, Y)\).

Relaxing the parametric form:

  • Kernel regression estimates \(\operatorname{E}\mathopen{}\left[Y|X = x\right]\mathclose{}\) by \(\sum_{i=1}^nw_h(x - X_i) Y_i \big/ \sum_{i=1}^nw_h(x - X_i)\), where the kernel \(w_h(z)\) is positive, maximal at \(z = 0\), and decreases to 0 as \(|z|\) grows, at a rate set by the bandwidth \(h\)
  • Generalized additive models replace \(\sum_i \theta_i X_i\) by a sum of smooth functions \(\sum_i f_i(X_i)\), fitted by backfitting

Kernel regression borrows information only from values of \(X\) near \(x\), so it is “nonparametric” in a different sense from this chapter’s nonparametric estimators, which used only individuals with \(X\) exactly equal to \(x\).

6 Summary

  1. Data cannot always speak for themselves: with sparse data, nonparametric (saturated) estimators of \(\operatorname{E}\mathopen{}\left[Y|A = a\right]\mathclose{}\) can be imprecise or undefined.
  2. Parametric models are a priori restrictions on the distribution of the data; they let us estimate otherwise inestimable quantities by borrowing information, but only correctly if the model is correctly specified.
  3. Saturated models impose no restrictions; Part I’s methods were nonparametric estimators under saturated models.
  4. Smoothing: fewer parameters, smoother fit, more borrowing of information.
  5. Bias-variance trade-off: more flexible models protect against misspecification bias at the cost of variance.

7 References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.