Chapter 10: Random Variability

Published

Last modified: 2026-10-09 13:21:38 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

So far we have treated study populations as if they were effectively infinite, so that the only obstacles to causal inference were the systematic biases of the previous chapters: confounding, selection bias, and measurement bias. Real studies are finite, and random variability is a second, qualitatively different reason why causal inferences may be wrong. This chapter, the last of Part I, introduces random variability and the problems it raises for causal inference.

This chapter is based on Hernán and Robins (2020, chap. 10, pp. 137-150). It closes Part I (Causal inference without models) and motivates Part II (Causal inference with models).

Example 1 (A study that is too small) The book opens with a toy randomized experiment (“does my looking up make other pedestrians look up?”) with only 4 pedestrians, 2 per arm. By chance, 1 of the 2 pedestrians in the “looking up” arm, and neither in the “looking straight” arm, was blind, so half of the treated group could not respond to treatment and the estimated effect was diluted. Nothing about this study is systematically biased: no confounding (randomized), no selection bias (all responses recorded), no measurement bias. The problem is purely the small size of the study population.

1 10.1 Identification Versus Estimation (pp. 137-139)


Definition 1 (Identification and estimation) Identification problems are problems in which we can act as if the study population were effectively infinite.

Estimation is the problem of learning about a causal effect from a finite sample, in the presence of random variability.

Part I reduced causal inference to an identification problem: under exchangeability, positivity, and consistency, can the average causal effect be computed from the data?

In Chapter 2 the twenty-person heart transplant study was treated as if each individual stood for 1 billion identical individuals. That device let us ignore random fluctuations and concentrate on systematic biases, but it is unrealistic: real study populations are not effectively infinite.

Identifiability was introduced in Section 3.1 and revisited in Sections 7.2 and 8.4 (Hernán and Robins 2020, chap. 10, p. 137). For background on statistics, the book points to Wasserman (2004) and Casella and Berger (2002).


1.1 Estimands, Estimators, and Estimates

Definition 2 (Super-population) A super-population is a population so large that we can regard it as infinite, from which the study individuals are a random sample.

Suppose the 20 individuals of the heart transplant study in Chapter 2 (Table 2.2) are a random sample from a super-population (Definition 2), and we want to make inferences about that super-population.

Definition 3 (Estimand, estimator, estimate)  

  • Estimand: the parameter of interest in the super-population, e.g. \(\Pr[Y = 1 \mid A = a]\).
  • Estimator: a rule that takes the data from any sample and produces a numerical value for the estimand.
  • Estimate: the value the estimator produces in a particular sample (a point estimate).

Example 2 (Sample proportion as an estimator)  

  • Estimand: \(\Pr[Y = 1 \mid A = 1]\) in the super-population.
  • Estimator: the sample proportion \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{}\).
  • Estimate: in the heart transplant study, 7 of the 13 treated individuals died, so \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} = 7/13 \approx 0.54\).

A different random sample of 20 individuals would give a different estimate.


1.2 Consistency of Estimators

Definition 4 (Consistent estimator) An estimator is consistent for an estimand if its estimates get arbitrarily close to the estimand as the sample size \(n\) grows.

This statistical property of an estimator is unrelated to the identifiability condition also called consistency, which links each individual’s observed outcome to the counterfactual outcome under the treatment actually received.

Example 3 (The sample proportion is a consistent estimator) The sample proportion \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\) is a consistent estimator of \(\Pr[Y = 1 \mid A = a]\): the larger \(n\), the smaller we expect \(\mathopen{}\left|\Pr[Y = 1 \mid A = a] - \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\right|\mathclose{}\) to be.

WarningA Consistent Estimator Does Not Protect a Small Study

Even a consistent estimator (Definition 4) can produce a point estimate far from the super-population value, especially in small samples. So we should have more confidence in estimates from larger studies, and we quantify that confidence with confidence intervals.


1.3 Wald Confidence Intervals

Definition 5 (Wald confidence interval) A 95% Wald confidence interval for a parameter \(\theta\) based on an estimator \(\hat\theta\) is

\[ \hat\theta\pm 1.96 \times \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat\theta\right)\mathclose{}, \]

where \(\mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat\theta\right)\mathclose{}\) estimates the (exact or large-sample) standard error of \(\hat\theta\), and 1.96 is the upper 97.5% quantile of the standard normal distribution.

Example 4 (Wald interval for the risk among the treated) Let \(p = \Pr[Y = 1 \mid A = 1]\) and \(\hat p = \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} = 7/13\), with \(n = 13\). The standard error of a binomial proportion is \(\sqrt{p(1-p)/n}\), so

\[ \begin{aligned} \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat p\right)\mathclose{} &= \sqrt{\frac{\hat p (1 - \hat p)}{n}} = \sqrt{\frac{(7/13)(6/13)}{13}} \\ &= \sqrt{\frac{42/169}{13}} = \sqrt{\frac{42}{2197}} \approx \sqrt{0.0191} \approx 0.138. \end{aligned} \]

The 95% Wald interval is

\[ \begin{aligned} \hat p \pm 1.96 \times 0.138 &= 0.538 \pm 0.271 \\ &= (0.27, 0.81). \end{aligned} \]


1.4 Calibrated, Conservative, and Valid Intervals

Definition 6 (Calibration and validity of a 95% confidence interval) A 95% confidence interval is

  • calibrated if it contains the estimand in 95% of random samples;
  • conservative if it contains the estimand in more than 95% of samples;
  • anticonservative otherwise.

It is valid if, for every value of the true parameter, it is calibrated or conservative (coverage of at least 95%).

Among valid intervals, we prefer the narrowest.

Definition 7 (Small-sample and large-sample validity) An interval is

  • small-sample valid if it is valid (Definition 6) at every sample size for which it is defined (small-sample calibrated intervals are sometimes called exact);
  • large-sample (asymptotically) valid if, for every value \(\theta\) of the true parameter, \(\liminf_{n \to \infty} \Pr_\theta[\theta \in \text{CI}_n] \ge 0.95\), where \(\text{CI}_n\) is the interval computed from a sample of size \(n\).

A large-sample valid interval may still have coverage below 95% at infinitely many sample sizes, provided the shortfall shrinks to nothing as \(n\) grows.

Remark 1 (What the book means by a 95% confidence interval). The Wald interval for \(p\) in Example 4 is large-sample valid (Definition 7) but not small-sample valid. In the book, “95% confidence interval” means a large-sample valid interval, such as a Wald interval, unless stated otherwise.

WarningWald Intervals Can Undercover in Small Samples

The Wald interval centered at \(\hat p\) is only guaranteed to be valid in large samples, and it can be badly anticonservative in small samples (Brown et al. 2001). The book assumes, for simplicity, that the sample is large enough for the Wald interval in the example to be valid. Small-sample valid intervals for \(p\) exist but are not often used in practice.


1.5 Interpreting a Confidence Interval

Remark 2 (Coverage over repeated studies). Validity is a property of the interval procedure, not of any one interval: over repeated samples from the super-population, the procedure covers its target with probability at least 95%. Picture every study that we and other researchers will ever run, each using a 95% procedure that is small-sample valid. If we idealize that collection as an unending sequence, the share of those intervals that contain their targets would be at least 95%. For the book’s large-sample valid intervals, such as Wald intervals, this guarantee holds only in the limit as samples grow; at a given sample size their coverage can fall well short of 95%. In every case, nothing in the data would reveal which intervals missed.

WarningA Confidence Interval Is Not a Probability Statement About the Estimand

A single 95% interval does not mean there is a 95% probability that the estimand lies in it. The estimand is fixed, so the probability that it lies in \((0.27, 0.81)\) is either 0 or 1.

A frequentist confidence level refers to how often the interval traps the unknown super-population quantity across a collection of studies, or across hypothetical repetitions of one study.

Remark 3 (Credible intervals). In contrast, a Bayesian 95% credible interval can be interpreted as “there is a 95% probability that the estimand is in the interval”, because for a Bayesian, probability is a degree of belief rather than a frequency. The book adopts the frequency definition of probability (see its Fine Point 11.2 for Bayesian intervals).

NoteFine Point 10.1: Honest Confidence Intervals

For a large-sample valid interval, how large the sample must be before coverage comes close to 95% may depend on the unknown true parameter. A large-sample valid interval is uniform or honest if there is a sample size \(n\) at which it covers the true parameter at least 95% of the time whatever the true parameter is. Without uniformity, at any finite \(n\) some data-generating distributions may have coverage well below 95%. Even for honest intervals, the smallest such \(n\) is generally unknown and hard to determine, even by simulation (Robins and Ritov 1997).

From here on, the book uses “valid confidence interval” to mean a large-sample honest interval. Any small-sample valid interval is automatically honest for all \(n\) at which it is defined.


1.6 Bias Means Failing to Center a Valid Wald Interval

Definition 8 (Unbiased estimator (informal)) An estimator is unbiased if it can center a valid Wald interval and biased if it cannot.

Not every consistent estimator can center a valid Wald interval, even in large samples. Most users of statistics use Definition 8, and the book adopts it for now.

WarningConfidence Intervals Ignore Systematic Bias

Confidence intervals quantify only random error, so they can overstate our confidence when systematic biases may be present.

NoteTechnical Point 10.1: Bias and Consistency in Statistical Inference

Suppose we observe \(n\) i.i.d. copies of a random vector whose distribution \(P\) lies in a model \(\mathcal{M}\).

  • \(\hat\theta_n\) is a consistent estimator (Definition 4) of \(\theta = \theta(P)\) in \(\mathcal{M}\) if, formally, for every \(P \in \mathcal{M}\) and every \(\varepsilon > 0\), \[\Pr_P\mathopen{}\left[\mathopen{}\left|\hat\theta_n - \theta(P)\right|\mathclose{} > \varepsilon\right]\mathclose{} \to 0 \text{ as } n \to \infty.\]
  • \(\hat\theta_n\) is exactly unbiased in \(\mathcal{M}\) if \(\operatorname{E}_{P}\mathopen{}\left[\hat\theta_n\right]\mathclose{} = \theta(P)\) for every \(P \in \mathcal{M}\); the exact bias under \(P\) is \(\operatorname{E}_{P}\mathopen{}\left[\hat\theta_n\right]\mathclose{} - \theta(P)\). For many parameters, such as the risk ratio \(\Pr[Y=1 \mid A=1]/\Pr[Y=1 \mid A=0]\), no exactly unbiased estimator exists.
  • A systematically biased estimator is neither consistent nor exactly unbiased.

Robins and Morgenstern (1987) argue that applied researchers call an estimator unbiased only if it can center a valid Wald interval, and show that this holds only for estimators that are uniformly asymptotically normal and unbiased (UANU): there exist sequences \(\sigma_n(P)\) such that

\[ \sup_{P \in \mathcal{M}} \mathopen{}\left|\Pr_P\mathopen{}\left[n^{1/2} \mathopen{}\left(\hat\theta_n - \theta(P)\right)\mathclose{} / \sigma_n(P) < t\right]\mathclose{} - \Phi(t)\right|\mathclose{} \to 0 \text{ as } n \to \infty \]

for every \(t \in \mathbb{R}\), where \(\Phi\) is the standard normal cumulative distribution function (Robins and Ritov 1997). All inconsistent estimators, and some consistent ones (examples in Chapter 18), are biased under this definition. In the book, “unbiased” without qualification means UANU.

2 10.2 Estimation of Causal Effects (pp. 140-141)


Suppose the heart transplant study were a marginally randomized experiment whose 20 individuals are a random sample from a near-infinite super-population.

Proposition 1 (Identification in a randomized super-population) Suppose that in a near-infinite super-population

  • everyone was randomly assigned to \(A = 1\) or \(A = 0\) with the same assignment probabilities for every individual (marginal randomization), each probability positive,
  • everyone adhered to their assigned treatment, and
  • consistency holds: \(Y = Y^a\) for everyone with \(A = a\).

Then exchangeability \(Y^a \perp\!\!\!\perp A\) holds in the super-population, and \(\Pr[Y^a = 1] = \Pr[Y = 1 \mid A = a]\) for \(a = 0, 1\), so

\[ \Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1] = \Pr[Y = 1 \mid A = 1] - \Pr[Y = 1 \mid A = 0]. \]

Proof. Marginal randomization of everyone in the super-population, with the same assignment probabilities for every individual, makes the assigned treatment independent of the counterfactual outcomes, and full adherence makes the treatment received equal to the treatment assigned, so \(Y^a \perp\!\!\!\perp A\). Because each arm has positive probability, the conditional probabilities given \(A = a\) are defined, and

\[ \begin{aligned} \Pr[Y^a = 1] &= \Pr[Y^a = 1 \mid A = a] && \text{(exchangeability, } Y^a \perp\!\!\!\perp A\text{)} \\ &= \Pr[Y = 1 \mid A = a] && \text{(consistency)}. \end{aligned} \]

Subtracting the equation for \(a = 0\) from the equation for \(a = 1\) gives the displayed risk difference.

Because the study population is a random sample, \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\) is an unbiased estimator of \(\Pr[Y = 1 \mid A = a]\), and, by exchangeability in the super-population (Proposition 1), also of \(\Pr[Y^a = 1]\).


2.1 Testing and Estimation in the Randomized Experiment

Example 5 (Testing and estimation in the heart transplant trial) Testing the causal null hypothesis \(\Pr[Y^{a=1} = 1] = \Pr[Y^{a=0} = 1]\) reduces to comparing the sample proportions \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} = 7/13\) and \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 0\right]\mathclose{} = 3/7\).

Standard methods give confidence intervals for the causal effects in the super-population, estimated by

  • causal risk difference: \(7/13 - 3/7\);
  • causal risk ratio: \((7/13)/(3/7)\).

In observational studies, slightly more involved but standard procedures give confidence intervals for standardized, IP weighted, or stratified association measures.


2.2 Randomizing the Sample Instead of the Super-Population

Now suppose only the 20 study individuals, not the whole super-population, were randomized.

WarningRandomizing the Sample Does Not Make the Sample Exchangeable

Because of sampling variability, exchangeability need not hold exactly in the sample.

Example 6 (Chance imbalance in prognosis) Classify each individual as good or bad prognosis at randomization, and call the arms exchangeable if they contain the same proportion with bad prognosis. By chance, 2 of the 13 individuals assigned to \(A = 1\) (\(\approx 15\%\)) and 3 of the 7 assigned to \(A = 0\) (\(\approx 43\%\)) might have bad prognosis. In a larger sample, the relative imbalance between arms would very probably shrink.

Remark 4 (Two targets of inference). There are then two possible targets of inference:

  1. The randomized sample itself (agnostic about any super-population): randomization-based inference (Robins 1988), whose technicalities are beyond the book’s scope.
  2. The super-population from which the sample was drawn: mathematically equivalent to the first conceptualization in this section (Proposition 1). Randomization followed by random sampling is equivalent to random sampling followed by randomization.

We are often not interested in the first target. In a trial comparing two first-line cancer treatments, learning which is better comes too late to help the trial participants; the study is meant to guide treatment for a larger group of individuals similar to those studied. So far we have assumed that this larger group is a super-population from which participants were randomly sampled.

NoteFine Point 10.2: Uncertainty from Systematic Biases

The width of a Wald-type interval depends on the estimator’s standard error, so it reflects only random error. Systematic biases (confounding, selection, measurement) are another source of uncertainty and do not shrink as the sample grows. Hence the larger the study, the more a usual Wald interval understates the total uncertainty.

Some authors therefore call these intervals compatibility intervals: they show which effect sizes are highly compatible with the data under our adjustment assumptions and methods (Amrhein et al. 2019; Greenland 2019), without claiming confidence that adjustment removed all systematic bias.

Discussion sections usually treat systematic bias informally. Quantitative bias analysis instead produces intervals that combine random and systematic uncertainty (Lash et al. 2009); Bayesian alternatives also exist.

3 10.3 The Myth of the Super-Population (pp. 142-143)


Chapter 1 named two sources of randomness: sampling variability and nondeterministic counterfactuals.

WarningThe Binomial Interval Assumes a Binomial Count

Nearly every investigator would report the binomial interval \(\hat p \pm 1.96 \sqrt{\hat p (1 - \hat p)/n}\) for \(p\) from \(\hat p = 7/13\). The interval’s variance term \(\hat p (1 - \hat p)/n\) estimates the binomial variance of \(\hat p\), which is the actual variance of \(\hat p\) when the number of individuals with the outcome is binomial. For any other count, using that variance term needs a separate justification.

When is the count binomial?


3.1 Two Scenarios That Give a Binomial Distribution

Remark 5 (Two scenarios that give a binomial count). Scenario 1 (super-population). The study population is a random sample from an essentially infinite super-population (sometimes called the source or target population), and the estimand is \(p = \Pr[Y = 1 \mid A = 1]\) in that super-population. Then, across repeated random samples of 13 treated individuals, the number who develop the outcome is binomial with success probability \(p\), and the Wald interval from Example 4 is asymptotically calibrated for \(p\).

Scenario 2 (stochastic counterfactuals, no super-population). Each of the 13 treated individuals \(i\) has a nondeterministic counterfactual probability \(p_i^{a=1}\), the observed outcome \(Y_i = Y_i^{a=1}\) occurs with probability \(p_i^{a=1}\), and \(p_i^{a=1}\) equals the same value \(p\) for all 13. Then the number of treated with the outcome is binomial with success probability \(p\), and the interval is again asymptotically calibrated for \(p\).

“i.i.d.” in Technical Point 10.1 means that the data were a random sample of size \(n\) from a super-population, i.e., Scenario 1. The book credits Robins (1988) with a more detailed discussion of the two scenarios.


3.2 Scenario 2 Is Untenable

Proposition 2 (Heterogeneous risks shrink the variance of a count) Let the outcomes of \(n\) individuals be independent Bernoulli(\(p_i\)), with mean risk \(\bar p = \sum_i p_i / n\). Their sum (a Poisson-binomial count) has variance \(\sum_i p_i(1-p_i) = n \bar p (1 - \bar p) - \sum_i (p_i - \bar p)^2\), which is smaller than the binomial variance \(n \bar p (1 - \bar p)\) at the same mean \(\bar p\) whenever the \(p_i\) are unequal.

Proof. The variance of a sum of independent variables is the sum of their variances, and a Bernoulli(\(p_i\)) variable has variance \(p_i(1-p_i)\). Then write

\[ \begin{aligned} \sum_i p_i(1-p_i) &= \sum_i p_i - \sum_i p_i^2 \\ &= n\bar p - \mathopen{}\left(n \bar p^2 + \sum_i (p_i - \bar p)^2\right)\mathclose{} \\ &= n \bar p (1 - \bar p) - \sum_i (p_i - \bar p)^2, \end{aligned} \]

using \(\sum_i p_i^2 = n\bar p^2 + \sum_i (p_i - \bar p)^2\). The subtracted sum of squares is positive unless all \(p_i\) equal \(\bar p\). (This derivation is ours, added to explain the book’s statement.)

Remark 6 (Heterogeneous risks rule out Scenario 2). The risks \(p_i^{a=1}\) almost certainly vary between individuals (for example, with genetic make-up). If they are not constant, the natural estimand for the study population is their average, say \(\bar p\), but the number of treated with the outcome is then not binomial with success probability \(\bar p\). By Proposition 2 the binomial formula overstates the variance of the count, by the sum of squared deviations \(\sum_i (p_i - \bar p)^2\). Suppose that \(\bar p\) stays bounded away from 0 and 1 as the number of individuals grows. Whether the overstatement then matters in large samples depends on the ratio of the average squared deviation \(\frac{1}{n} \sum_i (p_i - \bar p)^2\) to \(\bar p (1 - \bar p)\). If this ratio stays bounded away from zero, the binomial Wald interval is not asymptotically calibrated. Under the normal approximation it stays too wide, so it is asymptotically conservative, and for large enough samples its coverage stays above some level greater than 95%. If instead this ratio tends to zero (for example, with half the individuals at risk \(0.5 + 1/n\) and half at \(0.5 - 1/n\), for even \(n \ge 2\)), the overstatement becomes negligible relative to the binomial variance \(n \bar p (1 - \bar p)\), and the coverage still tends to 95%.

WarningA Binomial Interval Assumes a Super-Population

Anyone who reports a binomial interval for \(\Pr[Y = 1 \mid A = a]\) as if the count were binomial, while acknowledging between-individual variation in risk that persists as the study grows, is implicitly assuming Scenario 1. Under Scenario 1 the count is binomial whether the counterfactuals are deterministic or stochastic. If the target is instead the study population itself, and its risks vary in this persistent way (with \(\bar p\) bounded away from 0 and 1), Remark 6 shows that the interval is asymptotically conservative rather than calibrated: for large enough samples its coverage stays above some level greater than 95%.


3.3 Why Does the Myth Endure?

Remark 7 (Why the super-population myth endures). Advantage of the super-population: nothing hinges on whether the world is deterministic. Yet it is usually a fiction; individuals are rarely sampled at random from any near-infinite population. It endures because

  1. it leads to simple statistical methods, and
  2. it gives a simple route to generalization from the study population (e.g., our 20 individuals) to a large target population (e.g., all immortals in the Greek pantheon).
WarningSuper-Population Intervals Are What-If Statements

Because the random sampling is fictional, a 95% interval computed under Scenario 1 should be read as a what-if statement: it covers the super-population parameter had the study individuals been randomly sampled from a near-infinite super-population.

An investigator might reject Scenario 1 if the pool of potential treatment recipients is not much larger than the study population, or if the target population is believed to differ from the study population in ways that sampling variability cannot explain. The rest of the chapter accepts that individuals were randomly sampled from a super-population, and studies the consequences of random variability for causal inference, starting with a simple randomized experiment.

4 10.4 The Conditionality “Principle” (pp. 144-147)


Example 7 (A randomized trial, unadjusted (Table 10.1)) A trial of treatment \(A\) on 1-year death \(Y\) randomized 240 individuals, 120 per arm.

\(Y=1\) \(Y=0\)
\(A = 1\) 24 96
\(A = 0\) 42 78

Unadjusted risk difference:

\[ \frac{24}{120} - \frac{42}{120} = 0.20 - 0.35 = -0.15. \]

Estimated variance:

\[ \begin{aligned} \frac{(24/120)(96/120)}{120} + \frac{(42/120)(78/120)}{120} &= \frac{0.20 \times 0.80 + 0.35 \times 0.65}{120} \\ &= \frac{0.16 + 0.2275}{120} = \frac{0.3875}{120} = \frac{31}{9600}. \end{aligned} \]

95% Wald interval:

\[ -0.15 \pm 1.96 \sqrt{31/9600} = -0.15 \pm 1.96 \times 0.0568 = -0.15 \pm 0.111 = (-0.26, -0.04). \]

The investigators publish: treatment reduced the risk of death by 15 percentage points.

In a near-infinite super-population, randomization would make the arms exchangeable, \(Y^a \perp\!\!\!\perp A\), and the associational risk difference would equal the causal risk difference \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\). With only 240 individuals, randomization guarantees only that departures from exchangeability are due to chance, not that they are absent. Our ignorance of chance correlations between unmeasured baseline risk factors and \(A\) in the sample contributes to the length (0.22) of the interval.


4.1 A Chance Imbalance on Smoking

Example 8 (The same trial stratified by smoking \(L\) (Table 10.2))  

\(L = 1\) (smokers) \(Y=1\) \(Y=0\)
\(A = 1\) 4 76
\(A = 0\) 2 38
\(L = 0\) (nonsmokers) \(Y=1\) \(Y=0\)
\(A = 1\) 20 20
\(A = 0\) 40 40
  • Proportion treated: \(80/120\) among smokers, twice the \(40/120\) among nonsmokers.
  • Smokers: \(4/80 - 2/40 = 0.05 - 0.05 = 0\).
  • Nonsmokers: \(20/40 - 40/80 = 0.50 - 0.50 = 0\).

The adjusted (stratified) analysis finds no effect in either stratum, although the marginal estimate suggested a benefit.


Example 9 (The adjusted Wald interval) Assume the risk difference is the same in both strata of \(L\). The adjusted estimate is the maximum likelihood estimator (MLE) of the common risk difference, which in large samples weights the two stratum-specific estimates by their inverse variances. For a dichotomous \(L\), its estimated variance is \(\widehat V_0 \widehat V_1 / (\widehat V_0 + \widehat V_1)\) (Hernán and Robins 2020, 148), where \(\widehat V_l\) is the estimated variance of the risk difference in stratum \(L = l\). Computing each \(\widehat V_l\) from the observed number of individuals in each arm gives

\[ \begin{aligned} \widehat V_0 &= \frac{(20/40)(20/40)}{40} + \frac{(40/80)(40/80)}{80} = \frac{0.25}{40} + \frac{0.25}{80} = 0.009375, \\ \widehat V_1 &= \frac{(4/80)(76/80)}{80} + \frac{(2/40)(38/40)}{40} = \frac{0.0475}{80} + \frac{0.0475}{40} = 0.00178125, \\ \mathop{\widehat{\operatorname{Var}}}\nolimits\mathopen{}\left(\widehat{\text{sRD}}\right)\mathclose{} &= \frac{\widehat V_0 \widehat V_1}{\widehat V_0 + \widehat V_1} = \frac{0.009375 \times 0.00178125}{0.009375 + 0.00178125} = \frac{1.670 \times 10^{-5}}{0.01116} \approx 0.001497, \\ 1.96 \sqrt{0.001497} &= 1.96 \times 0.0387 \approx 0.076. \end{aligned} \]

The point estimate is 0 in both strata, so the pooled estimate is 0. The inverse-variance (MLE) Wald 95% interval centered at the adjusted estimate is therefore \((-0.076, 0.076)\), matching the book (Hernán and Robins 2020, 145).

Remark 8 (Should the article be retracted?). Malfeasance is ruled out, so the imbalance is “very very bad luck”. Should the article be retracted?

  • Investigator 1: no. Imbalances on unmeasured factors \(U\) may cancel the imbalance on \(L\), so the unadjusted estimate may still be closer to the truth.
  • Investigator 2: yes. The strong \(L\)-\(A\) association confounds the estimate; within levels of \(L\) we have mini randomized trials, whose intervals reflect possible \(U\)-\(A\) associations given \(L\).

4.2 Settling the Debate

Remark 9 (Conditional and unconditional properties of the two estimators). Assume a constant causal risk difference across strata of \(L\), and imagine running the trial trillions of times. Condition on the runs in which \(L\) and \(A\) are as strongly positively associated as observed. Within each level of \(L\), any pre-treatment risk factor \(U\) is positively associated with \(A\) in essentially as many of these runs as it is negatively associated (even if \(U\) and \(L\) are highly correlated). Therefore:

  • Conditionally (over those runs): the adjusted estimate is unbiased; the unadjusted estimate is greatly biased.
  • Unconditionally (over all runs): both are unbiased, and the adjusted estimate has smaller variance provided \(L\) is not independent of \(Y\) given \(A\).

Either way, the Wald interval centered on the adjusted estimator is the better analysis. Investigator 2 is right; the article should be retracted.

When \(L\) is not independent of \(Y\) given \(A\), the adjusted estimator is unconditionally more efficient, because, once \(L\) has been measured, it is the MLE of the risk difference (Hernán and Robins 2020, 146–47).


4.3 The Conditionality Principle

Definition 9 (Ancillary statistic) An ancillary statistic for a parameter is a statistic whose distribution does not depend on that parameter.

When the risk difference is constant across strata of \(L\), the observed \(L\)-\(A\) association is an ancillary statistic (Definition 9) for the causal risk difference.

Definition 10 (Conditionality principle) The conditionality principle is the principle that inference on a parameter should be performed conditional on ancillary statistics (Definition 9) for that parameter.

Definition 11 (Conditionally unbiased estimator) An estimator is conditionally unbiased if it can center a valid Wald interval conditional on ancillary statistics (Definition 9).

Investigators who intuitively follow the conditionality principle (Definition 10) call an estimator unbiased only if it is conditionally unbiased, a stricter definition than Definition 8.

To the frequentist objection that the unadjusted estimate is unbiased over all runs, they reply: why should \(L\)-\(A\) associations in hypothetical other studies matter, when in my study \(L\) is a confounder?

WarningConditionality Breaks Down With Many Confounders

The conditional reply (that the observed \(L\)-\(A\) association in my study is what matters) is convincing in randomized and observational studies as long as the number of measured confounders is not large. With many confounders, strictly following the conditionality principle stops being wise. The quotation marks around “principle” in the section title anticipate this point.

NoteTechnical Point 10.2: A Formal Statement of the Conditionality Principle

Take one dichotomous \(L\), exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), and a stratum-specific risk difference \(\text{sRD} = \Pr[Y = 1 \mid L = l, A = 1] - \Pr[Y = 1 \mid L = l, A = 0]\) known to be constant across strata and of interest. The likelihood factors as

\[ \prod_{i=1}^n f(Y_i \mid L_i, A_i; \text{sRD}, p_0) \times f(A_i \mid L_i; \alpha) \times f(L_i; \rho), \]

where \(p_{0l} = \Pr[Y = 1 \mid L = l, A = 0]\) is the risk among the untreated in stratum \(l\), \(p_0 = (p_{00}, p_{01})\) collects these two baseline risks, and \(p_0\), \(\alpha\), \(\rho\) are nuisance parameters.

The data on \((A, L)\) are S-ancillary for sRD: the conditional distribution of the data given \((A, L)\) depends on sRD, but the joint density of \((A, L)\) shares no parameters with \(f(Y \mid L, A; \text{sRD}, p_0)\). The conditionality principle says to condition on S-ancillary statistics, here \(\{A_i, L_i; i = 1, \ldots, n\}\). The same holds if the risk ratio, rather than the risk difference, is known to be constant. An exact ancillary is an S-ancillary statistic with known marginal distribution (here, \(\alpha\) and \(\rho\) known).

NoteTechnical Point 10.3: Approximate Ancillarity

If the stratum-specific risk differences \(\text{sRD}_l\) vary over \(L\), the population causal risk difference is the standardized risk difference

\[ \operatorname{RD}_{std} = \sum_l \mathopen{}\left[\Pr[Y = 1 \mid L = l, A = 1; \upsilon] - \Pr[Y = 1 \mid L = l, A = 0; \upsilon]\right]\mathclose{} f(l; \rho), \]

with \(\upsilon = \{\text{sRD}_l, p_{0l}; l = 0, 1\}\). In an unconditionally randomized experiment \(A \perp\!\!\!\perp L\), so \(\operatorname{RD}_{std}\) equals the associational risk difference. Because \(\operatorname{RD}_{std}\) depends on \(\rho\), \(\{A_i, L_i\}\) is no longer exactly ancillary, and no exact ancillary exists.

Let \(\widetilde S = \widehat{\operatorname{OR}}_{AL} - \operatorname{OR}_{AL}\), the sample minus the super-population \(A\)-\(L\) odds ratio, and \(\widehat S = \widetilde S / \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\widetilde S\right)\mathclose{}\), which is approximately standard normal in large samples. When \(\operatorname{OR}_{AL}\) is known (e.g., \(\operatorname{OR}_{AL} = 1\) in a randomized experiment), \(\widehat S\) is an approximate (large-sample) ancillary: it can be computed from the data, has approximately known distribution, is governed by the \(f(A \mid L; \alpha)\) factor of the likelihood only, and conditional on it the adjusted estimate of \(\operatorname{RD}_{std}\) is unbiased while the unadjusted one is biased.

Under a continuity principle (Buehler 1982), inferences should not change discontinuously after an arbitrarily small known change in the data-generating distribution. Accepting both conditionality and continuity implies conditioning on approximate ancillaries too: the extended conditionality principle.

NoteTechnical Point 10.4: Adjusted Versus Unadjusted Estimators

The adjusted estimator \(\widehat{\operatorname{RD}}_{MLE}\) replaces the population proportions in \(\operatorname{RD}_{std}\) by sample proportions; the unadjusted estimator is \(\widehat{\operatorname{RD}}_{UN} = \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} - \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 0\right]\mathclose{}\). Unconditionally, both are asymptotically normal and unbiased for \(\operatorname{RD}_{std}\).

Robins and Morgenstern (1987) show that

  • \(\widehat{\operatorname{RD}}_{MLE}\) has the same asymptotic distribution conditionally on \(\widehat S\) as unconditionally;
  • \(\operatorname{aVar}\mathopen{}\left(\widehat{\operatorname{RD}}_{MLE}\right)\mathclose{} = \operatorname{aVar}\mathopen{}\left(\widehat{\operatorname{RD}}_{UN}\right)\mathclose{} - \mathopen{}\left[\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{}\right]\mathclose{}^2\);
  • the conditional asymptotic bias of \(\widehat{\operatorname{RD}}_{UN}\) given \(\widehat S\) is \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} \, \widehat S\).

So \(\widehat{\operatorname{RD}}_{UN}\) is unconditionally inefficient if and only if it is conditionally biased (both iff \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} \neq 0\)). Moreover, \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} = 0\) iff \(L \perp\!\!\!\perp Y \mid A\), so when a measured risk factor for \(Y\) is available, \(\widehat{\operatorname{RD}}_{MLE}\) is preferred. The corrected estimator \(\widehat{\operatorname{RD}}_{UN} - \operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{}\widehat S\) has the same asymptotic distribution as \(\widehat{\operatorname{RD}}_{MLE}\) given \(\widehat S\).

NoteTechnical Point 10.5: Most Researchers Follow the Extended Conditionality Principle

With a constant sRD over a dichotomous \(L\), the estimated variance of the MLE of sRD is \(\widehat V_0 \widehat V_1 / (\widehat V_0 + \widehat V_1)\), where \(\widehat V_l\) is the estimated variance of the stratum-specific estimate \(\widehat{\operatorname{RD}}_l\).

Two choices for \(\widehat V_1\) (smokers, Table 10.2) differ only in the denominators:

\[ \begin{aligned} \widehat V_1^{obs} &= \frac{(4/80)(76/80)}{80} + \frac{(2/40)(38/40)}{40} = \frac{0.0475}{80} + \frac{0.0475}{40} \approx 1.78 \times 10^{-3}, \\ \widehat V_1^{exp} &= \frac{(4/80)(76/80)}{60} + \frac{(2/40)(38/40)}{60} = \frac{0.0475}{60} + \frac{0.0475}{60} \approx 1.58 \times 10^{-3}. \end{aligned} \]

\(\widehat V_1^{obs}\) uses the observed numbers per arm (80 and 40; observed information), \(\widehat V_1^{exp}\) the expected number (60 per arm under \(A \perp\!\!\!\perp L\); expected information). Nearly all researchers would choose \(\widehat V_1^{obs}\), the choice used in Example 9. Results of Efron and Hinkley (1978) and Robins and Morgenstern (1987) imply that this choice amounts to conditioning on the approximate ancillary \(\widehat S\), i.e., following the extended conditionality principle. (The conditional and unconditional variances of the MLE differ by order \(n^{-3/2}\); the observed-information estimator is closer than that to the conditional variance, the expected-information estimator closer than that to the unconditional variance.)

5 10.5 The Curse of Dimensionality (pp. 148-150)


The arguments of Section 10.4 rely on asymptotic theory in which the number of strata of \(L\) is small relative to the sample size.

Example 10 (A trial with 100 binary covariates) Now suppose the investigators had measured 100 binary pre-treatment variables, \(L = (L_1, \ldots, L_{100})\), with \(2^{100}\) strata: high-dimensional data.

For simplicity, assume no additive effect modification by \(L\): the super-population risk difference \(\Pr[Y = 1 \mid A = 1, L = l] - \Pr[Y = 1 \mid A = 0, L = l]\) is constant, and equal to 0, across all \(2^{100}\) strata.


5.1 The Adjusted Interval Becomes Uninformative

  • The investigators now accept the conditionality principle, so they consider the unadjusted \(-0.15\) conditionally biased.
  • But with \(2^{100}\) strata, far more than the 240 individuals, at most a few strata contain both a treated and an untreated individual.

Example 11 (One informative stratum) Suppose exactly one stratum contains one treated and one untreated individual, and no other stratum contains both. In that stratum the empirical risk difference can only be \(-1\), \(0\), or \(1\), depending on the two outcomes, so the stratum carries no usable information about the common risk difference. The 95% interval based on the adjusted estimator is \((-1, 1)\), the whole parameter space, so it is completely uninformative. The unadjusted interval is still \((-0.26, -0.04)\), since its width does not depend on how many covariates were measured.

This is the reverse of the one-covariate case. The adjusted estimator is guaranteed to be more efficient only when the ratio of individuals to unknown parameters is large.

TipIndividuals per Parameter

A common rule of thumb is at least 10 individuals per unknown parameter, though the minimum depends on the data.


5.2 Conditionality Should Not Be a Principle

Definition 12 (Curse of dimensionality) The curse of dimensionality is the situation in which the measured covariates \(L\) have many strata relative to the sample size \(n\), so that few strata contain both treated and untreated individuals and estimators conditional on all of \(L\) are uninformative.

In Example 10, the unadjusted estimator is still conditionally biased given all 100 covariates, but the estimator that conditions on them is now uninformative (Example 11).

WarningConditionality Should Not Be a Principle

Conditionality is compelling in simple examples, but it cannot be carried through in high-dimensional models, so it should not be raised to a principle (Robins and Ritov 1997). The same applies to observational studies.

Solving the curse of dimensionality is hard and an active research area; Chapter 18 reviews it. Chapters 11 through 17 supply the background on models for causal inference.

NoteTechnical Point 10.6: Can the Curse of Dimensionality Be Reversed?

With many strata of \(L\), informative inference conditional on the exact ancillary \(\{A_i, L_i\}\) is impossible whatever the estimator. Unconditional inference in a marginally randomized experiment is different: \(\widehat{\operatorname{RD}}_{UN}\) stays unbiased, with Wald interval width proportional to \(1/n^{1/2}\), regardless of the dimension of \(L\). It uses prior information the MLE does not: it is unbiased for the common risk difference only if \(A \perp\!\!\!\perp L\) is known to hold in the super-population. Even unconditionally, the MLE’s intervals remain uninformative. Can data on \(L\) give an unconditionally unbiased estimator more efficient than \(\widehat{\operatorname{RD}}_{UN}\)? Chapter 18 shows that this is sometimes possible.

NoteTechnical Point 10.7: Random Variability and Causal Discovery

Fine Point 6.3 showed that, under faithfulness and with infinite data, the causal structure can sometimes be learned. Suppose variables \(Z\), \(A\), \(Y\) occur in that order, the empirical \(Z\)-\(Y\) odds ratio is 1 at every level of \(A\), and all other odds ratios are far from 1. If \(Z \perp\!\!\!\perp Y \mid A\) held in the super-population, faithfulness would leave only \(Z \to A \to Y\) (possibly with a common cause of \(Z\) and \(A\)), so \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) would be the average causal effect.

But Robins et al. (2003) showed that faithfulness-based inferences are non-uniform: however large \(n\) is, some faithful distributions have conditional \(Z\)-\(Y\) odds ratios near enough to 1 that sample values of exactly 1 are to be expected, while \(A\) has no average causal effect on \(Y\). An honest 95% frequentist interval for that effect must therefore always include 0, however large the empirical risk difference is relative to its standard error (for instance, 0.2 at 30 standard errors).

A Bayesian who believes in faithfulness may obtain a narrow credible interval that excludes 0, but such inferences are very sensitive to the prior; Hernán and Robins argue that the prior probability of any observational-study diagram without a common cause of \(A\) and \(Y\) should be essentially zero. In finite samples, data alone cannot distinguish \(Z \perp\!\!\!\perp Y \mid A\) from near-independence, so causal discovery from observational data relies heavily on subject-matter knowledge.

6 Summary


  • Identification (Part I) assumes effectively infinite data; estimation deals with random variability in finite samples.
  • Estimands, estimators, and estimates; consistent estimators; Wald confidence intervals and their calibration, validity, and (frequentist) interpretation.
  • In the book, an estimator is unbiased if it can center a valid Wald interval (formally, uniformly asymptotically normal and unbiased, UANU); confidence intervals ignore systematic bias.
  • Standard binomial intervals implicitly assume random sampling from a super-population, usually a fiction, so they are what-if statements.
  • In a trial with a chance imbalance on a measured factor \(L\) that is not independent of \(Y\) given \(A\), and with a risk difference that is constant across strata of \(L\), the adjusted estimator is conditionally unbiased and unconditionally more efficient: the conditionality principle favors adjustment.
  • With high-dimensional \(L\), conditioning makes estimates uninformative while the unadjusted estimator stays conditionally biased: the curse of dimensionality, which motivates the models of Part II.

Part I established when causal effects are identified (under exchangeability, positivity, and consistency). Part II (Chapters 11 through 18) turns to estimation with models: why models are needed (Chapter 11), IP weighting and marginal structural models (Chapter 12), standardization and the parametric g-formula (Chapter 13), g-estimation of structural nested models (Chapter 14), outcome regression and propensity scores (Chapter 15), instrumental variable estimation (Chapter 16), causal survival analysis (Chapter 17), and variable selection and high-dimensional data (Chapter 18).

7 References


Amrhein, V, S Greenland, and B McShane. 2019. “Scientists Rise up Against Statistical Significance.” Nature 567: 305–7.
Brown, LD, T Cai, and A DasGupta. 2001. “Interval Estimation for a Binomial Proportion (with Discussion).” Statistical Science 16: 101–33.
Buehler, RJ. 1982. “Some Ancillary Statistics and Their Properties. Rejoinder.” Journal of the American Statistical Association 77: 593–94.
Casella, G, and RL Berger. 2002. Statistical Inference. 2nd ed. Duxbury Press.
Efron, B, and DV Hinkley. 1978. “Assessing the Accuracy of the Maximum Likelihood Estimator: Observed Versus Expected Fisher Information.” Biometrika 65: 457–87.
Greenland, S. 2019. “Valid P-Values Behave Exactly as They Should: Some Misleading Criticisms of P-Values and Their Resolution with S-Values.” The American Statistician 73 (supplement 1): 106–14.
Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Lash, TL, MP Fox, and AK Fink. 2009. Applying Quantitative Bias Analysis to Epidemiologic Data. Springer.
Robins, JM. 1988. “Confidence Intervals for Causal Parameters.” Statistics in Medicine 7: 773–85.
Robins, JM, and H Morgenstern. 1987. “The Foundations of Confounding in Epidemiology.” Computers & Mathematics with Applications 14 (9-12): 869–916.
Robins, JM, and Y Ritov. 1997. “Toward a Curse of Dimensionality Appropriate (CODA) Asymptotic Theory for Semiparametric Models.” Statistics in Medicine 16: 285–319.
Robins, JM, R Scheines, P Spirtes, and L Wasserman. 2003. “Uniform Consistency in Causal Inference.” Biometrika 90 (3): 491–515.
Wasserman, L. 2004. All of Statistics: A Concise Course in Statistical Inference. Springer.
Back to top