Chapter 10: Random Variability
So far we have treated study populations as if they were effectively infinite, so that the only obstacles to causal inference were the systematic biases of the previous chapters: confounding, selection bias, and measurement bias. Real studies are finite, and random variability is a second, qualitatively different reason why causal inferences may be wrong. This chapter, the last of Part I, introduces random variability and the problems it raises for causal inference.
This chapter is based on Hernán and Robins (2020, chap. 10, pp. 137-150). It closes Part I (Causal inference without models) and motivates Part II (Causal inference with models).
1 10.1 Identification Versus Estimation (pp. 137-139)
Part I reduced causal inference to an identification problem: under exchangeability, positivity, and consistency, can the average causal effect be computed from the data?
In Chapter 2 the twenty-person heart transplant study was treated as if each individual stood for 1 billion identical individuals. That device let us ignore random fluctuations and concentrate on systematic biases, but it is unrealistic: real study populations are not effectively infinite.
Identifiability was introduced in Section 3.1 and revisited in Sections 7.2 and 8.4 (Hernán and Robins 2020, chap. 10, p. 137). For background on statistics, the book points to Wasserman (2004) and Casella and Berger (2002).
1.1 Estimands, Estimators, and Estimates
Suppose the 20 individuals of the heart transplant study in Chapter 2 (Table 2.2) are a random sample from a super-population (Definition 2), and we want to make inferences about that super-population.
1.2 Consistency of Estimators
This statistical property of an estimator is unrelated to the identifiability condition also called consistency, which links each individual’s observed outcome to the counterfactual outcome under the treatment actually received.
Even a consistent estimator (Definition 4) can produce a point estimate far from the super-population value, especially in small samples. So we should have more confidence in estimates from larger studies, and we quantify that confidence with confidence intervals.
1.3 Wald Confidence Intervals
1.4 Calibrated, Conservative, and Valid Intervals
Among valid intervals, we prefer the narrowest.
A large-sample valid interval may still have coverage below 95% at infinitely many sample sizes, provided the shortfall shrinks to nothing as \(n\) grows.
The Wald interval centered at \(\hat p\) is only guaranteed to be valid in large samples, and it can be badly anticonservative in small samples (Brown et al. 2001). The book assumes, for simplicity, that the sample is large enough for the Wald interval in the example to be valid. Small-sample valid intervals for \(p\) exist but are not often used in practice.
1.5 Interpreting a Confidence Interval
A single 95% interval does not mean there is a 95% probability that the estimand lies in it. The estimand is fixed, so the probability that it lies in \((0.27, 0.81)\) is either 0 or 1.
A frequentist confidence level refers to how often the interval traps the unknown super-population quantity across a collection of studies, or across hypothetical repetitions of one study.
For a large-sample valid interval, how large the sample must be before coverage comes close to 95% may depend on the unknown true parameter. A large-sample valid interval is uniform or honest if there is a sample size \(n\) at which it covers the true parameter at least 95% of the time whatever the true parameter is. Without uniformity, at any finite \(n\) some data-generating distributions may have coverage well below 95%. Even for honest intervals, the smallest such \(n\) is generally unknown and hard to determine, even by simulation (Robins and Ritov 1997).
From here on, the book uses “valid confidence interval” to mean a large-sample honest interval. Any small-sample valid interval is automatically honest for all \(n\) at which it is defined.
1.6 Bias Means Failing to Center a Valid Wald Interval
Not every consistent estimator can center a valid Wald interval, even in large samples. Most users of statistics use Definition 8, and the book adopts it for now.
Confidence intervals quantify only random error, so they can overstate our confidence when systematic biases may be present.
Suppose we observe \(n\) i.i.d. copies of a random vector whose distribution \(P\) lies in a model \(\mathcal{M}\).
- \(\hat\theta_n\) is a consistent estimator (Definition 4) of \(\theta = \theta(P)\) in \(\mathcal{M}\) if, formally, for every \(P \in \mathcal{M}\) and every \(\varepsilon > 0\), \[\Pr_P\mathopen{}\left[\mathopen{}\left|\hat\theta_n - \theta(P)\right|\mathclose{} > \varepsilon\right]\mathclose{} \to 0 \text{ as } n \to \infty.\]
- \(\hat\theta_n\) is exactly unbiased in \(\mathcal{M}\) if \(\operatorname{E}_{P}\mathopen{}\left[\hat\theta_n\right]\mathclose{} = \theta(P)\) for every \(P \in \mathcal{M}\); the exact bias under \(P\) is \(\operatorname{E}_{P}\mathopen{}\left[\hat\theta_n\right]\mathclose{} - \theta(P)\). For many parameters, such as the risk ratio \(\Pr[Y=1 \mid A=1]/\Pr[Y=1 \mid A=0]\), no exactly unbiased estimator exists.
- A systematically biased estimator is neither consistent nor exactly unbiased.
Robins and Morgenstern (1987) argue that applied researchers call an estimator unbiased only if it can center a valid Wald interval, and show that this holds only for estimators that are uniformly asymptotically normal and unbiased (UANU): there exist sequences \(\sigma_n(P)\) such that
\[ \sup_{P \in \mathcal{M}} \mathopen{}\left|\Pr_P\mathopen{}\left[n^{1/2} \mathopen{}\left(\hat\theta_n - \theta(P)\right)\mathclose{} / \sigma_n(P) < t\right]\mathclose{} - \Phi(t)\right|\mathclose{} \to 0 \text{ as } n \to \infty \]
for every \(t \in \mathbb{R}\), where \(\Phi\) is the standard normal cumulative distribution function (Robins and Ritov 1997). All inconsistent estimators, and some consistent ones (examples in Chapter 18), are biased under this definition. In the book, “unbiased” without qualification means UANU.
2 10.2 Estimation of Causal Effects (pp. 140-141)
Suppose the heart transplant study were a marginally randomized experiment whose 20 individuals are a random sample from a near-infinite super-population.
Proof. Marginal randomization of everyone in the super-population, with the same assignment probabilities for every individual, makes the assigned treatment independent of the counterfactual outcomes, and full adherence makes the treatment received equal to the treatment assigned, so \(Y^a \perp\!\!\!\perp A\). Because each arm has positive probability, the conditional probabilities given \(A = a\) are defined, and
\[ \begin{aligned} \Pr[Y^a = 1] &= \Pr[Y^a = 1 \mid A = a] && \text{(exchangeability, } Y^a \perp\!\!\!\perp A\text{)} \\ &= \Pr[Y = 1 \mid A = a] && \text{(consistency)}. \end{aligned} \]
Subtracting the equation for \(a = 0\) from the equation for \(a = 1\) gives the displayed risk difference.
Because the study population is a random sample, \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\) is an unbiased estimator of \(\Pr[Y = 1 \mid A = a]\), and, by exchangeability in the super-population (Proposition 1), also of \(\Pr[Y^a = 1]\).
2.1 Testing and Estimation in the Randomized Experiment
In observational studies, slightly more involved but standard procedures give confidence intervals for standardized, IP weighted, or stratified association measures.
2.2 Randomizing the Sample Instead of the Super-Population
Now suppose only the 20 study individuals, not the whole super-population, were randomized.
Because of sampling variability, exchangeability need not hold exactly in the sample.
We are often not interested in the first target. In a trial comparing two first-line cancer treatments, learning which is better comes too late to help the trial participants; the study is meant to guide treatment for a larger group of individuals similar to those studied. So far we have assumed that this larger group is a super-population from which participants were randomly sampled.
The width of a Wald-type interval depends on the estimator’s standard error, so it reflects only random error. Systematic biases (confounding, selection, measurement) are another source of uncertainty and do not shrink as the sample grows. Hence the larger the study, the more a usual Wald interval understates the total uncertainty.
Some authors therefore call these intervals compatibility intervals: they show which effect sizes are highly compatible with the data under our adjustment assumptions and methods (Amrhein et al. 2019; Greenland 2019), without claiming confidence that adjustment removed all systematic bias.
Discussion sections usually treat systematic bias informally. Quantitative bias analysis instead produces intervals that combine random and systematic uncertainty (Lash et al. 2009); Bayesian alternatives also exist.
3 10.3 The Myth of the Super-Population (pp. 142-143)
Chapter 1 named two sources of randomness: sampling variability and nondeterministic counterfactuals.
Nearly every investigator would report the binomial interval \(\hat p \pm 1.96 \sqrt{\hat p (1 - \hat p)/n}\) for \(p\) from \(\hat p = 7/13\). The interval’s variance term \(\hat p (1 - \hat p)/n\) estimates the binomial variance of \(\hat p\), which is the actual variance of \(\hat p\) when the number of individuals with the outcome is binomial. For any other count, using that variance term needs a separate justification.
When is the count binomial?
3.1 Two Scenarios That Give a Binomial Distribution
“i.i.d.” in Technical Point 10.1 means that the data were a random sample of size \(n\) from a super-population, i.e., Scenario 1. The book credits Robins (1988) with a more detailed discussion of the two scenarios.
3.2 Scenario 2 Is Untenable
Proof. The variance of a sum of independent variables is the sum of their variances, and a Bernoulli(\(p_i\)) variable has variance \(p_i(1-p_i)\). Then write
\[ \begin{aligned} \sum_i p_i(1-p_i) &= \sum_i p_i - \sum_i p_i^2 \\ &= n\bar p - \mathopen{}\left(n \bar p^2 + \sum_i (p_i - \bar p)^2\right)\mathclose{} \\ &= n \bar p (1 - \bar p) - \sum_i (p_i - \bar p)^2, \end{aligned} \]
using \(\sum_i p_i^2 = n\bar p^2 + \sum_i (p_i - \bar p)^2\). The subtracted sum of squares is positive unless all \(p_i\) equal \(\bar p\). (This derivation is ours, added to explain the book’s statement.)
Anyone who reports a binomial interval for \(\Pr[Y = 1 \mid A = a]\) as if the count were binomial, while acknowledging between-individual variation in risk that persists as the study grows, is implicitly assuming Scenario 1. Under Scenario 1 the count is binomial whether the counterfactuals are deterministic or stochastic. If the target is instead the study population itself, and its risks vary in this persistent way (with \(\bar p\) bounded away from 0 and 1), Remark 6 shows that the interval is asymptotically conservative rather than calibrated: for large enough samples its coverage stays above some level greater than 95%.
3.3 Why Does the Myth Endure?
Because the random sampling is fictional, a 95% interval computed under Scenario 1 should be read as a what-if statement: it covers the super-population parameter had the study individuals been randomly sampled from a near-infinite super-population.
An investigator might reject Scenario 1 if the pool of potential treatment recipients is not much larger than the study population, or if the target population is believed to differ from the study population in ways that sampling variability cannot explain. The rest of the chapter accepts that individuals were randomly sampled from a super-population, and studies the consequences of random variability for causal inference, starting with a simple randomized experiment.
4 10.4 The Conditionality “Principle” (pp. 144-147)
In a near-infinite super-population, randomization would make the arms exchangeable, \(Y^a \perp\!\!\!\perp A\), and the associational risk difference would equal the causal risk difference \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\). With only 240 individuals, randomization guarantees only that departures from exchangeability are due to chance, not that they are absent. Our ignorance of chance correlations between unmeasured baseline risk factors and \(A\) in the sample contributes to the length (0.22) of the interval.
4.1 A Chance Imbalance on Smoking
4.2 Settling the Debate
When \(L\) is not independent of \(Y\) given \(A\), the adjusted estimator is unconditionally more efficient, because, once \(L\) has been measured, it is the MLE of the risk difference (Hernán and Robins 2020, 146–47).
4.3 The Conditionality Principle
When the risk difference is constant across strata of \(L\), the observed \(L\)-\(A\) association is an ancillary statistic (Definition 9) for the causal risk difference.
Investigators who intuitively follow the conditionality principle (Definition 10) call an estimator unbiased only if it is conditionally unbiased, a stricter definition than Definition 8.
To the frequentist objection that the unadjusted estimate is unbiased over all runs, they reply: why should \(L\)-\(A\) associations in hypothetical other studies matter, when in my study \(L\) is a confounder?
The conditional reply (that the observed \(L\)-\(A\) association in my study is what matters) is convincing in randomized and observational studies as long as the number of measured confounders is not large. With many confounders, strictly following the conditionality principle stops being wise. The quotation marks around “principle” in the section title anticipate this point.
Take one dichotomous \(L\), exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), and a stratum-specific risk difference \(\text{sRD} = \Pr[Y = 1 \mid L = l, A = 1] - \Pr[Y = 1 \mid L = l, A = 0]\) known to be constant across strata and of interest. The likelihood factors as
\[ \prod_{i=1}^n f(Y_i \mid L_i, A_i; \text{sRD}, p_0) \times f(A_i \mid L_i; \alpha) \times f(L_i; \rho), \]
where \(p_{0l} = \Pr[Y = 1 \mid L = l, A = 0]\) is the risk among the untreated in stratum \(l\), \(p_0 = (p_{00}, p_{01})\) collects these two baseline risks, and \(p_0\), \(\alpha\), \(\rho\) are nuisance parameters.
The data on \((A, L)\) are S-ancillary for sRD: the conditional distribution of the data given \((A, L)\) depends on sRD, but the joint density of \((A, L)\) shares no parameters with \(f(Y \mid L, A; \text{sRD}, p_0)\). The conditionality principle says to condition on S-ancillary statistics, here \(\{A_i, L_i; i = 1, \ldots, n\}\). The same holds if the risk ratio, rather than the risk difference, is known to be constant. An exact ancillary is an S-ancillary statistic with known marginal distribution (here, \(\alpha\) and \(\rho\) known).
If the stratum-specific risk differences \(\text{sRD}_l\) vary over \(L\), the population causal risk difference is the standardized risk difference
\[ \operatorname{RD}_{std} = \sum_l \mathopen{}\left[\Pr[Y = 1 \mid L = l, A = 1; \upsilon] - \Pr[Y = 1 \mid L = l, A = 0; \upsilon]\right]\mathclose{} f(l; \rho), \]
with \(\upsilon = \{\text{sRD}_l, p_{0l}; l = 0, 1\}\). In an unconditionally randomized experiment \(A \perp\!\!\!\perp L\), so \(\operatorname{RD}_{std}\) equals the associational risk difference. Because \(\operatorname{RD}_{std}\) depends on \(\rho\), \(\{A_i, L_i\}\) is no longer exactly ancillary, and no exact ancillary exists.
Let \(\widetilde S = \widehat{\operatorname{OR}}_{AL} - \operatorname{OR}_{AL}\), the sample minus the super-population \(A\)-\(L\) odds ratio, and \(\widehat S = \widetilde S / \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\widetilde S\right)\mathclose{}\), which is approximately standard normal in large samples. When \(\operatorname{OR}_{AL}\) is known (e.g., \(\operatorname{OR}_{AL} = 1\) in a randomized experiment), \(\widehat S\) is an approximate (large-sample) ancillary: it can be computed from the data, has approximately known distribution, is governed by the \(f(A \mid L; \alpha)\) factor of the likelihood only, and conditional on it the adjusted estimate of \(\operatorname{RD}_{std}\) is unbiased while the unadjusted one is biased.
Under a continuity principle (Buehler 1982), inferences should not change discontinuously after an arbitrarily small known change in the data-generating distribution. Accepting both conditionality and continuity implies conditioning on approximate ancillaries too: the extended conditionality principle.
The adjusted estimator \(\widehat{\operatorname{RD}}_{MLE}\) replaces the population proportions in \(\operatorname{RD}_{std}\) by sample proportions; the unadjusted estimator is \(\widehat{\operatorname{RD}}_{UN} = \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} - \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 0\right]\mathclose{}\). Unconditionally, both are asymptotically normal and unbiased for \(\operatorname{RD}_{std}\).
Robins and Morgenstern (1987) show that
- \(\widehat{\operatorname{RD}}_{MLE}\) has the same asymptotic distribution conditionally on \(\widehat S\) as unconditionally;
- \(\operatorname{aVar}\mathopen{}\left(\widehat{\operatorname{RD}}_{MLE}\right)\mathclose{} = \operatorname{aVar}\mathopen{}\left(\widehat{\operatorname{RD}}_{UN}\right)\mathclose{} - \mathopen{}\left[\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{}\right]\mathclose{}^2\);
- the conditional asymptotic bias of \(\widehat{\operatorname{RD}}_{UN}\) given \(\widehat S\) is \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} \, \widehat S\).
So \(\widehat{\operatorname{RD}}_{UN}\) is unconditionally inefficient if and only if it is conditionally biased (both iff \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} \neq 0\)). Moreover, \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} = 0\) iff \(L \perp\!\!\!\perp Y \mid A\), so when a measured risk factor for \(Y\) is available, \(\widehat{\operatorname{RD}}_{MLE}\) is preferred. The corrected estimator \(\widehat{\operatorname{RD}}_{UN} - \operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{}\widehat S\) has the same asymptotic distribution as \(\widehat{\operatorname{RD}}_{MLE}\) given \(\widehat S\).
With a constant sRD over a dichotomous \(L\), the estimated variance of the MLE of sRD is \(\widehat V_0 \widehat V_1 / (\widehat V_0 + \widehat V_1)\), where \(\widehat V_l\) is the estimated variance of the stratum-specific estimate \(\widehat{\operatorname{RD}}_l\).
Two choices for \(\widehat V_1\) (smokers, Table 10.2) differ only in the denominators:
\[ \begin{aligned} \widehat V_1^{obs} &= \frac{(4/80)(76/80)}{80} + \frac{(2/40)(38/40)}{40} = \frac{0.0475}{80} + \frac{0.0475}{40} \approx 1.78 \times 10^{-3}, \\ \widehat V_1^{exp} &= \frac{(4/80)(76/80)}{60} + \frac{(2/40)(38/40)}{60} = \frac{0.0475}{60} + \frac{0.0475}{60} \approx 1.58 \times 10^{-3}. \end{aligned} \]
\(\widehat V_1^{obs}\) uses the observed numbers per arm (80 and 40; observed information), \(\widehat V_1^{exp}\) the expected number (60 per arm under \(A \perp\!\!\!\perp L\); expected information). Nearly all researchers would choose \(\widehat V_1^{obs}\), the choice used in Example 9. Results of Efron and Hinkley (1978) and Robins and Morgenstern (1987) imply that this choice amounts to conditioning on the approximate ancillary \(\widehat S\), i.e., following the extended conditionality principle. (The conditional and unconditional variances of the MLE differ by order \(n^{-3/2}\); the observed-information estimator is closer than that to the conditional variance, the expected-information estimator closer than that to the unconditional variance.)
5 10.5 The Curse of Dimensionality (pp. 148-150)
The arguments of Section 10.4 rely on asymptotic theory in which the number of strata of \(L\) is small relative to the sample size.
5.1 The Adjusted Interval Becomes Uninformative
- The investigators now accept the conditionality principle, so they consider the unadjusted \(-0.15\) conditionally biased.
- But with \(2^{100}\) strata, far more than the 240 individuals, at most a few strata contain both a treated and an untreated individual.
This is the reverse of the one-covariate case. The adjusted estimator is guaranteed to be more efficient only when the ratio of individuals to unknown parameters is large.
A common rule of thumb is at least 10 individuals per unknown parameter, though the minimum depends on the data.
5.2 Conditionality Should Not Be a Principle
In Example 10, the unadjusted estimator is still conditionally biased given all 100 covariates, but the estimator that conditions on them is now uninformative (Example 11).
Conditionality is compelling in simple examples, but it cannot be carried through in high-dimensional models, so it should not be raised to a principle (Robins and Ritov 1997). The same applies to observational studies.
Solving the curse of dimensionality is hard and an active research area; Chapter 18 reviews it. Chapters 11 through 17 supply the background on models for causal inference.
With many strata of \(L\), informative inference conditional on the exact ancillary \(\{A_i, L_i\}\) is impossible whatever the estimator. Unconditional inference in a marginally randomized experiment is different: \(\widehat{\operatorname{RD}}_{UN}\) stays unbiased, with Wald interval width proportional to \(1/n^{1/2}\), regardless of the dimension of \(L\). It uses prior information the MLE does not: it is unbiased for the common risk difference only if \(A \perp\!\!\!\perp L\) is known to hold in the super-population. Even unconditionally, the MLE’s intervals remain uninformative. Can data on \(L\) give an unconditionally unbiased estimator more efficient than \(\widehat{\operatorname{RD}}_{UN}\)? Chapter 18 shows that this is sometimes possible.
Fine Point 6.3 showed that, under faithfulness and with infinite data, the causal structure can sometimes be learned. Suppose variables \(Z\), \(A\), \(Y\) occur in that order, the empirical \(Z\)-\(Y\) odds ratio is 1 at every level of \(A\), and all other odds ratios are far from 1. If \(Z \perp\!\!\!\perp Y \mid A\) held in the super-population, faithfulness would leave only \(Z \to A \to Y\) (possibly with a common cause of \(Z\) and \(A\)), so \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) would be the average causal effect.
But Robins et al. (2003) showed that faithfulness-based inferences are non-uniform: however large \(n\) is, some faithful distributions have conditional \(Z\)-\(Y\) odds ratios near enough to 1 that sample values of exactly 1 are to be expected, while \(A\) has no average causal effect on \(Y\). An honest 95% frequentist interval for that effect must therefore always include 0, however large the empirical risk difference is relative to its standard error (for instance, 0.2 at 30 standard errors).
A Bayesian who believes in faithfulness may obtain a narrow credible interval that excludes 0, but such inferences are very sensitive to the prior; Hernán and Robins argue that the prior probability of any observational-study diagram without a common cause of \(A\) and \(Y\) should be essentially zero. In finite samples, data alone cannot distinguish \(Z \perp\!\!\!\perp Y \mid A\) from near-independence, so causal discovery from observational data relies heavily on subject-matter knowledge.
6 Summary
- Identification (Part I) assumes effectively infinite data; estimation deals with random variability in finite samples.
- Estimands, estimators, and estimates; consistent estimators; Wald confidence intervals and their calibration, validity, and (frequentist) interpretation.
- In the book, an estimator is unbiased if it can center a valid Wald interval (formally, uniformly asymptotically normal and unbiased, UANU); confidence intervals ignore systematic bias.
- Standard binomial intervals implicitly assume random sampling from a super-population, usually a fiction, so they are what-if statements.
- In a trial with a chance imbalance on a measured factor \(L\) that is not independent of \(Y\) given \(A\), and with a risk difference that is constant across strata of \(L\), the adjusted estimator is conditionally unbiased and unconditionally more efficient: the conditionality principle favors adjustment.
- With high-dimensional \(L\), conditioning makes estimates uninformative while the unadjusted estimator stays conditionally biased: the curse of dimensionality, which motivates the models of Part II.
Part I established when causal effects are identified (under exchangeability, positivity, and consistency). Part II (Chapters 11 through 18) turns to estimation with models: why models are needed (Chapter 11), IP weighting and marginal structural models (Chapter 12), standardization and the parametric g-formula (Chapter 13), g-estimation of structural nested models (Chapter 14), outcome regression and propensity scores (Chapter 15), instrumental variable estimation (Chapter 16), causal survival analysis (Chapter 17), and variable selection and high-dimensional data (Chapter 18).