So far we have treated study populations as if they were effectively infinite, so that the only obstacles to causal inference were the systematic biases of the previous chapters: confounding, selection bias, and measurement bias. Real studies are finite, and random variability is a second, qualitatively different reason why causal inferences may be wrong. This chapter, the last of Part I, introduces random variability and the problems it raises for causal inference.
Definition 1 (Identification and estimation) Identification problems are problems in which we can act as if the study population were effectively infinite.
Estimation is the problem of learning about a causal effect from a finite sample, in the presence of random variability.
Part I reduced causal inference to an identification problem: under exchangeability, positivity, and consistency, can the average causal effect be computed from the data?
Definition 2 (Super-population) A super-population is a population so large that we can regard it as infinite, from which the study individuals are a random sample.
Suppose the 20 individuals of the heart transplant study in Chapter 2 (Table 2.2) are a random sample from a super-population (Definition 2), and we want to make inferences about that super-population.
Definition 3 (Estimand, estimator, estimate)
Example 2 (Sample proportion as an estimator)
A different random sample of 20 individuals would give a different estimate.
Definition 4 (Consistent estimator) An estimator is consistent for an estimand if its estimates get arbitrarily close to the estimand as the sample size \(n\) grows.
This statistical property of an estimator is unrelated to the identifiability condition also called consistency, which links each individual’s observed outcome to the counterfactual outcome under the treatment actually received.
Example 3 (The sample proportion is a consistent estimator) The sample proportion \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\) is a consistent estimator of \(\Pr[Y = 1 \mid A = a]\): the larger \(n\), the smaller we expect \(\mathopen{}\left|\Pr[Y = 1 \mid A = a] - \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = a\right]\mathclose{}\right|\mathclose{}\) to be.
A Consistent Estimator Does Not Protect a Small Study
Even a consistent estimator (Definition 4) can produce a point estimate far from the super-population value, especially in small samples. So we should have more confidence in estimates from larger studies, and we quantify that confidence with confidence intervals.
Definition 5 (Wald confidence interval) A 95% Wald confidence interval for a parameter \(\theta\) based on an estimator \(\hat\theta\) is
\[ \hat\theta\pm 1.96 \times \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat\theta\right)\mathclose{}, \]
where \(\mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat\theta\right)\mathclose{}\) estimates the (exact or large-sample) standard error of \(\hat\theta\), and 1.96 is the upper 97.5% quantile of the standard normal distribution.
Example 4 (Wald interval for the risk among the treated) Let \(p = \Pr[Y = 1 \mid A = 1]\) and \(\hat p = \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} = 7/13\), with \(n = 13\). The standard error of a binomial proportion is \(\sqrt{p(1-p)/n}\), so
\[ \begin{aligned} \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\hat p\right)\mathclose{} &= \sqrt{\frac{\hat p (1 - \hat p)}{n}} = \sqrt{\frac{(7/13)(6/13)}{13}} \\ &= \sqrt{\frac{42/169}{13}} = \sqrt{\frac{42}{2197}} \approx \sqrt{0.0191} \approx 0.138. \end{aligned} \]
The 95% Wald interval is
\[ \begin{aligned} \hat p \pm 1.96 \times 0.138 &= 0.538 \pm 0.271 \\ &= (0.27, 0.81). \end{aligned} \]
Definition 6 (Calibration and validity of a 95% confidence interval) A 95% confidence interval is
It is valid if, for every value of the true parameter, it is calibrated or conservative (coverage of at least 95%).
Among valid intervals, we prefer the narrowest.
Definition 7 (Small-sample and large-sample validity) An interval is
A large-sample valid interval may still have coverage below 95% at infinitely many sample sizes, provided the shortfall shrinks to nothing as \(n\) grows.
Remark 1 (What the book means by a 95% confidence interval). The Wald interval for \(p\) in Example 4 is large-sample valid (Definition 7) but not small-sample valid. In the book, “95% confidence interval” means a large-sample valid interval, such as a Wald interval, unless stated otherwise.
Remark 2 (Coverage over repeated studies). Validity is a property of the interval procedure, not of any one interval: over repeated samples from the super-population, the procedure covers its target with probability at least 95%. Picture every study that we and other researchers will ever run, each using a 95% procedure that is small-sample valid. If we idealize that collection as an unending sequence, the share of those intervals that contain their targets would be at least 95%. For the book’s large-sample valid intervals, such as Wald intervals, this guarantee holds only in the limit as samples grow; at a given sample size their coverage can fall well short of 95%. In every case, nothing in the data would reveal which intervals missed.
A Confidence Interval Is Not a Probability Statement About the Estimand
A single 95% interval does not mean there is a 95% probability that the estimand lies in it. The estimand is fixed, so the probability that it lies in \((0.27, 0.81)\) is either 0 or 1.
Fine Point 10.1: Honest Confidence Intervals
For a large-sample valid interval, how large the sample must be before coverage comes close to 95% may depend on the unknown true parameter. A large-sample valid interval is uniform or honest if there is a sample size \(n\) at which it covers the true parameter at least 95% of the time whatever the true parameter is. Without uniformity, at any finite \(n\) some data-generating distributions may have coverage well below 95%. Even for honest intervals, the smallest such \(n\) is generally unknown and hard to determine, even by simulation (Robins and Ritov 1997).
From here on, the book uses “valid confidence interval” to mean a large-sample honest interval. Any small-sample valid interval is automatically honest for all \(n\) at which it is defined.
Definition 8 (Unbiased estimator (informal)) An estimator is unbiased if it can center a valid Wald interval and biased if it cannot.
Not every consistent estimator can center a valid Wald interval, even in large samples. Most users of statistics use Definition 8, and the book adopts it for now.
Confidence Intervals Ignore Systematic Bias
Confidence intervals quantify only random error, so they can overstate our confidence when systematic biases may be present.
Technical Point 10.1: Bias and Consistency in Statistical Inference
Suppose we observe \(n\) i.i.d. copies of a random vector whose distribution \(P\) lies in a model \(\mathcal{M}\).
Robins and Morgenstern (1987) argue that applied researchers call an estimator unbiased only if it can center a valid Wald interval, and show that this holds only for estimators that are uniformly asymptotically normal and unbiased (UANU): there exist sequences \(\sigma_n(P)\) such that
\[ \sup_{P \in \mathcal{M}} \mathopen{}\left|\Pr_P\mathopen{}\left[n^{1/2} \mathopen{}\left(\hat\theta_n - \theta(P)\right)\mathclose{} / \sigma_n(P) < t\right]\mathclose{} - \Phi(t)\right|\mathclose{} \to 0 \text{ as } n \to \infty \]
for every \(t \in \mathbb{R}\), where \(\Phi\) is the standard normal cumulative distribution function (Robins and Ritov 1997). All inconsistent estimators, and some consistent ones (examples in Chapter 18), are biased under this definition. In the book, “unbiased” without qualification means UANU.
Suppose the heart transplant study were a marginally randomized experiment whose 20 individuals are a random sample from a near-infinite super-population.
Proposition 1 (Identification in a randomized super-population) Suppose that in a near-infinite super-population
Then exchangeability \(Y^a \perp\!\!\!\perp A\) holds in the super-population, and \(\Pr[Y^a = 1] = \Pr[Y = 1 \mid A = a]\) for \(a = 0, 1\), so
\[ \Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1] = \Pr[Y = 1 \mid A = 1] - \Pr[Y = 1 \mid A = 0]. \]
Proof. Marginal randomization of everyone in the super-population, with the same assignment probabilities for every individual, makes the assigned treatment independent of the counterfactual outcomes, and full adherence makes the treatment received equal to the treatment assigned, so \(Y^a \perp\!\!\!\perp A\). Because each arm has positive probability, the conditional probabilities given \(A = a\) are defined, and
\[ \begin{aligned} \Pr[Y^a = 1] &= \Pr[Y^a = 1 \mid A = a] && \text{(exchangeability, } Y^a \perp\!\!\!\perp A\text{)} \\ &= \Pr[Y = 1 \mid A = a] && \text{(consistency)}. \end{aligned} \]
Subtracting the equation for \(a = 0\) from the equation for \(a = 1\) gives the displayed risk difference.
Example 5 (Testing and estimation in the heart transplant trial) Testing the causal null hypothesis \(\Pr[Y^{a=1} = 1] = \Pr[Y^{a=0} = 1]\) reduces to comparing the sample proportions \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} = 7/13\) and \(\mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 0\right]\mathclose{} = 3/7\).
Standard methods give confidence intervals for the causal effects in the super-population, estimated by
In observational studies, slightly more involved but standard procedures give confidence intervals for standardized, IP weighted, or stratified association measures.
Now suppose only the 20 study individuals, not the whole super-population, were randomized.
Randomizing the Sample Does Not Make the Sample Exchangeable
Because of sampling variability, exchangeability need not hold exactly in the sample.
Example 6 (Chance imbalance in prognosis) Classify each individual as good or bad prognosis at randomization, and call the arms exchangeable if they contain the same proportion with bad prognosis. By chance, 2 of the 13 individuals assigned to \(A = 1\) (\(\approx 15\%\)) and 3 of the 7 assigned to \(A = 0\) (\(\approx 43\%\)) might have bad prognosis. In a larger sample, the relative imbalance between arms would very probably shrink.
Remark 4 (Two targets of inference). There are then two possible targets of inference:
Fine Point 10.2: Uncertainty from Systematic Biases
The width of a Wald-type interval depends on the estimator’s standard error, so it reflects only random error. Systematic biases (confounding, selection, measurement) are another source of uncertainty and do not shrink as the sample grows. Hence the larger the study, the more a usual Wald interval understates the total uncertainty.
Some authors therefore call these intervals compatibility intervals: they show which effect sizes are highly compatible with the data under our adjustment assumptions and methods (Amrhein et al. 2019; Greenland 2019), without claiming confidence that adjustment removed all systematic bias.
Discussion sections usually treat systematic bias informally. Quantitative bias analysis instead produces intervals that combine random and systematic uncertainty (Lash et al. 2009); Bayesian alternatives also exist.
Chapter 1 named two sources of randomness: sampling variability and nondeterministic counterfactuals.
The Binomial Interval Assumes a Binomial Count
Nearly every investigator would report the binomial interval \(\hat p \pm 1.96 \sqrt{\hat p (1 - \hat p)/n}\) for \(p\) from \(\hat p = 7/13\). The interval’s variance term \(\hat p (1 - \hat p)/n\) estimates the binomial variance of \(\hat p\), which is the actual variance of \(\hat p\) when the number of individuals with the outcome is binomial. For any other count, using that variance term needs a separate justification.
When is the count binomial?
Remark 5 (Two scenarios that give a binomial count). Scenario 1 (super-population). The study population is a random sample from an essentially infinite super-population (sometimes called the source or target population), and the estimand is \(p = \Pr[Y = 1 \mid A = 1]\) in that super-population. Then, across repeated random samples of 13 treated individuals, the number who develop the outcome is binomial with success probability \(p\), and the Wald interval from Example 4 is asymptotically calibrated for \(p\).
Scenario 2 (stochastic counterfactuals, no super-population). Each of the 13 treated individuals \(i\) has a nondeterministic counterfactual probability \(p_i^{a=1}\), the observed outcome \(Y_i = Y_i^{a=1}\) occurs with probability \(p_i^{a=1}\), and \(p_i^{a=1}\) equals the same value \(p\) for all 13. Then the number of treated with the outcome is binomial with success probability \(p\), and the interval is again asymptotically calibrated for \(p\).
Proposition 2 (Heterogeneous risks shrink the variance of a count) Let the outcomes of \(n\) individuals be independent Bernoulli(\(p_i\)), with mean risk \(\bar p = \sum_i p_i / n\). Their sum (a Poisson-binomial count) has variance \(\sum_i p_i(1-p_i) = n \bar p (1 - \bar p) - \sum_i (p_i - \bar p)^2\), which is smaller than the binomial variance \(n \bar p (1 - \bar p)\) at the same mean \(\bar p\) whenever the \(p_i\) are unequal.
Proof. The variance of a sum of independent variables is the sum of their variances, and a Bernoulli(\(p_i\)) variable has variance \(p_i(1-p_i)\). Then write
\[ \begin{aligned} \sum_i p_i(1-p_i) &= \sum_i p_i - \sum_i p_i^2 \\ &= n\bar p - \mathopen{}\left(n \bar p^2 + \sum_i (p_i - \bar p)^2\right)\mathclose{} \\ &= n \bar p (1 - \bar p) - \sum_i (p_i - \bar p)^2, \end{aligned} \]
using \(\sum_i p_i^2 = n\bar p^2 + \sum_i (p_i - \bar p)^2\). The subtracted sum of squares is positive unless all \(p_i\) equal \(\bar p\). (This derivation is ours, added to explain the book’s statement.)
Remark 6 (Heterogeneous risks rule out Scenario 2). The risks \(p_i^{a=1}\) almost certainly vary between individuals (for example, with genetic make-up). If they are not constant, the natural estimand for the study population is their average, say \(\bar p\), but the number of treated with the outcome is then not binomial with success probability \(\bar p\). By Proposition 2 the binomial formula overstates the variance of the count, by the sum of squared deviations \(\sum_i (p_i - \bar p)^2\). Suppose that \(\bar p\) stays bounded away from 0 and 1 as the number of individuals grows. Whether the overstatement then matters in large samples depends on the ratio of the average squared deviation \(\frac{1}{n} \sum_i (p_i - \bar p)^2\) to \(\bar p (1 - \bar p)\). If this ratio stays bounded away from zero, the binomial Wald interval is not asymptotically calibrated. Under the normal approximation it stays too wide, so it is asymptotically conservative, and for large enough samples its coverage stays above some level greater than 95%. If instead this ratio tends to zero (for example, with half the individuals at risk \(0.5 + 1/n\) and half at \(0.5 - 1/n\), for even \(n \ge 2\)), the overstatement becomes negligible relative to the binomial variance \(n \bar p (1 - \bar p)\), and the coverage still tends to 95%.
A Binomial Interval Assumes a Super-Population
Anyone who reports a binomial interval for \(\Pr[Y = 1 \mid A = a]\) as if the count were binomial, while acknowledging between-individual variation in risk that persists as the study grows, is implicitly assuming Scenario 1. Under Scenario 1 the count is binomial whether the counterfactuals are deterministic or stochastic. If the target is instead the study population itself, and its risks vary in this persistent way (with \(\bar p\) bounded away from 0 and 1), Remark 6 shows that the interval is asymptotically conservative rather than calibrated: for large enough samples its coverage stays above some level greater than 95%.
Remark 7 (Why the super-population myth endures). Advantage of the super-population: nothing hinges on whether the world is deterministic. Yet it is usually a fiction; individuals are rarely sampled at random from any near-infinite population. It endures because
Super-Population Intervals Are What-If Statements
Because the random sampling is fictional, a 95% interval computed under Scenario 1 should be read as a what-if statement: it covers the super-population parameter had the study individuals been randomly sampled from a near-infinite super-population.
Example 7 (A randomized trial, unadjusted (Table 10.1)) A trial of treatment \(A\) on 1-year death \(Y\) randomized 240 individuals, 120 per arm.
| \(Y=1\) | \(Y=0\) | |
|---|---|---|
| \(A = 1\) | 24 | 96 |
| \(A = 0\) | 42 | 78 |
Unadjusted risk difference:
\[ \frac{24}{120} - \frac{42}{120} = 0.20 - 0.35 = -0.15. \]
Estimated variance:
\[ \begin{aligned} \frac{(24/120)(96/120)}{120} + \frac{(42/120)(78/120)}{120} &= \frac{0.20 \times 0.80 + 0.35 \times 0.65}{120} \\ &= \frac{0.16 + 0.2275}{120} = \frac{0.3875}{120} = \frac{31}{9600}. \end{aligned} \]
95% Wald interval:
\[ -0.15 \pm 1.96 \sqrt{31/9600} = -0.15 \pm 1.96 \times 0.0568 = -0.15 \pm 0.111 = (-0.26, -0.04). \]
The investigators publish: treatment reduced the risk of death by 15 percentage points.
Example 8 (The same trial stratified by smoking \(L\) (Table 10.2))
| \(L = 1\) (smokers) | \(Y=1\) | \(Y=0\) |
|---|---|---|
| \(A = 1\) | 4 | 76 |
| \(A = 0\) | 2 | 38 |
| \(L = 0\) (nonsmokers) | \(Y=1\) | \(Y=0\) |
|---|---|---|
| \(A = 1\) | 20 | 20 |
| \(A = 0\) | 40 | 40 |
The adjusted (stratified) analysis finds no effect in either stratum, although the marginal estimate suggested a benefit.
Example 9 (The adjusted Wald interval) Assume the risk difference is the same in both strata of \(L\). The adjusted estimate is the maximum likelihood estimator (MLE) of the common risk difference, which in large samples weights the two stratum-specific estimates by their inverse variances. For a dichotomous \(L\), its estimated variance is \(\widehat V_0 \widehat V_1 / (\widehat V_0 + \widehat V_1)\) (Hernán and Robins 2020, 148), where \(\widehat V_l\) is the estimated variance of the risk difference in stratum \(L = l\). Computing each \(\widehat V_l\) from the observed number of individuals in each arm gives
\[ \begin{aligned} \widehat V_0 &= \frac{(20/40)(20/40)}{40} + \frac{(40/80)(40/80)}{80} = \frac{0.25}{40} + \frac{0.25}{80} = 0.009375, \\ \widehat V_1 &= \frac{(4/80)(76/80)}{80} + \frac{(2/40)(38/40)}{40} = \frac{0.0475}{80} + \frac{0.0475}{40} = 0.00178125, \\ \mathop{\widehat{\operatorname{Var}}}\nolimits\mathopen{}\left(\widehat{\text{sRD}}\right)\mathclose{} &= \frac{\widehat V_0 \widehat V_1}{\widehat V_0 + \widehat V_1} = \frac{0.009375 \times 0.00178125}{0.009375 + 0.00178125} = \frac{1.670 \times 10^{-5}}{0.01116} \approx 0.001497, \\ 1.96 \sqrt{0.001497} &= 1.96 \times 0.0387 \approx 0.076. \end{aligned} \]
The point estimate is 0 in both strata, so the pooled estimate is 0. The inverse-variance (MLE) Wald 95% interval centered at the adjusted estimate is therefore \((-0.076, 0.076)\), matching the book (Hernán and Robins 2020, 145).
Remark 8 (Should the article be retracted?). Malfeasance is ruled out, so the imbalance is “very very bad luck”. Should the article be retracted?
Remark 9 (Conditional and unconditional properties of the two estimators). Assume a constant causal risk difference across strata of \(L\), and imagine running the trial trillions of times. Condition on the runs in which \(L\) and \(A\) are as strongly positively associated as observed. Within each level of \(L\), any pre-treatment risk factor \(U\) is positively associated with \(A\) in essentially as many of these runs as it is negatively associated (even if \(U\) and \(L\) are highly correlated). Therefore:
Either way, the Wald interval centered on the adjusted estimator is the better analysis. Investigator 2 is right; the article should be retracted.
Definition 9 (Ancillary statistic) An ancillary statistic for a parameter is a statistic whose distribution does not depend on that parameter.
When the risk difference is constant across strata of \(L\), the observed \(L\)-\(A\) association is an ancillary statistic (Definition 9) for the causal risk difference.
Definition 10 (Conditionality principle) The conditionality principle is the principle that inference on a parameter should be performed conditional on ancillary statistics (Definition 9) for that parameter.
Definition 11 (Conditionally unbiased estimator) An estimator is conditionally unbiased if it can center a valid Wald interval conditional on ancillary statistics (Definition 9).
Investigators who intuitively follow the conditionality principle (Definition 10) call an estimator unbiased only if it is conditionally unbiased, a stricter definition than Definition 8.
To the frequentist objection that the unadjusted estimate is unbiased over all runs, they reply: why should \(L\)-\(A\) associations in hypothetical other studies matter, when in my study \(L\) is a confounder?
Conditionality Breaks Down With Many Confounders
The conditional reply (that the observed \(L\)-\(A\) association in my study is what matters) is convincing in randomized and observational studies as long as the number of measured confounders is not large. With many confounders, strictly following the conditionality principle stops being wise. The quotation marks around “principle” in the section title anticipate this point.
Technical Point 10.2: A Formal Statement of the Conditionality Principle
Take one dichotomous \(L\), exchangeability \(Y^a \perp\!\!\!\perp A \mid L\), and a stratum-specific risk difference \(\text{sRD} = \Pr[Y = 1 \mid L = l, A = 1] - \Pr[Y = 1 \mid L = l, A = 0]\) known to be constant across strata and of interest. The likelihood factors as
\[ \prod_{i=1}^n f(Y_i \mid L_i, A_i; \text{sRD}, p_0) \times f(A_i \mid L_i; \alpha) \times f(L_i; \rho), \]
where \(p_{0l} = \Pr[Y = 1 \mid L = l, A = 0]\) is the risk among the untreated in stratum \(l\), \(p_0 = (p_{00}, p_{01})\) collects these two baseline risks, and \(p_0\), \(\alpha\), \(\rho\) are nuisance parameters.
The data on \((A, L)\) are S-ancillary for sRD: the conditional distribution of the data given \((A, L)\) depends on sRD, but the joint density of \((A, L)\) shares no parameters with \(f(Y \mid L, A; \text{sRD}, p_0)\). The conditionality principle says to condition on S-ancillary statistics, here \(\{A_i, L_i; i = 1, \ldots, n\}\). The same holds if the risk ratio, rather than the risk difference, is known to be constant. An exact ancillary is an S-ancillary statistic with known marginal distribution (here, \(\alpha\) and \(\rho\) known).
Technical Point 10.3: Approximate Ancillarity
If the stratum-specific risk differences \(\text{sRD}_l\) vary over \(L\), the population causal risk difference is the standardized risk difference
\[ \operatorname{RD}_{std} = \sum_l \mathopen{}\left[\Pr[Y = 1 \mid L = l, A = 1; \upsilon] - \Pr[Y = 1 \mid L = l, A = 0; \upsilon]\right]\mathclose{} f(l; \rho), \]
with \(\upsilon = \{\text{sRD}_l, p_{0l}; l = 0, 1\}\). In an unconditionally randomized experiment \(A \perp\!\!\!\perp L\), so \(\operatorname{RD}_{std}\) equals the associational risk difference. Because \(\operatorname{RD}_{std}\) depends on \(\rho\), \(\{A_i, L_i\}\) is no longer exactly ancillary, and no exact ancillary exists.
Let \(\widetilde S = \widehat{\operatorname{OR}}_{AL} - \operatorname{OR}_{AL}\), the sample minus the super-population \(A\)-\(L\) odds ratio, and \(\widehat S = \widetilde S / \mathop{\widehat{\operatorname{se}}}\nolimits\mathopen{}\left(\widetilde S\right)\mathclose{}\), which is approximately standard normal in large samples. When \(\operatorname{OR}_{AL}\) is known (e.g., \(\operatorname{OR}_{AL} = 1\) in a randomized experiment), \(\widehat S\) is an approximate (large-sample) ancillary: it can be computed from the data, has approximately known distribution, is governed by the \(f(A \mid L; \alpha)\) factor of the likelihood only, and conditional on it the adjusted estimate of \(\operatorname{RD}_{std}\) is unbiased while the unadjusted one is biased.
Under a continuity principle (Buehler 1982), inferences should not change discontinuously after an arbitrarily small known change in the data-generating distribution. Accepting both conditionality and continuity implies conditioning on approximate ancillaries too: the extended conditionality principle.
Technical Point 10.4: Adjusted Versus Unadjusted Estimators
The adjusted estimator \(\widehat{\operatorname{RD}}_{MLE}\) replaces the population proportions in \(\operatorname{RD}_{std}\) by sample proportions; the unadjusted estimator is \(\widehat{\operatorname{RD}}_{UN} = \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 1\right]\mathclose{} - \mathop{\widehat{\Pr}}\nolimits\mathopen{}\left[Y = 1 \mid A = 0\right]\mathclose{}\). Unconditionally, both are asymptotically normal and unbiased for \(\operatorname{RD}_{std}\).
Robins and Morgenstern (1987) show that
So \(\widehat{\operatorname{RD}}_{UN}\) is unconditionally inefficient if and only if it is conditionally biased (both iff \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} \neq 0\)). Moreover, \(\operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{} = 0\) iff \(L \perp\!\!\!\perp Y \mid A\), so when a measured risk factor for \(Y\) is available, \(\widehat{\operatorname{RD}}_{MLE}\) is preferred. The corrected estimator \(\widehat{\operatorname{RD}}_{UN} - \operatorname{aCov}\mathopen{}\left(\widehat S, \widehat{\operatorname{RD}}_{UN}\right)\mathclose{}\widehat S\) has the same asymptotic distribution as \(\widehat{\operatorname{RD}}_{MLE}\) given \(\widehat S\).
Technical Point 10.5: Most Researchers Follow the Extended Conditionality Principle
With a constant sRD over a dichotomous \(L\), the estimated variance of the MLE of sRD is \(\widehat V_0 \widehat V_1 / (\widehat V_0 + \widehat V_1)\), where \(\widehat V_l\) is the estimated variance of the stratum-specific estimate \(\widehat{\operatorname{RD}}_l\).
Two choices for \(\widehat V_1\) (smokers, Table 10.2) differ only in the denominators:
\[ \begin{aligned} \widehat V_1^{obs} &= \frac{(4/80)(76/80)}{80} + \frac{(2/40)(38/40)}{40} = \frac{0.0475}{80} + \frac{0.0475}{40} \approx 1.78 \times 10^{-3}, \\ \widehat V_1^{exp} &= \frac{(4/80)(76/80)}{60} + \frac{(2/40)(38/40)}{60} = \frac{0.0475}{60} + \frac{0.0475}{60} \approx 1.58 \times 10^{-3}. \end{aligned} \]
\(\widehat V_1^{obs}\) uses the observed numbers per arm (80 and 40; observed information), \(\widehat V_1^{exp}\) the expected number (60 per arm under \(A \perp\!\!\!\perp L\); expected information). Nearly all researchers would choose \(\widehat V_1^{obs}\), the choice used in Example 9. Results of Efron and Hinkley (1978) and Robins and Morgenstern (1987) imply that this choice amounts to conditioning on the approximate ancillary \(\widehat S\), i.e., following the extended conditionality principle. (The conditional and unconditional variances of the MLE differ by order \(n^{-3/2}\); the observed-information estimator is closer than that to the conditional variance, the expected-information estimator closer than that to the unconditional variance.)
The arguments of Section 10.4 rely on asymptotic theory in which the number of strata of \(L\) is small relative to the sample size.
Example 10 (A trial with 100 binary covariates) Now suppose the investigators had measured 100 binary pre-treatment variables, \(L = (L_1, \ldots, L_{100})\), with \(2^{100}\) strata: high-dimensional data.
For simplicity, assume no additive effect modification by \(L\): the super-population risk difference \(\Pr[Y = 1 \mid A = 1, L = l] - \Pr[Y = 1 \mid A = 0, L = l]\) is constant, and equal to 0, across all \(2^{100}\) strata.
Example 11 (One informative stratum) Suppose exactly one stratum contains one treated and one untreated individual, and no other stratum contains both. In that stratum the empirical risk difference can only be \(-1\), \(0\), or \(1\), depending on the two outcomes, so the stratum carries no usable information about the common risk difference. The 95% interval based on the adjusted estimator is \((-1, 1)\), the whole parameter space, so it is completely uninformative. The unadjusted interval is still \((-0.26, -0.04)\), since its width does not depend on how many covariates were measured.
This is the reverse of the one-covariate case. The adjusted estimator is guaranteed to be more efficient only when the ratio of individuals to unknown parameters is large.
Individuals per Parameter
A common rule of thumb is at least 10 individuals per unknown parameter, though the minimum depends on the data.
Definition 12 (Curse of dimensionality) The curse of dimensionality is the situation in which the measured covariates \(L\) have many strata relative to the sample size \(n\), so that few strata contain both treated and untreated individuals and estimators conditional on all of \(L\) are uninformative.
In Example 10, the unadjusted estimator is still conditionally biased given all 100 covariates, but the estimator that conditions on them is now uninformative (Example 11).
Conditionality Should Not Be a Principle
Conditionality is compelling in simple examples, but it cannot be carried through in high-dimensional models, so it should not be raised to a principle (Robins and Ritov 1997). The same applies to observational studies.
Solving the curse of dimensionality is hard and an active research area; Chapter 18 reviews it. Chapters 11 through 17 supply the background on models for causal inference.
Technical Point 10.6: Can the Curse of Dimensionality Be Reversed?
With many strata of \(L\), informative inference conditional on the exact ancillary \(\{A_i, L_i\}\) is impossible whatever the estimator. Unconditional inference in a marginally randomized experiment is different: \(\widehat{\operatorname{RD}}_{UN}\) stays unbiased, with Wald interval width proportional to \(1/n^{1/2}\), regardless of the dimension of \(L\). It uses prior information the MLE does not: it is unbiased for the common risk difference only if \(A \perp\!\!\!\perp L\) is known to hold in the super-population. Even unconditionally, the MLE’s intervals remain uninformative. Can data on \(L\) give an unconditionally unbiased estimator more efficient than \(\widehat{\operatorname{RD}}_{UN}\)? Chapter 18 shows that this is sometimes possible.
Technical Point 10.7: Random Variability and Causal Discovery
Fine Point 6.3 showed that, under faithfulness and with infinite data, the causal structure can sometimes be learned. Suppose variables \(Z\), \(A\), \(Y\) occur in that order, the empirical \(Z\)-\(Y\) odds ratio is 1 at every level of \(A\), and all other odds ratios are far from 1. If \(Z \perp\!\!\!\perp Y \mid A\) held in the super-population, faithfulness would leave only \(Z \to A \to Y\) (possibly with a common cause of \(Z\) and \(A\)), so \(\operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}\) would be the average causal effect.
But Robins et al. (2003) showed that faithfulness-based inferences are non-uniform: however large \(n\) is, some faithful distributions have conditional \(Z\)-\(Y\) odds ratios near enough to 1 that sample values of exactly 1 are to be expected, while \(A\) has no average causal effect on \(Y\). An honest 95% frequentist interval for that effect must therefore always include 0, however large the empirical risk difference is relative to its standard error (for instance, 0.2 at 30 standard errors).
A Bayesian who believes in faithfulness may obtain a narrow credible interval that excludes 0, but such inferences are very sensitive to the prior; Hernán and Robins argue that the prior probability of any observational-study diagram without a common cause of \(A\) and \(Y\) should be essentially zero. In finite samples, data alone cannot distinguish \(Z \perp\!\!\!\perp Y \mid A\) from near-independence, so causal discovery from observational data relies heavily on subject-matter knowledge.