Chapter 2: Randomized Experiments

Published

Last modified: 2026-10-09 10:17:06 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Does your looking up at the sky make other pedestrians look up too? The question has the parts of any causal question: an action (your looking up), an outcome (others looking up), and a population (say, residents of Madrid in 2019). A natural design: flip a coin each time someone approaches, look up on heads and straight ahead on tails, repeat a few thousand times, and compare the proportions of pedestrians who look up within 10 seconds. Suppose 55% looked up when you did and 1% when you did not.

That design is a randomized experiment: an experiment because the investigator carries out the action, and randomized because a random device (the coin) decides whom to act on. A nonrandomized experiment, e.g. looking up only when a man approaches, would be much less convincing: critics could say that men and women differ in how often they look up, so the comparison is between “noncomparable” groups. This chapter explains why randomization supports convincing causal inferences.

This chapter is based on Hernán and Robins (2020, chap. 2, pp. 13-25).

1 2.1 Randomization (pp. 13-16)


In a real study we never know both of Zeus’s potential outcomes, only his observed outcome under the treatment he received. Table 2.1 adds the counterfactual columns to the observed data of Table 1.2; “?” marks the counterfactual outcome that is missing.

Table 2.1: Observed data with counterfactual outcomes, Zeus’s family (Hernán and Robins 2020, 14)

Name \(A\) \(Y\) \(Y^0\) \(Y^1\)
Rheia 0 0 0 ?
Kronos 0 1 1 ?
Demeter 0 0 0 ?
Hades 0 0 0 ?
Hestia 1 0 ? 0
Poseidon 1 0 ? 0
Hera 1 0 ? 0
Zeus 1 1 ? 1
Artemis 0 1 1 ?
Apollo 0 1 1 ?
Leto 0 0 0 ?
Ares 1 1 ? 1
Athena 1 1 ? 1
Hephaestus 1 1 ? 1
Aphrodite 1 1 ? 1
Polyphemus 1 1 ? 1
Persephone 1 1 ? 1
Hermes 1 0 ? 0
Hebe 1 0 ? 0
Dionysus 1 0 ? 0

With half the counterfactual outcomes missing, Table 2.1 by itself only supports association measures. Randomized experiments also have missing counterfactuals, but randomization ensures the values are missing by chance, so effect measures can be consistently estimated despite the missing data.

Neyman (1923) applied counterfactual theory to the estimation of causal effects via randomized experiments (Hernán and Robins 2020, 13).

1.1 An Ideal Randomized Experiment

Take the near-infinite population drawn as a diamond in Figure 1.1. For each individual, flip a coin with probability of heads below 50%:

  • tails: white group, given the treatment (\(A = 1\));
  • heads: grey group, given placebo (\(A = 0\)).

Five days later: \(\Pr[Y = 1 \mid A = 1] = 0.3\) and \(\Pr[Y = 1 \mid A = 0] = 0.6\), so the associational risk ratio is \(0.3/0.6 = 0.5\) and the associational risk difference is \(0.3 - 0.6 = -0.3\).

“Ideal” means: no loss to follow-up, full adherence to the assigned treatment for the duration of the study, a single version of treatment, and double-blind assignment (Chapter 9). Ideal experiments are unrealistic, but useful for introducing key concepts; later chapters treat more realistic ones (Hernán and Robins 2020, 14).

1.2 Exchangeability

Suppose the assistants mistakenly treated the grey group instead of the white group. Our conclusions would not change: the risk in the treated (now grey) would still be expected to be 0.3 and in the untreated 0.6. When group membership is randomized, which group got the treatment is irrelevant. The groups are exchangeable.

Definition 1 (Exchangeability) The treated and the untreated are exchangeable if the risk under each potential treatment value \(a\) is the same in both groups: \[ \Pr[Y^a = 1 \mid A = 1] = \Pr[Y^a = 1 \mid A = 0] \quad \text{for } a = 0 \text{ and } a = 1. \] Equivalently, the counterfactual outcome is independent of the actual treatment: \(Y^a \perp\!\!\!\perp A\) for all \(a\).

If a conditional risk is the same in every treatment group, it equals the marginal risk: \[ \Pr[Y^a = 1 \mid A = 1] = \Pr[Y^a = 1 \mid A = 0] = \Pr[Y^a = 1]. \] So the actual treatment \(A\) does not predict the counterfactual outcome \(Y^a\). Randomization is valued because it is expected to produce exchangeability.

When the treated and the untreated are exchangeable, treatment is sometimes called exogenous, so exogeneity is used as a synonym for exchangeability (Hernán and Robins 2020, 14).

1.3 In Ideal Randomized Experiments, Association Is Causation

The risk under treatment in the white (treated) group is not counterfactual at all, because that group was treated. Hence \[\begin{align} \Pr[Y^{a=1} = 1] &= \Pr[Y^{a=1} = 1 \mid A = 1] && \text{(exchangeability)} \\ &= \Pr[Y = 1 \mid A = 1] && \text{(consistency: } Y = Y^{a=1} \text{ when } A = 1\text{)} \\ &= 0.3 . \end{align}\] The same steps with \(a = 0\) give \(\Pr[Y^{a=0} = 1] = \Pr[Y = 1 \mid A = 0] = 0.6\). The causal risk ratio is \(0.3/0.6 = 0.5\) and the causal risk difference is \(0.3 - 0.6 = -0.3\), the same as the associational measures.

Another way to see why randomization gives \(Y^a \perp\!\!\!\perp A\): like one’s genetic make-up, \(Y^a\) can be thought of as a fixed characteristic that exists before treatment is randomly assigned, since it encodes what one’s outcome would be under \(a\) whatever treatment is later received. A randomized \(A\) is independent of both one’s genes and \(Y^a\). The difference is that \(Y^a\) can only be learned after treatment, and only if \(A = a\) (Hernán and Robins 2020, 15).

NoteTechnical Point 2.1: Full Exchangeability and Mean Exchangeability

Let \(\mathcal{A} = \{a, a', a'', \ldots\}\) be the set of treatment values in the population and \(Y^{\mathcal{A}} = \{Y^a, Y^{a'}, Y^{a''}, \ldots\}\) the set of all counterfactual outcomes. Randomization makes \(Y^{\mathcal{A}} \perp\!\!\!\perp A\), full exchangeability (for a dichotomous treatment, \(\{Y^{a=1}, Y^{a=0}\} \perp\!\!\!\perp A\)). Full exchangeability implies, but is not implied by, \(Y^a \perp\!\!\!\perp A\) for each \(a\).

Mean exchangeability is \(\operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) for all \(a, a'\). For a dichotomous outcome it is the same as exchangeability; for a continuous outcome, exchangeability implies mean exchangeability but not conversely (e.g., the variance of \(Y^a\) may depend on treatment).

Mean exchangeability is all that is needed for \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = a\right]\mathclose{}\): \[\begin{align} \operatorname{E}\mathopen{}\left[Y \mid A = a\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^a \mid A = a\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} && \text{(mean exchangeability)}. \end{align}\] (Hernán and Robins 2020, 15)

1.4 Caution: \(Y^a \perp\!\!\!\perp A\) Is Not \(Y \perp\!\!\!\perp A\)

  • \(Y^a \perp\!\!\!\perp A\) (exchangeability): the treated and the untreated would have had the same risk had they received the same treatment level.
  • \(Y \perp\!\!\!\perp A\): no association between observed treatment and observed outcome.

In a randomized experiment with exchangeability and a causal effect, \(Y \perp\!\!\!\perp A\) fails: treatment is associated with the observed outcome.

Why: if \(Y^{a=1} \neq Y^{a=0}\) for some individuals, then \(Y = Y^A\) is the counterfactual evaluated at the observed \(A\), which depends on \(A\), so \(Y\) is not independent of \(A\) (Hernán and Robins 2020, 15).

1.5 Does Exchangeability Hold in Table 2.1?

Check \(Y^a \perp\!\!\!\perp A\) for \(a = 0\), pretending we can see Table 1.1. Among the 13 treated, \(Y^{a=0} = 1\) for Poseidon, Ares, Athena, Persephone, Hermes, Hebe, and Dionysus; among the 7 untreated, for Kronos, Artemis, and Apollo: \[ \Pr[Y^{a=0} = 1 \mid A = 1] = \frac{7}{13} \approx 0.54 \;>\; \frac{3}{7} \approx 0.43 = \Pr[Y^{a=0} = 1 \mid A = 0]. \] The treated have a worse prognosis than the untreated: exchangeability does not hold (it also fails for \(a = 1\): \(7/13\) in the treated vs. \(3/7\) in the untreated, the latter from Rheia, Artemis, and Leto).

In the real world we only have Table 2.1, so we usually cannot check exchangeability. Even if we knew it failed, we could not conclude that the study was not randomized, for two reasons:

  1. 20 people are too few; sampling variability could explain almost anything (Chapter 10). Until then, each individual stands for 1 billion identical ones.
  2. A study can be randomized and still lack exchangeability in infinite samples, if investigators used more than one coin. Section 2.2 describes such experiments (Hernán and Robins 2020, 16).

Our discussion concerns population (average) causal effects; individual effects cannot generally be identified (but see Fine Point 2.1).

NoteFine Point 2.1: Crossover Experiments

Suppose Zeus calls a lightning strike (\(a = 1\)) yesterday and his blood pressure rises; today he refrains (\(a = 0\)) and it does not. Having seen “both” outcomes, can we conclude that lightning-bolt use affects his blood pressure? Only under strong assumptions.

In a crossover experiment, individual \(i\) receives treatment \(A_{it}\) in periods \(t = 0, 1\). Let \(Y_{i1}^{a_0, a_1}\) be the outcome at \(t = 1\) under \(a_0\) at \(t = 0\) and \(a_1\) at \(t = 1\), and \(Y_{i0}^{a_0}\) the outcome at \(t = 0\). The individual effect \(Y_{it}^{a_t = 1} - Y_{it}^{a_t = 0}\) is identified if:

    1. no carryover effect: \(Y_{i1}^{a_0, a_1} = Y_{i1}^{a_1}\);
    1. the individual effect does not depend on time: \(Y_{it}^{a_t = 1} - Y_{it}^{a_t = 0} = \alpha_i\) for \(t = 0, 1\);
    1. the outcome under no treatment does not depend on time: \(Y_{it}^{a_t = 0} = \beta_i\) for \(t = 0, 1\).

If \(A_{i1} = 1\) and \(A_{i0} = 0\), then \[\begin{align} Y_{i1} - Y_{i0} &= Y_{i1}^{a_1 = 1} - Y_{i0}^{a_0 = 0} && \text{(consistency and (i))} \\ &= \mathopen{}\left(Y_{i1}^{a_1 = 1} - Y_{i1}^{a_1 = 0}\right)\mathclose{} + \mathopen{}\left(Y_{i1}^{a_1 = 0} - Y_{i0}^{a_0 = 0}\right)\mathclose{} && \text{(add and subtract } Y_{i1}^{a_1 = 0}\text{)} \\ &= \alpha_i + (\beta_i - \beta_i) && \text{(by (ii) and (iii))} \\ &= \alpha_i . \end{align}\] Symmetrically, if \(A_{i1} = 0\) and \(A_{i0} = 1\), then \(Y_{i0} - Y_{i1} = \alpha_i\).

Condition (i) requires an outcome with abrupt onset that fully resolves by the next period, so crossover designs cannot study an irreversible action (heart transplant) on an irreversible outcome (death) (Hernán and Robins 2020, 16).

2 2.2 Conditional Randomization (pp. 17-18)


Table 2.2 adds a prognostic factor \(L\) (1 if in critical condition, 0 otherwise), measured before treatment was assigned.

Table 2.2: Heart transplant study with prognostic factor \(L\) (Hernán and Robins 2020, 17)

Name \(L\) \(A\) \(Y\)
Rheia 0 0 0
Kronos 0 0 1
Demeter 0 0 0
Hades 0 0 0
Hestia 0 1 0
Poseidon 0 1 0
Hera 0 1 0
Zeus 0 1 1
Artemis 1 0 1
Apollo 1 0 1
Leto 1 0 0
Ares 1 1 1
Athena 1 1 1
Hephaestus 1 1 1
Aphrodite 1 1 1
Polyphemus 1 1 1
Persephone 1 1 1
Hermes 1 1 0
Hebe 1 1 0
Dionysus 1 1 0

Two mutually exclusive designs could have produced these data:

  • Design 1: randomly select 65% of everyone for a transplant (one loaded coin, \(\Pr[\text{tails}] = 0.65\)); this would explain 13 of 20 treated.
  • Design 2: transplant a random 75% of those in critical condition (\(L = 1\)) and 50% of those in noncritical condition (\(L = 0\)) (two coins); this would explain 9 of 12 and 4 of 8 treated.

Design 1 is a marginally randomized experiment (one unconditional randomization probability for all). Design 2 is a conditionally randomized experiment (randomization probabilities depend on \(L\)).

2.1 Which Design Produced Table 2.2?

A marginally randomized experiment is expected to give exchangeability, \(Y^a \perp\!\!\!\perp A\). A conditionally randomized one generally does not, because each treatment group may have a different share of individuals with a bad prognosis.

In Table 2.2:

  • critical among the treated: \(9/13 \approx 69\%\);
  • critical among the untreated: \(3/7 \approx 43\%\).

The treated would have had a higher risk than the untreated had they remained untreated: \(A\) predicts \(Y^{a=0}\), so \(Y^a \perp\!\!\!\perp A\) fails. Since the study was randomized, it must have been randomized conditional on \(L\) (design 2).

2.2 Conditional Exchangeability

A conditionally randomized experiment is two marginally randomized experiments, one in \(L = 1\) and one in \(L = 0\). Within each stratum, the treated and the untreated are exchangeable: \[ \Pr[Y^a = 1 \mid A = 1, L = 1] = \Pr[Y^a = 1 \mid A = 0, L = 1], \quad \text{i.e. } Y^a \perp\!\!\!\perp A \mid L = 1, \] and likewise \(Y^a \perp\!\!\!\perp A \mid L = 0\).

Definition 2 (Conditional exchangeability) \(Y^a \perp\!\!\!\perp A \mid L\) for all \(a\): within every level \(l\) of \(L\), the counterfactual outcome is independent of the actual treatment, \(Y^a \perp\!\!\!\perp A \mid L = l\).

  • Marginal randomization (design 1) produces both marginal and conditional exchangeability.
  • Conditional randomization (design 2) produces only conditional exchangeability.

Missing-data view (Hernán and Robins 2020, 18): if \(A = 1\), \(Y^{a=0}\) is missing; if \(A = 0\), \(Y^{a=1}\) is missing.

  • Missing completely at random (MCAR): \(\Pr[A = a \mid L, Y^{a=1}, Y^{a=0}] = \Pr[A = a]\), which holds in a marginally randomized experiment.
  • Missing at random (MAR): the probability of \(A = a\) given the full data \((L, Y^{a=1}, Y^{a=0})\) depends only on the data that would be observed, \((L, Y^a)\), if \(A = a\). MAR implies \(\Pr[A = a \mid L, Y^{a=1}, Y^{a=0}] = \Pr[A = a \mid L]\), which holds in a conditionally randomized experiment: \(\Pr[A = 1 \mid L, Y^{a=1}, Y^{a=0}]\) cannot depend on \(Y^{a=0}\), and \(\Pr[A = 0 \mid \cdot] = 1 - \Pr[A = 1 \mid \cdot]\) cannot depend on \(Y^{a=1}\).

The terms MCAR, MAR, and MNAR (missing not at random) were introduced by Rubin (1976) and Marini, Olsen, and Rubin (1980).

2.3 Two Targets

If Table 2.2 had come from a marginally randomized experiment, the causal risk ratio would simply be the associational one: \[ \frac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = \frac{7/13}{3/7} = \frac{49}{39} \approx 1.26 . \] Under conditional randomization we have two options:

  1. Stratification: within each stratum association is causation, so the stratum-specific causal risk ratio \(\Pr[Y^{a=1} = 1 \mid L = l] / \Pr[Y^{a=0} = 1 \mid L = l]\) equals the stratum-specific associational risk ratio \(\Pr[Y = 1 \mid L = l, A = 1] / \Pr[Y = 1 \mid L = l, A = 0]\). If these differ across \(L\), there is effect modification by \(L\) (treatment effect heterogeneity), the topic of Chapter 4.
  2. The average causal effect in the entire population, \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\), e.g. when \(L\) will not be available for future patients, so treatment decisions cannot depend on it. Sections 2.3 and 2.4 show two ways to compute it.
NoteFine Point 2.2: Risk Periods

A risk is the proportion developing the outcome during a specified period (e.g., 5-day mortality); the book often states the period once and then omits it.

Why the period matters: in a randomized trial of antibiotics for elderly people with plague, one investigator reports a causal risk ratio of 0.05 and another a ratio of 1. Both are right: the first used 1-year risks, the second 100-year risks, which are 1 under either treatment. A treatment “having a causal effect on mortality” means death is delayed, not prevented (Hernán and Robins 2020, 19).

3 2.3 Standardization (pp. 19-20)


In our conditionally randomized study, hearts were assigned with probability 50% to the 8 individuals with \(L = 0\) and 75% to the 12 with \(L = 1\).

3.1 Stratum-Specific Risks

From Table 2.2:

\(A = 1\) \(A = 0\)
\(L = 0\) (8 individuals) \(\Pr[Y = 1 \mid L = 0, A = 1] = 1/4\) \(\Pr[Y = 1 \mid L = 0, A = 0] = 1/4\)
\(L = 1\) (12 individuals) \(\Pr[Y = 1 \mid L = 1, A = 1] = 6/9 = 2/3\) \(\Pr[Y = 1 \mid L = 1, A = 0] = 2/3\)

Because of conditional exchangeability, each observed stratum risk equals the corresponding counterfactual risk, e.g. \(\Pr[Y = 1 \mid L = 0, A = 1] = \Pr[Y^{a=1} = 1 \mid L = 0]\).

3.2 Weighting by Stratum Size

The risk if all 20 had been treated is a weighted average of \(1/4\) (in \(L = 0\)) and \(2/3\) (in \(L = 1\)), with weights \(\Pr[L = 0] = 8/20 = 0.4\) and \(\Pr[L = 1] = 12/20 = 0.6\): \[\begin{align} \Pr[Y^{a=1} = 1] &= \tfrac{1}{4} \times 0.4 + \tfrac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5, \\ \Pr[Y^{a=0} = 1] &= \tfrac{1}{4} \times 0.4 + \tfrac{2}{3} \times 0.6 = 0.1 + 0.4 = 0.5 . \end{align}\] The causal risk ratio is \(0.5/0.5 = 1\).

3.3 The General Formula

\[\begin{align} \Pr[Y^a = 1] &= \sum_l \Pr[Y^a = 1 \mid L = l] \Pr[L = l] && \text{(law of total probability)} \\ &= \sum_l \Pr[Y = 1 \mid L = l, A = a] \Pr[L = l] && \text{(conditional exchangeability, consistency)} \end{align}\] The left side is an unobserved counterfactual risk; the right side uses only observed quantities (\(L\), \(A\), \(Y\)).

Definition 3 (Identification) A counterfactual quantity is identified (identifiable) if it can be expressed as a function of the distribution of the observed data; otherwise it is unidentified.

Definition 4 (Standardization) The standardized mean of \(Y\) for treatment level \(a\), using the population as the standard, is \[ \sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \times \Pr[L = l]. \] Under conditional exchangeability it equals the counterfactual mean \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\), i.e. the mean that would have been observed had everyone in the population received \(a\).

So the causal risk ratio can be computed by standardization as \[ \frac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]} = \frac{\sum_l \Pr[Y = 1 \mid L = l, A = 1] \Pr[L = l]}{\sum_l \Pr[Y = 1 \mid L = l, A = 0] \Pr[L = l]} . \]

4 2.4 Inverse Probability Weighting (pp. 20-25)


The same causal risk ratio can be computed by inverse probability (IP) weighting.

4.1 The Tree (Figure 2.1)

Display Table 2.2 as a tree, from left to right:

  • first branching on \(L\): 8 with \(L = 0\) (\(\Pr[L = 0] = 0.4\)), 12 with \(L = 1\) (\(\Pr[L = 1] = 0.6\));
  • then on \(A\): in \(L = 0\), 4 untreated and 4 treated (\(\Pr[A = 0 \mid L = 0] = \Pr[A = 1 \mid L = 0] = 0.5\)); in \(L = 1\), 3 untreated and 9 treated (\(\Pr[A = 0 \mid L = 1] = 0.25\), \(\Pr[A = 1 \mid L = 1] = 0.75\));
  • then on \(Y\): e.g. in \((L = 0, A = 0)\), 3 survived and 1 died, so \(\Pr[Y = 1 \mid L = 0, A = 0] = 1/4\).

Figure 2.1 is a fully randomized causally interpreted structured tree graph, or FRCISTG (Robins 1986, 1987), representation of a conditionally randomized experiment. The book asks whether this wins “the prize for the worst acronym ever” (Hernán and Robins 2020, 20).

4.2 Everybody Untreated

  • \(L = 0\): 1 of the 4 untreated died. Had all 8 been untreated (twice as many), 2 would have died.
  • \(L = 1\): 2 of the 3 untreated died. Had all 12 been untreated (\(12 = 3 \times 4\)), \(2 \times 4 = 8\) would have died.

So \(\Pr[Y^{a=0} = 1] = (2 + 8)/(8 + 12) = 10/20 = 0.5\). This relies on exchangeability within each stratum: the treated, had they been untreated, would have had the same risk as the untreated in their stratum.

4.3 Everybody Treated

  • \(L = 0\): 1 of the 4 treated died; scaled to 8: 2 deaths.
  • \(L = 1\): 6 of the 9 treated died; scaled to 12 (a factor \(12/9 = 4/3\)): \(6 \times 4/3 = 8\) deaths.

So \(\Pr[Y^{a=1} = 1] = (2 + 8)/20 = 0.5\), and the causal risk ratio is \(0.5/0.5 = 1\), as with standardization.

4.4 The Pseudo-Population

The two simulated trees (Figure 2.2: everybody untreated; everybody treated) pooled together form a pseudo-population twice the size of the original, in which every individual appears once treated and once untreated (Figure 2.3).

In the pseudo-population, \(L\) is independent of \(A\), so under \(Y^a \perp\!\!\!\perp A \mid L\) in the original population the treated and untreated are unconditionally exchangeable. Hence the associational risk ratio in the pseudo-population equals the causal risk ratio in both the pseudo-population and the original population.

4.5 IP Weights

Each individual is weighted by the inverse of the probability of receiving the treatment they actually received:

Group \(\Pr[A = a \mid L]\) Weight Count Pseudo-population size
\(L = 0\), \(A = 0\) 0.5 \(1/0.5 = 2\) 4 8
\(L = 0\), \(A = 1\) 0.5 \(1/0.5 = 2\) 4 8
\(L = 1\), \(A = 0\) 0.25 \(1/0.25 = 4\) 3 12
\(L = 1\), \(A = 1\) 0.75 \(1/0.75 \approx 1.33\) 9 12

Definition 5 (IP weights) \[ W^A = \frac{1}{f[A \mid L]}, \] where \(f[a \mid l] = \Pr[A = a \mid L = l]\) for discrete \(A\) and \(L\). A treated individual with \(L = l\) gets \(1/\Pr[A = 1 \mid L = l]\); an untreated one with \(L = l'\) gets \(1/\Pr[A = 0 \mid L = l']\).

IP weighted estimators were proposed by Horvitz and Thompson (1952) for surveys in which subjects are sampled with unequal probabilities (see Technical Point 12.1) (Hernán and Robins 2020, 22).

NoteTechnical Point 2.2: Formal Definition of IP Weights

The weight’s denominator is the conditional density of \(A\) given \(L\) evaluated at the individual’s own values, i.e. at the random arguments \(A\) and \(L\): \(f[A \mid L]\). This notation is needed because \(\Pr[A = A \mid L = L]\) is tautologically 1. In a conditionally randomized experiment, \(f[a \mid l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\).

The mean outcome in the pseudo-population equals the IP weighted mean in the population. With \(I(A = a) = 1\) if \(A = a\) and 0 otherwise: \[\begin{align} \operatorname{E}_{ps}\mathopen{}\left[Y \mid A = a\right]\mathclose{} &= \frac{\operatorname{E}_{ps}\mathopen{}\left[Y\, I(A = a)\right]\mathclose{}}{\operatorname{E}_{ps}\mathopen{}\left[I(A = a)\right]\mathclose{}} && \text{(laws of probability)} \\ &= \frac{\operatorname{E}\mathopen{}\left[W^A Y\, I(A = a)\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[I(A = a)\, W^A\right]\mathclose{}} && \text{(definition of } \text{E}_{ps}\text{)} \\ &= \frac{\operatorname{E}\mathopen{}\left[Y\, I(A = a) / \Pr[A = a \mid L]\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[I(A = a) / \Pr[A = a \mid L]\right]\mathclose{}} && \text{(} I(A = a)/f[A \mid L] = I(A = a)/f[a \mid L]\text{)} \\ &= \operatorname{E}\mathopen{}\left[\frac{Y\, I(A = a)}{\Pr[A = a \mid L]}\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[I(A = a)/\Pr[A = a \mid L] \mid L\right]\mathclose{} = 1\text{)}. \end{align}\] (Hernán and Robins 2020, 23)

4.6 Standardization and IP Weighting Are Equivalent

Both gave a causal risk ratio of 1 here, and this is no coincidence (Technical Point 2.3). Both build a counterfactual tree in which everyone receives \(a\), from different pieces:

  • IP weighting uses \(\Pr[A = a \mid L = l]\);
  • standardization uses \(\Pr[L = l]\) and \(\Pr[Y = 1 \mid A = a, L = l]\).

Because both simulate what would have happened had \(L\) not been used to decide the probability of treatment, we say these methods adjust for \(L\). “Controlling for” \(L\) is a slight abuse of language: this analytic control is quite different from the physical control of a randomized experiment.

NoteTechnical Point 2.3: Equivalence of IP Weighting and Standardization

Let \(A\) be discrete with finitely many values and assume positivity: \(f[a \mid l] > 0\) for all \(l\) with \(\Pr[L = l] > 0\) (guaranteed in conditionally randomized experiments). Define

  • the standardized mean \(\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l]\);
  • the IP weighted mean \(\operatorname{E}\mathopen{}\left[\frac{I(A = a)\, Y}{f[A \mid L]}\right]\mathclose{}\).

They are equal (no counterfactuals needed): \[\begin{align} \operatorname{E}\mathopen{}\left[\frac{I(A = a)\, Y}{f[A \mid L]}\right]\mathclose{} &= \sum_l \frac{1}{f[a \mid l]} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} f[a \mid l] \Pr[L = l] && \text{(only } A = a \text{ contributes)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(cancel } f[a \mid l]\text{)}. \end{align}\] For continuous \(L\), replace the sum by an integral.

Under conditional exchangeability, both equal \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\). Standardized mean: \[\begin{align} \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid L = l\right]\mathclose{} \Pr[L = l] && \text{(law of total expectation)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y^a \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(conditional exchangeability, positivity)} \\ &= \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] && \text{(consistency)}. \end{align}\] IP weighted mean: \[\begin{align} \operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[A \mid L]} Y\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} Y^a\right]\mathclose{} && \text{(consistency)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} Y^a \;\middle|\; L\right]\mathclose{}\right]\mathclose{} && \text{(law of total expectation; } f[a \mid L] > 0\text{)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\frac{I(A = a)}{f[a \mid L]} \;\middle|\; L\right]\mathclose{} \operatorname{E}\mathopen{}\left[Y^a \mid L\right]\mathclose{}\right]\mathclose{} && \text{(conditional exchangeability)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y^a \mid L\right]\mathclose{}\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[I(A = a)/f[a \mid L] \mid L\right]\mathclose{} = 1\text{)} \\ &= \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} . \end{align}\]

With a continuous treatment the IP weighted mean is no longer equal to the standardized mean and is biased for \(\operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{}\) even under exchangeability: with \(f\) a density, \(\operatorname{E}\mathopen{}\left[I(A = a)/f(a \mid l) \mid L = l\right]\mathclose{} = 0\), not 1; with \(f\) a probability, \(f(a \mid L = l) = 0\) with probability 1 and positivity fails. Section 12.4 generalizes IP weighting to continuous treatments; Technical Point 3.1 shows the results fail without positivity even for discrete \(A\) (Hernán and Robins 2020, 25).

4.7 Why Not Finish the Book Here?

An ideal randomized experiment plus standardization or IP weighting gives average causal effects. But randomized experiments are often unethical, impractical, or untimely. For the heart transplant study:

  • ethics: hearts are scarce, and society prefers to give them to those most likely to benefit, not at random;
  • feasibility: double blinding is impossible, people assigned to medical treatment may not forgo a transplant, and compatible hearts may not be available;
  • timeliness: the study would take years, and decisions must be made in the meantime.

Frequently an observational study is the least bad option (Chapter 3).

5 Summary


  1. Randomization makes missing counterfactuals missing by chance and is expected to produce exchangeability, \(Y^a \perp\!\!\!\perp A\); in an ideal randomized experiment, association is causation.
  2. \(Y^a \perp\!\!\!\perp A\) (exchangeability) is different from \(Y \perp\!\!\!\perp A\) (no association).
  3. Conditional randomization (randomization probabilities depending on \(L\)) produces only conditional exchangeability, \(Y^a \perp\!\!\!\perp A \mid L\).
  4. Under conditional exchangeability, the population causal effect is identified by standardization, \(\sum_l \operatorname{E}\mathopen{}\left[Y \mid L = l, A = a\right]\mathclose{} \Pr[L = l]\), or equivalently by IP weighting with \(W^A = 1/f[A \mid L]\), which creates a pseudo-population in which \(L\) and \(A\) are independent.
  5. In Table 2.2 both methods give a causal risk ratio of 1, whereas the crude associational risk ratio is 1.26.

6 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top