Chapter 3: Observational Studies

Published

Last modified: 2026-10-09 09:48:03 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

In Chapter 2, randomization gave us exchangeability by design. Many scientific questions, however, are answered with observational studies, in which the investigators observe and record the relevant data but do not assign treatment. Much of what we know (evolution, plate tectonics, the fact that hot coffee can burn) comes from observation rather than experiment.

This chapter asks: under what conditions can an observational study support a valid causal inference?

This chapter is based on Hernán and Robins (2020, chap. 3, pp. 27-46).

The book’s opening example returns to the pedestrians: instead of randomizing who looks up, you watch pairs of pedestrians and record whether the second one looks up after the first does. A critic can always object that both looked up because of a common cause (a thunderclap, the first raindrops), an objection that does not apply to a randomized experiment.

1 3.1 Identifiability Conditions (pp. 27-29)


In an ideal randomized experiment, an associational risk ratio of, say, 0.7 is expected to equal the causal risk ratio, because randomization makes the treated and untreated exchangeable.

In an observational study of heart transplant where sicker patients were more likely to be transplanted, an associational risk ratio of 1.1 may be a compromise between:

  • a truly beneficial effect of transplant (pushing the ratio below 1), and
  • the higher baseline mortality of those who were transplanted (pushing it above 1).

The best explanation for an association in an observational study is not necessarily a causal effect. The common strategy is to analyze the data as if treatment had been randomly assigned conditional on measured covariates \(L\), knowing this is at best an approximation. Causal inference from observational data then rests on the hope that the study can be viewed as a conditionally randomized experiment.


An observational study can be conceptualized as a conditionally randomized experiment if:

  1. each individual’s observed outcome is her counterfactual outcome under the treatment she actually received (consistency, Chapter 1);
  2. given only the known, measured covariates \(L\), the counterfactual outcomes are independent of the treatment actually received (exchangeability, Chapter 2);
  3. the probability of receiving each treatment value conditional on \(L\) is greater than zero (positivity, Technical Point 2.3).

These are the identifiability conditions (or identifiability assumptions).

In ideal randomized experiments the identifiability conditions hold by design. In observational studies they must be assumed, and they are often heroic, which is why causal inferences from observational data are viewed with suspicion. Causal inference from observational data requires two elements: data and identifiability conditions.

For simplicity, this chapter considers only studies in which everyone stays under follow-up and adheres to the assigned treatment; Chapters 8 and 9 relax this.

The book notes that Rubin (1974, 1978) extended Neyman’s theory to observational studies, and that Rosenbaum and Rubin (1983) called the combination of exchangeability and positivity weak ignorability, and the combination of full exchangeability (Technical Point 2.1) and positivity strong ignorability (Hernán and Robins 2020, 28).

1.1 Table 3.1


Table 3.1 contains the same data as Table 2.2: \(L\) is critical condition at baseline (1: yes), \(A\) is heart transplant, and \(Y\) is death.

Name \(L\) \(A\) \(Y\) Name \(L\) \(A\) \(Y\)
Rheia 0 0 0 Leto 1 0 0
Kronos 0 0 1 Ares 1 1 1
Demeter 0 0 0 Athena 1 1 1
Hades 0 0 0 Hephaestus 1 1 1
Hestia 0 1 0 Aphrodite 1 1 1
Poseidon 0 1 0 Polyphemus 1 1 1
Hera 0 1 0 Persephone 1 1 1
Zeus 0 1 1 Hermes 1 1 0
Artemis 1 0 1 Hebe 1 1 0
Apollo 1 0 1 Dionysus 1 1 0

If these data came from an observational study in which the three identifiability conditions held, we would compute the same causal risk ratio as in Chapter 2: 1.

Example 1 (Standardized risks in Table 3.1) Stratum-specific risks from Table 3.1 (8 individuals with \(L = 0\), 12 with \(L = 1\)):

  • \(L = 0\): \(\Pr[Y = 1 \mid A = 1, L = 0] = 1/4\) and \(\Pr[Y = 1 \mid A = 0, L = 0] = 1/4\)
  • \(L = 1\): \(\Pr[Y = 1 \mid A = 1, L = 1] = 6/9 = 2/3\) and \(\Pr[Y = 1 \mid A = 0, L = 1] = 2/3\)

with \(\Pr[L = 0] = 8/20 = 0.4\) and \(\Pr[L = 1] = 12/20 = 0.6\). Standardizing,

\[ \begin{aligned} \Pr[Y^{a=1} = 1] &= \Pr[Y = 1 \mid A = 1, L = 0]\Pr[L = 0] + \Pr[Y = 1 \mid A = 1, L = 1]\Pr[L = 1] \\ &= (1/4)(0.4) + (2/3)(0.6) \\ &= 0.1 + 0.4 = 0.5, \end{aligned} \]

\[ \begin{aligned} \Pr[Y^{a=0} = 1] &= \Pr[Y = 1 \mid A = 0, L = 0]\Pr[L = 0] + \Pr[Y = 1 \mid A = 0, L = 1]\Pr[L = 1] \\ &= (1/4)(0.4) + (2/3)(0.6) \\ &= 0.1 + 0.4 = 0.5, \end{aligned} \]

so the causal risk ratio is \(0.5 / 0.5 = 1\).

1.2 Identifiability


Definition 1 (Identifiability) An average causal effect is (nonparametrically) identifiable under a set of assumptions if those assumptions imply that the distribution of the observed data is compatible with a single value of the effect measure. It is nonidentifiable if the observed data distribution is compatible with several values of the effect measure (Hernán and Robins 2020, Fine Point 3.1, p. 29).

NoteFine Point 3.1: Identifiability of causal effects

If Table 3.1 came from a conditionally randomized experiment, \(Y^a \perp\!\!\!\perp A \mid L\) holds by design and the causal risk ratio of 1 is identified with no further assumptions. If it came from an observational study, the risk ratio equals 1 only if we add the identifying assumption \(Y^a \perp\!\!\!\perp A \mid L\), which is external to the data. Without it, the same data are compatible with a causal risk ratio

  • below 1, if risk factors other than \(L\) are more frequent among the treated;
  • above 1, if risk factors other than \(L\) are more frequent among the untreated;
  • equal to 1, if all risk factors other than \(L\) are equally distributed, i.e., \(Y^a \perp\!\!\!\perp A \mid L\).

1.3 Other identifiability conditions

When any of the three conditions fails, the analogy with a conditionally randomized experiment breaks down. Other approaches rely on different identifiability conditions; for example, instrumental variable methods (Chapter 16) assume that a predictor of treatment behaves as if randomly assigned conditional on measured covariates.

Methods based on the analogy with a conditionally randomized experiment have traditionally been favored in disciplines where that analogy is often reasonable (e.g., epidemiology), and instrumental variable methods in disciplines where it often is not (e.g., economics). Until Chapter 16, the book focuses on the first family.

2 3.2 Exchangeability (pp. 29-31)


In a marginally randomized experiment, the treated and untreated are exchangeable (\(Y^a \perp\!\!\!\perp A\)) because randomization balances the independent predictors of the outcome between groups.

An independent predictor of the outcome is a covariate associated with \(Y\) within levels of treatment; for dichotomous outcomes these are often called risk factors.

In Table 3.1, 69% of the treated (\(9/13\)) but only 43% of the untreated (\(3/7\)) were in critical condition (\(L = 1\)), so marginal exchangeability does not hold.


Definition 2 (Conditional exchangeability) The treated and untreated are conditionally exchangeable given \(L\) when

\[Y^a \perp\!\!\!\perp A \mid L \quad \text{for all } a.\]

Equivalently, for a dichotomous outcome, \(\Pr[Y^a = 1 \mid A = 1, L = l] = \Pr[Y^a = 1 \mid A = 0, L = l]\) for all \(a\) and \(l\).

Conditional exchangeability holds in a conditionally randomized experiment because, within levels of \(L\), all other outcome predictors are equally distributed between the treated and the untreated.

The study in Table 3.1 can be described in two logically equivalent ways:

  1. an observational study in which \(\Pr[A = 1 \mid L = 1] = 9/12 = 0.75\) and \(\Pr[A = 1 \mid L = 0] = 4/8 = 0.50\), e.g., because doctors direct the scarce hearts to those in critical condition; or
  2. a (nonblinded) conditionally randomized experiment in which investigators assigned \(A = 1\) with probability 0.75 when \(L = 1\) and 0.50 when \(L = 0\).

Under either description, if \(L\) is the only outcome predictor unequally distributed between treatment groups, \(Y^a \perp\!\!\!\perp A \mid L\) holds and standardization or IP weighting identifies the causal effect. Chapter 7 calls such outcome predictors confounders.

2.1 Not every imbalanced variable must be in \(L\)


Heart transplants are allocated by HLA compatibility, so HLA genes are unequally distributed between treatment groups. But HLA genes do not predict mortality given \(L\) and \(A\), so treatment is still effectively random within levels of \(L\), and HLA need not be included in the analysis.

2.2 Unmeasured outcome predictors

Suppose, unknown to the investigators, doctors prefer to transplant hearts into nonsmokers. Let \(U\) denote smoking (an unmeasured variable). Within the stratum \(L = 1\), smokers (\(U = 1\)) are less likely to be treated, so smoking is less common among the treated than the untreated, and \(Y^a \perp\!\!\!\perp A \mid L\) fails.

Remark 1 (Conditional exchangeability cannot be verified). Conditional exchangeability \(Y^a \perp\!\!\!\perp A \mid L\) fails whenever there is an unmeasured independent predictor \(U\) of the outcome such that the probability of treatment depends on \(U\) within strata of \(L\). Even when it holds, it cannot be checked: verifying it would require comparing \(\Pr[Y^a = 1 \mid A = a, L = l]\) with \(\Pr[Y^a = 1 \mid A \neq a, L = l]\), but \(Y^a\) is unknown for individuals with \(A \neq a\).

Because unmeasured variables cannot be used for standardization or IP weighting, the causal effect is not identified when the measured \(L\) is insufficient for conditional exchangeability. Collecting data on smoking would not remove the possibility that other unknown outcome predictors remain imbalanced.

Investigators can use expert knowledge to make the assumption more plausible, for example by measuring many determinants of treatment that are also independent outcome predictors and assuming exchangeability within the strata they define. But no matter how many variables are in \(L\), the assumption cannot be tested. The validity of the inference rests on the investigators’ expert knowledge, encoded as conditional exchangeability, which supplements the data.

NoteFine Point 3.2: Crossover randomized experiments

Fine Point 2.1 identified individual effects in crossover experiments (periods \(t = 0, 1\)) under three conditions: (i) no carryover, \(Y_{i,t=1}^{a_0, a_1} = Y_{i,t=1}^{a_1}\); (ii) a time-constant individual effect, \(Y_{it}^{a_t = 1} - Y_{it}^{a_t = 0} = \alpha_i\); and (iii) a time-constant untreated outcome, \(Y_{it}^{a_t = 0} = \beta_i\).

Now drop (iii), and randomize the order: each individual gets \((A_{i0}, A_{i1}) = (0, 1)\) or \((1, 0)\) with probability 0.5. Let \(r_i \stackrel{\text{def}}{=}Y_{i1}^{a_1 = 0} - Y_{i0}^{a_0 = 0}\). Under (i), (ii), and consistency:

\[ \begin{aligned} \text{if } (A_{i0}, A_{i1}) = (0, 1):\quad Y_{i1} - Y_{i0} &= Y_{i1}^{a_1 = 1} - Y_{i0}^{a_0 = 0} && \text{(consistency)} \\ &= \mathopen{}\left(\alpha_i + Y_{i1}^{a_1 = 0}\right)\mathclose{} - Y_{i0}^{a_0 = 0} && \text{(ii)} \\ &= \alpha_i + r_i, && \text{(definition of } r_i\text{)} \\ \text{if } (A_{i0}, A_{i1}) = (1, 0):\quad Y_{i0} - Y_{i1} &= \mathopen{}\left(\alpha_i + Y_{i0}^{a_0 = 0}\right)\mathclose{} - Y_{i1}^{a_1 = 0} = \alpha_i - r_i. \end{aligned} \]

Individual effects are no longer identified because \(r_i\) is unknown. But, using \(A_{i0} = 1 - A_{i1}\),

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[(Y_{i1} - Y_{i0}) A_{i1} + (Y_{i0} - Y_{i1}) A_{i0}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[(\alpha_i + r_i) A_{i1} + (\alpha_i - r_i) A_{i0}\right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[\alpha_i (A_{i1} + A_{i0})\right]\mathclose{} + \operatorname{E}\mathopen{}\left[r_i (A_{i1} - A_{i0})\right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[\alpha_i\right]\mathclose{} + \operatorname{E}\mathopen{}\left[r_i\right]\mathclose{}\,\operatorname{E}\mathopen{}\left[A_{i1} - A_{i0}\right]\mathclose{} && \text{($A_{i1} + A_{i0} = 1$; randomization: } A \perp\!\!\!\perp r_i\text{)} \\ &= \operatorname{E}\mathopen{}\left[\alpha_i\right]\mathclose{} + \operatorname{E}\mathopen{}\left[r_i\right]\mathclose{}(0.5 - 0.5) = \operatorname{E}\mathopen{}\left[\alpha_i\right]\mathclose{}. \end{aligned} \]

If only (i) holds, the same mean estimates \((\operatorname{E}\mathopen{}\left[\alpha_{i1}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\alpha_{i0}\right]\mathclose{})/2\), the average of the period-specific average effects. The book concludes that, for the treatments and outcomes it studies, the no-carryover assumption is implausible (Hernán and Robins 2020, Fine Point 3.2, p. 31).

3 3.3 Positivity (pp. 32-33)


An experiment that assigned everyone to \(A = 1\) (or everyone to \(A = 0\)) could not estimate the average causal effect. Treatment must be assigned so that each treatment level has a positive probability.

Definition 3 (Positivity) \[\Pr[A = a \mid L = l] > 0 \quad \text{for all values } l \text{ with } \Pr[L = l] \neq 0 \text{ in the population of interest,}\]

for every treatment value \(a\) involved in the causal contrast.

Positivity is sometimes called the experimental treatment assumption. It is taken for granted in experiments: in marginally randomized experiments \(\Pr[A = 1]\) and \(\Pr[A = 0]\) are positive by design, and in both marginally and conditionally randomized experiments \(\Pr[A = a \mid L = l]\) is positive by design for every level \(l\) of \(L\). In Table 3.1, \(\Pr[A = 1 \mid L = 1] = 0.75\) and \(\Pr[A = 1 \mid L = 0] = 0.50\); neither is 0 or 1, so positivity holds.

3.1 Two refinements


  • Positivity is needed only for values \(l\) present in the population of interest: if the study were restricted to \(L = 1\), positivity in \(L = 0\) would be irrelevant.
  • Positivity is needed only for the variables \(L\) required for exchangeability. We do not ask whether people with blue eyes had a positive probability of transplant, because eye color is not needed to make the groups exchangeable.

3.2 Positivity in observational studies

Example 2 (A positivity violation) If doctors always transplanted a heart to individuals in critical condition, then \(\Pr[A = 0 \mid L = 1] = 0\) (book Figure 3.1). The data would contain no untreated individuals with \(L = 1\) who could stand in for what would have happened to the treated with \(L = 1\) had they been untreated.

Unlike exchangeability, positivity can sometimes be checked empirically (Chapter 12). In Table 3.1 there are treated and untreated individuals in both levels of \(L\).

NoteTechnical Point 3.1: Positivity for standardization and IP weighting

The standardized mean \(\sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l]\) is defined only if \(\operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{}\) is defined for every \(l\) with \(\Pr[L = l] \neq 0\), i.e., only under positivity; otherwise it is undefined.

For IP weighting, \(\operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[a \mid L]}\right]\mathclose{}\) is undefined without positivity (it involves \(0/0\)). The version \(\operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[A \mid L]}\right]\mathclose{}\) is always defined, because \(f[A \mid L]\) is never 0, but it no longer equals the counterfactual mean. Let \(Q(a) \stackrel{\text{def}}{=}\{l : \Pr[A = a \mid L = l] > 0\}\). Then

\[ \operatorname{E}\mathopen{}\left[\frac{I(A = a) Y}{f[A \mid L]}\right]\mathclose{} = \Pr[L \in Q(a)] \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l, L \in Q(a)\right]\mathclose{} \Pr[L = l \mid L \in Q(a)], \]

which under exchangeability equals \(\operatorname{E}\mathopen{}\left[Y^a \mid L \in Q(a)\right]\mathclose{} \Pr[L \in Q(a)]\). For binary \(A\) without positivity, \(Q(0) \neq Q(1)\), so the IP weighted contrast compares two different groups and has no causal interpretation even under exchangeability. Under positivity \(Q(0) = Q(1)\) and the contrast is the average causal effect if exchangeability holds.

4 3.4 Consistency: First, Define the Counterfactual Outcome (pp. 33-38)


Definition 4 (Consistency) The observed outcome of every treated individual equals her outcome had she received treatment, and the observed outcome of every untreated individual equals her outcome had she remained untreated:

\[Y = Y^A,\]

where \(Y^A\) is the counterfactual \(Y^a\) evaluated at the individual’s actual treatment \(A\).

The book takes the counterfactuals \(Y^a\) as primitives; the observed \(Y\) is a function of them. For binary \(A\):

\[ Y^A = A\,Y^{a=1} + (1 - A)\,Y^{a=0}, \]

which equals \(Y^{a=1}\) when \(A = 1\) (the second term vanishes) and \(Y^{a=0}\) when \(A = 0\) (the first term vanishes).

Consistency can look trivially true: if you take aspirin and die, your counterfactual outcome had you taken aspirin is death. But it cannot always be taken for granted. Consistency has two components:

  1. a precise definition of the counterfactual outcomes \(Y^a\), through the specification of the superscript \(a\) (this section);
  2. the linkage of the counterfactual outcomes to the observed outcomes (Section 3.5).

4.1 Well-defined interventions


An individual’s causal effect \(Y^{a=1} - Y^{a=0}\) can be well-defined only if her \(Y^a\) is well-defined for both \(a = 1\) and \(a = 0\). In turn, the average causal effect \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\) in a given population can be well-defined only if every member of that population has a well-defined individual causal effect.

How do we know the counterfactuals are well-defined? A natural sufficient condition: when \(a\) is itself a well-defined intervention, \(Y^a\) is well-defined, as the outcome that would have been seen had that intervention been carried out.

Example 3 (Two heart transplant trials) Two ideal randomized trials of heart transplant (\(a = 1\)) versus medical therapy (\(a = 0\)) in the same population:

  • Trial 1 writes a detailed protocol: \(a = 1\) means assignment to transplant (\(a_0 = 1\)) followed by specified preoperative procedures (\(a_1\)), anesthesia (\(a_2\)), and surgical technique (\(a_3\)).
  • Trial 2 is a pragmatic trial: \(a = 1\) means assignment to transplant (\(a_0 = 1\)) followed by whatever happens in routine care, i.e., \(a_1 = A_1\), \(a_2 = A_2\), \(a_3 = A_3\).

In both trials \(a = 1\) is well-defined by the protocol, so \(Y^{a=1}\) is well-defined. But the two counterfactuals, and hence the two causal effects, will likely differ.

Formally, \(Y^{a=1}\) is the joint counterfactual \(Y^{a_0 = 1, a_1, a_2, a_3}\) in Trial 1 and \(Y^{a_0 = 1, A_1^{a_0=1}, A_2^{a_0=1}, A_3^{a_0=1}}\) in Trial 2; for those assigned to transplant in Trial 2, this equals the observed \(Y\) (Technical Point 3.2).

The same treatment name is used with different meanings. Which effect is preferable is a question for the consumers of the research. Trial 1’s effect may vary less across populations, because its components are fixed, whereas the natural values \((A_1, A_2, A_3)\) in Trial 2 depend on the setting; so Trial 1’s result may be easier to transport (Chapter 4). This is a question of transportability, not of whether each effect is well-defined.

4.2 Interventions are never perfectly specified


Even Trial 1 did not specify the surgeon’s training. Because surgical experience affects post-transplant mortality, \(\Pr[Y^{a=1} = 1]\) depends on the (unknown) mix of surgeons. Had the protocol specified training, some other detail would still have been left unspecified or open to interpretation.

  • The average causal effect can therefore differ across populations that follow the same protocol: a community whose surgeons differ in experience from the trial’s will typically see a different effect.
  • The more precisely we define the interventions, the easier it generally is to transport the result.
  • Components with no effect on the outcome (the color of the surgeons’ scrubs) cannot affect transportability.

The phrase “no causation without manipulation” (Holland 1986) captures the idea that causal inference requires sufficiently well-defined interventions. Such interventions need not be feasible at the time: the effect of genetic variants on disease was considered well-defined before gene-editing technology existed.

4.3 Sufficiently well-defined interventions


In experiments, the protocol makes the interventions explicit. In observational studies, investigators who talk about “the effect of heart transplant” without defining the intervention may mean different things, so \(Y^a\) is ill-defined until the experts agree on one intervention at a time.

An intervention is sufficiently well-defined when, for all practical purposes, no meaningful vagueness remains for \(Y^a\).

How do we know no meaningful vagueness remains? We don’t. It is a matter of agreement among experts, based on the substantive knowledge available at a particular time, and experts who agree today may judge differently once new knowledge arrives.

NoteFine Point 3.3: Possible worlds

Philosophers (Stalnaker 1968, Lewis 1973) defined \(Y^a\) as the value of \(Y\) in the closest possible world in which the individual received \(a\). Since the closest world to the actual world is itself, \(Y^a = Y\) when \(A = a\), so consistency always holds under this definition. When \(A \neq a\), the closest possible world is vague. Robins and Greenland (2000) argued that well-defined interventions should replace closest possible worlds, because in observational studies counterfactuals are vague to the degree that the hypothetical interventions are not made precise.

4.4 Non-intervention variables


Biological states (blood pressure, LDL cholesterol, body weight) and social factors (socioeconomic status) can be changed only by intervening on their causes.

  • “The effect of becoming obese on myocardial infarction” has no meaningful interpretation: the counterfactual is ill-defined. Specifying the timing and the method of weight change (medication, surgery, diet, exercise) would turn it into the effect of those interventions, not of obesity.
  • Yet experts agree that “blood pressure causes stroke”, by synthesizing physics, in vitro studies, animal experiments, autopsies, and intervention studies.

Whether an effect is ill-defined also depends on the outcome. For obesity and job discrimination (whether an employer invites an applicant to interview after reviewing the application and photograph), the treatment is really whether the employer perceives the applicant as obese, and how the applicant became obese may be irrelevant. A margin note points to Hernán and Taubman (2008), a paper about the difficulties a tyrannical monarch and an inept head of state run into when weighing “the effect of obesity” (Hernán and Robins 2020, 37).


Definition 5 (Effective no-direct-effect (ENDE) intervention) One way to read the claim that blood pressure \(A\) affects this particular individual’s stroke outcome \(Y\) is as a belief that some intervention \(X\) exists (possibly not yet discovered) that can change this individual’s value of \(A\) to \(a\) and that affects her outcome \(Y\) only through \(A\), with no (direct) effect of its own. Such an \(X\) is an effective no-direct-effect intervention for \(a\), written ENDE(\(a\)). Its existence is logically equivalent to the statement that the counterfactuals \(Y^a\) and the joint counterfactuals \(Y^{x,a}\) are both well-defined and equal.

Antihypertensive drugs, exercise, and diet all change blood pressure. Some antihypertensive drugs might serve as ENDE interventions \(X\); diet and exercise cannot, because they affect stroke risk through pathways other than blood pressure. “\(X\) has no direct effect on \(Y\) except through \(A\)” means the same as “the effect of \(X\) on \(Y\) is completely mediated by \(A\)” (Chapter 23).

Belief in ENDE interventions is easier for variables whose mechanisms are well understood (blood pressure and stroke) than for others (socioeconomic status and myocardial infarction). Where experts disagree, a study design might support the claim that the effect of \(X\) on \(Y\) is completely mediated by \(A\) (Fine Point 6.4).

Blood pressure is really a vector \((A_{sys}, A_{dia})\). Accounting for this, the claim that \(X\) has no direct effect on \(Y\) except through \(A\) asserts that \(Y^{a_{sys}, a_{dia}}\) and \(Y^{x, a_{sys}, a_{dia}}\) are well-defined and equal.

4.5 Average effects of non-interventions


Because an average causal effect requires a well-defined individual effect for every member of the population, the average effect of a non-intervention \(A\), comparing \(a\) with \(a'\), is well-defined if every individual in the population has both an ENDE(\(a\)) and an ENDE(\(a'\)) intervention.

  • If, for some individuals, no ENDE intervention can move \(A\) to \(a\) and to \(a'\), the population effect is not defined.
  • A well-defined population effect therefore requires restricting attention to those individuals who have ENDE interventions available.

Considering small changes in \(A\) makes ENDE interventions more credible. For a change in blood pressure of \(\Delta\) mm Hg with \(\Delta\) near 0:

  • it is more plausible that (essentially) every individual has an effective intervention;
  • direct effects may appear only with large changes (an antihypertensive drug without direct effects at low doses may be toxic at high doses).

So \(Y^{\Delta}\) and \(\operatorname{E}\mathopen{}\left[Y^{\Delta} - Y^{\Delta = 0}\right]\mathclose{}\) are (essentially) well-defined for all individuals, where \(Y^{\Delta}\) is \(Y^a\) evaluated at \(a = A - \Delta\) and \(A\) is the individual’s blood pressure before the intervention begins.

NoteFine Point 3.4: An interventionist approach to causal inference

The book’s framework, sometimes called “interventionist”, rests on well-defined counterfactual outcomes, which are in turn defined through well-defined interventions. It supplies the formal language for precise discussion of causal inference about interventions and for interpreting the numbers that data analyses produce. Many published papers, however, attach causal interpretations to numerical quantities for questions that involve no recognizable intervention. The book therefore extends the interventionist approach to non-interventions by spelling out when quantitative causal inference about them is meaningful. The authors do not claim that this is the only valid philosophy of causality, but they know of no alternative framework that produces practically interpretable estimates of causal effects (Hernán and Robins 2020, Fine Point 3.4, p. 39).

6 3.6 The Target Trial (pp. 41-46)


Assuming the three identifiability conditions amounts to viewing an observational analysis as an attempt to emulate a hypothetical randomized experiment.

Definition 6 (Target trial) The target trial (or target experiment) is the hypothetical randomized experiment that would quantify the causal effect of interest. For each causal effect, we may (i) specify the protocol of the target trial that we would like to, but cannot, conduct, and (ii) describe how the observational data would be used to emulate it.

If the emulation were successful, the observational study and the target trial (had it been conducted) would give the same results.

The target trial, or its logical equivalents, has a long history: the book cites Dorn (1953), Wold (1954), Cochran (1972), Rubin (1974), Feinstein (1971), and Dawid (2000), and Robins (1986) generalized it to time-varying treatments. The key components were specified by Hernán and Robins (2016) and Hernán et al. (2025). Chapter 22 returns to the framework.

6.1 Key components of the protocol


  • eligibility criteria
  • interventions (or, generally, treatment strategies)
  • assignment
  • outcomes
  • start and end of follow-up
  • causal contrasts

Eligibility criteria must be chosen so that every eligible individual could in principle receive every treatment of interest; otherwise some individual causal effects, and hence the average effect, would be undefined. In Parts I and II, each individual meets the eligibility criteria at one time only, which becomes time zero (the start of follow-up). The acronym PICO (Population, Intervention, Comparator, Outcome; Richardson et al. 1995) summarizes some of these components.

6.2 Interventions and target trials


  • Experiments: in a fully adherent trial of low-dose aspirin \(A\) and stroke \(Y\), \(A\) is the investigators’ action, so \(Y^a\) is well-defined; the target trial is the trial actually conducted.
  • Observational studies of interventions: if \(A\) indicates prescription of aspirin, in a population without absolute indications or contraindications, most experts would accept “prescribed” and “not prescribed” as sufficiently well-defined; a target trial could have been conducted, and specifying its protocol characterizes the causal effect.
  • Non-interventions: the target trial randomizes an ENDE intervention \(X\) rather than \(A\), if experts believe one exists.

Whether a variable corresponds to an intervention is often debatable: assignment in a randomized trial uncontroversially is one, socioeconomic status uncontroversially is not, and there is a gradient in between.

6.3 Three non-intervention variables


  1. Genotype. \(Y^a\) is well-defined if ENDE interventions \(X\) could set each target-population member’s genotype to \(a\) at conception, so that \(Y^a\) and \(Y^{x,a}\) are equal for each of them. The interventions need not be known: decades ago most experts considered the effect of genotype well-defined, and gene editing was eventually developed.
  2. Systolic blood pressure. \(Y^a\) is well-defined for everyone in the population if ENDE interventions \(X\) could bring each individual’s blood pressure to \(a\). Many experts believe some drugs act as ENDE interventions in some populations, perhaps only over a limited range, so the effect of modest changes can be defined through a target trial that randomizes \(X\) (the small-change argument of Section 3.4).
  3. Socioeconomic status. Most experts doubt that interventions could place people at a given status without affecting \(Y\) through other means (they may even disagree on what the term means), so no target trial can be specified and the effect is not well-defined.

Remark 2 (Summary of the target trial for interventions and non-interventions). If \(A\) is an intervention, the target trial randomizes eligible individuals to values \(a\). If \(A\) is a non-intervention and we believe ENDE interventions \(X\) exist, the target trial randomizes eligible individuals to an \(X\) that sets \(A\) to \(a\), and the effect of \(A\) is equated with the effect of \(X\). Without that belief, \(Y^a\) and the effect of \(A\) are not well-defined.

NoteFine Point 3.6: When causal inference for non-interventions goes unquestioned

In a trial with assignment \(Z\) and received treatment \(A\), the intention-to-treat effect of \(Z\) is well-defined because \(Z\) is an intervention (Hernán and Robins 2020, Fine Point 3.6, p. 43). With perfect adherence and no direct effect of \(Z\), \(Z = A\) and \(Z\) is an ENDE intervention for \(A\). Without perfect adherence, \(Z\) is not an ENDE intervention for everyone, because it does not succeed in setting \(A\) to 1 or 0 for every individual (even if \(Z\) has no direct effect on \(Y\)). The effect of \(A\) in the whole population is then well-defined only if some ENDE intervention \(X\) on \(A\) exists (e.g., adding aspirin to food without the participants’ knowledge). Most experts find this so self-evident that “the effect of received treatment” is discussed without mentioning \(X\).

Whether or not such an \(X\) exists, \(Z\) is effective for one subset: individuals who would have \(A = 1\) under \(Z = 1\) and \(A = 0\) under \(Z = 0\). For them, \(Z\) is an ENDE intervention, so the effect of \(A\) in that subset is well-defined and equals the effect of \(Z\). Chapter 16 calls this subset the compliers; its members cannot be identified individually, but the effect of \(A\) among them can be identified under additional assumptions.

If \(Z\) has direct effects on \(Y\) not through \(A\), the effect of \(Z\) differs from that of \(X\) even under full adherence.

Some authors treat “the causal effect of \(A\) on \(Y\)” as well-defined even where many experts would not grant that any ENDE intervention \(X\) on \(A\) exists (Pearl 2009; Schwartz et al. 2016; Glymour and Spiegelman 2016) (Hernán and Robins 2020, 42).

6.4 What a valid emulation needs


  • enough information to map the protocol to the data: identify eligible individuals, classify them by intervention, and ascertain outcomes;
  • when using IP weighting or standardization, enough adjustment variables \(L\) for conditional exchangeability. For “the effect of blood pressure on stroke”, investigators who believe in an ENDE intervention would adjust for variables, such as diet and exercise, that they believe affect both blood pressure and stroke.

Chapter 16 considers alternative identifying conditions for emulating a target trial.

6.5 Anchoring analyses to realistic interventions


Example 6 (Body mass index and death) Comparing the risk of death between people with body mass index (BMI) 25 versus 30 at the start of follow-up implies a target trial of an instantaneous, possibly very large, weight change, for which no ENDE intervention is known. A modified analysis could emulate a more reasonable target trial, e.g., one assigning some individuals to a 5% reduction in BMI every year, starting at age 40, for as long as their BMI stays over 25 (a sustained strategy of the kind studied in Part III).

Experts may still disagree that ENDE interventions exist for gradual weight change, but such interventions are more plausible than instantaneous weight loss. One course of action is to proceed as if they existed, adjusting for prognostic factors \(L\), while accepting that the results lack a causal interpretation and may be inadequate for decisions. Such an analysis may still tell us something about what ENDE interventions would do were they to exist, or inspire technologies to create them. Danaei et al. (2016) considered unknown interventions producing progressive weight loss over several years in a study of weight loss and heart disease.

At the very least, specifying a target trial aligns causal questions with interventions that are feasible or not totally unrealistic, and helps decision makers recognize analyses that imply extreme or impossible interventions.

NoteFine Point 3.7: Some limits of target trial emulation

“The effect of heart transplant” can mean the average counterfactual outcome of the \(n\) eligible individuals under:

  1. an intervention in which all \(n\) individuals receive treatment concurrently; or
  2. an intervention in which each individual \(i\) receives treatment while everyone else receives the treatment they actually received.

Interpretation (ii), an average of \(n\) interventions, is the one relevant to a physician and patient deciding on transplant, and the one the book uses implicitly. Interpretation (i) is not well-defined: it does not say how the health system would be redesigned to supply the hearts and the capacity. If that redesign were specified, current observational data would be inadequate, because they were generated under the old system (e.g., a system that transplants everyone might accept lower-quality organs). This issue is related to interference (Fine Point 1.1), although the interference literature usually assumes interpretation (i) is well-defined. Observational data may be insufficient to characterize the effect of scaling up an intervention.

NoteTechnical Point 3.2: Recursive substitution

For chronologically ordered variables \(L, A, M, Y\) with well-defined interventions on \(L, A, M\), the one-step-ahead counterfactuals are \(L, A^l, M^{l,a}, Y^{l,a,m}\). All other factuals and well-defined counterfactuals are functions of them via recursive substitution, e.g., \(A = A^L\), \(M^a = M^{L,a}\), \(Y^a = Y^{L, a, M^a}\), and \(Y = Y^{L, A, M}\). Applied to the two transplant trials, recursive substitution shows why Trial 2’s outcome distribution is harder to transport: it changes whenever the distribution of the natural values \(A_1^{a_0=1}, A_2^{a_0=1}, A_3^{a_0=1}\) differs between populations. The argument applies equally to observational and randomized studies.

7 Summary


The identifiability conditions under which an observational study can be analyzed like a conditionally randomized experiment:

  1. Exchangeability: \(Y^a \perp\!\!\!\perp A \mid L\) (untestable; requires expert knowledge)
  2. Positivity: \(\Pr[A = a \mid L = l] > 0\) for all \(a\) in the contrast and all \(l\) with \(\Pr[L = l] \neq 0\) (sometimes checkable)
  3. Consistency: \(Y = Y^A\), which requires (i) sufficiently well-defined interventions and (ii) a link between those interventions and the observed data

Specifying the target trial makes the causal question, and what the data must supply to emulate it, explicit. For non-intervention variables, the target trial randomizes an ENDE intervention \(X\).

Looking ahead:

  • Chapter 4 discusses effect modification and the dependence of causal effects on the population
  • Chapter 6 introduces causal diagrams; Chapters 7-9 cover confounding, selection bias, and measurement bias
  • Chapter 16 describes instrumental variable methods, which rely on different identifiability conditions
  • Chapter 22 returns to target trial emulation

8 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top