Part I described causal inference from observational data as an attempt to emulate a hypothetical randomized trial, the target trial, but only for simple target trials comparing time-fixed treatments. With the g-methods of Chapters 19-21 in hand, we can now specify realistic target trials that compare sustained treatment strategies, and emulate them with either randomized or observational data.
Example 1 (A trial of vaccination plus an antiviral) Consider a randomized trial of a virus that can kill:
\(Z\) and \(A\) can differ: some people assigned to vaccine refuse it, and some assigned to no vaccine get vaccinated outside the study.
Fine Point 22.1: The Exclusion Restriction (Again)
An arrow \(Z \to Y\) means the exclusion restriction does not hold (see Technical Point 16.1 and Chapter 16). Investigators often try to remove that arrow by blinding: those assigned \(Z=1\) get the vaccine and those assigned \(Z=0\) get an identical placebo injection, so neither participants nor their doctors know the assignment (a double-blind placebo-controlled trial). Blinding is often infeasible (no convincing placebo exists for open heart surgery; side effects reveal who is treated), and it is not advisable when the goal is the effect of treatment in the real world, where there is no blinding or placebo (Hernán and Robins 2020, Fine Point 22.1, p. 306).
Definition 1 (Intention-to-Treat (ITT) Effect) The intention-to-treat effect is the causal effect of randomized assignment \(Z\), for example the causal risk ratio \[\frac{\Pr[Y^{z=1} = 1]}{\Pr[Y^{z=0} = 1]}.\] It is “the effect of having the intention of treating with \(A\),” not “the effect of treating with \(A\)” (Hernán and Robins 2020, 306).
Because \(Z\) is randomized, there are no backdoor paths from \(Z\) to \(Y\), so \(Y^z \perp\!\!\!\perp Z\).
Proposition 1 (The ITT effect equals the association between assignment and outcome) Suppose \(Z\) is randomized, so that \(Y^z \perp\!\!\!\perp Z\) for \(z = 0, 1\), that consistency holds for \(Z\) (\(Y = Y^z\) whenever \(Z = z\)), and that \(\Pr[Z = z] > 0\) for \(z = 0, 1\). Then, for \(z = 0, 1\), \[\Pr[Y = 1 \mid Z = z] = \Pr[Y^z = 1].\] Hence, whenever \(\Pr[Y = 1 \mid Z = 0] > 0\) (equivalently, \(\Pr[Y^{z=0} = 1] > 0\)), the associational risk ratio equals the ITT risk ratio: \[\frac{\Pr[Y = 1 \mid Z = 1]}{\Pr[Y = 1 \mid Z = 0]} = \frac{\Pr[Y^{z=1} = 1]}{\Pr[Y^{z=0} = 1]}.\]
Proof. For each \(z\), \(\Pr[Y = 1 \mid Z = z] = \Pr[Y^z = 1 \mid Z = z]\) by consistency, and \(\Pr[Y^z = 1 \mid Z = z] = \Pr[Y^z = 1]\) by \(Y^z \perp\!\!\!\perp Z\). This proves the per-arm equality. When \(\Pr[Y = 1 \mid Z = 0] > 0\), the per-arm equality makes the two denominators equal and nonzero, so taking the ratio of the \(z = 1\) and \(z = 0\) expressions gives the risk ratio equality.
Definition 2 (Intention-to-treat analysis) Estimating the ITT effect by the unadjusted associational risk ratio \(\Pr[Y = 1 \mid Z = 1] / \Pr[Y = 1 \mid Z = 0]\) (Proposition 1), or another unadjusted contrast of the randomized groups, is an intention-to-treat analysis.
Fine Point 22.2: Pseudo- and Modified Intention-to-Treat Analyses
An ITT analysis (Definition 2) is unbiased because it includes all randomized individuals; variations that include only a subset may be biased.
Definition 3 (Per-Protocol (PP) Effect) The per-protocol effect is the causal effect under full adherence, that is, if everyone had followed the protocol’s instructions for the treatment they were assigned: the contrast of \(\Pr[Y^{z=1,a=1} = 1]\) versus \(\Pr[Y^{z=0,a=0} = 1]\), or, under the exclusion restriction, of \(\Pr[Y^{a=1} = 1]\) versus \(\Pr[Y^{a=0} = 1]\).
Unlike the ITT effect, the PP effect is generally confounded.
Example 2 (Confounding of the per-protocol effect) Suppose \(U\) is high risk of infection, and high-risk people assigned \(Z=0\) seek vaccination outside the study. Then the backdoor path \(A \leftarrow U \rightarrow Y\) makes \(\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]\) differ from \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\).
Remark 1 (A trial viewed as an observational study). Estimating the PP effect requires viewing the trial as an observational study: adjustment under conditional exchangeability given measured covariates, or alternative assumptions such as those of instrumental variable estimation (Chapter 16).
Fine Point 22.3: Naive Per-Protocol Analyses
Both are observational analyses of a randomized experiment and require adjustment for confounding and selection bias (Hernán and Robins 2020, Fine Point 22.3, p. 308).
The ITT Effect Need Not Preserve the Null or Be Conservative
The ITT effect is privileged largely because it is unconfounded, not because it is the effect we want. Two common justifications deserve “a grain of salt” (Hernán and Robins 2020, 307):
Fine Point 22.4: More Misunderstandings About the ITT Effect
“ITT measures effectiveness in the real world; PP measures efficacy.” This reasoning is problematic because:
“ITT is always conservative.” Not if the effect is non-monotonic (Technical Point 5.2) and non-adherence is high. Even for monotonic effects, it can fail in head-to-head trials: in a trial of an expensive drug (\(Z = 1\)) versus ibuprofen (\(Z = 0\)) for severe pain at 1 year, both drugs are equally effective (PP risk ratio 1), but adherence to ibuprofen is lower because of an easily palliated side effect. The ITT comparison (Definition 2) then wrongly suggests ibuprofen is less effective (Hernán and Robins 2020, Fine Point 22.4, p. 309).
Definition 4 (Pragmatic trial) A pragmatic trial is a trial designed to estimate the effect of treatment strategies under conditions close to those of routine clinical care.
Because the goal is to emulate target trials with real-world data, we consider pragmatic trials with these features:
Example 3 (A Target Trial of Antiretroviral Therapy)
| Component | Specification |
|---|---|
| Eligibility criteria | HIV infection, age 18 or older, no AIDS, no previous antiretroviral therapy |
| Treatment strategies | \(g_1\): receive therapy (\(A_k = 1\)) continuously unless a contraindication or toxicity arises; \(g_0\): receive no therapy (\(A_k = 0\)) continuously |
| Assignment | Random, at baseline \(k = 0\); \(Z = 1\) if assigned to \(g_1\), \(Z = 0\) if assigned to \(g_0\) |
| Follow-up | From assignment until death, loss to follow-up, or 60 months, whichever comes first |
| Outcome | Death; \(D_k = 1\) if dead by month \(k\) |
| Causal contrasts | Intention-to-treat and per-protocol effects on the risk of death |
Here \(k = 0, 1, \ldots, K\) with \(K = 59\), and \(C_k\) indicates censoring by month \(k\), for \(k = 1, \ldots, K + 1\) (Hernán and Robins 2020, 309–10).
Definition 5 (ITT and PP effects for sustained strategies) With the strategies and notation of Example 3:
ITT effect at time \(k\): a contrast of the static strategies “be assigned to \(g_1\) (or \(g_0\)) at baseline, with no loss to follow-up”: \[\Pr\mathopen{}\left[D_k^{z=1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{z=0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]
PP effect at time \(k\): a contrast of “receive strategy \(g_1\) (or \(g_0\)) continuously between baseline and end of follow-up”: \[\Pr\mathopen{}\left[D_k^{g_1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{g_0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]
Both are defined as if nobody had been lost to follow-up through time \(k\) (\(\bar{c}_k = \bar{0}\)).
Remark 4 (Dynamic strategies). Sensible protocols do not mandate treatment no matter what: \(g_1\) requires stopping therapy when a contraindication or toxicity arises. So the PP effect generally compares dynamic strategies (“do this; if X happens, do this other thing”).
Stopping for Toxicity Is Adherence
An individual assigned to \(g_1\) who stops therapy because of toxicity is adhering to \(g_1\), not deviating from it, even if the protocol describes \(g_1\) loosely as “treat continuously.”
Example 4 (An alternative target trial) Suppose that during the trial a consensus emerges that \(g_0\) is inferior, and physicians start treating \(g_0\) participants once their CD4 count (\(L_k\)) first drops below 200 cells/µL. Many \(g_0\) participants then follow
\(g_0'\): “receive \(A_k = 0\) continuously, but switch to \(A_k = 1\) after \(L_k < 200\).”
The contrast of \(g_1\) versus \(g_0'\) is neither the ITT nor the original PP effect. It is the PP effect of another target trial, randomizing \(g_1\) versus \(g_0'\), which can be emulated with the actual trial’s data.
Technical Point 22.1: Controlled Direct Effects
The controlled direct effect of \(A\) on \(Y\) with mediator \(M\) set to \(m\) is \(\operatorname{E}\mathopen{}\left[Y^{a=1,m}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,m}\right]\mathclose{}\), for \(m = 0\) or \(m = 1\). It could be identified by a trial that randomizes \(A\) at baseline and \(M\) one month later, so that \(\Pr[Y^{a,m} = 1] = \Pr[Y = 1 \mid A = a, M = m]\), or by emulating such a trial when consistency, positivity, and exchangeability hold for both \(A\) and \(M\). It is just a contrast of sustained strategies: replace \(A\) and \(M\) by \(A_0\) and \(A_1\) (Chapter 19) (Hernán and Robins 2020, Technical Point 22.1, p. 311).
Technical Point 22.2: Pure and Principal Stratum Direct Effects
Chapter 23 presents yet another type of direct effect (Hernán and Robins 2020, Technical Point 22.2, p. 312).
When a pragmatic trial (Definition 4) is not possible, we emulate it with existing observational data; the trial is then the target trial of the observational analysis.
Definition 6 (Target trial protocol) At a minimum, the protocol of a target trial specifies:
Definition 7 (Observational analog of the ITT effect) The actual assignment is unknown in existing observational data, so a true ITT effect can rarely be emulated. The closest analog compares initiators of the strategies, “initiate \(A_0 = 1\) (or \(A_0 = 0\)) at baseline, with no loss to follow-up”: \[\Pr\mathopen{}\left[D_k^{a_0=1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{a_0=0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]
Remark 7 (Observational analog of the PP effect). The observational analog of the PP effect is defined exactly as in the target trial. Without a pre-specified protocol, every per-protocol effect corresponds to some target trial, so there is no distinction between the “original” PP effect and PP effects of alternative trials.
Only Strategies Followed in the Data Can Be Emulated
We can only emulate target trials whose strategies are actually followed by at least some individuals in the data, unless we are willing to extrapolate with models such as dose-response structural models.
Definition 8 (Time Zero) Time zero (baseline, start of follow-up) is the time at which eligibility criteria must be met (but not later) and after which outcomes begin to be counted (but not earlier). In a randomized trial, it is the time an eligible individual is assigned to a strategy.
Start Follow-Up as the Target Trial Would
Follow-up in the emulation should start when it would have started in the target trial.
Two Sources of Time Zero Errors
Two problems cause errors in emulating time zero:
Example 5 (Eligible once versus eligible many times)
Options for time zero: (a) the first eligible time, (b) a random eligible time, or (c) every eligible time.
Definition 9 (Sequential target trials) Emulating sequential target trials means starting a new emulated trial at every time an individual meets the eligibility criteria, as in option (c).
Example 6 (Data compatible with several strategies) Target trial: individuals whose CD4 count just dropped below 500 cells/µL are assigned to start therapy
Those who started at time zero follow strategy 1, but those who did not are compatible with both 2 and 3.
Algorithm 1 (Cloning, censoring, and weighting) Copy each individual whose data are compatible with more than one strategy (as in Example 6) into one clone per compatible strategy, and censor each clone when its data stop being consistent with its strategy. The likely informative censoring is corrected by IP weighting for time-varying factors.
Fine Point 22.5: Grace Periods
Therapy cannot be started on the very day it is assigned, so “immediate” initiation needs a grace period (say, 3 months) during which initiation still counts as immediate; otherwise the study would compare strategies that rarely occur or could not be implemented.
During the grace period an individual’s data are consistent with more than one strategy (e.g., someone who starts in month 3 is consistent with both “initiate within 3 months” and “never initiate” during months 1 and 2), so cloning and censoring are again used, with IP weighting to handle the censoring.
Consequences:
(Hernán and Robins 2020, Fine Point 22.5, p. 317)
The book’s chapter does not use the term, but misaligned time zero classically produces a bias that has its own name.
Definition 10 (Immortal time) Immortal time is a period of follow-up during which, by the way groups are defined, the treated cannot experience the outcome. The bias that results from misclassifying such a period as treated person-time, or from excluding it from the analysis, is immortal time bias.
Example 7 (Immortal Time Bias: Statins and Mortality) Define “statin users” as people who filled a statin prescription at any time during a one-year window, and compare their mortality over that year with “non-users.” Anyone who dies before filling a prescription is automatically a non-user, and users must survive until their first fill. The time from the start of follow-up to the first fill is “immortal” for users (Definition 10), which produces a spurious survival advantage. Starting users’ follow-up at their first fill instead, while non-users are followed from the start of the window, still favors users: the early deaths of people who would have filled a prescription stay in the non-user group. Starting follow-up when eligibility is met and strategies are assigned, as the target trial requires, avoids this bias (see Hernán, Sauer, Hernández-Díaz, Platt, and Shrier, 2016, listed in the book’s references).
Remark 11 (Target trial emulation as a unifying framework). Explicit target trial emulation brings together the two frameworks of this book, counterfactuals and causal diagrams, and grounds them in actionable causal inference:
Remark 12 (Randomization is the only difference). A randomized trial is a follow-up study with baseline randomization; observational longitudinal data form a follow-up study without it. Long-term trials of sustained strategies in real-world settings, with imperfect adherence and loss to follow-up, suffer the confounding and selection biases usually associated with observational studies.
“Time-varying confounding in observational studies is a bias with the same structure as nonrandom noncompliance in randomized trials” (Hernán and Robins 2020, 318).
Fine Point 22.6: How Do Randomized and Observational Data Differ?
In a randomized experiment:
An observational analysis can emulate (i) when a sufficient set of covariates is measured and adjusted for, and (ii) when the treatment model given the past is correctly specified. Feature (iii) is not needed for a per-protocol effect in either design, because efficient estimators ignore it: a trial’s assignment variable could be dropped from the data without losing the per-protocol effect, as long as a sufficient set of confounders had been measured. With dynamic strategies and full adherence, the covariates the strategies use to decide treatment form such a set (Robins, 1986) (Hernán and Robins 2020, Fine Point 22.6, p. 319).
ITT Analyses Can Suffer Selection Bias
ITT analysis (Definition 2, an unadjusted comparison of randomized groups): randomization rules out baseline and post-randomization confounding for the effect of assignment, but not selection bias from loss to follow-up. Valid ITT estimation may need adjustment for time-varying prognostic factors, e.g., g-methods if dropout depends on symptom onset.
Conventional Per-Protocol Analyses Are Questionable
Conventional per-protocol analysis (censor at the first deviation, no adjustment) is questionable for three reasons:
Randomized trials may be expensive, infeasible, unethical, or too slow for an urgent decision, so many decisions must be made without them.
Judge an Observational Analysis by Its Emulation
When we cannot run the trial that would answer our question, the observational analysis should explicitly emulate it and be judged by how well it emulates its target trial.
Fine Point 22.7: A Counterintuitive Comparison of a Trial and an Observational Study
A double-blind placebo-controlled trial of an over-the-counter treatment \(A\) enrolled a random 20% of people diagnosed with lung cancer; everyone adhered. The 60-month mortality was 550/1000 = 55% with \(A = 1\) and 450/1000 = 45% with \(A = 0\), so the regulator banned \(A\). An observational study of the other 80% found 0% mortality among both treated and untreated.
Classify individuals into types: doomed (\(Y^{a=0} = Y^{a=1} = 1\)), hurt (\(Y^{a=0} = 0\), \(Y^{a=1} = 1\)), helped (\(Y^{a=0} = 1\), \(Y^{a=1} = 0\)), immune (\(Y^{a=0} = Y^{a=1} = 0\)). Random sampling and randomization give the trial arms and the observational sample the same distribution of types. Then:
In the observational data every “hurt” person took \(A = 0\) and every “helped” person took \(A = 1\): everyone followed the optimal strategy. The trial compared “treat everyone” with “treat no one,” but the best strategy was “treat only those who benefit.” If, say, the type were determined by ethnic group and each group had learned from experience whether to take \(A\) (maximal effect modification and maximal confounding), the confounded observational study, not the unconfounded trial, revealed the correct policy (Hernán and Robins 2020, Fine Point 22.7, p. 321).
Fine Point 22.8: Generalizing Fine Point 22.7
Suppose lower \(Y\) is better and the observational mean \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\) is below the mean of both trial arms, so \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\). If \(U\) is a (possibly unknown) set of pre-treatment covariates with \(Y^a \perp\!\!\!\perp A \mid U\), then \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) for the random strategy \(g\) that assigns \(A = 1\) with probability \(\Pr[A = 1 \mid U]\), the strategy that generated the observational data. That strategy cannot be implemented without data on \(U\), but it can motivate measuring pre-treatment covariates \(V\) and using the trial data to find a deterministic dynamic strategy \(g^*\) whose mean, estimated from the trial, is below the observational \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\). Combining trial and observational data can thus be more informative than the trial alone, provided both are random samples of everyone eligible for the trial (Hernán and Robins 2020, Fine Point 22.8, p. 322).