Chapter 22: Target Trial Emulation
Part I described causal inference from observational data as an attempt to emulate a hypothetical randomized trial, the target trial, but only for simple target trials comparing time-fixed treatments. With the g-methods of Chapters 19-21 in hand, we can now specify realistic target trials that compare sustained treatment strategies, and emulate them with either randomized or observational data.
This chapter is based on Hernán and Robins (2020, chap. 22, pp. 305-322).
The chapter does three things:
- it builds a taxonomy of causal effects in trials: the intention-to-treat effect, the per-protocol effect, and per-protocol effects in alternative target trials;
- it defines observational analogs of those effects and describes how to emulate a target trial, including the choice of time zero;
- it argues that, aside from baseline randomization, randomized trials and observational studies of sustained strategies should be analyzed the same way.
What makes this more than a formal exercise is the existence of g-methods: if data on all important fixed and time-varying confounders are available, the effects of interest can be validly estimated.
1 22.1 Intention-to-Treat Effect and Per-Protocol Effect (pp. 305-309)
In the book’s Figure 22.1, \(Z\) affects \(Y\) through two pathways:
- \(Z \to A \to Y\): assignment changes the treatment received, which in turn affects mortality;
- \(Z \to Y\) directly: knowing one’s assignment can change behavior. For example, people who know they were assigned to vaccination plus a promising antiviral may become less careful about avoiding infection.
The effect of \(Z\) therefore depends on the strength of three arrows: \(A \to Y\) (the effect of the treatment itself), \(Z \to A\) (the degree of adherence), and \(Z \to Y\) (concurrent behavioral changes).
An arrow \(Z \to Y\) means the exclusion restriction does not hold (see Technical Point 16.1 and Chapter 16). Investigators often try to remove that arrow by blinding: those assigned \(Z=1\) get the vaccine and those assigned \(Z=0\) get an identical placebo injection, so neither participants nor their doctors know the assignment (a double-blind placebo-controlled trial). Blinding is often infeasible (no convincing placebo exists for open heart surgery; side effects reveal who is treated), and it is not advisable when the goal is the effect of treatment in the real world, where there is no blinding or placebo (Hernán and Robins 2020, Fine Point 22.1, p. 306).
1.1 The Intention-to-Treat Effect
Because \(Z\) is randomized, there are no backdoor paths from \(Z\) to \(Y\), so \(Y^z \perp\!\!\!\perp Z\).
Proof. For each \(z\), \(\Pr[Y = 1 \mid Z = z] = \Pr[Y^z = 1 \mid Z = z]\) by consistency, and \(\Pr[Y^z = 1 \mid Z = z] = \Pr[Y^z = 1]\) by \(Y^z \perp\!\!\!\perp Z\). This proves the per-arm equality. When \(\Pr[Y = 1 \mid Z = 0] > 0\), the per-arm equality makes the two denominators equal and nonzero, so taking the ratio of the \(z = 1\) and \(z = 0\) expressions gives the risk ratio equality.
An ITT analysis (Definition 2) is unbiased because it includes all randomized individuals; variations that include only a subset may be biased.
- Pseudo-intention-to-treat analysis: with loss to follow-up, the analysis is restricted to the uncensored, \(\Pr[Y = 1 \mid Z = 1, C = 0] / \Pr[Y = 1 \mid Z = 0, C = 0]\). Censoring can induce selection bias (Chapter 8) in either direction, so adjustment for selection bias is needed (Section 21.5).
- Modified intention-to-treat analysis: limited to those who started their assigned strategy at least once (for instance, took one or more pills). It usually needs adjustment for the risk factors of adherence.
1.2 The Per-Protocol Effect
Unlike the ITT effect, the PP effect is generally confounded.
- As-treated analysis: compares \(A = 1\) with \(A = 0\) regardless of \(Z\). It is confounded by unmeasured \(U\) (Figures 22.1 and 22.2); if measured factors \(L\) block all backdoor paths (Figure 22.3), it must adjust for \(L\).
- Conventional per-protocol (on-treatment) analysis: an ITT analysis (Definition 2) restricted to the “per-protocol population” with \(A = Z\). With selection indicator \(S\) (\(S = 1\) if \(A = Z\)), conditioning on \(S = 1\) opens the noncausal path \(Z \to A \leftarrow L \leftarrow U \to Y\) (Figure 22.4), so the analysis is biased unless it measures and adjusts for \(L\).
Both are observational analyses of a randomized experiment and require adjustment for confounding and selection bias (Hernán and Robins 2020, Fine Point 22.3, p. 308).
1.3 Two Justifications for the ITT Effect, Revisited
The ITT effect is privileged largely because it is unconfounded, not because it is the effect we want. Two common justifications deserve “a grain of salt” (Hernán and Robins 2020, 307):
- It preserves the null. Under the sharp causal null and the exclusion restriction, \(\Pr[Y = 1 \mid Z = 1] / \Pr[Y = 1 \mid Z = 0] = \Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1] = 1\). Without the exclusion restriction (no double-blind placebo control), the effect of \(A\) can be null while the effect of \(Z\) is not: erase \(A \to Y\) in Figure 22.1 and \(Z \to Y\) remains.
- It is conservative (between 1 and the PP risk ratio). This holds only if non-adherence attenuates the effect, which is not guaranteed. Even when it holds, a near-null ITT effect on an adverse outcome can wrongly suggest a harmful treatment is safe, because many assigned to \(Z = 1\) stopped treatment before the harm occurred.
“ITT measures effectiveness in the real world; PP measures efficacy.” This reasoning is problematic because:
- the ITT effect reflects adherence in that trial, which may differ from real life (close monitoring; adherence may rise once a treatment is shown to work);
- if real-world effectiveness were the goal, we should not run double-blind placebo-controlled trials, which remove the effects of assignment awareness that exist in practice;
- people who plan to adhere to their prescribed treatment care more about the PP effect.
“ITT is always conservative.” Not if the effect is non-monotonic (Technical Point 5.2) and non-adherence is high. Even for monotonic effects, it can fail in head-to-head trials: in a trial of an expensive drug (\(Z = 1\)) versus ibuprofen (\(Z = 0\)) for severe pain at 1 year, both drugs are equally effective (PP risk ratio 1), but adherence to ibuprofen is lower because of an easily palliated side effect. The ITT comparison (Definition 2) then wrongly suggests ibuprofen is less effective (Hernán and Robins 2020, Fine Point 22.4, p. 309).
2 22.2 A Target Trial with Sustained Treatment Strategies (pp. 309-313)
Because the goal is to emulate target trials with real-world data, we consider pragmatic trials with these features:
- treatment assignment is not blinded;
- nobody receives a placebo (strategies involve active treatments or no treatment);
- participants are monitored as often and as intensely as regular patients.
Unlike earlier chapters, where the outcome \(Y\) was measured at the end of follow-up, the outcome here is a failure time, time to death (see Technical Point 21.10).
2.1 ITT and PP Effects for Sustained Strategies
The ITT effect is agnostic about post-baseline protocol deviations (stopping treatment for no clinical reason, starting it in the \(g_0\) arm, non-approved concomitant treatments). Its magnitude can therefore depend heavily on the deviation patterns in each trial: two trials with the same protocol in different settings can have different ITT effects, and neither is biased. This limitation is why the ITT effect should be complemented with the PP effect.
2.2 Per-Protocol Strategies Are Usually Dynamic
An individual assigned to \(g_1\) who stops therapy because of toxicity is adhering to \(g_1\), not deviating from it, even if the protocol describes \(g_1\) loosely as “treat continuously.”
Ideally the protocol fully specifies the strategies, so that the per-protocol effect is well defined (Hernán and Robins, 2017, as cited in the book).
The PP effect is often the implicit target of inference. When investigators complain that the interventions implemented in a trial were not faithful to the protocol and call that “bias,” they are really interested in the PP effect: non-adherence after baseline cannot bias the effect of baseline assignment.
2.3 Per-Protocol Effects in Alternative Target Trials
Estimating per-protocol effects of sustained strategies in a trial raises the same issues as in an observational study: trial investigators need to collect post-randomization data on adherence and on time-varying prognostic factors associated with adherence.
The controlled direct effect of \(A\) on \(Y\) with mediator \(M\) set to \(m\) is \(\operatorname{E}\mathopen{}\left[Y^{a=1,m}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,m}\right]\mathclose{}\), for \(m = 0\) or \(m = 1\). It could be identified by a trial that randomizes \(A\) at baseline and \(M\) one month later, so that \(\Pr[Y^{a,m} = 1] = \Pr[Y = 1 \mid A = a, M = m]\), or by emulating such a trial when consistency, positivity, and exchangeability hold for both \(A\) and \(M\). It is just a contrast of sustained strategies: replace \(A\) and \(M\) by \(A_0\) and \(A_1\) (Chapter 19) (Hernán and Robins 2020, Technical Point 22.1, p. 311).
- Pure (natural) direct effect: \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{}\). It is a cross-world quantity, so it cannot be identified from any randomized experiment on \(A\), \(M\), or both, nor from observational data under an FFRCISTG model (Technical Point 6.2). Introduced by Robins and Greenland (1992); Pearl (2001) renamed it and showed it is identified for certain graphs under the NPSEM-IE model, which assumes untestable cross-world independencies.
- Principal stratum direct effect: the effect of \(A\) in the subset with \(M^{a=0} = M^{a=1} = m\). It equals \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid M^{a=0} = M^{a=1} = m\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid M^{a=0} = M^{a=1} = m\right]\mathclose{}\), a total effect in a subpopulation, so it needs no well-defined intervention on \(M\); but it has little policy relevance when \(A\) affects \(M\) in almost everyone. Introduced by Robins (1986) and popularized by Rubin (2004).
Chapter 23 presents yet another type of direct effect (Hernán and Robins 2020, Technical Point 22.2, p. 312).
3 22.3 Emulating a Target Trial with Sustained Strategies (pp. 313-315)
When a pragmatic trial (Definition 4) is not possible, we emulate it with existing observational data; the trial is then the target trial of the observational analysis.
Specifying the protocol precisely may require some exploration of the available data. For example, a target trial of individuals with HIV is a reasonable proposal only after confirming that the data include information on HIV diagnosis.
3.1 Observational Analog of the ITT Effect
3.2 Observational Analog of the PP Effect
We can only emulate target trials whose strategies are actually followed by at least some individuals in the data, unless we are willing to extrapolate with models such as dose-response structural models.
4 22.4 Time Zero (pp. 315-317)
Follow-up in the emulation should start when it would have started in the target trial.
In our antiretroviral therapy trial, follow-up does not start 2 years before or after assignment:
- before: the strategies have not been assigned and the eligibility criteria have not been met, or even defined;
- after: deaths in the first 2 years are excluded and short-term effects are missed. Worse, if treatment has a short-term effect, more susceptible individuals would have died by year 2 in the active arm but not in the other, destroying baseline comparability and opening the door to selection bias.
The same rules apply to observational analyses, for the same reasons. Errors in emulating time zero are nevertheless frequent. For example, the discrepancy between observational and randomized estimates of the effect of postmenopausal hormone therapy on heart disease was partly due to mishandling time zero in the observational studies (Hernán et al., 2008, as cited in Hernán and Robins (2020, 315)).
Two problems cause errors in emulating time zero:
- there may be no unique choice of time zero;
- the treatment strategies may not be uniquely assignable at time zero.
4.1 Problem 1: Multiple Eligible Times
Options for time zero: (a) the first eligible time, (b) a random eligible time, or (c) every eligible time.
The number of sequential trials depends on how often treatment and covariates are measured:
- with a fixed data-collection schedule (e.g., every two years in many cohorts), emulate a new trial at each scheduled time;
- with subject-specific schedules (e.g., electronic medical records), choose a time unit (day, week, month) and emulate a new trial at each unit.
The choice of time unit matters: if treatment and confounders change more than once a week for many individuals, a week or month unit introduces bias that a daily unit could eliminate; without daily data the bias cannot be fully corrected.
Option (c) (Definition 9) can be more efficient because it uses more of the data, but individuals contribute to multiple trials, so the variance must be adjusted, e.g., by bootstrapping the entire analysis.
4.2 Problem 2: Data Compatible with Several Strategies
Therapy cannot be started on the very day it is assigned, so “immediate” initiation needs a grace period (say, 3 months) during which initiation still counts as immediate; otherwise the study would compare strategies that rarely occur or could not be implemented.
During the grace period an individual’s data are consistent with more than one strategy (e.g., someone who starts in month 3 is consistent with both “initiate within 3 months” and “never initiate” during months 1 and 2), so cloning and censoring are again used, with IP weighting to handle the censoring.
Consequences:
- the ITT effect cannot be estimated, because almost everyone contributes a clone to every strategy, so groups defined by baseline assignment have essentially identical outcomes; such analyses target some form of per-protocol effect and need adjustment;
- a well-defined strategy with a grace period should specify the timing of initiation within the grace period (Cain et al., 2010).
(Hernán and Robins 2020, Fine Point 22.5, p. 317)
4.3 Supplement: Immortal Time Bias
The book’s chapter does not use the term, but misaligned time zero classically produces a bias that has its own name.
5 22.5 A Unified Approach to Answer What If Questions with Data (pp. 317-322)
All health and social scientists face the same fundamental task: articulating causal questions as contrasts of well-defined counterfactuals. The target trial helps by specifying the well-defined interventions that lead to well-defined counterfactuals.
5.1 Trials and Observational Studies Differ Only by Baseline Randomization
In a randomized experiment:
- no baseline confounding is expected;
- the randomization probabilities are known;
- each individual’s assigned strategy is known at baseline.
An observational analysis can emulate (i) when a sufficient set of covariates is measured and adjusted for, and (ii) when the treatment model given the past is correctly specified. Feature (iii) is not needed for a per-protocol effect in either design, because efficient estimators ignore it: a trial’s assignment variable could be dropped from the data without losing the per-protocol effect, as long as a sufficient set of confounders had been measured. With dynamic strategies and full adherence, the covariates the strategies use to decide treatment form such a set (Robins, 1986) (Hernán and Robins 2020, Fine Point 22.6, p. 319).
5.2 Conventional Trial Analyses, Revisited
ITT analysis (Definition 2, an unadjusted comparison of randomized groups): randomization rules out baseline and post-randomization confounding for the effect of assignment, but not selection bias from loss to follow-up. Valid ITT estimation may need adjustment for time-varying prognostic factors, e.g., g-methods if dropout depends on symptom onset.
Conventional per-protocol analysis (censor at the first deviation, no adjustment) is questionable for three reasons:
- selection bias from differential loss to follow-up;
- those remaining on protocol in each arm need not be exchangeable, so g-methods are needed for time-varying factors that affect staying on protocol (or instrumental variable methods, with their own strong assumptions; Technical Point 16.6);
- it ignores that the strategies are dynamic: censoring people who stop treatment because of toxicity or a contraindication (the “on-treatment” analysis) treats adherence as deviation.
Fine Point 22.2’s unadjusted ITT analysis (Definition 2) is the pseudo-intention-to-treat analysis, and Fine Point 22.3’s unadjusted per-protocol analysis is the naive per-protocol analysis.
For failure-time outcomes, g-methods are always needed when treatment affects the outcome, because \(A_k\) affects all later variables through its effect on \(D_{k+1}\) (Technical Point 21.10).
The upshot: when the goal is a per-protocol effect or its observational analog, randomized trials and observational studies should be analyzed identically. Any reason to adjust for time-varying confounding and selection bias in an observational study is equally a reason to adjust for them in a randomized trial (Hernán and Robins 2020, chap. 22, p. 320).
5.3 Why Observational Emulation Matters
Randomized trials may be expensive, infeasible, unethical, or too slow for an urgent decision, so many decisions must be made without them.
When we cannot run the trial that would answer our question, the observational analysis should explicitly emulate it and be judged by how well it emulates its target trial.
In very unusual situations, a decision informed by a well-conducted randomized trial can be worse than one informed by badly confounded observational data (Fine Points 22.7 and 22.8).
Further reading (listed in the book’s references): Hernán and Robins (2016), “Using big data to emulate a target trial when a randomized trial is not available,” American Journal of Epidemiology; Hernán and Robins (2017), “Per-protocol analyses of pragmatic trials,” New England Journal of Medicine.
A double-blind placebo-controlled trial of an over-the-counter treatment \(A\) enrolled a random 20% of people diagnosed with lung cancer; everyone adhered. The 60-month mortality was 550/1000 = 55% with \(A = 1\) and 450/1000 = 45% with \(A = 0\), so the regulator banned \(A\). An observational study of the other 80% found 0% mortality among both treated and untreated.
Classify individuals into types: doomed (\(Y^{a=0} = Y^{a=1} = 1\)), hurt (\(Y^{a=0} = 0\), \(Y^{a=1} = 1\)), helped (\(Y^{a=0} = 1\), \(Y^{a=1} = 0\)), immune (\(Y^{a=0} = Y^{a=1} = 0\)). Random sampling and randomization give the trial arms and the observational sample the same distribution of types. Then:
- 0% observational mortality means nobody is doomed;
- the trial’s treated arm gives \(\Pr[Y^{a=1} = 1] = \Pr[\text{doomed}] + \Pr[\text{hurt}] = 0 + \Pr[\text{hurt}] = 0.55\);
- the untreated arm gives \(\Pr[Y^{a=0} = 1] = \Pr[\text{doomed}] + \Pr[\text{helped}] = 0 + \Pr[\text{helped}] = 0.45\);
- so \(\Pr[\text{immune}] = 1 - 0 - 0.55 - 0.45 = 0\).
In the observational data every “hurt” person took \(A = 0\) and every “helped” person took \(A = 1\): everyone followed the optimal strategy. The trial compared “treat everyone” with “treat no one,” but the best strategy was “treat only those who benefit.” If, say, the type were determined by ethnic group and each group had learned from experience whether to take \(A\) (maximal effect modification and maximal confounding), the confounded observational study, not the unconfounded trial, revealed the correct policy (Hernán and Robins 2020, Fine Point 22.7, p. 321).
Suppose lower \(Y\) is better and the observational mean \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\) is below the mean of both trial arms, so \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\). If \(U\) is a (possibly unknown) set of pre-treatment covariates with \(Y^a \perp\!\!\!\perp A \mid U\), then \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) for the random strategy \(g\) that assigns \(A = 1\) with probability \(\Pr[A = 1 \mid U]\), the strategy that generated the observational data. That strategy cannot be implemented without data on \(U\), but it can motivate measuring pre-treatment covariates \(V\) and using the trial data to find a deterministic dynamic strategy \(g^*\) whose mean, estimated from the trial, is below the observational \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\). Combining trial and observational data can thus be more informative than the trial alone, provided both are random samples of everyone eligible for the trial (Hernán and Robins 2020, Fine Point 22.8, p. 322).
6 Summary
- The ITT effect is the effect of randomized assignment; it is unconfounded but is neither guaranteed to preserve the null (without the exclusion restriction) nor guaranteed to be conservative.
- The PP effect is the effect of adhering to the assigned strategies; it is often the more relevant estimand and generally requires adjustment, for sustained strategies via g-methods, in trials and observational studies alike.
- PP strategies are usually dynamic: stopping for toxicity is adherence, not deviation.
- Non-adherence in a trial allows emulation of alternative target trials.
- A target trial protocol specifies eligibility, follow-up, strategies, assignment procedures, outcomes, causal contrast, and analysis plan; observational analogs are the comparison of initiators (ITT) and the PP effect.
- Time zero must align eligibility, assignment, and start of follow-up; multiple eligible times call for sequential trials, and data compatible with several strategies call for cloning, censoring, and weighting.
- Aside from baseline randomization, randomized and observational studies of sustained strategies should be analyzed the same way.