Chapter 22: Target Trial Emulation

Published

Last modified: 2026-10-09 13:46:40 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~0% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Part I described causal inference from observational data as an attempt to emulate a hypothetical randomized trial, the target trial, but only for simple target trials comparing time-fixed treatments. With the g-methods of Chapters 19-21 in hand, we can now specify realistic target trials that compare sustained treatment strategies, and emulate them with either randomized or observational data.

This chapter is based on Hernán and Robins (2020, chap. 22, pp. 305-322).

The chapter does three things:

  • it builds a taxonomy of causal effects in trials: the intention-to-treat effect, the per-protocol effect, and per-protocol effects in alternative target trials;
  • it defines observational analogs of those effects and describes how to emulate a target trial, including the choice of time zero;
  • it argues that, aside from baseline randomization, randomized trials and observational studies of sustained strategies should be analyzed the same way.

What makes this more than a formal exercise is the existence of g-methods: if data on all important fixed and time-varying confounders are available, the effects of interest can be validly estimated.

1 22.1 Intention-to-Treat Effect and Per-Protocol Effect (pp. 305-309)


Example 1 (A trial of vaccination plus an antiviral) Consider a randomized trial of a virus that can kill:

  • \(Z\): assigned treatment (1: immediate vaccination plus an experimental antiviral if infected; 0: standard of care, with neither);
  • \(A\): received treatment (vaccinated or not);
  • \(Y\): death;
  • \(U\): unmeasured risk factors that influence the decision to get vaccinated.

\(Z\) and \(A\) can differ: some people assigned to vaccine refuse it, and some assigned to no vaccine get vaccinated outside the study.

In the book’s Figure 22.1, \(Z\) affects \(Y\) through two pathways:

  • \(Z \to A \to Y\): assignment changes the treatment received, which in turn affects mortality;
  • \(Z \to Y\) directly: knowing one’s assignment can change behavior. For example, people who know they were assigned to vaccination plus a promising antiviral may become less careful about avoiding infection.

The effect of \(Z\) therefore depends on the strength of three arrows: \(A \to Y\) (the effect of the treatment itself), \(Z \to A\) (the degree of adherence), and \(Z \to Y\) (concurrent behavioral changes).

NoteFine Point 22.1: The Exclusion Restriction (Again)

An arrow \(Z \to Y\) means the exclusion restriction does not hold (see Technical Point 16.1 and Chapter 16). Investigators often try to remove that arrow by blinding: those assigned \(Z=1\) get the vaccine and those assigned \(Z=0\) get an identical placebo injection, so neither participants nor their doctors know the assignment (a double-blind placebo-controlled trial). Blinding is often infeasible (no convincing placebo exists for open heart surgery; side effects reveal who is treated), and it is not advisable when the goal is the effect of treatment in the real world, where there is no blinding or placebo (Hernán and Robins 2020, Fine Point 22.1, p. 306).

1.1 The Intention-to-Treat Effect


Definition 1 (Intention-to-Treat (ITT) Effect) The intention-to-treat effect is the causal effect of randomized assignment \(Z\), for example the causal risk ratio \[\frac{\Pr[Y^{z=1} = 1]}{\Pr[Y^{z=0} = 1]}.\] It is “the effect of having the intention of treating with \(A\),” not “the effect of treating with \(A\)” (Hernán and Robins 2020, 306).

Because \(Z\) is randomized, there are no backdoor paths from \(Z\) to \(Y\), so \(Y^z \perp\!\!\!\perp Z\).

Proposition 1 (The ITT effect equals the association between assignment and outcome) Suppose \(Z\) is randomized, so that \(Y^z \perp\!\!\!\perp Z\) for \(z = 0, 1\), that consistency holds for \(Z\) (\(Y = Y^z\) whenever \(Z = z\)), and that \(\Pr[Z = z] > 0\) for \(z = 0, 1\). Then, for \(z = 0, 1\), \[\Pr[Y = 1 \mid Z = z] = \Pr[Y^z = 1].\] Hence, whenever \(\Pr[Y = 1 \mid Z = 0] > 0\) (equivalently, \(\Pr[Y^{z=0} = 1] > 0\)), the associational risk ratio equals the ITT risk ratio: \[\frac{\Pr[Y = 1 \mid Z = 1]}{\Pr[Y = 1 \mid Z = 0]} = \frac{\Pr[Y^{z=1} = 1]}{\Pr[Y^{z=0} = 1]}.\]

Proof. For each \(z\), \(\Pr[Y = 1 \mid Z = z] = \Pr[Y^z = 1 \mid Z = z]\) by consistency, and \(\Pr[Y^z = 1 \mid Z = z] = \Pr[Y^z = 1]\) by \(Y^z \perp\!\!\!\perp Z\). This proves the per-arm equality. When \(\Pr[Y = 1 \mid Z = 0] > 0\), the per-arm equality makes the two denominators equal and nonzero, so taking the ratio of the \(z = 1\) and \(z = 0\) expressions gives the risk ratio equality.

Definition 2 (Intention-to-treat analysis) Estimating the ITT effect by the unadjusted associational risk ratio \(\Pr[Y = 1 \mid Z = 1] / \Pr[Y = 1 \mid Z = 0]\) (Proposition 1), or another unadjusted contrast of the randomized groups, is an intention-to-treat analysis.

NoteFine Point 22.2: Pseudo- and Modified Intention-to-Treat Analyses

An ITT analysis (Definition 2) is unbiased because it includes all randomized individuals; variations that include only a subset may be biased.

  • Pseudo-intention-to-treat analysis: with loss to follow-up, the analysis is restricted to the uncensored, \(\Pr[Y = 1 \mid Z = 1, C = 0] / \Pr[Y = 1 \mid Z = 0, C = 0]\). Censoring can induce selection bias (Chapter 8) in either direction, so adjustment for selection bias is needed (Section 21.5).
  • Modified intention-to-treat analysis: limited to those who started their assigned strategy at least once (for instance, took one or more pills). It usually needs adjustment for the risk factors of adherence.

1.2 The Per-Protocol Effect


Definition 3 (Per-Protocol (PP) Effect) The per-protocol effect is the causal effect under full adherence, that is, if everyone had followed the protocol’s instructions for the treatment they were assigned: the contrast of \(\Pr[Y^{z=1,a=1} = 1]\) versus \(\Pr[Y^{z=0,a=0} = 1]\), or, under the exclusion restriction, of \(\Pr[Y^{a=1} = 1]\) versus \(\Pr[Y^{a=0} = 1]\).

Unlike the ITT effect, the PP effect is generally confounded.

Example 2 (Confounding of the per-protocol effect) Suppose \(U\) is high risk of infection, and high-risk people assigned \(Z=0\) seek vaccination outside the study. Then the backdoor path \(A \leftarrow U \rightarrow Y\) makes \(\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]\) differ from \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\).

Remark 1 (A trial viewed as an observational study). Estimating the PP effect requires viewing the trial as an observational study: adjustment under conditional exchangeability given measured covariates, or alternative assumptions such as those of instrumental variable estimation (Chapter 16).

NoteFine Point 22.3: Naive Per-Protocol Analyses
  • As-treated analysis: compares \(A = 1\) with \(A = 0\) regardless of \(Z\). It is confounded by unmeasured \(U\) (Figures 22.1 and 22.2); if measured factors \(L\) block all backdoor paths (Figure 22.3), it must adjust for \(L\).
  • Conventional per-protocol (on-treatment) analysis: an ITT analysis (Definition 2) restricted to the “per-protocol population” with \(A = Z\). With selection indicator \(S\) (\(S = 1\) if \(A = Z\)), conditioning on \(S = 1\) opens the noncausal path \(Z \to A \leftarrow L \leftarrow U \to Y\) (Figure 22.4), so the analysis is biased unless it measures and adjusts for \(L\).

Both are observational analyses of a randomized experiment and require adjustment for confounding and selection bias (Hernán and Robins 2020, Fine Point 22.3, p. 308).

1.3 Two Justifications for the ITT Effect, Revisited


WarningThe ITT Effect Need Not Preserve the Null or Be Conservative

The ITT effect is privileged largely because it is unconfounded, not because it is the effect we want. Two common justifications deserve “a grain of salt” (Hernán and Robins 2020, 307):

  1. It preserves the null. Under the sharp causal null and the exclusion restriction, \(\Pr[Y = 1 \mid Z = 1] / \Pr[Y = 1 \mid Z = 0] = \Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1] = 1\). Without the exclusion restriction (no double-blind placebo control), the effect of \(A\) can be null while the effect of \(Z\) is not: erase \(A \to Y\) in Figure 22.1 and \(Z \to Y\) remains.
  2. It is conservative (between 1 and the PP risk ratio). This holds only if non-adherence attenuates the effect, which is not guaranteed. Even when it holds, a near-null ITT effect on an adverse outcome can wrongly suggest a harmful treatment is safe, because many assigned to \(Z = 1\) stopped treatment before the harm occurred.

Remark 2 (Limits of exclusive reliance on the ITT effect). The argument against conservative ITT analyses (Definition 2) also applies to non-inferiority trials, whose goal is to show that one treatment is not inferior to another.

The book’s conclusion: exclusive reliance on ITT estimates is hard to justify for trials with substantial non-adherence and for trials of harms. The PP effect is often the more natural estimand for clinicians, patients, and other decision makers.

With sustained strategies, the probability of non-adherence increases greatly, the ITT effect becomes increasingly uninformative relative to the PP effect, and estimating the PP effect, in a trial or in an observational emulation, generally requires g-methods.

NoteFine Point 22.4: More Misunderstandings About the ITT Effect

“ITT measures effectiveness in the real world; PP measures efficacy.” This reasoning is problematic because:

  • the ITT effect reflects adherence in that trial, which may differ from real life (close monitoring; adherence may rise once a treatment is shown to work);
  • if real-world effectiveness were the goal, we should not run double-blind placebo-controlled trials, which remove the effects of assignment awareness that exist in practice;
  • people who plan to adhere to their prescribed treatment care more about the PP effect.

“ITT is always conservative.” Not if the effect is non-monotonic (Technical Point 5.2) and non-adherence is high. Even for monotonic effects, it can fail in head-to-head trials: in a trial of an expensive drug (\(Z = 1\)) versus ibuprofen (\(Z = 0\)) for severe pain at 1 year, both drugs are equally effective (PP risk ratio 1), but adherence to ibuprofen is lower because of an easily palliated side effect. The ITT comparison (Definition 2) then wrongly suggests ibuprofen is less effective (Hernán and Robins 2020, Fine Point 22.4, p. 309).

2 22.2 A Target Trial with Sustained Treatment Strategies (pp. 309-313)


Definition 4 (Pragmatic trial) A pragmatic trial is a trial designed to estimate the effect of treatment strategies under conditions close to those of routine clinical care.

Because the goal is to emulate target trials with real-world data, we consider pragmatic trials with these features:

  • treatment assignment is not blinded;
  • nobody receives a placebo (strategies involve active treatments or no treatment);
  • participants are monitored as often and as intensely as regular patients.

Example 3 (A Target Trial of Antiretroviral Therapy)  

Component Specification
Eligibility criteria HIV infection, age 18 or older, no AIDS, no previous antiretroviral therapy
Treatment strategies \(g_1\): receive therapy (\(A_k = 1\)) continuously unless a contraindication or toxicity arises; \(g_0\): receive no therapy (\(A_k = 0\)) continuously
Assignment Random, at baseline \(k = 0\); \(Z = 1\) if assigned to \(g_1\), \(Z = 0\) if assigned to \(g_0\)
Follow-up From assignment until death, loss to follow-up, or 60 months, whichever comes first
Outcome Death; \(D_k = 1\) if dead by month \(k\)
Causal contrasts Intention-to-treat and per-protocol effects on the risk of death

Here \(k = 0, 1, \ldots, K\) with \(K = 59\), and \(C_k\) indicates censoring by month \(k\), for \(k = 1, \ldots, K + 1\) (Hernán and Robins 2020, 309–10).

Unlike earlier chapters, where the outcome \(Y\) was measured at the end of follow-up, the outcome here is a failure time, time to death (see Technical Point 21.10).

2.1 ITT and PP Effects for Sustained Strategies


Definition 5 (ITT and PP effects for sustained strategies) With the strategies and notation of Example 3:

ITT effect at time \(k\): a contrast of the static strategies “be assigned to \(g_1\) (or \(g_0\)) at baseline, with no loss to follow-up”: \[\Pr\mathopen{}\left[D_k^{z=1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{z=0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]

PP effect at time \(k\): a contrast of “receive strategy \(g_1\) (or \(g_0\)) continuously between baseline and end of follow-up”: \[\Pr\mathopen{}\left[D_k^{g_1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{g_0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]

Both are defined as if nobody had been lost to follow-up through time \(k\) (\(\bar{c}_k = \bar{0}\)).

Remark 3 (The ITT effect as an effect of initiation). When assignment and initiation always coincide (everyone assigned to \(g_1\) starts treatment at time 0 and nobody assigned to \(g_0\) does, whatever happens later), the ITT effect is also the effect of initiation: \[\Pr\mathopen{}\left[D_k^{a_0=1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{a_0=0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]

WarningITT Effects Depend on Deviation Patterns

The ITT effect is agnostic about post-baseline protocol deviations (stopping treatment for no clinical reason, starting it in the \(g_0\) arm, non-approved concomitant treatments). Its magnitude can therefore depend heavily on the deviation patterns in each trial: two trials with the same protocol in different settings can have different ITT effects, and neither is biased. This limitation is why the ITT effect should be complemented with the PP effect.

2.2 Per-Protocol Strategies Are Usually Dynamic


Remark 4 (Dynamic strategies). Sensible protocols do not mandate treatment no matter what: \(g_1\) requires stopping therapy when a contraindication or toxicity arises. So the PP effect generally compares dynamic strategies (“do this; if X happens, do this other thing”).

WarningStopping for Toxicity Is Adherence

An individual assigned to \(g_1\) who stops therapy because of toxicity is adhering to \(g_1\), not deviating from it, even if the protocol describes \(g_1\) loosely as “treat continuously.”

TipSpecify the Strategies Fully

Ideally the protocol fully specifies the strategies, so that the per-protocol effect is well defined (Hernán and Robins, 2017, as cited in the book).

The PP effect is often the implicit target of inference. When investigators complain that the interventions implemented in a trial were not faithful to the protocol and call that “bias,” they are really interested in the PP effect: non-adherence after baseline cannot bias the effect of baseline assignment.

2.3 Per-Protocol Effects in Alternative Target Trials


Example 4 (An alternative target trial) Suppose that during the trial a consensus emerges that \(g_0\) is inferior, and physicians start treating \(g_0\) participants once their CD4 count (\(L_k\)) first drops below 200 cells/µL. Many \(g_0\) participants then follow

\(g_0'\): “receive \(A_k = 0\) continuously, but switch to \(A_k = 1\) after \(L_k < 200\).”

The contrast of \(g_1\) versus \(g_0'\) is neither the ITT nor the original PP effect. It is the PP effect of another target trial, randomizing \(g_1\) versus \(g_0'\), which can be emulated with the actual trial’s data.

Remark 5 (Non-adherence enables alternative target trials). Paradoxically, complete adherence would make this impossible: if everyone followed the original protocol, nobody in the data followed \(g_0'\). A fully adherent trial of CD4 thresholds for initiating therapy is of little use to emulate a trial of continuous treatment versus no treatment, and vice versa. It is precisely non-adherence that lets one trial’s data answer other, perhaps more relevant, causal questions.

TipCollect Post-Randomization Data

Estimating per-protocol effects of sustained strategies in a trial raises the same issues as in an observational study: trial investigators need to collect post-randomization data on adherence and on time-varying prognostic factors associated with adherence.

NoteTechnical Point 22.1: Controlled Direct Effects

The controlled direct effect of \(A\) on \(Y\) with mediator \(M\) set to \(m\) is \(\operatorname{E}\mathopen{}\left[Y^{a=1,m}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0,m}\right]\mathclose{}\), for \(m = 0\) or \(m = 1\). It could be identified by a trial that randomizes \(A\) at baseline and \(M\) one month later, so that \(\Pr[Y^{a,m} = 1] = \Pr[Y = 1 \mid A = a, M = m]\), or by emulating such a trial when consistency, positivity, and exchangeability hold for both \(A\) and \(M\). It is just a contrast of sustained strategies: replace \(A\) and \(M\) by \(A_0\) and \(A_1\) (Chapter 19) (Hernán and Robins 2020, Technical Point 22.1, p. 311).

NoteTechnical Point 22.2: Pure and Principal Stratum Direct Effects
  • Pure (natural) direct effect: \(\operatorname{E}\mathopen{}\left[Y^{a=1, M^{a=0}}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0, M^{a=0}}\right]\mathclose{}\). It is a cross-world quantity, so it cannot be identified from any randomized experiment on \(A\), \(M\), or both, nor from observational data under an FFRCISTG model (Technical Point 6.2). Introduced by Robins and Greenland (1992); Pearl (2001) renamed it and showed it is identified for certain graphs under the NPSEM-IE model, which assumes untestable cross-world independencies.
  • Principal stratum direct effect: the effect of \(A\) in the subset with \(M^{a=0} = M^{a=1} = m\). It equals \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid M^{a=0} = M^{a=1} = m\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid M^{a=0} = M^{a=1} = m\right]\mathclose{}\), a total effect in a subpopulation, so it needs no well-defined intervention on \(M\); but it has little policy relevance when \(A\) affects \(M\) in almost everyone. Introduced by Robins (1986) and popularized by Rubin (2004).

Chapter 23 presents yet another type of direct effect (Hernán and Robins 2020, Technical Point 22.2, p. 312).

3 22.3 Emulating a Target Trial with Sustained Strategies (pp. 313-315)


When a pragmatic trial (Definition 4) is not possible, we emulate it with existing observational data; the trial is then the target trial of the observational analysis.

Definition 6 (Target trial protocol) At a minimum, the protocol of a target trial specifies:

  • eligibility criteria;
  • start and end of follow-up;
  • treatment strategies;
  • assignment procedures;
  • outcomes of interest;
  • causal contrast;
  • data analysis plan.
TipExplore the Data Before Fixing the Protocol

Specifying the protocol precisely may require some exploration of the available data. For example, a target trial of individuals with HIV is a reasonable proposal only after confirming that the data include information on HIV diagnosis.

3.1 Observational Analog of the ITT Effect


Definition 7 (Observational analog of the ITT effect) The actual assignment is unknown in existing observational data, so a true ITT effect can rarely be emulated. The closest analog compares initiators of the strategies, “initiate \(A_0 = 1\) (or \(A_0 = 0\)) at baseline, with no loss to follow-up”: \[\Pr\mathopen{}\left[D_k^{a_0=1, \bar{c}_k = \bar{0}} = 1\right]\mathclose{} - \Pr\mathopen{}\left[D_k^{a_0=0, \bar{c}_k = \bar{0}} = 1\right]\mathclose{}.\]

Remark 6 (Grouping by initiation instead of assignment). This contrast groups people by initiation rather than assignment. Used in a randomized trial, it would put everyone who took no dose at baseline in the same group, whether assigned to \(g_1\) or \(g_0\). It is therefore equivalent to the modified ITT analysis of Fine Point 22.2. If initiation occurs shortly after assignment, it roughly preserves a key ITT feature: the contrast refers to interventions taken soon after baseline.

With prescription (rather than dispensing) data, a comparison by whether therapy was prescribed at baseline would come a little closer to the target trial’s ITT analysis (Definition 2).

3.2 Observational Analog of the PP Effect


Remark 7 (Observational analog of the PP effect). The observational analog of the PP effect is defined exactly as in the target trial. Without a pre-specified protocol, every per-protocol effect corresponds to some target trial, so there is no distinction between the “original” PP effect and PP effects of alternative trials.

WarningOnly Strategies Followed in the Data Can Be Emulated

We can only emulate target trials whose strategies are actually followed by at least some individuals in the data, unless we are willing to extrapolate with models such as dose-response structural models.

Remark 8 (Why explicit strategies help).

  • they prevent bias, by exposing analyses whose comparisons cannot be translated into a contrast between hypothetical interventions (Sections 3.5 and 3.6);
  • they add clarity. The labels “efficacy” (effect under perfect conditions) and “effectiveness” (effect under realistic conditions) are ambiguous: the ITT effect of a trial is sometimes called effectiveness and sometimes efficacy, and the PP effect of an emulation is sometimes called effectiveness. It is more helpful to accept that all causal effects lie somewhere on an effectiveness continuum, and to define explicitly the strategies that decision makers can choose between.

4 22.4 Time Zero (pp. 315-317)


Definition 8 (Time Zero) Time zero (baseline, start of follow-up) is the time at which eligibility criteria must be met (but not later) and after which outcomes begin to be counted (but not earlier). In a randomized trial, it is the time an eligible individual is assigned to a strategy.

TipStart Follow-Up as the Target Trial Would

Follow-up in the emulation should start when it would have started in the target trial.

WarningMisplaced Time Zero

In our antiretroviral therapy trial, follow-up does not start 2 years before or after assignment:

  • before: the strategies have not been assigned and the eligibility criteria have not been met, or even defined;
  • after: deaths in the first 2 years are excluded and short-term effects are missed. Worse, if treatment has a short-term effect, more susceptible individuals would have died by year 2 in the active arm but not in the other, destroying baseline comparability and opening the door to selection bias.

The same rules apply to observational analyses, for the same reasons. Errors in emulating time zero are nevertheless frequent. For example, the discrepancy between observational and randomized estimates of the effect of postmenopausal hormone therapy on heart disease was partly due to mishandling time zero in the observational studies (Hernán et al., 2008, as cited in Hernán and Robins (2020, 315)).

WarningTwo Sources of Time Zero Errors

Two problems cause errors in emulating time zero:

  1. there may be no unique choice of time zero;
  2. the treatment strategies may not be uniquely assignable at time zero.

4.1 Problem 1: Multiple Eligible Times


Example 5 (Eligible once versus eligible many times)  

  • Eligible once: follow-up starts at the only eligible time. Example: comparing initiation of therapy when CD4 first drops below 500 cells/µL versus below 350 cells/µL; follow-up starts when CD4 first drops below 500.
  • Eligible many times: e.g., initiation versus no initiation of hormone therapy in postmenopausal women with no chronic disease and no hormone therapy in the previous two years. A woman eligible continuously from age 51 to 65 could start follow-up at 51, 52, 53, …

Options for time zero: (a) the first eligible time, (b) a random eligible time, or (c) every eligible time.

Definition 9 (Sequential target trials) Emulating sequential target trials means starting a new emulated trial at every time an individual meets the eligibility criteria, as in option (c).

TipChoosing the Time Unit for Sequential Trials

The number of sequential trials depends on how often treatment and covariates are measured:

  • with a fixed data-collection schedule (e.g., every two years in many cohorts), emulate a new trial at each scheduled time;
  • with subject-specific schedules (e.g., electronic medical records), choose a time unit (day, week, month) and emulate a new trial at each unit.

The choice of time unit matters: if treatment and confounders change more than once a week for many individuals, a week or month unit introduces bias that a daily unit could eliminate; without daily data the bias cannot be fully corrected.

Option (c) (Definition 9) can be more efficient because it uses more of the data, but individuals contribute to multiple trials, so the variance must be adjusted, e.g., by bootstrapping the entire analysis.

4.2 Problem 2: Data Compatible with Several Strategies


Example 6 (Data compatible with several strategies) Target trial: individuals whose CD4 count just dropped below 500 cells/µL are assigned to start therapy

  1. immediately,
  2. when CD4 drops below 350, or
  3. when CD4 drops below 200.

Those who started at time zero follow strategy 1, but those who did not are compatible with both 2 and 3.

Algorithm 1 (Cloning, censoring, and weighting) Copy each individual whose data are compatible with more than one strategy (as in Example 6) into one clone per compatible strategy, and censor each clone when its data stop being consistent with its strategy. The likely informative censoring is corrected by IP weighting for time-varying factors.

Remark 9 (Why clone instead of randomly assigning). Assigning these individuals to one strategy at random would be statistically inefficient.

If an individual dies before either clone is censored, the death is counted under both strategies. This double allocation prevents the bias that would arise if events during the waiting period were systematically assigned to one strategy only.

Because individuals appear multiple times through their clones, the variance must again be adjusted by bootstrapping. Cloning, censoring, and weighting can be combined with sequential trial emulation when eligibility can be met at multiple times. See Robins et al. (2008) and Cain et al. (2010), as cited in the book; for related work, van der Laan and Petersen (2007).

NoteFine Point 22.5: Grace Periods

Therapy cannot be started on the very day it is assigned, so “immediate” initiation needs a grace period (say, 3 months) during which initiation still counts as immediate; otherwise the study would compare strategies that rarely occur or could not be implemented.

During the grace period an individual’s data are consistent with more than one strategy (e.g., someone who starts in month 3 is consistent with both “initiate within 3 months” and “never initiate” during months 1 and 2), so cloning and censoring are again used, with IP weighting to handle the censoring.

Consequences:

  • the ITT effect cannot be estimated, because almost everyone contributes a clone to every strategy, so groups defined by baseline assignment have essentially identical outcomes; such analyses target some form of per-protocol effect and need adjustment;
  • a well-defined strategy with a grace period should specify the timing of initiation within the grace period (Cain et al., 2010).

(Hernán and Robins 2020, Fine Point 22.5, p. 317)

4.3 Supplement: Immortal Time Bias


The book’s chapter does not use the term, but misaligned time zero classically produces a bias that has its own name.

Definition 10 (Immortal time) Immortal time is a period of follow-up during which, by the way groups are defined, the treated cannot experience the outcome. The bias that results from misclassifying such a period as treated person-time, or from excluding it from the analysis, is immortal time bias.

Example 7 (Immortal Time Bias: Statins and Mortality) Define “statin users” as people who filled a statin prescription at any time during a one-year window, and compare their mortality over that year with “non-users.” Anyone who dies before filling a prescription is automatically a non-user, and users must survive until their first fill. The time from the start of follow-up to the first fill is “immortal” for users (Definition 10), which produces a spurious survival advantage. Starting users’ follow-up at their first fill instead, while non-users are followed from the start of the window, still favors users: the early deaths of people who would have filled a prescription stay in the non-user group. Starting follow-up when eligibility is met and strategies are assigned, as the target trial requires, avoids this bias (see Hernán, Sauer, Hernández-Díaz, Platt, and Shrier, 2016, listed in the book’s references).

Remark 10 (New users). A related design point: eligibility criteria such as “no hormone therapy during the previous two years” restrict the target trial to new users. Section 20.5 explains why comparing current users with nonusers can be biased unless past treatment is adjusted for or the analysis is restricted to new users.

5 22.5 A Unified Approach to Answer What If Questions with Data (pp. 317-322)


Remark 11 (Target trial emulation as a unifying framework). Explicit target trial emulation brings together the two frameworks of this book, counterfactuals and causal diagrams, and grounds them in actionable causal inference:

  • organizing the analysis around a familiar concept, the experiment, helps articulate a well-defined causal question, from which design and analysis follow;
  • it applies across disciplines, whatever their vocabulary (economists’ “omitted variable bias” and “selection on observables” are confounding and conditional exchangeability);
  • it gives randomized and observational studies a common language.

All health and social scientists face the same fundamental task: articulating causal questions as contrasts of well-defined counterfactuals. The target trial helps by specifying the well-defined interventions that lead to well-defined counterfactuals.

5.1 Trials and Observational Studies Differ Only by Baseline Randomization


Remark 12 (Randomization is the only difference). A randomized trial is a follow-up study with baseline randomization; observational longitudinal data form a follow-up study without it. Long-term trials of sustained strategies in real-world settings, with imperfect adherence and loss to follow-up, suffer the confounding and selection biases usually associated with observational studies.

“Time-varying confounding in observational studies is a bias with the same structure as nonrandom noncompliance in randomized trials” (Hernán and Robins 2020, 318).

NoteFine Point 22.6: How Do Randomized and Observational Data Differ?

In a randomized experiment:

  1. no baseline confounding is expected;
  2. the randomization probabilities are known;
  3. each individual’s assigned strategy is known at baseline.

An observational analysis can emulate (i) when a sufficient set of covariates is measured and adjusted for, and (ii) when the treatment model given the past is correctly specified. Feature (iii) is not needed for a per-protocol effect in either design, because efficient estimators ignore it: a trial’s assignment variable could be dropped from the data without losing the per-protocol effect, as long as a sufficient set of confounders had been measured. With dynamic strategies and full adherence, the covariates the strategies use to decide treatment form such a set (Robins, 1986) (Hernán and Robins 2020, Fine Point 22.6, p. 319).

5.2 Conventional Trial Analyses, Revisited


WarningITT Analyses Can Suffer Selection Bias

ITT analysis (Definition 2, an unadjusted comparison of randomized groups): randomization rules out baseline and post-randomization confounding for the effect of assignment, but not selection bias from loss to follow-up. Valid ITT estimation may need adjustment for time-varying prognostic factors, e.g., g-methods if dropout depends on symptom onset.

WarningConventional Per-Protocol Analyses Are Questionable

Conventional per-protocol analysis (censor at the first deviation, no adjustment) is questionable for three reasons:

  1. selection bias from differential loss to follow-up;
  2. those remaining on protocol in each arm need not be exchangeable, so g-methods are needed for time-varying factors that affect staying on protocol (or instrumental variable methods, with their own strong assumptions; Technical Point 16.6);
  3. it ignores that the strategies are dynamic: censoring people who stop treatment because of toxicity or a contraindication (the “on-treatment” analysis) treats adherence as deviation.

Fine Point 22.2’s unadjusted ITT analysis (Definition 2) is the pseudo-intention-to-treat analysis, and Fine Point 22.3’s unadjusted per-protocol analysis is the naive per-protocol analysis.

WarningFailure-Time Outcomes Always Need G-Methods

For failure-time outcomes, g-methods are always needed when treatment affects the outcome, because \(A_k\) affects all later variables through its effect on \(D_{k+1}\) (Technical Point 21.10).

The upshot: when the goal is a per-protocol effect or its observational analog, randomized trials and observational studies should be analyzed identically. Any reason to adjust for time-varying confounding and selection bias in an observational study is equally a reason to adjust for them in a randomized trial (Hernán and Robins 2020, chap. 22, p. 320).

5.3 Why Observational Emulation Matters


Randomized trials may be expensive, infeasible, unethical, or too slow for an urgent decision, so many decisions must be made without them.

TipJudge an Observational Analysis by Its Emulation

When we cannot run the trial that would answer our question, the observational analysis should explicitly emulate it and be judged by how well it emulates its target trial.

In very unusual situations, a decision informed by a well-conducted randomized trial can be worse than one informed by badly confounded observational data (Fine Points 22.7 and 22.8).

Further reading (listed in the book’s references): Hernán and Robins (2016), “Using big data to emulate a target trial when a randomized trial is not available,” American Journal of Epidemiology; Hernán and Robins (2017), “Per-protocol analyses of pragmatic trials,” New England Journal of Medicine.

NoteFine Point 22.7: A Counterintuitive Comparison of a Trial and an Observational Study

A double-blind placebo-controlled trial of an over-the-counter treatment \(A\) enrolled a random 20% of people diagnosed with lung cancer; everyone adhered. The 60-month mortality was 550/1000 = 55% with \(A = 1\) and 450/1000 = 45% with \(A = 0\), so the regulator banned \(A\). An observational study of the other 80% found 0% mortality among both treated and untreated.

Classify individuals into types: doomed (\(Y^{a=0} = Y^{a=1} = 1\)), hurt (\(Y^{a=0} = 0\), \(Y^{a=1} = 1\)), helped (\(Y^{a=0} = 1\), \(Y^{a=1} = 0\)), immune (\(Y^{a=0} = Y^{a=1} = 0\)). Random sampling and randomization give the trial arms and the observational sample the same distribution of types. Then:

  • 0% observational mortality means nobody is doomed;
  • the trial’s treated arm gives \(\Pr[Y^{a=1} = 1] = \Pr[\text{doomed}] + \Pr[\text{hurt}] = 0 + \Pr[\text{hurt}] = 0.55\);
  • the untreated arm gives \(\Pr[Y^{a=0} = 1] = \Pr[\text{doomed}] + \Pr[\text{helped}] = 0 + \Pr[\text{helped}] = 0.45\);
  • so \(\Pr[\text{immune}] = 1 - 0 - 0.55 - 0.45 = 0\).

In the observational data every “hurt” person took \(A = 0\) and every “helped” person took \(A = 1\): everyone followed the optimal strategy. The trial compared “treat everyone” with “treat no one,” but the best strategy was “treat only those who benefit.” If, say, the type were determined by ethnic group and each group had learned from experience whether to take \(A\) (maximal effect modification and maximal confounding), the confounded observational study, not the unconfounded trial, revealed the correct policy (Hernán and Robins 2020, Fine Point 22.7, p. 321).

NoteFine Point 22.8: Generalizing Fine Point 22.7

Suppose lower \(Y\) is better and the observational mean \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\) is below the mean of both trial arms, so \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} < \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{}\). If \(U\) is a (possibly unknown) set of pre-treatment covariates with \(Y^a \perp\!\!\!\perp A \mid U\), then \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^g\right]\mathclose{}\) for the random strategy \(g\) that assigns \(A = 1\) with probability \(\Pr[A = 1 \mid U]\), the strategy that generated the observational data. That strategy cannot be implemented without data on \(U\), but it can motivate measuring pre-treatment covariates \(V\) and using the trial data to find a deterministic dynamic strategy \(g^*\) whose mean, estimated from the trial, is below the observational \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\). Combining trial and observational data can thus be more informative than the trial alone, provided both are random samples of everyone eligible for the trial (Hernán and Robins 2020, Fine Point 22.8, p. 322).

6 Summary


  • The ITT effect is the effect of randomized assignment; it is unconfounded but is neither guaranteed to preserve the null (without the exclusion restriction) nor guaranteed to be conservative.
  • The PP effect is the effect of adhering to the assigned strategies; it is often the more relevant estimand and generally requires adjustment, for sustained strategies via g-methods, in trials and observational studies alike.
  • PP strategies are usually dynamic: stopping for toxicity is adherence, not deviation.
  • Non-adherence in a trial allows emulation of alternative target trials.
  • A target trial protocol specifies eligibility, follow-up, strategies, assignment procedures, outcomes, causal contrast, and analysis plan; observational analogs are the comparison of initiators (ITT) and the PP effect.
  • Time zero must align eligibility, assignment, and start of follow-up; multiple eligible times call for sequential trials, and data compatible with several strategies call for cloning, censoring, and weighting.
  • Aside from baseline randomization, randomized and observational studies of sustained strategies should be analyzed the same way.

7 References


Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.
Back to top