---
title: "Chapter 8: Selection Bias"
format:
html: default
revealjs:
output-file: 08-selection-bias-slides.html
pdf:
output-file: 08-selection-bias-handout.pdf
docx:
output-file: 08-selection-bias.docx
preview-changed: true
---
{{< include ../latex-macros/macros.qmd >}}
::: {#exm-pedestrian-consent}
## Looking up at the sky, with consent
The chapter opens with a randomized experiment: an investigator randomly looks up at the sky in front of some pedestrians and records whether they look up too.
Randomization rules out confounding.
But only pedestrians who later consented to the use of their data were analyzed, and two kinds of pedestrians were less likely to consent: shy ones (who are also less likely to look up) and ones in front of whom the investigator looked up (who felt tricked).
Among the consenting pedestrians, those in front of whom the investigator looked up therefore tend to be less shy, and so more likely to look up.
The selection process alone produces an association between the investigator's looking up and the pedestrians' looking up, whether or not there is any causal effect.
:::
The pedestrian example shows an association created by the process that selects individuals into the analysis.
Such associations are the subject of this chapter.
::: {#def-selection-bias}
## Selection Bias
**Selection bias** is any bias produced by conditioning on a variable that is affected by two others,
where the first of those two is the treatment or one of its causes,
and the second is the outcome or one of its causes [@hernan2020causal, p. 111].
:::
::: {.notes}
A margin note in the book, crediting Hernán, Hernández-Díaz, and Robins (2004), widens the definition:
the two variables whose common effect is conditioned on need only be associated with treatment and with outcome, respectively,
rather than being treatment (or one of its causes) and outcome (or one of its causes) [@hernan2020causal, p. 111].
Conditioning on a descendant of such a collider has the same effect.
In the pedestrian example, consent is a common effect of the investigator's looking up (the treatment) and of shyness (a cause of the outcome).
:::
Selection bias needs no shared cause of treatment and outcome,
which sets it apart from confounding,
and randomization does not protect against it,
so observational studies and randomized experiments are both exposed.
What it does share with confounding is the underlying failure:
the treated and the untreated are not exchangeable.
The chapter defines selection bias and reviews methods to adjust for it.
::: {.notes}
This chapter is based on @hernan2020causal [Chapter 8, pp. 109-123].
:::
## 8.1 The Structure of Selection Bias (pp. 109-110)
---
"Selection bias" covers various biases that arise from how individuals are selected into the analysis.
::: {#def-selection-bias-under-null}
## Selection bias under the null
**Selection bias under the null** is selection bias (@def-selection-bias) that would arise even if treatment had no effect on the outcome ([Section 6.5](06-graphical-representation.qmd#bias-under-the-null)).
:::
The chapter focuses on selection bias under the null (@def-selection-bias-under-null).
Causal diagrams make the structure of selection bias explicit.
---
### Figure 8.1: Folic Acid and Cardiac Malformations
::: {#exm-folic-acid-selection-bias}
## Folic acid and cardiac malformations
Figure 8.1 depicts a study of folic acid supplements $A$, given to pregnant women shortly after conception, and the fetus's risk of a cardiac malformation $Y$ (1: yes, 0: no) during the first two months of pregnancy.
The variable $C$ is death before birth:
- a cardiac malformation increases mortality ($Y \rightarrow C$);
- folic acid lowers mortality by preventing non-cardiac malformations ($A \rightarrow C$);
- the study includes only fetuses who survived to birth, i.e., it conditions on $C = 0$ (the box around $C$).
$$A \rightarrow Y, \quad A \rightarrow C \leftarrow Y$$
The diagram shows two sources of association between treatment and outcome:
1. The open path $A \rightarrow Y$, the causal effect of $A$ on $Y$.
2. The path $A \rightarrow C \leftarrow Y$, which is opened by conditioning on the common effect $C$.
The association induced through the second path is **selection bias due to conditioning on $C$** (@def-selection-bias).
Because of it, the associational risk ratio $\Pr[Y = 1 \mid A = 1, C = 0] / \Pr[Y = 1 \mid A = 0, C = 0]$
generically (that is, for all parameter values outside a measure-zero set of special cases)
does not equal the causal risk ratio $\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]$.
Without conditioning on the collider $C$, the only open path would be $A \rightarrow Y$, and the associational risk ratio $\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]$ would equal the causal risk ratio.
:::
::: {.notes}
The book credits Pearl (1995) and Spirtes et al. (2000) with first using causal diagrams to describe bias from the selection of individuals [@hernan2020causal, p. 109].
:::
---
### Figure 8.2: Conditioning on a Descendant of the Collider
::: {#exm-parental-grief}
## Conditioning on parental grief
Figure 8.2 adds a node $S$, parental grief (1: yes, 0: no), which is affected by vital status at birth ($C \rightarrow S$).
If only nongrieving parents ($S = 0$) agreed to participate, the analysis conditions on $S = 0$.
As shown in Chapter 6, conditioning on a variable affected by the collider $C$ also opens the path $A \rightarrow C \leftarrow Y$.
In both Figures 8.1 and 8.2 the bias comes from conditioning on a common effect of treatment and outcome ($C$ or $S$).
Generically (for all but special parameter values), it is present whether or not there is an arrow $A \rightarrow Y$:
it is selection bias under the null (@def-selection-bias-under-null).
:::
::: {.notes}
A structure that produces bias under the null also produces bias when the treatment has a non-null effect.
Confounding (common causes of treatment and outcome) and selection bias (conditioning on common effects of treatment and outcome) are both biases under the null.
:::
---
### Figures 8.3-8.6: Differential Loss to Follow-Up
::: {#exm-differential-loss-to-follow-up}
## Differential loss to follow-up in an HIV study
Figure 8.3 depicts a follow-up study of individuals with HIV infection, estimating the effect of an antiretroviral treatment $A$ on 3-year risk of death $Y$.
The diagram has no arrow $A \rightarrow Y$: we assume the null of no treatment effect.
- $U$ (unmeasured): high level of immunosuppression (1: yes, 0: no);
$U = 1$ raises the risk of death.
- $C$: censoring, i.e., dropout or other loss to follow-up ($C = 1$).
Sicker individuals ($U = 1$) are more likely to be censored.
- $L$: symptoms (fever, weight loss, diarrhea, ...), CD4 count, and viral load, which mediate the effect of $U$ on $C$.
$L$ is treated as unmeasured here.
- $A \rightarrow C$: treated individuals have more side effects, which can lead them to drop out.
- The square around $C$: only uncensored individuals ($C = 0$) can have $Y$ ascertained, so the analysis is restricted to them.
By d-separation, conditioning on the collider $C$ opens the path
$$A \rightarrow C \leftarrow L \leftarrow U \rightarrow Y,$$
so, under the null (no arrow $A \rightarrow Y$), the associational risk ratio generically (for all but special parameter values) differs from 1, the causal value.
:::
::: {.notes}
Figure 8.3 is a transformation of Figure 8.1: the association between $C$ and $Y$ that came from the direct arrow $Y \rightarrow C$ now comes from a common cause $U$ of $Y$ and $C$.
:::
::: {#rem-direction-loss-to-follow-up-bias}
## Direction of the bias from loss to follow-up
Intuition for one possible direction: side effects of treatment push treated individuals toward dropping out.
A treated individual who stayed in the study ($C = 0$) stayed despite that push,
which makes it less likely that they also had another cause of dropout, such as $U = 1$.
So among those who stayed, the treated may carry $U = 1$ less often than the untreated.
That intuition suggests that $A$ and $U$ are inversely associated among the uncensored,
and so, because $U$ is positively associated with death $Y$,
that restricting to the uncensored induces an inverse association between $A$ and $Y$.
The diagram alone does not fix this sign.
[Faithfulness](06-graphical-representation.qmd#from-d-separation-to-independence) says only that $A$ and $U$ are dependent in at least one level of $C$;
it says nothing about a single stratum such as $C = 0$.
The sign, and even whether the uncensored stratum carries any association,
depends on how $A$ and $U$ act together on dropout.
For instance, if $\Pr[C = 0 \mid A, U]$ is a function of $A$ times a function of $U$,
then $A$ and $U$ stay independent among the uncensored.
:::
---
::: {#def-m-bias}
## M-bias
**M-bias** is selection bias that comes from conditioning on a common effect of a cause of treatment and a cause of the outcome,
or on a descendant of such a common effect [@hernan2020causal, pp. 97 and 111].
:::
The name refers to the M shape of the diagram when the two causes are drawn on top.
::: {#exm-loss-to-follow-up-variants}
## Further structures of differential loss to follow-up
Figure 8.3 is an example of selection bias from conditioning on $C$, a common effect of treatment $A$ and of a cause $U$ of the outcome, rather than a common effect of treatment and outcome themselves.
Three further structures lead to the same kind of bias through differential loss to follow-up:
- **Figure 8.4**: prior treatment $A$ directly affects symptoms $L$.
Restricting to the uncensored again conditions on a common effect $C$ of $A$ and $U$.
- **Figures 8.5 and 8.6**: variations of Figures 8.3 and 8.4 with an unmeasured common cause $W$ of treatment and another variable.
$W$ stands for lifestyle, personality, or educational factors that affect treatment ($W \rightarrow A$) and either attitudes toward attending study visits ($W \rightarrow C$, Figure 8.5) or the threshold for reporting symptoms ($W \rightarrow L$, Figure 8.6).
The book notes that Figures 8.5 and 8.6 are examples of M-bias (@def-m-bias),
with $W$ as the cause of treatment and $U$ as the cause of the outcome.
:::
::: {.notes}
In all of Figures 8.1-8.6, the bias results from selection on a collider, so each is selection bias in the sense of @def-selection-bias.
Figures 8.3-8.6 thus show that selection bias under the null (@def-selection-bias-under-null) can arise from more than conditioning on a common effect of treatment and outcome.
:::
## 8.2 Examples of Selection Bias (pp. 111-112)
---
::: {#exm-familiar-selection-biases}
## Familiar biases with the structure of selection bias
The structures in Figures 8.3-8.6 describe many familiar biases:
- **Differential loss to follow-up**, also called **bias due to informative censoring**: exactly the bias of Section 8.1.
- **Missing data bias, nonresponse bias**: $C$ can indicate missing outcome data for any reason, such as reluctance to provide information or missed study visits.
Restricting to complete cases ($C = 0$) may result in bias, whatever the reason the data are missing.
- **Healthy worker bias**: in a cohort of factory workers, $A$ is an occupational exposure (e.g., a chemical), $Y$ is death, $U$ is unmeasured true health status, and $C$ is being at work (1: no, 0: yes) when the outcome is ascertained;
$L$ could be blood tests and a physical exam.
The exposure lowers the chance of remaining at work either directly (e.g., disabling asthma;
Figures 8.3 and 8.4) or through a common cause $W$ (e.g., exposed jobs are eliminated and their workers laid off;
Figures 8.5 and 8.6).
- **Self-selection bias, volunteer bias**: $C$ is agreement to participate (1: no, 0: yes), $A$ is cigarette smoking, $Y$ is coronary heart disease, $U$ is family history of heart disease, $W$ is healthy lifestyle, and $L$ is a mediator between $U$ and $C$ such as awareness of heart disease.
Restricting to volunteers ($C = 0$) may introduce bias.
- **Selection affected by treatment received before study entry**: $C$ is selection into the study (1: no, 0: yes), and treatment $A$ happened before the study started.
If $A$ affects selection, bias is expected.
This generalizes self-selection bias and can arise whenever the treatment, or part of it, precedes the study, for example lifetime exposure to some factor in a study that recruits 50-year-olds.
Such studies may also have unmeasured confounding for the pre-study part of treatment if confounders were only measured during the study.
:::
::: {.notes}
The book attributes the structure of self-selection bias to Berkson (1955), and the use of causal diagrams for bias from pre-study treatments affecting selection to Robins, Hernán, and Rotnitzky (2007).
Causal diagrams have also characterized biases from attempts to eliminate ascertainment bias (Robins 2001), from estimating direct effects (Cole and Hernán 2002), and from conventional adjustment for variables affected by earlier treatment (Part III) [@hernan2020causal, pp. 112-113].
:::
---
### Selection Bias in Randomized and Observational, Prospective and Retrospective Studies
::: {#rem-selection-bias-study-designs}
## Selection bias across study designs
These examples show that selection bias (@def-selection-bias) can occur:
- in **retrospective** studies (treatment data collected after the outcome) and in **prospective** studies (treatment data collected before the outcome);
- in **observational studies** and in **randomized experiments**.
Figures 8.3 and 8.4 have no common causes of $A$ and any other variable, so they can depict a randomized trial.
Participants in a trial can be lost to follow-up before $Y$ is ascertained.
Then $\Pr[Y = 1 \mid A = a]$ cannot be computed, only $\Pr[Y = 1 \mid A = a, C = 0]$, and the uncensored may not be exchangeable with those who were lost.
:::
::: {.callout-warning title="Randomization Does Not Prevent Selection Bias"}
**Key difference from confounding**: randomization protects against confounding, but not against selection bias that occurs after randomization.
:::
---
::: {#rem-selection-before-randomization}
## Selection before randomization
Selection *before* treatment assignment does not bias a randomized trial.
Only volunteers are enrolled in trials, yet trials are not subject to volunteer bias, because participants are randomized only after agreeing to participate ($C = 0$).
None of Figures 8.3-8.6 can represent volunteer bias in a randomized trial:
- Figures 8.3 and 8.4 are ruled out because treatment cannot cause agreement to participate.
- Figures 8.5 and 8.6 are ruled out because random assignment leaves no common cause of treatment and any other variable.
:::
::: {.callout-note title="Fine Point 8.1: Selection Bias in Case-Control Studies"}
Figure 8.1 can represent selection bias in a case-control study.
An investigator wants the effect of postmenopausal estrogen treatment $A$ on coronary heart disease $Y$.
$C$ indicates whether a woman in the underlying cohort is selected into the case-control study (1: no, 0: yes).
The arrow $Y \rightarrow C$ (cases are more likely to be selected than noncases) is the defining feature of a case-control design.
Suppose controls ($Y = 0$) are selected preferentially among women with a hip fracture.
Because estrogens protect against hip fracture, treatment now affects selection ($A \rightarrow C$;
an intermediate node for hip fracture could be drawn but is not needed).
The odds ratio in a case-control study is by definition conditional on selection ($C = 0$), so this "inappropriate control selection" conditions on a common effect of treatment and outcome.
Heuristically: among the selected, controls are more likely than cases to have had a hip fracture, and so less likely to have used estrogens.
Hence the $A$-$Y$ odds ratio given $C = 0$ will be greater than the causal odds ratio in the population.
Other selection biases in case-control studies, including some described by Berkson (1946) and incidence-prevalence bias, can be represented by Figure 8.1 or modifications of it (Hernán, Hernández-Díaz, and Robins 2004).
:::
## 8.3 Selection Bias and Confounding (pp. 113-114)
---
::: {#rem-two-structures-nonexchangeability}
## Two structures of non-exchangeability
Chapters 7 and 8 describe two reasons the treated and the untreated may not be exchangeable:
1. **Common causes** of treatment and outcome: **confounding**.
2. **Conditioning on common effects** of treatment and outcome, or of causes of them: **selection bias** (@def-selection-bias).
This structural classification is clear-cut, though it does not always match the terminology of other disciplines.
Statisticians and econometricians often call both biases "selection bias," because both arise from selection: of individuals into the analysis (structural selection bias) or into a treatment (structural confounding).
The book's aim is not to police terminology but to stress that there are two distinct causal structures.
:::
::: {.notes}
For the same reason, social scientists often call unmeasured confounding "selection on unobservables."
Conditional exchangeability is also called "weak ignorability" or "ignorable treatment assignment" in statistics, "selection on observables" in the social sciences, and "no omitted variable bias" or "exogeneity" in econometrics [@hernan2020causal, pp. 111-113].
:::
---
### The Firefighter Example
::: {#exm-firefighters}
## Physical activity and heart disease among firefighters
A study restricted to firefighters estimates the effect of physical activity $A$ on heart disease $Y$ (Figure 8.7).
Assume, unknown to the investigators, that $A$ does not cause $Y$.
- Parental socioeconomic status $L$ affects becoming a firefighter $C$ and, through childhood diet, heart disease $Y$.
- An unmeasured attraction to physical activities $U$ affects becoming a firefighter and being physically active $A$.
- $U$ does not affect $Y$, and $L$ does not affect $A$.
There is no confounding: $A$ and $Y$ have no common cause, so in the full population $\Pr[Y = 1 \mid A = 1] / \Pr[Y = 1 \mid A = 0]$ is expected to equal $\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1] = 1$.
Among firefighters ($C = 0$), however, the two ratios generically (for all but special parameter values) differ:
conditioning on $C$, a common effect of causes of treatment ($U$) and of outcome ($L$),
opens $A \leftarrow U \rightarrow C \leftarrow L \rightarrow Y$.
:::
---
::: {#rem-structural-classification-advantages}
## Why classify biases by structure
For the investigators the label is moot: either way they must adjust for $L$ to make treated and untreated firefighters comparable.
For this reason many epidemiologists call any variable that needs adjustment a "confounder," whatever the structure.
Advantages of the structural classification anyway:
1. The structure guides the choice of method.
With time-varying treatments it reveals when stratification-based adjustment for confounding would introduce selection bias,
so that g-methods are needed (Part III).
2. It can inform design even when it does not change the analysis.
Investigators restricting to firefighters should collect joint risk factors for $Y$ and $C$,
as in the first example of [Section 7.1](07-confounding.qmd#the-structure-of-confounding-pp.-91-93).
3. Selection on pre-treatment variables (like being a firefighter) explains why a variable such as $L$
can act as a "confounder" in some studies but not in others,
e.g., studies not restricted to firefighters.
4. Causal diagrams improve communication and reduce misunderstandings.
:::
::: {.notes}
The book's margin uses Simpson's paradox as an example of a paradox that results from ignoring the difference between common causes and common effects;
Blyth (1972) misread Simpson's (1951) example as an extreme case of confounding (see Hernán, Clayton, and Keiding 2011).
:::
---
### Healthy Worker Bias Revisited
::: {#exm-healthy-worker-bias-two-structures}
## Two biases called healthy worker bias
The term "healthy worker bias" names two structurally different biases:
1. The bias of Section 8.2: conditioning on being at work $C$, a common effect of (a cause of) treatment and (a cause of) outcome.
This is **selection bias** (Figures 8.3-8.6).
2. The bias from comparing the risk in a group of workers with the risk in the general population.
With $L$ = health status, $A$ = membership in the group of workers, and $Y$ = outcome, health affects both job type and the outcome ($A \leftarrow L \rightarrow Y$, as in Figure 7.1).
This is **confounding**.
Drawing the causal diagram removes the confusion that comes from using one name for two sources of non-exchangeability.
:::
::: {.callout-note title="Selection Bias Is Not All or Nothing"}
The structural analysis of Sections 8.1-8.3 ignores the magnitude and direction of the bias.
A noncausal path opened by conditioning on a collider may be weak and induce little bias;
selection bias is not an "all or nothing" issue.
:::
---
::: {#def-frailty}
## Frailty
A **frailty** is an unmeasured cause of the outcome, often of death, that is marginally independent of treatment [@hernan2020causal, p. 114].
:::
::: {#def-noncollapsibility}
## Noncollapsibility
An effect measure is **noncollapsible** with respect to a covariate when its value in the whole population is not a weighted average of its values within the strata of that covariate;
this property is called **noncollapsibility**.
:::
::: {.callout-note title="Technical Point 8.1: The Built-in Selection Bias of Hazard Ratios"}
Figure 8.8 describes a randomized experiment of heart transplant $A$ and death at times 1 ($Y_1$) and 2 ($Y_2$).
Transplant lowers the risk of death at time 1 ($A \rightarrow Y_1$) but has no direct effect on death at time 2 (no arrow $A \rightarrow Y_2$).
An unmeasured haplotype $U$ lowers the risk of death at all times.
With no confounding, the associational risk ratios $\Pr[Y_1 = 1 \mid A = 1] / \Pr[Y_1 = 1 \mid A = 0]$ and $\Pr[Y_2 = 1 \mid A = 1] / \Pr[Y_2 = 1 \mid A = 0]$ are unbiased for the effects on death by times 1 and 2;
the second is less than 1 because it reflects total mortality through time 2.
In discrete time the hazard at time 1 is just the risk at time 1, but the hazard at time 2 conditions on surviving time 1:
$$\frac{\Pr[Y_2 = 1 \mid A = 1, Y_1 = 0]}{\Pr[Y_2 = 1 \mid A = 0, Y_1 = 0]}.$$
Surviving time 1 carries different information in the two arms.
For an untreated survivor, survival is more informative about the protective $U$,
because for a treated survivor the transplant offers a competing explanation.
That suggests $U$ is rarer among treated survivors than among untreated survivors,
which would leave treated survivors at higher risk of death at time 2.
The diagram does not guarantee this sign:
if $\Pr[Y_1 = 0 \mid A, U]$ is a function of $A$ times a function of $U$,
then $A$ and $U$ stay independent among time-1 survivors.
When the intuition does hold, the hazard ratio is below 1 at time 1 and above 1 at time 2,
so the hazards cross.
The time-2 hazard ratio is generically (for all but special parameter values) biased for the (null) direct effect,
because conditioning on $Y_1$, a common effect of $A$ and $U$, opens $A \rightarrow Y_1 \leftarrow U \rightarrow Y_2$.
The haplotype $U$ is a frailty (@def-frailty).
Within strata of $U$ the hazard ratio at time 2 is 1, because conditioning on the non-collider $U$ blocks that path.
Here $U$ is independent of $A$,
yet the time-2 hazard ratio ignoring $U$ generically (for all but special parameter values) differs from the ratio within each level of $U$.
A gap of that kind, with no confounding by $U$ to explain it,
is the noncollapsibility (@def-noncollapsibility) of the hazard ratio.
Since $U$ is unobserved, the data cannot tell whether $A$ has a direct effect on $Y_2$ (Figure 8.8 versus Figure 8.9).
This applies to observational studies and randomized experiments alike.
:::
## 8.4 Selection Bias and Censoring (pp. 115-116)
---
::: {#exm-wasabi-trial}
## The wasabi trial without censoring
A marginally randomized experiment estimates the effect of wasabi intake on one-year risk of death ($Y = 1$).
Of 60 participants, 30 were assigned to meals with wasabi ($A = 1$) until death or end of follow-up, and 30 to meals without wasabi ($A = 0$).
After one year, 17 had died in each group, so
$$\frac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = \frac{17/30}{17/30} = 1,$$
and, by randomization, the causal risk ratio $\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]$ is also expected to be 1.
(To set aside random variability, imagine 60 million participants instead of 60.)
:::
---
::: {#exm-wasabi-censoring}
## The wasabi trial with censoring
But many participants were lost to follow-up (censored, $C = 1$), more often those with heart disease at baseline ($L = 1$) and those assigned to wasabi.
Only 9 wasabi and 22 no-wasabi participants remained uncensored, with 4 and 11 observed deaths, respectively.
Among the uncensored,
$$\frac{\Pr[Y = 1 \mid A = 1, C = 0]}{\Pr[Y = 1 \mid A = 0, C = 0]} = \frac{4/9}{11/22} = \frac{0.444}{0.5} = 0.89.$$
The 0.89 among the uncensored differs from the causal risk ratio of 1: there is selection bias (@def-selection-bias) from conditioning on the common effect $C$.
:::
::: {.notes}
Figure 8.3 describes the wasabi trial, with $U$ now atherosclerosis, an unmeasured cause of both heart disease $L$ and death $Y$.
There is no common cause of $A$ and $Y$, as expected under marginal randomization, so no adjustment for confounding is needed for the effect of $A$.
But $C$ and $Y$ share a common cause, so the backdoor path $C \leftarrow L \leftarrow U \rightarrow Y$ is open: if we wanted the (null) effect of censoring $C$ on $Y$, we would need to adjust for confounding by $U$.
By the backdoor criterion, the measured $L$ suffices to block that path.
:::
---
### Counterfactual Outcomes Under No Censoring
Why discuss confounding for the effect of $C$, when the contrast $\Pr[Y^{a=1} = 1]$ versus $\Pr[Y^{a=0} = 1]$ does not involve $C$?
Because in the presence of censoring (or any selection) the causal contrast of interest should be redefined.
Selection bias would not exist if nobody were censored, so we target what would have happened without censoring.
::: {#def-counterfactual-no-censoring}
## Counterfactual outcome under no censoring
Let $Y^{a,c=0}$ be an individual's counterfactual outcome had they received treatment $a$ and remained uncensored ($C = 0$).
:::
The causal contrast of interest becomes
$$\Pr[Y^{a=1,c=0} = 1] \quad \text{versus} \quad \Pr[Y^{a=0,c=0} = 1],$$
for example the risk ratio $\E{Y^{a=1,c=0}} / \E{Y^{a=0,c=0}}$ or the risk difference $\E{Y^{a=1,c=0}} - \E{Y^{a=0,c=0}}$.
::: {.notes}
Using counterfactuals such as $Y^{a,c=0}$ presupposes, as in [Chapter 3](03-observational-studies.qmd),
that some no-direct-effect intervention could set censoring $C$ to 0 [@hernan2020causal, p. 116, margin note].
Censoring often has no causal effect on the outcome (an exception: loss to follow-up that prevents people from receiving further treatment).
One might then drop the superscript $c = 0$, but keeping it makes clear that confounding for the effect of $C$ becomes central when estimating the effect of $A$ under selection.
Even then, the observed $Y$ depends on $C$:
it equals $Y^{c=0}$ when $C = 0$ and is missing when $C = 1$.
So a causal diagram in which $C$ has no effect on the outcome can still replace $Y$ by $Y^{c=0}$ and add arrows $Y^{c=0} \rightarrow Y$ and $C \rightarrow Y$,
where $C \rightarrow Y$ encodes only that censoring determines whether the outcome is observed.
:::
---
::: {#rem-censoring-as-treatment}
## Censoring as another treatment
With the contrast written in terms of $Y^{a,c=0}$ (@def-counterfactual-no-censoring),
**censoring $C$ is just another treatment**,
and the goal is the effect of a joint intervention on $A$ and $C$.
To eliminate selection bias for the effect of $A$, we adjust for confounding for the effect of $C$.
That requires:
1. the identifiability conditions (exchangeability, positivity, consistency) to hold for $C$ as well as for $A$;
2. the same analytic methods we would use to estimate the effect of $C$.
Under these conditions, and with no measurement error or confounding for $A$, the effect of $A$ on $Y$ is identified.
:::
## 8.5 How to Adjust for Selection Bias (pp. 117-120)
---
::: {.callout-tip title="Avoid by Design, Correct by Analysis"}
Design can sometimes avoid selection bias (e.g., Fine Point 8.1), but loss to follow-up, self-selection, and missing data can happen however careful the investigator is.
Then the bias must be corrected in the analysis.
:::
---
### Inverse Probability Weighting for Selection Bias
IP weighting (or standardization) can do this.
::: {#def-censoring-ip-weights}
## IP weights for selection
Each selected individual ($C = 0$) gets a weight $W^C$ so that she represents herself and the unselected individuals ($C = 1$) with her values of $L$ and $A$:
$$W^C = \frac{1}{\Pr[C = 0 \mid L, A]}.$$
:::
::: {.notes}
IP weights for confounding are $W^A = 1/f(A \mid L)$;
IP weights for selection bias are $W^C = 1/\Pr[C = 0 \mid A, L]$ (@def-censoring-ip-weights).
When both biases exist, the product $W^A W^C$ adjusts for both simultaneously, under assumptions described in Chapter 12 and Part III.
:::
---
### The Wasabi Trial: IP Weighting in Practice
::: {#exm-wasabi-tree}
## The wasabi trial tree
The tree in Figure 8.10 shows the wasabi trial data.
Of 60 participants, 40 had heart disease ($L = 1$) and 20 did not ($L = 0$).
Everyone had a 50/50 chance of wasabi ($A = 1$) regardless of $L$, so 10 of the $L = 0$ group and 20 of the $L = 1$ group were treated (no arrow $L \rightarrow A$ in Figure 8.3).
The probability of remaining uncensored depends on $A$ and $L$ (arrows $A \rightarrow C$ and $L \rightarrow C$).
The book gives:
- $L = 0, A = 1$: 50% remained uncensored, so $0.5 \times 10 = 5$ individuals;
- $L = 1, A = 0$: 60% remained uncensored, so $0.6 \times 20 = 12$ individuals;
- $L = 1, A = 1$: 4 of 20 remained uncensored, so $\Pr[C = 0 \mid L = 1, A = 1] = 4/20 = 0.2$.
:::
---
::: {#exm-wasabi-missing-branch}
## The wasabi trial: completing the tree
The remaining branch follows from the totals in Section 8.4 (@exm-wasabi-censoring):
- treated uncensored: $5 + 4 = 9$, matching the 9 reported;
- untreated uncensored: $22 - 12 = 10$ in the $L = 0, A = 0$ branch, i.e., all 10, so $\Pr[C = 0 \mid L = 0, A = 0] = 10/10 = 1$.
:::
::: {#tbl-wasabi-censoring-weights}
| $L$ | $A$ | uncensored / total | $\Pr[C = 0 \mid L, A]$ | $W^C$ | weighted count |
|:---:|:---:|:---:|:---:|:---:|:---:|
| 0 | 0 | 10/10 | 1 | 1 | $10 \times 1 = 10$ |
| 0 | 1 | 5/10 | 0.5 | 2 | $5 \times 2 = 10$ |
| 1 | 0 | 12/20 | 0.6 | $5/3$ | $12 \times 5/3 = 20$ |
| 1 | 1 | 4/20 | 0.2 | 5 | $4 \times 5 = 20$ |
Probability of remaining uncensored, IP weight for censoring $W^C$, and weighted count in each branch of the wasabi trial (Figure 8.10).
:::
@tbl-wasabi-censoring-weights applies the weights $W^C = 1/\Pr[C = 0 \mid L, A]$ (@def-censoring-ip-weights)
to each branch of the completed tree.
The weighted counts reproduce the 60 original participants: $10 + 10 + 20 + 20 = 60$.
::: {.notes}
The book leaves this derivation of the branch weights to the reader.
The tree in Figure 8.10 also shows the deaths among the censored, which investigators could never observe in practice;
the book shows them only to document that $A$ is marginally independent of $Y$ and that $C$ is independent of $Y$ within levels of $L$, as Figure 8.3 implies.
With those numbers, the risk ratio in the entire population is 1 and the risk ratio among the uncensored is 0.89.
:::
---
### Identifiability Conditions for IP Weighting
::: {#prp-ipw-selection-identification}
## Identifiability conditions for IP weighting of selection
Let $L$ be discrete,
and let the pseudo-population be the original population with each individual weighted by $W^C$ (@def-censoring-ip-weights),
counting only the uncensored ($C = 0$).
Fix a treatment level $a$ and suppose that:
1. **Exchangeability for censoring**: $Y^{a,c=0} \ind C \mid A = a, L$.
2. **Positivity**: $\Pr[A = a] > 0$, and $\Pr[C = 0 \mid A = a, L = l] > 0$ for every $l$ with $\Pr[A = a, L = l] > 0$.
3. **Consistency**: $Y = Y^{a,c=0}$ for every individual with $A = a$ and $C = 0$.
This requires sufficiently well-defined interventions on $A$ and $C$,
and no measurement error in $A$ or $Y$.
4. **No confounding for $A$**: $Y^{a,c=0} \ind A$, as under marginal randomization.
Then the mean outcome among individuals with $A = a$ in the pseudo-population equals $\E{Y^{a,c=0}}$,
the mean outcome had everyone received $a$ and nobody been censored (@def-counterfactual-no-censoring):
$$\frac{\E{I(A = a)\, (1 - C)\, W^C\, Y}}{\E{I(A = a)\, (1 - C)\, W^C}} = \E{Y^{a,c=0}}.$$
So if the conditions hold for every $a$,
any contrast of the pseudo-population means, such as the risk ratio, equals the same contrast of the $\E{Y^{a,c=0}}$.
:::
::: {.proof}
Write $\pi(a, l) \eqdef \Pr[C = 0 \mid A = a, L = l]$, which is positive by positivity.
Conditioning on $A$ and $L$, and summing over the $l$ with $\Pr[A = a, L = l] > 0$:
$$
\begin{aligned}
\E{I(A = a)\, (1 - C)\, W^C\, Y}
&= \sum_l \Pr[A = a, L = l]\, \E{(1 - C)\, W^C\, Y \mid A = a, L = l} && \text{(law of total expectation)} \\
&= \sum_l \Pr[A = a, L = l]\, \frac{\E{(1 - C)\, Y \mid A = a, L = l}}{\pi(a, l)} && \text{(definition of } W^C\text{)} \\
&= \sum_l \Pr[A = a, L = l]\, \E{Y \mid A = a, C = 0, L = l} && \text{(conditioning on } C\text{)} \\
&= \sum_l \Pr[A = a, L = l]\, \E{Y^{a,c=0} \mid A = a, C = 0, L = l} && \text{(consistency)} \\
&= \sum_l \Pr[A = a, L = l]\, \E{Y^{a,c=0} \mid A = a, L = l} && \text{(exchangeability for censoring)} \\
&= \Pr[A = a]\, \E{Y^{a,c=0} \mid A = a} && \text{(law of total expectation)} \\
&= \Pr[A = a]\, \E{Y^{a,c=0}} && \text{(no confounding for } A\text{)}
\end{aligned}
$$
The third line uses $\E{(1 - C)\, Y \mid A = a, L = l} = \pi(a, l)\, \E{Y \mid A = a, C = 0, L = l}$.
Replacing $Y$ by 1 in the first three lines shows that the denominator is $\sum_l \Pr[A = a, L = l] = \Pr[A = a]$.
Dividing the two gives the result.
:::
---
::: {#exm-wasabi-pseudo-population}
## The wasabi trial pseudo-population
In the $L = 1, A = 1$ branch, the 16 censored individuals get weight 0 (they do not contribute)
and each of the 4 uncensored individuals gets weight $1/0.2 = 5$:
IP weighting replaces the 20 original individuals by 5 copies of each of the 4 uncensored ones.
Repeating this for every branch (Figure 8.11) creates a pseudo-population of the original size in which nobody is lost to follow-up.
In that pseudo-population the associational risk ratio is 1.
The wasabi trial meets the conditions of @prp-ipw-selection-identification:
- exchangeability for censoring holds because, in Figure 8.3, $A$ and $L$ block every backdoor path between $C$ and $Y$;
- positivity holds because every $\Pr[C = 0 \mid L, A]$ in @tbl-wasabi-censoring-weights is positive;
- consistency is assumed, with well-defined interventions on $A$ and $C$ and no measurement error;
- treatment was marginally randomized, so there is no confounding for $A$.
So the pseudo-population risk ratio of 1 equals $\Pr[Y^{a=1,c=0} = 1] / \Pr[Y^{a=0,c=0} = 1]$ (@def-counterfactual-no-censoring),
the risk ratio had nobody been censored.
:::
Exchangeability for censoring implies that the average counterfactual outcome $Y^{a,c=0}$ of the uncensored equals that of the censored with the same $A$ and $L$
(mean exchangeability, which is all the proof uses).
The condition is met when the model for $\Pr[C = 0 \mid L, A]$ includes treatment
together with every factor that is an independent predictor of selection and of the outcome alike.
Graphically, the condition is that
$A$ and $L$ jointly block every backdoor path from $C$ to $Y$.
::: {.callout-warning title="Exchangeability for Censoring Is Untestable"}
We can never be sure that $A$ and $L$ block every backdoor path between censoring and the outcome, so this assumption is untestable.
:::
The positivity needed here concerns staying uncensored:
$\Pr[C = 0 \mid L, A]$ must be positive in every stratum that occurs.
A stratum where censoring never happens causes no trouble,
because the target is a pseudo-population free of censoring,
never one in which censoring is universal.
For instance, Figure 8.10 has $\Pr[C = 1 \mid L = 0, A = 0] = 0$,
and that branch still receives a well-defined weight of 1.
::: {#def-competing-event}
## Competing event
A **competing event** is an event that prevents the outcome of interest from happening, typically death.
:::
Consistency, which here also demands that the interventions be sufficiently well defined,
lets us read the pseudo-population effect as the effect had nobody been censored.
Abolishing censoring is a fairly clear intervention
when censoring means loss to follow-up or nonresponse.
It is much less clear when the censoring is a competing event (@def-competing-event),
since it is hard to say what intervention would remove such an event.
---
::: {.callout-warning title="Competing Events Are Not Censoring"}
In a study of a treatment's effect on Alzheimer's disease, treating death from other causes (cancer, heart disease, and so on) as censoring would target a pseudo-population in which all other causes of death have been removed.
It is unclear what intervention could produce such a population, and no feasible intervention could remove one cause of death without affecting others.
:::
---
### Stratification vs. IP Weighting
Could we instead estimate effects within levels of $L$?
::: {#rem-stratification-figure-8-3}
## Stratification when $L$ only blocks a backdoor path
In Figure 8.3, conditioning on $L$ blocks the backdoor path from $C$ to $Y$,
so exchangeability for censoring holds within levels of $L$,
and treatment is randomized.
If positivity and consistency also hold, the conditional risk ratio
$$\frac{\Pr[Y = 1 \mid A = 1, C = 0, L = l]}{\Pr[Y = 1 \mid A = 0, C = 0, L = l]}$$
equals $\Pr[Y^{a = 1, c = 0} = 1 \mid L = l] / \Pr[Y^{a = 0, c = 0} = 1 \mid L = l]$,
the effect of treatment in the stratum $L = l$ had nobody been censored.
Among the diagram's variables, $Y^{a, c = 0}$ is a function only of $U$ (and of its own independent error),
and $U$ is independent of $A$ and $C$ given $L$.
So the same ratio is also the effect of treatment among the uncensored with $L = l$.
Stratification also works under Figure 8.5,
where treatment is not randomized,
because there too $U$ is independent of $A$ and $C$ given $L$.
:::
::: {.callout-warning title="Stratification Can Open a Collider Path"}
Stratification fails under Figures 8.4 and 8.6.
In Figure 8.4, conditioning on $L$ blocks the backdoor path from $C$ to $Y$ but opens $A \rightarrow L \leftarrow U \rightarrow Y$, because $L$ is a collider on that path;
even with a null effect, the $L$-conditional risk ratio generically (for all but special parameter values) differs from 1.
Figure 8.6 behaves the same way.
IP weighting adjusts correctly under all of Figures 8.3-8.6 because it estimates unconditional effects after reweighting by treatment and $L$, rather than conditioning on $L$.
:::
::: {.notes}
This is the book's first example of a setting in which stratification cannot validly estimate the causal effect even though exchangeability, positivity, and consistency hold.
Part III returns to similar structures with time-varying treatments.
:::
## 8.6 Selection Without Bias (pp. 121-123)
---
::: {#exm-surgery-haplotype}
## Surgery, haplotype, and death
Figure 8.12 depicts a study with dichotomous surgery $A$, genetic haplotype $E$, and death $Y$ ($A \rightarrow Y \leftarrow E$).
By d-separation, $A$ and $E$ are:
1. marginally independent: surgery is equally likely with and without the haplotype;
2. generically (that is, under faithfulness) associated conditionally on $Y$, i.e., in at least one stratum of $Y$.
:::
::: {#rem-collider-association-some-stratum}
## Collider stratification need not affect every stratum
Conditioning on a common effect $Y$ of two independent causes generically (that is, under faithfulness) induces an association between them in **at least one** stratum of $Y$ (say $Y = 1$).
But there is a special case in which they remain independent in the other stratum (say $Y = 0$).
:::
---
### When Collider Stratification Does Not Cause Bias
::: {#exm-separate-death-mechanisms}
## Separate mechanisms of death
Suppose $A$ and $E$ affect survival through completely separate mechanisms, so neither modifies the other's effect.
For example, surgery removes a tumor, while the haplotype raises LDL-cholesterol and hence the risk of heart attack (with or without a tumor).
Define three cause-specific death indicators:
- $Y_A$: death from tumor;
- $Y_E$: death from heart attack;
- $Y_O$: death from other causes.
Then $Y = 1$ if any of $Y_A$, $Y_E$, $Y_O$ equals 1, and $Y = 0$ if all three equal 0.
Figure 8.13 expands Figure 8.12 with these unmeasured variables ($A \rightarrow Y_A \rightarrow Y$, $E \rightarrow Y_E \rightarrow Y$, $Y_O \rightarrow Y$);
only $A$, $E$, $Y$ are recorded.
:::
::: {#def-augmented-causal-diagram}
## Augmented causal diagram
Adding to a causal diagram unmeasured variables, such as $Y_A$, $Y_E$, $Y_O$, that functionally determine one of its nodes, such as $Y$, gives an **augmented causal diagram**.
:::
Figure 8.13 is the augmented version of Figure 8.12 (@exm-separate-death-mechanisms).
::: {.notes}
Augmented causal DAGs were introduced by Hernán, Hernández-Díaz, and Robins (2004) and can be extended to represent the sufficient causes of Chapter 5 (VanderWeele and Robins 2007c).
:::
---
::: {#prp-independence-among-survivors}
## Independence among survivors
Suppose, as in @exm-separate-death-mechanisms, that:
- $Y = 1$ if at least one of $Y_A$, $Y_E$, $Y_O$ equals 1, and $Y = 0$ if all three equal 0;
- the causal diagram is Figure 8.13, whose only arrows are $A \rightarrow Y_A$, $E \rightarrow Y_E$, and $Y_A, Y_E, Y_O \rightarrow Y$,
so there are no arrows between the mechanisms and no common causes of any of $A$, $E$, $Y_A$, $Y_E$, $Y_O$;
- $\Pr[Y = 0] > 0$.
Then $A \ind E \mid Y = 0$.
:::
::: {.proof}
In Figure 8.13 the only path between $A$ and $E$ is $A \rightarrow Y_A \rightarrow Y \leftarrow Y_E \leftarrow E$.
Conditioning on the set $\{Y_A, Y_E, Y_O\}$ blocks this path at the non-collider $Y_A$ (and at $Y_E$),
and the collider $Y$ is neither in the set nor an ancestor of anything in it,
so $A$ and $E$ are d-separated given $\{Y_A, Y_E, Y_O\}$.
By the causal Markov property, $A \ind E \mid Y_A, Y_E, Y_O$,
and in particular $A$ and $E$ are independent given the event $Y_A = Y_E = Y_O = 0$.
Because $Y$ is a deterministic function of $Y_A$, $Y_E$, $Y_O$,
the event $Y = 0$ is the same event as $Y_A = Y_E = Y_O = 0$,
so $A \ind E \mid Y = 0$.
:::
::: {#rem-dependence-among-decedents}
## Dependence among the dead
Given $Y = 1$ the argument of @prp-independence-among-survivors does not apply:
the event $Y = 1$ is compatible with 7 combinations of $(Y_A, Y_E, Y_O)$, every combination except $(0, 0, 0)$,
so it does not determine $(Y_A, Y_E, Y_O)$,
and the step that replaced conditioning on $Y$ by conditioning on $(Y_A, Y_E, Y_O)$ does not go through.
Generically (for all but special parameter values), $A$ and $E$ are associated given $Y = 1$.
:::
---
::: {#def-multiplicative-survival-model}
## Multiplicative survival model
Treatment $A$ and another cause $E$ of $Y$ follow a **multiplicative survival model** if there are functions $g$ and $h$ such that, for all $a$ and $e$,
$$\Pr[Y = 0 \mid E = e, A = a] = g(e) h(a),$$
that is, the probability of survival is a factor depending only on $e$ times a factor depending only on $a$.
:::
When the data can be summarized by Figure 8.13, they follow a multiplicative survival model (@def-multiplicative-survival-model).
::: {.notes}
$A$ and $E$ are generically (for all but special parameter values) dependent within both strata of $Y$
(in particular $Y = 0$, unlike in Figure 8.13)
when any of these hold:
- $A$ and $E$ act through a common mechanism: an arrow from $A$ to $Y_E$ or from $E$ to $Y_A$ (Figure 8.14);
- $Y_A$ and $Y_E$ share a common cause $V$ (Figure 8.15);
- $Y_A$ and $Y_O$, and $Y_E$ and $Y_O$, share common causes $W_1$ and $W_2$ (Figure 8.16).
The augmented causal diagram (@def-augmented-causal-diagram) represents both the independence of $A$ and $E$ given $Y = 0$ and their generic dependence given $Y = 1$.
:::
---
::: {#rem-collider-stratification-without-bias}
## Selection without bias
**Summary of Section 8.6**: conditioning on a collider generically (under faithfulness) induces an association between its causes,
but the association may be confined to some levels of the collider.
Restricting the analysis to a single level of a common effect does not necessarily produce selection bias: collider stratification is not always a source of selection bias.
:::
::: {.callout-note title="Technical Point 8.2: Multiplicative Survival Model"}
Under a multiplicative survival model (@def-multiplicative-survival-model), $\Pr[Y = 0 \mid E = e, A = a] = g(e) h(a)$.
Equivalently, the survival ratio $\Pr[Y = 0 \mid E = e, A = a] / \Pr[Y = 0 \mid E = e, A = 0]$ equals $g(e)h(a) / [g(e)h(0)] = h(a)/h(0)$, which does not depend on $e$: there is no interaction between $A$ and $E$ for $Y = 0$ on the multiplicative scale.
Proof that Figure 8.13 implies this model:
$$
\begin{aligned}
\Pr[Y = 0 \mid E = e, A = a]
&= \Pr[Y_O = 0, Y_A = 0, Y_E = 0 \mid E = e, A = a] && \text{(determinism)} \\
&= \Pr[Y_O = 0] \Pr[Y_A = 0 \mid A = a] \Pr[Y_E = 0 \mid E = e] && \text{(DAG factorization)}
\end{aligned}
$$
Set $g(e) = \Pr[Y_E = 0 \mid E = e]$ and $h(a) = \Pr[Y_O = 0] \Pr[Y_A = 0 \mid A = a]$.
If survival follows $g(e)h(a)$, then mortality $\Pr[Y = 1 \mid E = e, A = a] = 1 - g(e)h(a)$ does not, in general, factor in the same way, i.e., it does not, in general, follow a multiplicative mortality model (it does only when $g$ or $h$ is constant).
Bayes' rule shows what this means for $A$ and $E$.
Because $A \ind E$ in Figure 8.13, for each $y$
$$\Pr[A = a, E = e \mid Y = y] \propto \Pr[A = a] \Pr[E = e] \Pr[Y = y \mid E = e, A = a].$$
For $y = 0$ the right side is $\{\Pr[A = a] h(a)\} \{\Pr[E = e] g(e)\}$, a function of $a$ times a function of $e$,
so $A$ and $E$ are conditionally independent given $Y = 0$.
For $y = 1$ the right side is $\Pr[A = a] \Pr[E = e] \{1 - g(e) h(a)\}$, which in general does not factor in this way,
so $A$ and $E$ are, generically, conditionally dependent given $Y = 1$.
:::
::: {.callout-note title="Fine Point 8.2: The Strength and Direction of Selection Bias"}
So far selection bias has been treated as present or absent.
In practice its direction and magnitude matter.
**Direction.**
The sign of the association between two marginally independent causes $A$ and $E$ within strata of their common effect $Y$ depends on how they interact to cause $Y$.
Suppose an undiscovered background factor $U$, unassociated with $A$ and $E$, must be present for either to cause death.
- **"Or" mechanism**: in the presence of $U$, either $A = 1$ or $E = 1$ is sufficient and necessary for death.
Among the dead ($Y = 1$), $A$ and $E$ are negatively associated: a decedent with $A = 0$ is more likely to have $E = 1$, since $E$ is then the more likely cause.
The log of the conditional odds ratio $\OR_{AE \mid Y = 1}$ approaches $-\infty$ as the prevalence of $U$ approaches 1.
- **"And" mechanism**: in the presence of $U$, having both $A = 1$ and $E = 1$ is sufficient and necessary for death.
Among the dead, those with $A = 1$ are more likely to have $E = 1$: $A$ and $E$ are positively associated.
A standard DAG such as Figure 8.12 cannot distinguish the two mechanisms;
causal DAGs with sufficient-causation structures (VanderWeele and Robins 2007c) can.
**Magnitude.**
A bias too small to change a study's conclusions can be ignored in practice, whatever its direction.
Generally, a large selection bias requires strong associations between the collider and both treatment and outcome.
Greenland (2003) studied the magnitude of selection bias under the null, which he called collider-stratification bias, in several scenarios.
:::
## Summary
---
This chapter examined **selection bias**: lack of exchangeability that arises from conditioning on common effects rather than from common causes.
**Key concepts**:
1. **Structure**: conditioning on a common effect (collider) of treatment (or a cause of it) and outcome (or a cause of it), or on a descendant of such a collider (Figures 8.1-8.6).
2. **Examples** (Figures 8.3-8.6):
- Differential loss to follow-up (informative censoring)
- Missing data and nonresponse bias
- Healthy worker bias
- Self-selection and volunteer bias
- Selection affected by pre-study treatment
3. **Selection bias vs. confounding**:
- Confounding: common causes ($A \leftarrow L \rightarrow Y$)
- Selection bias: conditioning on common effects (e.g., $A \rightarrow C \leftarrow Y$)
- Randomization prevents confounding but not selection bias that occurs after randomization
4. **Censoring as selection**: in the wasabi trial, restricting to the uncensored gives a risk ratio of 0.89 instead of the causal value of 1.
The target becomes $Y^{a,c=0}$, and censoring $C$ is treated as another treatment.
5. **Adjustment**: IP weighting with $W^C = 1 / \Pr[C = 0 \mid L, A]$ creates a pseudo-population with no censoring.
Stratification on $L$ works in Figures 8.3 and 8.5 but not in Figures 8.4 and 8.6.
6. **Selection without bias**: when $A$ and $E$ are marginally independent
and the data follow a multiplicative survival model (@def-multiplicative-survival-model, Figure 8.13),
$A$ and $E$ remain independent among survivors ($Y = 0$, by @prp-independence-among-survivors),
but are generically (for all but special parameter values) dependent among the dead ($Y = 1$).
::: {.notes}
**Practical implications**:
1. **Design**: selection before randomization does not bias a trial;
selection after it can, so minimize loss to follow-up and missing data.
2. **Analysis**: use causal diagrams to identify the structure of selection bias and choose an adjustment method.
3. **Key warning**: stratification on $L$ removes selection bias in some structures (Figure 8.3) but opens a collider path in others (Figure 8.4).
IP weighting handles all of Figures 8.3-8.6.
**Looking ahead**:
- **Chapter 9**: Measurement bias
- **Chapter 12**: IP weighting in detail, including weights for censoring
- **Part III**: time-varying treatments, where conventional adjustment can itself introduce selection bias
:::
## References
---
::: {#refs}
:::