So far we have focused on the average causal effect in an entire population. Many causal questions, however, are about subsets of the population: the effect of your looking up in city dwellers versus visitors, say.
Whether to target the whole population or a subset depends on the goal:
The theme of this chapter: there is no such thing as the causal effect of a treatment. The effect depends on the characteristics of the population under study.
Table 4.1 is Table 1.1 (the counterfactual outcomes in Zeus’s family) plus an indicator \(V\) for sex: \(V = 1\) for women, \(V = 0\) for men.
| Name | \(V\) | \(Y^{a=0}\) | \(Y^{a=1}\) | Name | \(V\) | \(Y^{a=0}\) | \(Y^{a=1}\) |
|---|---|---|---|---|---|---|---|
| Rheia | 1 | 0 | 1 | Kronos | 0 | 1 | 0 |
| Demeter | 1 | 0 | 0 | Hades | 0 | 0 | 0 |
| Hestia | 1 | 0 | 0 | Poseidon | 0 | 1 | 0 |
| Hera | 1 | 0 | 0 | Zeus | 0 | 0 | 1 |
| Artemis | 1 | 1 | 1 | Apollo | 0 | 1 | 0 |
| Leto | 1 | 0 | 1 | Ares | 0 | 1 | 1 |
| Athena | 1 | 1 | 1 | Hephaestus | 0 | 0 | 1 |
| Aphrodite | 1 | 0 | 1 | Polyphemus | 0 | 0 | 1 |
| Persephone | 1 | 1 | 1 | Hermes | 0 | 1 | 0 |
| Hebe | 1 | 1 | 0 | Dionysus | 0 | 1 | 0 |
Example 1 (Sex-specific effects in Table 4.1) Women (\(V = 1\), first 10 rows):
Men (\(V = 0\), last 10 rows):
A null average effect in the population does not imply a null effect in every subset. Here the two sex-specific effects are equal in magnitude, opposite in sign, and each sex is 50% of the population, so they cancel exactly.
Definition 1 (Effect modifier) \(V\) is a modifier of the effect of \(A\) on \(Y\) when the average causal effect of \(A\) on \(Y\) varies across levels of \(V\). For binary \(V\):
Only variables \(V\) not affected by treatment \(A\) are considered as effect modifiers.
In Table 4.1, sex modifies the effect on both scales, and the effects in the two strata go in opposite directions: this is qualitative effect modification.
With qualitative effect modification, additive implies multiplicative effect modification and vice versa. Without it, there can be modification on one scale but not the other.
Example 2 (Multiplicative but not additive effect modification) In a second study:
| \(\Pr[Y^{a=0} = 1 \mid V = v]\) | \(\Pr[Y^{a=1} = 1 \mid V = v]\) | risk difference | risk ratio | |
|---|---|---|---|---|
| \(V = 1\) | 0.8 | 0.9 | \(0.9 - 0.8 = 0.1\) | \(0.9/0.8 = 1.125\) |
| \(V = 0\) | 0.1 | 0.2 | \(0.2 - 0.1 = 0.1\) | \(0.2/0.1 = 2\) |
The risk differences are equal (no additive effect modification), but the risk ratios differ (multiplicative effect modification).
Because “effect modification” means nothing without naming the effect measure, some authors say effect-measure modification instead.
A stratified analysis is the natural way to identify effect modification by measured variables: compute the effect of \(A\) on \(Y\) in each level (stratum) of \(V\). For dichotomous \(V\), the stratified causal risk differences are
\[ \Pr[Y^{a=1} = 1 \mid V = 1] - \Pr[Y^{a=0} = 1 \mid V = 1] \quad\text{and}\quad \Pr[Y^{a=1} = 1 \mid V = 0] - \Pr[Y^{a=0} = 1 \mid V = 0]. \]
In real data we see \(A\) and \(Y\), not \(Y^{a=1}\) and \(Y^{a=0}\). How we stratify depends on the design.
If treatment is randomly and unconditionally assigned, exchangeability is expected in every subset of the population. So within each stratum the associational risk difference equals the causal one, e.g.,
\[ \Pr[Y = 1 \mid A = 1, V = 1] - \Pr[Y = 1 \mid A = 0, V = 1] = \Pr[Y^{a=1} = 1 \mid V = 1] - \Pr[Y^{a=0} = 1 \mid V = 1], \]
and a stratified analysis of association measures identifies effect modification.
A population of 40: transplant \(A\) randomly assigned with probability 0.75 if in severe condition (\(L = 1\)) and 0.50 otherwise (\(L = 0\)). Nationality \(V\): 20 Greeks (\(V = 1\), data in Table 3.1) and 20 Romans (\(V = 0\), Table 4.2).
Table 4.2 (Romans, \(V = 0\))
| Name | \(L\) | \(A\) | \(Y\) | Name | \(L\) | \(A\) | \(Y\) |
|---|---|---|---|---|---|---|---|
| Cybele | 0 | 0 | 0 | Latona | 1 | 0 | 0 |
| Saturn | 0 | 0 | 1 | Mars | 1 | 1 | 1 |
| Ceres | 0 | 0 | 0 | Minerva | 1 | 1 | 1 |
| Pluto | 0 | 0 | 0 | Vulcan | 1 | 1 | 1 |
| Vesta | 0 | 1 | 0 | Venus | 1 | 1 | 1 |
| Neptune | 0 | 1 | 0 | Seneca | 1 | 1 | 1 |
| Juno | 0 | 1 | 1 | Proserpina | 1 | 1 | 1 |
| Jupiter | 0 | 1 | 1 | Mercury | 1 | 1 | 0 |
| Diana | 1 | 0 | 0 | Juventas | 1 | 1 | 0 |
| Phoebus | 1 | 0 | 1 | Bacchus | 1 | 1 | 0 |
In the whole population of 40, \(\Pr[Y^{a=1} = 1] = 0.55\) and \(\Pr[Y^{a=0} = 1] = 0.40\): a risk difference of \(0.15\) and a risk ratio of \(0.55/0.40 = 1.375\).
To estimate \(\Pr[Y^{a} = 1 \mid V = v]\) in each stratum, use two stages:
Step 2 can be skipped when \(V\) is itself the set of variables needed for conditional exchangeability (Section 4.4).
Example 3 (Effect modification by nationality) Greeks (\(V = 1\)): as computed in Chapter 2, causal risk difference 0, risk ratio 1.
Romans (\(V = 0\)): from Table 4.2, 8 Romans have \(L = 0\) and 12 have \(L = 1\), so \(\Pr[L = 0 \mid V = 0] = 0.4\) and \(\Pr[L = 1 \mid V = 0] = 0.6\), and
\[ \begin{aligned} \Pr[Y^{a=1} = 1 \mid V = 0] &= \Pr[Y = 1 \mid A = 1, L = 0, V = 0](0.4) + \Pr[Y = 1 \mid A = 1, L = 1, V = 0](0.6) \\ &= (2/4)(0.4) + (6/9)(0.6) = 0.2 + 0.4 = 0.6, \\ \Pr[Y^{a=0} = 1 \mid V = 0] &= \Pr[Y = 1 \mid A = 0, L = 0, V = 0](0.4) + \Pr[Y = 1 \mid A = 0, L = 1, V = 0](0.6) \\ &= (1/4)(0.4) + (1/3)(0.6) = 0.1 + 0.2 = 0.3. \end{aligned} \]
Causal risk difference \(0.6 - 0.3 = 0.3\), risk ratio \(0.6/0.3 = 2\).
The effects differ between strata, so nationality modifies the effect on both the additive and multiplicative scales. The modification is not qualitative: the effect is harmful or null in both strata.
Nationality may only be a marker for the factor truly responsible. If heart surgery is better in Greece than in Rome, an intervention that improved surgery in Rome could remove the modification by passport-defined nationality.
“Effect modification by \(V\)” does not imply that \(V\) plays a causal role; some authors prefer the more neutral “effect heterogeneity across strata of \(V\)”. Chapter 5’s interaction does attribute a causal role to the variables involved.
Fine Point 4.1: Effect in the treated
The average causal effect can also be defined in one particular subset, the treated (\(A = 1\)). By consistency, the average causal effect in the treated is non-null if
\[\Pr[Y = 1 \mid A = 1] \neq \Pr[Y^{a=0} = 1 \mid A = 1].\]
The causal risk difference in the treated is \(\Pr[Y = 1 \mid A = 1] - \Pr[Y^{a=0} = 1 \mid A = 1]\), and the causal risk ratio in the treated, also called the standardized morbidity ratio (SMR), is \(\Pr[Y = 1 \mid A = 1] / \Pr[Y^{a=0} = 1 \mid A = 1]\); effects in the untreated replace \(A = 1\) by \(A = 0\) (book Figure 4.1).
The effect in the treated differs from the population effect if the distribution of individual effects differs between treated and untreated. Treatment group then acts as a marker for the true modifiers, but we would not say the effect of \(A\) is “modified by \(A\)”; the modification is by unidentified variables distributed differently between treatment groups. The book focuses on population effects because effects in the treated or untreated do not generalize directly to time-varying treatments (Part III).
Reasons to identify effect modification, and to collect pre-treatment descriptors \(V\) even in randomized experiments:
In Table 4.1 the effect is harmful in women and beneficial in men, and the population effect is null only because each sex is 50%. In a population with more women (e.g., graduating college students), the average effect would be harmful.
There is no such thing as “the average causal effect of \(A\) on \(Y\) (period)”, only “the average causal effect of \(A\) on \(Y\) in a population with a particular mix of causal effect modifiers.”
Extrapolating an effect from one population to another is transportability (lack of it is sometimes called lack of external validity). The effect in the population of Greeks and Romans may not transport to a population with a different mix of sex and nationality.
Fine Point 4.2: Transportability
Effects may be transported if the two populations are similar with respect to:
Stratum-specific effects can be combined in a weighted average to reconstruct the effect in a target population, with weights equal to the target population’s proportions in each stratum (e.g., Roman women, Greek women, Roman men, Greek men). The result still need not equal the true target effect because of unmeasured modifiers, interference patterns, and treatment definitions. Methods for transporting effects are discussed by Westreich et al. (2017), Rudolph and van der Laan (2017), and Dahabreh et al. (2020b).
Technical Point 4.1: Computing the effect in the treated
The population effect required \(Y^a \perp\!\!\!\perp A \mid L\) for both \(a = 0\) and \(a = 1\). The effect in the treated needs only partial exchangeability \(Y^{a=0} \perp\!\!\!\perp A \mid L\); the effect in the untreated needs only \(Y^{a=1} \perp\!\!\!\perp A \mid L\). Counterfactual means of the form \(\operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{}\) can be computed by
If physicians knew of the qualitative modification by sex in Table 4.1, then, lacking other information, they would treat the next patient only if he is a man.
In the second study (Example 2), treatment changes the risk by the same 0.1 in both strata, so treating everyone would change risk equally in both strata, despite the multiplicative effect modification.
Effect modification may give clues about biological, social, or other mechanisms: for example, a greater risk of HIV infection in uncircumcised compared with circumcised men. It can also be a first step towards characterizing interactions between two treatments, a related but distinct causal concept (Chapter 5).
Two different goals, two different tools:
In practice, stratification is often used instead of standardization to adjust for \(L\), so much so that many investigators treat “stratification” and “adjustment” as synonyms.
Example 4 (Stratification as adjustment in Table 2.2) Asked to “adjust for \(L\)”, an epidemiologist splits Table 2.2 (the same data as Table 3.1) into \(L = 0\) and \(L = 1\) and reports
\[\Pr[Y = 1 \mid A = 1, L = l] / \Pr[Y = 1 \mid A = 0, L = l] = 1 \quad \text{for } l = 0, 1\]
(\(1/4 \div 1/4\) for \(l = 0\) and \(2/3 \div 2/3\) for \(l = 1\)). Under \(Y^a \perp\!\!\!\perp A \mid L\), each is the causal risk ratio in its stratum, because then \(\Pr[Y = 1 \mid A = a, L = l] = \Pr[Y^a = 1 \mid L = l]\).
So stratification forces one to evaluate effect modification by all variables in \(L\), whether or not that is of interest. Stratifying by \(V\) and then standardizing or IP weighting for \(L\) handles exchangeability and effect modification separately.
Computing the effect in only some strata of \(L\) is restriction. Stratification is restriction applied to several comprehensive, mutually exclusive subsets, with exchangeability in each. When positivity fails in some strata, restriction limits the inference to the strata where positivity holds (Chapter 3).
The goal of matching is to build a subset of the population in which \(L\) has the same distribution among the treated and the untreated.
Example 5 (Matching in Table 2.2) For each untreated individual, randomly select a treated individual with the same \(L\) (the matching factor). One possible set of 7 matched pairs (untreated first):
All untreated but only some treated are selected. In the matched population, the proportion with \(L = 1\) is \(3/7\) in both groups by design. The risk is \(3/7\) in the treated (Zeus, Ares, Aphrodite died) and \(3/7\) in the untreated (Kronos, Artemis, Apollo died), so the causal risk ratio is \(1\).
Because the matched population is a subset of the original one, its distribution of effect modifiers generally differs, which is the topic of Section 4.6.
Standardization, IP weighting, stratification/restriction, and matching estimate different types of causal effects:
All four require exchangeability and positivity, but only in the subset to which the effect refers: for the effect among \(L = l\), only in that subset; for the marginal effect, in all levels of \(L\).
Without effect modification all four give the same answer: in Table 2.2 the causal risk ratio was 1 in the population, in each stratum of \(L\), and in the untreated.
Outcome \(Z\) = high blood pressure. Assume \(Z^a \perp\!\!\!\perp A \mid L\) and positivity.
| Name | \(L\) | \(A\) | \(Z\) | Name | \(L\) | \(A\) | \(Z\) |
|---|---|---|---|---|---|---|---|
| Rheia | 0 | 0 | 0 | Leto | 1 | 0 | 0 |
| Kronos | 0 | 0 | 1 | Ares | 1 | 1 | 1 |
| Demeter | 0 | 0 | 0 | Athena | 1 | 1 | 1 |
| Hades | 0 | 0 | 0 | Hephaestus | 1 | 1 | 1 |
| Hestia | 0 | 1 | 0 | Aphrodite | 1 | 1 | 0 |
| Poseidon | 0 | 1 | 0 | Polyphemus | 1 | 1 | 0 |
| Hera | 0 | 1 | 1 | Persephone | 1 | 1 | 0 |
| Zeus | 0 | 1 | 1 | Hermes | 1 | 1 | 0 |
| Artemis | 1 | 0 | 1 | Hebe | 1 | 1 | 0 |
| Apollo | 1 | 0 | 1 | Dionysus | 1 | 1 | 0 |
Example 6 (Four causal risk ratios from Table 4.3) Stratum-specific risks: \(\Pr[Z = 1 \mid A = 1, L = 0] = 2/4\), \(\Pr[Z = 1 \mid A = 0, L = 0] = 1/4\), \(\Pr[Z = 1 \mid A = 1, L = 1] = 3/9 = 1/3\), \(\Pr[Z = 1 \mid A = 0, L = 1] = 2/3\); \(\Pr[L = 0] = 0.4\).
Four different numbers, 0.8, 2.0, 0.5, and 1.0, and all of them are correct.
The explanation (leaving aside random variability) is qualitative effect modification by \(L\): treatment doubles the risk when \(L = 0\) and halves it when \(L = 1\).
The lesson: always specify the population, or subset, to which an effect measure corresponds.
Proposition 1 (When is the marginal risk ratio below 1? (Technical Point 4.3)) Let \(\operatorname{RR}_l \stackrel{\text{def}}{=}\Pr[Z^{a=1} = 1 \mid L = l] / \Pr[Z^{a=0} = 1 \mid L = l]\), \(p_l \stackrel{\text{def}}{=}\Pr[Z^{a=0} = 1 \mid L = l]\), and \(\pi_l \stackrel{\text{def}}{=}\Pr[L = l]\). With \(\operatorname{RR}_0 = 2\) and \(\operatorname{RR}_1 = 0.5\), the marginal risk ratio is below 1 if and only if
\[\frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}.\]
Proof. The marginal risk ratio is a weighted average of the conditional ones:
\[ \begin{aligned} \frac{\Pr[Z^{a=1} = 1]}{\Pr[Z^{a=0} = 1]} &= \frac{\sum_l \Pr[Z^{a=1} = 1 \mid L = l]\, \pi_l}{\Pr[Z^{a=0} = 1]} && \text{(law of total probability)} \\ &= \sum_l \operatorname{RR}_l \frac{p_l \pi_l}{\Pr[Z^{a=0} = 1]} && \text{(multiply and divide by } p_l\text{)} \\ &= \sum_l \operatorname{RR}_l\, w(l), && w(l) \stackrel{\text{def}}{=}\frac{p_l \pi_l}{\Pr[Z^{a=0} = 1]},\ \textstyle\sum_l w(l) = 1. \end{aligned} \]
Substituting \(\operatorname{RR}_0 = 2\), \(\operatorname{RR}_1 = 0.5\), and \(w(1) = 1 - w(0)\):
\[ \begin{aligned} 2\, w(0) + 0.5\,(1 - w(0)) < 1 &\iff 1.5\, w(0) < 0.5 \\ &\iff w(0) < 1/3 \\ &\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{; } \Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1\text{)} \\ &\iff 2 p_0 \pi_0 < p_1 \pi_1 \\ &\iff \frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}. \end{aligned} \]
In Table 4.3, \(w(0) = (1/4)(0.4)/0.5 = 0.2 < 1/3\), and indeed \(2(0.2) + 0.5(0.8) = 0.8\).
Chapter 3 argued that a well-defined causal effect is a prerequisite for meaningful causal inference. A well-characterized target population is another. Experiments that compare interventions in a population meeting a priori eligibility criteria have both automatically; observational studies must define them explicitly.
Fine Point 4.3: Collapsibility and the odds ratio
An effect measure is collapsible when its population value is a weighted average of its stratum-specific values. In follow-up studies the risk ratio and risk difference are collapsible, but the odds ratio (and the rarely used odds difference) is not (Greenland 1987).
Table 4.4: effect of moving to high altitude \(A\) (top of Mount Olympus) on depression \(Y\); \(V = 1\) for women. The decision to move was random, so \(Y^a \perp\!\!\!\perp A\).
| Name | \(V\) | \(A\) | \(Y\) | Name | \(V\) | \(A\) | \(Y\) |
|---|---|---|---|---|---|---|---|
| Rheia | 1 | 0 | 0 | Kronos | 0 | 0 | 0 |
| Demeter | 1 | 0 | 0 | Hades | 0 | 0 | 0 |
| Hestia | 1 | 0 | 0 | Poseidon | 0 | 0 | 1 |
| Hera | 1 | 0 | 0 | Zeus | 0 | 0 | 1 |
| Artemis | 1 | 0 | 1 | Apollo | 0 | 0 | 0 |
| Leto | 1 | 1 | 0 | Ares | 0 | 1 | 1 |
| Athena | 1 | 1 | 1 | Hephaestus | 0 | 1 | 1 |
| Aphrodite | 1 | 1 | 1 | Polyphemus | 0 | 1 | 1 |
| Persephone | 1 | 1 | 0 | Hermes | 0 | 1 | 0 |
| Hebe | 1 | 1 | 1 | Dionysus | 0 | 1 | 1 |
Example 7 (The odds ratio is not collapsible)
| risk if \(A = 1\) | risk if \(A = 0\) | risk ratio | odds ratio | |
|---|---|---|---|---|
| men (\(V = 0\)) | \(4/5\) | \(2/5\) | \(2\) | \(\dfrac{4/1}{2/3} = 6\) |
| women (\(V = 1\)) | \(3/5\) | \(1/5\) | \(3\) | \(\dfrac{3/2}{1/4} = 6\) |
| population | \(7/10\) | \(3/10\) | \(7/3 \approx 2.3\) | \(\dfrac{7/3}{3/7} = 49/9 \approx 5.4\) |
The population risk ratio (2.3) lies between the stratum-specific ones (2 and 3). The population odds ratio (5.4) is closer to the null than both stratum-specific odds ratios (6), even though there is no confounding.