Chapter 4: Effect Modification

So far we have focused on the average causal effect in an entire population. Many causal questions, however, are about subsets of the population: the effect of your looking up in city dwellers versus visitors, say.

Whether to target the whole population or a subset depends on the goal:

  • a nationwide water fluoridation program reaches every household, so the policy maker cares about the population-average effect;
  • variation across subsets matters when an intervention can be targeted, or when findings must be applied to other populations.

The theme of this chapter: there is no such thing as the causal effect of a treatment. The effect depends on the characteristics of the population under study.

1 4.1 Heterogeneity of Treatment Effects (pp. 47-48)

Table 4.1 is Table 1.1 (the counterfactual outcomes in Zeus’s family) plus an indicator \(V\) for sex: \(V = 1\) for women, \(V = 0\) for men.

Name \(V\) \(Y^{a=0}\) \(Y^{a=1}\) Name \(V\) \(Y^{a=0}\) \(Y^{a=1}\)
Rheia 1 0 1 Kronos 0 1 0
Demeter 1 0 0 Hades 0 0 0
Hestia 1 0 0 Poseidon 0 1 0
Hera 1 0 0 Zeus 0 0 1
Artemis 1 1 1 Apollo 0 1 0
Leto 1 0 1 Ares 0 1 1
Athena 1 1 1 Hephaestus 0 0 1
Aphrodite 1 0 1 Polyphemus 0 0 1
Persephone 1 1 1 Hermes 0 1 0
Hebe 1 1 0 Dionysus 0 1 0

Example 1 (Sex-specific effects in Table 4.1) Women (\(V = 1\), first 10 rows):

  • \(\Pr[Y^{a=1} = 1 \mid V = 1] = 6/10 = 0.6\) and \(\Pr[Y^{a=0} = 1 \mid V = 1] = 4/10 = 0.4\)
  • causal risk ratio \(0.6/0.4 = 1.5\); causal risk difference \(0.6 - 0.4 = 0.2\): transplant increases risk.

Men (\(V = 0\), last 10 rows):

  • \(\Pr[Y^{a=1} = 1 \mid V = 0] = 4/10 = 0.4\) and \(\Pr[Y^{a=0} = 1 \mid V = 0] = 6/10 = 0.6\)
  • causal risk ratio \(0.4/0.6 = 2/3\); causal risk difference \(0.4 - 0.6 = -0.2\): transplant decreases risk.

A null average effect in the population does not imply a null effect in every subset. Here the two sex-specific effects are equal in magnitude, opposite in sign, and each sex is 50% of the population, so they cancel exactly.

Effect modification

Definition 1 (Effect modifier) \(V\) is a modifier of the effect of \(A\) on \(Y\) when the average causal effect of \(A\) on \(Y\) varies across levels of \(V\). For binary \(V\):

  • additive effect modification: \(\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid V = 1\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid V = 0\right]\mathclose{}\)
  • multiplicative effect modification: \(\dfrac{\operatorname{E}\mathopen{}\left[Y^{a=1} \mid V = 1\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid V = 1\right]\mathclose{}} \neq \dfrac{\operatorname{E}\mathopen{}\left[Y^{a=1} \mid V = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid V = 0\right]\mathclose{}}\)

Only variables \(V\) not affected by treatment \(A\) are considered as effect modifiers.

In Table 4.1, sex modifies the effect on both scales, and the effects in the two strata go in opposite directions: this is qualitative effect modification.

Effect modification depends on the scale

With qualitative effect modification, additive implies multiplicative effect modification and vice versa. Without it, there can be modification on one scale but not the other.

Example 2 (Multiplicative but not additive effect modification) In a second study:

\(\Pr[Y^{a=0} = 1 \mid V = v]\) \(\Pr[Y^{a=1} = 1 \mid V = v]\) risk difference risk ratio
\(V = 1\) 0.8 0.9 \(0.9 - 0.8 = 0.1\) \(0.9/0.8 = 1.125\)
\(V = 0\) 0.1 0.2 \(0.2 - 0.1 = 0.1\) \(0.2/0.1 = 2\)

The risk differences are equal (no additive effect modification), but the risk ratios differ (multiplicative effect modification).

Because “effect modification” means nothing without naming the effect measure, some authors say effect-measure modification instead.

2 4.2 Stratification to Identify Effect Modification (pp. 49-51)

A stratified analysis is the natural way to identify effect modification by measured variables: compute the effect of \(A\) on \(Y\) in each level (stratum) of \(V\). For dichotomous \(V\), the stratified causal risk differences are

\[ \Pr[Y^{a=1} = 1 \mid V = 1] - \Pr[Y^{a=0} = 1 \mid V = 1] \quad\text{and}\quad \Pr[Y^{a=1} = 1 \mid V = 0] - \Pr[Y^{a=0} = 1 \mid V = 0]. \]

In real data we see \(A\) and \(Y\), not \(Y^{a=1}\) and \(Y^{a=0}\). How we stratify depends on the design.

Marginally randomized experiments

If treatment is randomly and unconditionally assigned, exchangeability is expected in every subset of the population. So within each stratum the associational risk difference equals the causal one, e.g.,

\[ \Pr[Y = 1 \mid A = 1, V = 1] - \Pr[Y = 1 \mid A = 0, V = 1] = \Pr[Y^{a=1} = 1 \mid V = 1] - \Pr[Y^{a=0} = 1 \mid V = 1], \]

and a stratified analysis of association measures identifies effect modification.

Conditionally randomized experiments: Greeks and Romans

A population of 40: transplant \(A\) randomly assigned with probability 0.75 if in severe condition (\(L = 1\)) and 0.50 otherwise (\(L = 0\)). Nationality \(V\): 20 Greeks (\(V = 1\), data in Table 3.1) and 20 Romans (\(V = 0\), Table 4.2).

Table 4.2 (Romans, \(V = 0\))

Name \(L\) \(A\) \(Y\) Name \(L\) \(A\) \(Y\)
Cybele 0 0 0 Latona 1 0 0
Saturn 0 0 1 Mars 1 1 1
Ceres 0 0 0 Minerva 1 1 1
Pluto 0 0 0 Vulcan 1 1 1
Vesta 0 1 0 Venus 1 1 1
Neptune 0 1 0 Seneca 1 1 1
Juno 0 1 1 Proserpina 1 1 1
Jupiter 0 1 1 Mercury 1 1 0
Diana 1 0 0 Juventas 1 1 0
Phoebus 1 0 1 Bacchus 1 1 0

In the whole population of 40, \(\Pr[Y^{a=1} = 1] = 0.55\) and \(\Pr[Y^{a=0} = 1] = 0.40\): a risk difference of \(0.15\) and a risk ratio of \(0.55/0.40 = 1.375\).

To estimate \(\Pr[Y^{a} = 1 \mid V = v]\) in each stratum, use two stages:

  1. stratify by \(V\);
  2. standardize by \(L\) (or, equivalently, IP weight with weights depending on \(L\)) within each stratum.

Step 2 can be skipped when \(V\) is itself the set of variables needed for conditional exchangeability (Section 4.4).

Example 3 (Effect modification by nationality) Greeks (\(V = 1\)): as computed in Chapter 2, causal risk difference 0, risk ratio 1.

Romans (\(V = 0\)): from Table 4.2, 8 Romans have \(L = 0\) and 12 have \(L = 1\), so \(\Pr[L = 0 \mid V = 0] = 0.4\) and \(\Pr[L = 1 \mid V = 0] = 0.6\), and

\[ \begin{aligned} \Pr[Y^{a=1} = 1 \mid V = 0] &= \Pr[Y = 1 \mid A = 1, L = 0, V = 0](0.4) + \Pr[Y = 1 \mid A = 1, L = 1, V = 0](0.6) \\ &= (2/4)(0.4) + (6/9)(0.6) = 0.2 + 0.4 = 0.6, \\ \Pr[Y^{a=0} = 1 \mid V = 0] &= \Pr[Y = 1 \mid A = 0, L = 0, V = 0](0.4) + \Pr[Y = 1 \mid A = 0, L = 1, V = 0](0.6) \\ &= (1/4)(0.4) + (1/3)(0.6) = 0.1 + 0.2 = 0.3. \end{aligned} \]

Causal risk difference \(0.6 - 0.3 = 0.3\), risk ratio \(0.6/0.3 = 2\).

The effects differ between strata, so nationality modifies the effect on both the additive and multiplicative scales. The modification is not qualitative: the effect is harmful or null in both strata.

Surrogate and causal effect modifiers

Nationality may only be a marker for the factor truly responsible. If heart surgery is better in Greece than in Rome, an intervention that improved surgery in Rome could remove the modification by passport-defined nationality.

  • surrogate effect modifier: nationality
  • causal effect modifier: quality of care

“Effect modification by \(V\)” does not imply that \(V\) plays a causal role; some authors prefer the more neutral “effect heterogeneity across strata of \(V\)”. Chapter 5’s interaction does attribute a causal role to the variables involved.

Fine Point 4.1: Effect in the treated

The average causal effect can also be defined in one particular subset, the treated (\(A = 1\)). By consistency, the average causal effect in the treated is non-null if

\[\Pr[Y = 1 \mid A = 1] \neq \Pr[Y^{a=0} = 1 \mid A = 1].\]

The causal risk difference in the treated is \(\Pr[Y = 1 \mid A = 1] - \Pr[Y^{a=0} = 1 \mid A = 1]\), and the causal risk ratio in the treated, also called the standardized morbidity ratio (SMR), is \(\Pr[Y = 1 \mid A = 1] / \Pr[Y^{a=0} = 1 \mid A = 1]\); effects in the untreated replace \(A = 1\) by \(A = 0\) (book Figure 4.1).

The effect in the treated differs from the population effect if the distribution of individual effects differs between treated and untreated. Treatment group then acts as a marker for the true modifiers, but we would not say the effect of \(A\) is “modified by \(A\)”; the modification is by unidentified variables distributed differently between treatment groups. The book focuses on population effects because effects in the treated or untreated do not generalize directly to time-varying treatments (Part III).

3 4.3 Why Care About Effect Modification (pp. 51-53)

Reasons to identify effect modification, and to collect pre-treatment descriptors \(V\) even in randomized experiments:

  1. the average effect differs between populations with different prevalence of \(V\) (transportability);
  2. it identifies who would benefit most from an intervention;
  3. it may help understand mechanisms.

1. Populations with different mixes of modifiers

In Table 4.1 the effect is harmful in women and beneficial in men, and the population effect is null only because each sex is 50%. In a population with more women (e.g., graduating college students), the average effect would be harmful.

  • With non-qualitative effect modification, the magnitude (but not the direction) of the average effect varies across populations, e.g., the effects of asbestos exposure (different in smokers and nonsmokers) or of universal health care (different in low- and high-income families).

There is no such thing as “the average causal effect of \(A\) on \(Y\) (period)”, only “the average causal effect of \(A\) on \(Y\) in a population with a particular mix of causal effect modifiers.”

Transportability

Extrapolating an effect from one population to another is transportability (lack of it is sometimes called lack of external validity). The effect in the population of Greeks and Romans may not transport to a population with a different mix of sex and nationality.

  • Conditional effects within strata of modifiers may transport better, but there is no guarantee: other unmeasured modifiers may be distributed differently.
  • Those unmeasured modifiers are not variables needed for exchangeability; they are just risk factors for the outcome, about which we have much less information.
  • So transportability is harder than identification within one population, and is an unverifiable assumption resting on subject-matter knowledge.

Fine Point 4.2: Transportability

Effects may be transported if the two populations are similar with respect to:

  • effect modification: the distribution of effect modifiers;
  • the definition of treatment: different versions of treatment can have different effects;
  • interference (Fine Point 1.1): e.g., a socially active person may bring friends along to exercise, so contact patterns affect the magnitude of an exercise intervention’s effect.

Stratum-specific effects can be combined in a weighted average to reconstruct the effect in a target population, with weights equal to the target population’s proportions in each stratum (e.g., Roman women, Greek women, Roman men, Greek men). The result still need not equal the true target effect because of unmeasured modifiers, interference patterns, and treatment definitions. Methods for transporting effects are discussed by Westreich et al. (2017), Rudolph and van der Laan (2017), and Dahabreh et al. (2020b).

Technical Point 4.1: Computing the effect in the treated

The population effect required \(Y^a \perp\!\!\!\perp A \mid L\) for both \(a = 0\) and \(a = 1\). The effect in the treated needs only partial exchangeability \(Y^{a=0} \perp\!\!\!\perp A \mid L\); the effect in the untreated needs only \(Y^{a=1} \perp\!\!\!\perp A \mid L\). Counterfactual means of the form \(\operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{}\) can be computed by

  • standardization: \(\operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{} = \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l \mid A = a']\);
  • IP weighting, with weights \(\Pr[A = a' \mid L] / f(A \mid L)\): \[ \operatorname{E}\mathopen{}\left[Y^a \mid A = a'\right]\mathclose{} = \frac{\operatorname{E}\mathopen{}\left[\dfrac{I(A = a)\, Y}{f(A \mid L)} \Pr[A = a' \mid L]\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[\dfrac{I(A = a)}{f(A \mid L)} \Pr[A = a' \mid L]\right]\mathclose{}}. \] For dichotomous \(A\), this equality was derived by Sato and Matsuyama (2003).

2. Who benefits most

If physicians knew of the qualitative modification by sex in Table 4.1, then, lacking other information, they would treat the next patient only if he is a man.

In the second study (Example 2), treatment changes the risk by the same 0.1 in both strata, so treating everyone would change risk equally in both strata, despite the multiplicative effect modification.

  • Additive, not multiplicative, effect modification is the scale for identifying the groups that would benefit most from intervention.
  • If there is a nonzero effect in at least one stratum and \(\Pr[Y^{a=0} = 1 \mid V = v]\) varies with \(v\), effect modification is guaranteed on at least one of the two scales.

3. Mechanisms

Effect modification may give clues about biological, social, or other mechanisms: for example, a greater risk of HIV infection in uncircumcised compared with circumcised men. It can also be a first step towards characterizing interactions between two treatments, a related but distinct causal concept (Chapter 5).

4 4.4 Stratification as a Form of Adjustment (pp. 53-55)

Two different goals, two different tools:

  • standardization (or IP weighting) adjusts for the variables \(L\) needed for conditional exchangeability;
  • stratification by \(V\) identifies effect modification.

In practice, stratification is often used instead of standardization to adjust for \(L\), so much so that many investigators treat “stratification” and “adjustment” as synonyms.

Example 4 (Stratification as adjustment in Table 2.2) Asked to “adjust for \(L\)”, an epidemiologist splits Table 2.2 (the same data as Table 3.1) into \(L = 0\) and \(L = 1\) and reports

\[\Pr[Y = 1 \mid A = 1, L = l] / \Pr[Y = 1 \mid A = 0, L = l] = 1 \quad \text{for } l = 0, 1\]

(\(1/4 \div 1/4\) for \(l = 0\) and \(2/3 \div 2/3\) for \(l = 1\)). Under \(Y^a \perp\!\!\!\perp A \mid L\), each is the causal risk ratio in its stratum, because then \(\Pr[Y = 1 \mid A = a, L = l] = \Pr[Y^a = 1 \mid L = l]\).

Conditional versus marginal effect measures

  • The two stratum-specific risk ratios are conditional effect measures; the risk ratio of 1 from Chapter 2 is a marginal (unconditional) one. All three equal 1 here only because \(L\) does not modify the effect.
  • Each stratum-specific measure describes a nonoverlapping subset; in general none of them is the population effect. That is why Chapter 2 used standardization and IP weighting rather than stratification.
  • Stratification as adjustment must be done within every combination of the variables in \(L\): Romans with \(L = 1\), Greeks with \(L = 1\), Romans with \(L = 0\), Greeks with \(L = 0\). The association in the stratum \(V = 0\) alone does not give the effect in Romans, because \(V\) by itself does not guarantee conditional exchangeability.

So stratification forces one to evaluate effect modification by all variables in \(L\), whether or not that is of interest. Stratifying by \(V\) and then standardizing or IP weighting for \(L\) handles exchangeability and effect modification separately.

Restriction

Computing the effect in only some strata of \(L\) is restriction. Stratification is restriction applied to several comprehensive, mutually exclusive subsets, with exchangeability in each. When positivity fails in some strata, restriction limits the inference to the strata where positivity holds (Chapter 3).

5 4.5 Matching as Another Form of Adjustment (pp. 55-56)

The goal of matching is to build a subset of the population in which \(L\) has the same distribution among the treated and the untreated.

Example 5 (Matching in Table 2.2) For each untreated individual, randomly select a treated individual with the same \(L\) (the matching factor). One possible set of 7 matched pairs (untreated first):

  • \(L = 0\): Rheia-Hestia, Kronos-Poseidon, Demeter-Hera, Hades-Zeus
  • \(L = 1\): Artemis-Ares, Apollo-Aphrodite, Leto-Hermes

All untreated but only some treated are selected. In the matched population, the proportion with \(L = 1\) is \(3/7\) in both groups by design. The risk is \(3/7\) in the treated (Zeus, Ares, Aphrodite died) and \(3/7\) in the untreated (Kronos, Artemis, Apollo died), so the causal risk ratio is \(1\).

Which population does matching describe?

  • The group used as the reference for finding matches defines the population in which the effect is computed. Matching each untreated individual to a treated one estimates the effect in the untreated; with fewer treated than untreated in every stratum of \(L\), one usually estimates the effect in the treated.
  • Matching need not be one-to-one (pairs); it can be one-to-many (matched sets).
  • With several variables in \(L\), matches come from the same stratum of the combination of all of them.
  • Frequency matching creates any chosen distribution of \(L\), e.g., selecting treated individuals so that 70% have \(L = 1\), then doing the same for the untreated.

Because the matched population is a subset of the original one, its distribution of effect modifiers generally differs, which is the topic of Section 4.6.

6 4.6 Effect Modification and Adjustment Methods (pp. 56-60)

Standardization, IP weighting, stratification/restriction, and matching estimate different types of causal effects:

  • standardization and IP weighting can compute marginal or conditional effects;
  • stratification/restriction and matching compute only conditional effects in certain subsets.

All four require exchangeability and positivity, but only in the subset to which the effect refers: for the effect among \(L = l\), only in that subset; for the marginal effect, in all levels of \(L\).

Without effect modification all four give the same answer: in Table 2.2 the causal risk ratio was 1 in the population, in each stratum of \(L\), and in the untreated.

Table 4.3: four correct answers

Outcome \(Z\) = high blood pressure. Assume \(Z^a \perp\!\!\!\perp A \mid L\) and positivity.

Name \(L\) \(A\) \(Z\) Name \(L\) \(A\) \(Z\)
Rheia 0 0 0 Leto 1 0 0
Kronos 0 0 1 Ares 1 1 1
Demeter 0 0 0 Athena 1 1 1
Hades 0 0 0 Hephaestus 1 1 1
Hestia 0 1 0 Aphrodite 1 1 0
Poseidon 0 1 0 Polyphemus 1 1 0
Hera 0 1 1 Persephone 1 1 0
Zeus 0 1 1 Hermes 1 1 0
Artemis 1 0 1 Hebe 1 1 0
Apollo 1 0 1 Dionysus 1 1 0

Example 6 (Four causal risk ratios from Table 4.3) Stratum-specific risks: \(\Pr[Z = 1 \mid A = 1, L = 0] = 2/4\), \(\Pr[Z = 1 \mid A = 0, L = 0] = 1/4\), \(\Pr[Z = 1 \mid A = 1, L = 1] = 3/9 = 1/3\), \(\Pr[Z = 1 \mid A = 0, L = 1] = 2/3\); \(\Pr[L = 0] = 0.4\).

  • Stratification: risk ratio \(\dfrac{2/4}{1/4} = 2.0\) in \(L = 0\) and \(\dfrac{1/3}{2/3} = 0.5\) in \(L = 1\).
  • Standardization / IP weighting (population): \[ \begin{aligned} \Pr[Z^{a=1} = 1] &= (2/4)(0.4) + (1/3)(0.6) = 0.2 + 0.2 = 0.4, \\ \Pr[Z^{a=0} = 1] &= (1/4)(0.4) + (2/3)(0.6) = 0.1 + 0.4 = 0.5, \end{aligned} \] so the risk ratio is \(0.4/0.5 = 0.8\).
  • Matching (the matched pairs of Section 4.5; effect in the untreated): \(\Pr[Z^{a=1} = 1 \mid A = 0] / \Pr[Z = 1 \mid A = 0] = (3/7)/(3/7) = 1.0\).

Four different numbers, 0.8, 2.0, 0.5, and 1.0, and all of them are correct.

The explanation (leaving aside random variability) is qualitative effect modification by \(L\): treatment doubles the risk when \(L = 0\) and halves it when \(L = 1\).

  • The population effect (0.8) is beneficial because the ratio \(\Pr[Z^{a=0} = 1 \mid L = 1] / \Pr[Z^{a=0} = 1 \mid L = 0] = (2/3)/(1/4) = 8/3\) exceeds 2 times the odds of being noncritical, \(2 \Pr[L = 0]/\Pr[L = 1] = 2(0.4/0.6) = 4/3\) (Technical Point 4.3).
  • The effect in the untreated is null (1.0) because the untreated contain a larger proportion of noncritical individuals (\(4/7\)) than the population (\(0.4\)).

The lesson: always specify the population, or subset, to which an effect measure corresponds.

Proposition 1 (When is the marginal risk ratio below 1? (Technical Point 4.3)) Let \(\operatorname{RR}_l \stackrel{\text{def}}{=}\Pr[Z^{a=1} = 1 \mid L = l] / \Pr[Z^{a=0} = 1 \mid L = l]\), \(p_l \stackrel{\text{def}}{=}\Pr[Z^{a=0} = 1 \mid L = l]\), and \(\pi_l \stackrel{\text{def}}{=}\Pr[L = l]\). With \(\operatorname{RR}_0 = 2\) and \(\operatorname{RR}_1 = 0.5\), the marginal risk ratio is below 1 if and only if

\[\frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}.\]

Proof. The marginal risk ratio is a weighted average of the conditional ones:

\[ \begin{aligned} \frac{\Pr[Z^{a=1} = 1]}{\Pr[Z^{a=0} = 1]} &= \frac{\sum_l \Pr[Z^{a=1} = 1 \mid L = l]\, \pi_l}{\Pr[Z^{a=0} = 1]} && \text{(law of total probability)} \\ &= \sum_l \operatorname{RR}_l \frac{p_l \pi_l}{\Pr[Z^{a=0} = 1]} && \text{(multiply and divide by } p_l\text{)} \\ &= \sum_l \operatorname{RR}_l\, w(l), && w(l) \stackrel{\text{def}}{=}\frac{p_l \pi_l}{\Pr[Z^{a=0} = 1]},\ \textstyle\sum_l w(l) = 1. \end{aligned} \]

Substituting \(\operatorname{RR}_0 = 2\), \(\operatorname{RR}_1 = 0.5\), and \(w(1) = 1 - w(0)\):

\[ \begin{aligned} 2\, w(0) + 0.5\,(1 - w(0)) < 1 &\iff 1.5\, w(0) < 0.5 \\ &\iff w(0) < 1/3 \\ &\iff 3 p_0 \pi_0 < p_0 \pi_0 + p_1 \pi_1 && \text{(definition of } w(0)\text{; } \Pr[Z^{a=0} = 1] = p_0 \pi_0 + p_1 \pi_1\text{)} \\ &\iff 2 p_0 \pi_0 < p_1 \pi_1 \\ &\iff \frac{p_1}{p_0} > \frac{2 \pi_0}{\pi_1}. \end{aligned} \]

In Table 4.3, \(w(0) = (1/4)(0.4)/0.5 = 0.2 < 1/3\), and indeed \(2(0.2) + 0.5(0.8) = 0.8\).

A well-characterized target population

Chapter 3 argued that a well-defined causal effect is a prerequisite for meaningful causal inference. A well-characterized target population is another. Experiments that compare interventions in a population meeting a priori eligibility criteria have both automatically; observational studies must define them explicitly.

Fine Point 4.3: Collapsibility and the odds ratio

An effect measure is collapsible when its population value is a weighted average of its stratum-specific values. In follow-up studies the risk ratio and risk difference are collapsible, but the odds ratio (and the rarely used odds difference) is not (Greenland 1987).

Table 4.4: effect of moving to high altitude \(A\) (top of Mount Olympus) on depression \(Y\); \(V = 1\) for women. The decision to move was random, so \(Y^a \perp\!\!\!\perp A\).

Name \(V\) \(A\) \(Y\) Name \(V\) \(A\) \(Y\)
Rheia 1 0 0 Kronos 0 0 0
Demeter 1 0 0 Hades 0 0 0
Hestia 1 0 0 Poseidon 0 0 1
Hera 1 0 0 Zeus 0 0 1
Artemis 1 0 1 Apollo 0 0 0
Leto 1 1 0 Ares 0 1 1
Athena 1 1 1 Hephaestus 0 1 1
Aphrodite 1 1 1 Polyphemus 0 1 1
Persephone 1 1 0 Hermes 0 1 0
Hebe 1 1 1 Dionysus 0 1 1

Example 7 (The odds ratio is not collapsible)  

risk if \(A = 1\) risk if \(A = 0\) risk ratio odds ratio
men (\(V = 0\)) \(4/5\) \(2/5\) \(2\) \(\dfrac{4/1}{2/3} = 6\)
women (\(V = 1\)) \(3/5\) \(1/5\) \(3\) \(\dfrac{3/2}{1/4} = 6\)
population \(7/10\) \(3/10\) \(7/3 \approx 2.3\) \(\dfrac{7/3}{3/7} = 49/9 \approx 5.4\)

The population risk ratio (2.3) lies between the stratum-specific ones (2 and 3). The population odds ratio (5.4) is closer to the null than both stratum-specific odds ratios (6), even though there is no confounding.

7 Summary

  • \(V\) modifies the effect of \(A\) on \(Y\) when the average causal effect varies across levels of \(V\); whether it does depends on the effect measure (additive vs. multiplicative). Qualitative modification (opposite directions) appears on both scales.
  • In a marginally randomized experiment, stratum-specific associations identify stratum-specific effects; otherwise, stratify by \(V\), then standardize or IP weight for \(L\).
  • Effect modifiers may be surrogate rather than causal.
  • Effect modification matters for transportability, for deciding who benefits (on the additive scale), and for mechanisms. Report absolute counterfactual risks in each stratum.
  • Stratification/restriction and matching adjust for \(L\) but estimate conditional effects (in strata, or in the treated/untreated); standardization and IP weighting can estimate marginal effects. Under effect modification these differ, and all can be correct: specify the target population.
  • The odds ratio is noncollapsible.

8 References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.