Chapter 4 introduced effect modification: the effect of a single treatment \(A\) varying across levels of another variable \(V\). Many causal questions, however, concern two or more treatments applied together. Roughly, two treatments interact if the causal effect of one depends on the value we set for the other; Section 5.1 makes this precise.
This chapter defines interaction between two treatments in two frameworks:
Definition 1 (Joint Intervention and Joint Counterfactual) A joint intervention fixes the values of several treatments simultaneously. For two treatments \(A\) and \(E\), an individual’s joint counterfactual outcome \(Y^{a,e}\) is their outcome in the hypothetical world where \(A\) is fixed at \(a\) and \(E\) at \(e\).
Example 1 (Transplant and Vitamins: Four Joint Counterfactuals) Extend the heart transplant example with a second treatment, assigned before the first:
A joint intervention on \(A\) and \(E\) assigns everyone to one of four treatment combinations, so each individual has four joint counterfactual outcomes: \(Y^{a=1,e=1}\), \(Y^{a=1,e=0}\), \(Y^{a=0,e=1}\), and \(Y^{a=0,e=0}\).
Remark 1 (Recursive Substitution). Suppose \(E\) is not affected by \(A\) (for example, because \(E\) is assigned before \(A\)), and that intervening to set \(E\) to the value it would have taken anyway does not change the outcome. Then an intervention on \(A\) alone leaves each individual’s \(E\) at its actual value, so the counterfactual \(Y^a\) equals the joint counterfactual evaluated at that value:
\[ Y^a = Y^{a,E}. \]
Combining this with consistency for \(A\) (\(Y = Y^A\)) gives \(Y = Y^A = Y^{A,E}\): consistency is the special case of the same substitution in which \(a\) is also the actual value \(A\).
Example 2 (Recursive Substitution for a Vitamin Taker) In Example 1, vitamins are assigned before transplant, so transplant cannot affect them. For an individual who actually took vitamins (\(E=1\)), Remark 1 gives \(Y^{a=1} = Y^{a=1,e=1}\): their outcome had they been transplanted equals their outcome had they been transplanted and given vitamins. For an individual who took no vitamins, \(Y^{a=1} = Y^{a=1,e=0}\) instead.
Definition 2 (Interaction Between Two Treatments) Let \(E\) be a dichotomous treatment. Treatments \(A\) and \(E\) interact for an outcome \(Y\) if the effect of \(A\) on \(Y\) when we also fix \(E\) at 1 is not the same as the effect of \(A\) on \(Y\) when we also fix \(E\) at 0. The definition is completed by choosing an effect measure to compare the two effects, such as the causal risk difference, and whether, and how, the two effects differ can depend on that choice.
Example 3 (Looking Up, Dressed or Naked (Hernán and Robins 2020, 61)) Chapter 2 asked whether your looking up at the sky makes other pedestrians look up too. Now randomize two treatments: whether you look up (\(A\)), and whether you are clothed or naked while you do it (\(E\)). If the effect of your looking up on other pedestrians when you are dressed differs from its effect when you are naked, then looking up and being dressed interact.
Definition 3 (Interaction on the Additive Scale) Let \(A\) and \(E\) be dichotomous treatments and \(Y\) a dichotomous outcome, and write \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) for the risk under the joint intervention that sets \(A\) to \(a\) and \(E\) to \(e\). There is interaction between \(A\) and \(E\) on the additive scale in the population if the causal risk difference for \(A\) when everybody receives \(E\) differs from the causal risk difference for \(A\) when nobody receives \(E\):
\[ \Pr[Y^{a=1,e=1}=1] - \Pr[Y^{a=0,e=1}=1] \neq \Pr[Y^{a=1,e=0}=1] - \Pr[Y^{a=0,e=0}=1], \]
that is, \(p_{11} - p_{01} \neq p_{10} - p_{00}\).
Example 4 (Transplant and Vitamins (Hernán and Robins 2020, 61–62)) Suppose the causal risk difference for transplant is
Because \(0.1 \neq 0.2\), transplant and vitamins interact on the additive scale.
Proposition 1 (Additive Interaction Is Symmetric in \(A\) and \(E\)) For dichotomous treatments \(A\), \(E\) and outcome \(Y\), with \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) (Definition 3), the causal risk difference for \(A\) differs between \(e = 1\) and \(e = 0\) if and only if the causal risk difference for \(E\) differs between \(a = 1\) and \(a = 0\):
\[ p_{11} - p_{01} \neq p_{10} - p_{00} \iff p_{11} - p_{10} \neq p_{01} - p_{00}. \]
Moreover, the two differences of risk differences are equal: \((p_{11} - p_{01}) - (p_{10} - p_{00}) = (p_{11} - p_{10}) - (p_{01} - p_{00})\).
Proof. Both differences of risk differences equal the same interaction contrast:
\[ \begin{aligned} (p_{11} - p_{01}) - (p_{10} - p_{00}) &= p_{11} - p_{01} - p_{10} + p_{00} \\ &= (p_{11} - p_{10}) - (p_{01} - p_{00}). \end{aligned} \]
So one is nonzero exactly when the other is.
Example 5 (The Vitamin Effect Depends on Transplant) In Example 4, \((p_{11} - p_{01}) - (p_{10} - p_{00}) = 0.1 - 0.2 = -0.1\). By Proposition 1, \((p_{11} - p_{10}) - (p_{01} - p_{00}) = -0.1\) too: the causal risk difference for vitamins is \(0.1\) lower had everybody been transplanted than had nobody been transplanted (Hernán and Robins 2020, 62).
Proposition 2 (No Additive Interaction Means Additive Joint Effects (Technical Point 5.1)) For dichotomous treatments \(A\), \(E\) and outcome \(Y\), with \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) (Definition 3), there is no interaction between \(A\) and \(E\) on the additive scale if and only if the joint effect of both treatments is the sum of the effect of \(A\) alone and the effect of \(E\) alone:
\[ p_{11} - p_{00} = (p_{10} - p_{00}) + (p_{01} - p_{00}). \]
Proof. Subtracting the right-hand side from the left-hand side gives
\[ \begin{aligned} (p_{11} - p_{00}) - \left[(p_{10} - p_{00}) + (p_{01} - p_{00})\right] &= p_{11} - p_{10} - p_{01} + p_{00} \\ &= (p_{11} - p_{01}) - (p_{10} - p_{00}). \end{aligned} \]
The displayed equality holds exactly when this difference is zero, which is exactly when \(p_{11} - p_{01} = p_{10} - p_{00}\), that is, when there is no additive interaction.
Definition 4 (Superadditive and Subadditive Interaction (Technical Point 5.1)) For dichotomous treatments \(A\), \(E\) and outcome \(Y\), with \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) (Definition 3), additive interaction between \(A\) and \(E\) is
Example 6 (Transplant and Vitamins Interact Subadditively) In Example 4, the proof of Proposition 2 shows that \((p_{11} - p_{00}) - [(p_{10} - p_{00}) + (p_{01} - p_{00})] = (p_{11} - p_{01}) - (p_{10} - p_{00}) = -0.1 < 0\), so the interaction is subadditive.
Definition 5 (Interaction on the Multiplicative Scale (Technical Point 5.1)) For dichotomous treatments \(A\), \(E\) and outcome \(Y\), with \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) (Definition 3) and \(p_{00} > 0\), there is interaction between \(A\) and \(E\) on the multiplicative scale if the causal risk ratio for both treatments is not the product of the causal risk ratios for each alone:
\[ \frac{p_{11}}{p_{00}} \neq \frac{p_{10}}{p_{00}} \times \frac{p_{01}}{p_{00}}. \]
The interaction is supermultiplicative if the left side is greater and submultiplicative if it is smaller.
Example 7 (Interaction on One Scale but Not the Other) Suppose \(p_{00} = 0.1\), \(p_{10} = 0.2\), \(p_{01} = 0.3\), and \(p_{11} = 0.4\). On the additive scale, \(p_{11} - p_{00} = 0.3\) and \((p_{10} - p_{00}) + (p_{01} - p_{00}) = 0.1 + 0.2 = 0.3\), so there is no additive interaction (Proposition 2). On the multiplicative scale, \(p_{11}/p_{00} = 4\) while \((p_{10}/p_{00}) \times (p_{01}/p_{00}) = 2 \times 3 = 6\), so the interaction is submultiplicative.
Name the Scale
“\(A\) and \(E\) interact” is incomplete without a scale: as Example 7 shows, the same four risks can show no interaction on one scale and interaction on another.
Recall: Effect Modification (Chapter 4)
A variable \(V\) that is not affected by \(A\) modifies the effect of \(A\) on \(Y\) on the additive scale if \(\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid V = 1\right]\mathclose{} \neq \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid V = 0\right]\mathclose{}\) (Chapter 4). The definition uses only the counterfactuals \(Y^a\) for interventions on \(A\).
| Effect modification by \(V\) | Interaction between \(A\) and \(E\) | |
|---|---|---|
| Counterfactuals involved | \(Y^a\) | \(Y^{a,e}\) |
| Interventions | on \(A\) only | joint, on \(A\) and \(E\) |
| Status of the two variables | unequal: \(V\) is not intervened on | equal |
Remark 2 (Unequal versus Equal Status). Effect modification concerns the causal effect of \(A\) only. In Chapter 4, sex modified the effect of transplant, but the effect of sex on death was never considered, so \(V\) and \(A\) were not on an equal footing. Interaction concerns the joint causal effect of \(A\) and \(E\), and Proposition 1 shows that its definition treats \(A\) and \(E\) symmetrically (Hernán and Robins 2020, 62).
Proposition 3 (Identifying Interaction by Standardization) Let \(A\) and \(E\) be dichotomous treatments, \(Y\) a dichotomous outcome, and \(L\) a discrete vector of measured covariates. Suppose that, for every \(a, e \in \{0, 1\}\):
Then each joint counterfactual risk is identified by standardization,
\[ p_{ae} = \Pr[Y^{a,e} = 1] = \sum_l \Pr[Y = 1 \mid A = a, E = e, L = l] \Pr[L = l], \]
so the contrasts in Definition 3 and Definition 5 are identified too.
Proof. For each \(a\) and \(e\), summing over the \(l\) with \(\Pr[L = l] > 0\):
\[ \begin{aligned} \Pr[Y^{a,e} = 1] &= \sum_l \Pr[Y^{a,e} = 1 \mid L = l] \Pr[L = l] && \text{(law of total probability)} \\ &= \sum_l \Pr[Y^{a,e} = 1 \mid A = a, E = e, L = l] \Pr[L = l] && \text{(exchangeability and positivity)} \\ &= \sum_l \Pr[Y = 1 \mid A = a, E = e, L = l] \Pr[L = l] && \text{(consistency)}. \end{aligned} \]
Positivity makes each conditional probability in the second and third lines well defined.
Treat the Pair as One Treatment
Proposition 3 is the single-treatment result of earlier chapters applied to the combined treatment \(AE\) with four levels (11, 01, 10, 00). Standardization or IP weighting for that four-level treatment estimates all four \(p_{ae}\), and any interaction contrast is then a contrast of those four risks.
Remark 3 (Sufficient, Not Necessary). The conditions of Proposition 3 are sufficient for identifying interaction. Part III of the book shows that they are not required once the two treatments are assigned at different times (Hernán and Robins 2020, 64). After this chapter, most of Parts I and II return to a single treatment \(A\).
Proposition 4 (Randomized \(E\): Interaction Equals Effect Modification by \(E\)) Let \(A\) and \(E\) be dichotomous treatments and \(Y\) a dichotomous outcome. Suppose that
Then \(\Pr[Y^{a,e} = 1] = \Pr[Y^a = 1 \mid E = e]\) for every \(a\) and \(e\). Hence there is interaction between \(A\) and \(E\) on the additive scale (Definition 3) if and only if \(E\) modifies the effect of \(A\) on the additive scale:
\[ \Pr[Y^{a=1}=1 \mid E=1] - \Pr[Y^{a=0}=1 \mid E=1] \neq \Pr[Y^{a=1}=1 \mid E=0] - \Pr[Y^{a=0}=1 \mid E=0]. \]
Proof. For each \(a\) and \(e\):
\[ \begin{aligned} \Pr[Y^{a,e}=1] &= \Pr[Y^{a,e}=1 \mid E=e] && \text{(exchangeability for } E\text{: } Y^{a,e} \perp\!\!\!\perp E\text{)} \\ &= \Pr[Y^{a,E}=1 \mid E=e] && \text{(}Y^{a,E} \text{ is } Y^{a,e} \text{ for those with } E=e\text{)} \\ &= \Pr[Y^{a}=1 \mid E=e] && \text{(recursive substitution, } Y^a = Y^{a,E}\text{)}. \end{aligned} \]
Substituting these four equalities into the inequality of Definition 3 gives the displayed inequality.
Example 8 (Randomized Vitamins) Suppose vitamins are randomly assigned, as in Proposition 4, and the risks are those of Example 4. Then among those who received vitamins, the causal risk difference for transplant is \(\Pr[Y^{a=1}=1 \mid E=1] - \Pr[Y^{a=0}=1 \mid E=1] = p_{11} - p_{01} = 0.1\), and among those who did not, it is \(p_{10} - p_{00} = 0.2\). So vitamins modify the effect of transplant, and the Chapter 4 methods for detecting effect modification by \(V\) apply with \(V\) replaced by \(E\) (Hernán and Robins 2020, 62–63).
Remark 4 (Effect Modification Needs Fewer Assumptions than Interaction). Suppose we are willing to assume exchangeability for \(A\) within levels of \(E\) (\(Y^a \perp\!\!\!\perp A \mid E\), with positivity in each stratum) but no identifying assumptions for \(E\), for example when \(A\) is randomized and we estimate its effect within subgroups defined by \(E\). Then we can generally assess effect modification by \(E\) but not interaction between \(A\) and \(E\): the effect of \(A\) within each stratum of \(E\) needs no identifying assumptions about \(E\), whereas the joint counterfactual risks \(p_{ae}\) do. Chapter 4 wrote such variables as \(V\), for which no identifying assumptions are made.
Example 9 (Nationality: Effect Modification Without Interaction) In Chapter 4, nationality \(V\) modified the effect of transplant \(A\), but it was argued to be a surrogate effect modifier: a marker for a factor, such as quality of care, that does act on \(Y\) and does interact with \(A\). Nationality itself does not act on \(Y\), so it does not interact with \(A\) (“no action, no interaction”). So the effect of \(A\) can be modified by a variable that does not interact with \(A\) (Hernán and Robins 2020, 64).
Assumptions for Sections 5.3 to 5.6
None of the results so far requires deterministic counterfactual outcomes. From here to the end of the chapter, counterfactual outcomes are deterministic, and treatments and outcome are dichotomous (Hernán and Robins 2020, 64).
Definition 6 (Response Type) With deterministic counterfactual outcomes, an individual’s response type (or response pattern) is the list of their counterfactual outcomes under every value of the treatment: \((Y^{a=0}, Y^{a=1})\) for one dichotomous treatment \(A\), and \((Y^{a=1,e=1}, Y^{a=0,e=1}, Y^{a=1,e=0}, Y^{a=0,e=0})\) for two dichotomous treatments \(A\) and \(E\).
| Type | \(Y^{a=0}\) | \(Y^{a=1}\) | Zeus’s family (Table 1.1) |
|---|---|---|---|
| Doomed | 1 | 1 | Artemis, Athena, Persephone, Ares |
| Helped | 1 | 0 | Hebe, Kronos, Poseidon, Apollo, Hermes, Dionysus |
| Hurt | 0 | 1 | Rheia, Leto, Aphrodite, Zeus, Hephaestus, Polyphemus |
| Immune | 0 | 0 | Demeter, Hestia, Hera, Hades |
Example 10 (Response Types in Zeus’s Family) With one dichotomous treatment and a dichotomous outcome there are four response types, and all four occur in Zeus’s family from Chapter 1 (Table 2): the doomed die whatever the treatment, the helped die only if untreated, the hurt die only if treated, and the immune survive whatever the treatment.
With two dichotomous treatments each individual has four counterfactual outcomes, so there are \(2^4 = 16\) response types (Miettinen 1982).
Definition 7 (Individual Interaction Contrast) For an individual with deterministic counterfactual outcomes \(Y^{a,e}\), the individual interaction contrast is
\[ c \stackrel{\text{def}}{=}Y^{a=1,e=1} - Y^{a=0,e=1} - Y^{a=1,e=0} + Y^{a=0,e=0}, \]
the individual’s effect of \(A\) when \(E\) is set to 1 minus their effect of \(A\) when \(E\) is set to 0. It depends only on the response type. Because expectations are linear, its population mean is the interaction contrast of Proposition 1: \(\operatorname{E}\mathopen{}\left[c\right]\mathclose{} = p_{11} - p_{01} - p_{10} + p_{00}\).
| Type | \(Y^{1,1}\) | \(Y^{0,1}\) | \(Y^{1,0}\) | \(Y^{0,0}\) | \(c\) |
|---|---|---|---|---|---|
| 1 | 1 | 1 | 1 | 1 | 0 |
| 2 | 1 | 1 | 1 | 0 | \(-1\) |
| 3 | 1 | 1 | 0 | 1 | 1 |
| 4 | 1 | 1 | 0 | 0 | 0 |
| 5 | 1 | 0 | 1 | 1 | 1 |
| 6 | 1 | 0 | 1 | 0 | 0 |
| 7 | 1 | 0 | 0 | 1 | 2 |
| 8 | 1 | 0 | 0 | 0 | 1 |
| 9 | 0 | 1 | 1 | 1 | \(-1\) |
| 10 | 0 | 1 | 1 | 0 | \(-2\) |
| 11 | 0 | 1 | 0 | 1 | 0 |
| 12 | 0 | 1 | 0 | 0 | \(-1\) |
| 13 | 0 | 0 | 1 | 1 | 0 |
| 14 | 0 | 0 | 1 | 0 | \(-1\) |
| 15 | 0 | 0 | 0 | 1 | 1 |
| 16 | 0 | 0 | 0 | 0 | 0 |
Example 11 (Individual Contrasts of Types 7 and 8) An individual with \((Y^{1,1}, Y^{0,1}, Y^{1,0}, Y^{0,0}) = (1, 0, 0, 1)\) (type 7 in Table 3) has \(c = 1 - 0 - 0 + 1 = 2\): transplant kills them if they take vitamins and saves them if they do not. An individual with \((1, 0, 0, 0)\) (type 8) has \(c = 1\): they die only if given both treatments.
Example 12 (Types 1 and 16: No Effect of Either Treatment) Type 1 dies under every joint treatment and type 16 survives under every joint treatment. If everyone were of type 1 or 16, the sharp causal null hypothesis would hold for the joint treatment \((A, E)\) (Hernán and Robins 2020, 65).
Proposition 5 (Six Types Without Interaction) Assume deterministic counterfactuals and dichotomous \(A\), \(E\), and \(Y\). For individuals of types 1, 4, 6, 11, 13, and 16 in Table 3, the effect of each treatment does not depend on the other:
If everyone in the population has one of these six types, there is no interaction between \(A\) and \(E\) on the additive scale.
Proof. The last column of Table 3 shows \(c = 0\) for each of the six types, so \(c = 0\) for every individual. By Definition 7, \(p_{11} - p_{01} - p_{10} + p_{00} = \operatorname{E}\mathopen{}\left[c\right]\mathclose{} = 0\), which is no additive interaction (Proposition 1).
Definition 8 (Interaction Classes) Greenland and Poole (1988) grouped the other ten response types of Table 3 into three classes:
Unlike the 16 types, the three classes do not change when the labels 0 and 1 of \(A\) or \(E\) are swapped.
Example 13 (Classifying Types 7 and 8) The type-8 individual of Example 11 dies only under \((a, e) = (1, 1)\), so belongs to class 1. The type-7 individual dies under \((1, 1)\) and \((0, 0)\): transplant kills them with vitamins and saves them without, and vitamins kill them with transplant and save them without, so they belong to class 2.
Proposition 6 (Additive Interaction and the Interaction Classes) Assume deterministic counterfactuals and dichotomous \(A\), \(E\), and \(Y\), let \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) be the population risk under the joint intervention \((a, e)\), and let \(\pi_k\) be the proportion of the population with response type \(k\) in Table 3. Then
\[ p_{11} - p_{01} - p_{10} + p_{00} = (\pi_3 + \pi_5 + \pi_8 + \pi_{15}) - (\pi_2 + \pi_9 + \pi_{12} + \pi_{14}) + 2 (\pi_7 - \pi_{10}). \]
Consequently:
Proof. By Definition 7, \(p_{11} - p_{01} - p_{10} + p_{00} = \operatorname{E}\mathopen{}\left[c\right]\mathclose{} = \sum_{k=1}^{16} \pi_k c_k\), where \(c_k\) is the value of \(c\) for type \(k\) in the last column of Table 3. Collecting the nonzero \(c_k\) gives the display. The types with \(c_k \neq 0\) are exactly the ten types of the three classes. If no one belongs to a class, every term is zero; so additive interaction requires some \(\pi_k > 0\) in a class.
Example 14 (Cancellation of Interaction Types) Let a fraction \(q\) of the population be type 8, a fraction \(q\) be type 12, and the rest type 16. Reading the risks off Table 3: \(p_{11} = q\) (type 8), \(p_{01} = q\) (type 12), \(p_{10} = 0\), \(p_{00} = 0\). Then
\[ \begin{aligned} p_{11} - p_{01} - p_{10} + p_{00} &= q - q - 0 + 0 \\ &= 0, \end{aligned} \]
so there is no interaction on the additive scale in the population, even though every type-8 and type-12 individual has an effect of \(A\) that depends on \(E\).
No Population Interaction Does Not Mean No Individual Interaction
As Example 14 shows, the absence of additive interaction in the population does not rule out individuals whose effect of \(A\) depends on \(E\).
Exercise 1 (Computing a Population Interaction Contrast) In a population, 30% are of type 2, 10% are of type 3, and the rest are of type 16. Is there additive interaction between \(A\) and \(E\), and if so, is it superadditive or subadditive?
Solution. By Proposition 6, \(p_{11} - p_{01} - p_{10} + p_{00} = \pi_3 - \pi_2 = 0.1 - 0.3 = -0.2\). This is nonzero, so there is additive interaction. It is negative, so by the proof of Proposition 2 the interaction is subadditive (Definition 4).
Definition 9 (Monotonic Effects (Technical Point 5.2)) For one dichotomous treatment \(A\), the causal effect of \(A\) on \(Y\) is monotonic if \(Y^{a=1} \geq Y^{a=0}\) for every individual, that is, if nobody is helped by treatment.
For two dichotomous treatments \(A\) and \(E\), the causal effects of \(A\) and \(E\) on \(Y\) are monotonic if every individual’s \(Y^{a,e}\) is nondecreasing in both \(a\) and \(e\). Equivalently, nobody has
Example 15 (Which Types Are Monotonic?) In Zeus’s family (Table 2), six individuals are helped by transplant, so the effect of transplant is not monotonic there. For two treatments, checking each row of Table 3 against the four forbidden pairs leaves only types 1, 2, 4, 6, 8, and 16: monotonic effects mean that everyone has one of these six types.
Proposition 7 (Conditions for Types 7 and 8 to Exist (Fine Point 5.1)) Assume deterministic counterfactuals and dichotomous \(A\), \(E\), and \(Y\), with \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\) (Definition 3).
Proof. For part 1, an individual with \(Y^{a=1,e=1}=1\) either has \(Y^{a=0,e=1}=1\), or has \(Y^{a=1,e=0}=1\), or is of type 7 or 8. So
\[ \begin{aligned} p_{11} &\leq \Pr[Y^{a=0,e=1}=1] + \Pr[Y^{a=1,e=0}=1] + \Pr[\text{type 7 or 8}] \\ &= p_{01} + p_{10} + \Pr[\text{type 7 or 8}], \end{aligned} \]
and \(p_{11} - p_{01} - p_{10} > 0\) forces \(\Pr[\text{type 7 or 8}] > 0\).
For part 2, under monotonicity only types 1, 2, 4, 6, 8, and 16 occur (Example 15), and their contrasts in Table 3 give \(p_{11} - p_{01} - p_{10} + p_{00} = \pi_8 - \pi_2\) (Proposition 6). Superadditive interaction makes this positive (proof of Proposition 2), so \(\pi_8 > \pi_2 \geq 0\).
Example 16 (Checking for Types 7 and 8 in a Trial) Suppose a randomized experiment on both \(A\) and \(E\), in which the conditions of Proposition 3 hold, gives \(p_{11} = 0.5\), \(p_{01} = 0.2\), and \(p_{10} = 0.1\). Then \(p_{11} - p_{01} - p_{10} = 0.2 > 0\), so by Proposition 7 some individuals would die if given both treatments but not if given either one alone.
Sufficient, Not Necessary
Both conditions in Proposition 7 are sufficient but not necessary. The first is so strong that it can fail in most populations in which types 7 or 8 exist (Hernán and Robins 2020, Fine Point 5.1, p. 67).
Exercise 2 (The Weaker Condition Under Monotonicity) Suppose the effects of \(A\) and \(E\) are monotonic and \(p_{00} = 0.1\), \(p_{10} = 0.2\), \(p_{01} = 0.3\), and \(p_{11} = 0.45\). Does part 1 of Proposition 7 apply? Does part 2?
Solution. Part 1 does not apply: \(p_{11} - p_{01} - p_{10} = 0.45 - 0.3 - 0.2 = -0.05\), which is not positive. Part 2 does: \(p_{11} - p_{01} = 0.15 > 0.1 = p_{10} - p_{00}\), so the interaction is superadditive, and under monotonicity some individuals are of type 8. In fact \(\pi_8 - \pi_2 = 0.45 - 0.3 - 0.2 + 0.1 = 0.05\), so at least 5% are of type 8.
The variety of response types shows that \(A\) is not the only determinant of \(Y\). The sufficient-component-cause framework represents the other determinants explicitly.
Definition 10 (Background Factor) In the sufficient-component-cause framework, a background factor is a dichotomous variable \(U\), other than the treatments, that helps determine the outcome. By definition, background factors cannot be intervened on and are not affected by treatment (Hernán and Robins 2020, 67, margin note).
Example 17 (Allergy to Anesthesia) In an oversimplified version of the heart transplant example, suppose the only way a transplant can cause death is through allergy to anesthesia. The indicator \(U_1\) of allergy to anesthesia is a background factor: we cannot intervene on it, and transplant does not change it.
Definition 11 (Minimal Sufficient Cause) A sufficient cause of an outcome is a collection of conditions, on the treatments and the background factors, such that anyone meeting all of them develops the outcome. It is minimal if no proper subset of its conditions is itself sufficient. Its conditions are its component causes; “sufficient-component causes” refers to the sufficient causes and their components together.
Example 18 (Three Sufficient Causes for One Treatment (Hernán and Robins 2020, 66–67)) Continue Example 17. Suppose also that the only way going without a transplant can cause death is through an ejection fraction below 20% (\(U_2 = 1\)), and that, apart from these two routes, the only cause of death is pancreatic cancer at study start (\(U_0 = 1\)), which kills whatever the treatment. Then death has three minimal sufficient causes:
| Sufficient cause | Treatment component | Background factor |
|---|---|---|
| \(A=1\) with \(U_1=1\) | transplant | allergy to anesthesia |
| \(A=0\) with \(U_2=1\) | no transplant | ejection fraction below 20% |
| \(U_0=1\) | none | pancreatic cancer at study start |
Figure 5.1 of the book draws each sufficient cause as a circle (“causal pie”) divided into its components.
Fine Point 5.2: From Response Types to Component Causes
In Example 18, each combination of background factors determines one response type:
| Type | \(Y^{a=0}\) | \(Y^{a=1}\) | Component causes |
|---|---|---|---|
| Doomed | 1 | 1 | \(U_0=1\) or \(\{U_1=1\) and \(U_2=1\}\) |
| Helped | 1 | 0 | \(U_0=0\), \(U_1=0\), \(U_2=1\) |
| Hurt | 0 | 1 | \(U_0=0\), \(U_1=1\), \(U_2=0\) |
| Immune | 0 | 0 | \(U_0=0\), \(U_1=0\), \(U_2=0\) |
Each combination of component causes gives exactly one response type, but a response type can arise from several combinations (e.g., “doomed” arises from any combination with \(U_0=1\), or with \(U_1=1\) and \(U_2=1\)).
Remark 5 (Exchangeability in Terms of Component Causes). In the setting of Example 18, exchangeability \(Y^a \perp\!\!\!\perp A\) for a dichotomous treatment and outcome means \(\Pr[Y^{a=1}=1 \mid A=1] = \Pr[Y^{a=1}=1 \mid A=0]\) and \(\Pr[Y^{a=0}=1 \mid A=1] = \Pr[Y^{a=0}=1 \mid A=0]\). By Table 5,
So exchangeability holds exactly when \(\Pr[U_0=1 \text{ or } U_1=1 \mid A=1] = \Pr[U_0=1 \text{ or } U_1=1 \mid A=0]\) and \(\Pr[U_0=1 \text{ or } U_2=1 \mid A=1] = \Pr[U_0=1 \text{ or } U_2=1 \mid A=0]\) (Hernán and Robins 2020, Fine Point 5.2, p. 70).
Example 19 (Prevalence of Allergy and the Size of the Effect) Take the sufficient causes of Example 18, and compare two populations that have the same joint distribution of \(U_0\) and \(U_2\), with \(\Pr[U_0 = 0] > 0\), but in which allergy to anesthesia (\(U_1=1\)) has prevalence
In each population, let \(U_1\) be independent of \((U_0, U_2)\). Since \(Y^{a=1}=1\) exactly when \(U_0=1\) or \(U_1=1\), the risk had everyone been transplanted is \(\Pr[Y^{a=1}=1] = \Pr[U_0=1] + \Pr[U_0=0]\Pr[U_1=1]\), while \(\Pr[Y^{a=0}=1] = \Pr[U_0=1 \text{ or } U_2=1]\) is the same in both populations. So the causal risk difference is larger in the second population by \(0.09 \Pr[U_0=0]\): the sufficient cause “\(A=1\) plus \(U_1=1\)” is ten times more common there. A randomized experiment in each population, assigning half to \(A=1\), would estimate this larger effect.
Remark 6 (Nine Possible Sufficient Causes for Two Treatments). With dichotomous treatments \(A\) and \(E\), there are 9 possible sufficient causes (Greenland and Poole 1988). Their treatment components are
Each also contains background factors, from \(U_1, \ldots, U_8\) and \(U_0\) (Figure 5.2 of the book).
Example 20 (Sufficient Causes That Are Absent) Not all 9 sufficient causes need be present. If vitamins (\(E=1\)) never kill anyone, whatever \(A\) is, then the 3 sufficient causes with component \(E=1\) are absent, in the sense that none of them is ever anyone’s only route to death. If one were, the individuals completing it (e.g., those whose only background factor is \(U_3=1\)) would be killed by vitamins, that is, saved by withholding them (Hernán and Robins 2020, 69).
The counterfactual definition of interaction (Definition 3) is a contrast of counterfactual risks. It can be identified in an ideal randomized experiment on \(A\) and \(E\) without any knowledge of the mechanisms by which the treatments act. A second concept of interaction refers to mechanisms directly.
Definition 12 (Sufficient Cause Interaction) There is a sufficient cause interaction between treatments \(A\) and \(E\) in a population if some sufficient cause (Definition 11) has a component involving \(A\) and a component involving \(E\), and at least one individual has all of its background-factor components, so that it would be completed under some joint intervention on \(A\) and \(E\).
Example 21 (Background Factor \(U_5\)) Suppose individuals with \(U_5=1\) develop the outcome when receiving both vitamins and transplant but not when receiving only one of them. Then a sufficient cause interaction exists if anyone has \(U_5=1\). Hence, if some individual has \(Y^{a=1,e=1}=1\) and \(Y^{a=0,e=1}=Y^{a=1,e=0}=0\), a sufficient cause interaction between \(A\) and \(E\) is present (Hernán and Robins 2020, 69–70).
Definition 13 (Synergism and Antagonism) A sufficient cause interaction between \(A\) and \(E\) is
Antagonism between \(A\) and \(E\) can be viewed as synergism between \(A\) and “no \(E\)” (or between “no \(A\)” and \(E\)).
Example 22 (Synergism and Antagonism Between Transplant and Vitamins) The sufficient cause of Example 21, with components \(A=1\), \(E=1\), and \(U_5=1\), is a synergism between transplant and vitamins. If, instead, some individuals died only when transplanted without vitamins, a sufficient cause with components \(A=1\) and \(E=0\) would be present: an antagonism between transplant and vitamins.
Proposition 8 (Detecting Synergism from Counterfactual Risks) Assume deterministic counterfactuals and dichotomous \(A\), \(E\), and \(Y\), and suppose the outcome is generated by sufficient causes: for every individual and every joint treatment \((a, e)\), \(Y^{a,e} = 1\) if and only if at least one sufficient cause is completed under \((a, e)\). If either condition of Proposition 7 holds (the second one together with monotonic effects), then there is synergism between \(A\) and \(E\) (Definition 13).
Proof. By Proposition 7, some individual has \(Y^{a=1,e=1}=1\) and \(Y^{a=0,e=1}=Y^{a=1,e=0}=0\). Under \((a, e) = (1, 1)\), some sufficient cause is completed for this individual. That sufficient cause cannot contain \(A=0\) or \(E=0\), because those components are absent under \((1, 1)\). It cannot lack a component involving \(A\), because it would then also be completed under \((0, 1)\), giving \(Y^{a=0,e=1}=1\). It cannot lack a component involving \(E\), because it would then also be completed under \((1, 0)\), giving \(Y^{a=1,e=0}=1\). So it contains both \(A=1\) and \(E=1\), which is synergism.
Example 23 (Synergism in a Trial) In Example 16, \(p_{11} - p_{01} - p_{10} = 0.2 > 0\). By Proposition 8, there is synergism between \(A\) and \(E\), even though nothing is known about the background factors involved.
Fine Point 5.3: “Biologic” Interaction
Sufficient cause interaction is often called biologic interaction (Rothman et al. 1980), but it need not involve treatments acting on each other.
Definition 14 (Compositional Epistasis) In genetics, compositional epistasis between two genetic factors \(A\) and \(E\) means that some individuals have response type 8 in Table 3: they develop the outcome when both factors are present, and not otherwise.
Example 24 (Two Alleles (Hernán and Robins 2020, Fine Point 5.3, p. 71)) Let \(A\) and \(E\) indicate a harmful mutation in each of the two copies of a gene needed to make a vital protein (VanderWeele and Robins 2007a). Infants with both mutations (\(A=1\), \(E=1\)) lack the protein and die within a week; those with one or no mutation survive. These infants are of type 8, so there is compositional epistasis. There is also synergism, since a sufficient cause of death contains \(A=1\) and \(E=1\), yet the two copies need not act on each other in any physical sense.
The two frameworks answer different questions:
| Sufficient-component causes | Counterfactuals | |
|---|---|---|
| Starts from | a particular effect | a particular cause or intervention |
| Central question | by what mechanisms the outcome came about | what the outcome would be under an intervention |
| Describes mechanisms | yes | no |
For estimating average causal effects of hypothetical interventions, the subject of the book, the counterfactual framework is the natural one.
Remark 7 (Useful for Teaching, Limited for Data Analysis). Sufficient-component causes are useful for teaching, because they illustrate
They are of limited use for data analysis, because in its classical form the framework
Fine Point 5.4: Attributable Fractions Do Not Add
Recall from Fine Point 3.5 that the excess fraction for a treatment is the share of observed cases that would not have occurred had everyone been untreated. Suppose the excess fraction is 75% for \(A\) and 50% for \(E\). A joint intervention cannot prevent \(75\% + 50\% = 125\%\) of cases: no intervention, single or joint, can prevent more than 100%.
Example 25 (Zeus Is Counted Twice) Suppose Zeus has background factor \(U_5=1\) (and no other background factors), so by Example 21 he dies only if he receives both treatments, and suppose he received \(A=1\) and \(E=1\) and died. Withholding either treatment alone would have averted his death. He therefore belongs both to the cases averted by setting \(a = 0\) (the 75% for \(A\)) and to those averted by setting \(e = 0\) (the 50% for \(E\)): he is counted in both excess fractions, which is why they can sum to more than 100%.
Example 26 (Phenylketonuria: 100% Genetic and 100% Environmental) Intellectual disability due to phenylketonuria occurs only in people who carry a genetic susceptibility and whose diet includes certain foods. Removing those foods from the diet would prevent every case, and so would replacing the susceptibility genes. So the excess fraction is 100% for the foods and 100% for the genes. When \(A\) and \(E\) can be components of the same sufficient cause, asking what fraction of disease is attributable to each separately makes little sense (Hernán and Robins 2020, Fine Point 5.4, p. 72).
Remark 8 (Monotonicity Rules Out Some Sufficient Causes (Technical Point 5.3)). Suppose the effects of \(A\) and \(E\) are monotonic (Definition 9), and the outcome is generated by sufficient causes as in Proposition 8. Consider an individual for whom a sufficient cause containing \(A=0\) is completed under some \((a=0, e)\), while no sufficient cause is completed under \((a=1, e)\), for example someone whose only background factor is \(U_2=1\). That individual would have \(Y^{a=0,e}=1\) but \(Y^{a=1,e}=0\): treating them would prevent their outcome, contradicting monotonicity. So under monotonicity, whoever completes a sufficient cause containing \(A=0\) under \((a=0, e)\) also develops the outcome under \((a=1, e)\), through some other sufficient cause. The same argument applies to \(E=0\). Monotonicity therefore rules out every sufficient cause containing \(A=0\) or \(E=0\) that would leave some individual’s outcome preventable by treatment; this is the sense in which such causes cannot be present.
Example 27 (Smoking and Physical Inactivity (Hernán and Robins 2020, Technical Point 5.3, pp. 73-74)) If smoking (\(A=1\)) never prevents heart disease and physical inactivity (\(E=1\)) never prevents heart disease, then, in the sense of Remark 8, no sufficient cause of heart disease contains “not smoking” (\(A=0\)) or “physically active” (\(E=0\)). Figure 5.3 of the book crosses out the sufficient causes excluded this way.