Chapter 5: Interaction
Chapter 4 introduced effect modification: the effect of a single treatment \(A\) varying across levels of another variable \(V\). Many causal questions, however, concern two or more treatments applied together. If the causal effect of one treatment depends on the value we set for the other, the two treatments interact.
This chapter defines interaction between two treatments in two frameworks:
- the counterfactual (potential outcomes) framework, which we have used so far;
- the sufficient-component-cause framework, which describes causal mechanisms.
This chapter is based on Hernán and Robins (2020, chap. 5, pp. 61-74).
The book motivates the chapter with the “looking up” experiment from Chapter 1: suppose we randomize not only whether you look up at the sky but also whether you stand in the street dressed or naked. If the effect of looking up differs between the dressed and naked conditions, the two treatments interact. When joint interventions are feasible, knowing about interaction tells us which combination of interventions works best.
1 5.1 Interaction Requires a Joint Intervention (pp. 61-62)
1.1 Joint Interventions
Extend the heart transplant example with a second treatment:
- \(A\): heart transplant (\(A=1\)) or not (\(A=0\));
- \(E\): multivitamin complex (\(E=1\)) or no vitamins (\(E=0\)).
Each individual now has four counterfactual outcomes, \(Y^{a=1,e=1}\), \(Y^{a=1,e=0}\), \(Y^{a=0,e=1}\), and \(Y^{a=0,e=0}\).
Recursive substitution (Hernán and Robins 2020, 61, margin note). The counterfactual \(Y^a\) under an intervention on \(A\) alone is the joint counterfactual \(Y^{a,e}\) evaluated at the value \(e\) that \(E\) actually takes: \(Y^a = Y^{a,E}\). Consistency is a special case of the same substitution: the observed outcome is \(Y = Y^A = Y^{A,E}\).
1.2 Interaction in the Counterfactual Framework
1.3 \(A\) and \(E\) Have Equal Status
Write \(p_{ae} \stackrel{\text{def}}{=}\Pr[Y^{a,e}=1]\). Moving terms across the inequality in Definition 2 shows that the same condition can be stated with the roles of \(A\) and \(E\) swapped:
\[ \begin{aligned} p_{11} - p_{01} &\neq p_{10} - p_{00} \\ p_{11} - p_{01} - p_{10} + p_{00} &\neq 0 && \text{(subtract } p_{10} - p_{00} \text{ from both sides)} \\ p_{11} - p_{10} &\neq p_{01} - p_{00} && \text{(add } p_{01} - p_{00} \text{ to both sides)} \end{aligned} \]
The last line says that the causal risk difference for vitamins \(E\) differs between “everybody transplanted” and “nobody transplanted”.
In Example 1, the middle line gives \(p_{11} - p_{01} - p_{10} + p_{00} = 0.1 - 0.2 = -0.1\), so the last line gives \(p_{11} - p_{10} = (p_{01} - p_{00}) - 0.1\): the risk difference for vitamins is \(0.1\) lower among the transplanted than among the untransplanted, which is the book’s statement that the inequality for \(A\) implies the corresponding inequality for \(E\) (Hernán and Robins 2020, 62).
No interaction on the additive scale means the joint effect is the sum of the two separate effects:
\[ p_{11} - p_{00} = (p_{10} - p_{00}) + (p_{01} - p_{00}). \]
When the two sides differ, the interaction is
- superadditive if the left side is greater (\(>\));
- subadditive if the left side is smaller (\(<\)).
On the multiplicative scale (causal risk ratios), \(A\) and \(E\) interact if
\[ \frac{p_{11}}{p_{00}} \neq \frac{p_{10}}{p_{00}} \times \frac{p_{01}}{p_{00}}, \]
with supermultiplicative (\(>\)) and submultiplicative (\(<\)) interaction defined analogously.
The additive form follows from the definition in two steps (Hernán and Robins 2020, Technical Point 5.1, p. 63). No interaction means \(p_{11} - p_{01} = p_{10} - p_{00}\). Adding \(p_{01}\) to both sides gives \(p_{11} = p_{10} - p_{00} + p_{01}\). Subtracting \(p_{00}\) from both sides gives \(p_{11} - p_{00} = (p_{10} - p_{00}) + (p_{01} - p_{00})\), the displayed equality.
Subtracting \((p_{10} - p_{00}) + (p_{01} - p_{00})\) from \(p_{11} - p_{00}\) leaves \(p_{11} - p_{10} - p_{01} + p_{00}\), so in Example 1 (where that quantity is \(-0.1\)) the interaction is subadditive.
1.4 Interaction versus Effect Modification
| Effect modification by \(V\) | Interaction between \(A\) and \(E\) | |
|---|---|---|
| Counterfactuals involved | \(Y^a\) | \(Y^{a,e}\) |
| Interventions | on \(A\) only | joint, on \(A\) and \(E\) |
| Status of the two variables | unequal: \(V\) is not intervened on | equal |
Effect modification is about the causal effect of \(A\) only. In Chapter 4, sex modified the effect of transplant, but the effect of sex on death was never considered, so \(V\) and \(A\) were not on an equal footing. Interaction concerns the joint causal effect of \(A\) and \(E\), and its two equivalent definitions treat \(A\) and \(E\) symmetrically (Hernán and Robins 2020, 62).
2 5.2 Identifying Interaction (pp. 62-64)
Interaction concerns a joint effect, so identifying it requires exchangeability, positivity, and consistency for both treatments.
2.1 When \(E\) Is Randomized, Interaction Is Effect Modification by \(E\)
If vitamins \(E\) are randomly and unconditionally assigned, then for each \(a\) and \(e\)
\[ \begin{aligned} \Pr[Y^{a,e}=1] &= \Pr[Y^{a,e}=1 \mid E=e] && \text{(exchangeability for } E\text{: } Y^{a,e} \perp\!\!\!\perp E\text{)} \\ &= \Pr[Y^{a,E}=1 \mid E=e] && \text{(}Y^{a,e} = Y^{a,E}\text{ when } E=e\text{)} \\ &= \Pr[Y^{a}=1 \mid E=e] && \text{(recursive substitution, } Y^a = Y^{a,E}\text{)}. \end{aligned} \]
Substituting into Definition 2, interaction on the additive scale becomes
\[ \Pr[Y^{a=1}=1 \mid E=1] - \Pr[Y^{a=0}=1 \mid E=1] \neq \Pr[Y^{a=1}=1 \mid E=0] - \Pr[Y^{a=0}=1 \mid E=0], \]
which is the definition of additive effect modification of \(A\) by \(E\).
So when \(E\) is randomized, interaction and effect modification coincide, and the Chapter 4 methods for detecting effect modification by \(V\) apply after replacing \(V\) by \(E\) (Hernán and Robins 2020, 62–63).
2.2 When \(E\) Is Not Randomized
We still need the four marginal risks \(\Pr[Y^{a,e}=1]\). Under the usual identifying assumptions for both treatments, they can be computed by standardization or IP weighting over the measured covariates.
Equivalently, treat \(AE\) as one treatment with four levels (11, 01, 10, 00). Identifying interaction is then the familiar problem of identifying the effect of a single treatment, with more treatment values and more counterfactual outcomes.
Hernán and Robins (2020, 64) states that exchangeability, positivity, and consistency for the joint treatment \((A, E)\) are a sufficient condition for identifying interaction. Part III of the book shows the condition is not necessary when the two treatments occur at different times. After this chapter, most of Parts I and II return to a single treatment \(A\).
2.3 Effect Modification Without Interaction
If we are willing to assume exchangeability for \(A\) but not for \(E\) (e.g., \(A\) is randomized and we look at subgroups defined by \(E\)), we can assess effect modification by \(E\) but generally not interaction between \(A\) and \(E\): computing the effect of \(A\) within strata of \(E\) requires no assumptions about the effect of \(E\).
- Chapter 4 wrote such variables as \(V\) (no identifying assumptions made for them).
- Nationality \(V\) was a surrogate effect modifier: it does not act on \(Y\), so it does not interact with \(A\) (“no action, no interaction”).
- \(V\) still modifies the effect of \(A\) because it is correlated with a variable that does act on \(Y\) and does interact with \(A\).
So there can be effect modification of \(A\) by a variable without interaction between \(A\) and that variable. The reverse, interaction between \(A\) and \(E\) without modification of the effect of \(A\) by \(E\), is logically possible but probably rare, because it requires dual effects of \(A\) and exact cancellations (Hernán and Robins 2020, 64, margin note citing VanderWeele 2009b).
The results so far do not require deterministic counterfactuals. Section 5.3 does: from there on the chapter assumes deterministic counterfactuals and dichotomous treatments and outcome (Hernán and Robins 2020, 64).
3 5.3 Counterfactual Response Types and Interaction (pp. 64-66)
3.1 One Treatment: Four Response Types
| Type | \(Y^{a=0}\) | \(Y^{a=1}\) | Zeus’s family (Table 1.1) |
|---|---|---|---|
| Doomed | 1 | 1 | Artemis, Athena, Persephone, Ares |
| Helped | 1 | 0 | Hebe, Kronos, Poseidon, Apollo, Hermes, Dionysus |
| Hurt | 0 | 1 | Rheia, Leto, Aphrodite, Zeus, Hephaestus, Polyphemus |
| Immune | 0 | 0 | Demeter, Hestia, Hera, Hades |
A combination of counterfactual responses is called a response pattern or response type.
3.2 Two Treatments: Sixteen Response Types
With two dichotomous treatments each individual has four counterfactual outcomes, so there are \(2^4 = 16\) response types (Miettinen 1982).
| Type | \(Y^{1,1}\) | \(Y^{0,1}\) | \(Y^{1,0}\) | \(Y^{0,0}\) |
|---|---|---|---|---|
| 1 | 1 | 1 | 1 | 1 |
| 2 | 1 | 1 | 1 | 0 |
| 3 | 1 | 1 | 0 | 1 |
| 4 | 1 | 1 | 0 | 0 |
| 5 | 1 | 0 | 1 | 1 |
| 6 | 1 | 0 | 1 | 0 |
| 7 | 1 | 0 | 0 | 1 |
| 8 | 1 | 0 | 0 | 0 |
| 9 | 0 | 1 | 1 | 1 |
| 10 | 0 | 1 | 1 | 0 |
| 11 | 0 | 1 | 0 | 1 |
| 12 | 0 | 1 | 0 | 0 |
| 13 | 0 | 0 | 1 | 1 |
| 14 | 0 | 0 | 1 | 0 |
| 15 | 0 | 0 | 0 | 1 |
| 16 | 0 | 0 | 0 | 0 |
3.3 Types Without Interaction
In six types the effect of each treatment does not depend on the other:
- type 1 (dies under every joint treatment) and type 16 (survives under every joint treatment);
- type 4: dies only if given vitamins (\(E\) alone matters);
- type 13: dies only if not given vitamins;
- type 6: dies only if transplanted (\(A\) alone matters);
- type 11: dies only if not transplanted.
If everyone has one of types 1, 4, 6, 11, 13, 16, there is no interaction between \(A\) and \(E\) on the additive scale.
If everyone were of type 1 or 16, the sharp causal null hypothesis would hold for the joint treatment \((A, E)\) (Hernán and Robins 2020, 65).
3.4 Additive Interaction Requires Interaction Types
Additive interaction implies that some individuals belong to at least one of three classes (Greenland and Poole 1988), each invariant to recoding \(A\) and \(E\):
- outcome under only one of the four joint treatments: types 8, 12, 14, 15;
- outcome under two joint treatments, with the effect of each treatment reversed across levels of the other: types 7, 10;
- outcome under three of the four joint treatments: types 2, 3, 5, 9.
No additive interaction implies either that nobody is in these classes or that equal deviations of opposite sign cancel exactly (e.g., equal proportions of types 7 and 10, or of types 8 and 12).
Miettinen’s 16 types are not invariant to switching the labels “0” and “1”; Greenland and Poole’s three classes are (Hernán and Robins 2020, 65, margin note). Greenland, Lash, and Rothman (2008) discuss such cancellations further.
For one treatment, \(Y^{a=0} > Y^{a=1}\) only for the “helped”. If nobody is helped, every individual has \(Y^{a=1} \geq Y^{a=0}\), and the causal effect of \(A\) on \(Y\) is monotonic.
For two treatments, the effects of \(A\) and \(E\) are monotonic if every \(Y^{a,e}\) is nondecreasing in both \(a\) and \(e\), i.e., nobody has
- \(Y^{a=1,e=1}=0\) and \(Y^{a=0,e=1}=1\);
- \(Y^{a=1,e=1}=0\) and \(Y^{a=1,e=0}=1\);
- \(Y^{a=1,e=0}=0\) and \(Y^{a=0,e=0}=1\);
- \(Y^{a=0,e=1}=0\) and \(Y^{a=0,e=0}=1\).
Source: Hernán and Robins (2020, Technical Point 5.2, p. 66).
Do individuals exist who develop the outcome under both treatments but under neither alone, i.e., with \(Y^{a=1,e=1}=1\) and \(Y^{a=0,e=1}=Y^{a=1,e=0}=0\) (types 7 and 8)? A sufficient condition (VanderWeele and Robins 2007a, 2008) is
\[ p_{11} - p_{01} - p_{10} > 0, \quad\text{equivalently}\quad p_{11} - p_{01} > p_{10}. \]
If the effects are monotonic, a weaker sufficient condition is superadditive interaction:
\[ p_{11} - p_{01} > p_{10} - p_{00}, \]
which then implies type 8 exists (monotonicity rules out type 7).
All three risks can be computed in a randomized experiment on \(A\) and \(E\), so the existence of types 7 and 8 can be checked empirically (Hernán and Robins 2020, Fine Point 5.1, p. 67). Because the first condition is sufficient but not necessary, it can fail even when types 7 and 8 exist, and it is strong enough to miss most such cases. The monotonic-case result was reported by Greenland and Rothman and appears in Greenland, Lash, and Rothman (2008).
In genetics, the existence of type-8 individuals is called compositional epistasis; VanderWeele (2010a) reviews tests for it.
4 5.4 Sufficient Causes (pp. 66-69)
The variety of response types shows that \(A\) is not the only determinant of \(Y\). The sufficient-component-cause framework represents the other determinants as background factors.
By definition, the dichotomous background factors \(U\) cannot be intervened on and cannot be affected by treatment \(A\) (Hernán and Robins 2020, 67, margin note).
4.1 One Treatment: Three Sufficient Causes
In the oversimplified heart transplant example (Hernán and Robins 2020, 66–67):
| Sufficient cause | Treatment component | Background factor (example) |
|---|---|---|
| \(A=1\) with \(U_1=1\) | transplant | allergy to anesthesia |
| \(A=0\) with \(U_2=1\) | no transplant | ejection fraction below 20% |
| \(U_0=1\) | none | pancreatic cancer at study start |
Figure 5.1 of the book draws each sufficient cause as a circle (“causal pie”) divided into its components.
4.2 Effect Size Depends on Background Factors
Compare two populations, identical except that \(U_1=1\) (allergy) has prevalence
- 1% in the first;
- 10% in the second.
Randomize half of each population to \(A=1\). The average causal effect of transplant on death is greater in the second population, because the sufficient cause “\(A=1\) plus \(U_1=1\)” is ten times more common there.
This is the Chapter 4 point that the magnitude of a causal effect depends on the distribution of effect modifiers, made visible with sufficient-component causes (Hernán and Robins 2020, 68).
4.3 Two Treatments: Nine Sufficient Causes
With treatments \(A\) and \(E\) there are 9 possible sufficient causes (Greenland and Poole 1988), with treatment components
- \(A=1\) only;
- \(A=0\) only;
- \(E=1\) only;
- \(E=0\) only;
- \(A=1\) and \(E=1\);
- \(A=1\) and \(E=0\);
- \(A=0\) and \(E=1\);
- \(A=0\) and \(E=0\);
- neither \(A\) nor \(E\).
Each also contains background factors from \(U_1, \ldots, U_8\) and \(U_0\) (Figure 5.2 of the book).
Not all 9 need exist. If vitamins \(E=1\) never kill anyone, whatever \(A\) is, then the 3 sufficient causes with component \(E=1\) are absent; their existence would mean some individuals (e.g., those with \(U_3=1\)) would be killed by vitamins, i.e., saved by withholding them (Hernán and Robins 2020, 69).
5 5.5 Sufficient Cause Interaction (pp. 69-71)
The counterfactual definition of interaction (Definition 2) is a contrast of counterfactual risks. It can be identified in an ideal randomized experiment on \(A\) and \(E\) without any knowledge of the mechanisms by which the treatments act.
A second concept of interaction refers to mechanisms directly.
5.1 Synergism and Antagonism
- Synergism: \(A=1\) and \(E=1\) are components of the same sufficient cause.
- Antagonism: \(A=1\) and \(E=0\) (or \(A=0\) and \(E=1\)) are components of the same sufficient cause.
Antagonism between \(A\) and \(E\) can be viewed as synergism between \(A\) and “no \(E\)” (or between “no \(A\)” and \(E\)).
Rothman (1976) described synergism and antagonism within the sufficient-component-cause framework (Hernán and Robins 2020, 71, margin note).
5.2 Detecting Synergism Without Knowing the Mechanisms
Sufficient cause interaction is defined through mechanisms, yet sometimes it can be detected with no knowledge of them: if the inequalities of Fine Point 5.1 hold, synergism between \(A\) and \(E\) exists.
This is not surprising, for two reasons (Hernán and Robins 2020, 71): response types correspond to sufficient causes (Fine Point 5.2), and the inequalities are sufficient but not necessary, so they can fail even when synergism exists.
| Type | \(Y^{a=0}\) | \(Y^{a=1}\) | Component causes |
|---|---|---|---|
| Doomed | 1 | 1 | \(U_0=1\) or \(\{U_1=1\) and \(U_2=1\}\) |
| Helped | 1 | 0 | \(U_0=0\), \(U_1=0\), \(U_2=1\) |
| Hurt | 0 | 1 | \(U_0=0\), \(U_1=1\), \(U_2=0\) |
| Immune | 0 | 0 | \(U_0=0\), \(U_1=0\), \(U_2=0\) |
Each combination of component causes gives exactly one response type, but a response type can arise from several combinations (e.g., “doomed” arises from any combination with \(U_0=1\), or with \(U_1=1\) and \(U_2=1\)).
Exchangeability in terms of component causes (Hernán and Robins 2020, Fine Point 5.2, p. 70). For a dichotomous treatment and outcome, \(Y^a \perp\!\!\!\perp A\) means \(\Pr[Y^{a=1}=1 \mid A=1] = \Pr[Y^{a=1}=1 \mid A=0]\) and \(\Pr[Y^{a=0}=1 \mid A=1] = \Pr[Y^{a=0}=1 \mid A=0]\).
- Those with \(Y^{a=1}=1\) are the doomed and the hurt, i.e., those with \(U_0=1\) or \(U_1=1\).
- Those with \(Y^{a=0}=1\) are the doomed and the helped, i.e., those with \(U_0=1\) or \(U_2=1\).
Substituting these sets, exchangeability holds when \(\Pr[U_0=1 \text{ or } U_1=1 \mid A=1] = \Pr[U_0=1 \text{ or } U_1=1 \mid A=0]\) and \(\Pr[U_0=1 \text{ or } U_2=1 \mid A=1] = \Pr[U_0=1 \text{ or } U_2=1 \mid A=0]\).
See Greenland and Brumback (2002), Flanders (2006), and VanderWeele and Hernán (2006); VanderWeele and Robins (2008) generalized some results to two or more treatments.
Sufficient cause interaction is often called biologic interaction (Rothman et al. 1980), but it need not involve treatments acting on each other.
6 5.6 Counterfactuals or Sufficient-Component Causes? (pp. 71-74)
The two frameworks answer different questions:
| Sufficient-component causes | Counterfactuals | |
|---|---|---|
| Starts from | a particular effect | a particular cause or intervention |
| Asks | “how does it happen?” | “what happens?” |
| Describes mechanisms | yes | no |
For estimating average causal effects of hypothetical interventions, the subject of the book, the counterfactual framework is the natural one.
The sufficient-component-cause model asks which events might have caused a given effect; the counterfactual model asks what would have happened had a factor been set to a different level than it was (Hernán and Robins 2020, 71). In philosophy the sufficient-component-cause framework goes back to Mackie (1965), whose INUS condition for \(Y\) is an Insufficient but Necessary part of a condition which is itself Unnecessary but exclusively Sufficient for \(Y\). A counterfactual framework of causation was already hinted at by Hume (1748).
6.1 Strengths and Limitations of Sufficient-Component Causes
Useful for teaching, because they illustrate
- why the size of a causal effect depends on the distribution of background factors (effect modifiers);
- how effect modification, interaction, and synergism relate.
Limited for data analysis, because in its classical form the framework
- is deterministic;
- gives conclusions that depend on the coding of the outcome;
- is restricted to dichotomous treatments and outcomes;
- requires large amounts of data to study the fine distinctions it makes.
Extensions to stochastic settings and to categorical or ordinal treatments might widen its use: VanderWeele (2010b) extended it to 3-level treatments, and VanderWeele and Robins (2012) related stochastic counterfactuals to stochastic sufficient causes (Hernán and Robins 2020, 72). The counterfactual framework will likely remain the one most often used, and apparent alternatives such as causal diagrams and decision theory are essentially equivalent to it (Chapter 6).
Suppose the excess fraction (Fine Point 3.5) is 75% for \(A\) and 50% for \(E\). A joint intervention cannot prevent \(75\% + 50\% = 125\%\) of cases: no intervention, single or joint, can prevent more than 100%.
The resolution: an individual like Zeus with \(U_5=1\), treated with \(A=1\) and \(E=1\), would not have been a case had either treatment been withheld, so Zeus is counted in both the 75% and the 50%.
When \(A\) and \(E\) can be components of the same sufficient cause, the fraction of disease attributable to each separately makes little sense (Hernán and Robins 2020, Fine Point 5.4, p. 72). Phenylketonuria example: mental retardation occurs in genetically susceptible individuals who eat certain foods. Removing the foods would prevent all cases, and so would replacing the susceptibility genes, so the cases are both “100% environmental” and “100% genetic”. See Rothman, Greenland, and Lash (2008).
If smoking (\(A=1\)) never prevents heart disease and physical inactivity (\(E=1\)) never prevents heart disease, then no sufficient cause can contain \(A=0\) or \(E=0\).
If a sufficient cause containing \(A=0\) existed, some individuals (e.g., those with \(U_2=1\)) would develop the outcome when unexposed, i.e., treating them (\(A=1\)) would prevent their outcome, contradicting monotonicity; the same argument applies to \(E=0\). Figure 5.3 of the book crosses out the sufficient causes excluded this way (Hernán and Robins 2020, Technical Point 5.3, pp. 73-74).
7 Summary
- Interaction between two treatments \(A\) and \(E\) is defined by joint counterfactuals \(Y^{a,e}\): on the additive scale, \(\Pr[Y^{1,1}=1] - \Pr[Y^{0,1}=1] \neq \Pr[Y^{1,0}=1] - \Pr[Y^{0,0}=1]\), a definition symmetric in \(A\) and \(E\); a multiplicative version uses risk ratios.
- Identification requires exchangeability, positivity, and consistency for the joint treatment \((A, E)\); if \(E\) is randomized, interaction coincides with effect modification by \(E\).
- Effect modification can occur without interaction (surrogate modifiers).
- With deterministic, dichotomous counterfactuals there are 16 response types; additive interaction requires individuals of interaction types, but their presence can cancel out.
- Sufficient cause interaction means \(A\) and \(E\) appear in the same sufficient cause; synergism can sometimes be detected from counterfactual risks alone (Fine Point 5.1).
- The counterfactual framework answers “what happens?” and is the one used in the rest of the book; sufficient-component causes answer “how does it happen?” and are mainly a teaching tool.
Looking ahead: Chapter 6 introduces causal diagrams, which the book describes as essentially equivalent to the counterfactual framework.