Example 1 (Looking Up at the Sky, with a Careless Recorder) Suppose a randomized experiment asks whether one’s looking up makes other pedestrians look up, and it finds only a weak association. Randomization rules out confounding, and every pedestrian’s response was recorded, so there is no selection bias. But the collaborator in charge of the records caught only half of the pedestrians who did look up, writing the rest down as “did not look up”, so even a strong effect would appear diluted.
In Example 1, the recording process, not the treatment, weakened the observed association (on the additive scale, as a worked example of ours at the end of Section 9.2 shows). Distortions of this kind, produced by the way the data are measured, are the subject of this chapter. Because measurement errors occur in any study design, randomized or observational, they always need to be considered. This chapter describes the structure of measurement error and then asks what the arrows of a causal diagram mean when some of its variables have no well-defined interventions (“noncausal” diagrams).
Consider an observational study of a cholesterol-lowering drug \(A\) on liver disease \(Y\). The analysis can only use what was recorded about drug use, which need not match the drug use that actually happened.
Definition 1 (Measurement Error) Let \(A\) be the true value of a variable for an individual and \(A^*\) (“A-star”) the value recorded in the study data, the measured (or mismeasured) version of \(A\). That individual’s measurement error of \(A\) is the difference \(A^* - A\) between the mismeasured and the true value. Measurement error of a discrete variable is called misclassification.
Example 2 (Drug Use Abstracted From Medical Records) In the cholesterol-drug study, treat drug use as binary: \(A = 1\) if the individual actually used the drug and \(A = 0\) otherwise. Suppose \(A^*\) is abstracted from medical records, so \(A^* = 1\) if the record shows the drug was prescribed. \(A^*\) can differ from \(A\) in several ways:
Because \(A\) is binary, each of these errors is a misclassification (Definition 1).
The causal diagram in Figure 9.1 has
The factors in \(U_A\) determine the magnitude and direction of the measurement error.
Remark 1 (Why Draw the Error Node). In Figure 9.1, the node \(U_A\) is not needed for reading off independencies: it is not a common cause of two other nodes, and nothing is conditioned on it. The book draws it anyway, so that the sources of measurement error are visible and the diagram can be compared directly with the ones that come next (Hernán and Robins 2020, 126).
Figure 9.2 adds the measured outcome \(Y^*\), with \(Y \rightarrow Y^*\) and \(U_Y \rightarrow Y^*\).
Proposition 1 (With Perfect Data, Association Is Causation in Figure 9.2) Let \(A\) and \(Y\) be binary, and suppose Figure 9.2 is the causal DAG, so that \(A\) has no parents and no backdoor path to \(Y\). Assume consistency and \(0 < \Pr[A = 1] < 1\). Then the associational risk ratio of the true variables equals the causal risk ratio:
\[ \frac{\Pr[Y = 1 \mid A = 1]}{\Pr[Y = 1 \mid A = 0]} = \frac{\Pr[Y^{a=1} = 1]}{\Pr[Y^{a=0} = 1]} \tag{1}\]
whenever \(\Pr[Y^{a=0} = 1] > 0\).
Proof. The empty set satisfies the backdoor criterion (Section 7.2), because there is no backdoor path between \(A\) and \(Y\). The backdoor criterion implies exchangeability even without faithfulness (Technical Point 7.1 in Chapter 7), so \(Y^a \perp\!\!\!\perp A\) for \(a = 0, 1\). For each \(a\), \(\Pr[Y = 1 \mid A = a] = \Pr[Y^a = 1 \mid A = a]\) by consistency, which equals \(\Pr[Y^a = 1]\) by exchangeability. Taking the ratio for \(a = 1\) and \(a = 0\) gives Equation 1.
In practice only \(A^*\) and \(Y^*\) are available, and the measured risk ratio
\[ \frac{\Pr[Y^* = 1 \mid A^* = 1]}{\Pr[Y^* = 1 \mid A^* = 0]} \tag{2}\]
need not equal the causal risk ratio on the right of Equation 1.
Definition 2 (Measurement Bias (Information Bias)) Suppose an analysis aims at the causal effect of a treatment \(A\) on an outcome \(Y\) but uses measured versions (Definition 1) of some of its variables: the treatment, the outcome, or covariates such as confounders. There is measurement bias, or information bias, when the resulting measure of association differs from the causal effect of \(A\) on \(Y\) because of measurement error, for example when the measured risk ratio \(\Pr[Y^* = 1 \mid A^* = 1] / \Pr[Y^* = 1 \mid A^* = 0]\) differs from the causal risk ratio \(\Pr[Y^{a=1} = 1] / \Pr[Y^{a=0} = 1]\).
Example 3 (Measurement Bias in Figure 9.2) In Figure 9.2, Proposition 1 says the risk ratio of \(A\) and \(Y\) is causal. If the analysis can only use the recorded drug use \(A^*\) of Example 2 and a recorded diagnosis \(Y^*\), the risk ratio of \(A^*\) and \(Y^*\) is not guaranteed to match it, and any gap between the two is measurement bias (Definition 2).
Identifiability Conditions Do Not Protect Against Measurement Bias
In our reading, exchangeability, positivity, and consistency concern the true \(A\) and \(Y\). When measurement bias is present, they no longer suffice to recover the effect of \(A\) on \(Y\) from \(A^*\) and \(Y^*\) (Hernán and Robins 2020, 126).
Technical Point 9.1: Independence and Nondifferentiality of Measurement Errors
The book defines two properties of the errors of treatment and outcome, which the next two definitions state (Hernán and Robins 2020, 126).
Definition 3 (Independent Measurement Errors) For each individual, define the measurement errors \(e_A = A^* - A\) and \(e_Y = Y^* - Y\), and let \(f(\cdot)\) denote a probability density (or mass) function. The errors \(e_A\) and \(e_Y\) are independent if their joint density factors into the product of the marginals:
\[ f(e_Y, e_A) = f(e_Y) f(e_A) \tag{3}\]
Otherwise they are dependent.
Definition 4 (Nondifferential Measurement Error) For each individual, let \(e_A = A^* - A\) and \(e_Y = Y^* - Y\) be the measurement errors of treatment and outcome, and let \(f(\cdot)\) denote a probability density (or mass) function, as in Definition 3.
An error that is not nondifferential is differential.
Example 4 (Error Types in the Pedestrian Experiment) In Example 1, assume the treatment is recorded perfectly, so \(e_A = 0\) for everyone. A constant error is independent of every other variable, so \(e_A\) and \(e_Y\) are independent (Definition 3) and \(e_A\) is nondifferential (Definition 4).
Now suppose the recorder catches each look-up with probability one half, whatever the investigator did, and never records a look-up that did not happen. Then \(e_Y = -1\) exactly when a pedestrian looked up and the recorder missed it, and \(e_Y = 0\) otherwise, so \(\Pr[e_Y = -1 \mid A = a] = \frac{1}{2} \Pr[Y = 1 \mid A = a]\). If \(\Pr[Y = 1 \mid A = 1] \neq \Pr[Y = 1 \mid A = 0]\), this probability differs between \(a = 0\) and \(a = 1\), and \(e_Y\) is differential in the sense of Definition 4, even though the recorder’s lapses have nothing to do with treatment. This example is ours, not the book’s.
Confounding has a single structure (common causes) and so does selection bias (conditioning on common effects), but there is no single structure for measurement error. The book classifies it by two properties, independence and nondifferentiality, read off the causal diagram through the error nodes.
Definition 5 (Graphical Classification of Measurement Error) Consider a causal diagram with a treatment \(A\), an outcome \(Y\), and their measured versions \(A^*\) and \(Y^*\). Let \(U_A\) and \(U_Y\) be the nodes standing for the factors, other than \(A\) and \(Y\), that determine \(A^*\) and \(Y^*\). The measurement errors are
Remark 2 (Error Nodes Versus Error Variables). Definition 5 classifies the error nodes \(U_A\) and \(U_Y\), while Definition 3 and Definition 4 classify the error variables \(e_A\) and \(e_Y\). If \(e_A\) is a function of \(U_A\) alone and \(e_Y\) a function of \(U_Y\) alone, then, under the causal Markov assumption, graphical independence implies independence of \(e_A\) and \(e_Y\), and graphical nondifferentiality implies nondifferentiality of the error variables. Otherwise the two classifications can disagree. Under misclassification of a binary variable, the error \(e_Y = Y^* - Y\) depends on \(Y\) as well as on \(U_Y\), so in Figure 9.2 it can be associated with \(A\) (through \(A \rightarrow Y\)) even though \(U_Y\) is not, as Example 4 showed. This remark is ours, not the book’s.
Example 5 (Independent Errors: Electronic Records) In Figure 9.2 the only path between \(U_A\) and \(U_Y\) is \(U_A \rightarrow A^* \leftarrow A \rightarrow Y \rightarrow Y^* \leftarrow U_Y\). It contains the colliders \(A^*\) and \(Y^*\), neither of which is conditioned on, so \(U_A\) and \(U_Y\) are d-separated and the errors are independent (Definition 5). This structure fits a study that takes drug use \(A\) as well as liver toxicity \(Y\) from electronic medical records whose data-entry mistakes occur haphazardly.
Example 6 (Dependent Errors: Phone Interviews) Suppose instead that drug use and liver toxicity are both obtained by interviewing participants by phone after the fact. A participant’s ability to recall her medical history, \(U_{AY}\), then affects the recorded values of both, which adds \(U_{AY} \rightarrow U_A\) and \(U_{AY} \rightarrow U_Y\) as in Figure 9.3. The path \(U_A \leftarrow U_{AY} \rightarrow U_Y\) is open, so the errors are dependent (Definition 5).
Definition 6 (Recall Bias) Recall bias is measurement bias (Definition 2) that arises when treatment is ascertained by asking participants to remember it, and the true outcome affects how well they remember (an arrow \(Y \rightarrow U_A\) into the error node \(U_A\) of Definition 5, Figure 9.4).
Example 7 (Dementia and Birth Defects) Two settings produce recall bias (Definition 6):
Definition 7 (Reverse Causation Bias) Reverse causation bias is measurement bias (Definition 2) that arises when the measured treatment \(A^*\) is a quantity that the outcome itself changes, giving the same structure as recall bias: an arrow from \(Y\) into the error node \(U_A\) of Definition 5.
Example 8 (Drug Levels Measured After Liver Toxicity) Suppose blood levels of the drug serve as \(A^*\), but they are measured after liver toxicity \(Y\) has developed. Liver toxicity itself alters the drug levels found in the blood, so the measured treatment depends on the outcome: reverse causation bias (Definition 7).
Example 9 (Closer Monitoring of Treated Patients) Physicians who suspect that the drug causes liver toxicity may monitor treated patients more closely than untreated ones. Toxicity is then more likely to be detected and recorded among the treated, so the true treatment affects the measurement of the outcome (an arrow \(A \rightarrow U_Y\) into the error node of Definition 5, Figure 9.5), a differential outcome error (Definition 4).
These settings can combine, as in Figures 9.6 and 9.7, giving errors that are both dependent and differential.
| Type | Figure |
|---|---|
| Independent nondifferential | 9.2 |
| Dependent nondifferential | 9.3 |
| Independent differential | 9.4, 9.5 |
| Dependent differential | 9.6, 9.7 |
Remark 3 (Correcting for Measurement Error). The structure of the error in Table 1 determines which correction methods apply; there is a large literature for independent nondifferential error. Correction methods generally combine modeling assumptions with validation samples, subsets of the data where the key variables are measured accurately. The book does not cover them; its point is that measuring variables, like selecting individuals, can introduce bias, so realistic causal diagrams must represent confounding, selection, and measurement simultaneously (Hernán and Robins 2020, 128).
Measure Better
The best defense against measurement bias is better measurement: improve how the variables are measured in the first place (Hernán and Robins 2020, 128).
Fine Point 9.1: The Strength and Direction of Measurement Bias
Measurement error generally causes bias, but there is one notable exception and no general rule for the direction. The boxes that follow give the exception and the direction and magnitude of the bias (Hernán and Robins 2020, 128). A worked example with the pedestrian experiment, which is ours, follows them.
Proposition 2 (Independent Nondifferential Error Preserves a Null) Suppose the causal DAG is Figure 9.2 without the arrow \(A \rightarrow Y\), so that its only arrows are \(A \rightarrow A^* \leftarrow U_A\) and \(Y \rightarrow Y^* \leftarrow U_Y\), and suppose the joint distribution satisfies the causal Markov assumption with respect to this DAG. Then \(A^* \perp\!\!\!\perp Y^*\) and \(A \perp\!\!\!\perp Y\): the true and the measured associations are both null.
Proof. The DAG has no edge between the sets \(\{A, U_A, A^*\}\) and \(\{Y, U_Y, Y^*\}\), so no path connects \(A^*\) with \(Y^*\), or \(A\) with \(Y\). d-separation then gives both independencies under the causal Markov assumption.
Remark 4 (Direction and Magnitude of Measurement Bias). Outside the setting of Proposition 2, the \(A^*\)-\(Y^*\) association may lie farther from the null, or nearer to it, than the \(A\)-\(Y\) association. The magnitude of the bias generally grows with the strength of the arrows \(U_A \rightarrow A^*\) and \(U_Y \rightarrow Y^*\). Causal diagrams encode no quantitative information, so they cannot describe the magnitude of the bias (Hernán and Robins 2020, 128).
Nondifferential Error Can Reverse a Trend
A common rule of thumb, that nondifferential error only biases toward the null, is not safe (this framing is ours). Even with independent nondifferential error and non-extreme bias, the \(A^*\)-\(Y^*\) and \(A\)-\(Y\) trends can go in opposite directions for ordinal (non-dichotomous) or continuous treatments. This reversal occurs when \(\operatorname{E}\mathopen{}\left[A^* \mid A\right]\mathclose{}\) does not move monotonically with \(A\) (Dosemeci, Wacholder, and Lubin 1990; Weinberg, Umbach, and Greenland 1994). VanderWeele and Hernán (2009) give a more general framework based on signed causal diagrams (Hernán and Robins 2020, 128).
Proposition 3 (Outcome Misclassification With Perfect Specificity) This proposition is ours, not the book’s. Let \(A\) and \(Y\) be binary, and let \(A\) be measured without error, so \(A^* = A\). Assume \(\Pr[Y = y, A = a] > 0\) for all \(y, a \in \{0, 1\}\), so that the conditional probabilities below are defined. The sensitivity of the outcome measurement in treatment group \(a\) is \(\Pr[Y^* = 1 \mid Y = 1, A = a]\) and its specificity is \(\Pr[Y^* = 0 \mid Y = 0, A = a]\). Suppose the sensitivity is \(s \in (0, 1]\) and the specificity is 1 in both treatment groups:
\[ \Pr[Y^* = 1 \mid Y = 1, A = a] = s, \quad \Pr[Y^* = 1 \mid Y = 0, A = a] = 0, \quad a = 0, 1. \tag{4}\]
Then, for \(a = 0, 1\), \(\Pr[Y^* = 1 \mid A = a] = s \Pr[Y = 1 \mid A = a]\). Because \(A^* = A\), conditioning on \(A^* = a\) is conditioning on \(A = a\), so the measured risk difference is \(s\) times the true associational risk difference, while the measured risk ratio equals the true associational risk ratio.
Proof. By the law of total probability and Equation 4,
\[\begin{align} \Pr[Y^* = 1 \mid A = a] &= \Pr[Y^* = 1 \mid Y = 1, A = a] \Pr[Y = 1 \mid A = a] + \Pr[Y^* = 1 \mid Y = 0, A = a] \Pr[Y = 0 \mid A = a] \\ &= s \Pr[Y = 1 \mid A = a] \end{align}\]
Subtracting the two treatment groups multiplies the risk difference by \(s\); dividing them cancels \(s\).
Example 10 (The Pedestrian Experiment in Numbers) Return to Example 1 with hypothetical numbers. Suppose that, in truth, 60% of pedestrians look up when the investigator looks up (\(A = 1\)) and 20% when she does not (\(A = 0\)). Randomization makes these associational risks causal (as in Proposition 1), so the causal risk difference is \(0.6 - 0.2 = 0.4\) and the causal risk ratio is \(0.6 / 0.2 = 3\).
The recorder misses half of the look-ups and never records a look-up that did not happen, in both groups: sensitivity \(s = 0.5\), perfect specificity. By Proposition 3 the recorded risks are \(0.3\) and \(0.1\), so the measured risk difference is \(0.2\), half the true one, while the measured risk ratio is still \(3\). So the “dilution” in Example 1 is real on the additive scale but absent on the ratio scale. These numbers and this conclusion are ours, not the book’s.
Mismeasured confounders can cause bias even when treatment and outcome are perfectly measured.
Example 11 (History of Hepatitis) In Figure 9.8, \(A\) is drug use, \(Y\) liver disease, and \(L\) history of hepatitis: people with prior hepatitis are less likely to be prescribed the drug and more likely to develop liver disease, so there is confounding through \(A \leftarrow L \rightarrow Y\). With \(L\) perfectly measured, \(L\) blocks that path and gives conditional exchangeability, so, with consistency and positivity, standardization or IP weighting by \(L\) recovers the causal risk ratio.
If hepatitis history is instead ascertained by questionnaire, some participants misreport it, and investigators have only the mismeasured \(L^*\). Conditioning on \(L^*\) does not generally block \(A \leftarrow L \rightarrow Y\), so the risk ratio standardized (or IP weighted) by \(L^*\) generally differs from the causal risk ratio: there is measurement bias (Hernán and Robins 2020, 128–29).
Example 12 (A Mismeasured Variable on a Longer Backdoor Path) In Figure 9.9, \(A\) is treatment, \(Y\) the outcome, \(L\) a covariate, and \(L^*\) its mismeasured version. As we read the figure, the confounding path is \(A \leftarrow L \leftarrow U \rightarrow Y\) with \(U\) unmeasured, and \(L^*\) is a child of \(L\). So \(L\) is not itself a common cause of \(A\) and \(Y\). Conditioning on the true \(L\) would block this path, but conditioning on \(L^*\) does not generally block it, so again there is measurement bias.
Remark 5 (Measurement Bias or Unmeasured Confounding?). Figure 9.8 is equivalent to Figure 7.8: one can view \(L\) as unmeasured and \(L^*\) as a surrogate confounder (Fine Point 7.3). Whether one calls the problem bias from a mismeasured confounder or unmeasured confounding makes no difference in practice (Hernán and Robins 2020, 129). In some settings, however, mismeasured variables are enough to adjust for confounding.
Mismeasured confounders can also create apparent effect modification.
Example 13 (Hepatitis Misreported in One Stratum Only) In the setting of Example 11 (drug use \(A\), liver disease \(Y\), true hepatitis history \(L\), reported history \(L^*\)), suppose the sharp null hypothesis holds (treatment has no effect on anyone’s liver disease), and
Then:
Reading both stratum-specific associations as effects, investigators would conclude that \(L^*\) modifies the effect of \(A\), although there is no effect at all (Hernán and Robins 2020, 129).
Measurement Error Can Masquerade as Effect Modification
When a confounder is measured more accurately in some strata than in others, stratum-specific associations can differ because residual confounding differs, even when the effect is the same in every stratum, as in Example 13. This general statement is ours. The book gives only that example (Hernán and Robins 2020, 129).
Example 14 (Conditioning on a Mismeasured Collider) A collider can be mismeasured too. In Figure 9.10 (equivalent to Figure 8.2) the arrows are \(A \rightarrow Y\), \(A \rightarrow C \leftarrow Y\), and \(C \rightarrow C^*\). For the effect of \(A\) on \(Y\), conditioning on the mismeasured \(C^*\) generally introduces selection bias, since \(C^*\) descends from the collider \(C\), which makes it a common effect of \(A\) and \(Y\) as well.
Fine Point 9.2: When Mismeasured Confounders Are Not a Problem
Error in a measured confounder often does no harm in clinical settings, because treatment decisions were themselves based on the measured value. The boxes that follow give two examples and the general principle (Hernán and Robins 2020, 130).
Example 15 (Treatment Decided on the Office Blood Pressure) High blood pressure \(L\) affects antihypertensive therapy \(A\) and stroke \(Y\), but treatment decisions are based on the office measurement \(L^*\), not on the true \(L\). In Figure 9.11 (structurally equivalent to Figure 7.2) the arrows are \(L \rightarrow L^* \rightarrow A\), \(L \rightarrow Y\), and \(A \rightarrow Y\): \(L^*\) carries the entire effect of \(L\) on \(A\), since any part of \(L\) not captured by \(L^*\) was unknown to decision makers and could not affect treatment. The only backdoor path, \(A \leftarrow L^* \leftarrow L \rightarrow Y\), is blocked by conditioning on either \(L\) or \(L^*\), and neither is a descendant of \(A\), so either one satisfies the backdoor criterion (Section 7.2).
Example 16 (Only the Measured Value Suffices) In Figure 9.12, \(A\) is antihypertensive therapy, \(Y\) stroke, \(L\) true blood pressure, and \(L^*\) its office measurement. As we read the figure, its arrows are \(L \rightarrow L^* \rightarrow A \rightarrow Y\), \(U_1 \rightarrow L\), \(U_1 \rightarrow Y\), \(U_2 \rightarrow L\), and \(U_2 \rightarrow L^*\), with \(U_1\) and \(U_2\) unmeasured. Conditioning on \(L^*\) blocks every backdoor path, because each one enters \(A\) through the non-collider \(L^*\). Conditioning on the true \(L\) instead leaves \(A \leftarrow L^* \leftarrow U_2 \rightarrow L \leftarrow U_1 \rightarrow Y\) open, because \(L\) is a collider on it. So data on \(L^*\) suffice to adjust for confounding and data on \(L\) do not (Hernán and Robins 2020, 130).
Remark 6 (Have What the Decision Makers Had). Example 15 and Example 16 illustrate a general point: if the other identifiability conditions hold, an effect can be identified when the data contain as much information as the decision makers used to assign treatment, whether or not that information was measured with error (Hernán and Robins 2020, 130).
Earlier chapters drew causal diagrams under two simplifying assumptions; this section makes the first explicit.
Remark 7 (Assumption 1: Every Variable Is Perfectly Measured). Earlier diagrams assumed that every variable on the diagram is measured without error. This is unrealistic. As this chapter showed, measurement error can
Should every diagram then show both true and measured values of every variable? Often the measurement error is believed, or known, to be negligible, and then a simpler diagram is preferable.
A Two-Step Approach
Confounding and selection bias that are present when every variable is measured perfectly typically persist under measurement error (with exceptions, as in Example 16). So first draw diagrams without measurement error to study confounding and selection bias, then add measurement error as an extra layer. Throughout the book, whenever the focus is confounding or selection, the distinction between true and measured values is omitted (Hernán and Robins 2020, 130–31).
The next section turns to the second assumption, which is fundamental to any causal diagram.
Example 17 (Obesity, Antiviral Treatment, and COVID-19 Death) Let \(A\) be antiviral treatment for COVID-19, \(Y\) death, and \(L\) obesity (body mass index above 30), all binary. Obese patients are more likely to be treated, and more likely to be hospitalized if untreated, so experts draw Figure 9.13 (equal to Figure 7.1): \(L \rightarrow A\), \(L \rightarrow Y\), \(A \rightarrow Y\). Assume treatment depends only on \(L\) and on physician preference, and that there is no measurement error.
Remark 8 (Assumption 2: Every Arrow Has a Well-Defined Intervention). Earlier diagrams also assumed that every arrow corresponds to a well-defined intervention.
Definition 8 (Causal DAG (Strict Sense)) Consider a DAG that meets the three conditions for a causal DAG in Chapter 6 (Technical Point 6.1), restated here:
Suppose also that the causal Markov assumption links the DAG to the data. The book calls it a causal DAG in the strict sense if, in addition, every arrow has a causal interpretation, that is, every arrow \(X \rightarrow W\) comes with well-defined interventions on \(X\) for its effect on \(W\). If at least one arrow lacks such an interpretation, the book calls the DAG a “noncausal” diagram (Hernán and Robins 2020, 132).
Example 18 (Figure 9.13 Is a “Noncausal” Diagram)
So Figure 9.13, given what is known today, is a “noncausal” diagram (Definition 8).
Fine Point 9.3: Whether Interventions Are Well-Defined Depends on the Outcome of Interest
If we know an intervention on \(L\) that changes \(A\), why can the same intervention not be used to learn about \(Y\)? The example below resolves this apparent contradiction (Hernán and Robins 2020, 132).
Example 19 (Body Weight Versus the Doctor’s Perception of It) The contradiction comes from using \(L\) for two things: the physical quantity body weight, and the doctor’s perception of that quantity. Let \(L\) be body weight and \(L^*\) the doctor’s perception of it. With perfect perception there is a deterministic arrow \(L \rightarrow L^*\), and the arrow into treatment is \(L^* \rightarrow A\): body weight affects treatment only once the doctor learns \(L^*\). Intervening on the perceived \(L^*\) while leaving body weight unchanged would change the doctor’s behavior just as intervening on \(L\) would. In this sense the counterfactuals \(A^l\) are well defined, while \(Y^l\) are not (Hernán and Robins 2020, 132). Our gloss on why: changing the doctor’s perception \(L^*\) leaves the patient’s actual body weight \(L\) untouched, so it cannot stand in for an intervention on \(L\) when the outcome is death, which body weight may affect by routes that do not pass through the doctor.
Example 20 (Adding Hidden Factors Behind Obesity) Experts who want a causal DAG that still accounts for the fact that obese patients die more often can instead propose Figure 9.14, which has the same structure as Figure 7.2. It adds a hidden node \(H\) that causes both \(L\) and \(Y\); \(H\) may contain unmeasured, possibly unknown, factors, for example genes, the metabolism of body fat, the microbiota, and factors not yet discovered. Figure 9.14 has
\[ H \rightarrow L \rightarrow A \rightarrow Y, \qquad H \rightarrow Y, \tag{5}\]
with no arrow \(H \rightarrow A\), which is reasonable if obesity \(L\) is the only information used in treatment decisions.
Remark 9 (When Is Figure 9.14 Causal?). The new arrows \(H \rightarrow L\) and \(H \rightarrow Y\) of Example 20 are causal exactly when the counterfactuals \(L^h\) and \(Y^h\) are well defined. That holds if the experts are willing to assume either
The precise intervention on \(H\) is unknown, but current knowledge does not exclude it, so both arrows out of \(H\) are tentatively justified and Figure 9.14 is treated as a causal diagram (Hernán and Robins 2020, 132–33).
Fine Point 9.4: “Noncausal” Diagrams With Well-Defined Statistical Interpretations
Some authors keep noncausal arrows by reading a DAG purely statistically. The remarks below describe that reading and the difficulties Richardson and Robins (2013) raised (Hernán and Robins 2020, 133).
Remark 10 (A Purely Statistical Reading of a DAG). Take a DAG such as Figure 7.14 (the front door graph), with treatment \(A\), mediator \(M\), and outcome \(Y\), and represent it as a finest fully randomized causally interpreted structured tree graph (FFRCISTG) model (Chapter 6) in which interventions are well defined for \(A\) alone, whose counterfactuals are \((M^a, Y^a)\), and the joint distribution factors according to the DAG. Its arrows need not be causal; they only encode, via d-separation, conditional independencies on the DAG and the associated SWIG. The same reading applies to Figure 9.13 of Example 17, where \(L\) is obesity, \(A\) antiviral treatment, and \(Y\) death, and only \(A\) has well-defined interventions. Under this reading, the arrow \(L \rightarrow Y\) in Figure 9.13 need not be removed merely because \(L\) lacks well-defined interventions.
Remark 11 (Difficulties With the Statistical Reading). Richardson and Robins (2013) raised serious objections to Remark 10: if arrows are not causal, nothing makes the distribution factor according to any incomplete DAG, and with unmeasured variables some of the implied independencies cannot even be checked. For example, Figure 7.14, the front door graph, implies \(Y \perp\!\!\!\perp A \mid M, U\), and it is hard to see a reason for postulating this other than believing every arrow is causal.
Alternatively, noncausal arrows can be read as a response by researchers skeptical that the counterfactuals \(Y^m\) exist: Figure 7.14 then states that \(Y \perp\!\!\!\perp A \mid M\) would hold in a future trial randomizing \(A\), and if it fails, the claim that \(M\)-counterfactuals exist is falsified together with the structure (Hernán and Robins 2020, 133).
Fine Point 9.5: A Connection to the Front Door Formula
Relabeling \(A\), \(M\), and \(U\) in the front door diagram of Figure 7.14 as \(L\), \(A\), and \(H\) turns it into Figure 9.14, so the front door formula identifies the effect of obesity on death through treatment. The proposition below states this, and the remark after it shows what changes when a direct arrow is added (Hernán and Robins 2020, 134).
Proposition 4 (The Front Door Formula in Figure 9.14) Suppose Figure 9.14 (Equation 5) is the causal DAG in the sense of Definition 8, so that the counterfactuals \(Y^l\) and \(Y^a\) are well defined, with \(H\) unmeasured, \(L\) and \(A\) discrete, and \(Y\) binary (so \(\operatorname{E}\mathopen{}\left[Y^l\right]\mathclose{} = \Pr[Y^l = 1]\)). Assume consistency and \(\Pr[A = a, L = l'] > 0\) for every \(a\) and \(l'\). Then, for every \(l\),
\[ \operatorname{E}\mathopen{}\left[Y^l\right]\mathclose{} = \sum_a \Pr[A = a \mid L = l] \sum_{l'} \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l'\right]\mathclose{} \Pr[L = l'] \tag{6}\]
Proof. Figure 7.14 has arrows \(A \rightarrow M\), \(M \rightarrow Y\), \(U \rightarrow A\), and \(U \rightarrow Y\), with \(U\) unmeasured. Figure 9.14 has arrows \(L \rightarrow A\), \(A \rightarrow Y\), \(H \rightarrow L\), and \(H \rightarrow Y\), with \(H\) unmeasured. Mapping \(L\) to \(A\), \(A\) to \(M\), and \(H\) to \(U\) turns one edge list into the other. In particular, \(L\) has no direct arrow to \(Y\), so \(A\) intercepts every directed path from \(L\) to \(Y\), and the strict-causal hypothesis makes the counterfactuals \(Y^l\) and \(Y^a\) well defined. The front door conditions also hold in Figure 9.14: the only backdoor path from \(L\) to \(A\), \(L \leftarrow H \rightarrow Y \leftarrow A\), is blocked at the collider \(Y\), and the only backdoor path from \(A\) to \(Y\), \(A \leftarrow L \leftarrow H \rightarrow Y\), is blocked by \(L\). The front door formula of Technical Point 7.4 (Chapter 7) therefore applies after substituting \(L, l, l'\) for \(A, a, a'\) and \(A, a\) for \(M, m\), which gives Equation 6.
Remark 12 (Adding a Direct Arrow From Obesity to Death). Suppose a researcher adds a direct arrow \(L \rightarrow Y\) to Figure 9.14, for example because patients with high \(L\) receive ancillary care (such as dietary advice) that is not in the records. Then \(Y^l - Y^{l'}\) becomes the total effect along both \(L \rightarrow Y\) and \(L \rightarrow A \rightarrow Y\), so Proposition 4 no longer applies, and the book states that, with that arrow added, \(\operatorname{E}\mathopen{}\left[Y^l\right]\mathclose{}\) is no longer identified. The effect along \(L \rightarrow A \rightarrow Y\) alone remains identified by the front door formula, which the book proves in Chapter 23 (Hernán and Robins 2020, 134).
Sometimes it does not.
Proposition 5 (Adjusting for Obesity in Figure 9.14) Let \(A\) be antiviral treatment, \(Y\) death, \(L\) obesity, and \(H\) the hidden factors of Example 20. Suppose the causal Figure 9.14 (Equation 5) is the true causal DAG, with \(L\) discrete. Assume consistency and \(\Pr[A = a \mid L = l] > 0\) for every \(l\) with \(\Pr[L = l] > 0\). Then
\[ \operatorname{E}\mathopen{}\left[Y^a\right]\mathclose{} = \sum_l \operatorname{E}\mathopen{}\left[Y \mid A = a, L = l\right]\mathclose{} \Pr[L = l] \tag{7}\]
Proof. In Figure 9.14 the only backdoor path from \(A\) to \(Y\) is \(A \leftarrow L \leftarrow H \rightarrow Y\). \(L\) is a non-collider on it, so conditioning on \(L\) blocks it, and \(L\) is not a descendant of \(A\). So \(\{L\}\) satisfies the backdoor criterion (Section 7.2), which implies \(Y^a \perp\!\!\!\perp A \mid L\) even without faithfulness (Technical Point 7.1 in Chapter 7). The standardization theorem (Chapter 7) then gives Equation 7.
An analyst who reads the “noncausal” Figure 9.13 as if it were causal reaches the same formula: its only backdoor path, \(A \leftarrow L \rightarrow Y\), is also blocked by \(L\). So, when Figure 9.14 is the true DAG, Figure 9.13 guides the analysis correctly, even though one of its arrows has no causal meaning.
This agreement is expected, because in neither DAG do unmeasured variables have arrows into \(A\) (compare Remark 6).
Remark 13 (Many Published DAGs Omit Their Hidden Variables). Many DAGs in the health and social sciences are noncausal because they omit the hidden variables \(H\) that would make them causal; a node whose effects on its descendants have no well-defined intervention makes the whole DAG noncausal. The identifying formula may still coincide with that from the causal DAG, as in Proposition 5.
Sometimes it does.
Example 21 (A Surrogate Mistaken for a Confounder) Figure 9.15 has arrows \(U \rightarrow L\), \(U \rightarrow A\), \(L \rightarrow Y\), and \(A \rightarrow Y\). Read causally, its direct arrow \(L \rightarrow Y\) claims well-defined interventions on \(L\) for \(Y\), and \(L\) blocks the backdoor path \(A \leftarrow U \rightarrow L \rightarrow Y\).
Suppose instead that \(L\) is a surrogate for hidden factors \(H\) for which well-defined interventions exist (Figure 9.16): \(U \rightarrow H\), \(U \rightarrow A\), \(H \rightarrow L\), \(H \rightarrow Y\), and \(A \rightarrow Y\). The backdoor path is now \(A \leftarrow U \rightarrow H \rightarrow Y\), and \(L\) is not on it, so \(L\) does not block it. Investigators unaware that \(L\) is only a surrogate confounder would wrongly conclude that adjusting for \(L\) suffices, the same trap as with the mismeasured confounder of Example 11.
Noncausal Arrows Give False Confidence
A diagram that draws a measured variable where a hidden one belongs can make an adjustment set look sufficient when it is not, as in Example 21.
Interrogate Every Arrow
When proposing a causal DAG, ask of each arrow \(X \rightarrow W\) whether there is a well-defined intervention on \(X\) for its effect on \(W\). The question is idle for an electrical circuit, where every intervention is well defined, but not in the health and social sciences (Hernán and Robins 2020, 135).
Remark 14 (Well-Defined Is a Matter of Degree). Chapter 3 noted that no intervention is free of all vagueness; what matters is whether scientists agree that an intervention is specified precisely enough. From here on, except for clearly sign-posted exceptions, the book assumes all DAGs are strictly causal (Definition 8): every arrow corresponds to an intervention that can be specified with no meaningful vagueness given current knowledge, while acknowledging that such beliefs, and the diagrams, may later prove wrong (Hernán and Robins 2020, 135).
Fine Point 9.6: From Noncausal Diagrams to Causal Diagrams
Replacing a noncausal arrow with a hidden variable can change what is identifiable, and the way the hidden variable is added matters. The examples below follow one set of investigators through that step, and then a variant of it (Hernán and Robins 2020, 136).
Example 22 (Replacing a Noncausal Arrow With a Hidden Variable) Investigators who want the effect of \(A\) on \(Y\) draw Figure 9.17, with arrows \(U \rightarrow A\), \(U \rightarrow L\), \(A \rightarrow L\), \(L \rightarrow Y\), and \(A \rightarrow Y\), from the temporal order of the variables and two facts: the measured \(L\) is associated with \(Y\), and known but unmeasured factors \(U\) affect \(A\) and are associated with \(L\). If Figure 9.17 were the true causal diagram, the effect would not be identifiable, because \(L\) is a descendant of \(A\) (Fine Point 7.4).
There is no well-defined intervention on \(L\) for \(Y\), so they replace \(L \rightarrow Y\) by \(H \rightarrow Y\) for a hidden \(H\), with \(L\) as a surrogate of \(H\), and redirect the arrows from \(U\) and \(A\) into \(L\) toward \(H\) (Figure 9.18). Now no measured variable can block the backdoor path \(A \leftarrow U \rightarrow H \rightarrow Y\), so the effect is not identifiable. An \(L\) fully determined by \(H\) would not help either: \(H\) is richer than \(L\), so knowing \(L\) does not fix \(H\), and conditioning on \(L\) leaves paths through \(H\) open.
Example 23 (A Hidden Variable That Should Not Inherit Every Arrow) Keep \(A\), \(Y\), \(U\), \(L\), and \(H\) as in Example 22. Letting \(H\) inherit all arrows into \(L\) is not always warranted. Suppose \(U\) affects \(L\) directly rather than \(H\), as in Figure 9.19, with arrows \(U \rightarrow A\), \(U \rightarrow L\), \(A \rightarrow H\), \(H \rightarrow L\), \(H \rightarrow Y\), and \(A \rightarrow Y\). Then the only backdoor path, \(A \leftarrow U \rightarrow L \leftarrow H \rightarrow Y\), is blocked by the collider \(L\), so the effect is identifiable without adjustment. Example: \(U\) is a physician’s decision to order a diagnostic test, \(L\) the test result, and \(H\) the biological determinants of the result (Hernán and Robins 2020, 136).