Every method in the previous chapters needs all the variables required to adjust for confounding and selection bias to be identified and measured correctly. Instrumental variable (IV) estimation replaces that untestable assumption with a different set of (also untestable) assumptions, built around a variable \(Z\), the instrument, that does not require measuring the confounders.
In the book’s Figure 16.1, \(Z \rightarrow A \rightarrow Y\), with \(U\) a common cause of \(A\) and \(Y\).
Adjustment methods (IP weighting, standardization, g-estimation, stratification, matching) need to block the backdoor path \(A \leftarrow U \rightarrow Y\), so they are biased if part of \(U\) is unmeasured or mismeasured. IV methods may identify the effect of \(A\) on \(Y\) without measuring \(U\).
Definition 1 (The Three Instrumental Conditions) A variable \(Z\) is an instrument for the effect of \(A\) on \(Y\) if it meets three conditions (Hernán and Robins 2020, 209):
In the double-blind trial, \(Z\) is an instrument: (i) holds because those assigned to treatment are more likely to take it, (ii) is expected from blinding, and (iii) is expected from random assignment.
Technical Point 16.1: The Instrumental Conditions, Formally
All of these are expected to hold in a double-blind randomized trial (Hernán and Robins 2020, 210).
In the book’s Figure 16.2, \(U_Z \rightarrow Z\) and \(U_Z \rightarrow A \rightarrow Y\), with \(U\) a common cause of \(A\) and \(Y\).
To estimate the effect of smoking cessation \(A\) on weight gain \(Y\), the book proposes \(Z = 1\) if the average price of a pack of cigarettes in the individual’s U.S. state of birth was greater than $1.50, and \(Z = 0\) otherwise (Hernán and Robins 2020, 210).
Fine Point 16.1: Candidate Instruments in Observational Studies
Three common categories (Hernán and Robins 2020, 211):
Under conditions (i)-(iii) alone, the average causal effect is not point identified; only upper and lower bounds are, and they are typically wide and often include the null (Hernán and Robins 2020, 212). In the smoking example, the bounds would only say that quitting can cause weight gain, weight loss, or no change.
Technical Point 16.2: Partial Identification (Bounds)
For a dichotomous \(Y\), \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\) lies a priori in \((-1, 1)\), an interval of width 2.
These bounds are often too wide to be informative. Sufficiently strong parametric assumptions about the effect of \(A\) on \(Y\) (Section 16.2 onward) collapse them to a single number (Hernán and Robins 2020, 212).
For a dichotomous instrument meeting (i)-(iii), plus a condition (iv) described in Section 16.3, the average causal effect \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) is identified and equals the usual IV estimand
\[ \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \]
For a dichotomous treatment, \(\operatorname{E}\mathopen{}\left[A \mid Z = z\right]\mathclose{} = \Pr[A = 1 \mid Z = z]\). For a continuous instrument, the usual IV estimand is \(\operatorname{Cov}\mathopen{}\left(Y, Z\right)\mathclose{} / \operatorname{Cov}\mathopen{}\left(A, Z\right)\mathclose{}\) (Hernán and Robins 2020, 213).
Both can be estimated without adjustment because \(Z\) is randomized. With perfect adherence the denominator is 1 and the IV estimand equals the intention-to-treat effect; as adherence worsens the denominator approaches 0 and the IV estimand is increasingly inflated relative to the intention-to-treat effect.
Example 1 (Usual IV Estimate in NHEFS) With \(Z\) = 1 for a state with high cigarette price (Hernán and Robins 2020, 213):
\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} &= 2.686 - 2.536 = 0.1503 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z = 0\right]\mathclose{} &= 0.2578 - 0.1951 = 0.0627 \\ \text{IV estimate} &= \frac{0.1503}{0.0627} \approx 2.4 \text{ kg}. \end{aligned} \]
Under (i)-(iv), this estimates the average causal effect of smoking cessation on weight gain.
The same estimate can be computed with two saturated linear models, \(\operatorname{E}\mathopen{}\left[A \mid Z\right]\mathclose{} = \alpha_0 + \alpha_1 Z\) (denominator) and \(\operatorname{E}\mathopen{}\left[Y \mid Z\right]\mathclose{} = \beta_0 + \beta_1 Z\) (numerator), or by two-stage least squares:
\(\hat \beta_1\) always equals the standard IV estimate: again 2.4 kg (Hernán and Robins 2020, 213).
Conditions (i)-(iii) do not make the IV estimand equal the average causal effect. A fourth condition, effect homogeneity (iv), is needed. The book describes four versions, in order of historical appearance (Hernán and Robins 2020, 214–16):
Define the effect modification by \(U\) of the effect of \(A\) on \(Y\) and of the \(Z\)-\(A\) association:
\[ e(U) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid U\right]\mathclose{}, \qquad t(U) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[A \mid Z = 1, U\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0, U\right]\mathclose{}. \]
Theorem 1 (General Homogeneity Condition) Under the causal diagram of Figure 16.1, if \(\operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} = 0\), then
\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{} = \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \]
Versions 3 and 4 are special cases: they say “\(e(U)\) is constant” and “\(t(U)\) is constant”, and the covariance of any variable with a constant is 0 (Hernán and Robins 2020, 216).
Proof (Proof (Technical Point 16.5)). In Figure 16.1, \(Y^a \perp\!\!\!\perp(A, Z) \mid U\) for each \(a\) and \(U \perp\!\!\!\perp Z\).
Step 1: the denominator is \(\operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}\). Because \(U \perp\!\!\!\perp Z\), \(f(u \mid z) = f(u)\), so
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[A \mid Z = z\right]\mathclose{} &= \textstyle\sum_u \operatorname{E}\mathopen{}\left[A \mid Z = z, U = u\right]\mathclose{} f(u \mid z) && \text{(law of total expectation)} \\ &= \textstyle\sum_u \operatorname{E}\mathopen{}\left[A \mid Z = z, U = u\right]\mathclose{} f(u) && (U \perp\!\!\!\perp Z) \\ \operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{} &= \textstyle\sum_u t(u) f(u) = \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}. && \text{(definition of } t) \end{aligned} \]
Step 2: the numerator is \(\operatorname{E}\mathopen{}\left[e(U)\,t(U)\right]\mathclose{}\). By consistency, \(Y = A (Y^{a=1} - Y^{a=0}) + Y^{a=0}\). For each \(z\) and \(u\),
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[A (Y^{a=1} - Y^{a=0}) \mid Z = z, U = u\right]\mathclose{} &= \Pr(A = 1 \mid Z = z, U = u)\, \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u\right]\mathclose{} \\ &= \Pr(A = 1 \mid Z = z, U = u)\, \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid U = u\right]\mathclose{} && (Y^a \perp\!\!\!\perp(A, Z) \mid U) \\ &= \Pr(A = 1 \mid Z = z, U = u)\, e(u), \\ \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = z, U = u\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a=0} \mid U = u\right]\mathclose{}. && (Y^{a=0} \perp\!\!\!\perp Z \mid U) \end{aligned} \]
Averaging over \(U\) (with \(f(u \mid z) = f(u)\)):
\[ \operatorname{E}\mathopen{}\left[Y \mid Z = z\right]\mathclose{} = \textstyle\sum_u \mathopen{}\left\{e(u) \Pr(A = 1 \mid Z = z, U = u) + \operatorname{E}\mathopen{}\left[Y^{a=0} \mid U = u\right]\mathclose{}\right\}\mathclose{} f(u), \]
so the second term cancels in the difference and
\[ \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} = \textstyle\sum_u e(u)\, t(u)\, f(u) = \operatorname{E}\mathopen{}\left[e(U)\, t(U)\right]\mathclose{} . \]
Step 3: use the zero covariance.
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[e(U)\, t(U)\right]\mathclose{} &= \operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} + \operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{} && (\operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} = 0) \\ &= \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{} \mathopen{}\left(\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}\right)\mathclose{}. && \text{(iterated expectations; Step 1)} \end{aligned} \]
Dividing the numerator (Step 2) by the denominator (Step 1) gives the result.
Technical Point 16.3: Additive Structural Mean Models and IV Estimation
Saturated additive structural mean model for dichotomous \(A\) and instrument \(Z\):
\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A = 1, Z\right]\mathclose{} = \beta_0 + \beta_1 Z, \quad\text{equivalently}\quad \operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid A, Z\right]\mathclose{} = A (\beta_0 + \beta_1 Z). \]
\(\beta_0\) is the effect in the treated with \(Z = 0\), \(\beta_0 + \beta_1\) the effect in the treated with \(Z = 1\); \(\beta_1\) measures additive effect modification by \(Z\). If we assume \(\beta_1 = 0\), then \(\beta_0\) equals the usual IV estimand (Hernán and Robins 2020, 215).
Proof. The instrumental conditions imply \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 0\right]\mathclose{}\). From the model, \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A, Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A, Z\right]\mathclose{} - A(\beta_0 + \beta_1 Z)\); averaging over \(A\) given \(Z\),
\[ \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid Z\right]\mathclose{} - (\beta_0 + \beta_1 Z)\, \operatorname{E}\mathopen{}\left[A \mid Z\right]\mathclose{}. \]
Setting \(\beta_1 = 0\) and equating the \(Z = 1\) and \(Z = 0\) values:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \beta_0 \operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} - \beta_0 \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{} \\ \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} &= \beta_0 \mathopen{}\left(\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}\right)\mathclose{} \\ \beta_0 &= \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \end{aligned} \]
Technical Point 16.4: Multiplicative Structural Mean Models and IV Estimation
Saturated multiplicative (log-linear) structural mean model:
\[ \frac{\operatorname{E}\mathopen{}\left[Y^{a=1} \mid A = 1, Z\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A = 1, Z\right]\mathclose{}} = \operatorname{exp}\mathopen{}\left\{\beta_0 + \beta_1 Z\right\}\mathclose{}, \quad\text{equivalently}\quad \operatorname{E}\mathopen{}\left[Y \mid A, Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid A, Z\right]\mathclose{} \operatorname{exp}\mathopen{}\left\{A (\beta_0 + \beta_1 Z)\right\}\mathclose{}. \]
Assuming \(\beta_1 = 0\) and no multiplicative effect modification by \(Z\) in the untreated either, the causal risk ratio is \(\operatorname{exp}\mathopen{}\left\{\beta_0\right\}\mathclose{}\) and
\[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}(1 - \operatorname{E}\mathopen{}\left[A\right]\mathclose{})\mathopen{}\left[\operatorname{exp}\mathopen{}\left\{\beta_0\right\}\mathclose{} - 1\right]\mathclose{} + \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[A\right]\mathclose{} \mathopen{}\left[1 - \operatorname{exp}\mathopen{}\left\{-\beta_0\right\}\mathclose{}\right]\mathclose{}. \]
This is identified but is not the usual IV estimand (Hernán and Robins 2020, 216).
Because the homogeneity conditions are often implausible, the book describes two ways around them (Hernán and Robins 2020, 216):
Technical Point 16.6: More General Structural Mean Models
With possibly continuous or multivariate \(A\) and \(Z\) and pre-instrument covariates \(V\), an additive structural mean model is
\[ \operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid Z, A, V\right]\mathclose{} = \gamma(Z, A, V; \beta), \]
with \(\gamma\) a known function, \(\beta\) an unknown parameter, and \(\gamma(Z, A = 0, V; \beta) = 0\). It models the effect of the observed treatment level versus level 0 among those with that \(Z\), \(V\), and observed \(A\). Its parameters are identified by g-estimation under \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 1, V\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 0, V\right]\mathclose{}\). The multiplicative analog is \(\operatorname{E}\mathopen{}\left[Y \mid Z, A, V\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z, A, V\right]\mathclose{} \operatorname{exp}\mathopen{}\left\{\gamma(Z, A, V; \beta)\right\}\mathclose{}\) (Hernán and Robins 2020, 218).
In the double-blind trial, let \(A^{z=1}\) and \(A^{z=0}\) be the treatment an individual would take if assigned to treatment or to no treatment.
Definition 2 (Compliance Types (Principal Strata))
These strata are not identified: someone with \(Z = 1\) and \(A = 1\) may be a complier or an always-taker; someone with \(Z = 1\) and \(A = 0\) may be a defier or a never-taker.
Definition 3 (Monotonicity) Monotonicity holds when there are no defiers, i.e., \(A^{z=1} \ge A^{z=0}\) for all individuals.
Theorem 2 (IV Estimand under Monotonicity) Under instrumental conditions (i)-(iii) with a dichotomous causal instrument that is randomly assigned, and monotonicity as condition (iv), the usual IV estimand equals the average causal effect in the compliers,
\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A^{z=1} = 1, A^{z=0} = 0\right]\mathclose{}, \]
not the average causal effect in the population (Hernán and Robins 2020, 218).
Sketch: the intention-to-treat effect is a weighted average over the four strata; it is zero in always-takers and never-takers (their \(A\) does not change with \(Z\), and \(Z\) acts only through \(A\)), and there are no defiers. So the numerator is the effect in compliers times the proportion of compliers, which is the denominator.
Proof (Proof (Technical Point 16.7)). Write the intention-to-treat effect as a weighted average over the strata (AT: always-takers, NT: never-takers, C: compliers, D: defiers):
\[ \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{} = \textstyle\sum_{s \in \{AT, NT, C, D\}} \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid s\right]\mathclose{}\, \Pr[s]. \]
Hence \(\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid C\right]\mathclose{}\, \Pr[C]\). In compliers \(A = Z\), so the effect of \(Z\) equals the effect of \(A\): \(\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid C\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid C\right]\mathclose{}\), and
\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid C\right]\mathclose{} = \frac{\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{}}{\Pr[C]} . \]
Random assignment implies \(Z \perp\!\!\!\perp\{Y^{a,z}, A^z;\ z = 0,1;\ a = 0,1\}\), so with consistency the numerator equals \(\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}\). For the denominator, \(\Pr[AT] = \Pr[A^{z=0} = 1] = \Pr[A = 1 \mid Z = 0]\) and \(\Pr[NT] = \Pr[A^{z=1} = 0] = \Pr[A = 0 \mid Z = 1]\); with no defiers,
\[ \begin{aligned} \Pr[C] &= 1 - \Pr[AT] - \Pr[NT] \\ &= 1 - \Pr[A = 1 \mid Z = 0] - \mathopen{}\left(1 - \Pr[A = 1 \mid Z = 1]\right)\mathclose{} \\ &= \Pr[A = 1 \mid Z = 1] - \Pr[A = 1 \mid Z = 0]. \end{aligned} \]
There is no assignment, but “compliers” still means \((A^{z=1} = 1, A^{z=0} = 0)\): in NHEFS, people who would quit smoking in a high-price state and not in a low-price state. If there are no defiers and the causal instrument is dichotomous, 2.4 kg estimates the effect in the compliers (Hernán and Robins 2020, 218–19).
Monotonicity was welcomed in the mid-1990s as the salvation of IV methods, but the book raises four problems (Hernán and Robins 2020, 219–20):
A proposed instrument may violate (ii) or (iii), or barely meet (i). Then IV estimation may be badly biased even if condition (iv) holds perfectly.
Fine Point 16.2: Defining Weak Instruments
The NHEFS instrument is weak by both: risk difference about 6%, F-statistic 0.8 (Hernán and Robins 2020, 221).
Three problems (Hernán and Robins 2020, 220–21):
A randomly generated \(Z\) has, in an infinite population, zero association with \(A\), so the IV estimand is undefined. In a finite sample, chance creates a small \(Z\)-\(U\), and hence \(Z\)-\(A\), association, so the denominator is small but not zero and the estimate can be wildly inflated.
Example 2 (The NHEFS Instrument Behaves Like Noise) Redefining \(Z\) as price above $1.60, $1.70, $1.80, or $1.90 gives IV estimates of 41.3, \(-40.9\), \(-21.1\), and \(-12.8\) kg, each with a huge confidence interval (Hernán and Robins 2020, 221).
IV estimation differs from the earlier methods in at least three ways (Hernán and Robins 2020, 223):
Fine Point 16.3: Regression Discontinuity Design
Example: an antiviral \(A\) for COVID-19 is given to everyone aged 65 or older (\(L \ge 65\)) and to no one younger, so \(\Pr[A = 1 \mid L < 65] = 0\) and \(\Pr[A = 1 \mid L \ge 65] = 1\). There is no positivity.
This is a sharp design; a fuzzy design, where the treatment probability jumps but not from 0 to 1, relies on monotonicity and estimates the effect in the compliers near the threshold (Hernán and Robins 2020, 225).