Chapter 16: Instrumental Variable Estimation

Published

Last modified: 2026-10-09 10:17:06 (UTC)

📝 Preview Changes: This page has been modified in this pull request (~6% of content changed).
🎨 Highlighting Legend: Modified text (yellow) shows changed words/phrases, added text (green) shows new content, and new sections (blue) highlight entirely new paragraphs.

Every method in the previous chapters needs all the variables required to adjust for confounding and selection bias to be identified and measured correctly. Instrumental variable (IV) estimation replaces that untestable assumption with a different set of (also untestable) assumptions, built around a variable \(Z\), the instrument, that does not require measuring the confounders.

This chapter is based on Hernán and Robins (2020, chap. 16, pp. 209-225).

Key theme: an instrument alone identifies only bounds on the average causal effect. A point estimate needs a fourth, untestable condition: either some form of effect homogeneity (giving the average causal effect in the population) or monotonicity (giving the effect in the “compliers” only). And because the IV estimator divides by the instrument-treatment association, small violations of any condition can produce large biases.

1 16.1 The Three Instrumental Conditions (pp. 209-212)


1.1 A Double-Blind Randomized Trial

  • \(Z\): randomization assignment (1: treatment, 0: placebo)
  • \(A\): treatment actually received (1: yes, 0: no); not everyone adheres
  • \(Y\): outcome
  • \(U\): factors, some unmeasured, that affect both adherence and the outcome

In the book’s Figure 16.1, \(Z \rightarrow A \rightarrow Y\), with \(U\) a common cause of \(A\) and \(Y\).

Adjustment methods (IP weighting, standardization, g-estimation, stratification, matching) need to block the backdoor path \(A \leftarrow U \rightarrow Y\), so they are biased if part of \(U\) is unmeasured or mismeasured. IV methods may identify the effect of \(A\) on \(Y\) without measuring \(U\).


Definition 1 (The Three Instrumental Conditions) A variable \(Z\) is an instrument for the effect of \(A\) on \(Y\) if it meets three conditions (Hernán and Robins 2020, 209):

  • (i) \(Z\) is associated with \(A\);
  • (ii) \(Z\) does not affect \(Y\) except through its potential effect on \(A\);
  • (iii) \(Z\) and \(Y\) do not share causes.

In the double-blind trial, \(Z\) is an instrument: (i) holds because those assigned to treatment are more likely to take it, (ii) is expected from blinding, and (iii) is expected from random assignment.

Condition (ii) would fail if, for example, side effects of treatment inadvertently unblinded participants (Hernán and Robins 2020, 209).


NoteTechnical Point 16.1: The Instrumental Conditions, Formally
  • (i) Relevance: \(Z\) and \(A\) are associated, i.e., \(Z \perp\!\!\!\perp A\) does not hold.
  • (ii) Exclusion restriction (“no direct effect of \(Z\) on \(Y\)”): at the individual level, \(Y_i^{z,a} = Y_i^{z',a} = Y_i^{a}\) for all \(z, z'\), all \(a\), and all individuals \(i\); some results only need the population-level version \(\operatorname{E}\mathopen{}\left[Y^{z,a}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{z',a}\right]\mathclose{}\).
  • (iii) as marginal exchangeability, \(Y^{a,z} \perp\!\!\!\perp Z\) for all \(a, z\); together with the individual-level (ii) it implies \(Y^a \perp\!\!\!\perp Z\). A stronger version is joint exchangeability, \(\{Y^{z,a};\ a \in \{0,1\},\ z \in \{0,1\}\} \perp\!\!\!\perp Z\) (dichotomous \(A\) and \(Z\)).

All of these are expected to hold in a double-blind randomized trial (Hernán and Robins 2020, 210).


1.2 Causal and Surrogate Instruments

  • A causal instrument has a causal effect on \(A\) (Figure 16.1).
  • A surrogate instrument \(Z\) has no effect on \(A\) but is associated with an unmeasured causal instrument \(U_Z\) (Figure 16.2). Condition (i) then holds through the shared cause \(U_Z\), and (iii) becomes “\(Z\) and \(Y\) share no causes other than \(U_Z\).”

In the book’s Figure 16.2, \(U_Z \rightarrow Z\) and \(U_Z \rightarrow A \rightarrow Y\), with \(U\) a common cause of \(A\) and \(Y\).

Both kinds of instrument can be used for IV estimation, with caveats discussed in Section 16.4. The book’s Figure 16.3 shows a more unusual surrogate instrument: \(Z\) and \(U_Z\) become associated because the analysis is restricted to a selected population, i.e., conditions on a common effect \(S\) of \(Z\) and \(U_Z\) (Hernán and Robins 2020, 210).


1.3 A Candidate Instrument in NHEFS

To estimate the effect of smoking cessation \(A\) on weight gain \(Y\), the book proposes \(Z = 1\) if the average price of a pack of cigarettes in the individual’s U.S. state of birth was greater than $1.50, and \(Z = 0\) otherwise (Hernán and Robins 2020, 210).

  • Condition (i) is the only one that can be checked: \(\Pr[A = 1 \mid Z = 1] = 25.8\%\) and \(\Pr[A = 1 \mid Z = 0] = 19.5\%\), a risk difference of about 6%. So \(Z\) is a weak instrument (Section 16.5).
  • Conditions (ii) and (iii) cannot be verified. Conditioning on \(A\) does not help: \(A\) is a collider on \(Z \leftarrow U_Z \rightarrow A \leftarrow U \rightarrow Y\).

“Price in state of birth” would meet condition (i) through the unmeasured causal instrument \(U_Z\) “price in place of residence”, so it is a surrogate instrument (Hernán and Robins 2020, 210).

Because only (i) is checkable, the book calls \(Z\) a proposed or candidate instrument: we argue for (ii) and (iii) with subject-matter knowledge, just as we argue for the identifying assumptions of other methods. Falsification tests exist that use data on \(Z\), \(A\), and \(Y\), but they can reject only a small subset of violations, and for most violations they have no power at any sample size (Hernán and Robins 2020, 211).


NoteFine Point 16.1: Candidate Instruments in Observational Studies

Three common categories (Hernán and Robins 2020, 211):

  • Genetic factors: a genetic variant associated with \(A\) and assumed to affect \(Y\) only through \(A\), e.g., a polymorphism affecting alcohol metabolism (such as ALDH2 in Asian populations) for the effect of alcohol on coronary heart disease. This is the framework of Mendelian randomization.
  • Preference: a physician’s preference for one treatment over another. Preference itself (\(U_Z\)) is unmeasured, so a surrogate such as “last prescription the physician issued before the current one” is used.
  • Access: distance or travel time to a facility, calendar period when access changed over time, or the price of treatment (the book’s example).

1.4 What an Instrument Alone Buys: Bounds

Under conditions (i)-(iii) alone, the average causal effect is not point identified; only upper and lower bounds are, and they are typically wide and often include the null (Hernán and Robins 2020, 212). In the smoking example, the bounds would only say that quitting can cause weight gain, weight loss, or no change.

NoteTechnical Point 16.2: Partial Identification (Bounds)

For a dichotomous \(Y\), \(\Pr[Y^{a=1} = 1] - \Pr[Y^{a=0} = 1]\) lies a priori in \((-1, 1)\), an interval of width 2.

  • Data alone: assigning each individual’s unobserved counterfactual its most extreme possible value halves the width, but the bounds still include 0.
  • Natural bounds (instrumental condition (ii) at the population level plus marginal exchangeability (iii)): width \(\Pr[A = 1 \mid Z = 0] + \Pr[A = 0 \mid Z = 1]\).
  • Sharp bounds: narrower still, under joint exchangeability.

These bounds are often too wide to be informative. Sufficiently strong parametric assumptions about the effect of \(A\) on \(Y\) (Section 16.2 onward) collapse them to a single number (Hernán and Robins 2020, 212).

Why the data alone halve the width: for a dichotomous \(Y\), each individual contributes one observed counterfactual, so only the unobserved half of each individual’s pair is unknown. For a continuous \(Y\), bounds require choosing a minimum and maximum possible outcome, and their width depends on that choice (Hernán and Robins 2020, 212).

2 16.2 The Usual IV Estimand (pp. 212-214)


For a dichotomous instrument meeting (i)-(iii), plus a condition (iv) described in Section 16.3, the average causal effect \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) is identified and equals the usual IV estimand

\[ \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \]

For a dichotomous treatment, \(\operatorname{E}\mathopen{}\left[A \mid Z = z\right]\mathclose{} = \Pr[A = 1 \mid Z = z]\). For a continuous instrument, the usual IV estimand is \(\operatorname{Cov}\mathopen{}\left(Y, Z\right)\mathclose{} / \operatorname{Cov}\mathopen{}\left(A, Z\right)\mathclose{}\) (Hernán and Robins 2020, 213).


2.1 Intuition from the Randomized Trial

  • Numerator: the effect of assignment \(Z\) on \(Y\), the intention-to-treat effect.
  • Denominator: the effect of \(Z\) on \(A\), a measure of adherence.

Both can be estimated without adjustment because \(Z\) is randomized. With perfect adherence the denominator is 1 and the IV estimand equals the intention-to-treat effect; as adherence worsens the denominator approaches 0 and the IV estimand is increasingly inflated relative to the intention-to-treat effect.

So the IV estimand avoids adjusting for confounders by inflating the effect of assignment. In observational studies the denominator is the effect of a causal instrument on \(A\) (Figure 16.1), or the noncausal association between a surrogate instrument and \(A\) (Figures 16.2 and 16.3) (Hernán and Robins 2020, 213).


Example 1 (Usual IV Estimate in NHEFS) With \(Z\) = 1 for a state with high cigarette price (Hernán and Robins 2020, 213):

\[ \begin{aligned} \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} &= 2.686 - 2.536 = 0.1503 \\ \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z = 0\right]\mathclose{} &= 0.2578 - 0.1951 = 0.0627 \\ \text{IV estimate} &= \frac{0.1503}{0.0627} \approx 2.4 \text{ kg}. \end{aligned} \]

Under (i)-(iv), this estimates the average causal effect of smoking cessation on weight gain.

The numerator difference is reported to four decimals (0.1503), presumably computed from unrounded means; the rounded means shown give \(2.686 - 2.536 = 0.150\). This ratio is also called the Wald estimator. The book excludes individuals with a missing outcome or instrument for simplicity; in practice IP weighting could adjust for the resulting selection bias (Hernán and Robins 2020, 213).


2.2 Two-Stage Least Squares

The same estimate can be computed with two saturated linear models, \(\operatorname{E}\mathopen{}\left[A \mid Z\right]\mathclose{} = \alpha_0 + \alpha_1 Z\) (denominator) and \(\operatorname{E}\mathopen{}\left[Y \mid Z\right]\mathclose{} = \beta_0 + \beta_1 Z\) (numerator), or by two-stage least squares:

  1. Fit the first-stage treatment model \(\operatorname{E}\mathopen{}\left[A \mid Z\right]\mathclose{} = \alpha_0 + \alpha_1 Z\) and compute predicted values \(\mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z\right]\mathclose{}\).
  2. Fit the second-stage outcome model \(\operatorname{E}\mathopen{}\left[Y \mid Z\right]\mathclose{} = \beta_0 + \beta_1 \mathop{\hat{\operatorname{E}}}\nolimits\mathopen{}\left[A \mid Z\right]\mathclose{}\).

\(\hat \beta_1\) always equals the standard IV estimate: again 2.4 kg (Hernán and Robins 2020, 213).


2.3 Precision and Weak Instruments

  • 95% confidence interval: \(-36.5\) to \(41.3\) kg.
  • Rule of thumb: an instrument is weak if the first-stage F-statistic is below 10; here it was 0.8 (Hernán and Robins 2020, 213–14).

The assumptions implicit in two-stage least squares can be made explicit with additive or multiplicative structural mean models (Technical Points 16.3 and 16.4), whose parameters can be estimated by g-estimation. When measured common causes \(L\) of the instrument and the outcome must be adjusted for, choosing between two-stage least squares and structural mean models involves trade-offs similar to those between outcome regression and structural nested models (Chapters 14 and 15) (Hernán and Robins 2020, 214).

All of these estimators estimate the average causal effect only if a fourth identifying condition holds.

3 16.3 A Fourth Identifying Condition: Homogeneity (pp. 214-217)


Conditions (i)-(iii) do not make the IV estimand equal the average causal effect. A fourth condition, effect homogeneity (iv), is needed. The book describes four versions, in order of historical appearance (Hernán and Robins 2020, 214–16):

  1. Constant effect of \(A\) on \(Y\) for every individual (e.g., everyone gains exactly 2.4 kg). This is additive rank preservation (Section 14.4): implausible for most treatments and impossible for dichotomous outcomes except under the sharp null or universal harm (or benefit).
  2. Equal average effect across levels of \(Z\) among the treated and among the untreated: \(\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid Z = 1, A = a\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid Z = 0, A = a\right]\mathclose{}\) for \(a = 0, 1\).
  3. No additive effect modification by \(U\): \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid U\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0} \mid U\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\).
  4. Constant \(Z\)-\(A\) association across \(U\): \(\operatorname{E}\mathopen{}\left[A \mid Z = 1, U\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0, U\right]\mathclose{} = \operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}\).

Version 1 was implicitly assumed in many early IV analyses that used two-stage least squares.

Version 2 is the one used in Technical Point 16.3. It is not automatically implied by condition (iii): even when \(Y^a \perp\!\!\!\perp Z\) holds, \(Y^a \perp\!\!\!\perp Z \mid A\) generally does not, so the treatment effect may vary with \(Z\) among the treated or the untreated (Hernán and Robins 2020, 214). It is also hard for subject-matter experts to argue for, which motivates versions stated in terms of the confounders \(U\).

Version 3 is often implausible because some unmeasured confounders are likely effect modifiers too, e.g., weight gain after quitting may depend on prior smoking intensity, itself a likely unmeasured confounder. Hernán and Robins showed that if \(U\) is an additive effect modifier, version 2 would not be a reasonable belief either (Hernán and Robins 2020, 215).


3.1 The General Homogeneity Condition

Define the effect modification by \(U\) of the effect of \(A\) on \(Y\) and of the \(Z\)-\(A\) association:

\[ e(U) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid U\right]\mathclose{}, \qquad t(U) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[A \mid Z = 1, U\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0, U\right]\mathclose{}. \]

Theorem 1 (General Homogeneity Condition) Under the causal diagram of Figure 16.1, if \(\operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} = 0\), then

\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{} = \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \]

Versions 3 and 4 are special cases: they say “\(e(U)\) is constant” and “\(t(U)\) is constant”, and the covariance of any variable with a constant is 0 (Hernán and Robins 2020, 216).


NoteTechnical Point 16.5: Proof of the General Homogeneity Condition

The proof of Theorem 1 rewrites the denominator and the numerator of the usual IV estimand as averages over the unmeasured confounders \(U\): the denominator becomes \(\operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}\) and the numerator \(\operatorname{E}\mathopen{}\left[e(U)\, t(U)\right]\mathclose{}\). Zero covariance then lets the numerator factor as \(\operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}\), and \(\operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\) is the average causal effect (Hernán and Robins 2020, 217). The proof below fills in each step.

Proof. In Figure 16.1, \(Y^a \perp\!\!\!\perp(A, Z) \mid U\) for each \(a\) and \(U \perp\!\!\!\perp Z\). Sums over \(u\) become integrals when \(U\) is continuous.

Step 1: the denominator is \(\operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}\). Because \(U \perp\!\!\!\perp Z\), \(f(u \mid z) = f(u)\), so

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[A \mid Z = z\right]\mathclose{} &= \textstyle\sum_u \operatorname{E}\mathopen{}\left[A \mid Z = z, U = u\right]\mathclose{} f(u \mid z) && \text{(law of total expectation)} \\ &= \textstyle\sum_u \operatorname{E}\mathopen{}\left[A \mid Z = z, U = u\right]\mathclose{} f(u) && (U \perp\!\!\!\perp Z) \\ \operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{} &= \textstyle\sum_u t(u) f(u) = \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{}. && \text{(definition of } t) \end{aligned} \]

Step 2: the numerator is \(\operatorname{E}\mathopen{}\left[e(U)\,t(U)\right]\mathclose{}\). By consistency, \(Y = A (Y^{a=1} - Y^{a=0}) + Y^{a=0}\). For each \(z\) and \(u\),

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[A (Y^{a=1} - Y^{a=0}) \mid Z = z, U = u\right]\mathclose{} &= \Pr(A = 1 \mid Z = z, U = u) \\ &\qquad \times \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A = 1, Z = z, U = u\right]\mathclose{} \\ &= \Pr(A = 1 \mid Z = z, U = u)\, \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid U = u\right]\mathclose{} \\ &= \Pr(A = 1 \mid Z = z, U = u)\, e(u), \\ \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = z, U = u\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y^{a=0} \mid U = u\right]\mathclose{}. \end{aligned} \]

The second equality uses \(Y^a \perp\!\!\!\perp(A, Z) \mid U\), and the last line uses \(Y^{a=0} \perp\!\!\!\perp Z \mid U\).

Averaging over \(U\) (with \(f(u \mid z) = f(u)\)):

\[ \operatorname{E}\mathopen{}\left[Y \mid Z = z\right]\mathclose{} = \textstyle\sum_u \mathopen{}\left\{e(u) \Pr(A = 1 \mid Z = z, U = u) + \operatorname{E}\mathopen{}\left[Y^{a=0} \mid U = u\right]\mathclose{}\right\}\mathclose{} f(u), \]

so the second term cancels in the difference and

\[ \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} = \textstyle\sum_u e(u)\, t(u)\, f(u) = \operatorname{E}\mathopen{}\left[e(U)\, t(U)\right]\mathclose{} . \]

Step 3: use the zero covariance.

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[e(U)\, t(U)\right]\mathclose{} &= \operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} + \operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[e(U)\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[t(U)\right]\mathclose{} && (\operatorname{Cov}\mathopen{}\left(e(U), t(U)\right)\mathclose{} = 0) \\ &= \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0}\right]\mathclose{} \mathopen{}\left(\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}\right)\mathclose{}. && \text{(iterated expectations; Step 1)} \end{aligned} \]

Dividing the numerator (Step 2) by the denominator (Step 1) gives the result.


NoteTechnical Point 16.3: Additive Structural Mean Models and IV Estimation

Saturated additive structural mean model for dichotomous \(A\) and instrument \(Z\):

\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A = 1, Z\right]\mathclose{} = \beta_0 + \beta_1 Z, \quad\text{equivalently}\quad \operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid A, Z\right]\mathclose{} = A (\beta_0 + \beta_1 Z). \]

\(\beta_0\) is the effect in the treated with \(Z = 0\), \(\beta_0 + \beta_1\) the effect in the treated with \(Z = 1\); \(\beta_1\) measures additive effect modification by \(Z\). If we assume \(\beta_1 = 0\), then \(\beta_0\) equals the usual IV estimand (Hernán and Robins 2020, 215).

Proof. The instrumental conditions imply \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 0\right]\mathclose{}\). From the model, \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A, Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A, Z\right]\mathclose{} - A(\beta_0 + \beta_1 Z)\); averaging over \(A\) given \(Z\),

\[ \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid Z\right]\mathclose{} - (\beta_0 + \beta_1 Z)\, \operatorname{E}\mathopen{}\left[A \mid Z\right]\mathclose{}. \]

Setting \(\beta_1 = 0\) and equating the \(Z = 1\) and \(Z = 0\) values:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \beta_0 \operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} - \beta_0 \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{} \\ \operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{} &= \beta_0 \mathopen{}\left(\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}\right)\mathclose{} \\ \beta_0 &= \frac{\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z = 0\right]\mathclose{}} . \end{aligned} \]

Why set \(\beta_1 = 0\)? The instrumental conditions give one equation in two unknowns (\(\beta_0\), \(\beta_1\)), and that equation exhausts the constraints they place on the data. A second constraint is needed and is necessarily untestable; \(\beta_1 = 0\) is an arbitrary choice (one could as well pick \(\beta_1 = 2\)). That is what “an instrument is insufficient to identify the average causal effect” means.

Even with \(\beta_1 = 0\), \(\beta_0\) is the effect in the treated. Interpreting it as the population effect \(\operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{}\) needs the further untestable assumption that the effect is the same in the treated and the untreated, i.e., that the corresponding \(Z\) parameter is also 0 in the model for \(A = 0\) (Hernán and Robins 2020, 215).


NoteTechnical Point 16.4: Multiplicative Structural Mean Models and IV Estimation

Saturated multiplicative (log-linear) structural mean model:

\[ \frac{\operatorname{E}\mathopen{}\left[Y^{a=1} \mid A = 1, Z\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[Y^{a=0} \mid A = 1, Z\right]\mathclose{}} = \operatorname{exp}\mathopen{}\left\{\beta_0 + \beta_1 Z\right\}\mathclose{}, \quad\text{equivalently}\quad \operatorname{E}\mathopen{}\left[Y \mid A, Z\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid A, Z\right]\mathclose{} \operatorname{exp}\mathopen{}\left\{A (\beta_0 + \beta_1 Z)\right\}\mathclose{}. \]

Assuming \(\beta_1 = 0\) and no multiplicative effect modification by \(Z\) in the untreated either, the causal risk ratio is \(\operatorname{exp}\mathopen{}\left\{\beta_0\right\}\mathclose{}\) and

\[ \operatorname{E}\mathopen{}\left[Y^{a=1}\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y^{a=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y \mid A = 0\right]\mathclose{}(1 - \operatorname{E}\mathopen{}\left[A\right]\mathclose{})\mathopen{}\left[\operatorname{exp}\mathopen{}\left\{\beta_0\right\}\mathclose{} - 1\right]\mathclose{} + \operatorname{E}\mathopen{}\left[Y \mid A = 1\right]\mathclose{}\, \operatorname{E}\mathopen{}\left[A\right]\mathclose{} \mathopen{}\left[1 - \operatorname{exp}\mathopen{}\left\{-\beta_0\right\}\mathclose{}\right]\mathclose{}. \]

This is identified but is not the usual IV estimand (Hernán and Robins 2020, 216).

So the estimate of the average causal effect depends on whether one assumes no additive or no multiplicative effect modification by \(Z\), and no amount of data can tell which (if either) is true, because the saturated models have more unknown parameters than equations (Hernán and Robins 2020, 216). The book refers to Robins (1989) and Hernán and Robins (2006b) for the proof of the formula.


3.2 Bypassing Homogeneity

Because the homogeneity conditions are often implausible, the book describes two ways around them (Hernán and Robins 2020, 216):

  1. Add baseline covariates to the IV models, preferably with structural mean models (fewer parametric assumptions than two-stage least squares). The effect in the treated may then vary with \(Z\), subject to constraints on how it varies within levels of the covariates.
  2. Replace homogeneity by a different condition (iv), monotonicity (Section 16.4), which gives the usual IV estimand a causal interpretation without identifying the population effect.
NoteTechnical Point 16.6: More General Structural Mean Models

With possibly continuous or multivariate \(A\) and \(Z\) and pre-instrument covariates \(V\), an additive structural mean model is

\[ \operatorname{E}\mathopen{}\left[Y - Y^{a=0} \mid Z, A, V\right]\mathclose{} = \gamma(Z, A, V; \beta), \]

with \(\gamma\) a known function, \(\beta\) an unknown parameter, and \(\gamma(Z, A = 0, V; \beta) = 0\). It models the effect of the observed treatment level versus level 0 among those with that \(Z\), \(V\), and observed \(A\). Its parameters are identified by g-estimation under \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 1, V\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z = 0, V\right]\mathclose{}\). The multiplicative analog is \(\operatorname{E}\mathopen{}\left[Y \mid Z, A, V\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=0} \mid Z, A, V\right]\mathclose{} \operatorname{exp}\mathopen{}\left\{\gamma(Z, A, V; \beta)\right\}\mathclose{}\) (Hernán and Robins 2020, 218).

Models also allow several proposed instruments at once, continuous treatments, and causal risk ratios for dichotomous outcomes (Hernán and Robins 2020, 216). Nested versions of these structural mean models extend IV methods to time-varying treatments and confounders (Hernán and Robins 2020, 218).

4 16.4 An Alternative Fourth Condition: Monotonicity (pp. 217-220)


In the double-blind trial, let \(A^{z=1}\) and \(A^{z=0}\) be the treatment an individual would take if assigned to treatment or to no treatment.

Definition 2 (Compliance Types (Principal Strata))  

  1. Always-takers: \(A^{z=1} = 1\) and \(A^{z=0} = 1\).
  2. Never-takers: \(A^{z=1} = 0\) and \(A^{z=0} = 0\).
  3. Compliers (cooperative): \(A^{z=1} = 1\) and \(A^{z=0} = 0\).
  4. Defiers (contrarians): \(A^{z=1} = 0\) and \(A^{z=0} = 1\).

These strata are not identified: someone with \(Z = 1\) and \(A = 1\) may be a complier or an always-taker; someone with \(Z = 1\) and \(A = 0\) may be a defier or a never-taker.

Definition 3 (Monotonicity) Monotonicity holds when there are no defiers, i.e., \(A^{z=1} \ge A^{z=0}\) for all individuals.


Theorem 2 (IV Estimand under Monotonicity) Under instrumental conditions (i)-(iii) with a dichotomous causal instrument that is randomly assigned, and monotonicity as condition (iv), the usual IV estimand equals the average causal effect in the compliers,

\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid A^{z=1} = 1, A^{z=0} = 0\right]\mathclose{}, \]

not the average causal effect in the population (Hernán and Robins 2020, 218).

Sketch: the intention-to-treat effect is a weighted average over the four strata; it is zero in always-takers and never-takers (their \(A\) does not change with \(Z\), and \(Z\) acts only through \(A\)), and there are no defiers. So the numerator is the effect in compliers times the proportion of compliers, which is the denominator.

The book’s main text refers to “Technical Point 16.6” for this proof; the proof is in Technical Point 16.7 (Hernán and Robins 2020, 218, 224).

This estimand is often called the complier average causal effect (CACE) or local average treatment effect (LATE): an effect in a subpopulation, not the population average. Greenland calls compliers “cooperative” and defiers “non-cooperative” to avoid confusion with observed compliance in trials (Hernán and Robins 2020, 219).


NoteTechnical Point 16.7: Monotonicity and the Effect in the Compliers

Imbens and Angrist (1994) showed that, for a dichotomous causal instrument \(Z\) and no defiers (monotonicity, condition (iv)), the usual IV estimand is the average causal effect in the compliers (Theorem 2); a closely related argument for binary outcomes is due to Baker and Lindeman (1994) (Hernán and Robins 2020, 224). The proof below assumes a causal instrument (Figure 16.1).

For a surrogate instrument \(Z\), Hernán and Robins (2006b) showed that the usual IV estimator still identifies the effect in the compliers (defined by \(U_Z\)), but only if \(Z\) is independent of \(A\) and \(Y\) given \(U_Z\) and \(U_Z\) is binary, and that independence is rarely plausible unless \(U_Z\) is continuous (Hernán and Robins 2020, 224). So when an observational study has at best a surrogate for \(U_Z\), interpreting the IV estimand as a complier effect is often doubtful.

Proof. Write the intention-to-treat effect as a weighted average over the strata (AT: always-takers, NT: never-takers, C: compliers, D: defiers):

\[ \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{} = \textstyle\sum_{s \in \{AT, NT, C, D\}} \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid s\right]\mathclose{}\, \Pr[s]. \]

  • In AT and NT, \(Z\) does not change \(A\), and by the individual-level exclusion restriction \(Z\) has no other effect on \(Y\), so \(\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid AT\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid NT\right]\mathclose{} = 0\).
  • Under monotonicity, \(\Pr[D] = 0\).

Hence \(\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid C\right]\mathclose{}\, \Pr[C]\). In compliers \(A = Z\), so the effect of \(Z\) equals the effect of \(A\): \(\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0} \mid C\right]\mathclose{} = \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid C\right]\mathclose{}\), and

\[ \operatorname{E}\mathopen{}\left[Y^{a=1} - Y^{a=0} \mid C\right]\mathclose{} = \frac{\operatorname{E}\mathopen{}\left[Y^{z=1} - Y^{z=0}\right]\mathclose{}}{\Pr[C]} . \]

Random assignment implies \(Z \perp\!\!\!\perp\{Y^{a,z}, A^z;\ z = 0,1;\ a = 0,1\}\), so with consistency the numerator equals \(\operatorname{E}\mathopen{}\left[Y \mid Z = 1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z = 0\right]\mathclose{}\). For the denominator, \(\Pr[AT] = \Pr[A^{z=0} = 1] = \Pr[A = 1 \mid Z = 0]\) and \(\Pr[NT] = \Pr[A^{z=1} = 0] = \Pr[A = 0 \mid Z = 1]\); with no defiers,

\[ \begin{aligned} \Pr[C] &= 1 - \Pr[AT] - \Pr[NT] \\ &= 1 - \Pr[A = 1 \mid Z = 0] - \mathopen{}\left(1 - \Pr[A = 1 \mid Z = 1]\right)\mathclose{} \\ &= \Pr[A = 1 \mid Z = 1] - \Pr[A = 1 \mid Z = 0]. \end{aligned} \]


4.1 Compliers in an Observational Study

There is no assignment, but “compliers” still means \((A^{z=1} = 1, A^{z=0} = 0)\): in NHEFS, people who would quit smoking in a high-price state and not in a low-price state. If there are no defiers and the causal instrument is dichotomous, 2.4 kg estimates the effect in the compliers (Hernán and Robins 2020, 218–19).


4.2 Criticisms of the Complier Effect

Monotonicity was welcomed in the mid-1990s as the salvation of IV methods, but the book raises four problems (Hernán and Robins 2020, 219–20):

  1. Relevance: compliers cannot be identified, and their proportion (the denominator, about 6% in NHEFS) varies across instruments and studies, so the effect is hard to use for decisions.
  2. Monotonicity may fail in observational studies, especially when treatment decisions weigh several criteria.
  3. Surrogate instruments: if the causal instrument \(U_Z\) is continuous, a dichotomous surrogate’s IV estimand is a weighted average of effects over everyone, not the effect in a subpopulation of compliers.
  4. Principal strata may be ill-defined.
  1. A mitigating factor: under strong assumptions, the distribution of observed variables among compliers can be characterized. Deaton (2010) likened focusing on the complier effect to choosing where the light falls and then claiming that whatever it illuminates is what we were looking for (Hernán and Robins 2020, 219).
  2. Monotonicity is safe in trials (participants are unlikely to consent in order to do the opposite of what they are asked), and it holds by design when those assigned to no treatment cannot get it (no always-takers or defiers); the effect in the compliers is then the effect in the treated. Example (Swanson and Hernán 2014): in a clinic, one physician usually prescribes the treatment except for patients with diabetes; the other usually does not, except for physically active patients. A physically active patient with diabetes is treated contrary to both physicians’ preferences, i.e., is a defier.
  3. For a continuous \(U_Z\), monotonicity means \(A^{u_z}\) is non-decreasing in \(u_z\).
  4. A stable partition requires, e.g., all physicians with the same preference level who could have seen a patient to treat the patient identically; it also requires deterministic counterfactuals, no interference, and stability across settings and instruments (Hernán and Robins 2020, 220).

Bottom line: if the complier effect is of interest, monotonicity is most promising in double-blind two-arm trials with all-or-nothing compliance, especially when one arm has full adherence by design; more caution is needed in complex settings and observational studies (Hernán and Robins 2020, 220).

5 16.5 The Three Instrumental Conditions Revisited (pp. 220-222)


A proposed instrument may violate (ii) or (iii), or barely meet (i). Then IV estimation may be badly biased even if condition (iv) holds perfectly.

5.1 Condition (i): Weak Instruments

NoteFine Point 16.2: Defining Weak Instruments
  1. Substantively weak: the true \(Z\)-\(A\) association (the denominator) is small.
  2. Statistically weak: the F-statistic for the observed \(Z\)-\(A\) association is small, typically below 10.

The NHEFS instrument is weak by both: risk difference about 6%, F-statistic 0.8 (Hernán and Robins 2020, 221).

Three problems (Hernán and Robins 2020, 220–21):

  1. Wide confidence intervals (NHEFS: \(-36.5\) to \(41.3\) kg).
  2. Bias amplification: any bias in the numerator (from a violation of (ii) or (iii)) is divided by the small denominator. In NHEFS it is multiplied by about \(1/0.0627 \approx 15.9\).
  3. Finite-sample bias even with a valid instrument and a large sample.

Definition 1 links to problem 2 (it holds even in an infinite sample); definition 2 links to problem 3.

In linear models, strong confounding guarantees a weak instrument: a strong \(A\)-\(U\) association leaves little room for a strong \(A\)-\(Z\) association (Martens et al. 2006) (Hernán and Robins 2020, 220).


5.2 Why Weak Instruments Are Biased in Finite Samples

A randomly generated \(Z\) has, in an infinite population, zero association with \(A\), so the IV estimand is undefined. In a finite sample, chance creates a small \(Z\)-\(U\), and hence \(Z\)-\(A\), association, so the denominator is small but not zero and the estimate can be wildly inflated.

Example 2 (The NHEFS Instrument Behaves Like Noise) Redefining \(Z\) as price above $1.60, $1.70, $1.80, or $1.90 gives IV estimates of 41.3, \(-40.9\), \(-21.1\), and \(-12.8\) kg, each with a huge confidence interval (Hernán and Robins 2020, 221).

Given the bias and variability weak instruments create, a strong proposed instrument that slightly violates (ii) or (iii) may be preferable to a weaker but less invalid one (Hernán and Robins 2020, 221). Bound, Jaeger and Baker (1995) documented this bias.


5.3 Condition (ii): Direct Effects of the Instrument

  • A direct arrow \(Z \rightarrow Y\) (Figure 16.8) adds to the numerator, and the denominator inflates it as if it were an effect of \(A\).
  • Coarsening the treatment can create a violation: if a continuous or multi-valued \(A\) is replaced by a dichotomized \(A^*\), the path \(Z \rightarrow A \rightarrow Y\) is a direct effect of \(Z\) not mediated by \(A^*\) (Figure 16.9). Coarsening is a problem for IV estimation but not necessarily for the methods of earlier chapters (Hernán and Robins 2020, 221–22).

5.4 Condition (iii): Confounding and Selection Bias

  • Common causes of \(Z\) and \(Y\) (which may or may not also cause \(A\); Figure 16.10) bias the numerator, and the bias is inflated by the denominator. In observational studies this possibility always exists.
  • Adjusting for measured pre-instrument covariates \(V\): assume no unmeasured confounding of \(Z\)-\(Y\) within levels of \(V\), then do IV estimation within strata of \(V\) and pool (assuming the effect is constant across \(V\)), or include \(V\) in the two-stage models. In NHEFS this reduced the effect estimate and widened its 95% confidence interval (Hernán and Robins 2020, 222).
  • Checking balance of measured confounders across levels of \(Z\) can give false security: small imbalances can cause large biases through amplification.
  • Selection bias: excluding individuals with some treatment levels (e.g., keeping only \(A = 1\) or \(A = 2\) when \(A \in \{0, 1, 2\}\)) can give a non-null IV estimate under the null, although it does not bias a non-IV comparison of \(A = 1\) versus \(A = 2\).

All these problems get worse when several proposed instruments are used together to compensate for weakness: the more proposed instruments, the more likely that some of them violate a condition (Hernán and Robins 2020, 222).

6 16.6 Instrumental Variable Estimation versus Other Methods (pp. 223-225)


IV estimation differs from the earlier methods in at least three ways (Hernán and Robins 2020, 223):

  1. Different assumptions: it replaces conditional exchangeability of treated and untreated with conditions (i)-(iv). The choice depends on whether it is easier to measure the confounders or to find an instrument and expect monotonicity or no relevant effect heterogeneity.
  2. Sensitivity to violations: because the denominator blows up the numerator, small violations or a weak instrument can cause large bias in either direction, possibly larger than an unadjusted estimate. Slight violations of the conditions for IP weighting or standardization tend to give only slight bias.
  3. Narrower applicability: standard IV estimation suits settings with lots of unmeasured confounding, a truly dichotomous time-fixed treatment, a strong (causal) instrument, and either effect homogeneity or (if the complier effect is of interest) monotonicity. So it is mostly used for point interventions and plays a small role in Part III (time-varying treatments).

Because it rests on different assumptions, IV estimation is useful for triangulation, but the wide confidence intervals typical of IV estimates often limit its added value, and the counterintuitive direction and size of its biases need special attention. Sensitivity analyses and transparent reporting are important (Hernán and Robins 2020, 223). Regression discontinuity (Fine Point 16.3) and difference-in-differences (Technical Point 7.3) also identify effects without conditional exchangeability.


NoteFine Point 16.3: Regression Discontinuity Design

Example: an antiviral \(A\) for COVID-19 is given to everyone aged 65 or older (\(L \ge 65\)) and to no one younger, so \(\Pr[A = 1 \mid L < 65] = 0\) and \(\Pr[A = 1 \mid L \ge 65] = 1\). There is no positivity.

  • Continuity assumption: \(\operatorname{E}\mathopen{}\left[Y^{a=1} \mid L\right]\mathclose{}\) and \(\operatorname{E}\mathopen{}\left[Y^{a=0} \mid L\right]\mathclose{}\) are continuous in \(L\) (at least around the threshold), and individuals just on either side are exchangeable.
  • Then a jump in the observed \(\operatorname{E}\mathopen{}\left[Y \mid L\right]\mathclose{}\) at \(L = 65\) estimates the effect of \(A\) among those with \(L\) near the threshold, e.g., by comparing mean outcomes within a chosen bandwidth, typically with regression models fit on each side.
  • The bandwidth trades precision against bias; data-adaptive procedures such as cross-validation can help choose it.

This is a sharp design; a fuzzy design, where the treatment probability jumps but not from 0 to 1, relies on monotonicity and estimates the effect in the compliers near the threshold (Hernán and Robins 2020, 225).

The design is biased if other treatments also change at the threshold (e.g., intensive care restricted to those 65 and older), or if people manipulate the running variable (e.g., falsify their age). The near-threshold effect can differ from the population effect if \(L\) is an effect modifier (Hernán and Robins 2020, 225).

7 Summary


  • An instrument meets (i) association with \(A\), (ii) no effect on \(Y\) except through \(A\), (iii) no shared causes with \(Y\); only (i) is testable.
  • Alone, an instrument identifies only (usually wide) bounds.
  • Usual IV estimand: \(\dfrac{\operatorname{E}\mathopen{}\left[Y \mid Z=1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y \mid Z=0\right]\mathclose{}}{\operatorname{E}\mathopen{}\left[A \mid Z=1\right]\mathclose{} - \operatorname{E}\mathopen{}\left[A \mid Z=0\right]\mathclose{}}\), computed directly or by two-stage least squares.
  • Adding homogeneity (iv) makes it the population average causal effect; adding monotonicity (iv) instead makes it the effect in the compliers.
  • NHEFS: estimate 2.4 kg (95% CI \(-36.5\) to \(41.3\)) from a weak instrument (risk difference 6%, F = 0.8).
  • Weak instruments widen intervals, amplify bias, and are biased in finite samples; violations of (ii) and (iii) are inflated by the small denominator.
Back to top

References

Hernán, Miguel A, and James M Robins. 2020. Causal Inference: What If. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook.