Variance and covariance

Published

Last modified: 2026-09-28 23:45:35 (PDT)

1 Deviation, error, and noise

Definition 1 (Deviation) A deviation is the difference between a value and a reference value. For any quantity \(z\) and reference value \(r\), the deviation of \(z\) from \(r\) is:

\[z - r\]

Remark. In probability and statistics, “deviation” usually means deviation from a population mean.

See: Wikipedia: Deviation (statistics)

Example 1 (Deviation from a reference weight) A newborn weighing 3,200 g deviates from a reference weight of 3,500 g by \(3200 - 3500 = -300\) g.

Definition 2 (Deviation from a population or subpopulation mean) The deviation from the mean of a random variable \(Y\) is its deviation from its expectation:

\[e(Y) \stackrel{\text{def}}{=}Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\]

For a realized observation \(y\):

\[e(y) \stackrel{\text{def}}{=}y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\]

Remark. Other sources often call this quantity an error or noise term. In regression settings, the reference mean is often conditional on covariates: \(e(y_i) \stackrel{\text{def}}{=}y_i - \operatorname{E}\mathopen{}\left[Y_i \mid X_i = x_i\right]\mathclose{}\).

These notes prefer “deviation” for this mean-deviation quantity; “error” and “noise” are common aliases. The Morrison Lab’s regression notes use “residual” (defined in their Linear regression chapter) for deviations from fitted values, write \(e(\cdot)\) for these model/data deviations, and reserve \(\varepsilon\mathopen{}\left(\cdot\right)\mathclose{}\) for estimator-to-estimand deviations (see Estimation).

See:

Example 2 (Deviation of a die roll from its mean) A fair die roll \(Y\) has \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} = 3.5\) (computed on the expectation page), so a roll of \(y = 5\) has deviation \(e(5) = 5 - 3.5 = 1.5\), and a roll of \(y = 2\) has deviation \(e(2) = 2 - 3.5 = -1.5\).

3 Covariance

Definition 9 (Covariance) The covariance of two random variables \(X\) and \(Y\) with \(\operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[Y^2\right]\mathclose{} < \infty\) is:

\[\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[(X - \operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right]\mathclose{}\]

Theorem 4 (Alternative formula for covariance) \[\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{} \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\]

Proof. By linearity of expectation, analogous to the proof of the simplified expression for variance:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[XY - X\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - Y\operatorname{E}\mathopen{}\left[X\right]\mathclose{} + \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\right]\mathclose{} && \text{(expand the product)} \\ &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\operatorname{E}\mathopen{}\left[X\right]\mathclose{} + \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(linearity of expectation; } \operatorname{E}\mathopen{}\left[X\right]\mathclose{}, \operatorname{E}\mathopen{}\left[Y\right]\mathclose{} \text{ are constants)} \\ &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(combine like terms)} \end{aligned} \]

Example 9 (Covariance of a binary exposure and outcome) For the joint PMF in Example 4, \(\operatorname{E}\mathopen{}\left[XY\right]\mathclose{} = \operatorname{P}(X = 1, Y = 1) = 0.4\), \(\operatorname{E}\mathopen{}\left[X\right]\mathclose{} = 0.5\), and \(\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} = 0.7\), so:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(alternative formula for covariance)} \\ &= 0.4 - 0.5 \cdot 0.7 && \text{(substitute)} \\ &= 0.05 && \text{(evaluate)} \end{aligned} \]

Theorem 5 (Independent random variables have zero covariance) If \(X\) and \(Y\) are independent, each discrete or continuous, with defined expectations, then:

\[\operatorname{E}\mathopen{}\left[XY\right]\mathclose{} = \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\]

If also \(\operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[Y^2\right]\mathclose{} < \infty\), then \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\).

Proof. Write \(f_X\) and \(f_Y\) for the PMFs or densities of \(X\) and \(Y\), with reference measures \(\mu_X\) and \(\mu_Y\) as in the joint-distribution form of Fubini–Tonelli (counting measure for a discrete variable, Lebesgue measure for a continuous one). Because \(X\) and \(Y\) are independent, \(f_{X,Y}(x, y) = f_X(x)\,f_Y(y)\) is their joint PMF, density, or density-mass function (the factorization in the notes to the definition of independence). First, with \(h(x, y) = \mathopen{}\left|x\right|\mathclose{}\mathopen{}\left|y\right|\mathclose{} \ge 0\) (condition (a)):

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left|XY\right|\mathclose{}\right]\mathclose{} &= \int\mathopen{}\left(\int \mathopen{}\left|x\right|\mathclose{}\mathopen{}\left|y\right|\mathclose{}\,f_X(x)\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(joint-distribution form of Fubini--Tonelli, condition (a))} \\ &= \int \mathopen{}\left|x\right|\mathclose{}\,f_X(x)\mathopen{}\left(\int \mathopen{}\left|y\right|\mathclose{}\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(} \mathopen{}\left|x\right|\mathclose{}\,f_X(x) \text{ does not depend on } y \text{)} \\ &= \int \mathopen{}\left|x\right|\mathclose{}\,f_X(x) \cdot\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{}\,d\mu_X(x) && \text{(LOTUS for } \mathopen{}\left|Y\right|\mathclose{} \text{)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left|X\right|\mathclose{}\right]\mathclose{} \cdot\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} && \text{(LOTUS for } \mathopen{}\left|X\right|\mathclose{} \text{)} \end{aligned} \]

which is finite, because \(X\) and \(Y\) have defined expectations. So condition (b) holds for \(h(x, y) = xy\), and the same steps without the absolute values give:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} &= \int\mathopen{}\left(\int xy\,f_X(x)\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(joint-distribution form of Fubini--Tonelli, condition (b))} \\ &= \int x\,f_X(x)\mathopen{}\left(\int y\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(} x\,f_X(x) \text{ does not depend on } y \text{)} \\ &= \int x\,f_X(x) \cdot\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\,d\mu_X(x) && \text{(definition of expectation)} \\ &= \operatorname{E}\mathopen{}\left[X\right]\mathclose{} \cdot\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(definition of expectation)} \end{aligned} \]

Finally, by the alternative formula for covariance:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(alternative formula for covariance)} \\ &= \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} = \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} \text{)} \\ &= 0 && \text{(subtract)} \end{aligned} \]

Example 10 (Two independent coin flips) For the independent coin flips \(X_1\) and \(X_2\) of the independence page’s example, \(X_1 X_2 = 1\) only when both flips are heads, so:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[X_1 X_2\right]\mathclose{} &= \operatorname{P}(X_1 = 1, X_2 = 1) && \text{(} X_1 X_2 \text{ is the indicator of two heads)} \\ &= \tfrac{1}{4} && \text{(each outcome has probability } \tfrac{1}{4} \text{)} \\ &= \tfrac{1}{2} \cdot\tfrac{1}{2} && \text{(factor)} \\ &= \operatorname{E}\mathopen{}\left[X_1\right]\mathclose{}\operatorname{E}\mathopen{}\left[X_2\right]\mathclose{} && \text{(each flip is } \operatorname{Ber}(1/2) \text{)} \end{aligned} \]

so \(\operatorname{Cov}\mathopen{}\left(X_1, X_2\right)\mathclose{} = 0\), as Theorem 5 requires.

Example 11 (Zero covariance without independence) The converse of Theorem 5 is false. Let \(X\) take the values \(-1\), \(0\), and \(1\) with probability \(1/3\) each, and let \(Y = X^2\). Then \(\operatorname{E}\mathopen{}\left[X\right]\mathclose{} = (-1 + 0 + 1)/3 = 0\), and \(XY = X^3 = X\), so:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(alternative formula for covariance)} \\ &= \operatorname{E}\mathopen{}\left[X\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(} XY = X^3 = X \text{ on } \mathopen{}\left\{-1, 0, 1\right\}\mathclose{} \text{)} \\ &= 0 - 0 \cdot\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[X\right]\mathclose{} = 0 \text{)} \\ &= 0 && \text{(multiply)} \end{aligned} \]

But \(X\) and \(Y\) are not independent: \(Y\) is a function of \(X\), and \(\operatorname{P}(X = 0, Y = 0) = \operatorname{P}(X = 0) = \tfrac{1}{3}\), while \(\operatorname{P}(X = 0)\,\operatorname{P}(Y = 0) = \tfrac{1}{3} \cdot\tfrac{1}{3} = \tfrac{1}{9}\).

Definition 10 (Correlation) The correlation of two random variables \(X\) and \(Y\) with finite, positive variances is their covariance divided by the product of their standard deviations:

\[\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \stackrel{\text{def}}{=}\frac{\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{}}\]

Remark. Dividing by the standard deviations removes the units of \(X\) and \(Y\). This population correlation is a property of a joint distribution; the sample (Pearson) correlation coefficient computed from data estimates it.

Example 12 (Correlation of a binary exposure and outcome) In Example 9, \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0.05\), where \(X \sim \operatorname{Ber}(0.5)\) has \(\operatorname{Var}\mathopen{}\left(X\right)\mathclose{} = 0.25\) (Example 3) and \(Y\) has \(\operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} = 0.21\) (Example 7), so:

\[ \begin{aligned} \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} &= \frac{\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{}} && \text{(definition of correlation)} \\ &= \frac{0.05}{\sqrt{0.25} \cdot\sqrt{0.21}} && \text{(substitute; } \operatorname{SD}\mathopen{}\left(\cdot\right)\mathclose{} = \sqrt{\operatorname{Var}\mathopen{}\left(\cdot\right)\mathclose{}} \text{)} \\ &\approx \frac{0.05}{0.5 \cdot 0.458} && \text{(evaluate the square roots)} \\ &\approx 0.218 && \text{(divide)} \end{aligned} \]

Definition 11 (Uncorrelated random variables) Random variables \(X\) and \(Y\) with \(\operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[Y^2\right]\mathclose{} < \infty\) are uncorrelated when their covariance is 0:

\[\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\]

Remark. Theorem 5 says independent random variables are uncorrelated, and Example 11 shows the converse fails.

Example 13 (Correlated and uncorrelated pairs) In Example 11, \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\), so \(X\) and \(Y = X^2\) are uncorrelated. In Example 9, \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0.05 \neq 0\), so the binary exposure and outcome are not uncorrelated.

Corollary 1 (Uncorrelated means zero correlation) If \(X\) and \(Y\) have finite, positive variances, then \(X\) and \(Y\) are uncorrelated if and only if \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} = 0\).

Proof. By Definition 10, \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} = \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} / \mathopen{}\left(\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{}\right)\mathclose{}\), and \(\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{} > 0\) because both variances are positive, so \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} = 0\) if and only if \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\), which is Definition 11.

Definition 12 (Conditional covariance) The conditional covariance of \(Y\) and \(Z\) given \(X = x\) is their covariance under their conditional distribution given \(X = x\):

\[\operatorname{Cov}\mathopen{}\left(Y,Z \mid X = x\right)\mathclose{} \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\mathopen{}\left(Y-\operatorname{E}\mathopen{}\left[Y \mid X = x\right]\mathclose{}\right)\mathclose{}\mathopen{}\left(Z-\operatorname{E}\mathopen{}\left[Z \mid X = x\right]\mathclose{}\right)\mathclose{} \mid X = x\right]\mathclose{}\]

Evaluating this function of \(x\) at \(X\) gives the random variable \(\operatorname{Cov}\mathopen{}\left(Y,Z \mid X\right)\mathclose{}\).

Example 14 (Conditional covariance of a variable with itself) Taking \(Z = Y\) in Definition 12 gives the conditional variance: \(\operatorname{Cov}\mathopen{}\left(Y,Y \mid X = x\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(Y \mid X = x\right)\mathclose{}\). In Example 4, for example, \(\operatorname{Cov}\mathopen{}\left(Y,Y \mid X = 0\right)\mathclose{} = 0.24\).

Theorem 6 (Law of total covariance) For random variables \(X\), \(Y\), and \(Z\) with \(\operatorname{E}\mathopen{}\left[Y^2\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[Z^2\right]\mathclose{} < \infty\):

\[\operatorname{Cov}\mathopen{}\left(Y,Z\right)\mathclose{} = \operatorname{E}\mathopen{}\left[\operatorname{Cov}\mathopen{}\left(Y,Z \mid X\right)\mathclose{}\right]\mathclose{} + \operatorname{Cov}\mathopen{}\left(\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}, \operatorname{E}\mathopen{}\left[Z \mid X\right]\mathclose{}\right)\mathclose{}\]

Proof. The proof follows the proof of the law of total variance, which is the special case \(Z = Y\). Write \(m_Y(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\), \(m_Z(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Z \mid X\right]\mathclose{}\), \(\mu_Y \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\), and \(\mu_Z \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Z\right]\mathclose{}\). Adding and subtracting \(m_Y(X)\) and \(m_Z(X)\) inside the deviations:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(Y,Z\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(Y - \mu_Y\right)\mathclose{}\mathopen{}\left(Z - \mu_Z\right)\mathclose{}\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\mathopen{}\left[Y - m_Y(X)\right]\mathclose{} + \mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\right)\mathclose{}\mathopen{}\left(\mathopen{}\left[Z - m_Z(X)\right]\mathclose{} + \mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right)\mathclose{}\right]\mathclose{} && \text{(add and subtract)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} \\&\quad + \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} && \text{(expand the product; linearity of expectation)} \end{aligned} \]

The two middle terms are 0, by the same steps as the cross term in the proof of Theorem 3: condition on \(X\) by the law of iterated expectations, factor out the function of \(X\) (pull-out property), and use \(\operatorname{E}\mathopen{}\left[Y - m_Y(X) \mid X\right]\mathclose{} = 0\) (or \(\operatorname{E}\mathopen{}\left[Z - m_Z(X) \mid X\right]\mathclose{} = 0\)), which follows from linearity of conditional expectation. For the first term:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{} \mid X\right]\mathclose{}\right]\mathclose{} && \text{(law of iterated expectations)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{Cov}\mathopen{}\left(Y,Z \mid X\right)\mathclose{}\right]\mathclose{} && \text{(definition of conditional covariance)} \end{aligned} \]

For the last term, the law of iterated expectations gives \(\operatorname{E}\mathopen{}\left[m_Y(X)\right]\mathclose{} = \mu_Y\) and \(\operatorname{E}\mathopen{}\left[m_Z(X)\right]\mathclose{} = \mu_Z\), so:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \operatorname{E}\mathopen{}\left[m_Y(X)\right]\mathclose{}\right]\mathclose{}\mathopen{}\left[m_Z(X) - \operatorname{E}\mathopen{}\left[m_Z(X)\right]\mathclose{}\right]\mathclose{}\right]\mathclose{} && \text{(substitute the means)} \\ &= \operatorname{Cov}\mathopen{}\left(m_Y(X), m_Z(X)\right)\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{Cov}\mathopen{}\left(\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}, \operatorname{E}\mathopen{}\left[Z \mid X\right]\mathclose{}\right)\mathclose{} && \text{(definitions of } m_Y, m_Z \text{)} \end{aligned} \]

Adding the four terms gives the result.

Remark. Alternate names include: the covariance decomposition formula and the conditional covariance formula.

Lemma 1 (The covariance of a variable with itself is its variance) For any random variable \(X\) with \(\operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} < \infty\):

\[\operatorname{Cov}\mathopen{}\left(X,X\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(X\right)\mathclose{}\]

Proof. By the alternative formula for covariance:

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,X\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XX\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[X\right]\mathclose{} && \text{(alternative formula for covariance, with } Y = X \text{)} \\&= \operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[X\right]\mathclose{}\right)^2\mathclose{} && \text{(} XX = X^2 \text{)} \\&= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} && \text{(simplified expression for variance)} \end{aligned} \]

Definition 13 (Variance/covariance of a \(p \times 1\) random vector) For a \(p \times 1\) dimensional random vector \(\tilde{X}\),

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{Cov}\mathopen{}\left(\tilde{X}\right)\mathclose{} \\ &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top}\right]\mathclose{} \end{aligned} \]

Theorem 7 (Elements of the variance-covariance matrix are pairwise covariances) For a \(p \times 1\) random vector \(\tilde{X}= {(X_1, \ldots, X_p)}^{\top}\), the \((i,j)\)-th element of \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{}\):

\[ \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}= \begin{pmatrix} \operatorname{Var}\mathopen{}\left(X_1\right)\mathclose{} & \operatorname{Cov}\mathopen{}\left(X_1, X_2\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_1, X_p\right)\mathclose{} \\ \operatorname{Cov}\mathopen{}\left(X_2, X_1\right)\mathclose{} & \operatorname{Var}\mathopen{}\left(X_2\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_2, X_p\right)\mathclose{} \\ \vdots & \vdots & \ddots & \vdots \\ \operatorname{Cov}\mathopen{}\left(X_p, X_1\right)\mathclose{} & \operatorname{Cov}\mathopen{}\left(X_p, X_2\right)\mathclose{} & \cdots & \operatorname{Var}\mathopen{}\left(X_p\right)\mathclose{} \end{pmatrix} \]

Proof. Let \(\mu_i = \operatorname{E}\mathopen{}\left[X_i\right]\mathclose{}\) for \(i = 1, \ldots, p\), so \(\operatorname{E}\tilde{X}= {(\mu_1, \ldots, \mu_p)}^{\top}\). By Definition 13:

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[ \mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[ \begin{pmatrix}X_1 - \mu_1 \\ \vdots \\ X_p - \mu_p\end{pmatrix} \begin{pmatrix}X_1 - \mu_1 & \cdots & X_p - \mu_p\end{pmatrix} \right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[ \begin{pmatrix} (X_1 - \mu_1)(X_1 - \mu_1) & \cdots & (X_1 - \mu_1)(X_p - \mu_p) \\ \vdots & \ddots & \vdots \\ (X_p - \mu_p)(X_1 - \mu_1) & \cdots & (X_p - \mu_p)(X_p - \mu_p) \end{pmatrix} \right]\mathclose{} \\ &= \begin{pmatrix} \operatorname{E}\mathopen{}\left[(X_1 - \mu_1)(X_1 - \mu_1)\right]\mathclose{} & \cdots & \operatorname{E}\mathopen{}\left[(X_1 - \mu_1)(X_p - \mu_p)\right]\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{E}\mathopen{}\left[(X_p - \mu_p)(X_1 - \mu_1)\right]\mathclose{} & \cdots & \operatorname{E}\mathopen{}\left[(X_p - \mu_p)(X_p - \mu_p)\right]\mathclose{} \end{pmatrix} \\ &= \begin{pmatrix} \operatorname{Cov}\mathopen{}\left(X_1, X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_1, X_p\right)\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{Cov}\mathopen{}\left(X_p, X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_p, X_p\right)\mathclose{} \end{pmatrix} \\ &= \begin{pmatrix} \operatorname{Var}\mathopen{}\left(X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_1, X_p\right)\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{Cov}\mathopen{}\left(X_p, X_1\right)\mathclose{} & \cdots & \operatorname{Var}\mathopen{}\left(X_p\right)\mathclose{} \end{pmatrix} \end{aligned} \]

where:

Theorem 8 (Alternate expression for variance of a random vector) \[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[\tilde{X}{\tilde{X}}^{\top}\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \end{aligned} \]

Proof. \[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[ \mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} && \text{(definition)} \\ &= \operatorname{E}\mathopen{}\left[ \tilde{X}{\tilde{X}}^{\top} - \tilde{X}{\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\tilde{X}}^{\top} + \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} && \text{(expand the product)} \\ &= \operatorname{E}\mathopen{}\left[\tilde{X}{\tilde{X}}^{\top}\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} + \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} && \text{(linearity, element-wise; } \operatorname{E}\tilde{X}\text{ is constant)} \\ &= \operatorname{E}\mathopen{}\left[\tilde{X}{\tilde{X}}^{\top}\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} && \text{(combine like terms)} \end{aligned} \]

Theorem 9 (Variance of a linear combination) For any vector of random variables \(\tilde{X}= (X_1, \ldots, X_n)\) and corresponding vector of constants \(\tilde{a}= (a_1, \ldots, a_n)\), the variance of their linear combination is:

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{a}\cdot \tilde{X}\right)\mathclose{} &= \operatorname{Var}\mathopen{}\left(\sum_{i=1}^na_i X_i\right)\mathclose{} \\ &= {\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a} \\ &= \sum_{i=1}^n\sum_{j=1}^n a_i a_j \operatorname{Cov}\mathopen{}\left(X_i,X_j\right)\mathclose{} \end{aligned} \]

Proof. Treat \(\tilde{a}\) and \(\tilde{X}\) as \(n \times 1\) column vectors, so \(\tilde{a}\cdot \tilde{X}= {\tilde{a}}^{\top}\tilde{X}= \sum_{i=1}^na_i X_i\), a scalar. By linearity of expectation, \(\operatorname{E}\mathopen{}\left[{\tilde{a}}^{\top}\tilde{X}\right]\mathclose{} = {\tilde{a}}^{\top}\,\operatorname{E}\tilde{X}\), so:

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left({\tilde{a}}^{\top}\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left({\tilde{a}}^{\top}\tilde{X}- {\tilde{a}}^{\top}\operatorname{E}\tilde{X}\right)\mathclose{}^2\right]\mathclose{} && \text{(definition of variance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left({\tilde{a}}^{\top}\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}\right)\mathclose{}^2\right]\mathclose{} && \text{(factor out } {\tilde{a}}^{\top} \text{)} \\ &= \operatorname{E}\mathopen{}\left[{\tilde{a}}^{\top}\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}{\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top}\tilde{a}\right]\mathclose{} && \text{(a scalar equals its transpose, so } s^2 = s\,{s}^{\top} \text{)} \\ &= {\tilde{a}}^{\top}\,\operatorname{E}\mathopen{}\left[\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}{\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top}\right]\mathclose{}\,\tilde{a} && \text{(linearity of expectation, element-wise)} \\ &= {\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a} && \text{(variance of a random vector)} \\ &= \sum_{i=1}^n\sum_{j=1}^n a_i a_j \operatorname{Cov}\mathopen{}\left(X_i,X_j\right)\mathclose{} && \text{(expand the quadratic form, using the elements of } \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \text{)} \end{aligned} \]

Corollary 2 (Variance of a sum of two random variables) For any two random variables \(X\) and \(Y\) and scalars \(a\) and \(b\):

\[\operatorname{Var}\mathopen{}\left(aX + bY\right)\mathclose{} = a^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + b^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2(a \cdot b) \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}\]

Proof. Apply Theorem 9 with \(n=2\), \(X_1 = X\), and \(X_2 = Y\):

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(aX+bY\right)\mathclose{} &= a^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + b^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2ab \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} \end{aligned} \]

Alternatively, by linearity of expectation:

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(aX+bY\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\mathopen{}\left(aX+bY - \operatorname{E}\mathopen{}\left[aX+bY\right]\mathclose{}\right)\mathclose{}^2\right]\mathclose{} && \text{(definition of variance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(a(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{}) + b(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right)\mathclose{}^2\right]\mathclose{} && \text{(linearity of expectation)} \\ &= \operatorname{E}\mathopen{}\left[a^2(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})^2 + 2(a \cdot b)(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}) + b^2(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(expand the square)} \\ &= a^2\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})^2\right]\mathclose{} + 2(a \cdot b)\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right]\mathclose{} + b^2\operatorname{E}\mathopen{}\left[(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(linearity of expectation)} \\ &= a^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + 2(a \cdot b) \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} + b^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(definitions of variance and covariance)} \end{aligned} \]

Corollary 3 (Variance of a sum of independent random variables) If \(X\) and \(Y\) are independent, each discrete or continuous, with \(\operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} < \infty\) and \(\operatorname{E}\mathopen{}\left[Y^2\right]\mathclose{} < \infty\), then:

\[\operatorname{Var}\mathopen{}\left(X + Y\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{}\]

Proof. By Theorem 5, \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\). Applying Corollary 2 with \(a = b = 1\):

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(X + Y\right)\mathclose{} &= 1^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + 1^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2(1 \cdot 1) \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(variance of a sum of two random variables)} \\ &= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(simplify)} \\ &= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0 \text{)} \end{aligned} \]

Remark. Corollary 2 is why two variables’ covariance matters for combining them: the covariance term is what separates the variance of their sum from the sum of their variances, and Corollary 3 is the case where that term is 0.

Theorem 10 (A variance matrix is symmetric and positive semidefinite) For a \(p \times 1\) random vector \(\tilde{X}= {(X_1, \ldots, X_p)}^{\top}\) with \(\operatorname{E}\mathopen{}\left[X_i^2\right]\mathclose{} < \infty\) for every \(i\), \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is symmetric and positive semidefinite: for every \(p \times 1\) vector of constants \(\tilde{a}\),

\[{\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a}\ge 0\]

Proof. By Theorem 7, the \((i,j)\)-th element of \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{}\), and by Definition 9, with \(\mu_i = \operatorname{E}\mathopen{}\left[X_i\right]\mathclose{}\):

\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[(X_i - \mu_i)(X_j - \mu_j)\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[(X_j - \mu_j)(X_i - \mu_i)\right]\mathclose{} && \text{(multiplication of numbers is commutative)} \\ &= \operatorname{Cov}\mathopen{}\left(X_j, X_i\right)\mathclose{} && \text{(definition of covariance)} \end{aligned} \]

so the \((i,j)\)-th and \((j,i)\)-th elements are equal, and \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is symmetric.

For positive semidefiniteness, let \(Y = {\tilde{a}}^{\top}\tilde{X}\). Then:

\[ \begin{aligned} {\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a} &= \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(variance of a linear combination)} \\ &= \operatorname{E}\mathopen{}\left[(Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(definition of variance)} \\ &\ge 0 && \text{(} (Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2 \ge 0 \text{)} \end{aligned} \]

The first step is Theorem 9. The last step holds because a random variable that is never negative has a non-negative expectation: in the discrete and continuous cases of the definition of expectation, every term of the sum, or the integrand, is non-negative (for the general case, see Billingsley (1995)).

Example 15 (Two variance matrices that differ only in sign) Let \(\tilde{X}= {(X_1, X_2)}^{\top}\) be equally likely to be each of the four points

\[D_1 = \mathopen{}\left\{(-5, 1),\ (0, -1),\ (0, 1),\ (5, -1)\right\}\mathclose{},\]

and let \(\tilde{X}'\) be equally likely to be each of the four points

\[D_2 = \mathopen{}\left\{(5, 1),\ (0, -1),\ (0, 1),\ (-5, -1)\right\}\mathclose{}.\]

Both sets have first coordinates \(\mathopen{}\left\{-5, 0, 0, 5\right\}\mathclose{}\) and second coordinates \(\mathopen{}\left\{1, -1, 1, -1\right\}\mathclose{}\), so \(\operatorname{E}\mathopen{}\left[X_1\right]\mathclose{} = \operatorname{E}\mathopen{}\left[X_2\right]\mathclose{} = 0\), \(\operatorname{E}\mathopen{}\left[X_1^2\right]\mathclose{} = (25 + 0 + 0 + 25)/4 = 12.5\), and \(\operatorname{E}\mathopen{}\left[X_2^2\right]\mathclose{} = (1 + 1 + 1 + 1)/4 = 1\), and likewise for \(\tilde{X}'\). They differ in the products of the coordinates:

\[ \begin{aligned} \operatorname{E}\mathopen{}\left[X_1 X_2\right]\mathclose{} &= \frac{(-5)(1) + (0)(-1) + (0)(1) + (5)(-1)}{4} = -2.5, \\ \operatorname{E}\mathopen{}\left[X_1' X_2'\right]\mathclose{} &= \frac{(5)(1) + (0)(-1) + (0)(1) + (-5)(-1)}{4} = 2.5. \end{aligned} \]

Since the means are 0, Theorem 7 and Theorem 4 give:

\[ \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} = \begin{pmatrix}12.5 & -2.5 \\ -2.5 & 1\end{pmatrix}, \qquad \operatorname{Var}\mathopen{}\left(\tilde{X}'\right)\mathclose{} = \begin{pmatrix}12.5 & 2.5 \\ 2.5 & 1\end{pmatrix}. \]

Both are symmetric. Theorem 9 with \(\tilde{a}= {(1, 1)}^{\top}\) and \(\tilde{a}= {(1, -1)}^{\top}\) gives:

\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(X_1 + X_2\right)\mathclose{} &= 12.5 + 1 + 2(-2.5) = 8.5, & \operatorname{Var}\mathopen{}\left(X_1 - X_2\right)\mathclose{} &= 12.5 + 1 - 2(-2.5) = 18.5, \\ \operatorname{Var}\mathopen{}\left(X_1' + X_2'\right)\mathclose{} &= 12.5 + 1 + 2(2.5) = 18.5, & \operatorname{Var}\mathopen{}\left(X_1' - X_2'\right)\mathclose{} &= 12.5 + 1 - 2(2.5) = 8.5. \end{aligned} \]

As a check, \(X_1 + X_2\) takes the values \(-4, -1, 1, 4\), each with probability \(1/4\), so its mean is 0 and its variance is \((16 + 1 + 1 + 16)/4 = 8.5\).

Remark. The two sets of points have the same spread along each coordinate axis, so the diagonals of their variance matrices agree. The sign of the covariance says which diagonal direction, \({(1, 1)}^{\top}\) or \({(1, -1)}^{\top}\), has the larger spread.

Example 16 (A variance matrix that is not positive definite) Let \(X_1\) have variance \(\sigma^2> 0\), and let \(X_2 = X_1\). Every covariance in \(\tilde{X}= {(X_1, X_2)}^{\top}\) is \(\operatorname{Cov}\mathopen{}\left(X_1, X_1\right)\mathclose{} = \sigma^2\) (Lemma 1), so:

\[ \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} = \sigma^2\begin{pmatrix}1 & 1 \\ 1 & 1\end{pmatrix}. \]

With \(\tilde{a}= {(1, -1)}^{\top}\), \({\tilde{a}}^{\top}\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\tilde{a}= \operatorname{Var}\mathopen{}\left(X_1 - X_2\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(0\right)\mathclose{} = 0\), so \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is positive semidefinite but not positive definite (compare this example). If \(X_1\) is continuous, \(\tilde{X}\) has no joint density.

Corollary 4 (Independent components give a diagonal variance matrix) Let \(\tilde{X}= {(X_1, \ldots, X_p)}^{\top}\) be a random vector whose components are each discrete or continuous, with:

  • \(\operatorname{E}\mathopen{}\left[X_i^2\right]\mathclose{} < \infty\) for every \(i\), and
  • \(X_i\) and \(X_j\) independent for every \(i \neq j\).

Then \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is diagonal, with diagonal elements \(\operatorname{Var}\mathopen{}\left(X_1\right)\mathclose{}, \ldots, \operatorname{Var}\mathopen{}\left(X_p\right)\mathclose{}\).

Proof. By Theorem 7, the \((i,j)\)-th element of \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{}\). For \(i \neq j\), \(X_i\) and \(X_j\) are independent, so \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{} = 0\) by Theorem 5. The \((i,i)\)-th element is \(\operatorname{Cov}\mathopen{}\left(X_i, X_i\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(X_i\right)\mathclose{}\) by Lemma 1.

Theorem 11 (Correlation lies between \(-1\) and \(1\)) If \(X\) and \(Y\) have finite, positive variances, then their correlation lies in \([-1, 1]\) (Casella and Berger 2002):

\[-1 \le \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \le 1\]

Proof. Write \(\sigma_X \stackrel{\text{def}}{=}\operatorname{SD}\mathopen{}\left(X\right)\mathclose{} > 0\) and \(\sigma_Y \stackrel{\text{def}}{=}\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{} > 0\), and take either sign \(\pm\) throughout. A variance is the expectation of a squared deviation, a non-negative random variable, so it is non-negative. Applying Corollary 2 with \(a = 1/\sigma_X\) and \(b = \pm 1/\sigma_Y\):

\[ \begin{aligned} 0 &\le \operatorname{Var}\mathopen{}\left(\frac{X}{\sigma_X} \pm \frac{Y}{\sigma_Y}\right)\mathclose{} && \text{(a variance is non-negative)} \\ &= \frac{\operatorname{Var}\mathopen{}\left(X\right)\mathclose{}}{\sigma_X^2} + \frac{\operatorname{Var}\mathopen{}\left(Y\right)\mathclose{}}{\sigma_Y^2} \pm \frac{2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\sigma_X \sigma_Y} && \text{(variance of a sum of two random variables)} \\ &= 1 + 1 \pm \frac{2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\sigma_X \sigma_Y} && \text{(definition of standard deviation: } \sigma_X^2 = \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} \text{, } \sigma_Y^2 = \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} \text{)} \\ &= 2 \pm 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(definition of correlation)} \end{aligned} \]

With the \(+\) sign, \(0 \le 2 + 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{}\) gives \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \ge -1\). With the \(-\) sign, \(0 \le 2 - 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{}\) gives \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \le 1\).

References

Billingsley, Patrick. 1995. Probability and Measure. 3rd ed. Wiley Series in Probability and Mathematical Statistics. Wiley.
Casella, George, and Roger Berger. 2002. Statistical Inference. 2nd ed. Cengage Learning. https://www.cengage.com/c/statistical-inference-2e-casella-berger/9780534243128/.
Back to top