Variance and covariance
1 Deviation, error, and noise
3 Covariance
Proof. By linearity of expectation, analogous to the proof of the simplified expression for variance:
\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[XY - X\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - Y\operatorname{E}\mathopen{}\left[X\right]\mathclose{} + \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\right]\mathclose{} && \text{(expand the product)} \\ &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\operatorname{E}\mathopen{}\left[X\right]\mathclose{} + \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(linearity of expectation; } \operatorname{E}\mathopen{}\left[X\right]\mathclose{}, \operatorname{E}\mathopen{}\left[Y\right]\mathclose{} \text{ are constants)} \\ &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(combine like terms)} \end{aligned} \]
Proof. Write \(f_X\) and \(f_Y\) for the PMFs or densities of \(X\) and \(Y\), with reference measures \(\mu_X\) and \(\mu_Y\) as in the joint-distribution form of Fubini–Tonelli (counting measure for a discrete variable, Lebesgue measure for a continuous one). Because \(X\) and \(Y\) are independent, \(f_{X,Y}(x, y) = f_X(x)\,f_Y(y)\) is their joint PMF, density, or density-mass function (the factorization in the notes to the definition of independence). First, with \(h(x, y) = \mathopen{}\left|x\right|\mathclose{}\mathopen{}\left|y\right|\mathclose{} \ge 0\) (condition (a)):
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left|XY\right|\mathclose{}\right]\mathclose{} &= \int\mathopen{}\left(\int \mathopen{}\left|x\right|\mathclose{}\mathopen{}\left|y\right|\mathclose{}\,f_X(x)\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(joint-distribution form of Fubini--Tonelli, condition (a))} \\ &= \int \mathopen{}\left|x\right|\mathclose{}\,f_X(x)\mathopen{}\left(\int \mathopen{}\left|y\right|\mathclose{}\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(} \mathopen{}\left|x\right|\mathclose{}\,f_X(x) \text{ does not depend on } y \text{)} \\ &= \int \mathopen{}\left|x\right|\mathclose{}\,f_X(x) \cdot\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{}\,d\mu_X(x) && \text{(LOTUS for } \mathopen{}\left|Y\right|\mathclose{} \text{)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left|X\right|\mathclose{}\right]\mathclose{} \cdot\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} && \text{(LOTUS for } \mathopen{}\left|X\right|\mathclose{} \text{)} \end{aligned} \]
which is finite, because \(X\) and \(Y\) have defined expectations. So condition (b) holds for \(h(x, y) = xy\), and the same steps without the absolute values give:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} &= \int\mathopen{}\left(\int xy\,f_X(x)\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(joint-distribution form of Fubini--Tonelli, condition (b))} \\ &= \int x\,f_X(x)\mathopen{}\left(\int y\,f_Y(y)\,d\mu_Y(y)\right)\mathclose{}\,d\mu_X(x) && \text{(} x\,f_X(x) \text{ does not depend on } y \text{)} \\ &= \int x\,f_X(x) \cdot\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\,d\mu_X(x) && \text{(definition of expectation)} \\ &= \operatorname{E}\mathopen{}\left[X\right]\mathclose{} \cdot\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(definition of expectation)} \end{aligned} \]
Finally, by the alternative formula for covariance:
\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(alternative formula for covariance)} \\ &= \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(} \operatorname{E}\mathopen{}\left[XY\right]\mathclose{} = \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{} \text{)} \\ &= 0 && \text{(subtract)} \end{aligned} \]
Proof. By Definition 10, \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} = \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} / \mathopen{}\left(\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{}\right)\mathclose{}\), and \(\operatorname{SD}\mathopen{}\left(X\right)\mathclose{}\,\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{} > 0\) because both variances are positive, so \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} = 0\) if and only if \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\), which is Definition 11.
Proof. The proof follows the proof of the law of total variance, which is the special case \(Z = Y\). Write \(m_Y(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\), \(m_Z(X) \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Z \mid X\right]\mathclose{}\), \(\mu_Y \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}\), and \(\mu_Z \stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[Z\right]\mathclose{}\). Adding and subtracting \(m_Y(X)\) and \(m_Z(X)\) inside the deviations:
\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(Y,Z\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(Y - \mu_Y\right)\mathclose{}\mathopen{}\left(Z - \mu_Z\right)\mathclose{}\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(\mathopen{}\left[Y - m_Y(X)\right]\mathclose{} + \mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\right)\mathclose{}\mathopen{}\left(\mathopen{}\left[Z - m_Z(X)\right]\mathclose{} + \mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right)\mathclose{}\right]\mathclose{} && \text{(add and subtract)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} \\&\quad + \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} && \text{(expand the product; linearity of expectation)} \end{aligned} \]
The two middle terms are 0, by the same steps as the cross term in the proof of Theorem 3: condition on \(X\) by the law of iterated expectations, factor out the function of \(X\) (pull-out property), and use \(\operatorname{E}\mathopen{}\left[Y - m_Y(X) \mid X\right]\mathclose{} = 0\) (or \(\operatorname{E}\mathopen{}\left[Z - m_Z(X) \mid X\right]\mathclose{} = 0\)), which follows from linearity of conditional expectation. For the first term:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[\mathopen{}\left[Y - m_Y(X)\right]\mathclose{}\mathopen{}\left[Z - m_Z(X)\right]\mathclose{} \mid X\right]\mathclose{}\right]\mathclose{} && \text{(law of iterated expectations)} \\ &= \operatorname{E}\mathopen{}\left[\operatorname{Cov}\mathopen{}\left(Y,Z \mid X\right)\mathclose{}\right]\mathclose{} && \text{(definition of conditional covariance)} \end{aligned} \]
For the last term, the law of iterated expectations gives \(\operatorname{E}\mathopen{}\left[m_Y(X)\right]\mathclose{} = \mu_Y\) and \(\operatorname{E}\mathopen{}\left[m_Z(X)\right]\mathclose{} = \mu_Z\), so:
\[ \begin{aligned} \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \mu_Y\right]\mathclose{}\mathopen{}\left[m_Z(X) - \mu_Z\right]\mathclose{}\right]\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left[m_Y(X) - \operatorname{E}\mathopen{}\left[m_Y(X)\right]\mathclose{}\right]\mathclose{}\mathopen{}\left[m_Z(X) - \operatorname{E}\mathopen{}\left[m_Z(X)\right]\mathclose{}\right]\mathclose{}\right]\mathclose{} && \text{(substitute the means)} \\ &= \operatorname{Cov}\mathopen{}\left(m_Y(X), m_Z(X)\right)\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{Cov}\mathopen{}\left(\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}, \operatorname{E}\mathopen{}\left[Z \mid X\right]\mathclose{}\right)\mathclose{} && \text{(definitions of } m_Y, m_Z \text{)} \end{aligned} \]
Adding the four terms gives the result.
Proof. By the alternative formula for covariance:
\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X,X\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[XX\right]\mathclose{} - \operatorname{E}\mathopen{}\left[X\right]\mathclose{}\operatorname{E}\mathopen{}\left[X\right]\mathclose{} && \text{(alternative formula for covariance, with } Y = X \text{)} \\&= \operatorname{E}\mathopen{}\left[X^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[X\right]\mathclose{}\right)^2\mathclose{} && \text{(} XX = X^2 \text{)} \\&= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} && \text{(simplified expression for variance)} \end{aligned} \]
Proof. Let \(\mu_i = \operatorname{E}\mathopen{}\left[X_i\right]\mathclose{}\) for \(i = 1, \ldots, p\), so \(\operatorname{E}\tilde{X}= {(\mu_1, \ldots, \mu_p)}^{\top}\). By Definition 13:
\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[ \mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[ \begin{pmatrix}X_1 - \mu_1 \\ \vdots \\ X_p - \mu_p\end{pmatrix} \begin{pmatrix}X_1 - \mu_1 & \cdots & X_p - \mu_p\end{pmatrix} \right]\mathclose{} \\ &= \operatorname{E}\mathopen{}\left[ \begin{pmatrix} (X_1 - \mu_1)(X_1 - \mu_1) & \cdots & (X_1 - \mu_1)(X_p - \mu_p) \\ \vdots & \ddots & \vdots \\ (X_p - \mu_p)(X_1 - \mu_1) & \cdots & (X_p - \mu_p)(X_p - \mu_p) \end{pmatrix} \right]\mathclose{} \\ &= \begin{pmatrix} \operatorname{E}\mathopen{}\left[(X_1 - \mu_1)(X_1 - \mu_1)\right]\mathclose{} & \cdots & \operatorname{E}\mathopen{}\left[(X_1 - \mu_1)(X_p - \mu_p)\right]\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{E}\mathopen{}\left[(X_p - \mu_p)(X_1 - \mu_1)\right]\mathclose{} & \cdots & \operatorname{E}\mathopen{}\left[(X_p - \mu_p)(X_p - \mu_p)\right]\mathclose{} \end{pmatrix} \\ &= \begin{pmatrix} \operatorname{Cov}\mathopen{}\left(X_1, X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_1, X_p\right)\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{Cov}\mathopen{}\left(X_p, X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_p, X_p\right)\mathclose{} \end{pmatrix} \\ &= \begin{pmatrix} \operatorname{Var}\mathopen{}\left(X_1\right)\mathclose{} & \cdots & \operatorname{Cov}\mathopen{}\left(X_1, X_p\right)\mathclose{} \\ \vdots & \ddots & \vdots \\ \operatorname{Cov}\mathopen{}\left(X_p, X_1\right)\mathclose{} & \cdots & \operatorname{Var}\mathopen{}\left(X_p\right)\mathclose{} \end{pmatrix} \end{aligned} \]
where:
- the step from the third to fourth line uses the expectation of a random matrix,
- the step from the fourth to fifth line uses Definition 9, and
- the last step uses Lemma 1.
Proof. \[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[ \mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} && \text{(definition)} \\ &= \operatorname{E}\mathopen{}\left[ \tilde{X}{\tilde{X}}^{\top} - \tilde{X}{\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\tilde{X}}^{\top} + \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} \right]\mathclose{} && \text{(expand the product)} \\ &= \operatorname{E}\mathopen{}\left[\tilde{X}{\tilde{X}}^{\top}\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} + \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} && \text{(linearity, element-wise; } \operatorname{E}\tilde{X}\text{ is constant)} \\ &= \operatorname{E}\mathopen{}\left[\tilde{X}{\tilde{X}}^{\top}\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{} {\mathopen{}\left(\operatorname{E}\tilde{X}\right)\mathclose{}}^{\top} && \text{(combine like terms)} \end{aligned} \]
Proof. Treat \(\tilde{a}\) and \(\tilde{X}\) as \(n \times 1\) column vectors, so \(\tilde{a}\cdot \tilde{X}= {\tilde{a}}^{\top}\tilde{X}= \sum_{i=1}^na_i X_i\), a scalar. By linearity of expectation, \(\operatorname{E}\mathopen{}\left[{\tilde{a}}^{\top}\tilde{X}\right]\mathclose{} = {\tilde{a}}^{\top}\,\operatorname{E}\tilde{X}\), so:
\[ \begin{aligned} \operatorname{Var}\mathopen{}\left({\tilde{a}}^{\top}\tilde{X}\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[\mathopen{}\left({\tilde{a}}^{\top}\tilde{X}- {\tilde{a}}^{\top}\operatorname{E}\tilde{X}\right)\mathclose{}^2\right]\mathclose{} && \text{(definition of variance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left({\tilde{a}}^{\top}\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}\right)\mathclose{}^2\right]\mathclose{} && \text{(factor out } {\tilde{a}}^{\top} \text{)} \\ &= \operatorname{E}\mathopen{}\left[{\tilde{a}}^{\top}\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}{\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top}\tilde{a}\right]\mathclose{} && \text{(a scalar equals its transpose, so } s^2 = s\,{s}^{\top} \text{)} \\ &= {\tilde{a}}^{\top}\,\operatorname{E}\mathopen{}\left[\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}{\mathopen{}\left(\tilde{X}- \operatorname{E}\tilde{X}\right)\mathclose{}}^{\top}\right]\mathclose{}\,\tilde{a} && \text{(linearity of expectation, element-wise)} \\ &= {\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a} && \text{(variance of a random vector)} \\ &= \sum_{i=1}^n\sum_{j=1}^n a_i a_j \operatorname{Cov}\mathopen{}\left(X_i,X_j\right)\mathclose{} && \text{(expand the quadratic form, using the elements of } \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \text{)} \end{aligned} \]
Proof. Apply Theorem 9 with \(n=2\), \(X_1 = X\), and \(X_2 = Y\):
\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(aX+bY\right)\mathclose{} &= a^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + b^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2ab \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} \end{aligned} \]
Alternatively, by linearity of expectation:
\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(aX+bY\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\mathopen{}\left(aX+bY - \operatorname{E}\mathopen{}\left[aX+bY\right]\mathclose{}\right)\mathclose{}^2\right]\mathclose{} && \text{(definition of variance)} \\ &= \operatorname{E}\mathopen{}\left[\mathopen{}\left(a(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{}) + b(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right)\mathclose{}^2\right]\mathclose{} && \text{(linearity of expectation)} \\ &= \operatorname{E}\mathopen{}\left[a^2(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})^2 + 2(a \cdot b)(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{}) + b^2(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(expand the square)} \\ &= a^2\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})^2\right]\mathclose{} + 2(a \cdot b)\operatorname{E}\mathopen{}\left[(X-\operatorname{E}\mathopen{}\left[X\right]\mathclose{})(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})\right]\mathclose{} + b^2\operatorname{E}\mathopen{}\left[(Y-\operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(linearity of expectation)} \\ &= a^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + 2(a \cdot b) \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} + b^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(definitions of variance and covariance)} \end{aligned} \]
Proof. By Theorem 5, \(\operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0\). Applying Corollary 2 with \(a = b = 1\):
\[ \begin{aligned} \operatorname{Var}\mathopen{}\left(X + Y\right)\mathclose{} &= 1^2 \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + 1^2 \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2(1 \cdot 1) \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(variance of a sum of two random variables)} \\ &= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} + 2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(simplify)} \\ &= \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} + \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(} \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{} = 0 \text{)} \end{aligned} \]
Proof. By Theorem 7, the \((i,j)\)-th element of \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{}\), and by Definition 9, with \(\mu_i = \operatorname{E}\mathopen{}\left[X_i\right]\mathclose{}\):
\[ \begin{aligned} \operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{} &= \operatorname{E}\mathopen{}\left[(X_i - \mu_i)(X_j - \mu_j)\right]\mathclose{} && \text{(definition of covariance)} \\ &= \operatorname{E}\mathopen{}\left[(X_j - \mu_j)(X_i - \mu_i)\right]\mathclose{} && \text{(multiplication of numbers is commutative)} \\ &= \operatorname{Cov}\mathopen{}\left(X_j, X_i\right)\mathclose{} && \text{(definition of covariance)} \end{aligned} \]
so the \((i,j)\)-th and \((j,i)\)-th elements are equal, and \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is symmetric.
For positive semidefiniteness, let \(Y = {\tilde{a}}^{\top}\tilde{X}\). Then:
\[ \begin{aligned} {\tilde{a}}^{\top} \operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{} \tilde{a} &= \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} && \text{(variance of a linear combination)} \\ &= \operatorname{E}\mathopen{}\left[(Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2\right]\mathclose{} && \text{(definition of variance)} \\ &\ge 0 && \text{(} (Y - \operatorname{E}\mathopen{}\left[Y\right]\mathclose{})^2 \ge 0 \text{)} \end{aligned} \]
The first step is Theorem 9. The last step holds because a random variable that is never negative has a non-negative expectation: in the discrete and continuous cases of the definition of expectation, every term of the sum, or the integrand, is non-negative (for the general case, see Billingsley (1995)).
Proof. By Theorem 7, the \((i,j)\)-th element of \(\operatorname{Var}\mathopen{}\left(\tilde{X}\right)\mathclose{}\) is \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{}\). For \(i \neq j\), \(X_i\) and \(X_j\) are independent, so \(\operatorname{Cov}\mathopen{}\left(X_i, X_j\right)\mathclose{} = 0\) by Theorem 5. The \((i,i)\)-th element is \(\operatorname{Cov}\mathopen{}\left(X_i, X_i\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(X_i\right)\mathclose{}\) by Lemma 1.
Proof. Write \(\sigma_X \stackrel{\text{def}}{=}\operatorname{SD}\mathopen{}\left(X\right)\mathclose{} > 0\) and \(\sigma_Y \stackrel{\text{def}}{=}\operatorname{SD}\mathopen{}\left(Y\right)\mathclose{} > 0\), and take either sign \(\pm\) throughout. A variance is the expectation of a squared deviation, a non-negative random variable, so it is non-negative. Applying Corollary 2 with \(a = 1/\sigma_X\) and \(b = \pm 1/\sigma_Y\):
\[ \begin{aligned} 0 &\le \operatorname{Var}\mathopen{}\left(\frac{X}{\sigma_X} \pm \frac{Y}{\sigma_Y}\right)\mathclose{} && \text{(a variance is non-negative)} \\ &= \frac{\operatorname{Var}\mathopen{}\left(X\right)\mathclose{}}{\sigma_X^2} + \frac{\operatorname{Var}\mathopen{}\left(Y\right)\mathclose{}}{\sigma_Y^2} \pm \frac{2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\sigma_X \sigma_Y} && \text{(variance of a sum of two random variables)} \\ &= 1 + 1 \pm \frac{2 \operatorname{Cov}\mathopen{}\left(X,Y\right)\mathclose{}}{\sigma_X \sigma_Y} && \text{(definition of standard deviation: } \sigma_X^2 = \operatorname{Var}\mathopen{}\left(X\right)\mathclose{} \text{, } \sigma_Y^2 = \operatorname{Var}\mathopen{}\left(Y\right)\mathclose{} \text{)} \\ &= 2 \pm 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} && \text{(definition of correlation)} \end{aligned} \]
With the \(+\) sign, \(0 \le 2 + 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{}\) gives \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \ge -1\). With the \(-\) sign, \(0 \le 2 - 2 \operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{}\) gives \(\operatorname{Cor}\mathopen{}\left(X,Y\right)\mathclose{} \le 1\).