Proof. Discrete case. When \(X\) and \(Y\) are discrete, applying Definition 1 to \(\operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\right]\mathclose{}\) and then the law of total probability applied to the countable partition \(\{X = x : x \in \mathcal{R}(X)\}\):
\[
\begin{aligned}
\operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\right]\mathclose{}
&= \sum_{x \in \mathcal{R}(X)} \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \cdot\operatorname{P}(X=x) && \text{(LOTUS for the function } x \mapsto \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \text{)}
\\&= \sum_{x \in \mathcal{R}(X)} \mathopen{}\left(\sum_{y \in \mathcal{R}(Y)} y \cdot\operatorname{P}(Y=y \mid X=x)\right)\mathclose{} \cdot\operatorname{P}(X=x) && \text{(definition of conditional expectation } \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \text{)}
\\&= \sum_{y \in \mathcal{R}(Y)} y \cdot\sum_{x \in \mathcal{R}(X)} \operatorname{P}(Y=y \mid X=x) \cdot\operatorname{P}(X=x) && \text{(exchange order of summation: Fubini, since } \operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} < \infty \text{)}
\\&= \sum_{y \in \mathcal{R}(Y)} y \cdot\operatorname{P}(Y=y) && \text{(law of total probability over } X \text{)}
\\&= \operatorname{E}\mathopen{}\left[Y\right]\mathclose{} && \text{(definition of expectation of } Y \text{)}
\end{aligned}
\]
Continuous case. When \(X\) and \(Y\) are continuous:
\[
\begin{aligned}
\operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\right]\mathclose{}
&= \int_{x \in \mathcal{R}(X)} \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \cdot\operatorname{p}(X=x)\, dx
&& \text{(LOTUS)}
\\&= \int_{x \in \mathcal{R}(X)} \mathopen{}\left(\int_{y \in \mathcal{R}(Y)} y \cdot\operatorname{p}(Y=y \mid X=x)\, dy\right)\mathclose{} \cdot\operatorname{p}(X=x)\, dx
&& \text{(definition of conditional expectation)}
\\&= \int_{x \in \mathcal{R}(X)} \int_{y \in \mathcal{R}(Y)} y \cdot\operatorname{p}(X=x, Y=y)\, dy\, dx
&& \text{(definition of conditional density: } \operatorname{p}(Y=y \mid X=x)\,\operatorname{p}(X=x) = \operatorname{p}(X=x, Y=y) \text{)}
\\&= \int_{y \in \mathcal{R}(Y)} y \cdot\mathopen{}\left(\int_{x \in \mathcal{R}(X)} \operatorname{p}(X=x, Y=y)\, dx\right)\mathclose{}\, dy
&& \text{(Fubini's theorem, since } \operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} < \infty \text{)}
\\&= \int_{y \in \mathcal{R}(Y)} y \cdot\operatorname{p}(Y=y)\, dy
&& \text{(marginal density from a joint density)}
\\&= \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}
&& \text{(definition of expectation)}
\end{aligned}
\]
Mixed case, \(X\) discrete and \(Y\) continuous. With the joint density-mass function \(\operatorname{p}(X = x,\, Y = y)\), Corollary 3 applies with \(h(x, y) = y\), counting measure for \(X\), and Lebesgue measure for \(Y\); its condition (b) holds because \(\operatorname{E}\mathopen{}\left[\mathopen{}\left|Y\right|\mathclose{}\right]\mathclose{} < \infty\). The sum runs over the \(x\) with \(\operatorname{P}(X = x) > 0\), where Definition 6 defines \(\operatorname{E}\mathopen{}\left[Y \mid X = x\right]\mathclose{}\):
\[
\begin{aligned}
\operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\right]\mathclose{}
&= \sum_{x} \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \cdot\operatorname{P}(X=x)
&& \text{(LOTUS, discrete case)}
\\&= \sum_{x} \mathopen{}\left(\int_{y} y \cdot\operatorname{p}(Y=y \mid X=x)\,dy\right)\mathclose{} \cdot\operatorname{P}(X=x)
&& \text{(definition of conditional expectation, mixed case)}
\\&= \sum_{x} \int_{y} y \cdot\operatorname{p}(Y=y \mid X=x) \cdot\operatorname{P}(X=x)\,dy
&& \text{(move the constant } \operatorname{P}(X=x) \text{ inside the integral)}
\\&= \sum_{x} \int_{y} y \cdot\operatorname{p}(X=x,\, Y=y)\,dy
&& \text{(definition of the conditional density, mixed case)}
\\&= \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}
&& \text{(joint-distribution form of Fubini--Tonelli, with } h(x, y) = y \text{)}
\end{aligned}
\]
Mixed case, \(X\) continuous and \(Y\) discrete. The same steps apply with the integral and the sum swapped, over the \(x\) with \(\operatorname{p}(X = x) > 0\); at any other \(x\), \(\operatorname{p}(X = x,\, Y = y) = 0\) for every \(y\), because these non-negative terms sum to \(\operatorname{p}(X = x)\):
\[
\begin{aligned}
\operatorname{E}\mathopen{}\left[\operatorname{E}\mathopen{}\left[Y \mid X\right]\mathclose{}\right]\mathclose{}
&= \int_{x} \operatorname{E}\mathopen{}\left[Y \mid X=x\right]\mathclose{} \cdot\operatorname{p}(X=x)\,dx
&& \text{(LOTUS, continuous case)}
\\&= \int_{x} \mathopen{}\left(\sum_{y} y \cdot\operatorname{P}(Y=y \mid X=x)\right)\mathclose{} \cdot\operatorname{p}(X=x)\,dx
&& \text{(definition of conditional expectation, mixed case)}
\\&= \int_{x} \sum_{y} y \cdot\operatorname{P}(Y=y \mid X=x) \cdot\operatorname{p}(X=x)\,dx
&& \text{(move the constant } \operatorname{p}(X=x) \text{ inside the sum)}
\\&= \int_{x} \sum_{y} y \cdot\operatorname{p}(X=x,\, Y=y)\,dx
&& \text{(definition of the conditional PMF, mixed case)}
\\&= \operatorname{E}\mathopen{}\left[Y\right]\mathclose{}
&& \text{(joint-distribution form of Fubini--Tonelli, with } h(x, y) = y \text{)}
\end{aligned}
\]