Posterior inference
Intuition
We have samples from \(m\) groups \(\{\boldsymbol{y}_1, \dots, \boldsymbol{y}_m\}\). Our task is to sample from the posterior distribution
\[
p(\theta_1, \dots, \theta_m, \mu, \tau^2, \sigma^2 \mid \boldsymbol{y}_1, \dots, \boldsymbol{y}_m)
\]
for which we’ll use a Gibbs sampler. We need to find the full conditionals for all unknown quantities above, which seems intimidating. In the Chapter 6 notes In the Chapter 6 notes (“a shortcut for thinking about full conditionals”), it should be mentioned that obtaining the full conditional for a single parameter is fairly straightforward by simply writing the entire joint posterior but then treating the other parameters as constants that can be discarded via proportionality.
To obtain the full joint posterior we will take use key independence assumptions between the parameters of our model. For example, given a group-specific mean \(\theta_j\), the corresponding random variables \(Y_{i, j}\) depend only on \((\theta_j, \sigma^2)\) and not on \(\mu\) or \(\tau^2\) (this is implied in Figure 8.3).
\[\begin{align}
& p(\theta_1, \dots, \theta_m, \mu, \tau^2, \sigma^2 \mid \boldsymbol{y}_1, \dots, \boldsymbol{y}_m) \\
\quad&\propto p(\mu, \tau^2, \sigma^2, \mu, \tau^2, \sigma^2) \times p(\boldsymbol{y}_1, \dots, \boldsymbol{y}_m \mid \theta_1, \dots, \theta_m, \mu, \tau^2, \sigma^2) & \text{Bayes' rule} \\
&= p(\mu, \tau^2, \sigma^2) \times p(\theta_1, \dots, \theta_m \mid \mu, \tau^2, \sigma^2) \times p(\boldsymbol{y}_1, \dots, \boldsymbol{y}_m \mid \theta_1, \dots, \theta_m, \mu, \tau^2, \sigma^2) & \text{Chain rule} \\
&= p(\mu) p(\tau^2) p(\sigma^2) \times p(\theta_1, \dots, \theta_m \mid \mu, \tau^2) \times p(\boldsymbol{y}_1, \dots, \boldsymbol{y}_m \mid \theta_1, \dots, \theta_m, \sigma^2) & \text{Indep.} \\
&= p(\mu) p(\tau^2) p(\sigma^2) \times \left[ \prod_{j = 1}^m p(\theta_j \mid \mu, \tau^2) \right] \times \left[ \prod_{j = 1}^m p(\boldsymbol{y}_j \mid \theta_j, \sigma^2) \right] & \text{de Finetti} \\
&= p(\mu) p(\tau^2) p(\sigma^2) \times \left[ \prod_{j = 1}^m p(\theta_j \mid \mu, \tau^2) \right] \times \left[ \prod_{j = 1}^m \left( \prod_{i = 1}^{n_j} p(y_{i, j} \mid \theta_j, \sigma^2) \right) \right] & \text{de Finetti 2x} \\
\end{align}\]
Now to evaluate a full conditional, for example that for \(\mu\), we take the full posterior and discard all terms that don’t depend on \(\mu\):
\[p(\mu \mid \theta_1, \dots, \theta_m, \tau^2, \sigma^2, \boldsymbol{y}_1, \dots, \boldsymbol{y}_m) \propto p(\mu) \prod_{j=1}^m p(\theta_j \mid \mu, \tau^2)\]
which in this case looks exactly like a standard one-sample Normal posterior from Chapter 6, so we borrow that result and replace the relevant variables from our priors. We can do this similarly for the other parameters.
Quantities
Since the work is quite tedious, the derivations for the full conditionals of the parameters. To summarize, given priors
\[\begin{align}
\sigma^2 &\sim \text{inverse-gamma}(\nu_0 / 2, \sigma_0^2 \nu_0 / 2) \\
\tau^2 &\sim \text{inverse-gamma}(\eta_0 / 2, \tau_0^2 \eta_0 / 2) \\
\mu^2 &\sim \mathcal{N}(\mu_0, \gamma_0^2)
\end{align}\]
the full conditionals are
\[\begin{align}
\{\mu \mid \theta_1, \dots, \theta_m, \tau^2\} &\sim \mathcal{N}\left(\frac{m\bar{\theta} / \tau^2 + \mu_0 / \gamma_0^2}{m/\tau^2 + 1 / \gamma_0^2}, \left[m / \tau^2 + 1 / \gamma_0^2 \right]^{-1}\right) \\
\{\tau^2 \mid \theta_1, \dots, \theta_m, \mu\} &\sim \text{inverse-gamma}\left(\frac{\eta_0 + m}{2}, \frac{\eta_0\tau_0^2 + \sum_{j = 1}^{m}(\theta_j - \mu)^2}{2} \right) \\
\{\theta_j \mid y_{1, j}, \dots, y_{n_{j}, j}, \sigma^2\} &\sim \mathcal{N}\left(\frac{n_j\bar{y}_j / \sigma^2 + 1 / \tau^2}{n_j / \sigma^2 + 1 / \tau^2}, \left[ n_j / \sigma^2 + 1 / \tau^2 \right]^{-1} \right) \\
\{\sigma^2 \mid \boldsymbol{\theta}, \boldsymbol{y_1}, \dots, \boldsymbol{y_n}\} &\sim \text{inverse-gamma}\left(\frac{1}{2}\left[ \nu_0 + \sum_{j = 1}^m n_j\right], \frac{1}{2}\left[\nu_0\sigma_0^2 + \sum_{j = 1}^{m} \left( \sum_{i = 1}^{n_j} (y_{i, j} - \theta)^2 \right) \right] \right)
\end{align}\]
It’s worth briefly discussing what these values represent. The full conditionals for \(\mu\) and \(\tau^2\) look like standard normal posteriors. Similarly, the full conditional for \(\theta_j\) looks like a normal posterior dependent only on the specific subgroup \(\boldsymbol{y}_j\) and the common variance \(\sigma^2\). Lastly (and most interestingly), notice that the posterior for \(\sigma^2\) looks like a standard inverse gamma posterior which depends on \(\sum \sum (y_{i, j} -
\theta)^2\) which is the pooled variance across all groups (see also \(\sum n_j\)).