Estimation
1 Scientific models
1.1 Models as approximations
…Essentially, all models are wrong, but some are useful. However, the approximate nature of the model must always be borne in mind.
[Box and Draper (1987), p. 424; emphasis added]
See also (Dunn and Smyth 2018, sec. 1.8).
1.2 Statistical analysis of scientific models
When we perform statistical analyses, we use data to help us choose between models; specifically, to determine which models best explain those data.
Physical processes do not produce data on their own. Data are only produced when scientists implement an observation process (that is, a scientific study), which is distinct from the underlying physical process. In some cases, the observation process and the physical process interact with each other; this interaction is called the “observer effect”.
To learn about the physical processes we are ultimately interested in, we often need to account for the observation process that produced the data we are analyzing. In particular, if some of the planned observations in the study design were not completed, we will likely need to account for the incompleteness of the resulting data set in our analysis. If we are not sure why some observations are incomplete, we may need to model the observation process in addition to the physical process we were originally interested in. For example, if some participants in a study dropped out part-way through, we may need to investigate why those participants dropped out, as opposed to other participants who completed the study.
These kinds of missing data issues are outside the scope of these notes; see Van Buuren (2018) for more details.
2 Estimands, estimates, and estimators
2.1 Estimands
In statistical contexts, most estimands are parameters of probabilistic models, or functions of model parameters.
Model parameters and other estimands are often symbolized using lower-case Greek letters: \(\alpha, \beta, \gamma, \delta\), etc.
2.2 Estimates
2.3 Estimators
When an estimator is applied to random variables \(X_1, \ldots, X_n\) rather than to their observed values, the result \(\hat\theta(X_1, \ldots, X_n)\) is also a random variable.
Estimators are often symbolized by placing a ^ (“hat”) symbol on top of the corresponding estimand; for example, \(\hat\theta\).
Usually, their dependence on the data is implicit:
\[\hat\theta\stackrel{\text{def}}{=}\hat\theta(x_1, \ldots, x_n)\]
2.4 Contrasting estimands, estimates, and estimators
It’s helpful to keep in mind the mathematical type of each estimation concept:
- estimands are numbers (or vectors of numbers);
- estimates are also numbers (or vectors);
- estimators are functions; an estimator applied to random data, \(\hat\theta(X_1, \ldots, X_n)\), is a random variable.
3 Accuracy of estimators
3.1 Accuracy
To determine which estimator is best, we need to define best. Accuracy is usually most important, and ease of computation is usually secondary.
3.2 Estimation error
The accuracy of an estimator has no single, agreed formal definition. The usual measures of accuracy, including the bias, mean squared error, and mean absolute error defined in this section, are all functions of the distribution of the estimator’s estimation error.
3.3 Residuals
See Linear-model residual definitions and terminology for residual definitions and for the relationship between residuals, model deviations, and estimation error.
3.4 Bias
Proof. \[ \begin{aligned} \operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\varepsilon\mathopen{}\left(\hat\theta\right)\mathclose{}\right]\mathclose{} && \text{(definition of bias)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta- \theta\right]\mathclose{} && \text{(definition of estimation error)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{} - \operatorname{E}\mathopen{}\left[\theta\right]\mathclose{} && \text{(linearity of expectation)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{} - \theta && \text{($\theta$ is a constant)} \end{aligned} \]
3.5 Mean squared error
Proof. Let’s start by expanding each term of the right-hand side. By Theorem 1, the squared bias is:
\[ \begin{aligned} \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} &= \mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{} - \theta\right)^2\mathclose{} && \text{(bias equals expectation minus truth)}\\ &= \mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\right)^2\mathclose{} - 2\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\theta+ \theta^2 && \text{(expand the binomial square)} \end{aligned} \]
The variance is (simplified expression for variance):
\[\operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} = \operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\right)^2\mathclose{}\]
Now, add them together and simplify:
\[ \begin{aligned} \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} &= \mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\right)^2\mathclose{} - 2\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\theta+ \theta^2 + \operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - \mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\right)^2\mathclose{} && \text{(substitute both expansions)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - 2\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\theta+ \theta^2 && \text{(cancel $\mathopen{}\left(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\right)^2\mathclose{}$)} \end{aligned} \]
Now let’s expand the left-hand side to reach the same expression:
\[ \begin{aligned} \operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} &\stackrel{\text{def}}{=}\operatorname{E}\mathopen{}\left[\mathopen{}\left(\varepsilon\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{}\right]\mathclose{} && \text{(definition of MSE)}\\ &= \operatorname{E}\mathopen{}\left[(\hat\theta- \theta)^2\right]\mathclose{} && \text{(definition of estimation error)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta^2 - 2\hat\theta\theta+ \theta^2\right]\mathclose{} && \text{(expand the binomial square)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - \operatorname{E}\mathopen{}\left[2\hat\theta\theta\right]\mathclose{} + \operatorname{E}\mathopen{}\left[\theta^2\right]\mathclose{} && \text{(linearity of expectation)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - 2\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\theta+ \theta^2 && \text{($\theta$ is a constant)} \end{aligned} \]
\(\operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{}\) and \(\mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{}\) both equal \(\operatorname{E}\mathopen{}\left[\hat\theta^2\right]\mathclose{} - 2\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{}\theta+ \theta^2\). Equality is transitive, so \(\operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{}\) and \(\mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{}\) are equal to each other:
\[\operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} = \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{}\]
3.6 Unbiased estimators
Proof. For Equation 3, apply Theorem 1:
\[ \begin{aligned} 0 &= \operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{(definition of unbiased)}\\ &= \operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{} - \theta && \text{(bias equals expectation minus truth)} \end{aligned} \]
Adding \(\theta\) to both sides gives \(\operatorname{E}\mathopen{}\left[\hat\theta\right]\mathclose{} = \theta\).
For Equation 4, apply Theorem 2:
\[ \begin{aligned} \operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} &= \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{(MSE equals bias squared plus variance)}\\ &= 0^2 + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{(definition of unbiased)}\\ &= \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{($0^2 = 0$)} \end{aligned} \]
3.7 Mean absolute error
3.8 Standard error
Proof. \[ \begin{aligned} \operatorname{Var}\mathopen{}\left(\varepsilon\mathopen{}\left(\hat\theta\right)\mathclose{}\right)\mathclose{} &= \operatorname{Var}\mathopen{}\left(\hat\theta- \theta\right)\mathclose{} && \text{(definition of estimation error)}\\ &= \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{(subtracting a constant does not change a variance)} \end{aligned} \]
Taking square roots of both sides, \(\operatorname{SD}\mathopen{}\left(\varepsilon\mathopen{}\left(\hat\theta\right)\mathclose{}\right)\mathclose{} = \operatorname{SD}\mathopen{}\left(\hat\theta\right)\mathclose{} = \operatorname{SE}\mathopen{}\left(\hat\theta\right)\mathclose{}\).
“Standard error” is a confusing name in two ways. It is defined through the estimator’s own spread, not through the estimation error (although Theorem 4 shows that the two spreads are equal). It is also a synonym for the standard deviation of an estimator, so it can look redundant. The name persists because standard errors are the building blocks of p-values and confidence intervals, so they come up often enough to deserve their own name.
Proof. Start from Theorem 2:
\[ \begin{aligned} \operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} &= \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} + \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} && \text{(MSE equals bias squared plus variance)}\\ \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} &= \operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} - \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} && \text{(subtract $\mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{}$ from both sides)}\\ \mathopen{}\left(\operatorname{SE}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} &= \operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} - \mathopen{}\left(\operatorname{Bias}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} && \text{($\operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{} = \mathopen{}\left(\operatorname{SD}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{} = \mathopen{}\left(\operatorname{SE}\mathopen{}\left(\hat\theta\right)\mathclose{}\right)^2\mathclose{}$)} \end{aligned} \]
Proof. By Equation 4, \(\operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{} = \operatorname{Var}\mathopen{}\left(\hat\theta\right)\mathclose{}\). Taking square roots of both sides, \(\sqrt{\operatorname{MSE}\mathopen{}\left(\hat\theta\right)\mathclose{}} = \operatorname{SD}\mathopen{}\left(\hat\theta\right)\mathclose{} = \operatorname{SE}\mathopen{}\left(\hat\theta\right)\mathclose{}\).