# one unbroken string, so that link checkers test the whole URL:
url <- "https://regression.ucsf.edu/sites/g/files/tkssra16191/files/wysiwyg/home/data/hersdata.dta" # nolint: line_length_linter.
hers <- haven::read_dta(url)Comparing Means
1 Introduction
This page reviews the standard methods for comparing the mean of a continuous outcome between groups: t-tests, confidence intervals for a difference in means, and one-way analysis of variance.
This page builds on three others:
- Exploratory Data Analysis defines the sample statistics these methods use;
- Statistical Inference defines hypotheses, test statistics, p-values, confidence intervals, and the t, chi-square, and F reference distributions;
- Estimation defines estimators and their standard errors.
The t-tests and ANOVA on this page are the exact, Gaussian-outcome tests in the table of exact and approximate tests on the Maximum Likelihood page, which pairs each one with its large-sample counterpart.
This page is adapted from Vittinghoff et al. (2012), Chapter 3.
2 The HERS data
The “heart and estrogen/progestin study” (HERS) was a clinical trial of hormone therapy for prevention of recurrent heart attacks and death among 2,763 post-menopausal women with existing coronary heart disease (CHD) (Hulley et al. 1998).
The trial was conducted at 20 US clinical centers. Participants were randomized to receive either conjugated equine estrogens (0.625 mg/day) plus medroxyprogesterone acetate (2.5 mg/day) or a matching placebo (Hulley et al. 1998). Women were followed for an average of 4.1 years (Hulley et al. 1998).
The primary outcome was nonfatal myocardial infarction or CHD death (Hulley et al. 1998).
The HERS data are distributed with Vittinghoff et al. (2012) on the book’s companion website, as a Stata file that R can read directly:
The rmb R package includes the same file, which these notes use so that rendering does not depend on the website. haven::as_factor() converts the Stata value labels to factors:
The examples on this page use these variables:
| Variable | Meaning |
|---|---|
HT |
Randomized treatment: placebo or hormone therapy |
age |
Age at baseline (years) |
raceth |
Race/ethnicity: White, African American, or Other |
exercise |
Exercises at least three times per week (no/yes) |
BMI |
Body mass index at baseline (kg/m2) |
SBP |
Systolic blood pressure at baseline (mmHg) |
glucose |
Fasting glucose at baseline (mg/dL) |
glucose1 |
Fasting glucose at the year-1 visit (mg/dL) |
hers |>
dplyr::select(HT, age, raceth, exercise, BMI, SBP, glucose, glucose1) |>
dplyr::glimpse()
#> Rows: 2,763
#> Columns: 8
#> $ HT <fct> placebo, placebo, hormone therapy, placebo, placebo, hormone …
#> $ age <dbl> 70, 62, 69, 64, 65, 68, 70, 69, 61, 62, 72, 73, 52, 57, 57, 6…
#> $ raceth <fct> African American, African American, White, White, White, Afri…
#> $ exercise <fct> no, no, no, no, no, no, no, yes, yes, no, no, no, no, no, yes…
#> $ BMI <dbl> 23.69, 28.62, 42.51, 24.39, 21.90, 29.05, 34.45, 23.16, 30.26…
#> $ SBP <dbl> 138, 118, 134, 152, 175, 174, 119, 178, 162, 111, 122, 158, 1…
#> $ glucose <dbl> 84, 111, 114, 94, 101, 116, 120, 95, 105, 98, 111, 95, 97, 10…
#> $ glucose1 <dbl> 94, 78, 98, 93, 92, 115, NA, 95, 113, 98, 96, 107, 90, 100, 1…3 Descriptive statistics
The Exploratory Data Analysis page defines the sample statistics used on this page, each with an example:
- the sample mean;
- the sample median;
- the sample variance;
- the sample standard deviation;
- the interquartile range;
- the sample proportion.
That page also defines the graphs used here, including box plots and scatter plots.
4 Comparing two groups: continuous outcomes
4.1 Hypotheses
The Statistical Inference page defines the null hypothesis, the alternative hypothesis, test statistics, p-values, and significance levels. In a comparison of two groups with means \(\mu_1\) and \(\mu_2\), the null hypothesis is usually \(H_0: \mu_1 = \mu_2\), and the two-sided alternative is \(H_1: \mu_1 \neq \mu_2\).
4.2 One-sample t-test
Proof. We use a standard fact about Gaussian samples (Hogg et al. 2019, sec. 5.5, pp. 203-205): \(\bar X\) and \(S^2\) are independent, and \(V \stackrel{\text{def}}{=}(n-1) S^2 / \sigma^2\) has the \(\chi^2_{n-1}\) distribution. Also, \(Z \stackrel{\text{def}}{=}(\bar X - \mu_0) / (\sigma / \sqrt{n})\) has the standard Gaussian distribution, and \(Z\) is independent of \(V\) because \(\bar X\) is independent of \(S^2\). So by the definition of the t-distribution, \(Z / \sqrt{V / (n-1)}\) has the \(t_{n-1}\) distribution, and this ratio equals \(T\):
\[ \begin{aligned} \frac{Z}{\sqrt{V/(n-1)}} &= \frac{(\bar X - \mu_0) / (\sigma / \sqrt{n})}{\sqrt{S^2 / \sigma^2}} && \text{(substitute $Z$ and $V$)}\\ &= \frac{(\bar X - \mu_0) / (\sigma / \sqrt{n})}{S / \sigma} && \text{($\sqrt{S^2} = S$, since $S \ge 0$)}\\ &= \frac{\bar X - \mu_0}{S / \sqrt{n}} && \text{(the factors of $\sigma$ cancel)}\\ &= T. \end{aligned} \]
When the observations are not Gaussian, \(T\) still has approximately the standard Gaussian distribution when \(n\) is large, by the central limit theorem, and then \(t_{n-1}\) is close to the standard Gaussian distribution (t quantiles approach Gaussian quantiles).
4.3 Paired t-test
4.4 Two-sample t-tests
Welch’s test does not assume that the two groups have equal variances. Even for Gaussian data, \(t_{\hat\nu}\) is only an approximation to the null distribution of \(t\). For large groups, the central limit theorem makes the statistic approximately standard Gaussian under \(H_0\), as in the WCGS cholesterol example. Welch’s test is the default in R’s t.test().
Theorem 2 needs equal variances in the two groups, and Welch’s test (Definition 3) does not, so these notes use Welch’s test by default.
4.5 Confidence intervals for the difference in means
5 One-way analysis of variance
Large values of \(F\) mean that the group means are spread out more than the variation within groups would explain. For the groups of Example 9, \(F = 13.5 / 1 = 13.5\).
The group standard deviations differ, from 36 to 44 mg/dL, so the equal-variance condition of Theorem 3 is questionable. oneway.test() performs Welch’s version of the F-test, which does not assume equal variances, and reaches the same conclusion:
oneway.test(glucose ~ raceth, data = hers)
#>
#> One-way analysis of means (not assuming equal variances)
#>
#> data: glucose and raceth
#> F = 12.49, num df = 2.0, denom df = 185.7, p-value = 8.17e-06One-way ANOVA is a special case of linear regression: it is the F-test comparing a linear regression model with a single categorical predictor to the model with an intercept only (Linear Models Overview). lm() gives the same F statistic as aov():
With \(k = 2\) groups, the ANOVA F statistic equals the square of the pooled t statistic (Definition 5), as the two treatment groups of Example 5 show:
