This page reviews tests for comparing groups on a categorical outcome: the chi-square test and Fisher’s exact test for contingency tables. It uses the chi-square reference distribution defined on the Statistical Inference page. This page is adapted from Vittinghoff et al. (2012), Chapter 3.
1 The HERS data
The examples on this page use the HERS data, which the Comparing Means page describes. The rmb R package includes the dataset; haven::as_factor() converts its Stata value labels to factors:
Example 1 (Exercise by treatment group in HERS)Table 1 cross-tabulates regular exercise at baseline by treatment group, as a contingency table with row percentages.
Definition 1 (Expected count under independence) Let a contingency table have \(r\) rows and \(c\) columns, with observed count \(O_{ij}\) in row \(i\) and column \(j\), row totals \(R_i\), column totals \(C_j\), and grand total \(n\). The expected count in cell \((i, j)\) under independence is
Example 2 (Expected count in a 2 x 2 table) In a table with \(n = 100\), first-row total \(R_1 = 30\), and first-column total \(C_1 = 40\), the expected count in cell \((1, 1)\) is \(E_{11} = 30 \cdot 40 / 100 = 12\).
Definition 2 (Pearson’s chi-square test of independence) With observed counts \(O_{ij}\) and expected counts\(E_{ij}\) in an \(r \times c\) contingency table, Pearson’s chi-square test of the null hypothesis that the row and column variables are independent uses the statistic
Its p-value is \(\Pr(W \ge X^2)\), where \(W\) has the \(\chi^2_{(r-1)(c-1)}\) distribution (chi-square distribution).
Theorem 1 (Large-sample null distribution of the chi-square statistic) Let \(n\) observations be sampled independently and classified by two categorical variables, and let the two variables be independent. Then as \(n \to \infty\), the distribution of \(X^2\) (Definition 2) converges to the \(\chi^2_{(r-1)(c-1)}\) distribution (Hogg et al. 2019, sec. 9.2, p. 440).
The chi-square approximation is poor when some expected counts are small. A common rule of thumb asks for every \(E_{ij}\) to be at least 5.
Example 3 (Chi-square test of exercise by treatment group in HERS) The expected counts and statistic of Definition 2 for Table 1:
Every expected count is large, so the \(\chi^2_1\) approximation of Theorem 1 is reasonable. chisq.test() with correct = FALSE reports the same values:
The p-value is large: the data give no evidence that exercise depends on treatment group, as randomization would lead us to expect.
For a \(2 \times 2\) table, chisq.test() applies Yates’ continuity correction by default, which subtracts 0.5 from each \(\mathopen{}\left|O_{ij} - E_{ij}\right|\mathclose{}\) before squaring, and so gives a smaller statistic than Definition 2. correct = FALSE turns the correction off.
2.3 Fisher’s exact test
Definition 3 (Fisher’s exact test) Take a \(2 \times 2\) contingency table with cells \(a\), \(b\), \(c\), and \(d\) as in the contingency table definition, and hold its row and column totals fixed. Under independence of the row and column variables, the probability that the top-left cell equals \(x\) is the hypergeometric probability
the total probability of the tables with the same totals that are no more probable than the observed table.
The p-value is exact: it comes from the null distribution itself, not from a large-sample approximation, so the test is valid even when expected counts are small, and it is often used for \(2 \times 2\) tables in which some expected count is below 5 (Theorem 1). Other two-sided versions exist; this one is the version that R’s fisher.test() computes.
Example 4 (Fisher’s exact test of exercise by treatment group in HERS) The p-value of Definition 3 for Table 1, computed from the hypergeometric probabilities with dhyper():
a<-observed[1, 1]row1<-sum(observed[1, ])row2<-sum(observed[2, ])col1<-sum(observed[, 1])x<-max(0, col1-row2):min(row1, col1)p_x<-dhyper(x, m =row1, n =row2, k =col1)p_a<-dhyper(a, m =row1, n =row2, k =col1)sum(p_x[p_x<=p_a*(1+1e-7)])#> [1] 0.72528
The tolerance 1e-7 keeps tables whose probability equals \(p(a)\) up to rounding error, as fisher.test() does. fisher.test() reports the same p-value:
fisher.test(hers$exercise, hers$HT)#> #> Fisher's Exact Test for Count Data#> #> data: hers$exercise and hers$HT#> p-value = 0.725#> alternative hypothesis: true odds ratio is not equal to 1#> 95 percent confidence interval:#> 0.879664 1.202192#> sample estimates:#> odds ratio #> 1.02836
With counts this large, the exact p-value is close to the chi-square p-value of Example 3.
2.4 Measures of association for \(2 \times 2\) tables
Tests of independence say whether two binary variables are associated, but not how strongly. Risk differences, risk ratios, and odds ratios measure the strength of the association; see Odds Ratios and Relative Risks.
References
Hogg, Robert V., Elliot A. Tanis, and Dale L. Zimmerman. 2019. Probability and Statistical Inference. Tenth edition. Pearson.
Vittinghoff, Eric, David V Glidden, Stephen C Shiboski, and Charles E McCulloch. 2012. Regression Methods in Biostatistics: Linear, Logistic, Survival, and Repeated Measures Models. 2nd ed. Springer. https://doi.org/10.1007/978-1-4614-1353-0.
---title: "Comparing Proportions"format: html: default revealjs: output-file: categorical-tests-slides.html pdf: output-file: categorical-tests-handout.pdf docx: output-file: categorical-tests-handout.docx---{{< include sds/_subfiles/shared-config.qmd >}}{{< include sds/_subfiles/basic-statistical-methods/_prose-intro-categorical.qmd >}}## The HERS data{{< include sds/_subfiles/basic-statistical-methods/_load-hers-short.qmd >}}## Comparing two groups: categorical outcomes {#sec-two-group-categorical}### Contingency tables {#sec-contingency-tables}<!-- No slidebreak here: the heading shares its slide with the exm-hers-crosstab div. -->{{< include sds/_subfiles/basic-statistical-methods/_exm-hers-crosstab.qmd >}}{{< slidebreak >}}### The chi-square test {#sec-chi-square}<!-- No slidebreak here: the heading shares its slide with the def-expected-count div. -->{{< include sds/_subfiles/basic-statistical-methods/_def-expected-count.qmd >}}{{< slidebreak >}}{{< include sds/_subfiles/basic-statistical-methods/_exm-expected-count.qmd >}}{{< slidebreak >}}{{< include sds/_subfiles/basic-statistical-methods/_def-chi-square-test.qmd >}}{{< slidebreak >}}{{< include sds/_subfiles/basic-statistical-methods/_thm-chi-square-null.qmd >}}{{< slidebreak >}}{{< include sds/_subfiles/basic-statistical-methods/_exm-hers-chisq.qmd >}}{{< slidebreak >}}### Fisher's exact test {#sec-fisher}<!-- No slidebreak here: the heading shares its slide with the def-fishers-exact div. -->{{< include sds/_subfiles/basic-statistical-methods/_def-fishers-exact.qmd >}}{{< slidebreak >}}{{< include sds/_subfiles/basic-statistical-methods/_exm-hers-fisher.qmd >}}{{< slidebreak >}}### Measures of association for $2 \times 2$ tables {#sec-2x2-measures}{{< include sds/_subfiles/basic-statistical-methods/_prose-2x2-measures.qmd >}}## References {.unnumbered}::: {#refs}:::