Last modified: 2026-09-29 00:05:17 (PDT)
These notes collect the statistics that data science courses assume.
The notes are organized in four groups of pages.
Describing data:
Inference:
Methods, illustrated with the HERS data:
Bayesian:
These notes began as the statistics appendices of Regression Models for Epidemiology, and keep those chapters’ file names, except that two long chapters are now split into several pages. Most results keep their rme #id anchors on the same page; the exceptions are:
exr-prac-*), which moved from the estimation page to the maximum likelihood page;def-sample-mean), so the exploratory-descriptive ids from rme def-mean, def-median, def-eda-variance, def-eda-sd, def-quantile, def-iqr, and def-correlation are retired;sec-two-group-categorical, sec-correlation, sec-simple-linear-regression, and sec-bootstrap-ci), which moved to their own pages, with their ids;sec-foundations through sec-dic) and the worked examples (sec-bayes-examples and the sections after it), which moved to the MCMC and JAGS pages, with their ids.Where these notes rely on probability or calculus, they link to the lab’s notes on probability and mathematics for data science, which are the canonical home for the probability and mathematics appendices of rme.
Course sites include these notes as a git submodule named sds at the site’s root, and include fragments with paths that start with sds/, for example {{< include sds/_subfiles/intro-MLEs/_def_mle.qmd >}}. This site includes its own fragments the same way, through an sds symlink that points at the repository root.
Quarto resolves @id cross-references only within one rendered page, so a host site that links to a result here uses an explicit link, such as [text](estimation.qmd#def-bias).