| symbol | meaning | LaTeX |
|---|---|---|
| \(\neg\) | not | \neg |
| \(\forall\) | all | \forall |
| \(\exists\) | some | \exists |
| \(\cup\) | union, “or” | \cup |
| \(\cap\) | intersection, “and” | \cap |
| \(\mid\) | given, conditional on | \mid, | |
| \(\sum\) | sum | \sum |
| \(\prod\) | product | \prod |
| \(\mu\) | mean | \mu |
| \(\operatorname{E}\) | expectation | \mathbb{E} |
| \(x^{\top}\) | transpose of \(x\) | x^{\top} |
| \('\) | transpose or derivative1 | ' |
| \(\perp\!\!\!\perp\) | independent | ⫫ |
| \(\therefore\) | therefore, thus | \therefore |
| \(\eta\) | linear component of a GLM | \eta |
| \(\mathopen{}\left\lfloor x\right\rfloor\mathclose{}\) | floor of \(x\): largest integer smaller than \(x\) | \lfloor x \rfloor |
| \(\mathopen{}\left\lceil x\right\rceil\mathclose{}\) | ceiling of \(x\): smallest integer larger than \(x\) | \lceil x \rceil |
| \(\mathbb{1}_{A}(x)\), \(\mathbb{1}\mathopen{}\left(P\right)\mathclose{}\) | indicator function (Section 4): \(1\) if condition holds, \(0\) otherwise | \indic{A}(x), \indicp{P} |
There is no consistency in the notation for observed and expected information matrices (see Table 2).
These notes currently have a mixture of notations, depending on my whims and what reference I had last looked at. Eventually, I will try to standardize my notation to \(I\) for observed information and \(\mathcal{I}\) for expected information.
The percent sign “%” is just a shorthand for “\(/100\)”. The word “percent” comes from the Latin “per centum”; “centum” is Latin for 100, so “percent” means “per hundred” (c.f., https://en.wikipedia.org/wiki/Percentage)
So, contrary to what you may have learned previously, \(10\% = 0.1\) is a true and correct equality, just as \(10 \text{kg} = 10,000 \text{g}\) is true and correct.
Proof. \[ \begin{aligned} 10\% &= 10 / 100 \\ &= \frac{10}{100} \\ &= 0.1 \end{aligned} \]
You are welcome to switch between decimal and percent notation freely; just make sure you execute it correctly.
We can use any of:
\therefore in LaTeX),\Rightarrow),\models)to denote logical entailments (deductive consequences).
Let’s save \(\rightarrow\) (\rightarrow) for convergence results.
See Proof Writing for general guidance on how to present proofs and derivations.
An indicator function is a mathematical function that signals whether an element belongs to a specified set, or whether a given logical condition is satisfied. In statistics and epidemiology, indicator functions are ubiquitous: they represent binary variables, censor and event indicators in survival analysis, membership in subpopulations, and domain restrictions in integrals and sums.
Despite their conceptual simplicity, notation for indicator functions varies substantially across textbooks, research papers, and subfields. This section summarizes the principal notational conventions.
Definition 1 (Indicator function) For any subset \(A \subseteq \Omega\) of a universal set \(\Omega\), the indicator function of \(A\) is the function \(\mathbb{1}_{A} : \Omega \to \{0, 1\}\) defined by:
\[ \mathbb{1}_{A}(x) \stackrel{\text{def}}{=}\begin{cases} 1, & x \in A \\ 0, & x \notin A \end{cases} \]
More generally, for any logical proposition or predicate \(P\), the indicator of \(P\) takes the value \(1\) when \(P\) is true and \(0\) when \(P\) is false:
\[ \mathbb{1}\mathopen{}\left(P\right)\mathclose{} \stackrel{\text{def}}{=}\begin{cases} 1, & \text{if } P \text{ is true} \\ 0, & \text{if } P \text{ is false} \end{cases} \]
Example 1 (Evaluating set and predicate indicators) Consider the real line \(\Omega = \mathbb{R}\), the set of nonnegative numbers \(A = [0, \infty)\), and a continuous random variable \(Y\).
The vast majority of indicator notations belong to one of two families: set notation or predicate notation.
In set notation, the indicator is tied to a set \(A\), which appears as a subscript:
When the function is viewed as a mathematical object in its own right (for instance, as an element of an \(L^p\) function space), authors often omit the argument \(x\), writing simply \(\mathbf{1}_A\), \(\mathbb{1}_A\), or \(I_A\).
In predicate notation, the indicator takes a logical condition, relation, or proposition \(P\) directly as its argument or subscript:
Predicate notation is especially common in applied statistics and survival analysis, where indicators frequently depend on inequalities involving random variables, such as \(\mathbb{1}\mathopen{}\left(T_i \le t\right)\mathclose{}\) (an event occurring before time \(t\)) or \(\mathbb{1}\mathopen{}\left(Y_i = 1\right)\mathclose{}\) (a binary outcome).
The two paradigms are connected by evaluating the predicate indicator at the membership statement \(x \in A\):
\[ \mathbf{1}_A(x) = \mathbb{I}(x \in A) \]
Set notation is more natural when the underlying set \(A\) has a standard name (such as the support of a distribution or a geometric region). Predicate notation is more natural when the condition involves compound inequalities, such as \(\mathbb{1}\mathopen{}\left(0 \le t \le u\right)\mathclose{}\).
In 1962, Kenneth Iverson introduced a compact notation in the programming language APL, later popularized in mathematics and computer science by Donald Knuth: the Iverson bracket.
The Iverson bracket encloses any mathematical predicate \(P\) inside square brackets:
\[ [P] \stackrel{\text{def}}{=}\begin{cases} 1, & \text{if } P \text{ is true} \\ 0, & \text{if } P \text{ is false} \end{cases} \]
Under this notation, set membership is written \([x \in A]\), and the Kronecker delta is simply \(\delta_{ij} = [i = j]\).
The primary advantage of the Iverson bracket is algebraic conciseness: it converts domain restrictions in sums and integrals into unrestricted operations. For example:
\[ \sum_{x \in A} f(x) = \sum_{x} f(x) [x \in A] \]
However, in statistics and epidemiology, square brackets are already heavily overloaded: they denote closed intervals \([a, b]\), conditional expectations \(\operatorname{E}[Y \mid X]\), and matrix delimiters. To prevent visual confusion with expectation brackets or intervals, statistical literature predominantly uses \(\mathbb{1}\) or \(I\) rather than the bare Iverson bracket.
Table 3 compares the major notations encountered across the literature.
| Notation style | Typical syntax | Primary fields | Notes and potential ambiguities |
|---|---|---|---|
| Blackboard bold 1 | \(\mathbb{1}_A(x)\), \(\mathbb{1}(P)\) | Modern probability, mathematical statistics | Unambiguous; distinct from matrices and scalars; standard in this book. |
| Bold numeral 1 | \(\mathbf{1}_A(x)\), \(\mathbf{1}(P)\) | Probability theory, measure theory | Can be confused with a vector of ones \(\mathbf{1} = (1, \dots, 1)^{\top}\). |
| Blackboard bold I | \(\mathbb{I}(x \in A)\), \(\mathbb{I}(P)\) | Econometrics, machine learning, statistics | Clear predicate notation; avoids confusion with numerals. |
| Letter \(I\) | \(I_A(x)\), \(I(P)\) | Classical statistics, epidemiology | Can be confused with the identity matrix \(I\) or Fisher information \(\mathcal{I}\). |
| Iverson bracket | \([P]\), \([x \in A]\) | Computer science, discrete mathematics | Very compact, but square brackets collide with intervals and expectation brackets. |
| Greek letter \(\chi\) | \(\chi_A(x)\) | Real analysis, measure theory | Often termed “characteristic function”; collides with the Fourier transform in probability. |
In these notes, we standardize on blackboard bold \(\mathbb{1}\) via the macros defined in latex-macros/macros.qmd:
\indic{A} produces \(\mathbb{1}_{A}\) (set subscript)\indicp{P} produces \(\mathbb{1}\mathopen{}\left(P\right)\mathclose{}\) (predicate in parentheses)\indiccb{P} produces \(\mathbb{1}\mathopen{}\left\{P\right\}\mathclose{}\) (predicate in curly braces)\1{P} produces \(\text{1}_{P}\) (text numeral with subscript, used in legacy formulas)Blackboard bold \(\mathbb{1}\) is preferred because it avoids all common collisions: it is visually distinct from the scalar \(1\), the identity matrix \(I\), and the information matrices (\(I\), \(\mathcal{I}\)).
Indicator functions translate logical operations on events into ordinary arithmetic on real numbers:
Intersection (“and”): \(\mathbb{1}\mathopen{}\left(A \cap B\right)\mathclose{} = \mathbb{1}\mathopen{}\left(A\right)\mathclose{} \cdot \mathbb{1}\mathopen{}\left(B\right)\mathclose{}\)
Union (“or”): \(\mathbb{1}\mathopen{}\left(A \cup B\right)\mathclose{} = \mathbb{1}\mathopen{}\left(A\right)\mathclose{} + \mathbb{1}\mathopen{}\left(B\right)\mathclose{} - \mathbb{1}\mathopen{}\left(A\right)\mathclose{} \cdot \mathbb{1}\mathopen{}\left(B\right)\mathclose{}\)
Complement (“not”): \(\mathbb{1}\mathopen{}\left(\neg A\right)\mathclose{} = 1 - \mathbb{1}\mathopen{}\left(A\right)\mathclose{}\)
Idempotence: \((\mathbb{1}\mathopen{}\left(A\right)\mathclose{})^2 = \mathbb{1}\mathopen{}\left(A\right)\mathclose{}\)
Expectation gives probability: For any event \(A\), the expectation of its indicator is the probability of the event:
\[ \operatorname{E}[\mathbb{1}\mathopen{}\left(A\right)\mathclose{}] = 0 \cdot \Pr(\neg A) + 1 \cdot \Pr(A) = \Pr(A) \]
This fundamental identity connects probability theory directly to linear expectation. It provides the mathematical foundation for empirical proportions, survival curve estimators, and regression models for binary outcomes.
The terms “stochastic”, “probabilistic”, and “random” are frequently used in statistics and probability theory, often interchangeably in everyday conversation, but they carry nuanced technical distinctions.
As noted in Wikipedia:
Stochasticity and randomness are technically distinct concepts: the former refers to a modeling approach, while the latter describes phenomena; in everyday conversation these terms are often used interchangeably.
Random describes something that occurs by chance, without a deterministic pattern. It is the most general term, used to describe variables or occurrences whose outcome cannot be predicted precisely, only probabilistically. For example, we speak of “random variables” and “random events”.
Note
The term “random” is sometimes used as shorthand for a uniform distribution (especially the discrete uniform distribution), but it can refer to any probability distribution.
Stochastic comes from the Greek “στόχος” (stókhos), meaning “aim” or “guess” (see etymology). In mathematics, a stochastic process is formally defined as a collection of random variables indexed by time or space. The term is almost always used in the context of processes or systems evolving in time or space under uncertain rules. Note that in probability theory, “stochastic process” and “random process” are synonyms (Adler and Taylor 2009; Stirzaker 2005; Kallenberg 2002).
Probabilistic refers to any model, reasoning, or method that explicitly involves probability theory. Probabilistic models assign probabilities to events or outcomes; they focus on quantifying and reasoning about uncertainty based on known or estimated distributions. While all stochastic models are probabilistic (since they use probabilities), not all probabilistic models need to describe processes evolving in time.
| Term | What it describes | Typical use | Example |
|---|---|---|---|
| Random | Single variable or event | Random variable, random outcome | Coin toss, die roll |
| Stochastic | System or process in time/space | Stochastic process | Stock price evolution, Markov chain |
| Probabilistic | Approach/model using probability | Probabilistic model/reasoning | Bayesian inference, regression |
While some sources treat “stochastic” and “random” as practically synonymous, the academic preference is to use “random” for variables and events, and “stochastic” for processes, especially to highlight temporal or spatial structure in the modeling.
In grad school, we are asked to learn from increasingly disorganized materials and lectures. Not coincidentally, as the amount of organization decreases, the amount of complexity increases, the amount of difficulty increases, the number of reliable references decreases, and the amount of inconsistency in notation and content increases (both between multiple references and within single references!). In other words, as you approach the cutting-edge of most fields, you start to encounter into content that hasn’t been fully thought through or standardized. This lack of clarity is unfortunate and undesirable, but it is understandable and inevitable.
It’s worth noting that calculus was formalized in the 1600s, elementary algebra was formalized around 820, and arithmetic even earlier. And calculus still has several competing notation systems. In contrast, the field of statistics only emerged in the late 1800s and early 1900s, so it’s not surprising that the notation and terminology is still developing. Generalized linear models were only formalized in 1972 (Nelder and Wedderburn (1972)), which is very recent in terms of the pace of scientific development.