Before reporting a finding, name the artifact the claim is about,
and confirm that is the one you read.

The failure this prevents is not lazy verification.
It is thorough verification of the wrong object.
Evidence gets gathered, it is genuine, the reasoning from it is sound,
and the conclusion is false --- so nothing along the way feels like a guess.
That is what separates this from an unchecked assertion,
and it is why care alone does not catch it:
the sensation of having checked is present and correct,
and it is attached to the wrong thing.

Worked-example case records for the rules below live in
[`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md),
moved out of the auto-loaded context.

## What this adds to the rules already covering one substitution each

Three fragments already name a *particular* adjacent artifact,
and each is worth reading on its own terms:

- [`metacognitive-monitoring`](metacognitive-monitoring.md)'s
  "Read the artifact that failed, not the one beside it"
  governs the **cause of a failure** read off a neighbouring step,
  a sibling job, or the log line above the error.
- [`fixtures-are-not-evidence`](fixtures-are-not-evidence.md)
  governs a **test fixture** standing in for the system it imitates.
- [`fact-check-prose`](../writing/fact-check-prose.md)'s
  "Confirm a rendered page carries your commit before reading anything off it"
  governs a **published build** standing in for the branch that produced it.

What none of them reaches is the substitution's generality.
Each is written as a fact about one situation --- diagnosing a failure,
reading a fixture, fetching a preview ---
so a claim outside those three situations matches none of them,
and the rule that would have caught it never loads.
The substitution is not a property of failure diagnosis.
It happens to a claim about **state**
(what a plugin currently contains),
about **location**
(where a file lives),
and about **mechanism**
(whether a cache is ever read),
in exactly the same shape.

**The boundary in the other direction is worth naming, because this fragment is where a reader lands first and the rule they need may be elsewhere.**
Every shape here begins with a substitution: you read A and the claim is about B.
The neighbouring failure has no substitution in it at all --- the artifact is the right one, it is read correctly, and the sentence after the reading answers a question that artifact does not address.
Nothing in this fragment fires on that, because there is no wrong object to name.
[`metacognitive-monitoring`](metacognitive-monitoring.md)'s "A sound measurement does not license the claim standing next to it" is the rule for it.
So when a check of yours came back clean and the claim still feels under-supported, ask which of the two is happening: whether you read the wrong thing, or read the right thing and then took a step.

- **Do:** send a claim to that section instead of this one when the artifact is the correct one and the doubt is about the step taken from it.
- **Don't:** read a shape here failing to match as evidence the claim is supported --- every substitution shape in this fragment covers substitutions only.

One case sits between the two, and has its own section below --- "A measurement of the right artifact can still be scoped narrower than the claim made from it": the artifact is right, the reading is right, and the claim is the *same* proposition at a wider scope than the measurement covered.
That is not a step taken from the measurement, so it is not the neighbouring rule either.

## The four shapes

Recognizable in advance, which is the point of enumerating them:

- **A cached copy for the origin.**
  A CDN-served page, a stale local checkout, a plugin cache.
  Go to the authoritative store instead:
  the branch's raw bytes, the install directory, the API.
- **A checkout for the run.**
  What a branch contains is not what a workflow checked out.
  An `issue_comment` trigger checks out the default branch,
  not the pull request's head.
- **One half of a mechanism for the whole.**
  A cache `save` with no `restore` caches nothing.
  A marketplace entry with no install installs nothing.
  Find the counterpart before asserting that the mechanism works.
- **A neighbour for the target.**
  A directory that happens to contain the files you expected
  is not thereby the path they are read from.
  It can coincide today and diverge tomorrow.

**A document that delegates carries claims about its delegate, and those are
the ones nobody checks.**
The shapes above are all about verifying a *claim you are making*.
This is about the claims a **delegating** document makes structurally, just by
saying "run X's steps 1 through 3" --- because that sentence quietly asserts
what those steps do.
Writing it feels like pointing rather than asserting, which is why no
claim-checking instinct fires on it.

The concrete failure: a skill built around "one confirmation, no mutation
before it" told the reader to run another skill's steps 1 through 3, and step 2
of that range ran a real `git worktree prune`.
The guarantee at the centre of the design was false, and every internal check
passed, because the skill was self-consistent --- the falsehood lived in the
*other* file.

(Measured 2026-08-21 on
[ai-config#1849](https://github.com/Morrison-Lab/ai-config/pull/1849),
[review comment](https://github.com/Morrison-Lab/ai-config/pull/1849#discussion_r3834408153).
The delegating skill is `skills/clean-git/SKILL.md`, added by that PR, and the
delegate is [`clean-worktrees`](../../skills/clean-worktrees/SKILL.md) step 2,
which runs `git worktree prune -v` rather than `--dry-run`.
Note that a grep of `main` cannot corroborate this while #1849 is unmerged,
since the delegating skill does not exist there yet --- which is the case for
citing the PR rather than the file.
Fixed in `c59ae986` by narrowing the gate's invariant to "nothing that can lose
work happens before confirmation".)

So when you delegate to a numbered range, open that range and read it.
Any property you assert about it --- that it mutates nothing, that it is
read-only, that it asks before acting --- is a claim about a file you did not
write, and it decays when that file changes without touching yours.

- **Do:** read every step you delegate to before describing what it does.
- **Do:** state the invariant that survives the delegate's actual behaviour,
  rather than the one you wish it had.
- **Don't:** treat "run X steps N through M" as a pointer; it is an assertion
  about N through M.

## The test

Confirming the claim against what you read cannot detect this,
because that is exactly what already happened.
Ask the falsifying question instead:
**what would have to be true for this claim to be false,
and could the artifact I read show me that?**

An artifact that cannot exhibit the claim's failure mode
has not tested the claim, however much it agrees with it.
A CDN copy cannot show that the branch differs from it.
A `save` step cannot show that no `restore` exists.
This is
[`fail-fast`](../principles/fail-fast.md)'s denominator move applied to
evidence:
a check whose passing and failing readings are indistinguishable
is not yet a check.

Where the answer is one command, run the command.
The cleanest measured case was settled by `grep -n "actions/cache"`
over a workflow directory:
a single line came back, `actions/cache/save@v6`,
which killed a claim that had already propagated
into an issue body, a second issue, a commit message,
and a pull request opened on its premise.

- **Do:** state which artifact a claim rests on, then read *that* one ---
  the branch's raw bytes over the rendered site,
  the install manifest over a guessed path,
  a search for the counterpart step over an inference from the one you found.
- **Do:** ask what would falsify the claim,
  and whether the artifact in hand could show it.
- **Don't:** treat "I checked something real and it supported the claim"
  as verification.
  The support is real and it is about a different object.
- **Don't:** read a specific, checkable-looking particular as a sign of rigour.
  Specificity is inherited from the artifact that was read,
  not from the one the claim is about.

## A working-directory checkout is another shape, and it stays silent

Shape 1 already names "a stale local checkout" among its examples, so this is a sharpening of that shape rather than a wholly new one.
What it adds is the *tell*, which shape 1 leaves implicit: its remedy, go to the authoritative store, presumes you already suspect the copy, and that presumption is exactly what fails here.
A working-directory read offers nothing to raise the suspicion: the path resolves, the file exists, and its contents are real bytes from a real commit.
Nothing distinguishes reading the default branch from reading a feature branch that happens to be checked out, so `cat <path>` in a repo you have open reads as consulting the repository rather than as consulting one revision of it.

Two properties make it worse than an ordinary stale read.

**The staleness is invisible in the direction that matters.**
A file missing from the branch errors, and an empty file is obviously wrong.
A file that is merely *older* returns a complete, coherent, plausible document --- frequently the document you remember, since a feature branch usually forked from a `main` you had already read.
So the failure mode is not confusion but false confidence.

**A shared checkout moves under you.**
Another session, or a `@claude` bot reacting to PR activity, can switch branches or pull between your read and your next command, so the branch you verified once is not the branch you are still on.
`git reflog` is what shows this after the fact; nothing shows it at the time.

The cheap check is one command, and it belongs *beside the read*, not once at session start:

```bash
git -C <repo> rev-parse --abbrev-ref HEAD     # which revision am I reading?
```

The better move is to skip the checkout altogether whenever the claim is about what the repository currently documents, and read the revision by name:

```bash
git -C <repo> fetch -q origin
DEF=$(git -C <repo> remote show origin | sed -n 's/.*HEAD branch: //p')
git -C <repo> show "origin/$DEF:<path>"
```

This names the revision in the command, so the bytes you read and the revision you cite cannot come apart, and it answers correctly whatever the checkout is doing.
Resolve the default branch from the repo rather than assuming `main`.

The `fetch` is load-bearing rather than tidiness.
`origin/<default-branch>` is a remote-tracking ref, so it is only as current as the last fetch --- without one, this substitutes a different stale artifact for the right one, which is this whole fragment's failure mode wearing the remedy's clothes.

The consequence generalizes past reading.
An assignment derived from a stale read is wrong in a way [`challenge-the-assignment`](challenge-the-assignment.md) cannot catch, because every premise check the recipient runs confirms a document that genuinely exists.
When a brief, an issue body, or a review finding asserts what a repository says, cite the revision alongside the path.

- **Do:** name the revision in the command when the claim is about what a repo currently says.
- **Do:** re-check the branch beside each read in a shared checkout, rather than trusting a verification from earlier in the session.
- **Don't:** treat a successful `cat` in a repo directory as evidence about that repo's default branch.
- **Don't:** read plausibility as freshness --- a feature branch forked from a `main` you already read returns exactly what you expect.

See [`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md), "A stale branch read that produced two issues and a config edit".

## A different endpoint is another shape, and the two names read as synonyms

Two APIs can describe overlapping but distinct populations under names that read as synonyms.
GitHub's `GET /repos/{owner}/{repo}/pages/builds` documents itself as listing *builds* of a Pages site ([REST API docs](https://docs.github.com/en/rest/pages/pages), read 2026-09-17), while `GET /repos/{owner}/{repo}/deployments?environment=github-pages` lists *deployments* to an environment.
Both answer to "how many times has this site deployed" in English, and they enumerate different objects, so a count taken from one is not comparable to a count taken from the other.

What distinguishes it is not that the substitution is silent --- [`A working-directory checkout is another shape, and it stays silent`](#a-working-directory-checkout-is-another-shape-and-it-stays-silent) says the same of a stale read, and says it first.
It is that there is no authoritative store to go to.
Every other shape's remedy presumes one of the two artifacts is the right one, so suspecting the substitution is most of the work of undoing it.
Here both endpoints are authoritative, each for its own population, and neither is the correct one in the abstract --- only the endpoint the original measurement used is comparable to the original measurement.
So the usual move, go and check against the real thing, does not terminate: whichever endpoint you reach for is a real thing.

Re-measuring a prior claim through the *other* endpoint and getting a different number is therefore evidence that the two endpoints disagree, not evidence that the original figure decayed.
A drift claim needs both readings taken the same way, and the endpoint is part of "the same way".

- **Do:** use the endpoint the original measurement named, and say which one it was.
- **Do:** use the other party's endpoint when checking someone else's number, before concluding drift.
- **Don't:** read a different number from a different endpoint as decay.
- **Don't:** treat two API paths as interchangeable because their names describe the same thing in English.

## A mechanism's prose is not the mechanism's definition

A hook's comment, a skill's description, a docstring: each one explains a
mechanism, and each is written by someone who had a specific case in mind
while writing it.
That case becomes the prose's running example, and the example is narrower
than the mechanism it illustrates almost by construction --- a comment
motivates a design decision by pointing at the situation that forced it, not
by re-deriving the mechanism's full scope from nothing.

Reading the prose and concluding the mechanism is scoped to the example is
this fragment's substitution again, in a form the four shapes do not name:
the artifact you read (the prose) is real, and the claim you draw from it (the
mechanism's boundary) is about a different artifact --- the mechanism's own,
separately-recorded definition, which the prose was never trying to state
exhaustively.

This differs from ["A summary read as its
source"](verify-the-right-artifact.cases.md) in what goes missing.
A summary drops a source document's hedges and caveats in the act of
restating it, so the fix is to open the source it claims to restate.
Here there may be no single document being restated at all --- the prose
motivates a design by its launching case, and the mechanism's actual boundary
lives in a separate, independently-queryable artifact (a label's own
description, a config file, a schema) that the prose never claims to
reproduce.
The fix is not to read the prose more carefully; it is to query the
mechanism directly.

Two measured instances, one on each side of the substitution:

**A guard's comment framed an exemption around its launching case, and the
label's own definition was broader.**
`hooks/no-unreviewed-pr.py`'s comment introduces `EXEMPT_LABEL = "no-ai-review"`
while discussing a redaction PR, and pairs it with an env var literally named
`ALLOW_UNREVIEWED_REDACTION_PR`.
Reading only the comment, applying the label to a non-redaction PR reads as
inventing an exemption the label was never meant to cover, and three separate
attempts to use it were refused on that reading.
The label's own description, on the repository, says otherwise:

```
$ gh label list -R Morrison-Lab/ai-config --json name,description \
    --jq '.[] | select(.name == "no-ai-review")'
{"name":"no-ai-review","description":"AI code review deliberately withheld on this PR; see the PR comment for the reason"}
```

General, and silent about redaction.
Redaction is the comment's motivating case, not the label's scope, and the one
command above settles which of the two governs.
(Morrison-Lab/ai-config#3304, 2026-09-06/07.)

**A written verification step tested a reconstruction of a string, not the
string a file actually contains.**
A memory entry claimed a single backslash in some source text makes a
substring check return `False`, and the claim was "verified" by rebuilding the
intended string with `chr(92)` and running the check against that
reconstruction.
That confirms the check behaves as expected on the string you *meant* to
write.
It says nothing about the string the file *actually holds*, because a
non-raw Python string literal has its escapes decoded by the parser before the
check ever sees it --- `"a\b"` in source is the two characters `a` and a
literal backslash-b sequence only if the parser reads it as written, and
pasting the file's real text and evaluating it in place (rather than
reconstructing what it "should" contain) is the only read that reports the
actual behaviour.
The general form: when the object under test is "does this file's text trip
this check", read the file's text, not a hand-rebuilt stand-in for it, however
carefully the stand-in is constructed.
(Measured in the same 2026-09-06/07 session as the label instance above,
while drafting a `memories/mistake-patterns.md` entry; caught before that
entry was committed, so no PR or issue number attaches to it.)

- **Do:** query a mechanism's own recorded definition (a label's description,
  a config's schema, a constant's value) before concluding its scope from a
  comment's motivating example.
- **Do:** evaluate a check against the artifact's actual bytes --- extracted
  from the file, not reconstructed from what you believe it should contain
  --- when the claim is about what that artifact's text does.
- **Don't:** read a comment's launching case as a boundary on the mechanism
  it explains; a comment motivates, it does not define.
- **Don't:** treat "I rebuilt the string and the check passed" as evidence
  about a file's own text; rebuilding tests the intent, not the artifact.

See [`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md), "A guard's comment and a label's own description disagreed about scope".

## A comparison's base is an artifact too, and it moves the scope in both directions

Every shape above concerns an artifact you **read**.
A diff is an artifact you **derive**, from two refs, and attention goes to the one you are interested in --- the branch under review.
The other ref is the base, and nothing about naming it feels like making a claim.

`git diff main...pr-98` reads as "the PR's changes".
It is not.
It is the changes since whatever commit your **local** `main` shares with `pr-98`, and a local branch is a cached copy of a remote branch, which is shape 1 exactly.
The three-dot form is what conceals it: a merge-base is a real computation over real history, so the range feels self-correcting, and the sensation of having used the careful form stands in for having checked the ref the careful form is computed from.
A merge-base is only ever as fresh as the ref you fed it.

**The error runs in both directions, and the quieter one is the worse.**
A base **behind** its remote moves the merge-base earlier, so the diff gets bigger.
The extra content is commits that already merged --- other people's work, already reviewed, already landed --- and a review run on it produces findings against code the author of this PR never wrote, spending the author's time and the reviewer's credibility at once.
A base carrying local commits the remote lacks --- **ahead** of it, or diverged from it --- where the head branch also carries those commits, moves the merge-base *later*, so the diff gets smaller.
That is the dangerous one.
An over-wide diff produces findings the author will dispute, so it announces itself within a round;
an under-wide one silently omits part of the change and comes back clean, and a clean verdict is the one nobody questions.

**Nothing in the output announces either direction.**
A 53-file diff and a 14-file diff are equally plausible artifacts.
Every finding derived from the wrong scope is individually well-formed, correctly quoted, and about a real line of real code.
So the usual detector --- a finding that looks wrong --- never fires, because none of them do.

A dispatched reviewer cannot catch it either, and [`challenge-the-assignment`](challenge-the-assignment.md) says why in the mirror: a brief must not assert what the author cannot query about the *recipient's* environment.
This is the inversion of that.
The brief asserts something about the author's **own** environment, which the author could have queried in one command and did not, and which the recipient cannot query at all.

**The falsifying question in "The test" above disposes of it, and its answer is that the diff cannot testify about itself.**
Ask what would have to be true for the base to be wrong, and whether the diff in hand could show it.
It could not.
The forge could, and it is one call:

```bash
git -C <repo> fetch -q <remote>
BASE=$(git -C <repo> merge-base <remote>/<default-branch> <pr-ref>)
git -C <repo> diff --shortstat "$BASE" <pr-ref>
gh pr view <N> --json changedFiles,additions,deletions
```

The two readings must agree.
A mismatch means one of the two refs is wrong, and the base is only the first place to look: the local copy of the *head* goes stale the same way, so a PR that has received commits since you fetched it disagrees with a perfectly correct base.
Re-fetch the head before concluding anything about the base. (Rename detection is a smaller third cause, since `diff.renames` is on by default and the forge counts renames its own way.)
This section is about attributing a discrepancy to the wrong artifact, so a single diagnosis for a symptom with several causes is the failure it describes rather than a shortcut past it.
Resolve the default branch from the repo rather than assuming `main`, and note the remote is not always `origin` --- a dual-forge repo has the PR's forge under a second remote name, and the fetch has to name that one.

The `fetch` is the load-bearing half, for the reason the working-directory section already gives: a remote-tracking ref is itself a cached copy, current only to the last fetch.
A fetch at session start does not cover a review dispatched an hour later, which is [`check-before-pushing`](check-before-pushing.md)'s point about a reading of a moment that has passed, moved from the push to the dispatch.

Report the base you resolved.
A review brief, or a review comment, that states the merge-base SHA and the file and insertion counts alongside its findings is one a reader can check;
one that says "the PR's diff" is not.

- **Do:** resolve a review diff's base from a remote-tracking ref, after fetching that remote, and state the merge-base SHA and the file and insertion counts beside it.
- **Do:** cross-check the derived counts against the forge's own (`gh pr view --json changedFiles,additions,deletions`) before dispatching, and treat any mismatch as a wrong base rather than as noise.
- **Don't:** pass a bare local branch name as a diff's base, in your own command or in a brief you hand a subagent.
- **Don't:** read the three-dot form as self-correcting --- it computes a merge-base from refs you supplied, and cannot know one of them is behind.
- **Don't:** wait for an implausible finding to reveal it;
  over-wide scope produces findings that are all individually sound.

`hooks/warn-stale-review-diff-base.py` is the instrument, per [`algorithmatize-checks`](algorithmatize-checks.md).
The rule it enforces is not "was the local ref fresh", which no hook can know, but "name a remote-tracking ref", which is lexical.
It warns and never blocks, because a bare local base is entirely correct for an ordinary local comparison and the hook cannot tell those apart.
It has no fetch-based discharge on purpose: [`keep-checkouts-fresh`](keep-checkouts-fresh.md) mandates a fetch at session start, so keying on one would silence the hook in exactly the sessions that follow the corpus.

See [`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md), "A stale local base that nearly quadrupled a review diff's file count".

**A base that is fresh, correct, and current can still be one the comparison is incapable of failing against.**

The error above is staleness, and every remedy it names is a freshness remedy.
This one survives all of them.
The ref is right, the fetch is current, the detector runs on every input, and the comparison still returns zero --- because the base has no behaviour of the kind being counted.

**The mechanism is already written down**, in
[`fixtures-are-not-evidence`](fixtures-are-not-evidence.md)'s
"Which ref to restore from, not only which file":
a base branch that lacks the structure under test cannot reproduce the
behaviour, so it returns a plausible result rather than an error, and the
remedy is to baseline against the previous round's head and to prefer a
three-way comparison over a two-way one.
Read that subsection for the argument;
this one adds two things to it, both about the **zero** rather than about the
baseline.

**A large case count weakens a zero rather than strengthening it.**
A differential run counts *transitions* between what two revisions decide, so
where the base classifies the whole family one way by default, the transition
has nothing to transition from.
Scaling that up multiplies the comparison's reach and not its capability:
60,000 cases returning zero reads as thorough while carrying exactly as much
information as one case would.
The count is the most reassuring number the run can print and the one least
entitled to reassure.

**So report a zero with the negative control's hit count beside it.**
A detector that never reached its read site and a detector that reached it
60,000 times and found nothing print the same zero.
Only an instrumented count of reads at the site separates them, and it is the
half that turns a zero into a measurement.

The section on baseline verdicts in
[`algorithmatize-checks`](algorithmatize-checks.md)
covers the opposite direction --- a baseline flag earned by coincidence and read as a regression the branch introduced.
That one produces a finding somebody argues with.
This one produces a clean run nobody questions, which is why it can repeat.

Measured on [ai-config#3635](https://github.com/Morrison-Lab/ai-config/pull/3635), whose own merged body records it:

> A 60,000-case differential fuzz claimed "zero BLOCK-to-allow against `main`".
> True and nearly vacuous --- `main` has no substitution scanner, so it already allows this whole family and no regression can appear in that comparison.
> **The baseline that can fail is the previous round.**

The consequence is in the same body's table.
`bash <(true; case b in b) echo "<m>";; esac)` reads `allow` on `main`, `BLOCK` at rounds 4 and 5, `allow` at round 6, and `BLOCK` at the head that merged.
Round 6 was a regression against its own two predecessors, and a comparison against `main` cannot see it by construction: the regressed revision and `main` both return `allow`, so the transition count is zero on precisely the input that shipped the defect.

The pre-merge gate on that PR did both, and **named the revisions rather than counting back from a moving head**:
29,813 strings scored at `994b975c` and its four predecessors --- `6449c317`, `d7169012`, `c8482025`, `73e95727` --- found 0 transitions, and a separate 240,000-scan comparison reported 0 diffs alongside an instrumented count of 4,640 reads at the site.
An earlier draft of this paragraph wrote that set as ``HEAD`` and ``HEAD~1``..``HEAD~4``, which names nothing once the branch merges and `HEAD` is somebody else's.
That is this fragment's own subject applied to a citation: a relative ref is a claim about the reader's checkout, and it resolves to a different artifact in every one.

The falsifying question in "The test" above settles it in one reading: ask what the base does with the family under test.
If the answer is that it has no opinion, the base cannot testify.

- **Do:** name what the baseline revision does with the construct under test, in the same sentence as the zero.
- **Do:** baseline a hardening branch against its own previous rounds, and say which ones **by SHA** --- a relative ref names a different commit in every checkout that reads it.
- **Do:** report the negative control's hit count beside a zero, so the zero distinguishes itself from a detector that never ran.
- **Don't:** read a large case count as strengthening a zero --- it multiplies the comparison's reach, not its capability.
- **Don't:** baseline against the default branch for a feature the default branch does not have;
  that arm agrees with every revision, including the regressed one.

## A measurement of the right artifact can still be scoped narrower than the claim made from it

Every shape above is a *substitution*: the thing read is not the thing the claim is about.
This one is about the **sentence** rather than the object: whatever was measured, the claim reported covers more than the measurement did.
Nothing about it feels like guessing, because the number really was derived by a real command.

The two failures overlap, and the four instances below show it.
Two are substitutions as well --- a build path the system never uses, and a baseline read that returned nothing --- so the shapes above would have caught them had anyone asked.
Two are not: the display-math and `microtype` cases read exactly the right document, and only the sentence overreached.
What they share is the tell, not the mechanism: a scope decision made once during setup, and never repeated in the sentence that reports the result.

[`metacognitive-monitoring`](metacognitive-monitoring.md)'s "A sound measurement does not license the claim standing next to it" already names the general gap between a measurement and a neighbouring claim, including cases where the claim is about a different proposition entirely.
What follows is the narrow case where the claim is about the *same* proposition as the measurement, just at a wider scope along one identifiable axis --- build path, ref, math subset, package set --- so the fix is naming that one axis rather than restating the whole claim.

Four instances from one session, all against the same PR, none of which felt like a guess at the time:

- **A build path that bypasses the real pipeline.**
  A harness extracted `$$...$$` math blocks from a `.qmd` chapter and ran `pdflatex` on them directly, to check whether the chapter's math compiles.
  The book never builds that way --- Pandoc reads the source, expands the `macros.qmd` LaTeX macros the chapter actually uses, and only then hands TeX to the renderer.
  Compiling the raw source with `pdflatex` measures an artifact the build never produces, so "the chapter's math does not compile" was a claim about a document nobody ships.
  Two issues were filed on that premise before the mismatch surfaced.
- **A submodule path read through `git show`, with the error thrown away.**
  The same harness fetched the comparison baseline's macros with `git show <ref>:latex-macros/macros.qmd`.
  `latex-macros` is a submodule, so that path is a gitlink in `<ref>`'s tree rather than a blob.
  Git says so, loudly: measured `rc=128` and
  `fatal: path 'latex-macros/macros.qmd' exists on disk, but not in 'HEAD'` on stderr.
  Only *stdout* was empty --- and the harness read stdout alone, discarding both the status and stderr,
  so the baseline arm compiled with zero macros defined and inflated every figure built against it
  (a "153pt worst case" that was really 47pt once the real baseline macros loaded).
  The lesson is the harness's, not git's: an empty read is only silent if you silence it.
- **Display math measured, inline math assumed included.**
  An overfull-box measurement scanned only `$$...$$` display blocks and was reported as covering "the chapter" --- it never touched the inline `$...$` math in the parent file, some of which also overflowed.
- **A required package left out of the harness, silently changing the answer.**
  The same overfull-box measurement ran without `microtype` loaded, reporting 0 overfull boxes where the `microtype`-loaded run of the same document reported 1 --- `microtype` changes line-breaking, so the count is not a rounding difference, it is a different measurement wearing the same label.
  Measured in-session on `d-morrison/rme#1138` rather than in a filed artifact, unlike the figures above: rme#1154 carries the corrected overfull table but records nothing about package configuration, so this arm is anchored here and nowhere else.

The shared shape: a scope decision --- which build path, which ref, which subset of the math, which packages --- gets made once while setting up the measurement, and then the sentence that reports the result names the whole claim ("the chapter's math", "0 overfull boxes") rather than the slice that was actually run --- or, where the slice was a comparison baseline, reports a difference against it ("a 153pt worst case") as though the baseline had loaded.
"The test" section above already supplies the fix for a substituted artifact;
the fix here is a stricter version of the same falsifying-question test, aimed at scope rather than identity: **what does this measurement cover, and is that the same thing the claim names?**

- **Do:** state a measurement's scope in the same sentence as its number --- which build path, which ref, which subset, which flags --- rather than in a paragraph the reader has to reconstruct.
- **Do:** confirm the artifact measured is produced by the same path the real system uses, not a hand-rolled shortcut that happens to consume the same source file.
- **Do:** check a read's exit status, not just whether it returned bytes --- the one empty read among these four instances carried a non-zero status and a `fatal:` message, both of which the harness discarded.
- **Don't:** report a subset measurement under the claim's full name without naming the subset.
  Display blocks only, reported as "the chapter's math".
  One package configuration, reported as "0 overfull boxes".
- **Don't:** trust a comparison baseline's absolute number without confirming its own inputs loaded --- an empty or under-configured baseline arm inflates every relative claim built on it.

(d-morrison/rme#1138, 2026-09-09: all four measured in one long session on the same PR.
The Pandoc-bypass and the empty-submodule-baseline are also written up in that PR's own thread and in d-morrison/rme#1154's "Two instrument traps" section;
the bypass produced a wrong "fix" and two issues filed on the false "math does not compile" premise, one of them d-morrison/macros#85, closed not-planned once the Pandoc-expansion mistake was found.)

## A sweep's PREDICATE is a choice too, and a named standard can stand in for the policy actually being applied

The section above keeps the predicate and narrows the scope: the right question was asked of too little.
This one inverts that.
The scope is complete --- every file read, every one of them scanned --- and the *question* is somebody else's.

The evidence is unusually strong here, which is the whole problem.
A sweep that visits every file and returns zero is a true statement about the predicate the sweep ran, and that predicate was exhaustively applied.
Nothing is missing from the coverage, so none of the width remedies above fire, and none of the emptiness remedies fire either --- a negative control would have confirmed the pattern works, because the pattern *does* work.
It answers a different question than the decision needed.

**The tell is that the sweep's terms came from a named standard rather than from the decision in hand.**
FERPA, HIPAA, PII, GDPR, an SPDX license list, a secrets-scanner ruleset: each is real, externally validated, and thorough about its own subject, so completing one reads as diligence in a way an improvised list never does.
That authority is exactly what suppresses the next question.
A standard is written for *its* decision, so its predicate and yours overlap rather than coincide, and the residue --- everything your policy forbids that the standard never contemplated --- is invisible by construction.

The check is one sentence, written before any pattern is typed: **state the predicate the decision actually turns on**, in the policy's own terms.
Then derive each search term from a clause of that sentence, and report which clause each pattern discharges.
A clause with no pattern beside it is the gap, and it is visible in the report rather than in the material.
The named standard then appears where it belongs --- as one clause among several, not as the sweep.

- **Do:** write the deciding policy's predicate as a sentence first, and derive every pattern from a clause of it.
- **Do:** report the clause each pattern discharges, so a clause nothing searched for shows up as a blank row rather than as silence.
- **Don't:** let a named compliance standard's checklist stand in for the policy's predicate --- it was written for a different decision and only overlaps yours.
- **Don't:** read a thorough, externally validated checklist's zero as clearing a decision that checklist was not written to make.

(Measured 2026-09-18, importing course material from a OneDrive folder into two sibling repos --- `Morrison-Lab/mln`, student-facing and intended to become public, and `Morrison-Lab/mlg`, private grading.
The decision being made was *which repo each file goes in*, whose predicate is "does this reveal anything a student is to be graded on".
The sweep that ran scanned every imported file for PII and for health keywords, found nothing, and reported the material clean.
It never searched for `Exercise Solution`, `answer`, or `solution`.
An adversarial review round then found a PowerPoint slide hidden with `show="0"`, titled `Exercise Solution:`, carrying worked answers to a graded exercise, in the repo intended to go public.
Both sweeps were sound.
Only one of them was about the decision.)

## A ref that resolves to a different commit than it did a moment ago

Most shapes above substitute one artifact for another.
A **moving ref** is the case where the operand you typed never changes and the commit it names does, so there is nothing to point at afterwards: the command line reads identically before and after, and the record of what you ran shows no error.

`FETCH_HEAD` is the everyday instance.
It is not a ref but a file that `git fetch` **replaces**, so a fetch run for an unrelated reason silently re-points an operand you set up earlier.

Four mechanics decide whether this bites, and three of them run opposite to the obvious guess.
Measured on git 2.43.0, in a throwaway clone:

- **A multi-ref fetch writes every ref, one per line, and `rev-parse` takes the FIRST.**
  ```console
  $ git fetch origin feat main && cat -A .git/FETCH_HEAD
  08bec512...^I^Ibranch 'feat' of /tmp/.../up$
  fae0bc12...^I^Ibranch 'main' of /tmp/.../up$
  $ git rev-parse FETCH_HEAD
  08bec512...        # feat, the first ref named
  ```
  Reversing the arguments reverses the answer.
  So a multi-ref fetch is *not* the hazard --- naming the ref you want first gets you that ref.
  Note the format from `cat -A`: each line is `<sha>` TAB `[not-for-merge]` TAB `<description>`, so there is a middle field between the SHA and the description, empty for a ref marked for merge.
  A `cut -f2` reaches that middle field, not the description.
- **A later fetch is the hazard**, because it replaces the file outright.
- **A fetch that FAILS truncates `.git/FETCH_HEAD` to zero bytes**, and `git rev-parse --verify FETCH_HEAD` then fails loudly rather than returning a stale value.
  That direction is safe, and worth knowing because the deleted-branch case (`couldn't find remote ref`, routine after a squash-merge with auto-delete) lands here rather than in the silent one.
- **`--dry-run` ANNOUNCES a write it does not make, which is this section's own failure arriving through a flag that looks like protection.**
  `--no-write-fetch-head` leaves the prior value intact too, but says nothing, so only `--dry-run` actively misleads:
  ```console
  $ git rev-parse FETCH_HEAD        # main
  fae0bc12...
  $ git fetch --dry-run origin feat
   * branch            feat       -> FETCH_HEAD
  $ git rev-parse FETCH_HEAD        # still main, not feat
  fae0bc12...
  ```
  Git reports that `feat` landed in `FETCH_HEAD`, and the next read answers the *previous* fetch.

Measured 2026-09-15 on [ai-config#3687](https://github.com/Morrison-Lab/ai-config/pull/3687), checking whether it collided with a peer PR.
The sequence matters, because the first comparison was correct and only the reads after the second fetch went wrong:

```bash
git fetch -q origin ums/qbt-session-2026-09-14 main
git merge-tree --write-tree HEAD FETCH_HEAD    # HEAD vs the PEER branch -- correct

git fetch -q origin main                       # FETCH_HEAD replaced

# the comparison that broke: two reads, one per ref, counted and compared
git show origin/main:path | grep -c '^- '      # main
git show FETCH_HEAD:path  | grep -c '^- '      # ALSO main, silently
```

Both operands of that count comparison resolved to `main`, so it returned equal numbers --- and equality is what agreement looks like.
Nothing exposed it until the two supposedly different revisions reported an identical count, which was implausible enough to force a re-read of what had actually been passed.

Note what the first draft of this very entry got wrong, since it is the same error one level up: it blamed the multi-ref fetch, on the guess that `FETCH_HEAD` keeps the last ref named.
It keeps the first.
The guess was never measured, and an adversarial review measured it.

This fragment's own test settles the class in one reading: ask what would have to be true for the claim to be **false**, and whether the artifact in hand could show it.
Two reads of one commit cannot show a difference, so equality between them is arithmetic rather than evidence --- the same emptiness [`A comparison's base is an artifact too`](#a-comparisons-base-is-an-artifact-too-and-it-moves-the-scope-in-both-directions) names for a base that classifies the whole family one way by default.

[`fully-clean`](fully-clean.md) already carries the remedy as working code: its pre-merge gate captures `tip=$(git rev-parse --verify FETCH_HEAD)` after the first fetch and a **separate** `head=$(...)` after the second, rather than trusting one capture across both.
The same site notes that a shell variable does not survive into a later tool call, which is exactly the multi-call shape this incident spanned --- so the pin has to be **printed and written down**, not left in `$tip`.
[`ums`](../../skills/ums/SKILL.md)'s worktree block chains its fetch with `&&` for a related but distinct reason it states itself: a failed fetch must stop the block rather than let it proceed on an older `FETCH_HEAD`.

- **Do:** resolve a fetched ref to a SHA immediately, print it, and use the printed SHA thereafter --- a pin you cannot quote is not a pin.
- **Do:** name the ref you actually want **first** when fetching several.
- **Do:** treat two refs producing *identical* results as a prompt to confirm which commits were compared, rather than as agreement.
- **Don't:** reuse `FETCH_HEAD` across any later fetch.
- **Don't:** read `* branch <ref> -> FETCH_HEAD` from a `--dry-run` as meaning a later read will return that ref;
  it will return whatever the last real fetch wrote.
- **Don't:** read "no conflict" or "no difference" as being about the pair you meant without confirming both operands resolved where you thought.

## A reviewer's counter-measurement needs the same check the claim it rebuts would have needed

[`A measurement of the right artifact can still be scoped narrower than the claim made from it`](#a-measurement-of-the-right-artifact-can-still-be-scoped-narrower-than-the-claim-made-from-it) is about the same artifact measured at a narrower scope than the claim names.
This one is a plain substitution, the kind the four shapes above describe --- a different document standing in for the one the claim is about --- and it is worth its own entry only because of *who* commits it: a **reviewer** refuting someone else's claim rather than an author supporting their own.
That is easy to miss, because a rebuttal reads as skepticism rather than as an assertion --- "I tested this and it isn't true" sounds like diligence applied, not like a new claim that itself owes [`dont-take-my-word-for-it`](../principles/dont-take-my-word-for-it.md).
A finding backed by a real command is not thereby a finding backed by the *right* command, and nothing about the reviewer's own confidence distinguishes the two.

A commit fixing a broken macro (`\def\v0`/`\def\v1` silently overriding `\renewcommand{\v}`) said only that it was "verified through `pandoc -t latex`" --- true, and unfalsifiable-looking to a reader with no further detail.
An `adversarial-reviewer` subagent, dispatched to check the fix, ran its own counter-test: appended `\v0` `\v1` `\v{x}` to a document containing **no macro definitions**, ran it through `quarto pandoc -t latex`, and observed every token pass through unexpanded.
From that it concluded pandoc does not expand TeX macros in math mode at all, so the stated verification could not possibly have discriminated a working macro file from a broken one --- and filed the fix's claim as unsubstantiated.

The reasoning was valid.
The measurement was real.
Both were about the wrong case: pandoc's `latex_macros` extension expands a macro only when it is **defined in the same document**, which the reviewer's test document was not.
Testing an empty document to ask "does pandoc expand macros" is the null case, indistinguishable in outcome whether the extension works or the extension is entirely absent --- [`fail-fast`](../principles/fail-fast.md)'s denominator move again: a test whose passing and failing readings look identical has not tested anything.
Re-running with the precondition restored (a document that actually defines `\v`) produces the discriminator the claim needed:

| | `\v0` | `\v1` | `\v{x}` |
| --- | --- | --- | --- |
| no definitions present | `\v0` | `\v1` | `\v{x}` |
| `macros.qmd` before the fix | `\v0` | `\tilde{1}` | `\v{x}` |
| `macros.qmd` after the fix | `\tilde{0}` | `\tilde{1}` | `\tilde{x}` |

The reviewer had measured the top row and read it as the whole truth table.
The middle row is the bug's actual signature --- only `\v1` expands, because a delimited `\def\v1` survived as the last definition of `\v` --- and the bottom row is the fix.
Nothing in the reviewer's transcript was fabricated;
the precondition the original claim depended on was simply never in the reviewer's own test.

**Two things follow, and both are needed --- one about re-measuring a finding, one about where the fix belongs.**

First: a rebuttal is a claim like any other, so the *rebutter* re-measures before publishing it, not only the party being rebutted.
[`address-every-comment`](address-every-comment.md)'s Rebut disposition already lets an author push back on a reviewer's finding;
the mirror obligation belongs to the reviewer before the finding is filed --- confirm the counter-test actually carries the precondition the original claim relied on, not merely a test that superficially exercises the same mechanism.

Second: the fix is not to win the rebuttal in a PR comment where it dies with the thread.
The original message's vagueness --- "verified through `pandoc -t latex`", true and giving the reader nothing to check --- is what invited a plausible wrong finding in the first place.
Amending the commit message to carry the three-row table above did both jobs at once: it rebutted the finding, and it left the next reader (human or reviewer) unable to repeat the reviewer's mistake, because the null row sits right next to the two rows that discriminate.
A durable artifact that states its own discriminator is [`quotable-findings`](quotable-findings.md)'s standard turned around --- a claim that names the exact measurement that would falsify it is the one nobody can plausibly misread.

- **Do:** treat a reviewer's own counter-test as a claim requiring the same re-derivation any other claim does, whichever side of the finding you are on.
- **Do:** when rebutting a finding, name the precondition the original claim relied on and confirm the counter-test carried it.
- **Do:** write the discriminating measurement --- including the null case that shows what a non-discriminating test looks like --- into the durable artifact (commit message, PR body) rather than only into a comment thread.
- **Don't:** read "the reviewer ran a command" as equivalent to "the reviewer ran the command that could have shown the claim false" --- a command that cannot exhibit the failure mode has not tested the claim, however real its output is.
- **Don't:** rebut by re-asserting the original claim against the counter-test's bare output;
  that answers confidence with confidence and settles nothing --- name the specific precondition the counter-test dropped.
- **Don't:** leave a verification claim as a bare tool invocation ("verified through X") with no stated discriminator;
  that vagueness is what makes a plausible-but-wrong counter-finding possible in the first place.

(Measured 2026-09-09 on d-morrison/macros#87: the reviewer's counter-test and its null-case conclusion are the measured half;
the general rule that a rebuttal is itself a claim requiring re-derivation, and that the fix belongs in the durable artifact rather than a comment, is the inferred half, extending [`address-every-comment`](address-every-comment.md)'s Rebut disposition to the reviewer's own side of it.)

## A summary is another shape, and the auto-loaded copy is the one you read

[`fact-check-prose`](../writing/fact-check-prose.md)'s "any condensation
of a verified source is a fresh claim" already names the psychology, and
[`citations`](../writing/citations.md) already names why an unquoted
attribution launders --- a paraphrase reports the source's conclusion in
your voice, with the source's authority attached.
What neither covers is the *retrieval* asymmetry that decides which copy
you consult at all.

Run `scripts/check-context-closure.py` for the current set: `CLAUDE.md`
and the fragments it `@`-imports are always in context, while `memories/`
and the rest of `shared/` are not.
So a rule restated in an auto-loaded file is the copy you will read, and
often the only one, while its source is a file you must decide to open.
The summary is not merely available --- it is already there, and it names
the source, which answers "where does this come from?" convincingly
enough that nothing prompts the read.

Two consequences follow that the neighbouring fragments do not draw.
Compression fails in a direction: hedges, caveats, and disambiguating
steps go first, because those read as qualifying detail rather than as
the claim itself.
And quotation is not the boundary.
A *characterization* --- "`X.md` records that ..." --- is the looser form
and the one no phrase-grep can check, so the attributions most likely to
be unfaithful are exactly the ones an instrument cannot see.
`hooks/remind-brief-premises.py` detects that sentence shape already, but
only on `Agent`/`Task`/`SendMessage` payloads, so the reader-side case is
currently uninstrumented.

- **Do:** open the cited file before asserting what it says, including
  when the passage in front of you names it.
- **Do:** suspect a dropped hedge first when a summary and its source
  disagree.
- **Don't:** treat "the auto-loaded file says so" as having consulted the
  fragment it cites.
- **Don't:** read an unquoted attribution as checkable --- it is
  precisely the form that is not.

See [`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md),
"A summary read as its source, in the session that fixed the summary".

**The same substitution runs over your own transcript,
and there it corrupts a measurement rather than a citation.**
The section above concerns a summary of a corpus *file*.
A context-window summary of the *conversation* is the other copy that is already in front of you,
and the reply it condenses is not.

That matters most when the text is being used as a **specimen** rather than as a source.
Designing a matcher --- a hook's regex, a grep, a classifier's word list ---
against the summary's rendering of a reply is validating it against paraphrase.
The summary's wording is written to be representative, so it matches readily,
and the measurement comes back clean;
the real reply's wording is what the matcher will actually meet in production,
and it need not match at all.
Nothing distinguishes the two outcomes,
because both are "the pattern fired on the text I tested it against".

Measured 2026-09-02/03 while drafting the `no-unverified-approval-claim` Stop hook on `Morrison-Lab/ai-config` (branch `hook/no-unverified-approval-claim`,
unpushed at the time of writing, so there is no PR to cite):
the matcher was designed and validated against a context-window summary of the session's own reply.
The summary's phrasing matched and the reply's phrasing did not,
so the design read as validated by its own measurement while not firing on the one case that motivated it.

- **Do:** pull the verbatim text out of the raw transcript when a matcher is being fitted to it,
  per [`get-under-the-hood`](../principles/get-under-the-hood.md)'s raw-log practice.
- **Do:** treat "the pattern matched my test string" as a claim about the test string until you can say where that string came from.
- **Don't:** fit a matcher to a summary of the thing it must match ---
  a paraphrase is the one specimen guaranteed to be cooperative.

## A drift claim is relational, so one read cannot settle it

Every shape above is one substitution: you read A and made a claim about B.
This is the case where the claim is about **both at once** ---
A has drifted from B, the install is stale, the two copies have diverged ---
and a two-place claim needs two reads.
Read one, supply the other from memory, and the check feels complete,
because the section's own Do line is singular:
"state which artifact a claim rests on, then read *that* one".
A drift claim rests on two, and reading either one satisfies that sentence.

Two proxies stand in for the missing read, and neither one reads the source.
That is the shared defect, and it is not "metadata rather than content" --- the
second proxy is a content read, of the wrong side.

**An mtime.**
It records when a write happened, not whether contents lag a source.
A file copied once and never needing to change carries an old mtime
and current contents,
which is the same reading a genuinely stale file gives.
[`keep-checkouts-fresh`](keep-checkouts-fresh.md) already warns against mtime
in the opposite direction --- spotting *local edits*, where `git` resets mtimes
on checkout --- and the same metadata is uninformative in this direction
for a different reason.

**An absence.**
Finding that the installed copy lacks some function,
and asserting the source "has carried" it,
checks the consumer and infers a property of the source.
The absence is equally consistent with the source never having had it.

The falsifying question in "The test" above disposes of both:
ask what the artifact would read in the case you are worried about,
not only in the case you expect.
An mtime reads identically for a current copy and a stale one,
so it discriminates nothing.

Where the two artifacts are files, the deciding query is a direct comparison,
and it is cheap: `cmp` in a loop, or `diff -r`.
Where an instrument already owns the comparison, run the instrument instead ---
a hand-rolled diagnosis of a property an instrument already computes
is re-deriving a verdict the repo already owns.
(The worked example here was `scripts/check-install.py`,
which compared `~/.claude` against the checkout and reported `stale` by name;
it was removed along with the symlink install it verified,
so read it as a historical illustration ---
see [ai-config#2229](https://github.com/Morrison-Lab/ai-config/pull/2229).)

**A diagnosis that resolves an irritation deserves the deriving query
before it deserves an issue.**
The false half here --- the claim that the copy was already stale, as
distinct from the true claim that a copy cannot track later merges --- was
satisfying twice over:
it explained an incident that had already cost three blocked turns,
and it assigned the cause to infrastructure rather than to
an unmerged fix of my own.
Neither feeling is evidence, and both suppress the impulse to run the command.

- **Do:** name both artifacts a drift, staleness, or parity claim compares,
  and read both before asserting the difference.
- **Do:** run the direct content comparison --- `cmp`, `diff -r`,
  or the instrument that owns it --- rather than inferring drift from metadata.
- **Do:** slow down on a diagnosis that explains something annoying
  or points away from your own work, and derive it before filing it.
- **Don't:** read an mtime as evidence about contents;
  it is evidence about when a write occurred.
- **Don't:** infer what a source contains from what a copy of it lacks ---
  read the source.

See
[`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md),
"A stale install diagnosed from an mtime and an absence".

**The interpreter's own defaults are another adjacent artifact, and a failed reproduction is where they stand in for the code.**
The four shapes above all substitute one *file* or *run* for another.
This one substitutes the **environment** for the program, and it arrives
disguised as diligence: you were told a defect exists, you tried to see it
yourself rather than taking the claim on faith, and nothing happened.

"Could not reproduce" is then a claim about the code, and the evidence
supports only a claim about the machine.
The direction is what makes it dangerous.
A failed reproduction closes the question --- the reported defect goes away,
the reporter is quietly downgraded, and no further check fires, because
there is no longer anything to check.

The tell is a defect whose trigger is a **default** rather than an input:
a decoder's error handler, a locale, a shell option, a umask, a timezone.
Those are configured outside the program, so the program is identical on
both machines and only one of them can show the bug.

The remedy is to force the condition rather than to hope for it.
Ask what setting has to hold for the report to be true, read that setting,
and set it explicitly before concluding anything.

- **Do:** read the setting the reported defect depends on, and say what it
  was on the machine where the reproduction failed.
- **Do:** force the condition and re-run before writing "cannot reproduce".
- **Don't:** treat a clean run as evidence about the code when the trigger
  is an ambient default.
- **Don't:** let a failed reproduction retire a report; it retires one
  environment.

**The mirror of that, for the person who made the report: a reviewer's failed
reproduction is not a refutation, and the remedy is to SCOPE the claim rather
than retract it.**
The section above tells whoever failed to reproduce what they owe.
It does not say what the original claimant owes when someone competent runs
the case, carefully, and sees nothing --- and the pull there is strongly
toward retracting, because the reviewer has evidence and you have a memory.

Both can be right, and usually are.
An observation made on one machine is true on that machine; stating it without
a scope is what makes it false everywhere else.
So the disagreement is rarely about whether the thing happened.
It is about a qualifier nobody wrote down.

Two moves settle it, in order.
**Re-run your own original case**, since a remembered failure is not evidence
and may simply have been a mistake.
Then, if it reproduces, **reduce it to a form that removes the suspected
explanation** --- the reviewer's runs are a hypothesis about the cause, so
build a case their hypothesis cannot account for.

What lands in the corpus afterwards is the scoped claim plus the environment
both parties measured, and the reviewer's null result belongs in the entry
rather than being discarded by it.
A rule stated unconditionally, true where it was written and false where it is
read, is [`timestamp-volatile-claims`](../writing/timestamp-volatile-claims.md)'s
failure with an environment in place of a date.

- **Do:** re-run your original case before conceding anything.
- **Do:** reduce to a form that rules out the reviewer's proposed cause, then
  publish the scope both of you measured.
- **Do:** record the null result in the entry; it is what tells the next reader
  to test rather than trust.
- **Don't:** retract a reproducible finding because someone else could not see
  it --- that discards a true observation to resolve a missing qualifier.
- **Don't:** defend it unscoped either; unconditional is the actual defect.

(Measured 2026-08-22 on
[ai-config#1926](https://github.com/Morrison-Lab/ai-config/pull/1926).
A `CLAUDE.md` entry claimed a doubled backslash collapses inside a quoted
heredoc.
The reviewer ran three cases in a Linux CI runner, could not reproduce any of
it, and said so with its evidence.
Re-running the original case reproduced it immediately; reducing it to a `cat`
heredoc writing two lines to a file --- no interpreter anywhere --- showed a
typed `\\` and a typed `\` both landing as one backslash, which the
reviewer's Python-parsing explanation could not account for.
Platform: Windows 11 / MINGW64 through the Claude Code Bash tool.
The entry now leads with that scope and carries the reviewer's null result.)


**A sixth: the fact that a check ran, standing in for what the check found.**
The five above all substitute one artifact for another --- a file, a run, an
environment.
This one keeps the right artifact and reads the wrong property off it.
A guard asks "was a measurement taken?" when the rule it enforces asks "is the
value you stated the one that was measured?", and the two questions come apart
the moment a measurement is taken and then departed from.

It is the hardest of the six to see from the inside, because the guard is
**correct on every case anyone thought to test**.
A session with no measurement fires, a session quoting its measurement stays
silent, and both are the intended behavior --- so the test matrix is green and
the missing case is the one nobody wrote, since it requires imagining evidence
being present and ignored rather than absent.

The tell is an instrument whose state is a **position, a flag, or a count**
where the rule is about a **value**.
An index recording that a reading occurred cannot distinguish quoting it from
contradicting it.
Neither can a boolean "the linter ran", a count of checks executed, or a
timestamp proving a job started.

The remedy is to ask what the guard would have to compare in order to be
wrong.
If the answer names a value the guard never captures, the guard is measuring
its own execution rather than the property.

- **Do:** capture the value a check produces, not only the fact that it ran,
  wherever the rule is stated in terms of that value.
- **Do:** write the test where the evidence is present and the claim departs
  from it, which is the case a green matrix is least likely to contain.
- **Don't:** read a passing guard as covering a rule whose subject it never
  reads.
- **Don't:** treat "the check is registered and did not fire" as evidence the
  claim was sound.

(Measured 2026-08-21, on this repo's own
[`hooks/no-unmeasured-clock-claim.py`](../../hooks/no-unmeasured-clock-claim.py).
It exists to catch a stated Pacific time nobody measured, was registered and
running, and stayed silent while a recap claimed `15:22 PDT` against an
injected reading of `14:48:23 PDT` --- 34 minutes ahead, and in the future at
the moment it was written.
Its `scan()` recorded each reading as a line *index* and never captured the
timestamp, so `main()` compared positions.
Sixteen existing tests passed, none of them a departing claim.
Tracked as [ai-config#1848](https://github.com/Morrison-Lab/ai-config/issues/1848).)

(Measured 2026-08-21 on
[ai-config#1784](https://github.com/Morrison-Lab/ai-config/pull/1784).
A review reported that `run_cli`'s `sys.stdin.read()` could raise
`UnicodeDecodeError` on a non-UTF-8 diff, the third of three reads in that
function and the one a prior round's `errors="replace"` fix had missed.
Piping non-UTF-8 bytes to the unfixed hook exited 0 and reported no
findings.
`sys.stdin.errors` in that container is `surrogateescape`, not `strict`, and
stayed `surrogateescape` under `PYTHONUTF8=0` with an explicit
`en_US.UTF-8` locale, so the crash could not occur there at all.
`PYTHONIOENCODING=utf-8:strict` reproduced it on the first try ---
`'utf-8' codec can't decode byte 0xff in position 47` --- and the fixed hook
exited 0 on the same input.
Stopping at the first attempt would have reported the finding
unreproducible, a true statement about the container and a false one about
the code.)

**Documentation for a capability describes a SURFACE, and your code may reach
that capability through a different one.**

The shapes above substitute a cached copy for an origin, a checkout for a run,
half a mechanism for the whole, a neighbour for the target, a different
endpoint for the same-sounding metric.
This is another: the documentation is correct, your reading of it is correct,
every quotation checks out --- and it describes the feature as reached through
a surface your code does not use.

The tell is a capability documented for a **settings file, an API, or a
library call** while your code exercises it through a **CLI flag, a wrapper, or
another entry point**.
That the second surface accepts the first's syntax is a separate proposition,
and documentation routinely leaves it unstated because it is obvious to whoever
implemented both.

It matters more than an ordinary unchecked assumption because of how it fails.
An unparsed rule is usually **dropped rather than rejected**, so the change is
a no-op that looks shipped: nothing errors, the diff reads correctly, review
passes, and the concern is retired.

Verifying it needs a **negative control**, and that is the part most likely to
be skipped, because running the thing and seeing no complaint feels like a
test.
It is not: a surface that silently ignores every unknown rule is quiet for the
same reason a working one is.
Pair the real input with one the documentation says must be refused.
If the refusal is announced and yours is not, the silence means something.

- **Do:** name the surface your code actually uses, and verify against it
  rather than against the one the docs describe.
- **Do:** run a known-bad input alongside, so silence is evidence rather than
  absence.
- **Don't:** read a correct quotation of a syntax as evidence that your caller
  parses it.
- **Don't:** ship a mechanism whose failure mode is a silent no-op on the
  strength of documentation alone.

(Measured 2026-08-21 on [gha#550](https://github.com/Morrison-Lab/gha/pull/550).
`code.claude.com/docs/en/permissions` documents `Tool(param:value)` deny rules;
every quotation in the diff was checked against the live page and all were
accurate.
The action passes its rules through `--disallowedTools`, which the page never
mentions.
On Claude Code 2.1.238 both rules were accepted with no warning --- and the
control is what made that informative, since `Bash(command:rm *)`, which the
same page says is ignored, answered
`targets command as a raw string and will not match` on that same CLI.)

**A seventh: a future state for the present one.**
The six above all substitute one artifact, environment, or property for
another that exists *now*.
This one substitutes a state of the **same** artifact at a *different time* ---
what the tree will contain once some other branch lands.

The figure is not wrong when written, which is what makes it durable.
It is wrong when read, because it is committed to a file whose own instrument
reports the present number a few lines below, so a reader compares the two
directly and the prose loses.

- **Do:** derive any number you commit against the state of the branch you are
  committing it to, not the state you expect after some other PR merges.
- **Do:** re-run the file's own checker and quote what it prints, when the file
  ships one.
- **Don't:** compute against a listing, count, or size that only exists on
  another branch.

(Measured 2026-08-21 on
[ai-config#1853](https://github.com/Morrison-Lab/ai-config/pull/1853).
A source comment claimed "about 21 skills of runway", derived from a listing of
8,070 --- the value once
[#1849](https://github.com/Morrison-Lab/ai-config/pull/1849)'s entry lands ---
in a file whose own validator printed `7998/9000` on that branch.
The reviewer re-derived it as `1002 // 43 = 23` and the figure was corrected in
`653dc9df`.)

**A PR body is prose no check reads, so it goes stale silently while the diff
moves underneath it.**

A near relative, and the one that bites after the work is finished rather than
during it.
Every claim a PR body makes about its own diff --- what is tested, what is
unchanged, what was verified --- was true when written and is unmaintained
thereafter.
No CI job compares the two, and a reviewer reads the body as context rather
than as a claim to check, so it is the one artifact in the review loop with no
detector at all.

Re-read the body at the moment you would report the PR ready, and treat each
of its factual claims as you would a line of the diff.

- **Do:** re-read and correct the body before reporting a PR ready, after the
  last push rather than before it.
- **Don't:** leave a body asserting a verification the diff has since outgrown.

**Re-reading each round does not fix it, and which half decays tells you what to write instead.**
The remedy above is a discipline, and a discipline applied every round on a PR with many rounds still loses, because the body has two kinds of content and they rot at different rates.

A body describing the **mechanism** is stale the moment the mechanism changes, which on a PR under revision is every round.
A body organized around the change's **invariants** --- what must remain true however it is built --- plus an **append-only history** survives, because neither is invalidated by a rewrite.

The residual after that restructuring is **counts**: a test total, a round number, a file tally.
They read as settled facts rather than as claims, so they escape the re-read that catches everything else, and they are wrong within one push.
The fix is not to check them harder but to give the deriving command beside each one, so a stale figure is repairable by running a line rather than by remembering what it was.

Two things follow.
Keep the mechanism in the module docstring or the code comment, which is version-controlled with the thing it describes and therefore cannot drift from it.
And prefer a figure a reader can re-derive to one you assert: `wc -l` beside a count costs nothing and converts an assertion into an instrument.

- **Do:** organize a body around invariants and an append-only history, not around how the change currently works.
- **Do:** print the deriving command beside every count.
- **Do:** put mechanism detail where the code lives, so it is versioned with what it describes.
- **Don't:** rely on re-reading each round --- it is necessary and it does not scale past a few rounds.
- **Don't:** treat a number as the safe part of a body.
  It is the part that reads as checked and is not.

(Same day, both PRs.
gha#550's body still read "the fixture suite unchanged and passing" after a
push that added two fixtures, and "there is no offline test" after a push that
added one.
ai-config#1833's body still carried, verbatim, all three prose defects that
three review rounds had just corrected in the file it was describing.)

**A suite's own summary line is the verdict; a count you derive from its log is
an adjacent artifact.**

The narrowest form of this substitution, and the one that survives review
because the derived number is *nearly* right.

A test suite typically reports two things: one line per failing case, and a
final line saying how many failed.
Counting the first with `grep -c` looks like reading the result and is not ---
the summary line is usually formatted like a finding (`::error::N of 14 cases
...` is itself an `::error::` line), so it counts as a case and every figure
comes out inflated by exactly one.

That constant offset is what makes it durable.
A wildly wrong number invites a second look; `5` where the truth is `4` reads
as plausible, stays plausible when re-derived the same way, and is not
checkable against anything else in the report.
Nothing in the suite's output contradicts it, because the suite never made the
claim --- you did.

It bites hardest on **mutation counts**, where the number is the entire
evidence for "this assertion is load-bearing".
A corpus that asks for counts to be confirmed by mutation rather than assumed
gets a count that was measured, from the wrong line.

- **Do:** read the suite's own summary line, or its exit status, as the
  verdict.
- **Do:** state which line you read it from when a count reaches prose someone
  will rely on.
- **Don't:** `grep -c` a suite's findings and call the result its failure
  count.
- **Don't:** treat a number as verified because you ran something to get it ---
  ask which artifact answered, and whether it was the one making the claim.

(Measured 2026-08-27 on [gha#687](https://github.com/Morrison-Lab/gha/pull/687).
Three mutation counts were documented as 5/8/4 in `gha`'s `CLAUDE.md`, read via
`grep -c '^::error::'`.
The reviewer independently reproduced 4/7/3 and flagged the mismatch at reduced
confidence, guessing the mutation implementations might differ.
They did not --- the suite's summary line was being counted as a fourteenth
case, which is why all three were off by the same amount.
Note the detector here was a second party re-running the measurement, not a
check: nothing in CI could have caught it.)

**Reading the right line does not settle what the line's number counts, and a wrong reading there survives every faithful copy that follows.**

The section above corrects a *derivation* error: grepping the wrong line instead of the summary.
This corrects an error in the summary line itself --- correctly read, correctly quoted, and still wrong, because its wording names one population while its number counts another.

A checker reporting "N display equation(s) missing a line break" is ambiguous between two claims: N *equations* are affected, or N *findings* occurred (an equation can be missing a break on either side, so one equation can produce two findings).
The tool's own sentence grammatically asserts the first while its counter tracks the second, and nothing about reading that line "correctly" resolves which one a downstream reader inherits.
Everyone who then repeats the number is being faithful to the source, which is exactly why the error survives: a copy cannot be more careful than the thing it copies, and each faithful copy looks like independent confirmation without being one.

- **Do:** before repeating a checker's summary count, read the noun in its own sentence and ask whether the counter beside it tracks that noun or a different one (a finding versus the distinct items it can occur on).
- **Do:** fix the wording at the tool once an ambiguity like this is found, not only the prose that repeated it --- the next reader inherits the tool's sentence, not your correction of it.
- **Don't:** treat "I read the suite's own summary line, not a grep-derived count" as sufficient;
  that rules out one substitution and not the one where the line itself names the wrong population.
- **Don't:** read several faithful copies of one number as corroboration --- they share a single source and inherit its error together.

(Measured 2026-09-09 on [Morrison-Lab/ai-config#3426](https://github.com/Morrison-Lab/ai-config/pull/3426): a per-file checker counting missing line breaks around display equations printed a summary of the shape "N display equation(s) missing a line break" where N counted findings (an equation missing breaks on both sides counts twice).
The true equation count was smaller.
The findings-count, under the equation-shaped sentence, was copied verbatim into a status report, a subagent brief, a filed issue, the corpus fragment recording the incident, a companion memory file, and the PR body describing all of it --- six artifacts, each a faithful transcription of the one before it, all wrong in the same direction.
A review round caught it at the sixth.
The fix reached the tool itself, which now prints both counts under distinct labels, rather than only correcting the prose that had repeated the conflated one.)

**An eighth: what a change TRANSFORMS, standing in for what it CONCLUDES.**

The shapes above substitute one artifact, environment, or property for another, and this one substitutes a property too --- so what distinguishes it is not *what* gets swapped but *where* the swap happens.
It happens inside a verification built specifically to catch the error it then misses, so the substitution arrives wearing the clothes of a parity proof, and the instrument's own clean number is what conceals it.

A change to a fail-closed instrument widens what it blanks before scanning, and the proof asks: does every character the new revision blanks and the old one did not lie inside a code span the change is meant to blank?
That question cannot come back non-zero for any implementation of that shape.
The extra-blanked set *is* the span set, so the metric restates the change's own definition and reports the restatement as evidence.
Its zero was truthful and worthless: two real fail-opens were live at the time, and both arose in the passes that run *after* the blanking, about which a metric over the blanking says nothing whatever.

The tell is that the metric's inputs are the two revisions' outputs at an **intermediate stage**, rather than the two revisions' **verdicts**.
An instrument exists to conclude something, and a change to it is safe when the conclusions match --- not when an intermediate buffer differs in the shape the change predicted.
So diff the **acceptance sets**: which bodies each revision calls clean, which it calls not clean, and which moved.
Let the transformation be whatever it needs to be.

Distinguish it from [`fail-fast`](../principles/fail-fast.md)'s fifth cause of a vacuous zero, which it superficially resembles.
There the check examines the right quantity and the *subject* absorbs its own failures through a designed fallback, so the remedy is to measure the fallback bucket.
Here the subject is fine and the check asks a question with only one possible answer, so no bucket exists to measure --- the metric has to be replaced rather than instrumented.

This is the general form of the trap [`mistake-patterns.md`](../../memories/mistake-patterns.md) Pattern 15 warns about.
Pattern 15 says to prove parity before widening a fail-closed exemption;
what it does not say, and what this section adds, is that a parity proof can be constructed over the wrong quantity and then cannot fail.
A proof that cannot fail is not a weak proof, it is an absent one wearing a number.

- **Do:** define a parity metric over what the two revisions *decide*, and name the decision function in the metric's own docstring.
- **Do:** ask of any verification metric what result would make you abandon the change --- and treat "none" as the finding.
- **Do:** report the acceptance-set delta in both directions, since a change that only narrows is still a change.
- **Don't:** measure the transformation a change performs and call the agreement a parity proof;
  that measures the diff against itself.
- **Don't:** read a zero as reassurance without checking that a non-zero was reachable for some implementation of the same shape.

(Measured 2026-08-28 on [ai-config#2515](https://github.com/Morrison-Lab/ai-config/pull/2515), fixing [#2449](https://github.com/Morrison-Lab/ai-config/issues/2449).
The first parity instrument compared what `strip_cited_finding_vocab` blanked across two revisions and reported 0 extra characters outside a code span.
The replacement, `scripts/check-verdict-scan-parity.py`, diffs what the two revisions conclude instead, triages each widening by offset, and runs a negative control first.
Only half of its discrimination claim is reproducible **from `main`**, and the entry says which half and how to reach the other.
Running it against the shipped design reports 0, which any reader can re-run.
The 3,924 / 108 / 270 / non-zero off-axis figures for the four rejected designs were recorded on that branch before #2515 was **squash-merged** as `07847b9`, so they are not reproducible from `main` --- which is the artifact a reader has.
They are not lost, though, and the difference matters: GitHub retains `refs/pull/<N>/head`, so `git fetch origin 'refs/pull/2515/head:refs/remotes/pr/2515'` restores the branch and all four designs (`c7ff646`, `4f9d3fc`, `68a14b9`, `a3251bf`) with it.
Name that route whenever you mark a figure unreproducible, since "unreachable" and "not on the default branch" are different claims and only the second is true here --- the first was asserted in this very section and refuted by one `git ls-remote`.)

**A ninth: a LOSSY CONVERSION of a document, standing in for the document.**

The shapes above substitute an artifact that is stale, partial, or adjacent.
This one substitutes an artifact that is current and complete for its own purpose, and **lossy by design**.
A conversion drops what its target format cannot carry,
so its omissions are the reason the tool is useful rather than a defect in it.

The tell is that the derived view answers the question you asked and cannot answer the question you meant.
`pandoc -t markdown` on a `.docx` reports the text a reader sees.
Asked whether a link is present, it can report nothing about a URL stored as a Word HYPERLINK field code,
and **more than one thing decides whether it does**.
All three arms below were measured on pandoc 3.1.3 under `--track-changes=accept`.
A `fldChar` HYPERLINK field in ordinary body text whose `instrText` sits in a single run
converts to a markdown link with its URL intact.
Splitting that same `instrText` across two runs, at the space before the quoted URL,
still emits a link and empties its target: `[anchor]()`.
The identical single-run field nested inside a `<w:ins>` tracked insertion drops the link entirely,
leaving bare text.
So `<w:ins>` is **sufficient** to lose the URL without being the determinant,
since the split-run arm loses it with no `<w:ins>` anywhere in the document.
The split-run arm is not a corner case either:
Word and Zotero routinely split `instrText` across runs,
which is the premise of [`memories/office-open-xml.md`](../../memories/office-open-xml.md)'s sibling entry
and the reason its `merge_runs.py` exists at all.
Say "single-run" rather than "contiguous", which this entry reached for first and which decides nothing:
a field split across *paragraphs*, its `instrText` still in one run, converts with the URL intact.
The negative control is what identifies each mechanism.
Without it the omission reads as "pandoc does not carry field-code links at all",
which is false and was written down that way once before the control was run.
A listing of `word/_rels/document.xml.rels` is no better:
it enumerates one of the two ways Word stores a hyperlink and is silent about the other.
Two independent readings then agree, and the agreement is a property of what both drop.

So before concluding a document does not contain something,
search the **stored form**: grep the source XML, the raw bytes, the file the application actually writes.
The converted view is evidence about what a reader sees, which is a different claim.

- **Do:** name which representation a negative is about --- rendered text, or stored source --- before reporting it.
- **Do:** grep the stored form (`word/document.xml`, the raw file) when the claim is that something is absent.
- **Do:** run a negative control on the conversion before naming a mechanism for what it dropped;
  an omission with no control behind it is a guess wearing a measurement.
- **Don't:** promote one mechanism to *the* determinant once a control has shown it sufficient;
  sufficient and necessary are different findings, and only a further arm separates them.
- **Don't:** read two derived views agreeing as corroboration when both drop the same class of content;
  that is [`grep-is-not-coverage`](grep-is-not-coverage.md)'s guaranteed-either-way null in a new surface.
- **Don't:** treat "lossy" as "stale" --- refetching a conversion returns the same omissions.

**Recurrence, 2026-09-09, on the same manuscript and the same link.**
Asked whether the Shiny-app link had survived an edit, a sweep collected
`w:hyperlink` elements and visible text, found neither, and reported to the
user that the manuscript advertised the app twice while linking to it
nowhere --- recommending the link be restored.
Both readings are derived views that drop a field code, which is the
agreement-by-shared-omission this section already names, and the Do-list
above already prescribes the fix: grep the stored form when the claim is
that something is absent.
The rule existed, named this exact link, and did not fire.

A second instance the same day generalizes it past hyperlinks.
Asked whether a script `L` was used anywhere, a search over Unicode
math-alphanumeric codepoints and over `m:sty` returned nothing, and the
report was that no script letter existed in any of the three files.
Word stores that styling as `m:rPr/m:scr val="script"`, a third
representation neither query touched, and the supplement had been using it
for the likelihood symbol all along.
So the failure is not specific to hyperlinks or to pandoc: any absence
claim about a `.docx` is a claim about the representations searched, and
OOXML stores most things more than one way.

- **Do:** enumerate which representations a `.docx` absence claim covers,
  and say so in the claim.
- **Don't:** report an absence from one attribute or element name when the
  format has a second spelling for the same thing.

(Measured 2026-09-01 while adding tracked changes and comments to three `.docx` files for a journal resubmission.
A manuscript's Shiny-app link was absent from the rels listing and absent from pandoc's markdown output,
and the conclusion that it had been deleted was written into a draft review finding.
It was present as a `fldChar` HYPERLINK field code, found by grepping `word/document.xml` for the URL.
In that manuscript the field carries exactly one `instrText` element and *is* wrapped in `<w:ins>`,
so the tracked insertion really is the operative cause there;
the split-run mechanism generalizes the entry rather than retracting its case.
The manuscript is private, but the pandoc behaviour needs no manuscript:
build a minimal `.docx` carrying the same HYPERLINK field three times ---
once in plain body text with its `instrText` in a single run,
once in plain body text with that `instrText` split across two runs at the space before the quoted URL,
and once single-run but wrapped in `<w:ins>` --- then convert with `--track-changes=accept`.
The three arms emit the URL, an empty target, and bare text respectively.
Read a `w:ins` grep over `word/document.xml` carefully when building the split arm:
`w:instrText` contains the substring `w:ins`, so a plain `grep -c` reports a match in a document that has none.
[`memories/office-open-xml.md`](../../memories/office-open-xml.md) carries the docx-specific mechanics,
including the two-pandoc-diff verification for a redlined document.)

**A tenth: a diff's changed lines, standing in for the file they changed.**

The shapes above substitute one file, run, or environment for another.
This one keeps the right file and reads only the fraction of it a diff highlighted.
A diff marks what changed;
it says nothing about what the surrounding text, including context the same diff adds, now means.
Grepping the diff for a keyword returns exactly the lines matching that keyword and nothing about the lines around them --- and those surrounding lines are the file's meaning as often as not, because a change is scoped by its neighbours.

The tell is a conclusion drawn from **removed or added lines alone**, when the same diff's own added context sits one hunk away and narrows what the removal actually licenses.

Case: `Morrison-Lab/gha#811` deleted three lines from `examples/quarto-publish.yml` --- `concurrency:` / `group: gh-pages` / `cancel-in-progress: false` --- with no replacement.
Reading that removal through `gh pr diff | grep` for `concurrency`/`group` lines supports one conclusion: the PR tells consumers to delete the block outright.
The same diff, earlier in the same file, adds a six-line NOTE the grep pattern never matched: "Do NOT declare a top-level `concurrency:` block naming `gh-pages`" in the caller workflow.
That sentence is scoped to the group's *name*, not to the presence of a block.
A caller-level group with a different name --- `website-publish-${{ github.ref }}`, the literal name `gha#667` gave a different workflow's group in the same repo, or `quarto-publish-${{ github.ref }}`, the name the consumer PR that motivated this fix later merged --- names something else and is not what the note forbids.
The removal and the addition are two edits inside one diff, and only reading the file whole, rather than the diff's hunks in isolation, shows that the second scopes the first.

**This survives the quote-the-passage check, which is what makes it worth recording rather than dismissing as ordinary carelessness.**
[`quotable-findings`](quotable-findings.md) requires a finding to quote the exact passage it is about, on the theory that a quotable finding is a checked one.
Quoting the three removed lines satisfies that rule to the letter: the passage exists, the quote is exact, the mechanical filter passes clean.
What the filter cannot check is whether the *file*, read whole, means what the quoted fragment suggests once its own neighbouring lines are included.
The check that would have caught this reads "open the file the passage lives in," not "quote the passage" --- a stricter requirement than [`quotable-findings`](quotable-findings.md) states, and this is the shape that shows the gap between them.

**Before writing a retraction, check whether the head moved.**
A wrong claim about a file can be wrong for two different reasons, and they produce different retractions.
The commit could have changed since the claim was written, in which case the honest statement is "this changed" and no misreading occurred.
Or the commit could be exactly the one that was read, in which case the honest statement is "I misread it" --- and conflating the two either lets a real misreading hide behind an invented edit, or accuses a PR of moving when it did not.
Settle it before writing either sentence: fetch the specific commit SHA the claim was written against (`gh api repos/<owner>/<repo>/contents/<path>?ref=<sha>`) and confirm it is unchanged, rather than assuming from the PR's current state.
In the case above, the branch head had in fact already moved by the time of both flagged comments --- a separate commit (`e34e03d5`) landed less than a minute before the first of them, fixing an unrelated review round --- but a direct compare (`gh api repos/<owner>/<repo>/compare/<cited-sha>...<later-sha>`) shows that commit never touched `examples/quarto-publish.yml`.
So the file the claim was about was genuinely unchanged, and the retraction's "I misread it" holds;
but the retraction described the cited commit as the one its comments were written against without checking the branch's own commit history, which shows a different commit was already the head by then.
A retraction is a claim like any other, and this one needed the same check.

- **Do:** read the file a diff's hunk lives in, not only the hunk, before concluding what a removal licenses or forbids.
- **Do:** treat added context in the *same* diff as evidence about scope, even when it sits in a different hunk than the lines a grep matched.
- **Do:** fetch the exact commit a wrong claim was written against before deciding whether to retract it as "I misread" or "this changed."
- **Don't:** treat a clean pass of [`quotable-findings`](quotable-findings.md)'s quote-the-passage check as evidence the file was read;
  it only proves the passage exists.
- **Don't:** infer a rule's scope from the lines a diff removed when the same diff also adds prose stating the scope.

(Measured 2026-09-02 on [`Morrison-Lab/gha#811`](https://github.com/Morrison-Lab/gha/pull/811).
Two comments on the PR, both derived from `gh pr diff | grep`-ing the changed `concurrency`/`group`/`cancel-in-progress` lines, argued that the deletion was the wrong fix and that a rename (as `gha#667` had already used elsewhere in the same repo) should have been kept instead;
a later comment restated the same position once a consumer PR merged its own renamed group, phrasing it as the stub telling consumers to delete the block while the consumer had shipped a rename.
The stub's own added NOTE at the cited commit (`856b8702`) read "Do NOT declare a top-level `concurrency:` block naming `gh-pages`," which scopes the prohibition to the group's name and forbids nothing about a renamed group.
Verified with `gh api "repos/Morrison-Lab/gha/contents/examples/quarto-publish.yml?ref=856b8702" --jq '.content' | base64 -d`.
The claim was retracted in the PR thread;
the retraction described `856b8702` as the commit its comments were written against, which a separate check against `gh api repos/Morrison-Lab/gha/pulls/811/commits` and `.../compare/856b8702...e34e03d5` did not confirm, though the file itself was confirmed unchanged.)

**An eleventh: an issue's OPEN state,
standing in for the behaviour it describes.**

The seventh shape above substitutes a *future* state for the present one.
This one substitutes a **stale past** one,
and it arrives with a citation attached,
which is what makes it the more persuasive of the two.
An open issue is a durable, linkable,
timestamped artifact that describes a defect precisely.
Everything about it reads as evidence.
What it actually records is that nobody has closed it,
and closing is a bookkeeping act performed by a person,
so the gap between "the defect exists" and "the issue is open" is exactly the set of fixes that landed without their issue being closed ---
which in a fast-moving repo is a large set.

The asymmetry runs against you.
A closed issue over-claims in the safe direction: you go and check.
An open issue under-claims in the dangerous one:
it confirms the belief you already had, from a source you can cite,
so nothing prompts the read.

The falsifying question from "The test" above disposes of it in one step:
*could this issue be open while the behaviour it describes is fixed?*
It always could.
So the issue can never settle the question, and only the code can.

- **Do:** read the code (or run the test) before asserting current behaviour,
  and cite the file and line rather than the issue.
- **Do:** cite the issue for the *history* --- that this was once broken,
  and is tracked --- which is the claim it can actually support.
- **Do:** check whether the fix landed and the issue simply was not closed,
  and close it (or say so) when it did.
- **Don't:** treat an open issue as a live measurement;
  its state is bookkeeping, not behaviour.

(Measured 2026-09-02, and re-checked 2026-09-03.
`hooks/flag-background-review-dispatch.py` ---
authored at commit `9009e787a` on the local `ai-config` branch `hook/flag-background-review`,
which `git ls-remote --heads origin hook/flag-background-review` confirms is unpushed,
so neither the commit nor the file is reachable from any clone but the author's ---
carried a docstring asserting that [`no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py) does not register verdicts arriving via background task notifications,
citing [ai-config#2483](https://github.com/Morrison-Lab/ai-config/issues/2483), which was ---
and as of 2026-09-03 still is --- open.
The fix had landed on 2026-09-01 in [#2820](https://github.com/Morrison-Lab/ai-config/pull/2820) (`0d78e04c`),
whose `is_task_notification` branch sits at `hooks/no-push-without-self-review.py:1400-1417` and is covered by a passing test.
`git log -L 1400,1417:hooks/no-push-without-self-review.py` names that commit in one command;
no such command was run, because an open issue looked like the answer.
The docstring was still uncorrected on that branch as this was written.)

**A twelfth: a pull request's check-run names, standing in for the branch's own workflow definitions.**

The seventh and eleventh shapes above are both substitutions across time, and the eleventh is unaware too, so what is new here is that **the artifact offers nothing to date**.
The seventh reads a future state for the present one knowingly, because the future state is the one being worked toward.
The eleventh reads a stale past one from a citable artifact: an issue carries a number and a timestamp, so its staleness is checkable by anyone who thinks to check, and what defeats it is that nothing prompts the read.
A check-run **name** carries neither.
`Spellcheck` is the same ten characters whenever it was produced, so there is no field to inspect and no version to compare --- the thing you would have to date is not in what you read.
A pull request's check runs are current, complete, and correct.
What they describe is the workflow definitions **in force when each run executed** --- resolved from the pushed commit for a `push` run, and from the head-into-base merge for a `pull_request` one.
Neither of those is the default branch as of now, and the name carries no trace of which moment or which resolution produced it.

The tell is a claim of the form "this repository emits X", derived from observing X somewhere.
A check run is produced by a workflow file, and a workflow file is versioned like any other, so a rename, a job restructuring, or a migration from inline jobs to a called reusable workflow changes every context string the repository publishes from that moment on.
A check run therefore records the definitions **in force when that run executed**, and its name carries no trace of when that was.

Two facts about a run decide which definitions it used, and a check-run name shows neither.
**When** it executed, and **which ref** it resolved the workflow file from.
For a `push` run that ref is the pushed commit;
for a `pull_request` run it is the merge of the head into the base, so a head that edits the workflow file overrides the base's copy while a head that does not simply gets whatever the base carries at that moment.
That second case is worth stating plainly, because the intuitive rule --- an old head publishes old names --- is false: an untouched workflow file follows the base, so a pull request opened long before a migration publishes the *new* names on its next run.

The staleness that does bite is therefore temporal rather than positional.
A name observed at time T is a fact about time T, and any later merge to the default branch retires it without touching the pull request you read.
Merged-ness is what conceals that.
A merged pull request feels like it *became* the branch, and in the ordinary case it did;
what it did not become is the branch as of any later moment, and nothing about a merged status says which moment you are reading.

The consequence for a required status check is unusually expensive, because it fails in the direction nothing reports.
A required context naming a check that no workflow emits does not error, does not turn red, and does not appear in any run.
It sits as `Expected`, and the only diagnosis is noticing that a check listed as required never appears at all.
How far that spreads is a question about runs rather than about settings.
An open pull request keeps whatever check runs it already has, so one that last ran before the workflow change still shows the old names and still looks satisfied.
Its next **push** re-resolves the workflow file through the current base and publishes the new names, after which the required context is unreportable on that pull request.
A *re-run* does not do this: GitHub re-runs reuse the original event's `GITHUB_SHA` and `GITHUB_REF`, so re-running a pre-change run republishes the old names and looks like evidence that nothing changed.
So the requirement is retired one pull request at a time, as each one is next pushed to, and nothing about that transition is announced.
Date the check runs you are reading before describing the blast radius;
a rollup showing a required context green may be showing a week-old run.

The authoritative artifact is the default branch's own workflow definitions, confirmed against a run **of that branch**:

```bash
gh api "repos/<o>/<r>/contents/.github/workflows?ref=<default-branch>" --jq '.[].name'
gh api "repos/<o>/<r>/contents/.github/workflows/<file>?ref=<default-branch>" --jq .content | base64 -d
gh api "repos/<o>/<r>/actions/runs/<run-id-on-that-branch>/jobs" --jq '.jobs[].name'
```

Read the definition rather than only the run, because a run answers only for the workflows that happened to trigger on that push.
A caller job `check:` invoking `uses: <org>/gha/.github/workflows/spellcheck.yml@v2`, in a workflow whose own `name:` is `Spellcheck`, publishes `check / spellcheck`.
The workflow-level `name:` does not appear at all.
What appears on each side of the slash is a **job** display name --- the caller job's, then the called workflow's inner job's --- and a job's display name is its `name:` where one is set and its key otherwise.
So the string is derivable from the file, and reading it off any observed check run is derivable from the wrong file.

This is also [`run-ums-proactively`](run-ums-proactively.md)'s false-*state*-claim case in its purest form.
No belief about reusable workflows was ever held and then corrected;
the wrong thing was simply looked up, so the reusable lesson is the query rather than the value.

One clause on the corroborating run, because the obvious reading of it is unsatisfiable.
The workflow definition on the default branch is the authority;
the run is corroboration that the definition composes the string you think it does.
A workflow triggered only by `pull_request` produces no run at all on an ordinary push to the default branch, so for that class the corroborating run is usually a pull-request run.
The exception is worth taking when it exists: a pull request opened *from* the default branch into some other base carries that branch's copy of the file on its head side, so its run reads what you want unless the base has diverged on that same file.
Failing that, a pull-request run resolved the workflow file from the merge of its head into its base, so it corroborates the **default branch's** copy only when two things hold together.
Its base is the default branch --- `gh pr view <N> --json baseRefName` --- since a stacked pull request or one targeting a release branch resolves the file from that other base instead, and would corroborate a different branch's copy while looking identical.
And its head does not touch that one file --- `gh pr diff <N> --name-only`, matching the single path rather than the `.github/workflows/` directory, since a head editing some other workflow is irrelevant.
The first condition is the one that goes unstated, and omitting it reinstates this shape's own substitution by way of its remedy.
Prefer a recent run, since an older one may predate the definition you just read.

- **Do:** derive a required-context string from the default branch's workflow definitions, and use a run only to corroborate how those definitions compose.
- **Do:** date every check-run observation by the commit its run executed, and say what has landed on the default branch since.
- **Don't:** accept a pull-request run as corroboration without reading its `baseRefName` --- a head that leaves the workflow file alone is necessary and not sufficient.
- **Don't:** read check names off a pull request, however recent, and generalize them to the repository.
- **Don't:** treat a pull request having merged as evidence its check names still describe the branch --- they described it at one instant, and a later merge can retire them without touching that pull request at all.

See [`verify-the-right-artifact.cases.md`](verify-the-right-artifact.cases.md), "A merged pull request's check names written into a live ruleset".

## A PR's `MERGED` status is another shape, and it is not corroboration of content

The sections above each name a claim about a PR that outlives the moment it was true.
This one is the claim made *at* the merge itself: that a PR reading `MERGED` is evidence your reviewed, verdict-clean diff reached the default branch.

It is not, for an ordinary reason that has nothing to do with the merge going wrong.
A PR branch can be merged while carrying a stale head --- another session, an `@claude` auto-sync, a rebase gone half-finished --- so the commit that lands on `main` is not the commit whose review you read.
Nothing about the merge fails: CI is green, the merge commit exists, GitHub reports success, and the PR page shows `MERGED` exactly as it would for a clean landing.
The status is real; it answers "did a merge happen", not "did my content land".
Confusing the two is the same substitution [`The four shapes`](#the-four-shapes) names elsewhere: the adjacent artifact (the PR's own state field) stands in for the one the claim is actually about (the tree at `origin/<default-branch>`).

The detector is cheap and belongs right after every merge you drive, not only when something looks wrong:

```bash
git merge-base --is-ancestor <your-last-pushed-sha> origin/<default-branch> && echo ok
git grep -c '<distinctive symbol from your diff>' origin/<default-branch> -- <path>
```

The ancestry check answers whether your commit is even in the merged history at all;
the content grep answers the sharper question, since a squash merge can be an ancestor-check false negative (the SHA changes on squash) while still needing the grep to confirm the actual lines survived.
Pick a symbol distinctive enough that a match means your specific change, not a coincidentally similar one nearby.

Recovery is not "push the stale branch again."
A branch that has drifted this far shows the merged base's *own* subsequent work as deletions when diffed against it, so reusing it re-proposes reverting content that was never yours to touch.
Cut a fresh branch off the current default branch and re-apply just the lost pieces instead.

- **Do:** after driving a PR to merge, grep `origin/<default-branch>` for a distinctive symbol from your diff and confirm your last pushed SHA is an ancestor of it.
- **Do:** cut a fresh branch off the current default branch to recover lost content, rather than reusing a branch that has drifted behind it.
- **Don't:** read `MERGED` as proof your content landed --- it is proof *a* merge happened, which is a claim about the PR's state field, not about the tree.
- **Don't:** diff a long-stale branch against the current default branch and treat what it shows as your own missing work --- some of it is the default branch's newer content read backwards.

(Measured 2026-09-04 on `Morrison-Lab/ai-config#3024`: the PR showed `MERGED`, but at another session's head commit rather than the one this session had pushed and had reviewed clean.
Three pieces of reviewed work were silently lost --- an enumeration, two corrected docstrings, and a test arm --- with nothing red anywhere.
Recovery was `Morrison-Lab/ai-config#3179`, cut fresh off `main` rather than off the stale branch, whose own diff against `main` showed `main`'s newer work as deletions.)

## Existence of a mechanism is not reachability of it

[`The four shapes`](#the-four-shapes) above names a counterpart that is **missing** ---
a cache `save` with no `restore`, a marketplace entry with no install.
This shape is the one where the counterpart is present, correct, and never reached.
The clearing branch is in the source, it does exactly what it should,
and the normal path never produces the input it reads ---
so the mechanism is real and the behaviour it promises is unavailable.

It is more convincing than the missing-counterpart case,
because finding the code that would have prevented a false positive feels like having explained the false positive.
Reading it produces a genuine and correct conclusion --- this guard clears on a terminal state ---
and that conclusion is about the source rather than about the run.
Existence and reachability are different claims,
and confirming the first is exactly what checking the second would feel like.

[`The test`](#the-test) above supplies the question, so ask it of reachability rather than of existence:
what would have to be true for this mechanism never to fire,
and does the normal path produce the input it matches on?
Trace the input backwards to whatever emits it.
Where the emitter is a command, read that command's actual output
rather than assuming it carries the fields the matcher wants.

- **Do:** name the producer of a mechanism's input, and read what that producer actually emits, before saying the mechanism works.
- **Do:** treat "the clearing branch exists" as an answer about the source and an open question about the run.
- **Do:** answer a guard's refusal from the guard --- read the property its message names,
  then its discharge condition in source --- before checking any state outside it.
- **Don't:** close an incident on the strength of having found the code that should have prevented it.
- **Don't:** read a matcher's field list as evidence those fields ever arrive --- a matcher is a claim about its input, not a supply of one.
- **Don't:** answer a guard's refusal by querying the forge.
  The refusal is a claim about the record the guard reads, not about the world it left you free to query;
  and a forge query nearly always returns something, so the wrong move feels like progress
  (2026-09-05: two responses spent on the forge ---
  the first confirming a review existed, the second that the PR had merged ---
  before anyone read the condition the guard actually consults).

(Measured 2026-09-03, and the record is the rule applied to itself three times.
A `Stop` hook demanded a per-HEAD reviewer request on an already-merged pull request.
The hook was read, a terminal-state matcher was found in it,
and the incident was written off as the guard behaving correctly given what it could see ---
a claim about the source presented as a claim about the run.
The first retraction asserted a *cause*:
that `gh pr merge`'s own success output carries none of that matcher's fields,
so merging without a later `--json state` probe would leave an obligation that can never discharge.
Adversarial review refuted that from the source of `hooks/no-unreviewed-pr.py`,
finding a second clearing branch --- `close_ident` ---
that discharges a merge structurally, from the command's argv and exit status,
with the terminal-state matcher reserved for a merge performed OUTSIDE the session.
The second retraction adopted that reading, and was wrong in the identical way,
because it too was reasoned from the code rather than run against the artifact.
Running the hook's own `close_ident` on the command the session actually issued settles it:

```python
close_ident("ALLOW_MERGE=1 gh pr merge 3101 -R Morrison-Lab/ai-config "
            "--squash --delete-branch 2>&1 | tail -3")
# -> (False, None, None, False)
```

Two independent defeats, either sufficient alone.
The environment-variable prefix makes the first token of the parsed argv something other than `gh`,
so the structural recogniser rejects the command outright;
and the merge is not the last simple command in the pipeline,
which the call site treats as ambiguous rather than as a discharge.
The idealized command with neither feature returns a clear,
which is the value both retractions were reasoning about.
So both clearing branches were reachable in principle and the run reached neither ---
one defeated by the command's shape, the other never fed its input.
[#3152](https://github.com/Morrison-Lab/ai-config/issues/3152) was filed on the first mistaken premise;
the correction is posted on its thread.
The lesson survives three wrong causes intact, and is sharper for them:
reading a matcher tells you what it would accept, and reading a recogniser tells you what it would recognize.
Neither tells you which branch this run took.
Only running the reader against the exact artifact does, and
[`mistake-patterns`](../../memories/mistake-patterns.md) Pattern 17 names that move ---
which is worth stating twice, because it was cited in the same change that failed to perform it.)

## A content diff verifies WHAT CHANGED, not whether the markup is valid

[`memories/office-open-xml.md`](../../memories/office-open-xml.md)'s "Two pandoc diffs verify a redlined docx" section already gives the standard content-level check for a tracked edit: diff the accept/reject conversions, and treat a clean pair as evidence the edit is right.
That is a check applied correctly and answering a narrower question than it looks like it answers -- related to the ninth shape above (a lossy conversion) and still distinct from it.
The ninth shape is about a conversion that *drops information it cannot carry*, such as a hyperlink target.
This is about a conversion that reports success over markup that is *outright invalid* -- the derived text comes out looking exactly right, and the file that produced it does not open.

The mechanism is specific to any format where an annotation is supposed to **gate** whether some content counts as present.
A tracked-change marker in OOXML gates a run's text under accept versus reject.
When the marker itself is malformed -- written as an empty child of a run's properties instead of wrapping those properties and the text (see `memories/office-open-xml.md`'s "Writing a NEW OMML tracked-change marker..." entry) -- a walker that simulates accept/reject has nothing to gate: the malformed marker sits *beside* the text rather than *around* it, so the same text is emitted whichever mode the walker simulates, and the diff between the edited file and the original comes out exactly as intended.
The check passes not because the markup is valid, but because content identity and markup validity are two different properties, and the check was only ever measuring the first.
A namespace defect can be just as invisible to the same diff for an unrelated reason: a part whose `mc:Ignorable` attribute names a prefix that part no longer declares (the same manuscript's second, independent defect) changes nothing about any run's text at all, so a text-level diff has no way to notice it regardless of how carefully it is read.

That is [`The test`](#the-test) above, applied to a check rather than to a claim: what would have to be true for this diff to be non-empty, and could a genuinely malformed file ever produce that?
For both defects here, no.
The malformed marker is symmetric under both readings, and the namespace defect touches no text a diff examines, so the diff cannot distinguish "the edit is correct" from "the edit corrupted the file's markup while leaving its rendered text (or its namespace-unrelated content) alone" -- the two states produce an identical diff.

- **Do:** run a structural/schema-level check (parse every part;
  verify the shapes an annotation is allowed to take;
  verify a prefix-list attribute against what is actually declared) *in addition to* a content diff, on any edit to a format where an annotation can be malformed without changing the content it annotates.
- **Do:** treat a clean content diff as evidence about content only, never as evidence that the file is well-formed or that its consuming application will open it.
- **Don't:** infer markup validity from a passing accept/reject (or any other rendered-content) comparison -- a malformed gate and a working one can render identically, and a namespace defect can sit entirely outside what the comparison looks at.
- **Don't:** trust a hand-rolled accept/reject walker's silence as confirmation;
  per [`memories/office-open-xml.md`](../../memories/office-open-xml.md)'s
  "A hand-built accept/reject simulator is itself an unverified instrument..."
  section, run it against a document you know is malformed and confirm it actually flags something, not only against documents you expect to pass.

(Measured 2026-09-09: a repair pass on a manuscript's tracked-change OMML equations swapped a `w:ins`/`w:del` marker from an invalid child-of-`w:rPr` position to the valid wrapping position across five successive delivered copies, while the verification in use throughout was `word/document.xml`'s accept/reject text diff (comparing paragraph text under each mode).
That diff reported the documents clean at every delivery -- the malformed marker was an empty element with no children, so neither the accept walk nor the reject walk treated it as gating anything, and the run's text simply always appeared.
A second, unrelated defect in the same manuscript -- `word/comments.xml`'s `mc:Ignorable` naming ten namespace prefixes it no longer declared, after a generic XML library re-serialized the part -- was equally invisible to the same text diff, for the unrelated reason that it touches no run text at all.
[`scripts/check-docx-tracked-changes.py`](../../scripts/check-docx-tracked-changes.py) in this repo is the structural check that would have caught both: it parses the actual XML and flags a `w:ins`/`w:del` sitting as a `w:rPr` child outside the one legal `w:pPr` exception, and a `mc:Ignorable` prefix with no matching namespace declaration in scope -- exactly the two properties a content diff cannot see.)

## A scripted edit's own PRINT and exit status, standing in for the file it changed

A heredoc'd or one-off patch script reports success two ways that are neither of them the artifact it was supposed to change: its exit status, and an `assert` or print statement it writes about its own progress.
Both describe the SCRIPT's control flow.
Neither describes the file on disk, because the write step, the encoding, or the target string the script matched against can each be wrong in a way that leaves the script's own report satisfied while the file is unchanged or corrupted.

Two heredoc'd Python patch scripts in one session printed a completion message and exited 0 while a re-grep of the target file immediately afterward showed the intended text unchanged.
The specific point of failure inside either script was never established --- only that the script's own report and the file's actual content disagreed, which is the fact this section is about regardless of which particular bug produced it in either case.

A third, in the same session, produced a more dangerous silence.
It replaced a one-line triple-quoted Python docstring, `"""..."""`, with replacement text that opened a new triple-quote delimiter without closing it.
The `str.replace` call itself succeeded exactly as instructed: the target string was found, and the swap at that one line looked sane on its own, since a single-line docstring edit is not the kind of change that reads as alarming.
What went wrong was not local to the edited line.
Every character of the file from that unclosed delimiter onward -- roughly 100 lines, an unrelated function among them -- silently became part of one Python string literal, until the next `"""` anywhere in the file happened to close it.
Counting quote marks would not have caught it either, since three triple-quote delimiters (balanced) is exactly what an *unclosed-then-reclosed-elsewhere* swallow also produces.
What caught it was `py_compile` raising at the point the file actually stopped parsing, together with a targeted re-grep for code that should have appeared in the swallowed region and did not.

- **Do:** after any scripted edit, verify from the artifact rather than from the script's own report -- grep the file for the new text and confirm the old text is gone, run a parser or compiler over it (`py_compile`, an R `parse()` call, the language's own syntax check), and run the relevant tests, before trusting that the change landed.
- **Do:** when a scripted edit replaces a delimited region (a docstring, a fenced block, a quoted string), verify the delimiters on both sides of the substitution are still balanced in context, not only that the substring search matched -- a replacement that opens a delimiter without closing it can leave the file syntactically parseable right up to the point it silently is not, with nothing at the edit site itself looking wrong.
- **Don't:** trust a script's own `assert`, print, or exit status as evidence the target file changed -- an assert can pass against the string it was handed without that string ever having matched the live file, and an exit 0 says only that the script's own control flow completed.
- **Don't:** read "the diff at the edit site looks fine" as sufficient for a delimiter-swap edit;
  the corruption in this shape is not local to the edited line, it is everything between the newly opened delimiter and wherever the file next happens to close one.

(Measured 2026-09-09: three heredoc'd Python patch scripts applied during one manuscript-review session.
Two reported success with the file unchanged on re-grep, cause unestablished.
The third swallowed roughly 100 lines of an unrelated function into a docstring by leaving a replacement's opening triple-quote unclosed;
`py_compile` and a targeted re-grep for the swallowed code were what caught it, not the script's own output.)

## A tracked-change DISPLAY VIEW, standing in for the resolved document a finding means

[`memories/office-open-xml.md`](../../memories/office-open-xml.md)'s "Two pandoc diffs verify a redlined docx" section already gives the producer-side use of accept/reject extraction: verify your own edit against both.
This is the same mechanism read from the other side --- a reviewer's finding, rather than an author's self-check --- and it is a different substitution from the ninth shape above.
The ninth shape is about a conversion that DROPS content it cannot carry.
This is about a rendering mode that SHOWS content that will not survive: Word's "All Markup" view (or an equivalent raw read of `word/document.xml` with no accept/reject simulation applied) displays a tracked insertion and the tracked deletion it replaces at once, stacked in the same place, which is exactly what a genuine stray duplicate would also look like.

Two review comments drafted for a manuscript told the author to repair a stray equation object and a doubled symbol.
Both were visible only in that display mode.
One sat inside a `<w:del>` the author had already used to remove it;
the other was the old half of a `<w:ins>`/`<w:del>` pair from an edit that replaced one symbol with another.
Extracting the resolved (accept-mode) text -- the same pandoc extraction the producer-side section already uses -- showed neither object survives: the equation and the doubled symbol are both absent once the tracked changes are resolved, and the finding was wrong.

The general shape: a display mode that shows pending edits inline is a genuine, CURRENT artifact of the file.
It is not stale, and it is not lossy in the ninth shape's sense.
It still is not the document a finding about "the document" is ordinarily understood to be about.
A reader who has not resolved the tracked changes is reading the union of two document states -- before the edits and after them -- and a finding drawn from that union has to say which state it is about before it means anything.

- **Do:** before reporting a finding about a redlined document, extract or view the RESOLVED (accept-mode) text and confirm the finding still holds there, not only in a display mode that shows pending changes inline.
- **Do:** when a finding is genuinely about the pre-edit or in-progress state -- a comment on the edit itself, not on its outcome -- say so explicitly ("in All Markup view", "before this deletion is accepted"), so the two states are never conflated silently.
- **Don't:** treat what an "All Markup" screen shows as the document a reader will eventually see;
  it is the union of two states, and "the document" defaults to the one a reader gets once changes are resolved.
- **Don't:** assume a stray-looking object or a doubled symbol found this way is a defect without first checking whether it is the visible half of a change the author already made.

(Measured 2026-09-09: two draft review comments for a manuscript resubmission named a stray equation object and a doubled symbol, both visible only in Word's "All Markup" display.
Extracting accept-mode and reject-mode text separately showed both were already-deleted tracked changes, and the accept-mode text was clean.
`memories/office-open-xml.md`'s "Two pandoc diffs verify a redlined docx" section gives the identical two extractions for a self-check on an edit;
this is the same mechanism applied to a finding about someone else's edit instead.)

## A correction's baseline is another artifact, and the nearest one in view is not it

["A drift claim is relational, so one read cannot settle it"](#a-drift-claim-is-relational-so-one-read-cannot-settle-it)
above already names the shape: a claim about two artifacts at once needs two
reads, and reading only one leaves the sentence feeling complete anyway.
Naming something a correction, a fix, a patch, or a workaround is that same
two-place claim, in the shape it takes most often in ordinary technical
writing rather than in an install or a config.
It names what changed AND what the change is against, and the artifact in
front of you, the corrected form, only ever supplies the first half.
The baseline lives somewhere else: an earlier paper, an earlier revision of
the same document, the pre-fix branch, last quarter's release.

That is what lets the failure survive careful reading.
The natural verification move is to reread the artifact you have, and doing
so genuinely confirms what the correction IS, its formula, its scope, its
effect.
It says nothing about what it corrects, because the document in hand was
never the baseline.
The nearest candidate actually in view, your own document's earlier draft, a
neighbouring equation, whatever you last edited, gets silently substituted
for the real one, because it is available and the real one takes a separate
retrieval to reach.

The check: before writing "X corrects/fixes/omits Y", name Y explicitly, say
where Y is written down, and read Y there.
If Y is not retrievable, describe what X does rather than what it corrects.

- **Do:** name the baseline a correction claim is against, cite where it is
  written down, and read it there before asserting the relationship.
- **Do:** describe what a correction term does, on its own, when the
  baseline it is said to correct cannot be retrieved and confirmed.
- **Don't:** confirm a correction's own content and treat that as having
  confirmed what it corrects.
- **Don't:** let the nearest document in view, your own earlier draft, a
  neighbouring section, stand in for a baseline a source names explicitly
  elsewhere.

(Measured 2026-09-09, in the same manuscript-resubmission session as the
shape above.
A response-to-reviewers letter and the manuscript's own supplement each
misnamed the baseline for a "correction term" in a cited paper (Teunis and
van Eijkeren, 2020, *Statistics in Medicine* 39:2799-2814).
That paper's "age dependent correction term" (p. 2801) corrects the age-free
density of an earlier 2012 Teunis et al. paper.
The letter instead said the term corrected an omission in the supplement's
own prior equation, which already carried the age restriction, an Iverson
indicator confining the relevant interval to the participant's age;
what that equation actually lacked was a different pair of factors.
The supplement separately mislabelled a term in the same formula: it called
one factor "the age-truncation term", when the truncation is a distinct
Iverson bracket and the named factor is instead the contribution of
inter-event intervals longer than the participant's age, vanishing as that
age grows.
Rereading the supplement, however closely, could confirm only what the
formula does; it could not show which paper's baseline the cited correction
was against, since that fact lives in the cited paper rather than in the
supplement.
Both were caught by the user asking "are you sure about that?", not by a
reread, which is the same discovery path
[`run-ums-proactively.cases.md`](run-ums-proactively.cases.md)'s "Are you
sure about that?" case record already names as invisible to a hook keyed on
a first-person admission: the wrongness surfaced as an answer to a question,
with no admission attached.)

## Naming a reference is not verifying it

Repairing a stale reference by making it durable and repairing it by making
it true are two different edits, and only the first one feels urgent when
the passage under repair is about references.

A positional cross-reference ("the 2nd occurrence above") breaks the moment
a record moves, which is exactly the failure this corpus's own
[`mistake-patterns.cases.md`](../../memories/mistake-patterns.cases.md)
header rules out by writing every cross-reference by name.
Swapping the position for a name is the correct fix for durability, and it
supplies none of the fix for accuracy: a named target is checkable, not
checked, and the check is a separate step that a reference-repair pass has
no built-in reason to take, since references are already the subject.
The named artifact still has to be opened and read against the specific
claim the reference is standing in for, the same substitution
[`The four shapes`](#the-four-shapes) already names for every other
adjacent-artifact case.

- **Do:** open the named target and confirm it contains the specific claim
  the reference stands in for, as a step separate from naming it.
- **Do:** treat "the reference is now durable" and "the reference is now
  true" as two claims needing two checks, even in a pass whose subject is
  references.
- **Don't:** replace a positional pointer with a named one and read the
  improvement in form as evidence of the content underneath.
- **Don't:** assume a reference-repair pass is exempt from this file's own
  rule merely because references, not facts, are what is being edited.

(Measured 2026-09-09, ai-config#3484: a stale positional reference reading
"the 2nd occurrence above" was replaced with a named pointer to "the
2026-09-03 occurrence recorded in this file" without opening that occurrence
to confirm it carried the claim being cited.
It did not.
The named occurrence was
[`mistake-patterns.cases.md`](../../memories/mistake-patterns.cases.md)'s
misidentified-hook-copy record, which carries no restart measurement at
all; the actual measurement lived in Pattern 43's own Fix step, in
[`mistake-patterns.md`](../../memories/mistake-patterns.md), a different
file entirely.
The repair converted an arguable pointer into a confidently false one, and
the confidence was new: a vague positional reference invites a reader to
check it, while a specific named one reads as already checked.)

## A diagnostic returning clean is evidence about the diagnostic, not the fault

A clean result from a targeted check answers "does this specific thing show
the problem", not "is the problem absent" --- and the gap between those two
questions is invisible exactly when every individual check was reasonable to
run.

The tell is a fault that keeps firing after every registration path a
diagnosis names comes back clean.
Each clean read gets spent arguing the fault must be elsewhere, when it is
equally consistent with the diagnosis having examined the wrong population:
a check that is sound on the artifact it reads says nothing about whether
that artifact is the one actually responsible.
Ruling out three registration paths in turn is real work and reads as
progress, but a fault that persists through all three is telling you about
the paths checked, not about the fault --- the same shape
[`fail-fast`](../principles/fail-fast.md) names for a pass path that isn't
provably disjoint from the failure path, applied here to a diagnostic
instead of to a guard.

The fix is not a sharper check on the same candidate set; it is capturing
the fault directly while it fires (a process sample, a live trace) rather
than continuing to deduce the culprit from registration files that have
already all read clean.

- **Do:** treat a clean result from every registration path checked so far
  as evidence about which paths were examined, not as evidence the fault
  sits elsewhere.
- **Do:** capture the fault live (a process sample taken while deliberately
  triggering it) once the obvious registration paths have all read clean,
  rather than adding a fourth path to the same deduction.
- **Don't:** read "every check I ran came back clean" as narrowing the
  search space --- it narrows the set of *checked* paths, not the set of
  *possible* ones.
- **Don't:** keep refining the diagnostic technique against a candidate set
  established by guesswork, when a direct capture would name the actual
  path without needing the set enumerated at all.

(ai-config#3141 is the worked incident, and it turned on the diagnostic twice
over.
Chasing which copy of `hooks/no-unreviewed-pr.py` was firing an expired
moratorium, a first pass reported the copy registered in
`~/.claude/settings.json` as current, all three `installed_plugins.json` pins as
containing no hook file, and `enabledPlugins` as `false` --- every registration
clean while the guard misbehaved.
That reading was itself an instrument artifact.
The probe initialised each pin's result to the string `(no hook file)` and
overwrote it only when a `MORATORIUM_END` line was found, so a pin whose hook
file exists but carries no moratorium constant printed as though the file were
missing.
Re-derived with the two conditions separated, one pin does hold the hook ---
user-scope, with no `MORATORIUM_END` at all, which is a copy predating the
moratorium and therefore one that demands the review unconditionally.
A registration did explain it.
So the incident supplies the rule twice: once for the diagnostic that returned
clean while the fault stood, and once for the probe whose defaulted variable
described a condition it never tested.
The record and its measurements live in
[`mistake-patterns.cases.md`](../../memories/mistake-patterns.cases.md)'s
Pattern 43 entry; this section states the transferable rule the incident
does not itself generalize.)

## A negative lookup cannot tell "never existed" from "no longer reachable"

The shapes above all substitute one artifact for another.
This one substitutes a *result* for a claim: an absence lookup returns nothing, and nothing is read as proof the thing was never there.

`git cat-file -t <sha>` answering `Not a valid object name` is the worked case.
It means only that the object is not in **this** clone **now**.
It does not distinguish an invented SHA from a real commit that has since become unreachable, and the difference is the whole finding: one is a fabrication to chase, the other a stale reference to re-point.

The trap is that the lookup feels like a *measurement* rather than an inference, so it escapes the claim-checking a stated fact would get.
It also arrives with the grammar of proof --- a command, a definite answer, no hedging --- which is exactly `grep-is-not-coverage`'s error one level down: there a search's silence is read as corpus coverage, here a lookup's silence is read as an artifact's nonexistence.

**GitHub's pull refs are the common generator.**
A `pull_request`-triggered workflow checks out `refs/pull/N/merge`, an ephemeral merge of the head into the base whose SHA is neither:

```console
$ git fetch origin 'refs/pull/N/merge:refs/remotes/origin/pr-merge'
$ git log -1 --format='%h parents: %p' origin/pr-merge
356e7cb5 parents: f3611051 4edc93a2
          ^ base    ^ PR head
```

That ref is replaced on every push, so a previous run's merge SHA is unreachable within minutes.
Anything a CI job reports about "the commit it ran on" is therefore unverifiable from an ordinary clone shortly afterwards --- and comes back looking fabricated.

The same shape covers a force-pushed commit, a deleted branch's tip, a dangling object past `gc`, and a rev in a shallow clone --- which is the sharpest, because the object exists on the remote and the local answer is still nothing.

The remedy is not a better lookup but a different question: ask what else would produce this exact silence, and whether the artifact you queried could hold the answer at all.
Where the reference is ephemeral, capture it **while it is current** rather than testing afterwards.
Where it is not, fetch the namespace that would carry it before concluding anything --- a clone that has never fetched `refs/pull/*` cannot see a pull ref, so its silence about one is a fact about the clone.

- **Do:** name what else explains the empty result, before reporting it as absence.
- **Do:** fetch the namespace or deepen the clone that would hold the object, and say which you did.
- **Do:** capture an ephemeral reference at the moment it is live.
- **Don't:** read `Not a valid object name`, a 404, or an empty query as evidence the thing never existed.
- **Don't:** treat a lookup as exempt from claim-checking because it ran a command --- the command measured this clone, and the claim was about the world.

(Measured 2026-09-10, ai-config#3508.
Three CI reviews on ai-config#3548 emitted a `commit_sha` that did not resolve locally, and I reported them on the tracking issue as SHAs that "do not exist" and abbreviate "nothing real".
The third carried a full 40-character value and prose saying "the merge commit introduces no further diff", which identified the mechanism: the job runs on the pull merge ref, so the JSON names a real commit that the next push made unreachable.
The clone I tested in had never fetched a pull ref, so it would have answered identically for every candidate explanation.
The defect is real and is a stale-reference one;
the fabrication reading was mine, and it pointed at the wrong fix.)

**Third occurrence, 2026-09-15, on ai-config#3635 --- and it supplies the one
check that needs no fetch at all.**

The prior two are the 2026-09-10 case just above and, earlier,
`memories/github-actions.md`'s "Detached HEAD on `pull_request` events" entry,
where a review's ancestry claim and merge SHA were read against the PR branch
on ai-config#2529 and were one edit from being published as a correction.
Here `git cat-file -t 341873f9` returned `Not a valid object name` and an
all-refs `git log` grep returned nothing, and a user-facing reply said the SHA
"does not exist in the repo at all".
`mcp__github__get_commit` resolved it at once:
`Merge d7169012 into 518ccc81`, committer `web-flow` --- the pull merge ref
this section already names as the common generator.

Two things the earlier records do not carry.

**The two local queries were one measurement.**
`git cat-file -t` and an all-refs `git log` grep both ask "is this object in
*this* clone", so their agreement is what
[`metacognitive-monitoring`](metacognitive-monitoring.md)'s
"count a sibling command that reads the same field as a second opinion" rules
out.
The clone was also shallow, which `git rev-parse --is-shallow-repository`
answers in one command --- `memories/git.md` already prescribes that pre-check
for ancestry and count queries, and it applies unchanged to an existence
query.

**The review carried the right SHA in its own body.**
Its structured `review-data` payload named the merge commit while its prose
trailer read `Reviewed commit: d7169012`, so the contradiction was visible in
the artifact already in hand and needed no lookup of any kind.
Read both fields before concluding anything about either.
On the producing side, a `pull_request`-triggered job wanting the head must
use `github.event.pull_request.head.sha`;
`github.sha` and a bare `git rev-parse HEAD` both give the merge commit.
Filed as [ai-config#3662](https://github.com/Morrison-Lab/ai-config/issues/3662).

Three occurrences of a rule written down in two places is the
[`deterministic-tools`](../principles/deterministic-tools.md) third-occurrence
bar, so the instrument is proposed rather than the sentence sharpened:
[ai-config#3666](https://github.com/Morrison-Lab/ai-config/issues/3666) is a
warn-only `Stop` guard for a reply calling a SHA nonexistent with no
forge-side lookup behind it.

- **Do:** read every SHA a review reports --- the structured payload and the
  prose trailer --- before treating either as the commit it reviewed.
- **Do:** run `git rev-parse --is-shallow-repository` before reading any
  local absence, existence queries included.
- **Don't:** count a second local query as corroboration; both answer for the
  clone, which was never the question.

## The invoking process is itself a member of the population a filter scopes, and reading the filter's prose does not check that

Every shape above substitutes one artifact for another.
This one substitutes a **claim about a passage** for a claim about the **mechanism the passage describes** --- distinct from ["A mechanism's prose is not the mechanism's definition"](#a-mechanisms-prose-is-not-the-mechanisms-definition) above, which is about a comment's motivating example being narrower than the mechanism it explains.
Here the passage is not narrow or ambiguous;
it states its remedy plainly and correctly as prose.
The gap is between reading that prose carefully and confirming it, and checking whether following the remedy actually produces the exclusion it is read as promising --- which needs the tool's own documented behaviour, not a second reading of the sentence describing it.

A companion memory entry, tracking the underlying `pkill -f` hazard as [ai-config#3427](https://github.com/Morrison-Lab/ai-config/issues/3427) and merged into [`memories/shell.md`](../../memories/shell.md) via [ai-config#3428](https://github.com/Morrison-Lab/ai-config/pull/3428), recommends: for a `pkill -f` scoped only by a shared script path, resolve candidates with `pgrep -f <pattern>`, then filter each one on its own working directory (`readlink /proc/<pid>/cwd`) against your own worktree, killing only a match.
Reading that passage and confirming it says what it says is not the same claim as "this filter cannot re-admit the shell that is running it."
The wrong belief was exactly that stronger claim: that a cwd filter, applied to `pgrep -f`'s candidate list, cannot reach the invoking session, because a non-matching session sits in a different working directory.
That reasons about *other* worktrees' processes and never asks whether the invoking shell itself belongs to the population the filter admits.

It does.
`pgrep -f <pattern>` matches by regex against every process's full command line, and excludes only its own PID --- not its ancestors.
The shell that typed the `pgrep`/`pkill` command has that pattern text in its own command line (you just typed it) and sits in your own worktree by construction, which is exactly the condition a cwd filter is built to accept.
So the shell running the search is a candidate the filter admits, not one it excludes, and nothing about re-reading the passage's prose --- however carefully --- would surface that, because the prose is a true description of the filter it proposes.
What it needed was a check against `pgrep`'s own documented behaviour: `pgrep --help` lists `-A, --ignore-ancestors` for precisely this case, which means the tool's own authors anticipated it and named the flag the passage's remedy omits.

- **Do:** before trusting a filter's stated coverage, ask whether the process performing the filtering is itself a member of the population being filtered --- not only whether the passage describing the filter is read correctly.
- **Do:** check a claim about what a filter excludes against the tool's own documented behaviour (a `--help` flag, a manpage clause), not against a second reading of the prose recommending it.
- **Don't:** treat "the filter's cwd check would not match a *different* worktree's process" as having shown it cannot match the *invoking* one --- those are different claims, and only the second is the one that matters for self-exclusion.
- **Don't:** read a remedy's prose as verified once it has been read carefully;
  a coherent, correctly-stated remedy can still fail to achieve what it is read as promising.

(Measured 2026-09-09 in this sandbox: `pgrep --help` lists `-A, --ignore-ancestors` as a documented flag, confirming `pgrep -f <pattern>` excludes only its own PID by default and matches an ancestor shell whose command line contains the pattern.
The corrected belief and its displacing fact are the same pair [`memories/shell.md`](../../memories/shell.md) (ai-config#3428) records for the underlying `pkill -f` hazard;
this entry is the general verification-method lesson the specific fix does not itself state --- that verifying the passage prescribing a remedy is not verifying the remedy holds.)

## An unresolved review thread is evidence about the review conversation, not about the code

The shapes above all substitute one artifact for another.
This one substitutes the state of a review conversation for the state of the codebase.
On 2026-09-12, six review threads were still unresolved when Morrison-Lab/ai-config pull requests [ai-config#3440](https://github.com/Morrison-Lab/ai-config/pull/3440) and [ai-config#3469](https://github.com/Morrison-Lab/ai-config/pull/3469) merged.
The natural reading was six open defects needing tracking issues.
Checking each against `origin/main` showed none of the six was an untracked open defect.
Four were fixed by the maintainer's final commits before merge.
One was addressed by a documenting note.
One was already recorded as a known gap in the approximation section of `hooks/flag-config-deletion-without-ref-check.py`.
Filing on thread state would have produced six issues, none of which named an untracked problem.

The remedy is to read the file an unresolved thread names on the default branch before concluding that thread represents a current defect.

- **Do:** read the file an unresolved thread names on the default branch (e.g., with `git show origin/main:<path>`), and file only what is still true there.
- **Don't:** treat an unresolved review thread on a merged pull request as evidence of an open defect in the code.

## When the check that would refute the claim is unavailable, the claim is unverified --- not merely caveated

The four shapes above all describe verifying the *wrong* artifact.
This one describes the case where the right artifact is identified correctly
and simply **cannot be reached** --- a blocked egress proxy, a missing
credential, a UI with no API behind it.

The failure is not that the check is skipped.
It is what happens to the claim afterwards.
The unreachable check gets demoted to a parenthetical, the claim is stated at
full confidence, and the caveat reads as thoroughness rather than as the
warning it is.
Nobody is deceived about the blocked check, because it is disclosed --- they
are deceived about the claim, which was never downgraded to match.

Measured 2026-09-15.
A Quarto site was rendered, deployed to `gh-pages`, and the branch confirmed to
hold 36 HTML pages, 21 PDFs and every asset directory.
Every one of those is a fact about the **branch**.
The claim made was that the site was *published*, with one routine settings
step left --- a fact about **serving**, which the session could not check
because its proxy blocked `github.io`, and which it noted in passing while
stating the claim anyway.
The setting did not exist: the repository was a private fork, and GitHub Pages
was unavailable to it entirely (see
[`github-repo-transfers`](../../memories/github-repo-transfers.md)).
The maintainer had to supply what the blocked check would have shown.

**The asymmetry to notice is that a blocked check removes evidence against the
claim while leaving every piece of evidence for it intact.**
So the remaining evidence looks unanimous, and confidence goes *up* exactly
when it should go down.
That inverts the usual relationship between missing information and certainty,
which is why disclosing the gap does not correct for it.

The test is the one this fragment already states, applied to reachability
rather than to identity: ask what would have to be true for the claim to be
false, then ask whether the artifact that would show it is one you can
actually reach.
When it is not, say the claim is unverified and name what would settle it.

- **Do:** state the claim at the confidence the reachable evidence supports,
  and say plainly which part is unverified.
- **Do:** name the specific check that would settle it, so whoever can run it
  knows what to run.
- **Don't:** disclose the blocked check and then assert the claim anyway --- a
  caveat beside a confident claim is read as rigour, not as doubt.
- **Don't:** treat unanimous surviving evidence as strong when the blocked
  check was the only thing that could have disagreed.

## A local test run is not the CI job's conclusion

This is the "a checkout for the run" substitution in its most available form, and the one least likely to register as a substitution: the local run is faster, it is under your hand, and it tests the same code.

Measured 2026-09-18 on [`Morrison-Lab/gha`](https://github.com/Morrison-Lab/gha) PR 883.
A session ran `python3 -m unittest discover -s antigravity-review/tests` at head `dd243dc`, got 42 of 42, and reported the `antigravity-tests` job green.
The check-runs query in that same turn reported the job `in_progress`.
It did conclude `success` a minute later --- so the conclusion was right and the evidence for it did not exist yet, which is the dangerous case, because nothing corrects it.

`hooks/no-stale-pr-status.py` caught it: it compares a clean-state assertion against the most recent status query in the transcript, so the gap was visible to an instrument even though the claim turned out true.

The two artifacts genuinely differ.
A CI job can diverge on runner OS, tool versions, steps wrapped around the suite, and inputs the job's own `with:` block overrides --- `gha`'s `CLAUDE.md` records `PHI_DETECTORS` and `NLB_GLOBS` doing exactly that, so a local run of the same script exercises a different configuration from the one CI runs.

- **Do:** use a local run to decide whether to push, never to report a job's state.
- **Do:** report a job by its `conclusion` field, and name the job id so the claim is checkable.
- **Do:** read `status` before `conclusion` --- `in_progress` has no conclusion, and an absent conclusion is not a pass.
- **Do:** read the job's own `with:` block before trusting a local invocation of the script it calls.
- **Don't:** characterize the PR when some checks are still running;
  say which jobs concluded and that others are in flight.
  A whole-PR claim is a scope claim over every check.

## A file's metadata is another shape, and it is the one that never feels like a substitution

Every shape above swaps one *document-like* artifact for another --- a cached copy, a checkout, an endpoint, a summary --- so each at least looks like the thing it stands in for.
This one swaps a document for its **filesystem metadata**, which resembles it not at all, and is easier to miss for exactly that reason: nothing about reading a modification date feels like reading the file, so no substitution registers as having happened.

Measured 2026-09-23 on [`Morrison-Lab/mlg`](https://github.com/Morrison-Lab/mlg) PR 20.
A vendored mirror of Stanford's CS229 was described as "Autumn 2008" in two repositories' READMEs, in two directory names (`cs229-stanford-2008`, `cs229-see-2008`), in three commit messages and in four issue comments.
No document said so.
`practice-midterm.pdf` heads itself "CS 229, Autumn 2007";
all four problem sets and all four solution keys head themselves "CS 229, Public Course" and name no term at all;
the mirror's own course page names only the instructor, and its sole `2007` and `2008` strings sit inside markup rather than visible text.

The year came from `ls -la` --- the files' October 2008 modification dates.
That is a real measurement of a real property, and it answers a different question: **when this copy was written**, not **which offering produced it**.
The two answers are both dates attached to the same file, which is what makes the substitution invisible;
and a publication date being later than the term it publishes is the normal case rather than a warning sign.

Three things made it durable rather than a passing slip.
The claim was written into directory *names*, so every later reference restated it as established fact.
The files had been handled extensively --- checksummed, sized, typed, moved --- so the session had every feeling of familiarity with them and had still never opened one.
And `file` was run, which reports a PDF's page-tree metadata: on a 26 MB, 1098-page book it said "3 pages", which is the same metadata-for-content confusion one level down.

The general form is worth stating because it is not specific to dates: **a property that a file *has* is not a property the file *asserts*.**
Size, mtime, permissions, path, and the filename itself are all facts about the copy in front of you.
Provenance --- who made it, when, for which offering, under what licence --- is a claim, and a claim has to be read out of the content or its source page.
A filename that encodes provenance (`Bishop-...-2006.pdf`) is somebody else's undocumented claim, not a source.

`hooks/flag-unsourced-term-attribution.py` is the instrument: it warns when a write pins a term to a year beside a document filename and no text-extraction command appears anywhere in the transcript.
It deliberately does not count a `WebFetch` as evidence --- the measured session fetched two course sites and still got the term wrong, because neither page stated one.

- **Do:** extract the document's own text (`pdftotext -f 1 -l 1 <file> -`) before writing any provenance claim about it, and quote what it returned.
- **Do:** treat a filename that encodes a date or a term as a claim needing the same check, not as the check.
- **Don't:** derive a provenance fact from an mtime, a size, or a path --- those describe the copy, not the work.
- **Don't:** read familiarity with a file as having read it;
  checksumming, moving and typing a file all leave its content unopened.
