Every self-review is an **adversarial review by a separate subagent**.
Whenever this corpus calls for reviewing your own work --- before a push, as the fallback when the external reviewer is down, or the project-conventions pass a clean verdict does not discharge --- dispatch it to a reviewer with its own context and an adversarial brief, and take its findings as findings.
The authoring session never reviews its own diff inline.

[`self-review-fallback`](self-review-fallback.md) governs *when* a self-review is owed and to what standard.
This governs *who performs it*, which that fragment left to the author by default.

## Why the authoring session cannot be the reviewer

A reviewer's job is to read what the diff **says**.
The author knows what it was **meant** to say, and that knowledge is not removable by care --- it is the context the session is made of.
So the author reads the artifact and recovers the intent, which is confirmation rather than review.

This corpus already names the shape.
[`verify-the-right-artifact`](verify-the-right-artifact.md) is about verifying an adjacent artifact thoroughly instead of the target one, and the adjacent artifact here is the intent in your own head --- the one artifact that is never wrong, because it is what the diff was written from.
Nothing about the check feels skipped: the reading is real, the standards are applied, and the answer comes back clean.

Two consequences worth stating separately, because they are easy to run together.

**A subagent buys independence of intent, not independence of vendor.**
A Claude subagent reviewing a Claude diff shares the training and therefore the blind spots, exactly as [`self-review-fallback`](self-review-fallback.md)'s cross-vendor section says.
What it does not share is the account of what the change was for.
Those are different independences, and this rule buys the second one only.

**So the subagent is the floor, not the ceiling.**
For merges, the next section adds a stricter gate still:
a second adversarial review from a different model and harness.

## Cross-model and cross-harness reviews are required for merging, and the harness list is concrete

(Directive from the user, 2026-08-25: "all reviews, even self-reviews, must be
adversarial; don't do them yourself, use a subagent, preferably using a
different model and harness".)

Two gates meet here, and they have different independence bars.

The **self-review duty** (gating a push) takes an adversarial subagent on any
harness, same-harness included --- that floor buys independence of intent,
which is what a push gate needs.
The **merge gate** (see [`fully-clean`](fully-clean.md)) requires more:
a reviewer differing from the authoring session in **both** model and harness,
the only configuration that also buys independence of blind spot.

The user's 2026-08-25 machine inventory names **cursor**, **agy** (CLI),
**opencode**, **claude**, and `codex` wherever installed
([`delegate-to-codex`](../../skills/delegate-to-codex/SKILL.md)).
From the authoring session's perspective the ladder filters itself:
any entry sharing your model or your harness does not qualify for this gate,
whatever the list says.
Dispatch in independence-and-availability order --- `agy` CLI or `opencode` first, then `codex` (UCDH projects only), then `claude` --- where each entry qualifies only if both its model and harness differ from the authoring session.
This review order serves independence and measured availability, overriding [`delegation.md`](../../memories/delegation.md)'s cost-first delegation order for general work.
A multi-backend harness qualifies only when both its harness
and its configured model differ from the authoring session.
`cursor` stays out of the active ladder until its headless dispatch
is probed here.
If no qualifying entry remains, autonomous merging waits ---
it never falls through to a same-model or same-harness reviewer.
A quota outage reroutes the dispatch --- it does not license skipping it.
Waiting does not overrule a human:
escalation to the repository owner per [`fully-clean`](fully-clean.md)'s
deadlock rule ends in their manual review and merge decision,
which is the one authority above this gate.

`agy` specifically: its API-dispatch route is retired, but the **agy CLI** is a
separate path and remains available --- see
[`delegation.md`](../../memories/delegation.md)'s delegate ladder.
A retired API never disqualifies a CLI harness
that operates on a separate path from it.

A second directive the same day sets the merge consequence:
"you must not merge, even with mwc enabled,
unless you have a 100% 'all clear' review verdict
from an adversarial review".

This **adds** a gate and replaces none.
Every requirement [`fully-clean`](fully-clean.md) already sets stands unchanged
--- including the external automated PR reviewer's clean verdict at head,
wherever a repo has one ---
and an author-dispatched subagent verdict never satisfies that external gate.
What is added: a merge additionally requires
the author-dispatched cross-model, cross-harness reviewer's
100% all-clear adversarial verdict at the shipping head.
A Needs-more-work verdict blocks until a compliant re-dispatch returns
all-clear at the new head.
A skip notice, a stub, or a stale-head verdict clears nothing.
A split --- one all-clear and another not-clean, nits included --- is not
100% all-clear, and `mwc` does not authorize merging it
(ai-config#2274).
ARD every item from every review, then request fresh reviews.
If no qualifying reviewer is reachable, the merge waits ---
"blocked on reviewer availability" is the honest status
only once it carries the per-provider enumeration
[`Availability is a per-route question`](#availability-is-a-per-route-question-and-command--v-answers-one-route)
requires ---
and arming an auto-merge while waiting is
[Pattern 12](../../memories/mistake-patterns.md).


The merge-side rules live with the gate they serve:

- **Do:** for any merge, use a reviewer on a **different model and harness**
  from your own
  (agy CLI, opencode, codex,
  claude only for sessions authored outside Claude,
  or cursor once its headless dispatch is measured),
  and report which harness produced each verdict.
- **Don't:** merge anything --- under any grant, `mwc` included ---
  without a 100% all-clear adversarial verdict at the shipping head.
  A skip notice, a stub, an older-head verdict,
  or a same-harness convenience pass clears nothing.
- **Don't:** reuse a passing same-harness pre-push verdict
  to satisfy the merge gate.
  A merge needs its own cross-model, cross-harness verdict
  evaluating the shipping head.


## Round repetition is a third axis, orthogonal to cost and independence

(Directive from the user, 2026-09-10, given while a session drove a `gha`
pull request through five straight rounds of push-gate self-review, each
dispatched to the same-harness `adversarial-reviewer` subagent on the
conductor's own Opus tier at roughly 350k subagent tokens per round: "always
use agy or other cheap subagents whenever feasible, to avoid draining claude
quota".)

[`when-to-orchestrate`](when-to-orchestrate.md)'s "Route each agent's
model/effort" section already names two axes for a dispatched call.
**Cost** says mechanical work gets a cheap tier and judgment-heavy work
inherits or escalates; a self-review is judgment-heavy, so this axis alone
argues for keeping it on the conductor's own tier by default.
**Independence** says a judgment-heavy verify stage wants a different model
family, not just a different prompt; the merge gate above already applies
this to review specifically.
The push-gate floor's own same-harness permission rests on neither axis: it
buys independence of intent, a subagent reading the diff without the
author's account of it, which any dispatched subagent supplies regardless of
its tier or family.

Neither axis, and not the push-gate's own intent-independence reasoning
either, prices in **repetition**.
ARDI drives a PR through however many rounds it takes to reach a clean
verdict, and each round dispatches the identical shape of call: the same
brief structure, the same reviewer persona, a diff that has usually only
shrunk.
The cost of that shape is the per-round cost times the round count, and the
round count is exactly the number nobody knows in advance.
Five rounds at roughly 350k tokens each is 1.75M tokens spent on one PR's
push-gate reviews alone, all of it against a task the push-gate floor never
required to run same-harness in the first place --- it only allowed it.

So repetition is a signal on its own, independent of whether any single
round is judgment-heavy: once a review-shaped dispatch is known to repeat
against the same PR, prefer the cheap or cross-family route from the first
round rather than the fifth.
[`delegation.md`](../../memories/delegation.md)'s "agy as a cheap
adversarial-review lane on macOS" measurement already shows this pays off
beyond cost: across nine rounds on two PRs, `agy --print` caught real defects
a same-family Sonnet round had missed, at no Claude quota cost.
The platform qualifier is part of the heading and is kept here deliberately,
since that measurement was taken on macOS and this section's own worked
example is a Windows checkout.

**This narrows `delegation.md`'s "Claude subagents are for reviewers only"
carve-out rather than repealing it.**
That entry (2026-09-09) reserves the `Agent` tool for the
`adversarial-reviewer` persona and routes every other subagent to `agy`.
Read on its own, the heading can sound like a standing preference for
Claude on review work specifically.
The same paragraph already says the opposite three sentences later: `main`'s
`hooks/no-push-without-self-review.py` accepts an `agy --print '<prompt>'`
discharge directly, so a current hook install needs no Claude reviewer for
the push gate at all.
The carve-out is a ceiling on non-review Claude dispatch, not a floor under
review dispatch --- check which of the two a given sentence in that entry
actually states before reading it either way.

Before assuming a Claude dispatch is required, check whether the pre-push
hook that would gate the push is actually current.
The installed copy on a given machine can be a symlink to a checkout that
has fallen behind `origin/main`, in which case it silently reverts to
requiring an `Agent`-tool reviewer regardless of what `main`'s own hook
source supports
([ai-config#3094](https://github.com/Morrison-Lab/ai-config/issues/3094)
tracks the drift).
When the installed hook is stale, satisfy it as it actually behaves rather
than as it should --- the fix for staleness belongs to that issue, not to
the push in front of you.

- **Do (from the user):** default to `agy` or another cheap, cross-family
  route for review-shaped dispatch whenever it is feasible, rather than
  reaching for the same-harness Claude subagent by habit.
- **Do (inferred):** treat a review known to repeat --- an ARDI loop driving
  a PR to clean, not a one-off pass --- as a stronger case for the cheap
  route than a single isolated review, since the cost is the per-round cost
  times the round count.
- **Do (inferred):** verify the active pre-push hook's actual behavior (the
  installed copy, not `main`'s source) before assuming it requires or
  forbids a given reviewer shape, and satisfy the hook as installed.
- **Don't (inferred):** read "Claude subagents are for reviewers only" as a
  reason to prefer Claude for review; it restricts non-review Claude
  dispatch and says nothing about preferring Claude over a cheaper
  discharge for review itself.
- **Don't (inferred):** keep dispatching the same-harness reviewer round
  after round on the strength of the push-gate floor's "any harness is
  fine" --- permitted is not preferred once the round count passes one.

## Availability is a per-route question, and `command -v` answers one route

The inventory above is a **machine** inventory:
every entry in it is a local binary,
so a `PATH` probe answers whether each entry is *installed*, and nothing else.
Installation is not availability:
an installed CLI can still be quota-blocked or unauthenticated,
which is the second question
[`Query all available providers sequentially`](#query-all-available-providers-sequentially)
asks when it requires every exclusion to be recorded with its reason.
A reviewer is reachable by any of three routes,
and only the first of them can ever appear on `PATH`:

1. **A local CLI**, probed with `command -v`.
   The binaries the orchestrator's model adapters
   ([`model_adapters.py`](../../scripts/orchestrator/model_adapters.py)) probe are derivable,
   so start from them rather than from a list copied into this sentence,
   per [`avoid-hardcoding-external-data`](../coding/avoid-hardcoding-external-data.md):
   `grep -o 'shutil.which("[a-z0-9-]*")' scripts/orchestrator/model_adapters.py | sort -u`.
   That set is a floor rather than the population:
   a CLI no adapter probes never appears in it,
   and `agy` is the worked case,
   since [`delegation.md`](../../memories/delegation.md)'s ladder routes dispatchable work to it
   while the adapter named after it probes `gemini`,
   so probe the union of that set with the 2026-08-25 machine inventory above.
2. **A forge-side bot**, which runs on the forge and so is invisible to `PATH` in principle.
3. **An API key** for a provider reachable without its CLI,
   probed in the environment rather than on `PATH`.
   The variables are every API-key variable the adapters read,
   so derive them from
   [`model_adapters.py`](../../scripts/orchestrator/model_adapters.py)
   rather than from a subset copied into this sentence,
   per [`avoid-hardcoding-external-data`](../coding/avoid-hardcoding-external-data.md):
   `grep -o '[A-Z_]*API_KEY' scripts/orchestrator/model_adapters.py | sort -u`.
   A subset written out here once left a reader probing fewer variables
   than the adapters read,
   which is this section's own failure one route over.
   Not every variable the derivation returns gates an adapter's `is_available()`,
   since some are read only when a call is made,
   so a hit names a provider to probe rather than one to record available.

Both derivations read a path that exists in `ai-config`'s own checkout.
Where that file is absent --- a consumer repository,
or the lab manual's transclusion of this fragment ---
each command returns nothing,
and that null is a missing source rather than an empty population.
Fall back to the machine inventory above for the local-CLI route,
reading it there as a floor rather than as that route's population,
and record the shortfall alongside the probe result.
Record the API-key route as underivable in that repository rather than as empty.
An underivable route is a recorded exclusion carrying its reason
rather than a satisfied enumeration,
extending [`Query all available providers sequentially`](#query-all-available-providers-sequentially)'s
requirement that every exclusion of a known provider be recorded with its reason
to the case where the providers themselves cannot be named ---
so the status names the route that could not be derived
rather than reading as a bare block.

A null `command -v` sweep is therefore evidence about `PATH` and about nothing else.
That is [`grep-is-not-coverage`](grep-is-not-coverage.md)'s shape,
with `command -v` in place of `grep`,
and [`verify-the-right-artifact`](verify-the-right-artifact.md)'s,
with the local machine standing in for the set of reachable providers.

The forge-side route is the one an inventory of binaries cannot see,
so its providers are enumerated here rather than left to be inferred:

| provider | how it is reached |
| --- | --- |
| Copilot | a reviewer request on the pull request, where the repository has Copilot review enabled |
| Jules | a mention comment from an account the workflow's `author_association` allowlist admits, where the repository carries a Jules review workflow |
| Antigravity | such a mention comment, or a manual workflow run, where the repository carries an Antigravity review workflow |

The concrete dispatch mechanism behind each row --- the endpoint, the tool name, the workflow file ---
varies by harness and by repository,
so it lives in [`claude-review-dispatch`](../../memories/claude-review-dispatch.md)
and [`gh-cli`](../../memories/gh-cli.md)
rather than here.

**Read the table as a starting list rather than as the population, and re-derive every row per repository.**
All three rows are repository-conditional, not only the two comment-triggered ones:
a review workflow is added or removed by one pull request,
so derive those rows from the repository's own `on:` blocks,
and Copilot review is switched on and off per repository by a ruleset rule
and per user by quota,
so derive that row from the repository's rulesets.

**A comment-triggered row also depends on who posts the mention, which is a property of the session rather than of the repository.**
Both comment-triggered workflows gate their job on
`author_association` being one of `OWNER`, `MEMBER`, or `COLLABORATOR`,
and a session whose comment writes land under a bot identity posts as `CONTRIBUTOR`,
so the job skips with no error and no verdict
([ai-config#1433](https://github.com/Morrison-Lab/ai-config/issues/1433),
and [`self-review-fallback`](self-review-fallback.md)'s own statement of the same gate).
A posted mention is therefore not a dispatch:
read the created comment's `author_association` back,
and read the workflow's own `if:` for the allowlist it applies.

**Copilot carries a state the other two rows do not: reachable and withheld.**
A standing maintainer directive forbids requesting Copilot code review
on any pull request in any repository while the moratorium stands.
Read its live expiry from the `MORATORIUM_END` constant
in [`no-unreviewed-pr.py`](../../hooks/no-unreviewed-pr.py),
never from a date copied into prose,
per [`avoid-hardcoding-external-data`](../coding/avoid-hardcoding-external-data.md);
that constant was still in the future when this section was written on 2026-09-03,
so the moratorium was live and the row's recorded state was "reachable, withheld".
The full statement and its measurements are in [`gh-cli`](../../memories/gh-cli.md).

Whether a forge-side reviewer's harness and model differ from the authoring session's
is a per-session question,
settled by the ladder above rather than by this table ---
a session authoring under the `agy` CLI and reviewed by an Antigravity workflow
shares a harness family with its reviewer,
and whether it shares a model is not readable from this table either.
Whether any of the three rows satisfies the **merge** gate is a separate question this section does not settle:
the ladder and its Do bullet above are unchanged by this inventory,
and the multi-backend rule stated there still has to be applied per provider,
since a caller that passes an empty model input resolves that model downstream
and the value has to be read where it is resolved rather than assumed here.

**So "blocked on reviewer availability" owes an enumeration, not a probe.**
That status is honest only after naming every known provider and the state it was found in,
which is what [`Query all available providers sequentially`](#query-all-available-providers-sequentially)
already requires
and what a single failed sweep never establishes.

- **Do:** run the availability check as a list of provider routes --- local CLI, forge bot, API key --- and record each provider's state.
- **Do:** name every known provider and its state before writing "blocked on reviewer availability",
  a provider that is reachable but withheld by policy included.
- **Do:** re-derive every forge-side row against the repository in hand,
  since a workflow, a ruleset rule, and a quota each turn one of them on or off.
- **Do:** settle a forge-side reviewer's merge-gate qualification against the ladder above,
  per provider and per authoring session.
- **Do:** derive the API-key route from the adapters that read those variables,
  since they are its source of truth and a copied list drifts from them.
- **Do:** derive the local-CLI route from the adapters' own probes,
  then probe the union of that set with the 2026-08-25 machine inventory above,
  since each source drops what the other carries ---
  the derivation drops `agy`, and the inventory drops `gemini`.
- **Do:** read the commenting identity's `author_association` back
  before recording a comment-triggered forge reviewer as dispatched.
- **Do:** read an empty local-CLI derivation as a missing source
  wherever `model_adapters.py` is not in the checkout,
  fall back to the machine inventory above as a floor,
  and record the shortfall alongside the probe result.
- **Do:** record the API-key route as underivable rather than as empty
  wherever that file is not in the checkout,
  since the machine inventory names no API-key variable.
- **Do:** record an underivable route as an explicit exclusion carrying its reason,
  by extension from [`Query all available providers sequentially`](#query-all-available-providers-sequentially),
  and name that route in the status line.
- **Don't:** read an empty derivation as an empty population;
  that is this section's own thesis failing on the section itself.
- **Don't:** read a null `command -v` sweep as an availability verdict;
  it reports what is on `PATH` and stops there.
- **Don't:** record a CLI as available on a `command -v` hit alone;
  a present binary can still be quota-blocked or unauthenticated,
  which is a state the probe never reports.
- **Don't:** enumerate the API-key route from a list written out in prose,
  here or anywhere else; that list is a subset the moment an adapter is added.
- **Don't:** treat either local-CLI source alone as that route's population
  where both are readable;
  the derivation drops `agy` and the inventory drops `gemini`,
  so a probe of one of them misses a route the other names.
- **Don't:** read the inventory-only fallback as that population either;
  where `model_adapters.py` is absent the inventory is a floor,
  so the shortfall is recorded rather than resolved.
- **Don't:** count an underivable route as an enumerated one;
  a route whose providers cannot be named is an exclusion,
  so a bare "blocked on reviewer availability" over it is the unenumerated claim again.
- **Don't:** record a comment-triggered forge reviewer as dispatched on a posted mention alone;
  an `author_association` allowlist skips the job with no error.
- **Don't:** read the three forge-side rows as that route's population;
  a repository can carry a review workflow this table does not name,
  so derive the rows from its own `on:` blocks and rulesets rather than from this list.
- **Don't:** treat the machine inventory above as the provider population, since a forge-side reviewer cannot appear in it.
- **Don't:** record a withheld provider as available,
  or read its row as licence to dispatch it.
- **Don't:** read a row in this table as evidence that its harness or its model
  differs from the authoring session's.

(Measured 2026-09-03 on ai-config#3105:
a session held two green PRs for roughly seven hours as
"blocked on reviewer availability",
on a `command -v` sweep over eight CLIs that correctly found none of them installed.
Jules was measured reachable, replying within a minute of a mention comment.
Antigravity carried a review workflow and was permitted, and was not probed.
Copilot was reachable and withheld:
the moratorium had been extended the previous day and was read as expired,
so the Copilot request made on #3084 that day was a breach of the directive
rather than a measurement, and is recorded here as the slippage it was.)

## What "separate" requires

**Its own context window.**
The near-miss is the pass performed in the same turn under a reviewer framing --- "now let me look at this adversarially" --- and it reads as compliance, because the prose that comes out is adversarial in tone and nothing in the output distinguishes it from a dispatched review.
The test is mechanical rather than tonal: an `Agent` call was made, or it was not.

**Foreground, not background.**
A background dispatch returns an agent id rather than a report, so the verdict is not the call's result and the work you are gating cannot wait on it.
This is the Agent tool's own criterion for `run_in_background: false` --- the very next action depends on the answer.
When the harness forces background execution (e.g. Remote Control active or background agent isolation), the tool result is an async launch stub (`agentId: ...` or `status: running`), and task `output_file` remains at 0 bytes while running.
Polling `output_file` for `Reviewed-Commit:` never matches while the subagent is in flight.
The verdict arrives only via a subagent hand-back message after the turn ends.
`hooks/no-unshipped-commit.py` stays quiet while a pre-push review of HEAD is in flight (ai-config#4109), allowing the turn to conclude cleanly so the background subagent can finish and deliver its report.

**Read-only.**
The reviewer reports; the author disposes.
A reviewer that can edit turns a finding into a silent fix, which loses the finding and the disposition together.

**Freshly dispatched, not resumed --- and this applies per round, not just to the first one.**
"Its own context window" above rules out reviewing in the author's own turn.
It does not by itself rule out a second failure with the same shape: resuming the *same* reviewer session across rounds (`SendMessage` back to an existing subagent) instead of dispatching a new one each time.
A resumed reviewer keeps its own context window separate from the author's, so it still satisfies the first bullet.
What it no longer has is independence from **itself** --- a later round of a resumed reviewer is reading the diff with every earlier round's own conclusions already in its context, which is [`learn-from-review-findings`](learn-from-review-findings.md)'s convergence pattern happening *inside one reviewer* rather than across a series of different ones.
Findings-per-round on a resumed reviewer characteristically decline toward zero, and the decline is not evidence the diff has actually gotten cleaner --- it is at least partly the reviewer running out of things it has not already told itself are fine, the same steerable-narrowing mechanism ["Narrowing severity is evidence about COVERAGE, not about the defect population"](#narrowing-severity-is-evidence-about-coverage-not-about-the-defect-population) describes for a series of rounds generally, here concentrated inside a single reviewer's own memory rather than spread across dispatches.

This is the same failure ["The PR's own review history is rationale you cannot withhold"](#the-prs-own-review-history-is-rationale-you-cannot-withhold) describes, one layer more direct: there a *fresh* reviewer inherits the narrowing by reading about prior rounds in the artifact, while here the reviewer does not need to read about its own prior rounds because it remembers making them.

A resumed reviewer is not useless.
It is the right tool for a narrower job: confirming that the specific findings *it already raised* were actually fixed, where continuity of context is exactly what makes it efficient.
What it cannot do is supply the go/no-go verdict that gates a push or a merge, because that verdict needs to be checking the diff against the standards, not against its own earlier self.

- **Do:** gate a push or merge's go/no-go verdict on a freshly dispatched reviewer with no prior context on this diff, every round, not only the first.
- **Do:** use a resumed reviewer for the narrower job of confirming that findings it already raised were fixed --- continuity is an asset there.
- **Don't:** treat a resumed reviewer's declining finding count, or a "ready for merge" it restates after several resumes, as the gating verdict.
- **Don't:** read "an `Agent` call was made" alone as satisfying independence --- a resumed call was made and still fails this bullet.

(Measured 2026-09-18 on `fix/1601-baseline-tolerance-band` / [Lacaedemon/sparta#1603](https://github.com/Lacaedemon/sparta/pull/1603), reconstructed from the PR's own commit messages rather than from session-internal reviewer state, which is not recoverable after the fact --- so the quoted headline in each item below is a direct quote from that commit's own message, not a reconstructed round-by-round narrative layered on top of it.
Round 1 (commit `d8fbc397`): "Six findings from the pre-push adversarial review, all addressed."
Round 2 (commit `67b58d26`): "Three findings from the second adversarial review."
A later round (commit `631c9216`) names the contrast this section is about directly: "A fresh adversarial review, run without the previous rounds' context, found two real defects the context-carrying reviewer had passed over."
Another (commit `574c1fa2`) repeats the same shape: "A third reviewer, dispatched with no knowledge of the earlier rounds, found a real bug two previous reviewers had passed over."
Three further rounds each self-label as a numbered, independently-dispatched reviewer and each found more: a corrupt-but-present baseline unpacked without raising, so the old gate read it as "no baseline" and routed the run into the bootstrap branch instead of reporting it as corrupt (commit `e21f2a4e`, "Three findings from a fourth independently-dispatched reviewer").
Only the *average* of two benchmark runs was validated, so a `-1.0` raw run paired with a normal one would have averaged to a plausible positive value and been committed (commit `42b78526`, "Five findings from a fifth independently-dispatched reviewer").
And `math.isfinite` raised `OverflowError` on an integer too large to convert to `float`, crashing the very predicate written to absorb corrupt input (commit `3453e113`, "Three findings from a sixth independently-dispatched reviewer").)

**No Agent tool, or no reviewer registered here?**
A separate CLI is the same move and a stronger one ---
[`delegate-to-codex`](../../skills/delegate-to-codex/SKILL.md),
[`delegate-to-opencode`](../../skills/delegate-to-opencode/SKILL.md),
or [`adv`](../../skills/adv/SKILL.md) (`pre-push-review.py`).
`adv` auto-detects and excludes the active agent harness
from rotation via session environment variables;
specify a target directly (`--engine <name>`)
or pass `--exclude-engine cursor` in alternate mode
until headless cursor dispatch is enabled.
The `adversarial-reviewer` persona also lives at `.claude/agents/` and `.opencode/agents/`, which are project agents: a session rooted in another repo may not be able to resolve it at all ([ai-config#1921](https://github.com/Morrison-Lab/ai-config/issues/1921) tracks shipping it alongside the guard).
The plugin does not close that gap either: its root is the ai-config repository root, which ships `skills/`, `commands/` and `hooks/hooks.json` and no `agents/` directory (checked at `c7201140`, 2026-09-15), so a consumer repo installs the guard and none of the personas it names.

**In that repo, use the fallback the guard already admits, and get its two conditions right.**
`no-push-without-self-review.py`'s `FALLBACK_AGENT_NAME` accepts `general-purpose`, `general`, `reviewer`, `code-reviewer`, `research` and `self` (with an optional `-`, `_` or space inside the two-word spellings), but only when the dispatch's own prompt matches `REVIEW_PROMPT_RE` --- `adversarial review`, `adversarial self-review`, `pre-push review`, or `self-review`, again with an optional separator.
"Review this adversarially" satisfies a reader and not the regex.
The report then has to meet the verdict-line contract, which "A verdict phrase separated from its heading by a line break is no verdict" and "Structured review data (JSON payload)" state below between them.

**Read which denial you got: the two messages fail at different stages, and the second has three causes.**
"No `adversarial-reviewer` subagent or recognized external reviewer ... was dispatched" means the dispatch was not recognized --- wrong persona name, or a prompt the regex missed.
"An `adversarial-reviewer` subagent was dispatched, but no verdict came back as that call's own result" means it *was* recognized and no verdict was extracted: a background dispatch, an errored result, or a verdict line that does not match that contract.
A missing `Reviewed-Commit:` is not one of them --- that has its own message, about a clean verdict that does not say which commit it read.
Where you chose to background the dispatch, the guard's own message gives the fix and it is a foreground re-dispatch.
Where the harness backgrounds it regardless --- #3045's two variants, below --- re-dispatching changes nothing and the `ALLOW_UNREVIEWED_PUSH=1` route below is the remedy.

- **Don't:** reach for `ALLOW_UNREVIEWED_PUSH=1` on the second message from a foreground dispatch that returned a report.
  It says the report was read and no verdict was found in it, which is a formatting fix, not a case where the guard cannot see a verdict at all.

(Measured on ucdavis/lbt, 2026-09-15, with the plugin enabled.
Dispatching `adversarial-reviewer` returns an errored result reading "Agent type 'adversarial-reviewer' not found.
Available agents: claude, claude-code-guide, Explore, general-purpose, Plan, statusline-setup".
The session's `general-purpose` fallbacks were refused anyway, on both messages, and pushed under the override.
Diagnosed in the same session by the step this file prescribes below.
One prompt opened "You are an adversarial reviewer" and nothing else in it named the review, which `REVIEW_PROMPT_RE` misses because the word does not end at `review`, so that dispatch was never recognized at all.
The prompts that were recognized produced reports with no `Verdict:` line in them: `read_latest_review` over that session's JSONL returns `(None, None, True)`, and `grep -c Reviewed-Commit` on the same file returns 0.
Both are the brief rather than the guard, which refused correctly each time.
A fallback dispatch has to name the review in words the regex matches, and has to *ask* for the verdict and fingerprint lines, because a persona file supplies them and a `general-purpose` prompt does not.)

**A subagent cannot discharge this guard at all, and the reason is where the transcript lives.**
Measured 2026-09-15 in the Claude Code desktop harness: an `Agent` dispatched *by a subagent* writes both the call and its report to `<session>/subagents/agent-<id>.jsonl`, while the guard reads the session JSONL at the top level.
Every conforming review of one branch, the clean one included, was reachable only through those per-subagent files, with `grep -c Reviewed-Commit` on the session transcript returning 0.
So the clean verdict exists, is about the right commit, and is unreadable to the guard by construction.
This is not #3045, whose dispatch never returns synchronously.
Here it returns synchronously to the subagent, which is not who the guard is reading, so it joins the cases below where the guard cannot see a verdict rather than the formatting case.
Note which denial that produces: with no top-level dispatch, `saw_reviewer_call` is false, so it is the **first** message, and the Don't above does not bite.
A subagent that reviews before pushing takes the override, and says in its report which reviews produced which verdicts and where they live.

Note what that CLI fallback does to the pre-push guard, since the two rules meet here and pull opposite ways.
A CLI's verdict never becomes an `Agent` call's `tool_result`, so the guard cannot see it however real the review was.
Prefix the push itself with `ALLOW_UNREVIEWED_PUSH=1` there, and say in the same reply which reviewer produced the verdict and why the subagent route was unavailable --- the override covers a push whose verdict the guard cannot check, not only a push with nothing to check.
The same applies to a session whose reviewer is registered from a stale definition, which is the case on any rollout of a change to the persona itself.
Where no second context is reachable at all, say so in the review itself rather than letting an inline pass be reported as a dispatched one.

**A harness that always backgrounds the `Agent` tool is a fourth case where the guard cannot see a verdict it should, alongside no reviewer being registered, a stale reviewer definition, and a subagent's own transcript above --- and the guard's partial fix for this one has a specific extraction bug.**
[`ai-config#3045`](https://github.com/Morrison-Lab/ai-config/issues/3045) tracks the general case: a harness whose `Agent` dispatch never returns synchronously, so the guard's own "dispatch in the foreground" remedy is unfollowable.
That issue's own report is one manifestation --- `run_in_background: false` explicitly set, and the dispatch still backgrounded.
A second, distinct manifestation is a harness whose `Agent` tool carries no `run_in_background` field in its schema at all, so there is nothing to set.
Confirmed directly for one Claude Code CLI session, not asserted as true of "Claude Agent SDK sessions" generally, since #3045's own report came from a session where the field did exist.
Both produce the identical symptom --- "Async agent launched successfully" with an agent id, and the verdict arriving later as a task-notification --- from different causes, and both are the same shape as [`memories/antigravity.md`](../../memories/antigravity.md)'s "Asynchronous subagent dispatch and pre-push self-review (`invoke_subagent`)" entry for Gemini CLI's `invoke_subagent`, which predates and independently confirms this is a cross-harness pattern rather than one build's quirk.

`read_latest_review()` in `hooks/no-push-without-self-review.py` already has a partial fix for exactly this: the "Genuine task notifications from tracked background reviewer dispatches" block, added by #2820 (closing #2544) on 2026-09-01, two days before #3045 was filed.
It tries to recover a task id from the dispatch's own tool result, then matches a later task-notification against that id and parses its text for a verdict.
The recovery step tries `json.loads()` first, and on failure falls back to a regex requiring the literal key `task[-_ ]?id` or `conversationId`.
One session's tool result read `agentId: a29a955ac15b38f72`, which matches neither alternative, so the id was never captured, the later notification never matched, and the guard refused the push on all four dispatches in that session despite each one returning a genuine, independently verified verdict.
Reproduced directly:

```python
re.search(r"\b(?:task[-_ ]?id|conversationId)[:=]\s*[`\"']?([\w-]+)",
           "agentId: a29a955ac15b38f72", re.I)
# -> None
```

Posted to #3045 with a proposed one-line fix: widen the key alternation to include `agentId`.

Until that lands, the remedy is the CLI-fallback one given above: run the review (the async dispatch still produces a real report, just not as the call's own synchronous result, and the guard's automatic matching cannot yet recover it either), confirm the reported `Reviewed-Commit` matches what the push will actually ship, and use `ALLOW_UNREVIEWED_PUSH=1` on the push itself, stating in the same reply which review produced the verdict and that the harness's dispatch could not satisfy the guard's own foreground check.
Re-dispatching the same reviewer again on the theory that a different `run_in_background` phrasing will change the outcome does not.
On the field-absent variant there is no field to change, and on #3045's own variant the harness ignored the field once already.

(Diagnosed 2026-09-14, driving `Morrison-Lab/ai-config#3684`: four `Agent` dispatches to `adversarial-reviewer` in that session, each with `isolation: "worktree"`, all returned "Async agent launched successfully" with no `run_in_background` field available on the call to begin with.
Each review's full report arrived only via a later task-notification, and each push attempt was refused until `ALLOW_UNREVIEWED_PUSH=1` was used on a push whose `Reviewed-Commit` matched the final CLEAN verdict's head.)

**The other, original #3045 manifestation --- `run_in_background: false` explicitly passed, and the dispatch still backgrounded --- recurred again on the same variant it was first reported under.**

(Diagnosed 2026-09-26, driving `Morrison-Lab/ai-config#4013`: three separate `Agent` dispatches to `adversarial-reviewer` in one session, each with `run_in_background: false` set on the call, each came back an async `agentId` rather than a synchronous result.
This is the field-present variant, not the field-absent one #3684 recorded above: the field existed, was set to the value that should have forced a synchronous call, and the harness ignored it three times in a row.
Confirming that both variants keep recurring independently, rather than one having been a one-off, is the point of recording this instance: `ALLOW_UNREVIEWED_PUSH=1` was used twice in that session with the review's own verdict and `Reviewed-Commit` stated in the same reply, per the remedy given above.)

**Remote Control and background agent isolation deliver the report as a hand-back message.**

(Diagnosed 2026-09-28/29, driving `Morrison-Lab/ai-config#4109`: on Windows under Claude Code with Remote Control active, every `adversarial-reviewer` launch through `Agent` returned an async-launch stub rather than the synchronous report, even with `run_in_background: false` explicitly passed and `isolation: "worktree"`.
The task's `output_file` stayed at 0 bytes while running, so polling it for `Reviewed-Commit:` never matches.
The report arrives only as a subagent hand-back message after the turn ends.
`hooks/no-unshipped-commit.py` was updated to recognize in-flight pre-push reviews of HEAD, staying quiet to allow the turn to conclude and let the subagent complete reactively rather than deadlocking against `no-push-without-self-review.py`.)

**Cursor Cloud has a subagent dispatch.**
On Cursor Cloud, when the session's `Task` tool lists
`adversarial-reviewer`, that is the dispatch
(measured 2026-08-25 PDT on a Grok conductor).
If `Task` is absent or does not list that persona,
that is the CLI-fallback case above.
Morrison-Lab/ai-config's Cursor adapter skips
`no-push-without-self-review.py` until
[#2241](https://github.com/Morrison-Lab/ai-config/issues/2241),
so `ALLOW_UNREVIEWED_PUSH=1` is inert for the adapter
under any reviewer
(see [`memories/cursor.md`](../../memories/cursor.md)).
Call `parse_report()` from the worktree's `hooks/no-push-without-self-review.py`
on the report recovered from the child's transcript
when the worktree hook script exists
(see [`memories/cursor.md`](../../memories/cursor.md)).
Do not import `~/.claude/hooks/`:
it is a different revision from the branch under review.
When the three-dot diff includes
`hooks/no-push-without-self-review.py`,
also parse with `origin/<default-branch>`'s copy, or obtain a CLI review.
If the worktree script is missing, obtain a CLI review.
Do not push unless the verdict is `clean` and the
fingerprint prefix-matches HEAD.
If there is no fingerprint
(including a stale-registered persona),
obtain a CLI review.
The empty `pr-on-claim` `--allow-empty` branch has no report to parse:
do not invent one,
do not refuse that push for lack of a verdict,
and say in the reply that the carve-out was used.
The carve-out is `git rev-list --count origin/<default-branch>..HEAD`
equal to 1 and `git diff --quiet HEAD^ HEAD` exit 0
in the checkout whose push follows.
Exit 1 means a diff; exit 128 means the command failed.
Both conditions passing is the `--allow-empty` pr-on-claim commit.
`git diff origin/<default-branch>...HEAD` empty
in the checkout whose push follows is tree equality,
not "this branch carries nothing".
A net-zero tree of other commits is not the carve-out.
If the dispatch errored, produced no report,
or produced a report whose fingerprint cannot be recovered
(including a stale-registered persona),
obtain a CLI review,
write that reviewer's report to a file under `/tmp`,
and call `parse_report()` on that file.
If Claude Code's native guard is also running, the prefix
is that guard's escape even when the adapter skip makes
it inert for the adapter.

## Brief it with the diff and the standards, never with your rationale

The instinct on writing the brief is to supply the context that makes the change make sense --- what the problem was, why this approach, what the alternatives were.
That re-imports precisely what dispatching was meant to exclude.
A reviewer holding your account of the change checks the diff against **it** rather than against the repo, and agrees, because the two were written by the same session minutes apart.

Give it the base ref, the paths, the standards that apply, and the question.
Where the change's own reasoning matters, it is in the diff --- a comment, a docstring, a fragment --- and the reviewer should be reading it there, where a later reader will.
Scope is not rationale: which branch, which base, where the tests live, and what is out of scope are facts the reviewer cannot derive and must be told.

**`<base>` is a claim, and it is the one nobody checks.**
The sentence above names the base as a fact the reviewer must be told, and stops there.
A base you resolve from a *local* branch name silently widens the diff whenever that branch is behind, so the reviewer spends its attention on already-merged work and returns findings against code this change never touched.
Resolve it from a remote-tracking ref after fetching that remote, and state the merge-base SHA and the file and insertion counts in the brief so the scope is checkable rather than asserted.
[`verify-the-right-artifact`](verify-the-right-artifact.md)'s "A comparison's base is an artifact too" carries the direction of the error and the forge cross-check that settles it.

- **Do:** hand over `git diff <base>...HEAD`, the applicable rules, and the question.
- **Do:** resolve `<base>` from a fetched remote-tracking ref, and state the merge-base SHA and the diff's counts alongside it.
- **Don't:** hand over the case for the change.
  If it is not persuasive from the diff alone, that is the finding.
- **Don't:** name a bare local branch as `<base>` --- one behind its remote widens the diff so the reviewer works on already-merged code, and one that is ahead of or diverged from its remote in commits the head branch also carries narrows it so part of the change is never reviewed.
- **Don't:** read a clean verdict as covering the whole change when the base was local;
  the narrowing direction produces exactly that.

### Brief reviewers to enumerate all findings in a single pass

Convergence slows down dramatically when a reviewer reports only a single finding per round (e.g. reporting one naming defect, then waiting for a fix to report casing in the next round, and style in another).
Brief reviewers to conduct an exhaustive pass across all categories (defects, factual claims, slop, repo conventions, semantic line breaks) and enumerate every single defect in a single pass to enable rapid convergence.

- **Do:** instruct the reviewer to enumerate all findings exhaustively across all categories in its first pass.
- **Don't:** allow a reviewer to trickle findings one at a time across multiple successive rounds.

### The PR's own review history is rationale you cannot withhold

The rule above governs the brief you write.
A reviewer reading the **PR** gets a second channel you never chose to open: the claim comment saying how many rounds ran, the round-by-round commit messages, and the `# round-2 review finding 8` markers a fix left behind in the test file.
Each of those is a true record, and none of them was written as an argument, which is why the effect is invisible from both ends --- nothing in the artifact reads as persuasion, and nothing in the verdict reads as deference.

Measured 2026-08-24 Pacific on [ai-config#2131](https://github.com/Morrison-Lab/ai-config/pull/2131).
The repo's own reviewer returned **Ready for merge** and named the history in its own justification: the PR's history and the round markers baked into the test file "show the near-misses I would normally look for [...] were already found and fixed in earlier rounds, and my independent probing did not surface anything beyond that."
A cross-vendor pass on that same head then returned 11 findings, 8 of them blocking (see [`self-review-fallback.cases.md`](self-review-fallback.cases.md), "A clean same-vendor verdict over eight blocking cross-vendor findings").

**State what is observable, which is narrower than it first looks.**
That verdict also lists nine verification steps it ran, so "it probed less" is a claim about effort that those steps weigh against, and at least one other explanation fits --- the reviewer probed normally and the diff was, in its own words, "almost entirely prose/documentation plus one well-isolated, warn-only hook with unusually thorough self-testing".
What is observable is that **the history entered the justification**: a reason for finding nothing was supplied by the artifact rather than derived from the diff.
That is enough, because a verdict resting partly on prior rounds is partly a re-reading of those rounds, so it is worth less as corroboration than its independence suggests --- however hard it worked.

**A `# round-N` marker is a changelog of past MISSES, not a certificate of coverage.**
It records that one defect was found there once.
It says nothing about the family that defect belonged to, and reading it as evidence of scrutiny inverts its meaning: the marker points at a line that was wrong, and a reader who takes it as a coverage claim reads it as a line that has been checked.

**The loop is self-reinforcing, which is what makes it a rule rather than a matter of care.**
More rounds produce a more reassuring history, which is available as a reason to stop, and a diff that accumulates many rounds is often one complex enough to need more.
So the effect is strongest exactly where it is most costly.
That is [`learn-from-review-findings`](learn-from-review-findings.md)'s convergence rule reaching a reviewer who never ran the earlier rounds: there a series narrows its own search space by inheriting findings, and here a *fresh* reviewer inherits the narrowing from the artifact instead.

- **Do:** quote back the sentences in a verdict that cite the PR's history rather than the diff, and say what independent evidence is left once they are set aside.
- **Do:** label a regression case with the property it pins rather than the round that found it, so the comment is a specification a reviewer can check instead of a report that scrutiny already happened.
- **Don't:** read a long visible review history as coverage --- it is a record of what was found, and every entry marks a place a defect once lived.
- **Don't:** count a verdict that cites prior rounds in its justification as a fully independent round;
  that much of it is a re-reading of the rounds it names, however hard the rest of it worked.

### Tell it to RUN the repo's validation, not only to read the diff

The rule above says which facts a brief must carry.
It does not say what the reviewer should **do** with the checkout, and the default answer --- read the diff against the standards --- is what every round does when the brief does not say otherwise.
Reading is the wrong instrument for a whole class of defect, because a diff-reading round can only find what is visible in the changed lines, and a registration gap, a broken import path, or a check that fails on files the diff never touched are all invisible there by construction.

So name the commands in the brief.
The repo's own checkers, its test suite, and whatever `check-install`-shaped verification exists are the ones that matter, and they are cheap for a reviewer already holding the checkout.

Recorded 2026-09-03 from [ai-config#3059](https://github.com/Morrison-Lab/ai-config/issues/3059), on a hook that took nine adversarial rounds.
Eight rounds read the diff.
The ninth ran the repository's own local validation, and found **two red CI gates that all eight prior rounds had missed** --- one of which would have shipped the hook completely **inert to plugin-path consumers** while every test in the suite passed.
The record gives no more detail than that.
It does not name the hook, the gates, or what was missing, so read it for the shape rather than for a mechanism --- and note that a diff-reading round cannot find a gate that fails on files the diff never touched, whatever the gate turns out to be.

Note what this is not.
It is not "Run every mechanical style instrument before dispatching, not after", which says *you* run the style instruments beforehand so the reviewer never spends attention on them.
That rule keeps mechanical noise out of the round.
This one puts a different instrument *into* the round, because the author's pre-dispatch run and the reviewer's own run answer different questions --- yours confirms the diff is clean, and the reviewer's confirms the repo is.

- **Do:** name the repo's *functional* checkers and test command in the brief --- not the style instruments, which you have already run --- and ask for their output rather than a judgement about them.
- **Do:** ask specifically whether the change is *reachable* --- registered, imported, wired into the path a consumer actually takes --- since that is the gap a diff read cannot show.
- **Don't:** assume a reviewer holding the checkout will run anything it was not asked to run.
- **Don't:** read a run of diff-reading rounds as having covered the repo;
  they covered the changed lines, which is a different population.

## Run every mechanical style instrument before dispatching, not after

The rule above is about what the reviewer sees.
This rule is about what should never reach the reviewer:
a defect the repo's own deterministic checker already catches.

Some of a prose diff's style classes have an instrument and some do not.
In this repo the instruments are
the vendored semantic line-break checker
(`scripts/vendor/gha-check-new-line-breaks.py`, which is what CI runs:
one sentence per line, plus a clause rule for a long line with a mid-line semicolon),
`markdownlint`,
`scripts/check-links.py`,
and the directional-word grep the [`fix-forward-references`](../../skills/fix-forward-references/SKILL.md) skill runs.
An ambiguous pronoun has no detector, per [`ambiguous-reference`](../writing/ambiguous-reference.md),
so that class stays with the reviewer,
though a grep for a pronoun that opens a clause after a comma or a conjunction,
the positional heuristic that fragment names,
narrows where to look.
Each instrument is cheap and deterministic,
and each runs in seconds.
That speed is not a reason to skip the adversarial round.
That speed is the reason the round should never be the first thing that finds a defect an instrument covers,
per [`algorithmatize-checks`](algorithmatize-checks.md).
An adversarial round costs real tokens and real time.
Spending a round on a defect a repo script would have caught for free
is the same waste `algorithmatize-checks` names for any check a human re-derives by hand:
reviewer judgment substituting for an instrument that already exists.

So run every available mechanical style instrument on the diff,
fix what those instruments report,
and only then hand the diff to the adversarial reviewer.
Brief the reviewer to report every finding in one round, style findings included.
The point of running the instruments first is to keep the round's own findings
down to what only judgment can catch,
not to teach the reviewer that style is someone else's job.

- **Do:** run the repo's own style checkers on the diff
  (the semantic line-break checker, `markdownlint`, the link checker,
  the forward-reference grep, or whatever the repo defines)
  and fix their output before the first adversarial dispatch.
- **Do:** brief the reviewer to report every finding in one round
  rather than holding style findings for a later pass.
- **Don't:** dispatch a diff to the adversarial reviewer
  before the repo's mechanical style checkers have run on that diff.
- **Don't:** brief the reviewer to leave style findings for a later pass.

[`ardi`](ardi.md)'s "Three or more review rounds" section carries the one exception, and it is narrow: on a **prose** diff that has already reached three finding-bearing rounds whose remaining findings are style preference, further trimming, or one more caveat, that section directs a later round's brief to withhold exactly those classes.
The rule above governs every other case, including the first round of any diff.

(Measured 2026-09-02 driving
[#3025](https://github.com/Morrison-Lab/ai-config/pull/3025),
a 20-line addition to `memories/reviewing-prs.md`.
Four adversarial-reviewer rounds ran, each costing roughly 210k tokens.
Round 1 found a misattributed citation plus word-wrapped lines.
Round 2 found an ambiguous "It" and a forward-pointing "below".
Round 3 found lines that ran several clauses together.
Round 4 was clean.
The word-wrapped and run-together lines were the semantic line-break checker's territory,
and the forward-pointing "below" was the forward-reference grep's;
only the citation and the pronoun needed a reader.
Running those two instruments first would have collapsed the four rounds to at most two.)

## Its findings are findings

[`self-review-fallback`](self-review-fallback.md) already rules out surfacing a defect in your own review and closing it on your own estimate of its blast radius.
A dispatched reviewer makes that concrete, because the finding now has an author who is not you: give each one Address, Rebut, or Defer-to-a-tracked-issue per [`ard`](../../skills/ard/SKILL.md), in writing, exactly as for a finding from the PR's own reviewer.
"I know why that is fine" is a Rebut, and a Rebut is something you would be willing to post.

## Review the instrument too, not only the change it verifies

The section above says a dispatched reviewer's findings are findings.
This says what to put in front of it, and the answer is wider than the change: **the verification artifacts are part of the diff and get reviewed as such.**

The reason is not symmetry.
A change is guarded by the suite and by the instrument, so a defect in it has two independent detectors.
A defect in the *instrument* has none --- the suite does not test the parity checker, the parity checker does not check itself, and a broken instrument's characteristic output is a reassuring number rather than an error.
That inverts the intuition that tooling is the low-risk part of a diff: it is the part with the fewest detectors, and therefore the part where an independent reader is worth the most.

The corollary is that a green suite is not a reason to shorten the round.
A suite reports that the assertions written so far pass, which is silent about an assertion that cannot fail, a control patching dead code, and a metric over the wrong quantity --- three defects a reader finds by reading and no run finds at all.

So brief the reviewer with the whole diff, naming the verification files explicitly rather than describing them as scaffolding, and ask specifically: what result would this instrument have to produce for the change to be abandoned?
An instrument with no such result is a finding on its own terms, per [`verify-the-right-artifact`](verify-the-right-artifact.md)'s transformation-for-conclusion section.

- **Do:** include tests, controls, harnesses, and parity checkers in the diff the reviewer sees, and name them in the brief.
- **Do:** ask what output would falsify the instrument, and treat "none" as a finding rather than as reassurance.
- **Do:** keep running rounds while findings keep landing;
  a round that finds something is evidence the next one will too.
- **Don't:** describe the verification files as scaffolding, or scope the review to "the actual change" --- that excludes the least-guarded code in the diff.
- **Don't:** read a green suite as a reason to stop early;
  the defects this section is about are invisible to it by construction.

(Measured 2026-08-28 on [ai-config#2515](https://github.com/Morrison-Lab/ai-config/pull/2515).
Five adversarial rounds each found real defects against a fully green suite, and three of the five found them in the verification tooling rather than in the change: a parity metric that could not fail, a negative control patching a function that had moved off the execution path, and an assertion comparing a function against itself.
The last of those had let a previously-rejected design pass 299 tests.)

## Give a docs-only diff describing an instrument a full round

The section above says a diff's verification artifacts are the least-guarded part of it.
Its limit case is a diff carrying no code at all: documentation describing how an instrument behaves.
That reads as the safest change available: nothing executes, no suite can break, and the round feels like a copy-edit.
Treat that reading as the risk rather than as a fact about relative rates --- one case cannot establish which diffs get cut short most often, and it does not have to, because the instruction is the same either way: review it at full depth.

The defect it carries is not new here.
[`fact-check-prose`](../writing/fact-check-prose.md)'s "Prose that distills code is a code claim, checked like code" already owns it, names the same psychology, and prescribes the same remedy;
its "condensation of the code that builds it" section extends the rule to a written-out command, and its fenced-block section to program output.
Read those for what the check is.
What this section adds is the **review-side** consequence, which none of them states: that a docs-only diff about instruments invites an early stop, and that the findings cluster rather than scatter when the round runs to depth.

The measured shape is worth carrying because it tells a reviewer where to aim.
The findings cluster, rather than scattering: a consumer described as reading one field when it falls back to another, a format called unparseable when the parser accepts it, a value called rejected when nothing validates it, a set of accepted forms given as two when the code accepts three.
None reads as a guess afterwards, because each is a claim about a file in the same repository, and knowing roughly what that file does feels like having read it.
Where the claim is about which branch fires, read the branch.
A negative claim --- *this form does not parse*, *nothing accepts this* --- is the one to execute rather than reason about.
Reading can settle it, when the parser is small and you read all of it;
what reading cannot tell you is whether you read all of it, and every refuted negative claim in this measurement was made by someone who believed they had.

- **Do:** run each consumer named in the prose against the input the prose describes, before writing the sentence about it.
- **Do:** treat a negative claim about a parser, guard, or matcher as owing an execution, not an argument.
- **Do:** let the round count be decided by whether findings are still landing --- the rule the section above already gives --- rather than by the diff's size or its lack of code.
- **Don't:** read "no code changed" as "nothing here can be wrong" --- the claims changed, and they have no suite.
- **Don't:** describe a fallback, a precedence rule, or an accepted-form list from the shape of the code;
  enumerate it from the code.

(Measured 2026-09-02 on [ai-config#3010](https://github.com/Morrison-Lab/ai-config/pull/3010), a documentation-only change carrying 52 insertions and 1 deletion across four files, of which 34 insertions are to this file.
Twelve adversarial rounds are recoverable from the session: 7, 8, 4, 3, 3, 3, 1, 1, 0, then 4, 1, 0 after the scope reopened.
Nine of those 35 findings were the one shape above, and five of the nine are recoverable, each a claim the named consumer disproves once read or run: that the three payload consumers read the payload and nothing else, when they fall back to prose;
that a bolded verdict phrase does not parse, when it does;
that demoting a disclosure marker changes `_reviewer_identity()`, when the Claude Code footer is deliberately excluded from `REVIEW_AGENT_MARKERS`;
that a non-conforming payload is rejected, when nothing validates it;
and that the pre-push guard accepts two verdict phrasings, when it accepts three.)

## Require detailed and holistic review passes

Reviewers must independently assess both detailed, evidence-backed implementation defects and the whole change:
requirements, intent, cross-file consistency, integration, regression risk, and validation.
A perfunctory scan of isolated diff hunks misses both subtle line-level bugs and systemic architectural drift.

The two passes evaluate complementary failure modes:

1. **Detailed implementation defect audit**:
   - Trace control flow, edge cases, error handling, syntax, regex greediness, and path-escaping at the line level.
   - Fact-check external tool behaviour and claims against direct documentation rather than trusting prose.
   - Detect placeholder comments, cargo-cult code, uninformative naming, and dead code.

2. **Holistic change assessment**:
   - Evaluate whether the implementation satisfies the stated requirements and broader intent.
   - Check cross-file and cross-module consistency across the entire repository.
   - Analyze architectural coherence, integration boundaries, downstream contract impacts, and regression risks.
   - Verify test suite adequacy and whether validation steps would actually fail if the underlying logic broke.

Review outputs must explicitly report both passes, even when one has no findings.
An explicit evaluation of the holistic assessment alongside an itemized findings list (or an affirmative clean declaration `No actionable findings identified.`) proves that both dimensions were thoroughly examined.

- **Do:** require reviewers to conduct and explicitly document both a detailed implementation defect audit and a holistic change assessment.
- **Do:** report the holistic assessment explicitly in review outputs, even when no architectural, integration, or regression issues are found.
- **Don't:** accept a review that stops at superficial surface checks without evaluating the systemic impact on requirements, architecture, cross-file consistency, and validation rigor.

## The posted fallback comment is the reviewer's report, not an author composite

When the self-review is posted as a PR comment, the comment body **is**
the dispatched reviewer's structured report, then the required
disclosure marker from
[`disclose-agent-authorship`](disclose-agent-authorship.md).
The marker is forge attribution, not author recap.
This form applies to every posted independent review, not only the fallback case: [Post every independent review on the PR or MR it reviewed](#post-every-independent-review-on-the-pr-or-mr-it-reviewed) points here for it.
Dispatching the reviewer and then writing a different review body is the
same failure as reviewing inline, one step later: the authoring session
still composed the text that readers treat as the review.

Measured 2026-08-25 on
[ai-config#2234](https://github.com/Morrison-Lab/ai-config/pull/2234#issuecomment-5415839535).
A foreground `Task` dispatch
(`bc-61fbadd0-7970-5b2d-8775-4924a28e09a1`, catalog name
"Final review HEAD f71c02ea") ran on `f71c02ea`.
The posted comment was author-assembled, labeled
"Fallback self-review", copied the child's
`### Verdict: Ready for merge` and `Reviewed-Commit` lines, and wrapped
them in a 16-item
"Round history that was Addressed, Rebutted, or Deferred" ledger.
That comment is the wrap, not the parent `Task` JSON.
How Cursor Cloud obtains the child's structured report is in
[`memories/cursor.md`](../../memories/cursor.md).

- **Do:** post the dispatched reviewer's structured report
  (Summary / Findings / Verdict / Reviewed-Commit) as the fallback comment,
  then append the required disclosure marker.
  How Cursor Cloud obtains that report is in
  [`memories/cursor.md`](../../memories/cursor.md).
- **Don't:** wrap the verdict in the authoring session's ARD round-history
  recap in the same comment.
- **Don't:** omit the disclosure marker, or treat that marker as license to
  add an ARD ledger.
- **Don't:** paraphrase a missing reviewer body as Ready for merge.

## The mechanism

[`hooks/no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py) gates the pre-push case on Claude Code, per [`algorithmatize-checks`](algorithmatize-checks.md).
It answers three questions rather than one, because provenance alone is not enough.

*Who said it*: a verdict is admitted from the `tool_result` of an `Agent` call whose `subagent_type` is the reviewer, and only when that result is not an error.
So an inline pass, a verdict quoted out of a file, the guard's own denial message, and a clean report from some other subagent all fail.

There is a **second** admitted provenance, which this paragraph read as the only one until the round-repetition section above was written: a `Bash` call matching the guard's own external-reviewer pattern, which today recognizes `agy --print` and not the other delegation CLIs.
Both statements have to live in one file, so read the paragraph above as the rule for a Claude-side review and this as the rule for the external lane, rather than as two populations of what the guard accepts.

**"A verdict quoted out of a file" has one narrow exception, added rather than relaxed.**
Claude Code sometimes delivers a dispatched subagent's report as a message from the subagent's own `SubagentHandback` call instead of inside the `Agent` tool's own result (ai-config#3945): the tool_result then carries only a pointer sentence and an `agentId`, and the report lives in a sibling `subagents/agent-<agentId>.jsonl` file the guard cannot see by reading the parent transcript alone.
The guard reads that ONE file, located by the `toolUseId` the original dispatch's own call id names (falling back to the `agentId` printed in the pointer sentence), and only after a `.meta.json` beside it independently confirms the subagent's `agentType` is an admitted reviewer -- the same persona check the dispatch itself already had to pass.
See `_handback_report_text` in [`hooks/no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py).
Nothing else about "who said it" changes: a phrase search over any OTHER file, or over this same file located any other way, still fails.

*What it said*: restricting provenance does not make a phrase search sound **inside** the admitted body, which is the same failure one layer in --- a review whose closing note quotes the clean verdict it is withholding would read as clean.
So the verdict is the last line that **is** a verdict line, anchored at line start, and a quotation mid-sentence is not one.

*What it was about*: the reviewer states the commit it read as a `Reviewed-Commit: <sha>` line after its verdict, and the guard resolves what the push would actually ship --- reading the refspec, not just `HEAD` --- and compares.
A push ships commits, so anything that changes what would be shipped --- a later commit, a `main` merge, a rebase, a commit a subagent made in a transcript the guard cannot see, or a branch other than the reviewed one --- fails the comparison.
That is also why the review comes **after** committing, which is where [`ardi`](ardi.md) already puts the pause point.

The other cases have no guard and are prose rules here.

- **Do:** dispatch [`adversarial-reviewer`](../../.claude/agents/adversarial-reviewer.md)
  (foreground, read-only) for the pre-push self-review gate,
  and report which agent produced the verdict.
- **Do:** at merge time, satisfy the separate cross-model, cross-harness
  gate defined under "Cross-model and cross-harness reviews are required
  for merging, and the harness list is concrete" above.
- **Do:** re-dispatch after fixing findings, so the clean verdict describes the tree you are shipping.
  Do not report a HEAD as reviewed until a dispatched review of **that** SHA has returned.
  If a fix already moved HEAD, re-dispatch on the current SHA before the next status report.
  (ai-config#2277, 2026-08-26: addressed two wording nits on `92c65d5c` and reported without a review of that SHA until asked.)
- **Don't:** perform a self-review inline under a reviewer framing --- that is the move this rule replaces, and it is indistinguishable from compliance in the output.
- **Don't:** brief the reviewer with the rationale for the change.
- **Don't:** count a subagent's clean verdict as the external verdict [`fully-clean`](fully-clean.md) requires.
  It is a self-review, performed properly.

## A verdict phrase separated from its heading by a line break is no verdict

`VERDICT_LINE` matches `Verdict[ \t]*:` and the phrase on the **same line**, optionally heading-prefixed (`### Verdict: Ready for merge`).
It does not match a heading naming the section with the phrase on the line that follows it:

```
## Verdict

Ready for merge

Reviewed-Commit: <sha>
```

`parse_report` returns no verdict for a report shaped that way, and `read_latest_review` then keeps whichever verdict it last successfully parsed from an **earlier** dispatch in the same transcript --- so a later, clean, correctly-formatted review can be invisible while an older `needs_work` review from before it stands as "the latest."
The refusal then reads "The latest adversarial self-review returned a blocking verdict," which is true about the parsed history and false about what the session actually did.
It is easy to misdiagnose as the transcript lagging a same-turn dispatch --- a plausible-sounding mechanical explanation that was never actually verified, and that ai-config#2444 was originally filed on before this parsing gap was found instead.

[`.claude/agents/adversarial-reviewer.md`](../../.claude/agents/adversarial-reviewer.md) already specifies the one-line form for its own persona.
The gap is any other brief that asks something to act as an adversarial reviewer --- a same-vendor fallback subagent, a CLI dispatch, a hand-written prompt --- without repeating that requirement.

- **Do:** put `Verdict: <phrase>` literally on one line in every review brief you compose, whatever is dispatching it --- `### Verdict: Ready for merge`, never a heading with the phrase on its own following line.
- **Do:** when a guard refuses a push on a verdict you believe is clean, run the guard's own reader (`read_latest_review`/`parse_report`) over the live transcript and print what it parsed per admitted result, before attributing the refusal to a cause like transcript lag.
- **Don't:** treat a heading form (`## Verdict` with the phrase on the next line) as equivalent to the required one-line form --- it parses as no verdict at all.
- **Don't:** file or accept a "the transcript lags the current turn" diagnosis for a refused push without first executing the parser against the actual transcript;
  the two failures produce an identical refusal message.

(ai-config#2444, 2026-08-27: filed on the lag diagnosis, which running `read_latest_review`/`parse_report` directly against the session transcript then refuted --- it returned the older `needs_work` verdict from a mid-session dispatch rather than a stale read of a same-turn one.
The issue's body was rewritten afterwards to lead with the corrected diagnosis and keep the lag theory behind a marked `<details>` block, so read it as the corrected account rather than the filed one.)

**A verdict line that is absent altogether falls into the same trap, not a different one.**
Measured 2026-09-04: a dispatched reviewer returned a full report ending "No findings." plus a `review-data` JSON block reading `"verdict": "CLEAN"`, with no `### Verdict:` line anywhere in the report.
`read_latest_review` found nothing to parse from this dispatch and kept the **previous** round's `needs_work`, so the guard refused the push reporting a blocking verdict over a review that had found nothing.
The fix is the same one this section already gives: state the required line explicitly in the brief, as a literal `### Verdict: Ready for merge` outside any code fence or HTML comment, and require the `review-data` payload to agree with it --- the two representations disagreeing (a `### Verdict: Ready for merge` line paired with a `review-data` payload naming findings) is itself a defect in the report, per this file's "Structured review data" section below.

- **Do:** treat a report with no verdict line at all as the identical failure to a heading-separated one --- both leave the guard holding a stale prior verdict.
- **Do:** leave the verdict's wording to the persona, or quote its phrases (`Ready for merge`, `Needs more work`) exactly when a brief has to mention them.
- **Do:** dispatch one reviewer per repository, and push each repository before dispatching the next review.
- **Don't:** assume a report that "sounds clean" (ends in "No findings.", carries a clean JSON payload) discharges the guard without the literal verdict line the parser requires.
- **Don't:** ask the reviewer for a verdict in your own vocabulary ("end with clean / not clean") --- the brief overrides the persona's format, the reviewer answers `### Verdict: clean`, and that parses as no verdict.
- **Don't:** review two repositories in one dispatch --- `parse_report` returns one `(verdict, Reviewed-Commit)` pair per report, so one report cannot clear both pushes.

(Measured 2026-09-24 on Morrison-Lab/mlg#41 and Morrison-Lab/mln#92: two clean rounds reporting `### Verdict: clean` were invisible to the guard, which kept an earlier `Needs more work`.
Both briefs asked for "clean / not clean" and covered both repositories.
One dispatch per repository, left to the persona's own format, cleared each once it was the latest verdict.)

**A separate, real constraint: the guard tracks one global latest verdict, not one per branch.**
`read_latest_review` scans the whole transcript and keeps overwriting a single `(verdict, reviewed_commit)` pair with whatever it parses next, with no branch scoping at all.
Reviewing branch A (clean, commit `X`) and then branch B (clean, commit `Y`) leaves `Y` as the global "latest" pair;
pushing branch A afterward compares its shipped commit `X` against the held `Y`, fails the SHA match, and refuses citing an unreviewed commit --- even though branch A's own review was genuinely clean.

- **Do:** review and push one branch before dispatching a review for a second branch in the same session, when driving more than one branch's push through this guard.
- **Don't:** read that refusal as a defect in branch A's review;
  the guard has no notion of "branch" to be defective about, and the SHA comparison is doing exactly what it is built to do.

**The remedy above is about pairing, not about ordering, and reading it as a sequencing preference is what lets the refusal happen anyway.**
What has to hold is that each dispatch is followed by its own push before anything else is dispatched.
Interleaving two branches at that granularity --- dispatch A, push A, dispatch B, push B --- satisfies it and converges fine.
What breaks it is a push *deferred* past the next branch's dispatch, which is easy to do without deciding to: a push waiting on a checker re-run, on a finding still being addressed, or on a report being written is a push that has not happened yet, and the next branch's round proceeds in the meantime.
So the operative question at each dispatch is not which branch to review next but whether the previous branch's verdict has already been spent.

**When one has been overwritten, the sanctioned override is the correct discharge, and re-dispatching is the expensive mistake.**
This file's own "What \"separate\" requires" section already draws that scope: `ALLOW_UNREVIEWED_PUSH=1` "covers a push whose verdict the guard cannot check, not only a push with nothing to check".
An overwritten slot is exactly the first case.
A genuine clean verdict for the exact commit was produced and is simply no longer the pair the guard holds, so the override reports the situation accurately rather than papering over an unreviewed push.
What licenses it is the mechanical evidence, not the recollection: run the guard's own `read_latest_review`/`parse_report` over the transcript, as this file's "A verdict phrase separated from its heading by a line break is no verdict" section already requires of any refusal you believe is wrong, and paste what it parsed alongside the retained report's own `Reviewed-Commit:` line.
An amend or a fixup between the review and the push is enough to make a confident narrative false.
Re-dispatching instead spends a full adversarial pass --- **about 125k subagent tokens for one round over a four-file, 105-line prose diff, measured 2026-09-03** --- to re-derive a verdict that already existed for that exact SHA, and lands in the slot the other branch will need next.

Locating the defect in the guard is the reading the pair above rules out, and it rules that reading out for the wrong reason.
It is right that the guard has no notion of branch to be defective about;
what it also has no notion of is *commit*, beyond the single most recent one.
Keying verdicts by SHA --- `{sha: verdict}` rather than one latest-verdict slot --- removes this failure, and subsumes [#3131](https://github.com/Morrison-Lab/ai-config/issues/3131)'s original report (a sibling subagent's verdict leaking into the pushing thread) without anyone having to reason about which session produced a given verdict.

- **Do:** push each branch on its own verdict before dispatching the next branch's review, treating a deferred push rather than an interleaved branch as the thing to avoid.
- **Do:** use the sanctioned override when a verdict for the exact commit was produced and overwritten, pasting the parser's output over the retained report rather than asserting the SHA from memory.
- **Don't:** re-dispatch to refill the slot;
  it costs a full pass and the verdict it buys is the one the next branch's round overwrites.
- **Don't:** read the refusal as saying the branch is unreviewed --- it says the guard is not holding that branch's verdict, which is a different claim.

(Measured 2026-09-03 across two worktrees in one session, recorded in [#3156](https://github.com/Morrison-Lab/ai-config/issues/3156).
Branch A reviewed clean;
branch B then reviewed not-clean, was fixed, and re-reviewed clean;
pushing A was then refused with "The clean verdict is for commit <B's sha>, but this push would ship <A's sha>", over a clean verdict for A's exact SHA that had been overwritten.
The reverse happened earlier in the same session.)

**A verdict from a resumed reviewer session may never reach the slot at
all, which is a narrower and less certain claim than the section above.**
The provenance chain `read_latest_review` walks is built from fresh
`AGENT_TOOLS` dispatch tool_use/tool_result pairs (or a retrieval call whose
`task_id` matches one such dispatch's own registered id) --- see the guard's
own `_is_reviewer_dispatch` and `TASK_OUTPUT_TOOLS` handling in
[`hooks/no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py).
A session that instead resumes an existing reviewer agent (sending it a
follow-up message through a mechanism other than a fresh `Agent`/`Task`
dispatch or a task-id-linked retrieval) is not obviously covered by that
chain, and one observed session saw exactly the symptom this predicts: a
resumed agent corrected its own previously-fabricated sha in its reply, and
the guard's next push attempt still cited the old, wrong value.
This half is **not** independently reproduced against the guard the way the
overwrite mechanism above is --- it is recorded as consistent with reading
the provenance code, and as matching one observed incident, rather than as
a traced execution.
This is a distinct failure from the overwrite above, so the sanctioned
override does not answer it the same way: an overwritten slot still holds
*a* genuine parsed verdict for a different commit, while a resumed-agent
correction may never have been parsed into the slot at all, and there is
nothing on record to paste over the guard's retained report in that case.

- **Do:** treat a correction from a resumed reviewer as unconfirmed until a
  fresh dispatch (or a push attempt) shows the guard picked it up.
- **Do:** re-dispatch fresh rather than resuming, when a prior review needs
  correcting and the correction must reach this guard.
- **Don't:** treat this as a confirmed guard defect on the strength of one
  observed incident and a code reading; file it for someone to trace with an
  actual resumed-session transcript before hardening the guard against it.

**The harness appends an `agentId:` trailer to a subagent's report, sometimes as its own block and sometimes concatenated onto the last line.**
Which of those is common is the question this section could not settle, and an earlier draft asserted an answer to it by generalizing from the two dispatches it happened to watch.

Measured 2026-09-02 over 565 stored session-transcript JSONL files, matched by a flat per-session glob.
A recursive walk of the transcript tree also reaches each session's nested subagent transcripts and roughly doubles the population.
The scope is stated because a fragment arguing "measured rather than assumed" should say what it measured:

```
trailer as its own content block : 334
trailer concatenated onto text   :   0
```

An independent re-run over the recursive set returned that same zero.

The concatenated form is known to exist.
It was seen directly, twice, within one session --- and the captured line is this:

```
--- end of report ---agentId: <id> (use SendMessage with to: '...')
```

**Note what that exhibit is: the trailer landing on a sentinel line, which is the safe case**, and the one the block below records as the mitigation [#3050](https://github.com/Morrison-Lab/ai-config/issues/3050) rejected.
The shape the hazard is actually about is a fingerprint line with the trailer glued to it:

```
Reviewed-Commit: <sha>agentId: <id> (use SendMessage with to: '...')
```

That second block is constructed to show the shape, not captured.

**The zero and the sightings are about different artifacts, and that is what has to be settled before either number means anything.**
The sweep read **stored** transcript JSONL.
The two sightings were **in-context renders**.
Three explanations fit, and the sweep as run distinguishes none of them:

1. The sweep's matcher cannot see the concatenated shape.
2. Those transcripts sit outside the tree it walked.
3. The concatenation is a render artifact that never reaches storage at all.

If the third holds, the zero is a **true** negative and the hazard does not reach the guard, because the guard reads storage: `read_latest_review` opens the transcript file and `_result_text` flattens its stored content blocks.
Under the first two the zero is uninformative, and reading it as confirmation would be the failure [`verify-the-right-artifact`](verify-the-right-artifact.md) names --- "running the thing and seeing no complaint feels like a test", and is not one, because a surface that silently ignores what it cannot see is quiet for the same reason a working one is.

So the two counts carry different weight.
The **334** establishes that the own-block form is common **in stored transcripts**, which is a lower bound rather than a ratio.
The **0** settles nothing on its own, since which of the three explanations holds decides whether it is evidence or an artifact.
Whoever next touches this section should settle that first, and the query has to be chosen carefully, because the obvious one cannot decide it.
Grepping a stored transcript for a `Reviewed-Commit:` line with a non-hex suffix cannot fire on a **conforming** report, which never puts the fingerprint last, so nothing can be glued to it.
That reason quantifies over conforming reports only, and this section is about **reordered** ones --- so the reason does not establish the null it appears to, and a zero from the grep stays uninformative either way.
The right *criterion* is whether `agentId:` ever appears in stored content **preceded by other text on the same line**.
The right *instrument* is not the sweep, and this is the part that is easy to get wrong: re-running the sweep's own matcher has no power against explanation 1, which says precisely that this matcher cannot see the shape --- under that explanation it returns zero whatever it is pointed at.
It is also already spent against the other two, since the counts above are that query, run twice, over both the flat set and the recursive superset.

So settling this needs something the section has not yet had: a **raw text scan** of the stored JSONL, written independently of the sweep, or a tree the sweep never reached.
Only a hit settles anything.
Another zero from the same matcher is the same uninformative zero, so do not read one as evidence for the rendering-quirk explanation.

In the own-block shape the trailer is a separate content block, and `_result_text` joins blocks with a newline, so the fingerprint line is untouched and the hazard does not arise at all.

**Two conditions have to hold together before any of this can bite, and a conforming report fails the second.**
The trailer must concatenate rather than arrive as its own block, AND the fingerprint must be the report's last line.
This file's own contract puts the JSON payload last, and [`.claude/agents/adversarial-reviewer.md`](../../.claude/agents/adversarial-reviewer.md) says to emit nothing after its closing marker --- so a conforming report ends with the payload, the trailer lands on that, and the fingerprint is never exposed.
The sighting that prompted this section was a report that put the verdict and fingerprint AFTER the payload, which is the ordering this file's "Structured review data" section rules out.
Whether the two sightings were two such reports or one report read twice is not recorded, so treat the shape as attested and its rate as unmeasured.

So read the rest of this section as what happens when a brief reorders the tail, not as a hazard of ordinary dispatch.

**Even then, for the mandated 40-character form nothing happens, and saying otherwise would be the easy overclaim here.**
`no-push-without-self-review.py`'s `REVIEWED_COMMIT` captures `([0-9a-fA-F]{7,40})`, so a full sha stops the capture exactly at the boundary and the suffix is never reached.
Run through the guard's own regex:

| fingerprint | captured |
|---|---|
| 40 chars, then `agentId: a3f5...` | the correct sha |
| 39 chars, then `agentId: a3f5...` | 40 chars ending in the `a` of `agentId` --- a **wrong sha, silently** |
| 7 chars, then `agentId: a3f5...` | 8 chars, wrong |

So the hazard is real and it is a **truncation** hazard rather than a suffix hazard.
`agentId` begins with a hex character, which is what turns a short fingerprint into a plausible-looking wrong one instead of a parse failure.
A wrong sha refuses the push with "the clean verdict is for commit X, but this push would ship Y" --- a message that reads as a stale verdict and is nothing of the kind.

The remedy below costs one line and reads as cheap insurance rather than a fix for a demonstrated break at 40 characters, which is how it was first written up here.
The block below is a report's TAIL, not a whole report.
It shows the REORDERED ordering described above --- verdict and fingerprint after the payload --- which is the shape the sentinel exists for and the shape this file's "Structured review data" section rules out.
So read the block as the shape the sentinel was proposed for, not as a template to copy.
[#3050](https://github.com/Morrison-Lab/ai-config/issues/3050) has since settled which of the two orderings a brief should mandate, and it settled it against the block below: the payload goes last, a conforming report carries no sentinel, and the fingerprint is the full sha.
The block stays as the record of the mitigation that decision rejected, so a reader meeting a sentinel in the wild can tell what it was for.
It is not a shape to copy, and it is not the remedy for a reordered tail either --- the full sha is, under every ordering.

```
-->
Verdict: <phrase>
Reviewed-Commit: <full sha>
--- end of report ---
```

The leading `-->` is the JSON payload's own closing marker, included so the ordering the prose describes is visible in the block rather than asserted over it: everything shown sits *after* the payload, which is what makes this a reordered tail and not a conforming one.

The sentinel puts a non-hex line between the fingerprint and anything the harness appends, so the fingerprint's length stops mattering.

It is not free, though, and both its cost and its protection belong to the **consumer** reading the report rather than to the report's shape.
Its cost falls unevenly across the two report contracts [`scripts/pre-push-review.py`](../../scripts/pre-push-review.py) validates,
and falls nowhere at all on [`hooks/no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py)'s own guard or on [`scripts/cursor-self-review-check.py`](../../scripts/cursor-self-review-check.py), the Cursor Cloud recovery gate that calls that same `parse_report`, neither of which runs a trailing-content check on either shape.
Those three are the whole consumer set as of 2026-09-04, derived rather than recalled: `grep -rln parse_report hooks scripts --include='*.py' | grep -v test`.
On that script's own four-section contract --- Summary Verdict, Critical Findings, Observations, Verification Steps --- it strips HTML comments and then scans whatever follows the last fingerprint.
What `_TRAILING_AFTER_FINGERPRINT` admits there is a status banner, an `=` rule, a disclosure footer, a stopping-point line, or a restated verdict.
`--- end of report ---` is none of those, so a sentinel is refused outright with "Reviewed-Commit fingerprint must be at the very end of the report".
A persona-contract report never reaches that scan.
`parse_review_verdict` routes it to `_parse_persona_verdict`, which runs no trailing-content check at all, so a sentinel there is tolerated rather than refused.
Measured 2026-09-03 through `parse_review_verdict` itself rather than against the regex in isolation, which is the substitution [`verify-the-right-artifact`](verify-the-right-artifact.md) rules out:
a persona report returned `(True, True, 'Clean (persona contract)')` payload-last, with a sentinel appended, and with the tail reordered alike, while the same sentinel on the local contract returned `(False, False, 'Reviewed-Commit fingerprint must be at the very end of the report.')`.
Every one of those reports carried the full sha, so they establish that the sentinel is accepted on the persona contract, not that it is unnecessary there.

That is refusal on one contract and tolerance on the other, and tolerance is not inertness.
Measured 2026-09-04 on the persona contract, over a report whose fingerprint was abbreviated to seven characters, whose tail was reordered so that fingerprint was the report's last line, and which carried an `agentId:` trailer glued to it:
without the sentinel `parse_review_verdict` returned a `Fingerprint SHA mismatch` refusal naming `b9dc14ba` against an expected sha beginning `b9dc14b0`, and with the sentinel sitting between the fingerprint and the trailer it returned `(True, True, 'Clean (persona contract)')`.
The same abbreviated, reordered, glued report, measured the same day against the pre-push guard's own `parse_report`, returned `('clean', 'b9dc14ba')` without the sentinel and `('clean', 'b9dc14b')` with the sentinel in that same position --- for the four-section report shape as well as the persona one, since that guard checks no trailing content under either.
`verify_review` then tests `c.startswith(reviewed_commit)`, so against a real `b9dc14b0...` commit the sentinel-free form refuses the push with the misparsed-sha message and the sentinel form passes it.
So the sentence above about the fingerprint's length ceasing to matter is right wherever no trailing-content check runs --- `pre-push-review.py`'s persona contract, the pre-push guard's own `parse_report`, and the Cursor Cloud recovery gate alike --- over a fingerprint that is both abbreviated and last.
What settled [#3050](https://github.com/Morrison-Lab/ai-config/issues/3050) against the sentinel is therefore not that it protects nothing.
It is that a conforming report never reaches the situation it protects: payload-last means the fingerprint is not the last line, and the mandated full sha stops `REVIEWED_COMMIT`'s capture at the boundary under either trailer shape.
What is left is the cost --- outright refusal on the four-section contract --- which a report gains nothing by taking on.

Two caveats.
The concatenation has been observed on Claude Code's `Agent` tool and nowhere else, so it is a claim about that harness on that date rather than about subagent dispatch generally.
And the **reordering** is what sits in tension with [`.claude/agents/adversarial-reviewer.md`](../../.claude/agents/adversarial-reviewer.md)'s own instruction to emit nothing after the JSON payload's closing `-->`.
The sentinel is not the source of that tension and does not add to it: a brief that puts the verdict and the fingerprint after the payload has already overridden the emit-nothing instruction, and the sentinel then joins a tail that exists either way.
The ordering the guard parsed successfully was that reordered one --- which is a statement about the guard, not an endorsement, since the same ordering fails this file's payload-last contract.
That decision is made, per #3050: a brief mandates the payload-last ordering, so a conforming report never puts the fingerprint last and the tension never arises in one.
The parser is what makes payload-last free rather than merely tidier --- `parse_report` blanks HTML comments before both of its regex searches, so a payload sitting last can neither supply nor displace the verdict LINE or the fingerprint.
Measured 2026-09-03 by running `parse_report` over a report whose Markdown fingerprint and payload `commit_sha` named different commits;
it returned the Markdown one.
That is a claim about those two searches and nothing wider.
The payload is read separately, from the raw text rather than the blanked copy, and it is authoritative: a payload listing any finding downgrades a clean verdict to `needs_work`.
Measured the same day --- a payload-last report whose Markdown line read `### Verdict: Ready for merge` and whose payload carried one finding parsed as `('needs_work', <sha>)`, and parsed as `('clean', <sha>)` once that findings array was emptied.
So a reviewer gains nothing by letting the two representations disagree, which is the failure the payload exists to catch.
The decision does not rest on how often the harness concatenates its trailer, so the re-measurement #3050 wanted first would not move it:
the full sha closes the truncation hazard under either trailer shape, and payload-last keeps the fingerprint off the report's last line however the trailer arrives.
It is adjacent to [#2483](https://github.com/Morrison-Lab/ai-config/issues/2483) and not the same item: that issue is about verdicts arriving via background task notifications going unregistered.

- **Do:** mandate the payload-last tail in every review brief you write --- verdict, then fingerprint, then payload, and nothing after it.
- **Do:** state the fingerprint as the **full 40-character** sha, which is what actually protects it.
- **Do:** instruct the reviewer to **derive its own fingerprint** with `git
  rev-parse HEAD` in its own worktree, and to confirm it resolves with `git
  rev-parse --verify --quiet <sha>^{commit}` before writing the
  `Reviewed-Commit:` line, rather than trusting the sha handed to it in the
  brief.
  This is a second, independent layer under the full-sha rule above, not a
  restatement of it: it catches an abbreviated sha the brief-writer sent by
  mistake (the reviewer's own `rev-parse` returns the correct full value
  regardless of what it was told), and it catches a reviewer that would
  otherwise transcribe an abbreviation's visible prefix and invent the rest
  to reach 40 characters --- a fabrication a length check alone cannot see,
  since the result is a well-formed 40-character hex string that simply does
  not exist.
  Every dispatch in one sweep briefed this way returned a correct
  fingerprint (Morrison-Lab/ai-config#3295, 2026-09-05); one dispatch briefed
  with an abbreviated sha and no derive instruction returned a fabricated
  tail whose first 8 characters matched the abbreviation it had been given
  and whose remaining 32 did not correspond to any real commit.
- **Do:** read a "verdict is for commit X, but this push would ship Y" refusal as possibly a *misparsed* fingerprint rather than only a stale one --- print what the guard captured before concluding.
- **Do:** fix the brief rather than keeping a sentinel you meet in the wild.
  It does protect an abbreviated fingerprint that is the report's last line, on every consumer that runs no trailing-content check on the shape it is handed --- `pre-push-review.py`'s persona contract, the pre-push guard's own `parse_report`, and [`scripts/cursor-self-review-check.py`](../../scripts/cursor-self-review-check.py), the Cursor Cloud recovery gate, which calls that same `parse_report` and then compares prefix-tolerantly.
  Measured on that gate 2026-09-04, over the same abbreviated, reordered, glued report: without the sentinel it printed `REFUSE: fingerprint b9dc14ba does not match expected head b9dc14b0...` and exited 1, and with the sentinel sitting between the fingerprint and the trailer it printed `PASS: clean verdict at the expected head` and exited 0.
  The payload-last full-sha tail removes those situations instead of insuring against them.
- **Don't:** write a brief that puts the verdict and the fingerprint after the payload, or that asks for a trailing sentinel.
- **Don't:** reach for the sentinel as insurance --- a conforming report puts the payload last, so the fingerprint is never the last line, and the full sha closes the truncation hazard under either trailer shape.
  It is refused outright on [`pre-push-review.py`](../../scripts/pre-push-review.py)'s four-section contract, whose trailing-content check admits no such line and which runs only when an expected sha is supplied --- as that script's own `main()` always supplies one.
- **Don't:** read the full-sha rule as mechanically enforced.
  `REVIEWED_COMMIT` still accepts `[0-9a-fA-F]{7,40}`, and [`pre-push-review.py`](../../scripts/pre-push-review.py) still clears a seven-character prefix match, so an abbreviated fingerprint is caught by nothing but the reviewer following the brief.
- **Don't:** claim the suffix breaks a 40-character fingerprint;
  run `REVIEWED_COMMIT` over the line before asserting either way.
- **Don't:** abbreviate the sha in a review brief's template, which is the input that turns the suffix into a silently wrong parse.
- **Don't:** let a reviewer transcribe its `Reviewed-Commit:` from the sha
  the brief handed it.
  That is the anti-pattern paired with the derive-your-own-fingerprint
  bullet above, and it is the half a reader skimming only the `Don't`
  list would otherwise miss.
  Transcription looks identical to derivation in the finished report ---
  both produce a 40-character hex string in the right place --- so the
  brief is the only place the difference can be established.
- **Don't:** read the sentinel as part of the payload-last contract.
  It is a mitigation for the ordering that contract rules out, so a conforming report needs none.

**Confirmed again, 2026-09-10, with the persona's existing "read that sha
yourself" instruction already in place and still not enough on its own.**
A dispatch against an unpushed commit reported `Reviewed-Commit:
b7f1d0c62d3a83c98d0cc4d17ac8ea7dbfe1ff67`, matching the real commit
(`b7f1d0c3f459eb7ac75b4453470d3e5b8046c298`) in its first 7 characters and
disagreeing in the remaining 33.
The persona file already carried "Read that sha yourself rather than taking
it from the brief," which names the *source* to avoid but not the *method*
that avoids it --- it does not say to run `git rev-parse HEAD` specifically,
and it does not forbid extending a short sha it already has (from `git log
--oneline`, or from the hook's own error text) out to 40 characters.
The review's surrounding content was independently verified and sound,
which is what made the fabricated fingerprint easy to miss: nothing else in
the report read as unreliable.
The pre-push guard's prefix-tolerant compare still caught it, because the two
strings share only 7 characters and neither is a prefix of the other beyond
that point, so `verify_review`'s mismatch check fired as designed.
The in-session fix was to re-dispatch with the real full sha supplied and an
explicit "do NOT invent or pad any SHA; report only a SHA a command you ran
printed in full, and paste that command's output," which held for every
later round.
[`.claude/agents/adversarial-reviewer.md`](../../.claude/agents/adversarial-reviewer.md)
now carries that instruction directly, so a future dispatch does not depend
on the brief-writer remembering to add it.

- **Do:** read this as confirmation that "read the sha yourself" needs the
  method spelled out (`git rev-parse HEAD`, verified with `git rev-parse
  --verify --quiet <sha>^{commit}`) and an explicit padding ban, not as a
  reason to distrust the prefix-tolerant guard --- the guard worked.
- **Don't:** treat a reviewer's otherwise sound, well-evidenced findings as
  proof its fingerprint is real; the two are independent, and a fabricated
  identifier can sit inside an accurate report undetected until something
  else (here, the guard) compares it.

**A third occurrence, `Lacaedemon/sparta`, 2026-09-20 --- relayed from the
session's own report rather than independently reconstructed (see the
closing note below), and worth recording anyway because it repeats the
FIRST occurrence's exact input mechanism after the fix aimed at the SECOND
occurrence was already shipped.**
A reviewer briefed with `d4691095` reported `Reviewed-Commit:
d469109590de1c1f8b4a5b8e5b4a6b3f8a9e0d4f`.
The real commit is `d46910953d5ea64ee58a295dd56a6f2b1ac55d80` --- the two
strings agree on the first 8 characters and diverge for the remaining 32,
matching the #3295 split above exactly (that occurrence's own abbreviated
sha, with no derive instruction, produced a fabricated tail whose first 8
characters matched the abbreviation and whose remaining 32 did not
correspond to any real commit) and differing from the 2026-09-10 pair,
which shares 7 characters and diverges for 33.
So the input mechanism here is not new: a brief supplying only a short
prefix is exactly what produced the very first occurrence, in violation of
the "Don't abbreviate the sha in a review brief's template" bullet above.
What is new is that it recurred after the OTHER remedy --- instructing the
reviewer to derive its own fingerprint rather than trust the brief --- was
already added to the persona file in response to the second occurrence.
That fix addressed the reviewer's half of the contract and left the
brief-writer's half exactly where the first occurrence found it: a persona
told to derive its own fingerprint was evidently still willing to pad the
one it was handed rather than discard it and run `git rev-parse`.

- **Do:** treat a brief that abbreviates the sha as the first thing to fix,
  before asking why the reviewer fabricated --- the derive-your-own
  instruction is a second layer under the full-sha rule, not a replacement
  for it, and this occurrence shows the first layer failing on its own.
- **Don't:** read the persona-file fix from 2026-09-10 as closing this
  failure mode; a brief-writer that abbreviates the sha is a distinct
  recurring point of failure the persona file cannot reach.

**A fourth failure shape, same session and relayed the same way (see the
closing note below) rather than independently reconstructed, with no
fabrication in it: the reviewer resolved a ref that had already moved.**
A review resolved `origin/<branch>` for its fingerprint while the local
branch already carried two commits not yet reflected in that remote ref, and
reported a blocking finding that those two commits had already addressed.
This is not the sha-fabrication failure above --- the fingerprint the
reviewer reported was a real, correctly-transcribed commit, just not the one
the push was about to ship.
It is the `git rev-parse HEAD`-in-its-own-worktree instruction succeeding at
exactly the wrong scope: `HEAD` in a worktree that has not fetched is a
faithful answer to "what does this worktree currently point at," and a
faithful answer to the wrong question.
- **Do:** brief a reviewer to `git fetch` (or otherwise confirm its working
  copy is current) before resolving its own fingerprint, not only to derive
  the fingerprint from whatever ref is already checked out.
- **Don't:** treat "the reviewer read its own `rev-parse` output" as
  sufficient; a correctly-derived fingerprint for a stale ref reports a real
  sha and a wrong verdict.
- **Don't:** read a stale-ref finding as evidence the guard's sha-comparison
  failed --- it did its job (the reported commit does not match what the
  push ships); the miss is upstream, in what the reviewer resolved before it
  ever wrote the fingerprint line.

(Both measured on `Lacaedemon/sparta`, 2026-09-20, during a session driving
several open PRs; the exact review transcripts were not preserved, so the
narrative above is relayed from the session's own report rather than
re-derived from a saved artifact.
The real commit's identity is independently confirmed here: `git rev-parse
d4691095` on that repository returns
`d46910953d5ea64ee58a295dd56a6f2b1ac55d80`.)

## Structured review data (JSON payload)

Every reviewer emits two representations of one verdict: the human-readable Markdown report, then a machine-readable JSON payload in a trailing HTML comment.

**"Every reviewer" means every review you post, not only one a dispatched persona wrote.**
The paragraphs around this section are mostly about a review dispatched to the [`adversarial-reviewer`](../../.claude/agents/adversarial-reviewer.md) persona before a push,
so the requirement reads as that persona's rather than as a rule about reviews.
It binds every review you post or produce: a forge comment, a local report, a review composed in-transcript.
The one you are likeliest to write without a payload is the review nobody dispatched and no push follows --- a review-only request on somebody else's PR --- since neither the persona nor the pre-push guard is in play there.

Nothing goes red when the payload is omitted, which is why this needs stating rather than more care.
Such a review still yields a verdict its consumers read, since Markdown parsing is every one of their fallbacks when no payload is present.
What a missing payload loses is the per-finding records, which no prose pattern recovers.
The near-miss is a thorough report --- findings reproduced, locations cited, the disclosure marker appended --- where every standard a human reader can see is met and the machine-readable half is the one absent.

The Markdown half owes a verdict line the consumers' own patterns match, and the forms are not interchangeable.
Measured 2026-09-02 against the three consumers, on `Ready for merge`:

| form | `check-pr-fully-clean.py` | `parse_report` | `enforce-mwc-review-gate.py` |
| --- | --- | --- | --- |
| `Verdict: <phrase>` | reads | reads | **no verdict** |
| `### Verdict` then `**<phrase>**` | reads | **no verdict** | reads |
| `### Verdict: <phrase>` | reads | reads | reads |

So write `### Verdict: Ready for merge` or `### Verdict: Needs more work` --- the heading and the phrase on one line, which is the only form all three read.
The phrase matters as much as the form: the pre-push guard accepts `Ready for merge`, `Needs more work`, and `Needs work` and nothing else, so `Verdict: Clean`, `Verdict: Approved`, and `Verdict: Ready` return no verdict there while `check-pr-fully-clean.py` accepts all of them.

- **Do:** append the payload to any review you post or produce, including one nobody dispatched and one no push follows.
- **Don't:** read this section as binding the persona alone --- the persona is where it is already implemented, not where it applies.
- **Don't:** treat a thorough, well-cited, correctly-disclosed report as complete without it;
  that combination is exactly what the omission looks like from the inside.

```html
Reviewed-Commit: <sha>

<!-- review-data:
{
  "schema_version": "1.0",
  "reviewer": "<agent/bot name>",
  "commit_sha": "<full sha>",
  "verdict": "CLEAN",
  "findings": []
}
-->
```

For a not-clean verdict, set `"verdict": "NOT_CLEAN"` and give `"findings"` one object per finding, each with the four keys `file`, `line`, `category`, and `message`.
State those keys in any brief you write, rather than only asking for "finding objects" --- a reviewer that guesses the key names produces `structured finding in unknown: ` as the reported blocking reason.

Name the target repository's schema version in that same brief when it differs from the template above, and its required fields with it --- a version string alone still leaves the reviewer emitting a payload that does not conform to the target's contract, which nothing there validates and so nothing there catches.
A review you compose yourself has no brief to carry that, so read the target's own reviewer prompt before copying the template: `Morrison-Lab/gha`'s asks for `1.1` and two fields this template does not have.
Nothing here reads `schema_version`, and this corpus emits `1.0` --- in that template, in both `adversarial-reviewer` persona files, and in `pre-push-review.py` --- while `Morrison-Lab/gha`'s reviewer prompt requires `1.1` with `detailed_assessment` and `holistic_assessment` fields (measured 2026-09-02, [ai-config#3006](https://github.com/Morrison-Lab/ai-config/issues/3006)).
A reviewer left to copy the template emits `1.0` into a repository asking for `1.1`.

Three rules govern how the payload is read, and each exists because its absence inverted a verdict:

- **The payload must be last, and the last one wins.**
  Last among `review-data` blocks, that is --- [`disclose-agent-authorship`](disclose-agent-authorship.md) still ends a posted comment with its marker, and the two do not conflict: `extract_structured_review` reads a payload the marker follows (verified 2026-09-02), and [`pre-push-review.py`](../../scripts/pre-push-review.py) already emits the report in that order.
  The authoritative payload follows the verdict and the `Reviewed-Commit` fingerprint.
  A reviewer who quotes the template above (it hardcodes `"verdict": "CLEAN"`) before writing its own would otherwise publish a `NOT_CLEAN` review that scored clean.
- **A payload inside a code region does not count.**
  Fences, inline code spans, and indented blocks are all excluded, so a comment that merely mentions the format is not a review of anything.
  This is the same rule `check-pr-fully-clean.py` already applies to quoted *finding* vocabulary (ai-config#2449), applied to a *verdict*.
- **Findings block regardless of the stated verdict.**
  A payload that enumerates findings and then labels itself `CLEAN` is contradicting itself, and the safe reading of a contradiction is the blocking one.

All three consumers read the payload through one extractor, [`scripts/lib/review_payload.py`](../../scripts/lib/review_payload.py): [`scripts/check-pr-fully-clean.py`](../../scripts/check-pr-fully-clean.py) for a comment posted to a PR, [`scripts/pre-push-review.py`](../../scripts/pre-push-review.py) for a report produced locally, and [`hooks/no-push-without-self-review.py`](../../hooks/no-push-without-self-review.py) for the pre-push guard.
They score the same artifact, so they must agree, and one extractor keeps them in agreement: a report whose payload says `NOT_CLEAN` or lists findings scores blocking across all three.
Markdown parsing remains the fallback when no payload is present.

## The review gates the push, not the work --- and it is one round, not a loop

The rule above is a gate on a **push**.
It is not a rule that the work must stop moving until the reviewer is happy,
and reading it that way turns one gate into an unbounded loop.

The failure runs like this.
A review returns findings, you fix them, and the fix needs its own review ---
correctly, since a fix is a diff nobody has read.
So far so good.
The trap is treating each round's fix as something that must clear review
*before the branch can be pushed at all*, because that condition never
arrives: every round produces new code, new code owes a review, and the
commits pile up locally while the loop runs.
Measured on [ai-config#1911](https://github.com/Morrison-Lab/ai-config/pull/1911):
five rounds on one file, every round finding something real, and nothing
pushed for hours.

It reads as rigour from the inside, which is what makes it worth a rule.
Each individual decision to hold is defensible, and the loop is invisible
because no single round is the one that went wrong.

**Pushing is not merging, and the costs run the other way.**
A pushed branch is where CI and the repo's own reviewer can see the code, so
pushing *adds* scrutiny rather than skipping it.
Holding subtracts it, and adds costs of its own: the work is invisible to
other sessions, it is unbacked-up, and its merge conflict with a moving base
grows the whole time.

So the gate is per push and the review is of what that push ships.
Under harnesses without a strict local push guard, once a round's fix is verified ---
its own tests, its own mutation control, its own measurement ---
push it, and let the next review run against the pushed head, which is the head that matters.

**Under Claude Code**, however, the `no-push-without-self-review.py` guard
enforces the loop strictly: a push is rejected by default if the latest review on the branch returned findings.
The primary remedy is not to bypass the guard by pushing mid-round (which the guard blocks without an override),
but to keep each round's scope strictly minimal.
Address the findings, get a clean verdict on that focused diff, and push immediately.

- **Do:** (Non-Claude harnesses) push a verified round and let the reviewer read the pushed head.
- **Do:** (Claude Code) keep the scope of each fix round small so you can achieve a clean verdict quickly, and push the verified round as soon as you obtain a clean local verdict on it.
- **Do:** batch a round's fixes into one push, per
  [`efficient-pr-babysitting`](efficient-pr-babysitting.md), rather than
  trickling or hoarding.
- **Don't:** hold a branch until some future round returns clean if the harness allows pushing ---
  that condition recedes with every fix.
- **Don't:** read a gate on pushing as a gate on shipping any of the work.

(Directive from the user, 2026-08-22, mid-session on #1911: "what are you
waiting for?".
Five commits were sitting unpushed behind a self-imposed review queue while
the branch's conflict with `main` had to be re-resolved twice.
The honest answer to the question was "nothing".)

## Query all available providers sequentially

When obtaining adversarial reviews,
you need a clean verdict from **every** available provider.
You must define the initial pinned quorum by performing an exhaustive discovery/availability check across all known providers (e.g., Cursor, OpenCode, Codex, Copilot, Claude, and the local `adversarial-reviewer` subagent).
Every provider found reachable at the start of the cycle must be included in the pinned quorum.
Any exclusion of a known provider must be recorded explicitly with its reason (e.g., quota blocked, CLI offline).
Do not stop after one provider returns clean.
Query them sequentially, one at a time.
Once one provider gives a clean review,
move on to the next one.
If any provider rejects the diff with findings,
you must address the feedback.
When you make fixes,
**do not hold the branch**:
push the verified fixes immediately.
Pushing the new commit naturally restarts the sequential query process against the new HEAD from the first provider.
When requesting review on the new push,
proactively carry forward any previously accepted rebuttals from earlier providers into your initial review request.
This ensures providers do not redundantly re-raise settled non-code issues on the new diff.
You must submit your rebuttal to the provider and request a new review.
This allows them to post a clean verdict at HEAD
that supersedes their previous findings.
Only after the provider posts a new clean verdict
may you continue to the next provider in the quorum.
Continue this iterative loop of review, fix, and push
until the current HEAD receives clean verdicts from the entire pinned quorum.

The set of required providers must be pinned at the start of the review cycle.
If a pinned provider drops offline or experiences transient operational failures (e.g. 500 errors, rate limits), you must wait and retry.
Alternatively, request explicit user permission to drop it from the quorum.
If the quorum size is zero at the start of the cycle, or drops to zero at any point during the cycle, you must fail closed and wait until at least one becomes reachable.
This applies if, for example, all external providers and the local fallback self-review subagent are offline or fail.
Alternatively, request explicit user permission to proceed.
Do not bypass the review gate.
If any provider (or combination of providers) creates an unbounded loop ---
whether through irreconcilably contradictory requirements,
self-contradictory oscillation,
or endless non-contradictory goalpost-moving ---
halt the review process and escalate to the user for a tie-breaking decision.

## A relayed not-clean round is a standing verdict under your login, so close it with a clean one on the new head

The findings a subagent review returns are posted to the PR as the report comment, per [Post every independent review on the PR or MR it reviewed](#post-every-independent-review-on-the-pr-or-mr-it-reviewed).
The report comment is posted under the account's own login, and `scripts/check-pr-fully-clean.py` reads it as that login's latest verdict.
"All addressed" does not clear it, because the instrument keys on the verdict phrase and on the reviewer, and a later all-clear from a *different* reviewer never supersedes a standing not-clean (ai-config#2274).
So the PR reads not-clean under `mwc` however many CLEAN bot rounds follow, until the same login posts a clean verdict on the current head.

The fix is cheap and it is the honest one anyway: once the findings are addressed, run the adversarial reviewer again on the new head and post its verdict, with a `### Verdict` line and the reviewed commit, under the same login.
That is the same-reviewer clean the rule asks for, and it is a real re-review rather than an edit to the old comment.

- **Do:** post a fresh adversarial verdict on the head that carries the fixes, in the same voice and login as the round that found them.
- **Don't:** rely on "all findings addressed in <sha>" inside the not-clean comment, or on later bot verdicts, to clear a not-clean you relayed.

(Measured 2026-09-01 on UCD-SERG/serocalculator#668: the relayed round on `065adf0` read as `d-morrison=not-clean` two CLEAN bot verdicts later, and a fresh Sonnet adversarial review of `2aa82df`, posted with its verdict, was what flipped the instrument to exit 0.)

## Post every independent review on the PR or MR it reviewed

Every independent review you run on a PR or MR goes on that PR or MR, not only the pre-push pass or the fallback review.
That covers a dispatched subagent review, an adversarial review, a referee read (an independent read of a rendered document, such as a manuscript, the way a journal referee would read it), and a review from another model or harness.
On GitLab, post it as an MR note;
for a branch with no PR yet, post it once the PR exists.
A review that returns a structured report with a verdict is posted in the form set by [The posted fallback comment is the reviewer's report](#the-posted-fallback-comment-is-the-reviewers-report-not-an-author-composite).
A referee read of a rendered document has no verdict line, so post its findings with page numbers.
Post the dispositions as [`ard`](../../skills/ard/SKILL.md) posts them, in a separate comment from the report, naming the fixing commit once it is pushed.
Each comment ends with the [`disclose-agent-authorship`](disclose-agent-authorship.md) marker.
A not-clean report you post is a standing verdict, so close it as [the relayed not-clean round section](#a-relayed-not-clean-round-is-a-standing-verdict-under-your-login-so-close-it-with-a-clean-one-on-the-new-head) directs.
In your reply to the user, link the posted comment itself, per [`link-forge-artifacts`](../writing/link-forge-artifacts.md).

A review that lives only in the session is invisible to the repository owner and lost when the session ends.
(Directive from the user, 2026-10-02, about a referee read of a manuscript MR: post the referee's report "for our records and so I can see what you're working on", and always give the link to it.)

- **Do:** post each independent review's report, and then its dispositions, on the PR or MR it reviewed, and link that comment in your reply.
- **Don't:** keep a review's findings in the chat, the session, or a scratch file and report only the fixes.

## Review every revision before it goes back to the user

Every revision you hand back to the user as a result gets one fresh independent review of the exact head you report.
That covers a rendered document, a draft, and a PR reported ready.
For a rendered document such as a manuscript, the review is a referee read of the new render, done by a reviewer other than the session that made the revision.
An earlier review covers the head it read and nothing after it, so a clean verdict on the previous revision does not carry over.
Post each review per [Post every independent review on the PR or MR it reviewed](#post-every-independent-review-on-the-pr-or-mr-it-reviewed), and link it in the reply that hands the revision back.

The pre-push review of that head counts.
A pre-push review of the exact head you report satisfies this rule, so an ordinary ARDI round needs nothing extra, and the same head never gets a second review.
This rule adds no push gate beyond the existing pre-push review.

The loop is bounded by the head you report:

- When the review of that head is clean, report it.
  A later fix needs a fresh review only if it changes what you report: the render for a rendered document, or the diff content for a PR.
- When the review returns findings, fix them and review the fix head before reporting it as ready.
- Alternatively, report the unreviewed head SHA explicitly as unreviewed, with the remaining findings listed, and not as ready.

(Directive from the user, 2026-10-02, after a manuscript revision went back without a new review: "did you get another adversarial peer review?
do that every time".)

- **Do:** review each head you report to the user once, by a reviewer other than its author, and link the posted report in that reply;
  where a head cannot be reviewed, report it explicitly as unreviewed, not as ready.
- **Don't:** reuse an earlier revision's review for a later head, report a revision as ready on the strength of your own reading, or re-review a head the pre-push review already covered.

## A reviewer handed nothing returns clean, so the brief must make an empty input an error

`git diff origin/<default-branch>...HEAD` reads the **commit graph**.
Work you have written but not committed is not in it.
So dispatching a review while the changes sit in the working tree hands the reviewer an empty diff, and an empty diff has no defects in it --- the verdict comes back clean, on time, in the usual shape, having examined nothing.

[`git-diffing`](../../memories/git-diffing.md)'s "A diff-scoped local check silently no-ops on an empty/uncommitted diff" states the same underlying fact for a **script** you run yourself, and its remedy is the same: commit first.
What that section cannot reach is the second half here.
A script you run reports its own zero, so the count is at least available to be read;
a dispatched reviewer reports a verdict and no count at all, and cannot supply one, because from inside it no input and no defects are the same observation.
That is why this case needs a fix in the **brief** rather than only in the habit.

Nothing about the result says so.
A clean review of an empty diff and a clean review of good work are the same artifact, and the second is what you were expecting, so the confirmation lands exactly where a check was supposed to be.
[`ardi.rationale.md`](ardi.rationale.md) records the sibling case, where the branch is empty **by design** because [`pr-on-claim`](pr-on-claim.md) opens the PR from an empty commit;
this one is worse to notice, because the work exists and you watched yourself write it.

**The specific move that produces it is switching branches with the work uncommitted.**
Git carries uncommitted changes across a checkout, so the edits follow you to whatever you switch to and the branch you meant to review keeps pointing at its old tip.
[`git`](../../memories/git.md)'s "commit before switching branches" section is the same hazard read from the other end.

Two fixes, and both are cheap enough that there is no reason to pick one.
**Commit before dispatching**, then read the diff's own stat line rather than assuming it.
And **tell the reviewer to refuse an empty input**: a brief that says which files the diff should touch, and that says to stop and report the emptiness rather than review, converts a silent false clean into a loud failure.
The reviewer cannot infer this --- from inside, no input and no defects are indistinguishable --- so it has to be stated.

The general form reaches past review.
Any instrument that reports on a population will report clean over an empty one, so **a verdict is only as good as the count of things examined**, which is [`batch-merge-and-resolve`](batch-merge-and-resolve.md)'s negative-control rule --- report how many pairs were examined and not only how many collided --- restated for a reviewer rather than a sweep.
[`algorithmatize-checks`](algorithmatize-checks.md) applies it to a combination sweep and credits that fragment as its source, which is where the wording comes from.
Measured 2026-09-03 within one hour, on two surfaces: a repo-wide scan reported zero invalid escapes while a hand-confirmed instance sat in the tree, because it wrote to `/dev/null`, every compile raised, and a bare `except: continue` swallowed all 238 ([ai-config#3114](https://github.com/Morrison-Lab/ai-config/issues/3114));
and an adversarial review returned clean over a branch carrying no commits ([ai-config#3118](https://github.com/Morrison-Lab/ai-config/issues/3118), which also tracks the mechanisms, since writing this rule did not prevent its second occurrence one message later).
Different surfaces, one shape, and neither announced itself.

- **Do:** commit, then confirm the diff is non-empty, before dispatching a review.
- **Do:** name the files the diff should touch in the brief, and instruct the reviewer to report an empty or unexpected diff instead of reviewing it.
- **Do:** read a clean verdict as a claim about a population, and ask what that population was.
- **Don't:** dispatch against a working tree and read the result as covering it.
- **Don't:** treat a clean result from any instrument as evidence until you know it examined something.

## Ask the reviewer plainly whether the work earns its place

The sections above brief the reviewer to find defects in a change.
This one asks a question they do not: **should this exist at all?**

Nothing in an ordinary round poses it.
A reviewer handed a diff reports what is wrong with the diff, so each round
returns a fix, the fix lands, and the next round finds the next thing --- a
loop that converges on a polished version of something that may never have
been worth building.
The author cannot break that loop, because by round three the sunk effort is
exactly what makes dropping feel like waste.

So put the question to the reviewer directly, in its own sentence: say
plainly whether this mechanism earns its place, ship or drop.
Then honour the answer.
A drop verdict is the cheapest finding available --- it retires the remaining
rounds along with the work --- and defending the change against it converts a
finished decision back into an open one.

- **Do:** ask for a ship-or-drop judgement outright when a mechanism has taken
  more than a round or two of polishing.
- **Do:** drop on a drop verdict, and say in the report that the reviewer's
  judgement is why.
- **Don't:** answer a drop verdict with the case for the change --- that is
  the rationale this fragment already rules out of a brief, arriving late.
- **Don't:** read successive rounds finding smaller things as evidence the
  work is converging; it is equally consistent with polishing something
  unjustified.

(Measured 2026-09-01, on a hook that had been a no-op three different ways
across three rounds.
Asked plainly, the reviewer said drop; dropping was right and ended the loop.
The companion half of that session --- concluding a silent subagent had
stalled when it was alive and twelve rounds ahead --- is recorded in
[`subagent-worktrees`](../../memories/subagent-worktrees.md), "A quiet worktree is not
evidence the session working it has stopped".)

### Ask it whether ANOTHER ROUND earns its place, which is a different question

The section above retires the **work**.
This one retires the **loop** while keeping the work, and the two need separating because a reviewer asked only the ship-or-drop question has no way to say "keep it, and stop iterating".

**The default rule is elsewhere in this file, and this question is for the case that rule does not reach.**
"Review the instrument too, not only the change it verifies" already gives the terminus: keep running rounds while findings keep landing, and "Give a docs-only diff describing an instrument a full round" restates it.
That is the right default and it is not withdrawn here.
What it assumes is that each round's findings are *about the change*.
Once a round's findings are about the **previous round's fix** --- the shape [`learn-from-review-findings`](learn-from-review-findings.md)'s "A later round can find a defect in the FIX" section documents --- findings still landing no longer distinguishes a series that is converging from one the fixes are feeding, so the default criterion returns "keep going" in both cases and stops discriminating.

So the expected-value question is not a competing terminus.
It is the tie-breaker for the case the findings-still-landing rule cannot separate, and it applies only there.
Put it to the reviewer in its own sentence, alongside the findings: **does another round have positive expected value, or is the remaining risk smaller than the risk a further fix introduces?**
A reviewer holding the round's own findings can weigh their severity against the observed rate at which fixes on this change have introduced new defects, which is a judgement the author cannot make about their own series.

Recorded 2026-09-03 from [ai-config#3059](https://github.com/Morrison-Lab/ai-config/issues/3059), across nine rounds on one hook.
Rounds 6, 7 and 8 each introduced a defect the next round found, which is exactly the condition above;
only round 9 introduced none.
Two things ended the series, per #3059: asking that question directly, and the validation-running round above.
Neither was a round happening to come back empty --- which, per the convergence rule, would not have been evidence it was finished anyway.

- **Do:** ask for a continue-or-stop judgement once a round's findings are about the previous round's *fix* rather than about the change, separately from the ship-or-drop question.
- **Do:** give the reviewer the rate at which this change's own fixes have introduced new defects, since that is the term it cannot derive from the diff.
- **Don't:** collapse the two questions.
  "Should this exist" and "should this iterate further" have different right answers, and a change worth shipping is the usual situation in which the continue-or-stop question arises at all.
- **Don't:** treat an empty round as the answer to either question;
  a converging series narrows its own search space, so the empty round is the least informative one.

### Narrowing severity is evidence about COVERAGE, not about the defect population

The two sections above give reasons a shrinking series might not mean what it looks like: the work may be unjustified, or the fixes may be feeding the findings.
The second already names the narrowing search space, in "a converging series narrows its own search space", and says an empty round is the least informative one.
This section takes that from a caution about the series' END to a rule about its STEERING: if the space narrows because each round returns to the last finding, the narrowing is steerable, and the dispatcher is the only party positioned to steer it.

A reviewer handed a change re-reads where the last finding landed.
So round N+1's search space is set by round N's result, and the severity curve across rounds is a record of **where attention went**, not of what remains.
A surface no round has opened contributes nothing to the curve however bad it is, and its absence from the findings is indistinguishable from its being clean.

Observed 2026-09-15 on `Morrison-Lab/ai-config`, fourteen adversarial rounds on one branch, and recorded as an unverified session account rather than as a measurement: no issue, PR or SHA anchors it, and neither named defect is greppable in the corpus today.
Read the round-by-round detail below as illustration of the mechanism, not as evidence for it --- the argument stands on why a reviewer's search space is set by the previous round's result, which is checkable from any review series, including this fragment's own.
One episode, not three.
The series ran fourteen rounds;
eleven of them ran checkers, which is the subset [`derive-dont-enumerate`](derive-dont-enumerate.md)'s eighth occurrence counts;
and thirteen pushes were refused across it, which is what [`get-under-the-hood`](../principles/get-under-the-hood.md)'s third refusal shape counts.
Fourteen is the figure in the round unit;
the other two are a subset of those rounds and a count of pushes, not competing totals.
Rounds 1 through 5 each found one stale-count defect, each less severe than the last, and read as convergent.
Round 6 was pointed deliberately at the files no earlier round had opened and immediately returned two defects that had been wrong for three rounds --- among them a function contract docstring naming the wrong regex.
Its own verdict named the mechanism: every round after the first had re-read the file round 1 landed in, so the apparent convergence was sampling bias.
Confirmed again at round 9, after three rounds returning only prose defects: steering at unswept surface found a real behavioural defect, a guard arm suppressed by any unrelated relocator in the command.

The remedy is bookkeeping rather than judgement, which is what makes it survivable across rounds: **track which files each round actually opened, and point the next round at the complement.**
That is a set the dispatcher can derive and the reviewer cannot.

Rounds 9 through 11 corroborated this from the other direction: every real defect in them came from executing a prediction taken from the prose rather than from re-reading prose against prose.
That half is already this file's, in "Tell it to RUN the repo's validation, not only to read the diff" and in "Give a docs-only diff describing an instrument a full round" --- the first of which closes on this section's own population point, that a set of rounds "covered the changed lines, which is a different population".
What is added here is only the steering rule, which neither of those gives.

- **Do:** record the files each round opened, and brief the next round at the ones no round has.
- **Do:** read a run of shrinking findings as "this surface is exhausted" rather than "this change is nearly clean" --- the two are the same observation about different populations.
- **Don't:** let a reviewer choose its own scope on a series of rounds;
  left alone it returns to the last finding, which is the one place already swept.
- **Don't:** count the severity trend as a stopping signal at all, separately from whether an empty round is one --- the trend and the empty round fail for the same reason, that the series narrows its own search space.

### Do not write to the tree a dispatched reviewer is reading

The reviewer reads the working tree, so any write to it moves the ground under a read already in progress.
The trigger is not `git checkout` specifically, which is the narrower form [`memories/subagent-worktrees.md`](../../memories/subagent-worktrees.md)'s "Switching a shared worktree's branch under a live dispatched reviewer breaks its reads" section records.
An ordinary in-place edit does it too: same worktree, same paths, different bytes underneath them mid-read.

The dispatcher cannot detect the damage afterwards, and the reviewer usually cannot either.
A reviewer that trusts a plain file read for the length of a long review reviews a mix of two states and reports no error at all, so the finding it returns may be about a line that no longer exists and the line it passes over may never have been in the tree it read.
Nothing in the resulting report says which.

The remedy is procedural rather than git-level, and it is the dispatcher's: dispatch the review, wait for it to report, then touch the files.
Fixing findings while the round is still running is what produces this, and it feels like promptness rather than a mistake.
On the reviewer's side, pin the target to a commit and read `git show <sha>:<path>` rather than the working copy.

- **Do:** treat "don't touch the tree under review" as covering every write to it, an edit and a `git add` and a formatter run alike.
- **Do:** have the reviewer read a pinned commit, so a dispatcher's slip degrades into a stale review rather than an incoherent one.
- **Do:** push a reviewed branch with `git push origin <local-branch>`, which ships that branch's tip without checking it out, so shipping a review's result never needs a branch switch in a tree something else may be reading.
- **Don't:** assume a live reviewer is safe from ordinary editing because no branch switch occurred.
- **Don't:** start fixing a round's findings before that round has reported.

(Measured 2026-09-09, `Morrison-Lab/ai-config`: an adversarial-reviewer subagent reading `scripts/check-docx-tracked-changes.py` reported the file "began changing under me (uncommitted)" while the dispatching session applied fixes to the same tree.
It recovered by comparing `git show HEAD:<path>` against a copy saved at the start of its read, and said so in its report;
that recovery is what surfaced the drift, not anything the dispatcher noticed.)
