Files
ersatztv/docs/guard-inventory.md
T
timothy 5ba442c11c fix(772,792): name the missing toolchain image, and stop a refusal leaving a verdict comment
#772 — the pinned CI toolchain image can be deleted out from under us, and when it was
(2026-08-11..13) all five `container:` jobs died at image pull, both required contexts
included, with the cause buried in each job's log. Root cause is registry-side and is now
established rather than guessed: an owner-level Gitea package cleanup rule (keep_count 15,
remove_days 1, remove_pattern `.*`, keep_pattern no 7-hex sha can match) deletes a sha tag
once 15 newer versions exist, and `ExecuteCleanupRules` ran nightly through the window. The
`ersatztv` package carries the same rule's fingerprint exactly — every sha tag older than
the 15-slot window is gone, every keep_pattern tag back to 26.3.1 survives. Version deletes
leave no audit row, so the specific run cannot be replayed; that limit is stated where the
claim is made. The durable fix belongs to the registry's repo: server-management#842.

What lands here is what a consumer of someone else's registry can do:

  * `toolchain-preflight`, a container-free job (a job consuming the image could not run to
    report it missing) resolving every pin against the registry and failing with a message
    that names the tag and the recovery. Not a `needs:` of the jobs it diagnoses — gating
    five jobs behind a checkout and one curl taxes every green run to speed up a rare red
    one, and they already fail fast.
  * Only HTTP 404 means gone. Everything else is could-not-tell, and rejected credentials
    fail rather than pass as unknown — "the check could not run" must never present as
    "the pin is fine".
  * A recovery path that does not need CI: rebuild the SAME tag from the commit it names
    and push it. The push half was verified against this registry on 2026-08-22 with a
    throwaway package (created, resolved 200, deleted).

#792 — the reported defect was the exit code, and re-measuring says that premise is false:
every no-status path already exits 1, and eight refusal modes now assert it against the real
predecessor, where they pass. The observed 0 came from the invocation, not the script. What
WAS broken is the half-state the issue describes second: the comment was written before the
status, so every refusal left `Review-verdict: MERGEABLE @ <head>` on a PR with no gating
status behind it. The two writes are now ordered status-then-comment, which makes the only
reachable half-state the safe one — a status with no comment leaves the merge hook's
condition (c) with nothing to classify, which is an `ask`. The refusals themselves are
untouched. Ordering rather than compensating deletion: an orphaned-comment cleanup needs a
Gitea call, and these refusals are usually caused by Gitea being unreachable.

Proof for the ordering is the split against origin/main's script: the 8 orphan/ordering
tests go red there, the 8 exit-code tests stay green.

fixes #772
fixes #792
Refs: server-management#842
Decisions-Edit: yes
2026-08-22 22:52:56 +02:00

423 lines
38 KiB
Markdown

# Guard inventory (ersatztv#774 / #775)
Every executable guard **file** in this repo, what it blocks, and whether it ships a proof it can go
red. `scripts/tests/test_guard_inventory.py` derives the population from the **git index** and the
workflow/hook call sites and asserts **set equality** against the `Guard` column, so a new guard
cannot be added without acquiring a row here, and a row cannot name a proof that does not exist.
The index rather than a filesystem walk since ersatztv#806 — the practical consequence is that a new
guard joins the population when it is **staged**, not when the file appears. Nothing local runs
these checks at all: `.husky/pre-commit` runs lint-staged, the decisions guard, the root-PNG check
and `dotnet format`, and `grep -rn pytest .husky/` returns nothing. The runner is
`pr-checks.yml::script-tests`, `on: pull_request`, so the red arrives in CI — plus
`docker-build.yml:723` on the release path, which re-runs two of these files
(`test_ci_dropped_step_guard.py` and `test_ci_release_path_scan_job.py`) as a `needs:` of `build`.
**Read `docs/decisions/records/testing/guard-derives-population-from-source.md` and
`…/guard-ships-with-mutation-proof.md` before editing a guard or adding a row.**
## Columns
- **Kind** — `GUARD` (it can block a commit, a push, a tool call, a merge or a CI job), `TOOLING`
(it does work but asserts nothing; listed so its absence from the guard set is a recorded
decision, not an oversight), or `PROOF` (a `scripts/tests/` file whose job is to prove another
guard). `PROOF` exists to stop a regress: once test files entered the population, every mutation
proof became a row wanting a proof of its own. A `scripts/tests/` file that enforces a repo
invariant with no separate guard behind it is a `GUARD`, graded normally, and may cite a mutation
case in its own file.
- **Proof** — `MUTATION`: a clause-level mutation **was executed** and this named test **was
witnessed red**. The test either performs the disarm itself (a monkeypatch, a deselection, a
removed marker) or a recorded incident shows it going red when the clause went away.
`BEHAVIOUR-ONLY`: a real test drives the guard through its real entry point with good and bad
input, but no one has demonstrated that removing the clause reddens it. `NONE`.
- **Proof ref** — `file.py::function`, verified to exist by the inventory test.
The distinction between `MUTATION` and `BEHAVIOUR-ONLY` is the whole point of #775 and is not a
grading curve. A behavioural test proves the guard reacts to its input; only a mutation proves the
guard is *load-bearing*. #685 shipped two guards on one condition where deleting either left the
suite green, and every behavioural test passed throughout.
**The definition above is the second one.** The first read "a named test disarms this guard's clause
and asserts red", and three rows were graded `MUTATION` against it that do not disarm anything —
they feed the real script an input only that clause rejects (empty stdin, a short page, a full first
page). Cold review caught it, and the objection is right twice over: those tests are the same
species as ones graded `BEHAVIOUR-ONLY` eight rows away, so the column was being applied as a curve
on the very day it was introduced. They are regraded. The surviving criterion is *witnessed*, not
*plausible* — "removing this clause would surely redden that test" is an argument, and this table
exists because arguments of that shape have been wrong here six times.
## Scope limit, stated rather than implied
This inventory covers guard **files**, discovered by reading `.claude/hooks/*.sh`, `.husky/*` and
`scripts/tests/test_*.py` **out of the git index**, plus every `scripts/…` path referenced by a
workflow or a hook. The classes below are outside that population. They are listed because the first
version of this section named only the first one, and cold review found that the very guards this
inventory shipped with were sitting in the gap:
1. **Guards inline in workflow YAML** — most importantly `pr-checks.yml:ci-image-pin`. "Which jobs
are guards" needs a judgement call per job the filesystem cannot supply. Two were audited under
#774 and one fixed; extending the population is tracked in #786.
2. **C# and TypeScript guards**`ErsatzTV.Mcp.Tests/ToolCatalogTests.cs` and
`web/src/api/pageSizeCallSites.guard.test.ts` are both structural guards and neither has a row.
3. **Mentions counted as call sites.** The `scripts/…` scrape matches any occurrence, including
inside a comment or an `::error::` string. `scripts/update-openapi.sh` is named in a
`pr-checks.yml` error message, so removing the step that runs it would leave its row intact.
4. **Nested and non-lowercase paths** beyond `scripts/tests/` — a guard under
`scripts/scripted-schedules/`, or with an uppercase name, is invisible to the scrape.
5. **Non-`.sh` hooks** — the hook pattern is `*.sh` only. (`.yaml` workflows are no longer in this
gap: the caller scan matches `*.yml` and `*.yaml` since #806, because Gitea accepts both.)
6. **Transitive calls** — a script invoked only by another script, rather than by a workflow or
hook, is not discovered.
7. **Non-`test_` modules under `scripts/tests/`** — the pattern is `test_*.py`, so `conftest.py`,
`mutation_harness_lib.py`, `mutation_manifest.py` and `tracked_files.py` are outside the
population and hold no rows. They are not guards (they assert nothing on their own), but the
middle two ARE what `test_mutation_harness.py` is made of, so gutting either would take that
guard with it. What catches that is the guard's own row: its declared mutation targets
`mutation_harness_lib.py`, and its proof test refuses to run if the clause it names has moved.
`tracked_files.py` has no such backstop — it is load-bearing for every module that imports it
(#806), and an edit to it is covered only by those modules' own proofs. Recorded rather than
force-fitted: a row for a library would need a `Kind` the vocabulary does not have.
8. **Nested workflow directories** — the workflow scope is direct children of `.gitea/workflows`, so
a tracked `.gitea/workflows/nested/x.yml` is invisible to the caller scan and to
`test_ci_image_pin_population.py`. Left as scope rather than widened: whether Gitea executes
nested workflow files was not verified here, and widening on an unverified premise risks a
permanent red on a correct tree, which is how a correct guard gets deleted.
Hook **wiring** is checked (`test_every_hook_file_is_actually_WIRED` reads `.claude/settings.json`
and the husky hooks with full-line comments stripped), so a hook file whose registration is deleted
fails rather than keeping a row that reads as coverage. Its limit, stated because the check reads
stronger than it is: **it is a substring test for the basename, not a parse of the invocation.**
`: # .claude/hooks/decisions-guard.sh disabled` still reads as wired, and conversely a hook invoked
through a wrapper or a constructed path reads as unwired. It catches deletion, which is the common
case; it does not catch deliberate disablement. The check does not extend to the `scripts/` half at
all.
## File populations and where they come from (ersatztv#806)
Every guard here whose members are FILES derives them from the **git index**, never a filesystem
walk. The disk is not an authoritative source: it reports build output and editor droppings and
differs per machine, so a guard derived from it asserts a different population in CI than on the
laptop of the person it is meant to stop. `scripts/tests/tracked_files.py` is the single derivation
and carries the full rationale; `scripts/tests/test_guard_populations_derive_from_git.py` proves it,
in both directions: removing EVERY member of each registered derivation from the index one at a time
and requiring it to disappear while still on disk, and watching for a directory LISTING issued while
the derivation runs (reading files stays allowed). Removal alone is blind to a source that
contributes only untracked members — an `rglob` reaching `.husky/_/` adds and never removes — and
any check phrased as "an untracked file must not enter" is itself machine-dependent, because the
untracked file has to exist. Watching for the call needs no arranged state.
That second check is a regression guard against the accidental shapes, not a boundary: what is
observed is any call that goes **through one of the spies**, whenever it happens — the spy records
into a list that outlives the patch, so a reference captured during the window and invoked after it
still counts. Whether the call goes through a spy is what decides, not when. The instance list — what
never reaches a spy at all — lives in the check's own docstring and is deliberately not restated
here, because a second copy of it drifted from the first within one commit.
One deliberate exception to "from the index" sits in the same file: the registration check lists
`scripts/tests/test_*.py` from disk on purpose, because it is a superset check over what pytest
collects — an untracked stray there makes it MORE demanding, never blind, whereas using the index
would let an unstaged new guard escape registration.
Its own limit, stated because a check described as complete stops being re-examined: it finds
derivations by PARSING each `test_*.py` for an import of the shared helper, so a module that derives
a file population some other way is invisible to it, and no mechanical check can close that (#774
reached the same conclusion about detecting filter-shaped guards by token).
The audit #806 asked for, recorded whichever way it came out, because "we looked and left it" and
"we never looked" are indistinguishable a year later:
| Guard | Population | Completeness claim over tracked files? | Outcome |
| --- | --- | --- | --- |
| `test_guard_inventory.py` | `.claude/hooks/*.sh`, `.husky/*`, `scripts/tests/test_*.py`, workflow/hook callers | **yes** — set equality against this table | converted to the index; `.husky/_/` had been excluded only because `_` is a directory, so the obvious "make it recursive" edit would have reintroduced #778's defect here |
| `test_hook_fire_log.py` | `.claude/hooks/*.sh` | **yes** — every hook must be instrumented | converted; an untracked scratch `.sh` used to demand instrumentation and redden the suite on that checkout alone |
| `test_ci_image_pin_population.py` | `.gitea/workflows/*.yml` + `*.yaml` | **yes** — "docker-build is the ONLY workflow pinning the toolchain image" | converted, and `*.yaml` added: Gitea accepts both spellings, so a `.yaml` workflow was structurally invisible while the test read as covering all of them |
| `test_pr_changed_files.py` | `.gitea/workflows/*.y*ml` | **yes** — "no OTHER workflow writes the review-verdict status" | converted. Not on #806's list and found by cold review on the pushed head: an untracked `.yaml` dropped in `.gitea/workflows/` reddened two guards while absent from the index — the issue's list of files to assess was a starting point, not the population |
| `web/src/api/pageSizeCallSites.guard.test.ts` | `web/src/**/*.{ts,tsx,mts,cts}` via `import.meta.glob` | **yes** — an unregistered discovered site fails | DEFERRED to #819, assessed not skipped. The glob is a documented workaround: `@types/node` is deliberately out of `tsconfig.app.json`, and wiring it in was tried and reverted (it leaked Node's `setTimeout` into the app project and broke three unrelated tests), so there is no `node:child_process` to reach the index from. Over-enumerates, so it fails loudly rather than going blind |
| `test_ci_release_path_scan_job.py` | `.gitea/workflows/*.y*ml` + `scripts/**` | **no** — a fixture assembling a tmp harness, asserted about behaviour not membership | takes its file LIST from the index anyway, for hermeticity not completeness: `shutil.copytree` copied whatever was on disk, including untracked files and `scripts/__pycache__`, into a tree whose behaviour the probes then measure. Content still comes from the working tree. The copy is not a git repo, so the two files this step RUNS may not use the helper — see the fixture docstring |
| `test_ci_dropped_step_guard.py` | the parsed workflow document | **no filesystem population at all** | unchanged; its residual is `MARKED_JOBS`, a SCOPE mirror of the required contexts on `main`, which #806 does not close |
`test_remote_state_inventory.py` fixed its own population under #778 and kept a private copy of the
derivation; #806 folded it onto the shared one, so that module is covered by the proof above like
the rest. Across the whole change, every module that derived a file population its own way now goes
through `tracked_files.py` — one implementation of the rule instead of one per module. The
registered derivations are listed in `DERIVATIONS` in
`scripts/tests/test_guard_populations_derive_from_git.py`; this page deliberately keeps no count
of them.
**Still on filesystem walks, deliberately out of scope:** the decisions corpus
(`scripts/decisions_lib.py`'s `active_files()`, and the suites over it). Its members are `docs/`
Markdown with no generated-file pressure and a different lifecycle, and folding it in here would
have been the reflex this milestone argues against — a change with no defect behind it. It is
recorded as unexamined rather than as cleared.
## Inventory
| Guard | Blocks | Kind | Proof | Proof ref |
| --- | --- | --- | --- | --- |
| `.claude/hooks/decisions-guard.sh` | a commit | GUARD | NONE | — |
| `.claude/hooks/design-sync-reminder.sh` | the first Stop after a UI change (one-shot, then allows) | GUARD | NONE | — |
| `.claude/hooks/posttooluse-worktree-marker.sh` | nothing (writes the marker the worktree guard reads) | GUARD | MUTATION | `test_worktree_ownership_guard.py::test_MUTATION_a_marker_hook_that_stops_WRITING_makes_the_guard_go_quiet` |
| `.claude/hooks/prepush-clean-worktree-check.sh` | a push with uncommitted changes in the pushed set | GUARD | NONE | — |
| `.claude/hooks/prepush-donewhen.sh` | a direct push to `main` with unticked Done-when boxes | GUARD | NONE | — |
| `.claude/hooks/prepush-rebase-check.sh` | a push from a branch behind `origin/main` | GUARD | BEHAVIOUR-ONLY | `test_prepush_rebase_check_tag_exemption.py::test_zero_ref_lines_does_not_exempt` |
| `.claude/hooks/pretooluse-agent-model.sh` | an Agent dispatch naming no model (asks) | GUARD | NONE | — |
| `.claude/hooks/pretooluse-agent-ram.sh` | an Agent dispatch under 10% free RAM | GUARD | NONE | — |
| `.claude/hooks/pretooluse-bash-guard.sh` | a Bash call setting `ETV_UPDATE_GOLDENS` | GUARD | NONE | — |
| `.claude/hooks/pretooluse-bom-guard.sh` | a commit/push carrying a BOM in a touched `.cs` | GUARD | MUTATION | `test_bom_guard_detection.py::test_DISARMING_the_BOM_comparison_stops_detection` |
| `.claude/hooks/pretooluse-merge-consent.sh` | a PR merge without derived consent | GUARD | BEHAVIOUR-ONLY | `test_merge_consent_exemption.py::test_protected_path_on_a_LATER_page_is_still_seen` |
| `.claude/hooks/pretooluse-nav-guard.sh` | a browser navigate to a streaming URL | GUARD | NONE | — |
| `.claude/hooks/pretooluse-worktree-guard.sh` | a commit/merge in a foreign worktree | GUARD | MUTATION | `test_worktree_ownership_guard.py::test_MUTATION_disarming_the_guards_MARKER_READ_stops_the_deny` |
| `.husky/commit-msg` | a commit with no `Co-Authored-By` trailer | GUARD | NONE | — |
| `.husky/pre-commit` | a commit failing lint-staged, decisions, root-PNG or format | GUARD | NONE | — |
| `.husky/pre-push` | a push failing any pre-push hook or the SPA gate | GUARD | MUTATION | `test_prepush_unsets_git_env.py::test_MUTATION_DELETING_the_unset_lets_drift_through_silently` |
| `scripts/build_decisions_catalog.py` | the `decisions-guard` job, on a stale catalog | GUARD | MUTATION | `test_build_catalog_check_path.py::test_MUTATION_disarming_the_stale_comparison_stops_detection` |
| `scripts/check-doc-narrative.py` | nothing, by design (advisory `::warning::` only, exits 0 on every path (the error/degradation arms are defensive and unproven — see the record) — `docs.no-session-narrative` says a string predicate over prose may not be load-bearing) | TOOLING | NONE | — |
| `scripts/check-kickoff-guard.sh` | the `decisions-guard` job, on a revived #237 reference | GUARD | NONE | — |
| `scripts/check-review-verdict.sh` | the merge-consent hook's verdict classification | GUARD | BEHAVIOUR-ONLY | `test_check_review_verdict.py::test_falseopen_token_must_be_a_whole_word` |
| `scripts/ci-detect-already-validated.sh` | nothing directly (feeds the skip gate) | GUARD | NONE | — |
| `scripts/ci-detect-docs-only.sh` | nothing directly (feeds the skip gate) | GUARD | NONE | — |
| `scripts/ci-peak-anon.sh` | nothing (samples container memory) | TOOLING | NONE | — |
| `scripts/ci-prove-ban-detects.sh` | the release path, if the delimiter ban is disarmed | GUARD | NONE | — |
| `scripts/ci-step-ran.sh` | the two required contexts, on a dropped step | GUARD | MUTATION | `test_ci_dropped_step_guard.py::test_dropping_ANY_single_step_FAILS_the_guard` |
| `scripts/ci-toolchain-image-resolves.sh` | the `toolchain-preflight` job, when the pinned CI toolchain image has been deleted from the registry | GUARD | MUTATION | `test_ci_toolchain_image_resolves.py::test_MUTATION_a_deleted_tag_is_reported_as_a_failure` |
| `scripts/decisions_validate.py` | the `decisions-guard` job, on a lifecycle fault | GUARD | MUTATION | `test_decisions_validate.py::test_main_actually_CALLS_the_wing_scan` |
| `scripts/e2e-functional.sh` | the Functional E2E job, on a failed HTTP contract assertion | GUARD | NONE | — |
| `scripts/e2e-local.sh` | nothing (boots a local instance) | TOOLING | NONE | — |
| `scripts/e2e-ui.sh` | nothing (drives the Playwright flows) | TOOLING | NONE | — |
| `scripts/hook-fire-log.sh` | nothing (records that each hook fired, and reports it) | TOOLING | NONE | — |
| `scripts/jq-preflight.sh` | the `script-tests` job, on a jq version change | GUARD | BEHAVIOUR-ONLY | `test_jq_preflight.py::test_below_the_floor_is_LOUD` |
| `scripts/post-review-verdict.sh` | nothing (writes the verdict status) | GUARD | BEHAVIOUR-ONLY | `test_post_review_verdict.py::test_never_retargets_the_verdict_at_the_new_head` |
| `scripts/pr-changed-files.sh` | the verdict exemption, on an incomplete enumeration | GUARD | BEHAVIOUR-ONLY | `test_pr_changed_files.py::test_a_SHORT_page_does_not_end_the_enumeration` |
| `scripts/prove-fix.sh` | the `prove-fix` job, on a commit whose `Proves:` trailer names a test that passes without the fix | GUARD | MUTATION | `test_prove_fix.py::test_MUTATION_disarming_the_UNPROVEN_clause_reddens_the_refusal_test` |
| `scripts/update-openapi.sh` | nothing (regenerates the spec) | TOOLING | NONE | — |
| `scripts/tests/test_bom_guard_detection.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_build_catalog.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_build_catalog_check_path.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_check_review_verdict.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_check_doc_narrative.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_ci_dropped_step_guard.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_ci_image_pin_population.py` | the `script-tests` job, when a container job loses its pin | GUARD | MUTATION | `test_ci_image_pin_population.py::test_a_single_job_losing_its_pin_is_DETECTED` |
| `scripts/tests/test_ci_release_path_scan_job.py` | the `script-tests` job, on a weakened release-path scan job | GUARD | NONE | — |
| `scripts/tests/test_ci_toolchain_image_resolves.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_decisions_lib.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_decisions_validate.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_guard_inventory.py` | the `script-tests` job, on an unclassified guard or a stale proof ref | GUARD | MUTATION | `test_guard_inventory.py::test_the_inventory_covers_exactly_the_guards_that_exist` |
| `scripts/tests/test_guard_populations_derive_from_git.py` | the `script-tests` job, on a guard whose file population admits a file git does not track | GUARD | MUTATION | `test_guard_populations_derive_from_git.py::test_no_derivation_admits_an_untracked_file` |
| `scripts/tests/test_hook_fire_log.py` | the `script-tests` job, on a hook that stops reporting that it fired, or whose reporting changes what the harness sees | GUARD | MUTATION | `test_hook_fire_log.py::test_a_hook_that_LOSES_its_instrumentation_is_DETECTED` |
| `scripts/tests/test_jq_preflight.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_merge_consent_base_change.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_merge_consent_exemption.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_merge_consent_required_check.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_migration_equivalence.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_mutation_harness.py` | the `script-tests` job, when a `MUTATION` row's declared clause no longer reddens the test the row names | GUARD | MUTATION | `test_mutation_harness.py::test_MUTATION_disarming_the_DIAGNOSTIC_gate_accepts_a_red_for_the_wrong_reason` |
| `scripts/tests/test_post_review_verdict.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_prepush_unsets_git_env.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_pr_changed_files.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_prepush_rebase_check_tag_exemption.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_prove_fix.py` | the `script-tests` job | PROOF | NONE | — |
| `scripts/tests/test_remote_state_inventory.py` | the `script-tests` job, on an executable that talks to a remote service with no row in `docs/remote-state-inventory.md` | GUARD | MUTATION | `test_remote_state_inventory.py::test_every_in_scope_file_has_a_row_and_every_row_names_a_real_file` |
| `scripts/tests/test_worktree_ownership_guard.py` | the `script-tests` job | PROOF | NONE | — |
## The `MUTATION` column is executed, not asserted
Every row graded `MUTATION` carries a DECLARED clause mutation in `scripts/tests/mutation_manifest.py`.
`scripts/tests/test_mutation_harness.py` applies each one to an isolated copy of this repository and
requires that row's own named test to go red; the manifest and the `MUTATION` rows are compared for
set equality in both directions, so the grade and the mutation cannot drift apart. Adding a row
graded `MUTATION` without declaring its clause fails the suite. Full contract and its limits:
`docs/decisions/records/testing/mutation-claims-are-executed.md`.
Two things that column still does not say. It does not say the declared clause is the ONLY thing the
guard hangs on — for three rows (measured 2026-08-22) the redness arrives through the
proof test's own "this clause has moved, RETARGET it" assertion rather than through changed
behaviour, which catches the recorded proof going stale but not much else. And one row is graded `DETECTOR` rather than `CLAUSE` in the
manifest: `instrumentation_faults` accumulates from four independent arms and a stripped hook trips
three at once, so no single-arm disarm reddens its proof. That finer mutation is carried as data and
re-run every suite, and must keep surviving — if it starts reddening, the guard has become
clause-provable and the entry is regraded.
## What the numbers say
37 guards, 6 tooling scripts, 20 proof files. **15 guards carry a mutation proof; 6 are
behaviour-only; 16 have none.** These figures are asserted against the table by
`test_the_summary_counts_match_the_table` — they were wrong in the first draft (28/4/6/3/19 against
a table holding 27/5/6/3/18), because a hand-maintained summary of a table is a second copy of it,
which is the duplication family this change argues against. Both cold reviewers found the error
independently.
**One guard was WITHDRAWN rather than shipped**, and it is worth a line here because the inventory
is where a future session will look for it. A `test_review_verdict_vocabulary_parity.py` asserted
set equality between the H10 verdict vocabulary's write side and read side by extracting both from
shell source with regexes. Six cold-review rounds each found another shell construction that either
escaped it or made it red on a correct tree. It was deleted rather than patched a seventh time: the
`testing.guard-derives-population-from-source` record says a weak detector is itself the
symptom-keyed mistake, and this was the empirical demonstration. The underlying duplication is real
and is now UNMITIGATED — tracked in #788, which fixes it by removing the second copy.
**What that file asserted, and where each part went** — enumerated because the first attempt at this
withdrawal silently dropped one and claimed otherwise, which is
`process.enumerate-workaround-behaviors-before-deleting` failing on its own removal:
| Invariant it asserted | Where it is now |
|---|---|
| Bidirectional write/read vocabulary parity | **LOST.** The reason for the withdrawal; #788 |
| No READ-side word is both positive and negative | **LOST** as a universal property; #788 |
| No WRITE-side word is both success and failure | **LOST** as a universal property; #788 |
| Strict extractor consumes every case arm | Retired with the extractor; #788 |
| `POS_RE`/`NEG_RE` each assigned exactly once | Retired with the extractor; #788 |
| Both extractors find non-empty sets (anti-vacuity) | Retired with the extractor; #788 |
| A wrong write source makes the extractors disagree | Retired with the extractor; #788 |
Two guards were ADDED in its place. They are **not** replacements for rows 2 and 3 and the table
above deliberately does not list them as such — an earlier draft did, and cold review was right that
relabelling a lost invariant as a narrower surviving one is the exact failure this table exists to
prevent, committed by the table:
| Added | What it pins |
|---|---|
| `test_check_review_verdict.py::test_each_verdict_word_retains_its_established_polarity` | the five established tokens still READ with their established polarity |
| `test_post_review_verdict.py::test_each_verdict_word_posts_its_established_polarity` | the five established tokens still POST with their established polarity |
**Polarity is strictly weaker than disjointness, and the gap is demonstrable.** Put `MERGEABLE` in
BOTH `case` arms of `post-review-verdict.sh`: the success arm wins, every polarity assertion stays
green, and the withdrawn disjointness test failed. What the two added guards do catch is the
dangerous direction — a token that a reviewer means as `BLOCKED` silently reading or posting as
approval, which writes a green `review-verdict/h10`. A token added to one script and not the other
is untested by either, which is the whole of #788.
**A guard was WITHDRAWN from `test_hook_fire_log.py` after four iterations**, recorded here because
this is where a future session will look for it. A `test_no_redirection_in_the_sink_PRECEDES_its_stderr_redirect`
generalised the `exec … 2>/dev/null` rule to any command by pattern-matching shell source. Each
iteration fixed a false positive or false negative the previous one introduced — `>/dev/null 2>&1`
flagged wrongly, then operators preceded by a digit (`exec 0<`) missed entirely, then multi-command
lines flagged wrongly — and cold review then constructed more of both (`[[ "$x" < "$y" ]] 2>/dev/null`,
a backslash continuation, a `>` inside a quoted string). Deleted rather than patched a fifth time,
on the same reasoning as the vocabulary-parity withdrawal above.
**`test_hook_fire_log.py` proves THREE clauses, and its row claims two of them.** Coverage (every
hook reports that it fired) via `test_a_hook_that_LOSES_its_instrumentation_is_DETECTED`;
transparency (the wrapper changes nothing the harness can see) via
`test_instrumentation_changes_NOTHING_the_harness_can_see` with its mutation proof
`test_DELETING_the_replay_makes_the_differential_go_RED`; and placement (`etv_hook_fire_begin` must
precede the stdin read) via `test_begin_placed_AFTER_the_stdin_read_is_DETECTED`. The row's `Blocks`
column covers the first two — "stops reporting that it fired, **or whose reporting changes what the
harness sees**" — and the `Proof ref` column holds one ref because the column holds one, not because
the second is unproven. Placement is proved and unclaimed.
**Its `test_the_suite_does_not_write_to_the_PRODUCTION_log` is narrower than its docstring**, which
says `conftest.py` "must isolate every test, not just this file's". What it checks is that the
fixture set `ETV_HOOK_FIRE_LOG_DIR` *for the test currently running* and that a hook it drives in
its own sandbox does not touch the real log. It cannot see another suite that snapshots
`os.environ` at IMPORT time — before the autouse fixture runs — and hands that stale mapping to its
subprocesses. That suite's hooks then write to `$HOME/.cache/ersatztv/hook-fire/` while every
assertion stays green, because the fire-log library is fail-open. #785 shipped exactly that defect,
inside the file added to prove those very hooks.
The reproduction is the part worth keeping: restore a module-level `{**os.environ}` snapshot — or
leave the helper correct and point a single `env=` argument at one — then run that file and count
records for its synthetic session ids. **58 per run**, measured identically on macOS/git 2.55 and
Linux/git 2.47.3. (The accumulated total observed before the fix ran to four figures across many
runs; that is an observation rather than a reproducible measurement, so check the per-run figure.)
The pin is `test_worktree_ownership_guard.py::test_driving_a_hook_LANDS_its_records_in_the_ISOLATED_dir`,
which asserts the EFFECT — records land in the fixture's dir — rather than the shape of the fix. Its
predecessor asserted the helper's return value, and cold review showed that leaves the call site
unguarded: `_env()` correct, one `env=` reverted, all tests green, records still leaking. Pinning to
a hand-written revert rather than to the property is `verify-against-the-REAL-predecessor`.
It is still narrower than the property: it guards the launch path it drives. A second launcher in
the same file that passed a stale environment would leave it green — measured, 18 records — because
the hooks it drives would still log correctly. Every hook in that file goes through one helper
today, which is what makes it sufficient there. Generalising it is #809, and the reason that is hard
is that the obvious version races against a real session's hooks firing during the run.
**That unproven set carries no number here on purpose.** It restated the count, drifted the moment
the BOM guard was regraded, and `test_the_summary_counts_match_the_table` cannot see it — the parser
checks the formatted summary sentence and nothing else. A second hand-maintained copy of a number is
the duplication family this file argues against, so the copy is removed rather than corrected: the
set is the eight `.claude/hooks/` guards enumerated in the standing list below, and the count is in
the summary above. Naming them by event is what made the earlier wording wrong twice — they do not
share one event, and `design-sync-reminder.sh` is registered on **both** `PreToolUse` and `Stop` in
`.claude/settings.json`, so any "every X hook" phrasing double-counts it.
**They are now observable but still unproven, and the two words carry different weight.** Observable:
every hook records its own execution through `scripts/hook-fire-log.sh`, so "did this hook fire, and
what did it decide" is a measurement — run `scripts/hook-fire-log.sh report` (#776). Unproven:
nobody has demonstrated any of them is load-bearing — the first group in the standing list below.
Observability tells you a guard ran; only a mutation tells you it would have caught anything. The
BOM guard is the case that shows why the distinction matters — it was firing on every commit the
whole time it was fail-open.
The gaps are not uniform in cost, and the ranking that matters is *what a silent failure would let
through*, not test count. **All four ranked entries now carry clause-level mutation proofs (#785);
they are kept here with what each mutation established, because the ranking is the reusable part and
because two of them turned out to be worse than the ranking predicted.**
1. ~~`pretooluse-bom-guard.sh`~~**proven, and it was fail-open the whole time.** Ranked first
because the defect it guards has recurred three times (#311, #402, #405); that ranking turned out
to be right for a worse reason than intended. It detected a BOM with `xxd -p`, and `xxd` ships
with vim and is **absent on the Linux CI runner**, so the comparison never matched and every BOM
was allowed in silence. `od` now. The lesson for the rows below: an unproven guard is not merely
untested, it is a guard whose *current* behaviour nobody has established.
2. ~~`pretooluse-worktree-guard.sh` + `posttooluse-worktree-marker.sh`~~**proven as a pair.**
Four clauses were disarmed and witnessed red: the guard's marker read; the guard's ownership
comparison (inverted, because disarming it the other way only makes the guard deny more and every
deny assertion stays green); the `commit|merge` alternation, whose `merge` half every other case
in the file left untested while guarding the plumbing-merge path; and — the one that could not
exist while the halves were tested apart — the *marker hook's write*, asserted against the
*guard's* decision. Both hooks are
deliberately fail-open, so an absent mechanism and a working one produce the identical "commit
allowed".
3. ~~`.husky/pre-push:11`'s `unset GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE`~~ — **proven, and the case
is the normal one rather than an edge.** Git exports `GIT_DIR` to `pre-push` **when the push
comes from a worktree** and not from the main tree; `process.shared-tree-readonly` makes the
worktree the mandated way to work here, so every push takes the exposed path. With `GIT_DIR` set
and `GIT_WORK_TREE` unset git stops discovering the repo and treats the *current directory* as
the work tree, so `cd web && npm run check:api`'s `git diff --exit-code` compares against index
paths that do not exist and reports no diff. Both the deletion and the relocation are proved.
4. ~~`scripts/build_decisions_catalog.py`~~**proven, including the wiring.** The `--check`
comparison is mutated directly, and a separate subprocess case runs the command *derived from
`pr-checks.yml`* against a copied corpus. That second case is not redundant: replacing
`raise SystemExit(main())` with a bare `main()` leaves the script printing
`docs/decisions/README.md is stale` on stderr while exiting **0**, and the workflow step reads
nothing but the exit code. Only the subprocess case reddens — the #751/#719 shape.
**What a file-level `MUTATION` grade does and does not claim, because three of these four rows are
multi-purpose files.** The grade covers *the clause the cited case actually mutates*, per
`testing.guard-ships-with-mutation-proof`; it is not a statement about every line in the file.
`.husky/pre-push` runs four other things, and its row asserts only that line 11 is load-bearing —
the three hooks it invokes carry their own rows and their own grades. Reading the row as "pre-push
is tested" is the same relabelling error the withdrawn-guard table above exists to prevent.
Clause-level grading is tracked in #790.
**The remaining unproven guards, and why each is still `NONE` rather than merely unattended** — the
third `## Done-when` box of #785, answered here rather than left implicit. They fall into three
groups, and the split is the point: the first two are backlogs, the third is not.
- **The eight remaining hook guards** — genuinely unproven, and a real backlog. Four are
`PreToolUse` (`pretooluse-agent-model.sh`, `pretooluse-agent-ram.sh`, `pretooluse-bash-guard.sh`,
`pretooluse-nav-guard.sh`), two are pre-push (`prepush-clean-worktree-check.sh`,
`prepush-donewhen.sh`), one is pre-commit (`decisions-guard.sh`) and one is registered on both
`PreToolUse` and `Stop` (`design-sync-reminder.sh`); they are grouped by their status, not by
their event, because the events do not partition them. They are now
*observable* (`scripts/hook-fire-log.sh report`, #776), which is a weaker claim than proven and is
stated as such above.
- **`.husky/commit-msg` and `.husky/pre-commit`** — unproven, and each carries its own clause, so
they are a backlog too rather than "covered by what they call". `pre-commit` dispatches to
`lint-staged` and `decisions-guard.sh`, but the root-level-`*.png` refusal and the
`dotnet format whitespace --verify-no-changes` block are its own. `commit-msg` is not a dispatcher
at all: it is one `grep -q '^Co-Authored-By:'` plus a `MERGE_HEAD` exemption, and nothing tests
that the exemption fires only for merges. Neither is covered by what it calls — the trap is to read
"it dispatches to guards" as "its own clauses are guarded".
- **`ci-detect-already-validated.sh`, `ci-detect-docs-only.sh`, `ci-prove-ban-detects.sh`,
`e2e-functional.sh`, `check-kickoff-guard.sh`, `test_ci_release_path_scan_job.py`** — not
unattended either, but each needs its own harness rather than a proof written to clear a row.
`ci-prove-ban-detects.sh` is the exception noted in `testing.guard-ships-with-mutation-proof`: it
runs its own mutation at CI time, because what it proves is disarmable from inside pytest.
That list is prose and nothing checks it, which is the honest limit — a guard moving out of a group
will not redden anything. It is here so the next session inherits *why* a row is `NONE`, which is
the distinction #785 asked for; the machine-checked half remains the table.
Filling the rest is tracked rather than done in one pass, deliberately: a mutation proof written to
close a row is the kind of test that passes for the wrong reason.