Commit Graph
3 Commits
Author SHA1 Message Date
timothyandClaude Opus 5 0c5938dd60 fix(774): the wiring check counted a MENTION, reproducing inside the fix the defect it closed
Self-audit before re-review, and it found one. `wired_hook_files()` was added to stop hook
EXISTENCE standing in for hook WIRING — but it substring-matched the filename against the
whole husky text, and `.husky/pre-commit:7` reads

    # CI where a base ref exists). Fail-open shim — see .claude/hooks/decisions-guard.sh.

one line above the real invocation. Delete line 8, keep line 7, and the hook still reads as
wired. That is mention-for-invocation, which is the exact substitution the function exists
to prevent, one line inside the fix for it. Comment lines are now stripped from the husky
hooks first; settings.json needs no stripping because JSON has no comments.

Proven both ways: with the invocation removed and the comment left, the guard names
decisions-guard.sh as unwired; clean tree stays green.

Also verified rather than assumed, since a fix round is where adjacent defects live:
  - a stale SELF-referencing proof ref is still caught (the self-reference escape hatch
    skips only the PROOF-row classification check, not the def-existence check);
  - a reworded summary is LOUD, not vacuous — an unparsed summary fails with a message
    saying so, rather than silently checking nothing.

ruff clean, pyright clean (0 errors) on the three new files. Deliberately NOT ruff-format-ed:
the pre-existing scripts/tests corpus is not formatted either, so reformatting only these
three would diverge them from every sibling and bake in a format derived from an
un-versioned config on one machine — which is the divergence #780 exists to settle.

580 script-tests pass.

Refs #774

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 22:26:25 +02:00
timothyandClaude Opus 5 3473a6c889 fix(774,775): close the cold-review findings — including three the change inflicted on itself
Two independent cold reviews (Codex GPT-5.6 cross-family; Fable 5 on the patch) both
returned BLOCKED. They agreed on the counts error and the extractor hole; each found
things the other did not. Fixes, with what each was:

THE INVENTORY DID NOT COVER ITS OWN NEW GUARDS. `_SCRIPT_REF` matched `scripts/name.py`
but not `scripts/tests/*.py`, so the three guard files this change introduced had no rows
and the completeness check stayed green. A completeness guard blind to its author's new
guards is precisely the defect being legislated against. The population now globs
`scripts/tests/test_*.py` — which is how they actually run, since pr-checks.yml invokes
the directory. 32 rows -> 48.

That forced a third Kind. Once test files are in the population, every mutation proof
becomes a row wanting a proof of its own, forever. `PROOF` marks a file whose job is to
prove another guard; a scripts/tests file enforcing a repo invariant with no separate
guard behind it stays GUARD and may cite a mutation case in its own file.

HOOK EXISTENCE WAS STANDING IN FOR HOOK WIRING. Deleting a hook's registration from
.claude/settings.json left the population and the table unchanged, so the row went on
describing a guard that no longer ran — #631's shape one level down. Now derived from
settings.json plus the husky hooks.

THE SUMMARY COUNTS WERE A HAND-KEPT MIRROR AND WERE WRONG ON ARRIVAL: "28 guards, 4
tooling ... 19 have none" against a table holding 27/5/6/3/18. Both reviewers found it
independently. The prose is now parsed and asserted against the table.

TWO FALSE MUTATION GRADES, each with a concrete disarm:
  - test_full_first_page_alone_does_not_end_enumeration sends 50 docs paths then one more
    docs path; disarm pagination to treat a full page as final and it is still all-docs,
    still exempt, still green. Re-pointed at test_protected_path_on_a_LATER_page_is_still_seen,
    which does go red under that mutation.
  - test_the_scan_job_runs_the_out_of_pytest_positive_control asserts only that the script
    exists, is executable, is referenced and is marked; replace its logic with `exit 0` and
    all four pass. ci-prove-ban-detects.sh regraded NONE.

The MUTATION column was also being applied as a curve: three rows graded MUTATION fed the
real script an input only that clause rejects, which is what the rows eight lines away are
graded BEHAVIOUR-ONLY for. Definition sharpened to *witnessed* rather than plausible, and
those regraded. 5 MUTATION / 6 BEHAVIOUR-ONLY / 21 NONE across 32 guards.

THE VOCABULARY EXTRACTOR COULD RETURN A PARTIAL SET. `[A-Z|-]` cannot match `SHIP*)`, so
adding that arm leaves the extracted set non-empty AND equal to the read side — parity
green while the gate desyncs. Emptiness checks cannot see partial degradation. A loose
counterpart now asserts the strict pattern consumed every arm; proven red on exactly that
attack and green on a clean tree. Also: each verdict pattern must be assigned once, since
the extractor unions assignments while the classifier runs the last.

Also: docker-build.yml was itself an unchecked scope mirror (now asserted to be the only
workflow with toolchain container jobs, by parsing container.image rather than grepping —
ci-image.yml names the image because it builds it); the mutant floor is an equality;
e2e-functional.sh reclassified GUARD (it exits 1 on a failed contract assertion);
design-sync-reminder.sh does block the first Stop. The doc now states all six excluded
classes instead of one.

Not done here, filed instead: workflow-owned execution-class metadata to replace
TOOLCHAIN_JOBS, a single shared verdict vocabulary, and an executable clause-level
mutation harness. Each touches a CI-gating or merge-gate path and wants its own review.

580 script-tests pass.

Refs #774
Refs #775

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 22:20:51 +02:00
timothyandClaude Opus 5 0bd59b0b6e feat(774,775): one rule for guard populations, one for guard proofs — both enforced
#773's analysis found that the largest recorded failure family is reasoning about a
representative instead of the population (39% of process-failure records), and that the
most common is a check that never ran at all (25%). Both rules had been reinvented
repeatedly and written down nowhere.

Two decision records:

  testing.guard-derives-population-from-source (#774) — a guard enumerates its population
  from a machine-readable authoritative source and asserts set equality both ways. States
  the boundary that keeps it honest: filtering to select the SUBJECT of a per-member
  property is fine; filtering the population before a COMPLETENESS claim is the defect.
  Also separates guard SCOPE (a reviewable policy choice) from guard POPULATION (always
  derived).

  testing.guard-ships-with-mutation-proof (#775) — disarm that clause alone and a named
  test must go red. Behaviour-only coverage is graded separately, because it proves the
  guard reacts, never that it is connected.

Audit findings fixed:

  ci-image-pin stated an invariant it did not check. Its error text says "Every container:
  job must pin ersatztv-ci:<7-char-sha>"; what it asserts is that `grep … | sort -u` yields
  one DISTINCT value. Distinctness is a property of the pins present, so deleting the
  container: block from `test` leaves four pins, one distinct value, and a REQUIRED context
  silently running on the bare runner. test_ci_image_pin_population.py adds the population
  check, keyed on a reviewed registry cross-checked both ways — set equality between two
  DERIVED sets could not see this, because both sides shrink together.

  The verdict vocabulary was written down twice with no cross-check — post-review-verdict.sh
  (write) and check-review-verdict.sh (read). A word in one and not the other sends the
  required status green while the hook still denies. Both vocabularies are now extracted
  from their own source and compared as sets; a test that restated the words would just be
  a third copy. The write side's comment pointing at pretooluse-merge-consent.sh was also
  stale — the hook carries no copy and delegates.

Mechanical enforcement, answered explicitly for both:

  No to a filter-shaped-guard lint. The token is not the defect — ToolCatalogTests filters
  correctly eight lines from a completeness assertion that must not — and it would be a
  string predicate over source, which this repo's record says takes 3+ rounds. Building it
  would be #774 violating #774.

  Yes to enforcing the bookkeeping. docs/guard-inventory.md classifies all 32 guard files;
  test_guard_inventory.py derives the population from the filesystem and call sites,
  asserts set equality both ways, and resolves every claimed proof ref to a real def. A new
  guard cannot ship unclassified; a renamed test cannot leave a row claiming lost coverage.
  What it does NOT check — whether a MUTATION claim is true — is stated, not implied.

Measured: 28 guards, 4 tooling. 6 mutation-proved, 3 behaviour-only, 19 unproven.

Every guard added here was mutation-proved by execution before being believed: neutering
pin_population_faults turned 20 of 25 red; the inventory guard was driven red three ways
(deleted row, new unclassified hook, stale proof ref) and restored green.

573 script-tests pass. Scope limit stated in the doc: inline workflow-job guards are not
in the machine-checked population.

Refs #774
Refs #775

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 21:56:52 +02:00