499dd348abc371d94f0839bdfd8de8e1edec72df
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c5938dd60 |
fix(774): the wiring check counted a MENTION, reproducing inside the fix the defect it closed
Self-audit before re-review, and it found one. `wired_hook_files()` was added to stop hook
EXISTENCE standing in for hook WIRING — but it substring-matched the filename against the
whole husky text, and `.husky/pre-commit:7` reads
# CI where a base ref exists). Fail-open shim — see .claude/hooks/decisions-guard.sh.
one line above the real invocation. Delete line 8, keep line 7, and the hook still reads as
wired. That is mention-for-invocation, which is the exact substitution the function exists
to prevent, one line inside the fix for it. Comment lines are now stripped from the husky
hooks first; settings.json needs no stripping because JSON has no comments.
Proven both ways: with the invocation removed and the comment left, the guard names
decisions-guard.sh as unwired; clean tree stays green.
Also verified rather than assumed, since a fix round is where adjacent defects live:
- a stale SELF-referencing proof ref is still caught (the self-reference escape hatch
skips only the PROOF-row classification check, not the def-existence check);
- a reworded summary is LOUD, not vacuous — an unparsed summary fails with a message
saying so, rather than silently checking nothing.
ruff clean, pyright clean (0 errors) on the three new files. Deliberately NOT ruff-format-ed:
the pre-existing scripts/tests corpus is not formatted either, so reformatting only these
three would diverge them from every sibling and bake in a format derived from an
un-versioned config on one machine — which is the divergence #780 exists to settle.
580 script-tests pass.
Refs #774
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
3473a6c889 |
fix(774,775): close the cold-review findings — including three the change inflicted on itself
Two independent cold reviews (Codex GPT-5.6 cross-family; Fable 5 on the patch) both returned BLOCKED. They agreed on the counts error and the extractor hole; each found things the other did not. Fixes, with what each was: THE INVENTORY DID NOT COVER ITS OWN NEW GUARDS. `_SCRIPT_REF` matched `scripts/name.py` but not `scripts/tests/*.py`, so the three guard files this change introduced had no rows and the completeness check stayed green. A completeness guard blind to its author's new guards is precisely the defect being legislated against. The population now globs `scripts/tests/test_*.py` — which is how they actually run, since pr-checks.yml invokes the directory. 32 rows -> 48. That forced a third Kind. Once test files are in the population, every mutation proof becomes a row wanting a proof of its own, forever. `PROOF` marks a file whose job is to prove another guard; a scripts/tests file enforcing a repo invariant with no separate guard behind it stays GUARD and may cite a mutation case in its own file. HOOK EXISTENCE WAS STANDING IN FOR HOOK WIRING. Deleting a hook's registration from .claude/settings.json left the population and the table unchanged, so the row went on describing a guard that no longer ran — #631's shape one level down. Now derived from settings.json plus the husky hooks. THE SUMMARY COUNTS WERE A HAND-KEPT MIRROR AND WERE WRONG ON ARRIVAL: "28 guards, 4 tooling ... 19 have none" against a table holding 27/5/6/3/18. Both reviewers found it independently. The prose is now parsed and asserted against the table. TWO FALSE MUTATION GRADES, each with a concrete disarm: - test_full_first_page_alone_does_not_end_enumeration sends 50 docs paths then one more docs path; disarm pagination to treat a full page as final and it is still all-docs, still exempt, still green. Re-pointed at test_protected_path_on_a_LATER_page_is_still_seen, which does go red under that mutation. - test_the_scan_job_runs_the_out_of_pytest_positive_control asserts only that the script exists, is executable, is referenced and is marked; replace its logic with `exit 0` and all four pass. ci-prove-ban-detects.sh regraded NONE. The MUTATION column was also being applied as a curve: three rows graded MUTATION fed the real script an input only that clause rejects, which is what the rows eight lines away are graded BEHAVIOUR-ONLY for. Definition sharpened to *witnessed* rather than plausible, and those regraded. 5 MUTATION / 6 BEHAVIOUR-ONLY / 21 NONE across 32 guards. THE VOCABULARY EXTRACTOR COULD RETURN A PARTIAL SET. `[A-Z|-]` cannot match `SHIP*)`, so adding that arm leaves the extracted set non-empty AND equal to the read side — parity green while the gate desyncs. Emptiness checks cannot see partial degradation. A loose counterpart now asserts the strict pattern consumed every arm; proven red on exactly that attack and green on a clean tree. Also: each verdict pattern must be assigned once, since the extractor unions assignments while the classifier runs the last. Also: docker-build.yml was itself an unchecked scope mirror (now asserted to be the only workflow with toolchain container jobs, by parsing container.image rather than grepping — ci-image.yml names the image because it builds it); the mutant floor is an equality; e2e-functional.sh reclassified GUARD (it exits 1 on a failed contract assertion); design-sync-reminder.sh does block the first Stop. The doc now states all six excluded classes instead of one. Not done here, filed instead: workflow-owned execution-class metadata to replace TOOLCHAIN_JOBS, a single shared verdict vocabulary, and an executable clause-level mutation harness. Each touches a CI-gating or merge-gate path and wants its own review. 580 script-tests pass. Refs #774 Refs #775 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0bd59b0b6e |
feat(774,775): one rule for guard populations, one for guard proofs — both enforced
#773's analysis found that the largest recorded failure family is reasoning about a representative instead of the population (39% of process-failure records), and that the most common is a check that never ran at all (25%). Both rules had been reinvented repeatedly and written down nowhere. Two decision records: testing.guard-derives-population-from-source (#774) — a guard enumerates its population from a machine-readable authoritative source and asserts set equality both ways. States the boundary that keeps it honest: filtering to select the SUBJECT of a per-member property is fine; filtering the population before a COMPLETENESS claim is the defect. Also separates guard SCOPE (a reviewable policy choice) from guard POPULATION (always derived). testing.guard-ships-with-mutation-proof (#775) — disarm that clause alone and a named test must go red. Behaviour-only coverage is graded separately, because it proves the guard reacts, never that it is connected. Audit findings fixed: ci-image-pin stated an invariant it did not check. Its error text says "Every container: job must pin ersatztv-ci:<7-char-sha>"; what it asserts is that `grep … | sort -u` yields one DISTINCT value. Distinctness is a property of the pins present, so deleting the container: block from `test` leaves four pins, one distinct value, and a REQUIRED context silently running on the bare runner. test_ci_image_pin_population.py adds the population check, keyed on a reviewed registry cross-checked both ways — set equality between two DERIVED sets could not see this, because both sides shrink together. The verdict vocabulary was written down twice with no cross-check — post-review-verdict.sh (write) and check-review-verdict.sh (read). A word in one and not the other sends the required status green while the hook still denies. Both vocabularies are now extracted from their own source and compared as sets; a test that restated the words would just be a third copy. The write side's comment pointing at pretooluse-merge-consent.sh was also stale — the hook carries no copy and delegates. Mechanical enforcement, answered explicitly for both: No to a filter-shaped-guard lint. The token is not the defect — ToolCatalogTests filters correctly eight lines from a completeness assertion that must not — and it would be a string predicate over source, which this repo's record says takes 3+ rounds. Building it would be #774 violating #774. Yes to enforcing the bookkeeping. docs/guard-inventory.md classifies all 32 guard files; test_guard_inventory.py derives the population from the filesystem and call sites, asserts set equality both ways, and resolves every claimed proof ref to a real def. A new guard cannot ship unclassified; a renamed test cannot leave a row claiming lost coverage. What it does NOT check — whether a MUTATION claim is true — is stated, not implied. Measured: 28 guards, 4 tooling. 6 mutation-proved, 3 behaviour-only, 19 unproven. Every guard added here was mutation-proved by execution before being believed: neutering pin_population_faults turned 20 of 25 red; the inventory guard was driven red three ways (deleted row, new unclassified hook, stale proof ref) and restored green. 573 script-tests pass. Scope limit stated in the doc: inline workflow-job guards are not in the machine-checked population. Refs #774 Refs #775 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |