Build ErsatzTV Image / CI toolchain image resolves (pull_request) Successful in 35s
Build ErsatzTV Image / Delimiter ban (release path) (pull_request) Successful in 57s
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 37s
PR Gates / Docs update reminder (pull_request) Successful in 1m0s
PR Gates / decisions lifecycle (pull_request) Successful in 20s
PR Gates / Fix proofs (Proves trailers) (pull_request) Successful in 17s
review-verdict/h10 Review-verdict: MERGEABLE @ a7d91bf (base: main)
Review verdict / Set review-verdict status (pull_request_target) Successful in 45s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 9m25s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 6m17s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Skipped
PR Gates / Script lint and tests (ruff + pytest) (pull_request) Successful in 19m27s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 6m4s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 8s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 7s
`docs.no-session-narrative` reaches every durable artifact, but its detector scanned only `docs/**/*.md` and root markdown, and nothing had ever swept the rest. The issue named four sites from one grep and called them a floor. Deriving the population instead — a whitespace-joined sweep over every tracked file outside the detector, for the detector's own phrasings plus the attribution and review-round class #812 found — gave 453 sites in 108 files at `fb5592971`, and a second pass for phrasings the first list missed (hyphenated `round-N`, "an earlier version", "the reviewer proved") added residuals in the same files. Every site was classified with #812's three dispositions (CUT / SEVER / KEEP with its sub-kind) under the who-benefits test; the per-site manifests are on the PR. The rejected designs, tested-and-rejected fixtures, measurements and traps stay; the attribution of who found them and the round in which they were found go. The detector's population grows to `.claude/`, `.gitea/`, `.husky/` and `scripts/` regardless of extension, minus the detector and its own test (whose fixtures ARE the phrasings) and minus `scripts/tests/fixtures/` (test data, including decision-record copies — the same reasoning as the records' own exemption, and what keeps the record's depth measurement true), and `--all` lists tracked REGULAR files only — a symlink's content is its target and a gitlink has none. The #812 argument for leaving `docs/superpowers/**` in the population runs the other way here: `--diff` sees only ADDED lines, and 287 of the 453 sites were under 30 days old — this corpus is where narrative is being added, so the advisory nudge has reach. Density agrees: 56 line-mode hits over the 113 regular files the predicate admits, against 9 over 66 docs files before #812. `web/` and C# stay out on the same measurement (3 of 74 PATTERNS-matching sites, ~4,600 files). The predicate did not grow: PATTERNS matched 74 of 453 sites, and widening the word list to the attribution class is the treadmill the withdrawn parity test ran on. The population oracle is restated over segments with the new arms, the synthetic cross product gains the process heads and non-markdown extensions, a fixture witnesses that a tracked symlink is neither scanned nor counted, a `.py.bak` axis separates a by-name exemption from a `startswith` over the same tuple, and eight mutants (drop the process arm, drop the by-name exemption, exempt by `startswith`, drop or add a prefix, drop the fixtures exemption, list only markdown, drop the symlink filter, test the mode per row instead of per path) each redden it. A pre-existing silent drop in `--diff` goes with it: git tab-terminates a `+++` filename that contains a space, and the kept tab made `is_scanned_path` refuse the file with no notice — fixed, with a positive control and its own mutant. Code is unchanged by construction, measured per file type against `origin/main`: Python modules are AST-equal with docstrings stripped, except `#` lines inside the embedded fixture programs (string literals) of three test modules; workflows differ only in `#` lines inside `run:` block scalars; shell, C#, TypeScript and jq are equal with comment lines stripped. The stated exceptions: the detector and its test, 26 vitest titles that carried review-round or severity labels or a reviewer attribution (call sites whose title changed — every changed title line walked back to its `it(` / `it.each(...)(` anchor, so a `' + '` concatenation counts once), two registry note strings and the mutation manifest's prose fields. scripts/tests: 1565 passed. Web: lint, typecheck, 1319 tests green. Closes #876. Decisions-Edit: yes Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PEcBoFw7ctrf3Nb7R7x7wk
547 lines
37 KiB
Python
547 lines
37 KiB
Python
"""The DECLARED clause mutations, one per `MUTATION`-graded row of `docs/guard-inventory.md`.
|
|
|
|
Data only. The machinery that applies these is `mutation_harness_lib.py`; the checks that keep this
|
|
file honest are `test_mutation_harness.py`.
|
|
|
|
Every entry is declared by hand and none is inferred, which is the whole design constraint from
|
|
ersatztv#790: "a harness that guesses which clause of a 90-line hook is *the* guard would manufacture
|
|
exactly the confident-but-empty coverage this is meant to prevent". Where a proof test already names
|
|
its own clause in source — the BOM guard's `= "efbbbf" ]; then`, `UNSET_CLAUSE`, `prove-fix.sh`'s
|
|
`if [ "$RC" -eq 0 ]; then` — the entry reuses THAT string rather than inventing a second one, so a
|
|
retarget in either place is caught by the other.
|
|
|
|
WHY AN ENTRY'S `target` MAY DIFFER FROM ITS `guard`. Some guards here ARE tests
|
|
(`scripts/tests/test_*.py`). Disarming such a guard makes it ABSENT rather than red, so
|
|
`testing.guard-ships-with-mutation-proof`'s checker-guard exception applies: the mutation goes into
|
|
the guarded ARTIFACT — a deleted row, a planted phantom row — and the check must report it. Mutating
|
|
a checker's own POPULATION instead is a trap that looks identical and is not: a shrunken population
|
|
makes every real row report as PHANTOM, so the proof reddens on a false positive while saying
|
|
nothing about the missing-row detection the row claims. `why` states per entry which shape applies
|
|
and why; no count is kept here, because a count of the entries below is a second copy of them.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from scripts.tests.mutation_harness_lib import Mutation
|
|
|
|
CLAUSE = Mutation.CLAUSE
|
|
DETECTOR = Mutation.DETECTOR
|
|
|
|
MUTATIONS: tuple[Mutation, ...] = (
|
|
Mutation(
|
|
guard=".claude/hooks/posttooluse-worktree-marker.sh",
|
|
target=".claude/hooks/posttooluse-worktree-marker.sh",
|
|
clause='printf \'%s\\n\' "$me" > "$abs/.claude-worktree-owner" 2>/dev/null || true',
|
|
replacement="true",
|
|
proof="test_worktree_ownership_guard.py::test_MUTATION_a_marker_hook_that_stops_WRITING_makes_the_guard_go_quiet",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED pair did not deny",
|
|
why="The marker write is the hook's entire job; without it the guard has nothing to read and "
|
|
"fails open. The clause string is the one the proof test itself passes to its `_mutate` helper.",
|
|
),
|
|
Mutation(
|
|
guard=".claude/hooks/pretooluse-bom-guard.sh",
|
|
target=".claude/hooks/pretooluse-bom-guard.sh",
|
|
clause='= "efbbbf" ]; then',
|
|
replacement='= "deadbeef" ]; then',
|
|
proof="test_bom_guard_detection.py::test_DISARMING_the_BOM_comparison_stops_detection",
|
|
granularity=CLAUSE,
|
|
expect="the BOM comparison has moved",
|
|
why="The BOM comparison is the guard's only detection logic. Same clause the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard=".claude/hooks/pretooluse-worktree-guard.sh",
|
|
target=".claude/hooks/pretooluse-worktree-guard.sh",
|
|
clause='marker="$root/.claude-worktree-owner"',
|
|
replacement='marker="$root/.claude-worktree-owner-NOTHING-WRITES-THIS"',
|
|
proof="test_worktree_ownership_guard.py::test_MUTATION_disarming_the_guards_MARKER_READ_stops_the_deny",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED guard did not deny",
|
|
why="The marker read is what the ownership decision hangs on. Same clause the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard=".husky/pre-push",
|
|
target=".husky/pre-push",
|
|
clause="unset GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE",
|
|
replacement=": # clause removed by the mutation harness",
|
|
proof="test_prepush_unsets_git_env.py::test_MUTATION_DELETING_the_unset_lets_drift_through_silently",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED pre-push did not catch the drift",
|
|
why="Without the unset, every git call the pre-push chain makes is aimed at the repository git "
|
|
"exported the environment for, not the one being pushed. `UNSET_CLAUSE` in the proof test.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/build_decisions_catalog.py",
|
|
target="scripts/build_decisions_catalog.py",
|
|
clause="want.strip() != have.strip()",
|
|
replacement="False",
|
|
proof="test_build_catalog_check_path.py::test_MUTATION_disarming_the_stale_comparison_stops_detection",
|
|
granularity=CLAUSE,
|
|
expect="the stale-detection clause has moved or been reworded",
|
|
why="`main()`'s only stale-detection logic, per the proof test's own docstring, which uses this "
|
|
"exact clause and this exact replacement.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-detect-docs-only.sh",
|
|
target="scripts/ci-detect-docs-only.sh",
|
|
clause='if [ "$is_shallow" = "true" ]; then',
|
|
replacement="if true; then",
|
|
proof="test_docs_only_detector_clone_depth.py::test_the_push_arm_leaves_a_COMPLETE_clone_complete",
|
|
granularity=CLAUSE,
|
|
expect="GRAFTED the complete clone shallow",
|
|
why="The shallow test is the whole of ersatztv#836's fix: disarmed, `--depth=2` goes back to "
|
|
"every checkout including `build`'s complete one, which grafts it and makes the `git describe` "
|
|
"in the next step find no reachable tag. This row was UNDECLARED until #836 on the stated "
|
|
"grounds that the script 'feeds the skip gate, so its effect is visible only in a workflow "
|
|
"run' — untrue of this clause, whose effect is the shallow flag on a real clone and is "
|
|
"observable in-process, which is what the proof test asserts.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-step-ran.sh",
|
|
target="scripts/ci-step-ran.sh",
|
|
clause='if ! grep -qxF "$key" "$marker" 2>/dev/null; then',
|
|
replacement="if false; then",
|
|
proof="test_ci_dropped_step_guard.py::test_dropping_ANY_single_step_FAILS_the_guard",
|
|
granularity=CLAUSE,
|
|
expect="never having executed",
|
|
why="The per-key membership test is what turns a dropped step into a red job; disarmed, every "
|
|
"expected key reads as present and the guard passes a run in which nothing executed.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_ci_status_context_uniqueness.py",
|
|
target=".gitea/workflows/docker-build.yml",
|
|
clause=" name: API docs in sync (OpenAPI + endpoint index)",
|
|
replacement=" name: Build & test (.NET)",
|
|
proof="test_ci_status_context_uniqueness.py::test_no_two_jobs_synthesize_the_SAME_status_context",
|
|
granularity=CLAUSE,
|
|
expect="these status contexts can be produced by more than one job",
|
|
why="THE CHECKER-GUARD SHAPE: this guard IS a test, so disarming its own assertion makes it "
|
|
"absent rather than red. Per `testing.guard-ships-with-mutation-proof` the mutation goes into "
|
|
"the guarded ARTIFACT instead — here a workflow job renamed to collide with the REQUIRED "
|
|
"`Build & test (.NET)` job, which is exactly the evasion the guard exists to catch: two jobs "
|
|
"synthesizing one status context that branch protection cannot tell apart, so the required "
|
|
"check could be satisfied by the producer whose steps carry no execution markers.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/check-required-contexts.sh",
|
|
target="scripts/check-required-contexts.sh",
|
|
clause="($rule.status_check_contexts | sort | unique) == $snap",
|
|
replacement="true",
|
|
proof="test_check_required_contexts.py::test_MUTATION_disarming_the_set_comparison_stops_every_drift_report",
|
|
expect="the set-comparison clause has moved or been reworded",
|
|
granularity=CLAUSE,
|
|
why="The set comparison IS the finding: it is the only thing that turns a live required-check "
|
|
"list differing from the committed snapshot into `drift`. Witnessed: unmutated reports "
|
|
"`drift` on an added context, the mutant reports `match` — a permanent no-op that would "
|
|
"confirm the snapshot fresh forever. The proof asserts the mutant's EXACT verdict "
|
|
"`(0, 'match')` rather than merely 'not drift', and ships a positive control for its own "
|
|
"tmp layout: a copy of the script without the classifier it loads beside "
|
|
"itself exits 2 with empty stdout, and 'not drift' is then satisfied by a copy "
|
|
"that never ran. The clause string is the one the proof test asserts on before mutating.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/decisions_validate.py",
|
|
target="scripts/decisions_validate.py",
|
|
clause="wing_faults=record_wing_faults() + yaml_faults,",
|
|
replacement="wing_faults=yaml_faults,",
|
|
proof="test_decisions_validate.py::test_main_actually_CALLS_the_wing_scan",
|
|
granularity=CLAUSE,
|
|
expect="a wing fault must fail the validator",
|
|
why="The wiring the proof test exists for: deleting this call left the whole suite green while "
|
|
"a real block-scalar record vanished under `decisions-validate: OK` (#609).",
|
|
),
|
|
Mutation(
|
|
guard="scripts/prove-fix.sh",
|
|
target="scripts/prove-fix.sh",
|
|
clause='if [ "$RC" -eq 0 ]; then',
|
|
replacement="if false; then",
|
|
proof="test_prove_fix.py::test_MUTATION_disarming_the_UNPROVEN_clause_reddens_the_refusal_test",
|
|
granularity=CLAUSE,
|
|
expect="the clause under mutation is gone",
|
|
why="The UNPROVEN branch: a named test that passes WITHOUT the fix must be refused. Same clause "
|
|
"the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_ci_image_paths_pin_agreement.py",
|
|
target="scripts/tests/test_ci_image_paths_pin_agreement.py",
|
|
clause="for path in sorted(published - required):",
|
|
replacement="for path in sorted(set()):",
|
|
proof="test_ci_image_paths_pin_agreement.py::test_a_diverging_list_is_DETECTED",
|
|
granularity=CLAUSE,
|
|
expect="the agreement check accepted a workflow pair whose publish paths and pin pathspec",
|
|
why="The two set-difference clauses fail differently and neither implies the other, so each "
|
|
"is separately load-bearing: this one catches a path that PUBLISHES an image the pin never "
|
|
"tracks (silent and green), while `required - published` catches a pathspec entry that "
|
|
"never publishes (a red on a blocking job until an image for it is published). Disarming "
|
|
"this clause alone leaves "
|
|
"the `publish-path-added` mutant undetected while the other four cases stay green, which "
|
|
"is why the grade is CLAUSE rather than DETECTOR.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_ci_image_pin_population.py",
|
|
target="scripts/tests/test_ci_image_pin_population.py",
|
|
clause="for name in sorted(expected - set(pinned)):",
|
|
replacement="for name in sorted(set()):",
|
|
proof="test_ci_image_pin_population.py::test_a_single_job_losing_its_pin_is_DETECTED",
|
|
granularity=CLAUSE,
|
|
expect="the population check accepted a workflow in which a container job no longer runs",
|
|
why="The against-the-DECLARED-CLASS direction, and the one the other two clauses cannot "
|
|
"cover: a job that loses its `container:` block leaves `declared` and `pinned` equal, so "
|
|
"only this comparison notices it has moved to the bare runner. This is the clause #790 "
|
|
"asked for instead of neutering `pin_population_faults` wholesale. `expected` is "
|
|
"`toolchain_declared(doc)` since ersatztv#789 replaced the `TOOLCHAIN_JOBS` literal this "
|
|
"clause used to name; the comparison and its role are unchanged.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_guard_populations_derive_from_git.py",
|
|
target="scripts/tests/test_guard_inventory.py",
|
|
clause="if ref in tracked:",
|
|
replacement="if (REPO_ROOT / ref).exists():",
|
|
proof="test_guard_populations_derive_from_git.py::test_no_derivation_admits_an_untracked_file",
|
|
granularity=CLAUSE,
|
|
expect="after git stopped tracking them",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — one of the "
|
|
"derivations it watches — rather than into the checker, per the checker-guard exception. "
|
|
"The clause is the exact defect this guard was written after: `derived_guard_files` read "
|
|
"its CALLERS from the index and then admitted the paths they name on `Path.exists()`, so a "
|
|
"tracked workflow naming a script that exists on one machine only entered the population "
|
|
"there, red on that checkout and green in CI (#778's third shape, "
|
|
"inside #806 itself). Note what this mutation does NOT do: on a clean tree the mutated set "
|
|
"is identical, so `test_guard_inventory.py`'s own assertions stay green — only narrowing "
|
|
"the index, which is what the proof does, separates them. That is why the proof has to "
|
|
"remove EVERY member rather than sample one.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_optional_request_members.py",
|
|
target="scripts/tests/test_optional_request_members.py",
|
|
clause='"ArtworkContentTypeModel": (',
|
|
replacement='"ArtworkContentTypeModelRENAMED": (',
|
|
proof="test_optional_request_members.py::test_every_droppable_request_schema_has_a_stated_disposition",
|
|
granularity=CLAUSE,
|
|
expect="no disposition written down",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — here the "
|
|
"DISPOSITIONS registry the checker maintains, the same shape as the deleted "
|
|
"guard-inventory row below. Renaming the key rather than deleting the entry keeps the "
|
|
"module importable, so the red is a real set-equality failure and not an ImportError "
|
|
"reddening for the wrong reason. The rename fires BOTH directions — MISSING for the real "
|
|
"schema and PHANTOM for the renamed key — which is the correct behaviour and worth stating, "
|
|
"since `expect` names only the MISSING half. "
|
|
"`ArtworkContentTypeModel` is the right key to name: it is the exact schema #807's "
|
|
"hand-written table omitted, because `...Model` reads as a response model while it is in "
|
|
"fact reachable from the full-replace PUT /channels/{id}.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_workflow_job_guards.py",
|
|
target="docs/guard-inventory.md",
|
|
clause="| `pr-checks.yml::ci-image-pin` |",
|
|
replacement="| `pr-checks.yml::ci-image-pin-RENAMED` |",
|
|
proof="test_workflow_job_guards.py::test_the_inventory_covers_exactly_the_guard_JOBS_that_exist",
|
|
granularity=CLAUSE,
|
|
expect="guard/report-only but have NO row in",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — the workflow-job "
|
|
"table — rather than into the checker, per the checker-guard exception. RENAMED rather than "
|
|
"deleted, so the row count is unchanged and the red cannot come from an empty table: the "
|
|
"anti-vacuity test still sees rows, and what fails is the set equality itself, in BOTH "
|
|
"directions at once (MISSING for the real job, PHANTOM for the renamed key). "
|
|
"`ci-image-pin` is the right row to name — it is the guard job #774 found stating an "
|
|
"invariant it did not check, and the case #786 was filed to bring into a population.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_guard_inventory.py",
|
|
target="docs/guard-inventory.md",
|
|
clause="| `.claude/hooks/decisions-guard.sh` | a commit | GUARD | NONE | — |\n",
|
|
replacement="",
|
|
proof="test_guard_inventory.py::test_the_inventory_covers_exactly_the_guards_that_exist",
|
|
granularity=CLAUSE,
|
|
expect="these guard files exist but have no row in guard-inventory.md",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT rather than into the "
|
|
"checker — disarming a checker makes it absent, not red, and mutating its population instead "
|
|
"would only demonstrate a false POSITIVE (a shrunken population reports every real row as "
|
|
"phantom) while proving nothing about the missing-row detection the row claims. A deleted "
|
|
"row is the defect this guard exists to catch, and it is one of the mutations #774 witnessed "
|
|
"by hand.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_hook_fire_log.py",
|
|
target="scripts/tests/test_hook_fire_log.py",
|
|
clause=" if mentions != [CANONICAL_SINK_ASSIGNMENT, CANONICAL_SINK_SOURCE]:",
|
|
replacement=" if False:",
|
|
proof="test_hook_fire_log.py::test_a_LATER_reassignment_the_regex_cannot_see_is_DETECTED",
|
|
granularity=CLAUSE,
|
|
expect="reassignment of the SOURCED sink passed clean",
|
|
why="REGRADED FROM `DETECTOR` BY ersatztv#891, because the guard acquired a clause that IS "
|
|
"load-bearing on its own. `instrumentation_faults` used to accumulate only from arms a "
|
|
"stripped hook trips several of at once, so disarming any one left the rest answering and no "
|
|
"clause-level mutation could redden the proof; the entry carried the finest surviving "
|
|
"mutation (`if not _SOURCES_SINK.search(text):` -> `if False:`) as its honesty check. #891 "
|
|
"replaced the lexical rule over the sink assignment with BYTE-IDENTITY of the preamble's two "
|
|
"lines, and that comparison decides alone: disarming it reddens this proof on every hook, "
|
|
"while nothing else catches an indented or `export`ed reassignment. The old canary had to go "
|
|
"with it rather than be carried on: byte-identity necessarily matches the sourcing line, so "
|
|
"`_SOURCES_SINK` can no longer decide anything (measured: 0 cases in 13 hooks x 8 "
|
|
"perturbations) and a mutation required to KEEP surviving would have been guaranteed to. "
|
|
"THE GENERAL SHAPE, because it is not specific to this entry: a `survived_clause` asserts "
|
|
"that a finer mutation still survives, so its precondition is that the guard is deliberately "
|
|
"coarse there. WIDENING the main clause can remove that precondition, and the canary then "
|
|
"does not fail — it becomes a tautology, which reads exactly like a passing proof. When a "
|
|
"clause is widened, re-derive every proof calibrated against the narrow one; a widened clause "
|
|
"can retire its own evidence.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_remote_state_inventory.py",
|
|
target="docs/remote-state-inventory.md",
|
|
clause="| `scripts/post-review-verdict.sh` — commit-status write |",
|
|
replacement="| `scripts/DELETED-BY-THE-MUTATION-HARNESS.sh` — not a real path |",
|
|
proof="test_remote_state_inventory.py::test_every_in_scope_file_has_a_row_and_every_row_names_a_real_file",
|
|
granularity=CLAUSE,
|
|
expect="in scope but absent from docs/remote-state-inventory.md",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT: a real executable's "
|
|
"row is renamed away, which is the MISSING-row defect the row's Blocks column claims — an "
|
|
"in-scope file with no classification. Two shapes were tried and rejected. Emptying the "
|
|
"guard's `git ls-files` derivation reddens the proof with an IndexError over an empty "
|
|
"population: a crash, not a detection. Planting a PHANTOM row reddens "
|
|
"`test_MUTATION_PROOF_a_dropped_row_and_a_phantom_row_are_both_detected` by contaminating "
|
|
"the fixture that test builds for itself, and proves the opposite direction from the one the "
|
|
"row claims. Renaming the row exercises both directions of the production set comparison at "
|
|
"once and is matched on the missing half.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_mutation_harness.py",
|
|
target="scripts/tests/mutation_harness_lib.py",
|
|
clause=" if mutation.expect not in diagnostic:",
|
|
replacement=" if False:",
|
|
proof="test_mutation_harness.py::test_MUTATION_disarming_the_DIAGNOSTIC_gate_accepts_a_red_for_the_wrong_reason",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED verdict already accepted it, so the mutant proves nothing",
|
|
why="The target is not the guard for a structural reason: the harness keeps its machinery in "
|
|
"`mutation_harness_lib.py` so a clause of it can be disarmed in an isolated copy at all. The "
|
|
"clause is the DIAGNOSTIC gate — the check that a failing proof failed with the diagnostic "
|
|
"its row declares. Disarmed, a red for any unrelated reason is certified as a guard doing "
|
|
"its job, which is the shape that made two rows in this very file measure nothing. The other "
|
|
"gate, the one requiring pytest exit code 1, carries its own proof in "
|
|
"`test_MUTATION_disarming_the_EXIT_STATUS_gate_accepts_a_run_that_NEVER_RAN_A_TEST`; the "
|
|
"inventory holds one ref per row, so this entry names the stronger of the two.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-toolchain-image-resolves.sh",
|
|
target="scripts/ci-toolchain-image-resolves.sh",
|
|
clause=" 404)",
|
|
replacement=" 4040)",
|
|
proof="test_ci_toolchain_image_resolves.py::test_MUTATION_a_deleted_tag_is_reported_as_a_failure",
|
|
granularity=CLAUSE,
|
|
expect="a deleted tag was not reported as GONE",
|
|
why="404 is the ONE answer that establishes the pinned toolchain image is gone; every other "
|
|
"code means the check could not run. Both fail the job, so the EXIT CODE does not separate "
|
|
"them and the mutation is caught by the DIAGNOSTIC instead: retargeting the arm sends the "
|
|
"real outage down the could-not-verify path, which sends an operator to the registry's "
|
|
"health rather than to the rebuild that fixes it.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_review_verdict_vocabulary.py",
|
|
target="scripts/check-review-verdict.sh",
|
|
clause="""POS_RE='^review-verdict:[[:space:]]*('"$POS_ALTERNATION"')([[:space:]@]|$)'""",
|
|
replacement="""POS_RE='^review-verdict:[[:space:]]*(mergeable|approved|lgtm)([[:space:]@]|$)'""",
|
|
proof="test_review_verdict_vocabulary.py::test_a_word_added_to_the_shared_source_reaches_BOTH_sides",
|
|
granularity=CLAUSE,
|
|
expect="restating the word list rather than deriving it",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — the read side "
|
|
"whose derivation it watches — rather than into the checker, per the checker-guard "
|
|
"exception. The replacement is not an invented mutant: it is the LITERAL pre-#788 line, the "
|
|
"second hand-written copy this change removed, so the proof is taken against the real "
|
|
"predecessor. Mutating the shared vocabulary instead would be the trap that looks "
|
|
"identical: emptying or corrupting the word list reddens the proof through the fail-closed "
|
|
"validator, which says nothing about whether the read side still DERIVES from it.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_workflow_persist_credentials.py",
|
|
target=".gitea/workflows/ci-image.yml",
|
|
clause=""" persist-credentials: false
|
|
# only docker/ci/Dockerfile is needed; no git describe/log here""",
|
|
replacement=""" # only docker/ci/Dockerfile is needed; no git describe/log here""",
|
|
proof="test_workflow_persist_credentials.py::test_every_actions_checkout_DROPS_the_persisted_credential",
|
|
granularity=CLAUSE,
|
|
expect="persists a write-capable Authorization header into .git/config",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT rather than into "
|
|
"the checker, per the checker-guard exception. The defect is the one the guard exists for: "
|
|
"an `actions/checkout` step that simply omits the key, which is what a newly added job gets "
|
|
"by default — not an explicit `true`, which nobody writes. `ci-image.yml` carries exactly "
|
|
"one checkout, so the clause is unambiguous there; the trailing comment line is part of it "
|
|
"only to make the match unique within the file.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_mcp_smoke.py",
|
|
target="scripts/mcp_smoke.py",
|
|
clause=" id_init = secrets.randbelow(2**31 - 1000) + 1000",
|
|
replacement=" id_init = 1",
|
|
proof="test_mcp_smoke.py::test_MUTATION_a_PRE_ANSWERED_id_is_refused_because_the_request_ids_are_UNGUESSABLE",
|
|
granularity=CLAUSE,
|
|
expect="a pre-answered id was ACCEPTED",
|
|
why="The target is not the guard for the same structural reason as the harness entry above: "
|
|
"`scripts/mcp_smoke.py` is a checker reached only transitively, through "
|
|
"`scripts/check-local-lsp.sh`, so it holds no inventory row of its own and the clause has to "
|
|
"be disarmed in it directly (`testing.verification-code-needs-its-own-proof`). The clause is "
|
|
"the unguessable request id, which AT THE `initialize` STAGE is the only thing refusing a "
|
|
"server that answers before it is asked: the pending-registration cannot help there, because "
|
|
"that id is already in flight when the pre-answer arrives, which is why #793 replaced "
|
|
"the lock rather than tightening it. Disarmed, the stub's pre-answer is "
|
|
"accepted at `initialize` and the run dies one stage later at `tools/list`, so the proof "
|
|
"asserts the STAGE (rc 9 and the initialize diagnostic) rather than mere failure: the mutant "
|
|
"still fails, and a test checking only that would stay green over a real vulnerability. THE "
|
|
"`id_tools` TWIN IS DELIBERATELY NOT THIS CLAUSE, measured rather than assumed: replacing "
|
|
"`id_tools` alone leaves every case green, because the reader keeps only the reply matching "
|
|
"the id in flight, so a `tools/list` frame emitted while `initialize` is pending is dropped "
|
|
"whatever its id. Disarming that retention clause alone is green too. Only BOTH together "
|
|
"produce the false green (rc 0), so the tools stage is held by two mechanisms that mask each "
|
|
"other and neither is singly detectable — #685's shape. `id_init` is the one place a "
|
|
"pre-answer IS singly exploitable, which is why it is the declared clause; the tools stage is "
|
|
"covered behaviourally in the same file.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_complete_annotation_dispositions.py",
|
|
target="web/src/api/completeAnnotations.guard.test.ts",
|
|
clause="MultiCollectionItemRequest: {\n disposition: 'ANNOTATED',",
|
|
replacement="MultiCollectionItemRequest: {\n disposition: 'CREATE',",
|
|
proof="test_complete_annotation_dispositions.py::test_the_two_dispositions_AGREE",
|
|
granularity=CLAUSE,
|
|
expect="this file says COVERED, the SPA guard says CREATE",
|
|
why="This guard IS a test, so disarming it makes it absent rather than red — the checker-guard "
|
|
"exception applies and the mutation goes into the guarded ARTIFACT, the SPA guard's disposition "
|
|
"table. The VALUE is the clause: flipping this row from ANNOTATED "
|
|
"to CREATE retires the requirement that MultiCollectionItemRequest be annotated, so deleting the "
|
|
"Complete<...> from MultiCollectionsScreen.toItemRequest then leaves every suite green with "
|
|
"#807's silent weight reset live again. The row is named in full rather than by the bare "
|
|
"disposition line, which occurs three times and would identify none of them.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/post-review-verdict.sh",
|
|
target=".gitea/workflows/review-verdict.yml",
|
|
clause='H10_REVIEWERS="timothy"',
|
|
replacement='H10_REVIEWERS="somebody-who-is-not-a-reviewer"',
|
|
proof="test_post_review_verdict.py::test_a_verdict_posted_by_an_ALLOWLISTED_account_is_accepted",
|
|
granularity=CLAUSE,
|
|
expect="the verdict writer REFUSED an allow-listed account",
|
|
why="THE COUPLING, NOT THE COMPARISON. The guard is the writer; the clause lives in the GATE, "
|
|
"because what ersatztv#845 is about is the two drifting apart. Mutating the gate's own "
|
|
"allow-list while the posting account stays fixed reddens the accept path only if BOTH hold: "
|
|
"the writer reads the list LIVE from the workflow (a hand-copied list would not move), and "
|
|
"the membership comparison actually gates the outcome (a no-op comparison would not care "
|
|
"that it moved). Disarming the comparison inside the script instead would prove the second "
|
|
"and say nothing about the first, which is the half that was missing.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_image_build_delegates_the_spa_suite.py",
|
|
target="docker/Dockerfile",
|
|
clause="RUN npm run lint && npm run typecheck && npm run build",
|
|
replacement="RUN npm run lint && npm run typecheck && npm test -- --run && npm run build",
|
|
proof="test_image_build_delegates_the_spa_suite.py::test_every_SPA_CARRYING_STAGE_runs_exactly_its_pinned_commands",
|
|
granularity=CLAUSE,
|
|
expect="docker/Dockerfile::web-build does not run its pinned commands",
|
|
why="The target is not the guard, per the checker-guard exception: the guard IS a test, so "
|
|
"disarming it makes it absent rather than red, and the mutation goes into the guarded "
|
|
"ARTIFACT. The clause is the web-build stage's whole command line, and the replacement is "
|
|
"the defect itself rather than a caricature of it — ersatztv#883 put a vitest run back into "
|
|
"a gitless stage, and every image build failed from that commit until #887. WHAT THIS PROVES "
|
|
"IS NARROWER THAN IT LOOKS, and saying so is the point: the guard no longer decides whether "
|
|
"a command RUNS the suite (that predicate was wrong nine times), "
|
|
"it compares the stage's commands against a pin. So this mutation proves the pin is "
|
|
"compared and reported — not that any particular spelling is recognised, because none needs "
|
|
"to be. The mutant is deliberately the UNFILTERED spelling: the filtered one is what broke, "
|
|
"but a pin rejects both identically and choosing the narrower case would suggest otherwise.",
|
|
),
|
|
)
|
|
|
|
|
|
# ------------------------------------------------------------------------------------------------
|
|
# THE OTHER GUARDS — stated per guard, and compared for SET EQUALITY against the inventory
|
|
# ------------------------------------------------------------------------------------------------
|
|
#
|
|
# #790's third Done-when box asks that guards whose mutation cannot be declared be STATED. A reason
|
|
# keyed on the row's GRADE would be cheaper and is tautological: a new guard graded NONE inherits one
|
|
# automatically and nobody ever looks at that particular guard. A count of them is no better — it
|
|
# moves only on net change, so adding one undeclared guard while promoting another leaves it at 22.
|
|
#
|
|
# So this is keyed on the guard, and `test_every_GUARD_row_is_either_DECLARED_or_STATED_here` asserts
|
|
# set equality against the inventory's GUARD rows in both directions. That makes it the same kind of
|
|
# hand-maintained-but-machine-checked table as `docs/guard-inventory.md` itself: a new guard cannot
|
|
# arrive without someone writing a line here about why it carries no mutation, and a line cannot
|
|
# outlive the row it is about.
|
|
#
|
|
# WHAT `NONE` ACTUALLY MEANS, because the wording matters here: the row nominates no proof ref. It
|
|
# does NOT mean the guard is untested. `scripts/ci-prove-ban-detects.sh` is graded NONE and is driven
|
|
# end to end by `test_ci_release_path_scan_job.py`. Nominating a proof is a judgement about which
|
|
# test is THE proof, which is #775's scope; this file can only verify one afterwards.
|
|
|
|
UNDECLARED: dict[str, str] = {
|
|
# NO GROUPING. Sorting these into "driven through their deciding path" and
|
|
# "not driven at all" was wrong twice — in both
|
|
# directions, over entries whose own text said the opposite. A category above a list is a second
|
|
# classification of the same facts, and it drifts the moment one entry's situation changes. Each
|
|
# entry states its own case instead.
|
|
#
|
|
# THE TWO THINGS THAT GO MISSING ARE DIFFERENT, and which one it is decides where the work goes.
|
|
# A guard may be DRIVEN — `test_hook_fire_log.py` executes most hooks through their real deciding
|
|
# branch, its matrix asserting instrumentation TRANSPARENCY (the wrapped and unwrapped runs
|
|
# agree), never that the decision is right or that a particular clause produced it — and still
|
|
# have no NOMINATED proof and no witnessed clause disarm. Nominating one is a judgement about
|
|
# which test is THE proof, which is #775's scope; this file can only verify one afterwards. A
|
|
# guard nothing executes at all needs the test first.
|
|
#
|
|
# `NONE` in the inventory means the row nominates no proof ref. It does NOT mean untested.
|
|
".claude/hooks/decisions-guard.sh": "Driven to a block and to a pass by test_hook_fire_log.py's "
|
|
"constructed cases, which assert transparency rather than the decision. No nominated proof, and "
|
|
"no clause disarmed.",
|
|
".claude/hooks/prepush-clean-worktree-check.sh": "Driven with a file both modified in the tree "
|
|
"and present in the pushed set, by test_hook_fire_log.py, for transparency. No nominated proof, "
|
|
"and no clause disarmed.",
|
|
".claude/hooks/prepush-donewhen.sh": "Driven against a stub Gitea by test_hook_fire_log.py, so "
|
|
"its real blocking path is reached — for transparency. No nominated proof, and no clause "
|
|
"disarmed.",
|
|
".claude/hooks/prepush-rebase-check.sh": "BEHAVIOUR-ONLY. A named test drives it and "
|
|
"test_hook_fire_log.py reaches its behind-origin block, but which clause carries that decision "
|
|
"has not been established by disarming one.",
|
|
".claude/hooks/pretooluse-agent-ram.sh": "Driven at 5% and 15% free memory through a stubbed "
|
|
"`memory_pressure`, by test_hook_fire_log.py, for transparency. No nominated proof, and no "
|
|
"clause disarmed.",
|
|
".claude/hooks/pretooluse-agent-model.sh": "Driven with and without a `model` in the payload by "
|
|
"test_hook_fire_log.py's matrix, for transparency. No nominated proof, and no clause disarmed.",
|
|
".claude/hooks/pretooluse-bash-guard.sh": "Driven with an ETV_UPDATE_GOLDENS command and a "
|
|
"harmless one by test_hook_fire_log.py's matrix, for transparency. No nominated proof, and no "
|
|
"clause disarmed.",
|
|
".claude/hooks/pretooluse-nav-guard.sh": "Driven with an `/iptv/` URL by test_hook_fire_log.py's "
|
|
"matrix, for transparency. No nominated proof, and no clause disarmed.",
|
|
".claude/hooks/design-sync-reminder.sh": "Driven by test_hook_fire_log.py, which gives it its "
|
|
"start/finish arguments and works around its self-throttle — for transparency, and never to the "
|
|
"one-shot branch that fires on the first Stop after a UI change and then allows. A proof has to "
|
|
"model that state transition rather than a single invocation.",
|
|
".claude/hooks/pretooluse-merge-consent.sh": "BEHAVIOUR-ONLY. Several suites execute it, but consent "
|
|
"is derived from several independent conditions, so which one a given red belongs to has to be "
|
|
"established before a clause can be named.",
|
|
".husky/commit-msg": "Outside the hook-fire population (that globs `.claude/hooks/*.sh`) and "
|
|
"executed by no test: the repositories the suite builds are fresh `git init`s that never install "
|
|
"husky, so the hook is absent rather than bypassed. A proof has to install or invoke it.",
|
|
".husky/pre-commit": "Runs lint-staged, the decisions guard, the root-PNG check and the format "
|
|
"gate. Only the decisions guard has an inventory row of its own — root-PNG and format are INLINE "
|
|
"here, so this one row is the whole classification of both, and neither has a proof. Executed by "
|
|
"no test, and nothing observes the dispatch itself.",
|
|
"scripts/check-kickoff-guard.sh": "Nothing drives it. A proof needs a tree carrying a revived "
|
|
"#237 reference, which is cheap and simply not written.",
|
|
"scripts/check-review-verdict.sh": "BEHAVIOUR-ONLY. Its named test feeds the real script an "
|
|
"input only one clause rejects, which proves it reacts, not that the clause is load-bearing.",
|
|
"scripts/ci-detect-already-validated.sh": "Blocks nothing directly — it feeds the skip gate. The "
|
|
"consequence a mutation would have to be observed through is a job that skips, which is visible "
|
|
"only in a workflow run.",
|
|
"scripts/ci-prove-ban-detects.sh": "Driven end to end by test_ci_release_path_scan_job.py, and "
|
|
"it runs a mutation of its own at CI time. Grading it here needs a decision about what a second "
|
|
"mutation would add; the row nominates no ref today.",
|
|
"scripts/e2e-functional.sh": "Needs a running instance. Its clauses are HTTP contract "
|
|
"assertions, so a proof means booting the app — `scripts/e2e-local.sh`'s job, not this "
|
|
"harness's.",
|
|
"scripts/jq-preflight.sh": "BEHAVIOUR-ONLY. Its named test drives the real script below the "
|
|
"version floor; no clause has been disarmed to show the floor comparison is what refuses.",
|
|
"scripts/pr-changed-files.sh": "BEHAVIOUR-ONLY. Its named test feeds a short page to the real "
|
|
"enumeration; the pagination clause has not been disarmed.",
|
|
"scripts/tests/test_ci_release_path_scan_job.py": "A GUARD that is a test, so the mutation would "
|
|
"have to go into the guarded artifact — the release-path scan job in the workflow. Which "
|
|
"weakening of that job is THE defect it exists to catch has not been settled.",
|
|
}
|