Round 4, and the third cold review found the same mechanism failing again, so it is
removed rather than patched a tenth time.
WHAT KEPT BREAKING. Three versions of this guard asked "does this command RUN the suite,
and can its failure be swallowed?" of arbitrary shell text. That predicate was wrong NINE
times across three review rounds, and twice a clause added to remove a FALSE RED opened a
FALSE GREEN on the guard's headline assertion:
* heredoc bodies were skipped as data, but BuildKit EXECUTES `RUN <<EOF` — and the
opener regex also fired inside quotes (`echo "tags<<__EOT__"`), which blinded the
whole-file scan over the last 303 lines of docker-build.yml. Wrong in both directions
at once, and measurably live on this tree.
* `shlex.shlex` does not clear `commenters` the way `shlex.split` does, so `#`
truncated a command mid-word — including the live `${#reports[@]}` idiom — and made
this file's own stated residual false.
* compound punctuation (`);`) welded two commands into one segment.
* `npm t`, `./node_modules/.bin/vitest`, `pnpm vitest`, `yarn vitest`,
`node …/vitest.mjs`, `timeout …`, `su -c …`, `if npm test; then` — all invisible.
* `true || npm test` counted as the gating run while never executing it.
* `continue-on-error: ${{ … }}` passed a check written against two literals — a
presence test that cannot see polarity, fail-OPEN in the one direction that matters.
WHAT REPLACES IT. Nothing in the file decides what a command means any more. The commands
that may run in the two risky places are PINNED as text: the `RUN` lines of every
SPA-carrying Dockerfile stage, and the gating step's `run:` body and `if:`. A suite run
re-added in ANY spelling is simply not equal to its pin — the pin does not have to
recognise a spelling in order to reject it. A pin cannot produce a false green, only a
false red, and a false red is a human reading a diff they should have read anyway.
The population/pin split is the load-bearing distinction, and it is now stated in the
inventory: a POPULATION decides what is CHECKED, so a hand-written one goes silently
short; a PIN decides what is EXPECTED, so a stale one goes loudly red. Only the second is
safe to write by hand. Populations stay derived from the git index.
Two premises that were prose are now assertions: the publish step keeps its own
`docs_only` gate (without it, a docs-only push skips the suite and publishes anyway), and
no step other than the pinned one mentions the suite — a SUBSTRING sweep, deliberately
not a predicate, whose failure mode is a false red asking someone to look.
41 mutants, 0 missed, including all nine spellings above and the three from the previous
round. Exactly ONE is declared in `mutation_manifest.py` and re-executed every suite; the
other 40 were witnessed during development and are NOT standing — stated in the row
rather than left to be assumed.
Also fixed: the truncated sentence the round-2 rewrite left in the Dockerfile comment,
and the `web/src/api/*.guard.test.ts` glob, which over-claimed — it matches three files
and only two of them need git.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
522 lines
35 KiB
Python
522 lines
35 KiB
Python
"""The DECLARED clause mutations, one per `MUTATION`-graded row of `docs/guard-inventory.md`.
|
|
|
|
Data only. The machinery that applies these is `mutation_harness_lib.py`; the checks that keep this
|
|
file honest are `test_mutation_harness.py`.
|
|
|
|
Every entry is declared by hand and none is inferred, which is the whole design constraint from
|
|
ersatztv#790: "a harness that guesses which clause of a 90-line hook is *the* guard would manufacture
|
|
exactly the confident-but-empty coverage this is meant to prevent". Where a proof test already names
|
|
its own clause in source — the BOM guard's `= "efbbbf" ]; then`, `UNSET_CLAUSE`, `prove-fix.sh`'s
|
|
`if [ "$RC" -eq 0 ]; then` — the entry reuses THAT string rather than inventing a second one, so a
|
|
retarget in either place is caught by the other.
|
|
|
|
WHY AN ENTRY'S `target` MAY DIFFER FROM ITS `guard`. Some guards here ARE tests
|
|
(`scripts/tests/test_*.py`). Disarming such a guard makes it ABSENT rather than red, so
|
|
`testing.guard-ships-with-mutation-proof`'s checker-guard exception applies: the mutation goes into
|
|
the guarded ARTIFACT — a deleted row, a planted phantom row — and the check must report it. Mutating
|
|
a checker's own POPULATION instead is a trap that looks identical and is not: a shrunken population
|
|
makes every real row report as PHANTOM, so the proof reddens on a false positive while saying
|
|
nothing about the missing-row detection the row claims. `why` states per entry which shape applies
|
|
and why; no count is kept here, because a count of the entries below is a second copy of them.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from scripts.tests.mutation_harness_lib import Mutation
|
|
|
|
CLAUSE = Mutation.CLAUSE
|
|
DETECTOR = Mutation.DETECTOR
|
|
|
|
MUTATIONS: tuple[Mutation, ...] = (
|
|
Mutation(
|
|
guard=".claude/hooks/posttooluse-worktree-marker.sh",
|
|
target=".claude/hooks/posttooluse-worktree-marker.sh",
|
|
clause='printf \'%s\\n\' "$me" > "$abs/.claude-worktree-owner" 2>/dev/null || true',
|
|
replacement="true",
|
|
proof="test_worktree_ownership_guard.py::test_MUTATION_a_marker_hook_that_stops_WRITING_makes_the_guard_go_quiet",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED pair did not deny",
|
|
why="The marker write is the hook's entire job; without it the guard has nothing to read and "
|
|
"fails open. The clause string is the one the proof test itself passes to its `_mutate` helper.",
|
|
),
|
|
Mutation(
|
|
guard=".claude/hooks/pretooluse-bom-guard.sh",
|
|
target=".claude/hooks/pretooluse-bom-guard.sh",
|
|
clause='= "efbbbf" ]; then',
|
|
replacement='= "deadbeef" ]; then',
|
|
proof="test_bom_guard_detection.py::test_DISARMING_the_BOM_comparison_stops_detection",
|
|
granularity=CLAUSE,
|
|
expect="the BOM comparison has moved",
|
|
why="The BOM comparison is the guard's only detection logic. Same clause the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard=".claude/hooks/pretooluse-worktree-guard.sh",
|
|
target=".claude/hooks/pretooluse-worktree-guard.sh",
|
|
clause='marker="$root/.claude-worktree-owner"',
|
|
replacement='marker="$root/.claude-worktree-owner-NOTHING-WRITES-THIS"',
|
|
proof="test_worktree_ownership_guard.py::test_MUTATION_disarming_the_guards_MARKER_READ_stops_the_deny",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED guard did not deny",
|
|
why="The marker read is what the ownership decision hangs on. Same clause the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard=".husky/pre-push",
|
|
target=".husky/pre-push",
|
|
clause="unset GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE",
|
|
replacement=": # clause removed by the mutation harness",
|
|
proof="test_prepush_unsets_git_env.py::test_MUTATION_DELETING_the_unset_lets_drift_through_silently",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED pre-push did not catch the drift",
|
|
why="Without the unset, every git call the pre-push chain makes is aimed at the repository git "
|
|
"exported the environment for, not the one being pushed. `UNSET_CLAUSE` in the proof test.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/build_decisions_catalog.py",
|
|
target="scripts/build_decisions_catalog.py",
|
|
clause="want.strip() != have.strip()",
|
|
replacement="False",
|
|
proof="test_build_catalog_check_path.py::test_MUTATION_disarming_the_stale_comparison_stops_detection",
|
|
granularity=CLAUSE,
|
|
expect="the stale-detection clause has moved or been reworded",
|
|
why="`main()`'s only stale-detection logic, per the proof test's own docstring, which uses this "
|
|
"exact clause and this exact replacement.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-detect-docs-only.sh",
|
|
target="scripts/ci-detect-docs-only.sh",
|
|
clause='if [ "$is_shallow" = "true" ]; then',
|
|
replacement="if true; then",
|
|
proof="test_docs_only_detector_clone_depth.py::test_the_push_arm_leaves_a_COMPLETE_clone_complete",
|
|
granularity=CLAUSE,
|
|
expect="GRAFTED the complete clone shallow",
|
|
why="The shallow test is the whole of ersatztv#836's fix: disarmed, `--depth=2` goes back to "
|
|
"every checkout including `build`'s complete one, which grafts it and makes the `git describe` "
|
|
"in the next step find no reachable tag. This row was UNDECLARED until #836 on the stated "
|
|
"grounds that the script 'feeds the skip gate, so its effect is visible only in a workflow "
|
|
"run' — untrue of this clause, whose effect is the shallow flag on a real clone and is "
|
|
"observable in-process, which is what the proof test asserts.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-step-ran.sh",
|
|
target="scripts/ci-step-ran.sh",
|
|
clause='if ! grep -qxF "$key" "$marker" 2>/dev/null; then',
|
|
replacement="if false; then",
|
|
proof="test_ci_dropped_step_guard.py::test_dropping_ANY_single_step_FAILS_the_guard",
|
|
granularity=CLAUSE,
|
|
expect="never having executed",
|
|
why="The per-key membership test is what turns a dropped step into a red job; disarmed, every "
|
|
"expected key reads as present and the guard passes a run in which nothing executed.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_ci_status_context_uniqueness.py",
|
|
target=".gitea/workflows/docker-build.yml",
|
|
clause=" name: API docs in sync (OpenAPI + endpoint index)",
|
|
replacement=" name: Build & test (.NET)",
|
|
proof="test_ci_status_context_uniqueness.py::test_no_two_jobs_synthesize_the_SAME_status_context",
|
|
granularity=CLAUSE,
|
|
expect="these status contexts can be produced by more than one job",
|
|
why="THE CHECKER-GUARD SHAPE: this guard IS a test, so disarming its own assertion makes it "
|
|
"absent rather than red. Per `testing.guard-ships-with-mutation-proof` the mutation goes into "
|
|
"the guarded ARTIFACT instead — here a workflow job renamed to collide with the REQUIRED "
|
|
"`Build & test (.NET)` job, which is exactly the evasion the guard exists to catch: two jobs "
|
|
"synthesizing one status context that branch protection cannot tell apart, so the required "
|
|
"check could be satisfied by the producer whose steps carry no execution markers.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/check-required-contexts.sh",
|
|
target="scripts/check-required-contexts.sh",
|
|
clause="($rule.status_check_contexts | sort | unique) == $snap",
|
|
replacement="true",
|
|
proof="test_check_required_contexts.py::test_MUTATION_disarming_the_set_comparison_stops_every_drift_report",
|
|
expect="the set-comparison clause has moved or been reworded",
|
|
granularity=CLAUSE,
|
|
why="The set comparison IS the finding: it is the only thing that turns a live required-check "
|
|
"list differing from the committed snapshot into `drift`. Witnessed: unmutated reports "
|
|
"`drift` on an added context, the mutant reports `match` — a permanent no-op that would "
|
|
"confirm the snapshot fresh forever. The proof asserts the mutant's EXACT verdict "
|
|
"`(0, 'match')` rather than merely 'not drift', and ships a positive control for its own "
|
|
"tmp layout: an earlier draft copied the script without the classifier it loads beside "
|
|
"itself, so the mutant exited 2 with empty stdout and 'not drift' was satisfied by a copy "
|
|
"that never ran. The clause string is the one the proof test asserts on before mutating.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/decisions_validate.py",
|
|
target="scripts/decisions_validate.py",
|
|
clause="wing_faults=record_wing_faults() + yaml_faults,",
|
|
replacement="wing_faults=yaml_faults,",
|
|
proof="test_decisions_validate.py::test_main_actually_CALLS_the_wing_scan",
|
|
granularity=CLAUSE,
|
|
expect="a wing fault must fail the validator",
|
|
why="The wiring the proof test exists for: deleting this call left the whole suite green while "
|
|
"a real block-scalar record vanished under `decisions-validate: OK` (#609).",
|
|
),
|
|
Mutation(
|
|
guard="scripts/prove-fix.sh",
|
|
target="scripts/prove-fix.sh",
|
|
clause='if [ "$RC" -eq 0 ]; then',
|
|
replacement="if false; then",
|
|
proof="test_prove_fix.py::test_MUTATION_disarming_the_UNPROVEN_clause_reddens_the_refusal_test",
|
|
granularity=CLAUSE,
|
|
expect="the clause under mutation is gone",
|
|
why="The UNPROVEN branch: a named test that passes WITHOUT the fix must be refused. Same clause "
|
|
"the proof test names.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_ci_image_pin_population.py",
|
|
target="scripts/tests/test_ci_image_pin_population.py",
|
|
clause="for name in sorted(expected - set(pinned)):",
|
|
replacement="for name in sorted(set()):",
|
|
proof="test_ci_image_pin_population.py::test_a_single_job_losing_its_pin_is_DETECTED",
|
|
granularity=CLAUSE,
|
|
expect="the population check accepted a workflow in which a container job no longer runs",
|
|
why="The against-the-DECLARED-CLASS direction, and the one the other two clauses cannot "
|
|
"cover: a job that loses its `container:` block leaves `declared` and `pinned` equal, so "
|
|
"only this comparison notices it has moved to the bare runner. This is the clause #790 "
|
|
"asked for instead of neutering `pin_population_faults` wholesale. `expected` is "
|
|
"`toolchain_declared(doc)` since ersatztv#789 replaced the `TOOLCHAIN_JOBS` literal this "
|
|
"clause used to name; the comparison and its role are unchanged.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_guard_populations_derive_from_git.py",
|
|
target="scripts/tests/test_guard_inventory.py",
|
|
clause="if ref in tracked:",
|
|
replacement="if (REPO_ROOT / ref).exists():",
|
|
proof="test_guard_populations_derive_from_git.py::test_no_derivation_admits_an_untracked_file",
|
|
granularity=CLAUSE,
|
|
expect="after git stopped tracking them",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — one of the "
|
|
"derivations it watches — rather than into the checker, per the checker-guard exception. "
|
|
"The clause is the exact defect this guard was written after: `derived_guard_files` read "
|
|
"its CALLERS from the index and then admitted the paths they name on `Path.exists()`, so a "
|
|
"tracked workflow naming a script that exists on one machine only entered the population "
|
|
"there, red on that checkout and green in CI (#778's third shape, found by cold review "
|
|
"inside #806 itself). Note what this mutation does NOT do: on a clean tree the mutated set "
|
|
"is identical, so `test_guard_inventory.py`'s own assertions stay green — only narrowing "
|
|
"the index, which is what the proof does, separates them. That is why the proof has to "
|
|
"remove EVERY member rather than sample one.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_optional_request_members.py",
|
|
target="scripts/tests/test_optional_request_members.py",
|
|
clause='"ArtworkContentTypeModel": (',
|
|
replacement='"ArtworkContentTypeModelRENAMED": (',
|
|
proof="test_optional_request_members.py::test_every_droppable_request_schema_has_a_stated_disposition",
|
|
granularity=CLAUSE,
|
|
expect="no disposition written down",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — here the "
|
|
"DISPOSITIONS registry the checker maintains, the same shape as the deleted "
|
|
"guard-inventory row below. Renaming the key rather than deleting the entry keeps the "
|
|
"module importable, so the red is a real set-equality failure and not an ImportError "
|
|
"reddening for the wrong reason. The rename fires BOTH directions — MISSING for the real "
|
|
"schema and PHANTOM for the renamed key — which is the correct behaviour and worth stating, "
|
|
"since `expect` names only the MISSING half. "
|
|
"`ArtworkContentTypeModel` is the right key to name: it is the exact schema #807's "
|
|
"hand-written table omitted, because `...Model` reads as a response model while it is in "
|
|
"fact reachable from the full-replace PUT /channels/{id}.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_workflow_job_guards.py",
|
|
target="docs/guard-inventory.md",
|
|
clause="| `pr-checks.yml::ci-image-pin` |",
|
|
replacement="| `pr-checks.yml::ci-image-pin-RENAMED` |",
|
|
proof="test_workflow_job_guards.py::test_the_inventory_covers_exactly_the_guard_JOBS_that_exist",
|
|
granularity=CLAUSE,
|
|
expect="guard/report-only but have NO row in",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — the workflow-job "
|
|
"table — rather than into the checker, per the checker-guard exception. RENAMED rather than "
|
|
"deleted, so the row count is unchanged and the red cannot come from an empty table: the "
|
|
"anti-vacuity test still sees rows, and what fails is the set equality itself, in BOTH "
|
|
"directions at once (MISSING for the real job, PHANTOM for the renamed key). "
|
|
"`ci-image-pin` is the right row to name — it is the guard job #774 found stating an "
|
|
"invariant it did not check, and the case #786 was filed to bring into a population.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_guard_inventory.py",
|
|
target="docs/guard-inventory.md",
|
|
clause="| `.claude/hooks/decisions-guard.sh` | a commit | GUARD | NONE | — |\n",
|
|
replacement="",
|
|
proof="test_guard_inventory.py::test_the_inventory_covers_exactly_the_guards_that_exist",
|
|
granularity=CLAUSE,
|
|
expect="these guard files exist but have no row in guard-inventory.md",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT rather than into the "
|
|
"checker — disarming a checker makes it absent, not red, and mutating its population instead "
|
|
"would only demonstrate a false POSITIVE (a shrunken population reports every real row as "
|
|
"phantom) while proving nothing about the missing-row detection the row claims. A deleted "
|
|
"row is the defect this guard exists to catch, and it is one of the mutations #774 witnessed "
|
|
"by hand.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_hook_fire_log.py",
|
|
target="scripts/tests/test_hook_fire_log.py",
|
|
clause=" return faults\n\n\ndef strip_instrumentation",
|
|
replacement=" return []\n\n\ndef strip_instrumentation",
|
|
proof="test_hook_fire_log.py::test_a_hook_that_LOSES_its_instrumentation_is_DETECTED",
|
|
granularity=DETECTOR,
|
|
expect="left the check GREEN. The check is not load-bearing",
|
|
why="NO CLAUSE-LEVEL MUTATION REDDENS THIS ONE, and that is a finding rather than a shortcut. "
|
|
"`instrumentation_faults` accumulates from four independent arms and a stripped hook trips "
|
|
"three of them at once (no sink source, no ETV_HOOK_FIRE_LIB assignment, no begin call), so "
|
|
"disarming any single arm leaves the other two answering and the proof test stays green. The "
|
|
"whole detector is therefore the smallest mutation this proof can witness — and the surviving "
|
|
"single-arm mutation below is re-run every time so that claim is checked, not recited.",
|
|
survived_clause=" if not _SOURCES_SINK.search(text):",
|
|
survived_replacement=" if False:",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_remote_state_inventory.py",
|
|
target="docs/remote-state-inventory.md",
|
|
clause="| `scripts/post-review-verdict.sh` — commit-status write |",
|
|
replacement="| `scripts/DELETED-BY-THE-MUTATION-HARNESS.sh` — not a real path |",
|
|
proof="test_remote_state_inventory.py::test_every_in_scope_file_has_a_row_and_every_row_names_a_real_file",
|
|
granularity=CLAUSE,
|
|
expect="in scope but absent from docs/remote-state-inventory.md",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT: a real executable's "
|
|
"row is renamed away, which is the MISSING-row defect the row's Blocks column claims — an "
|
|
"in-scope file with no classification. Two shapes were tried and rejected. Emptying the "
|
|
"guard's `git ls-files` derivation reddens the proof with an IndexError over an empty "
|
|
"population: a crash, not a detection. Planting a PHANTOM row reddens "
|
|
"`test_MUTATION_PROOF_a_dropped_row_and_a_phantom_row_are_both_detected` by contaminating "
|
|
"the fixture that test builds for itself, and proves the opposite direction from the one the "
|
|
"row claims. Renaming the row exercises both directions of the production set comparison at "
|
|
"once and is matched on the missing half.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_mutation_harness.py",
|
|
target="scripts/tests/mutation_harness_lib.py",
|
|
clause=" if mutation.expect not in diagnostic:",
|
|
replacement=" if False:",
|
|
proof="test_mutation_harness.py::test_MUTATION_disarming_the_DIAGNOSTIC_gate_accepts_a_red_for_the_wrong_reason",
|
|
granularity=CLAUSE,
|
|
expect="the UNMUTATED verdict already accepted it, so the mutant proves nothing",
|
|
why="The target is not the guard for a structural reason: the harness keeps its machinery in "
|
|
"`mutation_harness_lib.py` so a clause of it can be disarmed in an isolated copy at all. The "
|
|
"clause is the DIAGNOSTIC gate — the check that a failing proof failed with the diagnostic "
|
|
"its row declares. Disarmed, a red for any unrelated reason is certified as a guard doing "
|
|
"its job, which is the shape that made two rows in this very file measure nothing. The other "
|
|
"gate, the one requiring pytest exit code 1, carries its own proof in "
|
|
"`test_MUTATION_disarming_the_EXIT_STATUS_gate_accepts_a_run_that_NEVER_RAN_A_TEST`; the "
|
|
"inventory holds one ref per row, so this entry names the stronger of the two.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/ci-toolchain-image-resolves.sh",
|
|
target="scripts/ci-toolchain-image-resolves.sh",
|
|
clause=" 404)",
|
|
replacement=" 4040)",
|
|
proof="test_ci_toolchain_image_resolves.py::test_MUTATION_a_deleted_tag_is_reported_as_a_failure",
|
|
granularity=CLAUSE,
|
|
expect="a deleted tag was not reported as GONE",
|
|
why="404 is the ONE answer that establishes the pinned toolchain image is gone; every other "
|
|
"code means the check could not run. Both fail the job, so the EXIT CODE does not separate "
|
|
"them and the mutation is caught by the DIAGNOSTIC instead: retargeting the arm sends the "
|
|
"real outage down the could-not-verify path, which sends an operator to the registry's "
|
|
"health rather than to the rebuild that fixes it.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_review_verdict_vocabulary.py",
|
|
target="scripts/check-review-verdict.sh",
|
|
clause="""POS_RE='^review-verdict:[[:space:]]*('"$POS_ALTERNATION"')([[:space:]@]|$)'""",
|
|
replacement="""POS_RE='^review-verdict:[[:space:]]*(mergeable|approved|lgtm)([[:space:]@]|$)'""",
|
|
proof="test_review_verdict_vocabulary.py::test_a_word_added_to_the_shared_source_reaches_BOTH_sides",
|
|
granularity=CLAUSE,
|
|
expect="restating the word list rather than deriving it",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT — the read side "
|
|
"whose derivation it watches — rather than into the checker, per the checker-guard "
|
|
"exception. The replacement is not an invented mutant: it is the LITERAL pre-#788 line, the "
|
|
"second hand-written copy this change removed, so the proof is taken against the real "
|
|
"predecessor. Mutating the shared vocabulary instead would be the trap that looks "
|
|
"identical: emptying or corrupting the word list reddens the proof through the fail-closed "
|
|
"validator, which says nothing about whether the read side still DERIVES from it.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_workflow_persist_credentials.py",
|
|
target=".gitea/workflows/ci-image.yml",
|
|
clause=""" persist-credentials: false
|
|
# only docker/ci/Dockerfile is needed; no git describe/log here""",
|
|
replacement=""" # only docker/ci/Dockerfile is needed; no git describe/log here""",
|
|
proof="test_workflow_persist_credentials.py::test_every_actions_checkout_DROPS_the_persisted_credential",
|
|
granularity=CLAUSE,
|
|
expect="persists a write-capable Authorization header into .git/config",
|
|
why="THE GUARD IS A TEST, so the mutation goes into the guarded ARTIFACT rather than into "
|
|
"the checker, per the checker-guard exception. The defect is the one the guard exists for: "
|
|
"an `actions/checkout` step that simply omits the key, which is what a newly added job gets "
|
|
"by default — not an explicit `true`, which nobody writes. `ci-image.yml` carries exactly "
|
|
"one checkout, so the clause is unambiguous there; the trailing comment line is part of it "
|
|
"only to make the match unique within the file.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_mcp_smoke.py",
|
|
target="scripts/mcp_smoke.py",
|
|
clause=" id_init = secrets.randbelow(2**31 - 1000) + 1000",
|
|
replacement=" id_init = 1",
|
|
proof="test_mcp_smoke.py::test_MUTATION_a_PRE_ANSWERED_id_is_refused_because_the_request_ids_are_UNGUESSABLE",
|
|
granularity=CLAUSE,
|
|
expect="a pre-answered id was ACCEPTED",
|
|
why="The target is not the guard for the same structural reason as the harness entry above: "
|
|
"`scripts/mcp_smoke.py` is a checker reached only transitively, through "
|
|
"`scripts/check-local-lsp.sh`, so it holds no inventory row of its own and the clause has to "
|
|
"be disarmed in it directly (`testing.verification-code-needs-its-own-proof`). The clause is "
|
|
"the unguessable request id, which AT THE `initialize` STAGE is the only thing refusing a "
|
|
"server that answers before it is asked: the pending-registration cannot help there, because "
|
|
"that id is already in flight when the pre-answer arrives, which is why #793 round 5 replaced "
|
|
"the lock rather than tightening it. Disarmed, the stub's pre-answer is "
|
|
"accepted at `initialize` and the run dies one stage later at `tools/list`, so the proof "
|
|
"asserts the STAGE (rc 9 and the initialize diagnostic) rather than mere failure: the mutant "
|
|
"still fails, and a test checking only that would stay green over a real vulnerability. THE "
|
|
"`id_tools` TWIN IS DELIBERATELY NOT THIS CLAUSE, measured rather than assumed: replacing "
|
|
"`id_tools` alone leaves every case green, because the reader keeps only the reply matching "
|
|
"the id in flight, so a `tools/list` frame emitted while `initialize` is pending is dropped "
|
|
"whatever its id. Disarming that retention clause alone is green too. Only BOTH together "
|
|
"produce the false green (rc 0), so the tools stage is held by two mechanisms that mask each "
|
|
"other and neither is singly detectable — #685's shape. `id_init` is the one place a "
|
|
"pre-answer IS singly exploitable, which is why it is the declared clause; the tools stage is "
|
|
"covered behaviourally in the same file.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_complete_annotation_dispositions.py",
|
|
target="web/src/api/completeAnnotations.guard.test.ts",
|
|
clause="MultiCollectionItemRequest: {\n disposition: 'ANNOTATED',",
|
|
replacement="MultiCollectionItemRequest: {\n disposition: 'CREATE',",
|
|
proof="test_complete_annotation_dispositions.py::test_the_two_dispositions_AGREE",
|
|
granularity=CLAUSE,
|
|
expect="this file says COVERED, the SPA guard says CREATE",
|
|
why="This guard IS a test, so disarming it makes it absent rather than red — the checker-guard "
|
|
"exception applies and the mutation goes into the guarded ARTIFACT, the SPA guard's disposition "
|
|
"table. The VALUE is the clause: cold review demonstrated that flipping this row from ANNOTATED "
|
|
"to CREATE retires the requirement that MultiCollectionItemRequest be annotated, so deleting the "
|
|
"Complete<...> from MultiCollectionsScreen.toItemRequest then leaves every suite green with "
|
|
"#807's silent weight reset live again. The row is named in full rather than by the bare "
|
|
"disposition line, which occurs three times and would identify none of them.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/post-review-verdict.sh",
|
|
target=".gitea/workflows/review-verdict.yml",
|
|
clause='H10_REVIEWERS="timothy"',
|
|
replacement='H10_REVIEWERS="somebody-who-is-not-a-reviewer"',
|
|
proof="test_post_review_verdict.py::test_a_verdict_posted_by_an_ALLOWLISTED_account_is_accepted",
|
|
granularity=CLAUSE,
|
|
expect="the verdict writer REFUSED an allow-listed account",
|
|
why="THE COUPLING, NOT THE COMPARISON. The guard is the writer; the clause lives in the GATE, "
|
|
"because what ersatztv#845 is about is the two drifting apart. Mutating the gate's own "
|
|
"allow-list while the posting account stays fixed reddens the accept path only if BOTH hold: "
|
|
"the writer reads the list LIVE from the workflow (a hand-copied list would not move), and "
|
|
"the membership comparison actually gates the outcome (a no-op comparison would not care "
|
|
"that it moved). Disarming the comparison inside the script instead would prove the second "
|
|
"and say nothing about the first, which is the half that was missing.",
|
|
),
|
|
Mutation(
|
|
guard="scripts/tests/test_image_build_delegates_the_spa_suite.py",
|
|
target="docker/Dockerfile",
|
|
clause="RUN npm run lint && npm run typecheck && npm run build",
|
|
replacement="RUN npm run lint && npm run typecheck && npm test -- --run && npm run build",
|
|
proof="test_image_build_delegates_the_spa_suite.py::test_every_SPA_CARRYING_STAGE_runs_exactly_its_pinned_commands",
|
|
granularity=CLAUSE,
|
|
expect="docker/Dockerfile::web-build does not run its pinned commands",
|
|
why="The target is not the guard, per the checker-guard exception: the guard IS a test, so "
|
|
"disarming it makes it absent rather than red, and the mutation goes into the guarded "
|
|
"ARTIFACT. The clause is the web-build stage's whole command line, and the replacement is "
|
|
"the defect itself rather than a caricature of it — ersatztv#883 put a vitest run back into "
|
|
"a gitless stage, and every image build failed from that commit until #887. WHAT THIS PROVES "
|
|
"IS NARROWER THAN IT LOOKS, and saying so is the point: the guard no longer decides whether "
|
|
"a command RUNS the suite (that predicate was wrong nine times across three review rounds), "
|
|
"it compares the stage's commands against a pin. So this mutation proves the pin is "
|
|
"compared and reported — not that any particular spelling is recognised, because none needs "
|
|
"to be. The mutant is deliberately the UNFILTERED spelling: the filtered one is what broke, "
|
|
"but a pin rejects both identically and choosing the narrower case would suggest otherwise.",
|
|
),
|
|
)
|
|
|
|
|
|
# ------------------------------------------------------------------------------------------------
|
|
# THE OTHER GUARDS — stated per guard, and compared for SET EQUALITY against the inventory
|
|
# ------------------------------------------------------------------------------------------------
|
|
#
|
|
# #790's third Done-when box asks that guards whose mutation cannot be declared be STATED. A reason
|
|
# keyed on the row's GRADE would be cheaper and is tautological: a new guard graded NONE inherits one
|
|
# automatically and nobody ever looks at that particular guard. A count of them is no better — it
|
|
# moves only on net change, so adding one undeclared guard while promoting another leaves it at 22.
|
|
#
|
|
# So this is keyed on the guard, and `test_every_GUARD_row_is_either_DECLARED_or_STATED_here` asserts
|
|
# set equality against the inventory's GUARD rows in both directions. That makes it the same kind of
|
|
# hand-maintained-but-machine-checked table as `docs/guard-inventory.md` itself: a new guard cannot
|
|
# arrive without someone writing a line here about why it carries no mutation, and a line cannot
|
|
# outlive the row it is about.
|
|
#
|
|
# WHAT `NONE` ACTUALLY MEANS, because the wording matters here: the row nominates no proof ref. It
|
|
# does NOT mean the guard is untested. `scripts/ci-prove-ban-detects.sh` is graded NONE and is driven
|
|
# end to end by `test_ci_release_path_scan_job.py`. Nominating a proof is a judgement about which
|
|
# test is THE proof, which is #775's scope; this file can only verify one afterwards.
|
|
|
|
UNDECLARED: dict[str, str] = {
|
|
# NO GROUPING. An earlier version sorted these into "driven through their deciding path" and
|
|
# "not driven at all", and the sort was wrong twice in successive review rounds — in both
|
|
# directions, over entries whose own text said the opposite. A category above a list is a second
|
|
# classification of the same facts, and it drifts the moment one entry's situation changes. Each
|
|
# entry states its own case instead.
|
|
#
|
|
# THE TWO THINGS THAT GO MISSING ARE DIFFERENT, and which one it is decides where the work goes.
|
|
# A guard may be DRIVEN — `test_hook_fire_log.py` executes most hooks through their real deciding
|
|
# branch, its matrix asserting instrumentation TRANSPARENCY (the wrapped and unwrapped runs
|
|
# agree), never that the decision is right or that a particular clause produced it — and still
|
|
# have no NOMINATED proof and no witnessed clause disarm. Nominating one is a judgement about
|
|
# which test is THE proof, which is #775's scope; this file can only verify one afterwards. A
|
|
# guard nothing executes at all needs the test first.
|
|
#
|
|
# `NONE` in the inventory means the row nominates no proof ref. It does NOT mean untested.
|
|
".claude/hooks/decisions-guard.sh": "Driven to a block and to a pass by test_hook_fire_log.py's "
|
|
"constructed cases, which assert transparency rather than the decision. No nominated proof, and "
|
|
"no clause disarmed.",
|
|
".claude/hooks/prepush-clean-worktree-check.sh": "Driven with a file both modified in the tree "
|
|
"and present in the pushed set, by test_hook_fire_log.py, for transparency. No nominated proof, "
|
|
"and no clause disarmed.",
|
|
".claude/hooks/prepush-donewhen.sh": "Driven against a stub Gitea by test_hook_fire_log.py, so "
|
|
"its real blocking path is reached — for transparency. No nominated proof, and no clause "
|
|
"disarmed.",
|
|
".claude/hooks/prepush-rebase-check.sh": "BEHAVIOUR-ONLY. A named test drives it and "
|
|
"test_hook_fire_log.py reaches its behind-origin block, but which clause carries that decision "
|
|
"has not been established by disarming one.",
|
|
".claude/hooks/pretooluse-agent-ram.sh": "Driven at 5% and 15% free memory through a stubbed "
|
|
"`memory_pressure`, by test_hook_fire_log.py, for transparency. No nominated proof, and no "
|
|
"clause disarmed.",
|
|
".claude/hooks/pretooluse-agent-model.sh": "Driven with and without a `model` in the payload by "
|
|
"test_hook_fire_log.py's matrix, for transparency. No nominated proof, and no clause disarmed.",
|
|
".claude/hooks/pretooluse-bash-guard.sh": "Driven with an ETV_UPDATE_GOLDENS command and a "
|
|
"harmless one by test_hook_fire_log.py's matrix, for transparency. No nominated proof, and no "
|
|
"clause disarmed.",
|
|
".claude/hooks/pretooluse-nav-guard.sh": "Driven with an `/iptv/` URL by test_hook_fire_log.py's "
|
|
"matrix, for transparency. No nominated proof, and no clause disarmed.",
|
|
".claude/hooks/design-sync-reminder.sh": "Driven by test_hook_fire_log.py, which gives it its "
|
|
"start/finish arguments and works around its self-throttle — for transparency, and never to the "
|
|
"one-shot branch that fires on the first Stop after a UI change and then allows. A proof has to "
|
|
"model that state transition rather than a single invocation.",
|
|
".claude/hooks/pretooluse-merge-consent.sh": "BEHAVIOUR-ONLY. Several suites execute it, but consent "
|
|
"is derived from several independent conditions, so which one a given red belongs to has to be "
|
|
"established before a clause can be named.",
|
|
".husky/commit-msg": "Outside the hook-fire population (that globs `.claude/hooks/*.sh`) and "
|
|
"executed by no test: the repositories the suite builds are fresh `git init`s that never install "
|
|
"husky, so the hook is absent rather than bypassed. A proof has to install or invoke it.",
|
|
".husky/pre-commit": "Runs lint-staged, the decisions guard, the root-PNG check and the format "
|
|
"gate. Only the decisions guard has an inventory row of its own — root-PNG and format are INLINE "
|
|
"here, so this one row is the whole classification of both, and neither has a proof. Executed by "
|
|
"no test, and nothing observes the dispatch itself.",
|
|
"scripts/check-kickoff-guard.sh": "Nothing drives it. A proof needs a tree carrying a revived "
|
|
"#237 reference, which is cheap and simply not written.",
|
|
"scripts/check-review-verdict.sh": "BEHAVIOUR-ONLY. Its named test feeds the real script an "
|
|
"input only one clause rejects, which proves it reacts, not that the clause is load-bearing.",
|
|
"scripts/ci-detect-already-validated.sh": "Blocks nothing directly — it feeds the skip gate. The "
|
|
"consequence a mutation would have to be observed through is a job that skips, which is visible "
|
|
"only in a workflow run.",
|
|
"scripts/ci-prove-ban-detects.sh": "Driven end to end by test_ci_release_path_scan_job.py, and "
|
|
"it runs a mutation of its own at CI time. Grading it here needs a decision about what a second "
|
|
"mutation would add; the row nominates no ref today.",
|
|
"scripts/e2e-functional.sh": "Needs a running instance. Its clauses are HTTP contract "
|
|
"assertions, so a proof means booting the app — `scripts/e2e-local.sh`'s job, not this "
|
|
"harness's.",
|
|
"scripts/jq-preflight.sh": "BEHAVIOUR-ONLY. Its named test drives the real script below the "
|
|
"version floor; no clause has been disarmed to show the floor comparison is what refuses.",
|
|
"scripts/pr-changed-files.sh": "BEHAVIOUR-ONLY. Its named test feeds a short page to the real "
|
|
"enumeration; the pagination clause has not been disarmed.",
|
|
"scripts/tests/test_ci_release_path_scan_job.py": "A GUARD that is a test, so the mutation would "
|
|
"have to go into the guarded artifact — the release-path scan job in the workflow. Which "
|
|
"weakening of that job is THE defect it exists to catch has not been settled.",
|
|
}
|