A verification harness is itself unproven code — five false greens, each created by the fix for the last #796
Closed
opened 2026-08-14 17:07:27 +02:00 by timothy
·
4 comments
No Branch/Tag Specified
main
renovate/meziantou.analyzer-3.x
release/v26.15.0-notes
fix/830-add-items-error-surface
renovate/lucene.net
renovate/cliwrap-3.x
issue-806-guard-populations
renovate/dotnet-monorepo
scratch/767b-poisoned
scratch/767b-control
release/v26.14.0-notes
release/v26.14.0
renovate/sqlitepclraw.bundle_e_sqlite3-3.x
docs/510-skill-logo-bug-policy
fix/510-watermark-resolution-policy
fix/629-verdict-classifier-falseopens
fix/609-decisions-edit-token-scope
issue-135-clear-to-none
release/v26.12.0-notes
fix/409b-lastscan-api-parity
fix/401-updatechannel-mirror-422
fix/327-playlist-rename-validation
fix/410-scancancel-log-level
fix/409-447-librariesscreen-neverscanned
fix/338-zap-exit-code
fix/367-plex-budget-message
fix/310-debom-legacy-cs
ci/604-lane-rebalance
feat/388-design-mirror
feat/247-test-ownership
feat/247-primary-action
feat/357-player-owned-playback
feat/357-jellyfin-plugin-poc
fix/289-mcp-hardening
issue58-mcp
feat/244-channels-extract
ci/auto-bump-prod-compose
feat/multi-rerun-collections-api
feat/collections-api
feat/quick-wins
feat/185-docs-part2
feat/140-collections-screen
feat/146-channel-edit
feat/147-classic-ui-link
issue22-renovate-dashboard
feat/91-cutover
feat/63-composite-create
feat/65-library-browse
feat/85-epg
feat/86-schedule-editor
feat/109-dashboard-data
feat/99-session-tracking
fix/dockerfile-node-tag
feat/59-spa-foundation
docs/59-ui-redesign-brief
feat/102-json-guide
feat/111-schedule-durations
feat/104-artwork-upload
feat/103-media-sources-api
feat/playouts-read-api
feat/108-health-api
feat/105-picker-list-endpoints
issue-97-channel-state-api
issue42-jellyfin-musicvideos
issue46-rest-api-error-contract
dependabot/nuget/ErsatzTV.FFmpeg.Tests/multi-d307a2e06f
qsv-improvements
hdr-vulkan-cuda-test
v26.15.0
v26.14.0
v26.13.0
v26.12.0
v26.11.0
v26.10.0
v26.9.0
v26.8.0
v26.7.0
blazor-final
v26.6.0
v26.5.0
v26.4.0
v26.3.1
v26.3.0
v26.2.0
v26.1.1
v26.1.0
v25.9.0
v25.8.0
v25.7.1
v25.7.0
v25.6.0
v25.5.0
v25.4.0
v25.3.1
v25.3.0
v25.2.0
v25.1.0
v0.8.8-beta
v0.8.7-beta
v0.8.6-beta
v0.8.5-beta
v0.8.4-beta
v0.8.3-beta
v0.8.2-beta
v0.8.1-beta
v0.8.0-beta
v0.7.9-beta
v0.7.8-beta
v0.7.7-beta
v0.7.6-beta
v0.7.5-beta
v0.7.4-beta
v0.7.3-beta
v0.7.2-beta
v0.7.1-beta
v0.7.0-beta
v0.6.9-beta
v0.6.8-beta
v0.6.7-beta
v0.6.6-beta
v0.6.5-beta
v0.6.4-beta
v0.6.3-beta
v0.6.2-beta
v0.6.1-beta
v0.6.0-beta
v0.5.8-beta
v0.5.7-beta
v0.5.6-beta
v0.5.5-beta
v0.5.4-beta
v0.5.3-beta
v0.5.2-beta
v0.5.1-beta
v0.5.0-beta
v0.4.5-alpha
v0.4.4-alpha
v0.4.3-alpha
v0.4.2-alpha
v0.4.1-alpha
v0.4.0-alpha
v0.3.8-alpha
v0.3.7-alpha
develop
v0.3.6-alpha
v0.3.5-alpha
v0.3.4-alpha
v0.3.3-alpha
v0.3.2-alpha
v0.3.1-alpha
v0.3.0-alpha
v0.2.5-alpha
v0.2.4-alpha
v0.2.3-alpha
v0.2.2-alpha
v0.2.1-alpha
v0.2.0-alpha
v0.1.5-alpha
v0.1.4-alpha
v0.1.3-alpha
v0.1.2-alpha
v0.1.1-alpha
v0.1.0-alpha
v0.0.62-alpha
v0.0.61-alpha
v0.0.60-alpha
v0.0.59-alpha
v0.0.58-alpha
v0.0.57-alpha
v0.0.56-alpha
v0.0.55-alpha
v0.0.54-alpha
v0.0.53-alpha
v0.0.52-alpha
v0.0.51-alpha
v0.0.50-alpha
v0.0.49-prealpha
v0.0.48-prealpha
v0.0.47-prealpha
v0.0.46-prealpha
v0.0.45-prealpha
v0.0.44-prealpha
v0.0.43-prealpha
v0.0.42-prealpha
v0.0.41-prealpha
v0.0.40-prealpha
v0.0.39-prealpha
v0.0.38-prealpha
v0.0.37-prealpha
v0.0.36-prealpha
v0.0.35-prealpha
v0.0.34-prealpha
v0.0.33-prealpha
v0.0.32-prealpha
v0.0.31-prealpha
v0.0.30-prealpha
v0.0.29-prealpha
v0.0.28-prealpha
v0.0.27-prealpha
v0.0.26-prealpha
v0.0.25-prealpha
v0.0.24-prealpha
v0.0.23-prealpha
v0.0.22-prealpha
v0.0.21-prealpha
v0.0.20-prealpha
v0.0.19-prealpha
v0.0.18-prealpha
v0.0.17-prealpha
v0.0.16-prealpha
v0.0.15-prealpha
v0.0.14-prealpha
v0.0.13-prealpha
v0.0.12-prealpha
v0.0.11-prealpha
v0.0.10-prealpha
v0.0.9-prealpha
v0.0.8-prealpha
v0.0.7-prealpha
v0.0.6-prealpha
v0.0.5-prealpha
v0.0.4-prealpha
v0.0.3-prealpha
v0.0.2-prealpha
v0.0.1-prealpha
Labels
Clear labels
ad-hoc
api
bug
ci-cd
content
dependencies
enhancement
frontend
in-progress
jellyfin
parked
priority: high
priority: low
priority: medium
review
security
One-off / ad-hoc work not tracked by a dedicated issue
REST API / HTTP endpoints
Something isn't working
Build, test, deploy pipeline
Channel content / schedules / playlists
Dependency updates (Renovate)
New feature or improvement
ChicoryTV React SPA frontend
Claimed by an active session — do not pick up
Jellyfin tuner / IPTV integration
Excluded from automatic queue pickup; work only when explicitly selected
Adversarial review finding
Security / vulnerability fix
No labels
priority: medium
Milestone
No items
No Milestone
Projects
Clear projects
No projects
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: timothy/ersatztv#796
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Spawned by #777, whose PR (#793) produced the evidence at first hand. Adjacent to #785 (mutation proofs for unproven guards), #790 (executable clause-level mutation harness) and #776 (hooks cannot be observed firing) — this is the same family measured from a different angle, so it should be read alongside them rather than as a new class.
The measurement
#793 added one small operator-run check. Five independent cold review rounds found five false greens in the verification code itself, and — the part that matters — each was introduced by the fix for the previous one:
[ -x "$cmd" ]-x;[ -x /bin ]is true{"id":1}, both passedperl -e 'alarm shift; exec @ARGV'wrapperexecfails → prints PASS having run nothingNone was hypothetical: each was reproduced with a control before being fixed.
The three transferable claims
1. The guard rules apply to the code that verifies, not only to the code under test.
testing.guard-ships-with-mutation-proofsays a guard ships with a proof it can fail. Every one of the five above was a guard without one — the harness was assumed correct because it was small and because it was the thing doing the checking. #785's 19 unproven guards are the same debt; this is evidence the debt extends to harness code nobody counted as a guard.2. The exit is usually subtraction. Round 4's fix was deleting the caps added in round 3, not guarding them — three defects had one cause. Round 5's fix was replacing lock-based ordering with unguessable random ids: strictly less code, and it removed the race rather than policing it. Compare
fix-the-boundary-not-the-site. When round N's fix creates round N+1's defect, adding a layer is the wrong move; the signal is to remove the layer that keeps generating them.3. Read-the-construct fails; execute-it works. #3 above is indistinguishable from correct by reading. It is the same shape as #776's finding that
exec … 2>/dev/nullsilences the shell and a trap'sexit "$?"invents a status — where shellcheck caught 0 of 5. Two independent sessions, same week, same conclusion: for shell and process-level constructs, static reading and linting do not substitute for running the failure path. Note the trap for a future reader: my own memory recommended that exactperl … execidiom as the macOStimeoutreplacement. Following documented advice verbatim produced the false green. The correct form isexec @ARGV or die.Scope
testing.verification-code-needs-its-own-proof), explicitly stating it covers harnesses/wrappers/timeouts, not just files named like guardsperl … execidiom correction into wherever the macOS-timeout recipe is documented, since the wrong form is currently the recommended onescripts/*.pyhelpers that are not in the guard population (scripts/mcp_smoke.pyis one: it is a checker, wired to no CI, and outsidedocs/guard-inventory.mdby design)Done-when
Independent corroboration from #776, and the
perlidiom is fixed at its sourceParallel session, same week, arrived at claims 2 and 3 from a different change. Adding the measurements rather than agreement, since two sessions agreeing is only worth something if the evidence is independent — and it is: different subsystem, different reviewers, no shared code.
The
perl … execrecipe (scope item 2) — correctedThe wrong form came from my session memory (
no-timeout-on-macos.md), which recommendedperl -e 'alarm shift; exec @ARGV'verbatim and had it marked "Verified 2026-08-05". Re-measured today:Memory file and its index line are corrected, with the failure mode stated rather than just the new form. It is not in any repo doc —
rg 'alarm shift'over the tree returns nothing — so session memory was the only carrier and that scope item is now closed unless you want it written intodocs/deliberately.Worth noting for the record: that file already documented this exact failure once —
gtimeoutexiting 127 and its empty output reading as "no block, correct". The replacement advice reintroduced the same shape. A correction is a fix, and fixes introduce adjacent defects; the correction was not itself proven. That is your claim 1 applied to a memory file.Claim 2 ("the exit is usually subtraction") — #776 is 3-for-3
Six review rounds. Every area patched kept producing defects; every area removed went quiet and stayed quiet:
INT/TERM/HUP)main. Silent in rounds 5 and 6exec … 2>/dev/nullsilencing the shellSo the trigger in your scope item 4 has a second dataset. My proposed threshold from #776 is lower than "N consecutive rounds": a mechanism that has produced two defects gets withdrawn rather than patched a third time. Both withdrawals here would have fired at that threshold and saved a round each.
Claim 3 (read-the-construct fails) — measured, 0 of 6
ShellCheck 0.11.0, default and
-o all, against #776's six shell defects: 0 detected, output wasSC2034unused-variable noise and anSC2250brace nit. Table is in a comment on #773. Two corpora now, both negative — §5.5's strike-through stands.Cross-link, so the two halves are not built twice
#794 (filed from #776) is adjacent but not the same rule, and the boundary is worth stating:
They meet at #785/#790. If both land, #794's prover is a natural way to discharge part of #796's rule — a harness change is a fix like any other, and the same revert-and-assert-red mechanism applies to it. Suggest whoever picks either one reads both first.
Claiming for this session (Claude Code / Opus 5, orchestrator).
Not bundling with #786/#789 — those were claimed by a different parallel session ~19:11 today, and their branch rewrites
docs/guard-inventory.mdand the guard-population tests. #809 and #819 are excluded here for the same reason (both require editingdocs/guard-inventory.md). This issue's only overlap with that branch is the generateddocs/decisions/README.mdcatalog, which is regenerated rather than hand-resolved after a rebase.Working the four Done-when boxes as scoped: the decision record (
testing.verification-code-needs-its-own-proof), the macOS-timeout idiom correction wherever the recipe is documented in-repo, and a stated position on mutation proofs for non-guard checker scripts.Progress and any scope cuts recorded here.
Progress — scope grew by one deliverable, on review evidence
The first commit did what the issue asked: recorded the rule and stated the position that operator-run
checkers get no mutation proof. Cold review blocked it, and was right. The position rested on
"
scripts/mcp_smoke.pycannot participate — it needs the gitignored.mcp.jsonand a live languageserver". That is false: it takes its config path and server name as positional arguments, so a test
can hand it a synthetic config pointing at a stub responder. The record had failed its own headline
rule (
READING THE CONSTRUCT IS NOT EXECUTING IT) on the one claim its decision rested on.The split was also drawn on the wrong axis. "CI cannot execute it" equally describes every
.claude/hooks/*.shand.husky/pre-push, four of which carryMUTATIONrows because a pytest proofdrives them in a sandbox. The discriminator is whether a proof test can drive the construct
hermetically.
So rather than restate an exemption a five-second experiment refutes, this now ships the proof:
scripts/tests/test_mcp_smoke.py(stub JSON-RPC responder; cases for impostor identity, a body carryingonly an id, a missing expected tool, and a pre-answered request; plus an anti-vacuity positive control),
and a declared clause in
mutation_manifest.pytargetingmcp_smoke.py's unguessable request id.Witnessed red rather than asserted:
id_init = secrets.randbelow(...)→id_init = 1makes thepre-answer accepted at
initialize, the run dies one stage later attools/list(rc 9 → 10), and theproof reddens with its declared diagnostic. Only that one test reddens — the mutation is clause-scoped.
mcp_smoke.pyitself still gets no inventory row (measured: it is rejected as a phantom, since thepopulation derives from workflow/hook call sites). The row goes to the test file, which joins the
population automatically — the
guard=test / target=scriptshape already used formutation_harness_lib.py.No follow-up issue is needed for this — the proof is in the PR rather than deferred.
Docs updated in the same PR:
docs/README.md(three guard rules → four) anddocs/guard-inventory.mdscope-limit item 6.
scripts/tests: 1103 passed, 2 skipped.Closing record
Outcome: Shipped
testing.verification-code-needs-its-own-proof— the proof obligation follows theverdict, not the file, so it binds harnesses, wrappers, timeouts and checkers, not only the files
docs/guard-inventory.mdderives. The issue asked for a position on whether non-guard checker scripts getmutation proofs. The position as first written was refuted by execution during review, so the PR ships the
proof instead:
scripts/tests/test_mcp_smoke.py(hermetic stub server, six cases) plus a declared clause inmutation_manifest.pytargetingscripts/mcp_smoke.py. PR #871.Root cause: Not a code bug — a verification-discipline gap.
mcp_smoke.pycarried three of #793's fivefalse greens; each was reproduced with a control at the time and none of those controls survived as a
test, so nothing re-ran them. The structural cause: the guard rules are scoped by a population of guard
files, and verification code sits outside it by construction — a
perlwrapper inside anifis not a file,and
mcp_smoke.pyis reached only transitively throughcheck-local-lsp.sh.Decisions/conventions changed: Added
testing.verification-code-needs-its-own-proof(active,supersedes: none— additive totesting.guard-ships-with-mutation-proofandtesting.mutation-claims-are-executed).Catalog regenerated.
Reusable knowledge:
mcp_smoke.pytakes its config path and server name as positional arguments, so a synthetic config pointing at a stub
responder drives it hermetically — no gitignored
.mcp.json, no language server..claude/hooks/*.shor
.husky/pre-pushas hooks either, yet several carryMUTATIONrows, because a pytest proof drivesthem in a sandbox.
rejected as a phantom). Name the proof test as
guardand the checker astarget; the test joins thepopulation automatically and carries the row. Precedents:
test_mutation_harness.py/mutation_harness_lib.py,test_review_verdict_vocabulary.py/check-review-verdict.sh.git adda new decision record before runningscripts/tests. The mutation sandbox takes its filelist from the git index with working-tree content (
mutation_harness_lib.py:22), so an untrackedrecord yields a sandbox whose catalog lists N+1 records against N files — "catalog is stale", and the
proof test then reddens for the wrong reason.
mcp_smoke.py'stools/liststage is held by thereader's pending-registration and an unguessable
id_tools: disarming either alone leaves every testgreen, only both together give rc 0. That is why
id_init— singly exploitable — is the declared clause.count each introduced a new wrong count, in the same direction. What fixed it was deleting the counts and
pointing at a single enumeration, and preferring an executing-site claim ("the place that RUNS it") over
a repo-wide uniqueness claim — the latter is falsified by the record quoting the idiom itself.
Verification:
scripts/tests1104 passed, 2 skipped.ruff format --check+ruff checkclean over all47 tracked Python files — a formatting red on the new file was caught by review, not by me, and would have
failed
script-tests.decisions_validate.pyOK; catalog regenerated, not hand-edited. Declared mutationwitnessed red with its declared diagnostic and clause-scoped; the harness re-applies it every run. Runtime
claims executed rather than read (
perl … execrc 0 vs rc 2 withor die;[ -x /bin ]true; notimeout/gtimeout; ShellCheck 0.11.0 flags neither shell defect). Rebased onto #868 mid-session and thegate re-run. Five cold-review rounds: four BLOCKED, round 5 MERGEABLE.
Deferred: None. The follow-up drafted for this (build the
mcp_smoke.pyproof) is in the PR rather thandeferred. Stated-not-implied gaps:
mcp_smoke.py's usage, config, command-resolution,--project, spawn,initialize-returned-an-error and zero-tools stages are not driven, and the unbounded raw line buffer is left
uncovered deliberately — the caps that would bound it were themselves defect 4.
Docs updated:
docs/decisions/records/testing/verification-code-needs-its-own-proof.md(new),docs/decisions/README.md(regenerated),docs/guard-inventory.md(row + summary counts + scope-limit item6),
docs/README.md(task-signal map).