Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
47 KiB
Recurring defect shapes across the closed-issue corpus (ersatztv#773)
Root-cause analysis over every closed issue in the tracker, the detectors that would have caught each class, and an audit of the tooling already configured in this repo. This is an analysis artifact, not a rulebook: the rules it argues for land as decision records and follow-up issues, which are linked per class below.
Measured 2026-08-13 against origin/main at f9f8f65ce. The issue corpus is a snapshot drawn at
19:45 that day — a boundary that matters, because #767 closed nineteen minutes later and is
therefore absent from every count here (§3.7).
1. Method, and what these numbers can and cannot support
Corpus. All 349 closed issues (state=closed&type=issues, 7 pages, 349 distinct numbers,
#1–#757 — pull requests excluded). Of those, 95 carry a ## Closing record comment. Each
record was classified by one of five independent agents against a written taxonomy, with
NEW:<name> available so the taxonomy could not be self-fulfilling. A sixth agent re-rated a
stratified 15-issue sample blind, as an inter-rater control. A seventh sampled 25 of the 254
pre-convention closed issues from their raw bodies and comments, to test whether the shapes exist
outside the era that writes records.
Three limits, stated up front because they bound every number below.
- Record coverage is not uniform. The
## Closing recordconvention starts at #520 (2026-07-21). Coverage is 31/32 for #600–699 and 4/69 for #1–99. So frequencies over the 95 records are frequencies over recent practice. - Corpus composition shifts underneath the measurement. The share of records describing no process failure at all falls from 12/19 (#60–#438) to 0/19 (#616–#671). Early issues are feature work; recent issues are largely CI/process hardening. Guard-shaped defects therefore dominate the recent records partly by construction — we have lately been building guards, so our defects are in guards.
- A closing record is the closing author's self-report. Raters were told to flag records whose own claims outrun what they describe, and several did (#60, #651, #691, #720). But the corpus systematically cannot contain a defect nobody noticed.
What limit 2 does not undermine. The obvious worry — that these classes are an artifact of
recent guard-building — was tested and refuted. The backward sample found high-confidence
instances of the same shapes well before the convention existed: twin-missed at #215 (REST playout
mutations skip the build-lock gating the Blazor path had) and #403 (an unhandled enum silent at 5 of
6 dispatch sites); wrapper drift at #198 and #287 (OpenAPI spec diverged from runtime JSON);
masking guards at #232, whose SPA PENDING_GRACE_TICKS "existed only to paper over the resulting
dishonest 200", and #234, where a blind finally-unlock was "papering over the stranded lock, so
the two had to land together"; overclaim at #1 ("the earlier 'confirmed' was based on one curl test
and several unverified assertions"). The convention changed the density and legibility of the
evidence, not the existence of the shapes.
Classification of pre-convention issues is harder but mostly possible — of the five the backward
rater flagged as unclassifiable, four (#94, #172, #265, #377) do carry a closing comment with a
clear root cause on re-inspection, and only #157 (an unresolved flaky-test report with a single
comment) is genuinely evidence-starved. The insufficient-evidence verdict is therefore rarer than
first reported, which if anything strengthens the backward result.
2. What the corpus actually says
95 records; 69 describe a process failure, 26 describe none. Percentages are of the 69.
| Rank | Class | n | % | Shape |
|---|---|---|---|---|
| 1 | Vacuous verification | 17 | 25% | A check was green having done no work |
| 2 | Twin-missed | 14 | 20% | The fix hit one instance and missed its structural twin |
| 3 | Symptom-keyed guard | 9 | 13% | The guard keyed on the symptom seen, not the defect's mechanism |
| 4 | (new classes — see §3) | 8 | 12% | Proposed by raters as fitting nothing above |
| 5 | Environment divergence | 7 | 10% | Behaviour differs by interpreter/provider/hardware, untested there |
| 6 | Wrapper drift | 5 | 7% | A hand-maintained mirror drifted from the contract it mirrors |
| 7 | Overclaim / stale claim | 3 | 4% | An assertion stronger than its evidence, or since gone false |
| 7 | String-predicate churn | 3 | 4% | A regex/grep predicate needing repeated rounds — but see §3.7: the rounds are cross-cutting, only the substrate is string-shaped |
| 9 | Identity-not-capability | 2 | 3% | Authorization scoped by who, not by what is being changed |
| 10 | Masking guards | 1 | 1% | Two guards on one condition; one hid the other's total failure |
This reorders the issue's own ranking. #773 put twin-missed first (9 instances) and vacuous verification second (8), from a 50-issue sample. Over the full record corpus they swap: vacuous verification is the most common single shape. Overclaim also drops sharply as a primary cause (4%) — raters overwhelmingly assigned it as a secondary. It is better understood as a modifier that rides on another failure than as a class with its own detector.
Inter-rater control (measured, not asserted). A sixth rater re-classified a stratified 15-issue sample blind. Raw agreement 12/15 (80%) across an 11-category scheme. The three disagreements are informative rather than noise:
- #177 (twin-missed → no-failure) and #521 (no-failure → twin-missed) disagree at the "did a process failure occur" boundary and point in opposite directions, so that boundary is noisy but not systematically biased. The 26 clean-record count should be read as ±3.
- #649 was rated twin-missed by one and wrapper-drift by the other — which is precisely the blur that Family C in §3 exists to resolve. Both raters named the other class as secondary.
The top-three separation is far larger than the disagreement. Treat ±1 rank as noise.
3. Consolidation: four families, one of which the issue's taxonomy missed
Family A — reasoning about a representative instead of the population
#773's central hypothesis, confirmed, and it extends further than proposed. The issue asked whether twin-missed and vacuous verification are the same underlying error. They are — and symptom-keyed guards belong with them. All three are the same mistake at different targets:
- twin-missed — the fix was applied to a sample of the population,
- vacuous-by-sampling — the verification sampled an empty or unrepresentative subset,
- symptom-keyed guard — the guard matched only the member that was noticed.
Size: 27 of 69 (39%) — 9 twin-missed, 9 symptom-keyed, 9 vacuous-by-sampling. See §3.6 for the full partition; the number is smaller than the naive merge because records that belong to Families C and D are not double-counted here.
The merge is not merely rhetorical; it predicts a single detector, and the corpus already contains that detector, reinvented several times without anyone noticing it was the same rule:
"the by-id handler covered 4 of 10 media types… Sweep by FIELD, not by the call site the issue names" — #671 "#644's guard keyed on the symptom (an inflated pageSize) rather than the defect… an at-cap request was structurally invisible to it" — #650 "filtered on tools that already declared
pageNum, so a tool wrapping a paged endpoint with no paging args escaped it entirely" — #616 "the pre-existing test filtersWhere(t => t.QueryParameters is {Count: > 0}), so a tool that lost its query parameters escaped it entirely" — #757
Four sessions, four reinventions of one rule → detector A in §4.
Scope limit, because it bounds the highest-value detector in this document. The distinction is what kind of thing the population is made of.
- A population of values — enum members, OpenAPI operations, registered tools — always has an
external authoritative source, and detector A applies directly.
Enum.GetValues<T>()is that source for an enum, which is precisely #503's fix for a catch-alldefaultthat silently absorbed an unhandled value. - A population of sites in code — the places that dispatch on a value, rather than the values
themselves — has no external list to assert against. #403 is the case:
PlaybackOrderis an enum, so its values are enumerable, but the defect was that 5 of 6 dispatch sites failed to handle one, and no artifact anywhere enumerates those sites. That residue needs find-all-references tooling, which is broken here and tracked in #777, not a set-equality assertion.
Family B — the check never ran at all
Size: 9 of 69 (13%) — the 8 vacuous records that are not sampling errors, plus the one masking-guards
record. Here the check was dead, not mis-aimed. scripts/tests/
invoked by no CI job (#631); a validator whose call site could be deleted with the suite still green
(#621); a ${{ }} in a shell comment silently dropping a step while the job reported success in
6s (#751); if ! cmd; then status=$? reading bash's logical negation, so a failing spec run exited 0
(#445); new pre-push logic never wired to receive stdin, "dead code that every unit test still passes
over" (#719). A sampling detector does nothing here. This needs detector B: proof the guard can
go red.
Family C — two copies of one thing
Size: 9 of 69 (13%) — the 5 wrapper-drift records plus the 4 twin-missed records whose twin is literally a second copy (#510, #649, #711, #756). Those 4 are counted here and not in Family A.
.claude/ and its byte-identical .codex/ mirror where only one was in the PROTECTED list (#711);
an advisory local hook and the server-enforced workflow where four rounds of hardening landed on the
copy with the lower stakes (#649); a hand-maintained MCP schema drifting from the generated OpenAPI
DTO by one field and silently clearing it (#754, again #757); a hand-rolled frontmatter parser
accepting YAML that PyYAML rejects (#674); a dropped-step guard added to one required workflow and
not its fail-open twin (#756).
The preferred fix is deletion — one implementation, called from both places, as #649 did by
extracting scripts/pr-changed-files.sh and recording ci.shared-pr-file-enumeration. (#649 also
shows why the duplication is dangerous: before the extraction the guard was "safe only by
redundancy," and a failing exit status "previously left the entire suite green.")
But dedup is a preference, not a law, and #711
is the counterexample that proves it. That record deliberately kept its enumerative list:
"the list stays ENUMERATIVE rather than derived — a derived rule would have to be evaluated against
the very file list being classified, putting more moving parts inside a security predicate to save
one line per new tooling directory." Deriving a security predicate from the input it judges is
worse than maintaining two entries. So: dedup by construction where the duplication is incidental;
keep the enumeration where deriving it would feed the judged input back into the judge.
Family D — check-and-use race over mutable state (the taxonomy missed this entirely)
Size: 5 of 69 (7%). Three raters proposed overlapping new classes — unprompted by the taxonomy,
which had no bucket for this, though they shared a corpus and a brief and so were not strictly
independent. The stronger evidence is the blind re-rater, who saw none of their output and
proposed the same class again for #536 under its own name (toctou-partial-atomicity vs
half-atomic-toctou).
Four of the five (#536, #622, #706, #707) were tagged NEW by their raters. #632 was rated
twin-missed as its primary, with the race named as its secondary; it sits here because the
mechanism is the race, and it is counted here rather than in Family A.
| # | The race |
|---|---|
| 536 | Check-then-act split by an await; the write side was atomic, the read side a stale Volatile.Read |
| 622 | Merge consent bound to a head sha, then Gitea's async auto-merge evaluated against a later head |
| 632 | Head sha bound, but retargeting the PR base changes the diff while moving neither sha nor status |
| 706 | Status writes are read-then-write with no compare-and-set; an older run can finish last and win |
| 707 | Paged file enumeration diffs each page against the base's live tip, so a mid-paging advance drops rows |
Five instances — larger than identity-not-capability, string-predicate churn and overclaim individually, and #773's taxonomy had no bucket for it. The unifying property: a check and the action it authorizes are separated in time over state that can change in between, with nothing pinning a version. #622's own record puts it exactly: "The gate was never bypassed — it was satisfied against a snapshot that stops being true."
Two further cross-era classes, from the pre-convention sample
- Fail-open by default — a surface that defaults to permissive when config is absent or a step
vanishes: API writes open when
Api:WriteKeyis unset (#280); a dropped step is fail-CLOSED inreview-verdict.ymlbut fail-OPEN indocker-build.yml(#751, #756, #768); a jq-1.6-inert classifier (#647). Distinct from identity-not-capability: nothing is mis-scoped, the default is wrong. - Destructive replace — a full-replace path that silently destroys state a reconcile-by-id would have kept: a schedule PUT resetting fill-group progression, which #252 fixed with a "scoped positional/no-op reconcile" so that a no-op PUT-back stops being destructive; a full-replace wrapper missing one field and clearing it (#754); a dedup fix that turned a duplicate row into permanent data loss because the remove filter used a different key (#500).
3.6 The partition, stated exactly
The families overlap conceptually — a twin-missed record can be read as "another member of a population" or as "a second copy" — so membership is given as an explicit partition, mechanically verified to cover all 69 records exactly once. Read a family's n as a disjoint count, not a tally of everything the family's description could fit:
Family letters A–E are the ones §4 names detectors for, and those letters correspond. This table assigns no letters below E. §4's detectors F and G address fail-open and destructive-replace, which are cross-era classes drawn from the pre-convention sample and are not members of this 69-record partition — so an F or G in §4 refers to nothing in this table.
| Family | n | % | Composition |
|---|---|---|---|
| A population reasoned about via a sample | 27 | 39% | 9 twin-missed + 9 symptom-keyed + 9 vacuous-by-sampling |
| B the check never ran | 9 | 13% | 8 vacuous-by-non-execution + 1 masking-guards |
| C two copies of one thing | 9 | 13% | 5 wrapper-drift + 4 twin-missed-by-duplication |
| E environment divergence | 7 | 10% | |
| D check-and-use race over mutable state | 5 | 7% | 4 rated NEW + #632 rated twin-missed with the race secondary |
| — overclaim / stale claim | 3 | 4% | no detector — see §4 |
| — string-predicate churn | 3 | 4% | no detector proposed; see §3.7 — the round-churn it is named for is cross-cutting |
| — identity-not-capability | 2 | 3% | |
| — unmerged singletons | 4 | 6% | #60, #503, #586, #688 — real, but one instance each |
| Total | 69 | — | the n column is exact; percentages are rounded to the nearest point and sum to 99% |
The partition is checked mechanically rather than by eye: the assertion is that the union of the families equals the set of process-failure records, with no overlap. That check immediately caught a transcription slip (#438 omitted from the control set) — a small live demonstration of detector A applied to this analysis's own numbers.
Two caveats travel with these percentages. The Family A/C boundary is a judgment about whether a twin is "a second copy" or "another member of a population", and #649 is exactly the record the two raters split on. And per §1 limit 2, these are proportions over a recent, guard-heavy corpus — the backward sample establishes that the shapes existed before, not that they occurred in these ratios.
3.7 Round-churn is a property, not a class — and the fix text is the fourth sampling target
The ranked table lists string-predicate churn as a class, inherited from #773's own taxonomy ("a parser / string-matching predicate needing repeated rounds"). That name identifies a substrate when the evidence identifies a cross-cutting property. Counting records that narrate three or more review rounds:
| Family | n | records with ≥3 rounds |
|---|---|---|
| A — population via a sample | 10 | #460, #496, #616, #633, #644, #650, #671, #684, #726, #743 |
| B — the check never ran | 5 | #445, #620, #631, #685, #751 |
| C — two copies of one thing | 5 | #510, #649, #754, #756, #757 |
| E — environment divergence | 5 | #491, #643, #647, #648, #668 |
| D — check-and-use race | 3 | #622, #632, #706 |
| identity-not-capability | 2 | #697, #698 |
| string-predicate churn | 2 | #578, #629 |
| overclaim | 1 | #651 |
| Total | 33 of 69 | every family represented |
Method: every one of the 69 records was read for round language — 48 mention "round" at all, 33 state an explicit count of three or more. Two attributions are loose: #647 narrates four rounds that happened on #643, and #629 counts a round it deliberately did not run.
The class named after the phenomenon holds 2 of the 33. Round-churn should therefore be read the way §2 reads overclaim — a modifier riding on another failure, not a class with its own detector. #697 settles it independently: three BLOCKED rounds on credential scoping, no parser anywhere near it; #698 ran six.
One honest qualification against over-correcting. String predicates are disproportionately round-prone — 2 of that class's 3 records, against 10 of Family A's 27. The naming error is not that the association is false; it is that the class was defined by the substrate where the property was noticed, so 31 instances outside it had nowhere to be counted.
Why the rounds happen, which is the part that generalises. §3 Family A names three things that get sampled instead of enumerated — the fix, the verification, the guard. There is a fourth: the review scope. A re-review briefed to "check the reported finding" samples; the population is the whole changed artifact, and the text written to fix the last round is unreviewed by construction at the moment it is written. That is why each round's defect lands in the newest prose rather than in the text under review.
Two limits on this section, since it is the one making a claim about its own production.
-
33 is a floor, not a total. It counts records that state a round count; a record that ran four rounds without narrating them is invisible here. Do not read the 36 remaining records as single-round work.
-
Counting this by pattern rather than by enumeration under-reports it by roughly half. A regex over the corpus finds 18 of these 33 — the misses narrate their rounds in the same words as the hits. If you re-derive this number, read all 69.
-
#767 — the case #773 describes as taking eight review rounds — is excluded by snapshot timing, not by absence. The corpus was drawn at 19:45 on 2026-08-13; #767 closed at 20:04, nineteen minutes later, and now carries a full
## Closing record. So the worst instance of the phenomenon is missing from every count in this document, including the 33 above. It is named here rather than quietly folded in, because re-drawing the corpus would move every denominator in §2 and §3.6 and the snapshot boundary has to sit somewhere.Reading its record changes nothing structurally and confirms Family A twice over — it reinvents detector A independently, for the fifth and sixth times in this repo: "Refuting two variants of a channel is not clearing the channel" (the working attacks through that channel were never tried), and "Enumerating shapes of a command loses. Nine disarms across two rounds; running the command settles them together." That last sentence is detector A in one line — stop sampling the shapes you thought of, execute the population.
4. Detectors
Ranked by instances covered per unit of build cost. "Rule status" distinguishes a docs problem (no rule exists) from a hooks problem (the rule exists and nothing enforces it) — they have different fixes, and conflating them is why several of these recurred.
| # | Detector | Covers | Rule status |
|---|---|---|---|
| A | A guard derives its expected set from the authoritative source and asserts set equality — never filter, never a sample. The population comes from the enum / the OpenAPI doc / the workflow YAML / the provider list, and the assertion is equality against it. A filter over the population cannot see the member that is missing. |
Family A: 27 records (39%), minus the code-structural residue noted in §3 | rule-missing as a general rule; instantiated ad hoc in #616, #644→#650, #671, #757 |
| B | Every guard ships with a proof it can go red: delete or disarm that guard alone and the suite must fail. Not "a test exists" — a mutation. | Family B: 9 records (13%) | rule-present-unenforced — stated in #685's record and in project memory, enforced nowhere |
| C | Dedup by construction. When two copies of one rule exist, delete one and have both callers invoke it — unless deriving the list would feed the judged input back into the judge, which is #711's reasoned exception. | Family C: 9 records (13%) | rule-missing |
| D | Anything read-then-written against live Gitea/remote state pins a version or uses compare-and-set. See the honesty note below — this is a fix pattern, and its detector is detector A applied to an enumerated inventory of such sites. | Family D: 5 records (7%) | rule-missing (partially addressed for the merge path by the per-sha required check from #622) |
| E | Run the check under the interpreter/provider it will actually run under, and preflight-log the version. | Environment divergence: 7 records | rule-present and working — scripts/jq-preflight.sh + ci.jq-version-contract closed the jq axis after #643/#647/#648. Unclosed axes: SQLite-vs-MySQL query semantics (#668), GPU generation (#505), CI-VM speed (#512) |
| F | Test the DENY path with the production config value. #756's lesson generalised: a fixture that omits a field tests only the default, so a fail-open in the production value stays invisible. Parametrise the whole matrix, including "config absent". | Fail-open-by-default: ~5 records | rule-missing |
| G | Full-replace endpoints assert their complete field list in a test, and prefer reconcile-by-id over delete-and-reinsert where child state exists. | Destructive replace: ~3 records | rule-missing |
Detector A is the highest-value single change in this analysis — one convention, 39% of the recorded process failures over this corpus, and it is already proven four times in this repo under four different names. Two honesty notes on that ranking, because it drives #774's priority:
- The 39% is measured over a recent, guard-heavy corpus (§1 limit 2). The backward sample shows the shape predates that corpus; it does not show the proportion holds across eras. Read it as "the largest family in the work we have been doing lately," which is still the right basis for prioritising the next change, but not as a timeless property of the project.
- Detector A also applies to review scope, which is free. Per §3.7, a re-review briefed to "check the reported finding" samples the artifact. Briefing it to sweep the whole changed artifact enumerates it. This costs one sentence in a review brief and is the only detector here with no build step at all.
- Detector D is weaker than its neighbours in this table and is listed anyway. A, B, F and G each name a check that fails when violated. D names a fix pattern with no general lint — you cannot mechanically spot "this code should have pinned a sha." What makes it actionable is that the population is small and enumerable: the handful of scripts that touch live remote state. So D's real detector is detector A applied to that inventory, which is why #778's scope is "enumerate every such site and mark each pinned / CAS / knowingly-unsafe" rather than "write a linter."
Classes where no mechanical detector is plausible — stated rather than papered over
The issue explicitly asked for this, and inventing a weak detector here would itself be the symptom-keyed-guard mistake.
- Overclaim / stale claim. No check can tell that a sentence is stronger than the evidence
behind it. Partial mitigations exist and should not be oversold:
stale-afterfrontmatter (#603) dates a claim, and #578's retracted-term grep catches a specific known retraction propagating into generated artifacts. Neither detects a fresh overclaim. This stays a review responsibility. - Per-task review blindness (#60). A review scoped to one task's diff structurally cannot see a defect that only exists once several tasks compose — five Important findings at #60 were invisible to six per-task reviews and surfaced only in a whole-branch pass. The fix is a process step (a whole-branch review before close), not a check.
- Omitted brief constraint (#586). A delegated brief silent on a hazard gets a plausible-but-wrong default — "an omitted rule isn't an unenforced rule, it's a rule replaced by whatever default the agent reaches for." A brief lint is conceivable but would be a keyword matcher, i.e. exactly the string-predicate class. A hazards checklist in the brief template is the honest ceiling.
- Guard-parity by verb (#458), and re-deriving an inherited exemption against a new failure mode (#484). Both need judgment about whether a prior rationale still applies.
A meta-finding about this repo's own knowledge base
docs/decisions/records/ holds 189 active records, and the two most relevant to this analysis
(testing.enumerating-guard-identity-not-position, ci.required-job-step-execution-markers) are
extraordinarily detailed — each a full account of one incident. The corpus grows one record per
instance. That is Family A operating on our own process: we are enumerating cases rather than
removing the mechanism. Detector A, C and B are class-level rules precisely because the per-instance
record has already been tried 189 times.
5. Part 2 — audit of the tooling already configured
Judged against the measured classes above, per the issue's sequencing. Everything here was run, not assumed.
5.1 The LSPs: two of three are broken, and the one that matters most is the most broken
#773 asked whether routine LSP use would catch Family A (find all references rather than the one in front of you). It would help — and it is not available. Measured in this session:
| LSP | State | Evidence |
|---|---|---|
csharp-lsp |
Broken — cannot initialize | findReferences on ChannelPlaylist.ToM3U() → System.InvalidOperationException: .NET SDK cannot be resolved, because libhostfxr.dylib cannot be found inside /opt/homebrew/Cellar/dotnet/10.0.302/bin/host/fxr. That directory does not exist — Homebrew's dotnet layout is not what MSBuildLocator expects. dotnet --version works (10.0.302), so builds are fine; only the language server is dead. |
typescript-lsp |
Broken at the repo root | Could not find a valid TypeScript installation… ensure that the "typescript" dependency is installed in the workspace. typescript lives in web/node_modules, not at the workspace root. Configuration problem, not a missing dependency. |
pyright-lsp |
Works | documentSymbol on scripts/decisions_lib.py returned the full symbol tree; findReferences on active_files correctly returned 4 references across 3 files, including cross-file hits in decisions_validate.py and migrate_decisions_split.py. |
Two conclusions, and the second is the sharper one.
- The LSP that would help most is the one that is dead. The residue detector A cannot reach —
populations of sites in code rather than of values — is overwhelmingly C#: #403 (5 of 6
dispatch sites) and #671 (a by-id handler covering 4 of 10 media types) are "find every site"
problems, and
csharp-lspcannot answer a single query. (#510 is not an example here despite looking like one: its record names duplication as the root cause and its switch keys on a plain enum, so it belongs to Family C andEnum.GetValues<T>()is its authoritative source.) - A dispatched subagent could not reach the LSP tool at all. The agent tasked with testing the
three LSPs reported
ToolSearchreturning "No matching deferred tools found" for every query, while the same tool resolved immediately in this main session. So the standing note "workflow agents must use csharp-lsp" is doubly rotten: the server is broken and the agents it addresses cannot invoke it even when it works. This is a live instance of overclaim/stale claim (§2 rank 7) sitting in our own guidance.
Resolved 2026-08-14 (#777), and one claim above needs correcting. Both servers now work and both
answer a real cross-file query in this repo; the full setup, traps and verification are in
docs/local-lsp-tooling.md, and the rule is session.local-code-intelligence.
| Row above | Resolution |
|---|---|
csharp-lsp |
MSBuildLocator needs a dotnet root that owns host/fxr, which Homebrew's bin does not and its libexec does. Fixed by env.DOTNET_ROOT in .claude/settings.local.json. findReferences on ChannelPlaylist.ToM3U() returns the declaration plus its 5 call sites, excluding the mention of the name in a comment that grep matches. |
typescript-lsp |
No configuration lever exists — v5 dropped --tsserver-path, and a plugin lspServers entry cannot pass initializationOptions. Fixed by making the package resolvable from the workspace root (a gitignored root node_modules/typescript link). Returns 20 references across 7 files for canLeaveCurrentScreen. |
The correction to conclusion 2, which matters because the guidance it judges is still in use: the
note "workflow agents must use csharp-lsp" names the csharp-lsp MCP server tools
(csharp_set_workspace, csharp_diagnostics, csharp_references, …), not the LSP tool. MCP tools
are reachable from a subagent. The subagent measurement above is correct about the LSP tool and
was generalised one step too far: the note was unsatisfiable because that MCP server's .mcp.json
entry named a dotnet install that no longer existed, so it never started — not because agents cannot
invoke it. With the entry repaired the server serves 16 tools, so the guidance becomes satisfiable —
inferred from the MCP boundary generally (326 subagent MCP calls across four other servers), not yet
measured for csharp-lsp from a subagent. What must be briefed explicitly is which surface: pointing a subagent at the LSP tool
is still an instruction it cannot obey.
5.2 Python: the global instruction and this repo disagree, and the repo is silent
~/.claude/CLAUDE.md instructs ruff check, ruff format --check, and pyright after modifying
Python. Measured against reality:
| Claim | Measured |
|---|---|
| ruff/pyright run in CI or hooks | No. grep -rn "ruff|pyright" .gitea/workflows/ .husky/ returns nothing. |
| What does run on Python | pr-checks.yml → decisions-guard (decisions_validate.py, build_decisions_catalog.py --check) and script-tests (pytest scripts/tests). Neither Husky hook touches Python. |
ruff check scripts/ |
47 errors — but 28 are E702 (semicolons) and 7 S105 "hardcoded password" are false positives on test stubs (env["ETV_GITEA_TOKEN"] = "stub"). The real cleanup is small. |
ruff format --check scripts/ |
10 of 20 files would be reformatted. |
pyright scripts/ |
2 errors, both reportMissingImports for etv_client in scripts/scripted-schedules/entrypoint.py — a package resolvable only in that script's deploy environment. Effectively clean. |
| Repo-level ruff config | None. No ruff.toml/pyproject.toml. Ruff silently falls back to whichever ~/.config/ruff/ruff.toml the operator's machine happens to have. |
The last row is the finding that matters, and it is Family E (environment divergence) in our own toolchain: Python lint behaviour here is a function of an un-versioned file on one laptop. A second machine lints differently, or not at all. One of the two must move — either the repo adopts a committed ruff config and enforces it, or the global instruction stops claiming this repo enforces something it does not.
5.3 Configured vs actually invoked
Measured over the session transcript corpus (811 files under
~/.claude/projects/-Users-timothy-ersatztv/, this session excluded).
| Tool | Configured | Actually invoked | Verdict |
|---|---|---|---|
gitea MCP |
project .mcp.json and user scope (duplicate, identical values) |
Heavy, same-day | Keep; de-duplicate the config |
ssh-mcp |
project .mcp.json |
16 transcripts, same-day | Keep |
mempalace MCP |
user scope | 10 transcripts, last hit 2026-07-25 (19 days) | Keep, but see below |
LSP (all three servers) |
3 plugins enabled + csharp-lsp in enabledMcpjsonServers |
0 calls in 811 transcripts | Broken and unused — both C# and TS servers repaired 2026-08-14 (#777); the zero is the baseline a future measurement is compared against |
context7 |
enabled, CONTEXT7_API_KEY set |
0 calls | Dead |
playwright |
enabled | 145 calls, last 2026-07-25 | Keep |
superpowers |
enabled | 29 transcripts, last 2026-08-04 | Keep |
feature-dev, ralph-loop, security-guidance |
enabled | 0 skill invocations | Dead as configured |
codex plugin |
enabled | 0 skill invocations — but the documented workflow is codex exec via Bash, which this search cannot see |
Not dead; measured the wrong surface |
nuget, docker-mcp |
in .mcp.json, explicitly disabled |
0 (expected) | Correctly off |
The zero for LSP is verified rather than assumed: the same query shape returns 23,661 Bash calls
and 5,102 Read calls over the same corpus, so the search was demonstrably capable of finding hits.
Re-derive it with that positive control, not on its own — rg -c over multiple files prints
path:count, so the obvious way to total it sums the paths and returns zero for everything.
Hooks are the good news. All 13 scripts in .claude/hooks/ are wired from either
.claude/settings.json or .husky/*, and no settings entry points at a missing path — there are
no dead hook scripts, contrary to the issue's suspicion. Husky hooks have genuine fired-output
evidence in transcripts (husky - dotnet format found whitespace/BOM issues, husky - commit message missing Co-Authored-By trailer, husky - refusing to commit root-level screenshot(s)).
Wiring was the strongest claim available for the hook rows when this table was written. It is no
longer the ceiling — §5.4 is now measured — so read the hook rows against
scripts/hook-fire-log.sh report, not against this paragraph.
5.4 Whether our own hooks fire — now measured (#776 closed this)
The finding as originally recorded. In this harness version only Stop hooks emit a structured
record (stop_hook_summary/hookInfos). PreToolUse and PostToolUse hooks — which is every
guard that matters here: merge consent, worktree, BOM, agent model/RAM — leave no durable execution
trace. What this audit could count for those hooks was filename mentions in settings dumps, which is
not evidence of execution. So the guards this repo relies on most were exactly the ones whose
execution could only be inferred: Family B (the check never ran) applied to the hook layer itself,
after this project had already paid for it twice at the CI layer (#751, #756).
What replaced the inference. Every hook now records its own execution through one shared sink,
scripts/hook-fire-log.sh (testing.hook-reports-its-own-execution): a fire record on entry and
an exit record carrying the status and the decision, where the decision is parsed from the bytes
the hook actually emitted rather than declared by its author. Read it with
scripts/hook-fire-log.sh report [--all].
Measured 2026-08-14 — a scripted headless session (a plain Bash call, a Bash call carrying
ETV_UPDATE_GOLDENS=1, a Write to a SPA file, and one Agent dispatch) plus a real git commit
and a git push --dry-run in a worktree. Snapshot boundary: this is one deliberately-constructed
window, not a corpus statistic.
| Hook | Event | Fires | Decisions observed |
|---|---|---|---|
pretooluse-bash-guard |
PreToolUse/Bash | 2 | deny 1, no-op 1 |
pretooluse-bom-guard |
PreToolUse/Bash | 2 | no-op 2 |
pretooluse-worktree-guard |
PreToolUse/Bash | 2 | no-op 2 |
posttooluse-worktree-marker |
PostToolUse/Bash | 1 | no-op 1 |
design-sync-reminder |
PreToolUse/Write + Stop | 4 | context 1, block 1, no-op 2 |
pretooluse-agent-model |
PreToolUse/Agent | 2 | no-op 2 |
pretooluse-agent-ram |
PreToolUse/Agent | 2 | no-op 2 |
decisions-guard |
git pre-commit | 2 | pass 2 |
prepush-clean-worktree-check |
git pre-push | 1 | pass 1 |
prepush-donewhen |
git pre-push | 1 | pass 1 |
prepush-rebase-check |
git pre-push | 1 | pass 1 |
pretooluse-nav-guard |
PreToolUse/navigate | 0 | not exercised — needs a live browser session |
pretooluse-merge-consent |
PreToolUse/PR write | 0 | not exercised — needs a real merge attempt |
11 of 13 hooks are confirmed firing, with the decision each reached. The deny row is the load-
bearing one: pretooluse-bash-guard did not merely run, it blocked the ETV_UPDATE_GOLDENS=1
probe, so at least one guard in this set is demonstrably live rather than merely present.
Re-verified against the shipped implementation. The table was first measured against an early
version of the sink, and the classifier changed materially afterwards, so the run was repeated
against the final code: pretooluse-bash-guard deny+no-op, pretooluse-bom-guard and
pretooluse-worktree-guard no-op, posttooluse-worktree-marker no-op, design-sync-reminder
context+block+no-op — identical decisions. The repeat run exercised the five Claude hooks a
Bash/Write session reaches; the Agent pair and the four git hooks are carried over from the
original run and were not re-measured.
Reproducing this table. It is a constructed window, not a corpus statistic, and it is not
reproducible from a reader's own report output — running the test suite alone would not produce
it, and before scripts/tests/conftest.py landed, running the suite actively polluted the default
log with synthetic fires. To re-derive: point ETV_HOOK_FIRE_LOG_DIR at an empty directory, run a
headless session exercising the four tool paths above, then a git commit and a
git push --dry-run, and read scripts/hook-fire-log.sh report --all --dir <that directory>.
A zero is two findings wearing one number — a hook that is broken, and a hook whose trigger did not occur — and the report says so rather than presenting a zero as a verdict.
For the two zeroes here, what is established and what is not, kept apart deliberately.
Established: both script bodies work. pretooluse-nav-guard and pretooluse-merge-consent are
each driven to their deciding branch in scripts/tests/test_hook_fire_log.py — a deny on an
/iptv/ URL and an ask on a merge call — emitting the correct decision with the instrumentation
in place. Not established: that the harness would dispatch to them. Those tests invoke the
scripts directly, so they bypass registration and matcher dispatch entirely; a typo in
.claude/settings.json, a settings file that never loaded, or a matcher that does not match would
leave both tests green while the production zero still meant broken wiring. A working script is a
necessary condition, not the finding.
So these two zeroes remain genuinely ambiguous, and neither can be resolved without the thing that
resolves it: a live browser session, or a real merge attempt. Manufacturing a merge to observe the
merge guard is a worse idea than the gap it would close. Distinguishing the two meanings of zero
stays a human judgement — testing.hook-reports-its-own-execution says so, and an earlier draft of
this paragraph quietly contradicted it by treating "the script works" as "the wiring works".
The measurement earned its keep on its first run. pretooluse-agent-model was recorded deciding
no-op on an Agent dispatch where ask was expected. Previously that would have been an
unanswerable suspicion about a guard nobody could observe; instead the real payload was dumped and
it carried model: haiku, so no-op was correct and there was no defect. The general form is the
standing #756 lesson — make the system report it rather than infer it — and note what it
displaced: an argument from the hook's source about what it "must" do, which is the reasoning shape
§3 measures going wrong repeatedly.
5.5 Tools we do not have that would address a named class
Each is justified against specific issues, not general merit; anything that could not be tied to recorded instances is left out rather than padded in.
| Tool | Class it addresses | Justification |
|---|---|---|
| Stryker.NET (mutation testing for C#) | Family B — detector B, mechanised | Detector B currently relies on an author remembering to write a mutation proof. Stryker generates them. #621 (a guard whose call site could be deleted with the suite green), #685 (two guards where deleting either left the suite green) and #719 (logic never wired to stdin) are all surviving-mutant detections by construction. Cost is real (mutation runs are slow), so scope it to the guard/validator projects rather than the whole solution. |
shellcheck |
(would have been Family B, shell half) | The obvious candidate for #445 (if ! cmd; then status=$?, where bash sets $? to the logical negation so a failing run exited 0). Measured: it does not catch it. ShellCheck 0.11.0 on that exact construct reports nothing, and even -o all returns only an unrelated brace-style nit; run against our 13 hook scripts it yields 2 SC2034 unused-variable warnings. Recorded here as a negative result so it is not proposed again on plausibility. |
A duplicate-code detector (jscpd or equivalent) |
Family C | #440 shipped two byte-identical Lucene escapers in different files; #711 and #649 are the same shape at the config/CI layer. A duplication report would have surfaced all three at authoring time — subject to #711's exception, so it should advise, not block. |
An enum-exhaustiveness analyzer rule for C# switch statements |
Family A, code-structural residue | #503's catch-all default silently absorbed an unhandled enum value and produced byte-identical output to a handled one. Switch expressions already warn; statements do not. This is a rule to enable in the analyzers we already run (Meziantou), not a new dependency — the cheapest item here. |
Explicitly not proposed: a "brief linter" for omitted delegation constraints (#586) and any overclaim detector. §4 argues both are implausible, and inventing them here to look thorough would be the symptom-keyed mistake. Nothing in the corpus suggests a tool we lack would beat detector A for Family A — the four in-repo reinventions show the convention works when applied; the failure is that it was never written down once.
Evidence status of these rows, which differs and matters. Only the shellcheck row was executed — and running it refuted the reason it had been proposed, which is why it is struck through. The three surviving rows are reasoned from precedent: each names issues whose recorded mechanism the tool addresses, but none has been run against this repo to confirm it fires. That is a weaker warrant and should be discharged before any of them is adopted — a tool justified by "it plausibly catches this class" is the same species of claim as the guards this document criticises. Treat these three as candidates to test, not findings.
On MemPalace, the 19-day gap deserves care rather than a verdict: CLAUDE.md makes it the
documented discovery path for decisions, so either sessions are skipping the documented path, or
they are correctly reading docs/decisions/README.md directly (which the same contract permits, and
which is authoritative). The measurement cannot distinguish those, and I am not going to guess.
6. Follow-ups
This issue is analysis; it spawns implementation rather than doing it.
| Issue | Detector / finding | Covers | Priority |
|---|---|---|---|
| #774 | A — a guard derives its population from the authoritative source and asserts set equality | Family A, 27 records (39%) | high |
| #775 | B — every guard ships a mutation proof: delete that guard alone, see red | Family B, 9 records (13%) | high |
| #776 | Make hooks report that they fired — PreToolUse/PostToolUse execution is unobservable | the whole hook layer | high — DONE, §5.4 is measured |
| #777 | csharp-lsp and typescript-lsp are broken; the "workflow agents must use csharp-lsp" note is stale |
partial mitigation for Family A | medium |
| #778 | D — pin a version or use compare-and-set for read-then-write against live remote state | Family D, 5 records | medium |
| #779 | F + G — test the deny path with the production config value; assert full-replace field lists | fail-open + destructive-replace, ~8 records | medium |
| #780 | Reconcile Python tooling — commit a ruff config and enforce it, or stop claiming we do | environment divergence in our own toolchain | medium |
| #781 | Retire enabled-but-never-invoked plugins/MCP servers; de-duplicate the gitea server |
config that reads as coverage | low |
Detector C (dedup by construction) has no issue of its own on purpose. It is not a thing to build; it is the shape the fixes in #774 and #778 should take when they find two copies of one rule — with #711's exception carried along, since that record deliberately keeps its enumeration and is a counterexample rather than a supporting case.
Not filed, deliberately: overclaim/stale-claim, per-task review blindness, and omitted brief constraints. §4 argues no plausible mechanical detector exists for these, and filing an issue for each would produce exactly the weak enumerating guard this analysis recommends against.