# Recurring defect shapes across the closed-issue corpus (ersatztv#773) Root-cause analysis over every closed issue in the tracker, the detectors that would have caught each class, and an audit of the tooling already configured in this repo. This is an **analysis artifact**, not a rulebook: the rules it argues for land as decision records and follow-up issues, which are linked per class below. Measured 2026-08-13 against `origin/main` at `f9f8f65ce`. **The issue corpus is a snapshot drawn at 19:45 that day** — a boundary that matters, because #767 closed nineteen minutes later and is therefore absent from every count here (§3.7). --- ## 1. Method, and what these numbers can and cannot support **Corpus.** All 349 closed issues (`state=closed&type=issues`, 7 pages, 349 distinct numbers, #1–#757 — pull requests excluded). Of those, **95 carry a `## Closing record`** comment. Each record was classified by one of five independent agents against a written taxonomy, with `NEW:` available so the taxonomy could not be self-fulfilling. A sixth agent re-rated a stratified 15-issue sample blind, as an inter-rater control. A seventh sampled 25 of the 254 pre-convention closed issues from their raw bodies and comments, to test whether the shapes exist outside the era that writes records. **Three limits, stated up front because they bound every number below.** 1. **Record coverage is not uniform.** The `## Closing record` convention starts at #520 (2026-07-21). Coverage is 31/32 for #600–699 and 4/69 for #1–99. So frequencies over the 95 records are frequencies over *recent* practice. 2. **Corpus composition shifts underneath the measurement.** The share of records describing no process failure at all falls from 12/19 (#60–#438) to 0/19 (#616–#671). Early issues are feature work; recent issues are largely CI/process hardening. Guard-shaped defects therefore dominate the recent records *partly by construction* — we have lately been building guards, so our defects are in guards. 3. **A closing record is the closing author's self-report.** Raters were told to flag records whose own claims outrun what they describe, and several did (#60, #651, #691, #720). But the corpus systematically cannot contain a defect nobody noticed. **What limit 2 does *not* undermine.** The obvious worry — that these classes are an artifact of recent guard-building — was tested and **refuted**. The backward sample found high-confidence instances of the same shapes well before the convention existed: twin-missed at #215 (REST playout mutations skip the build-lock gating the Blazor path had) and #403 (an unhandled enum silent at 5 of 6 dispatch sites); wrapper drift at #198 and #287 (OpenAPI spec diverged from runtime JSON); masking guards at #232, whose SPA `PENDING_GRACE_TICKS` "existed only to paper over the resulting dishonest 200", and #234, where a blind `finally`-unlock was "papering over the stranded lock, so the two had to land together"; overclaim at #1 ("the earlier 'confirmed' was based on one curl test and several unverified assertions"). The convention changed the **density and legibility** of the evidence, not the existence of the shapes. Classification of pre-convention issues is harder but mostly possible — of the five the backward rater flagged as unclassifiable, four (#94, #172, #265, #377) do carry a closing comment with a clear root cause on re-inspection, and only **#157** (an unresolved flaky-test report with a single comment) is genuinely evidence-starved. The `insufficient-evidence` verdict is therefore rarer than first reported, which if anything strengthens the backward result. --- ## 2. What the corpus actually says 95 records; **69 describe a process failure**, 26 describe none. Percentages are of the 69. | Rank | Class | n | % | Shape | |---|---|---|---|---| | 1 | **Vacuous verification** | 17 | 25% | A check was green having done no work | | 2 | **Twin-missed** | 14 | 20% | The fix hit one instance and missed its structural twin | | 3 | **Symptom-keyed guard** | 9 | 13% | The guard keyed on the symptom seen, not the defect's mechanism | | 4 | *(new classes — see §3)* | 8 | 12% | Proposed by raters as fitting nothing above | | 5 | **Environment divergence** | 7 | 10% | Behaviour differs by interpreter/provider/hardware, untested there | | 6 | **Wrapper drift** | 5 | 7% | A hand-maintained mirror drifted from the contract it mirrors | | 7 | **Overclaim / stale claim** | 3 | 4% | An assertion stronger than its evidence, or since gone false | | 7 | **String-predicate churn** | 3 | 4% | A regex/grep predicate needing repeated rounds — but see §3.7: the *rounds* are cross-cutting, only the *substrate* is string-shaped | | 9 | **Identity-not-capability** | 2 | 3% | Authorization scoped by *who*, not by *what is being changed* | | 10 | **Masking guards** | 1 | 1% | Two guards on one condition; one hid the other's total failure | **This reorders the issue's own ranking.** #773 put twin-missed first (9 instances) and vacuous verification second (8), from a 50-issue sample. Over the full record corpus they swap: **vacuous verification is the most common single shape.** Overclaim also drops sharply as a *primary* cause (4%) — raters overwhelmingly assigned it as a *secondary*. It is better understood as a modifier that rides on another failure than as a class with its own detector. **Inter-rater control (measured, not asserted).** A sixth rater re-classified a stratified 15-issue sample blind. **Raw agreement 12/15 (80%)** across an 11-category scheme. The three disagreements are informative rather than noise: - **#177** (twin-missed → no-failure) and **#521** (no-failure → twin-missed) disagree at the "did a process failure occur" boundary and point in **opposite directions**, so that boundary is noisy but not systematically biased. The 26 clean-record count should be read as ±3. - **#649** was rated twin-missed by one and wrapper-drift by the other — which is precisely the blur that Family C in §3 exists to resolve. Both raters named the other class as secondary. The top-three separation is far larger than the disagreement. Treat ±1 rank as noise. --- ## 3. Consolidation: four families, one of which the issue's taxonomy missed ### Family A — reasoning about a representative instead of the population **#773's central hypothesis, confirmed, and it extends further than proposed.** The issue asked whether twin-missed and vacuous verification are the same underlying error. They are — and symptom-keyed guards belong with them. All three are the same mistake at different targets: - **twin-missed** — the *fix* was applied to a sample of the population, - **vacuous-by-sampling** — the *verification* sampled an empty or unrepresentative subset, - **symptom-keyed guard** — the *guard* matched only the member that was noticed. **Size: 27 of 69 (39%)** — 9 twin-missed, 9 symptom-keyed, 9 vacuous-by-sampling. See §3.6 for the full partition; the number is smaller than the naive merge because records that belong to Families C and D are **not** double-counted here. The merge is not merely rhetorical; it predicts a single detector, and the corpus already contains that detector, reinvented several times without anyone noticing it was the same rule: > "the by-id handler covered 4 of 10 media types… **Sweep by FIELD, not by the call site the issue > names**" — #671 > "#644's guard keyed on the *symptom* (an inflated pageSize) rather than the *defect*… an at-cap > request was structurally invisible to it" — #650 > "filtered on tools that *already* declared `pageNum`, so a tool wrapping a paged endpoint with no > paging args escaped it entirely" — #616 > "the pre-existing test filters `Where(t => t.QueryParameters is {Count: > 0})`, so a tool that > lost its query parameters escaped it entirely" — #757 Four sessions, four reinventions of one rule → **detector A** in §4. **Scope limit, because it bounds the highest-value detector in this document.** The distinction is what kind of thing the population is made of. - **A population of *values*** — enum members, OpenAPI operations, registered tools — always has an external authoritative source, and detector A applies directly. `Enum.GetValues()` is that source for an enum, which is precisely #503's fix for a catch-all `default` that silently absorbed an unhandled value. - **A population of *sites in code*** — the places that dispatch on a value, rather than the values themselves — has no external list to assert against. #403 is the case: `PlaybackOrder` is an enum, so its *values* are enumerable, but the defect was that 5 of 6 **dispatch sites** failed to handle one, and no artifact anywhere enumerates those sites. That residue needs find-all-references tooling, which is broken here and tracked in #777, not a set-equality assertion. ### Family B — the check never ran at all **Size: 9 of 69 (13%)** — the 8 vacuous records that are not sampling errors, plus the one masking-guards record. Here the check was *dead*, not mis-aimed. `scripts/tests/` invoked by no CI job (#631); a validator whose call site could be deleted with the suite still green (#621); a `${{ }}` in a shell **comment** silently dropping a step while the job reported success in 6s (#751); `if ! cmd; then status=$?` reading bash's logical negation, so a failing spec run exited 0 (#445); new pre-push logic never wired to receive stdin, "dead code that every unit test still passes over" (#719). A sampling detector does nothing here. This needs **detector B**: proof the guard can go red. ### Family C — two copies of one thing **Size: 9 of 69 (13%)** — the 5 wrapper-drift records plus the 4 twin-missed records whose twin is literally a second copy (#510, #649, #711, #756). Those 4 are counted **here and not in Family A**. `.claude/` and its byte-identical `.codex/` mirror where only one was in the `PROTECTED` list (#711); an advisory local hook and the server-enforced workflow where four rounds of hardening landed on the copy with the lower stakes (#649); a hand-maintained MCP schema drifting from the generated OpenAPI DTO by one field and silently clearing it (#754, again #757); a hand-rolled frontmatter parser accepting YAML that PyYAML rejects (#674); a dropped-step guard added to one required workflow and not its fail-open twin (#756). The preferred fix is **deletion** — one implementation, called from both places, as #649 did by extracting `scripts/pr-changed-files.sh` and recording `ci.shared-pr-file-enumeration`. (#649 also shows why the duplication is dangerous: before the extraction the guard was "safe only by redundancy," and a failing exit status "previously left the *entire* suite green.") **But dedup is a preference, not a law, and #711 is the counterexample that proves it.** That record deliberately *kept* its enumerative list: *"the list stays ENUMERATIVE rather than derived — a derived rule would have to be evaluated against the very file list being classified, putting more moving parts inside a security predicate to save one line per new tooling directory."* Deriving a security predicate from the input it judges is worse than maintaining two entries. So: dedup by construction where the duplication is incidental; keep the enumeration where deriving it would feed the judged input back into the judge. ### Family D — check-and-use race over mutable state *(the taxonomy missed this entirely)* **Size: 5 of 69 (7%).** Three raters proposed overlapping new classes — *unprompted by the taxonomy*, which had no bucket for this, though they shared a corpus and a brief and so were not strictly independent. The stronger evidence is the **blind** re-rater, who saw none of their output and proposed the same class again for #536 under its own name (`toctou-partial-atomicity` vs `half-atomic-toctou`). Four of the five (#536, #622, #706, #707) were tagged `NEW` by their raters. **#632 was rated twin-missed as its primary**, with the race named as its secondary; it sits here because the mechanism is the race, and it is counted here rather than in Family A. | # | The race | |---|---| | 536 | Check-then-act split by an `await`; the write side was atomic, the read side a stale `Volatile.Read` | | 622 | Merge consent bound to a head sha, then Gitea's async auto-merge evaluated against a later head | | 632 | Head sha bound, but retargeting the PR base changes the diff while moving neither sha nor status | | 706 | Status writes are read-then-write with no compare-and-set; an older run can finish last and win | | 707 | Paged file enumeration diffs each page against the base's *live* tip, so a mid-paging advance drops rows | **Five instances — larger than identity-not-capability, string-predicate churn and overclaim individually**, and #773's taxonomy had no bucket for it. The unifying property: a check and the action it authorizes are separated in time over state that can change in between, with nothing pinning a version. #622's own record puts it exactly: *"The gate was never bypassed — it was satisfied against a snapshot that stops being true."* ### Two further cross-era classes, from the pre-convention sample - **Fail-open by default** — a surface that defaults to permissive when config is absent or a step vanishes: API writes open when `Api:WriteKey` is unset (#280); a dropped step is fail-CLOSED in `review-verdict.yml` but fail-OPEN in `docker-build.yml` (#751, #756, #768); a jq-1.6-inert classifier (#647). Distinct from identity-not-capability: nothing is mis-scoped, the *default* is wrong. - **Destructive replace** — a full-replace path that silently destroys state a reconcile-by-id would have kept: a schedule PUT resetting fill-group progression, which #252 fixed with a "scoped positional/no-op reconcile" so that a no-op PUT-back stops being destructive; a full-replace wrapper missing one field and clearing it (#754); a dedup fix that turned a duplicate row into permanent data loss because the remove filter used a different key (#500). ### 3.6 The partition, stated exactly The families overlap conceptually — a twin-missed record can be read as "another member of a population" or as "a second copy" — so membership is given as an explicit partition, **mechanically verified to cover all 69 records exactly once**. Read a family's n as a disjoint count, not a tally of everything the family's description could fit: Family letters A–E are the ones §4 names detectors for, and those letters correspond. **This table assigns no letters below E.** §4's detectors F and G address fail-open and destructive-replace, which are cross-era classes drawn from the pre-convention sample and are *not* members of this 69-record partition — so an F or G in §4 refers to nothing in this table. | Family | n | % | Composition | |---|---|---|---| | **A** population reasoned about via a sample | 27 | 39% | 9 twin-missed + 9 symptom-keyed + 9 vacuous-by-sampling | | **B** the check never ran | 9 | 13% | 8 vacuous-by-non-execution + 1 masking-guards | | **C** two copies of one thing | 9 | 13% | 5 wrapper-drift + 4 twin-missed-by-duplication | | **E** environment divergence | 7 | 10% | | | **D** check-and-use race over mutable state | 5 | 7% | 4 rated `NEW` + #632 rated twin-missed with the race secondary | | — overclaim / stale claim | 3 | 4% | no detector — see §4 | | — string-predicate churn | 3 | 4% | no detector proposed; see §3.7 — the round-churn it is named for is cross-cutting | | — identity-not-capability | 2 | 3% | | | — unmerged singletons | 4 | 6% | #60, #503, #586, #688 — real, but one instance each | | **Total** | **69** | — | the **n** column is exact; percentages are rounded to the nearest point and sum to 99% | The partition is checked mechanically rather than by eye: the assertion is that the union of the families equals the set of process-failure records, with no overlap. That check immediately caught a transcription slip (#438 omitted from the control set) — a small live demonstration of detector A applied to this analysis's own numbers. Two caveats travel with these percentages. The Family A/C boundary is a judgment about whether a twin is "a second copy" or "another member of a population", and #649 is exactly the record the two raters split on. And per §1 limit 2, these are proportions **over a recent, guard-heavy corpus** — the backward sample establishes that the shapes existed before, not that they occurred in these ratios. ### 3.7 Round-churn is a property, not a class — and the fix text is the fourth sampling target The ranked table lists **string-predicate churn** as a class, inherited from #773's own taxonomy ("a parser / string-matching predicate needing repeated rounds"). That name identifies a *substrate* when the evidence identifies a *cross-cutting property*. Counting records that narrate three or more review rounds: | Family | n | records with ≥3 rounds | |---|---|---| | A — population via a sample | 10 | #460, #496, #616, #633, #644, #650, #671, #684, #726, #743 | | B — the check never ran | 5 | #445, #620, #631, #685, #751 | | C — two copies of one thing | 5 | #510, #649, #754, #756, #757 | | E — environment divergence | 5 | #491, #643, #647, #648, #668 | | D — check-and-use race | 3 | #622, #632, #706 | | identity-not-capability | 2 | #697, #698 | | **string-predicate churn** | 2 | #578, #629 | | overclaim | 1 | #651 | | **Total** | **33 of 69** | every family represented | *Method: every one of the 69 records was read for round language — 48 mention "round" at all, 33 state an explicit count of three or more. Two attributions are loose: #647 narrates four rounds that happened on #643, and #629 counts a round it deliberately did not run.* **The class named after the phenomenon holds 2 of the 33.** Round-churn should therefore be read the way §2 reads overclaim — a modifier riding on another failure, not a class with its own detector. #697 settles it independently: three BLOCKED rounds on credential scoping, no parser anywhere near it; #698 ran six. One honest qualification against over-correcting. String predicates *are* disproportionately round-prone — 2 of that class's 3 records, against 10 of Family A's 27. The naming error is not that the association is false; it is that the class was defined by the substrate where the property was noticed, so 31 instances outside it had nowhere to be counted. **Why the rounds happen, which is the part that generalises.** §3 Family A names three things that get sampled instead of enumerated — the fix, the verification, the guard. There is a fourth: **the review scope.** A re-review briefed to "check the reported finding" samples; the population is the whole changed artifact, and the text written to *fix* the last round is unreviewed by construction at the moment it is written. That is why each round's defect lands in the newest prose rather than in the text under review. **Two limits on this section, since it is the one making a claim about its own production.** - **33 is a floor, not a total.** It counts records that *state* a round count; a record that ran four rounds without narrating them is invisible here. Do not read the 36 remaining records as single-round work. - **Counting this by pattern rather than by enumeration under-reports it by roughly half.** A regex over the corpus finds 18 of these 33 — the misses narrate their rounds in the same words as the hits. If you re-derive this number, read all 69. - **#767 — the case #773 describes as taking eight review rounds — is excluded by snapshot timing, not by absence.** The corpus was drawn at 19:45 on 2026-08-13; #767 closed at 20:04, nineteen minutes later, and now carries a full `## Closing record`. So the worst instance of the phenomenon is missing from every count in this document, including the 33 above. It is named here rather than quietly folded in, because re-drawing the corpus would move every denominator in §2 and §3.6 and the snapshot boundary has to sit somewhere. Reading its record changes nothing structurally and confirms Family A twice over — it reinvents detector A independently, for the fifth and sixth times in this repo: *"Refuting two variants of a channel is not clearing the channel"* (the working attacks through that channel were never tried), and *"Enumerating **shapes** of a command loses. Nine disarms across two rounds; running the command settles them together."* That last sentence is detector A in one line — stop sampling the shapes you thought of, execute the population. --- ## 4. Detectors Ranked by instances covered per unit of build cost. "Rule status" distinguishes a **docs problem** (no rule exists) from a **hooks problem** (the rule exists and nothing enforces it) — they have different fixes, and conflating them is why several of these recurred. | # | Detector | Covers | Rule status | |---|---|---|---| | **A** | **A guard derives its expected set from the authoritative source and asserts set equality — never `filter`, never a sample.** The population comes from the enum / the OpenAPI doc / the workflow YAML / the provider list, and the assertion is equality against it. A filter over the population cannot see the member that is missing. | Family A: 27 records (39%), minus the code-structural residue noted in §3 | **rule-missing** as a general rule; instantiated ad hoc in #616, #644→#650, #671, #757 | | **B** | **Every guard ships with a proof it can go red: delete or disarm *that guard alone* and the suite must fail.** Not "a test exists" — a mutation. | Family B: 9 records (13%) | **rule-present-unenforced** — stated in #685's record and in project memory, enforced nowhere | | **C** | **Dedup by construction.** When two copies of one rule exist, delete one and have both callers invoke it — *unless* deriving the list would feed the judged input back into the judge, which is #711's reasoned exception. | Family C: 9 records (13%) | **rule-missing** | | **D** | **Anything read-then-written against live Gitea/remote state pins a version or uses compare-and-set.** See the honesty note below — this is a fix pattern, and its *detector* is detector A applied to an enumerated inventory of such sites. | Family D: 5 records (7%) | **rule-missing** (partially addressed for the merge path by the per-sha required check from #622) | | **E** | **Run the check under the interpreter/provider it will actually run under**, and preflight-log the version. | Environment divergence: 7 records | **rule-present** and working — `scripts/jq-preflight.sh` + `ci.jq-version-contract` closed the jq axis after #643/#647/#648. **Unclosed axes**: SQLite-vs-MySQL query semantics (#668), GPU generation (#505), CI-VM speed (#512) | | **F** | **Test the DENY path with the production config value.** #756's lesson generalised: a fixture that omits a field tests only the default, so a fail-open in the production value stays invisible. Parametrise the whole matrix, including "config absent". | Fail-open-by-default: ~5 records | **rule-missing** | | **G** | **Full-replace endpoints assert their complete field list in a test**, and prefer reconcile-by-id over delete-and-reinsert where child state exists. | Destructive replace: ~3 records | **rule-missing** | **Detector A is the highest-value single change in this analysis** — one convention, 39% of the recorded process failures over this corpus, and it is already proven four times in this repo under four different names. Two honesty notes on that ranking, because it drives #774's priority: - **The 39% is measured over a recent, guard-heavy corpus** (§1 limit 2). The backward sample shows the shape predates that corpus; it does not show the *proportion* holds across eras. Read it as "the largest family in the work we have been doing lately," which is still the right basis for prioritising the next change, but not as a timeless property of the project. - **Detector A also applies to review scope, which is free.** Per §3.7, a re-review briefed to "check the reported finding" samples the artifact. Briefing it to sweep the whole changed artifact enumerates it. This costs one sentence in a review brief and is the only detector here with no build step at all. - **Detector D is weaker than its neighbours in this table and is listed anyway.** A, B, F and G each name a check that fails when violated. D names a *fix pattern* with no general lint — you cannot mechanically spot "this code should have pinned a sha." What makes it actionable is that the population is small and enumerable: the handful of scripts that touch live remote state. So D's real detector is detector A applied to that inventory, which is why #778's scope is "enumerate every such site and mark each pinned / CAS / knowingly-unsafe" rather than "write a linter." ### Classes where no mechanical detector is plausible — stated rather than papered over The issue explicitly asked for this, and inventing a weak detector here would itself be the symptom-keyed-guard mistake. - **Overclaim / stale claim.** No check can tell that a sentence is stronger than the evidence behind it. Partial mitigations exist and should not be oversold: `stale-after` frontmatter (#603) dates a claim, and #578's retracted-term grep catches a *specific known* retraction propagating into generated artifacts. Neither detects a fresh overclaim. This stays a review responsibility. - **Per-task review blindness** (#60). A review scoped to one task's diff structurally cannot see a defect that only exists once several tasks compose — five Important findings at #60 were invisible to six per-task reviews and surfaced only in a whole-branch pass. The fix is a *process* step (a whole-branch review before close), not a check. - **Omitted brief constraint** (#586). A delegated brief silent on a hazard gets a plausible-but-wrong default — "an omitted rule isn't an unenforced rule, it's a rule replaced by whatever default the agent reaches for." A brief lint is conceivable but would be a keyword matcher, i.e. exactly the string-predicate class. A hazards checklist in the brief template is the honest ceiling. - **Guard-parity by verb** (#458), and re-deriving an inherited exemption against a new failure mode (#484). Both need judgment about whether a prior rationale still applies. ### A meta-finding about this repo's own knowledge base `docs/decisions/records/` holds **189 active records**, and the two most relevant to this analysis (`testing.enumerating-guard-identity-not-position`, `ci.required-job-step-execution-markers`) are extraordinarily detailed — each a full account of one incident. The corpus grows **one record per instance**. That is Family A operating on our own process: we are enumerating cases rather than removing the mechanism. Detector A, C and B are class-level rules precisely because the per-instance record has already been tried 189 times. --- ## 5. Part 2 — audit of the tooling already configured Judged against the measured classes above, per the issue's sequencing. Everything here was run, not assumed. ### 5.1 The LSPs: two of three are broken, and the one that matters most is the most broken #773 asked whether routine LSP use would catch Family A (find *all* references rather than the one in front of you). It would help — **and it is not available.** Measured in this session: | LSP | State | Evidence | |---|---|---| | `csharp-lsp` | **Broken — cannot initialize** | `findReferences` on `ChannelPlaylist.ToM3U()` → `System.InvalidOperationException: .NET SDK cannot be resolved, because libhostfxr.dylib cannot be found inside /opt/homebrew/Cellar/dotnet/10.0.302/bin/host/fxr`. That directory **does not exist** — Homebrew's dotnet layout is not what MSBuildLocator expects. `dotnet --version` works (10.0.302), so builds are fine; only the language server is dead. | | `typescript-lsp` | **Broken at the repo root** | `Could not find a valid TypeScript installation… ensure that the "typescript" dependency is installed in the workspace`. `typescript` lives in `web/node_modules`, not at the workspace root. Configuration problem, not a missing dependency. | | `pyright-lsp` | **Works** | `documentSymbol` on `scripts/decisions_lib.py` returned the full symbol tree; `findReferences` on `active_files` correctly returned 4 references across 3 files, including cross-file hits in `decisions_validate.py` and `migrate_decisions_split.py`. | **Two conclusions, and the second is the sharper one.** 1. The LSP that would help most is the one that is dead. The residue detector A cannot reach — populations of *sites in code* rather than of values — is overwhelmingly C#: #403 (5 of 6 dispatch sites) and #671 (a by-id handler covering 4 of 10 media types) are "find every site" problems, and `csharp-lsp` cannot answer a single query. (#510 is *not* an example here despite looking like one: its record names duplication as the root cause and its switch keys on a plain enum, so it belongs to Family C and `Enum.GetValues()` is its authoritative source.) 2. **A dispatched subagent could not reach the LSP tool at all.** The agent tasked with testing the three LSPs reported `ToolSearch` returning "No matching deferred tools found" for every query, while the same tool resolved immediately in this main session. So the standing note *"workflow agents must use csharp-lsp"* is doubly rotten: the server is broken **and** the agents it addresses cannot invoke it even when it works. This is a live instance of overclaim/stale claim (§2 rank 7) sitting in our own guidance. **Resolved 2026-08-14 (#777), and one claim above needs correcting.** Both servers now work and both answer a real cross-file query in this repo; the full setup, traps and verification are in `docs/local-lsp-tooling.md`, and the rule is `session.local-code-intelligence`. | Row above | Resolution | |---|---| | `csharp-lsp` | `MSBuildLocator` needs a dotnet root that owns `host/fxr`, which Homebrew's `bin` does not and its `libexec` does. Fixed by `env.DOTNET_ROOT` in `.claude/settings.local.json`. `findReferences` on `ChannelPlaylist.ToM3U()` returns the declaration plus its 5 call sites, excluding the mention of the name in a comment that `grep` matches. | | `typescript-lsp` | No configuration lever exists — v5 dropped `--tsserver-path`, and a plugin `lspServers` entry cannot pass `initializationOptions`. Fixed by making the package resolvable from the workspace root (a gitignored root `node_modules/typescript` link). Returns 20 references across 7 files for `canLeaveCurrentScreen`. | **The correction to conclusion 2**, which matters because the guidance it judges is still in use: the note *"workflow agents must use csharp-lsp"* names the **`csharp-lsp` MCP server** tools (`csharp_set_workspace`, `csharp_diagnostics`, `csharp_references`, …), not the `LSP` tool. MCP tools **are** reachable from a subagent. The subagent measurement above is correct about the `LSP` tool and was generalised one step too far: the note was unsatisfiable because that MCP server's `.mcp.json` entry named a dotnet install that no longer existed, so it never started — not because agents cannot invoke it. With the entry repaired the server serves 16 tools, so the guidance becomes satisfiable — inferred from the MCP boundary generally (326 subagent MCP calls across four other servers), not yet measured for `csharp-lsp` from a subagent. What must be briefed explicitly is **which surface**: pointing a subagent at the `LSP` tool is still an instruction it cannot obey. ### 5.2 Python: the global instruction and this repo disagree, and the repo is silent `~/.claude/CLAUDE.md` instructs `ruff check`, `ruff format --check`, and `pyright` after modifying Python. Measured against reality: | Claim | Measured | |---|---| | ruff/pyright run in CI or hooks | **No.** `grep -rn "ruff\|pyright" .gitea/workflows/ .husky/` returns **nothing**. | | What *does* run on Python | `pr-checks.yml` → `decisions-guard` (`decisions_validate.py`, `build_decisions_catalog.py --check`) and `script-tests` (`pytest scripts/tests`). Neither Husky hook touches Python. | | `ruff check scripts/` | **47 errors** — but 28 are `E702` (semicolons) and **7 `S105` "hardcoded password" are false positives** on test stubs (`env["ETV_GITEA_TOKEN"] = "stub"`). The real cleanup is small. | | `ruff format --check scripts/` | **10 of 20 files would be reformatted.** | | `pyright scripts/` | **2 errors**, both `reportMissingImports` for `etv_client` in `scripts/scripted-schedules/entrypoint.py` — a package resolvable only in that script's deploy environment. Effectively clean. | | Repo-level ruff config | **None.** No `ruff.toml`/`pyproject.toml`. Ruff silently falls back to whichever `~/.config/ruff/ruff.toml` the operator's machine happens to have. | **The last row is the finding that matters**, and it is Family E (environment divergence) in our own toolchain: Python lint behaviour here is a function of an un-versioned file on one laptop. A second machine lints differently, or not at all. One of the two must move — either the repo adopts a committed ruff config and enforces it, or the global instruction stops claiming this repo enforces something it does not. **Resolved 2026-08-21 (#780): the repo moved.** `ruff.toml` is committed at the root and the `script-tests` job runs `ruff check` + `ruff format --check` under a pinned `ruff==0.12.11` over an explicit population from `git ls-files` — not `ruff check .`, which an `exclude` in the right config scope silently empties into a green run. The pre-fix state, reproducible rather than asserted — against `706674272`, the base this landed on, with the committed config dropped in: ``` mkdir -p ~/scratch/m780 && git archive 706674272 | tar -x -C ~/scratch/m780 git show 01f7a89e8:ruff.toml > ~/scratch/m780/ruff.toml # a sha: the file is not on main pre-merge cd ~/scratch/m780 && ruff check . ; ruff format --check . # ruff 0.12.11 # -> Found 74 errors. / 20 files would be reformatted, 13 files already formatted ``` (`;` not `&&` — `ruff check` exits 1, which would swallow the second command. Not `/tmp`: macOS purges it. Redirect the config into place *before* running anything: an empty `ruff.toml` is valid, so a failed `git show` leaves ruff silently using its own defaults and printing a different number.) The row above measured **47** eight days earlier against the operator's global config; the tree grew and the configs differ, so the two numbers are not comparable and neither supersedes the other. Two of the 74 are `RUF100` on suppressions that were already in the tree before this change — they exist in this count only because the committed config enables that rule. Of the 74, **57 were fixed in code** (most of them by the `ruff format` pass itself, which splits the 40 semicolon statements) and **17 carry a per-site `# noqa` with its reason inline**: 8 `S105` on stub credentials handed to the real hooks by `scripts/tests`, 9 `E501` on one-line JSON and shell fixtures. The `S105`s are deliberately per-site rather than a directory exemption, so a real credential pasted into a fixture later still reddens the gate. Only `S101` is exempted directory-wide for `scripts/tests/**`, because a test suite asserts. `RUF100` is selected, which is what keeps that split honest: a `# noqa` that suppresses nothing still reads as a suppression, and it is invisible without this rule. Three were live when it was switched on: one on a site that had already been fixed in code, plus the two counted above — one whose rule had stopped firing, one for a rule this config never enables. `pyright` stayed ungated: its only findings are the `etv_client` imports in the row above, and gating it would put a node toolchain on the git-only `small` lane to find nothing. Rationale, the exemption list and the measured exclude matrix: `ci.python-lint-ruff-config-committed`. ### 5.3 Configured vs actually invoked **Measured 2026-08-28 (#781/#799).** Population derived from config, not listed by hand: `~/.claude/settings.json` `enabledPlugins`, `~/.claude.json` `mcpServers` (plus its per-project `projects[].mcpServers`), the project `.mcp.json`, and `.claude/settings.local.json`. Corpus: **486 transcripts** under `~/.claude/projects/` via `rglob` (a top-level glob sees 209 and manufactures false zeros), this session excluded. Positive control on the same query shape: 22,927 `Bash`, 1,420 `Read`, 1,116 `Edit`. **A count is only meaningful against the surface a thing is actually used through**, so each row below names its surface. Do not reason from a surface *count* — enumerate the event shapes and say which you measured, because the whole failure here is an audit that checked one and reported zeros for the rest. Measured: MCP tool calls (`mcp____*`), `Skill` invocations, `Agent` `subagent_type`, user-typed slash commands (``), `LSP` tool calls, hooks (which appear as `hookInfos` inside a `stop_hook_summary` record and *never* as a tool call), and CLIs invoked inside `Bash`. Measured at zero everywhere: MCP **resource** tools (`ListMcpResourcesTool`, `ReadMcpResourceTool`). **Not** separately identified: MCP prompts, and the `attributionSkill` records that also appear in transcripts — so a zero below is a zero on the shapes listed, not on every shape that exists. | Entry | Scope / surface | Measured | Verdict | |---|---|---|---| | `gitea` MCP | user `~/.claude.json` | **975 calls**, 105 transcripts, last 2026-08-28 | Keep — heaviest tool in the corpus | | `codex` plugin | `codex exec` inside `Bash`, plus `SessionStart`/`SessionEnd`/`Stop` hooks | **109–126 `Bash` inputs invoke it at command position** across ~21–26 transcripts, last 2026-08-28 — 109 under a quote- and heredoc-aware parse, rising toward 126 if wrappers such as `nohup` are admitted. The range is stated rather than resolved: three successive parsers gave three answers, which is the point at which a hand-written command-position grammar stops being worth another round. **258** inputs merely *contain* the string — `pgrep -f "codex exec"` polling loops and heredoc bodies mention it — so a substring count overstates it roughly twofold | Keep — heavily used, and invisible to a skill-name search | | `playwright` | plugin MCP | 64 calls, last 2026-08-27 | Keep | | `security-guidance` | **hooks** (`SessionStart`, `UserPromptSubmit`, `PostToolUse`, `Stop`) | **1,086 execution records** in `hookInfos` naming `security_reminder_hook`, across 78 transcripts / 23 UTC days, last 2026-08-28 | Keep — **not** dead; it exposes no tool at all | | `ssh-mcp` | project `.mcp.json` | 16 calls, 9 transcripts, last 2026-08-22 | Keep | | `superpowers` | `Skill` | 13 invocations (`systematic-debugging` 6, `brainstorming` 5, …), last 2026-08-26 | Keep | | `mempalace` | user `~/.claude.json` | 11 calls, 7 transcripts, last 2026-08-27 | Keep — in active use | | `feature-dev` | **`Agent` subagent_type** | `feature-dev:code-reviewer` ×2, last 2026-08-21 | Keep — **not** dead; used via agents, not skills | | `serena` | plugin MCP | **Disabled until 2026-08-14T12:31Z**; 0 invocations in the 14 days since it was enabled; **demonstrated working 2026-08-28** | Keep — see #799 and `docs/local-lsp-tooling.md` | | `LSP` tool (`csharp-lsp`, `pyright-lsp`, `typescript-lsp` plugins) | `LSP` | 13 calls, last 2026-08-14 — all §5.1/#777 probes | Repaired but unadopted; no post-repair use | | `csharp-lsp` MCP server | project `.mcp.json` | **0** since it became startable on 2026-08-14 | Reachable, unadopted | | `ralph-loop` | **hooks** (`Stop`) + slash command | Its `Stop` hook ran **1,086 times** across 78 transcripts / 23 days — the same triple as `security-guidance` at this snapshot, consistent with both registering a `Stop` hook in their `hooks.json`. The `/ralph-loop` command itself: **0** | Keep — the hook executes and exits when no loop state exists; the *feature* is unused, the plugin is not inert | | `context7` | plugin MCP | **0**, and it was enabled across the whole retained window | No observed use in the window. **Deliberately kept** — operator decision 2026-08-28 | | `code-review`, `code-simplifier`, `commit-commands`, `claude-md-management`, `skill-creator`, `frontend-design` | `Skill` / `Agent` / slash command | **0** — but all six were **disabled** until 2026-08-14T12:41Z, so the zero covers only the 14 days since | Not yet exercised rather than rejected. **Deliberately kept** — operator decision 2026-08-28 | | `nuget` | project `.mcp.json`, explicitly disabled | 0 (expected); `command` names `/usr/local/share/dotnet/dotnet`, **verified absent** | Entry removed 2026-08-28 | | `docker-mcp` | project `.mcp.json`, explicitly disabled | 0 (expected) | Correctly off | | `homeassistant` | `~/.claude.json` `projects[/Users/timothy/server-management]` | 129 calls | Not in this repo's population — scoped to another project | `claude-in-chrome` also appears in the corpus (154 calls) and is **not** configured anywhere in these files: it is harness-provided. An audit deriving its population from config alone will not see it. **Why "retire what is enabled and never invoked" was not carried out.** #781 assumed a zero means useless. Across three re-measurements the zeros here split **six** ways, and none of them is grounds for removal until you know which one you are looking at: 1. **Broken** — `csharp-ls` and the `csharp-lsp` MCP server named a dotnet root this machine does not have (#777). Unreachable, not unwanted. 2. **Measured on the wrong surface** — `codex` (used through `Bash`), `security-guidance` and `ralph-loop` (hooks, which never appear as a tool call), `feature-dev` (agents, not skills). Four separate instances of one mistake, and each looked like a confident zero. 3. **Not enabled** — `serena` was `false` until 2026-08-14T12:31Z, and six of the plugins above until 12:41Z the same day. *Installed* is not *enabled*: `installedAt` says March for all of them, and only `enabledPlugins` says whether a thing could have been used at all. 4. **Enabled but never called** — `serena` since 2026-08-14, and the `csharp-lsp` MCP server since it became startable. This is the only shape that resembles the issue's premise. 5. **Evidence aged out** — the corpus is a rolling window: 964 transcripts on 2026-08-14, **486** on 2026-08-28, reaching back only to 2026-07-29. A zero can mean the calls expired, and an older measurement's counts are not reproducible from a later corpus. 6. **No observed use in the window** — `context7`. Note the wording: the corpus starts 2026-07-29 and these plugins date to March, so the evidence cannot distinguish "never used" from "not used lately", and it cannot support "excess" as a historical claim. Establishing *which* of the six applies is the precondition for acting, and doing so cost three sessions here. What was demonstrably wrong with this configuration was **duplication** and **breakage**; excess was never shown. **`serena` was disabled, and #799's observation was correct (#799).** The plugin was `false` in `~/.claude/settings.json` until **2026-08-14T12:31Z**, when a *concurrent* session enabled it — `~/.claude/settings.json.bak-doctor`, written moments earlier, still records `serena: false`. Its tools appear in **no transcript before 2026-08-14T12:42:54Z**, then in 6 transcripts that day and 186 across the 10 days since, always in a harness-emitted availability listing rather than prose (counts as of the 2026-08-28 snapshot; the corpus is live and grows). The #799 session had started at 12:01Z **and never reloaded**, so its tool set never contained them; its `ToolSearch` probe at 13:41Z correctly reported them absent, and that probe *did* carry a positive control — two other `ToolSearch` queries resolved in the same session minutes before. **Enabling a plugin did not reach a session that was already running; a reload did.** Both sessions started *before* the 12:31Z toggle, so start time alone separates nothing — what separates them is the reload. The enabling session ran `/reload-plugins` at **12:54:32.694Z**, printed `Reloaded: 17 plugins`, and its next tool delta added serena **19 seconds later at 12:54:51.402Z**. The #799 session ran for 3h54m, emitted exactly one tool delta (12:01:27Z, no serena), and never reloaded. Its author read `serena: true` from the settings file and `serena: absent` from their own tool set, and both were true simultaneously. Two sessions are not a proof of the harness's contract, so treat this as the operating rule rather than a mechanism: **do not assume a plugin a sibling session enabled is available to yours — reload or restart, then confirm the tool resolves.** Two things follow for any future audit. **Read enablement as of the evidence, not as of now** — an audit that compares *current* config against *historical* usage will call a plugin unused for a window in which it was switched off. And **date transcript evidence by record timestamp, never by file mtime**: mtime is when a session last wrote, which can be hours or days after the event being proved, and here it is what made a first-appearance look like a pre-existing state. The `activate_project` requirement is a separate, real trap, reachable only once the tools are present: without it the first query returns a "no active project" error naming the projects it does know. Both of #799's phrasings resolve today, and the demonstrated query is in `docs/local-lsp-tooling.md`. `.mcp.json` is **gitignored**, so none of its content is reviewable from this repo and the de-duplication recorded below happened on one machine only. Two scopes defined `gitea` (the user entry additionally carrying `"type": "stdio"`); the **project** copy was removed rather than the user one, because `server-management` and `homelab-docs` have no `.mcp.json` and depend on user scope. Verified afterwards by starting the surviving server: 54 tools. **Hooks are the good news.** All 13 scripts in `.claude/hooks/` are wired from either `.claude/settings.json` or `.husky/*`, and no settings entry points at a missing path — there are **no dead hook scripts**, contrary to the issue's suspicion. Husky hooks have genuine fired-output evidence in transcripts (`husky - dotnet format found whitespace/BOM issues`, `husky - commit message missing Co-Authored-By trailer`, `husky - refusing to commit root-level screenshot(s)`). Wiring was the strongest claim available for the hook rows when this table was written. It is no longer the ceiling — §5.4 is now measured — so read the hook rows against `scripts/hook-fire-log.sh report`, not against this paragraph. ### 5.4 Whether our own hooks fire — now measured (#776 closed this) **The finding as originally recorded.** In this harness version only `Stop` hooks emit a structured record (`stop_hook_summary`/`hookInfos`). `PreToolUse` and `PostToolUse` hooks — which is *every* guard that matters here: merge consent, worktree, BOM, agent model/RAM — leave no durable execution trace. What this audit could count for those hooks was filename mentions in settings dumps, which is not evidence of execution. So the guards this repo relies on most were exactly the ones whose execution could only be inferred: Family B (the check never ran) applied to the hook layer itself, after this project had already paid for it twice at the CI layer (#751, #756). **What replaced the inference.** Every hook now records its own execution through one shared sink, `scripts/hook-fire-log.sh` (`testing.hook-reports-its-own-execution`): a `fire` record on entry and an `exit` record carrying the status and the decision, where **the decision is parsed from the bytes the hook actually emitted** rather than declared by its author. Read it with `scripts/hook-fire-log.sh report [--all]`. **Measured 2026-08-14** — a scripted headless session (a plain Bash call, a Bash call carrying `ETV_UPDATE_GOLDENS=1`, a `Write` to a SPA file, and one `Agent` dispatch) plus a real `git commit` and a `git push --dry-run` in a worktree. Snapshot boundary: this is one deliberately-constructed window, not a corpus statistic. | Hook | Event | Fires | Decisions observed | |---|---|---|---| | `pretooluse-bash-guard` | PreToolUse/Bash | 2 | `deny` 1, `no-op` 1 | | `pretooluse-bom-guard` | PreToolUse/Bash | 2 | `no-op` 2 | | `pretooluse-worktree-guard` | PreToolUse/Bash | 2 | `no-op` 2 | | `posttooluse-worktree-marker` | PostToolUse/Bash | 1 | `no-op` 1 | | `design-sync-reminder` | PreToolUse/Write + Stop | 4 | `context` 1, `block` 1, `no-op` 2 | | `pretooluse-agent-model` | PreToolUse/Agent | 2 | `no-op` 2 | | `pretooluse-agent-ram` | PreToolUse/Agent | 2 | `no-op` 2 | | `decisions-guard` | git pre-commit | 2 | `pass` 2 | | `prepush-clean-worktree-check` | git pre-push | 1 | `pass` 1 | | `prepush-donewhen` | git pre-push | 1 | `pass` 1 | | `prepush-rebase-check` | git pre-push | 1 | `pass` 1 | | `pretooluse-nav-guard` | PreToolUse/navigate | **0** | not exercised — needs a live browser session | | `pretooluse-merge-consent` | PreToolUse/PR write | **0** | not exercised — needs a real merge attempt | **11 of 13 hooks are confirmed firing, with the decision each reached.** The `deny` row is the load- bearing one: `pretooluse-bash-guard` did not merely run, it *blocked* the `ETV_UPDATE_GOLDENS=1` probe, so at least one guard in this set is demonstrably live rather than merely present. **Re-verified against the shipped implementation.** The table was first measured against an early version of the sink, and the classifier changed materially afterwards, so the run was repeated against the final code: `pretooluse-bash-guard` `deny`+`no-op`, `pretooluse-bom-guard` and `pretooluse-worktree-guard` `no-op`, `posttooluse-worktree-marker` `no-op`, `design-sync-reminder` `context`+`block`+`no-op` — identical decisions. The repeat run exercised the five Claude hooks a Bash/Write session reaches; the `Agent` pair and the four git hooks are carried over from the original run and were not re-measured. **Reproducing this table.** It is a constructed window, not a corpus statistic, and it is not reproducible from a reader's own `report` output — running the test suite alone would not produce it, and before `scripts/tests/conftest.py` landed, running the suite actively polluted the default log with synthetic fires. To re-derive: point `ETV_HOOK_FIRE_LOG_DIR` at an empty directory, run a headless session exercising the four tool paths above, then a `git commit` and a `git push --dry-run`, and read `scripts/hook-fire-log.sh report --all --dir `. **A zero is two findings wearing one number** — a hook that is broken, and a hook whose trigger did not occur — and the report says so rather than presenting a zero as a verdict. For the two zeroes here, what is established and what is not, kept apart deliberately. **Established:** both script bodies work. `pretooluse-nav-guard` and `pretooluse-merge-consent` are each driven to their deciding branch in `scripts/tests/test_hook_fire_log.py` — a `deny` on an `/iptv/` URL and an `ask` on a merge call — emitting the correct decision with the instrumentation in place. **Not established:** that the harness would dispatch to them. Those tests invoke the scripts directly, so they bypass registration and matcher dispatch entirely; a typo in `.claude/settings.json`, a settings file that never loaded, or a matcher that does not match would leave both tests green while the production zero still meant broken wiring. A working script is a necessary condition, not the finding. So these two zeroes remain genuinely ambiguous, and neither can be resolved without the thing that resolves it: a live browser session, or a real merge attempt. Manufacturing a merge to observe the merge guard is a worse idea than the gap it would close. Distinguishing the two meanings of zero stays a human judgement — `testing.hook-reports-its-own-execution` says so, and treating "the script works" as "the wiring works" is the way that judgement goes wrong. **The measurement earned its keep on its first run.** `pretooluse-agent-model` was recorded deciding `no-op` on an `Agent` dispatch where `ask` was expected. Previously that would have been an unanswerable suspicion about a guard nobody could observe; instead the real payload was dumped and it carried `model: haiku`, so `no-op` was correct and there was no defect. The general form is the standing #756 lesson — **make the system report it rather than infer it** — and note what it displaced: an argument from the hook's source about what it "must" do, which is the reasoning shape §3 measures going wrong repeatedly. ### 5.5 Tools we do *not* have that would address a named class Each is justified against specific issues, not general merit; anything that could not be tied to recorded instances is left out rather than padded in. | Tool | Class it addresses | Justification | |---|---|---| | **Stryker.NET** (mutation testing for C#) | Family B — detector B, mechanised | Detector B currently relies on an author remembering to write a mutation proof. Stryker generates them. #621 (a guard whose call site could be deleted with the suite green), #685 (two guards where deleting either left the suite green) and #719 (logic never wired to stdin) are all surviving-mutant detections by construction. Cost is real (mutation runs are slow), so scope it to the guard/validator projects rather than the whole solution. | | ~~`shellcheck`~~ — **tested and rejected** | (would have been Family B, shell half) | The obvious candidate for #445 (`if ! cmd; then status=$?`, where bash sets `$?` to the *logical negation* so a failing run exited 0). **Measured: it does not catch it.** ShellCheck 0.11.0 on that exact construct reports nothing, and even `-o all` returns only an unrelated brace-style nit; run against our 13 hook scripts it yields 2 `SC2034` unused-variable warnings. Recorded here as a negative result so it is not proposed again on plausibility. | | **A duplicate-code detector** (`jscpd` or equivalent) | Family C | #440 shipped two byte-identical Lucene escapers in different files; #711 and #649 are the same shape at the config/CI layer. A duplication report would have surfaced all three at authoring time — subject to #711's exception, so it should advise, not block. | | **An enum-exhaustiveness analyzer rule** for C# `switch` *statements* | Family A, code-structural residue | #503's catch-all `default` silently absorbed an unhandled enum value and produced byte-identical output to a handled one. Switch *expressions* already warn; statements do not. This is a rule to enable in the analyzers we already run (Meziantou), not a new dependency — the cheapest item here. | **Explicitly not proposed:** a "brief linter" for omitted delegation constraints (#586) and any overclaim detector. §4 argues both are implausible, and inventing them here to look thorough would be the symptom-keyed mistake. Nothing in the corpus suggests a tool we lack would beat detector A for Family A — the four in-repo reinventions show the convention works when applied; the failure is that it was never written down once. **Evidence status of these rows, which differs and matters.** Only the shellcheck row was *executed* — and running it refuted the reason it had been proposed, which is why it is struck through. The three surviving rows are **reasoned from precedent**: each names issues whose recorded mechanism the tool addresses, but none has been run against this repo to confirm it fires. That is a weaker warrant and should be discharged before any of them is adopted — a tool justified by "it plausibly catches this class" is the same species of claim as the guards this document criticises. Treat these three as candidates to test, not findings. **On MemPalace**, the 19-day gap deserves care rather than a verdict: `CLAUDE.md` makes it the documented discovery path for decisions, so either sessions are skipping the documented path, or they are correctly reading `docs/decisions/README.md` directly (which the same contract permits, and which is authoritative). The measurement cannot distinguish those, and I am not going to guess. --- ## 6. Follow-ups This issue is analysis; it spawns implementation rather than doing it. | Issue | Detector / finding | Covers | Priority | |---|---|---|---| | #774 | **A** — a guard derives its population from the authoritative source and asserts set equality | Family A, 27 records (39%) | high — **CLOSED** | | #775 | **B** — every guard ships a mutation proof: delete that guard alone, see red | Family B, 9 records (13%) | high — **CLOSED** | | #776 | Make hooks report that they fired — PreToolUse/PostToolUse execution is unobservable | the whole hook layer | high — **CLOSED**, §5.4 is measured | | #777 | `csharp-lsp` and `typescript-lsp` are broken; the "workflow agents must use csharp-lsp" note is stale | partial mitigation for Family A | medium — **CLOSED**, §5.1 | | #778 | **D** — pin a version or use compare-and-set for read-then-write against live remote state | Family D, 5 records | medium — **CLOSED** | | #779 | **F + G** — test the deny path with the production config value; assert full-replace field lists | fail-open + destructive-replace, ~8 records | medium — **CLOSED** | | #780 | Reconcile Python tooling — commit a ruff config and enforce it, or stop claiming we do | environment divergence in our own toolchain | medium — **CLOSED**, §5.2 | | #781 | Retire enabled-but-never-invoked plugins/MCP servers; de-duplicate the `gitea` server | config that reads as coverage | low — **DONE**, §5.3 is re-measured from a derived population; the retirement half was answered *no* on evidence | | #799 | `serena` is enabled but never invoked — establish why, then use it or disable it | one row of §5.3 | low — **DONE**, it had been *disabled* until 2026-08-14 and is now adopted (`docs/local-lsp-tooling.md`) | **Detector C (dedup by construction) has no issue of its own on purpose.** It is not a thing to build; it is the shape the fixes in #774 and #778 should take when they find two copies of one rule — with #711's exception carried along, since that record deliberately keeps its enumeration and is a counterexample rather than a supporting case. Not filed, deliberately: overclaim/stale-claim, per-task review blindness, and omitted brief constraints. §4 argues no plausible mechanical detector exists for these, and filing an issue for each would produce exactly the weak enumerating guard this analysis recommends against.