Two independent cold reviews (Codex/GPT-5.6 cross-family, and a cold Claude reviewer in
its own worktree) converged on the same class: paths where "this job cannot establish
what is on the head" still resolved by leaving the head alone, which protects a real
verdict and leaves a forged one.
Behaviour:
1. The four page-2 completeness refusals now replace the unknown state too. They were
excluded on the reasoning that the probe fires when NO row for this context was on page
1, so there is no green of any provenance to leave standing — self-contradictory, since
the only reason page 2 is read is that the row may be beyond page 1, which the probe's
own message says. Accepted cost, stated in the record: a head with more CONTEXTS than
the 50-row cap stalls every run; measured 2026-08-29, this repo puts 8 on a `main` head,
and that case already stalled with an ABSENT check.
2. The no-mark downgrade covers every re-derivable write, not only `success`. Restricting
it analysed the wrong PR: the damaging case is one that IS exemptible and got the
generic `pending` only from a transient enumeration failure. That description carries no
marker, nothing verifies it without a mark, and the next run re-derives it into the
exemption with the human row below its own mark — route 2's damage through route 1's
condition. `$REPAIR_DESC` stays exempt, being stronger and not re-derivable.
3. The fence branch that cannot trust its retarget count while holding a derived `success`
writes the sentinel instead of abstaining. It is reached only after the classification
DECLINED to inherit the row the head carries, so posting nothing left that row current;
the message said the context "stays absent", true only of a head that had none.
4. Reconciliation needs a WITNESS: it may clear only over a complete history containing the
sentinel's own row. `ex_unverified` means the combined endpoint just returned that row
and `/statuses/{sha}` keeps one per POST, so a complete-but-empty history contradicts a
write that demonstrably happened — and `page_statuses` accepts an empty page 1 as
complete, which is what made it reachable. Both reviewers reproduced the clear-then-exempt
outcome. The shipped positive test used exactly that impossible fixture, so it was
pinning the defect; it now seeds the sentinel row, and an impossible-empty negative plus
a witness mutation proof were added.
5. The mid-run "did this row change" comparison now includes the row ID. The two sentinels
are byte-identical by design, so a mid-run replacement of one by another was invisible to
a state/creator/description triple. Measured 2026-08-29 (Gitea 1.27.1, head 736649b3):
the COMBINED endpoint carries `id` on every row, ids 14..30 ascending — the job had only
ever read ids from `/statuses/{sha}`. Where a server omits it both sides are empty and
the comparison degrades to the pre-existing text test.
6. The repair has a FLOOR — it may never write a description weaker than the one this run
decided — and is skipped when it would rewrite what is already there. Widening the gate
to every write meant a transient post-write read could rewrite a correct `$REPAIR_DESC`
carry-forward with the machine-clearable sentinel, reversing the ordering rule the
classification chain states.
Writing the sentinel and failing the job are separate decisions, which is why
`replace_unknown_state` and `replace_unknown_and_die` are two functions: the read refusals
were already non-zero exits on `main` and stay red; the fence branch exited 0 there and
still does, because an unreadable timeline is an ordinary hiccup and reddening every one is
noise this file elsewhere refuses to add.
Prose corrected where it now overclaimed: "the green never stands" after the post-POST
re-check is wrong — it is live between the POST and the repair, so the check makes a
permanent green TRANSIENT; "a later run reconciles this automatically" is wrong in the one
case where the replacement costs anything, since finding a masked verdict UPGRADES to the
human-only sentinel; and the mutation-proof framing claimed every mutant restores the exact
predecessor, when two do, one restores the shape #742 withdrew, and the rest disarm clauses
that have no predecessor. The quiet-timeline positive control now counts timeline walks,
because a single POST is also what a skipped re-check produces.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
docs/ — task-signal map
Purpose: route a fresh contributor/agent to the minimal set of docs for the task at hand, instead of a mandatory front-to-back read. Update this doc in the same PR that adds, removes, or retitles a doc below, or that changes which sections a task signal points to.
Start here, always
CLAUDE.md(repo root) — project intro: architecture, layout, dev commands, conventions, Task Completion Protocol.docs/contributing.md— established code patterns (CQRS/MediatR, LanguageExt, the ChicoryTV SPA, EF Core dual-provider migrations, FFmpeg pipeline, analyzers, testing). Read before any non-trivial change.
Task signal → minimal sections
| Signal | Read |
|---|---|
| Session startup / "what's next" (no issue named) | docs/handoffs/chicorytv-issue-queue.md (standing kickoff — two concurrent tracks: orientation ‖ scripts/select-queue.sh 5) |
| Named-issue pickup | Skip queue selection; go straight to focused retrieval — see "Knowledge retrieval" below, then the issue body |
Adding/changing a /api/* endpoint |
docs/api-conventions.md checklist + docs/endpoint-index.md |
| Adding a ChicoryTV SPA screen | docs/spa-conventions.md |
| Explaining a consequential settings field in the SPA (summary → hover/tap panel → docs link) | docs/spa-conventions.md §15 — use the shared FieldHelp trigger and put the copy in the screen's own FIELD_HELP record; the icon, the gesture and the a11y contract are fixed |
Graphics element / overlay work (text bug, On Now / Next, watermark-vs-[vge]) |
docs/graphics-elements.md, then decisions catalog rows keyed graphics.* |
| Scheduling / playout engine work | docs/domain-model.md + decisions catalog rows keyed sched.* (docs/decisions/README.md) |
Adding or changing a paged list handler (a page plus a TotalCount) |
Resolve api.paged-count-matches-page-query via docs/decisions/README.md — for an EF-backed filtered list, count the SAME query you page, with includes appended to the page chain only; where the count and the page are separate methods, a test pins their agreement. Then api.paging-zero-based for the pageNum/pageSize contract |
| Concurrency / optimistic-locking work | docs/api-conventions.md §7a/b/c + docs/decisions/optimistic-concurrency.md |
| Auth / security-surface work | docs/decisions/api-auth-security.md |
| CI / release pipeline work | docs/ci-cd.md + docs/decisions/release-ci-governance.md |
| Proposing a new guard / CI check / regression test convention | docs/defect-shapes-773.md §4 (detector menu + the classes where no detector is plausible), then the rules every guard must satisfy: docs/decisions/records/testing/guard-derives-population-from-source.md, …/guard-ships-with-mutation-proof.md and …/mutation-claims-are-executed.md (a MUTATION grade carries a DECLARED clause mutation that is re-run every suite) — plus …/verification-code-needs-its-own-proof.md, which extends the same obligation BEYOND guards to the harness, wrapper or checker doing the checking, and says where its proof lives when the checker holds no row |
| Adding or bounding a consequential numeric config field (an FFmpeg profile tunable, a pipeline knob) | docs/api-conventions.md §3d — reject out of range with a 422 naming the bound and its consequence, never accept-then-rewrite; validate against the constants the renderer reads, keep the render-time clamp for pre-existing rows, and let an UNCHANGED legacy value through on update. Then api.ffmpeg-profile-numeric-bounds |
| Testing a surface gated by config / an env var / a credential | docs/decisions/records/testing/deny-path-at-production-config-value.md — cover the setting absent, at its production value, and each opt-out, and assert the DENY branch |
| Touching a full-replace write path or a hand-built request object | docs/decisions/records/testing/full-replace-asserts-field-list.md — derive the field list from the DTO and assert set equality; reconcile by id where child state exists. In the SPA the same rule is enforced by the type system: docs/spa-conventions.md §4b — build the body as Complete<T>, annotating BOTH the wrapper parameter and every construction site |
| Writing or editing any doc, or answering a review finding in prose | docs/decisions/records/docs/no-session-narrative.md — the doc records the END STATE; the path to it goes in the commit message. Apply the who-benefits test, and read the carve-out before you cut (dated measurements, stated snapshot boundaries and tested-and-rejected results stay) |
| Adding, renaming or removing a workflow JOB | docs/ci-cd.md → "Per-job declarations" — every job declares env.CI_JOB_ROLE (and, in docker-build.yml, env.CI_EXECUTION_CLASS); a missing or unknown value fails scripts/tests/test_workflow_job_guards.py / …/test_ci_image_pin_population.py, and a guard or report-only job also needs a row in docs/guard-inventory.md → "Workflow-job guards". Rationale: docs/decisions/records/testing/workflow-declares-its-own-job-metadata.md |
| Adding / changing / deleting a guard file | docs/guard-inventory.md — every guard's row is machine-checked by scripts/tests/test_guard_inventory.py, so a new guard must acquire a row before the suite goes green, and a row graded MUTATION must also acquire a declared clause in scripts/tests/mutation_manifest.py |
| Writing code that reads live Gitea/remote state and then acts on it | docs/decisions/records/process/check-and-use-pins-a-version.md, then docs/remote-state-inventory.md — a new executable under scripts/ (excluding scripts/tests/), .claude/hooks/, .husky/ or .gitea/workflows/ must acquire a row there before scripts/tests/test_remote_state_inventory.py goes green |
| Finding every site that references a symbol (multi-site fix/sweep) | docs/local-lsp-tooling.md — which of the three surfaces answers, and why a delegated agent must be pointed at an MCP server (csharp-lsp, or serena after an activate_project) rather than the LSP tool, which no subagent has been observed to reach |
| Live local run / Playwright-MCP verification | docs/e2e-local.md + scripts/e2e-local.sh |
| Adding/changing a UI-E2E browser flow | docs/e2e-local.md → "UI-E2E harness" + scripts/e2e-ui.sh |
| What does a test suite cover | docs/testing.md |
| Writing a test whose behaviour is PROVIDER-SPECIFIC (collation, a value converter, data-migration DML) | docs/testing.md → "Provider-parity fixtures (opt-in MySQL)" — run one fixture body against both providers via ETV_TEST_MYSQL_CONNECTION; without it the MySQL arm Assert.Ignores visibly, and CI does not currently run it (ersatztv#627) |
| Legacy Blazor route lookup | docs/blazor-route-parity.md (historical #91 phase (b) inventory) |
| "Why do we do X this way" / challenging a convention | Catalog-first: docs/decisions/README.md (active rows) → follow the row's link to docs/decisions/records/<area>/<topic>.md for full rationale. docs/decisions/archive/<area>/ only for "what did the rule used to be." |
Knowledge retrieval (MemPalace + catalog + Gitea)
These four rules are the seam agreed with server-management#642 (the Gitea→MemPalace exporter). They apply whether the question comes up via MemPalace, a grep, or a stale comment:
- Current conventions/decisions → catalog-first. Start at
docs/decisions/README.md; discover via theErsatzTV-Decisionswing (active) /ErsatzTV-Decisions-Archive(superseded/retired). Resolve by topic/key, never by chasing a file path. - Issue history → evidence, not authority. The
Gitea-ErsatzTVwing is historical narrative that may be stale; it never overrides current Markdown. - The breadcrumb rule (the crux behavior change). A file path named inside a historical issue
comment (e.g. "grep
docs/decisions.md2026-07-17", "see …") is a breadcrumb, not a live pointer. Find the current rule via the catalog / active wing by concept; do not treat the named path as current. (Why it's safe: still-current → in the active wing, breadcrumb resolves; superseded → the active wing returns the successor and a literal follow lands on a record that announces its ownstatus: superseded; retired → the active wing returns nothing, which is itself the signal. The validator-enforced move-to-archive/is what prevents the catastrophic "superseded rule read as current" case.) - Fallback when MemPalace is stale/down:
docs/decisions/README.mdcatalog, thenrg '^`key: <dotted.key>`' docs/decisions/. MemPalace is never authority nor sole fallback.
MemPalace is candidate discovery only — every passage is verified against its cited Markdown/Gitea
source before use. Never derive live queue state from MemPalace, #237, or historical comments; queue
state is live Gitea state, retrieved via scripts/select-queue.sh (see
docs/handoffs/chicorytv-issue-queue.md). Full retrieval contract (altitude/precedence, staleness
bounds, what's mined per issue): docs/handoffs/chicorytv-issue-queue.md → "Knowledge retrieval".
Also present in docs/
docs/domain-model.md— what the app IS: entity glossary, channel→playout→schedule/block concept map, where each concept is edited in the SPA.docs/api-conventions.md— checklist for adding/changing a/api/*endpoint (controllers, DTOs, error mapping, auth, OpenAPI regen, tests).docs/spa-conventions.md— playbook for adding a screen to the ChicoryTV React SPA.docs/e2e-local.md(+scripts/e2e-local.sh) — how to run a live local instance for manual or Playwright-MCP verification.docs/local-lsp-tooling.md— the code-intelligence surfaces (theLSPtool's three servers, thecsharp-lspMCP server, andserena): how each is configured, which ones a subagent can actually reach, the traps (a cold server answers the first query with a confidently partial result; serena needs anactivate_projectper directory), andscripts/check-local-lsp.shto verify the preconditions. Read before briefing an agent to find every site referencing a symbol.docs/testing.md— testing map: what each*.Testsproject /websuite covers, golden-file nets, the timezone-independence rule, how to run subsets, the per-PR verification gate.docs/blazor-route-parity.md— historical record of the completed #91 phase (b) cutover: the Blazor Server UI is removed and every legacy route now 302-redirects to its SPA equivalent (or falls through to the catch-all →/app). Read it for the full legacy→SPA route inventory.docs/decisions/records/<area>/<topic>.md— one active decision record per file, YAML frontmatter (key/title/status/since/supersedes/superseded-by, plus optionalstale-after/sources— ersatztv#603), rationale prose in the body. The filename is the key, so one-active-record-per-key is a filesystem property (ersatztv#610).docs/decisions.mdand the topic files remain as the lifecycle-schema narrative plus a "Records formerly in this file" index, which is what keeps older date-based pointers resolvable. Generated active view:docs/decisions/README.md(catalog / task router) — start there. Superseded/retired records live indocs/decisions/archive/and are read only for history, never for "what is the current rule."docs/ci-cd.md— build/test/release pipeline, versioning, dependency management.docs/rest-api.md— REST API design doc for ersatztv#2 (goals, conventions, per-slice plan). Largely superseded day-to-day bydocs/api-conventions.md; read this for the original rationale.docs/mcp.md— theErsatzTV.Mcpstdio JSON-RPC MCP server (#58): how it wraps/api/v1as read + cautious-write tools, its config/env vars, auth, security posture, and the tool catalog.docs/graphics-elements.md— graphics element (overlay) schema reference: how elements are discovered and attached, the YAML parsing traps (an unknown key disables the element outright), the full text-element field table including the #732 background box, and why[vge]in a filter graph does not imply a graphics element is bound.docs/channels.md— Channel entity field reference.docs/m3u-xmltv.md— M3U/XMLTV generation overview (ChannelPlaylist,GetChannelGuideHandler).docs/fork-strategy.md— divergence policy vs upstream ErsatzTV.docs/design-sync.md— Claude Design ↔ repo screen workflow (#92).docs/endpoint-index.md— generated REST endpoint index (method/path/operationId/summary per OpenAPI tag). Do not edit by hand; regenerated byscripts/generate-endpoint-index.py/scripts/update-openapi.sh.docs/handoffs/chicorytv-issue-queue.md— static session kickoff prompt + workflow lore. Queue state is live Gitea state, retrieved each session viascripts/select-queue.sh— see that file's standing kickoff for the two concurrent tracks (orientation ‖ selection). ersatztv#237 is a closed, archival historical tracker (superseded bystartup.parallel-orientationindocs/decisions.md) — not a live pointer.docs/defect-shapes-773.md— root-cause analysis of the recurring defect shapes across the whole closed-issue corpus (#773): the measured class ranking, the four families they consolidate into, the cheapest mechanical detector per class, the classes where no detector is plausible, and an audit of which configured hooks/MCP servers/LSPs are actually invoked. Read it before proposing a new guard or CI check — §4 is the detector menu, and it argues against enumerating cases one incident at a time.docs/remote-state-inventory.md— every executable inscripts/(excludingscripts/tests/),.claude/hooks/,.husky/and.gitea/workflows/that reads live remote state and acts on that read, classifiedPINNED/CAS/UNSAFE-KNOWN/N/Awith the window and what bounds it. Code outside those directories — C#/TypeScript guards,web/, and the test suites themselves — is out of scope, and the doc states that rather than implying coverage. The population is derived fromgit ls-filesand compared for set equality byscripts/tests/test_remote_state_inventory.py, so a new script that talks to a remote service cannot ship unclassified. Read it withprocess.check-and-use-pins-a-version; it is that record's detector, since the class has no plausible linter (docs/defect-shapes-773.md§4 detector D).docs/guard-inventory.md— every executable guard file, what it blocks, whether it is aGUARDorTOOLING, and whether it ships a mutation proof (MUTATION/BEHAVIOUR-ONLY/NONE) with afile::functionref. The population is derived from the GIT INDEX (not a filesystem walk, since ersatztv#806) and the workflow/hook call sites, and compared for set equality byscripts/tests/test_guard_inventory.py, so a new guard cannot ship unclassified and a renamed test cannot leave a row claiming coverage it has lost. Guards implemented inline in workflow YAML are deliberately outside that population — the doc states the limit rather than implying coverage.docs/tracker-retrofit-triage-237.md— audit trail for the #524 triage of ersatztv#237's 111 comments (method, per-comment classification, totals). Evidence for thedocs.tracker-comment-retrofitdecision; read it only when triaging another over-cap tracker.docs/handoffs/rest-api.md— original handoff prompt for kicking off the REST API work (#2).