feat(610): split the decision corpus into one YAML-frontmatter file per record
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped

168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
This commit is contained in:
2026-07-25 19:45:09 +02:00
parent 8578dc1ca7
commit fba5233caf
190 changed files with 7193 additions and 5987 deletions
@@ -0,0 +1,41 @@
---
key: ffmpeg.work-ahead-slot-atomic
title: 2026-07-21 — Work-ahead slots are claimed atomically by the caller, released by the transcode it hands them to (#536)
status: active
since: '2026-07-21'
supersedes: none
superseded-by: none
rule: '`workAheadSegmenterLimit` is enforced by a single compare-exchange claim on a shared `WorkAheadSlots` pool taken by the *caller* of `Transcode`, which then passes ownership in and gets the release in `Transcode`''s `finally` — never a `Volatile.Read` compare in one place and an `Interlocked.Increment` in another.'
signals: 'work-ahead slot, workAheadSegmenterLimit, check-then-act across an await, unthrottled tune-in, TOCTOU · paths: `ErsatzTV.Application/Streaming/WorkAheadSlots.cs`, `HlsSessionWorker.Run`/`Transcode`, `ErsatzTV.Core.Tests/Streaming/WorkAheadSlotsTests.cs` · issues: #536, #350, #529, #231, #250'
mechanics: '`WorkAheadSlots.TryAcquire(limit)` CAS loop; `Transcode(bool ownsWorkAheadSlot, …)`'
---
**The defect was structural, not a missing `Interlocked`.** The write side already used
`Interlocked.Increment`, which is why the code read as thread-safe. The read side was a separate,
earlier `Volatile.Read(ref _workAheadCount) < await GetWorkAheadLimit(...)` in `Run`, and the
increment happened later inside `Transcode` — with at least one `await` (a DB-backed config read) in
between. `Interlocked` on one half of a check-then-act buys nothing. Three simultaneous tune-ins on
prod with `workAheadSegmenterLimit: 1` all observed `0 < 1` and all ran with no `-readrate`.
**Why the caller claims and the callee releases.** Moving the whole acquire/release pair inside
`Transcode` would be more symmetric, but `Run` needs the outcome *before* the call: it sets
`_state = SeekAndWorkAhead | SeekAndRealtime`, and `Transcode` reads that state on entry
(`wasSeekAndWorkAhead`) to decide whether the item starts at `DateTimeOffset.Now` or at
`_transcodedUntil`. Folding acquisition inward would have silently changed that branch. Instead the
parameter was inverted — `Transcode(bool ownsWorkAheadSlot, …)` with `realtime = !ownsWorkAheadSlot`
derived on the first line — so the contract is stated in the signature rather than implied by a
double negative, and there is exactly one release site guarded by the same flag.
**Compare-exchange rather than increment-then-back-off.** The issue proposed
`Interlocked.Increment` followed by a decrement when the post-increment value exceeds the limit.
That is correct on holder count, but the counter transiently overshoots, so a concurrent reader of
`Count` can observe a value above the limit. The CAS loop never publishes a state that violates the
invariant, which matters because the QSV hardware-frame pool sizing (`ffmpeg.qsv-extra-hw-frames-floor`)
is derived from that bound.
**Testing shape is inherited from #231/#250.** A single `Barrier(N)` + `Task.WhenAll` round does not
reliably collide on this hardware; the tests hammer 8 threads × 20 000 rounds with a per-round
barrier that validates the winner count and resets the pool. The negative control is documented in
the test file: reinstate the check-then-act body (**not** `if (true)`, which trips CS0219 under
warnings-as-errors and leaves `--no-build` running a stale, still-fixed dll). Verified: with the
pre-fix shape, 15 912 of 20 000 rounds over-claimed.