Files
ersatztv/docs/decisions.md
T
timothyandClaude Opus 4.8 aa32fd78dd
Build ErsatzTV Image / CI image pin matches docker/ci (pull_request) Successful in 5s
Build ErsatzTV Image / Docs update reminder (pull_request) Successful in 9s
Build ErsatzTV Image / decisions.md append-only (pull_request) Successful in 7s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 7m23s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 7s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 7s
Build ErsatzTV Image / Functional E2E (curl contracts) (pull_request) Successful in 15m20s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 18m18s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
fix(477): guard media-server library sweeps against successful-but-empty fetches
A successful fetch returning zero items made existing.Except([]) flag the ENTIRE
library FileNotFound in one scan — feeding EmptyTrashHandler's permanent delete
and emptying every affected collection (dead channels). Add a shared
MediaServerReconciliationGuard that skips (and logs a Warning) the sweep when
incoming==0 while items exist, wired into the three library-level sweeps
(Television shows / Movie / OtherVideo).

An empty incoming set is indistinguishable at scan time from a mid-restore /
emptied-upstream error (both report a zero total), so this deliberately overrides
#476's degenerate "last item removed => empty incoming => flag" case. #476's
cascade still fires for partial deletions (survivors present); its characterization
test moves from an empty incoming to a survivor+removed partial-deletion case.

Tests: policy table (MediaServerReconciliationGuardTests) + per-scanner integration
proving the wiring (empty incoming + non-empty existing flags/reindexes nothing).
Proven non-vacuous by neutralizing the guard. Nested TV season/episode sweeps left
unguarded (bounded blast radius); ratio-threshold + projection-failure detection
deferred to a follow-up. docs/decisions.md updated.

Fixes #477

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 23:28:27 +02:00

2189 lines
185 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Decisions — append-only log
Purpose: why the codebase does what it does, so agents don't "fix" an established convention or
relitigate a settled call. Append new entries at the bottom in date order (plus a TOC line in the
Index); **update this doc in the same PR that changes any fact below** (or that establishes a new
convention worth recording).
**Append-only, enforced (ersatztv#303 H9).** A commit or PR that *deletes or modifies* an existing
line is blocked by the Husky `commit-msg` hook and the CI `decisions-guard` job; insertions anywhere
are always allowed. The block is lifted only by the **`[decisions-edit]`** token in the commit
message, for the two legitimate reasons to touch history:
- **Fix a factual error** in a past entry.
- **Supersede a reversed decision** — add the new dated entry, then prepend a
`> **Superseded YYYY-MM by <new entry title>.**` banner to the old entry and tag its Index line
`(superseded)`. Keep the old entry — *why we changed our mind* is the point; never silently rewrite it.
**Consolidation** (prune/merge superseded entries, refresh the Index) is primarily a step in the
release checklist (`docs/ci-cd.md` → Versioning & releases), done at each release/milestone with
`[decisions-edit]` — not ad hoc mid-arc. A **size floor** backstops it between releases: the
`decisions-guard` CI job warns (non-blocking) once this file exceeds **1800 lines** — the point where
it no longer fits one default 2000-line agent Read — so append-only can't grow past what an agent can
read in a pass. The metric is the file's read cost (line count), not the entry count. Together these
keep append-only from accreting stale, contradictory, or unreadably-large history.
---
## Index
Decisions are split between the chronological log **in this file** and four **topic files** under
`docs/decisions/` — large same-topic clusters extracted at the v26.9.0 consolidation (full
rationale preserved). Check the relevant topic file below for its subject; otherwise scan the
in-file entries.
**Topic files:**
- [`decisions/optimistic-concurrency.md`](decisions/optimistic-concurrency.md) — ETag / If-Match / `Version` optimistic concurrency (#253, #259, #265, #269). Mechanics: `api-conventions.md` §7a–§7c.
- [`decisions/api-auth-security.md`](decisions/api-auth-security.md) — API/SPA auth & security posture (#197 bundles + #279 headers, #206, #283, #292, #295, #301, #319, #330).
- [`decisions/release-ci-governance.md`](decisions/release-ci-governance.md) — release / CI / merge governance (#303 H3/H4/H5/H6/H9/H10, #311, #314, #315, #335).
- [`decisions/spa-modularization.md`](decisions/spa-modularization.md) — App.tsx screen/shell extraction epic #243 (#244, #245, #247).
**In this file:**
- [2026-06 — REST API wraps existing MediatR handlers 1:1, no service layer](#2026-06--rest-api-wraps-existing-mediatr-handlers-11-no-service-layer)
- [2026-06 — UI rebuild is a React SPA (ChicoryTV) on the REST API, not a Blazor reskin](#2026-06--ui-rebuild-is-a-react-spa-chicorytv-on-the-rest-api-not-a-blazor-reskin)
- [2026-07 — Response DTOs live in `ErsatzTV.Core/Api`, file-scoped `#nullable enable`](#2026-07--response-dtos-live-in-ersatztvcoreapi-file-scoped-nullable-enable)
- [2026-07 — PUT-replace list endpoints derive `Index` from array order; alternate-schedules last row = catch-all default](#2026-07--put-replace-list-endpoints-derive-index-from-array-order-alternate-schedules-last-row--catch-all-default)
- [2026-07 — Templates editor in the SPA is a table, not Blazor's drag-calendar](#2026-07--templates-editor-in-the-spa-is-a-table-not-blazors-drag-calendar)
- [2026-07-07 — API artwork contract: rooted URLs produced server-side](#2026-07-07--api-artwork-contract-rooted-urls-produced-server-side)
- [2026-07-07 — Decode-style endpoints take a row id and look up server-side](#2026-07-07--decode-style-endpoints-take-a-row-id-and-look-up-server-side)
- [2026-07-07 — Season/episode/music-video drill-in via `parentId`, not new child-listing endpoints](#2026-07-07--seasonepisodemusic-video-drill-in-via-parentid-not-new-child-listing-endpoints)
- [2026-07-07 — Convention docs read at session start, updated in-PR](#2026-07-07--convention-docs-read-at-session-start-updated-in-pr)
- [2026-07-09 — Playback-troubleshooting completion feedback: poll status, no push channel](#2026-07-09--playback-troubleshooting-completion-feedback-poll-status-no-push-channel)
- [2026-07-09 — datetime-local instead of Chronic natural-language start parsing](#2026-07-09--datetime-local-instead-of-chronic-natural-language-start-parsing)
- [2026-07-09 — SPA gates Download Media Sample while a session is active](#2026-07-09--spa-gates-download-media-sample-while-a-session-is-active)
- [2026-07-09 — OpenAPI spec mirrors the runtime Newtonsoft serializer (#198)](#2026-07-09--openapi-spec-mirrors-the-runtime-newtonsoft-serializer-198)
- [2026-07-09 — YAML playout validator: paste-textarea instead of a server file path](#2026-07-09--yaml-playout-validator-paste-textarea-instead-of-a-server-file-path)
- [2026-07-09 — Channel numbers: prompt-driven sequential renumber instead of drag-to-reorder](#2026-07-09--channel-numbers-prompt-driven-sequential-renumber-instead-of-drag-to-reorder)
- [2026-07-09 — "Table, not calendar" convention also covers the deco-templates editor](#2026-07-09--table-not-calendar-convention-also-covers-the-deco-templates-editor)
- [2026-07-11 — Trash "See all" reuses library-browse paging; search stays capped per kind (#213)](#2026-07-11--trash-see-all-reuses-library-browse-paging-search-stays-capped-per-kind-213)
- [2026-07-11 — Logs page-size is a client-local preference, not a server ConfigElement](#2026-07-11--logs-page-size-is-a-client-local-preference-not-a-server-configelement)
- [2026-07-11 — Logs column sorting: allow-listed `sortField`/`sortDirection` on `GET /api/logs`](#2026-07-11--logs-column-sorting-allow-listed-sortfieldsortdirection-on-get-apilogs)
- [2026-07-09 — Per-playout "Schedule reset" button dropped; Reset uses the server-default build mode](#2026-07-09--per-playout-schedule-reset-button-dropped-reset-uses-the-server-default-build-mode)
- [2026-07-09 — Collection custom order: move up/down buttons, any-kind collections](#2026-07-09--collection-custom-order-move-updown-buttons-any-kind-collections)
- [2026-07-10 — Shared "Add to…" layer lives in `web/src/media/addTo/`; select-mode is an explicit toggle](#2026-07-10--shared-add-to-layer-lives-in-websrcmediaaddto-select-mode-is-an-explicit-toggle)
- [2026-07-10 — Schedule-item GET returns a flat, non-polymorphic DTO (`ScheduleItemResponseModel`)](#2026-07-10--schedule-item-get-returns-a-flat-non-polymorphic-dto-scheduleitemresponsemodel)
- [2026-07-10 — Playout API mutations return 409 while the build lock is held (#215)](#2026-07-10--playout-api-mutations-return-409-while-the-build-lock-is-held-215)
- [2026-07-11 — Schedules SPA editor: draft/explicit-Save over instant-persist; Copy includes multi/smart/rerun; shuffled-GET normalization preserved](#2026-07-11--schedules-spa-editor-draftexplicit-save-over-instant-persist-copy-includes-multismartrerun-shuffled-get-normalization-preserved)
- [2026-07-11 — Channel editor: bare-create entry point + external-logo mutual exclusion (#212)](#2026-07-11--channel-editor-bare-create-entry-point--external-logo-mutual-exclusion-212)
- [2026-07-11 — Queue state lives in the pinned Gitea tracker (#237), not in the handoff file](#2026-07-11--queue-state-lives-in-the-pinned-gitea-tracker-237-not-in-the-handoff-file)
- [2026-07-11 — EntityLocker: atomic flags + single-owner release discipline, no owner tokens (#231)](#2026-07-11--entitylocker-atomic-flags--single-owner-release-discipline-no-owner-tokens-231)
- [2026-07-11 — Media-source management REST write API + SPA (#202)](#2026-07-11--media-source-management-rest-write-api--spa-202)
- [2026-07-11 — Legacy→SPA redirect matcher: exact map + ordered segment-template patterns (#204)](#2026-07-11--legacyspa-redirect-matcher-exact-map--ordered-segment-template-patterns-204)
- [2026-07-11 — Pre-removal Blazor rollback tag `blazor-final` (#205)](#2026-07-11--pre-removal-blazor-rollback-tag-blazor-final-205)
- [2026-07-11 — Async-op API contract normalization + playout build observability + F9 scan endpoints (#235)](#2026-07-11--async-op-api-contract-normalization--playout-build-observability--f9-scan-endpoints-235)
- [2026-07-11 — Post-commit side effects run on `CancellationToken.None` (generalized from #251 to #254)](#2026-07-11--post-commit-side-effects-run-on-cancellationtokennone-generalized-from-251-to-254)
- [2026-07-12 — External-collections scans get an authoritative status surface (#271); the SPA timeout is retired](#2026-07-12--external-collections-scans-get-an-authoritative-status-surface-271-the-spa-timeout-is-retired)
- [2026-07-11 — Blazor Server UI removed (#91 phase b)](#2026-07-11--blazor-server-ui-removed-91-phase-b)
- [2026-07-12 — Live-E2E is a required step for API write-path handler changes (#303)](#2026-07-12--live-e2e-is-a-required-step-for-api-write-path-handler-changes-303)
- [2026-07-12 — TopBar primary-action button: wire creates, drop the rest (#238)](#2026-07-12--topbar-primary-action-button-wire-creates-drop-the-rest-238)
- [2026-07-13 — API versioning: the whole `/api` surface is mounted at `/api/v1`, additive-only after freeze (#286)](#2026-07-13--api-versioning-the-whole-api-surface-is-mounted-at-apiv1-additive-only-after-freeze-286)
- [2026-07-13 — Scheduling API hardening: null-name 500s, duplicate template items, unreachable 404 (#172)](#2026-07-13--scheduling-api-hardening-null-name-500s-duplicate-template-items-unreachable-404-172)
- [2026-07-16 — Functional-E2E CI harness: advisory curl-contract job over an app booted from source (#299)](#2026-07-16--functional-e2e-ci-harness-advisory-curl-contract-job-over-an-app-booted-from-source-299)
- [2026-07-16 — Optional advertised IPTV base URL (`iptv.base_url`) resolved centrally in the two generators (#340)](#2026-07-16--optional-advertised-iptv-base-url-iptvbase_url-resolved-centrally-in-the-two-generators-340)
- [2026-07-16 — Auto-tuning enumerates via EF, persists via SmartCollection; additive coexistence (#69)](#2026-07-16--auto-tuning-enumerates-via-ef-persists-via-smartcollection-additive-coexistence-69)
- [2026-07-16 — Per-playout reshuffle = scoped Reset build; seed surfaced (#71)](#2026-07-16--per-playout-reshuffle--scoped-reset-build-seed-surfaced-71)
- [2026-07-17 — Clock-boundary schedule padding already exists (FillerMode.Pad); #77 verified, convenience toggle deferred](#2026-07-17--clock-boundary-schedule-padding-already-exists-fillermodepad-77-verified-convenience-toggle-deferred)
- [2026-07-17 — Shuffle-source construction extracted to `ShuffleSourceBuilder`; per-family seam, not a god-factory (#380)](#2026-07-17--shuffle-source-construction-extracted-to-shufflesourcebuilder-per-family-seam-not-a-god-factory-380)
- [2026-07-17 — Seasonal / date-conditional scheduling already exists (alternate schedules / playout templates); #73 closed as implemented](#2026-07-17--seasonal--date-conditional-scheduling-already-exists-alternate-schedules--playout-templates-73-closed-as-implemented)
- [2026-07-17 — Auto-Tune DetailPanel member list = live search-index roll-up, not EF enumeration (#384)](#2026-07-17--auto-tune-detailpanel-member-list--live-search-index-roll-up-not-ef-enumeration-384)
- [2026-07-17 — No persistent compiler servers in CI; every `services:` container gets an explicit cap; #390's small-lane move reversed (#406)](#2026-07-17--no-persistent-compiler-servers-in-ci-every-services-container-gets-an-explicit-cap-390s-small-lane-move-reversed-406)
- [2026-07-17 — Channel health on the API = the raw `PlayoutCount` fact on the list DTO, not a derived status enum (#72)](#2026-07-17--channel-health-on-the-api--the-raw-playoutcount-fact-on-the-list-dto-not-a-derived-status-enum-72)
- [2026-07-17 — Weighted / fair-share distribution is a new `WeightedShuffle` order; `ShuffleInOrder` is anti-clumping, not fair-share (#70)](#2026-07-17--weighted--fair-share-distribution-is-a-new-weightedshuffle-order-shuffleinorder-is-anti-clumping-not-fair-share-70)
- [2026-07-17 — Auto-Tune per-channel overrides reuse the Channel Builder advanced-options DTO; weights + bug-colour logo split out to #425 (#385)](#2026-07-17--auto-tune-per-channel-overrides-reuse-the-channel-builder-advanced-options-dto-weights--bug-colour-logo-split-out-to-425-385)
- [2026-07-17 — Health-check remediation is server-declared `{Kind, Target}` on an additive DTO; the SPA acts on it (#164)](#2026-07-17--health-check-remediation-is-server-declared-kind-target-on-an-additive-dto-the-spa-acts-on-it-164)
- [2026-07-18 — Auto-Tune DetailPanel SPA: reusable `SlideOver` + shared advanced-options model; decorative panes dropped to match the backend (#386)](#2026-07-18--auto-tune-detailpanel-spa-reusable-slideover--shared-advanced-options-model-decorative-panes-dropped-to-match-the-backend-386)
- [2026-07-18 — SmartCollection rule builder: compile-only closed subset, no stored AST, one-level nesting (#176)](#2026-07-18--smartcollection-rule-builder-compile-only-closed-subset-no-stored-ast-one-level-nesting-176)
- [2026-07-18 — Auto-Tune per-source weights ride #70's MultiCollection machinery; created at tune time, not a post-hoc PUT (#425)](#2026-07-18--auto-tune-per-source-weights-ride-70s-multicollection-machinery-created-at-tune-time-not-a-post-hoc-put-425)
- [2026-07-18 — Search all-items is paged to cap DoS exposure; SPA add-all pages to completeness (#293)](#2026-07-18--search-all-items-is-paged-to-cap-dos-exposure-spa-add-all-pages-to-completeness-293)
- [2026-07-18 — Unsupported PlaybackOrder is loud at build time; a declared support matrix and tripwire test make new orders safe by construction (#403)](#2026-07-18--unsupported-playbackorder-is-loud-at-build-time-a-declared-support-matrix-and-tripwire-test-make-new-orders-safe-by-construction-403)
- [2026-07-18 — CI build-once was measured and rejected; keep the #420 tree-skip](#2026-07-18--ci-build-once-was-measured-and-rejected-keep-the-420-tree-skip)
- [2026-07-18 — Never-scanned `LastScan` surfaces as null at the API boundary, not the 0001-01-01 MinValue sentinel (#409)](#2026-07-18--never-scanned-lastscan-surfaces-as-null-at-the-api-boundary-not-the-0001-01-01-minvalue-sentinel-409)
- [2026-07-19 — WeightedShuffle SPA: weights edited on the multi-collection, order offered only on classic MultiCollection schedule items; fair-share is a reset not a mode (#404)](#2026-07-19--weightedshuffle-spa-weights-edited-on-the-multi-collection-order-offered-only-on-classic-multicollection-schedule-items-fair-share-is-a-reset-not-a-mode-404)
- [2026-07-19 — Health-check results are TTL-cached; `?refresh=true` forces a fresh run (#431)](#2026-07-19--health-check-results-are-ttl-cached-refreshtrue-forces-a-fresh-run-431)
- [2026-07-19 — CI `test` job reports a sampled true peak-anon, not cache-inflated `memory.peak` (#412)](#2026-07-19--ci-test-job-reports-a-sampled-true-peak-anon-not-cache-inflated-memorypeak-412)
---
## 2026-06 — REST API wraps existing MediatR handlers 1:1, no service layer
The REST API (#2, `docs/rest-api.md`) is thin controllers over the existing MediatR
Create/Update/Delete handlers — no new service/business-logic layer was introduced, since nearly
every handler already returns `Either<BaseError, T>`, which maps cleanly to HTTP status codes.
Latent handler bugs (missing existence checks, `KeyNotFoundException` risk, etc.) are fixed **at
the handler**, converting what would have 500'd into a proper 404/422 — not papered over in the
controller. Established across the #2a#2e gap-issue PRs. Deep FK ids nested inside item-list
request bodies (e.g. a schedule item's `CollectionId`) are deliberately **not** existence-checked at
that depth, to avoid N+1 validation queries — precedent set by the schedules endpoints (#172); see
`docs/api-conventions.md` §3 for the up-to-date statement of this rule.
## 2026-06 — UI rebuild is a React SPA (ChicoryTV) on the REST API, not a Blazor reskin
#59 committed to a full SPA rebuild rather than reskinning Blazor Server pages. Blazor removal is
split into two phases under #91: **(a)** root-flip (SPA becomes `/`) + legacy-route redirects —
DONE, merged via PR #148 (`ErsatzTV/LegacyUiRedirects.cs`, `feat/91-cutover` → main). **(b)** full
Blazor removal — gated on every route having an SPA equivalent; tracked route-by-route in
`docs/blazor-route-parity.md`.
## 2026-07 — Response DTOs live in `ErsatzTV.Core/Api`, file-scoped `#nullable enable`
New REST response DTOs go in `ErsatzTV.Core/Api/<Domain>/*ResponseModel.cs` and mirror the shape of
the corresponding Application-layer ViewModel — controllers never expose VM types directly. Because
`ErsatzTV.Core.csproj` sets `<Nullable>disable</Nullable>` project-wide, any response-model file
with an optional member needs its own `#nullable enable` pragma at the top (most already have one).
`ErsatzTV.Application` has no nullable context at all — do not add `?` annotations to types living
there; that's a Core/Api-layer-only convention. Full detail: `docs/api-conventions.md` §2.
## 2026-07 — PUT-replace list endpoints derive `Index` from array order; alternate-schedules last row = catch-all default
For "replace the whole list" endpoints (PUT over a collection — schedule items, template items,
etc.), the item's `Index` is derived from its position in the request array, not from a
client-supplied index/order field — established by `ReplaceScheduleItemsRequest.ToCommand`
(`Items.Select((item, index) => item.ToReplaceCommand(index))`). Separately, `ProgramScheduleAlternate`
and `PlayoutTemplate` rows (both `IAlternateScheduleItem`) are evaluated in `Index` order,
first-match-wins; the convention is to place the least-conditional (or unconditional) row **last**
so it acts as the catch-all default. Established by the alternate-schedules work (PR #179,
`AlternateScheduleSelector.cs`).
## 2026-07 — Templates editor in the SPA is a table, not Blazor's drag-calendar
The legacy Blazor `TemplateEditor.razor` used a drag-and-drop day-grid calendar UI. The SPA
equivalent (`/app/templates/{id}`, PR #173) renders the same day/block assignment as a table
instead. This is an accepted, deliberate parity deviation — don't "fix" it to match Blazor's
interaction model without discussing it first.
## 2026-07-07 — API artwork contract: rooted URLs produced server-side
API response DTOs return artwork as rooted, directly-usable URLs (`/artwork/posters/...`,
`/artwork/thumbnails/...`, `/artwork/fanart/...`), plus passthrough for absolute `http(s)://` URLs
and Jellyfin/Emby proxy variants. Established by PR #181
(`ErsatzTV.Application/LibraryBrowse/Queries/GetLibraryBrowseItemsHandler.cs`, private `Artwork(...)`
helper — comment: *"Returns a rooted, directly-usable artwork URL for the SPA's `<img src>`... the
SPA [needs it pre-rooted]"*), then generalized into the reusable `ApiArtwork` helper
(`ErsatzTV.Core/Api/ApiArtwork.cs`, PR #183). Root cause: the SPA has no `<base href>`, unlike
Blazor, so relative artwork paths that worked for Blazor pages 404 in the SPA. Do **not** reuse the
Application-layer Mappers used by Blazor (e.g. `MediaCards`/`Television` mappers) for new API
DTOs — those still return old Blazor-convention relative paths; map from the domain/VM directly and
root the path via `ApiArtwork`.
## 2026-07-07 — Decode-style endpoints take a row id and look up server-side
Endpoints that decode/expand opaque stored state accept a database row id and resolve server-side,
rather than accepting client-supplied serialized state to decode. Established by
`GET /api/playouts/history/{id}` (`PlayoutController.GetHistoryDetails`, PR #182) — the row's raw
JSON (`Key`/`Details`) is decoded server-side into `PlayoutHistoryDetailsResponseModel`, the client
never round-trips the raw payload itself.
## 2026-07-07 — Season/episode/music-video drill-in via `parentId`, not new child-listing endpoints
Rather than adding dedicated child-listing endpoints per media kind (e.g. "list episodes of a
season"), the library-browse endpoint takes an optional `parentId` query param and the SPA drills
in by re-querying with it. Established across PRs #181/#183 (library-picker season drill-in, then
media-detail's season/episode/artist/music-video browsing). Avoids a combinatorial explosion of
per-kind child endpoints.
## 2026-07-07 — Convention docs read at session start, updated in-PR
`docs/api-conventions.md`, `docs/spa-conventions.md`, `docs/e2e-local.md`,
`docs/blazor-route-parity.md`, `docs/domain-model.md`, `docs/decisions.md`, and `docs/README.md`
are the standing reference set every ChicoryTV session should read before starting work, and each
one carries an explicit "update this doc in the same PR" rule rather than deferring doc updates to
a follow-up. These docs **replace per-session recon** — an agent reads the index
(`docs/README.md`) and the relevant convention doc instead of re-deriving conventions from the code
each time it starts API/SPA/E2E/parity work. A testing map and a generated-endpoint index are
tracked as still-to-come under #185. Drafting this doc set also surfaced a drift in
`ApiControllerSecurityTests.cs`'s hardcoded controller registry (several controllers under
`ErsatzTV/Controllers/Api/` are missing from it — see `docs/api-conventions.md` §6) — tracked as a
follow-up under #184 rather than fixed inline, since it's a pre-existing gap, not something this
doc-drafting pass caused.
## 2026-07-09 — Playback-troubleshooting completion feedback: poll status, no push channel
The SPA playback-troubleshooting screen (`PlaybackTroubleshootingScreen.tsx`, #145) reports FFmpeg
completion by **polling `GET /api/troubleshoot/playback/status` every ~2s** while a session is
running (plus one poll on mount so a session started elsewhere still gates Play), rather than a
server push. The status endpoint returns `{ state, exitCode, speed, logs }`; the screen captures the
running→completed/failed transition in local component state and surfaces a completion notice
(success on exit 0, warning otherwise) — the SPA equivalent of the Blazor page's MediatR
`ICourier`/`ISnackbar` `PlaybackTroubleshootingCompletedNotification`. Chosen over SignalR/SSE
because the SPA has **no push channel** and troubleshooting sessions are short and user-initiated, so
a lightweight poll (started on Play, stopped on settle/unmount) is simpler than standing up a new
real-time transport. Speed thresholds and the "(Speed: Nx)" badge colors are copied verbatim from the
Blazor `GetSpeedClass` (red <0.9, green >1.1, amber otherwise).
## 2026-07-09 — datetime-local instead of Chronic natural-language start parsing
The channel-mode "Date and Time" input in the SPA playback-troubleshooting screen uses a native
`<input type="datetime-local">`, a **deliberate deviation** from the Blazor page, which parsed a
free-text field with `Chronic.Core.Parser` (natural language like "yesterday at 8pm"). The SPA has no
Chronic dependency and a picker is unambiguous; the selected local datetime is sent to
`playback.m3u8` as an ISO-8601 `start` param via `new Date(value).toISOString()`, which the
controller binds to `DateTimeOffset?` exactly as the Blazor round-trip (`"o"`) format did.
## 2026-07-09 — SPA gates Download Media Sample while a session is active
Minor intentional deviation: the SPA playback-troubleshooting screen disables **Download Media
Sample** (alongside Download Results) while a troubleshooting session is starting/running; Blazor
only gated Download Results. Both downloads compete with the live transcode for I/O and the sample
archiver reads the same media file, so gating both during a session is strictly safer and costs
nothing (sessions are short).
## 2026-07-09 — OpenAPI spec mirrors the runtime Newtonsoft serializer (#198)
The generated OpenAPI document is made to follow the **runtime** JSON contract, not the reverse. Runtime
`/api/*` responses are serialized by Newtonsoft via `CustomContractResolver`/`CustomNamingStrategy`
(camelCase + a `FFmpegProfileId``ffmpegProfileId` special case + `[JsonProperty]` overrides such as
`ChannelResponseModel.FFmpegProfile``ffmpegProfile`), while `Microsoft.AspNetCore.OpenApi` generates the
spec from System.Text.Json metadata, whose camelCase drifted (`fFmpegProfileId`, `fFmpegProfile`). That
drift fed the SPA the wrong key. Rather than hand-patch the spec or change the wire format (breaking clients),
we added `NewtonsoftSchemaNamingTransformer` — an OpenAPI schema transformer registered on all three
documents that renames each schema property through the *same* Newtonsoft contract resolver the runtime uses,
so the spec matches the wire format by construction. A contract test
(`OpenApiSerializerContractTests`) serializes representative DTOs through the real runtime settings and pins
the spec property sets to them. Decision: **the wire format is the source of truth; the spec follows it via the
real contract resolver.** This also fixed a latent SPA bug (the channel-list "FFmpeg profile" column read
`fFmpegProfile` and always showed "Unassigned"). Issue #198.
## 2026-07-09 — YAML playout validator: paste-textarea instead of a server file path
The legacy Blazor YAML playout validator took a **server-side file path** (read directly off the
container's filesystem). The SPA's `YamlValidatorScreen.tsx` instead uses a paste `<textarea>` — a
deliberate deviation, not an oversight. The SPA runs entirely client-side against `/api/*` and has
no access to the server's filesystem, so a file-path field would either need a new
filesystem-browsing endpoint or silently fail; pasting the YAML directly is simpler and matches how
every other SPA editor already round-trips content through the API instead of the disk.
## 2026-07-09 — Channel numbers: prompt-driven sequential renumber instead of drag-to-reorder
The legacy Blazor channel list let you drag-and-drop rows to reorder channel numbers. The SPA
(`App.tsx`) instead offers a "Renumber" action that walks the list and asks for each channel's new
number via a sequence of native `prompt()` calls. Deliberate deviation: drag-to-reorder needs a
dedicated drag library and a bespoke reorder-persistence endpoint; a sequential prompt reuses the
existing per-channel update call and needs no new UI dependency. Revisit only if channel counts grow
large enough that prompt-per-channel becomes tedious.
## 2026-07-09 — "Table, not calendar" convention also covers the deco-templates editor
Extends the 2026-07 "Templates editor in the SPA is a table, not Blazor's drag-calendar" entry
above (not editing that entry — this generalizes it): the deco-templates editor
(`DecoTemplatesScreen.tsx`) follows the same convention, rendering its day/deco assignment as a
table rather than reproducing Blazor's drag-and-drop calendar grid. Same rationale, same
accepted-deviation status — don't "fix" either editor to match Blazor's interaction model without
discussing it first.
## 2026-07-11 — Trash "See all" reuses library-browse paging; search stays capped per kind (#213)
`GET /api/v1/search` still returns at most 100 items per media kind, which is the cheap first page for
the common case. For an overflowing kind, the SPA's "See all N …" action pages `GET
/api/v1/library/browse` with `query=state:FileNotFound`, `mediaType`, `pageNum`, and `pageSize=100`, then
appends the results client-side. This reuses the same `GetLibraryBrowseItems` query behind search,
adds no API surface, and only pays for follow-up requests when a kind exceeds the first-page cap.
## 2026-07-11 — Logs page-size is a client-local preference, not a server ConfigElement
The legacy Blazor Logs page persisted the user's chosen rows-per-page via
`ConfigElementKey.LogsPageSize` (`SaveConfigElementByKey`/`GetConfigElementByKey`), a
per-server-instance setting stored in the DB. `LogsScreen.tsx` instead persists it to
`window.localStorage` under `ctv-logs-page-size` (same wrapped-`Storage` pattern as
`designSystem.ts`'s theme preference: try/catch getter, validated against the known option set,
falls back to a default) and restores it on mount. Deliberate deviation: this is a per-browser UI
preference, not server/business state — no other client should see or be affected by it, so there
is no reason to round-trip it through the API and grow a new `/api/*` surface (or reuse the
generic config-element endpoints) just to store a page-size number. Follows the existing SPA
localStorage convention (`designSystem.ts` theme, `auth.ts` token) rather than introducing a new
persistence mechanism.
## 2026-07-11 — Logs column sorting: allow-listed `sortField`/`sortDirection` on `GET /api/logs`
Parity for `Logs.razor`'s `MudTableSortLabel` columns (Timestamp, Level — Message was never
sortable in Blazor either). `LogsController.GetLogs` adds `sortField` (`timestamp` | `level`,
default `timestamp`) and `sortDirection` (`asc` | `desc`, default `desc`) query params, normalized
server-side the same way `pageNum`/`pageSize` are clamped rather than rejected with a 422: an
unrecognized `sortField` silently falls back to `timestamp`, an unrecognized `sortDirection` falls
back to `desc` — the pre-existing default behavior (newest-first) is unreachable to break via a bad
query string. `LogsScreen.tsx` renders the two sortable headers as buttons with a chevron
indicating the active field/direction; clicking the active column toggles direction, clicking the
other column switches to it ascending.
## 2026-07-09 — Per-playout "Schedule reset" button dropped; Reset uses the server-default build mode
Blazor's playouts page had both a per-playout **Reset** and a separate **Schedule Reset** control
(setting the daily rebuild time). The SPA keeps Reset — `POST
/api/channels/{channelNumber}/playout/reset` with no `mode` param, so the server picks the same
default Blazor used (Classic → Refresh, everything else → Reset) — but drops the dedicated
"Schedule reset" button: the daily rebuild time is already editable through the playout's
Edit-details flow, so a second entry point would duplicate an existing capability. Deliberate
deviation, not a lost capability. Issue #210.
## 2026-07-09 — Collection custom order: move up/down buttons, any-kind collections
Blazor reordered collection items with SortableJS drag-and-drop and only enabled custom ordering
for movies-only collections. The SPA (`CollectionsScreen.tsx`) uses per-row **Move up / Move
down** buttons in an explicit reorder mode instead of drag (no new drag dependency; the mode loads
ALL items first because `PUT /api/collections/{id}/custom-order` replaces the whole order from
array position — submitting a partial page would scramble the rest), and does **not** replicate
the movies-only gate: the API and the playback-side `CustomOrderCollectionEnumerator` sort by
`CustomIndex` regardless of item kind, so the SPA offers reorder for any manual collection with
custom order enabled. Issue #211.
## 2026-07-10 — Shared "Add to…" layer lives in `web/src/media/addTo/`; select-mode is an explicit toggle
The media mutation surface (#208/#209) is built on one reusable component group,
`web/src/media/addTo/` (media-domain components, like `MediaPosterCard` — not generic
`components/`): `AddToCollectionDialog` (existing-collection Select + inline "(New collection)"
create, Blazor `AddToCollectionDialog.razor` parity), `AddToPlaylistDialog` (group → playlist
Selects, no inline create), `AddToScheduleDialog` (schedule Select; payload replicates Blazor's
`AddProgramScheduleItem.ForMediaItem` defaults — see `addTo/scheduleItem.ts`),
`SaveAsSmartCollectionDialog`, and `AddToMenu` (the drop-in popover for cards/detail pages via
`MediaPosterCard`'s `actions` slot). New screens wanting add-to affordances use this layer —
don't build screen-local pickers. Two deliberate deviations from Blazor, applied consistently on
the search and browse screens: **(1) multi-select is an explicit screen-level "Select" toggle**
(off = cards open, on = cards select) rather than Blazor's always-on corner-select, because
`MediaPosterCard`'s select handler takes over the card's single click gesture; **(2) the
per-card menu offers collection/playlist for a single item of any kind, plus schedule only for
shows/seasons/artists** — collection/playlist is a superset of Blazor's per-card collection-only
menu, while the schedule target is gated to exactly the kinds Blazor's
`AddProgramScheduleItem.ForMediaItem` call sites offer, because the server validator
(`ProgramScheduleItemCommandBase.CollectionTypeMustBeValid`) accepts only the
TelevisionShow/TelevisionSeason/Artist per-media-item CollectionTypes and 422s the rest.
"Add All" (query-wide) mirrors Blazor's two-step: materialize ids
via `GET /api/search/all-items`, then reuse the id-list add endpoints — no query-based add
command exists server-side. Issues #208/#209.
## 2026-07-10 — Schedule-item GET returns a flat, non-polymorphic DTO (`ScheduleItemResponseModel`)
`GET/POST/PUT /api/schedules/{id}/items` return `ScheduleItemResponseModel` /
`ScheduleItemsResponseModel` (`ErsatzTV.Core/Api/Scheduling/`), **not** the Application-layer
`ProgramScheduleItemViewModel` hierarchy (One/Flood/Multiple/Duration subtypes). The polymorphic VM
only described its base shape in OpenAPI, so the SPA couldn't see the subtype fields (issue #126).
The flat DTO promotes every subtype field to a nullable top-level member — `multipleMode`,
`multipleCount` (renamed from the VM's `Count`), `playoutDuration`, `tailMode`,
`discardToFillAttempts` — mapped by pattern-matching the concrete VM in
`ScheduleItemResponseMapper` (`ErsatzTV.Application/ProgramSchedules/`). Its **mutation fields are
named 1:1 with `ScheduleItemRequest`** so a GET maps losslessly back to a PUT/POST
(`ScheduleItemResponseRoundTripTests` is the release gate proving the fixed point). It also carries
picker-hydration fields the editor needs: `collectionName`/`smartCollectionName`/…/`playlistName`,
`playlistGroupId` (to preselect the playlist's group), per-filler names, `watermarks` /
`graphicsElements` as `NamedIdResponseModel` lists, the computed `name`, and `durationEstimate`.
`GetProgramScheduleItemsHandler.EnforceProperties` still rewrites StartType→Dynamic, Flood→One and
Playlist/Rerun→PlaybackOrder None when `ShuffleScheduleItems` is on — that lossy normalization is
deliberate and lives on the read side (documented + tested). New shared `NamedIdResponseModel`
(`ErsatzTV.Core/Api/`) is the generic `{id, name}` embed for API responses. Issues #126/#207/#212.
## 2026-07-10 — Playout API mutations return 409 while the build lock is held (#215)
Blazor disabled per-playout Reset/Erase/Delete/Edit while a `BuildPlayout` was in flight
(`EntityLocker.IsPlayoutLocked`, `Playouts.razor` + per-kind editors); the REST API had no
equivalent, so a client could race an in-flight build with a destructive `ExecuteDelete` and leave
a half-built playout. Adversarial-reviewer#18 promoted this to a #91-phase-(b) removal gate: after
Blazor is deleted the invariant would vanish entirely.
Decision: enforce the invariant **server-side** on the API rather than re-implementing a live push
channel. `PlayoutController` and `ChannelController` inject `IEntityLocker`; every id-keyed mutation
`PUT /api/playouts/{id}`, `.../deco`, `.../alternate-schedules`, `.../templates`,
`POST .../erase-items`, `.../erase-items-and-history`, `DELETE /api/playouts/{id}`, and
`POST /api/channels/{channelNumber}/playout/reset` — checks `IsPlayoutLocked(id)` first and returns
**409 Conflict** (`ApiResults.ConflictProblem`, new shared helper mirroring `NotFoundProblem`) while
the build lock is held. The PUTs are gated too (not just the destructive ops): the target invariant
is "no mutation during a build", matching Blazor's edit-disable.
- **The guard is advisory check-then-act, not mutual exclusion** — same posture as Blazor's disabled
buttons. A `BuildPlayout` already sitting in the worker queue can take the lock a few milliseconds
after the check passes, so the original race is *narrowed*, not eliminated; consequences remain
self-healing (the next rebuild corrects a half-mutated playout). True prevention — having each
mutation acquire the playout lock for its duration — was deliberately not taken: `LockPlayout`
publishes `PlayoutUpdatedNotification` (UI churn per mutation) and would make mutations block
builds, a semantics change out of scope for restoring Blazor parity.
- **`reset-all` is deliberately NOT gated** — it stays 202. `ResetAllPlayoutsHandler` already
*silently skips* locked playouts, which matches Blazor and the handler semantics; a fire-and-forget
bulk enqueue always accepts.
- **SPA mirrors the lock via data, not a push channel** — `PlayoutListItemResponseModel` gains an
`IsLocked` bool (set from `IsPlayoutLocked` in the controller's list projection). The playouts
screen disables Reset/Erase/Erase-and-history/Delete for a locked row and shows a "Building…"
Badge; on a 409 from any mutation it surfaces the error and calls `query.refresh()` so the row
picks up the flag. No new polling was added (the existing 30s channel-state poll is unchanged).
Precedent for the 409 shape: `TraktController` (left as-is with its own private `ConflictProblem()`
to keep the diff small). Convention recorded in `api-conventions.md` §3a.
## 2026-07-11 — Schedules SPA editor: draft/explicit-Save over instant-persist; Copy includes multi/smart/rerun; shuffled-GET normalization preserved
The ChicoryTV schedules editor (`web/src/screens/SchedulesScreen.tsx` + `web/src/schedules/`,
issue #207) rebuilds the schedule-item lineup to full mutation parity with the legacy Blazor editor.
Three deliberate decisions:
**(a) Draft model with one explicit Save, replacing instant-persist.** All add/edit/copy/remove/
reorder mutate a **local draft list only**; a single **Save** issues one
`PUT /api/schedules/{id}/items` (the replace endpoint). This is intentional because that PUT is
**destructive server-side** — it deletes and recreates every item row (new ids) and triggers playout
rebuilds — so batching edits into one flush (vs. the old per-action POST/DELETE/PUT) minimizes churn
and gives the user a Discard/dirty affordance. On 422/network error the draft is kept and the error
surfaced; on success the draft is replaced with the server response. A **Discard** action and a dirty
guard (native `confirm()` on schedule-switch + in-app nav via `navigationGuard.ts`, plus a
`beforeunload` listener) protect the draft. See `docs/spa-conventions.md` §8. This retires the old
inline `ScheduleScreen` in `App.tsx` and its instant-persist add/delete/reorder tests.
**(b) Copy item deep-copies ALL source references, including multi/smart/rerun collections.** Blazor's
`CopyItem` omitted the multi-collection / smart-collection / rerun-collection references when
duplicating an item (copying only the plain collection/media-item/playlist refs) — a latent bug. The
SPA's `copyDraftItem` (`web/src/schedules/itemRules.ts`) copies every source field + display name, so
copying a MultiCollection/SmartCollection/Rerun item preserves its source. Deliberate deviation
fixing the Blazor omission.
**(c) The shuffled-schedule GET normalization (`EnforceProperties`) is preserved lossiness, matching
Blazor.** When a schedule has `ShuffleScheduleItems`, `GET .../items` still rewrites startType→Dynamic,
Flood→One, and Playlist/Rerun playbackOrder→None (and zeroes discardToFillAttempts for non-random
Duration items). The SPA does **not** fight this — it hides the Fixed start type and Flood playout
mode from the option lists (and disables the reorder arrows) for shuffled schedules, mirroring Blazor,
so a GET→edit→PUT round-trip stays consistent with the server's read-side normalization. Issue #207.
## 2026-07-11 — Channel editor: bare-create entry point + external-logo mutual exclusion (#212)
Two Blazor-parity decisions closing the channel-editor gaps (`ChannelEditor.razor` +
`ChannelEditViewModel`):
1. **Bare-channel create lives on the channels list, not a form-first route.** Blazor's
`/channels` (no `Id`) is a full add form; the SPA instead adds a "New blank channel" action next
to "Add Channel" on `ChannelsScreen` (`web/src/App.tsx`) that POSTs `CreateChannelRequest` with
Blazor's computed add-mode defaults (`(max existing int-parsed channel number) + 1`, `name: "New
Channel"`, `group: "ErsatzTV"`, `ffmpegProfileId` = `GET /api/settings/ffmpeg`'s
`defaultFFmpegProfileId`, `streamingMode: "TransportStreamHybrid"`, `isEnabled`/`showInEpg: true`,
every other field at its C# `default(T)` — see `ChannelEditor.razor`'s `else` branch for the
source of truth) directly, then navigates to `/app/edit-channel/{id}` for the rest of the fields.
This is a deliberate equivalent, not a parity gap: it reuses the existing full editor instead of
duplicating its ~20 fields into a second form. "Add Channel" (`/app/new-channel`, the
library-to-lineup `ChannelBuilder` flow) is unrelated and untouched.
2. **External logo URL wins over an uploaded logo, mirroring
`ChannelEditViewModel.ToUpdate/ToCreate`.** The channel's logo is `ArtworkContentTypeModel {
path, contentType, isExternalUrl }`; on hydration, `isExternalUrl: true` populates a separate
"External logo URL" field and the uploaded-logo draft state is treated as empty (`EMPTY_LOGO =
{ path: '', contentType: '' }` in `ChannelEditScreen.tsx`) so the two never disagree. On submit,
a non-blank URL always wins: `logo: { path: url, contentType: '' }`, exactly matching
`ExternalLogoUrl`'s precedence in the C# view model. Uploading a file clears the URL field (the
last-set field wins, Blazor parity via `UploadLogo`'s `_model.ExternalLogoUrl = null`). The URL
is validated as http(s) client-side before save is enabled (Blazor has no equivalent validation;
added because the field is free text with no server-side format check surfaced to the SPA).
Also landed with #212: `preferredAudioLanguageCode`/`preferredSubtitleLanguageCode`,
`musicVideoCreditsTemplate`, and `streamSelector` moved from free-text `Input`s to `Select`s fed by
`GET /api/languages` / `/api/channels/music-video-credits-templates` /
`/api/channels/stream-selectors`. Each keeps the channel's currently-stored value selectable even if
it's absent from the reference list (`optionsKeepingCurrent` in `ChannelEditScreen.tsx`) so loading
an existing channel never silently changes the value out from under an unmanaged language code or a
template/selector file removed from disk since save.
## 2026-07-11 — Queue state lives in the pinned Gitea tracker (#237), not in the handoff file
With multiple sessions/agents working the repo in parallel, the old protocol — every session
wholesale-rewrites `docs/handoffs/chicorytv-issue-queue.md` on main (session state + queue +
next-session prompt) — became a last-writer-wins race. New protocol: **volatile queue state
moved to Gitea**, which is concurrency-safe by construction. Pinned tracker issue **#237**
holds the goal + ordered arc in its body (edited rarely, only on arc changes, re-read before
edit) and an append-only session-comment log (fixed template: Closed / Filed / Triage /
Arc change / Recommended next). Milestone `Blazor removal (#91 phase b)` + the `review` and
`in-progress` labels are the machine-queryable view. Sessions **claim** an issue before working
it (`in-progress` label + claim comment; the tiny read→claim race window is accepted, later
claimant backs off; stale claims — no commits/comments ~48h — may be taken over with a comment).
Every new issue gets an explicit end-of-session triage verdict — gate-blocker (milestone + arc
slot) or backlog (label only) — so review findings adjust the queue only through that step and
the arc doesn't drift. The handoff file keeps only the **static kickoff prompt** and the
**append-only Lessons lore** (per-session prompts are gone; task context lives in issue bodies).
## 2026-07-11 — EntityLocker: atomic flags + single-owner release discipline, no owner tokens (#231)
`EntityLocker` (process-wide singleton, `ErsatzTV.Infrastructure/Locking/EntityLocker.cs`) is the
advisory "operation on entity X is in progress" mutex layer. Adversarial review (#20/F5, →#231)
found three defects: six plain-`bool` flags with a non-atomic check-then-set (two threads could both
acquire and both return `true`), tokenless `Unlock*` letting any caller release another owner's lock
(e.g. `BuildPlayoutHandler` ignored `LockPlayout`'s return then unconditionally unlocked in
`finally`), and one batch taking a single `LockLibrary` released after the *first* of two enqueued
scan units (so the second ran unlocked).
**Model chosen: tokenless atomic flag + documented single-owner release discipline.** The six bools
became `int`s guarded by `Interlocked.CompareExchange` (0/1), so `Lock*`'s `bool` return is now a
reliable "this call won the transition" signal and the change event fires exactly once per
transition. The contract (XML-doc'd on `IEntityLocker`): a `true` from `Lock*` confers ownership of
exactly one release — performed either in the acquiring scope (`finally`, gated on the captured
bool) or by the single designated releaser the acquirer hands off to (the consumer of the message
enqueued while holding the lock, with a compensating unlock if the enqueue throws — the established
`TraktController.EnqueueWithTraktLock` / scheduler `unlock: last` tail pattern). Callers must never
release a lock they did not acquire. `Unlock*` on an already-unlocked slot returns `false`, fires no
event, and logs a Warning — the loud tripwire for double-release bugs; it does not throw or
`Debug.Assert`, because an advisory flag must stay safe to release in `finally` paths. The three
`ConcurrentDictionary`-backed kinds (Library/Playout/RemoteMediaSource) were already atomic
(`TryAdd`/`TryRemove`) and kept their semantics (the redundant `ContainsKey` pre-checks were dropped
as tidy-up); the interface signature is unchanged across its ~40 call sites.
**Rejected:** owner tokens/leases — the acquirer and releaser for Library/Trakt/Plex/Collections
locks are different code correlated only by entity id across in-memory `Channel<T>` queues, so a
token would have to travel inside ~8 background-request message types for a defect that discipline
plus the now-trustworthy atomic return value already prevents. Counted/reentrant locks — wrong
semantics: these are exclusive in-progress flags; two holders is the failure mode, not a feature.
Call-site fixes this model prescribes land separately: #232 (scan lifecycle — enqueue only after a
successful lock, compensating unlock on enqueue failure, one release per acquisition in batches) and
#234 (BuildPlayout/subtitles gate their `finally` unlock on the captured acquire result).
## 2026-07-11 — Media-source management REST write API + SPA (#202)
Replaced the Blazor `/media/sources/{local,plex,jellyfin,emby}/...` pages (14 routes) with SPA
screens under `/app/libraries/*` over new write controllers (`LocalLibrariesController`,
`Plex|Jellyfin|EmbyMediaSourcesController`), wrapping existing MediatR commands 1:1 (no new
commands, no DB migration). Full design + adversarial-review reconciliation:
`docs/handoffs/` session record and issue #202. The design surfaced and fixed several pre-existing
Application/Infrastructure bugs newly reachable from a programmatic client; each is recorded here
because it changes documented behavior, not just adds a route.
**Secure `apiKey` contract (Jellyfin/Emby connection).** The connection GET
(`RemoteConnectionResponseModel`) returns `{ address, hasApiKey }` — the stored key **never**
leaves the server, closing a leak where the old design would have served the raw key from an
unauthenticated GET under any-origin CORS. On the connection PUT, a blank/omitted `apiKey` means
*retain the existing key*; a non-blank value sets a new one; the key is **required on first
connect** (no existing secret) → 422. Rationale: GETs aren't behind `X-Api-Key`
(`ApiKeyAuthorizationFilter` only guards mutating verbs), so a secret-bearing GET is a real
exposure regardless of how obscure the route is. **Stated explicitly as an input to #197** (the
planned read-side-auth review for secret-bearing GETs) — #197 should treat "does any GET return a
credential" as one of its checks, not just this one instance.
**Three list-replace identity contracts, not one uniform one.** An earlier draft assumed a single
"`Id<1`=add / missing=delete / id-preserved" contract across all three PUT-replace families; source
inspection proved that false for two of them:
- *Remote library sync preferences* (`PUT .../{id}/libraries`) — the command carries no source id
and the handler toggles only the ids present in the body; a row **absent** from the request is
left untouched, not deleted (libraries are sync-discovered, never created via this PUT, so there
are no `Id=0` adds either). The controller validates the submitted id set against
`Get{Family}LibrariesBySourceId(id)` (422 on any id not owned by the route's source — closes a
cross-source hole). Identity is **not stable across a disable**: `Disable{Family}LibrarySync`
removes and re-adds the row with a fresh id, so the SPA keys its draft to `(name, mediaKind)`,
never to `Id`, and refetches after every save (the PUT returns the reloaded list).
- *Path replacements* (`PUT .../{id}/path-replacements`) — id-based (existing `Id`=update,
`Id<1`=add, absent=delete) as documented, **but** the repo UPDATE SQL had no source-id predicate
(`WHERE Id = @id`, no `AND {Family}MediaSourceId = @id`), so a PUT to source A could silently
overwrite source B's row with the same numeric id. Fixed with a handler-level ownership guard
(reject any incoming positive id not in this source's current set → 422, no partial mutation)
**and** the repo SQL predicate itself (defense-in-depth for any other caller of that repo method).
- *Local library paths* (`PUT /api/libraries/local/{id}`) — identity is the **normalized path
string** (full path, trailing-separator/case-insensitive), not `Id`; `Id` in the request is
advisory. Renaming a path is delete-old+add-new under the hood (its `LibraryPath.Id` changes).
Kept as-is (matches the entrenched, tested Blazor behavior and how the SPA edits by value); not
rewritten to id-based identity, which would be a bigger, riskier change out of #202's scope.
**Plex pin-flow as REST: poll until the lock releases, exception-safe non-handoff unlock.** The
SPA polls `GET /api/media-sources/plex` rather than a per-pin status resource (no pin-addressable
server state exists to expose; SSE/push was already rejected, 2026-07-09). Polling contract is
`isLocked && !isAuthorized` = waiting on the user; `isLocked && isAuthorized` = finalizing
(discovering servers — **do not** stop here, the server list is still empty);
`!isLocked && isAuthorized` = success; `!isLocked && !isAuthorized` = timed out/abandoned. This
required fixing a **latent lock-leak bug**: `TryCompletePlexPinFlowHandler` threw
`OperationCanceledException` on its 2-minute timeout instead of returning `false`, and nothing
unlocked on that path — an abandoned sign-in wedged the Plex lock until restart or manual sign-out.
Fix releases `UnlockPlex()` on the timeout-throw, a poll-exception, and an enqueue-exception — but
**deliberately not** in an unconditional `finally`: on success the lock is handed off to
`SynchronizePlexMediaSources`, the sole releaser after server discovery; a blanket `finally` would
double-release and release *before* discovery completes, re-opening the same race the fix closes.
This is the same non-owner-token discipline as the #231 `EntityLocker` model (2026-07-11 entry
above), applied to the pin-flow's handoff-vs-terminal distinction specifically.
**404 comes from the controller pre-check, not the handler.** `Apply`/`ToEitherAsync` both
`.Join()` errors, which flattens any `NotFoundError` inside a joined `Validation` down to a plain
422. So every id-taking endpoint's real 404 is a controller-side pre-check
(`Get...ById(id)`-is-`None``ApiResults.NotFoundProblem`, the `TemplateController.DeleteGroup`
pattern), not a handler-level conversion — converting the joined validators to `NotFoundError`
would be dead code, since the join discards the distinction anyway. This is check-then-act (a
delete racing between the pre-check and the command falls through to the handler's own 422, not a
404); accepted and tested against the actual runtime error type rather than a hoped-for handler 404.
**"Scan All" dropped, not implemented.** The disabled SPA header button on the libraries hub was
speculative UI with **no Blazor equivalent** (`Libraries.razor` only ever supported per-library
scan). Removed rather than backed with a new bulk-scan endpoint; per-library scan (#232) and the
new per-source refresh-libraries endpoints (P8/J9/E9) cover the real capability set.
**App-owned popstate for guarded sub-path routes.** `LocalLibraryEditScreen` and the other
`/app/libraries/*` editors are the first screens to both register a dirty-navigation guard *and*
track their own sub-path pathname — the combination `spa-conventions.md` §8 had flagged as
unvalidated. React commits child passive effects before parent ones, so a sub-path wrapper that
self-registers `popstate` would fire (and switch sub-screen) *before* App's guard-restore listener
could veto. Resolution: App owns `popstate` centrally for the `libraries` route and only pushes an
approved sub-path down to the wrapper (which no longer self-registers `popstate`); on a vetoed pop
App re-pushes the pre-pop URL and the wrapper never sees the rejected path. This is scoped to the
`libraries` route only (gated on `activeRoute === 'libraries'`) so unguarded sub-path routes
(Playouts, Media) stay byte-identical. See `spa-conventions.md` §2/§8 for the updated exemplar
list and the resolved caveat text.
## 2026-07-11 — Legacy→SPA redirect matcher: exact map + ordered segment-template patterns (#204)
`LegacyUiRedirects.TryGetRedirect` grew from a single exact-path dictionary to a **two-tier matcher**
behind the unchanged `(PathString, out string)` signature. **Tier 1** is the existing
`OrdinalIgnoreCase` `Map` (now 52 entries — the parameterless (A)/(B) routes plus the (E-base) browse
roots whose *targets* carry `?kind=…`). **Tier 2** is an ordered `IReadOnlyList<PatternRule>` of 36
segment-template rules ((C)/(C2)/(D)/(E-page)), consulted only on a Tier-1 miss; declaration order is
match order (first-match-wins).
Template tokens are minimal: `{id}` matches a **strict positive integer**
(`int.TryParse(seg, NumberStyles.None, InvariantCulture, out id) && id > 0` — rejects signs,
whitespace, separators, `0`, negatives, and overflow like `999999999999`; the raw segment text,
e.g. `007`, is substituted, not re-formatted); `{any}` matches any non-empty segment and is dropped
(only `/playouts/add/{any}`); everything else is a literal compared `OrdinalIgnoreCase`. The request
path is split with `StringSplitOptions.None` and **empty segments are rejected** (load-bearing: so
`/channels//5` cannot match `/channels/{id}`); templates themselves use `RemoveEmptyEntries`. The
existing single-trailing-slash normalization runs before both tiers, so `/channels/5/` matches.
The set is **collision-free by construction** — exact-before-pattern plus strict numeric `{id}` means
no two tiers/rules can match the same path. **Guard invariant** (comment + `Map`-keys meta-test): no
Tier-1 key or Tier-2 template may begin with `/api`, `/artwork`, `/docs`, `/openapi`, `/iptv`, `/app`,
or `/media/sources`; rules are always full, specific templates — **never prefix wildcards** (a bare
`/media/{any}` rule is forbidden). The blazor branch does not prefix-guard `/api|/artwork|/docs|
/openapi`, so the matcher's specificity is part of their protection.
**Query-string merge**: the incoming request query is now merged into the target via a new public
`AppendQueryString(target, QueryString)` helper (one-line Startup change:
`context.Request.PathBase + LegacyUiRedirects.AppendQueryString(target, context.Request.QueryString)`).
A target that already carries `?` (the `?kind=…` browse roots) is `&`-joined instead of producing a
malformed double `?`; plain targets keep verbatim-append behavior byte-for-byte. A duplicated key
after a merge (`?kind=movies` + incoming `?kind=shows`) is first-wins in the SPA
(`URLSearchParams.get` returns the first value) — acceptable. Extracting the merge into
`LegacyUiRedirects` keeps it unit-testable without a TestServer while preserving the PathBase
re-application invariant the Startup source-text test protects.
**Rejected**: regex pairs (harder to audit for the `/api`/`/artwork` greediness invariant, noisier
tests, no benefit — every parameterized route here is "fixed segments + one variable segment");
ASP.NET `TemplateMatcher`/`RouteMatcher` (pulls routing machinery into a static helper for 36 rules);
a single unified rule list (loses the O(1) dictionary hit for the ~52 exact routes that dominate real
traffic). No `/api` change, no OpenAPI regen, no SPA change.
## 2026-07-11 — Pre-removal Blazor rollback tag `blazor-final` (#205)
Removal-gate item #205: the removal PR deletes both the Blazor reference implementation and the
`/system/health` escape hatch, so a post-deletion parity gap would otherwise be an archaeology exercise
(guessing which release tag still matches `main` minus Blazor). Decision + procedure, to run as the **first
action of the Step 2 deletion PR merge** (not before — `main` moves until then):
1. On the `main` commit **immediately preceding** the removal merge (the last commit that still contains
`ErsatzTV/Pages/**`), cut an annotated tag and push it:
`git tag -a blazor-final -m "Last commit with the legacy Blazor Server UI (pre-#91-phase-b removal)"`
then `git push origin blazor-final`. The tag name is **`blazor-final`** (not `v*`) so it does **not**
trigger the `v*` prod-release build in `.gitea/workflows/docker-build.yml`.
2. **Restore path** (if a gap surfaces post-removal): `git checkout blazor-final` → `docker build -f
docker/Dockerfile -t ersatztv:blazor-final .` → pin the **test** container to that image while the gap is
fixed forward on `main`. Alternatively `git revert` the single deletion merge commit (keep the deletion as
one squash/merge commit specifically to make this a one-liner).
3. Document the tag + restore path in the removal PR body; update this entry with the tag's commit sha when
cut.
Not cut this session — `main` still carries Blazor and will advance before the removal PR.
## 2026-07-11 — Async-op API contract normalization + playout build observability + F9 scan endpoints (#235)
Reviewer#20 F7/F8/F9. Normalizes the queue-triggering `/api/*` endpoints onto one contract, closes the two
F9 `Libraries.razor` parity gaps, and hardens the Trakt batch-lock lifecycle. Much of the F8 surface was
**already normalized** by #232 (library scan → `QueueLibraryScanResult` 202/404/409/422) and #215 (per-id
playout mutations + reset → 409 lock guard) — this issue finished the remaining outliers.
**Normalized async-op contract** (queue-triggering endpoints): **202 Accepted** = work queued; **404
ProblemDetails** = entity missing (controller pre-check); **409 ProblemDetails** = lock held (the running
job, or a mutation racing it — §3a/§3b); **422 ProblemDetails** = domain precondition (sync disabled /
unsupported / start failed). Trakt was the reference implementation. Changes made:
- `MaintenanceController.EmptyTrash` — error path **500 text/plain → 404/422 ProblemDetails** (`ToErrorResult`).
- `MaintenanceController.CleanArtwork` — silent **200 → 202** (fire-and-forget enqueue). No SPA consumer.
- `LibrariesController.ScanShow` — conflated **400 `{error}` → 202/404/409/422** via a new
`QueueShowScanResult` enum (6 outcomes incl. an honest `ScanFailed`→422, distinct from `Unsupported`).
- `ChannelController.ResetPlayout` — **200 → 202** (queue-triggering; 404/409 unchanged).
- `PlayoutController.ResetAll` — **202 (no body) → 202 + `ResetAllPlayoutsResponseModel`** reporting
`queuedPlayoutIds` / `skippedLocked` / `skippedUnsupported` (replaces the silent skip; still 202, still
skips locked/ExternalJson by design per §3a — now it *reports* what it skipped).
- `TroubleshootController.TroubleshootPlayback` — bare body-less `NotFound()` → **404/422 ProblemDetails**
with distinguishing detail. **Status codes the SPA HLS player depends on were preserved** — verified
`HlsPlayer.tsx` never branches on this endpoint's status (playback state comes from the separate
`/api/troubleshoot/playback/status` poll); only the error *body* was enriched.
**Playout build observability**: the list endpoint (`GET /api/playouts`) already stamped `isLocked` +
`BuildStatus` on `PlayoutListItemResponseModel` (#215); this issue adds **`isLocked` to the single-playout
`GET /api/playouts/{id}`** (`PlayoutResponseModel`), so the detail poll surface carries the §3a lock flag
too. No dedicated `GET /api/playouts/{id}/status` push channel was added — the flag on the existing GETs is
the HTTP-observable substitute for Blazor's live lock event, matching the `GET /api/trakt/status` precedent.
**F9 parity endpoints** (the `Libraries.razor` deletion gate — #202 did NOT close these):
- **Deep scan**: `POST /api/libraries/{id}/scan` gains `?deep=false`, threaded through
`QueueLibraryScanByLibraryId(LibraryId, DeepScan=false)` into `ForceSynchronize{Plex,Jellyfin,Emby}LibraryById(id, deep)`
(was hardcoded `false`). Non-breaking: existing callers omit it.
- **External-collections scan**: new `POST /api/media-sources/{plex|jellyfin|emby}/{id}/scan-collections?deep=false`
on the three #202 media-source controllers, dispatching `Synchronize{X}Collections(id, ForceScan:true, deep)`.
Each pre-checks source existence (404), acquires the per-source **collections** lock (`Lock{X}Collections()` —
the lock *is* the running scan, so a false = **409**), then enqueues and returns 202; the controller
compensating-unlocks in a `catch` if the enqueue throws (§3b), and `ScannerService` releases in its `finally`.
Thin SPA clients shipped (`scanLibrary(id, deep)`, `scanCollections`); **the SPA deep-scan / collections
buttons are the removal PR's remaining parity work** (parity doc §5).
**F7 Trakt batch-lock leak fix**: the global Trakt lock was released only when the *terminal* batch message
(`Unlock: true`) was processed; a `WorkerService` shutdown/cancellation before that message leaked the lock
permanently (subsequent Trakt ops 409 until restart — same class as #231/#233/#234). Fix: `WorkerService`
now releases the Trakt lock in a `finally` on read-loop exit if still held. Non-vacuous regression test proven
against an inverted-condition control.
**Accepted-by-design** (per the issue's decision-record ask): the worker's channels are **unbounded** and
there is **no shutdown drain** — messages still queued at process exit are dropped. This is acceptable because
the entity locks are **in-memory singletons that die with the process**, so a dropped message can't strand a
lock across restarts (the F7 `finally` covers the *within-process* shutdown-break leak, which is the only way
a lock outlives its batch while the process keeps running). Adding a bounded-channel backpressure / graceful
drain is out of scope and would not fix a correctness bug.
## 2026-07-11 — Post-commit side effects run on `CancellationToken.None` (generalized from #251 to #254)
Audit #22 (adversarial-reviewer) found ~20 command handlers threading the request `cancellationToken`
into work that runs **after** `SaveChangesAsync` commits — the post-commit `WriteAsync` enqueue that
rebuilds/refreshes the affected entity, `mediator.Publish`, search-index reindex, cache refresh. A late
HTTP-client disconnect cancels that token, so the *already-committed* mutation throws on the way out
**and silently drops its side effect** (the playout rebuild is never queued → the persisted edit never
takes visible effect until a manual Reset). #251 fixed this for the deco handlers; #254 generalizes the
policy across the codebase.
**Decision.** Once a mutation has committed, the *entire* compensating side effect — enqueues, publishes,
reindexes, cache refreshes, and any post-commit lookup that **gates** one of those enqueues — runs on
`CancellationToken.None`. The commit is the point of no return; past it the side effect must not be
half-abortable. Full convention + the two boundaries in `docs/api-conventions.md` §7b.
**Scope of the #254 sweep (this PR).** Swept the single-`SaveChanges` handlers under `MediaCollections/`,
`ProgramSchedules/`, `Playouts/`, `Channels/` (20 handlers). Deliberately **excluded**:
- **`BuildPlayoutHandler`** — a background/worker handler; its token is the worker shutdown token, not a
client-disconnect token, so its downstream enqueues *correctly* honor cancellation.
- **`UpdateFFmpegSettingsHandler` + the two `Configuration/` settings handlers** — they commit via several
sequential `IConfigElementRepository.Upsert` calls with an interleaved enqueue; "when is it committed"
is a partial-commit-under-cancellation question broader than the clean single-`SaveChanges` F4 pattern.
Left for a separate follow-up.
- **Response-projection reloads** (`ReplaceProgramScheduleItemsHandler` / `AddProgramScheduleItemHandler`
post-commit graph reload that builds the *returned* view model) keep the request token — a cancelled
response after a durable commit + `None`-enqueue loses nothing.
- Handlers a no-token `WriteAsync()` already makes behaviorally correct (`default` == `None`) were left
alone (explicit-`None` there is cosmetic).
Also folded in the two other #254 items on the same handlers: the channel-guide `{number}.xml` delete in
`DeleteChannelHandler`/`DeletePlayoutHandler` now routes through `IFileSystem.File.Delete` (observable
under `MockFileSystem`) **before** the commit (a post-commit delete orphans the xml on a crash; the guide
xml is regenerable on demand, so a pre-commit delete is the safe ordering), and
`ReplacePlayoutAlternateScheduleItemsHandler` now rejects an empty item list in the handler (not only at
the controller pre-guard) so a direct caller can't trip the `Max()`-on-empty crash.
**Coordination note for #253 PR2PR4.** Those PRs add `Version++` (pre-commit) to the mutating handlers of
the versioned aggregates — several of which this sweep also touched (post-commit token, a different line
region). Low git-conflict risk, but merge `main` in and expect to see the `CancellationToken.None`
convention already present on the post-commit enqueues.
## 2026-07-12 — External-collections scans get an authoritative status surface (#271); the SPA timeout is retired
**External Collections rows stay client-derived.** The SPA derives them from `GET
/api/v1/media-sources`: remote sources expose only sync-enabled libraries, so a non-empty `libraries` list is
equivalent to Blazor's `Libraries.Any(ShouldSyncItems)` filter. No second listing endpoint or fetch is needed.
The first scan-button implementation used a bounded optimistic timeout because collections locks had no HTTP
mirror. #271 replaces that temporary ceiling with the authoritative status contract below.
**Decision 1 — one family-global status endpoint, reading `IEntityLocker`.** New
`GET /api/v1/media-sources/collections-scan-status` (`MediaSourcesController` → `GetCollectionsScanStatus`
handler) returns one `CollectionsScanStatusResponseModel { family }` entry per media-source family
(`plex`/`jellyfin`/`emby`) whose collections lock is currently held, and only for active ones — the direct
counterpart to `GET /api/v1/libraries/scan-status`. It reads `IEntityLocker.Are{X}CollectionsLocked()` (the lock
*is* the running scan — the scan-collections controllers acquire it before enqueueing and the scanner releases
it on completion), analogous to how the library endpoint reads `IScannerProxyService.GetActiveScans()`. Two
deliberate shape differences from libraries: (a) **family-global, not per-source** — the collections lock takes
no source id (`LockPlexCollections()`), so an entry means *every* source of that family is scanning, matching
Blazor's all-rows-disabled behavior (per the PR #272 review note); (b) **no percent** — collections scans
expose only a boolean lock, not progress.
**Decision 2 — the SPA reconciles authoritatively; the fixed timeout is removed.** `useCollectionsScan` now
polls the new endpoint (seeding on mount, so a scan already running when the screen opens disables the buttons
immediately — the old timeout couldn't) and reconciles optimistic pending against the active-family set using
the **same `pruneGraceExpiredPending` grace-tick helper** the library hook uses (now generic over the pending
key type). A row shows "Scanning" when its family is in the active set **or** it has a still-in-grace optimistic
pending key. The grace window is kept (not the old wholesale timeout) to absorb the click→observed-active lag and
the fast-scan-between-polls race — the same bounded-pending discipline #232/#230 established for library scans.
`COLLECTIONS_PENDING_TIMEOUT_MS` is gone.
## 2026-07-11 — Blazor Server UI removed (#91 phase b)
The #91 phase (b) removal PR deletes the legacy Blazor Server UI now that the ChicoryTV SPA has parity
(all gates cleared: #145, #151/#152/#153/#155, #202, #207, #212/#213, #235 F9). The SPA is the only UI.
**Deleted.** `ErsatzTV/Pages/**` (all `.razor`, incl. `_Host.cshtml`, `FragmentNavigationBase.cs`,
`MultiSelectBase.cs`), `ErsatzTV/Shared/**` (all `.razor` + `_Favicons.cshtml`), `ErsatzTV/ViewModels/**`
(39 Blazor edit-form VMs), `ErsatzTV/Validators/**` (10 Blazor edit-VM FluentValidation validators),
`App.razor`, `_Imports.razor`, `ErsatzTV/Locals/Shared/**` + `ErsatzTV/Locals/Pages/**` (Blazor
localization resx — `ErsatzTV/Locals/Resources.*` is KEPT), `ErsatzTV/wwwroot/css/**` (site.css),
`ErsatzTV/wwwroot/lib/**` (jquery, jqueryui, sortablejs, hls, media-chrome, roboto), `libman.json`, and
`ErsatzTV.Tests/Pages/MultiSelectBaseTests.cs`.
**9 packages pruned** (from both `Directory.Packages.props` and `ErsatzTV/ErsatzTV.csproj`; each verified
to have zero remaining consumers after the Blazor deletion): **MudBlazor**, **Heron.MudCalendar**,
**Blazored.FluentValidation**, **BlazorSortable** — unambiguous Blazor UI; **MediatR.Courier.DependencyInjection**
— `ICourier` was consumed only by the deleted pages, and the app's `mediator.Publish` notification path is
plain MediatR (unaffected by the `AddCourier` removal); **Markdig**, **HtmlSanitizer**, **Chronic.Core**,
**NaturalSort.Extension** — verified zero non-Blazor consumers post-deletion.
**Startup surgical reduction.** Removed `AddRazorPages` (+`AuthorizeFolder("/")`), `AddServerSideBlazor`,
`AddMudServices`, `AddSortable`, `AddCourier`, the Blazor-attached `UseAuthentication`/`UseAuthorization`,
`MapBlazorHub`, and `MapFallbackToPage("/_Host")`. The former "blazor" `MapWhen` branch (lambda param
renamed `blazor`→`legacy`) is KEPT — it still co-hosts `MapControllers()`, `/docs` (Scalar), dev
`MapOpenApi()`, and the `LegacyUiRedirects` middleware. `MapFallbackToPage("/_Host")` is REPLACED by a
catch-all `endpoints.MapFallback(...)` that 302-redirects any unmatched path to `PathBase + "/app"`
EXCEPT paths under `/api`, `/artwork`, `/docs`, `/openapi` (those get a genuine 404, per #204's design).
**KEPT** (not removed): the OIDC/JWT/API-key SERVICE registrations (inert unless configured; real auth is
#197), `ConditionalIptvAuthorizeFilter` (`/iptv/*`), and `ApiKeyAuthorizationFilter` (mutating `/api/*`) —
per the #206 auth-posture sign-off (deleting the Blazor page challenged nothing beyond phase (a)).
**LegacyUiRedirects.** Added redirects for all 14 `/media/sources/*` routes → their `/app/libraries/*`
SPA screens (7 Tier-1 exact + 7 Tier-2 `{id}` patterns; the last Section-2 rows in
`blazor-route-parity.md`), and LIFTED the #204-era `/media/sources` forbidden-prefix guard (its Blazor
pages were replaced by #202's SPA screens). The forbidden-prefix guard now covers only `/api`, `/artwork`,
`/docs`, `/openapi`, `/iptv`, `/app`.
Also removed the now-dead ersatztv#25 razor-Sonar `<NoWarn>S6966;S3267;…</NoWarn>` line from
`ErsatzTV.csproj` — those Sonar rules only needed suppression in `.razor` `@code`; on `.cs` they run at
`suggestion` via `.editorconfig`, so removal is safe (closes part of #25's burn-down).
**Rollback.** The tag `blazor-final` was cut on the pre-removal `main` commit as the first step (see the
2026-07-11 "Pre-removal Blazor rollback tag `blazor-final` (#205)" entry above for the exact command +
restore path). Not a `v*` tag → no prod release build.
## 2026-07-12 — Live-E2E is a required step for API write-path handler changes (#303)
**A PR that changes an API write-path handler (a `POST`/`PUT`/`DELETE` `/api/*` command that mutates
state and reloads it through the read path) MUST include a live-E2E pass** — driving the real endpoint
or its SPA screen and confirming the mutation round-trips through a subsequent read — not only unit /
characterization tests. Rationale: this class has a **correlated blind spot** unit and characterization
tests share. A green fixed-point test passed while the write-path returned a production 500 because the
handler returned a *lazy* LanguageExt `Map` the test never enumerated (#229; reload-through-read-path
mechanics in `api-conventions.md` §7); the failure surfaces only when the result is materialised, which
the SPA does and the test did not. Live driving is the only reliable net for it.
Non-write-path (pure-SPA/read-only) and docs PRs don't need it. The requirement is auditable, not
silent: the PR/close comment states that live-E2E ran, or — for a non-write-path change — that it
wasn't required (the same stated-exemption discipline as the review skip rubric). Recipe +
"When live-E2E is required": `docs/e2e-local.md`. This formalizes the #229 lore bullet ("live E2E
remains the only net for this class") into a standing convention.
## 2026-07-12 — TopBar primary-action button: wire creates, drop the rest (#238)
The shell TopBar rendered a prominent top-right primary-action button (Plus icon) for **every** screen, but
only `SchedulesScreen` had ever subscribed to its `ctv:primary-action` event — so on every other screen the
button was **dead** (either a labelled no-op like "Save Changes"/"Add Channel", or, for the ~10 routes whose
`primaryAction` was `''`, a bare labelless "+"; the button rendered unconditionally). Issue #238 (from a #229
live-E2E finding, pre-existing since the TopBar's introduction).
**Decision — the Plus-icon button is a "create new item" affordance; keep+wire it only where that fits:**
- **TopBar renders the button only when the active route declares a non-empty `primaryAction`** (was:
unconditional). This alone removes every empty-`''` dead "+".
- **WIRE** (8 list screens with a single unambiguous create flow) via a shared `usePrimaryAction(routeId,
handler)` hook (`web/src/primaryAction.ts`), each delegating to the same top-level create handler its in-body
control uses (`navigateToPath('/app/new-channel')`, `setEditing({kind:'new'})`, `navigateToPath('…/add')`,
etc.): channels, schedules (refactored onto the hook), multiCollections, rerunCollections, traktLists,
fillerPresets, ffmpegProfiles, watermarks.
- **DROP** (`primaryAction: ''`, no button) everywhere else, for one of four reasons: (1) the action isn't a
create so the "+" is wrong and a correct in-body control already exists — editChannel (savebar), settings
(savebar), guide (Jump to now), playouts (Reset all), logs/troubleshooting/blockPlayoutTroubleshooting
(Refresh), playbackTroubleshooting (Play), yamlValidator (Validate); (2) ambiguous — collections (two create
types behind tabs); (3) a silent no-op — builder ("Create Channel" is disabled until the form is valid),
playlists (needs a group first); (4) semantically misplaced — dashboard (status page), libraries ("Scan" is
per-library-row, no global target).
**Why not wire everything** (the SchedulesScreen precedent): the recon found every screen already carries a
correct, disabled-state-aware, context-aware in-body control, and several banner actions are unreachable
without refactoring handlers out from under early returns, or would render a silent no-op — reintroducing the
very dead-button class #238 fixes. Wiring only the single-create-flow screens gives one explainable rule
("primary button = create a new item on a list screen") and needs no risky refactors.
**Also fixed here:** the `apiKey` route's `primaryAction: 'Save key'` + its description/comment were stale
post-#295 (the screen displays the machine key + changes the local password; there is nothing to "save") —
relabelled to reflect the #295 reality. Convention + hook documented in `spa-conventions.md` §10. The
implicit route-label ⟺ screen-subscription coupling is the kind of thing #247 (shell extraction) will
formalize; #238 keeps it a documented convention guarded by a data-driven `App.test.tsx` test over the URL-navigating
create screens (a typo'd route id → the banner navigates nowhere → red). Refs #238 #247.
## 2026-07-13 — API versioning: the whole `/api` surface is mounted at `/api/v1`, additive-only after freeze (#286)
The #197 cold review's C1 **BLOCKER**: `/api/*` was entirely unversioned (`info.version` was cosmetic), so the
first breaking change would silently break the SPA and any external/MCP client with no negotiation path. This is
the Phase-2 contract-freeze gate — versioning can't be added compatibly *after* the contract ossifies, so it
lands before freeze.
**What changed.** Every route under `/api` was swept to `/api/v1` — all 251 controller route attributes, the ~24
`Location`-header literals, the scanner callback URL (`CallLibraryScannerHandler.GetBaseUrl`), and the
`Startup` request-log path literal. This is **uniform**: the machine JSON API, the browser-session auth surface
(`/api/v1/auth/*`, still `IgnoreApi`), the internal loopback callbacks (`/api/v1/scan/*`) and the scripted-build
surface (`/api/v1/scripted/*`) are all versioned, so there is no unversioned corner and the compat rewrite needs
no exclusion list. The OpenAPI `v1.json` (160 paths), `endpoint-index.md`, and the SPA (945 request literals +
its test mocks, incl. regex/positional URL parsers) were regenerated/swept in lockstep. **No wire-DTO or
status-code change** — only the path prefix moved.
**Legacy compat = an in-pipeline rewrite, NOT a redirect** (`ApiVersionRewriteMiddleware`, sequenced before
`UseRouting` in the API branch). A legacy caller hitting an unversioned `/api/foo` has its request *path*
rewritten to `/api/v1/foo` and continues in-pipeline — method, body, auth headers and query string all survive,
so curl / the future MCP server / bookmarked URLs keep working with no round-trip (a 307/308 redirect would have
been fragile for non-GET + custom-header clients). Rewritten (legacy) responses carry RFC 8594 `Deprecation: true`
+ `Link: </docs>; rel="deprecation"`, and a `Sunset` header when `Api:LegacyRoutesSunset` is configured. An
already-versioned path (`/api/v1/*`) passes through untouched; a future `/api/v2/*` is **not** forced back to v1
(the middleware only fills in a *missing* version).
**Freeze semantics (owner decisions):** once shipped, `/api/v1` is **additive-only** — new endpoints/optional
fields are fine; renaming/removing/retyping an existing one requires a new `/api/v2`, never an in-place break.
The legacy-rewrite compat shim has a **2-release sunset window** (owner-chosen) before removal; the actual removal
is a tracked Phase-3 follow-up, not this PR. Existing pre-freeze warts (e.g. channel `{id}` vs `{channelNumber}`,
the synthesized negative-id "(none)" group rows) are frozen as-is per their own prior decisions.
**Route-convention standardization (#286, owner-requested).** The leading-slash inconsistency (238 absolute
`"/api/…"` method routes vs 13 relative `"api/…"`) is resolved: the standard is a **leading-slash absolute route
on each method's `[Http*]` attribute, no class-level `[Route]`** — except the two controllers where many actions
share a parametrized prefix (`ScannerController` `{scanId}`, `ScriptedScheduleController` `{buildId}`, ~40
methods), which keep a leading-slash absolute **class** `[Route("/api/v1/…")]` with relative method segments (the
right tool for a shared prefix). Enforced by `ApiRouteVersioningTests`: it reflects over every `[ApiController]`
in `Controllers.Api`, computes each action's *effective* route (ASP.NET's class+method combination rule), and
asserts it matches `^/api/v\d+/` — so a new controller that drifts (relative or unversioned) fails CI, the
"fix-it-while-you're-in-the-file" gate the `dotnet format` rules use. Browser-nav endpoints deliberately outside
`/api` (e.g. `GET /auth/oidc/login`) are out of scope for the test. Docs: `api-conventions.md` §1/§9. Refs #286 #197.
## 2026-07-13 — Scheduling API hardening: null-name 500s, duplicate template items, unreachable 404 (#172)
Cleared the still-live findings from issue #172 (consolidated non-blocking nits from the #144 S1/S2
reviews). Most of the 2026-07-07 list had already been ratified deliberate (§8 "(none)" synthesized
rows; §3b deep-FK non-existence-check) or fixed since (the unauthenticated `/api/logs` +
`/api/troubleshoot/info` GETs now carry `[RequiresAuthentication]` per §9; the Trakt matched-items link
points at the live `/app/search`; `GET /api/search` already fans out via `Task.WhenAll`). Three were
genuinely live:
- **Null/empty `name` → 500 (10 handlers).** Create + Replace/Update handlers for Block, Template,
DecoTemplate, Deco (8, all genuine 500s), plus `UpdateFFmpegProfile` (genuine 500; `CreateFFmpegProfile`
was already guarded) and `CreatePlaylist` (its DTO coalesces `null`→`""`, so an empty-name persist, not a
500) all did `if (request.Name.Length > 50)` on a client-nullable `string Name` → unhandled
`NullReferenceException`. Fixed to `if (string.IsNullOrWhiteSpace(request.Name) || request.Name.Length >
50)` — kills the NRE, and also rejects empty/whitespace names (matching the group-create handlers'
`NotEmpty` behavior, closing a latent "block/template named ''" gap). Chose the one-line guard over
refactoring each handler onto the `NotEmpty`/`NotLongerThan` combinator to keep the blast radius tiny and
preserve each handler's existing error message + 422 mapping. Convention captured in `api-conventions.md`
§3b handler-hardening checklist.
- **Exact-duplicate template items bypassed overlap validation.** `ReplaceTemplateItemsHandler`'s O(n²)
overlap loop skipped on `item == otherItem`, but `BlockTemplateItem` is a `record`, so two value-identical
items (same BlockId + StartTime → same computed EndTime) were value-equal and skipped — both persisted
unvalidated. Switched to index-based iteration (`i != j`) so identical items at distinct positions are
compared and register as a (self-)intersection → rejected 422. (The SPA's index-based check already caught
this client-side; it was an API-only gap.)
- **Unreachable 404 on create-group actions.** `POST /api/blocks/groups` and `POST /api/templates/groups`
declared `[ProducesResponseType(ProblemDetails, 404)]` copied from precedent, but a create has no parent
lookup that can 404 (only 201/422). Trimmed — OpenAPI spec regenerated.
Deliberately **not** fixed (documented as accepted): the §8 "(none)" synthetic rows, the §3b deep-FK
non-existence-check, and the missing `Name=` on `PlayoutController` Create/Delete/Update (moot — the
"v1"-doc `OperationIdOpenApiTransformer` (#197 Bundle C) already synthesizes stable operationIds for
`Name=`-less ops, and adding `Name=` would risk renaming generated SPA client methods). Refs #172 #197.
## 2026-07-16 — Functional-E2E CI harness: advisory curl-contract job over an app booted from source (#299)
The manual live-E2E curl flows sessions had been re-running by hand (and leaving only as PR/issue
comments) are now a CI regression net. Two decisions shaped it:
**Boot from source + `dotnet run`, not the built image.** The only "E2E" in CI before this was the
smoke test in the `build` job, which runs against the *pushed* image — so it exists only on `main`/`v*`
(the image isn't built on PRs) and would test a stale image, not the PR's code. To gate PRs on the PR's
own code, the `functional-e2e` job builds the SPA + solution and launches `dotnet ErsatzTV.dll` via the
same `scripts/e2e-local.sh` used locally (parameterized with `ETV_BUILD_CONFIG=Release`). The assertions
live in `scripts/e2e-functional.sh`, so the identical harness runs by hand and in CI — which is the
point of the issue (stop re-deriving the flows each session).
**Advisory, not blocking — separate job, not a `build` dependency, not a required check.** Per the
issue's "a functional-E2E flake must not block the unit-test gate." A boot-the-app job has more moving
parts (background process, port, readiness wait) than a pure unit test, so it starts advisory and gets
promoted to a required check / `build` dependency once proven reliable — the same staged rollout the
`migrations` job used. SQLite is the default provider, so it needs no DB service container.
**Scope is curl-only and deterministic; the racy/interactive flows are explicitly deferred.** The first
cut asserts the legacy→SPA redirect sweep (+ `/api`/`/artwork` never-redirect exemption), the
auth/CSRF/security-stamp flow, the library-scan status contract (404/202/`scan-status`), and the
`If-Match`/412 round-trip — all exercisable without seeded media, ffmpeg-transcode, or a browser (an
empty local library still enqueues `202`; an empty collection drives the concurrency editor). The 409
"already-scanning" re-trigger (needs a long-running scan to be non-racy), the playout-build lock 409,
and the genuinely UI-interactive Playwright flows are deferred as #299 follow-ups rather than shipped
flaky. Assertions were written against a real running instance, not the source — which caught that
`/artwork/*` returns `400` (not the `404` a static read suggested); extend the harness the same way.
## 2026-07-16 — Optional advertised IPTV base URL (`iptv.base_url`) resolved centrally in the two generators (#340)
ErsatzTV's absolute IPTV URLs (M3U stream/logo/guide URLs, XMLTV `<icon>`/artwork) were always
derived from the incoming request's `Scheme`/`Host`/`PathBase`, so any client that fetched with a
host downstream consumers can't resolve (the historical Gitea #1 `localhost:8409` symptom) baked
that host into the output. #340 adds an **optional** advertised base URL to pin those URLs to a
fixed public origin. Several deliberate decisions shaped it:
**(a) Resolve the override centrally in the two generation handlers, via a pure Core helper — not
in the controller.** A new `ErsatzTV.Core/Iptv/AdvertisedBaseUrl.cs` exposes `TryParse(string) →
Option<(scheme, host, baseUrl)>` and `Resolve(configured, requestScheme, requestHost,
requestBaseUrl)`. `GetChannelPlaylistHandler` (M3U, via `ChannelPlaylist`) and
`GetChannelGuideHandler` (both XMLTV `{RequestBase}` substitution sites) call `Resolve` and use its
result instead of the raw request values. Keeping the logic in a pure, allocation-free Core helper
(not the thin controller) keeps controllers dumb, makes the parse/validate/resolve rules unit-testable
in isolation, and — because the fallback path returns the exact request-derived values — leaves the
M3U/XMLTV golden tests (`ChannelPlaylistGoldenTests`, `ChannelGuideGoldenTests`) untouched.
**(b) Validation rules, with blank/invalid → request-derived so "unset" is byte-identical.**
`TryParse` accepts only an **absolute http(s)** URL with **no credentials, query, or fragment**; it
**preserves the port and any path prefix** and **normalizes a trailing slash** off. Anything failing
these rules → `None`, and `Resolve` then falls back to the request-derived scheme/host/base. So an
unset or malformed value produces output byte-for-byte identical to today's request-derived behaviour
— the override is strictly opt-in and can never silently corrupt the default path.
**(c) A NEW `iptv` settings group, not folded under `xmltv`.** The base URL affects **both** the M3U
playlist and the XMLTV guide, so it does not belong under the existing XMLTV-only settings. `GET`/`PUT
/api/v1/settings/iptv` on `SettingsController` (tier `[RequiresAuthentication]`) with body `{ baseUrl }`,
backed by `GetIptvSettings`/`UpdateIptvSettings` handlers + `IptvSettingsViewModel` /
`IptvSettingsResponseModel` / `UpdateIptvSettingsRequest`, and a new "IPTV" section on the SPA Settings
screen. Follows the existing settings GET/PUT pattern — no new API convention.
**(d) Blank clears the key; non-blank malformed → 422.** A blank/whitespace `baseUrl` on PUT
**deletes** the `ConfigElement` (reverting to request-derived); a non-blank value that fails
`AdvertisedBaseUrl.TryParse` is rejected with **422** rather than being silently stored and ignored at
generation time — the error surfaces at the point of configuration.
**(e) Scoped to M3U + XMLTV, deliberately NOT HDHomeRun.** #340 covers only the M3U playlist and XMLTV
guide generators. The HDHomeRun lineup URLs still echo the request host; extending the override there
was explicitly out of scope.
**(f) Distinct from `ETV_BASE_URL`.** The `ETV_BASE_URL` environment variable only sets the ASP.NET
Core `PathBase` (request routing/prefix); it does not advertise a scheme+host. `iptv.base_url` is the
separate, DB-stored (`ConfigElementKey.IptvBaseUrl`, **no EF migration**) advertised origin for IPTV
output. See `docs/m3u-xmltv.md` → "IPTV base URL (#340)".
## 2026-07-16 — Auto-tuning enumerates via EF, persists via SmartCollection; additive coexistence (#69)
Auto-tuning (#69) generates channels from library metadata (TV Show / TV Genre / Movie Genre).
Enumeration for the preview uses EF distinct+count queries (exact counts drive the min-items
threshold and preview display); each created channel is backed by a newly-created **SmartCollection**
(live Lucene query) so channels keep tracking the library as it grows. Query authorship is
server-side only — the client passes `{axis, value}`, never a Lucene string. Coexistence is additive:
the batch gets a reserved starting channel number (skipping taken numbers), a proposed channel whose
name already exists is flagged and de-selected by default, and existing channels are never mutated.
Bulk create loops the #63 `CreateChannelFromLineup` primitive via `ISender` and returns a per-channel
Created/Skipped/Failed outcome. Known MVP limitation: the generated SmartCollection is named after the
channel; a name collision with an existing SmartCollection surfaces as a per-channel Failed outcome.
## 2026-07-16 — Per-playout reshuffle = scoped Reset build; seed surfaced (#71)
- The shuffle seed (`Playout.Seed`) + per-collection `CollectionEnumeratorState` already persist a stable
shuffled order across rebuilds. #71's real gap was a **user-triggered per-playout reshuffle** (a Classic
playout's seed is otherwise only reseeded by a full Reset, and `reset-all` uses `Refresh` for Classic, so
it never reseeds Classic) plus visibility.
- `POST /api/v1/playouts/{id}/reshuffle` first runs `ErasePlayoutHistory` (reseeds `Playout.Seed` + clears
anchors/rerun-history) and then enqueues `BuildPlayout(id, PlayoutBuildMode.Reset)` to rebuild. `Reset`
alone only reseeds the seed for **Classic** playouts (`PlayoutBuilder`); Block/Sequential/Scripted rebuild
deterministically from the existing seed, so `ErasePlayoutHistory` is the one primitive that reseeds +
clears the derived per-collection enumerator state for all four resettable kinds — reshuffle runs it
first so the reshuffle is never a no-op for non-Classic playouts. Named `/reshuffle` (not `/reset`) to (a)
match user intent and (b) avoid the "reset-one reseeds Classic while reset-all refreshes Classic" naming
clash. It is deliberately more aggressive than `reset-all` for Classic: an explicit single-channel action
rolls a new order; the bulk action stays non-disruptive.
- `Playout.Seed` is surfaced on the playout list + detail DTOs so the SPA can show it — its purpose is
**visible confirmation** (the seed changes after a reshuffle).
## 2026-07-17 — Clock-boundary schedule padding already exists (FillerMode.Pad); #77 verified, convenience toggle deferred
#77 asked to "pad/snap channel schedules to clock boundaries (:00/:15/:30) using filler" — a Tunarr
"Pad Times" analogue for broadcast-looking guides. Investigation found the **core is already
implemented** (upstream ErsatzTV machinery carried into the fork), so #77 was scoped to *verification +
documentation* rather than a new feature:
- **Classic playouts**: a `FillerPreset` with `FillerMode.Pad` + `PadToNearestMinute = N` snaps each
content item up to the next N-minute clock mark, filling the gap with the preset's filler collection
(falling back to the item/channel fallback filler). The boundary math is a ceiling-to-the-next-multiple
on the local minute with seconds zeroed (`PlayoutModeSchedulerBase.AddFiller`, the `FillerMode.Pad`
block). Live SPA editor at `/app/filler-presets` (increment options 5/10/15/30).
- **Sequential (YAML) playouts**: the `pad_to_next: N` and `pad_until: "HH:mm"` instructions do the same,
reusing the identical boundary formula.
- **Block playouts**: no per-item Pad, and none is needed — Block is inherently clock-anchored (each
`Block` starts at a fixed `TemplateItem.StartTime`), so it already yields clean guide times.
- **EPG reflects the padding with no extra code**: `ChannelGuideProjector.ProjectFlood` coalesces trailing
PostRoll/Tail/GuideMode/Fallback/DecoDefault filler into the programme window, so a programme's `Stop` is
the padded boundary (the filler run's finish), not the raw content finish.
**What this decision fixes in the record:** the issue's premise ("reverses an earlier not-adopting call")
predates the machinery; do not re-implement pad-to-boundary from scratch. Two deterministic tests lock the
behaviour end-to-end (chosen over a one-off live-E2E because #77 changed **no production code**, so the
live-E2E's write-path/500 justification does not apply): `PlayoutBuildGoldenTests.Classic_clock_padded`
(schedule → `PlayoutItem`s land on :15 through the real `PlayoutBuilder`) and
`ChannelGuideProjectorClockPadTests` (those items → guide programmes stop on :15).
**Deferred (UI-gated):** the only real residual is *convenience* — there is no one-click per-channel /
per-schedule "clock-align this whole channel" toggle; today you hand-build a Pad `FillerPreset` and wire it
into a schedule item's roll slot. That toggle is mostly SPA work and is deferred to a follow-up **blocked on
the #388 design-system sync epic** (all UI work is currently gated on #388), which also covers adding a
60-minute increment to the filler-preset editor's options (the backend already accepts any integer).
## 2026-07-17 — Shuffle-source construction extracted to `ShuffleSourceBuilder`; per-family seam, not a god-factory (#380)
`PlaybackOrder` is turned into an enumerator at five independent sites (one per schedule kind, dispatched
by `BuildPlayoutHandler` on `ScheduleKind`): Classic (`PlayoutBuilder`), Block (`BlockPlayoutBuilder`,
`PlayoutHistory` rotation), Scripted (`SchedulingEngine.EnumeratorForContent`), Sequential/YAML
(`EnumeratorCache`), and the Playlist leaf (`PlaylistEnumerator.Create`, consumed by the others). These
are **not** accidental duplication — they map the same enum to different enumerator *families* keyed on
different state models (Block's history cursor vs the stateless-index `CollectionEnumeratorState`
enumerators). The real smell was one **cross-engine reach-in**: `PlaylistEnumerator` called
`PlayoutBuilder.GetGroupedMediaItemsForShuffle` / `GetCollectionItemsForShuffleInOrder` as **statics** —
one engine reaching into another engine's class.
- **Fix (a) shipped, not (b).** The two shuffle-source helpers moved verbatim to a new
`public static class ShuffleSourceBuilder` in `ErsatzTV.Core/Scheduling` (sibling to the also-static
`MultiCollectionGrouper` / `MultiPartEpisodeGrouper`; dependencies passed as parameters, **not** a DI
service — `PlaylistEnumerator.Create` is itself a static factory). Both Classic and Playlist now consume
it, so the reach-in is gone and shuffle-source construction lives in one directly-unit-tested place.
- **One intentional signature change.** `GetGroupedMediaItemsForShuffle` now takes
`bool keepMultiPartEpisodesTogether, bool treatCollectionsAsShows` instead of a `ProgramSchedule`
(verified those are the only two properties it read). This deletes `PlaylistEnumerator`'s fake
`new ProgramSchedule { KeepMultiPartEpisodesTogether = false }` (its TODO becomes an honest
`false, false`) and gives future callers with no `ProgramSchedule` (#176 PseudoTV, #70 distribution) a
schedule-entity-free entry point. Behavior is unchanged.
- **Rejected: a unified Classic+Playlist enumerator factory (b).** That would be the god-factory the issue
forbids — three concrete hazards: Classic's `Marathon` arm *produces* a `PlaylistEnumerator` (a factory
consumed by `PlaylistEnumerator.Create` that also builds one is a dependency cycle); Playlist hardcodes
`KeepMultiPartEpisodesTogether=false`/`randomStartPoint=false` where Classic reads schedule flags (a
merge silently changes behavior); and Classic's `CustomOrder`/`Rerun`/`MultiEpisodeShuffle`-template arms
gate on instance deps Playlist doesn't have. Engine separation is deliberately preserved.
- **Block / Scripted / YAML left alone.** Block is a distinct family; the Scripted≡YAML construction
duplication (`EnumeratorForContent` ≡ `EnumeratorCache.GetEnumerator`, both using the *different*
`BlockPlayoutShuffledMediaCollectionEnumerator` for shuffle) is a real but separate cleanup — deferred to
a follow-up gated on #381 (those paths have no golden coverage yet).
- **Net first, because #163's was too narrow.** `PlayoutBuildGoldenTests` only pinned
`PlaybackOrder.Chronological`; the moved shuffle paths (incl. the entire Playlist reach-in) were
uncovered, so "goldens green" would have been a false signal. The PR adds a `classic-shuffle` golden
(pinned `Playout.Seed` + `Continue` build, since `Reset` randomizes the seed) and a
`PlaylistEnumerator.Create` reach-in characterization (Shuffle + ShuffleInOrder items) **before** the
move, plus direct `ShuffleSourceBuilder` unit tests for the branches goldens don't reach
(multi-collection vs fake-multi-collection lookup; multi-part grouping on/off).
## 2026-07-17 — Seasonal / date-conditional scheduling already exists (alternate schedules / playout templates); #73 closed as implemented
#73 asked for "seasonal / date-conditional channels & schedule rules" — a PseudoTV holiday-channel analogue
— on the stated premise that *"ErsatzTV has no native date-conditional scheduling today; you'd hand-build it
seasonally."* **That premise is false.** The feature is first-class, shipped, SPA-reachable, unit-tested and
already documented (upstream machinery carried into the fork). #73 is closed as already-implemented; the only
deliverable was the docs recipe below.
- **The predicate**: `IAlternateScheduleItem` (`DaysOfWeek`, `DaysOfMonth`, `MonthsOfYear`,
`LimitToDateRange`, `StartMonth/StartDay/StartYear?`, `EndMonth/EndDay/EndYear?`), implemented by
`ProgramScheduleAlternate` (Classic — picks the `ProgramSchedule` for a date) and
`Scheduling/PlayoutTemplate` (Block — picks the `Template` + optional `DecoTemplate` for a date).
- **The evaluator**: `AlternateScheduleSelector.GetScheduleForDate` — first match in `Index` order,
broadest/unconditional row last = catch-all. Wired into `PlayoutBuilder` (:572, :956), `EffectiveBlock`,
`BlockPlayoutFillerBuilder`, and `DecoSelector`. Wrap-around (Nov→Feb) and invalid/leap dates (Feb 31
clamped) are already handled; `AlternateScheduleSelectorTests` + `playoutTemplateCalendar.test.ts` pin it.
- **SPA**: `PlayoutScheduleEditors.tsx` ships the "Limit to date range" control for both editors, with a
per-date calendar preview (`playoutTemplateCalendar.ts`).
- **Nullable years are the seasonal switch** (`AlternateScheduleSelector.cs:32-40`): years default to the
*queried date's* year, so blank years = repeats every year. The override branch requires **both**
`StartYear` and `EndYear` (only one → silently still yearly), and explicit years force `reverse = false`,
disabling wrap-around. This was the single most load-bearing undocumented behaviour and is now in the docs.
- **Correcting the issue's Deco hypothesis**: the issue proposed "a Deco-like modifier". Decos carry **no**
dates — `DecoTemplateItem.StartTime/EndTime` are time-of-day. The date dimension lives *exclusively* on the
`PlayoutTemplate` / `ProgramScheduleAlternate` join row. Composition is **PlayoutTemplate (date range) →
Template (time-of-day grid) → Block (content)**.
**Why closed rather than #77-style "re-scope to verify + document":** #77 had two real gaps (undocumented AND
untested). #73 has neither — tests and docs both already existed on main — so even the #77 re-scope target was
already met. Under the close-don't-park convention (same date), holding #73 open for a hypothetical future
build is precisely the parking that convention forbids.
**Rejected asks, and why:**
- **Per-schedule-item date predicate** (the issue's literal ask; shipped design swaps the *whole* schedule per
date instead). Rejected as-designed: functionally equivalent for every seasonal use case — clone a schedule,
date-gate it, order it above the catch-all — at a cost of "clone a schedule". Moving the predicate onto
`ProgramScheduleItem` rows would touch the classic playout builder's core loop for **zero new expressible
behaviour**, carrying regression risk in the exact engine #163/#380 just built golden nets around.
- **"Prioritize collection X during a date range"** — routed to **#70**, not built here. The ask decomposes as
*date scoping* (shipped) × *soft prioritization* (the weighting primitive #70 is actively building across the
`PlaybackOrder` switch sites). Once #70 lands, this is pure composition with no date-aware code in the
weighting path. Building a second weighting primitive under #73 while #70 was mid-flight would have been a
direct collision in the same milestone.
- **Date-gating a playout's default `Playout.DecoId`** — rejected: the gated path already exists
(`PlayoutTemplate.DecoTemplateId`), and a date-conditional *fallback* is a contradiction in terms.
- **UI rename** of "Alternate Schedules" / "Playout Templates" to something that reads as "seasonal" — rejected:
inherited upstream vocabulary baked into routes, API paths and parity docs; churn for marginal gain. The
discoverability gap is addressed with docs instead.
**Known non-goal:** the issue's "optionally, auto-surface a seasonal channel in the lineup during its window"
is genuinely absent — M3U/XMLTV lineup visibility is not date-conditional. It was speculative, has no concrete
use case (the practical equivalent: air a date-gated schedule year-round, or toggle `ShowInEpg`), and per
close-don't-park no placeholder issue was filed. Refile if a real need appears.
**The real gap was discoverability, and it was a docs gap.** The mechanism was documented as a *mechanism*
("Alternate Schedules… useful for seasonal programming") but never as a *task* — a user thinking "holiday
channel" has no reason to look under "Alternate Schedules", and on finding that line gets no worked steps.
Fixed with a task-shaped **"Recipe: seasonal / holiday programming"** section in `channels.md` (both engines,
plus the gotchas above) and a `domain-model.md` glossary row. No production code changed, so no live-E2E
(same reasoning as #77).
## 2026-07-17 — Auto-Tune DetailPanel member list = live search-index roll-up, not EF enumeration (#384)
The Auto-Tune DetailPanel (#383) shows, per proposed channel, the distinct **content sources** its
generated SmartCollection resolves to (a genre channel's shows/movies), each weightable in the #385
write path. `GET /api/v1/channels/auto-tune/members?axis=&value=` backs that list.
**Why the search index, not an EF distinct+count query** — even though #69's *preview* enumeration uses
EF (`PreviewAutoTuneChannelsHandler.EnumerateTvShows/…`). The created channel's playout is built from a
**SmartCollection** whose members come from `ISearchIndex.Search` (`MediaCollectionRepository.GetSmartCollectionItems`).
For the DetailPanel to faithfully preview *what the built channel will actually contain*, the member list
must run the **same** query through the **same** index — so the handler calls the server-owned
`AutoTuneAxisMap.GenerateQuery(axis, value)` (client never sends Lucene, per #69 PR1) and rolls the matching
leaf items up to their distinct parents. #69's preview is a different granularity (enumerating candidate
axis *values* with EF exact counts to drive the min-items threshold); this is enumerating the *members of one
value*, where index-fidelity matters more than count-exactness. The two coexist deliberately.
**Roll-up + shape.** Episode axes (TvShow/TvGenre) → distinct parent shows (Episode→Season→ShowId), with
`ItemCount` = the **query-matching** episode count, not the show's total (a show contributes only its matching
episodes to a genre channel). Movie-genre axis → the matching movies are themselves the sources. Both reuse
the existing `PagedLibraryBrowseItemsResponseModel` / `LibraryBrowseItemResponseModel` DTOs and
`LibraryBrowseItemMapper.GetShows/GetMovies` (no new schema), ordered by title then id, paged in memory
(the distinct-source set is bounded — dozens for a genre). Search pulls up to the 10k cap
`GetSmartCollectionItems` already uses; a value resolving to >10k leaf items could under-report sources past
the cap — the same staleness bound the smart-collection path accepts. Read-only, catalog-read tier (no
`[RequiresAuthentication]`), so a cold review sufficed. Sibling backend child #385 (write-path overrides +
weights) and SPA child #386 remain open under the #383 milestone.
---
## 2026-07-17 — No persistent compiler servers in CI; every `services:` container gets an explicit cap; #390's small-lane move reversed (#406)
Three CI changes, all downstream of one incident: on 2026-07-17 bumblebee (the **prod media** Docker
host, 25 GiB / 12 cores) hit load **340** with **21 GiB swapped** and ~238 MiB free, taking prod
ersatztv/jellyfin down until reboot. Prod ersatztv itself was healthy at 168 MiB throughout — the
thrash was **CI-induced**, triggered by a burst of parallel merges to `main`. Infra sizing is
server-management#604's boundary; these three levers live in this repo and each does more than any
capacity knob.
**1. No persistent compiler servers.** `VBCSCompiler` is a *persistent* Roslyn server: it outlives
the `dotnet build` that started it and holds its heap for the next one. Measured at **7.8 GB RSS**
live on bumblebee — the single largest consumer on the box, and the actual reason each job needed a
10 GiB cap. In CI it buys **nothing**: each job container is torn down at the end of the run, so
there is never a "next build" to warm. The workflow's top-level `env:` now sets
`UseSharedCompilation=false`, `DOTNET_CLI_USE_MSBUILD_SERVER=0`, `MSBUILDDISABLENODEREUSE=1` —
MSBuild properties set as env vars so they apply to every `dotnet` call without touching each call
site (MSBuild surfaces env vars as properties; `UseSharedCompilation` only defaults to `true` when
empty, so the env var wins). Verified locally: a default build leaves 1 `VBCSCompiler` alive, the
same build under these vars leaves **0**, and `ErsatzTV.sln` still builds clean (0 errors).
**Why the Dockerfile also sets them** — and this is the part the issue's suggestion would have
missed: the workflow `env:` reaches the *runner-side* dotnet jobs only. The `build` job compiles
inside `docker build`, where it does not propagate, so the SDK stage of `docker/Dockerfile` sets the
same three as `ENV`. That is precisely the job server-management#570 measured pegging **5.999/6
GiB** — the one that most needs it. Build-stage only; the final image is `FROM runtime-base`, so
nothing lands in the shipped image or affects runtime.
**Trade-off accepted**: without the shared server each project's `csc` is a fresh process, which
costs some build time. Worth it — the memory spike is what takes prod down, and the cap sizing that
spike forces is what starves the lanes.
**2. `services:` containers do not inherit the runner's cap.** A runner's `container.options`
(`--cpus=4 --memory=10g`) applies to the **job container only**. Verified by inspecting a live
`migrations` job: the job container reported `HostConfig.Memory=10737418240`, its `mysql:8.4`
service reported `mem=0 nanocpus=0` — **unbounded**. Every migrations run was adding an uncapped
MySQL to an already-tight host. Now `--memory=2g --memory-swap=2g --cpus=2`.
**Why `--memory-swap` is not redundant** (cold-review catch, and the sharpest thing in this change):
Docker defaults an unset `--memory-swap` to **twice** `--memory`, so `--memory=2g` alone grants 2g
RAM **plus 2g of swap** — verified live: `--memory=2g` → `memory.max=2147483648` **and**
`memory.swap.max=2147483648`; `--memory=2g --memory-swap=2g` → `memory.swap.max=0`. Capping RAM
while silently permitting swap is close to the worst outcome **on the host whose swap thrash is the
entire reason for the cap**, and a swapping mysqld mid-DDL is exactly the pathology behind the known
`Command Timeout expired` migrations flake — i.e. the naive cap could have made that flake worse.
**Standing rule: prefer a loud OOM over silent swapping.** An OOM is an unambiguous "raise the cap"
signal; swapping just degrades everything and blames something else.
**The same 2× applies to the runners' 10g job slots** — each is really 10 GiB RAM *plus* 10 GiB
swap, so the "60 GiB promise on a 25 GiB host" understates by 2× and is a plausible direct mechanism
for the incident's 21 GiB swapped. That is server-management#604's boundary; reported there.
**On 2g, honestly**: 543 MiB is init+idle, **not** the 787-migration replay (which grows caches idle
never touches), so 2g is a measured *floor* plus headroom, not a measured ceiling — the `migrations`
job going green is what validates it. `--cpus=2` has no measurement behind it at all; 787 sequential
DDL statements on one connection are ~1-core-bound, so it is judgement. Recording that rather than
dressing a guess as measurement — the same failure this entry criticises below.
**Standing rule: any new `services:` container needs its own explicit cap** — it will not inherit
one, and `--memory` without `--memory-swap` silently grants 2× in swap.
**3. #390's `small`-lane move for `api-docs`/`format` reversed.** #390 moved them to dodge a ~29 min
`ubuntu-latest` queue. The queue was real, but the lane was the wrong fix, and **#390's own comment
flagged why**: *"on an API-touching PR this job does a full `dotnet build`, so it is not always a
'small' job; capacity 4 absorbs that."* "Capacity 4 absorbs that" held only because **nothing
enforces the sum** of the lanes' caps — 6 slots × 10 GiB on a 25 GiB host is a 60 GiB promise. These
are not small jobs: a live `docker stats` caught the `format` job container at **3.95 GiB**, which
#604's re-sized 2 GiB `small` lane would OOM-kill outright. #604 grows `ubuntu-latest` to 5 slots
(48 GiB ci-runner at capacity 4 + a bumblebee overflow slot) and fixes the queue at the source. This
one is **order-coupled with #604**: the small lane's caps can't tighten until it lands.
**Measurement is now continuous, not a one-off.** The `test` job's **last** step reports the
cgroup's `memory.peak` plus an `anon`/`file` breakdown (`continue-on-error`, tolerates absence).
#604 sizes both runners' caps on that number, and until now it was *inherited* rather than measured
— the 10g cap traces back to #570 observing a different job entirely. Two things that look like
details but are the whole point: it must run **last** (`memory.peak` read at step N reports the peak
only up to N, so an earlier placement silently excludes the job's later workload), and
`continue-on-error` — not `if: always()` — is what makes it advisory (`always()` controls whether a
step *runs*, not whether its failure fails the job, and `defaults.run.shell: bash` means `-e` is on).
**And then the instrument taught us the lesson twice, both times at our own expense.** The first
reading came back `peak 8305 MiB`. `memory.peak` is the high-water mark of `memory.current`, which
charges **page cache** as well as anonymous memory — proven: a container with `anon=0` that merely
reads an 800 MB file reports `memory.peak=826 MiB`, `file=800 MiB`. Page cache is *reclaimed* under
a tighter cap, not OOM-killed, so a large peak that is mostly `file` is **not** evidence a cap must
stay high. `anon` is what forces an OOM; size caps on it. The step now prints the split (end-of-job,
so indicative rather than peak-instant); a true peak-anon sample is **#412**.
**Then we made the same mistake in the opposite direction.** Having established that peak
*overstates*, the docs (and a report to #604) leaned to "so this number is probably mostly cache."
That was a guess about *magnitude* dressed in a verified fact about *mechanism* — and an independent
probe killed it: a full solution build in the CI image with shared compilation off measured `peak
9457 MiB` / **`anon 7134 MiB`** / `file 421 MiB`. **Anon dominated.** So a 6g cap looks *unsafe*,
#570's "6g proved too tight" is the rule rather than an outlier, and **#406's premise ("if this
brings peak RSS well under 6 GiB, the whole budget loosens") is looking dead** — the 7134 MiB was
measured *with* shared compilation already off. The switches are still right (no persistent 7.8 GB
server between builds); the looser budget they were supposed to buy is not.
**What is NOT established:** there is still no pre-change baseline from this instrument — the 7.8 GB
`VBCSCompiler` figure was measured host-wide across concurrent jobs, not inside one job container —
and the anon figure above is one probe, not the `test` job. #412 covers the real A/B. What *is*
established: no persistent compiler server survives a build, and `migrations` is green with mysql
capped at 2g with swap disabled.
The lesson generalizes, and note it bit *this* change twice — once in the issue's premise and once
in our own instrument: this repo's CI perf work keeps stating numbers from plausibility rather than
measurement (see #390's "24min" apt-ffmpeg estimate; real 110s, and not load-bearing). Measuring
the wrong quantity precisely is the same failure wearing a lab coat.
**What this does NOT shrink**: `format`. `dotnet format` loads Roslyn in-process via
MSBuildWorkspace and never spawns `csc`, so its measured 3.95 GiB is untouched by any of this. Do
not size the `small` lane expecting otherwise.
**Not addressed here**: the 1235 min queue waits (server-management#604) and the redundant
triple-build (#398).
## 2026-07-17 — Channel health on the API = the raw `PlayoutCount` fact on the list DTO, not a derived status enum (#72)
#72 asks for per-channel status in the Channels list, "especially channels that will fail to play".
**Where it lives: `ChannelResponseModel` (the lean list DTO), not a new endpoint and not
`/channels/state`.** Channel health is *config-derived* — it changes when someone edits a playout, not
tick to tick — whereas `/channels/state` is the fast-poll runtime-liveness feed (`OnAir` = someone is
streaming *right now*). Folding health into the polled feed would recompute rarely-changing data every
tick and mix two cadences in one DTO; a third endpoint is over-engineering for one derived integer on a
list whose consumer already reads it. Precedent: `api-conventions.md` §3a stamps the server-derived
`IsLocked` onto `PlayoutListItemResponseModel` for exactly this reason. **It is also free**: `GetAll`
already `Include`s `Playouts` and `MirrorSourceChannel.Playouts` and was discarding them, so no extra
query and no N+1 on a large lineup.
**A raw fact (`int PlayoutCount`), not a `ChannelHealth` enum.** v1 has exactly one trustworthy negative
signal, and the auto-tune arc (#383/#384) is about to churn the status taxonomy (origin, auto-tune
outcomes) — freezing a server-side enum now guarantees a breaking rev of a frozen-additive `/api/v1`
surface. `PlayoutCount` mirrors the long-standing `ChannelViewModel.PlayoutCount` 1:1, is a fact rather
than a policy, and leaves the SPA to derive `0 ⇒ "No playout"` in one predicate.
**What v1 deliberately does NOT compute** — each was considered and ruled out, so don't "finish" them
without reading this:
- **Empty schedule behind an existing playout.** `EmptyScheduleHealthCheck` only understands **Classic**
`ProgramSchedule` playouts. Block, Sequential, Scripted and ExternalJson channels have no
`ProgramSchedule` at all, so a badge driven off that query would be silently absent or wrong for four of
the five schedule kinds — the #71 "verify a shared primitive covers ALL variants" trap. Needs a
per-kind emptiness notion first.
- **Broken / missing source.** `FileNotFound`/`Unavailable` are server-wide media-item counts with no
channel attribution; mapping media → collection → schedule → channel is a project, not a field.
- **User-defined vs auto-generated origin** (#72 scope item a). No honest signal exists:
`Channel` has no origin column, and `ChannelPlayoutSource.Generated` is a *playout-strategy* value that
SPA-created blank channels also carry, so it would mislabel them. Requires a new column + a dual-provider
migration, and provenance belongs to the auto-tune arc that stamps it at creation. A join through the
`"Channel Lineups"` system playlist group was **rejected**: it is a heuristic that breaks the moment a
user edits the channel. Deferred to a follow-up blocked on the auto-tune backend.
Keeping all three out held #72 to a **read-path-only** change: no migration, no write-handler live-E2E.
**Corrected in passing:** `ChannelRepository.GetChannel` never included `Playouts`, so
`GET /api/v1/channels/{id}` reported `playoutCount: 0` for every channel, which silently disabled the
channel editor's playout-source guard. Both call sites now share `Mapper.GetPlayoutsCount` (Mirror-aware).
The related *silent* server-side coercion of Mirror→Generated (a 200 that discards the caller's intent,
against the §3 "surface it, don't silently filter" rule) is filed as **#401**, not fixed here.
## 2026-07-17 — Weighted / fair-share distribution is a new `WeightedShuffle` order; `ShuffleInOrder` is anti-clumping, not fair-share (#70)
`PlaybackOrder.WeightedShuffle = 9` ships the two behaviors #70 asked for — "air Show A 70% / Show B 30%"
and "no single show dominates" — as **one** order, because fair-share is the equal-weights case of weighted.
- **`ShuffleInOrder` does NOT already do fair-share, despite looking like it.** Its balanced shuffle
(keyj) pads every source to the longest with `Option<MediaItem>.None` spacers — but **spacers emit
nothing** (`ShuffleInOrderCollectionEnumerator.cs`, `result.AddRange(maybeItem)` over an `Option`). One
cycle therefore plays **every item exactly once**, so a 200-episode show still takes 10× the airtime of a
20-episode one. What it buys is **anti-clumping**: the small show is spread evenly instead of arriving in
bursts, and each show plays chronologically. That is a different product from "airs as often". Recorded
because the padding reads as equalization and this was misread once during #70's own design pass — the
distinction is the entire justification for the issue.
- **Plain `Shuffle` is already implicitly weighted by collection size** (FisherYates over flattened items),
which is why "just use Shuffle" isn't fair-share either.
- **One new enum value, not two, and not a separate setting.** Fair-share = `WeightedShuffle` with weights
left at their default of 1, so the UI's "fair-share toggle" needs no second mechanism. A separate
non-enum "distribution" setting was rejected: it would add an orthogonal axis every `PlaybackOrder`
dispatch site must *also* consult, multiplying the blast radius below. Retrofitting weights onto
`ShuffleInOrder` was rejected: it would silently change shipped users' output.
- **Weights live on `MultiCollectionItem` AND `MultiCollectionSmartItem`** (`int Weight`, DB default 1,
dual-provider migration `AddMultiCollectionItemWeight`). The mirror is mandatory — omitting the smart
item un-weights smart-collection members silently. The DB default must be **1, not 0**: a 0 backfill would
hand every pre-existing member to the rotation carrying a weight that means nothing on a share-of-airtime
scale. (The enumerator clamps such rows to the floor so they rotate fair-share rather than vanish — but the
backfill should be right at the source.)
Rejected carriers: `PlaylistItem` (sequential engine; its `Count`/`PlayAll` already mean "how much of this
source", and drain-N-consecutively `A A A B` is a different product from smooth interleave `A A B A`);
`CollectionItem` (wrong grain — per media item, and the largest table).
- **Applied on the `ShuffleInOrder`-shaped path only, because that is where source identity survives.**
`MultiCollectionGrouper` collapses sources into `GroupedMediaItem` + `.Distinct()`, so by the time the
`Shuffle` path has its list, *which source an item came from* is gone. `CollectionWithItems` gains
`int Weight = 1`; `WeightedShuffleCollectionEnumerator` consumes
`ShuffleSourceBuilder.GetCollectionItemsForShuffleInOrder` **unchanged** — the schedule-entity-free entry
point #380 reserved for this issue (no signature change, no god-factory, engines stay separate).
- **Stateless.** Smooth weighted round-robin (`acc += weight`; richest wins; pays the total) is a pure
function of (`Seed`, `Index`), so it restores by replay like its siblings and needs no per-source counters
(`CollectionEnumeratorState` has nowhere to put them). A rotation is sized so the source needing the most
picks works through all its items at its share; smaller sources loop inside it — that looping is what makes
equal weights mean equal airtime. **Weight ratios are data, not state**: a `Reset` build re-randomizes
`Playout.Seed` (fresh within-source shuffle and start point) but the ratios hold identically.
- **Ties break to the earliest source in list order** (strict `>`), keeping the sequence deterministic.
- **`ScheduleAsGroup` is not read by this order**, unlike `ShuffleInOrder`, which merges every
non-`ScheduleAsGroup` collection into ONE source. Here each `CollectionWithItems` is its own weighted
source, which is the entire point — per-source weights are meaningless if sources are merged. Recorded
because it means a persisted per-item flag quietly has no effect under this order.
- **Within-source order is shuffled, not chronological** (custom-ordered collections are still honored).
Chronological inner order would make every rotation byte-identical, since source selection is
deterministic — the reseed on wrap would change nothing and the avoid-an-immediate-repeat retry could
never succeed. "Weighted **in order**" would be a separate order if wanted.
- **Cross-engine exposure is closed by write-path validation, not by touching the silent fallbacks.** An
unhandled `PlaybackOrder` fails **silently at five of six dispatch sites**: Classic substitutes
`RandomizedMediaCollectionEnumerator` (`PlayoutBuilder`'s `default:`, `// TODO: handle this error case
differently?`); `PlaylistEnumerator` has **no `default:` arm**, so the enumerator stays null and the item is
dropped from the playlist (and null is *legitimate* there for `SeasonEpisode` with `Count == 0`, so nothing
flags it); `BlockPlayoutBuilder` filters block items against an allow-list and `continue`s past the rest;
YAML/Scripted return `Option.None`, which their callers read as "no content". Only `MultiCollectionGroup`
throws. **A weighted order degrading to unweighted random is the worst case, because the output is supposed
to look arbitrary.** So the write path **rejects** `WeightedShuffle` everywhere it isn't handled — if it
can't be persisted where it isn't handled, the silent sites never see it — and YAML/Scripted (which address
orders by *name*, so `Enum.Parse` accepts it regardless) log a warning when an order falls through
unhandled. Making those pre-existing fallbacks loud is a real but separate defect class: **#403**, an
explicit non-goal here.
- **Exactly two writers persist a CALLER-SUPPLIED `PlaylistItem.PlaybackOrder`, and the second is not
obvious.** `ReplacePlaylistItems` is the expected one; **`CreateChannelFromLineup`** is the other — a
lineup of 2+ entries is stored as a `Playlist`, and its own guard only covered *MultiCollection* entries,
so a multi-entry lineup of plain collections slipped `WeightedShuffle` straight through to
`PlaylistEnumerator`'s null-drop. Both are gated. The full writer set, since "persisting writer" alone is
the wrong axis: `Add*ToPlaylist` and `TraktCommandBase` **do** persist the field but hardcode it
(`Shuffle`/`Chronological`), so no caller value reaches them; `ReplaceBlockItems` writes the *different*
field `BlockItem.PlaybackOrder` (also gated); and `PreviewPlaylistPlayout`, `PreviewBlockPlayout` and
`Engine/PlaylistHelper` build in memory without persisting.
**This bullet has now been wrong three times, each time in the same shape**: it recorded the perimeter as
complete when the gate covered only the writers already in hand (missing `CreateChannelFromLineup`); the
first correction miscounted by conflating `PlaylistItem` with `BlockItem`; the second said "two persisting
writers" when the true predicate is *two writers that persist a caller-supplied order*. The failure is
always **enumerating from a list instead of re-deriving from a grep** — which is also how the
"0-weight is filtered out" claims survived their own correction. Before trusting any "this order can't
reach that engine" claim, grep every writer of **each field separately** and classify each as
persists-caller-value / persists-hardcoded / in-memory. The non-obvious composite handler is the one that
gets missed.
- **Weight is bounded at the write path** (`MultiCollectionItemWeight`, 1..1000) and clamped again in the
enumerator (`EffectiveWeight`). They do **different** jobs, which is why both stay: **EF's
`HasDefaultValue(1)` substitutes 1 for a `0` on INSERT** (0 reads as "not set") **but an UPDATE writes the
0 through**, so create and update disagreed on the same input — the gate makes them agree and refuses
values that mean nothing on a share-of-airtime scale (0, negative, or absurdly large). The **clamp** is
what makes the rotation arithmetic safe for rows that predate the gate: it bounds every weight before the
sum, so `Sum(weights)` cannot overflow regardless of what is stored. Historical note: an earlier revision
of this PR *filtered* `Weight > 0` instead of clamping, which silently deleted a 0-weight source from the
channel; the clamp replaced it, and any comment still describing that filter is stale.
- **Weight is returned on the multi-collection GET, not just accepted on write.** The update replaces the
item list, so a client that reads, edits a name, and writes back would silently reset every weight to the
default if the read didn't carry it. Edits ride the existing `Version` token (If-Match/412, §7a).
- **Scope:** Classic engine only; no SPA (the schedule-editor weight UI is gated on #388). Plain
single/smart collections (the `GroupIntoFakeCollections` path) carry no persisted weights — fair-share
applies there by construction, and per-source weighting requires a multi collection. Auto-tune's
per-member weights (#385/#386) need their own carrier: an auto-tuned channel is backed by one live
SmartCollection, so its "sources" are members of a single collection, and a `PlaylistItem` cannot even
reference a search query. Nothing here forecloses it — the enumerator reads weights off
`CollectionWithItems` source-agnostically.
## 2026-07-17 — Docs-only CI skip gates STEPS in always-running required jobs, never `if:`-skips them (#416)
A change touching only `docs/**` or `*.md` ran the entire `docker-build.yml` matrix (`test`,
`migrations` incl. its `mysql:8.4` service, `functional-e2e`, `format`, `api-docs`) — ~9 min of warm
CI to validate Markdown. `docker-build.yml` had no path filtering.
**Why not `paths-ignore` / an `if:`-skipped job — the trap.** `main`'s branch protection requires
two contexts *by name* (`Build & test (.NET)`, `EF migration integrity (SQLite + MySql)`). If a
docs-only PR produced **no run** for them, those contexts never report and the PR can **never merge**
— the naive fix bricks docs PRs instead of speeding them. A probe (throwaway PR #418) confirmed that
on Gitea **1.25.4** an `if:`-skipped job reports commit-status state **`skipped`** (a distinct state,
not `success`); how branch protection treats a `skipped` *required* context is not something we rely
on.
**The decision.** Each heavy job (`test`, `migrations`, `functional-e2e`, `build`) runs
`scripts/ci-detect-docs-only.sh` as its first post-checkout step (`id: detect`) and gates every real
step on `if: steps.detect.outputs.docs_only != 'true'`. The job **always runs** and reports
`success` in seconds on docs-only — so the two required contexts report unconditionally (safe by
construction). Non-required jobs may skip freely (production proves a `skipped` non-required context
doesn't block merge — `build` is `skipped` on every PR), so `build` skips its image steps on a
**docs-only push to `main`** (docs aren't in the image); tag builds force `docs_only=false` so a
release is never skipped. `api-docs`/`format` already self-short-circuit; `docs-reminder`/
`decisions-guard`/`ci-image-pin` keep running.
**Detection biases toward running MORE.** `docs_only=true` only when *every* changed path is docs;
any code path, a tag build, a non-merge push, or an undeterminable diff → `false` (run everything). A
false `true` would skip real tests on a code change (a correctness bug); a false `false` merely wastes
CI. Trade-off accepted: `migrations`' `mysql` service still starts on a docs-only run (a `services:`
container starts with the job regardless of step `if:`), but the 787-migration replay — the expensive
part — is skipped.
Two adjacent redundancies are deliberately **out of scope**: the whole matrix re-running on a PR and
again on the merge-to-`main` over identical code (#420), and the within-run triple `dotnet build`
(#398). Full mechanism in `ci-cd.md` → "Docs-only skip".
## 2026-07-17 — Docs-only detect must be shallow-checkout safe: FETCH_HEAD + two-dot, not origin/main + three-dot (#416 follow-up)
The #416 docs-only skip shipped (#422) safe but **ineffective**: every docs-only PR still ran the full
matrix. Root cause — `test`/`migrations` check out `fetch-depth: 1`, and in a shallow clone
`origin/<base>` has no remote-tracking ref and there is no merge-base, so the detect script's
three-dot `git diff origin/main...HEAD` errored → `|| true` → empty diff → the fail-safe returned
`docs_only=false` → full matrix. Confirmed in a real shallow `file://` clone (`origin/main` did not
resolve; three-dot errored; `git diff FETCH_HEAD HEAD` returned the changed files correctly).
Decision: the detect diffs against **`FETCH_HEAD`** (always written by `git fetch`, resolves in a
shallow clone) with a **two-dot** tree diff (no merge-base). `api-docs`/`format` were unaffected only
because they use `fetch-depth: 0` — a difference the first cut missed. Meta-lesson reinforced: a CI
gating change can pass every local test and merge green while being a complete no-op in CI; only
real-PR verification that **measures the effect** (job durations, not just a green check) catches it —
which is exactly what #416's Done-when demanded. Fixed in the #416 follow-up PR.
## 2026-07-17 — Pre-push guard: don't push a file whose working-tree copy is uncommitted (H13, #416 session)
A review fix (`--no-renames`) was edited into the working file and empirically verified there, but a
`git reset --soft` + `git add` then committed the *stale index*, leaving the fix as an uncommitted
working-tree diff. The push, the CI run, and a cold reviewer each saw a **different tree**: CI/push
had the OLD code; the reviewer read the working file and "confirmed" a fix that never shipped. A PR
went out still carrying the bug the review had cleared. Root cause: local build/test/review all
operate on the working tree, but what ships is the *committed* tree — nothing enforced that they
match.
Decision: a fail-open **pre-push hook** (`.claude/hooks/prepush-clean-worktree-check.sh`, wired into
`.husky/pre-push` after the H11 rebase check) blocks a push when a file that is part of the branch's
diff vs `origin/main` **also** has uncommitted working-tree or index changes. Scope is deliberately
precise — only files in the pushed diff, so unrelated uncommitted scratch (or untracked files) never
false-block. Fail-open on anything undecidable (not a repo, offline, no `origin/main`); deliberate
escape `ETV_ALLOW_DIRTY_PUSH=1`. This is the mechanized half of the working-tree-vs-committed lesson;
the review-process half (point reviewers at `git show <sha>:<file>`, never the bare working file) stays
guidance. Same hook family as H11 (rebase-not-merge) / H12 (issue-qualification), same "#303 make the
rule a hook, not prose to remember" throughline. Verified: dirty PR-file → block; clean tree → allow;
dirty non-PR file → allow; escape hatch → allow.
## 2026-07-17 — Auto-Tune per-channel overrides reuse the Channel Builder advanced-options DTO; weights + bug-colour logo split out to #425 (#385)
The Auto-Tune DetailPanel (#383) makes each proposed channel individually editable before bulk-create.
The backend for that (#385) split cleanly along a "structural cost" line, and only the additive half
shipped here; the rest is deliberately deferred rather than forced.
- **Shipped now (additive, over the frozen `/api/v1`):** `POST /api/v1/channels/auto-tune`'s
`AutoTunedChannelRequest` gained three optional per-channel fields — `templateId`, `advanced`, and an
uploaded `logo` image. Omitting each keeps PR1 behavior exactly (batch template, axis-derived playback
order, on-the-fly fallback logo), so the change is backward-compatible and the older positional
`{axis, value, name, number}` form still binds.
- **`advanced` reuses the manual Channel Builder's `CreateChannelFromLineupAdvancedOptionsRequest`
verbatim** — the same 24-field override set, the same `ToCommand()`, and the same
`advanced.X ?? template.X` stamp-at-create contract that `POST /api/v1/channels/from-lineup` already
honors. Auto-tune is now just a second caller of that wire contract; we did **not** mint a parallel
24-field DTO. The one auto-tune-specific rule: the axis default (`SeasonEpisode` for a single show,
`Shuffle` for a genre) fills `PlaybackOrder` only when the caller left it null, so an explicit
per-channel order wins but an unset one is never dropped.
- **`logo` is the uploaded channel image**, forwarded to `CreateChannelFromLineup.Logo` (persisted as an
`ArtworkKind.Logo` artwork) and run through `ArtworkContentTypeModel.Sanitized()` at the request
boundary — the same stored-XSS defense (#283) the Channel Builder uses; a hostile content type never
reaches the command.
- Per-channel independence is preserved: `templateId`/`advanced`/`logo` are resolved per channel inside
`CreateAutoTunedChannelsHandler.CreateOne`, so one channel's bad override still yields a per-channel
`Failed`/`Skipped` outcome without aborting the batch.
- **Deferred to #425 — per-source rotation weights + query corrections (exclude / add-untagged).** This is
a structural redesign, not effort: #70's `WeightedShuffle` reads weights only off
`MultiCollectionItem`/`MultiCollectionSmartItem` rows, but an auto-tuned channel is backed by a **single
SmartCollection** (→ `GetFakeMultiCollectionCollections`, every fake group forced to `Weight = 1`), which
gives fair-share but **cannot express differing per-source weights**. Honoring them requires auto-tune to
build a **MultiCollection of per-source SmartCollections** (one query per weighted source), which
contradicts the documented "one live query per channel" design and lands new structure in the scheduling
subsystem. That is the `[PLAN-MODE]` design call the #385 body flagged; it belongs in its own issue with
a recorded design decision, so it moved to **#425** (consumed by the SPA #386).
- **Deferred — bug initials + fallback colour generated logo.** The generated logo (`/iptv/logos/gen`) is a
stateless on-the-fly render used by both the XMLTV `<icon>` and the FFmpeg `WatermarkSelector` bug, with
**no persisted bug-initials/colour state on `Channel`**. Making it configurable needs either new `Channel`
columns (a dual-provider migration) wired through the watermark pipeline, or create-time artwork-file
materialization — neither is the additive plumbing this slice was scoped to, and it is a channel-wide
capability rather than an auto-tune concern. Left for its own issue / the SPA branding work; the uploaded
`logo` above already covers the on-screen bug for channels that supply an image.
## 2026-07-17 — Health-check remediation is server-declared `{Kind, Target}` on an additive DTO; the SPA acts on it (#164)
#164 asked to make the ~14 health checks *actionable* — the Dashboard health panel showed problems
with no way to investigate or fix them. Two structural decisions came out of it.
**Remediation is server-declared metadata, not SPA-derived.** Each check that has a fix knows where the
fix lives, so the *check* declares it. The domain `HealthCheckLink` grew from `(string Link)` to
`(string Target, HealthCheckLinkKind Kind)` with `Kind ∈ {ExternalDoc, AppRoute}` and two factories
(`HealthCheckLink.ExternalDoc(url)` / `HealthCheckLink.AppRoute("/app/...")`). Only the 4 checks that
built links and the API mapper touched `.Link`, so the widening was local. The SPA then *acts* on the
kind: `AppRoute` → client-side `navigateToPath(target)` button; `ExternalDoc` → new-tab anchor. The
human label is derived SPA-side from the route (a small lookup + prettified fallback) rather than sent
over the wire — keeping the DTO minimal.
**The DTO evolved additively (`/api/v1` is frozen-additive, #286).** `HealthCheckResponseModel` kept
its existing `Detail` and gained `Brief` (← the domain `BriefMessage` the old mapper silently dropped)
and `Remediation { Kind, Target }` (a nested `HealthCheckRemediationResponseModel`). The old flat
`string? Link` is **kept and still populated** (mirrors `Remediation.Target`) but documented deprecated —
we don't remove a frozen field, and existing consumers keep working. `Remediation.Kind` is a plain
string ("ExternalDoc"/"AppRoute") mapped in the Application `Mapper` exactly like `Status`
("pass"/"fail"/…), not a wire enum — matching the established pattern for that DTO.
**Three defects the audit surfaced, fixed here.** (1) The Application `Mapper.GetStatus` threw
`ArgumentOutOfRangeException` on `NotApplicable`; the handler filters `NotApplicable` before mapping so
it was latent, but the mapper is now **total** (defense-in-depth — a future caller that skips the filter
can't 500 the endpoint). `InternalsVisibleTo("ErsatzTV.Tests")` was added to the Application assembly
(mirroring Core's precedent) to unit-test that totality directly. (2) Two checks linked to **stale
Blazor routes** (`media/trash`, `search?query=…`) — repointed to the SPA `/app/trash` and
`/app/search?query=…` as `AppRoute`s. (3) A dead `Open Classic UI` → `/system/health` link lingered in
`SettingsScreen` (a #91b leftover that just 302'd to `/app`); removed (see `blazor-route-parity.md`
Section 4 correction).
**Actionable checks that had no link gained an `AppRoute`** (metadata → `/app/libraries`, empty
schedules → `/app/schedules`, HW-accel / VAAPI → `/app/ffmpeg-profiles`, FFmpeg reports → `/app/settings`).
Pure-noise / no-clean-action checks (UnifiedDocker, MacOsConfigFolder, FFmpegCapabilities, the Info-tier
nags) were left untouched — semantic-tier changes (e.g. adding a Pass path, demoting a nag) were
deliberately **not** bundled into a remediation-UX PR.
**Deferred (own issue): a TTL cache for `PerformHealthChecks`** (#108 — every `GET /api/v1/health`
re-runs all 14 checks, 4 shelling out to ffmpeg, and the existing summary cache is write-only dead
code). Orthogonal to the UX; filed separately so a SPA-polled health panel gets a cache before it
polls.
## 2026-07-18 — Auto-Tune DetailPanel SPA: reusable `SlideOver` + shared advanced-options model; decorative panes dropped to match the backend (#386)
The SPA half of the Auto-Tune per-channel editor. It builds only what the shipped `/api/v1` surface
(#384 members read, #385 per-channel `templateId`/`logo`/`advanced`) can actually carry, so the panel
never presents a control with nowhere to send its value.
- **New reusable primitive `SlideOver`** (`web/src/components/overlay.tsx`), sharing one
`useOverlayBehavior` hook with `Dialog` (focus/scroll-lock/Escape/scrim). Right-edge panels are a
recurring need; this is the seam, not a screen-local drawer. See spa-conventions §11.
- **Advanced-options logic is shared, not re-implemented** — the `AdvancedPanel` internals in the 84 KB
`ChannelBuilder.tsx` were extracted to `web/src/builder/advancedOptions.tsx` (enum catalogs,
`ADVANCED_KEYS`, `effectiveValue`, and the INHERIT/omit `useAdvancedOverrides` hook). ChannelBuilder now
imports them with zero behavioral change (its test suite passes byte-for-byte); the DetailPanel writes
its own field JSX over the same hook. The #135 "inherit = omit, no None option" contract therefore has
a single implementation. Field *layout* stays per-screen (presentation, not logic) — the deliberate,
stated deviation from full component reuse.
- **Shipped panes:** identity (name/number + `uploadArtwork(_, 'logo')` dropzone), Playback (Shuffle →
`advanced.playbackOrder`, Always-playing → `advanced.playoutMode`, merged into `advanced` only when
diverged from the template), per-channel template picker, the Advanced disclosure, a lean read-only
Query&size (order-from-axis + est items + effective streaming mode), and a read-only Content-sources
member list via `GET /members`.
- **Dropped as backend-less decoration** (the prototype had them; the wire contract does not): the MiniEpg
"example schedule", the bug-initials/colour generator, and the "generated smart-collection query" text —
the proposal DTO carries no query string, so rendering one would be fabricated. **Deferred to #425** (its
own end-to-end slice): the per-source rotation-weight steppers and exclude/add-untagged corrections; the
Content-sources pane shows a "#425" hint so its read-only state reads as intentional.
- **Unsaved-changes guard is screen-scoped** (spa-conventions §8/§11): per-channel edits live in
`AutoTuneScreen` draft state until bulk-create, so closing the panel keeps them (an "Edited" row badge
makes that visible) and only screen navigation / full-page unload with uncommitted edits confirms.
## 2026-07-18 — SmartCollection rule builder: compile-only closed subset, no stored AST, one-level nesting (#176)
The SmartCollection create/edit dialog gained a visual rule builder (`web/src/builder/rules/`)
alongside the existing raw-Lucene textarea. **The SmartCollection still stores a plain Lucene query
string — no new stored rule AST, no schema change.** The builder compiles its in-memory rule tree
into a **closed subset** of the Lucene grammar (`compile.ts`) and parses exactly that subset back out
(`parse.ts`, the exact inverse — returns `null`, not a best-effort guess, for anything outside the
subset); escaping is total, so any builder-authored query round-trips losslessly, proven by a 500-tree
property test (`roundtrip.test.ts`, including Lucene special characters). Opening an existing
SmartCollection tries the parse first and falls back to raw-text mode on `null` (fuzzy queries,
boosts, mixed AND/OR at one nesting level, or nesting deeper than one level).
**Why compile-only over persisting an authoritative rule AST**: an AST would still need a
Lucene→rules parser to open every *pre-existing* free-text query — including every query the
Auto-Tune feature (#69) generates — so the AST would buy almost nothing (it still can't represent
arbitrary hand-written Lucene) while costing a dual-provider EF migration and a second source of
truth to keep in sync with the Lucene grammar. Compile-only keeps the query string as the single
source of truth and treats the builder as a structured *editor* over it, not a new storage model.
**One-level-nesting "Kodi" model.** `types.ts` defines a top `Group` (`match: all|any`) over `Rule`s
and/or **one level** of sub-`Group`s — enough to express `type:movie AND (genre:Horror OR
genre:Thriller)`, which covers the smart-playlist patterns Kodi-style rule builders are known for.
Arbitrary/recursive nesting was scoped out as YAGNI; revisit only if a real query needs it.
**Field vocabulary comes from a new catalog endpoint, not a hardcoded list.** `GET
/api/v1/search/fields` (read-only, MCP-introspectable; see `api-conventions.md`) returns the curated,
typed, labeled field set derived from `LuceneSearchIndex` — name/label/type/group/values — and is the
single source of truth the builder's field pickers (`fieldCatalog.ts`'s `useSearchFields` hook) and
operator/value-input choices are driven from, so the builder's vocabulary can't drift from what the
index actually supports.
**Deferred as separate follow-up issues** (explicitly out of scope for #176): facet-value typeahead
for value inputs, relative-date operators, nesting deeper than one level, and inline adoption of
`RuleBuilder` by ChannelBuilder / Auto-Tune (it was built reusable for exactly that reuse — see
`spa-conventions.md` §12).
## 2026-07-18 — Auto-Tune per-source weights ride #70's MultiCollection machinery; created at tune time, not a post-hoc PUT (#425)
Per-source rotation weights (`3× Show A, 1× Show B`) and query corrections (exclude / add-untagged) for
an auto-tune channel are supplied **at bulk-create time** — an optional `sources: [{sourceId, weight,
excluded}]` list on each `AutoTunedChannelRequest` (the same DetailPanel-wizard surface #385 added its
per-channel `advanced`/`templateId`/`logo` to). There is **no** post-hoc `PUT .../weights`, no lazy
upgrade/downgrade, and no idempotent desired-state handler: the auto-tune DetailPanel (#383/#386) is a
create-wizard, and #425's Done-when is "create → built playout". (An earlier plan draft assumed a PUT on an
existing channel; the shipped flow is create-time, matching #385.)
**Structure — Option A (reuse #70), not a new fake-collection weight path.** When any source is customized
(a non-default weight, an exclusion, or an added out-of-axis id), the channel is backed by a system-owned
`MultiCollection` of per-source `SmartCollection`s, weights on `MultiCollectionSmartItem.Weight`, and its
single-item lineup points at the `MultiCollectionId` with `PlaybackOrder.WeightedShuffle` — the exact path
`WeightedShuffleCollectionEnumerator` already consumes (`GetMultiCollectionCollections` forwards the per-row
weight). Rejected: threading a weight map through `GroupIntoFakeCollections` (the fake-collection path a lone
SmartCollection takes, which hardcodes weight 1) — it derives its group keys at runtime, is shared with the
unrelated `FillWithGroupMode`, and would need a parallel weight-persistence home. When **no** source is
customized (all weights 1, nothing excluded/added) the channel keeps the #69 single-SmartCollection
fair-share shape — cheaper, and identical output for TV since the fake path already groups per show.
**Discriminators.** A source's member query is discriminator-only (membership is fixed at tune time): TV uses
`type:episode AND show_title:"X"` (episode docs carry no parent-show id in the index — `show_title` is the
only per-show field, the same one #69's TvShow axis already uses; a post-create rename empties the member
until re-tuned), and movies use the stable, rename-proof `type:movie AND id:{mediaItemId}` (a movie is the
played item). The remainder is `({base}) AND NOT ({d1} OR {d2} …)` over every materialized excluded
discriminator — a partition of the base set, so no item is counted twice or dropped. Exclude = omit the
member **and** keep it in the NOT-list (else its items leak back through the remainder); add-untagged = an
ordinary member whose id isn't in the base set (no genre clause, so it just airs).
**Materialization bound is axis-dependent — "weight N" means N× *each* other source.** TV materializes
**every** base show as its own member (un-weighted shows keep per-show fair-share; a single merged remainder
would regress them to item-proportional, so a 200-episode show would swamp a 20-episode one) plus one live
remainder at weight 1 for shows/episodes added after tune-in (empty at create → harmlessly skipped by
`WeightedShuffleCollectionEnumerator`'s `ActiveSources` filter). MovieGenre materializes **only** the touched
movies (weighted or added) and leaves every un-touched base movie in ONE remainder whose weight = its member
count — exactly equivalent to materializing each movie individually, because the fake path already pools
movies uniformly, without hundreds of rows. Cost note: a large TV genre materializes one SmartCollection per
show (dozenshundreds), each a cheap stored term query; a future optimization could cap/warn.
**Ownership + lifecycle.** The MultiCollection and its member/remainder SmartCollections carry a nullable
`OwnedByChannelId` (dual-provider migration `AddCollectionOwnedByChannelId`, indexed). Owned rows are hidden
from the user-facing collection lists and cascade-cleaned when the channel is deleted (the
`ProgramScheduleItem → MultiCollection` cascade removes the dangling flood item). Names embed a per-create
token (`at-mc:{token}`, `at:{token}:{n}`) because both names are unique `varchar(50)` and the channel id
isn't known until `CreateChannelFromLineup` runs; ownership is stamped immediately after. Non-weighted (#69)
channels set no ownership, so their pre-existing orphan-on-delete behavior is unchanged. The create is
non-atomic across the two handlers (mirrors #69) with best-effort rollback of the artifacts on channel-create
failure.
## 2026-07-18 — Search all-items is paged to cap DoS exposure; SPA add-all pages to completeness (#293)
`GET /api/v1/search/all-items` (`SearchController.SearchAllItems` → `QuerySearchIndexAllItemsHandler`) fired
ten index searches with **`limit: 0`** (= "return every hit", `LuceneSearchIndex` line ~244), so a single
broad query (e.g. one matching the whole library) materialized *every* matching doc across all ten media
kinds into ten `List<int>` buckets and serialized them in one response — unbounded work per request. #285
closed the original *unauthenticated* exposure (the endpoint is now behind `Api:RequireKeyForReads`, default
true); the residual was DoS-hardening against an **authenticated** caller with a very broad query. Deferred
from #285 because the SPA "add all to collection/playlist" flow materializes the full id set before the add
POST, so a naive hard cap would silently truncate "add all".
**Decision (issue option (a), operator-confirmed): paginate the endpoint and teach the SPA add-all flow to
page to completeness** — rather than option (b) (a generous cap + truncation signal). Chosen because
"add all" must stay complete for real use, and it matches the sibling `GET /api/v1/search` /
`GET /api/v1/channels/auto-tune/members` (#384) paging convention already in the codebase.
- **Endpoint (additive).** `SearchAllItems` gains optional `pageNum` (0-based) + `pageSize`, clamped exactly
like the §1 Logs / sibling `Search` precedent: `pageSize = Math.Clamp(pageSize, 1, MaxAllItemsPageSize)`
with `MaxAllItemsPageSize = 1000`, `DefaultAllItemsPageSize = 500`. `pageNum` is clamped
`Math.Clamp(pageNum, 0, MaxAllItemsPageNum)` with `MaxAllItemsPageNum = 2_000_000` — the upper bound keeps
`pageNum * pageSize` (the search skip) inside `int` range so an absurd page number can't overflow to a 500
(the sibling `Search` only floors at 0; the all-items endpoint hardens the upper bound too since this is a
DoS-hardening change). The clamp is applied per media kind (a page returns ≤ `pageSize` ids of
*each* of the ten kinds), so one response is bounded to ≤ 10 × `pageSize` ids. `QuerySearchIndexAllItems`
carries `PageNum`/`PageSize`; the handler passes `skip = PageNum × PageSize`, `limit = PageSize` into
`ISearchIndex.Search` (native skip/limit) and reads `SearchResult.TotalCount` (the true total, free) per
kind.
- **Response (additive, frozen-v1-safe).** The ten `…Ids` buckets are unchanged; a new non-null nested
`Totals` (`SearchResultAllItemsTotalsResponseModel`, ten `…Count` ints) is added so a client knows how many
ids exist per kind and can page to completeness. Nothing is removed or retyped (#286 additive-only holds).
- **Deliberate default-behavior change.** A caller that sends no `pageSize` now gets one page (default 500 /
kind) plus `Totals`, not the entire id set. This is the security change the issue asks for; it is safe here
because the only in-repo consumer is the SPA (updated in the same PR) and any external/MCP caller can read
`Totals` and page. Recorded as intentional, not a regression.
- **SPA pages to completeness.** `web/src/api/search.ts` `getSearchAllItems(query, pageNum, pageSize)` gains
the params; a new `getAllSearchItemIds(query)` loops pages (requesting `pageSize = 1000`, the server max),
accumulating every bucket until each kind has collected its `Totals` count (with an empty-page safety break
against total-count drift), and returns the merged `SearchAllItemIds`. `SearchScreen.addAll` calls it
instead of the single-shot fetch; the #221 stale-query guard and the single add POST are unchanged.
- **Out of scope (unchanged):** the add POST itself still accepts the full merged id set in one request body
— bounding *that* surface is a separate concern (see #308 for the add path); #293 is the GET.
## 2026-07-18 — Collapsible sidebar + nav-group accordions: two `ctv-sidebar-*` localStorage keys, labeled groups default-collapsed (#396)
The shell sidebar (`web/src/app/AppShell.tsx`) gained (a) a header toggle that collapses it to a 60px
icon rail and (b) collapsible accordions per **labeled** nav group (Media, System); the unlabeled
**Primary** group is always open. Mirrors the Claude Design prototype's updated `Sidebar`.
- **State lives in a small hook, not App.** `web/src/app/sidebarState.ts` `useSidebarState()` owns both
pieces of state + persistence; `AppShell` consumes it (nothing else needs it) and stamps
`ctv-app-shell-collapsed` on the shell root so the collapse is CSS-driven from one class.
- **Persistence keys use the established `ctv-` hyphen convention, NOT the prototype's dotted names.**
The issue quoted `ctv.sidebar.collapsed` / `ctv.sidebar.groups`, but every existing client-local pref
is hyphenated (`ctv-theme`, `ctv-logs-page-size` — spa-conventions §5d), so we use
**`ctv-sidebar-collapsed`** (`"1"`/`"0"`) and **`ctv-sidebar-groups`** (JSON `{groupKey: boolean}`,
boolean = *collapsed*). Deliberate deviation from the issue's literal key text in favour of the repo
convention the issue itself points to; helpers validate/parse defensively (bad JSON / non-boolean
values → default).
- **Labeled groups default to COLLAPSED** (absent `ctv-sidebar-groups` entry ⇒ collapsed), so a fresh
load shows only Primary — matching the prototype ("default-collapsed, leaving only Primary visible").
A behavior change for existing users; `App.test.tsx`'s shell/nav suite seeds the two groups open
because it clicks Media/System nav links directly (the accordion behavior is covered in its own
describe).
- **Accordions apply only in the expanded sidebar.** In the rail, group-collapse is ignored — every
item renders as an icon (label kept in the a11y tree via an sr-only span so the accessible name/tests
survive; surfaced as a native `title` tooltip), groups separated by a hairline divider, numeric
badges shown as a corner dot. The active-route indicator (left rail bar + active background) works in
both states.
- **Group keys are explicit + stable** (`SidebarNavGroupDefinition.key`: `'media'`, `'system'`) rather
than derived from the label, so renaming a label doesn't silently orphan persisted state.
- No route/screen was added or redirected (shell-chrome only), so no `blazor-route-parity.md` change.
## 2026-07-18 — Unsupported PlaybackOrder is loud at build time; a declared support matrix and tripwire test make new orders safe by construction (#403)
`#70` closed the *persistence* hole for `WeightedShuffle` (the write path rejects it on the engines that
can't handle it) and made **YAML + Scripted** log a warning; `MultiCollectionGroup` already threw. It left
the three still-**silent** build-time dispatch sites — the ones this issue names — as an explicit non-goal.
`#403` makes those loud and adds the by-construction net.
- **Loud, but NON-FATAL, at every build site.** On an order it doesn't handle, each site now logs a
`Warning` (naming the order + engine + the fallback taken) instead of silently substituting/dropping/
skipping: Classic's `PlayoutBuilder` `default:` arm (was `// TODO`, silently returned Random), the
`PlaylistEnumerator.Create` switch (had **no** `default:` arm → null enumerator → item dropped), and
`BlockPlayoutBuilder`'s allow-list `continue` (silent skip). The actual fallback is **preserved** — a
live channel airing wrong-but-something beats going dark on one misconfigured item — so scheduler goldens
do not move. This matches the log-and-continue posture `#70` already shipped for YAML/Scripted.
- `PlaylistEnumerator.Create` was `static` with **no logger**, which is *why* the playlist drop was
unreportable. It gained an optional `Option<ILogger> logger = default` (mirrors the existing
`GetStartTimeAfter(..., Option<ILogger>.None)` pattern); the three callers that have a logger pass it,
the rest default to `None`.
- `BlockPlayoutBuilder`'s `GetEnumerator` switch gained an explicit `PlaybackOrder.Random` arm. Random is
in Block's allow-list but previously reached an enumerator only via the switch's coincidental `_ =>`
Random fallback; the fallback is now a loud helper (defensive — the allow-list should make it
unreachable).
- **Safe-by-construction net: a declared support matrix + a tripwire test.** `PlaybackOrderSupport`
(`ErsatzTV.Core/Scheduling/`) declares, per `SchedulingEngineKind` (Classic / Playlist / Block / Yaml /
Scripted), BOTH a `Supported` and an `Unsupported` set. `PlaybackOrderSupportTests` asserts the two sets
**partition** every `PlaybackOrder` value (union total, intersection empty) for every engine — so adding a
new enum value lands in neither set and **fails the test until it is consciously classified** and wired
into the matching dispatch switch. Both sets are hand-maintained on purpose: deriving `Unsupported` as
"everything not supported" would let a new order fall through silently, which is the exact defect being
killed. The matrix is not pure test scaffolding — `BlockPlayoutBuilder` consumes it at runtime for its
allow-list (replacing the previously duplicated inline list).
- **Write-path rejection is deliberately UNCHANGED.** `#70`'s per-engine guards already close the real
persistence exposure, and the decisions log records that the "which order reaches which engine" perimeter
has been wrong three times — always by enumerating from a list instead of grepping every writer.
Broadening write-path rejection here would risk newly-`400`ing configs that currently save-and-fall-back,
for no safety gain, so it was left alone. This PR hardens the *build-time* sites and adds the tripwire.
- **Reverse-mappings deferred (a different axis).** The three `_ => PlaybackOrder.None` arms that map an
enumerator *class* back to a `PlaybackOrder` for `PlayoutHistory` were left alone: that path is reached in
normal operation by legitimately-supported enumerator types (`ShuffleInOrderCollectionEnumerator`,
`SeasonEpisodeMediaCollectionEnumerator`) that simply aren't reverse-mapped, so making it "loud" would
emit false-positive warnings. Completing that reverse map is a separate concern from "an unsupported
*order* degrades silently" and is not part of #403's scope.
## 2026-07-18 — CI build-once was measured and rejected; keep the #420 tree-skip
Build-once (a `compile` job producing a single artifact, consumed by `test`/`migrations`/
`functional-e2e` via `--no-build`) was fully implemented and went **green on CI** (PR #455, run
830), then **rejected on measurement**: it traded a ~12% slot-occupancy saving for a ~4085%
**per-run wall-clock regression**.
- **Why it regressed.** The `bin`+`obj` artifact is 2.5 GB raw / 972 MB gz; tar alone costs ~82s CPU
plus ~180s transport, consuming most of the ~465s the shared compile was meant to save. Worse,
`compile` serializes **before** `migrations`' long ef-replay, which is pure DB work a shared build
cannot shorten — the bottleneck was never the redundant compiles.
- **Incidental finding worth recording.** `actions/upload-artifact@v4` does not work on this Gitea
instance — it throws `GHESNotSupportedError`, because the `@actions/artifact` v2 client library
rejects any non-`github.com` host. `@v3` is required for any future artifact use here.
- **Kept: the #420 cross-run tree-identity skip** (`docs/ci-cd.md` → Cross-run tree-identity skip).
It is independent of build-once — it only *skips* redundant work on identical-tree main pushes, at
zero wall-clock cost, rather than trying to share a build across jobs. Don't re-attempt build-once
unless the runner's artifact storage or network changes materially.
Refs: #398 (closed), #420, PR #455.
## 2026-07-18 — Never-scanned `LastScan` surfaces as null at the API boundary, not the 0001-01-01 MinValue sentinel (#409)
`Library.LastScan` / `LibraryPath.LastScan` are `DateTime?`; a never-scanned library is `null` at
runtime for a freshly-created row. But the `0001-01-01 00:00:00` MinValue sentinel still appears in the
DB from **two** sources — and the second is ongoing, not historical:
1. Old `Reset_*` migrations wrote it via raw SQL (`UPDATE Library SET LastScan = '0001-01-01 00:00:00'`).
2. **Live code still writes it today**: `MediaSourceRepository` sets `library.LastScan =
SystemTime.MinValueUtc` on the Plex/Jellyfin/Emby remove-and-recreate (disable-sync) flows
(`MediaSourceRepository.cs:480/611/976`). So the sentinel keeps being written during normal use.
`GetAllMediaSourcesForApiHandler` projected `l.LastScan` straight onto its `DateTime?` DTO, so that
sentinel leaked to API/MCP clients as a fake midnight timestamp — the SPA papered over it with a
client-side year<1900 heuristic (#409 first pass). Decision: the API is the right place to be honest, so
**never-scanned reports as null** for API/MCP parity with the UI, via two layers:
- **Read-boundary coercion — the load-bearing, ongoing guard** (provider/history-independent):
`GetAllMediaSourcesForApiHandler.NormalizeLastScan` maps any `< 2000-01-01` value to null. Because
source #2 above keeps writing the sentinel, this coercion is *permanent*, not a stopgap — a
migration-only fix would regress the next time a user toggles a library's sync off.
- **Data migration** (`NullOutNeverScannedLastScan`, dual-provider): a one-time cleanup of the historical
residue — `UPDATE Library/LibraryPath SET LastScan = NULL WHERE LastScan IS NOT NULL AND LastScan <
'2000-01-01'`. Data-only (empty `Up`/`Down` otherwise, both `TvContextModelSnapshot.cs` byte-identical).
The `< '2000-01-01'` predicate matches the `0001-01-01` sentinel robustly on both providers (ISO-text
compare on SQLite, whatever the out-of-range zero date stored on MySQL — where no `Reset_*LastScan`
migration ever ran, so it's a safe no-op there) without depending on the exact stored bytes; no real
scan predates ErsatzTV. `Down` is a no-op — the original sentinel is unrecoverable and worthless.
`GetAllMediaSourcesForApiHandler` is the **only** API/MCP-facing consumer of `LastScan` (grepped
`LastScan` under `ErsatzTV.Application/**/Queries` and `ErsatzTV/Controllers`); the other reads are
internal scanner code that coalesces to `SystemTime.MinValueUtc` for its own non-nullable
`DateTimeOffset` scan-comparison needs and never serializes it to a client. The DTO field was already
`DateTime?`, so the OpenAPI schema is unchanged (no regen). The SPA's `hasScanned` heuristic in
`LibrariesScreen.tsx` was removed — both call sites revert to a plain null/truthy check now that the API
is honest. (Follow-up option, not done here: have `MediaSourceRepository` write `null` instead of
`MinValue` so the data is clean at rest too; the read coercion makes that non-urgent.)
## 2026-07-19 — WeightedShuffle SPA: weights edited on the multi-collection, order offered only on classic MultiCollection schedule items; fair-share is a reset not a mode (#404)
The UI half of #70 (backend + API shipped in PR #402). No new endpoint or DTO — `weight` was already on
`MultiCollectionItemRequest`/`…ResponseModel` and `WeightedShuffle` already in the `PlaybackOrder` enum; this
is purely SPA (+ docs).
- **Two surfaces, split by where each datum lives.** The per-source *weights* are a property of the
`MultiCollection`, so they are edited in the multi-collection editor (`/app/multi-collections`), not on the
schedule item. The *order* (`WeightedShuffle`) is a property of the schedule item, so it is added to the
classic schedule item's Playback Order options. Putting weights on the schedule item was rejected: the same
multi-collection can be referenced by many schedule items, and weights are shared, not per-reference.
- **`WeightedShuffle` is offered ONLY for `collectionType === 'MultiCollection'`** (`itemRules.ts`
`MULTI_COLLECTION_ORDERS`) — a *meaningfulness* scope, not a rejection mirror. Weights only exist on
multi-collection members, so the order only does something with 2+ weighted sources; on a single collection
the classic engine still **accepts** it (verified by live-E2E: `PUT /schedules/{id}/items` with a plain
Collection + `WeightedShuffle` → 200) but it degrades to fair-share (≈ `Shuffle`), a confusing no-op — so
the SPA doesn't offer it there. The *rejection* the #70 entry describes is on the separate **playlist and
block** write paths (different engines); the playlist/block editors keep their own order lists (not
`MULTI_COLLECTION_ORDERS`) and already omit `WeightedShuffle`, so nothing extra was needed there.
- **`WeightedShuffle` joins `ShuffleInOrder` in the `fillWithGroup` exclusion.** `fillWithGroupModeEligible`
already excluded `ShuffleInOrder`; WeightedShuffle is excluded for the same structural reason —
`PlayoutBuilder` splits a fill-with-group item into per-group enumerators scheduled one group at a time,
which is incompatible with WeightedShuffle's *whole-multi-collection* per-source share of airtime. Offering
the combination would silently not do what the name says.
- **Fair-share is a "Reset to fair share" button, not a stateful toggle.** decisions.md 2026-07-17 established
that fair-share is not a separate mode — it is `WeightedShuffle` with all weights left at the default 1. So
there is nothing to toggle *on*: the button just sets every source's weight to 1 (disabled when already
there). A persistent toggle was rejected because "fair share" has no stored representation distinct from
"the weights all happen to be 1".
- **Weights round-trip through the editor draft or they silently reset.** The multi-collection PUT replaces
the whole item list, so `weight` is read in `itemsFromMultiCollection` and written in `toItemRequest`;
dropping it would reset every source to 1 on the next unrelated save (rename, add a source). Covered by a
round-trip test. This is the general replace-all-DTO trap now recorded in `spa-conventions.md` §4.
- **Percentages are display-only; the wire format is the integer weight.** A 3:1 weight shows as 75% / 25%,
computed from the clamped weights; the request always carries the integer relative share (1..1000). The
weight field is held as a string in the editor draft so it edits smoothly, clamped to the API validator's
1..1000 on blur and again at save — an out-of-range value never reaches the server as a raw 400. `Input`
gained `min`/`max`/`inputMode`/`onBlur` passthroughs for this (reusable by #425's weight UI).
## 2026-07-19 — Health-check results are TTL-cached; `?refresh=true` forces a fresh run (#431)
`HealthCheckService.PerformHealthChecks` re-ran all 14 checks on **every** call, four of which shell out to
`ffmpeg`/`ffprobe` via CliWrap — so a bare `GET /api/v1/health` spawned ~4 subprocesses per request. The
existing `HealthCheckSummary` cache was **write-only** (populated + published, never read back to short-circuit
a re-run). Harmless while the SPA Dashboard health panel refreshes on-demand only, but a real cost the moment
anything *polls* health (a status widget, an MCP client, monitoring). Split out of #164 as the orthogonal
performance half.
- **A short TTL cache of the full result list lives inside `HealthCheckService`.** A `_memoryCache` entry
(`"healthcheck.results"`, `TimeSpan.FromSeconds(30)`) holds the last `List<HealthCheckResult>`; a non-forced
call returns it directly on a hit, skipping both the 14 checks and the summary `Publish`. Chosen over
"make the existing summary cache read-through" because the API returns the full per-check list, not the
2-int summary — the summary entry (`"healthcheck.summary"`, read by `GetHealthCheckSummary`) is kept as-is
(un-expiring) so its fallback behavior is unchanged.
- **`PerformHealthChecks` gained a `bool forceRefresh` first parameter** (interface signature change; one
implementer, 3 live callers). `forceRefresh: true` bypasses the cache and repopulates it.
- **The refresh surface is an optional `?refresh=` query param on the existing GET**, following the
`?deep=` bool-query-param exemplar (`api-conventions.md` §1/§3b) — additive, backward-compatible, no new
endpoint. `[FromQuery] bool refresh` → `GetAllHealthCheckResultsForApi(Refresh)` → `PerformHealthChecks(request.Refresh, …)`.
The SPA "Refresh health" button calls `/api/v1/health?refresh=true`; the initial/poll load calls the bare
path (cached). A separate `POST …/refresh` endpoint was rejected as unnecessary surface for a read.
- **Who forces vs. who reads the cache:** the API GET poll path reads the cache; the **startup**
`RunHealthChecksService` and the **troubleshooting** support bundle force a fresh run (both want current
state — startup is a cold cache anyway, and a diagnostic bundle should reflect *now*, not a ≤30s-old poll).
The legacy `GetAllHealthCheckResults` handler is dead (no senders) and reads the cache.
- **Thundering-herd on a cold cache was left out of scope** (no request-coalescing lock): polling is sequential
per client and the TTL collapses steady-state load, so at most a handful of exactly-simultaneous cold callers
re-run — a once-per-30s edge, not the repeated per-request cost the issue targets. Recorded here so a later
reviewer doesn't read the absence of a `SemaphoreSlim` as an oversight.
## 2026-07-19 — The `format` gate runs `dotnet format whitespace . --folder`, not the full solution format (#469)
The blocking `format` CI job (and the matching Husky pre-commit hook) verify changed `.cs` files with
`dotnet format whitespace . --folder --verify-no-changes --include <files>` instead of the previous
`dotnet format ErsatzTV.sln --no-restore --verify-no-changes --include <files>`.
- **Why.** `--include` narrows *which* files are checked, never what gets loaded. The full recipe loaded
the whole ~10-project MSBuild workspace and built a Roslyn compilation per project before checking a
single line — a fixed cost independent of how few files changed. Measured **~480s** for a whole-solution
`dotnet format` locally (matching the issue's "7+ min"). `--folder` treats the tree as a plain folder of
files and skips MSBuild/Roslyn entirely: **~0.5s**, and it needs no `dotnet restore`, so the job's
NuGet-cache + Restore steps were deleted. It also drops the job's ~3.95 GiB Roslyn heap (the #406
memory note about `format` not shrinking is now moot).
- **Coverage is unchanged, not merely "good enough".** Folder mode reads `.editorconfig` and enforces
exactly the two things this gate exists for — **whitespace** (indent/EOL/trailing/final-newline) and
**charset** (no UTF-8 BOM). Proven non-vacuous: exits non-zero with `error WHITESPACE` on an injected
trailing-whitespace line and `error CHARSET` on a prepended BOM; exits 0 on a clean file. What it drops
is the style/analyzer pass — but the *full* gate never enforced that either: a probe injecting a
`warning`-severity naming violation (`local_constants` not `ALL_UPPER`) **passed** the full solution
format (exit 0): the only `.editorconfig` rule above `:suggestion`/`:none` severity is that one naming
rule, and naming violations have no `dotnet format` batch code-fixer, so `--verify-no-changes` reports
no change regardless of severity. The analyzers that must block (`NU1904`,
`S3981`) are enforced at compile time via `WarningsAsErrors` in `Directory.Build.props`, never by this
job.
- **Fix command for a violation:** `dotnet format whitespace . --folder --include <files>`. The full
`dotnet format ErsatzTV.sln --include <files>` is a superset (also applies style) and still works, so
existing muscle memory and the #311 lore's `dotnet format --include` guidance are not broken.
- **Lane left on `ubuntu-latest`.** The job is now seconds-long and low-memory, so it could move to a
lighter lane, but that re-touches the per-lane memory-cap accounting (#406/#604) and is a
server-management capacity call — deliberately out of scope here.
## 2026-07-19 — CI `test` job reports a sampled true peak-anon, not cache-inflated `memory.peak` (#412)
**Decision.** The `test` job's memory instrument (added in #411) now reports a **sampled high-water
mark of the cgroup's `anon` memory** as the headline figure, produced by `scripts/ci-peak-anon.sh`
(a `start` step before the dotnet Build/Test/Coverage, a `report` step last). `memory.peak` and the
end-of-job `anon`/`file` split stay in the output as a cache-inflated ceiling and a reference.
**Why not just `memory.peak`.** `memory.peak` is the high-water mark of `memory.current`, which
charges reclaimable **page cache** to the cgroup alongside anon. A build does heavy
NuGet/npm/obj/bin/coverage I/O, so cache can dominate the peak — and page cache is *reclaimed* under
a tighter cap, not OOM-killed. Sizing a per-job cap (server-management#604) off `memory.peak`
therefore **inverts the decision**: a big, mostly-`file` peak reads like "the cap must stay high"
when it isn't. The OOM-forcing quantity is peak **anon**. The kernel exposes `memory.peak` but has
**no peak-anon counter**, and the end-of-job `anon` is the composition *then*, not at the peak
instant (a job that peaks mid-`dotnet test` then frees reports a misleadingly low `anon`) — so it
must be **sampled**. Details + the sampler's robustness rationale: `scripts/ci-peak-anon.sh` header
and `docs/ci-cd.md` → "CI build memory".
**Implementation note (do not "simplify" back to `memory.peak`).** The sampler is a detached
`nohup` poller that survives step-boundary re-execs (reparents to the container's PID 1) and is
reaped at container teardown; a TERM trap + `sleep & wait` stops it at once on `report`. Both steps
are `continue-on-error` with a fail-open script, so the instrument can never redden a green build.
Validated on bumblebee: it catches a transient 2.5 GiB anon spike that the end-of-job snapshot
reports as 0.
**Compiler-server A/B verdict (refines the #406 entry's "premise looking dead" read).** Measured
in the CI image, swap-off, sampled peak-anon, n=2 interleaved: **OFF (the CI config) ≈ 5.84 GiB,
consistent; ON (defaults) 6.37.6 GiB, always higher, + a ~3 GiB resident `VBCSCompiler`.**
Disabling the servers is worth it (consistent reduction, no resident server), but OFF sits *right at
6 GiB for the build phase alone* and the `test` job adds test + coverage on top — so #406's premise
("disabling brings peak *well under* 6 GiB → the budget loosens") is **not supported**. Size the cap
off the live test-job peak-anon this instrument now reports, not off the build-only A/B. The older
#411 probe (`anon 7134 MiB`) read higher than these swap-off sampled numbers and is superseded
(swap/read-method move the figure >1 GiB).
## 2026-07-19 — Media-server remote-stream URLs are probed before use: a redirected 404 fails closed, everything else fails open, no toggle (#473)
`GetPlayoutItemProcessByChannelNumberHandler.ValidatePlayoutItemPath` now probes the Plex/Jellyfin/Emby
remote-stream URL via the new `IRemoteStreamProber` seam before returning it, and on a 404 **from the media
server** returns the new `PlayoutItemNotAvailableFromMediaServer` error instead of a playable path.
- **The bug this closes.** The method's local branch checked `_fileSystem.File.Exists(path)`, but the
three remote-stream branches returned `http://localhost:{port}/media/{plex,jellyfin,emby}/{id}`
*unconditionally*. When the media was gone from the media server too, validation "succeeded", ffmpeg was
launched against a URL that 404s and died with **exit 8** — landing in `HlsSessionWorker`'s generic
ffmpeg-failure path, which sizes its error card to the failed **44s work-ahead chunk** and then
re-selects the *same* broken item, so a dead item produces a repeating error card for its whole slot
(~22 min for an episode). The fix restores the method's own invariant — every `PlayoutItemWithPath` it
returns has been checked for existence — so the failure now returns an **error** rather than a playable
URL, and the handler's error path sizes the card to run until the **next** playout item.
- **What the new `case` label does and does not do** (a review corrected an earlier draft of this entry):
the skip-to-next-item sizing comes from `maybeDuration`/`finish`, computed *before* the switch — the
`default:` arm already had it. Adding `case PlayoutItemNotAvailableFromMediaServer:` alongside
`PlayoutItemDoesNotExistOnDisk` changes only the **caption** on the card (the real error text instead of
"Channel is Offline"). Worth having, but not where the fix lives. A handler test asserts that message
specifically, so the label cannot silently decay into dead weight.
- **Scope: the generated-playout path only.** All three remote branches of `ValidatePlayoutItemPath` are
covered. `ExternalJsonPlayoutItemProvider` builds its own `/media/plex/...` URL and its result is
assigned *without* passing through `ValidatePlayoutItemPath`, so external-JSON channels keep the old
behaviour — tracked as #480 rather than silently claimed as fixed. The troubleshooting and
subtitle-extraction paths build the same URLs and are deliberately **left unprobed**: troubleshooting
should surface the raw ffmpeg failure, and a subtitle-extraction miss is not a channel outage.
- **Fail-open on everything except a *redirected* 404.** A timeout, a 5xx, an auth error or a transport
failure returns *available* — and so does a **404 that arrived without a redirect**. ErsatzTV's own
`InternalController` returns `NotFound` when the media source is unconfigured or momentarily missing, so
honouring that would fail *closed* for every item on that source; the first draft probed for "any 404",
which made this contract untrue. Only a 404 reached *after* the redirect to the media server is evidence
the item is gone. Pinned by tests (`Should_Fail_Open_*`,
`Should_Fail_Open_On_404_That_Was_Not_Redirected`) so a later refactor can't quietly invert it. Caller
cancellation is **not** swallowed — it propagates, because a shutdown is a genuine signal, not a probe
failure.
- **No `ConfigElement` toggle.** The behaviour strictly dominates the status quo, and the probe is bounded
by a 2s linked-CTS timeout (not `HttpClient.Timeout`, which defaults to 100s). A toggle would be surface
for a switch nobody has a reason to flip. **Not yet measured:** on an install without path replacement
every item selection pays this probe, and a slow media server could cost up to the full 2s. That is a
worst case rather than a typical one, but it has not been measured against #350's cold-start budget —
measure before assuming it is noise.
- **Rejected — resizing the `HlsSessionWorker` retry loop to skip the item on any ffmpeg failure.** It buys
the viewer nothing the above doesn't (the slot is dead either way) and cannot distinguish "this item is
dead" from "this transcoder hiccupped". Prod shows real *transient* channel-wide failures (VAAPI
`hwupload -22` / exit 234), and retrying those is correct; converting them into whole-item blackouts is a
regression. The retry loop is deliberately left intact — it also remains the backstop for the probe's
TOCTOU window (media can vanish between a 200 probe and ffmpeg's own request).
- **Rejected — writing `MediaItemState` from the streaming path.** Only the scanner writes `State`, and
breaking that ownership wouldn't even have fixed this: the failing item is `RemoteOnly`, which
`PlayoutBuilder`'s `PlayoutSkipMissingItems` skip does not exclude. On Jellyfin (`ServerSupportsRemoteStreaming`)
`RemoteOnly` is the *normal* state for every item when path replacement isn't configured, so it carries no
signal about playability — never key availability decisions off it.
- **Deferred — HEAD instead of GET.** The probe sends `GET` with `Range: bytes=0-0`, which on
`/Videos/{id}/stream?static=true` is a real (if minimal) playback request and may register a session or
touch play-state on some media-server versions. `GET` was chosen because HEAD support varies across
media servers — that is a reasonable prior but an *unverified* one. A HEAD-with-GET-fallback would avoid
the side effect; deferred rather than guessed at, since it trades a known-working request for an
untested one.
## 2026-07-19 — A media-server library sweep refuses to flag when a successful fetch returns zero items, rather than nuking the whole library (#477)
Each media-server scanner reconciles "gone upstream" as `existing.Except(incoming)` and flags the result
`FileNotFound`. If a *successful* fetch returns **zero** items — the server is up but mid-restore /
mid-rebuild, or the library was genuinely emptied upstream — then `existing.Except([])` is **every** item,
so one scan flags the entire library. That is data-loss-adjacent: `EmptyTrashHandler` permanently deletes
`state:FileNotFound` rows (a user clicking Empty Trash after a bad scan), and `PlayoutSkipMissingItems`
empties every affected collection (dead channels). There was no zero-count / ratio / server-total guard;
the only thing that stopped a mid-*pagination* failure was an exception unwinding past the flag step —
protection by accident of control flow, not by design.
- **The guard is a single shared policy.** `MediaServerReconciliationGuard.ShouldFlagMissing(logger,
libraryName, incomingCount, existingCount)` returns false (and logs a Warning) **only** when
`incomingCount == 0 && existingCount > 0`; every other combination sweeps normally. Wired into the three
**library-level** sweeps — `MediaServerTelevisionLibraryScanner` (shows), `MediaServerMovieLibraryScanner`,
`MediaServerOtherVideoLibraryScanner`. One place owns the invariant so the policy can't drift between
scanners.
- **An empty incoming set is genuinely ambiguous, so we choose the non-destructive branch.** "User removed
every item" and "server returned empty erroneously" are **indistinguishable** at scan time — both report
a total of zero (the paginator computes `pages` from `TotalRecordCount`, so a 0 total is a clean empty
enumeration, not an error). Given the blast radius, skipping wins: the cost of *not* flagging a
legitimately-emptied library (stale rows persist until an item returns or the library is removed by hand)
is far smaller than a one-scan permanent wipe of a live library.
- **This partially overrides #476 for the degenerate case, on purpose.** #476 cascades a removed show's
flag to its seasons/episodes. Its common path — some items removed while **survivors are present**
(incoming non-empty) — still flags and cascades exactly as before. Only the degenerate "the last item was
removed, so incoming is empty" case now skips instead of flagging. The #476 characterization test was
rewritten from an empty incoming to a survivor-plus-removed partial deletion so it exercises the cascade
without tripping the guard.
- **Scope: library-level sweeps only; the nested TV season/episode sweeps are deliberately left unguarded.**
Their blast radius is one show's seasons / one season's episodes (not the whole library), a per-parent
empty is a more plausible legitimate state there, and the #476 descendant cascade already handles a fully
removed parent. Guarding them would alter #476's per-parent behaviour for little safety gain.
- **Deferred — ratio threshold and projection-failure detection.** The issue also floated "skip if the
missing fraction exceeds a threshold" and "distinguish a silently-dropped projection failure from a real
deletion." A ratio threshold risks suppressing a legitimate bulk deletion and needs a tunable, telemetry-
backed policy; projection-failure detection needs a dropped-count threaded out of `JellyfinApiClient`
through to the scanner (a cross-layer change). The deterministic zero-count guard has **no** false
positives and covers the reported catastrophic case, so both are deferred to a follow-up rather than
guessed at here.
- **Tests.** `MediaServerReconciliationGuardTests` pins the policy table (only `(0, N>0)` skips-and-warns;
`(0,0)`, `(3,5)`, `(3,0)` all sweep). Per-scanner integration tests
(`MediaServer{Television,Movie,OtherVideo}LibraryScannerTests`) prove the wiring: empty incoming +
non-empty existing flags nothing and reindexes nothing. Proven non-vacuous by neutralizing the guard and
watching all four anti-nuke assertions fail while the `(0,0)` no-op case stays green.