b1d5fbefcba02fdc6c19fef85cec1c4e82fc8dea
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e132c422bb |
fix(511): bound remote graphics-engine image fetches
`ImageElementBase.LoadImage` fetched http(s) images with a throwaway `new HttpClient()` + `GetStreamAsync`: no timeout override (the 100s default), no size cap, unbounded redirects, no pooling — all inside stream startup, while ffmpeg waits on the pipe. #502 routed ordinary channel-logo watermarks onto that path, widening a pre-existing weakness. Introduce `IRemoteImageFetcher` / `HttpRemoteImageFetcher`, modelled on the neighbouring `IRemoteStreamProber`: - deadline covers headers AND body (linked CTS + `CancelAfter`, client `Timeout = InfiniteTimeSpan`) — under `ResponseHeadersRead` the body read falls outside `HttpClient.Timeout` (the #289 lesson) - 10 MiB cap enforced during the copy; `Content-Length` is only a cheap early reject, since it can be absent or a lie - permissive content-type check (rejects an HTML error page, allows a missing type and octet-stream) - pooled via `IHttpClientFactory`; redirects capped at 3, not 50 A byte cap does NOT bound decoding, so `DecodeRemoteImage` additionally reads declared dimensions + frame count from the header and rejects before `Image.LoadAsync` allocates (50 MP / 600 frames). A 4 KB PNG declaring 30000x30000 costs ~3.6 GB to decode and passes every wire-size check — caught by adversarial review of the first version of this change, which capped bytes and wrongly claimed that was decode-bomb protection. Not cached and SSRF not mitigated — both deliberate, with the reasoning recorded in docs/decisions.md. fixes #511 |
||
|
|
dc5ceb5a14 |
fix(473): bound the drain, correct the interface contract, quiet graceful cancels
Build ErsatzTV Image / CI image pin matches docker/ci (pull_request) Successful in 8s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 12s
Build ErsatzTV Image / Docs update reminder (pull_request) Successful in 13s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / decisions.md append-only (pull_request) Failing after 12m6s
Build ErsatzTV Image / Functional E2E (curl contracts) (pull_request) Successful in 15m47s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 18m54s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 20m46s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
Second review pass returned BLOCKED on two findings introduced by the first fix commit. Both were right. BLOCKER 1 — the drain added for "return the connection to the pool" was unbounded. `response.Content.ReadAsByteArrayAsync()` buffers the WHOLE body, and it ran for every non-404 response. A server that ignores `Range: bytes=0-0` answers 200 with the entire file, so this would download at line rate into a byte[] on the streaming hot path for up to the 2s timeout -- strictly worse than the aborted socket it replaced, and it defeated the ResponseHeadersRead the probe deliberately uses. Now the single byte is read only on 206 (where the server honoured the range and the body really is one byte); any other status aborts the socket, which is much the cheaper evil. Two tests pin both directions; verified non-vacuous (restoring the unbounded drain fails the 200-with-body test). BLOCKER 2 — IRemoteStreamProber's doc-comment still described pre-fix behaviour. I had told the reviewer it was updated; it was not -- only the implementation's <remarks> had been. It claimed `false` on any 404 (now only a redirected one) and that every other outcome returns `true` (caller cancellation throws). Both clauses corrected, and the throwing contract is now documented with <exception>. Also fixed the reviewer's own follow-on finding: the cancellation rethrow it asked for reached HlsSessionWorker's catch-all, which logs a channel-level ERROR with a stack trace. The graceful TaskCanceledException/OperationCanceledException handler at :662 wraps only the inner ffmpeg block, not the mediator sends, so every client disconnect on a remote-streaming channel would have produced a spurious ERROR -- in exactly the logs a #350 cold-start investigation reads. Added a cancellation filter on the outer try that logs Information instead. Nit: stale SeedAll doc-comment now mentions the emby case. Deferred, per reviewer's explicit agreement: Plex-branch handler coverage (follow-up), and HEAD-with-GET-fallback. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a10c0325eb |
fix(473): compare parsed Uris in the redirect check
Defense in depth on the redirect detector: Uri.Equals compares normalized components, so an escaping/casing difference can't be mistaken for a redirect and fail CLOSED -- the exact failure the check exists to prevent. A plex key can contain spaces or unicode. Honest note: this is NOT a fix for an observed bug. I wrote a test claiming to pin it, then ran the negative control and the test passed against the string comparison too -- Uri.ToString() unescapes, so both forms agree for our machine-generated URLs. The test was vacuous as written. It is kept, retitled and re-commented to describe what it actually guards (an un-redirected 404 on an escaping-sensitive url fails open), and the code comment says plainly that this is defense in depth rather than a repair. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
bad19f8d26 |
fix(473): review fixes — only a redirected 404 fails closed
Adversarial review of PR #479 found the stated fail-open contract was not what the code measured, plus four smaller gaps. All fixed here as a follow-up commit (no amend/force-push). High — a 404 from ErsatzTV's OWN endpoint was treated as "media gone". /media/{provider}/... is served by InternalController, which returns NotFound when the media source is unconfigured or momentarily missing (a media-source edit that deletes+reinserts connections, a restore, a partially-configured server). Probing for "any 404" therefore failed CLOSED for every item on that source -- exactly the case the fail-open contract exists to prevent. A media-server 404 always arrives after a redirect, so an un-redirected 404 is now treated as available. Medium — the new switch label was untested and its benefit overstated. maybeDuration/finish are computed before the switch, so `default:` already sized the error card to the next playout item; the label only changes the caption. The handler test asserted call counts only, so deleting the label still passed. It now asserts the error message, and removing the label fails the test (verified). Medium — Plex/Emby branches changed but had no coverage. Added an Emby handler test asserting the probe is called with the emby URL. Low — caller cancellation was swallowed and pinned as desired behaviour. A shutdown / client disconnect is a genuine signal, not a probe failure; it now propagates, and only the probe's own 2s timeout fails open. Low — the response stream was disposed unread, aborting the connection instead of returning it to the pool. The one requested byte is drained. Nit — fully-qualified RangeHeaderValue replaced with a using. docs/decisions.md corrected where it overstated: the switch label's role, the "fixes the class for all three media servers" claim (external-JSON channels bypass ValidatePlayoutItemPath entirely -- filed as #480), and the unmeasured latency assertion. Deferred HEAD-instead-of-GET recorded with its reason rather than silently dropped. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
90dc864ce5 |
fix(473): probe media-server remote streams before handing the URL to ffmpeg
Tuning a channel intermittently hard-failed with ffmpeg exit 8 and
`Server returned 404 Not Found` on /media/jellyfin/{itemId}.
Root cause: ValidatePlayoutItemPath checked `File.Exists` on the local
branch, but the three media-server remote-stream branches returned
`http://localhost:{port}/media/{plex,jellyfin,emby}/{id}` unconditionally.
When the media was gone from the media server too, validation "succeeded"
and ffmpeg was launched against a URL that 404s.
That bypassed the good error path the handler already had
(PlayoutItemDoesNotExistOnDisk renders an error card sized to run until
the NEXT playout item, so the dead item is skipped) and instead landed in
HlsSessionWorker's generic ffmpeg-failure path, which sizes its error card
to the failed 44s work-ahead chunk and then re-selects the SAME broken
item -- a repeating error card for the item's whole slot (~22 min).
Restore the method's own invariant: every PlayoutItemWithPath it returns
has been checked for existence. A definitive 404 now returns the new
PlayoutItemNotAvailableFromMediaServer error, handled in the same switch
arm as PlayoutItemDoesNotExistOnDisk.
The probe is deliberately fail-open: only a 404 reports the media gone.
A timeout, 5xx, auth error or transport failure reports available, so a
probe that cannot answer can never break a tune that would have worked.
That contract is pinned by tests so a later refactor cannot invert it.
Rejected alternatives (see docs/decisions.md): resizing the
HlsSessionWorker retry loop (cannot distinguish a dead item from a
transient transcoder failure -- prod has live VAAPI hwupload -22 failures
that must keep retrying), and writing MediaItemState from the streaming
path (breaks scanner ownership, and would not have fixed this: the item
is RemoteOnly, which PlayoutBuilder's skip does not exclude).
Scanner-side follow-ups filed separately: #476 (FileNotFound does not
cascade show -> episodes, the reason dead items keep being scheduled),
fixes #473
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|