Files
ersatztv/docs/decisions/records/ffmpeg/hls-cold-start-burst.md
T
timothy fba5233caf
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
feat(610): split the decision corpus into one YAML-frontmatter file per record
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
2026-07-25 19:45:09 +02:00

6.0 KiB
Raw Blame History

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
ffmpeg.hls-cold-start-burst 2026-07-20 — HLS cold start is fixed with `-readrate_initial_burst`, not by raising the work-ahead limit (#350) active 2026-07-20 none none HLS cold-start latency is fixed with a bounded `-readrate_initial_burst` (gated on FFmpeg ≥6.1 capability detection), not by raising `work_ahead_limit`, which would remove the concurrency guarantee it exists for. HLS cold start, readrate, work_ahead_limit, HlsSessionWorker, FFmpegKnownOption capability gate · paths: `HlsSessionWorker`, `SetRealtimeInput`, `FFmpegPlaybackSettingsCalculator`, `FFmpegKnownOption`/`HasOption` · issues: #350 `SetRealtimeInput` readrate-burst option; `FFmpegKnownOption.HasOption` version-capability gate

Correction (2026-07-21, #529): the claim below that the burst is "bounded" is true only in seconds of input — it is not a bound on memory or hardware surfaces. -readrate was also incidentally bounding how fast decoded frames enter the filter graph, and removing that on a QSV pipeline whose profile stores qsvExtraHardwareFrames: 0 exhausts the upload pool: the graph fails with -12 (Cannot allocate memory), h264_qsv never opens, and zero segments are written. The burst is not the root cause (a work-ahead start takes no -readrate at all and was already failing the same way in production), but it removed the throttle on every realtime session and so made the failure near-deterministic. See ffmpeg.qsv-extra-hw-frames-floor; the decision recorded here still stands, with that floor in place.

Correction (2026-07-21, #536): the bullet below stating that "every concurrent tune-in falls back to the throttled path" described the intent, not the behaviour. The slot check and its increment straddled an await, so N simultaneous tune-ins all read 0 < limit and all started unthrottled; the "one winner, two throttled" observation held only because those tunes were effectively staggered. The measurement and the conclusions drawn from it stand — slot availability was the variable behind the bimodality — but the guarantee itself was not enforced until ffmpeg.work-ahead-slot-atomic.

  • -readrate throttles from the first read, so it sets a floor on time-to-first-segment. The realtime playback path pins input reading to 1.05× wall clock so a channel behaves like live TV. Since OutputFormatHls.SegmentSeconds is 4 and the segmenter serves the playlist only once the first segment exists, the playlist cannot appear sooner than ~4/1.05 ≈ 3.8 s. Measured on real media: time-to-first-playlist 5369/5344 ms with -readrate 1.05, 648/649 ms with an initial burst.
  • The cold-start bimodality was never about the media. HlsSessionWorker grants an unthrottled SeekAndWorkAhead start only while _workAheadCount < ffmpeg.segmenter.work_ahead_limit (prod: 1); every concurrent tune-in falls back to the throttled path. Three concurrent tunes on prod: the one that won the slot reached firstGop in 866 ms, the other two in 3845 ms and 6357 ms. This is why the earlier rounds found no correlation with subtitle burn-in, GOP length, or source file — the variable was slot availability, and the same channel could differ 7.2× between tunes.
  • Two earlier hypotheses are falsified, not deferred. Accurate-seek decode-discard (the issue's ranked #1 driver) costs 30100 ms on real prod media, and capping -probesize/-analyzeduration buys 2050 ms. Neither can account for seconds. Recorded here so they are not re-proposed.
  • Burst rather than a bigger work-ahead budget. Raising work_ahead_limit would fix latency by deleting the guarantee that limit exists for — it caps how many unthrottled transcodes N viewers can start at once. The burst is bounded (SegmentSeconds * 2 = 8 s of input, enough for the first segments at the default InitialSegmentCount of 1), after which live pacing resumes. An operator who raises InitialSegmentCount above 2 gets less of the benefit; that is a deliberate trade.
  • The burst is per ffmpeg process — i.e. per playout item — not per session. SetRealtimeInput runs on every pipeline build and HlsSessionWorker spawns a process per item, so each item boundary bursts too; this is not only the session's cold start. Two consequences, both accepted: item transitions get the same head start (a benefit), and on a channel whose items are shorter than the burst every item transcodes unthrottled, so the instantaneous-concurrency guarantee that work_ahead_limit provides is weaker than before — weaker, not absent, because HlsSessionWorker's transcodedBuffer <= 1min gate still stops the loop at a 60 s buffer, leaving average CPU unchanged. Making the burst strictly cold-start-only would mean plumbing a "first process of this session" flag through FFmpegState; that complexity was not judged worth a bounded peak.
  • Still images are excluded. Their video input is paced by the realtime filter and takes no readrate at all, so a burst would only run the audio input ahead of the video for songs and offline filler, with no cold-start gain to show for it.
  • Non-HLS realtime outputs (TransportStream, HLS-Direct) burst too, since FFmpegPlaybackSettingsCalculator makes them unconditionally realtime. That is untested by the benchmark, which was segmenter-only; it is kept because the same first-read throttle delays those clients identically, and the outer WrapSegmenter/Concat processes still pace at readrate 1.0 with no burst.
  • Gated on runtime capability, not on a parsed version. -readrate_initial_burst needs FFmpeg ≥ 6.1, and FFmpegKnownOption/HasOption already existed for exactly this (its AllOptions list had simply been empty). Detection parses ffmpeg -h long, so an older binary silently keeps today's behavior instead of failing to start — the same fail-safe posture as the other capability gates, and cheaper to reason about than the version-string parsing in NvidiaHardwareCapabilities.