168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key, so one-active-record-per-key becomes a filesystem property rather than a validator check, and supersession becomes a `git mv`. WHY: the monolith was a concurrency problem before an aesthetic one. A 3,900-line append target made parallel sessions collide -- PR #605 and PR #614 both hit append-vs-append conflicts during routine rebases, and hand-resolving those inside the corpus is exactly the operation the rationale-rewrite guard exists to police. HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness does not rest on reading it. The parser was taught BOTH formats first, so the body-diff guard parses the old form at the merge-base and the new form at head -- the migration validates itself, no bypass. The proof is a field-level equivalence harness: 168 records before and after, zero lost, zero gained, zero field mismatches, zero rationale bodies differing. Reviewers should scrutinise the harness; it is the actual evidence. What measuring caught that reading would not have: - ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each topic file's preamble, mostly the only copy. Source files are kept and stripped, never deleted. They also cannot be filed per-area: topic files hold several areas and 4 of 23 areas span several files. - Archive discovery was a non-recursive glob; after the split it found ZERO archived records, surfacing as four bogus "supersedes points to unknown key" errors rather than an obvious failure. - ~32 live docs point into the corpus BY DATE, which the split dangles. Each stripped file now ends with a generated "Records formerly in this file" index, which also rescues the identical breadcrumbs in old issue comments. - decisions.md's "In this file:" list was 97 same-file anchor bullets that the split makes WRONG, not merely stale. Dropped; the generated index replaces them with links that resolve. The equivalence harness now runs against a checked-in FIXTURE, not the live corpus. The earlier version migrated the real tree, which made it a one-shot: the moment the migration landed there was nothing left to move and the tests failed for reasons unrelated to the code. A fixture keeps them testing the SCRIPT rather than the repo's current state. Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain two directories for one concept. Renaming a key is not a move -- it changes identity, breaks the equivalence proof, and invalidates MemPalace's per-key drawers. Taxonomy normalisation is separate work. refs #610
6.0 KiB
key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
| key | title | status | since | supersedes | superseded-by | rule | signals | mechanics |
|---|---|---|---|---|---|---|---|---|
| ffmpeg.hls-cold-start-burst | 2026-07-20 — HLS cold start is fixed with `-readrate_initial_burst`, not by raising the work-ahead limit (#350) | active | 2026-07-20 | none | none | HLS cold-start latency is fixed with a bounded `-readrate_initial_burst` (gated on FFmpeg ≥6.1 capability detection), not by raising `work_ahead_limit`, which would remove the concurrency guarantee it exists for. | HLS cold start, readrate, work_ahead_limit, HlsSessionWorker, FFmpegKnownOption capability gate · paths: `HlsSessionWorker`, `SetRealtimeInput`, `FFmpegPlaybackSettingsCalculator`, `FFmpegKnownOption`/`HasOption` · issues: #350 | `SetRealtimeInput` readrate-burst option; `FFmpegKnownOption.HasOption` version-capability gate |
Correction (2026-07-21, #529): the claim below that the burst is "bounded" is true only in seconds of input — it is not a bound on memory or hardware surfaces.
-readratewas also incidentally bounding how fast decoded frames enter the filter graph, and removing that on a QSV pipeline whose profile storesqsvExtraHardwareFrames: 0exhausts the upload pool: the graph fails with-12 (Cannot allocate memory),h264_qsvnever opens, and zero segments are written. The burst is not the root cause (a work-ahead start takes no-readrateat all and was already failing the same way in production), but it removed the throttle on every realtime session and so made the failure near-deterministic. Seeffmpeg.qsv-extra-hw-frames-floor; the decision recorded here still stands, with that floor in place.
Correction (2026-07-21, #536): the bullet below stating that "every concurrent tune-in falls back to the throttled path" described the intent, not the behaviour. The slot check and its increment straddled an
await, so N simultaneous tune-ins all read0 < limitand all started unthrottled; the "one winner, two throttled" observation held only because those tunes were effectively staggered. The measurement and the conclusions drawn from it stand — slot availability was the variable behind the bimodality — but the guarantee itself was not enforced untilffmpeg.work-ahead-slot-atomic.
-readratethrottles from the first read, so it sets a floor on time-to-first-segment. The realtime playback path pins input reading to 1.05× wall clock so a channel behaves like live TV. SinceOutputFormatHls.SegmentSecondsis 4 and the segmenter serves the playlist only once the first segment exists, the playlist cannot appear sooner than ~4/1.05 ≈ 3.8 s. Measured on real media: time-to-first-playlist 5369/5344 ms with-readrate 1.05, 648/649 ms with an initial burst.- The cold-start bimodality was never about the media.
HlsSessionWorkergrants an unthrottledSeekAndWorkAheadstart only while_workAheadCount < ffmpeg.segmenter.work_ahead_limit(prod: 1); every concurrent tune-in falls back to the throttled path. Three concurrent tunes on prod: the one that won the slot reachedfirstGopin 866 ms, the other two in 3845 ms and 6357 ms. This is why the earlier rounds found no correlation with subtitle burn-in, GOP length, or source file — the variable was slot availability, and the same channel could differ 7.2× between tunes. - Two earlier hypotheses are falsified, not deferred. Accurate-seek decode-discard (the issue's
ranked #1 driver) costs 30–100 ms on real prod media, and capping
-probesize/-analyzedurationbuys 20–50 ms. Neither can account for seconds. Recorded here so they are not re-proposed. - Burst rather than a bigger work-ahead budget. Raising
work_ahead_limitwould fix latency by deleting the guarantee that limit exists for — it caps how many unthrottled transcodes N viewers can start at once. The burst is bounded (SegmentSeconds * 2= 8 s of input, enough for the first segments at the defaultInitialSegmentCountof 1), after which live pacing resumes. An operator who raisesInitialSegmentCountabove 2 gets less of the benefit; that is a deliberate trade. - The burst is per ffmpeg process — i.e. per playout item — not per session.
SetRealtimeInputruns on every pipeline build andHlsSessionWorkerspawns a process per item, so each item boundary bursts too; this is not only the session's cold start. Two consequences, both accepted: item transitions get the same head start (a benefit), and on a channel whose items are shorter than the burst every item transcodes unthrottled, so the instantaneous-concurrency guarantee thatwork_ahead_limitprovides is weaker than before — weaker, not absent, becauseHlsSessionWorker'stranscodedBuffer <= 1mingate still stops the loop at a 60 s buffer, leaving average CPU unchanged. Making the burst strictly cold-start-only would mean plumbing a "first process of this session" flag throughFFmpegState; that complexity was not judged worth a bounded peak. - Still images are excluded. Their video input is paced by the realtime filter and takes no readrate at all, so a burst would only run the audio input ahead of the video for songs and offline filler, with no cold-start gain to show for it.
- Non-HLS realtime outputs (
TransportStream, HLS-Direct) burst too, sinceFFmpegPlaybackSettingsCalculatormakes them unconditionally realtime. That is untested by the benchmark, which was segmenter-only; it is kept because the same first-read throttle delays those clients identically, and the outerWrapSegmenter/Concatprocesses still pace atreadrate 1.0with no burst. - Gated on runtime capability, not on a parsed version.
-readrate_initial_burstneeds FFmpeg ≥ 6.1, andFFmpegKnownOption/HasOptionalready existed for exactly this (itsAllOptionslist had simply been empty). Detection parsesffmpeg -h long, so an older binary silently keeps today's behavior instead of failing to start — the same fail-safe posture as the other capability gates, and cheaper to reason about than the version-string parsing inNvidiaHardwareCapabilities.