The ban that keeps `build`'s `Smoke + IPTV E2E` from being silently dropped was enforced
only by a pytest in `script-tests` — `on: pull_request`, and not a required context. Nothing
re-checked it on a `v*` tag push, which is exactly when the candidate image is published and
`DeployStack jazz-media` promotes it. A delimiter that reached `main` would drop `Smoke` on
the tag build, publish an unsmoked candidate, and report green.
A `scan` job now runs the PyYAML-based ban test, and `build` lists it in `needs:`. That edge
is the whole property: a red `scan` skips `build` outright, so the image is never built.
TWO DESIGNS WERE TRIED AND THE FIRST ONE'S FAILURES ARE RECORDED, because both are easy to
re-invent. The first cut put a bespoke stdlib scanner in `build` itself, as an unconditional
step before `Build and push`. Two independent cold reviews rejected it:
* A guard STEP cannot protect the job it lives in. `build` is what publishes, so a dropped
guard step there fails OPEN — and "the guard's own body has no opener, so it cannot be
dropped" is circular when the only thing enforcing that property is the same PR-only test
being backstopped. A `needs:` edge is not circular.
* The hand-written YAML parser had ~10 false NEGATIVES in one review round (flow mappings,
a quoted `"run":` key, aliases, multiline quoted scalars) — strictly WEAKER than the check
it backstopped, in the only direction that matters for a security gate. Deleted rather
than patched: running the existing test needs no second definition of "what is a `run:`
body", so there is no drift surface at all.
No third marker bucket was needed. The deferral assumed the answer had to be markers on
`build`, modelling `Smoke`'s publish-ref `if:`. The delimiter class is a STATIC property of
the workflow text, so a job that reads the text catches it without modelling any `if:`.
The `scan` job's own steps carry #756 markers and a trailing assert, so a drop inside it is
caught too — moving the terminal assumption rather than removing it: to fail open you must
now drop the pytest step AND the assert step.
Every way of disarming the gate was mutation-tested to a red: removing the `needs:` edge,
adding a job-level `if:`, marking a step `continue-on-error`, injecting a delimiter into a
scan body, dropping the ban test from the pytest invocation, removing a marker, and deleting
the assert step. The guard's real command line is also driven against the steps' real marker
lines with each key dropped in turn.
Refs: #767
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
138 KiB
CI/CD for the ErsatzTV Fork
The fork builds its own Docker image via Gitea Actions on the homelab and pushes to the Gitea container registry. Runner + registry were provisioned in server-management#172; the build pipeline is ersatztv#4; test/prod containers are server-management#481.
Hosts (read this before trusting a hostname below)
| Host | IP | Role |
|---|---|---|
| jazz | 192.168.1.29 | Docker host for the media transcoders — prod ersatztv (8409), ersatztv-test (8410), Jellyfin. Release scans (security-scan.sh) and the prod-copy migration-smoke.sh run here. |
| bumblebee | 192.168.1.99 | CI runners (bumblebee-runner, small-runner), plus every other Docker stack. All the memory/lane measurements below were taken here. |
| ci-runner | VM 127 (pve4) | The other ubuntu-latest CI runner. |
The transcoders moved bumblebee → jazz on 2026-07-20 (server-management#633).
Name-reuse trap.
jazzwas an earlier name for the .99 host. Anything written before 2026-07-20 that says "jazz" means today's bumblebee — resolve hostnames by IP, not by name, and don't "fix" a historical bumblebee reference into jazz.
Versioning & releases
The fork inherits upstream ErsatzTV's scheme: vYY.<release-seq>.<patch> (lightweight, v-prefixed git tags).
YY— two-digit year.<release-seq>— a sequential release counter within the year, reset at each year boundary. It is not the calendar month. (Evidence:v25.2.0shipped in June 2025,v25.5.0in Sep,v26.3.0in Feb 2026 — minors don't track months; andv25.9.0→v26.1.0shows the year-reset.)<patch>— a small follow-up/hotfix on the same release line (e.g.v26.1.0→v26.1.1, days later).
Upstream's final release was v26.3.0 (archived). Our line continues from there:
| Tag | Meaning |
|---|---|
v26.3.1 |
Upstream 26.3.0 rebuilt on our infra (Gitea CI/registry, fork ffmpeg base) — no application changes. A patch bump, because nothing functional changed. |
v26.4.0 |
First fork release carrying application changes. Later 2026 releases continue 26.5.0, 26.6.0, …; a new year resets to 27.1.0. |
v26.7.0 |
Blazor-removal release: ChicoryTV became the only UI. |
v26.8.0 |
Secured/versioned ChicoryTV SPA + REST API go-live release (#335). |
v26.9.0 |
Configurable advertised IPTV base URL for M3U/XMLTV (#340) + SPA shell/routing + playouts modularization (#247/#245); on-air/Plex/library-path fixes (#99/#345/#371); coverage + functional-E2E CI (#15/#299). |
v26.10.0 |
Auto-Tune channel workflow (#69) + weighted content distribution (#70); scheduling refactors, health-check remediation UX (#164), HLS cold-start instrumentation (#350), security hardening (#293/#376/#308). |
v26.11.0 |
QSV profiles decode via VA-API — QsvPreferNativeDecoder, default on, fixes ~50% channel cold-start failures on Intel (#498); unified logo/on-screen bug via a shared watermark preset (#67). Media-scanner resilience: Jellyfin mixed-content libraries (#489), music-video scan correctness (#488/#494/#497), remote-stream probing before ffmpeg (#473/#480); weighted-distribution SPA (#404). First release deployed to jazz (server-management#633). |
v26.12.0 |
ErsatzTV.Mcp MCP server — read + cautious-write over /api/v1, ERSATZTV_ALLOW_WRITES-gated (#58). External channel-logo URLs download + cache at save time (#525), with the on-screen bug now rendered for external-URL logos (#502). HLS cold-start hardening: burst-read the first segments so start isn't -readrate-bound (#350) and floor QSV extra hardware frames so an unthrottled read can't exhaust the pool (#529); remote graphics-engine image fetches bounded — timeout, size cap, decode cap, redirects, pooling (#511). Decision-lifecycle tooling + parallel-orientation startup rewrite (#520/#521); CI docker build lane rebalance (#508). |
v26.13.0 |
RuleBuilder maturation — arbitrary-depth group nesting (#436), inline smart-query authoring in Channel Builder (#437), DB-sourced facet typeahead + relative-date operators + validation (#434/#435/#438), and an artist typeahead covering music-video/song credits with album_artist no longer 404ing (#578). Per-channel On Now/Next transient overlay (#74/#570) and per-schedule clock-boundary padding (#392); in-browser channel preview (#60); Auto-Tune per-source weight steppers + exclude/add-untagged (#440). Library-browse pickers now resolve by search instead of a 100-row window, closing several silent at-cap truncations (#644/#650/#651/#634). Correctness: one watermark resolver for all four attachment points, incl. MiddleCenter (#503/#510); QSV HDR tonemaps through OpenCL because vpp_qsv=tonemap is a silent no-op (#505); LibraryFolder unique index + concurrent-insert tolerance (#491); per-library music-video identity with soft trash (#496); Jellyfin Album/Track music-video projection (#177); metadata-collection dedup (#500); accented facet values via a registered Unicode fold on SQLite (#668); WorkAheadSlots atomic slot claim, never a negative count (#536/#539); on-demand guide rebuild on thaw (#68). Process/CI: the H10 review-verdict gate became a sha-bound required commit status and was hardened through its false-open chain (#622/#629/#632/#648/#649/#672/#698), the decision corpus split to one YAML-frontmatter record per file (#610/#620), and headless Playwright UI-E2E flows landed (#445/#533). Five dual-provider migrations. |
v26.14.0 |
Live TV no longer starves on embedded bitmap subtitles — -readrate paces an input off its furthest-behind stream, and a PGS/DVD subtitle read through the video's own -i is sparse enough to drag the whole process to 0.53x realtime against the 1.0x a client consumes, draining the buffer until the channel stalls. Fixed with a capability-gated -readrate_catchup (ffmpeg 8.0+) on realtime inputs, keeping -readrate on the frame-producing path so the ffmpeg.qsv-extra-hw-frames-floor bound is untouched; measured 0.533x → 1.067x on QSV and software, with a 240s QSV soak clean of allocation errors (#726). Affects items carrying an embedded bitmap subtitle matching the channel's subtitle mode — 3,182 of 24,646 media versions on prod, and a property of the item, not the channel, which is why the stall presented as random. Process/CI: the H10 review-verdict gate's repair sentinel became a fixed point and its write is now fenced on the timeline retarget count, closing a raced-sentinel false-open (#706/#707/#711). The decisions validator now cross-checks its dependency-free frontmatter parse against PyYAML and reports both the truncating unquoted # and the scalar-closing bare apostrophe as errors, so a record whose rule: silently halves under PyYAML fails the local gate instead of CI (#674/#688) — the ceiling-calibration claim was also split so the suite pins what the derivation MEANS rather than live-corpus order statistics. Dependencies: CliWrap 3.10.4, JetBrains.ReSharper.GlobalTools 2025.3.5. |
Before cutting a release — sweep docs/decisions.md + docs/decisions/ (ersatztv#521, supersedes
the ersatztv#303 H9 append-only ritual). Supersession/retirement is now a same-PR act (add the new
active record, relocate the predecessor to docs/decisions/archive/ with reciprocal
supersedes/superseded-by links), not a release-boundary batch job — most of the old "consolidate"
step is now continuous. The release boundary is instead where you:
-
Run
PYTHONPATH=. python3 scripts/decisions_validate.py— confirms lifecycle metadata is well-formed and everysupersedes/superseded-bylink resolves both ways. Since ersatztv#674 it also cross-checks its dependency-free frontmatter parse against PyYAML when PyYAML is importable, failing on any file PyYAML rejects (a bare apostrophe in a single-quoted value) or reads differently (an unquoted#, which YAML truncates as a comment). Where PyYAML is absent — thedecisions-guardjob, the Husky hooks — the cross-check is skipped with a::notice::and every other check still runs; the read path stays dependency-free. -
Confirm every record already classified
superseded/retiredactually lives underdocs/decisions/archive/(the validator fails this, but eyeball it at the boundary too). -
Regenerate the active catalog:
PYTHONPATH=. python3 scripts/build_decisions_catalog.pyand commit any drift. -
Read the corpus size signals. Since ersatztv#620 these are two separate things:
- a per-record prose ceiling (
decisions_validate.py --record-ceiling <n>, default 60) — a non-blocking::warning::naming every record over it. This is the actionable signal: it points at a file. The 60 is derived from the distribution, not picked as a round number. Its calibration is guarded in two pieces of different robustness (ersatztv#688), because four earlier single-assertion versions all failed — the first two by being vacuous or accepting an absurd ceiling, the last two by ratcheting:- blocking (
script-tests) — only the coarse property that the ceiling flags a meaningful minority of records (0.02 <= fraction_over <= 0.25). One record moves a fraction by at most 1/N, so no SINGLE ordinary addition can cross it. This is measured headroom, not immunity: from today's 18/183 it takes 38 consecutive over-ceiling additions to breach the cap, 718 short ones to dilute below the floor, or — the tightest arm — consolidating 15 of the 18 offenders away. The floor is a fraction rather than "at least one record", which would accept any ceiling up to 229 on the live corpus; as a fraction the accepted range is 43..180. - reported, never asserted against the LIVE corpus — the fine claim that the ceiling sits
between the 90th and 95th percentile, i.e. at the tail boundary.
main()prints a::notice::when it drifts; the tests assert it only on distributions they own. It is an order statistic over a sparse distribution, so a single new record could move p90 by 21 lines and red the blocking job for whoever wrote it; a ceiling going out of date is the passage of corpus growth, not a defect in the commit under test, so it is treated likestale-after. Re-derive the constant when the notice says so.
- blocking (
- the aggregate prose total, printed every run as an unthresholded
::notice::trend. It has no pass/fail. A total over a monotonically growing corpus can only ratchet: the old 4800→5600 budget went quiet at 5228 after #610 changed the metric and was back over at 5658 three and a half hours later the same evening, with nobody consolidating anything — the "permanently red = no signal" failure, not in slow motion at all. It reports record prose and non-record scaffolding separately, because they are not the same unit. The generated catalog is no longer counted at all — it gains one row per record and cannot be consolidated away.
Being listed by the ceiling is an invitation to check for redundancy, not an instruction to cut. A long record that is entirely distinct findings is a legitimate decline — say so in the record and move on. (
--budgetis still accepted and ignored, so old invocations keep working.) - a per-record prose ceiling (
-
Report the remaining
legacy-unmigratedcount (the validator prints it as a::notice::) so the backlog is visible, even though it isn't required to hit zero before a release. A genuine rationale-prose rewrite still needs aDecisions-Edit: yesgit trailer on a non-merge commit in the range (see thedecisions.mdheader) — routine lifecycle metadata writes above do not.
Cutting a release: keep build and promotion as two explicit phases (#335):
- Confirm
mainCI is green; run the full local gate plusdotnet list package --vulnerable --include-transitive; then push avYY.N.Ptag on that exactmaincommit. - Wait for tag CI to build
:prod+ the immutable:<version>+:<sha>images. Runscripts/security-scan.shon jazz against the immutable:<version>image, not a moving tag, and triage every ZAP/semgrep finding. - Only after the candidate passes, manually
DeployStack jazz-mediaand observe its pre-deploy output. Prod's compose deliberately follows floating:prod(Timothy's 2026-07-11 decision), so no CI push or pin bump is needed.
The Komodo stack is
jazz-media, notmedia-servers(verified live 2026-07-20 during the v26.11.0 cut). The compose project is stillmedia-servers— which is what the container labels show — but the Komodo stack name changed with the move to jazz. A stack namedmedia-serversstill exists on bumblebee and isunhealthy(the stopped migration leftovers), soDeployStack media-serverssilently targets the dead stack. Confirm with/read ListStacksbefore deploying.There is no Global Auto Update fallback anymore:
jazz-mediahasauto_update: false(poll_for_updates: trueonly), so nothing promotes:prodon a timer — promotion is manual, full stop. The old "don't cut a tag near the 03:00 run" caveat no longer applies.
server-management#585 source-confirmed that Global Auto Update invokes the same DeployStack
execution as a manual promotion, and extended the #553 pre-deploy hook to detect a floating-tag
digest change. Either path now takes the fail-closed prod backup; server-management#589 then
wired migration-smoke.sh into that hook, against the exact candidate and the backup it just made.
A backup, fetch, or migration-smoke failure aborts before the live container is recreated. See
homelab-docs/Docker/ErsatzTV.md for the operational evidence and rollback procedure.
Gotcha: never put a [skip ci] token in a commit you intend to tag — Gitea reads skip-ci from the tagged commit and will suppress the release build. (Also, workflow_dispatch on a tag ref isn't supported on this Gitea version, so the tag push must do the triggering.) Release commits, and anything you'll tag, must not contain skip-ci.
Also avoid firing several pushes back-to-back (e.g. a [skip ci] commit, then main, then a tag, all within ~1s). Observed once on this Gitea instance: the later events were silently dropped — no ActionRun records created at all, even though the runner was online and the workflow active. Pushing again, spaced out, created the runs normally. If a push/tag doesn't produce a run, re-push (or push an empty commit) rather than assuming the runner is broken.
The workflow: .gitea/workflows/docker-build.yml
Single workflow. Gating jobs test + migrations run in parallel and gate build; a
non-blocking docs-reminder job runs on PRs only (see below). Prod deploy is not a CI
job — it's Komodo Global Auto Update off the :prod tag (see "Cutting a release").
Triggers & tags
| Trigger | test job |
build job |
Image tags pushed |
|---|---|---|---|
pull_request |
✅ | — (skipped) | none |
push to main |
✅ | ✅ | :latest + :<short-sha> |
push tag v* |
✅ | ✅ | :prod + :<version> + :<short-sha> |
workflow_dispatch |
✅ | ✅ | only if ref is main/v*, else build-only (no push) |
A docs-only change (see "Docs-only skip" below) reduces every ✅ above to a seconds-long no-op that still reports its status.
:latest is the test/dev channel (every main commit). Prod's compose follows the
floating :prod tag (reverted from the 2026-07-07 version pin on 2026-07-11) — never
:latest. Both :prod and :<version> are produced by pushing a v* tag; prod tracks
:prod and is redeployed by Komodo Global Auto Update (see "Cutting a release"). The
immutable :<version> tags remain for reproducible rollback (docker run …:26.6.0).
Concurrency is scoped per event+ref (group: ersatztv-build-${{ github.event_name }}-${{ github.ref }},
cancel-in-progress for PRs): PR runs parallelize across PRs, a new sync auto-cancels
its superseded run, and image builds still serialize within their own ref. Do NOT push
main and a v* tag simultaneously — those are separate groups but share the
:buildcache tag and the smoke container name; tag only after the main build is green.
(History: originally one global group serializing ALL runs for the single runner —
with three runners that starved the queue; changed 2026-07-11, server-management#574.)
Four runners serve the fork (server-management#570/#574/#639):
| Runner | Host | Label | Slots | Per-job cap |
|---|---|---|---|---|
ci-runner |
VM 127 (pve4) — no prod workload | ubuntu-latest |
4 | --cpus=4 --memory=10g |
bumblebee-runner |
bumblebee — prod media | ubuntu-latest |
2 | --cpus=4 --memory=10g --cpu-shares=256 |
small-runner |
bumblebee — prod media | small |
2 | --cpus=1 --memory=1g --cpu-shares=256 |
jazz-small-runner |
jazz — prod media (#633) | small |
2 | --cpus=1 --memory=1g --cpu-shares=128 |
The small lane exists because Gitea dispatches a job as a runner task even when its
if skips it, and those skip-tasks used to wait behind long builds (observed 31 min),
stalling every PR run. --cpu-shares below the default 1024 is what makes a runner on a
prod media host acceptable: under contention CI loses to the transcoders (ersatztv 1536 /
jellyfin), which are the reason those hosts exist.
That same "dispatched even when if skips it" behavior is why the git-only PR gates live in
their own PR gates workflow (pr-checks.yml, on: pull_request)
rather than in docker-build.yml — see that section (ersatztv#535).
small is git-only, and that is load-bearing (server-management#639). Everything in
the lane is a checkout plus a git diff: decisions-guard, ci-image-pin,
docs-reminder — plus script-tests, which is a checkout plus a pytest run needing only
pytest and pyyaml (ersatztv#631; it is NOT stdlib-only — that assumption is what turned the
job red on its first CI run, see below). Nothing there runs a compiler or a docker build, which is why the lane
can be capped at 1 GiB per job. The lightweight-Python jobs are the deliberate edge of the
"git-only" rule, not an exception to it: setup-python + pip install pytest + a suite whose
heaviest allocation is a handful of temp-dir git repos stays far under the cap. Route a heavy job here and it will OOM — give it
ubuntu-latest, or its own label on ci-runner, the only host with no prod workload.
Lane assignment (ersatztv#390). Slot counts below are as-of 2026-07-17; the table above is
current. At the time, the ubuntu-latest lane had 4 slots (2 + 2) and the
small lane 4. A 2026-07-17 audit of the Actions API found the ubuntu-latest lane
saturated and the small lane idle — queue wait exceeded every job's runtime:
| Job | Runtime | Queue wait | Lane |
|---|---|---|---|
test |
354s | 1363s | ubuntu-latest |
migrations |
639s | 1428s | ubuntu-latest |
functional-e2e |
520s | 1447s | ubuntu-latest |
api-docs |
5s | 1722s | ubuntu-latest → small → reverted to ubuntu-latest (#406) |
format |
37s → ~0.5s (#469) | 1731s | ubuntu-latest → small → reverted to ubuntu-latest (#406) |
docs-reminder / decisions-guard |
10s | 5s | small |
api-docs and format moved to small because the queue wait dwarfed their runtime. Both lanes
run the identical runner-images:ubuntu-latest base, so small was a label with spare capacity,
not a different capability — a move only possible because those jobs now run in the CI toolchain
image (below) and no longer need the runner image to supply .NET/Node.
Reverted 2026-07-17 (ersatztv#406 / server-management#604). #390's own caveat — "on an
API-touching PR api-docs does a full dotnet build, so it is not always small" — turned out to
be the deciding factor, and "capacity 4 absorbs that" held only because nothing enforces the sum
of the lanes' per-job caps. Each job container is correctly capped (--memory=10g), but 6 slots ×
10 GiB = 60 GiB on a 25 GiB host that also runs prod media; on 2026-07-17 bumblebee hit load
340 with 21 GiB swapped. These were not small jobs — a live docker stats caught the format job
container at 3.95 GiB, which the re-sized 2 GiB small lane would OOM-kill outright. #604 fixes
the queue at the source instead (ubuntu-latest grown to 5 slots: a 48 GiB ci-runner at capacity 4
plus a bumblebee overflow slot), so the small lane can be reserved for genuinely-tiny shell jobs.
(The format half of this is now moot: ersatztv#469 moved it to dotnet format whitespace . --folder, which loads no Roslyn workspace — the job's 3.95 GiB heap and multi-minute runtime are
gone, so it is no longer a reason to keep the lane large. api-docs on an API-touching PR still is.)
Queue wait is still a dominant cost and capacity is server-management's boundary — tracked in server-management#604. The redundant triple-build behind those runtimes is ersatztv#398.
The other failure mode: setup-phase starvation (server-management#639, 2026-07-20). The
table above measures queue wait — time before a job is dispatched. A saturated lane also
produces a second, much more confusing symptom: a job that is dispatched, sits in_progress
for >10 minutes, writes no log file at all (OpenLogs … .log.zst: file does not exist),
and then fails — wedged in act's job-setup phase, before Checkout. Same-config siblings
that started 90s earlier finished in seconds; a concurrent job's log showed a normally-fast
compile taking a 7-minute gap between projects. This is the origin of the "decisions.md is a
known flake, just rerun it" folklore: the rerun succeeds only because it lands after load
clears, so the guard's logic gets blamed for a capacity problem.
The fix was not more capacity for its own sake. small was stuck at one slot because it
still held two heavy jobs — docker-build.yml's image build and ci-image.yml's toolchain
buildx (the latter reads as lightweight because it is "docker-only", but it is the heaviest
thing that ran in the lane) — and their 10 GiB requirement set the lane's per-job cap, which
on a 25 GiB host permits exactly one slot. Moving both to ubuntu-latest made the lane
genuinely tiny, so it could widen to 4 slots across two hosts while committing less RAM to
CI than the single slot did. docker-build.yml's build does not re-create #574's
skip-task queueing, because needs: [test, migrations] means it cannot be dispatched until
the lane it would queue behind has already drained.
CI build memory: no persistent compiler servers (ersatztv#406)
Roslyn's VBCSCompiler is a persistent compiler server — it outlives the dotnet build that
started it and keeps its heap warm for the next one. Locally that is a genuine speedup; in CI it
buys nothing, because each job container is torn down at the end of the run and there is never a
"next build" to warm. It was measured at 7.8 GB RSS on bumblebee — the single largest consumer
on the host, and the reason each job needed a 10 GiB cap in the first place.
So the workflow's top-level env: disables the servers for every runner-side dotnet job:
| Variable | Effect |
|---|---|
UseSharedCompilation=false |
no persistent VBCSCompiler; csc runs per project and exits |
DOTNET_CLI_USE_MSBUILD_SERVER=0 |
no persistent MSBuild server process |
MSBUILDDISABLENODEREUSE=1 |
MSBuild worker nodes exit with the build instead of lingering |
These are MSBuild properties set as environment variables so they apply to every dotnet
invocation without touching each call site (MSBuild surfaces env vars as properties, and
UseSharedCompilation is only defaulted to true when empty, so the env var wins).
What this does and does not shrink. It helps the jobs that compile — test, migrations,
api-docs on an API-touching PR, and the in-Docker build. It never helped format: dotnet format loaded Roslyn in-process via MSBuildWorkspace and never spawned csc, so the compiler-server
env vars left its measured 3.95 GiB untouched. (Moot since ersatztv#469 switched format to
dotnet format whitespace . --folder, which skips the MSBuild/Roslyn workspace entirely — the job is
now a ~0.5 s, low-memory whitespace/BOM check with no Roslyn heap. See Static analysis &
formatting → Formatting below.)
The same three are repeated in dependency-scan.yml; workflow env: does not cross workflow files.
That one is lower-stakes (restore/list are MSBuild-driven, so it's lingering worker nodes rather
than a 7.8 GB VBCSCompiler) but it runs unattended on a cron against the prod media host.
The workflow env: does not reach the build job's compilation, which happens inside
docker build — the same three are set as ENV in the SDK stage of docker/Dockerfile. That is
the job server-management#570 measured pegging 5.999/6 GiB, so it is the one that most needs this.
Build-stage only; the final image is FROM runtime-base, so nothing lands in the shipped image.
The test job samples its own memory and reports it every run. Two steps (continue-on-error, so
they never fail a build), driven by scripts/ci-peak-anon.sh (ersatztv#412):
- Start peak-anon sampler (before the dotnet Build/Test/Coverage steps) launches a detached background poller that tracks the high-water mark of the cgroup's anon memory every 2 s.
- Report peak container memory (the job's last step) stops the sampler and prints the sampled
peak anon — the headline number — alongside
memory.peakand the end-of-jobanon/filesplit, to the log and the job step summary.
Read the peak anon off a recent run to size a cap — not memory.peak, and here is why:
⚠️
memory.peakis not "peak RSS". It is the high-water mark ofmemory.current, which charges page cache to the cgroup as well as anonymous memory. Proven on bumblebee: a container withanon=0that merely reads an 800 MB file reportsmemory.peak=826 MiB, of whichfile=800 MiB.This matters because the naive reading inverts the decision: page cache is reclaimed under a tighter cap, not OOM-killed, so a big peak that is mostly
fileis not evidence that the cap must stay high.anonis the part that actually forces an OOM. Size caps on peakanon, not onpeak.Why a sampler and not just the end-of-job split: the kernel exposes
memory.peak(peak of anon+cache) but has no peak-anon counter, and the end-of-jobanonis the composition then, not at the peak instant — a job that peaks mid-dotnet testand then frees reports a misleadingly lowanon. The 2 s background sampler catches the true peak-anon instant;memory.peakand the end-of-job split stay in the report as a cache-inflated ceiling and a reference. (Before #412 the instrument printed onlymemory.peak+ the end-of-job split — see #411.)
The compiler-server A/B (ersatztv#412). Measured on bumblebee in the CI toolchain image,
swap-off (--memory-swap == --memory), full-solution dotnet build --no-incremental, peak anon
sampled by this instrument, servers shut down between arms (n=2 each, interleaved):
| Arm | shared-compilation env | peak anon (2 runs) | resident after build |
|---|---|---|---|
| OFF (the CI config) | disabled | 5818 / 5854 MiB (~5.84 GiB, tight) | none |
| ON (dotnet defaults) | enabled | 6305 / 7604 MiB (~7.0 GiB, noisy) | ~3 GiB VBCSCompiler |
Two things are solid: OFF is consistently ~5.84 GiB and ON is always higher (mean delta
~1.1 GiB, up to ~1.8 GiB), so disabling the servers is worth it; and ON leaves a ~3 GiB
VBCSCompiler resident after the build — the host-between-jobs cost #406 removed. Don't read a
precise delta into the ON peak: it is noisy because a parallel build's peak depends on how many
csc/project compilations overlap at the peak instant.
#406's premise — "if disabling shared compilation brings peak RSS well under 6 GiB, the whole
budget loosens" — is NOT supported. OFF sits at ~5.84 GiB for the build phase alone — right at
the 6 GiB line, not well under it — and the test job adds dotnet test + coverlet +
reportgenerator on top. Disabling the compiler servers stays right (consistent reduction, no 3 GiB
resident server) but do not bank a looser cap budget on it: size the cap off the live
test-job peak anon this instrument now reports (build + test + coverage), not off this build-only
A/B.
The earlier PR #411 probe (
peak 9457 / anon 7134 / file 421 MiB) read higher than these swap-off sampled numbers. Swap settings and read-method (end-of-job snapshot vs sampled peak) move these figures by >1 GiB (#406), so treat the committed instrument's sampled peak-anon as authoritative and that probe as superseded.
What is established: no persistent compiler server survives a build, migrations is green with
mysql capped at 2g swap-off, and the test job now self-reports a true peak-anon every run.
services: containers are capped explicitly (ersatztv#406)
A runner's container.options (--cpus=4 --memory=10g) applies to the job container only, not
to services:. Verified on a live migrations job: the job container reported
HostConfig.Memory=10737418240; its mysql:8.4 service reported mem=0 nanocpus=0 — unbounded.
So migrations runs added an uncapped MySQL to an already-tight host.
The mysql service now sets --memory=2g --memory-swap=2g --cpus=2.
--memory-swap is the part that matters, and it is easy to get wrong. Docker defaults an unset
--memory-swap to twice --memory, so --memory=2g alone grants 2g RAM plus 2g of swap.
Verified on bumblebee:
| Options | memory.max |
memory.swap.max |
|---|---|---|
--memory=2g |
2147483648 | 2147483648 ← 2 GiB of swap |
--memory=2g --memory-swap=2g |
2147483648 | 0 ← swap disabled |
Setting --memory-swap equal to --memory disables swap for the container. On this host that is
the whole point: swap thrash is what took prod down, and a swapping mysqld mid-DDL is precisely the
pathology behind the known Command Timeout expired migrations flake. Prefer a loud OOM over
silent swapping — an OOM is a clear signal to raise the cap; swapping just degrades everything.
⚠️ The same 2× applies to the runners' container.options: --memory=10g — each job slot is
really 10 GiB RAM plus 10 GiB swap. The "6 slots × 10 GiB = 60 GiB on a 25 GiB host" framing
understates the promise by 2×, and it is a plausible direct mechanism for the incident's 21 GiB of
swap. Fixing that is server-management#604's call (reported there).
On the 2g figure, honestly: a mysql:8.4 container with this exact env peaked at 543 MiB
during init and settled at 481 MiB idle (probed on bumblebee, 2026-07-17) — but that is init+idle,
not the 787-migration replay, which grows table/definition caches idle never touches. So 2g is a
measured floor plus headroom, not a measured ceiling; the migrations job going green is what
validates it. --cpus=2 has no measurement behind it at all — 787 sequential DDL statements on one
connection are ~1-core-bound, so it is judgement; revisit if the apply step's tail latency grows.
Any new services: container needs its own explicit cap — it will not inherit one, and it needs
--memory-swap set alongside --memory or it silently gets 2× in swap.
test job
dotnet restore → strip the Scanner project ref (sed -i '/Scanner/d', matching the
Docker build) → dotnet build -c Release → dotnet test -c Release --no-build. Gates
the image build.
- Code coverage (ersatztv#15):
dotnet testruns with--collect:"XPlat Code Coverage" --settings coverlet.runsettings --results-directory ./coverage, socoverlet.collector(referenced by every*.Testsproject) emits a Cobertura report per project. A follow-up Coverage summary step merges them with ReportGenerator (TextSummaryto the log,MarkdownSummaryGithubto the job step summary). No floor is enforced yet ("decide on a floor later" — #15); the step iscontinue-on-error: true, so a missing report or a transient tool install never blocks a build.coverlet.runsettingsexcludes generated EF migration code (**/Migrations/*.cs, ~2.59M generated lines vs ~200k authored). Instrumenting it OOM-killed the sharedtestjob (exit 137); excluding it cuts the instrumented surface ~126× (2.5M→20k coverable lines in the whole-solutionArchitecture.Testsprocess) and makes the percentage reflect authored code.
- Shallow checkout:
fetch-depth: 1(ersatztv#190) — this job never runsgit describe/git log, onlybuildneeds full history/tags for version computation, sotestandmigrationsboth check out shallow.build's checkout staysfetch-depth: 0. - NuGet package cache: both
testandmigrationscache~/.nuget/packagesviaactions/cache@v4, keyed onhashFiles('Directory.Packages.props', 'global.json')with arestore-keysOS-level fallback (ersatztv#190). Avoids a from-scratchdotnet restoreon every run; the key only changes when the central package manifest or SDK pin changes.
build job
- Compute
INFO_VERSION(git describe+ short sha onmain; tag version onv*). docker/setup-buildx-actionwithbuildkitd-config-inlinesettinghttp = truefor192.168.1.95:3000— BuildKit does not inherit the host daemon'sinsecure-registries, so without this, cache/base-image/push over the HTTP registry fails (http: server gave HTTP response to HTTPS client).docker/login-actionwith repo secretsREGISTRY_USER/REGISTRY_PASSWORD.REGISTRY_PASSWORDis a scoped PAT (write:package+read:repository), not an account password — deliberately, so head-resolved PR code cannot use it to forge a commit status (ci.actions-credential-scoping, ersatztv#697). If a job ever fails withtoken does not have at least one of required scope(s), the fix is to narrow what the job does, never to widen the token towrite:repositoryor to put the admin password back. Note what the scope still reaches:write:packagecoversersatztv:prod(the tag prod's stack follows) andersatztv-ci:<sha>(the toolchain image fivecontainer:jobs execute), so this is the deployment supply chain, not an inert endpoint — seeci.actions-credential-scoping.docker/build-push-action@v6: amd64-only,docker/Dockerfile,INFO_VERSIONbuild-arg, registry layer cache (type=registry,ref=…:buildcache,cache-to … ignore-error=true).- Smoke + IPTV E2E test: pull the just-pushed
:<sha>, run it, poll for HTTP readiness (docker exec … python3→http://localhost:8409/), then assert the real Jellyfin-facing surfaces on the freshly built image (ersatztv#16):/iptv/channels.m3ureturns 2xx containing#EXTM3U, and/iptv/xmltv.xmlreturns 2xx containing a<tvroot.xmltv.xmlneedschannels.xml(written by the scheduler a few seconds after boot), so each endpoint is polled with a deadline. Unique container name +trap … EXITcleanup; dumps container logs on failure. Catches routing / base-URL (#1) / migration regressions that leave the app "up" but serving broken output.
functional-e2e job (advisory; PR + main)
Boots the app from source and drives the manual live-E2E curl flows sessions have historically
re-run by hand, turning them into a CI regression net (ersatztv#299). It is the automatable half of
docs/e2e-local.md; every script it chains runs identically locally and in CI:
npm ci+npm run build(SPA),dotnet build ErsatzTV.sln -c Release. (ffmpeg comes from the CI toolchain image; the oldapt-get install ffmpegstep is gone — see below.)ETV_BUILD_CONFIG=Release scripts/e2e-local.sh <fresh-config>— copieswwwroot, launchesdotnet ErsatzTV.dllin the background (logging to a file so the launch step returns once the app is ready), printsPID/CONFIG_DIR.scripts/e2e-functional.sh http://localhost:8409 <config>— asserts. Mostly curl-only: the legacy→SPA redirect sweep (+ the/api,/artworknever-redirect exemption), the auth/CSRF/security-stamp flow (setup-claim → read-gate 401/200 → re-claim 409 → CSRF 403 → login 401/200 → logout 403/204 → post-logout stamp-revocation 401), the library-scan status contract (404 unknown / 202 queued /scan-status200), and the If-Match/412 round-trip onrerun-collections. Since ersatztv#363 (extended by #444) it also asserts three lock-contention 409s that aren't curl-only — see below. Atrapkills the instance on step exit.ETV_UI_PORT=8410 scripts/e2e-ui.sh— boots a second, fresh instance and runs the headless Playwright UI flows (ersatztv#445). See "UI-E2E step" below.
Lock-contention 409s (ersatztv#363). The harness now seeds rows the API can't create — a
LibraryPath and a Jellyfin media-source — directly into the running instance's SQLite DB (via
python3's stdlib sqlite3, whose busy-timeout retry serializes behind the app's writer), and synthesizes media with the
image's ffmpeg, to exercise two IEntityLocker contracts deterministically (it only fires the
racing request once the lock is provably held, never a sleep-and-hope): (a) the library-scan
"already scanning" 409 — seed ~60 tiny clips into the built-in Shows library so the scanner
subprocess runs a few seconds, poll GET /libraries/scan-status until the library shows active (that
window is a strict subset of the scan lock's held window), then a second POST .../scan is a 409
(deterministic bar a tiny residual TOCTOU gap the multi-second scan covers); (b) the
external-collections "already scanning" 409 — the per-family lock is taken
synchronously before the 202, so the 202 proves it held, and pointing the seeded source at a
non-routable address keeps the background sync hung so the window stays open; and (c) the
playout-build "build in progress" 409 + isLocked projection (#215/#444) — a build is enqueued
onto the single-consumer WorkerService and the trigger returns before the handler locks, so poll
GET /playouts/{id} until isLocked:true (not the accepted trigger), then fire. Seed a Classic Flood
schedule over a few short episodes and crank PlayoutDaysToBuild (playout.days_to_build) so one
build is wide enough to observe (~5 days ≈ 43k items ≈ ~1s locally, wider on slower CI); assert PUT /playouts/{id} 409, POST .../playout/reset 409, the list isLocked:true, then post-build the same
PUT 200s. No new CI step or dependency: ffmpeg ships in the toolchain image, and python3 was already
a harness dependency (json_field). The scan + build flows self-skip (advisory) if ffmpeg is ever
absent, and the build flow also self-skips if the build is never observed locked (never asserts an
unproven race).
Advisory, by design (the issue's "keep it a separate job so a functional-E2E flake can't block the
unit-test gate"): it is not a needs: of build and not (yet) a required check, so a flake
blocks nothing. Promote it to a required check / build dependency once it's proven reliable — the
same staged rollout the migrations job used. SQLite is the default provider, so unlike migrations
it needs no DB service container. Runs on PRs and on main (regression net); skipped for v* tag
builds. The playout-build lock 409 + isLocked projection (#215) landed in #444.
UI-E2E step (ersatztv#445). The last step of this same job (item 4 in the summary above; the
Run UI-E2E Playwright flows (headless) entry in the YAML) runs the headless-browser flows the curl
harness structurally cannot express — client-side form validation, AuthGate's rendered states, the
session cookie authenticating the SPA's own /api XHRs, and sign-out via the UserMenu:
ETV_BUILD_CONFIG=Release ETV_UI_PORT=8410 scripts/e2e-ui.sh
e2e-ui.sh owns the lifecycle (fresh config dir → boot → playwright test → always kill the server)
and exits with Playwright's status. Three deliberate choices:
- This job, not a new one. The dominant cost here is
npm ci+ the Release build, both already done; a separate job would duplicate them to add ~5s of browser work. The browser is baked into the toolchain image, so the step installs nothing. - Its own fresh instance on port 8410. The first spec asserts the one-shot Setup gate, which step 3's auth section has already claimed on its config dir; a separate port also keeps this step independent of step 3's teardown timing.
retries: 0,serial, single worker. #445 asked for deterministic flows and a retry would let a flaky flow merge looking green (measured: 4 consecutive clean runs, ~2s each). Full rationale and the rule for extending the suite:docs/e2e-local.md→ "UI-E2E harness".
Docs-only skip (ersatztv#416)
A change that touches only docs/** or *.md (anywhere: README.md, CLAUDE.md, handoff
files) has nothing for the heavy jobs to validate. Before this, such a change ran the entire matrix
— test, migrations (with its mysql:8.4 service), functional-e2e, format, api-docs —
~9 min of warm CI for a Markdown edit.
The mechanism, and why it is shaped this way. Each heavy job (test, migrations,
functional-e2e, build) runs scripts/ci-detect-docs-only.sh as its first post-checkout step
(id: detect), which emits docs_only=true|false to $GITHUB_OUTPUT. Every real step in the job
is gated if: steps.detect.outputs.docs_only != 'true'. On a docs-only change the job runs only
checkout + detect and reports success in seconds.
Shallow-checkout safe (the change-set diff). test/migrations check out fetch-depth: 1, and
a shallow clone has no origin/<base> tracking ref and no merge-base — so a three-dot
origin/main...HEAD diff errors, the fail-safe returns docs_only=false, and the skip silently
never fires (the first cut shipped this bug — every docs-only PR still ran the full matrix; caught by
#416's "verify on a real PR" box). The script therefore fetches the base and diffs against
FETCH_HEAD (always written by git fetch, resolves in a shallow clone) with a two-dot tree
diff (git diff --no-renames FETCH_HEAD HEAD) — no merge-base required. (api-docs/format avoided
the bug only because they check out fetch-depth: 0.)
The jobs are not if:-skipped. That is deliberate and it is the whole trap of this issue:
main's branch protection requires two checks by name —Build ErsatzTV Image / Build & test (.NET) (pull_request)andBuild ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request). If a docs-only PR produced no run for those (a workflow-levelpaths-ignore, or anif:-skipped job), those contexts would never report and the PR could never merge — the naive fix bricks docs PRs rather than speeding them up.- On Gitea 1.25.4 an
if:-skipped job reports commit-status stateskipped, a distinct state (verified with a throwaway probe, PR #418) — notsuccess. We do not rely on how branch protection treats askippedrequired context. Keeping the job running and gating its steps makes the required context reportsuccessunconditionally, which is safe by construction. - Non-required jobs may skip freely: production already proves a
skippednon-required context does not block merge (buildisskippedon every PR). Sobuildskips its image steps on a docs-only push tomain(docs are not in the image, so there is nothing to rebuild); tag builds forcedocs_only=falsein the script so a release is never skipped.
The detection biases toward running more: docs_only=true only when every changed path is
docs; any code path, a tag build, a non-merge push, or an undeterminable diff resolves to false
(run the full matrix). A false true would skip real tests on a code change — a correctness bug —
so every ambiguous case runs everything. The migrations job's mysql service still starts on a
docs-only run (a services: container starts with the job regardless of step if:), but the
expensive 787-migration replay is skipped; the service is capped and idle for seconds.
api-docs and format already short-circuit on docs-only changes via their own path detection (no
API path / no .cs changed → they pass in ~5s), so they needed no change. docs-reminder,
decisions-guard, ci-image-pin and script-tests keep running on docs-only changes — the first
two are about docs and must, and script-tests is unconditional by design (ersatztv#631).
Not in scope: the within-run triple dotnet build (ersatztv#398; measured and rejected as
build-once — see docs/decisions.md). The separate redundancy of running the whole matrix on a
PR and again on the merge-to-main over identical code (ersatztv#420) is addressed below.
Cross-run tree-identity skip (ersatztv#420)
A merge to main re-runs the entire matrix over code the PR's last run already validated — the
same redundancy as docs-only, but for identical code rather than docs. On a push to main that
is a real merge commit, test/migrations/functional-e2e each run
scripts/ci-detect-already-validated.sh as a second detect step (id: revalidate, right after
the docs-only detect), and every heavy step gains an added && steps.revalidate.outputs.skip != 'true' to its existing if:.
Skip condition — all four required, else fail-safe skip=false:
- the event is a push to
refs/heads/main; HEADhas a second parentHEAD^2(a real merge commit — the PR head CI already validated; squash, rebase, fast-forward, or a direct push have noHEAD^2, so they run);git rev-parse HEAD^{tree}equalsHEAD^2^{tree}— main did not advance since the PR's last run, a byte-identical tree;HEAD^2has a green Gitea combined commit status, queried via the API withETV_STATUS_AUTH. Trusting the aggregate.stateis sound: askippedcontext does not drag the combined state belowsuccess(verified live against this instance — a real merge commit with fourskippedPR-only contexts still reported.state == success), and the two required jobs never reportskipped(they always run and report a realsuccess/failure), so.state == successimplies they were green.
Those three jobs check out fetch-depth: 2 so HEAD^2 and its tree resolve.
Why it's safe: build is not gated. The three heavy jobs skip their steps (same
required-context reasoning as docs-only — they still run and report success in seconds), but
build always runs on main, ungated, building and pushing the image from that identical,
already-validated tree. No image ships from unvalidated source. The required contexts are
unchanged (Build & test (.NET), EF migration integrity (SQLite + MySql)) — no branch-protection
change.
Fail-safe bias. Any uncertainty — not a main push, no HEAD^2, a differing tree, a
missing/failing/non-success status, missing auth — resolves to skip=false and runs the full
matrix. A false skip could ship an under-validated image, so every ambiguous case runs everything.
Honest limitation — this fires rarely here, by design. The tree is identical only on a
fast-forward-equivalent merge: main did not advance since the PR's last green run and the PR
head was not rebased at merge time. Two routine patterns defeat it in this repo: (1) under parallel
merges main usually advances; and (2) — the bigger one — the standard workflow rebases a PR
before merging to resolve the docs/decisions.md lifecycle conflict (see MEMORY: the
"decisions.md conflict treadmill"), which mints a new head SHA whose tree was never itself
CI-validated, so the tree-match check correctly declines. So the skip is a genuine but occasional
win (clean, up-to-date, un-rebased merges in quiet periods) — correct-but-conservative by
construction, not a general dedup. It never fires unsafely; when in doubt it runs the full matrix.
Dropped-step guard on the required jobs (ersatztv#756)
test and migrations write the only two docker-build.yml contexts branch protection requires on
main. A step the runner declines to interpolate is dropped, and the job still concludes
success (ersatztv#751), so in these two jobs that failure is fail-OPEN: a required check
reports green having done no work. In review-verdict.yml the same drop is fail-closed — the status
is simply absent and the merge is blocked — which is why #751 fixed the safe direction first.
Two independent mechanisms hold it, and neither is redundant:
-
A static ban on expression delimiters in any
run:body oftest,migrationsandbuild. The drop mechanism requires an opener in the scalar, so this makes the class unreachable rather than merely detected — and it is the raw${{opener that is banned, not a well-formed pair, because an unclosed one triggers the same rewrite. When a step genuinely needs a value, pass it through the step'senv:block, which is interpolated per value, so a bad payload there cannot take the body with it.Why
buildis in the ban although it is not a required context. Its one delimiter-bearing body wasSmoke + IPTV E2E, which runs afterBuild and push— so on av*tag the image is already in the registry as the release candidate and that step is what decides whether the candidate was ever booted. A drop there publishes an unsmoked candidate, reports green, andDeployStack jazz-mediapromotes exactly that image. Its two payloads moved into the step'senv:, so the ban cost nothing.The ban is re-checked on the release path itself (ersatztv#767). It used to be enforced only by the
script-testsjob, which lives inpr-checks.yml(on: pull_request) and is not a required context — a review-time check on the PR that would introduce a delimiter, not a gate on the release.pr-checks.ymldoes not run on av*tag push at all, so a delimiter that ever reachedmainwould still dropSmokeon the tag build and go green;mainbeing PR-only (#743) meant such a change had to pass through a PR wherescript-testsreddens, but a red on a non-required check does not block the merge server-side.There is now a
scanjob (Delimiter ban (release path)) that runs the PyYAML-based ban test, andbuildlists it inneeds:. That single edge is the fail-closed property: a redscanmeansbuildis skipped outright, so the image is never built, let alone pushed.Why a job and not a step inside
build. A step cannot protect the job it lives in.buildis what publishes, so a guard step there fails open if the runner drops it — and the defence ("the guard's own body has no opener, so it cannot be dropped") is circular when the only thing enforcing that property is the same PR-only test being backstopped. This was the first design and two independent reviews rejected it for exactly that.Why it runs the real pytest and not a bespoke scanner. The same first cut hand-parsed the workflow YAML in stdlib Python, to avoid provisioning PyYAML on
build's bare runner. Review found ~10 false negatives in that parser in one round — flow mappings ({run: …}), a quoted"run":key, aliases, multiline quoted scalars — making it strictly weaker than the check it backstopped, in the only direction that matters for a security gate. Running the existing test needs no second definition of "what is arun:body", so it has no drift surface at all.scanruns onsmalland provisions Python the same wayscript-testsdoes.The wiring is held by
scripts/tests/test_ci_release_path_scan_job.py—builddepends on it, it carries no job-levelif:(one that excluded the tag push would restore the hole; one that skipped the job would skipbuildtoo), no step iscontinue-on-error, it actually invokes the ban test, and every one of its ownrun:bodies is delimiter-free. Its steps also carry #756 markers and a trailing assert, so a drop inside this job is caught as well.What this does not claim: that no step can ever fail to run for a reason other than the interpolation drop. It moves the terminal assumption — to fail open you must now drop the pytest step and the assert step, rather than either one alone.
Measuring a guard on this path does not require cutting a release, and an earlier draft here claiming it would was simply wrong:
buildruns on every push tomain(if: github.event_name != 'pull_request'), and aworkflow_dispatchon any other ref runs the job whileBuild and pushpublishes nothing (itspush:is gated onmain/v*). That is how #767 was verified — see the decision record for the run ids.functional-e2eis delimiter-free too but is deliberately not banned: it is advisory by declaration, and the rule is "ban where a drop is consequential", not "ban wherever it is currently free".api-docsandformatkeep onegithub.base_refeach in a detect step and gate nothing that ships. -
Runtime per-step markers, for a step that fails to run for any other reason. Every
run:step that is notcontinue-on-error: truecalls"$GITHUB_WORKSPACE/scripts/ci-step-ran.sh" mark <key>as its first act, and the job's last step callsci-step-ran.sh assert --always … --gated …, which fails the job when an expected key was never recorded.
Per step, not per job. A marker written by the first step only proves the job began, which was
never in doubt. The drop that costs something is Test, Build or a migration replay — all well
past step one — so a job-level marker would have been a guard that cannot see the case it exists for.
The guard carries no if:, and that is deliberate. The #751 guard uses if: always() because its
job has one real step. These have a dozen, and a genuine failure in an early step legitimately skips
every later one — an always() guard would then announce a false "these steps never executed:
typecheck web-test build dotnet-test" on top of every ordinary red build, and a guard that cries wolf
gets deleted. (migrations is smaller — six marked steps — but the same argument applies, and its
guard comment is worded for its own keys rather than copied from test's.) The default if: is success(), which is the wanted condition, and the invariant that
makes relying on it safe rather than lucky is: the guard is skipped only when an earlier step
failed, and that already fails the job. So guard skipped ⇒ job red, and every path to a green job
runs the guard. A dropped step is invisible precisely because it concludes success — which keeps
the job green and therefore reaches the guard.
That invariant has one path where it could plausibly be false and where being wrong would be silent:
a step marked continue-on-error: true that FAILS. If that flipped success(), the guard would be
skipped on a job that still concluded green — the guard rendered a no-op by exactly the failure mode
it exists to catch, with no signal. The test job has three continue-on-error steps and two of them sit immediately before the guard,
so this is a live path, not a theoretical one. Measured (scratch PR #766, run 1913 job 8075): the last
advisory step was made to exit 1, the log carries ❌ Failure - Main Report peak container memory,
and the guard still ran, reported All 12 expected step(s) executed, and the job concluded
success. A failing continue-on-error step does not flip success() on this runner, so the
invariant holds where it mattered most. That run is also the test job's full twelve-key positive
control on the build lane.
Adding a step to either job? Mark it, and add its key to that job's guard list in the right
bucket (--always for the two detect steps, --gated for anything carrying the docs-only /
already-validated if:). scripts/tests/test_ci_dropped_step_guard.py derives the expected set from
the workflow, so an unmarked step or a bucket mismatch is a red — it does not rely on anyone
remembering. One caveat, since this section is careful about it elsewhere: that red is script-tests,
the same non-required, PR-only check discussed above. For the delimiter ban on test/migrations
that hardly matters, because the runtime guard is the fail-closed backstop — but a newly added,
unmarked step is caught by the static test alone, since the runtime guard cannot expect a key
nobody declared.
The marker path is keyed on job + run id + attempt — and be precise about why, because the
obvious justification is a #751 measurement that does not transfer. #751 found RUNNER_TEMP to be
/tmp and called it "not a private per-job directory", but that was taken on review-verdict.yml,
which runs without a container:. These two jobs run inside the CI toolchain image, so their
/tmp is the job container's own and starts empty. The fresh container is therefore what actually
rules out a stale marker here; the keying is defence in depth against a lane change nobody would
think to re-check this against. GITHUB_JOB and GITHUB_RUN_ID are measured present and the
script refuses without them rather than falling back to a name other runs share.
GITHUB_RUN_ATTEMPT is required too — but how that was established is the part worth keeping,
because the first two attempts at it were both worthless. Grepping a job log for the variable name
proves nothing: logs do not dump the environment. Inferring it from the absence of the script's
"not set" warning proves nothing either, because that warning goes to stderr, and whether step
stderr reaches a job log here was itself never established — the control offered for that turned out
to be an ::error:: line this script writes to stdout. So the script was made to report its
resolved identity on stdout, where capture is not in question, and the answer was simply read off
this change's own run: Marker identity: job=test run=1916 attempt=1 (from the runner), and the same
for migrations. Both required jobs, on the lane that matters.
That measurement is what promoted it from warn-and-default to required, and it is why the residual this paragraph used to describe — a rerun inheriting attempt 1's markers — no longer exists. The identity line stays, as the standing evidence a future reader checks first if the keying is ever doubted again.
The premise was re-measured on the build lane. The whole thing rests on the runner still executing
a later step after dropping an earlier one. #751 established that on the small lane; these jobs run
in a container: on ubuntu-latest, so it was measured there rather than assumed — scratch PR #765
(Gitea 1.27.1, 2026-08-10) reintroduced the exact #751 defect in the test job's revalidate step.
Recorded outcome (job test, run 1910, 20:03:49→20:13:31Z — a full 9m42s heavy run, so Build and
Test really executed):
Unable to interpolate expression 'format('# PROBE ONLY … {0}\n…', pr number)'at 20:04:06 — the step was dropped, exactly as #751 describes, and it reported conclusionsuccess.- Every other marked step still ran — eleven markers were recorded, ten of them AFTER the drop
(
restore npm-ci check-api lint typecheck web-test web-build strip-scanner build dotnet-test),detectbeing the eleventh and earlier. The premise holds on this lane. - The guard ran at 20:13:29, reported
These steps of job 'test' never executed: revalidate, and was the only ❌ in the entire job log — every other step succeeded. Without it this run would have concludedsuccesshaving never executed that step, which is precisely the fail-open being closed. - Incidental but kept: the dropped step's output arrived as
ETV_REVALIDATE_SKIP:empty, notfalse— the case the guard must read as "widen what is required", never as a skip.
The positive control is the same run's migrations job, which the probe did not touch: it marked
all six steps, the guard reported All 6 expected step(s) executed: detect revalidate restore build sqlite mysql, and the job concluded success. So one run demonstrates both directions on the build
lane — a drop caught and reddened, and a clean job passing. The twelve-step test positive control is
this change's own CI run.
Full rationale: docs/decisions/records/ci/required-job-step-execution-markers.md.
docs-reminder job (non-blocking, PR-only — in pr-checks.yml)
A lightweight nudge that enforces the CLAUDE.md "docs-update is part of done" rule for the
one case that's easy to forget and easy to detect: a PR that touches a SPA screen
(web/src/screens/*.tsx) or ErsatzTV/LegacyUiRedirects.cs but does not update
docs/blazor-route-parity.md. It diffs the PR against its base branch and emits a
::warning:: annotation (never fails the build — it's a reminder, not a gate; prose-doc
gates get gamed with token edits). Deliberately has no setup-dotnet/setup-node (and
thus no actions/cache), so it can't hit the cache-save hangs seen on the VM-127 runner
(server-management#570). It does not cover the remaining doc obligations in the CLAUDE.md table
(domain-model, spa-conventions) — those stay on the author. (The API contract is mechanized by the
blocking api-docs job, and docs/decisions.md by the blocking decisions-guard job below.)
decisions-guard job (decisions lifecycle, blocking, PR-only — in pr-checks.yml)
Enforces decision-record lifecycle invariants (ersatztv#521, supersedes the ersatztv#303 H9
append-only mechanic): well-formed 5-field metadata, exactly one active record per key,
reciprocal supersedes/superseded-by links, no record vanishing from the active set without an
archive copy, no rationale-prose rewrite without a Decisions-Edit: yes trailer on a non-merge commit
in the range (ersatztv#609), a structural per-path check that every *.md under
docs/decisions/records/** and docs/decisions/archive/** parses to exactly one keyed
record (ersatztv#621 — without it, a file the dependency-free frontmatter reader cannot parse,
such as one using a YAML block scalar, yields [] and vanishes from the corpus with every check
still reporting green; a file directly in archive/ is exempt only when it really is a stripped index — one keyless
record with a known generated heading — never merely by its location; the single further exemption,
archive/README.md, is by exact relative path, never by basename, which would otherwise exempt the
same filename in the active wing),
and the generated active catalog (docs/decisions/README.md) in sync with source. Two steps:
scripts/decisions_validate.py --base origin/<base> --head HEAD (the merge-base diff checks, which
need a base/head range — CI-only) and scripts/build_decisions_catalog.py --check (catalog drift).
The same validator backs the Husky pre-commit hook (.claude/hooks/decisions-guard.sh, no
base/head there — structural checks only, over the working tree), so local and CI enforcement can't
drift on the rules that don't need a range. python3 isn't guaranteed on the bare small lane, so
the job adds actions/setup-python@v5 before invoking it; that install is lightweight (no
compiler/docker build), so it doesn't violate the "small is git-only" lane rule. Like
docs-reminder, otherwise a seconds-long git diff + parse with no dotnet/node setup
(runs-on: small).
script-tests job (Script tests (pytest), PR-only — in pr-checks.yml)
Reddens the run on failure, but like the other
pr-checks.ymlgates it is not one of the three required status checks onmain(Build & test (.NET),EF migration integrity,review-verdict/h10). Promoting it to required is a branch-protection change, tracked separately.
Runs the repository's Python test suite: PYTHONPATH=. python3 -m pytest scripts/tests -q
(~190 tests at time of writing, ~10s; the suite grows, so treat the figure as indicative). It covers the decision-corpus parser/validator/catalog builder, the ersatztv#610
migration-equivalence harness, the merge-consent exemption logic and the ersatztv#622 review-verdict
poster.
Until ersatztv#631, nothing ran these tests. No workflow and no Husky hook invoked pytest.
decisions-guard executes decisions_validate.py and build_decisions_catalog.py directly — it
exercises that code but never its tests — and the test job is dotnet test only. The suite
guarding our merge-gating machinery was therefore local-only, and a test added "for CI enforcement"
was decorative.
Why it is its own job, not a step inside decisions-guard. decisions-guard is covered by the
standing ci.decisions-lifecycle-flake rule: a lone decisions lifecycle red is a known infra
flake and sessions are instructed not to investigate it. Adding the suite there would make a
genuine pytest regression surface as precisely the red everyone is told to wave through — the same
"reports success while doing nothing" failure mode ersatztv#631 exists to close. A distinct job
name keeps a real failure unambiguous.
Why it runs unconditionally rather than behind a scripts/** path filter: the suite's true
input set spans more than one directory — test_post_review_verdict.py and
test_merge_consent_exemption.py execute the real scripts/post-review-verdict.sh and
.claude/hooks/pretooluse-merge-consent.sh — so a scripts/** filter would silently miss a
.claude/hooks/** edit. At ~10s, a filter buys nothing but drift.
Dependencies: pytest and pyyaml — the complete third-party set across scripts/, established
by an AST import scan rather than by reading the files that looked relevant. PyYAML does not
contradict the dependency-free decisions read path: decisions_lib._read_frontmatter is
hand-written exactly so validation runs where nothing is installed, but the one-shot write path
migrate_decisions_split.py uses PyYAML by design, and test_migration_equivalence.py imports that
module. (The first cut of this job claimed "pure stdlib + pytest", passed locally on a machine that
happened to have PyYAML installed, and went red in CI on a ModuleNotFoundError at collection —
which is itself a small demonstration of why the suite needed to run in CI at all.) Like the other
small-lane Python jobs it adds actions/setup-python@v5 first. Checkout is at default depth: every git call in the suite runs
against a temp repo it creates itself, never this repository's history.
Two preflight steps run before the suite. The first asserts git is on PATH; the second runs
scripts/jq-preflight.sh --expect 1.6, which checks jq's version, not merely its presence (see
"The jq contract" below). Those two tests exec the real shell scripts, which shell out to jq ~26
times; the tests shim curl on PATH but not jq, so a runner image without it would surface as ~20
opaque assertion failures instead of one diagnosis. Both deliberately check rather than install —
ersatztv#390 removed run-time apt-get from CI; the fix for a genuine miss is to bake the tool into
the runner image.
The jq contract (ersatztv#648)
Full rationale:
docs/decisions/records/ci/jq-version-contract.md.
Every shell gate in this repo — decisions-guard, script-tests's own harness,
pretooluse-merge-consent.sh, review-verdict.yml, scripts/pr-changed-files.sh — is authored and
tested on a developer Mac shipping jq 1.8.x. The CI runner ships jq 1.6. Author to the
1.6-compatible subset; three concrete constructs diverge between the two and each one produced a real
bug when it hit CI for the first time:
jq -eover EMPTY input. Exits 4 on jq >= 1.7, but 0 on jq 1.6. A guard that infers "transport failure" from that exit status silently passes an empty/failed page on 1.6.contains("\u0000")(or any NUL literal). The NUL escape truncates to""on jq 1.6, so the containment test is vacuously true for every string, not just ones containing a NUL. Useexplode | index(0)instead — it is version-stable.- Parse-error exit code.
jq emptyexits 5 on jq >= 1.7 but 4 on jq 1.6 — the same code 1.6 uses for "no output produced". Reading that exit code as a specific failure mode conflates garbage input with an empty-but-valid response.
scripts/jq-preflight.sh makes the running version observable in every gate job's log (it prints
the parsed version and asserts a floor of 1.6) so a future divergence can be diagnosed from the log
alone instead of guessing at the runner image.
Pin vs floor is deliberately asymmetric. scripts/jq-preflight.sh --expect 1.6 additionally pins
the version and fails loudly if it drifts, but that mode is used only by script-tests
(.gitea/workflows/pr-checks.yml) — advisory, not a required check. review-verdict.yml runs the
no-args floor-only mode and never pins, because that workflow writes review-verdict/h10, the
branch-protection-required status check on main: a hard pin there would mean the day the
runner's jq version changes (a base-image bump, a host reimage — nothing this repo controls), every
PR on main stops merging until someone notices and re-pins. A required merge gate cannot fail
because an upstream package manager did its job. The narrower pin on script-tests exists precisely
because that job is the suite's only 1.6 coverage — if the runner's jq silently changed, that coverage
would evaporate with no signal, so failing loudly there forces a human decision instead.
Baking a pinned jq into docker/ci/Dockerfile was considered and rejected: review-verdict.yml is
runs-on: small with no toolchain-image pin, and per ci.small-lane-git-only the small lane is
git-only, so it gets the host's jq regardless of what the toolchain image contains — a pin in the
image provably cannot reach the gate that broke. This was checked against the running binary, not
assumed.
PR gates workflow
File: .gitea/workflows/pr-checks.yml — on: pull_request only.
The four git-only PR gates — ci-image-pin, docs-reminder, decisions-guard, script-tests
(all described above) — live here, not in docker-build.yml, and that separation is the fix
for ersatztv#535.
Why they are split out. All three are pure checkout + git diff gates on the small lane
(no container:) and are PR-only (if: github.event_name == 'pull_request'). While they lived in
docker-build.yml — which also triggers on push to main and on v* tags — Gitea still
dispatched them as runner tasks on every such push to evaluate the skip (the small-lane
behavior documented above: a job is dispatched even when its if skips it). On the v26.12.0
release tag those dispatched skip-tasks wedged in act's setup phase and were killed by a runner
restart mid-setup, so they reported failure (no logs) and reddened the tag's overall commit
status even though the release built, scanned, and deployed fine. The two PR-only jobs on
ubuntu-latest (api-docs, format) carry the identical if: and skipped cleanly on the
same tag — the job logic was never the problem; the kill lands in the dispatch window, before any
step or if:-skip runs, so tweaking the if:/step logic could not fix it.
Why a separate workflow fixes it. Gitea evaluates a workflow's trigger before creating any
job, so a pull_request-only workflow produces zero jobs on a tag/main push: no dispatch, no
kill, no spurious red — for the whole class, permanently. The per-job if: guards are kept as
belt-and-suspenders (they also encode "these steps need a PR base_ref").
What stays put and why. These three carry no CI toolchain image pin, so ci-image-pin's
grep of docker-build.yml still validates the five pin-bearing jobs
(test/migrations/functional-e2e/api-docs/format) that remain there. api-docs and
format stay in docker-build.yml because they carry the shared-image container: + pin and run
on the healthy ubuntu-latest lane (where they skipped correctly). None of the three moved jobs is
a required check — branch protection requires Build & test (.NET), EF migration integrity
and review-verdict/h10 (next section) — so relocating them (their status-context prefix changes
from Build ErsatzTV Image / … to PR Gates / …) does not affect merges. The file declares
defaults: run: shell: bash because ci-image-pin uses mapfile/set -o pipefail.
Review-verdict gate (review-verdict/h10, required — .gitea/workflows/review-verdict.yml)
A required status check named review-verdict/h10, written per-sha, is what actually stops an
unreviewed commit from merging (ersatztv#622). It is not produced by a job's success/failure; it
is a commit status that scripts/post-review-verdict.sh POSTs onto one specific sha.
main is PR-only AND admin-override-proof, and it takes both to make the check load-bearing
(ersatztv#743, release.main-direct-push-disabled). Gitea evaluates status_check_contexts when it
merges a PR — a direct git push origin HEAD:main never consults them. So until 2026-08-05 the
entire gate was skippable with no forgery at all, which was cheaper than every route enumerated in
#697. main now carries two fields, and citing either alone is a mistake:
enable_push: false— a direct push is refused server-side at pre-receive (Not allowed to push to protected branch main), for every account including a site admin. The contents API is refused too — measured, HTTP 403user cannot commit to repo. The web editor, upload, apply-patch, revert and cherry-pick paths share that sameCanUserPushpredicate and are therefore expected to refuse as well, but were not probed (source-attested only).block_admin_merge_override: true— without it (the default isfalse), a repo admin couldPOST /pulls/{n}/mergewithforce_merge: trueand merge straight past a missing or redreview-verdict/h10. Disabling push alone just moves the bypass from the push path to the merge path, sincetimothyis admin and is the identity every session already uses. Source-attested, not probed (Gitea 1.27CanBypassBranchProtection): verifying it by experiment means merging an unreviewed PR, so the field was set rather than measured. Setting it is safe under either semantics; re-confirming the bypass itself rides with ersatztv#747.
Operator recovery when a required context gets stuck. block_admin_merge_override: true removes
the "Merge (admin)" / force_merge: true escape that used to unstick a PR whose required context was
absent or wrongly red — a recurring situation here (a killed run overwriting a newer green, an
advisory red counted into the combined status, a gate workflow that cannot post). That escape is gone
by design: it was also the bypass. The supported recovery is to fix the status
(re-run the job, or re-post the verdict with scripts/post-review-verdict.sh); the last resort is to
PATCH .../branch_protections/main setting block_admin_merge_override: false, merge, and set it
straight back. Do the last one deliberately and say so in the PR — it is the one action that
re-opens the hole this section exists to close.
Practical consequences: every change to main goes through a PR, including a one-line docs fix;
and the client-side Husky guards (H6/H11/H13) remain useful friction but were never the control —
they are fail-open and --no-verify bypasses them. Tag pushes are unaffected (separate mechanism;
tag_protections is empty), so the release cut in "Cutting a release" still works unchanged.
The hole it closes. pretooluse-merge-consent.sh proves its three consent conditions at the
moment the merge tool is called. Pass merge_when_checks_succeed=true and Gitea performs the merge
later, against whatever head is green then — while the Done-when and review-verdict checks were
proven against the head at scheduling time. Every commit pushed in between merges unreviewed.
This was demonstrated as a controlled A/B rather than inferred (ci/fake stands in for a slow CI
check so Gitea waits, as it really does): review head A → post its verdict → schedule auto-merge →
push an unreviewed commit B → CI greens on B. Without the required verdict context, B merged.
With it, the same sequence was refused, and merged only once B itself was reviewed.
Note the motivating anecdote in ersatztv#622 — "PR #619 merged 263 insertions with no verdict" —
is wrong: #619 does carry Review-verdict: MERGEABLE @ 02c82b35, posted six seconds before the
merge, explicitly re-reviewing the follow-up commits. It was filed from an API read that lagged.
The gap is real anyway, and structural: nothing forced that re-review inside the 45-minute window
where Gitea would have merged whatever went green. This turns diligence into construction.
Why a commit status fixes it and a smarter hook cannot. A status belongs to exactly one sha, so a new commit cannot inherit it: the required context is simply absent on the new head, Gitea's merge-requirement check reads that as not-passing, and the scheduled auto-merge refuses to fire. The invariant self-invalidates — nothing has to notice the push. It also covers merge paths the hook never sees (Gitea UI, raw API, another agent's session).
Posting a verdict. After reviewing a PR's current head:
ETV_GITEA_BASICAUTH=user:pass scripts/post-review-verdict.sh <pr> MERGEABLE [note...]
That posts both the Review-verdict: … @ <sha> comment (the human-readable artifact, and the
hook's condition (c)) and the review-verdict/h10 status on the same sha. BLOCKED /
NOT-MERGEABLE post a failure status instead. The script re-reads the head after commenting: if
a commit landed mid-flight it writes no status and exits non-zero rather than retargeting your
verdict at a commit you never read.
The status description also records the base branch — Review-verdict: MERGEABLE @ abc1234 (base: main) — and the merge-consent hook denies when that no longer matches the PR's live base.ref
(ersatztv#632). Retargeting a PR changes the effective diff without moving the head sha, so the
per-sha binding alone cannot see it. This is detection on the hook path only: a commit status
carries no base of its own, so a merge driven through the Gitea UI or API is unaffected. The
comparator is the base branch, never its tip sha — a base that merely advances is ordinary churn,
and comparing tips would invalidate every open verdict on every unrelated merge to main.
Exemptions are handled by review-verdict.yml on every pull_request_target event, which posts the
status as success for Renovate-authored PRs (it uses platformAutomerge: true, so a required
verdict with no exemption would stall every dependency bump) and for docs-only PRs, and as
pending for everything else so the block has a visible reason. Both exemptions are void when the
PR touches .claude/, .codex/, .gitea/, .husky/, scripts/ or docker/ci/ — a PR that can weaken the
gate must not be able to exempt itself from the gate. That includes Renovate's docker/ci base
bumps, which already need the manual publish-then-pin two-step anyway.
The Renovate exemption additionally requires every changed path to be a dependency manifest —
Directory.Packages.props or .config/dotnet-tools.json, and only those (ersatztv#698). The npm
manifests are deliberately excluded: renovate.json enables only nuget/github-actions/dockerfile,
so npm is unmanaged here, while package.json scripts are executed by CI (npm ci, npm run build)
— exempting it would put a code-execution path inside the allow-list for no benefit. An author match alone is not enough, because pull_request.user.login is the PR's
immutable creator while its head is not: pushing application code onto an open Renovate branch
leaves the PR still "authored by renovate" and, previously, still exempt. A Renovate PR touching
anything else — a .csproj, a source file — is not blocked, it just needs a real verdict. If a
dependency PR is unexpectedly asking for a verdict, this is why; the status description says so.
The two exemptions are evaluated as independent predicates, never as an elif chain: a Renovate
PR touching only docs/ still gets the docs-only exemption on its own merits.
An existing review-verdict/h10 on the head is only left alone when it is positively identifiable
as a human verdict — a non-null .creator.login and a Review-verdict: description, which is what
post-review-verdict.sh writes. Anything else, including any shape the workflow does not recognise, is
re-derived rather than inherited. (Measured: a status POSTed with a user credential carries a
creator; one POSTed by an Actions job carries "creator": null.) Without this, an exemption obtained
once was accepted unchanged on every later run. This is a provenance check, not an authentication
one — someone who can POST statuses directly can still impersonate a verdict (ersatztv#697). That
provenance asymmetry is why the credential scoping in ci.actions-credential-scoping mattered: a
forgery through a user credential inherits as a human verdict, while one through a job's
GITEA_TOKEN carries creator: null and is re-derived, so it must win a race. CI's registry secret
was a user credential — the admin account — and no longer carries status-write. RENOVATE_TOKEN
still is one (write:repository, a real bot account), and secrets are a per-repo store any
PR-added workflow can reference, so that route is narrowed rather than closed; tightening this check
from "non-null creator" to an allow-list of approved reviewers is what would close it
(ersatztv#742). A collaborator's own personal token still can, and no repo-side change closes that.
Note also that re-derivation is not a race the attacker can lose: it fires only on the trigger's
types, and posting a status is not one of them, so a POST timed after the last PR event stands
until the next one.
Deciding either exemption requires the PR's complete changed-file list, which the workflow does
not compute itself: it calls scripts/pr-changed-files.sh, the single shared implementation also
used by the advisory hook .claude/hooks/pretooluse-merge-consent.sh (ersatztv#649). The workflow
reads that script's exit status — a non-zero exit means "could not tell" and withholds the
exemption; its stdout is meaningless on any failure path and is never consumed.
Never write a classification guard as producer | grep -q… here. Under set -o pipefail, grep -q
exits at its first match, the producer takes SIGPIPE (141), and a MATCH is reported as a failed
pipeline — inverting the guard for any PR whose path list exceeds the pipe buffer. That let a large PR
be classified docs-only, and let one editing .gitea/ skip the protected-path check entirely. A
here-string is also wrong (bash spills a large one to temp storage, which fails the same way when
temp is full). Count instead — grep -c drains stdin over an ordinary pipe — evaluate the counts
once at top level rather than inline in an if, and fail closed on a non-numeric result. Full detail:
ci.grep-q-pipefail-inversion.
That script takes the expected base branch as a required 5th argument and refuses to enumerate when
the PR's live base does not match it, checked both before and after paging (ersatztv#698).
/pulls/{n}/files diffs against the PR's live base, so retargeting changes the answer without moving
the head sha — a PR opened into main and retargeted mid-run was granted a docs-only exemption while
its diff against main carried a C# file. The workflow passes the base from the pull_request_target
payload, which a retarget cannot rewrite, and edited is in types: so a retarget reclassifies.
edited gives detection, not atomicity: runs are not serialized, so a stale run could still post
success after the reclassifying run posted pending.
That residual is now fenced (ersatztv#706). Runs are still not serialized — instead a run that was
overtaken declines to write. The job counts change_target_branch events on the PR's issue timeline
at start and again immediately before its POST, and posts nothing if the count moved. The count is
the key precisely because the branch name is ABA-vulnerable: main → scratch → main reads main at
both ends, which is how the forged exemption was obtained in the first place. Abstaining never strands
a PR, because every retarget fires edited — the event that makes one run abstain has already queued
its successor.
If the count can't be established (unreadable timeline, paging that never reached a validated empty
page), only the exemption success is withheld; pending still posts, since pending cannot turn an
unreviewed head green and withholding it would strand ordinary PRs for nothing. If an exempt PR is
unexpectedly missing its status after a retarget, this is why — the job log names the counts.
Worth knowing before reaching for the obvious alternative: a concurrency group does not work here,
measured rather than assumed. Gitea 1.25.4 auto-cancels superseded push runs on a branch, but not
pull_request_target runs — two runs for one PR genuinely overlap, and adding
concurrency: {…, cancel-in-progress: false} changed nothing (probe runs still overlapped by 36s).
cancel-in-progress: true is deliberately untried, because a cancelled run leaves an exempt PR
statusless with nothing left to re-trigger it. Full measurements and the two surviving residuals:
ci.verdict-write-retarget-fence.
Separately, after posting an exemption success the job re-reads the per-POST status history and, if
a human Review-verdict: row appeared during the write window, overwrites its own status with
pending and logs an error — so a human BLOCKED can never be silently turned green. The repair is
pending, never a copy of the human's verdict, which would attribute a human decision to the job.
Three properties of this workflow are security-relevant and are structurally asserted by tests in
scripts/tests/test_pr_changed_files.py — those tests pin the workflow's shape, which is not the same
as establishing that the gate cannot be forged (see the residual below, and ersatztv#697/#698):
- The trigger is
pull_request_target, scoped tobranches: [main]— never plainpull_request(ersatztv#672). Gitea resolves apull_requestworkflow definition from the PR's own head, so under that trigger a PR editingreview-verdict.ymlran its own rewritten copy and could postreview-verdict/h10=successfor itself. The base-ref checkout below binds the scripts this job runs; only the trigger binds the definition. Thebranchesfilter is half the fix, not a refinement of it: base resolution means the base branch supplies the gate, so an unfiltered trigger merely moves the rewrite to an attacker-pushed base — and a status forged there is inherited by any later PR carrying the same head sha (ersatztv#663).pull_request_targetis safe here only because this job never checks out or executes head-supplied code. Verified on this instance with four scratch PRs rather than inferred from GitHub; full rationale indocs/decisions/records/ci/gate-trigger-base-resolved.md. This closes the rewrite route through this workflow, not the class:docker-build.ymlis also head-resolved and must stay onpull_requestbecause it builds the PR's code, so it got the read-only status identity instead — itsETV_STATUS_AUTHis now a PAT scopedwrite:package+read:repository, which the status endpoint refuses (ci.actions-credential-scoping, ersatztv#697). The inventory was never that one workflow, though: Gitea injects a write-capableGITEA_TOKENinto every job and branch protection binds the context, not its issuer. Gitea >=1.26 with the Actions default set to Restricted (server-management#714) binds the injected token, but does not close the class either — not against a personal token, and not againstRENOVATE_TOKEN(ersatztv#742). And none of it was necessary: direct pushes tomainwere server-side permitted, so the gate could be skipped without any forgery (ersatztv#743). That is now CLOSED —maincarriesenable_push: falseandblock_admin_merge_override: true, so it is reachable only through the PR merge path, the one path on which Gitea evaluatesstatus_check_contexts, and an admin cannotforce_mergepast them (release.main-direct-push-disabled— neither field is citable alone). Note the fix is disabling push, not whitelisting it: a push whitelist namingtimothywas measured to still admit the push, andtimothyis the identity every session, PAT and injectedGITEA_TOKENalready acts as, so the whitelist form would have closed nothing. The block binds a site admin at pre-receive but not a credential that can first PATCH branch protection off — an accepted residual, recorded in that decision. The exemption path has separate defects of its own (ersatztv#698). One operational consequence of the trigger change: a PR whose base is notmainnow gets noreview-verdict/h10at all. That is fail-closed.editedis now among the trigger'stypes(ersatztv#698), so a PR retargeted ontomainreclassifies instead of staying statusless until its next push — but note that only gives detection: runs are not serialized, so a stale run can still postsuccessafter the reclassifying run postspending(ersatztv#706). - The checkout takes the PR's BASE ref,
ref: ${{ github.event.pull_request.base.sha }}withpersist-credentials: false— never the head. This job judges the PR, so the PR must not supply the code that judges it; a head checkout would let a PR edit the enumeration to return an empty list and exempt itself. scripts/jq-preflight.shruns in floor-only mode, never--expect. This job writes a branch-protection-required status, so an exact version pin would turn any jq upgrade on the runner into a repo-wide merge deadlock.
A PR whose base predates ersatztv#658 has no such script on its base ref; that case posts pending
with the reason rather than dying with no status at all.
⚠️ Changing review-verdict.yml itself: it is not exercised by its own PR. Base resolution cuts
both ways — the PR editing this workflow runs the version already on main, so an edit goes live
only on merge, repo-wide, having never run. A broken edit merges green and then breaks the gate
for every subsequent PR, and the PR that would repair it is gated by the same broken workflow. Do not
trust the editing PR's own checks. Verify the way ersatztv#672 did:
- Push a scratch base branch carrying the candidate workflow.
- Open a throwaway PR from a scratch head into that base, so the candidate is the definition that
runs. Have it post a probe-named context (e.g.
review-verdict/h10-PROBE), never the realreview-verdict/h10— a probe must not be able to forge the gate it is testing. - Read the resulting commit statuses to see which definition actually ran, then delete both branches.
The same shape is what makes a branches:/types: change verifiable at all, since neither can be
observed from the editing PR. Note step 2 requires the scratch base's own branches: filter to
name that base — the definition comes from the base, so a base the filter does not admit produces no
run at all.
⚠️ Never write an expression delimiter inside a run: body here — a comment is NOT inert
(ersatztv#751, ci.workflow-run-body-no-expressions). A run: body is not shell when the runner
reads it. The runner scans the whole scalar for the expression opener and, on finding one, rewrites
the entire body into a single format(...) call so the result can be spliced back in. That
rewrite is all-or-nothing: a payload that does not evaluate fails the interpolation of the whole
scalar, and the runner then drops the step and concludes the job success.
That is not hypothetical. From 8f6d4f443 (2026-08-03) to 2026-08-06 the classify step never ran.
The #706 note above, explaining why a concurrency group does not work here, quoted a concurrency:
snippet containing a PR-number expression as an illustration, in a shell comment. pr number is not
a valid expression. So review-verdict/h10 was posted by nothing but a human hand for three days,
both exemption classes silently stopped working, and every run reported success. The prose documenting
a fix disabled the fix.
The silent green is the real defect. An absent required status reads as "not reviewed yet", which is indistinguishable from the correct pending state — so an ordinary PR looked ordinary while the gate was dead, and the cost landed only where no human was in the loop. PR #739 (docs-only) merged 2026-08-05 with zero commit statuses on its head, and got in only because admin force-merge was still enabled; ersatztv#743 removed that escape the next day, so a docs-only or Renovate-manifest PR arriving after that would simply have been stuck with no bypass. The two Renovate PRs in the window escaped by timing, merging minutes before the bad commit.
Three things now hold the line, and they are deliberately different in kind:
- The prose names expressions instead of quoting them — write "a
github.event.pull_request.numberexpression", not the delimiters. Pass values in through the step'senv:block, which is interpolated per value, so a bad payload there cannot take the body with it. - A start-marker guard turns a dropped step RED. The classifier writes a marker as its first act
and an
if: always()step fails the job when it is missing. It asserts execution started, never that it completed — the classifier has several legitimateexit 0abstention paths. The guard's own body must stay expression-free, or the mechanism it guards against can delete the guard too, and that absence would be silent as well. - Two static guards, in
scripts/tests/test_pr_changed_files.py: no delimiter in anyrun:body of this file (absolute — a dropped step here is a dead merge gate, and its bodies are ~700 lines of prose), and repo-wide, every expression payload's head token must name a context or function the runner can resolve (permissive, because the other workflows interpolate intorun:legitimately — 5 occurrences today, inci-image.yml,docker-build.yml'sapi-docs/formatandpr-checks.yml's two git-diff gates; #756 removedbuild's two and banned that job as well, so the ban now coverstest,migrationsandbuild). Be precise about the second one's reach: it catches the historical defect (pr number) and a nonexistent context, but not a syntactically invalid payload whose tokens are all known (${{ github.ref == }}passes), nor a renamed output (steps.metadata.outputs.shortshapasses — every token after the first is preceded by.and is skipped), nor an unclosed opener. Catching those needs an expression parser. An earlier draft of this section claimed it caught "a payload that cannot evaluate, wherever it sits"; that was false, and the corrected claim is the one to rely on.
Worth knowing why nothing caught this for three days: every other workflow-shape test in that file
reads _code_lines(), which strips comments. That is correct for what it was for, but it encodes the
assumption this bug falsifies. The strict test reads the raw scalar, and must never adopt
_code_lines.
⚠️ A page past the end of /issues/{n}/timeline is JSON null, not [] — and this instance is
not consistent between endpoints (/issues/{n}/comments returns [] when empty). The retarget
fence's count_retargets gated on type == "array", so it read the real terminator as unreadable:
the walk never reached a validated empty page, rt_ok was never yes for any PR, and the fence
therefore withheld every exemption success. Renovate and docs-only PRs got no status at all —
the same user-visible outcome as the dropped step above, by a completely unrelated route. So fixing
the interpolation alone would not have restored the exemptions.
Two things kept it invisible, and both are worth generalising:
- It shipped in the same commit (
8f6d4f443) that stopped the step executing, so the fence had never once run in production. A guard's first real execution is not the same event as its merge. - The test double asserted the wrong shape while claiming measured fidelity. Its comment read
"Real shapes, measured on this instance and deliberately mirrored" and it printed
[]for a page past the end. Every fence test was green against a response the server never produces, so thearray-only gate was never exercised by the suite either. With the double corrected and the old gate restored, most of the fence suite fails — 18 tests when first measured atc710db4a1, 21 once three more fence-dependent tests existed. The invariant is the point, not the count: they had all been passing for the wrong reason. (Given as a range on purpose — an earlier draft cited a bare "18", which was stale two commits later, inside a section about stale claims.) When a double claims fidelity, that claim is a test assertion and needs re-measuring like any other.
The type is now read as a value (case over jq -r 'type') rather than through jq -e, whose
exit-status semantics already bit this workflow once at jq 1.6, and both null and [] terminate the
walk. The regression test is parameterised over both shapes because both are live on this server.
null is accepted as exhaustion only from page 2 on — every real PR's first page carries events
(spot-checked non-empty across #752/#753/#749/#739/#717; the counts are deliberately not recorded here
because timelines grow and an earlier draft's five figures were stale within days), so a null first
page is anomalous rather
than empty, and the walk should not certify "no retarget happened" from a response it cannot explain.
The same nil-slice shape bites /commits/{sha}/status — a third instance, found by cold review of
the fix for the second. A head with no statuses yet returns
{"state":"pending","total_count":0,"statuses":null} (measured on PR #739's head). read_existing_verdict
gated on .statuses | type == "array", so it hit its exit 1 and posted nothing at all — fail-closed,
same user-visible outcome. null is now accepted there only when total_count is 0, so a body that
merely lost its array is still refused and an existing verdict is still protected from a transient
error. scripts/pr-changed-files.sh was swept and is unaffected (pulls/{n}/files returns []).
The generalisable rule: a nil Go slice serialises to null, so every list-shaped field on this API
is suspect and only a per-endpoint measurement settles it.
Establishing that "no verdict exists" needs a second page, and both arithmetic guards for it are
no-ops here. read_existing_verdict concluding absence is what licenses posting an exemption over a
verdict the job cannot see, so that conclusion has to be earned. Two obvious checks were tried and both
proved empty:
.statuses | lengthvs.total_count—total_countis the count for the page returned, not for the commit. Measured at 1.27.1 on3aed43c6(6 contexts):?limit=1returnslen=1, total_count=1,?limit=3returnslen=3, total_count=3. Equal by construction, so the check reads as a completeness proof while proving nothing.- "refuse when the page comes back full at the requested
limit=100" — this instance capslimitat the server-wideMAX_RESPONSE_ITEMS, measured at 50 (/issues?limit=100returns 50). A response can therefore never carry 100 rows, and the comparison was dead code. The repo already documented that cap inscripts/pr-changed-files.sh, two test files andci.script-tests-job; the guard was written against 100 anyway, and a cold review caught it. Hardcoding 50 instead would re-break the day the setting changes.
So the job asks the server, and only when it matters: if the review-verdict/h10 row is on page 1
there is nothing further to learn (this endpoint returns the latest status per context, and a context
cannot recur on a later page). When the row is absent it reads page 2 — any rows there mean the list
runs longer than one page and a verdict could be beyond it, so it refuses instead of concluding absence.
Cap-independent by construction. Paging is real here: measured ?limit=3&page=2 returning three further
rows, and page=9 returning the same statuses: null terminator.
The total_count zero-check also requires the JSON type to be a number: jq -r renders 0 and
"0" identically, so a text compare would accept a schema-corrupted "total_count": "0" as "no
statuses".
The repo-wide expression guard scans PARSED scalars, not raw file text. A delimiter in an ordinary
top-level YAML comment is inert — the runner never evaluates it — so redding on it is a false positive,
and this file has now produced that false red twice. PyYAML drops those comments. A run: body is
itself a scalar and keeps its shell comments, which is the point: inside a run: scalar a comment is
not inert. Verified both directions by mutation — an inert top-level comment passes; the same payload
in a run-body comment still reds.
CLAUDE.md and AGENTS.md are now PROTECTED paths. DOCS_ONLY matched them, so the documents
that define the completion protocol, the merge-consent convention and the H10 rule were themselves
docs-only-exemptible while .claude/ was protected — the same self-exemption the gate rules out, one
directory over. Driving the real classify body with a lone CLAUDE.md change produced
review-verdict/h10=success. It is fixed here rather than deferred because restoring the exemptions is
what makes it reachable: no exemption success was writable at all while the classify step was
dropped. README.md is deliberately not listed — ordinary prose, no enforcement. For the same reason,
#706's known residual returns with the working fence: while rt_ok was never yes, route 1 was closed
by accident.
That gap is now closed — docker-build.yml's test and migrations jobs are also required
contexts, and there a dropped step is fail-OPEN: the required check goes green having done no work,
which is strictly worse than an absent status (compare #684). ersatztv#756 gave those two jobs
per-step execution markers and extended the delimiter ban to them; see
"Dropped-step guard on the required jobs" above.
It lives in its own workflow file on purpose: pr-checks.yml sets cancel-in-progress: true,
and a cancelled run there would leave an exempt PR with no status and no further push to
re-trigger it. Its own job context (Review verdict / Set review-verdict status) is not the
required check — a workflow must not satisfy the gate merely by running successfully.
Full rationale: docs/decisions/records/release/verdict-status-check.md and
docs/decisions/records/ci/shared-pr-file-enumeration.md.
CI toolchain image (docker/ci/Dockerfile, .gitea/workflows/ci-image.yml)
The jobs that need a toolchain — test, migrations, functional-e2e, api-docs, format —
run inside a shared image via container: instead of installing their toolchain per run
(ersatztv#390). They therefore carry no setup-dotnet, no setup-node, no apt-get,
and no dotnet tool install.
What it ships: .NET 10 SDK, Node 22, prod-identical ffmpeg/ffprobe, git/python3/jq/zstd,
the dotnet-ef + dotnet-reportgenerator-globaltool global tools (which the migrations and
test jobs used to install on every run — bump those versions in the Dockerfile, not the
workflow), and headless Chromium for the UI-E2E flows (below). Project dependencies (NuGet/npm)
are deliberately not baked in — they change per commit and stay on actions/cache
(~/.nuget/packages, ~/.npm).
Headless Chromium for UI-E2E (ersatztv#445). PLAYWRIGHT_BROWSERS_PATH=/ms-playwright holds
chromium-headless-shell, installed with --with-deps at image build time so the functional-e2e
job installs no browser per run. Measured on this exact base: the headless shell is 267M where
full chromium is 656M, and chromium.launch() resolves to the shell anyway because
web/playwright.config.ts never asks for headed — the accepted tradeoff being that a headed run
inside this image would fail. Also verified on the real base rather than assumed: Chromium launches
as root inside a container with no --no-sandbox/chromiumSandbox:false opt-out, so the config
carries no sandbox workaround. The Dockerfile's build-time smoke test actually launches the browser,
so a missing system library fails the image build rather than a CI run.
⚠️ ARG PLAYWRIGHT_VERSION must equal web/package.json's @playwright/test pin, which is
deliberately EXACT (no caret): Playwright ties a browser revision to the package version, so a
mismatch leaves no usable browser. Renovate bumps the npm pin but cannot know about this ARG — when it
does, bump the ARG, let ci-image.yml publish the new :<sha>, then update all five container pins.
scripts/e2e-ui.sh guards the drift by launching a browser up front and failing with exactly that
instruction (it probes by launch, not by path, because chromium.executablePath() reports the
full-chromium path that a headless-shell-only image deliberately lacks).
How it's layered: FROM ersatztv-ffmpeg:8.1.2 + COPY --from=mcr.microsoft.com/dotnet/sdk:10.0-noble-amd64 /usr/share/dotnet — the same pattern docker/Dockerfile uses for the prod image. Our ffmpeg base is
ghcr.io/linuxserver/baseimage-ubuntu:noble, the same Ubuntu release as the SDK image, so the copied
SDK matches the base's glibc/ICU. Keep the ffmpeg tag on that FROM equal to the one
docker/Dockerfile pins, so CI's ffmpeg stays prod-identical — that fidelity is what the
ersatztv#299 seeded-media/scanner E2E follow-ups will need.
Bumping the toolchain is a deliberate two-step. The jobs pin an immutable :<sha>, never
:latest, so a bad toolchain push cannot break every job at once:
- Merge a
docker/ci/Dockerfilechange.ci-image.ymlpublishesersatztv-ci:<sha>(+:latestfrommainonly — a human pointer; jobs must never consume it). - In a follow-up PR, update the pin in
docker-build.yml— all five jobs together. That PR's own CI is what proves the new image works. The pin is repeated per job becausejobs.<id>.container.imagecannot read the workflowenvcontext.
The tag is exactly 7 hex chars — get the length right, not just the commit (ersatztv#594).
ci-image.yml tags with git rev-parse --short HEAD under fetch-depth: 1, and that shallow clone
holds few enough objects that git always abbreviates to 7. A full local clone abbreviates to 8,
so the natural command prints one character too many:
git rev-parse --short HEAD # 8 chars in a full clone — WRONG, no such registry tag
git rev-parse --short=7 HEAD # 7 chars — what ci-image.yml publishes. Use this.
An 8-char pin names the right commit but no existing image: it satisfies a resolve-and-compare
check, then every container: job dies at image-pull with manifest unknown, which reads like a
registry outage rather than a one-character pin error. ci-image-pin therefore checks the pin's
length as an invariant separate from its correctness, and prints the exact tag to use.
Caveat worth knowing before you trust the 7:
ci-image.ymlstill tags with a plain--short, whose length git auto-scales to the object count. 7 is therefore an empirical property of today's shallow clone, not an enforced invariant — if that count ever crosses git's threshold, the publisher emits 8, the correct pin becomes 8, and the gate's hardcoded 7 goes permanently red demanding a tag with no image behind it. Making the publisher emit--short=7is tracked as ersatztv#597.
ci-image.yml triggers on pushes touching docker/ci/**, workflow_dispatch, and a weekly Monday
05:00 UTC cron (base-image security updates; Gitea registers schedule only from main). It runs on
ubuntu-latest — it was on small until server-management#639, where "docker-only" was found to be
a poor proxy for "small": this is a full buildx of the .NET toolchain image, the heaviest job in that
lane. Like docker-build.yml, it needs BuildKit's inline http = true for the HTTP
registry. Renovate tracks the Dockerfile's image pins (dockerfile manager, see renovate.json).
Three container-specific gotchas — worth knowing if you add a job or a step:
sh, not bash, is the default shell inside acontainer:. act_runner runs steps assh -e {0}(dash) because it can't assume bash exists in an arbitrary image — even though ours has it. Every bashism (set -o pipefail, arrays,shopt,mapfile) then dies instantly withset: Illegal option -o pipefail.docker-build.ymltherefore declares a workflow-leveldefaults: run: shell: bash. If you add a workflow with containerized jobs, do the same — outside a container the shell defaults to bash, so this failure only appears once you containerize and it looks nothing like a shell problem (it surfaced as themigrationsjob dying in 0.13s).actions/checkoutclones as root into a mounted workspace, which trips git's "detected dubious ownership" guard and breaks everygitcall in a step. Fixed in the Dockerfile withgit config --global --add safe.directory '*'.- (defensive, not load-bearing) The ffmpeg base sets
ENTRYPOINT ["ffmpeg"]because it ships as an ffmpeg CLI, so the Dockerfile resetsENTRYPOINT/CMD. act overrides the entrypoint anyway (entrypoint=["/bin/sleep" "10800"]), so this is belt-and-braces for anyone running the image by hand — unlike the two above, which are real.
⚠️ A REBASE invalidates the pin. The pin must equal the short sha of the commit that touched
docker/ci/**, and a rebase rewrites that commit's sha — so ci-image-pin goes red on a branch
that was green before, with a pin that still resolves to a real (now-orphaned) commit and an image
that still exists in the registry. Worse, the force-push usually does not rebuild: ci-image.yml
filters on paths: docker/ci/**, and a rebase that doesn't change the Dockerfile's content
produces no diff for that path, so nothing republishes. And you cannot simply re-dispatch it —
ci-image.yml tags git rev-parse --short HEAD, i.e. whatever the branch HEAD is when it runs, not
the commit that touched docker/ci. Those two coincide only when the docker/ci commit is HEAD.
Recovery (ersatztv#445 hit this): make the docker/ci commit be HEAD again — push a commit that
really does change docker/ci/**, let ci-image.yml publish :<its short sha>, then bump the pin in
a follow-up commit. That is the same two-step below, just re-run after the rebase. The cheapest way to
avoid it entirely is to land a toolchain-image change on its own, before the work that consumes
it, so the consuming branch never carries the docker/ci commit through a rebase.
Bumping the pin is enforced, not remembered. The ci-image-pin job (blocking, PR-only; defined
in pr-checks.yml, but it greps docker-build.yml where the pins live) fails if
docker-build.yml's pin isn't the short sha of the last commit to touch docker/ci/** or
ci-image.yml, if that pin isn't exactly 7 chars long (see above), or if the five jobs ever pin
different tags. This exists because Renovate manages
docker/ci/Dockerfile's base pins but cannot bump an opaque :<sha> in container.image — so a
Renovate base bump would otherwise publish a new image, test the old one, and merge with the
Dockerfile disagreeing with the pin. A red ci-image-pin means: let ci-image.yml publish the new
:<sha>, then update all five pins to it.
What it is and isn't worth. Measured honestly (ersatztv#390): the image saves ~15–40s per job
(setup-dotnet is 8–19s, setup-node 2–5s cached, the two tool installs ~9s) plus the 110s
apt-ffmpeg step — roughly 3–8% of runtime. It is not where CI time goes; see the lane table above
(queue wait, server-management#604) and ersatztv#398 (742s of redundant compilation). Its durable
value is prod-identical ffmpeg, a pinned/consistent toolchain, and making jobs runner-agnostic — the
last is what allowed the lane rebalance.
Dockerfile notes (docker/Dockerfile)
- Base image:
192.168.1.95:3000/timothy/ersatztv-ffmpeg:8.1.2(our Gitea fork of the archivedghcr.io/ersatztv/ersatztv-ffmpeg). FFmpeg 8 base image work landed in ersatztv-ffmpeg#4; app-side compatibility work landed in ersatztv#9. - Copies
Directory.Build.props,Directory.Build.targets,Directory.Packages.props,global.json,.editorconfigbeforedotnet restoreso the image build uses the same MSBuild config, central package versions, SDK pin, and analyzer severities as local/CI builds (it previously copied only*.sln).Directory.Packages.propsis required here: under Central Package Management the csproj carry no inline versions, so the image's restore fails (NU1015) without the central manifest. - amd64-only (the runner/build host is x86_64). No arm32/arm64, no DMG/exe artifacts, no GHCR/DockerHub.
- openapi-generator jar layer ordering (ersatztv#190): the
wgetfor the openapi-generator-cli jar runs before theCOPYofErsatzTV/wwwroot/openapi/, so the ~30MB download layer is cached independently of the openapi spec. Previously the jar was downloaded after thatCOPY, so any PR touching the spec (e.g.v1.json) busted the download layer too and re-fetched the jar on every such change. Codegen itself still runs after the specCOPY, since it needs both the jar and the spec files.
Dependency management (Central Package Management + scans)
Central Package Management (CPM) — package versions live in a single repo-root
Directory.Packages.props (ManagePackageVersionsCentrally=true); the per-project
csproj reference packages by name only (no Version=). One source of truth, atomic
one-line bumps, and cross-project version drift is structurally impossible. To add or
change a dependency, edit the <PackageVersion> entry centrally — never put a Version=
back on a <PackageReference> (that trips NU1008). The Docker build must copy this file
before restore (see Dockerfile notes). The .mcp/ vendored tool (gitignored, not in the
solution) keeps inline versions via a local-only .mcp/Directory.Packages.props
opt-out (ManagePackageVersionsCentrally=false). (ersatztv#14)
NuGet audit — .NET 10 runs NuGet audit on restore. Several projects set
TreatWarningsAsErrors=true, so vulnerable transitive packages failed the build.
Directory.Build.props demotes low/moderate/high advisories (NU1901-1903) to warnings
and promotes NU1904 (critical) to an error in every project via WarningsAsErrors.
The advisories that prompted this were resolved in ersatztv#8 (NCalcSync→6.x; SQLitePCLRaw
bundle 3.x) and ersatztv#314 (Microsoft.OpenApi 2.0.0→2.7.5, GHSA-v5pm-xwqc-g5wc High —
direct-pinned in ErsatzTV.csproj over the 2.0.0 that Microsoft.AspNetCore.OpenApi +
Scalar.AspNetCore pull transitively; the SQLitePCLRaw override pattern; regenerates the
OpenAPI doc byte-identically). The NU1901-1903 demotion is kept by design: criticals (NU1904)
still hard-block, while low/moderate/high advisories surface as warnings + via the weekly scan and
Renovate security PRs, rather than breaking unrelated PRs the moment a new transitive
advisory drops.
Scheduled vulnerability scan — .gitea/workflows/dependency-scan.yml runs weekly
(cron 0 6 * * 1) + on workflow_dispatch: dotnet list package --vulnerable --include-transitive over the full solution (incl. Scanner, which the image build
strips). dotnet list exits 0 even with findings, so the step (bash -euo pipefail)
greps for the "has the following vulnerable packages" marker and fails the run if present.
Detection only — it surfaces advisories on a schedule, a Gitea-native stand-in for
Dependabot; it does not open update PRs (that's Renovate — server-management#484).
Gitea registers schedule triggers only from the default branch, so the cron starts
after merge to main; use workflow_dispatch to run on demand. It went green once
ersatztv#8 cleared the NCalcSync/SQLitePCLRaw advisories — a red run now means a new
advisory has appeared. (ersatztv#14, ersatztv#8)
Renovate (automated update PRs) — .gitea/workflows/renovate.yml runs self-hosted
Renovate weekly (cron 0 3 * * 1) + on workflow_dispatch,
as a renovate/renovate:43 container job on the shared act_runner. This is the proposing
layer the scan above deliberately omits: it opens grouped dependency-update PRs and
OSV-driven vulnerability-fix PRs against main, and maintains a Dependency Dashboard
issue listing the full backlog. Config is the repo-root renovate.json — managers nuget
(via CPM), github-actions, and dockerfile (scoped to the built docker/Dockerfile; it reads the
HTTP-only Gitea registry for the ersatztv-ffmpeg base via a RENOVATE_HOST_RULES host rule —
insecureRegistry + registry read creds, set in the workflow env, not the committed config). The
docker-compose manager is unused (repo compose files are build:-only). Auth: a dedicated
renovate Gitea bot (Write
collaborator) via repo Actions secrets RENOVATE_TOKEN (bot PAT) + GH_COM_TOKEN (no-scope
github.com PAT for changelogs — named GH_, not GITHUB_, a prefix Gitea reserves).
Patch bumps to test/dev-only packages (NUnit*, NSubstitute, Shouldly, coverlet,
Microsoft.NET.Test.Sdk, Testably.Abstractions*, threading analyzer) auto-merge once
the Build & test (.NET) check passes — branch protection on main requires that context;
everything else is manual review (ersatztv is prod-bearing). Range-pinned packages (e.g. EF
Core [9.0.x,10)) are respected — no v10 jump. PR volume is throttled (prConcurrentLimit
5 + config:recommended's prHourlyLimit 2); tick a dashboard checkbox or raise the limits
to drain faster. workflow_dispatch defaults to a safe dry run. Cross-repo rollout
tracked in server-management#484. (server-management#484)
Security scanning — black-box DAST + SAST (scripts/security-scan.sh, ersatztv#314)
Every other security check we run is in-ecosystem / white-box — SonarAnalyzer, NetArchTest, the
adversarial fork + Codex review passes, the api-docs/format/decisions CI gates, dotnet list package --vulnerable — so they share our blind spots. scripts/security-scan.sh is the out-of-ecosystem,
black-box complement and a #197 exit criterion (HARD GATE before remote exposure): it drives the
running product from outside our C#/review stack.
- What it does. Boots a throwaway container from the image under test (fresh empty config volume;
never the deployed prod/test container — the authenticated active scan sends attack payloads to write
endpoints), reads the generated machine key, and runs an authenticated OWASP ZAP API scan
(
zap-api-scan.py) that imports the static/openapi/v1.jsonso it exercises every declared/api/v1operation, injectingX-Api-Keyon every request via a ZAP replacer rule so it reaches the[RequiresAuthentication]+RequireKeyForReadssurface (not just the/appshell an unauthenticated spider sees). Then a semgrep SAST cross-check (p/security-audit+p/secrets+p/csharp). The container is torn down on exit. - Where/when. Runs on the docker host (jazz — the Mac has no docker), like
migration-smoke.sh:scripts/security-scan.sh [IMAGE] [PORT](defaults…:latest/8411). It is a manual release-gate, deliberately not a per-PR CI job — it needs docker + a booted image, takes several minutes, and is noisy (expect to tune, not take raw). The continuous layer is the per-PR white-box gates + the weeklydependency-scan; this is the per-release black-box pass. Re-run it each release and before any change to the exposure posture. - Exit-code contract (ersatztv#338).
zap-api-scan.py's raw exit code is NOT a simple pass/fail — it conflates a clean run with a warnings-only run unless you know its wrapper contract: 0 clean (no FAIL or WARN alerts), 2 WARN-only (triage required, but not release-blocking), 1 FAIL (at least one FAIL-level alert — release-blocking), 124 the script's owntimeoutwrapper killed a hung post-scan cleanup (the report written before the hang is still usable — triage it), any other code means the scanner/tool itself errored (not a scan result at all).scripts/security-scan.shencodes this inclassify_zap_exit()and prints an unambiguous==> ZAP result: <PASS|WARN|FAIL| TIMEOUT|TOOL ERROR> ...line; the script's own exit status reflects that classification (0 for clean/WARN, 1 for FAIL/timeout/tool-error) rather than ZAP's raw code, so a warnings-only run no longer reads as a failed scan. Found when the v26.8.0 release scan (#335) returned raw exit 2 for a report withFAIL-NEW: 0and two known/expected warning classes — the shell result looked like a failure though the release gate had actually passed. Runscripts/security-scan.sh --selftestfor a docker-free regression check of the classification logic. - Triage. Triage each WARN/FAIL finding false-positive vs real. Real, in-scope, go-live-blocking findings get fixed (e.g. the security headers from the #319 baseline; the Microsoft.OpenApi pin above); LAN-expected noise (Private-IP disclosure) is revisited only for genuine remote exposure. nuclei (template-based CVE fingerprinting) is an optional third pass — deferred while its template fetch is blocked in the runner env (pre-seed a template volume to add it); ZAP covers the DAST baseline and semgrep the SAST, so it is not on the critical path.
Static analysis & formatting
Analyzers — Directory.Build.props enables the SDK analyzers at latest-All and turns on
Microsoft.VisualStudio.Threading.Analyzers for every centrally managed project.
Directory.Build.targets also references
Roslynator, SonarAnalyzer.CSharp, Meziantou.Analyzer, and AsyncFixer repo-wide (versions
central via CPM). All analyzer package references are guarded on ManagePackageVersionsCentrally, so the
gitignored .mcp tool—which deliberately uses inline package versions—does not inherit versionless
references.
They are introduced incrementally (ersatztv#15). eng/analyzers/sdk-all-suggestion.globalconfig
enumerates the .NET 10 SDK All inventory at suggestion; this exact-ID baseline is necessary because
the SDK's generated latest-All severities outrank .editorconfig bulk settings. .editorconfig keeps
the threading and curated-pack baselines at suggestion. Diagnostics remain visible to IDEs and
dotnet format analyzers, but do not create a wall of failures (a direct latest-All trial activated
455 existing errors in the TWAE projects).
Promotion is the enforcement — set a reviewed rule to warning in .editorconfig and append its ID
to the central WarningsAsErrors list in Directory.Build.props. The explicit list makes the rule block
in every project, including test projects that do not otherwise use TWAE. On a major SDK upgrade,
regenerate the checked-in SDK baseline from analysislevel_<major>_all.globalconfig, preserve SDK none
entries, and review newly introduced rules before accepting the snapshot.
Promoted rules are recorded here so the blocking subset stays intentional and reviewable:
-
Sonar
S3981—warning+WarningsAsErrors(ersatztv#15): rejects collection-count comparisons that are constant regardless of collection size. Its first finding exposedWorkers.Count >= 0, which permanently classified scheduled memory releases as busy and skipped the intended aggressive idle collection. -
StyleCop.Analyzers is intentionally excluded: its latest stable (1.1.118) crashes (
AD0001) on C#recorddeclarations, and its rules overlap the existing.editorconfig/Roslynator. Revisit via the record-compatible1.2.0-betaonly if specifically wanted. -
The former Blazor
.razorcaveat is retired: Blazor removal deleted the Razor sources and their temporary SonarNoWarnlist. The.razor/.cshtmlsuggestion scopes remain in.editorconfigonly as a defensive default if server-rendered view code is ever reintroduced.
Formatting — the inherited tree still contains legacy UTF-8 BOM/whitespace debt, so the standing
policy is format as you touch, not a mass rewrite (ersatztv#311). The Husky pre-commit hook and
the blocking format CI job both run dotnet format whitespace . --folder --verify-no-changes --include <changed .cs> — scoped to the files the commit/PR touches. Untouched legacy files remain
outside the gate; .gitattributes pins line endings. A one-time full-tree normalization remains a
separate, unmade decision.
Why whitespace . --folder, not the full dotnet format <sln> (ersatztv#469) — the gate only
needs to enforce .editorconfig whitespace (indent/EOL/trailing/final-newline) and charset
(no UTF-8 BOM). The old recipe (dotnet format ErsatzTV.sln --no-restore --verify-no-changes --include) loaded the entire ~10-project MSBuild workspace and built a Roslyn compilation per
project before checking a single file — --include narrows which files are checked, never what
gets loaded. Measured whole-solution dotnet format ran ~480s locally; folder mode runs in
~0.5s and needs no dotnet restore (the NuGet-cache + Restore steps were removed from the job).
--folder treats the tree as a plain folder of files, skipping MSBuild/Roslyn entirely, and still
reads .editorconfig. Verified non-vacuous: it exits non-zero on an injected trailing-whitespace
line (error WHITESPACE) and on a prepended UTF-8 BOM (error CHARSET), and exits 0 on a clean file.
No coverage was lost: the full dotnet format gate did not enforce the style/analyzer pass
either — a probe injecting a warning-severity naming violation (local_constants not ALL_UPPER)
passed the full solution format (exit 0): the only .editorconfig rule above :suggestion/:none
severity is that one naming rule, and naming violations have no dotnet format batch code-fixer, so
--verify-no-changes reports no change regardless of severity. The analyzers that must block
(NU1904, S3981) are enforced at compile time via
WarningsAsErrors in Directory.Build.props, not by this job. Devs fix a violation with dotnet format whitespace . --folder --include <files> (the full dotnet format ErsatzTV.sln --include <files> is a superset and also works).
Migration integrity (EF Core, both providers)
TvContext (ErsatzTV.Infrastructure/Data/TvContext.cs) has two migration sets — one per
provider project: ErsatzTV.Infrastructure.Sqlite/Migrations and
ErsatzTV.Infrastructure.MySql/Migrations, each with its own TvContextModelSnapshot. A model
change needs a migration in BOTH. Add them with scripts/add-migration.sh <Name> (runs the EF CLI
for each provider). The EF CLI pattern (provider selected by the post--- arg, which Startup
reads as the provider config key):
dotnet ef <cmd> --context TvContext --startup-project ErsatzTV \
--project ErsatzTV.Infrastructure.{Sqlite|MySql} -- --provider {Sqlite|MySql}
The migrations job in docker-build.yml runs on every push/PR and, for each provider:
dotnet ef migrations has-pending-model-changes— fails if an entity changed without a matching migration (model drift), so a forgotten migration can't merge.dotnet ef database updateagainst a fresh empty DB — applies all migrations in order and fails on any broken/un-orderable one.
- SQLite (the prod provider) uses a throwaway file (
ETV_CONFIG_FOLDER=$(mktemp -d)); no service needed. Validated: 787 migrations → 139 tables. - MySql uses
ServerVersion.AutoDetect, which connects at config time, so the job needs a reachable server — provided by aservices: mysql:8.4container (the act_runner uses Docker execution with an auto-created per-job network — service reachable asmysql:3306(the old bumblebee runner pinned networkdownloadswarm; relocated in server-management#570)). Connection string viaMySql__ConnectionString(→ config keyMySql:ConnectionString). Validated: 305 migrations → 137 tables. It's an independent gate (not yet aneeds:of the image build) so the new MySql-service dependency can't block image builds until it's proven; promote it to a required check once stable.
Caveat — non-transactional operations: some migrations (e.g. SQLite PRAGMA foreign_keys) run
outside a transaction and warn at startup; they can't be rolled back mid-migration, so review such
migrations carefully (this is part of what motivated the apply-to-fresh check before the prod
cutover, server-management#481).
Resilience — the MySql apply is retried (concurrent-runner contention, not a model bug): both
runners (ci-runner VM 127 + bumblebee-runner) serve ubuntu-latest, and when two migration jobs
land on the same host at once (common when several PRs push together), each spins its own
mysql:8.4 service container and they starve each other — producing intermittent Command Timeout expired or mid-replay MySqlEndOfStreamException (dropped connection) on the MySql
apply-to-fresh-DB step. This is pure infra flakiness — has-pending-model-changes (the actual model
check) still passes, and the same commit passes on a quieter host. The job hardens against it two
ways: the connection string sets DefaultCommandTimeout=300 (up from MySqlConnector's 30s default),
and the apply is wrapped in a 3× retry that resumes from __EFMigrationsHistory (EF commits each
migration in its own transaction, so an interrupted one rolls back and the retry continues). A real
migration failure fails deterministically on every attempt, so the retry never masks it. If a run
still flakes past the retry, re-trigger (Gitea has no rerun API on this version — push, or the run
drains); don't treat a lone MySql-apply red as a code problem without checking the failure mode.
Migration-on-prod-copy smoke — release path (scripts/migration-smoke.sh, ersatztv#315)
The migrations job above only proves a migration is well-formed against a fresh, empty DB. It
can't prove it applies cleanly to the accumulated prod SQLite — real row volume, historical values,
and the post-migration data steps ErsatzTV runs on startup: DatabaseMigratorService (a
BackgroundService) applies pending migrations, then DbInitializer.Initialize + PopulatePathHashes
(an UPDATE over the real MediaFile table). A migration green on a fresh DB can still fail or corrupt
on prod, and today you'd only find out mid-deploy after the container recreates.
scripts/migration-smoke.sh rehearses it on a throwaway copy of the latest prod backup — it never
touches the live DB:
scripts/migration-smoke.sh --image <ref-about-to-be-promoted> [--db <backup.sqlite3>] [--timeout 180]
It copies the backup into a temp config dir, boots the new image against it (ETV_CONFIG_FOLDER), and
gates PASS on the Done applying database migrations log line — not merely on HTTP readiness, since
the migrator runs concurrently with Kestrel, so the web server can serve before/while migrations run.
FAIL = the container exits before finishing, a migration exception appears in the logs, migrations
don't finish within --timeout, or the app won't serve /iptv/channels.m3u afterwards. The smoke
container, the DB copy, and the temp dir are always torn down on exit (the ErsatzTV image runs as root,
so cleanup deletes its root-owned config files from inside a throwaway root container — otherwise each
run would leak the multi-hundred-MB copy). Exit 0 = clean, 1 = migration/boot failure, 2 = usage error.
--dbdefault: the newestersatztv.sqlite3under$ETV_BACKUP_DIR(default~/downloadswarm/ersatztv-backups— where the host-side pre-deploy backup hook writes timestamped snapshots). Pass--image= the version tag about to be promoted.- Where it runs: it's meant to run on the docker host as a Komodo pre-deploy step (which already produces the backup — see "Cutting a release" and the #553 pre-deploy backup caveat), so a bad migration aborts the promote before the live container recreates. Wiring it into that hook is a server-management concern (cross-repo — this repo owns the script + docs, server-management owns the Komodo hook). Until wired, run it by hand before cutting a migration-bearing release.
- Validated live 2026-07-12:
:latestagainst a copy of the 283 MB prod backup → migrations applied cleanly, app booted and served, temp dir removed.
Pre-commit hooks (web/)
The repo uses husky git hooks (installed via web/'s lint-staged + npm) to catch
lint/format/type/API-drift errors locally, before they reach CI. Because the git root and
the npm project dir differ (monorepo: no root package.json, the JS/TS project lives
entirely in web/), the wiring is:
husky+lint-stagedare devDependencies ofweb/package.json(not a root package — there isn't one).- The committed hook scripts live at the repo root:
.husky/pre-commit,.husky/pre-push,.husky/commit-msg. web/package.json'spreparescript (cd .. && husky) runs onnpm installinsideweb/and points git at the repo-root.huskydir (git config core.hooksPath .husky/_— the_subdir is husky's generated internal dir, gitignored via its own.husky/_/.gitignore; only the hook scripts themselves are committed). This works because npm keepsweb/node_modules/.binonPATHfor thepreparescript even after itcd ..s to the repo root (which husky's init requires — it hard-checks for.gitin the current directory).
The four hooks:
pre-commit— (a)cd web && npx lint-staged: runseslint --fixon stagedweb/src/**/*.{ts,tsx}files, then a project-widenpm run typecheck(tsc -bisn't file-scoped, so it runs the full check, but only when a.ts/.tsxfile is staged); (b) back at the repo root, if any*.csfiles are staged,dotnet format ErsatzTV.sln --verify-no-changes --include <staged .cs>— a formatting violation blocks the commit. The .cs step is skipped entirely when no .cs is staged, so web-only commits don't pay the sln-load cost; when it does run it's scoped to the staged files (~6-7s wall in practice, dominated by the workspace load); (c) H3 (ersatztv#303) — refuses a staged root-level*.png(git diff --cached --name-only | grep -E '^[^/]+\.png$'), belt-and-suspenders with the.gitignorescreenshot rule so a forcedgit add -fstill can't land a review/debug screenshot at the repo root. Nested*.png(real assets) pass; (d) decision lifecycle validator (ersatztv#521, supersedes the ersatztv#303 H9 append-only mechanic) — runs.claude/hooks/decisions-guard.sh(no args; a fail-open shim aroundscripts/decisions_validate.py), the structural checks over the working tree (metadata well-formedness, one active record per key, reciprocal links). It has no base/head here, so the body-diff/no-vanish checks it also knows about are skipped locally and only run in the CIdecisions lifecyclejob, which has a PR base to diff against.pre-push— CI-parity gate:cd web && npm run check:api && npm run lint && npm run typecheck && npm run build.check:apiguards generated-OpenAPI drift (ErsatzTV/wwwroot/openapi/v1.json→web/src/api/generated/v1.d.ts); the full lint/typecheck/build catch a staged change that breaks an unstaged file (lint-staged only sees staged files). Any failure blocks the push.commit-msg— enforces the CLAUDE.md protocol: the message must carry aCo-Authored-By:trailer, else the commit is rejected (merge commits are exempt, detected viagit rev-parse --verify MERGE_HEAD). The decision-lifecycle check lives inpre-commit(above), not here — theDecisions-Edit:trailer is read from the commit message, but only by the CIdecisions lifecyclejob's body-diff step (rangemode over the PR's merge-base diff), which is the only place a base/head range exists to diff against.
- Worktree/subdir gotcha: git exports
GIT_DIR(and friends) while running hooks. In a worktree or any subdir, an explicitGIT_DIRmakes nestedgitcommands mislocate the working tree —pre-push'scheck:api(git diff --exit-code, run fromweb/) then silently reports "no diff" and lets drift through.pre-pushthereforeunsetsGIT_DIR GIT_WORK_TREE GIT_INDEX_FILEfirst. (pre-commit's.cscollection usesgit diff --cached, index-vs-HEAD, which needs onlyGIT_DIRand is unaffected.) - Practical effect: a fresh
web/npm install(after cloning or pulling this change) installs all four hooks automatically — no separate setup step. Commits that touch only non-web/, non-.csfiles skip linting/formatting (lint-staged no-ops with nothing to run, the.csstep is skipped).
Registry
Gitea Packages, HTTP-only at 192.168.1.95:3000. the ci-runner VM's Docker daemon (192.168.1.127) has it as an
insecure-registry (server-management#172; runner relocated off bumblebee in #570). Images: 192.168.1.95:3000/timothy/ersatztv:<tag>.
Test / prod environments
Container/compose wiring lives in server-management (project boundary): test
ersatztv-test on 8410 (:latest), prod ersatztv on 8409 (:prod). See
server-management#481 for the full spec (registry pull on the docker host, volumes, Jellyfin
isolation for test, Watchtower/manual promotion).
Retired upstream workflows
The upstream .github/workflows/ (ci.yml, docker.yml, artifacts.yml,
release.yml, pr.yml, issue-stale.yml) were removed — they targeted
GHCR/DockerHub + Azure/Apple signing and called reusable workflows at dead
ersatztv/ersatztv@main paths, and ran as noise (incl. a daily stale-issue cron) on
the Gitea runner. Upstream is archived, so there are no future merges to preserve them
for. The dead .github/dependabot.yml and FUNDING.yml (upstream-pointed) were also
removed.
Known follow-ups
- Pin third-party actions to commit SHAs (currently floating major tags cloned from github.com at runtime) — low priority for a homelab; tracked informally.