Build ErsatzTV Image / CI image pin matches docker/ci (pull_request) Successful in 6s
Build ErsatzTV Image / Docs update reminder (pull_request) Successful in 8s
Build ErsatzTV Image / decisions.md append-only (pull_request) Successful in 6s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 14s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 15s
Build ErsatzTV Image / Functional E2E (curl contracts) (pull_request) Successful in 5m18s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 18m36s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 19m7s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
Refs #416
770 lines
55 KiB
Markdown
770 lines
55 KiB
Markdown
# CI/CD for the ErsatzTV Fork
|
||
|
||
The fork builds its own Docker image via **Gitea Actions** on the homelab and pushes
|
||
to the **Gitea container registry**. Runner + registry were provisioned in
|
||
server-management#172; the build pipeline is ersatztv#4; test/prod containers are
|
||
server-management#481.
|
||
|
||
## Versioning & releases
|
||
|
||
The fork inherits upstream ErsatzTV's scheme: **`vYY.<release-seq>.<patch>`** (lightweight, `v`-prefixed git tags).
|
||
|
||
- **`YY`** — two-digit year.
|
||
- **`<release-seq>`** — a sequential release counter **within the year**, reset at each year boundary. It is **not** the calendar month. (Evidence: `v25.2.0` shipped in June 2025, `v25.5.0` in Sep, `v26.3.0` in Feb 2026 — minors don't track months; and `v25.9.0` → `v26.1.0` shows the year-reset.)
|
||
- **`<patch>`** — a small follow-up/hotfix on the *same* release line (e.g. `v26.1.0` → `v26.1.1`, days later).
|
||
|
||
Upstream's final release was **`v26.3.0`** (archived). Our line continues from there:
|
||
|
||
| Tag | Meaning |
|
||
|-----|---------|
|
||
| `v26.3.1` | Upstream 26.3.0 **rebuilt on our infra** (Gitea CI/registry, fork ffmpeg base) — **no application changes**. A patch bump, because nothing functional changed. |
|
||
| `v26.4.0` | First fork release carrying application changes. Later 2026 releases continue `26.5.0`, `26.6.0`, …; a new year resets to `27.1.0`. |
|
||
| `v26.7.0` | Blazor-removal release: ChicoryTV became the only UI. |
|
||
| `v26.8.0` | Secured/versioned ChicoryTV SPA + REST API go-live release (#335). |
|
||
| `v26.9.0` | Configurable advertised IPTV base URL for M3U/XMLTV (#340) + SPA shell/routing + playouts modularization (#247/#245); on-air/Plex/library-path fixes (#99/#345/#371); coverage + functional-E2E CI (#15/#299). |
|
||
|
||
**Before cutting a release — consolidate `docs/decisions.md`.** The log is append-only between releases
|
||
(ersatztv#303 H9), so a release boundary is where superseded entries get pruned/merged and the Index
|
||
refreshed. Fold and drop any entry marked `> **Superseded …**`, then commit with `[decisions-edit]` in
|
||
the message (the append-only guard blocks history edits otherwise). Mark-and-keep during the arc,
|
||
consolidate here — or sooner if the `decisions-guard` job's 1800-line consolidation-floor warning fires.
|
||
|
||
**Cutting a release:** keep build and promotion as two explicit phases (#335):
|
||
|
||
1. Confirm `main` CI is green; run the full local gate plus `dotnet list package --vulnerable
|
||
--include-transitive`; then push a `vYY.N.P` tag on that exact `main` commit.
|
||
2. Wait for tag CI to build `:prod` + the immutable `:<version>` + `:<sha>` images. Run
|
||
`scripts/security-scan.sh` on bumblebee against **the immutable `:<version>` image**, not a
|
||
moving tag, and triage every ZAP/semgrep finding.
|
||
3. Only after the candidate passes, manually `DeployStack media-servers` and observe its pre-deploy
|
||
output. Prod's compose deliberately follows floating `:prod` (Timothy's 2026-07-11 decision), so
|
||
no CI push or pin bump is needed. Global Auto Update (`auto_update: true` on `media-servers`) is
|
||
the daily fallback, not the pre-scan promotion mechanism; do not cut a tag close enough to its
|
||
03:00 run that an unscanned digest could be promoted first.
|
||
|
||
server-management#585 source-confirmed that Global Auto Update invokes the **same** `DeployStack`
|
||
execution as a manual promotion, and extended the #553 pre-deploy hook to detect a floating-tag
|
||
**digest** change. Either path now takes the fail-closed prod backup; server-management#589 then
|
||
wired `migration-smoke.sh` into that hook, against the exact candidate and the backup it just made.
|
||
A backup, fetch, or migration-smoke failure aborts before the live container is recreated. See
|
||
`homelab-docs/Docker/ErsatzTV.md` for the operational evidence and rollback procedure.
|
||
|
||
**Gotcha:** never put a `[skip ci]` token in a commit you intend to tag — Gitea reads skip-ci from the *tagged* commit and will **suppress the release build**. (Also, `workflow_dispatch` on a tag ref isn't supported on this Gitea version, so the tag *push* must do the triggering.) Release commits, and anything you'll tag, must not contain skip-ci.
|
||
|
||
**Also avoid firing several pushes back-to-back** (e.g. a `[skip ci]` commit, then `main`, then a tag, all within ~1s). Observed once on this Gitea instance: the later events were silently dropped — no `ActionRun` records created at all, even though the runner was online and the workflow `active`. Pushing again, spaced out, created the runs normally. If a push/tag doesn't produce a run, re-push (or push an empty commit) rather than assuming the runner is broken.
|
||
|
||
## The workflow: `.gitea/workflows/docker-build.yml`
|
||
|
||
Single workflow. Gating jobs `test` + `migrations` run in parallel and gate `build`; a
|
||
non-blocking `docs-reminder` job runs on PRs only (see below). Prod deploy is **not** a CI
|
||
job — it's Komodo Global Auto Update off the `:prod` tag (see "Cutting a release").
|
||
|
||
### Triggers & tags
|
||
|
||
| Trigger | `test` job | `build` job | Image tags pushed |
|
||
|---------|:----------:|:-----------:|-------------------|
|
||
| `pull_request` | ✅ | — (skipped) | none |
|
||
| push to `main` | ✅ | ✅ | `:latest` + `:<short-sha>` |
|
||
| push tag `v*` | ✅ | ✅ | `:prod` + `:<version>` + `:<short-sha>` |
|
||
| `workflow_dispatch` | ✅ | ✅ | only if ref is `main`/`v*`, else build-only (no push) |
|
||
|
||
A **docs-only** change (see "Docs-only skip" below) reduces every ✅ above to a seconds-long no-op
|
||
that still reports its status.
|
||
|
||
`:latest` is the **test/dev** channel (every `main` commit). Prod's compose **follows the
|
||
floating `:prod`** tag (reverted from the 2026-07-07 version pin on 2026-07-11) — never
|
||
`:latest`. Both `:prod` and `:<version>` are produced by pushing a `v*` tag; prod tracks
|
||
`:prod` and is redeployed by Komodo Global Auto Update (see "Cutting a release"). The
|
||
immutable `:<version>` tags remain for reproducible rollback (`docker run …:26.6.0`).
|
||
|
||
Concurrency is scoped **per event+ref** (`group: ersatztv-build-${{ github.event_name }}-${{ github.ref }}`,
|
||
`cancel-in-progress` for PRs): PR runs parallelize across PRs, a new sync auto-cancels
|
||
its superseded run, and image builds still serialize within their own ref. Do NOT push
|
||
`main` and a `v*` tag simultaneously — those are separate groups but share the
|
||
`:buildcache` tag and the smoke container name; tag only after the main build is green.
|
||
(History: originally one global group serializing ALL runs for the single runner —
|
||
with three runners that starved the queue; changed 2026-07-11, server-management#574.)
|
||
|
||
Three runners serve the fork (server-management#570/#574): `ci-runner` (VM 127 pve4,
|
||
`ubuntu-latest`, 2 slots), `bumblebee-runner` (bumblebee, `ubuntu-latest`, 2 slots,
|
||
jobs capped `--cpus=4 --memory=10g` so CI can't starve prod media playback), and
|
||
`small-runner` (bumblebee, label **`small`**, 4 slots) — the small-jobs lane. The
|
||
`build` and `docs-reminder` jobs use `runs-on: small`: Gitea dispatches a job as a
|
||
runner task even when its `if` skips it, and those skip-tasks used to wait behind
|
||
long builds (observed 31 min) stalling every PR run.
|
||
|
||
**Lane assignment (ersatztv#390).** The `ubuntu-latest` lane has **4 slots** (2 + 2) and the
|
||
`small` lane has 4. A 2026-07-17 audit of the Actions API found the `ubuntu-latest` lane
|
||
saturated and the `small` lane idle — **queue wait exceeded every job's runtime**:
|
||
|
||
| Job | Runtime | Queue wait | Lane |
|
||
|---|---|---|---|
|
||
| `test` | 354s | 1363s | ubuntu-latest |
|
||
| `migrations` | 639s | 1428s | ubuntu-latest |
|
||
| `functional-e2e` | 520s | 1447s | ubuntu-latest |
|
||
| `api-docs` | 5s | **1722s** | ubuntu-latest → `small` → **reverted to ubuntu-latest** (#406) |
|
||
| `format` | 37s | **1731s** | ubuntu-latest → `small` → **reverted to ubuntu-latest** (#406) |
|
||
| `docs-reminder` / `decisions-guard` | 10s | **5s** | small |
|
||
|
||
`api-docs` and `format` moved to `small` because the queue wait dwarfed their runtime. Both lanes
|
||
run the *identical* `runner-images:ubuntu-latest` base, so `small` was a label with spare capacity,
|
||
not a different capability — a move only possible because those jobs now run in the **CI toolchain
|
||
image** (below) and no longer need the runner image to supply .NET/Node.
|
||
|
||
**Reverted 2026-07-17 (ersatztv#406 / server-management#604).** #390's own caveat — *"on an
|
||
API-touching PR `api-docs` does a full `dotnet build`, so it is not always small"* — turned out to
|
||
be the deciding factor, and "capacity 4 absorbs that" held only because **nothing enforces the sum**
|
||
of the lanes' per-job caps. Each job container is correctly capped (`--memory=10g`), but 6 slots ×
|
||
10 GiB = **60 GiB on a 25 GiB host** that also runs prod media; on 2026-07-17 bumblebee hit load
|
||
340 with 21 GiB swapped. These are not small jobs — a live `docker stats` caught the `format` job
|
||
container at **3.95 GiB**, which the re-sized 2 GiB `small` lane would OOM-kill outright. #604 fixes
|
||
the queue at the source instead (`ubuntu-latest` grown to 5 slots: a 48 GiB ci-runner at capacity 4
|
||
plus a bumblebee overflow slot), so the `small` lane can be reserved for genuinely-tiny shell jobs.
|
||
|
||
Queue wait is still a dominant cost and capacity is server-management's boundary — tracked in
|
||
**server-management#604**. The redundant triple-build behind those runtimes is **ersatztv#398**.
|
||
|
||
### CI build memory: no persistent compiler servers (ersatztv#406)
|
||
|
||
Roslyn's `VBCSCompiler` is a **persistent** compiler server — it outlives the `dotnet build` that
|
||
started it and keeps its heap warm for the next one. Locally that is a genuine speedup; in CI it
|
||
buys nothing, because each job container is torn down at the end of the run and there is never a
|
||
"next build" to warm. It was measured at **7.8 GB RSS** on bumblebee — the single largest consumer
|
||
on the host, and the reason each job needed a 10 GiB cap in the first place.
|
||
|
||
So the workflow's top-level `env:` disables the servers for every runner-side dotnet job:
|
||
|
||
| Variable | Effect |
|
||
|---|---|
|
||
| `UseSharedCompilation=false` | no persistent `VBCSCompiler`; `csc` runs per project and exits |
|
||
| `DOTNET_CLI_USE_MSBUILD_SERVER=0` | no persistent MSBuild server process |
|
||
| `MSBUILDDISABLENODEREUSE=1` | MSBuild worker nodes exit with the build instead of lingering |
|
||
|
||
These are MSBuild properties set as environment variables so they apply to every `dotnet`
|
||
invocation without touching each call site (MSBuild surfaces env vars as properties, and
|
||
`UseSharedCompilation` is only defaulted to `true` when empty, so the env var wins).
|
||
|
||
**What this does and does not shrink.** It helps the jobs that *compile* — `test`, `migrations`,
|
||
`api-docs` on an API-touching PR, and the in-Docker `build`. It does **not** help `format`: `dotnet
|
||
format` loads Roslyn in-process via MSBuildWorkspace and never spawns `csc`, so its measured 3.95
|
||
GiB is unaffected. Don't size the `small` lane expecting `format` to have shrunk — it hasn't.
|
||
|
||
The same three are repeated in `dependency-scan.yml`; workflow `env:` does not cross workflow files.
|
||
That one is lower-stakes (restore/list are MSBuild-driven, so it's lingering worker nodes rather
|
||
than a 7.8 GB VBCSCompiler) but it runs unattended on a cron against the prod media host.
|
||
|
||
**The workflow `env:` does not reach the `build` job's compilation**, which happens inside
|
||
`docker build` — the same three are set as `ENV` in the SDK stage of `docker/Dockerfile`. That is
|
||
the job server-management#570 measured pegging 5.999/6 GiB, so it is the one that most needs this.
|
||
Build-stage only; the final image is `FROM runtime-base`, so nothing lands in the shipped image.
|
||
|
||
The `test` job's **last** step, **Report peak container memory** (`continue-on-error`, never fails a
|
||
build), prints the cgroup's `memory.peak` plus an `anon`/`file` breakdown to the log and the job
|
||
step summary. Read it off a recent run instead of re-deriving it by hand — but read it correctly:
|
||
|
||
> ⚠️ **`memory.peak` is not "peak RSS".** It is the high-water mark of `memory.current`, which
|
||
> charges **page cache** to the cgroup as well as anonymous memory. Proven on bumblebee: a
|
||
> container with `anon=0` that merely reads an 800 MB file reports `memory.peak=826 MiB`, of which
|
||
> `file=800 MiB`.
|
||
>
|
||
> This matters because the naive reading **inverts the decision**: page cache is *reclaimed* under a
|
||
> tighter cap, not OOM-killed, so a big peak that is mostly `file` is **not** evidence that the cap
|
||
> must stay high. `anon` is the part that actually forces an OOM. **Size caps on `anon`, not on
|
||
> `peak`.** The `anon`/`file` split is read at end-of-job, so it is indicative, not the composition
|
||
> at the peak instant — sampling a true peak-anon is **ersatztv#412**.
|
||
|
||
**First measurement (PR #411, `test` job):** `peak 8305 MiB`, cache share **unquantified** by that
|
||
run. Do not fill the gap with a guess in either direction — an independent probe (full solution
|
||
`dotnet build`, same CI image, shared compilation off) came back `peak 9457 MiB` / **`anon 7134
|
||
MiB`** / `file 421 MiB`, i.e. **anon dominated and cache did not**. On that evidence a 6g cap looks
|
||
*unsafe* rather than safe, and #570's "6g proved too tight" reads as the rule rather than an
|
||
outlier.
|
||
|
||
That also means **#406's premise — *"if disabling shared compilation brings peak RSS well under 6
|
||
GiB, the whole budget loosens"* — is looking dead**: the 7134 MiB anon was measured *with* shared
|
||
compilation already off. Disabling the compiler servers remains right (no persistent 7.8 GB server
|
||
sits on the host between builds) but **do not bank a looser cap budget on it**.
|
||
|
||
Treat all of the above as strong-but-not-final: it is one probe, not the `test` job, and there is
|
||
**no pre-change baseline from this instrument** (the 7.8 GB `VBCSCompiler` figure was measured
|
||
host-wide across concurrent jobs, not inside one job container). #412 covers the real A/B.
|
||
What *is* established: no persistent compiler server survives a build, and `migrations` is green
|
||
with mysql capped at 2g with swap off.
|
||
|
||
### `services:` containers are capped explicitly (ersatztv#406)
|
||
|
||
A runner's `container.options` (`--cpus=4 --memory=10g`) applies to the **job container only, not
|
||
to `services:`**. Verified on a live `migrations` job: the job container reported
|
||
`HostConfig.Memory=10737418240`; its `mysql:8.4` service reported `mem=0 nanocpus=0` — unbounded.
|
||
So `migrations` runs added an uncapped MySQL to an already-tight host.
|
||
|
||
The mysql service now sets `--memory=2g --memory-swap=2g --cpus=2`.
|
||
|
||
**`--memory-swap` is the part that matters, and it is easy to get wrong.** Docker defaults an unset
|
||
`--memory-swap` to **twice** `--memory`, so `--memory=2g` *alone* grants 2g RAM **plus 2g of swap**.
|
||
Verified on bumblebee:
|
||
|
||
| Options | `memory.max` | `memory.swap.max` |
|
||
|---|---|---|
|
||
| `--memory=2g` | 2147483648 | **2147483648** ← 2 GiB of swap |
|
||
| `--memory=2g --memory-swap=2g` | 2147483648 | **0** ← swap disabled |
|
||
|
||
Setting `--memory-swap` equal to `--memory` disables swap for the container. On this host that is
|
||
the whole point: swap thrash is what took prod down, and a swapping mysqld mid-DDL is precisely the
|
||
pathology behind the known `Command Timeout expired` migrations flake. **Prefer a loud OOM over
|
||
silent swapping** — an OOM is a clear signal to raise the cap; swapping just degrades everything.
|
||
|
||
⚠️ **The same 2× applies to the runners' `container.options: --memory=10g`** — each job slot is
|
||
really 10 GiB RAM *plus* 10 GiB swap. The "6 slots × 10 GiB = 60 GiB on a 25 GiB host" framing
|
||
understates the promise by 2×, and it is a plausible direct mechanism for the incident's 21 GiB of
|
||
swap. Fixing that is server-management#604's call (reported there).
|
||
|
||
**On the 2g figure, honestly:** a `mysql:8.4` container with this exact env peaked at **543 MiB**
|
||
during init and settled at 481 MiB idle (probed on bumblebee, 2026-07-17) — but that is init+idle,
|
||
**not** the 787-migration replay, which grows table/definition caches idle never touches. So 2g is a
|
||
measured *floor* plus headroom, not a measured ceiling; the `migrations` job going green is what
|
||
validates it. `--cpus=2` has no measurement behind it at all — 787 sequential DDL statements on one
|
||
connection are ~1-core-bound, so it is judgement; revisit if the apply step's tail latency grows.
|
||
|
||
**Any new `services:` container needs its own explicit cap** — it will not inherit one, and it needs
|
||
`--memory-swap` set alongside `--memory` or it silently gets 2× in swap.
|
||
|
||
### `test` job
|
||
|
||
`dotnet restore` → strip the Scanner project ref (`sed -i '/Scanner/d'`, matching the
|
||
Docker build) → `dotnet build -c Release` → `dotnet test -c Release --no-build`. Gates
|
||
the image build.
|
||
|
||
- **Code coverage** (ersatztv#15): `dotnet test` runs with `--collect:"XPlat Code Coverage"
|
||
--settings coverlet.runsettings --results-directory ./coverage`, so `coverlet.collector`
|
||
(referenced by every `*.Tests` project) emits a Cobertura report per project. A follow-up
|
||
**Coverage summary** step merges them with ReportGenerator (`TextSummary` to the log,
|
||
`MarkdownSummaryGithub` to the job step summary). No floor is enforced yet ("decide on a
|
||
floor later" — #15); the step is `continue-on-error: true`, so a missing report or a
|
||
transient tool install never blocks a build.
|
||
- **`coverlet.runsettings` excludes generated EF migration code** (`**/Migrations/*.cs`,
|
||
~2.59M generated lines vs ~200k authored). Instrumenting it OOM-killed the shared `test`
|
||
job (exit 137); excluding it cuts the instrumented surface ~126× (2.5M→20k coverable
|
||
lines in the whole-solution `Architecture.Tests` process) and makes the percentage reflect
|
||
authored code.
|
||
- **Shallow checkout**: `fetch-depth: 1` (ersatztv#190) — this job never runs `git
|
||
describe`/`git log`, only `build` needs full history/tags for version computation, so
|
||
`test` and `migrations` both check out shallow. `build`'s checkout stays `fetch-depth: 0`.
|
||
- **NuGet package cache**: both `test` and `migrations` cache `~/.nuget/packages` via
|
||
`actions/cache@v4`, keyed on `hashFiles('Directory.Packages.props', 'global.json')` with
|
||
a `restore-keys` OS-level fallback (ersatztv#190). Avoids a from-scratch `dotnet restore`
|
||
on every run; the key only changes when the central package manifest or SDK pin changes.
|
||
|
||
### `build` job
|
||
|
||
1. Compute `INFO_VERSION` (`git describe` + short sha on `main`; tag version on `v*`).
|
||
2. `docker/setup-buildx-action` with `buildkitd-config-inline` setting `http = true`
|
||
for `192.168.1.95:3000` — **BuildKit does not inherit the host daemon's
|
||
`insecure-registries`**, so without this, cache/base-image/push over the HTTP
|
||
registry fails (`http: server gave HTTP response to HTTPS client`).
|
||
3. `docker/login-action` with repo secrets `REGISTRY_USER` / `REGISTRY_PASSWORD`.
|
||
4. `docker/build-push-action@v6`: amd64-only, `docker/Dockerfile`, `INFO_VERSION`
|
||
build-arg, registry layer cache (`type=registry,ref=…:buildcache`,
|
||
`cache-to … ignore-error=true`).
|
||
5. **Smoke + IPTV E2E test**: pull the just-pushed `:<sha>`, run it, poll for HTTP
|
||
readiness (`docker exec … python3` → `http://localhost:8409/`), then assert the real
|
||
Jellyfin-facing surfaces on the freshly built image (ersatztv#16): `/iptv/channels.m3u`
|
||
returns 2xx containing `#EXTM3U`, and `/iptv/xmltv.xml` returns 2xx containing a `<tv`
|
||
root. `xmltv.xml` needs `channels.xml` (written by the scheduler a few seconds after
|
||
boot), so each endpoint is polled with a deadline. Unique container name + `trap … EXIT`
|
||
cleanup; dumps container logs on failure. Catches routing / base-URL (#1) / migration
|
||
regressions that leave the app "up" but serving broken output.
|
||
|
||
### `functional-e2e` job (advisory; PR + main)
|
||
|
||
Boots the app **from source** and drives the manual live-E2E curl flows sessions have historically
|
||
re-run by hand, turning them into a CI regression net (ersatztv#299). It is the automatable half of
|
||
`docs/e2e-local.md`; the two scripts it chains run identically locally and in CI:
|
||
|
||
1. `npm ci` + `npm run build` (SPA), `dotnet build ErsatzTV.sln -c Release`. (ffmpeg comes from the
|
||
CI toolchain image; the old `apt-get install ffmpeg` step is gone — see below.)
|
||
2. `ETV_BUILD_CONFIG=Release scripts/e2e-local.sh <fresh-config>` — copies `wwwroot`, launches
|
||
`dotnet ErsatzTV.dll` in the background (logging to a file so the launch step returns once the app
|
||
is ready), prints `PID`/`CONFIG_DIR`.
|
||
3. `scripts/e2e-functional.sh http://localhost:8409 <config>` — asserts, all curl-only and
|
||
deterministic (no seeded media, no browser): the **legacy→SPA redirect sweep** (+ the `/api`,
|
||
`/artwork` never-redirect exemption), the **auth/CSRF/security-stamp** flow (setup-claim →
|
||
read-gate 401/200 → re-claim 409 → CSRF 403 → login 401/200 → logout 403/204 → post-logout
|
||
stamp-revocation 401), the **library-scan status contract** (404 unknown / 202 queued /
|
||
`scan-status` 200), and the **If-Match/412** round-trip on `rerun-collections`. A `trap` kills the
|
||
instance on step exit.
|
||
|
||
**Advisory, by design** (the issue's "keep it a separate job so a functional-E2E flake can't block the
|
||
unit-test gate"): it is **not** a `needs:` of `build` and **not (yet) a required check**, so a flake
|
||
blocks nothing. Promote it to a required check / `build` dependency once it's proven reliable — the
|
||
same staged rollout the `migrations` job used. SQLite is the default provider, so unlike `migrations`
|
||
it needs **no** DB service container. Runs on PRs and on `main` (regression net); skipped for `v*` tag
|
||
builds. Out of scope for this first cut (need the scanner subprocess + seeded media, or a browser, to
|
||
be deterministic — tracked as ersatztv#299 follow-ups): the racy 409 "already-scanning" re-trigger,
|
||
the playout-build lock 409, and the genuinely UI-interactive Playwright flows.
|
||
|
||
### Docs-only skip (ersatztv#416)
|
||
|
||
A change that touches **only** `docs/**` or `*.md` (anywhere: `README.md`, `CLAUDE.md`, handoff
|
||
files) has nothing for the heavy jobs to validate. Before this, such a change ran the entire matrix
|
||
— `test`, `migrations` (with its `mysql:8.4` service), `functional-e2e`, `format`, `api-docs` —
|
||
~9 min of warm CI for a Markdown edit.
|
||
|
||
**The mechanism, and why it is shaped this way.** Each heavy job (`test`, `migrations`,
|
||
`functional-e2e`, `build`) runs `scripts/ci-detect-docs-only.sh` as its first post-checkout step
|
||
(`id: detect`), which emits `docs_only=true|false` to `$GITHUB_OUTPUT`. Every real step in the job
|
||
is gated `if: steps.detect.outputs.docs_only != 'true'`. On a docs-only change the job runs only
|
||
checkout + detect and **reports `success` in seconds**.
|
||
|
||
**Shallow-checkout safe (the change-set diff).** `test`/`migrations` check out `fetch-depth: 1`, and
|
||
a shallow clone has **no `origin/<base>` tracking ref and no merge-base** — so a three-dot
|
||
`origin/main...HEAD` diff *errors*, the fail-safe returns `docs_only=false`, and the skip silently
|
||
never fires (the first cut shipped this bug — every docs-only PR still ran the full matrix; caught by
|
||
#416's "verify on a real PR" box). The script therefore fetches the base and diffs against
|
||
**`FETCH_HEAD`** (always written by `git fetch`, resolves in a shallow clone) with a **two-dot** tree
|
||
diff (`git diff --no-renames FETCH_HEAD HEAD`) — no merge-base required. (`api-docs`/`format` avoided
|
||
the bug only because they check out `fetch-depth: 0`.)
|
||
|
||
The jobs are **not** `if:`-skipped. That is deliberate and it is the whole trap of this issue:
|
||
|
||
- `main`'s branch protection requires two checks **by name** — `Build ErsatzTV Image / Build & test
|
||
(.NET) (pull_request)` and `Build ErsatzTV Image / EF migration integrity (SQLite + MySql)
|
||
(pull_request)`. If a docs-only PR produced **no run** for those (a workflow-level `paths-ignore`,
|
||
or an `if:`-skipped job), those contexts would never report and the PR could **never merge** — the
|
||
naive fix *bricks* docs PRs rather than speeding them up.
|
||
- On Gitea **1.25.4** an `if:`-skipped job reports commit-status state **`skipped`**, a distinct
|
||
state (verified with a throwaway probe, PR #418) — not `success`. We do **not** rely on how branch
|
||
protection treats a `skipped` **required** context. Keeping the job running and gating its *steps*
|
||
makes the required context report `success` unconditionally, which is safe by construction.
|
||
- Non-required jobs may skip freely: production already proves a `skipped` **non-required** context
|
||
does not block merge (`build` is `skipped` on every PR). So `build` skips its image steps on a
|
||
**docs-only push to `main`** (docs are not in the image, so there is nothing to rebuild); tag
|
||
builds force `docs_only=false` in the script so a release is never skipped.
|
||
|
||
The detection **biases toward running more**: `docs_only=true` only when *every* changed path is
|
||
docs; any code path, a tag build, a non-merge push, or an undeterminable diff resolves to `false`
|
||
(run the full matrix). A false `true` would skip real tests on a code change — a correctness bug —
|
||
so every ambiguous case runs everything. The `migrations` job's `mysql` service still starts on a
|
||
docs-only run (a `services:` container starts with the job regardless of step `if:`), but the
|
||
expensive 787-migration replay is skipped; the service is capped and idle for seconds.
|
||
|
||
`api-docs` and `format` already short-circuit on docs-only changes via their own path detection (no
|
||
API path / no `.cs` changed → they pass in ~5s), so they needed no change. `docs-reminder`,
|
||
`decisions-guard` and `ci-image-pin` keep running on docs-only changes — the first two are *about*
|
||
docs and must.
|
||
|
||
Not in scope: the separate redundancy of running the **whole matrix on a PR and again on the
|
||
merge-to-`main`** over identical code (ersatztv#420), and the within-run triple `dotnet build`
|
||
(ersatztv#398).
|
||
|
||
### `docs-reminder` job (non-blocking, PR-only)
|
||
|
||
A lightweight nudge that enforces the CLAUDE.md "docs-update is part of done" rule for the
|
||
one case that's easy to forget and easy to detect: a PR that touches a SPA screen
|
||
(`web/src/screens/*.tsx`) or `ErsatzTV/LegacyUiRedirects.cs` but **does not** update
|
||
`docs/blazor-route-parity.md`. It diffs the PR against its base branch and emits a
|
||
`::warning::` annotation (never fails the build — it's a reminder, not a gate; prose-doc
|
||
gates get gamed with token edits). Deliberately has **no** `setup-dotnet`/`setup-node` (and
|
||
thus no `actions/cache`), so it can't hit the cache-save hangs seen on the VM-127 runner
|
||
(server-management#570). It does not cover the remaining doc obligations in the CLAUDE.md table
|
||
(domain-model, spa-conventions) — those stay on the author. (The API contract is mechanized by the
|
||
blocking `api-docs` job, and `docs/decisions.md` by the blocking `decisions-guard` job below.)
|
||
|
||
### `decisions-guard` job (blocking, PR-only)
|
||
|
||
Enforces the `docs/decisions.md` append-only convention (ersatztv#303 H9): fails a PR whose
|
||
merge-base diff **deletes or modifies** any existing line of that file, unless a commit in the range
|
||
carries the `[decisions-edit]` token (for a factual fix or a documented supersession — see the
|
||
`decisions.md` header). Pure insertions (a normal new entry: TOC line + appended block) pass. The job
|
||
also emits a **non-blocking** consolidation nudge once the file exceeds **1800 lines** (the read-cost
|
||
floor — one default agent Read caps at 2000 lines), so append-only can't outgrow what agents read. It runs
|
||
the same `.claude/hooks/decisions-guard.sh` the Husky `commit-msg` hook uses, so the *detection logic*
|
||
is shared and can't drift. Granularity differs, deliberately: Husky checks **each commit** (`staged`
|
||
mode, that commit's own message must carry the token); CI checks the **PR-wide** net diff (`range` mode,
|
||
accepts the token in *any* commit of the range). The local hook is therefore the stricter, primary
|
||
gate; CI is the backstop for direct pushes or bypassed hooks. Like `docs-reminder`, it's a
|
||
seconds-long `git diff` with no dotnet/node setup (`runs-on: small`).
|
||
|
||
## CI toolchain image (`docker/ci/Dockerfile`, `.gitea/workflows/ci-image.yml`)
|
||
|
||
The jobs that need a toolchain — `test`, `migrations`, `functional-e2e`, `api-docs`, `format` —
|
||
run **inside** a shared image via `container:` instead of installing their toolchain per run
|
||
(ersatztv#390). They therefore carry **no** `setup-dotnet`, **no** `setup-node`, **no** `apt-get`,
|
||
and **no** `dotnet tool install`.
|
||
|
||
**What it ships:** .NET 10 SDK, Node 22, prod-identical ffmpeg/ffprobe, `git`/`python3`/`jq`/`zstd`,
|
||
and the `dotnet-ef` + `dotnet-reportgenerator-globaltool` global tools (which the `migrations` and
|
||
`test` jobs used to install on every run — **bump those versions in the Dockerfile, not the
|
||
workflow**). Project dependencies (NuGet/npm) are deliberately **not** baked in — they change per
|
||
commit and stay on `actions/cache` (`~/.nuget/packages`, `~/.npm`).
|
||
|
||
**How it's layered:** `FROM ersatztv-ffmpeg:8.1.2` + `COPY --from=mcr.microsoft.com/dotnet/sdk:10.0-noble-amd64
|
||
/usr/share/dotnet` — the same pattern `docker/Dockerfile` uses for the prod image. Our ffmpeg base is
|
||
`ghcr.io/linuxserver/baseimage-ubuntu:noble`, the same Ubuntu release as the SDK image, so the copied
|
||
SDK matches the base's glibc/ICU. Keep the ffmpeg tag on that `FROM` equal to the one
|
||
`docker/Dockerfile` pins, so CI's ffmpeg stays prod-identical — that fidelity is what the
|
||
ersatztv#299 seeded-media/scanner E2E follow-ups will need.
|
||
|
||
**Bumping the toolchain is a deliberate two-step.** The jobs pin an immutable `:<sha>`, never
|
||
`:latest`, so a bad toolchain push cannot break every job at once:
|
||
|
||
1. Merge a `docker/ci/Dockerfile` change. `ci-image.yml` publishes `ersatztv-ci:<sha>` (+ `:latest`
|
||
from `main` only — a human pointer; **jobs must never consume it**).
|
||
2. In a follow-up PR, update the pin in `docker-build.yml` — **all five jobs together**. That PR's own
|
||
CI is what proves the new image works. The pin is repeated per job because
|
||
`jobs.<id>.container.image` cannot read the workflow `env` context.
|
||
|
||
`ci-image.yml` triggers on pushes touching `docker/ci/**`, `workflow_dispatch`, and a weekly Monday
|
||
05:00 UTC cron (base-image security updates; Gitea registers `schedule` only from `main`). It runs on
|
||
the `small` lane and, like `docker-build.yml`, needs BuildKit's inline `http = true` for the HTTP
|
||
registry. Renovate tracks the Dockerfile's image pins (`dockerfile` manager, see `renovate.json`).
|
||
|
||
**Three container-specific gotchas** — worth knowing if you add a job or a step:
|
||
- **`sh`, not bash, is the default shell inside a `container:`.** act_runner runs steps as
|
||
`sh -e {0}` (dash) because it can't assume bash exists in an arbitrary image — even though ours
|
||
has it. Every bashism (`set -o pipefail`, arrays, `shopt`, `mapfile`) then dies instantly with
|
||
`set: Illegal option -o pipefail`. `docker-build.yml` therefore declares a workflow-level
|
||
`defaults: run: shell: bash`. **If you add a workflow with containerized jobs, do the same** —
|
||
outside a container the shell defaults to bash, so this failure only appears once you containerize
|
||
and it looks nothing like a shell problem (it surfaced as the `migrations` job dying in 0.13s).
|
||
- `actions/checkout` clones as root into a mounted workspace, which trips git's *"detected dubious
|
||
ownership"* guard and breaks every `git` call in a step. Fixed in the Dockerfile with
|
||
`git config --global --add safe.directory '*'`.
|
||
- *(defensive, not load-bearing)* The ffmpeg base sets `ENTRYPOINT ["ffmpeg"]` because it ships as an
|
||
ffmpeg CLI, so the Dockerfile resets `ENTRYPOINT`/`CMD`. act overrides the entrypoint anyway
|
||
(`entrypoint=["/bin/sleep" "10800"]`), so this is belt-and-braces for anyone running the image by
|
||
hand — unlike the two above, which are real.
|
||
|
||
**Bumping the pin is enforced, not remembered.** The `ci-image-pin` job (blocking, PR-only) fails if
|
||
`docker-build.yml`'s pin isn't the short sha of the last commit to touch `docker/ci/**` or
|
||
`ci-image.yml`, or if the five jobs ever pin different tags. This exists because **Renovate manages
|
||
`docker/ci/Dockerfile`'s base pins but cannot bump an opaque `:<sha>`** in `container.image` — so a
|
||
Renovate base bump would otherwise publish a new image, test the *old* one, and merge with the
|
||
Dockerfile disagreeing with the pin. A red `ci-image-pin` means: let `ci-image.yml` publish the new
|
||
`:<sha>`, then update all five pins to it.
|
||
|
||
**What it is and isn't worth.** Measured honestly (ersatztv#390): the image saves ~15–40s per job
|
||
(`setup-dotnet` is 8–19s, `setup-node` 2–5s cached, the two tool installs ~9s) plus the 110s
|
||
apt-ffmpeg step — roughly 3–8% of runtime. It is **not** where CI time goes; see the lane table above
|
||
(queue wait, server-management#604) and ersatztv#398 (742s of redundant compilation). Its durable
|
||
value is prod-identical ffmpeg, a pinned/consistent toolchain, and making jobs runner-agnostic — the
|
||
last is what allowed the lane rebalance.
|
||
|
||
## Dockerfile notes (`docker/Dockerfile`)
|
||
|
||
- Base image: **`192.168.1.95:3000/timothy/ersatztv-ffmpeg:8.1.2`** (our Gitea fork of
|
||
the archived `ghcr.io/ersatztv/ersatztv-ffmpeg`). FFmpeg 8 base image work landed in
|
||
ersatztv-ffmpeg#4; app-side compatibility work landed in ersatztv#9.
|
||
- Copies `Directory.Build.props`, `Directory.Build.targets`, `Directory.Packages.props`,
|
||
`global.json`, `.editorconfig` before `dotnet restore` so the image build uses the same
|
||
MSBuild config, central package versions, SDK pin, and analyzer severities as local/CI
|
||
builds (it previously copied only `*.sln`). `Directory.Packages.props` is **required**
|
||
here: under Central Package Management the csproj carry no inline versions, so the
|
||
image's restore fails (`NU1015`) without the central manifest.
|
||
- amd64-only (the runner/build host is x86_64). No arm32/arm64, no DMG/exe artifacts, no GHCR/DockerHub.
|
||
- **openapi-generator jar layer ordering** (ersatztv#190): the `wget` for the
|
||
openapi-generator-cli jar runs *before* the `COPY` of `ErsatzTV/wwwroot/openapi/`, so the
|
||
~30MB download layer is cached independently of the openapi spec. Previously the jar was
|
||
downloaded after that `COPY`, so any PR touching the spec (e.g. `v1.json`) busted the
|
||
download layer too and re-fetched the jar on every such change. Codegen itself still runs
|
||
after the spec `COPY`, since it needs both the jar and the spec files.
|
||
|
||
## Dependency management (Central Package Management + scans)
|
||
|
||
**Central Package Management (CPM)** — package versions live in a single repo-root
|
||
`Directory.Packages.props` (`ManagePackageVersionsCentrally=true`); the per-project
|
||
csproj reference packages by name only (no `Version=`). One source of truth, atomic
|
||
one-line bumps, and cross-project version drift is structurally impossible. To add or
|
||
change a dependency, edit the `<PackageVersion>` entry centrally — never put a `Version=`
|
||
back on a `<PackageReference>` (that trips `NU1008`). The Docker build must copy this file
|
||
before restore (see Dockerfile notes). The `.mcp/` vendored tool (gitignored, not in the
|
||
solution) keeps inline versions via a **local-only** `.mcp/Directory.Packages.props`
|
||
opt-out (`ManagePackageVersionsCentrally=false`). (ersatztv#14)
|
||
|
||
**NuGet audit** — .NET 10 runs NuGet audit on restore. Several projects set
|
||
`TreatWarningsAsErrors=true`, so vulnerable transitive packages failed the build.
|
||
`Directory.Build.props` demotes low/moderate/high advisories (NU1901-1903) to warnings
|
||
and promotes NU1904 (critical) to an error in **every** project via `WarningsAsErrors`.
|
||
The advisories that prompted this were resolved in ersatztv#8 (NCalcSync→6.x; SQLitePCLRaw
|
||
bundle 3.x) and ersatztv#314 (**Microsoft.OpenApi 2.0.0→2.7.5**, GHSA-v5pm-xwqc-g5wc High —
|
||
direct-pinned in `ErsatzTV.csproj` over the 2.0.0 that `Microsoft.AspNetCore.OpenApi` +
|
||
`Scalar.AspNetCore` pull transitively; the SQLitePCLRaw override pattern; regenerates the
|
||
OpenAPI doc byte-identically). The NU1901-1903 demotion is **kept by design**: criticals (NU1904)
|
||
still hard-block, while low/moderate/high advisories surface as warnings + via the weekly scan and
|
||
Renovate security PRs, rather than breaking unrelated PRs the moment a new transitive
|
||
advisory drops.
|
||
|
||
**Scheduled vulnerability scan** — `.gitea/workflows/dependency-scan.yml` runs weekly
|
||
(cron `0 6 * * 1`) + on `workflow_dispatch`: `dotnet list package --vulnerable
|
||
--include-transitive` over the **full** solution (incl. Scanner, which the image build
|
||
strips). `dotnet list` exits 0 even with findings, so the step (`bash -euo pipefail`)
|
||
greps for the "has the following vulnerable packages" marker and fails the run if present.
|
||
Detection only — it surfaces advisories on a schedule, a Gitea-native stand-in for
|
||
Dependabot; it does **not** open update PRs (that's Renovate — server-management#484).
|
||
Gitea registers `schedule` triggers only from the default branch, so the cron starts
|
||
after merge to `main`; use `workflow_dispatch` to run on demand. It went **green** once
|
||
ersatztv#8 cleared the NCalcSync/SQLitePCLRaw advisories — a red run now means a **new**
|
||
advisory has appeared. (ersatztv#14, ersatztv#8)
|
||
|
||
**Renovate (automated update PRs)** — `.gitea/workflows/renovate.yml` runs self-hosted
|
||
[Renovate](https://docs.renovatebot.com) weekly (cron `0 3 * * 1`) + on `workflow_dispatch`,
|
||
as a `renovate/renovate:43` container job on the shared act_runner. This is the *proposing*
|
||
layer the scan above deliberately omits: it opens grouped dependency-update PRs and
|
||
OSV-driven vulnerability-fix PRs against `main`, and maintains a **Dependency Dashboard**
|
||
issue listing the full backlog. Config is the repo-root `renovate.json` — managers `nuget`
|
||
(via CPM), `github-actions`, and `dockerfile` (scoped to the built `docker/Dockerfile`; it reads the
|
||
HTTP-only Gitea registry for the `ersatztv-ffmpeg` base via a `RENOVATE_HOST_RULES` host rule —
|
||
`insecureRegistry` + registry read creds, set in the workflow env, not the committed config). The
|
||
docker-compose manager is unused (repo compose files are `build:`-only). Auth: a dedicated
|
||
**`renovate` Gitea bot** (Write
|
||
collaborator) via repo Actions secrets `RENOVATE_TOKEN` (bot PAT) + `GH_COM_TOKEN` (no-scope
|
||
github.com PAT for changelogs — named `GH_`, **not** `GITHUB_`, a prefix Gitea reserves).
|
||
**Patch** bumps to test/dev-only packages (NUnit\*, NSubstitute, Shouldly, coverlet,
|
||
`Microsoft.NET.Test.Sdk`, `Testably.Abstractions*`, threading analyzer) **auto-merge** once
|
||
the `Build & test (.NET)` check passes — branch protection on `main` requires that context;
|
||
everything else is manual review (ersatztv is prod-bearing). Range-pinned packages (e.g. EF
|
||
Core `[9.0.x,10)`) are respected — no v10 jump. PR volume is throttled (`prConcurrentLimit`
|
||
5 + `config:recommended`'s `prHourlyLimit` 2); tick a dashboard checkbox or raise the limits
|
||
to drain faster. `workflow_dispatch` defaults to a safe **dry run**. Cross-repo rollout
|
||
tracked in server-management#484. (server-management#484)
|
||
|
||
## Security scanning — black-box DAST + SAST (`scripts/security-scan.sh`, ersatztv#314)
|
||
|
||
Every other security check we run is **in-ecosystem / white-box** — SonarAnalyzer, NetArchTest, the
|
||
adversarial fork + Codex review passes, the api-docs/format/decisions CI gates, `dotnet list package
|
||
--vulnerable` — so they share our blind spots. `scripts/security-scan.sh` is the **out-of-ecosystem,
|
||
black-box** complement and a **#197 exit criterion** (HARD GATE before remote exposure): it drives the
|
||
running product from outside our C#/review stack.
|
||
|
||
- **What it does.** Boots a **throwaway** container from the image under test (fresh empty config volume;
|
||
never the deployed prod/test container — the authenticated *active* scan sends attack payloads to write
|
||
endpoints), reads the generated machine key, and runs an **authenticated OWASP ZAP** API scan
|
||
(`zap-api-scan.py`) that imports the static `/openapi/v1.json` so it exercises **every** declared
|
||
`/api/v1` operation, injecting `X-Api-Key` on every request via a ZAP replacer rule so it reaches the
|
||
`[RequiresAuthentication]` + `RequireKeyForReads` surface (not just the `/app` shell an unauthenticated
|
||
spider sees). Then a **semgrep** SAST cross-check (`p/security-audit` + `p/secrets` + `p/csharp`). The
|
||
container is torn down on exit.
|
||
- **Where/when.** Runs on the **docker host** (bumblebee — the Mac has no docker), like
|
||
`migration-smoke.sh`: `scripts/security-scan.sh [IMAGE] [PORT]` (defaults `…:latest` / `8411`). It is a
|
||
**manual release-gate**, deliberately **not** a per-PR CI job — it needs docker + a booted image, takes
|
||
several minutes, and is noisy (expect to tune, not take raw). The continuous layer is the per-PR
|
||
white-box gates + the weekly `dependency-scan`; this is the per-release black-box pass. Re-run it each
|
||
release and before any change to the exposure posture.
|
||
- **Triage.** ZAP exits non-zero on any FAIL-level alert; the tooling is noisy, so triage each finding
|
||
false-positive vs real. Real, in-scope, go-live-blocking findings get fixed (e.g. the security headers
|
||
from the #319 baseline; the Microsoft.OpenApi pin above); LAN-expected noise (Private-IP disclosure) is
|
||
revisited only for genuine remote exposure. nuclei (template-based CVE fingerprinting) is an optional
|
||
third pass — deferred while its template fetch is blocked in the runner env (pre-seed a template volume
|
||
to add it); ZAP covers the DAST baseline and semgrep the SAST, so it is not on the critical path.
|
||
|
||
## Static analysis & formatting
|
||
|
||
**Analyzers** — `Directory.Build.props` enables the SDK analyzers at `latest-All` and turns on
|
||
`Microsoft.VisualStudio.Threading.Analyzers` for every centrally managed project.
|
||
`Directory.Build.targets` also references
|
||
**Roslynator**, **SonarAnalyzer.CSharp**, **Meziantou.Analyzer**, and **AsyncFixer** repo-wide (versions
|
||
central via CPM). All analyzer package references are guarded on `ManagePackageVersionsCentrally`, so the
|
||
gitignored `.mcp` tool—which deliberately uses inline package versions—does not inherit versionless
|
||
references.
|
||
They are introduced **incrementally** (ersatztv#15). `eng/analyzers/sdk-all-suggestion.globalconfig`
|
||
enumerates the .NET 10 SDK `All` inventory at `suggestion`; this exact-ID baseline is necessary because
|
||
the SDK's generated `latest-All` severities outrank `.editorconfig` bulk settings. `.editorconfig` keeps
|
||
the threading and curated-pack baselines at `suggestion`. Diagnostics remain visible to IDEs and
|
||
`dotnet format analyzers`, but do not create a wall of failures (a direct `latest-All` trial activated
|
||
455 existing errors in the TWAE projects).
|
||
|
||
**Promotion is the enforcement** — set a reviewed rule to `warning` in `.editorconfig` and append its ID
|
||
to the central `WarningsAsErrors` list in `Directory.Build.props`. The explicit list makes the rule block
|
||
in every project, including test projects that do not otherwise use TWAE. On a major SDK upgrade,
|
||
regenerate the checked-in SDK baseline from `analysislevel_<major>_all.globalconfig`, preserve SDK `none`
|
||
entries, and review newly introduced rules before accepting the snapshot.
|
||
|
||
Promoted rules are recorded here so the blocking subset stays intentional and reviewable:
|
||
- **Sonar `S3981` — `warning` + `WarningsAsErrors`** (ersatztv#15): rejects collection-count comparisons that are constant
|
||
regardless of collection size. Its first finding exposed `Workers.Count >= 0`, which permanently
|
||
classified scheduled memory releases as busy and skipped the intended aggressive idle collection.
|
||
|
||
- **StyleCop.Analyzers is intentionally excluded**: its latest stable (1.1.118) crashes (`AD0001`) on
|
||
C# `record` declarations, and its rules overlap the existing `.editorconfig`/Roslynator. Revisit via
|
||
the record-compatible `1.2.0-beta` only if specifically wanted.
|
||
- **The former Blazor `.razor` caveat is retired**: Blazor removal deleted the Razor sources and their
|
||
temporary Sonar `NoWarn` list. The `.razor`/`.cshtml` suggestion scopes remain in `.editorconfig` only
|
||
as a defensive default if server-rendered view code is ever reintroduced.
|
||
|
||
**Formatting** — the inherited tree still contains legacy UTF-8 BOM/whitespace debt, so the standing
|
||
policy is **format as you touch**, not a mass rewrite (ersatztv#311). The Husky pre-commit hook runs
|
||
`dotnet format --verify-no-changes` for staged C# files, and the blocking `format` CI job repeats that
|
||
check for C# files changed by the PR. Untouched legacy files remain outside the gate; `.gitattributes`
|
||
pins line endings. A one-time full-tree normalization remains a separate, unmade decision.
|
||
|
||
## Migration integrity (EF Core, both providers)
|
||
|
||
`TvContext` (`ErsatzTV.Infrastructure/Data/TvContext.cs`) has **two** migration sets — one per
|
||
provider project: `ErsatzTV.Infrastructure.Sqlite/Migrations` and
|
||
`ErsatzTV.Infrastructure.MySql/Migrations`, each with its own `TvContextModelSnapshot`. **A model
|
||
change needs a migration in BOTH.** Add them with `scripts/add-migration.sh <Name>` (runs the EF CLI
|
||
for each provider). The EF CLI pattern (provider selected by the post-`--` arg, which `Startup`
|
||
reads as the `provider` config key):
|
||
|
||
```
|
||
dotnet ef <cmd> --context TvContext --startup-project ErsatzTV \
|
||
--project ErsatzTV.Infrastructure.{Sqlite|MySql} -- --provider {Sqlite|MySql}
|
||
```
|
||
|
||
The **`migrations` job** in `docker-build.yml` runs on every push/PR and, for **each** provider:
|
||
|
||
1. `dotnet ef migrations has-pending-model-changes` — fails if an entity changed without a matching
|
||
migration (**model drift**), so a forgotten migration can't merge.
|
||
2. `dotnet ef database update` against a **fresh empty DB** — applies *all* migrations in order and
|
||
fails on any broken/un-orderable one.
|
||
|
||
- **SQLite** (the prod provider) uses a throwaway file (`ETV_CONFIG_FOLDER=$(mktemp -d)`); no service
|
||
needed. Validated: 787 migrations → 139 tables.
|
||
- **MySql** uses `ServerVersion.AutoDetect`, which **connects at config time**, so the job needs a
|
||
reachable server — provided by a `services: mysql:8.4` container (the act_runner uses Docker
|
||
execution with an auto-created per-job network — service reachable as `mysql:3306` (the old
|
||
bumblebee runner pinned network `downloadswarm`; relocated in server-management#570)). Connection
|
||
string via `MySql__ConnectionString` (→ config key `MySql:ConnectionString`). Validated: 305
|
||
migrations → 137 tables. It's an **independent gate** (not yet a `needs:` of the image build) so
|
||
the new MySql-service dependency can't block image builds until it's proven; promote it to a
|
||
required check once stable.
|
||
|
||
**Caveat — non-transactional operations**: some migrations (e.g. SQLite `PRAGMA foreign_keys`) run
|
||
outside a transaction and warn at startup; they can't be rolled back mid-migration, so review such
|
||
migrations carefully (this is part of what motivated the apply-to-fresh check before the prod
|
||
cutover, server-management#481).
|
||
|
||
**Resilience — the MySql apply is retried (concurrent-runner contention, not a model bug)**: both
|
||
runners (`ci-runner` VM 127 + `bumblebee-runner`) serve `ubuntu-latest`, and when two migration jobs
|
||
land on the **same host at once** (common when several PRs push together), each spins its own
|
||
`mysql:8.4` service container and they starve each other — producing intermittent `Command Timeout
|
||
expired` **or** mid-replay `MySqlEndOfStreamException` (dropped connection) on the MySql
|
||
apply-to-fresh-DB step. This is pure infra flakiness — `has-pending-model-changes` (the actual model
|
||
check) still passes, and the *same commit* passes on a quieter host. The job hardens against it two
|
||
ways: the connection string sets `DefaultCommandTimeout=300` (up from MySqlConnector's 30s default),
|
||
and the apply is wrapped in a **3× retry** that resumes from `__EFMigrationsHistory` (EF commits each
|
||
migration in its own transaction, so an interrupted one rolls back and the retry continues). A *real*
|
||
migration failure fails deterministically on every attempt, so the retry never masks it. If a run
|
||
still flakes past the retry, re-trigger (Gitea has no rerun API on this version — push, or the run
|
||
drains); don't treat a lone MySql-apply red as a code problem without checking the failure mode.
|
||
|
||
### Migration-on-prod-copy smoke — release path (`scripts/migration-smoke.sh`, ersatztv#315)
|
||
|
||
The `migrations` job above only proves a migration is **well-formed against a fresh, empty DB**. It
|
||
can't prove it applies cleanly to the **accumulated prod SQLite** — real row volume, historical values,
|
||
and the post-migration data steps ErsatzTV runs on startup: `DatabaseMigratorService` (a
|
||
`BackgroundService`) applies pending migrations, then `DbInitializer.Initialize` + `PopulatePathHashes`
|
||
(an `UPDATE` over the real `MediaFile` table). A migration green on a fresh DB can still fail or corrupt
|
||
on prod, and today you'd only find out mid-deploy after the container recreates.
|
||
|
||
`scripts/migration-smoke.sh` rehearses it on a **throwaway copy** of the latest prod backup — it never
|
||
touches the live DB:
|
||
|
||
```
|
||
scripts/migration-smoke.sh --image <ref-about-to-be-promoted> [--db <backup.sqlite3>] [--timeout 180]
|
||
```
|
||
|
||
It copies the backup into a temp config dir, boots the new image against it (`ETV_CONFIG_FOLDER`), and
|
||
**gates PASS on the `Done applying database migrations` log line** — not merely on HTTP readiness, since
|
||
the migrator runs concurrently with Kestrel, so the web server can serve before/while migrations run.
|
||
FAIL = the container exits before finishing, a migration exception appears in the logs, migrations
|
||
don't finish within `--timeout`, or the app won't serve `/iptv/channels.m3u` afterwards. The smoke
|
||
container, the DB copy, and the temp dir are always torn down on exit (the ErsatzTV image runs as root,
|
||
so cleanup deletes its root-owned config files from inside a throwaway root container — otherwise each
|
||
run would leak the multi-hundred-MB copy). Exit 0 = clean, 1 = migration/boot failure, 2 = usage error.
|
||
|
||
- **`--db` default**: the newest `ersatztv.sqlite3` under `$ETV_BACKUP_DIR`
|
||
(default `~/downloadswarm/ersatztv-backups` — where the host-side pre-deploy backup hook writes
|
||
timestamped snapshots). Pass `--image` = the version tag about to be promoted.
|
||
- **Where it runs**: it's meant to run **on the docker host** as a **Komodo pre-deploy step** (which
|
||
already produces the backup — see "Cutting a release" and the #553 pre-deploy backup caveat), so a
|
||
bad migration aborts the promote before the live container recreates. Wiring it into that hook is a
|
||
server-management concern (cross-repo — this repo owns the script + docs, server-management owns the
|
||
Komodo hook). Until wired, run it by hand before cutting a migration-bearing release.
|
||
- Validated live 2026-07-12: `:latest` against a copy of the 283 MB prod backup → migrations applied
|
||
cleanly, app booted and served, temp dir removed.
|
||
|
||
## Pre-commit hooks (web/)
|
||
|
||
The repo uses **husky** git hooks (installed via `web/`'s **lint-staged** + npm) to catch
|
||
lint/format/type/API-drift errors locally, before they reach CI. Because the git root and
|
||
the npm project dir differ (monorepo: no root `package.json`, the JS/TS project lives
|
||
entirely in `web/`), the wiring is:
|
||
|
||
- `husky` + `lint-staged` are devDependencies of `web/package.json` (not a root package —
|
||
there isn't one).
|
||
- The committed hook scripts live at the repo root: `.husky/pre-commit`, `.husky/pre-push`,
|
||
`.husky/commit-msg`.
|
||
- `web/package.json`'s `prepare` script (`cd .. && husky`) runs on `npm install` inside
|
||
`web/` and points git at the repo-root `.husky` dir (`git config core.hooksPath
|
||
.husky/_` — the `_` subdir is husky's generated internal dir, gitignored via its own
|
||
`.husky/_/.gitignore`; only the hook scripts themselves are committed). This works
|
||
because npm keeps `web/node_modules/.bin` on `PATH` for the `prepare` script even after
|
||
it `cd ..`s to the repo root (which husky's init requires — it hard-checks for `.git`
|
||
in the *current* directory).
|
||
|
||
**The four hooks:**
|
||
|
||
1. **`pre-commit`** — (a) `cd web && npx lint-staged`: runs `eslint --fix` on staged
|
||
`web/src/**/*.{ts,tsx}` files, then a project-wide `npm run typecheck` (`tsc -b` isn't
|
||
file-scoped, so it runs the full check, but only when a `.ts`/`.tsx` file is staged);
|
||
(b) back at the repo root, if any **`*.cs`** files are staged, `dotnet format
|
||
ErsatzTV.sln --verify-no-changes --include <staged .cs>` — a formatting violation
|
||
blocks the commit. The .cs step is **skipped entirely when no .cs is staged**, so
|
||
web-only commits don't pay the sln-load cost; when it does run it's scoped to the staged
|
||
files (~6-7s wall in practice, dominated by the workspace load); (c) **H3** (ersatztv#303) —
|
||
refuses a staged **root-level `*.png`** (`git diff --cached --name-only | grep -E '^[^/]+\.png$'`),
|
||
belt-and-suspenders with the `.gitignore` screenshot rule so a forced `git add -f` still can't land
|
||
a review/debug screenshot at the repo root. Nested `*.png` (real assets) pass.
|
||
2. **`pre-push`** — CI-parity gate: `cd web && npm run check:api && npm run lint && npm run
|
||
typecheck && npm run build`. `check:api` guards generated-OpenAPI drift
|
||
(`ErsatzTV/wwwroot/openapi/v1.json` → `web/src/api/generated/v1.d.ts`); the full
|
||
lint/typecheck/build catch a staged change that breaks an *unstaged* file (lint-staged
|
||
only sees staged files). Any failure blocks the push.
|
||
3. **`commit-msg`** — (a) enforces the CLAUDE.md protocol: the message must carry a
|
||
`Co-Authored-By:` trailer, else the commit is rejected (merge commits are exempt, detected
|
||
via `git rev-parse --verify MERGE_HEAD`); (b) **H9** (ersatztv#303) — runs
|
||
`.claude/hooks/decisions-guard.sh staged "$1"`, which blocks the commit if it deletes/modifies
|
||
an existing line of `docs/decisions.md` unless the message carries the `[decisions-edit]` token.
|
||
Append-only enforcement; see `docs/decisions.md` header for the supersession/consolidation rules.
|
||
The same script backs the blocking **`decisions-guard`** CI job (`range` mode over the PR's
|
||
merge-base diff) so local and CI enforcement can't drift.
|
||
|
||
- **Worktree/subdir gotcha**: git exports `GIT_DIR` (and friends) while running hooks. In a
|
||
worktree or any subdir, an explicit `GIT_DIR` makes nested `git` commands mislocate the
|
||
working tree — `pre-push`'s `check:api` (`git diff --exit-code`, run from `web/`) then
|
||
silently reports "no diff" and lets drift through. `pre-push` therefore `unset`s
|
||
`GIT_DIR GIT_WORK_TREE GIT_INDEX_FILE` first. (`pre-commit`'s `.cs` collection uses
|
||
`git diff --cached`, index-vs-HEAD, which needs only `GIT_DIR` and is unaffected.)
|
||
- **Practical effect**: a fresh `web/` `npm install` (after cloning or pulling this change)
|
||
installs all four hooks automatically — no separate setup step. Commits that touch only
|
||
non-`web/`, non-`.cs` files skip linting/formatting (lint-staged no-ops with nothing to
|
||
run, the `.cs` step is skipped).
|
||
|
||
## Registry
|
||
|
||
Gitea Packages, HTTP-only at `192.168.1.95:3000`. the `ci-runner` VM's Docker daemon (192.168.1.127) has it as an
|
||
insecure-registry (server-management#172; runner relocated off jazz in #570). Images: `192.168.1.95:3000/timothy/ersatztv:<tag>`.
|
||
|
||
## Test / prod environments
|
||
|
||
Container/compose wiring lives in **server-management** (project boundary): test
|
||
`ersatztv-test` on 8410 (`:latest`), prod `ersatztv` on 8409 (`:prod`). See
|
||
server-management#481 for the full spec (registry pull on jazz, volumes, Jellyfin
|
||
isolation for test, Watchtower/manual promotion).
|
||
|
||
## Retired upstream workflows
|
||
|
||
The upstream `.github/workflows/` (`ci.yml`, `docker.yml`, `artifacts.yml`,
|
||
`release.yml`, `pr.yml`, `issue-stale.yml`) were removed — they targeted
|
||
GHCR/DockerHub + Azure/Apple signing and called reusable workflows at dead
|
||
`ersatztv/ersatztv@main` paths, and ran as noise (incl. a daily stale-issue cron) on
|
||
the Gitea runner. Upstream is archived, so there are no future merges to preserve them
|
||
for. The dead `.github/dependabot.yml` and `FUNDING.yml` (upstream-pointed) were also
|
||
removed.
|
||
|
||
## Known follow-ups
|
||
|
||
- Pin third-party actions to commit SHAs (currently floating major tags cloned from
|
||
github.com at runtime) — low priority for a homelab; tracked informally.
|