Build ErsatzTV Image / Delimiter ban (release path) (push) Successful in 19s
Build ErsatzTV Image / Build & test (.NET) (push) Successful in 9m13s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (push) Successful in 6m40s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (push) Successful in 6m30s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (push) Skipped
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (push) Skipped
Build ErsatzTV Image / Build & push image (amd64) (push) Successful in 4m27s
The delimiter ban protecting `build`'s `Smoke + IPTV E2E` was enforced only by a pytest in `script-tests` — `on: pull_request`, not a required context — so nothing re-checked it on a `v*` tag push, which is exactly when the candidate image is published. A `scan` job now runs the ban test and `build` lists it in `needs:`, so a red `scan` skips `build` and no image is built. Measured both directions without cutting a release: run 1928 (poisoned Smoke) → scan failed, `Build & push` skipped; run 1929 (control) → scan green, build ran. The gate rests on three different KINDS of check, because each single kind was defeated in review: the ban test; an execution probe against a poisoned copy with all three `env:` tiers layered; and `scripts/ci-prove-ban-detects.sh`, which is not a test — it poisons the real checkout and vouches only for the ban test's `build` parametrisation failing. Eight review rounds; rounds 1-5 each found a real defect in the previous fix. Refs: #767 Decisions-Edit: yes Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
3.1 KiB
3.1 KiB
key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
| key | title | status | since | supersedes | superseded-by | rule | signals | mechanics |
|---|---|---|---|---|---|---|---|---|
| ci.small-lane-git-only | 2026-07-20 — `runs-on: small` means git-only; the two `docker build` jobs move to `ubuntu-latest` (server-management#639) | active | 2026-07-20 | none | none | `runs-on: small` is defined by what a job does (git-only), not its usual runtime; the two `docker build` jobs (docker-build.yml, ci-image.yml) move to `ubuntu-latest` because their worst-case memory, not median runtime, was pinning the small lane's per-slot cap. | CI lane definition, per-job memory cap, small lane widening, memory cap vs capacity, act setup-phase hang, docker build placement · paths: `.gitea/workflows/docker-build.yml`, `.gitea/workflows/ci-image.yml`, runner config · issues: server-management#639, #406, #604, #574 | sum-of-caps rule (#406/#604); second jazz runner at `--cpu-shares=128`; `docs/ci-cd.md` |
- The
smalllane is defined by what a job does, not by how long it usually takes. Both jobs removed from it here were justified as small on a runtime argument that only held in the common case:docker-build.yml'sbuildis a 1-second skip on PR runs (but a real image build on main/tags), andci-image.yml'sbuildwas reasoned about as "docker-only, no toolchain needed — it builds the toolchain", which is true and yet describes the single heaviest job in the lane. The lane's per-job memory cap is set by its worst member, not its median, so both of these forced--memory=10g. - That cap, not a capacity decision, is what pinned the lane at one slot. 10 GiB per slot on a 25 GiB host that also runs prod media permits exactly one — the sum-of-caps rule from #406/#604 (6 slots × 10 GiB on a 25 GiB host produced load 340 and 21 GiB of swap). So "widen the lane" and "keep the heavy jobs" were never simultaneously available; the earlier note in the runner config had parked the widening indefinitely behind moving the lane to a different host.
- Fixing the cap dominates fixing the capacity. With both builds on
ubuntu-latest,smallis a checkout plus agit diff, cappable at 1 GiB, so it widened from 1 slot to 4 across two hosts while committing less RAM to CI than the single slot did. A second runner was added on jazz at--cpu-shares=128— CI on a prod media host is only acceptable while it loses every scheduling contest to the transcoders. - The symptom this fixes is not queue wait. A saturated lane also wedges dispatched jobs in act's
setup phase: >10 min
in_progress, no log file written at all, then failure, before Checkout runs. That produced the standing "decisions.mdis a known flake, just rerun it" belief — the rerun works only because it lands after load clears, so a capacity problem read as a bug in the guard. A job that fails with zero log output is evidence about the runner, not about the job. - #574's skip-task queueing does not return by moving
buildback toubuntu-latest:needs: [test, migrations, scan]means it cannot be dispatched until the jobs it would have queued behind have already finished.