--- key: ci.small-lane-git-only title: '2026-07-20 — `runs-on: small` means git-only; the two `docker build` jobs move to `ubuntu-latest` (server-management#639)' status: active since: '2026-07-20' supersedes: none superseded-by: none rule: '`runs-on: small` is defined by what a job does (git-only), not its usual runtime; the two `docker build` jobs (docker-build.yml, ci-image.yml) move to `ubuntu-latest` because their worst-case memory, not median runtime, was pinning the small lane''s per-slot cap.' signals: 'CI lane definition, per-job memory cap, small lane widening, memory cap vs capacity, act setup-phase hang, docker build placement · paths: `.gitea/workflows/docker-build.yml`, `.gitea/workflows/ci-image.yml`, runner config · issues: server-management#639, #406, #604, #574' mechanics: sum-of-caps rule (#406/#604); second jazz runner at `--cpu-shares=128`; `docs/ci-cd.md` --- - **The `small` lane is defined by what a job *does*, not by how long it usually takes.** Both jobs removed from it here were justified as small on a runtime argument that only held in the common case: `docker-build.yml`'s `build` is a 1-second skip on PR runs (but a real image build on main/tags), and `ci-image.yml`'s `build` was reasoned about as "docker-only, no toolchain needed — it *builds* the toolchain", which is true and yet describes the single heaviest job in the lane. The lane's per-job memory cap is set by its worst member, not its median, so both of these forced `--memory=10g`. - **That cap, not a capacity decision, is what pinned the lane at one slot.** 10 GiB per slot on a 25 GiB host that also runs prod media permits exactly one — the sum-of-caps rule from #406/#604 (6 slots × 10 GiB on a 25 GiB host produced load 340 and 21 GiB of swap). So "widen the lane" and "keep the heavy jobs" were never simultaneously available; the earlier note in the runner config had parked the widening indefinitely behind moving the lane to a different host. - **Fixing the cap dominates fixing the capacity.** With both builds on `ubuntu-latest`, `small` is a checkout plus a `git diff`, cappable at 1 GiB, so it widened from 1 slot to **4 across two hosts while committing less RAM to CI than the single slot did**. A second runner was added on jazz at `--cpu-shares=128` — CI on a prod media host is only acceptable while it loses every scheduling contest to the transcoders. - **The symptom this fixes is not queue wait.** A saturated lane also wedges *dispatched* jobs in act's setup phase: >10 min `in_progress`, **no log file written at all**, then failure, before Checkout runs. That produced the standing "`decisions.md` is a known flake, just rerun it" belief — the rerun works only because it lands after load clears, so a capacity problem read as a bug in the guard. A job that fails with zero log output is evidence about the runner, not about the job. - **#574's skip-task queueing does not return** by moving `build` back to `ubuntu-latest`: `needs: [test, migrations]` means it cannot be dispatched until the jobs it would have queued behind have already finished.