Files
ersatztv/docs/decisions/records/ci/small-lane-git-only.md
T
timothy fba5233caf
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
feat(610): split the decision corpus into one YAML-frontmatter file per record
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
2026-07-25 19:45:09 +02:00

3.1 KiB
Raw Blame History

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
ci.small-lane-git-only 2026-07-20 — `runs-on: small` means git-only; the two `docker build` jobs move to `ubuntu-latest` (server-management#639) active 2026-07-20 none none `runs-on: small` is defined by what a job does (git-only), not its usual runtime; the two `docker build` jobs (docker-build.yml, ci-image.yml) move to `ubuntu-latest` because their worst-case memory, not median runtime, was pinning the small lane's per-slot cap. CI lane definition, per-job memory cap, small lane widening, memory cap vs capacity, act setup-phase hang, docker build placement · paths: `.gitea/workflows/docker-build.yml`, `.gitea/workflows/ci-image.yml`, runner config · issues: server-management#639, #406, #604, #574 sum-of-caps rule (#406/#604); second jazz runner at `--cpu-shares=128`; `docs/ci-cd.md`
  • The small lane is defined by what a job does, not by how long it usually takes. Both jobs removed from it here were justified as small on a runtime argument that only held in the common case: docker-build.yml's build is a 1-second skip on PR runs (but a real image build on main/tags), and ci-image.yml's build was reasoned about as "docker-only, no toolchain needed — it builds the toolchain", which is true and yet describes the single heaviest job in the lane. The lane's per-job memory cap is set by its worst member, not its median, so both of these forced --memory=10g.
  • That cap, not a capacity decision, is what pinned the lane at one slot. 10 GiB per slot on a 25 GiB host that also runs prod media permits exactly one — the sum-of-caps rule from #406/#604 (6 slots × 10 GiB on a 25 GiB host produced load 340 and 21 GiB of swap). So "widen the lane" and "keep the heavy jobs" were never simultaneously available; the earlier note in the runner config had parked the widening indefinitely behind moving the lane to a different host.
  • Fixing the cap dominates fixing the capacity. With both builds on ubuntu-latest, small is a checkout plus a git diff, cappable at 1 GiB, so it widened from 1 slot to 4 across two hosts while committing less RAM to CI than the single slot did. A second runner was added on jazz at --cpu-shares=128 — CI on a prod media host is only acceptable while it loses every scheduling contest to the transcoders.
  • The symptom this fixes is not queue wait. A saturated lane also wedges dispatched jobs in act's setup phase: >10 min in_progress, no log file written at all, then failure, before Checkout runs. That produced the standing "decisions.md is a known flake, just rerun it" belief — the rerun works only because it lands after load clears, so a capacity problem read as a bug in the guard. A job that fails with zero log output is evidence about the runner, not about the job.
  • #574's skip-task queueing does not return by moving build back to ubuntu-latest: needs: [test, migrations] means it cannot be dispatched until the jobs it would have queued behind have already finished.