168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key, so one-active-record-per-key becomes a filesystem property rather than a validator check, and supersession becomes a `git mv`. WHY: the monolith was a concurrency problem before an aesthetic one. A 3,900-line append target made parallel sessions collide -- PR #605 and PR #614 both hit append-vs-append conflicts during routine rebases, and hand-resolving those inside the corpus is exactly the operation the rationale-rewrite guard exists to police. HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness does not rest on reading it. The parser was taught BOTH formats first, so the body-diff guard parses the old form at the merge-base and the new form at head -- the migration validates itself, no bypass. The proof is a field-level equivalence harness: 168 records before and after, zero lost, zero gained, zero field mismatches, zero rationale bodies differing. Reviewers should scrutinise the harness; it is the actual evidence. What measuring caught that reading would not have: - ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each topic file's preamble, mostly the only copy. Source files are kept and stripped, never deleted. They also cannot be filed per-area: topic files hold several areas and 4 of 23 areas span several files. - Archive discovery was a non-recursive glob; after the split it found ZERO archived records, surfacing as four bogus "supersedes points to unknown key" errors rather than an obvious failure. - ~32 live docs point into the corpus BY DATE, which the split dangles. Each stripped file now ends with a generated "Records formerly in this file" index, which also rescues the identical breadcrumbs in old issue comments. - decisions.md's "In this file:" list was 97 same-file anchor bullets that the split makes WRONG, not merely stale. Dropped; the generated index replaces them with links that resolve. The equivalence harness now runs against a checked-in FIXTURE, not the live corpus. The earlier version migrated the real tree, which made it a one-shot: the moment the migration landed there was nothing left to move and the tests failed for reasons unrelated to the code. A fixture keeps them testing the SCRIPT rather than the repo's current state. Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain two directories for one concept. Renaming a key is not a move -- it changes identity, breaks the equivalence proof, and invalidates MemPalace's per-key drawers. Taxonomy normalisation is separate work. refs #610
3.5 KiB
key, title, status, since, supersedes, superseded-by, stale-after, rule, signals, mechanics, sources
| key | title | status | since | supersedes | superseded-by | stale-after | rule | signals | mechanics | sources |
|---|---|---|---|---|---|---|---|---|---|---|
| ci.peak-anon-measurement | 2026-07-19 — CI `test` job reports a sampled true peak-anon, not cache-inflated `memory.peak` (#412) | active | 2026-07-19 | none | none | 2027-03-15 | The `test` job's headline memory figure is a sampled high-water mark of cgroup `anon`, produced by `scripts/ci-peak-anon.sh`; `memory.peak` and the end-of-job `anon`/`file` split are kept only as a cache-inflated reference. | CI memory measurement · paths: `scripts/ci-peak-anon.sh`, `.gitea/workflows/*.yml` · issues: #412, #411 (prose-only predecessor, no standalone record — its `memory.peak`-headline approach is superseded by this record) | `scripts/ci-peak-anon.sh` header; `docs/ci-cd.md` → CI build memory | `scripts/ci-peak-anon.sh` (the sampler itself) · cgroup v2 `memory.peak` vs `memory.stat` `anon` accounting, measured on the bumblebee runners #412 |
Decision. The test job's memory instrument (added in #411) now reports a sampled high-water
mark of the cgroup's anon memory as the headline figure, produced by scripts/ci-peak-anon.sh
(a start step before the dotnet Build/Test/Coverage, a report step last). memory.peak and the
end-of-job anon/file split stay in the output as a cache-inflated ceiling and a reference.
Why not just memory.peak. memory.peak is the high-water mark of memory.current, which
charges reclaimable page cache to the cgroup alongside anon. A build does heavy
NuGet/npm/obj/bin/coverage I/O, so cache can dominate the peak — and page cache is reclaimed under
a tighter cap, not OOM-killed. Sizing a per-job cap (server-management#604) off memory.peak
therefore inverts the decision: a big, mostly-file peak reads like "the cap must stay high"
when it isn't. The OOM-forcing quantity is peak anon. The kernel exposes memory.peak but has
no peak-anon counter, and the end-of-job anon is the composition then, not at the peak
instant (a job that peaks mid-dotnet test then frees reports a misleadingly low anon) — so it
must be sampled. Details + the sampler's robustness rationale: scripts/ci-peak-anon.sh header
and docs/ci-cd.md → "CI build memory".
Implementation note (do not "simplify" back to memory.peak). The sampler is a detached
nohup poller that survives step-boundary re-execs (reparents to the container's PID 1) and is
reaped at container teardown; a TERM trap + sleep & wait stops it at once on report. Both steps
are continue-on-error with a fail-open script, so the instrument can never redden a green build.
Validated on bumblebee: it catches a transient 2.5 GiB anon spike that the end-of-job snapshot
reports as 0.
Compiler-server A/B verdict (refines the #406 entry's "premise looking dead" read). Measured
in the CI image, swap-off, sampled peak-anon, n=2 interleaved: OFF (the CI config) ≈ 5.84 GiB,
consistent; ON (defaults) 6.3–7.6 GiB, always higher, + a ~3 GiB resident VBCSCompiler.
Disabling the servers is worth it (consistent reduction, no resident server), but OFF sits right at
6 GiB for the build phase alone and the test job adds test + coverage on top — so #406's premise
("disabling brings peak well under 6 GiB → the budget loosens") is not supported. Size the cap
off the live test-job peak-anon this instrument now reports, not off the build-only A/B. The older
#411 probe (anon 7134 MiB) read higher than these swap-off sampled numbers and is superseded
(swap/read-method move the figure >1 GiB).