PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key, so one-active-record-per-key becomes a filesystem property rather than a validator check, and supersession becomes a `git mv`. WHY: the monolith was a concurrency problem before an aesthetic one. A 3,900-line append target made parallel sessions collide -- PR #605 and PR #614 both hit append-vs-append conflicts during routine rebases, and hand-resolving those inside the corpus is exactly the operation the rationale-rewrite guard exists to police. HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness does not rest on reading it. The parser was taught BOTH formats first, so the body-diff guard parses the old form at the merge-base and the new form at head -- the migration validates itself, no bypass. The proof is a field-level equivalence harness: 168 records before and after, zero lost, zero gained, zero field mismatches, zero rationale bodies differing. Reviewers should scrutinise the harness; it is the actual evidence. What measuring caught that reading would not have: - ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each topic file's preamble, mostly the only copy. Source files are kept and stripped, never deleted. They also cannot be filed per-area: topic files hold several areas and 4 of 23 areas span several files. - Archive discovery was a non-recursive glob; after the split it found ZERO archived records, surfacing as four bogus "supersedes points to unknown key" errors rather than an obvious failure. - ~32 live docs point into the corpus BY DATE, which the split dangles. Each stripped file now ends with a generated "Records formerly in this file" index, which also rescues the identical breadcrumbs in old issue comments. - decisions.md's "In this file:" list was 97 same-file anchor bullets that the split makes WRONG, not merely stale. Dropped; the generated index replaces them with links that resolve. The equivalence harness now runs against a checked-in FIXTURE, not the live corpus. The earlier version migrated the real tree, which made it a one-shot: the moment the migration landed there was nothing left to move and the tests failed for reasons unrelated to the code. A fixture keeps them testing the SCRIPT rather than the repo's current state. Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain two directories for one concept. Renaming a key is not a move -- it changes identity, breaks the equivalence proof, and invalidates MemPalace's per-key drawers. Taxonomy normalisation is separate work. refs #610
3.3 KiB
3.3 KiB
key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
| key | title | status | since | supersedes | superseded-by | rule | signals | mechanics |
|---|---|---|---|---|---|---|---|---|
| api.healthcheck-ttl-cache | 2026-07-19 — Health-check results are TTL-cached; `?refresh=true` forces a fresh run (#431) | active | 2026-07-19 | none | none | Health-check results are held in a 30s TTL cache inside `HealthCheckService`; a non-forced `GET /api/v1/health` returns the cached list, and `?refresh=true` (or a forced internal caller) bypasses it to run fresh. | health check caching, TTL, refresh query param · paths: `HealthCheckService._memoryCache`, api-conventions.md §1/§3b · issues: #431, #164 | `PerformHealthChecks(forceRefresh, ...)`; `GET /api/v1/health?refresh=true` |
HealthCheckService.PerformHealthChecks re-ran all 14 checks on every call, four of which shell out to
ffmpeg/ffprobe via CliWrap — so a bare GET /api/v1/health spawned ~4 subprocesses per request. The
existing HealthCheckSummary cache was write-only (populated + published, never read back to short-circuit
a re-run). Harmless while the SPA Dashboard health panel refreshes on-demand only, but a real cost the moment
anything polls health (a status widget, an MCP client, monitoring). Split out of #164 as the orthogonal
performance half.
- A short TTL cache of the full result list lives inside
HealthCheckService. A_memoryCacheentry ("healthcheck.results",TimeSpan.FromSeconds(30)) holds the lastList<HealthCheckResult>; a non-forced call returns it directly on a hit, skipping both the 14 checks and the summaryPublish. Chosen over "make the existing summary cache read-through" because the API returns the full per-check list, not the 2-int summary — the summary entry ("healthcheck.summary", read byGetHealthCheckSummary) is kept as-is (un-expiring) so its fallback behavior is unchanged. PerformHealthChecksgained abool forceRefreshfirst parameter (interface signature change; one implementer, 3 live callers).forceRefresh: truebypasses the cache and repopulates it.- The refresh surface is an optional
?refresh=query param on the existing GET, following the?deep=bool-query-param exemplar (api-conventions.md§1/§3b) — additive, backward-compatible, no new endpoint.[FromQuery] bool refresh→GetAllHealthCheckResultsForApi(Refresh)→PerformHealthChecks(request.Refresh, …). The SPA "Refresh health" button calls/api/v1/health?refresh=true; the initial/poll load calls the bare path (cached). A separatePOST …/refreshendpoint was rejected as unnecessary surface for a read. - Who forces vs. who reads the cache: the API GET poll path reads the cache; the startup
RunHealthChecksServiceand the troubleshooting support bundle force a fresh run (both want current state — startup is a cold cache anyway, and a diagnostic bundle should reflect now, not a ≤30s-old poll). The legacyGetAllHealthCheckResultshandler is dead (no senders) and reads the cache. - Thundering-herd on a cold cache was left out of scope (no request-coalescing lock): polling is sequential
per client and the TTL collapses steady-state load, so at most a handful of exactly-simultaneous cold callers
re-run — a once-per-30s edge, not the repeated per-request cost the issue targets. Recorded here so a later
reviewer doesn't read the absence of a
SemaphoreSlimas an oversight.