Files
ersatztv/docs/decisions/records/api/search-allitems-paging.md
T
timothy fba5233caf
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
feat(610): split the decision corpus into one YAML-frontmatter file per record
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
2026-07-25 19:45:09 +02:00

4.3 KiB
Raw Blame History

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
api.search-allitems-paging 2026-07-18 — Search all-items is paged to cap DoS exposure; SPA add-all pages to completeness (#293) active 2026-07-18 none none `GET /api/v1/search/all-items` is paginated (capped page size, `Totals` field) to bound DoS exposure; the SPA add-all flow pages to completeness instead of relying on an unbounded response. search all-items, pagination, DoS hardening · paths: `SearchController.SearchAllItems`, `LuceneSearchIndex`, `web/src/api/search.ts` · issues: #293, #285, #308, #384 `MaxAllItemsPageSize`/`DefaultAllItemsPageSize` clamps; `getAllSearchItemIds`

GET /api/v1/search/all-items (SearchController.SearchAllItemsQuerySearchIndexAllItemsHandler) fired ten index searches with limit: 0 (= "return every hit", LuceneSearchIndex line ~244), so a single broad query (e.g. one matching the whole library) materialized every matching doc across all ten media kinds into ten List<int> buckets and serialized them in one response — unbounded work per request. #285 closed the original unauthenticated exposure (the endpoint is now behind Api:RequireKeyForReads, default true); the residual was DoS-hardening against an authenticated caller with a very broad query. Deferred from #285 because the SPA "add all to collection/playlist" flow materializes the full id set before the add POST, so a naive hard cap would silently truncate "add all".

Decision (issue option (a), operator-confirmed): paginate the endpoint and teach the SPA add-all flow to page to completeness — rather than option (b) (a generous cap + truncation signal). Chosen because "add all" must stay complete for real use, and it matches the sibling GET /api/v1/search / GET /api/v1/channels/auto-tune/members (#384) paging convention already in the codebase.

  • Endpoint (additive). SearchAllItems gains optional pageNum (0-based) + pageSize, clamped exactly like the §1 Logs / sibling Search precedent: pageSize = Math.Clamp(pageSize, 1, MaxAllItemsPageSize) with MaxAllItemsPageSize = 1000, DefaultAllItemsPageSize = 500. pageNum is clamped Math.Clamp(pageNum, 0, MaxAllItemsPageNum) with MaxAllItemsPageNum = 2_000_000 — the upper bound keeps pageNum * pageSize (the search skip) inside int range so an absurd page number can't overflow to a 500 (the sibling Search only floors at 0; the all-items endpoint hardens the upper bound too since this is a DoS-hardening change). The clamp is applied per media kind (a page returns ≤ pageSize ids of each of the ten kinds), so one response is bounded to ≤ 10 × pageSize ids. QuerySearchIndexAllItems carries PageNum/PageSize; the handler passes skip = PageNum × PageSize, limit = PageSize into ISearchIndex.Search (native skip/limit) and reads SearchResult.TotalCount (the true total, free) per kind.
  • Response (additive, frozen-v1-safe). The ten …Ids buckets are unchanged; a new non-null nested Totals (SearchResultAllItemsTotalsResponseModel, ten …Count ints) is added so a client knows how many ids exist per kind and can page to completeness. Nothing is removed or retyped (#286 additive-only holds).
  • Deliberate default-behavior change. A caller that sends no pageSize now gets one page (default 500 / kind) plus Totals, not the entire id set. This is the security change the issue asks for; it is safe here because the only in-repo consumer is the SPA (updated in the same PR) and any external/MCP caller can read Totals and page. Recorded as intentional, not a regression.
  • SPA pages to completeness. web/src/api/search.ts getSearchAllItems(query, pageNum, pageSize) gains the params; a new getAllSearchItemIds(query) loops pages (requesting pageSize = 1000, the server max), accumulating every bucket until each kind has collected its Totals count (with an empty-page safety break against total-count drift), and returns the merged SearchAllItemIds. SearchScreen.addAll calls it instead of the single-shot fetch; the #221 stale-query guard and the single add POST are unchanged.
  • Out of scope (unchanged): the add POST itself still accepts the full merged id set in one request body — bounding that surface is a separate concern (see #308 for the add path); #293 is the GET.