Files
ersatztv/docs/decisions/records/api/search-field-values.md
T
timothy fba5233caf
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
feat(610): split the decision corpus into one YAML-frontmatter file per record
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
2026-07-25 19:45:09 +02:00

4.4 KiB

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
api.search-field-values 2026-07-23 — Facet-value typeahead is a new endpoint, allow-listed to text fields, no caching (#434) active 2026-07-23 none none `GET /api/v1/search/fields/{name}/values?q=&limit=` returns distinct WHOLE values from the database for one of a narrow allow-list of catalog fields (not the Lucene term dictionary — analyzed `TextField`s store lowercased word tokens, e.g. "Science Fiction" → `science`/`fiction`, useless as a typeahead suggestion), 404 for an unknown field, a non-`text` field, or a `text` field with no distinct-value source; case-insensitive prefix-filtered on `q`, `limit` clamped to `[1, 50]` (default 50). facet-value typeahead, rule builder value combobox, distinct field values, GetSearchFieldValues, text field allow-list, DB-sourced distinct values, content_rating split · paths: `ErsatzTV/Controllers/Api/SearchController.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValues.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValuesHandler.cs`, `web/src/api/search.ts` · issues: #434, #176 `SearchController.GetSearchFieldValues`; `GetSearchFieldValuesHandler`; api-conventions.md; spa-conventions.md §12

Enum fields (e.g. type, content_rating group) already ship their allowed values inline on SearchFieldResponseModel from the existing GET /api/v1/search/fields catalog (spa.smartcollection-rule-builder, #176), so they need no endpoint — a client already has the full value set. Text fields (title, studio, genre-as-free-text, etc.) don't: their values are whatever strings the library actually contains, so the rule builder's value input for a text field needs a live lookup rather than a fixed list. The handler allow-lists on field.Type != "text" (matching the same SearchFieldCatalog.Fields the /fields endpoint serves) and returns Option.None → 404 for anything else, rather than silently returning an empty list for a field that will never have values — a 404 tells a caller "wrong field kind," an empty 200 would look like "no matches yet."

DB-sourced, not the search index. The handler injects IDbContextFactory<TvContext> and resolves an explicit per-field-name IQueryable<string> (or, for a few special cases, an in-memory list) rather than querying ISearchIndex: genre/show_genreSet<Genre>(), studioSet<Studio>(), directorSet<Director>(), writerSet<Writer>(), actorActors, artistArtistMetadata.Title (entity artists only — free-text music-video/song artist credits are a known, intentionally-uncovered gap), tagSet<Tag>() excluding Tag.NfoCountryTypeId/Tag.PlexNetworkTypeId (reapplying the indexer's own exclusions so country/network strings don't leak in as tags), networkSet<Tag>() filtered to Tag.PlexNetworkTypeId, collectionCollections, video_codecMediaStreams filtered to MediaStreamKind.Video, albumMusicVideoMetadata.Album concatenated with SongMetadata.Album. Every DB-sourced field runs the same pipeline: .Where(v => v.ToLower().StartsWith(qLower)).Distinct().OrderBy(v => v).Take(limit), translated to SQL by EF for both SQLite and MySQL. Two fields are computed in memory instead of queried: state (the fixed 4-value MediaItemState enum) and video_dynamic_range (the literal ["hdr", "sdr"]). content_rating is special-cased: the DB stores an unsplit "PG-13/TV-14" string across MovieMetadata/ShowMetadata/OtherVideoMetadata/RemoteStreamMetadata, so the handler pulls the distinct raw strings then Split('/')s, trims, and dedupes in memory before the same prefix-filter/sort/take — this matches what search actually matches on, rather than surfacing the compound string as one facet value. title, show_title, album_artist are explicitly NOT supported (404, free-text fallback): title/ show_title are near-unique free-text fields spanning ~9 metadata tables where a distinct list of every title isn't a useful facet; album_artist backs onto SongMetadata.AlbumArtists, a value-converted IList<string> column EF can't translate into a server-side distinct query.

Why a thin query, not a cache. No result cache, no debounce on the server side (the SPA combobox debounces the keystroke) — each per-field query is a bounded, indexed Distinct/Take; adding a cache before there's a measured cost would be premature.