Files
ersatztv/docs/decisions/archive/api/search-field-values.md
T
timothy 1641ca8305 fix(578): the LIKE prefilter under-matched every accented artist; make the superset provable
Review of 1b78dc9e found the pre-filter's correctness claim was false, and the claim was in the
decision record as well as the code.

F1 (high). The pattern JSON-encoded the whole query prefix on the reasoning that the stored text
escapes non-ASCII, so encoding the prefix the same way would line up. It does not: SQL LOWER()
lowercases the *escape text* (`É` -> `é`); it cannot case-fold the codepoint that escape
denotes. So `q=é` built `%"é%`, the stored `Édith Piaf` never matched, and the row was
discarded before the in-memory filter could accept it. Every accented artist — Beyoncé, Björk,
Sigur Rós, Édith Piaf — was silently unsuggestable, which in a music library is the common case.

The invariant that was missing, now stated in the code: the SQL pre-filter is an OPTIMIZATION. It
may over-match; it must never under-match. Correctness lives in the in-memory filter. So the pattern
now narrows only on the leading run of characters the JSON writer stores verbatim and stops at the
first character it cannot prove — `q=Beyoncé` still narrows on `beyonc`, `q=é` narrows on nothing
and leans on the row cap. Soundness rests on two facts now asserted by exhaustive computation rather
than argued: no non-ASCII codepoint in U+0080..U+10FFFF OrdinalIgnoreCase-equals a printable ASCII
character (false for InvariantCultureIgnoreCase, which folds ~190 — the choice of Ordinal is
load-bearing), and the exact set of ASCII the encoder escapes.

F1b. `UseRequestLocalization` honours Accept-Language, so the culture was caller-controlled and
`ToLower()` plus the default linguistic `StartsWith(string)` let a header change the answer.
Comparison is now OrdinalIgnoreCase and ordering StringComparer.Ordinal throughout — including the
shared FilterSortTake that state/video_dynamic_range/content_rating also use. Sets unchanged,
order now ordinal rather than culture-dependent.

F2. The merge comment asserted an exactness the code does not have: sources truncate by their own
ordering (DB collation / primary key), not the merge's, so a dropped value can outrank a survivor.
Comment and record now say best-effort, exact only below the truncation points.

F3/F4. The cap now rides `ORDER BY Id` rather than the JSON column: MySQL sorts TEXT by only
max_sort_length bytes, so the old ordering was not deterministic there, and sorting the whole
matching set was avoidable work. What the cap still does NOT bound is the scan — a leading-wildcard
LIKE cannot seek an index — so that cost is now documented as accepted, with a normalized
`SongArtist` table named as the follow-up candidate rather than left implicit.

Every clause above is covered by a test verified to FAIL when that clause is mutated (old pattern
builder: 5 red; culture chain: 3 red; cap=3 / cap=limit / ORDER BY json / no cap: red each).

F5. Converted to a proper supersession. The old record did not merely hold a stale fact — it
recorded song/music-video credits as an "intentionally-uncovered gap" and album_artist as
unsupported, and this reverses that call, which `docs.decision-lifecycle` says is never a
line-edit. `api.search-field-values` is archived with its original prose restored, and
`api.search-field-values-sources` replaces it carrying the whole endpoint contract.
2026-07-27 03:10:35 +02:00

4.5 KiB

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
api.search-field-values 2026-07-23 — Facet-value typeahead is a new endpoint, allow-listed to text fields, no caching (#434) superseded 2026-07-23 none api.search-field-values-sources@2026-07-26 (superseded) `GET /api/v1/search/fields/{name}/values?q=&limit=` returns distinct WHOLE values from the database for one of a narrow allow-list of catalog fields (not the Lucene term dictionary — analyzed `TextField`s store lowercased word tokens, e.g. "Science Fiction" → `science`/`fiction`, useless as a typeahead suggestion), 404 for an unknown field, a non-`text` field, or a `text` field with no distinct-value source; case-insensitive prefix-filtered on `q`, `limit` clamped to `[1, 50]` (default 50). facet-value typeahead, rule builder value combobox, distinct field values, GetSearchFieldValues, text field allow-list, DB-sourced distinct values, content_rating split · paths: `ErsatzTV/Controllers/Api/SearchController.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValues.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValuesHandler.cs`, `web/src/api/search.ts` · issues: #434, #176 superseded by `api.search-field-values-sources` (ersatztv#578), which keeps this endpoint contract and reverses the "no distinct-value source" call for the list-valued music fields

Enum fields (e.g. type, content_rating group) already ship their allowed values inline on SearchFieldResponseModel from the existing GET /api/v1/search/fields catalog (spa.smartcollection-rule-builder, #176), so they need no endpoint — a client already has the full value set. Text fields (title, studio, genre-as-free-text, etc.) don't: their values are whatever strings the library actually contains, so the rule builder's value input for a text field needs a live lookup rather than a fixed list. The handler allow-lists on field.Type != "text" (matching the same SearchFieldCatalog.Fields the /fields endpoint serves) and returns Option.None → 404 for anything else, rather than silently returning an empty list for a field that will never have values — a 404 tells a caller "wrong field kind," an empty 200 would look like "no matches yet."

DB-sourced, not the search index. The handler injects IDbContextFactory<TvContext> and resolves an explicit per-field-name IQueryable<string> (or, for a few special cases, an in-memory list) rather than querying ISearchIndex: genre/show_genreSet<Genre>(), studioSet<Studio>(), directorSet<Director>(), writerSet<Writer>(), actorActors, artistArtistMetadata.Title (entity artists only — free-text music-video/song artist credits are a known, intentionally-uncovered gap), tagSet<Tag>() excluding Tag.NfoCountryTypeId/Tag.PlexNetworkTypeId (reapplying the indexer's own exclusions so country/network strings don't leak in as tags), networkSet<Tag>() filtered to Tag.PlexNetworkTypeId, collectionCollections, video_codecMediaStreams filtered to MediaStreamKind.Video, albumMusicVideoMetadata.Album concatenated with SongMetadata.Album. Every DB-sourced field runs the same pipeline: .Where(v => v.ToLower().StartsWith(qLower)).Distinct().OrderBy(v => v).Take(limit), translated to SQL by EF for both SQLite and MySQL. Two fields are computed in memory instead of queried: state (the fixed 4-value MediaItemState enum) and video_dynamic_range (the literal ["hdr", "sdr"]). content_rating is special-cased: the DB stores an unsplit "PG-13/TV-14" string across MovieMetadata/ShowMetadata/OtherVideoMetadata/RemoteStreamMetadata, so the handler pulls the distinct raw strings then Split('/')s, trims, and dedupes in memory before the same prefix-filter/sort/take — this matches what search actually matches on, rather than surfacing the compound string as one facet value. title, show_title, album_artist are explicitly NOT supported (404, free-text fallback): title/ show_title are near-unique free-text fields spanning ~9 metadata tables where a distinct list of every title isn't a useful facet; album_artist backs onto SongMetadata.AlbumArtists, a value-converted IList<string> column EF can't translate into a server-side distinct query.

Why a thin query, not a cache. No result cache, no debounce on the server side (the SPA combobox debounces the keystroke) — each per-field query is a bounded, indexed Distinct/Take; adding a cache before there's a measured cost would be premature.