Files
ersatztv/docs/decisions/archive/api/search-field-values.md
T
timothy 1641ca8305 fix(578): the LIKE prefilter under-matched every accented artist; make the superset provable
Review of 1b78dc9e found the pre-filter's correctness claim was false, and the claim was in the
decision record as well as the code.

F1 (high). The pattern JSON-encoded the whole query prefix on the reasoning that the stored text
escapes non-ASCII, so encoding the prefix the same way would line up. It does not: SQL LOWER()
lowercases the *escape text* (`É` -> `é`); it cannot case-fold the codepoint that escape
denotes. So `q=é` built `%"é%`, the stored `Édith Piaf` never matched, and the row was
discarded before the in-memory filter could accept it. Every accented artist — Beyoncé, Björk,
Sigur Rós, Édith Piaf — was silently unsuggestable, which in a music library is the common case.

The invariant that was missing, now stated in the code: the SQL pre-filter is an OPTIMIZATION. It
may over-match; it must never under-match. Correctness lives in the in-memory filter. So the pattern
now narrows only on the leading run of characters the JSON writer stores verbatim and stops at the
first character it cannot prove — `q=Beyoncé` still narrows on `beyonc`, `q=é` narrows on nothing
and leans on the row cap. Soundness rests on two facts now asserted by exhaustive computation rather
than argued: no non-ASCII codepoint in U+0080..U+10FFFF OrdinalIgnoreCase-equals a printable ASCII
character (false for InvariantCultureIgnoreCase, which folds ~190 — the choice of Ordinal is
load-bearing), and the exact set of ASCII the encoder escapes.

F1b. `UseRequestLocalization` honours Accept-Language, so the culture was caller-controlled and
`ToLower()` plus the default linguistic `StartsWith(string)` let a header change the answer.
Comparison is now OrdinalIgnoreCase and ordering StringComparer.Ordinal throughout — including the
shared FilterSortTake that state/video_dynamic_range/content_rating also use. Sets unchanged,
order now ordinal rather than culture-dependent.

F2. The merge comment asserted an exactness the code does not have: sources truncate by their own
ordering (DB collation / primary key), not the merge's, so a dropped value can outrank a survivor.
Comment and record now say best-effort, exact only below the truncation points.

F3/F4. The cap now rides `ORDER BY Id` rather than the JSON column: MySQL sorts TEXT by only
max_sort_length bytes, so the old ordering was not deterministic there, and sorting the whole
matching set was avoidable work. What the cap still does NOT bound is the scan — a leading-wildcard
LIKE cannot seek an index — so that cost is now documented as accepted, with a normalized
`SongArtist` table named as the follow-up candidate rather than left implicit.

Every clause above is covered by a test verified to FAIL when that clause is mutated (old pattern
builder: 5 red; culture chain: 3 red; cap=3 / cap=limit / ORDER BY json / no cap: red each).

F5. Converted to a proper supersession. The old record did not merely hold a stale fact — it
recorded song/music-video credits as an "intentionally-uncovered gap" and album_artist as
unsupported, and this reverses that call, which `docs.decision-lifecycle` says is never a
line-edit. `api.search-field-values` is archived with its original prose restored, and
`api.search-field-values-sources` replaces it carrying the whole endpoint contract.
2026-07-27 03:10:35 +02:00

47 lines
4.5 KiB
Markdown

---
key: api.search-field-values
title: 2026-07-23 — Facet-value typeahead is a new endpoint, allow-listed to text fields, no caching (#434)
status: superseded
since: '2026-07-23'
supersedes: none
superseded-by: api.search-field-values-sources@2026-07-26
rule: '(superseded) `GET /api/v1/search/fields/{name}/values?q=&limit=` returns distinct WHOLE values from the database for one of a narrow allow-list of catalog fields (not the Lucene term dictionary — analyzed `TextField`s store lowercased word tokens, e.g. "Science Fiction" → `science`/`fiction`, useless as a typeahead suggestion), 404 for an unknown field, a non-`text` field, or a `text` field with no distinct-value source; case-insensitive prefix-filtered on `q`, `limit` clamped to `[1, 50]` (default 50).'
signals: 'facet-value typeahead, rule builder value combobox, distinct field values, GetSearchFieldValues, text field allow-list, DB-sourced distinct values, content_rating split · paths: `ErsatzTV/Controllers/Api/SearchController.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValues.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValuesHandler.cs`, `web/src/api/search.ts` · issues: #434, #176'
mechanics: superseded by `api.search-field-values-sources` (ersatztv#578), which keeps this endpoint contract and reverses the "no distinct-value source" call for the list-valued music fields
---
Enum fields (e.g. `type`, `content_rating` group) already ship their allowed values inline on
`SearchFieldResponseModel` from the existing `GET /api/v1/search/fields` catalog (`spa.smartcollection-rule-builder`,
#176), so they need no endpoint — a client already has the full value set. **Text** fields (title, studio,
genre-as-free-text, etc.) don't: their values are whatever strings the library actually contains, so the
rule builder's value input for a text field needs a live lookup rather than a fixed list.
The handler allow-lists on `field.Type != "text"` (matching the same `SearchFieldCatalog.Fields` the
`/fields` endpoint serves) and returns `Option.None` → 404 for anything else, rather than silently returning
an empty list for a field that will never have values — a 404 tells a caller "wrong field kind," an empty
200 would look like "no matches yet."
**DB-sourced, not the search index.** The handler injects `IDbContextFactory<TvContext>` and resolves an
explicit per-field-name `IQueryable<string>` (or, for a few special cases, an in-memory list) rather than
querying `ISearchIndex`: `genre`/`show_genre``Set<Genre>()`, `studio``Set<Studio>()`, `director`
`Set<Director>()`, `writer``Set<Writer>()`, `actor``Actors`, `artist``ArtistMetadata.Title` (entity
artists only — free-text music-video/song artist credits are a known, intentionally-uncovered gap), `tag`
`Set<Tag>()` excluding `Tag.NfoCountryTypeId`/`Tag.PlexNetworkTypeId` (reapplying the indexer's own
exclusions so country/network strings don't leak in as tags), `network``Set<Tag>()` filtered to
`Tag.PlexNetworkTypeId`, `collection``Collections`, `video_codec``MediaStreams` filtered to
`MediaStreamKind.Video`, `album``MusicVideoMetadata.Album` concatenated with `SongMetadata.Album`. Every
DB-sourced field runs the same pipeline: `.Where(v => v.ToLower().StartsWith(qLower)).Distinct().OrderBy(v =>
v).Take(limit)`, translated to SQL by EF for both SQLite and MySQL. Two fields are computed in memory instead
of queried: `state` (the fixed 4-value `MediaItemState` enum) and `video_dynamic_range` (the literal
`["hdr", "sdr"]`). `content_rating` is special-cased: the DB stores an unsplit `"PG-13/TV-14"` string across
`MovieMetadata`/`ShowMetadata`/`OtherVideoMetadata`/`RemoteStreamMetadata`, so the handler pulls the distinct
raw strings then `Split('/')`s, trims, and dedupes in memory before the same prefix-filter/sort/take — this
matches what search actually matches on, rather than surfacing the compound string as one facet value.
**`title`, `show_title`, `album_artist` are explicitly NOT supported** (404, free-text fallback): `title`/
`show_title` are near-unique free-text fields spanning ~9 metadata tables where a distinct list of every
title isn't a useful facet; `album_artist` backs onto `SongMetadata.AlbumArtists`, a value-converted
`IList<string>` column EF can't translate into a server-side distinct query.
**Why a thin query, not a cache.** No result cache, no debounce on the server side (the SPA combobox
debounces the keystroke) — each per-field query is a bounded, indexed `Distinct`/`Take`; adding a cache
before there's a measured cost would be premature.