Review of 1b78dc9e found the pre-filter's correctness claim was false, and the claim was in the decision record as well as the code. F1 (high). The pattern JSON-encoded the whole query prefix on the reasoning that the stored text escapes non-ASCII, so encoding the prefix the same way would line up. It does not: SQL LOWER() lowercases the *escape text* (`É` -> `é`); it cannot case-fold the codepoint that escape denotes. So `q=é` built `%"é%`, the stored `Édith Piaf` never matched, and the row was discarded before the in-memory filter could accept it. Every accented artist — Beyoncé, Björk, Sigur Rós, Édith Piaf — was silently unsuggestable, which in a music library is the common case. The invariant that was missing, now stated in the code: the SQL pre-filter is an OPTIMIZATION. It may over-match; it must never under-match. Correctness lives in the in-memory filter. So the pattern now narrows only on the leading run of characters the JSON writer stores verbatim and stops at the first character it cannot prove — `q=Beyoncé` still narrows on `beyonc`, `q=é` narrows on nothing and leans on the row cap. Soundness rests on two facts now asserted by exhaustive computation rather than argued: no non-ASCII codepoint in U+0080..U+10FFFF OrdinalIgnoreCase-equals a printable ASCII character (false for InvariantCultureIgnoreCase, which folds ~190 — the choice of Ordinal is load-bearing), and the exact set of ASCII the encoder escapes. F1b. `UseRequestLocalization` honours Accept-Language, so the culture was caller-controlled and `ToLower()` plus the default linguistic `StartsWith(string)` let a header change the answer. Comparison is now OrdinalIgnoreCase and ordering StringComparer.Ordinal throughout — including the shared FilterSortTake that state/video_dynamic_range/content_rating also use. Sets unchanged, order now ordinal rather than culture-dependent. F2. The merge comment asserted an exactness the code does not have: sources truncate by their own ordering (DB collation / primary key), not the merge's, so a dropped value can outrank a survivor. Comment and record now say best-effort, exact only below the truncation points. F3/F4. The cap now rides `ORDER BY Id` rather than the JSON column: MySQL sorts TEXT by only max_sort_length bytes, so the old ordering was not deterministic there, and sorting the whole matching set was avoidable work. What the cap still does NOT bound is the scan — a leading-wildcard LIKE cannot seek an index — so that cost is now documented as accepted, with a normalized `SongArtist` table named as the follow-up candidate rather than left implicit. Every clause above is covered by a test verified to FAIL when that clause is mutated (old pattern builder: 5 red; culture chain: 3 red; cap=3 / cap=limit / ORDER BY json / no cap: red each). F5. Converted to a proper supersession. The old record did not merely hold a stale fact — it recorded song/music-video credits as an "intentionally-uncovered gap" and album_artist as unsupported, and this reverses that call, which `docs.decision-lifecycle` says is never a line-edit. `api.search-field-values` is archived with its original prose restored, and `api.search-field-values-sources` replaces it carrying the whole endpoint contract.
47 lines
4.5 KiB
Markdown
47 lines
4.5 KiB
Markdown
---
|
|
key: api.search-field-values
|
|
title: 2026-07-23 — Facet-value typeahead is a new endpoint, allow-listed to text fields, no caching (#434)
|
|
status: superseded
|
|
since: '2026-07-23'
|
|
supersedes: none
|
|
superseded-by: api.search-field-values-sources@2026-07-26
|
|
rule: '(superseded) `GET /api/v1/search/fields/{name}/values?q=&limit=` returns distinct WHOLE values from the database for one of a narrow allow-list of catalog fields (not the Lucene term dictionary — analyzed `TextField`s store lowercased word tokens, e.g. "Science Fiction" → `science`/`fiction`, useless as a typeahead suggestion), 404 for an unknown field, a non-`text` field, or a `text` field with no distinct-value source; case-insensitive prefix-filtered on `q`, `limit` clamped to `[1, 50]` (default 50).'
|
|
signals: 'facet-value typeahead, rule builder value combobox, distinct field values, GetSearchFieldValues, text field allow-list, DB-sourced distinct values, content_rating split · paths: `ErsatzTV/Controllers/Api/SearchController.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValues.cs`, `ErsatzTV.Application/Search/Queries/GetSearchFieldValuesHandler.cs`, `web/src/api/search.ts` · issues: #434, #176'
|
|
mechanics: superseded by `api.search-field-values-sources` (ersatztv#578), which keeps this endpoint contract and reverses the "no distinct-value source" call for the list-valued music fields
|
|
---
|
|
|
|
Enum fields (e.g. `type`, `content_rating` group) already ship their allowed values inline on
|
|
`SearchFieldResponseModel` from the existing `GET /api/v1/search/fields` catalog (`spa.smartcollection-rule-builder`,
|
|
#176), so they need no endpoint — a client already has the full value set. **Text** fields (title, studio,
|
|
genre-as-free-text, etc.) don't: their values are whatever strings the library actually contains, so the
|
|
rule builder's value input for a text field needs a live lookup rather than a fixed list.
|
|
The handler allow-lists on `field.Type != "text"` (matching the same `SearchFieldCatalog.Fields` the
|
|
`/fields` endpoint serves) and returns `Option.None` → 404 for anything else, rather than silently returning
|
|
an empty list for a field that will never have values — a 404 tells a caller "wrong field kind," an empty
|
|
200 would look like "no matches yet."
|
|
|
|
**DB-sourced, not the search index.** The handler injects `IDbContextFactory<TvContext>` and resolves an
|
|
explicit per-field-name `IQueryable<string>` (or, for a few special cases, an in-memory list) rather than
|
|
querying `ISearchIndex`: `genre`/`show_genre` → `Set<Genre>()`, `studio` → `Set<Studio>()`, `director` →
|
|
`Set<Director>()`, `writer` → `Set<Writer>()`, `actor` → `Actors`, `artist` → `ArtistMetadata.Title` (entity
|
|
artists only — free-text music-video/song artist credits are a known, intentionally-uncovered gap), `tag` →
|
|
`Set<Tag>()` excluding `Tag.NfoCountryTypeId`/`Tag.PlexNetworkTypeId` (reapplying the indexer's own
|
|
exclusions so country/network strings don't leak in as tags), `network` → `Set<Tag>()` filtered to
|
|
`Tag.PlexNetworkTypeId`, `collection` → `Collections`, `video_codec` → `MediaStreams` filtered to
|
|
`MediaStreamKind.Video`, `album` → `MusicVideoMetadata.Album` concatenated with `SongMetadata.Album`. Every
|
|
DB-sourced field runs the same pipeline: `.Where(v => v.ToLower().StartsWith(qLower)).Distinct().OrderBy(v =>
|
|
v).Take(limit)`, translated to SQL by EF for both SQLite and MySQL. Two fields are computed in memory instead
|
|
of queried: `state` (the fixed 4-value `MediaItemState` enum) and `video_dynamic_range` (the literal
|
|
`["hdr", "sdr"]`). `content_rating` is special-cased: the DB stores an unsplit `"PG-13/TV-14"` string across
|
|
`MovieMetadata`/`ShowMetadata`/`OtherVideoMetadata`/`RemoteStreamMetadata`, so the handler pulls the distinct
|
|
raw strings then `Split('/')`s, trims, and dedupes in memory before the same prefix-filter/sort/take — this
|
|
matches what search actually matches on, rather than surfacing the compound string as one facet value.
|
|
**`title`, `show_title`, `album_artist` are explicitly NOT supported** (404, free-text fallback): `title`/
|
|
`show_title` are near-unique free-text fields spanning ~9 metadata tables where a distinct list of every
|
|
title isn't a useful facet; `album_artist` backs onto `SongMetadata.AlbumArtists`, a value-converted
|
|
`IList<string>` column EF can't translate into a server-side distinct query.
|
|
|
|
**Why a thin query, not a cache.** No result cache, no debounce on the server side (the SPA combobox
|
|
debounces the keystroke) — each per-field query is a bounded, indexed `Distinct`/`Take`; adding a cache
|
|
before there's a measured cost would be premature.
|