Files
ersatztv/docs/decisions/records/process/per-agent-model-routing.md
T
timothy fba5233caf
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 11s
PR Gates / Docs update reminder (pull_request) Successful in 16s
PR Gates / decisions lifecycle (pull_request) Failing after 23s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m17s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m29s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m5s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 16m5s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 17m6s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
feat(610): split the decision corpus into one YAML-frontmatter file per record
168 records -> docs/decisions/records/<area>/<topic>.md (163 active, 23 dirs) and
docs/decisions/archive/<area>/<topic>.md (5 archived). The filename IS the key,
so one-active-record-per-key becomes a filesystem property rather than a
validator check, and supersession becomes a `git mv`.

WHY: the monolith was a concurrency problem before an aesthetic one. A
3,900-line append target made parallel sessions collide -- PR #605 and PR #614
both hit append-vs-append conflicts during routine rebases, and hand-resolving
those inside the corpus is exactly the operation the rationale-rewrite guard
exists to police.

HOW IT IS VERIFIED: a ~170-file diff cannot be meaningfully read, so correctness
does not rest on reading it. The parser was taught BOTH formats first, so the
body-diff guard parses the old form at the merge-base and the new form at head --
the migration validates itself, no bypass. The proof is a field-level equivalence
harness: 168 records before and after, zero lost, zero gained, zero field
mismatches, zero rationale bodies differing. Reviewers should scrutinise the
harness; it is the actual evidence.

What measuring caught that reading would not have:

- ~500 lines sit OUTSIDE any record -- decisions.md's lifecycle schema and each
  topic file's preamble, mostly the only copy. Source files are kept and
  stripped, never deleted. They also cannot be filed per-area: topic files hold
  several areas and 4 of 23 areas span several files.
- Archive discovery was a non-recursive glob; after the split it found ZERO
  archived records, surfacing as four bogus "supersedes points to unknown key"
  errors rather than an obvious failure.
- ~32 live docs point into the corpus BY DATE, which the split dangles. Each
  stripped file now ends with a generated "Records formerly in this file" index,
  which also rescues the identical breadcrumbs in old issue comments.
- decisions.md's "In this file:" list was 97 same-file anchor bullets that the
  split makes WRONG, not merely stale. Dropped; the generated index replaces
  them with links that resolve.

The equivalence harness now runs against a checked-in FIXTURE, not the live
corpus. The earlier version migrated the real tree, which made it a one-shot:
the moment the migration landed there was nothing left to move and the tests
failed for reasons unrelated to the code. A fixture keeps them testing the
SCRIPT rather than the repo's current state.

Keys preserved verbatim, warts included: `sched` (12) and `scheduling` (1) remain
two directories for one concept. Renaming a key is not a move -- it changes
identity, breaks the equivalence proof, and invalidates MemPalace's per-key
drawers. Taxonomy normalisation is separate work.

refs #610
2026-07-25 19:45:09 +02:00

5.0 KiB

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
process.per-agent-model-routing 2026-07-25 — Name the model tier for every dispatched agent; a PreToolUse gate makes the silent default visible (#583) active 2026-07-25 none none State the model tier (and effort, where the client exposes it) in the dispatch itself for every delegated agent — bounded recon → cheapest fast tier at `low`; mechanical slice against a documented contract → mid tier; judgment-heavy work → orchestrator tier; independent review → a different model family than the implementer. subagent model routing · `model` omitted · silent tier inheritance · orchestrator tier for a mechanical slice · capability routing · prose rule vs HARD CONSTRAINT · dispatch-time checkpoint · paths: `.claude/hooks/pretooluse-agent-model.sh`, `docs/handoffs/chicorytv-issue-queue.md` · issues: #583, #436, #440 `Agent` tool `model` parameter; PreToolUse hook on the existing `Agent|Task` matcher in `.claude/settings.json`.

On 2026-07-25 a session dispatched two implementers (#436, #440) with model omitted on both calls; both silently inherited the Opus orchestrator tier. #440 was a mechanical SPA slice against an already-shipped backend contract — a plausible mid-tier candidate.

The interesting part is why, because it wasn't forgetfulness. Every rule in HARD CONSTRAINTS was followed in that same session — worktree off origin/main, one committing agent per worktree, parallelize on disjoint slices, local gate before push. Routing was the one instruction living only in a prose paragraph, and it was the one that got defaulted. Treat that as the general lesson: in a kickoff doc that is pasted into every session, the bulleted imperative list is what actually functions as the checklist, and prose above it is read as background. A rule you want followed belongs in the list, keyed, or it is advisory in practice.

Three aggravating factors, all worth checking when writing any future rule here:

  • Scope gap. The low-cost-routing paragraph is written entirely about queue selection and recon ("bounded searches, inventories, log triage, report drafting"). It never named implementers, and gave no default for the bounded-but-not-trivial case — so the largest-cost dispatch fell in a gap.
  • The wrong default is the silent one. Omitting model produces no artifact. Nothing in the session report revealed the tier; the operator had to ask. Contrast the BOM trap (process.bom-format-detection-recipe), where a hook fires because a memory describing the trap demonstrably failed to prevent it twice in one day.
  • Distance from the decision point. The rule sits ~line 84 of the kickoff; dispatch happens after orientation, claiming, the bundle scan and doc reading.

Hence the two-part fix: a keyed HARD CONSTRAINT that requires the tier to be stated out loud in the dispatch (a self-correcting mechanism — it turns an invisible omission into visible output), plus .claude/hooks/pretooluse-agent-model.sh, which asks whenever an agent is dispatched with no explicit model. It exempts only fork, whose model override the tool ignores by design, so a prompt there could not be acted on. It is ask, never deny: routing is a judgment call with no derivable right answer, unlike the H6/H10 merge gate (release.merge-consent-autogrant), which derives a verifiable state and can therefore grant or refuse outright.

The first cut was narrower, and review killed it — worth recording, because the reasoning was seductive. It fired only when the prompt text matched implementer signals (git commit, worktree, fixes #), on the theory that a gate firing on every fan-out trains one-shot dismissal. Independent review confirmed the heuristic both over- and under-fired: a read-only recon brief merely mentioning "worktree" nagged, while "author the change and open a PR", "land this on the branch" and "make the changes and commit them" all passed silently — it missed precisely the case it existed to catch. Prompt prose is not a reliable signal for authority, and a gate with an unreliable catch rate is worse than none, because it gets trusted.

Two further reasons the broad form is right, both of which the narrow version had backwards:

  • It now matches the rule it enforces. The HARD CONSTRAINT says "every dispatched agent"; a hook gating only implementer-looking dispatches contradicted its own rule.
  • Routing matters MOST for the cheap cases. The old exemption list justified itself as "read-only, so routing barely matters" — but bounded recon is exactly what should be explicitly routed down to a fast tier. The premise was also false: Explore, Plan and claude-code-guide all carry Bash, so none of them provably "cannot commit".

The noise objection is answered by the escape hatch rather than by scoping: naming a tier costs one parameter and the hook never fires again. The prompt is self-eliminating for anyone following the rule, which is the habit being built.