api-conventions.md now says which half of an id-list validator a sibling field may
gate (existence, never the raw-count cap) and that a lost-race recovery re-asks the
whole validator set rather than the fields whoever wrote the catch remembered.
Those, with the bound and the field-named 422, are one convention with residuals, so
they get a record -- api.top-level-id-list-validation -- and a task-signal row.
The record states what #568 does NOT settle: three validators on two DTOs is a
per-field constant, not the repo-wide rule #917 owns, and it says to expect #917 to
replace the mechanism.
graphics-elements.md: rows 35-44 re-measured against the whole ErsatzTV.Tests
project on this tree, because the fix moved five of their red sets -- Validate is
now also what the recovery path re-runs, so removing a validator from it reddens
that handler's race test too. Rows 45-47 are new and measured the same way. The
"redden more than one test" figure is recounted from the table (21 -> 24); the
cross-fixture set is unchanged at five.
The negative discriminator rows now carry a stated seeding rule: vary one half of
the identity and hold the other at the seeded value. Varying both leaves the row
rejected by the pre-#568 predicate as well, so a composite revert to it would pass
every test at that site.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two holes the review round found in the previous fix, both of the same shape: a
guard that names its own fields instead of deriving them.
The deco validators short-circuited the entire Validators.IdsMustExist call when
the DecoMode does not consume the ids, which took the 512-item raw-count cap with
it -- an arbitrarily large array under Inherit/Disable parsed and materialized
with nothing bounding it. Only the EXISTENCE half is the apply path's business,
so the mode predicate is now a required argument of the shared validator and gates
that half alone; the cap runs under every mode.
The channel recovery path rechecked GraphicsElementIdsMustExist alone, so a
watermark deleted between validation and SaveChangesAsync still surfaced as the
unhandled 500 the fix exists to remove -- WatermarkId, FFmpegProfileId,
FallbackFillerId and MirrorSourceChannelId are all written by the same save and
lose the same race. Both handlers now re-ask the whole of Validate on
DbUpdateException, so a validator added later is covered without editing the
recovery path.
The API-site outside-folder discriminator test seeded an Image row, so the Kind
conjunct rejected it whatever the path comparison did: a composite revert to
Path.GetFileName(path) == filename && kind == Text passed every API test. It now
carries the seeded Kind, mirroring the seeder-site twin, so only the path half can
reject it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Nine rows of the mutation table in docs/graphics-elements.md name a clause that this
branch's last commit moved, merged or gave new callers, and a red set is a measurement of
the tree it ships in. All were re-taken against the whole ErsatzTV.Tests project
(2139 tests, 6 skipped) on the code as it now stands, and four new rows added for the
clauses the fix introduced.
What moved and why the numbers changed:
- Row 10 was the seeded-path filter alone. `Kind` now lives inside `IsOnNowNext`, so
dropping the lookup's `Where` drops both halves at once and reddens four tests, not
three.
- Row 18 was the seeder's SQL `Kind == Text` filter and is now the `kind` conjunct of the
shared predicate, so it reddens the API site too — a second cross-fixture row.
- Rows 35-37 pick up the count-cap tests, since the cap rides in the validator they
disarm. Rows 33, 34, 38, 39 re-measured unchanged.
- Row 40's mutation text follows the API call's new two-argument shape; it reddens the new
wrong-kind test as well.
- Rows 41-44 are the new clauses: the raw-count cap (one clause, three call sites, which
is what its red set shows), the diagnostic-id truncation, and the two lost-race catches.
The two self-counted figures above the table were recounted from the table itself rather
than adjusted: twenty-one multi-test rows and five cross-fixture ones (13, 18, 22, 33, 41).
The deco lost-race test is renamed so no two rows cite the same test name.
api-conventions.md gains the three rules the fix establishes for any write path with a
top-level FK id list — bound the raw list, name the field, translate a lost check-then-write
race — in the handler-hardening checklist where they belong rather than as a #568 anecdote.
Refs #568
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Three of the four findings standing on the 2026-09-05 16:24 review verdict, which the
branch had not answered.
The count cap is the blocking one. The three id-list validators took whatever the
request carried, so the only bound on `graphicsElementIds`/`watermarkIds` was the
Kestrel body cap -- a transport limit, not a collection limit. The earlier disposition
deferred it to #917 on the grounds that `ApplyUpdateRequest` reconciles the same list
uncapped anyway; that is true and does not answer the ask, because the reconcile is
downstream of a validator that can refuse the request outright. One shared
`Validators.IdsMustExist` now carries the cap for all three, counted on the RAW list
before `Distinct` (a million copies of one id costs the same to parse and materialize
whatever the distinct count is) and before any database work.
The same helper is where the field name and the diagnostic cap now live. The 422 said
"Graphics element(s) do not exist: 999" without naming which request field carried the
999, and echoed every rejected id -- an oversized request answered with an oversized
response. Both fixed once, in the shared place, so the three sites cannot drift.
`Kind` moves into `GraphicsElementDefaults.IsOnNowNext`. The seeder required
`Kind == Text` and the API's `builtIn` did not, so an Image row at the exact seeded path
was `builtIn:true` on the wire while `GetBuiltInElementId` refused to treat it as the
built-in element -- two sites disagreeing about one row, which is the shape #568 exists
to close. Identity is now one predicate applied whole at both sites; the seeder's SQL
`Kind` filter is gone rather than kept as a duplicate, since a duplicate guard would mask
the predicate's own clause.
Also the fourth finding, the check-then-write race: `RefreshGraphicsElements` can delete a
validated element between `Validate` and `SaveChangesAsync`, handing the join insert the
FK violation the validator exists to prevent. A transaction does not close it -- neither
provider locks rows the validator merely read -- so both handlers catch `DbUpdateException`,
re-ask the existence question on a fresh context, and return the validator's own 422 when
an id has since gone; anything else keeps its own exception. Foreign keys are off in
`InMemoryTvContext`, so the trigger is simulated by an armed save-failure interceptor while
the recovery itself runs against real post-delete state.
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The sentence explaining why the wrong-kind-and-wrong-folder case is not shipped said
"no single-clause mutation can let it through" — an unbounded quantifier over a
population nothing here measures. What is actually established is narrower and is
established: rows 10 and 18 are the two clauses of `GetBuiltInElementId`, each measured,
and dropping either leaves the other rejecting such a row. The claim now says that, and
names those rows as its evidence.
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The §8 Channel-graphics aside cited the deep-FK-in-a-nested-list exception as "below";
that exception is §3b line 314 and the aside is line ~944, so the pointer sent the reader
the wrong way. It now names the section (§3b above) rather than a direction alone, so a
later reflow cannot invert it again. Re-wrapped the same passage so `deep-FK-in-a-nested-list`
no longer straddles a soft line break — Markdown joins those with a space and the term
rendered with a stray gap mid-word.
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`Ignores_A_Same_Named_Element_Of_A_Different_Kind` seeded an Image row at
`/templates/image/on-now-next.yml` and asserted the backfill ignores it. Under the
bare-filename lookup that row was rejected by the `Kind == Text` filter alone, which is
what its comment described. Under the full-path predicate the path rejects it first, so
neither clause is load-bearing for it: dropping the `Kind` filter reddens only
`A_Row_Of_Another_Kind_At_The_Seeded_Path_Does_Not_Suppress_The_Built_In_Row` (row 18) and
dropping `IsOnNowNext` reddens only the three tests of row 10. The test survived both and
its comment claimed a mechanism it no longer exercised. Its scenario is the conjunction of
two already-pinned negatives and is strictly weaker than
`Ignores_A_Same_Named_Same_Kind_Element_Outside_The_Seeded_Folder`, so it is retired rather
than reshaped, and graphics-elements.md now says why the combination is deliberately not
shipped — otherwise the next reader re-adds it.
Row 40 records the API-side half of the discriminator, which had a measured red and no row.
Measured whole-project on this tree, `dotnet test ErsatzTV.Tests/ErsatzTV.Tests.csproj`:
baseline `Failed: 0, Passed: 2123, Skipped: 6, Total: 2129`; with
`BuiltIn = GraphicsElementDefaults.IsOnNowNext(e.Path)` reverted to
`Path.GetFileName(e.Path) == GraphicsElementDefaults.OnNowNextFileName`,
`Failed: 1, Passed: 2122`, the sole red being
`GetAllGraphicsElementsForApi_Should_Not_Mark_Same_Filename_Outside_Seeded_Folder_As_BuiltIn`.
The table's two self-counts were recounted from the table after adding the row and both
still hold: sixteen rows redden more than one test, three of those span two fixture classes
(13, 22, 33).
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Five test comments asserted "reddens if <validator> alone is removed" with
nothing binding the sentence to a measurement -- the shape
testing.mutation-claims-are-executed refuses, and the shape whose CLAIMS half of
the manifest cannot reach a .NET proof. The repo's record for those is the
mutation table, so each claim got a row: all five mutated in turn against this
tree with the whole ErsatzTV.Tests project re-run (the tuple-arity fix included,
since a mutation that does not compile is not a result).
35 GraphicsElementIdsMustExist out of UpdateChannelHandler.Validate -> 2 red;
36/37 the deco graphics/watermark validators out of UpdateDecoHandler.Validate
-> 1 red each; 38/39 the two Consumes* mode gates -> 1 red each. Sixteen rows now
redden more than one test; three still span two fixture classes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The comment asserted an outcome ("also reddens if the Kind filter is dropped")
with nothing tying it to a measurement -- the shape testing.mutation-claims-are-
executed exists to refuse. Both clauses it covers are rows of the mutation table
in docs/graphics-elements.md, measured against this tree; cite them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The seeder now resolves the built-in row through GetBuiltInElementId, which puts
that lookup on a second call path, so every row whose clause the new call can
reach was re-run against this tree: 10 and 33 unchanged, 18 reinstated (the
Kind == Text filter has a red now that a wrong-kind row at the seeded path can
suppress the row the lookup needs), 21 unchanged, 22 gains a third red, and 34
is new (the existence check re-derived as SQL instead of asking the lookup).
Two stale measurements went with it. The "known clauses with no red" bullet for
the Kind filter quoted 2121 passed against a tree that produces 2123, having been
taken before the branch's last two tests existed -- the whole bullet is gone now
that the clause has a red. And the per-fixture-filter trap counted thirteen
multi-test rows with two spanning two fixture classes, true on origin/main and
false here since the branch added rows: fifteen and three, both recounted from
the table, with a note that they are.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
EnsureBuiltInElementRow decided whether the built-in row already existed with its
own `AnyAsync(e => e.Path == target)` -- the one discriminator site left comparing
in SQL after 28827a7d3 moved the rest in memory. Two ways it could answer
differently from GetBuiltInElementId, each leaving the built-in element
undiscoverable for the life of the install: string equality in SQL is the
provider's collation to decide, so on MySQL's normally case-insensitive default a
case-variant row satisfied the check and the canonical row was never created; and
it ignored Kind, so a row of another kind at the seeded path suppressed the Text
row the lookup resolves.
Ask GetBuiltInElementId instead, so the existence question and the resolution
question are the same code. The wrong-kind half is observable under SQLite and is
now pinned; the collation half is not (BINARY and an ordinal comparison agree on
every input) and stays held by keeping the comparison out of SQL.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Its graphics twin seeds an element, attaches it, then saves with Inherit and an
unknown id, so "the join is empty afterwards" distinguishes a cleared attachment
from one that was never there. The watermark half asserted the same emptiness on
a deco that had no watermarks to begin with -- true of the fixture regardless of
what the handler did, which is a fixture that omits the field it means to test.
Seed a ChannelWatermark, attach it under Override, then save with Disable and
watermarkId 777. Re-measured 2026-09-05 with the ConsumesWatermarkIds guard
removed alone from the committed tree: 1 failed / 4 passed, the failure being
Should_Ignore_An_Unknown_WatermarkId_When_The_Mode_Does_Not_Consume_It. Whole
project green with the guard in place: 2123 passed, 6 skipped, 0 failed.
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two review findings, both about docs the branch already rewrote.
The mutation-coverage table in docs/graphics-elements.md ended up with two rows
numbered 24 -- the new IsOnNowNext row was inserted after 23 without checking
what followed -- while 18 was vacated when the Kind-filter row moved to the
"no red" list. The section's own prose cites rows by number ("the IsOnNowNext
clause (row 10)"), so a duplicate id makes a citation ambiguous. The new row
becomes 33, the next unused number, and the rule that made it 24 in the first
place is now written down: a row number is an identity, not a position, so a new
row takes the next unused number, nothing is renumbered, and a retired clause
leaves its number vacant rather than having it reused under a new meaning. Both
row claims were re-measured and are unchanged; only the id moves.
#568's second half is titled "builtIn discriminator is filename-only,
case-sensitive, folder-agnostic", and the branch removes the first and third
while deliberately keeping case sensitivity -- which reads like two thirds of a
done-when box. It is not: the remedy the same box prescribes, "full seeded
relative path", is exactly as case-sensitive as the filename match it replaces,
so the three adjectives describe one predicate rather than name three separable
demands. Read the other way the box would be unsatisfiable by its own remedy.
The reason case sensitivity is kept -- a case-INsensitive test hands the built-in
identity to a user element differing from the seeded path only in case -- lived
only in GraphicsElementDefaults.cs, where a reader arriving from the issue title
would not find it. It is now in the active record that owns the discriminator.
Refs #568
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Review found the new UpdateDecoHandler FK validators ran unconditionally while
ApplyUpdateRequest reads either id list ONLY under DecoMode.Override or Merge --
under Inherit/Disable it Clear()s the join and ignores the field. So the branch
turned a previously-succeeding save into a 422 over ids that were about to be
discarded, and the SPA reaches that shape: DecosScreen's toReplaceRequest sends
watermarkIds/graphicsElementIds from the draft whatever the mode selector says,
while the picker itself is disabled off-Override. RefreshGraphicsElementsHandler
deletes rows whose template file is gone (cascading the join away), so a stale
editor draft could be locked out of saving a deco back to Inherit, with a 422
naming an element the disabled UI does not even show.
Measured before the fix on the review's E2E instance: PUT /api/v1/decos/1 with
graphicsElementsMode=Inherit and graphicsElementIds=[999] returned 422
"Graphics element(s) do not exist: 999".
The mode predicate is now named once per collection -- ConsumesWatermarkIds /
ConsumesGraphicsElementIds -- and read by both the apply path and its validator,
rather than the apply path holding one copy and the validator implying another.
A second copy is what let the two disagree in the first place.
Two tests pin the gate, one per collection, each reddening when its guard alone
is removed:
Should_Ignore_An_Unknown_GraphicsElementId_When_The_Mode_Does_Not_Consume_It
Should_Ignore_An_Unknown_WatermarkId_When_The_Mode_Does_Not_Consume_It
Measured 2026-09-05, each guard removed alone from the committed tree: 1 failed
/ 4 passed, and the failure is exactly the test named for that guard. Both
assert the apply-path outcome as well as the accept, so a validator that stopped
rejecting for some other reason would not satisfy them.
Refs #568
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The mutation-coverage table's claims were measured against an earlier shape of
GetBuiltInElementId and are re-taken here, because the lookup changed twice on
this branch (filename -> seeded path, then SQL -> in-memory IsOnNowNext) and a
claim about which tests a mutation reddens does not survive either move on its
own.
Measured 2026-09-05, each mutation applied alone to the committed tree:
- Row 10, the seeded-path filter removed: 3 red, not the 2 the row listed.
Ignores_A_Case_Variant_Of_The_Seeded_Path joins the two already named,
because without the filter every Text row resolves as the built-in one.
- Row 24 is new: IsOnNowNext loosened from Ordinal to OrdinalIgnoreCase reddens
exactly the two case-variant tests, 2 failed / 77 passed. One row covers both
discriminator sites because they now share the predicate.
- The Kind==Text filter's "no red" bullet is re-measured across the WHOLE
ErsatzTV.Tests project -- 2121 passed, 6 skipped, 0 failed -- rather than the
11 tests of the one file that names GetBuiltInElementId. ChannelGraphicsDefaults
reaches the lookup from the channel-create handlers as well, so the narrower
population could not have seen a red there. The conclusion is unchanged; what
changes is that it is now measured over the population that could falsify it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The branch moved the `builtIn` discriminator from a bare filename to the full
seeded path, but split how the two sites evaluate it: the API handler compares
in memory (ordinal) while GetBuiltInElementId's new `.Where(e => e.Path ==
OnNowNextSeededPath)` compares in SQL. GraphicsElement.Path takes no explicit
collation -- TvContext.OnModelCreating pins one only on the listed name/title
columns -- so SQLite answers that case-sensitively and MySQL uses the server
default, which is normally case-INsensitive. On MySQL the two discriminators
could therefore disagree about the same row: AttachOnNowNextByDefault would
resolve a case-variant user element as the built-in one while the API reported
builtIn:false for it.
Collapse both onto GraphicsElementDefaults.IsOnNowNext, ordinal, applied in
memory. GetBuiltInElementId goes back to loading the Text candidates and
filtering in memory (the shape it had before this branch), keeping only the
`Kind` enum filter in SQL.
The prose claimed more than the code did. "A filename-only comparison is
case-sensitive-by-accident" appeared in four places as a defect the full-path
fix removed; a full-path comparison is exactly as case-sensitive, so the clause
said nothing and implied a fix that had not happened. Case sensitivity is now
deliberate and stated as such -- the built-in element is the exact file the
seeder wrote, at the exact path it wrote it to -- and the reason the comparison
is kept out of SQL is recorded where the predicate lives.
docs/decisions/records/graphics/channel-level-attachment.md said BuiltIn was
"computed by comparing the row's `Path` to GraphicsElementDefaults.
OnNowNextFileName", which was true of neither the pre-#568 rule (filename to
filename) nor the current one; an active record resolved by key now states the
current predicate in its own sentence rather than in a parenthetical.
Two tests pin the ordinal rule against a loosening to OrdinalIgnoreCase, one
per site. Measured: OrdinalIgnoreCase reddens exactly
GetAllGraphicsElementsForApi_Should_Not_Mark_A_Case_Variant_Of_The_Seeded_Path_As_BuiltIn
and Ignores_A_Case_Variant_Of_The_Seeded_Path, 2 failed / 77 passed of the 79
graphics tests. They do NOT pin provider independence -- under SQLite's BINARY
collation an equivalent SQL comparison answers identically, so no test in this
suite can distinguish the two. That is stated at each site rather than left for
a reader to assume the tests cover it.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Review round on #568 found the branch changed builtIn identity from a bare
filename comparison to the full seeded path (GraphicsElementDefaults.
OnNowNextSeededPath) but left several places still asserting the old rule:
- docs/decisions/records/graphics/channel-level-attachment.md and
on-now-next-on-by-default.md (both status: active) still described a
filename-only match; corrected in place and cross-referenced.
- docs/api-conventions.md §8 quoted the retired
`Path.GetFileName(element.Path) == OnNowNextFileName` expression verbatim;
replaced with the current OnNowNextSeededPath comparison and a note on the
UpdateChannelHandler 422 hardening.
- Three in-code comments (GraphicsElementDefaults.cs, GraphicsElementSeeder.cs,
ChannelGraphicsDefaults.cs) still said "identity is the filename".
- docs/graphics-elements.md's mutation-coverage table (row 10, row 18) named
clauses that no longer exist or no longer redden any test post-#568;
re-measured directly (removing the seeded-path check reddens
Ignores_A_Non_Built_In_Element_With_A_Different_Filename and
Ignores_A_Same_Named_Same_Kind_Element_Outside_The_Seeded_Folder; removing
the Kind==Text filter alone reddens nothing, so it moves to the "known
clauses with no red" list with that measurement dated).
Also closed the should-fix twin: UpdateDecoHandler's graphicsElementIds and
watermarkIds are top-level ReplaceDecoRequest fields in the same position as
UpdateChannelRequest.graphicsElementIds (not the deep-FK-in-a-nested-list
carve-out), and the reconcile in ApplyUpdateRequest blindly Added a join row
for any incoming id -- the identical FK-constraint-to-500 defect #568 fixed
on the channel path. Added GraphicsElementIdsMustExist/WatermarkIdsMustExist
validators mirroring UpdateChannelHandler's, pinned by
UpdateDecoGraphicsElementsTests (reddens when either validator alone is
removed -- verified).
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
UpdateChannelHandler.Validate never checked incoming graphicsElementIds against
GraphicsElements, so PUT /api/v1/channels/{id} with a non-existent id hit
FK_ChannelGraphicsElement_GraphicsElement_GraphicsElementId at SaveChangesAsync
and surfaced as an unhandled 500. Add GraphicsElementIdsMustExist, following the
existing FFmpegProfileMustExist/WatermarkMustExist/FillerPresetMustExist shape,
so an unknown id now returns 422 for parity with every other FK field on this
full-replace DTO.
GetAllGraphicsElementsForApiHandler and GraphicsElementSeeder.GetBuiltInElementId
keyed builtIn off Path.GetFileName(e.Path) == OnNowNextFileName -- folder-agnostic,
so a user element named exactly on-now-next.yml in any other template folder would
also report builtIn:true. Both now compare against
GraphicsElementDefaults.OnNowNextSeededPath, the full path the seeder actually
writes to.
Follow-up from the #74 whole-branch review (2026-07-22), deferred as
data-safe/not SPA-reachable.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The thesis measurement quoted three numbers side by side. Two are verbatim from the
corpus — `docs/guard-inventory.md:179` "wrong NINE times", and
`scripts/tests/test_image_build_delegates_the_spa_suite.py:772` "defeated seven measured
ways". The third, "a lexical rule over a hook preamble **five**", was reachable only from
#901's own issue body ("Five spellings, one mechanism"); nothing in the repo re-derives
it, and both artifacts the record cites for #891 —
`docs/decisions/records/process/hook-resolves-inputs-from-repo-root.md:64-68` and
`scripts/tests/test_hook_fire_log.py:144` — count the SAME sequence as three ("Three
successive lexical rules over this line each fell"). Counted as spellings it is eight or
nine; five is neither basis. A number in prose that no artifact re-derives is this
record's own subject matter, and "Measured 2026-08-30" invites trust rather than
re-derivation.
Answered by subtraction, not by new prose:
- body: the third clause states what both cited artifacts state — three successive
lexical rules, each defeated by the next shape.
- `rule:` drops the copied numeric triple ("nine, seven and five times") for a pointer to
the sequences below, so the counts live in one place (`dont-keep-a-copy-of-a-set`).
- `signals:` swaps the unsourced token for the sourced one.
Same class, found while checking the neighbours: "would have licensed the parser above
through most of nine rounds" (`rule:` and body) counted #887's NINE DEFECTS as rounds —
guard-inventory records them as nine defects across THREE cold-review rounds. Now "most
of those nine defects".
Body stays 58 lines; the paragraph is reflowed at the file's existing width.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`dc3f158bc` corrected a false universal ("this arm's errors are refusals, never
acceptances") and, in the same sentence, appended a second unmeasured
fail-direction claim: "and the `split(\"\\n\")` residual below is the second
acceptance of the same kind". That clause inverts the paragraph it points at.
MEASURED, by loading the module and calling its own primitives with
`CANONICAL_SINK_ASSIGNMENT + \x0c + "rm -rf /tmp/nothing"` on one physical line
followed by `CANONICAL_SINK_SOURCE`:
instrumentation_faults(...) -> "the sink preamble is not
the canonical two lines"
splitlines() selection == [CANON, SOURCE] -> True (ACCEPTS)
split("\n") selection == [CANON, SOURCE] -> False (REFUSES)
The shipped `split("\n")` refuses exactly where the rejected `splitlines()`
accepts, which is what the pre-existing paragraph six lines below already says.
There is no second acceptance below; the acceptance belongs to the alternative
that was NOT shipped. `\x0c` is not a line terminator for bash either, so
`split("\n")` matches the shell the checker models and has no hole of this kind.
The FIRST half of the sentence stands and was re-measured: a hook carrying the
two canonical lines plus `v="ETV_HOOK_FIRE""_LIB=/tmp/evil.sh"; eval "$v"`
produces no preamble fault, so the selector really does accept a writer that
never spells the literal. Answered by SUBTRACTION per the brief: the sentence
ends at "not all refusals." and no replacement prose is written.
`dc3f158bc`'s message argues from the same inversion ("the comment already
conceded an acceptance hole of the same kind six lines later"); that half of its
reasoning is withdrawn here. Its correction of the universal is unaffected — the
selector escape it names is structural and independently measured, above.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round seven's blocker was the record's own thesis sentence, and the fix is to
delete rather than to re-argue. Six rounds each replaced a refuted absolute with a
fresh one; this one states only what was measured.
THE THESIS (body). Struck: "a pin cannot be defeated by a respelling of what it
COMPARES" and "nothing [stays exposed] for a whole-file pin". Both are false for
the record's own worked whole-file pin — measured against the guard's own
primitives:
_normalise_lines("const flag = '--run --reporter=x';")
== _normalise_lines("const flag = '--run --reporter=x';") -> True
_normalise('RUN npm ci && echo "a b"')
== _normalise('RUN npm ci && echo "a b"') -> True
Both pairs differ in bytes and both are respellings of what the pin compares.
`PINNED_VITE_CONFIG` — the pin the record calls "pinned WHOLE" — is compared
through `_normalise_lines`, so its normalisation is a second exposure axis beside
the selector's. The record already refuted itself twice: `rule:` ends "A pin also
declares its NORMALISATION and what the normalisation cannot see", and `mechanics:`
says the whitespace collapse "including inside a QUOTED STRING" belongs to both
TEXT pins. What replaces the sentence is the fail DIRECTION alone — a pin's is a
false RED, a shape-matcher's a false GREEN — plus the declaration obligation the
rule already carries. No new universal is written in its place.
MECHANICS. "named once so the checker, the mutation proofs and the hooks cannot
come to mean different strings" is deleted, not repaired: the assignment string is
written out at THREE sites in `test_hook_fire_log.py` (measured by walking the AST
and comparing each assembled string to `CANONICAL_SINK_ASSIGNMENT` — the constant,
and the `current` local of `test_an_ENV_VAR_resolved_sink_path_is_DETECTED` and of
`test_the_NEXT_env_var_to_be_invented_is_DETECTED`). The source comment making the
same claim is corrected in place, and its correction is STRUCTURAL: it names the
three sites and the `current in text` assertion each proof carries, and asserts no
mutation outcome, because "an edit here faults loudly there" would be a `CLAIMS`
entry under `testing.mutation-claims-are-executed` — wherever it is written — or it
is not written. Same reason `4d5bd0dbb` removed the outcome claim from `mechanics:`
rather than binding it.
THE QUALIFYING-GRAMMAR LIST now exists once, in the record's `rule:`.
`docs/README.md` and `docs/guard-inventory.md` state the operative test — an
artifact with a grammar the predicate does not implement — and point at the record.
The copies had already disagreed inside the commit that wrote them:
`docs/README.md` carried five of the six members, omitting JSON5, which is the
member the issue's own correction comment names as the one the narrow "shell or
config TEXT" framing would have let through (`dont-keep-a-copy-of-a-set`, #869).
Rider 1 drops its copy of the vite `DEFAULT_CONFIG_FILES` ordering the same way,
deferring to the dated reading in the `test_image_build_delegates_the_spa_suite.py`
row.
Gate: `pytest scripts/tests -q` 1599 passed, 3 skipped (all pre-existing by-design
skips); `decisions_validate.py` OK with the record off the >60-line list; catalog
regenerated with no diff; ruff clean; `check-doc-narrative --diff origin/main` 0
warnings.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`scripts/tests/test_hook_fire_log.py:154` ended "this arm's errors are refusals,
never acceptances". That is false for the SELECTION, which is the half this branch's
record is about: `mentions` is chosen by the literal `ETV_HOOK_FIRE_LIB`, so a writer of
that variable which never spells the literal is outside the compared set altogether, and
the checker never sees it. The comment already conceded an acceptance hole of the same
kind six lines later, about `split("\n")` versus `splitlines()`, so the universal was
contradicted inside its own block.
The correction stays STRUCTURAL — it reads the selector one line below and says what the
selector reaches — and makes no claim about what a mutated hook would return. An outcome
claim would be a `CLAIMS` entry in `scripts/tests/mutation_manifest.py` or nothing, per
`testing.mutation-claims-are-executed` as amended by #881, which is the same reason
`mechanics:` states the selector rather than a measured result (`4d5bd0dbb`).
WHY THIS LANDS NOW: #885 closed while this branch was in review (merged as #919,
`c30847204`), which frees `scripts/tests/`. Of the two obligations `47619320e` recorded
as owed, that message is superseded here:
- DISCHARGED: this one, inline, above.
- NOT OWED, and the reason is not the blocker: the `CLAIMS` entry for the hook-preamble
selector. `4d5bd0dbb` removed the outcome claim from `mechanics:` rather than binding
it, so the record asserts no mutation outcome and the rule it invokes has nothing to
bind. Re-adding a claim in order to bind it would reverse a review-mandated change; an
executed GREEN entry (target `.claude/hooks/decisions-guard.sh`, clause = the canonical
sink assignment, replacement = that line plus the `eval` spelling, plus the mandatory
`reach_replacement`/`reach_expect`) remains available as an ENRICHMENT of the structural
fact, and belongs to whoever wants the fact executed rather than argued.
Rebased onto `366a0f904..c30847204` on the way: the `docs/guard-inventory.md` conflict is
two rows, resolved by taking #885's newer `test_workflow_persist_credentials.py` row
(it gained a second invariant) and this branch's `test_image_build_delegates_the_spa_suite.py`
row (it dates the vite `DEFAULT_CONFIG_FILES` reading to 8.1.3 and to the lockfile).
`docs/decisions/README.md` regenerated, not merged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`798e5ee40` stopped at 65 prose lines and gave a reason: the remaining candidates for
removal were the *why* behind non-obvious choices, which `docs.no-session-narrative`
says to keep. That reason is sound about the WORDS and wrong about the LINE COUNT,
because the two are not the same quantity. `decisions_validate.record_prose_lines` counts
PHYSICAL lines, and this record was the narrowest thing in the set being measured:
max body width, 222 active records: median 106; 28 at <=100, 154 at 101-120, 40 >120
this record: 100. The sibling it cites by key, process.hook-resolves-inputs-from-repo-root: 116.
The 60-line ceiling was derived at #620 from that distribution, so measuring a
100-column record against it charges the record for a wrap width the corpus does not use.
Re-wrapping the seven body paragraphs at 116 — the exact width of the neighbour record —
takes the body from 66 physical lines to 58, and the validator now reports 58, off the
over-ceiling list (50 records over -> 49, and the key no longer appears).
The reflow removes NOTHING: the script asserted word count equal before and after (893)
and whitespace-normalised body text byte-identical, and refused to write otherwise. What
it buys is that the issue's `## Done-when` box "The record is under the 60-line advisory
prose ceiling" is satisfiable as written, so `pretooluse-merge-consent.sh` is not asked
to derive consent from a box ticked falsely or left standing. The advisory itself was
never breached — the ceiling is a `::warning::`, the validator exits 0, and 49 of 222
records are over it inside the 2-25% CEILING_MINORITY band.
Also: the file was the only one of 222 records with no final newline. Fixed in the same
commit; `record_prose_lines` is `splitlines()`, so it does not move the count.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`mechanics:` opened with a universal — "a residual read off one of them does not
transfer to the others" — and then attributed the collapse-whitespace-inside-a-string
residual to `_normalise_lines` alone. It does transfer, between exactly the two pins
that sentence separates: `_normalise_lines` is `_normalise` applied per line
(`scripts/tests/test_image_build_delegates_the_spa_suite.py:327`), and the stage-command
pin calls `_normalise` directly (lines 386/389, compared at line 586), so it cannot see a
whitespace change inside a quoted shell string either. For SHELL text that is the more
consequential of the two residuals, which is the opposite of what the old ordering
implied.
The sentence now says NEED NOT transfer, names the one that does, and puts the
string-literal blindness on `_normalise` where it originates. The vite-only fact that
survives is the blank-line drop, and the claim that the test STATES its residual is
narrowed to the vite test, which is the only one of the two that does.
The hole is dormant rather than live — no entry in `PINNED_STAGE_COMMANDS` carries a
quote character — but the defect was the prose universal, which the record's own rule
("a pin also declares its NORMALISATION and what the normalisation cannot see") is what
this paragraph exists to demonstrate. This is the third finding read off this one
sentence: `b8dc321af` corrected its fault-message half and `4d5bd0dbb` its
outcome-claim half.
TWO OBLIGATIONS ARE OWED to the closing record, both blocked on #885 (open, so this
branch does not touch `scripts/tests/`), and both freed together when it closes:
1. The `CLAIMS` entry in `scripts/tests/mutation_manifest.py` for the hook-preamble
selector — target `.claude/hooks/decisions-guard.sh`, clause = the canonical sink
assignment, replacement = that line plus the `eval` spelling, proof = a
`test_hook_fire_log.py` node, outcome=GREEN with the mandatory
`reach_replacement`/`reach_expect`. Until then `mechanics:` states the SELECTOR as a
structural fact and makes no outcome claim (`4d5bd0dbb`).
2. The comment at `scripts/tests/test_hook_fire_log.py:153-154`, which ends "this arm's
errors are refusals, never acceptances". The structural fact this record ships — the
compared set is the non-comment-led lines containing the literal `ETV_HOOK_FIRE_LIB`,
so a reassignment that never spells the literal is outside the selection — is an
ACCEPTANCE by that arm, and the same comment block concedes the class five lines later
("so it is an acceptance hole, not only a stricter refusal"). The sentence is owed a
correction; the record documents the residual beside it in the meantime.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The two corrections added four prose lines to a record already one line over
the 60-line advisory ceiling. Recover what can be recovered without losing
substance: reflow the paragraphs, drop the padding ("and it fails silently"
→ ", silently"; "the count rises" → "and rises"), and cut one restatement.
It lands at 65 lines, not 60. That is a deliberate stop: the remaining
candidates are the *why* behind non-obvious choices — which the repo's own
docs rule says to keep — and `decisions_validate` reports the constant itself
as drifted from the distribution it is supposed to mark the tail of (p90=104,
p95=142, 50 of 222 records over it). The validator passes.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two prose overclaims in the new record, both of the shape the record itself
exists to police.
The thesis sentence generalised over the two worked pins and held for one.
`PINNED_VITE_CONFIG` compares a whole file, so no respelling escapes it; the
hook-preamble pin compares a SELECTION (`test_hook_fire_log.py:161` keeps only
the lines containing the literal `ETV_HOOK_FIRE_LIB` that are not comment-led),
so it must recognise a line before it can reject it — which is exactly the
residual `mechanics:` documents four lines later. Bound the immunity to what a
pin COMPARES and name the leftover exposure as the selector's reach.
The neighbour citation claimed `docs/defect-shapes-773.md` §4 "argues the
general form". §4 is a ranked table of detectors A-G — none of them this rule,
and A is already assigned to `guard-derives-population-from-source` by the
preceding clause. The doc's nearest class is `string-predicate churn` in the
§3.6 partition, marked `no detector proposed`, and §3.7 argues the class away as
a cross-cutting property (2 of 33 round-churn records). Only §4's closing
meta-finding — class-level rules beat one record per instance — supports
anything here, and it supports the FORM, not the content. Say that.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round three, blocking. The record's `mechanics:` field asserted a measured
mutation outcome — that appending an `eval` which composes `ETV_HOOK_FIRE_LIB`
at runtime leaves `instrumentation_faults` returning `[]`, identical to the
unmutated baseline — with no `CLAIMS` entry in scripts/tests/mutation_manifest.py.
`testing.mutation-claims-are-executed` as amended by #881 puts a prose claim
about a mutation's outcome under the executed-claim rule wherever it is written,
a decision record included, and this branch's own docs/README.md row restates
that. A record whose rule text requires "each defeat the matcher claims to catch
is a DECLARED, executed mutation" cannot itself carry an undeclared one.
scripts/tests/ is held by #885, which is open, so the entry cannot be added
here. What replaces the outcome claim is the structural fact that carries the
same point and needs no execution: the compared set is the lines containing the
literal `ETV_HOOK_FIRE_LIB` that are not comment-led, so the pin reaches exactly
the two preamble lines and a later reassignment which never spells the literal
is outside the selection — whatever the checker then returns. The `CLAIMS` entry
is owed once scripts/tests/ is free.
Three more from the same round:
- `_normalise_lines` was a universal over three pins that holds for one. The
stage commands compare through `_normalise` (continuations joined, whitespace
within one command collapsed); the script map is dict equality over parsed
JSON and compares no text; only PINNED_VITE_CONFIG uses `_normalise_lines`.
The three are now stated separately, with the note that a residual read off
one does not transfer.
- docs/guard-inventory.md's pointer restated the "shell or config TEXT" framing
the record exists to reject. It now says what the record says: an artifact
with a grammar the predicate does not implement.
- Rider 1 stated vite's DEFAULT_CONFIG_FILES ordering unbound to a version. The
ordering belongs to a release and expires with one, so the record cites the
guard-inventory row rather than keeping a second copy, and that row now dates
the reading and names the release web/package-lock.json pins.
Refs #901, #891, #887, #881, #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The record's own rule ends: a pin also declares its NORMALISATION and what the
normalisation cannot see, because "pinned whole" invites a reader to assume
byte equality. The `mechanics:` field did that for the image-build pin (it
names `_normalise_lines` and what it drops) and not for the hook pin, which it
described only as pinning two lines byte for byte.
The comparison at scripts/tests/test_hook_fire_log.py:161-167 is over a
SELECTED set — lines containing the literal `ETV_HOOK_FIRE_LIB` that are not
comment-led — so byte-identity holds within the selection and says nothing
about a writer of that variable spelled without the literal.
MEASURED 2026-09-05 on this branch against `.claude/hooks/decisions-guard.sh`:
appending
eval "$(printf %s%s=/dev/null ETV_HOOK_FIRE _LIB)"
after the canonical assignment leaves `instrumentation_faults(text,
'decisions-guard')` == `[]`, byte-identical to the unmutated baseline `[]`,
while running those two lines under bash prints `final=/dev/null` — the sink is
repointed and the checker is silent. A false GREEN, which is the failure
direction this record exists to argue about.
The residual is #891 code and is not introduced here; no code changes. What
changes is that a record citing this file as one of its two worked pins now
states the residual instead of implying byte equality over the file.
Refs #901, #891
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The record is a rule about the SHAPE of a guard's predicate — the third member
of the set docs/guard-inventory.md:15-16 tells a guard author to read before
editing a guard or adding a row — and that pointer named only
`guard-derives-population-from-source` and `guard-ships-with-mutation-proof`.
Before this commit, `grep -rln guard-pins-the-artifact-not-a-shape docs/`
outside the record and the generated catalog returned docs/README.md alone, so
the record was reachable from the task-signal map and by topic but not from the
inventory a guard author already has open.
That is the reachability failure the record itself names: its body says
`docs/guard-inventory.md` carries its precedent per incident, findable only
from inside one. A record about topic-resolvability that is missing from the
entry point of its own topic reproduces it.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two review findings, both the class this record is about: prose that claims
more than the artifact says.
The neighbour-record citation presented a paraphrase in quotation marks.
`testing.guard-derives-population-from-source` line 96-97 reads "when the
authoritative source is missing, the answer is to create one, never to
approximate it with a predicate over text"; the record quoted it without "the
answer is to" and without the second "to", so a fixed-string lookup of the
quoted span found nothing anywhere under docs/ — #812's second defect, which
`testing.mutation-claims-are-executed` names explicitly. The quote marks are
gone rather than repaired: the source sentence spans a line break, so any
single-line verbatim quotation of it would still not resolve by grep, and an
open paraphrase claims only what it is.
The `mechanics:` field said the image-build pins fail with a message naming
the constant to update. Measured against
scripts/tests/test_image_build_delegates_the_spa_suite.py: PINNED_STAGE_COMMANDS
(line 589) and PINNED_VITE_CONFIG (line 798) name themselves,
PINNED_PACKAGE_SCRIPTS (lines 757-763) does not — it names the FILE and prints
both maps. 423bf94e7 corrected this same sentence for the hook half after
verifying it and left the image-build half an unverified universal.
The first replacement drafted here read "all four faults ask for the reason in
the same commit", which is false a second way: the hook-preamble fault
(test_hook_fire_log.py lines 162-167) asks for no reason at all, it reports
got={mentions}. The shipped sentence is scoped to the three pins whose fault
messages were read.
Refs #901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The catalog renders `rule:` alone, so a reader resolving this by topic saw the two
decisions without the alternatives they rejected — which is what stops a rejected
option being re-proposed on plausibility.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The record asserted both worked pins fail with a message naming the constant to update
and asking for a reason. True of `test_image_build_delegates_the_spa_suite.py`'s three
pins; false of `test_hook_fire_log.py`, whose byte-identity fault reports the divergent
`mentions` list and names no constant. Verified against both files rather than inferred
from the neighbouring one.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two predicates over artifacts with a real grammar were each defeated by successive
spellings and withdrawn in favour of pinning the artifact whole: #887's shell parse
(nine defects from one mechanism, then seven more against a partial match of
web/vite.config.ts) and #891's lexical rule over the hook preamble (five spellings).
Both incidents carry a record; neither is resolvable by topic before round three, which
is what this class-level record adds.
Decides the two questions #901 left open:
- DEFAULT, not remedy. A shape-matcher's failure is a false GREEN, so the defeat that
would trigger a remedy policy is found by a reviewer or an incident and never by the
guard: "not defeated yet" measures who has looked. Rejected: write the matcher and
pin after the first defeat — it also understates its bill, since a withdrawal costs
the rounds spent AND the proofs calibrated against the narrow clause.
- The exception argument carries FOUR things: the grammar and its parser; the input
space as a closed enumeration with the reason it is closed; the fail direction
measured as a declared, executed mutation; and what it buys priced in a cost the pin
charges. Rejected: a numeric "survives N spellings" bar (measures the reviewer's
imagination) and a reviewer sign-off bar (depends on the signal that arrives late).
Records both riders (a pin assumes it pins the artifact that still DECIDES; widening a
clause turns a survived-clause canary into a tautology) and states the threshold as the
moment the NEXT spelling is found by the reviewer rather than the author.
docs/README.md's guard-convention task-signal row points at the record; catalog
regenerated; 59 prose lines, under the 60-line advisory ceiling.
fixes#901
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round-nine's bare-if fix table paired each secret name with only one level
(job for REGISTRY_PASSWORD, step for RENOVATE_TOKEN), so the exact
REGISTRY_PASSWORD-at-step-level and RENOVATE_TOKEN-at-job-level cases the
finding named were never driven. All four combinations now run.
The decision record's anonymous-layer-download closure read as reporting a
past run without saying who ran it. Re-measured directly this session
(2026-09-05, no stored credential): anonymous token -> pinned manifest's
first layer -> 200/32991280 bytes, same GET with no token -> 401. Record
updated to say the leg was re-confirmed, not merely "measured...since".
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The block claimed the ordinary-string read — `secret_refs`, how `secret_name_counts` routed
an `if:` value before `condition_refs` — "is asserted empty on every row", while asserting it
on the job-level condition only. The step-level row's own string went unchecked, so a
predecessor that happened to see it would have left the row proving nothing. Both rows now
run the same three assertions from one loop.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The `if:` clause reads the one site whose reference the evaluator resolves without
materialising anything into the job environment, so a reader can reasonably ask why it
faults. Both the function and the record now say: the predicate is "names a stored secret",
never "exports one" — on the head-authored route the contributor picks the comparison, which
makes a condition an oracle over the value, and a predicate about exposure would have to
model what each site does with its reference and give up the structure-blindness that saw
`toolchain-preflight`'s step `env:` when a `container:`-shaped predicate did not.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`scripts/ci-toolchain-image-resolves.sh` and `docs/ci-cd.md` both stated that the #772
container jobs "died after 1-2s", and the header used the same number to argue the preflight
needs no `needs:` gate. Nobody measured it, and it cannot be measured from a working session
without reproducing a deleted-tag incident. What the number stood for is structural and IS
known: a container job that cannot pull its image fails AT the pull, before it runs a step,
so it wastes no work waiting to be told and the argument against serialising the five jobs
survives intact.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two corrections to `ci.pr-route-carries-no-stored-credential`, both about claims that read
as checked and were not.
The `rule:` said the detector reads every spelling "only inside a `${{ }}` span", and the
body enumerated `secrets: inherit` as the ONE shape left uncovered. An unwrapped `if:` is a
second, and it is a shape this repo writes: both now name it, and say the value of an `if:`
is read whole.
`mechanics:` listed an anonymous LAYER download among two things the daemon probe did not
exercise. Measured 2026-09-05 from a workstation holding no registry credential: the
anonymous pull token reads the pinned manifest's first layer
`sha256:179c68a720750ab4d354f6b55c0a9f551d4fd7bde93606dd0be79ba16493a39e` -> HTTP 200,
32991280 bytes, and the same GET with no token -> 401. act_runner's own pull call path is
the one leg still unexercised.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The `${{ }}` span scoping added in a6225e4ee was correct about prose and wrong about one
real spelling: `if:` is the only key the expression grammar lets omit the delimiters in, so
`if: secrets.REGISTRY_PASSWORD != ''` named a stored secret in a document holding no `${{`
at all, and the collector reported it clean. Measured on the previous head e35e1b772:
`secret_refs("secrets.REGISTRY_PASSWORD != ''")` -> `[]`, and the same string as a
job-level or step-level `if:` on a synthetic `pull_request` job -> `stored_secret_faults(...)
== []`. That an unwrapped condition is evaluated is not inferred — `docker-build.yml`'s own
`build` job carries `if: github.event_name != 'pull_request'` bare, and `PR_EXCLUDING_IFS`
pins that exact string.
`condition_refs` reads an `if:` value as one span with the delimiters neutralised to a
SPACE (deleting them collapses `${{ secrets.A }}${{ secrets.B }}` into the single identifier
`secrets.Asecrets`, losing a reference), and `secret_name_counts` routes the value there
instead of onto the stack, so a wrapped condition still counts once. Everywhere else the
scoping stands and the English `# We pass no secrets. Then …` still costs nothing.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The code banner and the test docstring say that whether act_runner on this
instance resolves `workflow_call` + `secrets: inherit` was not probed, and why
that is acceptable — it governs reachability today, not the guard's silence. The
record stated the clause without that bound, so a reader who meets the rule
through the catalog rather than through the file met a confidence claim the
source deliberately does not make.
Refs #885
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The detector was exhaustive over the `${{ }}` expression grammar and blind to
`jobs.<id>.secrets: inherit` on a `uses:` job, which passes the caller's whole
store to the called workflow while naming nothing. `secret_refs` reads only
inside expression spans — correctly, since outside one `secrets.` is a full stop
— and `inherit` is a plain scalar, so such a job was put in the derived
population by `pull_request_jobs`, walked, and reported CLEAN. Measured against
the predecessor:
stored_secret_faults('synthetic.yml', {True: {'pull_request': None},
'jobs': {'reused': {'uses': './.gitea/workflows/reusable.yml',
'secrets': 'inherit'}}}) -> []
... the same job with secrets: {TOK: '${{ secrets.RENOVATE_TOKEN }}'} -> 1 fault
so the miss was specific to the VALUE SHAPE, not the key. That is the failure the
done-condition names — a new job joining the population unprotected without
reddening anything — in a guard whose stated selling point is exhaustiveness over
the grammar and no exemption list.
`opaque_secret_handovers` now faults a `secrets:` key whose value is not a mapping
of names, under the existing `secrets.*` whole-context sentinel, and both fault
sites read through one `held_secret_names` so the workflow scope and the job
subtree cannot drift on which references are forgiven. The test is on the value
shape and not on the word `inherit`, for the reason the residue counter is not a
match on `toJSON`: any non-mapping value hands over a set the guard cannot
enumerate, a spelling act_runner grows later included.
Both halves of the predecessor measurement are re-derived every run rather than
left as prose: the new test asserts `secret_names(job) - INJECTED_SECRETS` — the
collector verbatim as it read before this clause — empty on the same fixtures it
asserts the fault on, and asserts the job is in the population. Reverting
`held_secret_names` to that expression reddens that test and only that test
(measured: 1 failed, 18 passed).
The clause reads the DOCUMENT only and the text-versus-walk cross-check cannot
cover it — there is no expression for its half to match, which is a stronger
reason than the shared-blind-spot one the cross-check already discloses. Said at
the definition, in the cross-check's "STRUCTURALLY CANNOT REPORT" paragraph, and
in the record, rather than left to be discovered; it does not redden the
cross-check either, since the clause feeds the fault collector and not
`secret_name_counts`.
Whether Gitea 1.27.1 / act_runner resolves `workflow_call` + `secrets: inherit`
on this instance was NOT probed — that affects reachability today, not the
guard's silence, and the direction is the one the spelling rows already take.
No tracked workflow uses a job-level `uses:`, so nothing reddens.
Refs #885
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`release.main-direct-push-disabled` was edited on this branch to note the new `v*`
tag protection, and picked up two of the defects the round was hunting elsewhere.
Its `rule:` said the `renovate` bot "can no longer push a tag that publishes
`:prod`" as settled fact, while `release.tag-protection-v-star` records that exact
claim as NOT VERIFIED and `docs/ci-cd.md` was already corrected to EXPECTED,
UNVERIFIED. Only the `timothy` credential exists in a working session, so neither a
real release cut nor a refused bot push has been exercised; all three now agree on
confidence.
Its `mechanics:` still read "`GET .../tag_protections` returns `[]`" in the present
tense — the one fact this branch changed, and the one site an otherwise complete
sweep left behind. Read back live today the endpoint returns one rule, `v*`
whitelisted to `timothy`. The clause is now past tense and bound to its probe date,
with the current state named.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The two `docs/guard-inventory.md` rows and the module docstring described
`test_workflow_persist_credentials.py` as the `actions/checkout` guard only. The
deferral to #909 rested on `docs/guard-inventory.md` being held by the session
working #881; that issue is closed and its PR is the commit this branch is rebased
onto, so the file is free and the edit belongs here under docs-update-is-part-of-done.
`MUTATIONS` keys at most one declared clause mutation per guard FILE
(`test_the_manifest_covers_exactly_the_MUTATION_rows` asserts `len(MUTATIONS) ==
len(declared)`), and the grading row's proof-ref column is compared against it, so
the route invariant cannot take a second `MUTATIONS` row. It takes a `CLAIMS` entry
instead — the population #881 widened this file to carry — bound to the inventory
sentence that states it: deleting `build`'s `if: github.event_name != 'pull_request'`
from the shipped `docker-build.yml` is applied to a sandbox copy every run and the
named proof is required to redden with the collector's own wording.
That grows the `CLAIMS` population from three entries to four, which invalidates the
cost span `testing.mutation-claims-are-executed` measured over three. Re-taking it
here produced 54.3s/149.5s, 81.6s/78.7s and 114.3s/84.2s across three A/B pairs with
other builds on the host — two inverted, so the load dominates the signal. The record
now says the span is a lower bound and that a re-measurement is owed on a quiet
machine, rather than carrying a scaled or invented number.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
"unaffected for the release operator" asserted the outcome the rest of the
paragraph then marks unverified. Both places now say what is intended and what is
measured, and the blockquote is rewrapped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`pattern="$pattern[$char…]"` is SC1087 — shellcheck reads `$pattern[` as an array
expansion and errors out. It concatenates correctly here because `pattern` is a
plain string, so this is a lint stop rather than a runtime defect; braced, the
character class is unambiguous to reader and linter alike.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`review-verdict.yml`'s residual list said the injected `GITEA_TOKEN` on the
`pull_request` route is BOUNDED by `docker-build.yml`'s workflow-level
`permissions: code: read`. That block lives in the head-supplied file on exactly
that route: a PR author deletes it, and with the owner-level Actions default at
`permissive` that alone yields a write-capable token. It is NARROWED for the
committed file, and it stays in the residual set the paragraph exists to enumerate
— which is what `release.verdict-status-check` and `test_pr_changed_files.py`
already say. The same reword lands in `ci.pr-route-carries-no-stored-credential`,
where the allow-list reason is now the store the token is not in rather than a
bound.
The "dies at image pull in 1-2s" figure was never measured on this branch — the
1-2s in `ci-toolchain-image-resolves.sh`'s header is an observation from the #772
incident, not a property of this change. The loud/silent asymmetry is what carries
the argument, so the claim is now that a container job dies at image pull before it
runs a step, which is true by construction.
`ci.actions-credential-scoping`'s reworded `mechanics:` said "all three are now
confined to the `build` job". `build` declares no `container:` at all; the
buildcache write and the base-image pull are what it confines, and the `container:`
pull is credential-free everywhere.
`docs/ci-cd.md` asserted the `renovate` bot can no longer push a `v*` tag while
`release.tag-protection-v-star` records that as NOT VERIFIED. The rule is read back
live and real; what is unmeasured is Gitea honouring it against an account only the
operator can test. Both docs now say expected, unverified.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Header field names (RFC 9110 §5.1) and auth-param names (RFC 7235 §2.1) are both
case-insensitive, so `WWW-AUTHENTICATE: Bearer REALM="…"` is the same challenge this
registry sends in mixed case today. The preflight matched the header name in a fixed
case for all but four letters and the directive name in lowercase only, so that
spelling fell into the "named no realm" arm: the job fails — the safe direction —
but names a cause that is not the real one and points an operator at a token
endpoint that is healthy.
The header line is now selected by an `awk` comparison on the lowercased field name,
which leaves the value's case alone (a realm URL is case-sensitive), and the
directive name is matched through a character class generated from the key. The new
test drives the whole anonymous read end to end against an all-caps challenge rather
than testing the parser, so the token leg and the authenticated re-read both have to
survive the spelling.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`_SECRET_REF` ran over the whole string while the comment above it claimed every
pattern was confined to `${{ }}` spans. Executed against the shipped module,
`secret_refs("# We deliberately pass no secrets. Then the pull is anonymous.")`
returned `['Then']`, and fed through the real collector that is one fault reading
"job `j` names stored secret(s) on the pull_request route: Then" — a fabricated
name, on a PR-route job, for its own comment. The same comment separately reddened
the text-versus-walk cross-check, because the line-level strip removes a `#` line
from the text half only.
The existing negative control passed for a reason that does not generalise: no `.`
follows the word in `"no secrets are used here"`. A sentence ENDING in "secrets."
is the likeliest thing to be written into a PR-route `run:` block on this branch's
own subject, so the trap was self-inflicted.
`secret_refs` now resolves names per `${{ }}` span, so every spelling is scoped the
way the residue counter already was. The added rows drive the real predecessor —
`_SECRET_REF` applied to the whole string — and assert it read a name where the
scoped reader reads none, so reverting the scoping reddens them.
The `INJECTED_SECRETS` comment stops calling the injected `GITEA_TOKEN` "bounded by
the workflow's own `permissions:`": on this route the head supplies that file and
can delete the block. Allow-listing it is a claim about the store it is not in, not
about a bound.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The same-subject sweep that past-tensed `ETV_STATUS_AUTH` and this record's own
`rule:` field left the record BODY saying `toolchain-preflight` takes the
registry credential "via `ETV_REGISTRY_AUTH`" in the present tense — a symbol
this branch removes from every workflow, so the body contradicted the `rule:`
field of the same record. Container-free and `runs-on: small` are still true
today and stay in the present tense; only the credential clause moves to the
past, matching the `rule:` field's "took the credential through
`ETV_REGISTRY_AUTH`".
Body-only, so the generated catalog is unchanged (regenerated to confirm).
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The header comment and the record's `rule:` said the registry reads and the
commit-status reads "both depend on `timothy/ersatztv` and its `ersatztv-ci`
package staying PUBLIC; making either private fails those jobs loudly at image
pull, never silently". Wrong in both clauses for the repo half, and this branch
has already been sent back twice for exactly this shape of mechanism claim.
Measured 2026-09-05: the `ersatztv-ci` package is linked to no repository (every
version reports `"repository": null`), so the repo's visibility does not gate the
anonymous pull token at all; and the only thing it does gate — the combined-status
GET — fails in the opposite direction, because `ci-detect-already-validated.sh`
answers a failed `curl -sf` with `emit false; exit 0`. That job stays GREEN and
the #420 cross-run skip silently stops firing. So the two dependencies are now
stated apart, each with its own failure direction, in `docker-build.yml`, in the
preflight's header, in the record and in the `ci-cd.md` outcome table; the
preflight's own 401/403 messages stop sending an operator to the repo's
visibility when it is the package's.
`token_leg_done` was set once per RUN, before the attempt, so a token endpoint
that could not be reached failed the preflight with no retry while an identical
blip on the manifest read got three. The stated reason — "a registry genuinely
refusing anonymous reads is asked once rather than once per pin" — is a per-pin
argument that never covered the per-attempt axis. It is now sorted by what the
endpoint SAID rather than by which leg it happened on: an answer (no token in the
body, a challenge naming no realm, no challenge at all) settles the question and
is asked once per run; an endpoint that could not be reached, or that answered
5xx, settled nothing and is retried on the same `ETV_CI_ATTEMPTS` budget as the
manifest read, because a red here denies a merge (the consent hook reads the
COMBINED status, #598) and the two legs of one read must not have opposite flake
tolerances. The token-leg message now reports the attempts it actually made.
Driven against the SHIPPED predecessor rather than a hand-written mutant: the
three new behavioural assertions are red on it (1 token call where 3 are
required, and a blip shorter than the budget failing the run), while the two that
pin the property the retry must not cost pass on both.
Refs #885
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`ETV_STATUS_AUTH` is gone from `test`, `migrations` and `functional-e2e`, and three
same-subject sites were reworded to match. `review-verdict.yml`'s comment was the fourth
and still named the symbol as a live thing; it now names the credential by what it is, and
says the PR route materialises none to refuse.
`docs/remote-state-inventory.md`'s row for `ci-toolchain-image-resolves.sh` listed "an
unusable credential" among the shapes that fail the job — that script holds no credential
any more. The row names the three refused-anonymous-read shapes the shipped script
actually has instead, and re-confirms the `UNSAFE-KNOWN` grade against the anonymous
script: the tag it reads is mutable either way. That is #909's first half; its other half,
`docs/guard-inventory.md`, stays with the session holding that file.
`ci.pr-route-carries-no-stored-credential`'s `mechanics:` carried one self-declared
unmeasured claim — whether act_runner's daemon performs the credential-free `container:`
pull. Measured 2026-09-05 on the runner host 192.168.1.99, which runs both act_runner
containers and creates every job container on its own docker socket: a `docker pull` of
the pinned tag with a scratch docker config holding only `{}` exits 0. The two things that
run did not exercise — an anonymous layer download, and act_runner's own pull call path —
replace the open unknown rather than being dropped, and `docs/ci-cd.md` cites both
measurements.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`secrets['REGISTRY_PASSWORD']` is the same reference to the expression evaluator as
`secrets.REGISTRY_PASSWORD`, and the guard shipped here could see only the dot form. Measured
2026-09-05 against the predecessor of this commit, a `pull_request`-route job whose `env:` read
`"${{ secrets['REGISTRY_PASSWORD'] }}"` produced `stored_secret_faults(...) == []` AND
`walk_versus_text_faults(...) == []` — the text cross-check cannot report the gap, because both of
its halves resolve references through the one pattern, so a spelling it does not know is a shared
blind spot they agree at zero on rather than a disagreement they name
(`proof-sharing-with-subject-proves-nothing`).
The file enumerated four other blind spots it has — composite actions, reusable workflows, nested
directories, both directions of the comment strip — and not this one, which is what made the
omission read as coverage.
`secret_refs` is now the single entry point for both halves, and it matches the dot form, both index
forms and a case-varied context, then counts the RESIDUE: any `secrets` token inside a `${{ }}` span
that yielded no literal name is reported under the sentinel `secrets.*`. Counting the residue rather
than pattern-matching `toJSON(secrets)` and a computed index is what makes it exhaustive over the
grammar — a spelling nobody has written yet still faults, in the fail-closed direction. The bare word
is read as the context only inside an expression, because in prose it is ordinary English; this file
and four workflows discuss "secrets" in comments.
`test_the_collector_sees_every_SPELLING_of_a_secret_reference` drives the six spellings through the
collector and the cross-check and asserts each is invisible to the real predecessor, so reverting the
widening reddens it. Whether act_runner resolves each spelling against this instance was not probed
from here (that needs a live run); the direction makes that acceptable — a spelling the runner does
not support costs a spurious demand on a job nobody has written, the omission cost a live
write-capable credential on the head-authored route.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Three docstrings dated their measurement to `59003d5a3`, this branch's head before
it was rebased onto `main` after ersatztv#907 landed. That commit is unreachable
from the branch and will never be in `main`, so `git show` on it fails for every
later reader — a citation that cannot be followed is worse than none, because it
reads as checkable. Each now names what it measured against ("the predecessor of
this commit", and for the guard, "as it walked `jobs.<id>` only and compared
per-file NAME SETS"), which is what the reader actually needs and what survives
any rebase.
Refs #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`test_the_cross_check_COUNTS_locations_rather_than_collecting_NAMES` asserted the
property on the two COLLECTORS, not on the check that uses them. The comparison
itself was inlined in a loop over the real workflows, which agree under either
mechanism — so changing `walked != scanned` back to a name-set comparison, restoring
the exact blind spot this branch exists to close, left all 17 tests passing. A guard
whose distinguishing mechanism has no mutation proof is the shape
`testing.guard-ships-with-mutation-proof` names.
The per-file half is now `walk_versus_text_faults(name, text)`, driven on a
text/walk pair whose NAME SETS AGREE: a second `${{ secrets.REGISTRY_PASSWORD }}` in
a trailing comment, which the line-level strip leaves in the text half and the YAML
walk cannot reach. Counting reports it; the set comparison the branch replaced
reports nothing, and the test asserts BOTH halves of that so the contrast is the
assertion rather than a comment.
Measured: with `walked == scanned` mutated to `set(walked) == set(scanned)`,
1 failed / 16 passed; restored, 17 passed.
Refs #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
"WITHOUT issuing a Bearer challenge" is a claim about the registry's response that
this script never checks. `probe` enters the token leg on a `401` only, so a `403`
carrying a perfectly good `Www-Authenticate` would be refused with that sentence
having never looked at the header — the same defect one branch over, in the message
written to fix it.
It now says NO TOKEN WAS EVER REQUESTED, which is a fact about the run: the token
leg was not entered, and this answer was never followed as a challenge. The
assertion and the outcome-table row move with it, and the comment says why the
weaker claim is the honest one.
Refs #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The row added for the third refusal shape was written as "401 / 403 carrying NO
`Www-Authenticate` challenge at all", and the script comment beside it made the same
binding. Both are wrong for the 401 half: `probe` enters the token leg on a 401, so
a challenge-less 401 DOES call `acquire_token`, which sets `token_leg_done=1` and
abandons for want of a realm — it reports `could NOT OBTAIN an anonymous pull
token`, the row above. Only a FIRST-READ 403 reaches the never-asked arm. The
parametrised test already drives both codes and asserts exactly that split; the
prose beside them did not match it.
The three rows now bind one shape each: a refusal surviving a bearer the run really
obtained, a 401 whose token leg yielded none (no challenge header, no realm, or no
token in the answer), and a first-read 403 that asked for nothing.
Prose between arms regenerates mis-bindings — which is why the arms are stated as
one self-binding row apiece rather than as a category sentence covering two.
Refs #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`ci.workflow-dispatch-ref-unrestricted` and `ci.actions-credential-scoping` are the
two records the issue requires be updated to match, and both restated the invariant
as "every job of a `pull_request`-triggered workflow that names a `secrets.*`". That
was the shipped predicate when they were written and is now narrower than what the
guard holds: the workflow scope outside `jobs:` is judged too, because a root `env:`
or `defaults:` is materialised into every job and no job-level `if:` reaches it. A
record that understates its own guard is the failure this repo grades worst — it
reads as a checked description and stops the next reader looking.
`ci.pr-route-carries-no-stored-credential` also names the two inventory rows that
this issue made incomplete and did not edit, because both files are held by
concurrent changes: `docs/remote-state-inventory.md` still lists "an unusable
credential" among the shapes that fail the preflight, and `docs/guard-inventory.md`
still describes `test_workflow_persist_credentials.py` as the `actions/checkout`
guard alone. Neither goes red — both suites assert set equality over FILES and both
files were already listed — so the carry is tracked as #909 rather than left to be
discovered.
Refs #885, #909
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two follow-ons from the fixes in this branch, both of the class the branch is about.
`docs/ci-cd.md`'s preflight outcome table listed two token-leg rows and now needs
three: a `401`/`403` carrying no `Www-Authenticate` at all never reaches the token
leg, and the table is what an operator reads to decide where a red preflight sends
them. The paragraph after it named "the two token-leg rows" and now says why the
three are worded apart at all — a message naming a step the run skipped is evidence
for a diagnosis nobody performed.
`secret_name_counts` was added beside `secret_names` as a second traversal with a
different accumulator. That is a copy of a mechanism, free to drift from the one the
assertion runs on — the guard reproducing, inside itself, the defect it was just
widened to catch. There is now ONE walk: the counting one, with `secret_names`
derived from it, which is the lossless direction. Re-witnessed after the refactor —
the workflow-scope hoist into the shipped `docker-build.yml` still reports 3 failed,
the clean tree 17 passed.
Refs #885
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`probe` enters `acquire_token` on a `401` only. A registry answering `403` on the
first read — or a `401` with no `Www-Authenticate` — therefore leaves
`token_leg_done=0` and `token=""`, the guard at the `401|403` arm is false, and the
run fell through to the message that says the read was refused "even after a Bearer
token was obtained". Probed 2026-09-05 with a curl shim answering `403` and dumping
only `HTTP/1.1 403 Forbidden`: that message is printed, EXIT=1, and no token was
ever requested. The fail direction was safe; the diagnosis was not. It sends an
operator to package visibility on evidence that does not exist
(`dont-narrate-mechanisms-you-didnt-measure`) — in a script whose whole design is
that its refusal messages are worded apart on purpose.
The arm now branches on what actually ran, `token` first so the never-asked case
cannot borrow either other mechanism:
* `token` non-empty -> refused after a GOOD bearer (an answer about the PACKAGE)
* token leg attempted -> challenged but produced no token (about the TOKEN ENDPOINT)
* neither -> refused with no challenge at all (about ACCESS)
The pre-existing `403` test could not reach this: `CURL_SHIM` answered `401` + a
challenge to every unauthenticated read regardless of the configured code, so the
`403` parameter was only ever observable AFTER the token leg. The shim grew a
challenge-less behaviour (`CHALLENGE=none`, `REFUSAL=403|401`) rather than the
assertion being written against the old one, and both codes are driven because they
take different paths — the challenge-less `401` still enters and abandons the token
leg. Witnessed red on the predecessor script (2 failed) and green on the fix.
`docs/ci-cd.md`'s "Cutting a release" runbook — the section an operator reads at cut
time — gains the `v*` tag protection, the account it whitelists, the fact that its
positive half is unverified, and the `DELETE .../tag_protections/1` unblock. The
tag-protection note already in this file sits inside the `main`-direct-push
discussion, which is not where a release cut is driven from, and
`release.tag-protection-v-star` names its own failure mode as a cut that will not
push.
`ci.pr-route-carries-no-stored-credential` records that
`docs/remote-state-inventory.md`'s row for the preflight still lists "an unusable
credential" among the shapes that fail the job, which this issue deleted. That file
is held by a concurrent change, so the one-clause edit is tracked as #909 rather
than made here.
Refs #885, #909
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The shipped guard walked `jobs.<id>` only and leaned on a text-versus-walk
cross-check to catch anything the walk could not reach. That cross-check compared
per-file NAME SETS, and the two halves cancelled on the one file the invariant is
about: measured 2026-09-05 at 59003d5a3, hoisting
env:
ETV_REGISTRY_AUTH: ${{ secrets.REGISTRY_USER }}:${{ secrets.REGISTRY_PASSWORD }}
into `.gitea/workflows/docker-build.yml`'s root `env:` — which materialises into
EVERY job on the head-authored PR route — left `pytest
scripts/tests/test_workflow_persist_credentials.py -q` at `14 passed`, rc=0. The
same hoist in `pr-checks.yml` reddened, because no job there already names those
secrets. The guard could only ever see a name NO job used; a second copy of a
reference `build` legitimately keeps naming changed no set. That is
`dont-keep-a-copy-of-a-set` / `proof-sharing-with-subject-proves-nothing`: the
proof shared its accumulator with its subject and cancelled.
Two changes, because the cross-check was being asked to do the assertion's job:
* the workflow scope (everything outside `jobs:`) is now judged in its own right
by the same structure-blind collector — it is a second entry site on equal
footing with the job subtree, not an edge case, since no job-level `if:` can
take a root `env:`/`defaults:` off the route;
* the cross-check walks the whole document and compares occurrence COUNTS. A
duplicate at an unreachable location now reddens: probed 2026-09-05, a trailing
`# ${{ secrets.REGISTRY_PASSWORD }}` on a root `env:` line reports `walk
[('REGISTRY_PASSWORD', 1)] vs text [('REGISTRY_PASSWORD', 2)]` where the set
version agreed. Under counting the comment strip becomes load-bearing rather
than the no-op the old docstring admitted it was.
Driven by a mutation on the SHIPPED `docker-build.yml`, the way the `build`-loses-
its-`if:` mutation already is, plus a direct assertion on the two collectors that
a duplicated reference changes the count and not the names. Witnessed red with the
hoist in the tree (3 failed) and green without it (17 passed).
The decision record's own claims were false in the same way and are corrected:
`rule:` said "NO job ... may name a stored secret" (a root `env:` is not a job) and
the prose said "a text-versus-walk cross-check reports any reference the walk
cannot reach".
Refs #885
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`docker-build.yml` triggers on `pull_request:`, which Gitea resolves from the PR HEAD, so that
run executes contributor-authored YAML and every `secrets.*` it names is materialised into it.
Six jobs held `REGISTRY_PASSWORD` that way — `toolchain-preflight`, `test`, `migrations`,
`functional-e2e`, `api-docs`, `format` — two of them branch-protection required contexts.
The read-only pull PAT the issue asked to cost first was REJECTED, and the measurement is the
reason: this registry already issues an anonymous pull token for `timothy/ersatztv-ci`
(`GET /v2/token?scope=repository:timothy/ersatztv-ci:pull` -> 200), that token reads the pinned
manifest and its config blob (200/200), and the combined-status GET answers 200 unauthenticated.
A read-only PAT would grant exactly what anonymity grants while adding one more credential to the
store head-supplied YAML reaches. So the stronger form was implemented instead: no PR-route job
names a stored secret at all.
- `.gitea/workflows/docker-build.yml`: the five `container: credentials:` blocks, the
`ETV_REGISTRY_AUTH` step env and the three `ETV_STATUS_AUTH` step envs are gone. `build` keeps
the PAT; it is gated `if: github.event_name != 'pull_request'`.
- `scripts/ci-toolchain-image-resolves.sh`: reads `realm` out of the `Www-Authenticate` challenge,
exchanges it once per run for an anonymous pull token, retries with the bearer. Every refusal
direction is preserved — a 401/403 after the token leg, a token endpoint yielding no token, and
one that cannot be reached all `fail` rather than degrading to could-not-tell — and the message
now names the cause an operator can act on (the repo or package has stopped being public).
- `scripts/ci-detect-already-validated.sh`: the status GET is anonymous. No credential override is
kept: the URL names one instance, that instance is public, and an unusable `":"` would draw a 401
and turn a working read into a permanent skip=false.
- `scripts/tests/test_workflow_persist_credentials.py`: the invariant, derived from the git index by
"every job of a `pull_request`-triggered workflow that names a `secrets.*`" — never the six-name
list, and never "every `container:` job", which names five of six because `toolchain-preflight` is
container-free. Witnessed red against the unfixed workflow naming all six jobs; green after.
Live tag protection applied and read back: `POST /repos/timothy/ersatztv/tag_protections`
`{"name_pattern": "v*", "whitelist_usernames": ["timothy"]}` -> id 1. A non-`v*` probe tag pushed
and deleted proves tag pushes still work at all. The POSITIVE release-cut verification is DEFERRED
to the operator's next real cut: pushing a `v*` tag publishes the `:prod` image, which is a release,
not a verification step.
What this does not close, stated so the records are not cited as a boundary: `REGISTRY_PASSWORD`
stays in the Actions store for `build`, and head YAML can still name it, `RENOVATE_TOKEN` or
`SERVERMGMT_DEPLOY_KEY`. Blast radius, not the route.
New records `ci.pr-route-carries-no-stored-credential` and `release.tag-protection-v-star`;
`ci.workflow-dispatch-ref-unrestricted`, `ci.actions-credential-scoping` and
`release.main-direct-push-disabled` updated to match; catalog regenerated. Closes#885.
Decisions-Edit: yes
Proves: scripts/tests/test_workflow_persist_credentials.py::test_no_PULL_REQUEST_route_job_names_a_STORED_secret
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The runner observer reset the shared counter to zero one line before reading
it, so its assertion held under the very mutant it existed to reject and the
fallback assertion caught that mutant for the wrong reason. Completions are
now keyed by the round in each agent's own label; no stub resets shared
state. Measured: re-serialising the runner reddens the runner assertion
(expected [2] to equal [0]) in both scripts; moving the fallback beside the
lenses reddens the fallback assertion and the round-one-failure case.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
A substitute that failed in round one said nothing about the tree that lands
after round two, yet the flag was sticky and doomed the run; the xfamily
string was never reset either, so clearing the stickiness alone would have let
a stale "substitute ALSO failed" sentence into the PR body. Both reset at the
top of review(). The harness runner is round-aware (ran per round, its own
counter reset) and a two-round case pins the fix; restoring the sticky flag
reddens it in both scripts. Step 4 of the mechanics page tells the referee to
read cross_family, not only error, before posting on a rubric-class PR.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The runner-beside-the-lenses design is now what Done-when box 1 asks for (body
amended). A rubric round whose runner and worktree fallback both fail returns
an error before the push instead of landing a PR whose body claims a substitute
reviewed it. The harness records lens count at the runner's start too (expects
0, so a re-serialised runner reddens), resets its counter per round, and has a
case for the double failure.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The runner builds nothing, so serialising it only added its wait to the
critical path; the worktree-isolated fallback is what must follow the lenses,
and the harness case now records lens count at the FALLBACK's start alone.
setTimeout in the harness is globalThis.setTimeout (the .mjs lint config has
ES builtins only). head_sha carries the same description in both scripts and
every fixer/implementer prompt asks for the worktree HEAD, not a PR head. The
mechanics page says why the cap stays at one after the serialisation and
restores the 20%/10% RAM thresholds by key; the record says "several", not
"three".
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
One .NET slot's review round ran two worktree-isolated reviewers at once — the
correctness lens and, on a rubric change, the Codex fallback — and took swap
from 6.8 GB to 10.8 GB in three minutes on the 16 GB host; three slots reached
load 82. review() now awaits the lenses, then the Codex runner, then its
fallback. The finisher's fix attribution is the sha range the fixer's report
head advances (head_sha is required on every report), replacing a line-set
difference over free text that listed all eleven #563 commits as fixes. The
harness gains a case that records how many lenses were still in flight when
the cross-family agents started (must be zero); moving the fallback back into
the parallel batch reddens it in both scripts. The mechanics page and the
standing prompt state the measured cap: one .NET-building slot at a time.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The paragraph added a commit ago pointed at "the commit messages that ran them" as the home of the
binder and `trim` mutant outcomes. A squash merge writes its own message and drops the bodies it
squashes, so that pointer can go stale the moment this branch lands. The issue and its pull request
survive it, and #563 is where the round-by-round measurements already are.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`testing.mutation-claims-are-executed` was amended on main while this branch was in review (#881,
merged as #914): a sentence asserting that a specific mutation reddens — or does not redden — a named
test is now either a `CLAIMS` entry in `scripts/tests/mutation_manifest.py` that executes every run,
or it is not written. This branch carried six such sentences and none of them can be declared:
`Claim.node_id` resolves a proof to `scripts/tests/<node id>` and `run_pytest` invokes pytest, so an
NUnit proof has no representation in that harness at all.
Durable prose now states the mechanism each test is built on — which serializer difference, which
engine branch — which a reader re-checks by reading the code rather than by trusting a remembered
outcome. The record says that in one paragraph, so the limit is stated rather than papered over.
The outcomes themselves are here. Re-measured 2026-09-05 on this branch's tree (the commit before
this one), each mutant applied to the working tree and restored from the index between runs, tree
verified clean afterwards:
positive control ScriptedScheduleControllerTests Passed: 9, Failed: 0
OpenApiSerializerContractTests Passed: 4, Failed: 0
Bind<T> -> System.Text.Json with JsonSerializerDefaults.Web
Failed: 2, Passed: 7 — Production_Body_Binder_Ignores_Required_Members,
Production_Body_Binder_Keeps_Declared_Defaults_Over_An_Explicit_Null
BodyBinderSettings = ApiJsonSettings.Create() -> new JsonSerializerSettings()
Failed: 1, Passed: 8 —
Production_Body_Binder_Keeps_Declared_Defaults_Over_An_Explicit_Null
OpenApiSerializerContractTests RuntimeSettings -> new JsonSerializerSettings()
Failed: 4, Passed: 0 — all four cases, on PascalCase keys
ScriptedScheduleController AddDuration(..., request.Trim, ...) -> false
Failed: 1, Passed: 8 — Committed_Script_Fixture_Produces_The_Pinned_Snapshot
ScriptedScheduleController PadUntilExact(..., request.Trim, ...) -> false
Failed: 1, Passed: 8 — Committed_Script_Fixture_Produces_The_Pinned_Snapshot
The last one is the round-two finding closed and re-witnessed: before the fixture's pad target moved
off the content boundary, that mutant left all nine green.
A squash merge writes its own message, so these figures also belong in the PR description.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two pointers still named #563 as the open reason scripted playout has no
coverage. The ContentEnumeratorBuilderTests header now points at the
successor record and docs/testing.md instead of the issue this branch
closes. docs/decisions.md carried an orphaned fragment, "external-process
pipeline remains #563's", in the residual block under ## Index -- with the
pipeline now permanently outside the automated suite rather than deferred,
the fragment states something false and has no recoverable subject to
rewrite it around, so it goes.
The record's rule gains the two things measurement settled: that
ApiJsonSettings shares production's configuration and never MVC's settings
object (MaxDepth 32, the two ProblemDetails converters, pinned by
ApiJsonSettingsTests), and that a fixture must aim every trimming
instruction between two content boundaries or that action's trim argument
is witnessed by nothing. Its body loses the mutant table and the
extraction paragraph, which the mechanics doc its own frontmatter points
at carries verbatim; what remains is the conclusion plus the reasoning
that exists nowhere else. 70 prose lines to 56, under the advisory ceiling
without dropping a distinct finding.
Refs #563
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The fixture's pad_until_exact targeted 10:00, which the 15- and 30-minute
items reached exactly, so the engine's trim branch never ran and mutating
`engine.PadUntilExact(..., request.Trim, ...)` to `false` left every
controller test green -- measured on the pre-change fixture, `Passed! -
Failed: 0, Passed: 9`. The target moves to 09:55, off every content
boundary: Movie 01 is now trimmed from 30 minutes to 25, the snapshot is
re-pinned around it, and the same mutant fails
Committed_Script_Fixture_Produces_The_Pinned_Snapshot while the sibling
add_duration mutant still does. The trimmed span and OutPoint are asserted
directly rather than resting on the snapshot alone, and the fixture and the
snapshot comment both record that landing a trimming instruction on a
content boundary is what silences its trim flag.
ApiJsonSettings.Create() was documented as a standalone serializer
configured the way MVC's is, which measurement refutes: Apply runs against
a bare JsonSerializerSettings rather than the one MvcNewtonsoftJsonOptions
pre-configures, so MaxDepth stays at Newtonsoft's 64 instead of MVC's 32
and ProblemDetailsConverter and ValidationProblemDetailsConverter are
absent (MissingMemberHandling, TypeNameHandling and DateParseHandling do
match). Neither gap can reach a scripted request body -- two levels of
nesting, never a ProblemDetails -- so this was overstated prose, not a
broken test. Restating the delta everywhere parity was claimed would leave
four copies to rot, so ApiJsonSettingsTests pins it in both directions and
the prose points at the pin.
Also clears the three nullable warnings the replayer helpers introduced
(CS8600/CS8604 on the action string, CS8603 on Bind<T>) and corrects the
ExpectedSnapshot comment, whose last column is built from MediaItemId
rather than looked up from the seeded title.
Refs #563
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The branch asserted that binding fixture bodies through ApiJsonSettings makes "a swap to a
lookalike serializer" redden. Measured, only half of that was true: replacing
ScriptedScheduleControllerTests' BodyBinderSettings with a plain `new JsonSerializerSettings()`
-- a Newtonsoft lookalike that has lost the production configuration -- left all 8 tests green.
Only the System.Text.Json swap reddened. So the production edits the branch makes for that
coupling (ErsatzTV/Serialization/ApiJsonSettings.cs and the Startup rewrite) were justified in
four places by a hazard no test could see.
Both halves are now real. Production_Body_Binder_Keeps_Declared_Defaults_Over_An_Explicit_Null
binds `{"order": null}` and posts it to AddCollection: NullValueHandling.Ignore keeps
ContentCollection.Order at its declared "shuffle" and the call is a 200, where Newtonsoft's own
Include default writes the null through and AddCollection's Enum.TryParse returns a 400. That is
a behaviour difference a script would see, not a settings-shape assertion, so it is not a second
copy of the settings list.
Mutants, run 2026-09-05 over the 9-test fixture:
Bind -> System.Text.Json web defaults 2 red
BodyBinderSettings -> new() 1 red (was 0 before this commit)
OpenApi RuntimeSettings -> new() 4 red (write side, naming strategy)
What still nothing observes is Startup.ConfigureServices itself: re-inlining the
AddNewtonsoftJson lambda as a hand-copy of Apply reddens no test, because a byte-equal mirror is
behaviourally indistinguishable. ApiJsonSettings removes the duplicate rather than detecting its
drift, and docs/testing.md, the decision record and all four docstrings now say that instead of
claiming a detector. Drift confined to ReferenceLoopHandling or the StringEnumConverter is
witnessed by neither suite; that is stated rather than left implied.
Also files the 401 blind spot the record had described as "tracked separately" while nothing
tracked it. ersatztv#913 records the chain, verified from source: the filter is registered
globally, EndpointRequiresKey fail-closes every mutating verb, ScriptedScheduleController carries
no [SkipApiAuthorization], and neither ScriptedPlayoutBuilder nor entrypoint.py supplies a
credential.
Refs ersatztv#563 and ersatztv#913.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The previous commit said a body omitting a `required` member "reaches the action" in production. That
overreaches what was measured: MVC adds an implicit required check for non-nullable reference types
(ErsatzTV.Core.Nullable has <Nullable>enable</Nullable>, and Startup configures no ApiBehaviorOptions,
so the [ApiController] automatic 400 is live), which would very likely reject that body before the
action. What is measured is the SERIALIZER: Newtonsoft deserializes it to a default, System.Text.Json
throws. The prose in the test, ApiJsonSettings, the record and docs/testing.md now stops there and puts
MVC model validation on the uncovered-wrapper list where it belongs.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The docstring, the decision record and docs/testing.md all claimed the replay covered everything
except two hops. MVC model binding was a third: production binds /api/* bodies with Newtonsoft
(Startup -> AddNewtonsoftJson -> CustomContractResolver + StringEnumConverter) while the replay
deserialized with System.Text.Json. Measured on this tree: for the fixture's own bodies the two
agree, but for a body omitting the `required` member "collection" they diverge -- System.Text.Json
throws, Newtonsoft binds Collection = null and the action runs. So the fixture's stated purpose
("field names and casing match what the HTTP body binder accepts") was asserted by nothing, and a
fixture production would bind differently could still go green.
Rather than only widening the residue list, bind the way production binds. The registration moves
into ErsatzTV/Serialization/ApiJsonSettings.cs, Startup applies it from there, and both
OpenApiSerializerContractTests (which had its own mirror of the settings) and the scripted replay
now call that same function -- one definition, no copies to drift.
Production_Body_Binder_Ignores_Required_Members asserts both halves of the divergence THROUGH the
replay's own Bind helper, so pointing the replayer at another serializer reddens; the fixture's own
bodies cannot witness that swap.
The residue is now named honestly in all four places: the binding WRAPPER (input formatter, the
[ApiController] automatic 400 before an action runs) is uncovered, the serializer inside it is not.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The deferral record described ScriptedScheduleController as a "1:1 pass-through"
to SchedulingEngine. It is not: an unparseable playback order is a 400, an
unparseable filler kind SILENTLY degrades to FillerKind.None, an unknown build id
is a 404, and the engine's no-progress InvalidOperationException is translated to
a 400. Carrying that wording forward would have shipped a false statement, so the
successor states a thin adapter with named mappings, each pinned by a test.
- new record testing.scripted-engine-in-process-net (active, since 2026-09-05)
- predecessor testing.scripted-playout-golden-deferred git mv'd to
docs/decisions/archive/testing/ with frontmatter retargeted only; body prose
byte-identical, so no Decisions-Edit trailer
- docs/decisions.md Index line retargeted to the archive path plus a new dated
line for the successor
- catalog regenerated with scripts/build_decisions_catalog.py
- docs/testing.md: the Golden-file nets paragraph now points at the new coverage
instead of "tracked in ersatztv#563"; a new "Scripted playout coverage" section
states what is covered where and what is deliberately not covered (Cli.Wrap
launch, Kestrel + Startup middleware, ApiAuthorizationFilter), dated
2026-09-05; Timezone independence records the per-call TZ audit that decided
which engine instructions the fixtures may use.
Refs #563
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Removing SchedulingEngine.AddCountInternal's _state.AdvanceGuideGroup() left both
guide-group assertions green: the locked-group test only compared inside/outside
the group, and the snapshot only compared item 2 to item 0. Assert the actual
sequence instead.
Refs #563
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
#381 deferred Scripted from the playout golden net because ScriptedPlayoutBuilder
shells out via Cli.Wrap to a user-authored program that drives SchedulingEngine
over HTTP loopback. #563 offered two arms: a full process+Kestrel integration
harness, or expanded engine coverage with the shell-out scoped out.
This takes the second arm, but delivers the first arm's "documented in-process
stand-in" so the scope-out is a measured claim rather than a prose one:
- SchedulingEngineTests grows from 1 test to 21, covering AddCollection/AddCount/
AddAll/AddDuration/PadUntilExact, EPG guide-group locking, per-item history,
the 20-call no-progress halt and its reset, and the anchor round-trip a
Continue build restores from. Unknown-content-key cases assert false AND that
nothing was scheduled.
- ScriptedScheduleControllerTests replays Fixtures/scripted-build.json through
the real ScriptedScheduleController + ScriptedPlayoutBuilderService.MockSession
+ SchedulingEngine and pins a 13-item snapshot in the golden line format, so
there is exactly one action->engine mapping under test — the production one. It
also pins the three mappings the predecessor record's "1:1 pass-through"
wording hides: 404 on an unknown build id, 400 on an unparseable playback
order, a SILENT fall back to FillerKind.None on an unparseable filler kind, and
the InvalidOperationException -> 400 translation.
MockSession was declared on IScriptedPlayoutBuilderService with zero callers;
it is the seam this needs and now has one.
Both fixtures are TZ-independent by construction (Chronological order plus only
instant-preserving instructions) and verified passing, not skipping, under
TZ=UTC, America/New_York, Australia/Lord_Howe and Asia/Kathmandu.
Refs #563
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The previous commit replaced "the order is load-bearing" with "and that is the
whole of what the order buys". What was measured is narrower: with the relevance
gate moved first, the three gates are each still witnessed refusing alone. That
does not establish the order buys nothing else - the reset placement is a second
candidate, unmeasured either way - so the absolute is gone from both sites and
what stays is the cost reason, which is readable from the control flow.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`verify_claim`'s docstring and the record's `mechanics` both said the relevance
gate must run LAST or the status and vacuity gates could never be witnessed
failing alone. Executed at ff65e5e7c: moving the `reset_sandbox` + reach
`verify_mutation` + `if not reach.ok:` block ahead of both earlier gates and
changing nothing else, then running `pytest scripts/tests/test_mutation_harness.py
-p no:randomly -k "GREEN_EXIT_STATUS or GREEN_VACUITY or GREEN_RELEVANCE or
UNKNOWN_outcome"` gives 4 passed. Two of those four assert the reasons the LATER
gates produce ("exited 1", "NOTHING PASSED"), so with the relevance gate first
both earlier gates were still read and still witnessed refusing alone. It cannot
hold, and the branch already said why one line away: the reach mutation injects a
failing test, so it reddens in every fixture except the relevance one, which is
what `_inert_claim_sandbox`'s own docstring states. What the order actually buys
is cost - the relevance gate is the only one of the three that costs a second run
of the proof - and that is what both sites now say. #881's own defect shape,
inside the record that establishes the rule against it.
Second, the record twice gave line-wrapping as the reason a sentence was
paraphrased rather than quoted. The branch's own first `CLAIMS` entry quotes a
sentence that spans a comment line break, embedding the `# ` continuation, and
the harness resolves it exactly once - so a wrapped sentence is quotable by this
very mechanism. The real reason at the calibration site is the replacement
itself: that sentence is not in the tree any more, measured 2026-09-05 by a
fixed-string search over `git ls-files`, which returns no file. At the second
site the referent (`docs/defect-shapes-773.md` section 4) is present and
quotable, so the causal clause is dropped and only the paraphrase marker stays.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The `Claim` docstring motivated the GREEN direction by enumerating three corpus
sites and restating what each asserts. The middle restatement said
`scripts/check-doc-narrative.py` "says removing its `/dev/null` arm reddens no
test" — the universal a9341d841 removed from that file when it narrowed the
comment to the scope the harness actually executes
("`test_check_doc_narrative.py` stays green with this arm removed"). The
docstring was written before that narrowing and kept re-asserting the wider
claim, attributed to a file that no longer makes it: `git grep -F "reddens no
test"` returned exactly one hit, the line asserting it. That is #881's own
defect #2 reproduced inside the fix.
Restating an outcome is what makes it drift, so the enumeration now names the
three sites and the mutation each describes, states the shape they share, and
says why the outcome wording is not repeated. The only two copies of that
outcome left in the tree are the site comment and the `CLAIMS` quote bound to
it, which is the binding by construction. Both other members were re-checked
today and hold: `.gitea/workflows/review-verdict.yml:2411` and
`scripts/tests/hook_fire_isolation.py:82`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The residual paragraph called it "a stronger mutation", which is a judgement
about size; what the mechanism requires is a second mutation of the same clause,
declared and required to redden the proof. One vocabulary across the rule field,
the manifest and the body, so a reader does not have to decide whether two
descriptions are the same thing.
Refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The claim the shipped GREEN entry binds — what removing the `+++ /dev/null` arm
does — was written twice: in `check-doc-narrative.py`, where the entry binds it,
and again in `test_a_DELETED_doc_is_not_reported_as_added_content`'s docstring,
where nothing does. That is the copy-of-an-outcome shape this rule forbids, in a
site class the rule names, found while reading the proof for the residual below.
The docstring now points at the manifest entry and keeps its rationale (a
deletion yields no `+` lines either way), which is the half the carve-out
protects.
The residual paragraph is also made exact rather than general. The reach
mutation proves the proof depends on the clause through the `b/` stripping every
scanned header goes through, not through the `/dev/null` arm itself, so in
general such a green cannot separate "no test feeds that input" from "the arm
changes nothing". For this entry it can, by reading the proof: the deleted-doc
test deletes a tracked file, and a deletion diff under the flags `run_diff` pins
carries a `+++ /dev/null` header — probed rather than reasoned.
Refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round three found the one half of the new mechanism with no relevance gate.
`verify_claim`'s GREEN path read exactly two things — the run exited 0, and
something PASSED — and both are satisfied by a proof that never touches the
mutated file at all. Reproduced before fixing: retargeting the shipped GREEN
entry's proof from `test_check_doc_narrative.py` to `test_bom_guard_detection.py`
changed nothing, and the entry still reported verified. The RED direction never
had this hole, because a proof that ignores the mutation stays green and is
refused as "the clause is not load-bearing".
So a GREEN entry now declares a `reach_replacement` and its `reach_expect`: a
SECOND mutation of the SAME clause, required to REDDEN the same proof, executed
through `verify_mutation` so its red is read through the diagnostic gate rather
than on exit status. The shipped entry declares `path = p` — dropping the `b/`
stripping every scanned diff header goes through — and the run then scans
NOTHING, which is what the declared diagnostic reads. The same retarget now
fails, naming the reach verdict.
The gate runs LAST of the three: run first it would refuse before the status and
vacuity gates were read and neither could be witnessed failing alone (#685), and
the sandbox is reset between a claim's two proof runs for the reason it is reset
between mutations. It has its own disarm proof, and the two synthetic claim
sandboxes are now real git repositories so `reset_sandbox` has a baseline;
`_lib_with` shares the baseline registry, since a copied module's own starts
empty.
Also from that round:
- The record no longer counts the mutation-outcome claims in the pinned
proposal-3 scan. A third of the same shape sits in the same result set
(`test_a_verdict_BEYOND_A_SHORT_PAGE_is_still_found`), and which side of the
line a sentence falls on is a judgement, so an exact count is a figure the
next reader re-derives differently — the failure this record is about.
- The calibration paragraph no longer restates the post-review-verdict outcome
as a dated witnessing. It points at the `CLAIMS` entry that executes it, which
is the form the rewritten shell comment beside it demands.
- The comment in `check-doc-narrative.py` claimed a universal ("reddens no
test") while one file is executed. It now names that file, so the quote binds
an outcome no wider than what is checked.
- Proposal 4 from the issue is dispositioned explicitly: rejected as a rule
here, on the issue's own argument that an exhortation does not fire at the
moment of least slack.
- `docs/README.md`'s task-signal parenthetical now names the `CLAIMS`
population; the file was owned by another slot when this branch started.
Cost re-measured 2026-09-05, three baseline/branch pairs: the `CLAIMS` half adds
31.7-43.6%, up from the 12.7-16.6% measured before the gate existed. The old
figure is retired rather than scaled — growing the population invalidates the
measurement that described it.
Refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Self-review of the previous commit, on the same class it fixes.
"none of them asserts a mutation outcome", written about the 58 lines the
proposal-3 scan returns, is false. Two do: the mutation table at
`docs/decisions/records/ffmpeg/watermark-resolution-unified.md` line 104, which
names a dropped discriminator and the single test that catches it, and
`web/src/screens/AutoTuneScreen.test.tsx` line 174, which says what a revert to
the old flex row can redden. Both read in full at `efadbec29` rather than from
the truncated grep line - the truncation is how the first pass missed them.
Three kinds of sentence under one pattern is a better argument than the one the
false claim was making: it is not that the pattern finds only rationale, it is
that it finds rationale, state anchors and mutation-outcome claims side by side
and nothing in the string tells them apart.
Two smaller ones in the same commit. The new test's docstring said the gate is
"the one gate the others cannot cover" and the manifest said "the one PRE-FLIGHT
refusal a red proof cannot be told apart from" - both assert uniqueness among
the pre-flight refusals that neither measured, and a clause occurring zero times
also leaves the text identical. Narrowed to what the mutant demonstrates: no
later gate stands in for it. And the fixture comment glossed `verify_claim`'s
GREEN refusal in quote marks, which under this record's own proposal-2 clause
reads as a quotation of the library; it is not one, so the marks are gone.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Four review findings, all of the class this branch is about - prose asserting
what a mechanism does, with nothing binding it to the mechanism.
The record's `mechanics:` said a GREEN claim over an already-red proof "would
otherwise be satisfied by the redness it is supposed to rule out". Inverted:
`verify_claim`'s GREEN branch REFUSES any non-zero exit, so redness refutes a
GREEN claim and can never satisfy one - which is what
test_MUTATION_disarming_the_GREEN_EXIT_STATUS_gate_accepts_a_proof_that_WENT_RED
asserts. The hazard the control removes is the same one it removes for the
rows, and it runs in both directions: an already-red proof satisfies a RED
claim with redness its mutation did not cause, and refuses a GREEN one for a
reason unrelated to its mutation. Both the record and the fixture comment now
say that, and both say what the control CANNOT do - its assertions are over the
aggregate of every proof ref, so a single ref collecting nothing is invisible to
it and is caught per-claim by the vacuity gate instead.
The manifest's `why` had widened a scoped sentence into "THE OTHER GATES EACH
CARRY THEIR OWN PROOF" and then enumerated them, which made the enumeration a
completeness claim it could not meet: the identical-replacement refusal carried
no proof at all. The review measured that at 8adf21eff - `if mutated ==
original:` disarmed, whole file 49 passed 1 skipped. That gate is the one a red
proof cannot be told apart from: the mutant is byte-identical, so the proof runs
against the original tree and an already-red one reddens exactly like a
detection. Disarmed, the harness certifies it as "the named test went red under
the declared mutation, with the declared diagnostic" - witnessed here on the
real library, restored after. So the measurement above no longer holds, by
construction: the gate now has a disarm proof, and the sentence says explicitly
that naming the gates is not a claim the list is closed.
Proposal 3's rejection quoted "16 lines" with no predicate - the defect the
record's own body names three paragraphs later, where the population scan is
pinned verbatim for exactly that reason. The figure is not reproducible from the
text. Replaced by a pinned `git grep` over the same corpus at the same sha
(`32 files, 58 lines`), with what reading all 58 shows: they are rationale, the
class the rule carves out, and the few real state anchors among them are not
separable by pattern, because the difference is whether the sentence explains or
asserts.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
"one of the three entries below asserts that a mutation is NOT noticed" is a
count of the entries it sits above, and the preamble's last line - and the file's
own docstring - say no count is kept here, because a count of the entries is a
second copy of them. The previous wording said "three", which was also wrong: one
entry is GREEN. Corrected to "three" would have been an accurate second copy;
the sentence now states the shape and counts nothing.
refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`docs/defect-shapes-773.md` §4's sentence was quoted verbatim-looking but with a
lowered initial capital, and the source wraps it across a line at
`evidence`/`behind`, so neither the written form nor the corrected one is
findable by grep. Same treatment as the `post-review-verdict.sh` one: paraphrase
without quote marks, keep the section reference, say why.
The section reference itself was checked - the sentence is at line 351, under
`## 4. Detectors`.
refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Two more in the same class, both in prose this branch wrote.
The calibration paragraph quoted `post-review-verdict.sh` as asserting "the
absent-entry check catches every case on its own". That string occurs in no file:
the comment wraps it across a line break at `catches`/`every`, so `git grep` for
it finds exactly one hit - the record asserting it. That is ersatztv#812's second
defect reproduced inside the record written to end it. Paraphrased without quote
marks and pinned to lines 316-317 at `efadbec29`, which is what this record's own
proposal-2 clause prescribes for a quotation that cannot be checked.
The closing paragraph called the scan's hits "the remaining population" and "a
backlog", three paragraphs after establishing that both figures are CANDIDATE
counts and that reading them as a backlog of real claims overstates them. The
closing text now says what is actually known: a place to look, with nobody having
established how many are claims.
refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The sentence explaining why the scan pins a sha quoted a second figure for
`HEAD`, which is exactly the anchor-to-a-moving-state shape this record settles:
correct on this branch, wrong the moment anything else lands. The reason it was
supporting is checkable without a number - the paragraph's own prose, the pinned
command included, matches the pattern.
The first draft of that replacement said "twice over". Three lines of the
paragraph match, so the count is dropped rather than corrected; a count of
matching lines in a paragraph nobody will re-measure is the same defect one size
smaller.
refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The generated catalog embeds each record's `rule`, so rewording the quote-scope
clause in 38bdf7bd0 left `docs/decisions/README.md` behind the record. Nine tests
red on it - the four `build_catalog_check_path` reformat cases, its stale-catalog
CLI proof, two `decisions_validate` main() cases, and the two mutation-harness
entries whose positive control runs that validator.
refs #881
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Self-review over the fix commit, on the same predicate the record states.
The cost band said 13-17% over measurements spanning 12.7% to 16.6% - a band
that ROUNDS is still a band the data does not support at its lower edge. It now
states 12.7-16.6%, which is the span itself, with the per-pair figures beside it.
The quote-scope rule was stated as an OUTCOME claim ("free to be rewritten under
a green harness", "with the entry still reporting the red as verified") in the
record's `rule`, in the manifest preamble and beside the entry. That is a
mutation-outcome claim about the harness with no `CLAIMS` entry behind it -
manufactured by the sentence that introduces the rule against it. All three now
state the STRUCTURE, which is what a reader can check by looking: the assertion
and the test it names are outside the binding.
The manifest preamble said "three of the entries below assert that a mutation is
NOT noticed". `CLAIMS` holds three entries and exactly ONE is GREEN; the three
the `Claim` docstring names are CORPUS sites, not entries. Corrected to one, and
"the most common shape prose actually takes" - a frequency nothing measured -
dropped rather than quantified.
Two claims in the new record prose were themselves overstated. The 69-line green
narrowing was described as the negative direction rather than as candidates for
it: sampling the hits shows `green` in this corpus is as often a CI job's colour
as a mutation's outcome, so both figures are now labelled CANDIDATE counts. And
the "the number moves under the commit that records it" sentence now carries the
figure that shows it - the same command with `HEAD` in place of the sha prints
`107 files, 344 lines`, measured on the committed tree.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Four review findings, every one the defect class this record is about: a prose
assertion with nothing binding it to what it asserts.
THE DERIVED POPULATION WAS NOT REPRODUCIBLE. "128 lines across 65 files" and
"47 of the 128 carry the NEGATIVE direction" cannot be reached from the predicate
the record described, while the record told the reader to re-run it there. A
review swept ~40 readings of that description at efadbec29 and none returns
either number. The scan is now pinned VERBATIM as the command that produced it,
and the figures are what that command prints at efadbec29 on 2026-09-05:
104 files, 311 lines - candidates
39 files, 69 lines - the same command with the outcome half narrowed to
`green`, i.e. the negative direction
Both were re-run by extracting the fenced command from the committed file and
executing it, so the text and the numbers cannot have diverged. This supersedes
the 47/128 figures quoted in 05992bec7's message.
CLAIMS[0]'s QUOTE BOUND THE WRONG HALF. It stopped at the comma after the
mutation, leaving "so `test_a_readback_whose_statuses_array_is_NULL_is_refused`
reddens" outside the binding - the words that make the sentence a claim. The
harness counts occurrences of the quote alone, so the outcome could be rewritten,
or the test renamed in the prose, with the entry still reporting the red as
verified. The quote now spans both halves, and the scope rule is stated in the
record's `rule` and in the manifest beside the entries, where the next one is
written.
"BOTH WERE CORRECTED IN THE SAME CHANGE" WAS FALSE. Only the shell comment was:
the Python test has asserted the shape diagnostic since 5d955000f (#889), and
this branch does not touch it. What this change adds beside the rewritten comment
is the binding.
THE COST BAND CONTRADICTED ITS OWN MEASUREMENTS. `mechanics` said 13-15% over
three pairs spanning 12.7%, 16.6% and 14.4%. It now states 13-17% and the
per-pair figures, since the record tells the reader to carry the percentage
forward rather than the seconds.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`test_mutation_harness.py` now imports the shared index derivation, which
`test_every_index_derived_module_is_registered` requires be registered or exempt.
It is exempt: its population is `CLAIMS`, and it consults the index only per
member, to answer whether a declared `site` is a path git tracks. The exemption
list's own docstring counted its entries, so that count and its review date move
with it.
Three claims written by this change were falsified by this change, which is the
shape it exists to catch:
- the binding test's docstring said membership comes from the index "not from
`Path.is_file`", while the same test now asserts existence with `is_file`;
- the record quoted the manifest docstring this change rewrites — an anchor to a
state the commit moves, which the record itself rejects. It now anchors to
`efadbec29`;
- the `Claim` docstring quoted three files without naming them. They are named.
refs #881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`testing.mutation-claims-are-executed` was the right rule scoped to its first
site: `mutation_manifest.py` declared itself "one per `MUTATION`-graded row of
`docs/guard-inventory.md`", so the same claim written in a code comment, a test
docstring or a decision record was outside it by construction. That is where all
four of ersatztv#812's consecutive review-round defects lived.
Extend the rule in place rather than adding a sibling record: a sibling would
recreate the exact shape (a rule per site class, with the next site class outside
both) that #773, #784 and #743 each are. The subject is unchanged; only the
population widens.
Mechanism: `CLAIMS` in `scripts/tests/mutation_manifest.py`, keyed on the PROSE.
Each entry carries the tracked `site` and the verbatim `quote`, checked every run,
so a reworded sentence reports as a retarget instead of drifting from the entry
that justifies it — this is proposal 2 (a quotation of another file is a claim
about that file) adopted where the referent is declared. Each entry also declares
RED or GREEN and is executed in the existing sandbox. GREEN is new: 47 of the 128
candidate lines the corpus grep returns at efadbec29 assert that a mutation is NOT
noticed, and no `MUTATION` row can express that, so the rule was unsatisfiable for
them. The green direction is read by two separate clauses (exited 0, and something
actually passed) so neither can mask the other, and each carries its own disarm
proof.
Proposal 3 (never anchor prose to a state your own commit moves) is rejected as a
DETECTOR and kept as a phrasing rule: measured 2026-09-04, the only plausible
pattern set for it matched 16 lines across the scanned corpus and every one was
legitimate rationale prose.
The seed set falsified a shipped claim on its first run: `post-review-verdict.sh`
asserted that disarming its array-TYPE read-back test left the suite green. It
does not — jq refuses to iterate a `null` `.statuses` and the script dies with the
parse message, reddening `test_a_readback_whose_statuses_array_is_NULL_is_refused`.
Comment corrected, entry graded RED.
fixes#881
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
A round in which every lens returned null read as a clean pass and would now
have been quoted verbatim into the PR body; it is an error before the push,
in the review loop and in the post-rebase round, which also gains the same
blocking-or-should-fix filter; a fixer that dies or stops (no done) is an
error too, the same test the implementer already gets. The history entry now
carries the fix that answered that round and only the commits that fix added
(a line-set difference against the previous branch log — a fixer that
reformats or rebases mid-loop defeats it, which is why the finisher is told
to read git show, not the list). An empty fix-commit set is described as
"answered without a new commit" when a fix round ran, and as "round one was
clean" only when none did.
web/scripts/orchestration-workflow-loop.test.mjs compiles the committed script
bodies with stubbed agent/parallel and pins eleven paths per script (22
tests). Measured: reverting the loop condition to blocking-only reddens six
cases per script (every case that needs a should-fix round to reach the
fixer); deleting any of the three zero-lens guards, the fixer guard or its
done half, or the empty-fix sentence branch reddens its own case, in both
scripts. web/vite.config.ts is untouched: it is pinned whole by
test_image_build_delegates_the_spa_suite.py, comments included.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
"its button renders only before the player has ever mounted" is loose: a channel
switch to a forced channel re-renders the button after a player has mounted for the
previous channel. The load-bearing fact is the one the code comment states — the
button renders only while `started` is false, and the only thing that un-starts the
panel is the channel reset that clears the flag in the same batch.
Refs #554
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The hint's single-guard invariant was stated as "every path back to `starting` clears
the flag itself", enumerating Retry and a channel change. `onOptIn` is a third such
path and clears nothing, so the sentence was false as written — in the code comment,
in the Retry test's comment, and in docs/spa-conventions.md §5b.
Adding a clear to `onOptIn` would be dead code no test could distinguish, which is the
exact shape this branch removed from `onPlaying`. The omission is correct for a reason
none of the three places stated: the opt-in button renders only while `started` is
false, `started` only goes false in the render-phase reset that clears the flag two
lines later, and no player exists to set the flag while `started` is false. State that
exception, and pin the reachability premise it rests on with a test that fails if the
opt-in button outlives the mounted player.
Refs #554
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Dropping the onPlaying clear made the `state === 'starting'` render guard the
hint's only guard, which moves the burden onto the two paths back to
`starting`: each has to clear `autoplayBlocked` itself. Both were assertable
but unasserted — either `setAutoplayBlocked(false)` could be deleted with the
whole panel suite green, so the invariant the code comment and
docs/spa-conventions.md §5b both state was unpinned in both of its named paths.
Add one test per path (Retry; a channel switch), each asserting the hint is
gone while the panel is back at `starting` — so the render guard cannot be
what hid it. Each also asserts the player really re-mounted (loadSource count
/ last URL, plus a non-null <video>), so the hint cannot be absent merely
because the `resolvedSrc` block is unrendered.
Measured on this tree, each mutation caught by exactly one test:
deleting the onRetry clear reds only 'clicking Retry clears the
autoplay-blocked hint' (Tests 1 failed | 25 passed); deleting the
render-phase reset clear reds only 'clears the autoplay-blocked hint when
switching to a different channel' (1 failed | 25 passed); replacing
`autoplayBlocked && state === 'starting' &&` with `autoplayBlocked &&` reds
one test too. Unmutated: 26 passed.
Refs #554
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round one measured three defects in the first commit.
onAutoplayBlocked fired on ANY rejected video.play(), so the panel could
say "autoplay was blocked" when it was not. A play() interrupted by
teardown rejects with AbortError — which is exactly what the panel's own
Retry produces while the MANIFEST_PARSED play() is still pending — and
because the element was muted, a genuine NotAllowedError is the rare
case, so the realistic firings were the mislabelled ones. Report only a
DOMException named NotAllowedError, on both the MSE and native paths.
`muted` was applied to the shared player unconditionally, which silently
muted the playback-troubleshooting screen — the tool whose job includes
verifying the audio side of an FFmpeg profile, and which the legacy
Blazor player never muted. Make it an opt-in `muted` prop defaulting to
false; the channel preview passes it, troubleshooting does not, and a
test on each side pins its own value.
The two clauses hiding the hint once playback starts masked each other:
removing either alone left the panel suite green. Every path back to
'starting' (Retry, a channel change) already clears the flag itself, so
the clear in onPlaying could never be the load-bearing guard — drop it
and let the `state === 'starting'` render guard be the single pinned one.
Also cover the native-HLS (Safari) branch, which no test had ever
executed: its play() kick, its playing/error wiring, and both autoplay
rejection names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The fixer brief already said "fix every blocking and should-fix one"; the loop
condition alone disagreed, so a merge-worded round with real defects skipped
the fixer and the finisher attested to fixes it never saw (#554 / PR #910).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
HlsPlayer's manifest GET can block until segments exist (unbounded
maxTimeToFirstByteMs), so MANIFEST_PARSED can arrive past the
browser's transient user-activation window and video.play() gets
rejected as blocked autoplay — the channel preview panel then sat at
"starting" over a black frame with no hint the operator just needed
to press play.
Render the <video> element muted (browsers permit autoplay of muted
media without user activation) so the common case starts on its own,
and add an optional onAutoplayBlocked callback for the residual case
(stricter policy/extension) that the channel preview panel wires to a
"press play" hint shown only while still starting.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Review measured M6: swapping `reportFailure(...)` for `onAddFailed(...)` in
`AddItemsDialog`'s catch left all 490 tests across the 42 `src/screens` files
green. Under that mutant a failure that lands while the dialog is STILL OPEN
renders into `ManualItemsView`'s screen banner, which sits behind the dialog's
`createPortal` panel with `aria-modal="true"` — covered for sighted users,
hidden from AT, and the surface the user is actually looking at stays blank.
That is the exact shape the decision record calls "its own defect", and the
whole gap was the call-site wiring: the hook's inline branch is pinned at unit
level in `hooks.test.tsx`, but a unit test of the hook cannot see which arm a
consumer reaches.
Adds the integration assertion: fail the POST with the dialog still up, assert
the message is inside `[role="dialog"]` and appears exactly once in the tree.
Re-executed the mutation with it in place — 1 failed / 490 passed, and the red
is this test alone.
Records the new pin as mechanics (5) on
`spa.dismissible-write-failure-reporting` and the general form in
`spa-conventions.md` §3c, so the next site wired to the hook pins both arms at
its call site rather than inheriting the hook's unit coverage.
Refs #830
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
Round 6 verdict was MERGEABLE with two follow-ups; both are one-liners on lines this branch just
touched, so they are in rather than deferred.
L11. I believed the suite covered "a normal add still closes the dialog". Review MEASURED that it
did not: deleting the `onClose()` call entirely -- so a successful add leaves the picker open
forever -- kept the whole suite green, 1277/1277. The new #830 test only pins the NEGATIVE direction
(do not close when unmounted), so a future edit dropping the call, believing the guard had made it
dead, would have shipped silently. The Song add test now asserts the dialog closes; with that line,
the same deletion reddens. Both directions of the report/dismiss split are pinned.
Worth naming the shape: I asserted coverage from plausibility rather than from a mutation, in the
same PR whose whole subject is claims that were written down before they were measured.
N12. The guards test's "exactly ONE post-await write to state THIS component owns" is still true,
but it now reads as a census of `mountedRef` reads, and `submit` has two -- the success path's
guarded `onClose()` is the other, which that failure-path test never reaches. Added the clause so
nobody derives the guard population from that number.
1277 tests green, tsc/eslint/build clean, pytest 1228 passed, validator OK, catalog no drift.
The two red CI contexts on the previous head are runner flakes, not this branch: both failed inside
`Post Checkout` with `Cannot find module '/var/run/act/actions/<hash>/dist/index.js'`, their logs
are timestamped 19:18 (before that head existed), this branch touches no CI or docker/ci file, and
both contexts were green on its earlier heads.
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Round 5 found that §3c and the `rule:` field now instruct readers to gate the DISMISS request, while
`AddItemsDialog` -- the one site this record names as "the shape fixed here" -- called `onClose()`
unguarded, with a comment arguing that was correct. So a reader following the convention wrote the
gate and a reader copying the reference implementation did not.
The defect is pre-existing; the CONTRADICTION is mine, introduced when round 3 withdrew H1's code
but kept the convention it produced. I checked the docs against the withdrawn addTo code and did not
re-check them against the exemplar that stayed.
Measured at this site: submit, Escape mid-request, reopen the picker to retry, first POST returns
204 -> the stale instance's `onClose()` (`() => setPickerOpen(false)`) closes the dialog the user
just reopened, discarding the selection they rebuilt. Identical mechanism to the addTo clobber.
Unlike the addTo layer, the one-line gate IS sufficient here, and that difference is the point:
`AddItemsDialog`'s parent has no competing closer (`onAdded` is `load`, which never touches
`pickerOpen`), whereas `AddToMenu.handleAdded` closes its dialog itself. That is now stated in the
record as the concrete reason one half shipped and the other went to #877.
- `onAdded()` stays unguarded -- it REPORTS, and the parent's list reload must survive dismissal
- `onClose()` is guarded -- it REQUESTS A DISMISSAL, and after dismissal it aims at whatever the
user opened next
- comment rewritten to say which is which and why, instead of defending both as "belong to the
still-mounted PARENT"
Pinned, and nothing pinned it before: "a late SUCCESS does not close the dialog the user reopened
after dismissing (#830)". It carries an anti-vacuity check that the late response was actually
processed -- `onAdded` is `load`, so a second GET of the items endpoint must have happened -- because
otherwise "the dialog is still open" holds trivially. Executed: deleting the `if (mountedRef.current)`
around `onClose()` reddens it alone.
Also rewrapped five record body lines left ragged by earlier splices.
1277 tests green, tsc/eslint/build clean, pytest 1228 passed, validator OK, catalog no drift.
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Round 4 found one clause left, and it is a good example of the thing this record is about. The
one-consumer paragraph said `hooks.test.tsx` "is the ONLY thing pinning the diverted branch".
Measured at the previous head, disarming `reportRef.current(message)` reddens THREE tests -- both
hooks.test.tsx divert tests AND the CollectionsScreen integration test -- which is exactly what the
`mechanics:` field of the same record says 58 lines earlier. So the record asserted a coverage fact
and then contradicted itself.
The concrete harm is not the inconsistency: a future session pruning tests reads "hooks.test.tsx is
the only pin", concludes the CollectionsScreen #830 test is redundant, and deletes the only
end-to-end pin of the whole path -- the one that actually drives Escape-dismissal through the real
dialog. Clause dropped; the argument the paragraph needed (the hook's shape earns its own unit
tests) survives without it.
The clause originated in the reviewer's round-3 wording and I transcribed it without checking it
against a field I had written myself two rounds earlier. Worth recording: a review finding is not
exempt from verification just because it came from the reviewer.
Also:
- the `AddToMenu` clobber sentence now splits what was MEASURED (a late success closes a reopened
dialog) from what was READ (both parents call `clearSelection()` unconditionally, so the wipe
follows). On a record whose subject is over-attributing measurements, that distinction has to hold
in its own prose.
- §3c now carries the same "nothing diverts to those screens today" disclaimer the record's limit
(2) has, so the two artifacts say the same thing
- rewrapped one 141-char comment line left ragged by the previous round's splice
Docs only, plus one comment rewrap. 1276 tests green, tsc/eslint/build/validator/catalog clean.
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Round 3 verified the withdrawal itself is clean -- the eight reverted files are byte-identical to
origin/main, no orphans, and the mutation gives the stated three reds -- but found docs still
asserting the withdrawn change shipped. That is the stale-comment failure in its usual form: after a
retraction, the retracted WORDING has to be swept, not just the code.
- `hooks.ts` said "#830 removed that gate", flatly false at this head, in the hook's own doc comment
right above the export -- the first thing a maintainer reads. It also carried round 2's framing
("both halves of the outcome") as the hook's purpose, when what ships carries only the failure
half. Rewritten to the present tense of the shipped tree.
- The `rule:` field still said the surviving surface "differs per screen", naming MediaBrowseScreen
and SearchScreen as wired. They wire nothing. This one matters beyond an ordinary sentence:
`rule:` is the canonical summary, it is what the catalog row shows, and it is what gets mirrored
per-key into MemPalace -- so it is the version a future session retrieves WITHOUT opening the
file. Now: exactly one wired screen, the Toast pair named as a CANDIDATE.
- The "two limits" bullet described a failure being diverted to those same screens and announced
politely. Nothing can divert there -- they receive no reporting callback. Restated as the limit
the second surface will have when it is wired.
- `onFailed` in a hooks.ts comment was a dangling identifier; the real prop is `onAddFailed`.
Also added the caveat the reviewer asked for rather than leaving it to be discovered: this is a
shared hook with exactly ONE consumer. It earns that shape (directly unit-tested, and those tests
are the only thing pinning the diverted branch; prescribed by §3c; #877 queued as a second
consumer) -- but #877 may land a shared reporting SURFACE instead of a per-site prop, in which case
the second consumer never arrives. Accepted risk, now written down.
Docs only. No code change, 1276 tests still green, tsc/eslint/validator/catalog clean.
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Round 2 of adversarial review found that my round-1 fix introduced an adjacent defect, and it is the
same mechanism both times: `onAdded?.()` and `onClose()` sat behind ONE unmount gate in the four
`media/addTo/` dialogs, and those two callbacks do not mean the same thing.
- Gate both (origin/main): a write that SUCCEEDS after dismissal reports nothing. Measured on
`AddToCollectionDialog` -- `onAdded` called 0 times after dismissal. That was round 1's finding.
- Un-gate both (my round-1 fix): a late success closes a dialog the user REOPENED to retry, and on
SearchScreen/MediaBrowseScreen `clearSelection()` wipes a multi-select they rebuilt. Measured
against the real `AddToMenu`. That was round 2's finding, and I introduced it.
- Gate only `onClose`: still wrong on its own, because `AddToMenu.handleAdded` nulls the dialog
itself. Needs three coupled edits across five files -- plus a genuine product question nobody has
answered: should `clearSelection()` fire for a write the user walked away from?
That is a design change, not a bug fix, and #830 never asked for it -- the issue is about
`AddItemsDialog`. Two defects from one mechanism in two rounds is the signal to stop widening, so
the `media/addTo/` extension is REVERTED here and moves to #877 with every measurement attached
(#877 comment). What ships is the thing the issue asked for, proved:
- `useDismissSafeError` + `AddItemsDialog` + `CollectionsScreen` wiring
- the witnessed red is unchanged: disarming `reportRef.current(message)` reddens the integration
test on `Unable to find an element with the text: Request failed with status 500`
Docs now describe what is actually true rather than what I hoped:
- the record says ONE A1 site is fixed and explains why the other four were withdrawn, keeping the
wrong first claim visible because "one site read, four assumed" is the lesson
- §3c splits the rule the round-2 defect came from: report the OUTCOME unguarded, gate the DISMISS
request separately -- the earlier text lumped `onClose` in with `onAdded` and would have
propagated the clobber to the next screen that adopted it
- §5c no longer tells authors to wire an `onFailed` that the addTo layer does not have; it says the
layer has no failure channel at all and points at #877
- the reporting prop is REQUIRED where the host has a surface (`AddItemsDialog.onAddFailed`), which
is what the docs now say instead of calling it optional
- `mechanics:` no longer implies the shared clause reddens one test; it reddens three, so re-running
the mutation should expect three
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Adversarial review (cold, worktree-isolated) blocked the first commit on its central prose claim,
correctly. I wrote "at every one of these sites SUCCESS already outlives dismissal" into three
durable artifacts -- the decision record, spa-conventions §3c and the hooks.ts header -- after
reading ONE site. `AddItemsDialog` does report success past dismissal and says so in a comment; I
generalised from it. The review probed the other four instead and MEASURED `onAdded` called 0 times
after dismissal: all four `media/addTo/` dialogs gated `onAdded?.()`/`onSaved?.()` behind their own
`activeRef`, exactly like the failure path.
So after the first commit those four were still asymmetric, just inverted: dismiss-then-fail loud,
dismiss-then-succeed silent -- and additionally leaving the caller's selection state stale, because
SearchScreen's `onAddedToSelectionTarget` never ran to clear it. The record's own advice ("add the
failure counterpart") followed literally would have reproduced it.
Fixes, each proved by execution:
- the `activeRef` gate above `onAdded?.()`/`onSaved?.()` is removed in all four dialogs; those two
statements belong to the still-mounted PARENT, which is the reasoning AddItemsDialog already had
- `AddToCollectionDialog.test.tsx` covers the media/addTo half in BOTH directions. It had NO
coverage before: reverting `reportFailure` to `setInlineError` in all four left the whole suite
green. Restoring the success gate reddens the SUCCESS test alone; disarming
`reportRef.current(message)` reddens the FAILURE test alone
- the three prose sites now say what was measured, and the record keeps the wrong first version
visible, because "one site read, four assumed, written down before measuring" is the finding
Also from the review:
- hooks.ts said "React 18"; package.json pins 19.2.7. Now "React 18+"
- the "nothing better to do" comment overclaimed: diversion reaches ONE level, so Back out of a
collection mid-add still drops the message. Stated, with where it would be fixed
- recorded two limits rather than leaving them to be rediscovered: useIsMountedRef clears in a
PASSIVE effect cleanup, leaving a narrow window where the message renders inline into a detached
tree (useLayoutEffect would close it, but that hook is shared by every async caller -- #877, not a
bug fix); and Toast is role="status" with one last-writer-wins slot, so it is not equivalent to
CollectionsScreen's role="alert"
- §5c now cross-links §3c, since that is the section a screen author reads before wiring AddToMenu
- the sweep count is 67 caller-owned + ConfirmDialog's own internal <Dialog>
Three other media/addTo dialogs remain unpinned; they are identical in shape to the covered one,
which is a reason to expect the same behaviour, not evidence of it. Said so in the record.
refs #830, #877
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
`AddItemsDialog.submit` POSTed to `/api/v1/collections/{id}/items` and reported failure into a
banner rendered from its OWN state. The dialog is dismissible mid-request through three paths that
never consult `adding` -- Escape and a backdrop click (both `useOverlayBehavior`) and the header
close button -- and the caller remounts it on `key={`add-${pickerOpen}`}`, so dismissal genuinely
unmounts it. Select 12 items, Add, press Escape, the request fails: nothing surfaces, the list
reloads unchanged, and the user believes 12 items were added.
drop deliberate rather than accidental. A deliberate drop is still a user who is told nothing.
The asymmetry is the finding: SUCCESS already outlived dismissal everywhere here, because it is
reported through a parent callback (`onAdded`/`onDone`, which the screens turn into a `Toast`).
Only failure died with the surface. So this is not a new notification system -- it routes failure
through the channel success already uses. `AddToMenu` had `onDone` and no counterpart at all.
`useDismissSafeError` (`web/src/hooks.ts`) renders the message INLINE while the surface is mounted
-- the better surface, since it keeps the user's selections and context -- and diverts to a
caller-supplied `onFailed` once it is gone. The surviving surface belongs to the parent and differs
per screen (a `role="alert"` banner on CollectionsScreen, `notice`+`Toast` on MediaBrowse/Search),
so it is a prop contract rather than a rendering decision. Gating dismissal on the busy flag was
considered and rejected: it traps the user behind an in-flight request with no cancel path, and
would not cancel the write anyway.
Applied to the A1 shape -- where the surface owns the error state and is really unmounted:
AddItemsDialog plus the four `web/src/media/addTo/` dialogs, whose failures previously could not
reach the screen Toast that already showed their successes.
Proofs, executed rather than described:
- deleting `reportRef.current(message)` alone reddens the new CollectionsScreen test on
`Unable to find an element with the text: Request failed with status 500`
- `hooks.test.tsx` pins both branches directly, plus that the report goes through the LATEST
callback rather than the one captured on first render
- `CollectionsScreen.guards.test.tsx`'s is-mounted read count moves 2 -> 1 because the catch's
guard migrated into the hook (its `...actual` module mock cannot see the hook's internal
`useIsMountedRef()`); removing the surviving `finally` guard takes it to 0 and reddens, so the
anti-masking property that count was added for is intact
Scope is stated rather than implied. A sweep of all 68 Dialog/ConfirmDialog/SlideOver call sites
found three shapes; only A1 is fixed here. A2 -- error state that survives but whose render site is
gated by the same condition dismissal clears, mostly delete-confirm flows -- is left open in #877
because its right answer is probably a shared surface, not twenty prop threads.
fixes#830
refs #877, #740, #685
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XE2tF2aUasK2hWPmBRsrMY
Embedding the slot's gate text in the reviewer brief carried the slot worktree
into the one prompt that forbids it, and a substring port substitution could
rewrite a path containing the same digits; gateFor(port, where) renders each
brief for its own tree and port. The port guard accepted "", null and false
through Number(); it now requires a JS integer in (1024, 65000). The finisher
schema requires only patch_changed, and a done report without a PR URL or head
sha is an error rather than a placeholder.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The marker-overwrite instruction asserted an antecedent no agent can verify
and, since same-session worktrees carry no marker, could only fire in another
session's worktree; the scripts now stop and report. `ran` and `patch_changed`
move into required keys of their own schemas so a missing field cannot read
as a successful cross-family review or an unchanged patch. A blocking finding
in the post-rebase round now returns an error like every other failure path.
Each reviewer lens gets its own E2E port; the gate text travels with the
reviewer brief. The standing prompt no longer contradicts the substitution
the scripts perform; the record names patch-id, the mechanism the scripts use.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The mkdir lock around scripts/e2e-local.sh serialised the launch, not the run
(the launcher returns with the server up), and its stale-holder path double-
acquired in 4 of 91 measured races; the launcher's documented conflict is its
per-worktree wwwroot, so slots now run on their own port and the lock is gone
with its inventory row. process.orchestrated-session records the two scopings
the harness needed: a rebase pushed with --force-with-lease as the one sanctioned
rewrite, and the referee as the only agent that ticks Done-when boxes. Scripts:
required-arg guard, per-issue claim probe, reviewer fetch recipe, codex fallback
to a cold review-only agent with the substitution stated in the PR body, rebase
before the review loop with a patch-id check at the push, non-interactive
squash recipe, Land phase. README bullets re-parented; kickoff bullet keyed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
The single-issue kickoff stays as it is; this adds the layer that runs several of
those under one referee. docs/handoffs/orchestration.md owns roles and sizing,
one worktree per issue under ~/orca/workspaces/ersatztv/, the landing order with
the review loop inside the worktree before the single push, and the merge through
the consent hook. Three Workflow scripts encode it: a picker over
scripts/select-queue.sh with two refuters, an issue-build pipeline (claim, recon,
implement, gate, cold review with a cross-family runner for the rubric's risk
classes, fix loop, finisher), and a resume pipeline for a paused branch.
scripts/e2e-gate.sh serialises live-E2E across worktrees because e2e-local.sh
refuses concurrent runs.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
`docs.no-session-narrative` reaches every durable artifact, but its detector scanned only
`docs/**/*.md` and root markdown, and nothing had ever swept the rest. The issue named four sites
from one grep and called them a floor. Deriving the population instead — a whitespace-joined sweep
over every tracked file outside the detector, for the detector's own phrasings plus the attribution
and review-round class #812 found — gave 453 sites in 108 files at `fb5592971`, and a second pass
for phrasings the first list missed (hyphenated `round-N`, "an earlier version", "the reviewer
proved") added residuals in the same files. Every site was classified with #812's three
dispositions (CUT / SEVER / KEEP with its sub-kind) under the who-benefits test; the per-site
manifests are on the PR. The rejected designs, tested-and-rejected fixtures, measurements and
traps stay; the attribution of who found them and the round in which they were found go.
The detector's population grows to `.claude/`, `.gitea/`, `.husky/` and `scripts/` regardless
of extension, minus the detector and its own test (whose fixtures ARE the phrasings) and minus
`scripts/tests/fixtures/` (test data, including decision-record copies — the same reasoning as
the records' own exemption, and what keeps the record's depth measurement true), and `--all`
lists tracked REGULAR files only — a symlink's content is its target and a gitlink has none. The #812
argument for leaving `docs/superpowers/**` in the population runs the other way here: `--diff`
sees only ADDED lines, and 287 of the 453 sites were under 30 days old — this corpus is where
narrative is being added, so the advisory nudge has reach. Density agrees: 56 line-mode hits over
the 113 regular files the predicate admits, against 9 over 66 docs files before #812. `web/` and C# stay out on the same
measurement (3 of 74 PATTERNS-matching sites, ~4,600 files). The predicate did not grow: PATTERNS
matched 74 of 453 sites, and widening the word list to the attribution class is the treadmill
the withdrawn parity test ran on. The population oracle is restated over segments with the new
arms, the synthetic cross product gains the process heads and non-markdown extensions, a fixture
witnesses that a tracked symlink is neither scanned nor counted, a `.py.bak` axis separates a
by-name exemption from a `startswith` over the same tuple, and eight mutants (drop the process
arm, drop the by-name exemption, exempt by `startswith`, drop or add a prefix, drop the fixtures
exemption, list only markdown, drop the symlink filter, test the mode per row instead of per
path) each
redden it. A pre-existing silent drop in `--diff` goes with it: git tab-terminates a `+++`
filename that contains a space, and the kept tab made `is_scanned_path` refuse the file with no
notice — fixed, with a positive control and its own mutant.
Code is unchanged by construction, measured per file type against `origin/main`: Python modules
are AST-equal with docstrings stripped, except `#` lines inside the embedded fixture programs
(string literals) of three test modules; workflows differ only in `#` lines inside `run:` block
scalars; shell, C#, TypeScript and jq are equal with comment lines stripped. The stated
exceptions: the detector and its test, 26 vitest titles that carried review-round or severity
labels or a reviewer attribution (call sites whose title changed — every changed title line
walked back to its `it(` / `it.each(...)(` anchor, so a `' + '` concatenation counts once), two
registry note strings and the mutation manifest's prose fields. scripts/tests: 1565 passed.
Web: lint, typecheck, 1319 tests green. Closes#876.
Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PEcBoFw7ctrf3Nb7R7x7wk
Finishes the 1.25.4-dated CI claim sweep #747 deliberately left incomplete (#869), and answers #893 from the Gitea v1.27.1 source instead of inferring it from a header. Docs and comments only - zero non-comment changes in scripts/ and .gitea/.
Population derived with `git ls-files`, not from the issue's item list: 17 files, 43 occurrences of `1.25.4`, against the 4 items #869 named.
Re-established on 1.27.1: the `creator`-attribution claim the H10 allow-list rests on (4 merged heads, both endpoints); the scope enum (no `status` scope); the `reqRepoWriter(unit.TypeCode)` gate; the `write:package` 403 (live probe with a read control 200 and a write control 201, throwaway repo, artifacts deleted); the absence of any REST cancel route (from source, which a 404 alone cannot establish); and `pull_request`/`pull_request_target` definition resolution.
#893: `/statuses/{sha}` does NOT drop rows after pagination. `getCommitStatuses` appends unconditionally and its only filter is a SQL WHERE in the same query as the LIMIT/OFFSET, so an empty page really is the end, `page_statuses` terminating on its first empty page is safe, and the asymmetry with `count_pr_mutations` is correct - recorded with its reason and a date so it is not tidied away.
Corrected rather than re-dated: the `--depth=1` no-merge-base claim was filed against the wrong axis (a git property, re-probed on git 2.55.0), and `enable_bypass_allowlist` postdating 1.25.4 had an issue body as its only provenance.
Five cold review rounds plus a cross-family Codex pass. They caught a wrong MECHANISM for `creator: null` (it is `CreatorID == -2`, not `== 0`), an evidence count that straddled the upgrade, and a reason for not re-probing MCP `cancel_run` that was invented - all fixed, final verdict CLEAN.
fixes#869fixes#893
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Whe75djeAEuZpdNk6KU7No
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Every hook under `.claude/hooks/` assigned `ETV_HOOK_FIRE_LIB` from `${CLAUDE_PROJECT_DIR:-<self>}`
and then `. `-SOURCED it. Sourcing is execution, so a file of that name in an env-designated tree ran
as code inside the hook before stdin was read and before it could decide anything. Measured on the
merge gate before #858 fixed that one hook: a decoy tree's copy printed an `allow` and exited 0.
Reachable without an attacker, because husky is a different launcher: `.husky/pre-push` invokes
`./.claude/hooks/…` relative to the PUSHED tree, independent of the variable, so a push from one
worktree while the environment names another sources the other tree's code into a gate.
Sweeps the remaining twelve hooks together (population derived from `git ls-files`), reconciles the
second resolution inside `scripts/hook-fire-log.sh` itself, and requires the root to OWN the sink
(`-ef`, not `-e`). The static guard pins the preamble BYTE-FOR-BYTE — a withdrawal, after a lexical
rule was defeated by five successive shapes.
Also pins two arms of the checker that were unsubsumed AND unpinned: the begin call's presence and
its missing stdout-mode token. `…_LOSES_its_instrumentation_…` looked like their proof and was not —
it asserts only that the fault list is non-empty, and a stripped hook trips four arms, so deleting
either left the suite green.
fixes#891
refs #858, #859, #776
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UYNbVwgVszv6Pum7ZuGd75
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
`ci-image.yml`'s `on.push.paths` decides which pushes to `main` publish a toolchain image;
`ci-image-pin`'s `git log` pathspec decides what the pin must name. #744 removed the shared
self-reference that kept them in step, leaving the agreement carried by three prose comments, and
divergence is silent and green in the dangerous direction.
The guard derives both lists from the workflow documents and compares them for set equality in both
directions. The comparison is deliberately narrow: it accepts a publish entry spelled exactly
`<dir>/**` against a pathspec entry spelled exactly `<dir>`, segments restricted to
`[A-Za-z0-9._-]`, and raises on every other spelling rather than deciding what that spelling would
have selected.
That narrowness is the substance. Measured against Gitea 1.27.1's own in-tree compiler
(`modules/actions/workflowpattern` -> `modules/glob.CompileWorkflow`) and real git: a bare
`docker/ci` in `paths:` compiles to an anchored `^docker/ci` and selects none of the directory's
contents while the git pathspec `docker/ci` selects all of them; `<file>/**` matches nothing while
the pathspec `<file>` tracks the file; a leading `/` is literal to Gitea while git refuses it
outright. A canonicaliser mapping the two dialects onto one string form was built twice and defeated
twice, each repair surfacing another spelling, so it was deleted rather than extended per
`testing.verification-code-needs-its-own-proof`.
The guard also asserts from the git index that each named path really is a directory, since
`<file>/**` and the pathspec `<file>` spell the same string; takes the pathspec from the `git log`
assignment rather than any `git log` in the job; refuses a `<<` token on a code line (a herestring
excluded) and a second bare `--`; and treats an absent and an empty `paths:` alike, because Gitea's
`Skip` returns false on an empty sequence, so `paths: []` filters nothing and every push publishes.
The docstring states the boundaries rather than implying coverage: the guard compares the pathspec
the pin job writes and does not establish that the staleness comparison consumes it, and a descendant
whose path below `<dir>` contains a newline is matched by the git pathspec but not by the publish
pattern.
Verified by nine independent cold-review rounds, none of which found a false green; the last fuzzed
27,720 publish/pathspec pairs against a port of the deployed compiler and real `git ls-files`.
fixes#855
`docker/Dockerfile`'s web-build stage ran the SPA vitest suite with two hand-written
`--exclude`d spec paths. The stage is gitless twice over, and the suite has members
needing a git checkout or the git binary, so that list was a population nothing derives.
#883 added a third member without updating it; because `Build & push image (amd64)` is
`if: github.event_name != 'pull_request'`, the red was unreachable on a PR and landed on
`main` and the `v*` tag path. Every image build has failed since.
The list is removed rather than extended: the stage builds the SPA and does not test it,
and the suite runs once, unfiltered, in the `test` job that `build` already `needs:`.
`scripts/tests/test_image_build_delegates_the_spa_suite.py` holds the invariant in three
parts, because the first two together still certify a publish on which the suite never
ran. It PINS command text rather than parsing it: three earlier versions asked what a
command MEANS and were wrong nine times, and a partial match of `web/vite.config.ts` was
then defeated seven more ways, so both mechanisms were withdrawn rather than respelled.
The transferable rule, recorded in the guard and the decision record: a pin assumes it is
pinning the artifact that still DECIDES. Every route found was authority moving where the
pin was not looking — another file, another occurrence, another workflow, or a hook the
pinned command invokes.
Nine independent cold-review rounds, eight BLOCKED. 73-mutant development battery, 0
missed; one declared clause mutation harness-executed per suite.
fixes#887
Round 9 returned MERGEABLE with BLOCKER and HIGH empty. Every remaining item was a
sentence, and every one erred by UNDERSTATING the guard — which is the safe direction and
still worth fixing, because the decision record is what CLAUDE.md routes convention
lookups to.
The record's `rule:` still listed "`web/vite.config.ts`'s `test:` block" among the pinned
things — the very mechanism the previous commit withdrew — and named only `vitest.config.*`
as the outranking family, omitting `vite.config.js`/`.mjs`, which is the MEASURED attack
from round 7 (a `web/vite.config.js` ran the suite in the gitless stage with 1411 tests
green). That family went short in round 7 and again in round 8. This is
`enumerate-CLAUSES-to-close-a-sweep`: the survivors were phrased in a different category
(WHAT is pinned) from the retracted claim (HOW it is extracted), so sweeping for the
retracted words missed them.
Also: "any edit to this file reddens, including a comment" was an absolute and is
refutable — a reindent, added blank lines, tabs, and a form feed all stay green, because
`_normalise_lines` collapses whitespace. Restated as what is actually true (a line's TOKEN
sequence, a comment's words included) plus the reason the tolerance is currently inert:
this file has no template literal and no ASI-sensitive token outside a comment. And a YAML
single-quote escape had leaked from the frontmatter into the markdown BODY, where `''`
renders literally.
No code change; the guard is unchanged and still 73/0.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Adds the `v26.15.0` row to the release table in `docs/ci-cd.md`.
The tag goes on `736649b3b`, NOT on this commit and not on `main`'s head. Every
`Build & push image (amd64)` since `e8f80c42c` fails: that commit added
`web/src/api/completeAnnotations.guard.test.ts`, a third importer of
`virtual:etv-tracked-source-files`, without adding it to the hand-maintained
`--exclude` list in the Dockerfile's `web-build` stage — and that stage has no
git index, by construction (#887, claimed and in progress elsewhere). The guard
is behaving correctly; it refuses to fall back to a filesystem walk. Measured:
run 2459 on `736649b3b` ran the image job for 6m45s and published; run 2515 on
`cf5f42edf` died in web-build after 86s. `736649b3b` is therefore the newest
commit on `main` that can produce a release image.
Consequence recorded in the row itself: #880 (scheduling recurrence) slips to
the next release, since it merged after the break.
Release-boundary sweep (docs/ci-cd.md -> "Before cutting a release"):
- `decisions_validate.py` -> OK; 0 legacy-unmigrated records remain
- `build_decisions_catalog.py` -> no drift
- record ceiling: 45/218 over 60 lines (fraction 0.21, inside the blocking
0.02-0.25 band). The validator notes the 60 has drifted below the tail
boundary (p90=104, p95=142) and asks for re-derivation when convenient —
a maintenance signal about the constant, not a blocker for this cut.
`_normalise_lines` drops blank lines and collapses whitespace WITHIN a line, so a reflow,
an indentation change, and a change to the spacing inside a STRING LITERAL are invisible.
The first two carry no meaning; the third could, and does not here. Line order and any
token change are caught. All four measured.
Stated because the phrase 'pinned whole' invites a reader to assume byte equality, and a
reader who assumes that will not check the one case where it matters.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Round 8. Three consecutive rounds had each closed one SPELLING of the same marker match,
which is `every-blocker-was-one-mechanism-so-delete-it` at the count where it says
withdraw. Round 6 pinned a block; round 7 fixed `test: {` (two spaces); round 8 defeated
the repaired matcher four more ways — `test : {`, `"test": {`, and the same two for
`plugins:` — plus two that never touched the marker at all:
plugins: [react(), trackedSourceFilesPlugin()].concat([evil])
test: { …pinned… }, ...moreTest
`defineConfig` is the identity function in BOTH vite and vitest (read from the installed
tree), so a spread AFTER the pinned span simply replaces what the pin matched. No
respelling of the marker could ever have caught those: the defect was partial matching,
not the pattern.
So the file is pinned whole. 48 lines, nothing generates it, no marker to respell and
nothing after the span. One assertion replaces a bracket walk, a block extractor and two
uniqueness assertions — and catches all seven measured routes. Stated cost, which is the
same one every other pin here carries: any edit to that file reddens, a comment included.
This also retires a claim I made in a commit message AND in the inventory row: that the
two pins "share one bracket walk and cannot drift apart again". It was false when
written — the block extractor had its own inline copy and never called the shared helper.
Verified by spying on the call: the `test:` pin made zero calls to it.
Also closed: the second exec-form branch in `copies_spa_source` was DEAD code (measured by
poisoning it with a `raise` — every test stayed green), since the round-7 reorder rewrites
`remainder` before it; the exec-form `--from` case that reorder was written for had no
STANDING test, only the development battery, and now has three; and
`OUTRANKING_CONFIG_PREFIXES` matched nested paths, so a `web/e2e/vite.config.ts` — and
`web/e2e/` is a real directory — would have reddened with a message asserting it outranks
the pinned config, which is false. Direct children of `web/` only.
Battery 68 -> 73, 0 missed.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
It counts brackets without understanding string literals, so a `]` inside one ends the
span early. The direction is what matters and it was measured: the truncated span does
not equal the pin, so the outcome is a false RED, never a false green. Stated rather than
fixed — parsing TypeScript to do better is exactly the predicate this file withdrew.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Round 7. BLOCKER empty; three real holes, two of them introduced by round six's own
fixes, which is this repo's recorded `each-fix-round-introduces-an-adjacent-defect`
happening inside a single commit.
**`web/vite.config.js` outranks the pinned `.ts` for `vite build`.** Read from the pinned
tarball rather than remembered: vite 8.1.3's `DEFAULT_CONFIG_FILES` is
`["vite.config.js", "vite.config.mjs", "vite.config.ts", …]`. So a `vite.config.js` whose
plugin shells out to the suite runs it in the gitless stage — measured with all 1411
tests green. The governing rule was already written down one family over ("pinning one
config is worthless while a second can outrank it") and had been applied to
`vitest.config.*` and not to this one. The refusal now covers both families.
**The decoy fix counted one SPELLING of the marker.** `text.count("test: {")` is defeated
by writing the real block `test: {` with two spaces beside a decoy that matches exactly:
count is 1, the comparison takes the decoy, and the live block filters out precisely the
specs #883 broke on. Now `re.finditer(r"\btest:\s*\{")`.
**The plugin pin added in that same commit shipped the identical decoy hole** it was
written next door to fix — a raw `text.count(PINNED_VITE_PLUGINS) == 1` with a decoy
above `defineConfig`. Both `vite.config.ts` pins now share ONE bracket walk and ONE
whitespace-tolerant uniqueness rule, so they cannot drift apart again.
PROSE, and this one is a false completion claim in my own previous commit message: I said
the `PUBLISH_ACTION`/anti-vacuity sentence and the singular "only an `ENV`" residual were
corrected. They were — in the record and the inventory row, and NOT in the guard
docstring, which is the artifact a code reader hits first. Both are now fixed there too,
the route COUNT is removed from the docstring and the record and kept in ONE place, and
the residual that stated its own false version before retracting it now states the
boundary once.
Also: the `--from=` branch never reached the JSON exec-form parser, so
`COPY --from=web-build ["/source/web", "/dest"]` left the receiving stage unpinned; the
revalidate arm of the gating `if:` is now described as a DEPENDENCY on
`ci-detect-already-validated.sh` (graded `MUTATION: NONE`) rather than as something
asserted here, since only the `docs_only` arm is; and the plugin-bodies residual now says
there are TWO plugins, `react()`'s being third-party and unmitigated.
Battery 64 -> 68, 0 missed. One of those four exists because the battery itself briefly
reported NOTHING and exited 0 after a bad splice deleted its `main()` — it now carries an
anti-vacuity assert on its own mutant count.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Hunting a fifth instance of "the pin assumes it is pinning the thing that still decides"
turned up two candidates that look like routes and are not, both measured rather than
argued:
* `setupFiles` is pinned by NAME while its CONTENT is not, which reads like the
package.json hole one level down. It is fail-NOISY: `process.exit(0)` at the top of
`src/setupTests.ts` makes vitest report `121 failed (121)`, not a green.
* `tsconfig*.json` shapes what `tsc -b` compiles, not what vitest collects.
Recorded because a reader who spots either will otherwise spend the same probe to reach
the same answer — and because the honest residual beside them is the one that IS open: a
dependency's own install script, reached through `npm ci` and `web/package-lock.json`.
That is a supply-chain concern wider than this guard, and it is named rather than claimed
covered.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Round 6 found three more false greens and named the class they share, which is worth
more than any of the three fixes:
* `web/vitest.config.ts` OUTRANKS the pinned `vite.config.ts` — closed in the previous
commit, found by probing vitest rather than reading about it.
* A DECOY first `test: {` block. The comparison took `text.index("test: {")`, so a copy
of the pin placed above `defineConfig` satisfied it while the real block was narrowed.
Exactly one is now required — the same assertion this file already made about the
gating step's NAME, for the same reason, not carried across.
* A `needs:` edge matched by bare job id. `needs:` resolves within its own workflow, so
a SECOND workflow publishing this Dockerfile while needing its own unrelated job
called `test` satisfied it. Now bound to `GATING_WORKFLOW`. (The reviewer downgraded
this to MEDIUM on measuring that `test_remote_state_inventory.py` forces a human to
classify any new workflow — so the hole is "the guard is blind", not "silent". The
forced review asks about remote state, not about whether the image is gated, so the
one-line fix stands.)
* A vite PLUGIN can shell out to the suite from `buildStart()`. The plugin ARRAY is
pinned; the plugin BODIES are a stated residual, mitigated because
`trackedSourceFilesPlugin` is deliberately lazy — a fact its own comment now marks as
LOAD-BEARING for the image build rather than leaving as an optimisation note.
THE CLASS: **a pin assumes it is pinning the artifact that still decides.** Every route
found so far is authority moving where the pin is not looking — to another FILE, another
OCCURRENCE in the same file, another WORKFLOW, or a HOOK the pinned command invokes. That
question is now written down for the next person adding a pin, because a list of four
instances is not what generalises.
Prose, all refuted by execution: the residual naming the uncovered COPY shapes was wrong a
THIRD time at the same site (`/source/web /elsewhere` IS recognised — only the destination
is renamed — and the file's own test 700 lines below said so); "only an `ENV` is
unmodelled" was an absolute and is now a list; "Reach: N mutants, 0 missed" is restated as
a DEVELOPMENT BATTERY, since it is not in the repo, nothing re-derives it, and an
independent battery found misses against an earlier head; and `PUBLISH_ACTION` was claimed
covered by anti-vacuity, which proves the selector is non-empty and cannot prove it
complete.
Battery 61 -> 64, 0 missed.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Pinning `web/vite.config.ts`'s `test:` block is worthless while a file that takes
precedence over it can simply be added. Vitest resolves `vitest.config.*` (and
`vitest.workspace.*` / `vitest.projects.*`) BEFORE `vite.config.*`.
MEASURED, not read: dropping a `web/vitest.config.ts` carrying
`include: ['nope/**'], passWithNoTests: true` beside the pinned file made `npx vitest run`
report "No test files found, exiting with code 0". The gating step would be green having
run NOTHING — worse than the filtered run ersatztv#887 removed, because a filtered suite
at least reports on what it ran.
The construct is refused rather than modelled: no such file exists, so the guard asserts
none appears. Its population is the git INDEX, which is right and worth stating — an
untracked config does not exist in a CI checkout either, so the mutant proving this has
to STAGE the file. It failed to redden until it did, which is the correct behaviour
demonstrating itself.
Route count five -> six -> seven -> eight, wrong at every previous count, so it stays a
running total with its history attached. Battery 60 -> 61, 0 missed.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Round 5 found a BLOCKER, and it is the sharpest kind: I made the exact mistake I had
described one screen earlier. `test_only_the_PINNED_npm_SCRIPTS_run_vitest` SELECTED the
scripts to pin by asking whether their body contained the literal `vitest` — a selector,
the category this file calls the worst-behaved because going short is silent — and then
its docstring claimed "going short is caught by the equality below", which is false: a
script the substring misses is absent from the compared map, so the equality still holds.
Four one-line `web/package.json` edits, none of which spells `vitest`, each put the suite
back into the gitless stage with the whole guard green: `"build": "npm run test -- --run
&& …"`, the same via `npm t`, and the `prebuild` / `preinstall` LIFECYCLE HOOKS, which
npm runs for `npm run build` and `npm ci` without anything naming them. That is #883
verbatim, through the route round 4 identified and the previous commit reported closed.
The fix is the one the file's own vocabulary prescribes: pin the WHOLE script map. A
script that does not exist cannot be a lifecycle hook, and one that changes is not equal.
The category disappears rather than being widened by two entries.
ALSO CLOSED, all measured:
* `web/vite.config.ts`'s `test:` block is now pinned. `npm test -- --run` collects what
that file says, so `test.exclude` is where a filter would now naturally be written —
it is the only place left after this change removed the Dockerfile's. Three mutants
narrowed the gating suite through it with the step's own command unchanged.
* A step-level `shell:` and a job-level `defaults:` each override the pinned workflow
default. Both forbidden.
* `test_no_run_BODY_builds_or_pushes_an_image` is RESTORED — I dropped it in the parser
withdrawal, and a job publishing via `run: docker build … && docker push …` was then
outside the action-derived population with anti-vacuity none the wiser.
* A leading-slash context copy (`COPY /web/. ./web/`) was not recognised.
* The sweep gains `yarn test`, `pnpm test`, `bun test`.
The residual naming the uncovered COPY shapes was wrong for the SECOND consecutive round —
it named `COPY --from=X /source/web /elsewhere`, which is covered (only the destination is
renamed). The real gaps are an ANCESTOR source (`/source` brings `/source/web` along) and
`/source/.`. Both measured.
Route count: five, then six, now seven. It has been wrong at every count, so it is now
stated as a running total with that history attached rather than as an enumeration.
Battery 51 -> 60, 0 missed.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The previous commit said "five such routes" in three places. Checking rather than
restating found six, and the sixth is one this guard must NOT close itself: the gating
job runs in a `container:`, whose image decides which `npm` exists at all. That is
already pinned by `test_ci_image_pin_population.py`, so it is CITED — two guards on one
condition mask each other (ersatztv#685), and the way to find that out is to delete one
and look for a red, which nobody does.
Also measured rather than assumed: an INDIRECT script chain (`"test": "npm run inner"`
with `inner` running vitest) needs no clause of its own. The set-equality against
`PINNED_VITEST_SCRIPTS` reddens on it, because `inner` mentions vitest and `test` no
longer does — verified across four scenarios, three red and one green.
An enumeration is a claim like any other. This one was written from memory of what had
been fixed rather than from the code, and it was short by one.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
A fourth cold review attacked the pin itself. BLOCKER empty, the mechanism upheld, and
every finding was prose claiming more than I had measured — plus one one-token gap that
was live.
THE ABSOLUTES, refuted by execution and now corrected in all three places they appeared
(guard docstring, `docs/guard-inventory.md`, the decision record's `rule:`):
* "A pin cannot produce a false green." True only in the trivial reading. A pin is
immune to a different SPELLING of the command — the entire class that defeated the
parser nine times — and is NOT immune to the same text MEANING something else. Two
mutants re-armed ersatztv#887 through `web/package.json` alone: `RUN npm run build`
executes whatever that file says, so `"build": "vitest run && …"` puts the suite back
into the gitless stage with every pin still matching, and `"prepare"` does it via
`npm ci`. Now pinned: exactly one script may mention vitest, and its body is fixed.
* Residual (1), "a stage that does not carry the SPA source is unpinned — correct,
since without `web/` there is no suite there". False. The boundary is what
`copies_spa_source` RECOGNISES, which is narrower than "has the suite available".
Restated, with the case still outside it named: a stage copy that RENAMES the tree.
* The substring sweep's "never a false green". Its reported failures are false reds;
what it fails to REPORT is not. `SUITE_MENTIONS` is a hand-written SELECTOR — a third
category beside population and pin, and the worst-behaved, because a population going
short is caught by an equality and a stale pin reddens loudly, while a selector going
short is silent. It was short by exactly one entry: `npm t`, npm's own alias, which
this guard already names among the spellings that defeated the parser. A stage
running `npm t -- --run` escaped it. Fixed, and the category is now named.
ALSO CLOSED: `run: |` -> `run: >` folded the two-line body into one command whose
whitespace-normalised text was byte-identical to the pin, so the marker script swallowed
the suite as its arguments — the body is now compared LINE BY LINE, since a newline
separates two commands. A SECOND step named `Test SPA` inherited the exemption both the
pin lookup and the sweep key on; exactly one is now required. And `COPY web*/` — a glob
that matches `web/` — was read as not carrying the source, leaving the receiving stage
unpinned.
Battery 45 -> 51, 0 missed. The remaining meaning-change route, an `ENV` rewriting `PATH`
so a pinned `RUN` resolves a different `npm`, is not modelled and is recorded as a
residual rather than implied away.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
A pin rejects any command it does not equal, so the remaining attack is to change what
the pinned text MEANS. Four such vectors, found by attacking the new mechanism rather
than reading it; the first was MEASURED escaping and the rest are the same class:
* A NEW stage taking the SPA source across with `COPY --from=web-build /source/web`
and running the suite there. `copies_spa_source` excluded every `--from=` copy on the
grounds that a stage copy is not a context copy — true, but it can still carry the
SOURCE TREE from a stage that has it. The receiving stage was therefore unpinned and
unchecked, which is exactly the false-NEGATIVE direction this file's own residual
warns about. A stage copy now counts when its SOURCE has a whole `web` path segment,
which keeps the built-artifact copy this repo actually makes
(`/source/ErsatzTV/wwwroot/app/.`) correctly out.
* `working-directory:` moved off `web` — `npm test` somewhere else runs a different
package, or none, with the pinned command text unchanged. Now pinned.
* `defaults.run.shell` changed from `bash` to `sh`. `bash` here is `bash -e`, which is
what makes a failing command fail the step; changing it changes whether a red suite
blocks the image without touching the step at all. Now pinned.
* A `SHELL` instruction in a pinned stage, which redefines what every later `RUN`
executes. Refused outright rather than modelled — there is none in this repo, so the
honest move is to reject the construct, not to reason about a replacement
interpreter.
Battery 41 -> 45, 0 missed; the stage-copy mutant reddens three assertions. The residual
list gains the two cases that remain in this class and are NOT covered: a stage copy that
RENAMES the tree on the way in (no `web` segment in its source), and an `ENV` altering
`PATH` so a pinned `RUN` resolves a different `npm`. Naming them is the point — the
previous rounds' residual lists read as exhaustive while omitting the largest holes.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Round 4, and the third cold review found the same mechanism failing again, so it is
removed rather than patched a tenth time.
WHAT KEPT BREAKING. Three versions of this guard asked "does this command RUN the suite,
and can its failure be swallowed?" of arbitrary shell text. That predicate was wrong NINE
times across three review rounds, and twice a clause added to remove a FALSE RED opened a
FALSE GREEN on the guard's headline assertion:
* heredoc bodies were skipped as data, but BuildKit EXECUTES `RUN <<EOF` — and the
opener regex also fired inside quotes (`echo "tags<<__EOT__"`), which blinded the
whole-file scan over the last 303 lines of docker-build.yml. Wrong in both directions
at once, and measurably live on this tree.
* `shlex.shlex` does not clear `commenters` the way `shlex.split` does, so `#`
truncated a command mid-word — including the live `${#reports[@]}` idiom — and made
this file's own stated residual false.
* compound punctuation (`);`) welded two commands into one segment.
* `npm t`, `./node_modules/.bin/vitest`, `pnpm vitest`, `yarn vitest`,
`node …/vitest.mjs`, `timeout …`, `su -c …`, `if npm test; then` — all invisible.
* `true || npm test` counted as the gating run while never executing it.
* `continue-on-error: ${{ … }}` passed a check written against two literals — a
presence test that cannot see polarity, fail-OPEN in the one direction that matters.
WHAT REPLACES IT. Nothing in the file decides what a command means any more. The commands
that may run in the two risky places are PINNED as text: the `RUN` lines of every
SPA-carrying Dockerfile stage, and the gating step's `run:` body and `if:`. A suite run
re-added in ANY spelling is simply not equal to its pin — the pin does not have to
recognise a spelling in order to reject it. A pin cannot produce a false green, only a
false red, and a false red is a human reading a diff they should have read anyway.
The population/pin split is the load-bearing distinction, and it is now stated in the
inventory: a POPULATION decides what is CHECKED, so a hand-written one goes silently
short; a PIN decides what is EXPECTED, so a stale one goes loudly red. Only the second is
safe to write by hand. Populations stay derived from the git index.
Two premises that were prose are now assertions: the publish step keeps its own
`docs_only` gate (without it, a docs-only push skips the suite and publishes anyway), and
no step other than the pinned one mentions the suite — a SUBSTRING sweep, deliberately
not a predicate, whose failure mode is a false red asking someone to look.
41 mutants, 0 missed, including all nine spellings above and the three from the previous
round. Exactly ONE is declared in `mutation_manifest.py` and re-executed every suite; the
other 40 were witnessed during development and are NOT standing — stated in the row
rather than left to be assumed.
Also fixed: the truncated sentence the round-2 rewrite left in the Dockerfile comment,
and the `web/src/api/*.guard.test.ts` glob, which over-claimed — it matches three files
and only two of them need git.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Both were measured, not reasoned, and both are the shape this repo keeps recording — the
fix round introducing an adjacent defect, and a unit test using a simpler input shape
than the real file has.
`npm test -- --run && echo ok || true` reported NO suppression. `&&`/`||` chain across a
whole list, so when the suite fails the `&&` right-hand side is skipped and the `||`
right-hand side runs: the list exits 0 and the suite's failure is swallowed even though
the `||` is not adjacent to it. The detector looked only at the separator IMMEDIATELY
after the suite segment. It is now scoped to the `;`-delimited list, which also catches a
backgrounded `npm test &` (status never awaited) and `( npm test ) || true`. A `;` ends
the list and resets, so `npm test; other || true` stays clean — that `||` is about the
other command.
`--exclude 2 > log` reported `['--exclude']`, losing the filter's own value: stripping
redirections as a PRE-PASS let the file-descriptor rule claim the `2` before the flag
could. Redirections are now consumed inside the walk, after flag values are taken.
The mutant battery grew from 17 to 24 and is 0-missed. The `docs/guard-inventory.md` row
now states the count and, explicitly, the grading: exactly ONE of the 24 is declared in
`mutation_manifest.py` and re-executed every suite; the other 23 were witnessed by hand
and are NOT standing. That is the same footing `pageSizeCallSites.guard.test.ts` states
for its nine, and saying so is the difference between evidence for the reach and a claim
of a per-run proof.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Both independent reviews (Codex GPT-5.6 cross-family, and a cold Opus agent in an
isolated worktree) returned BLOCKED. Both independently confirmed the CI path itself is
sound — neither found a route that publishes an image on which the suite never ran — so
every finding is about the guard's reach, plus one factual error in the prose.
THE STRUCTURAL ONE. The guard asserted a `needs:` edge EXISTS, never that it is load
bearing. Since this change deletes the in-image run, that edge is the only remaining
layer, so `continue-on-error: true`, `if: false`, a job-level `if:`, `npm test … || true`,
a pipe into `tee`, and `set +e` each certified a publish over a red suite with every
assertion green. `test_the_gating_suite_run_is_NOT_ADVISORY` closes all six.
A filter written into `web/package.json`'s script body was invisible at the call site:
`"test": "vitest --exclude x"` with a workflow saying `npm test -- --run` is a filtered
gating run reading as clean — the removed defect, one level down. `vitest_scripts()` now
derives each script's own narrowing arguments and `suite_args` prepends them.
PARSER REACH, every case measured rather than argued. `shlex.split` yields `lint&&npm` as
one token, so unspaced `&&` and `;` re-adds were invisible; `shlex` in punctuation_chars
mode splits them. Added: `sh -c` payload expansion, `npm --prefix`/`npx -p` flag skipping,
`xargs`, heredoc bodies as DATA (a `cat > f <<'EOF' … npm test … EOF` block counted as a
real run), `ADD`/JSON-form/no-trailing-slash `COPY` in `carries_spa_source`, and
redirections no longer read as spec filters. `--root` and `--config` moved to the
narrowing set: both change which specs vitest collects.
A FACTUAL ERROR, in five places including the mutation `expect`: "the build context is
`web/` + `design-system/`, so there is no `.git`". The context is the repository root
(`context: .`) and `.dockerignore` does not exclude `.git`. The true statement is about
the STAGE, which copies only those two directories. The conclusion survives — bookworm
slim has no git binary either — but a reader who checked would have found `.git` in the
context and concluded the note was stale.
ONE FINDING WAS MINE, from the mutant battery rather than from either review, and it is
the reason the battery exists: `failure_suppressions` tokenised the whole multi-line
`run:` body at once. A newline is not a shell separator, so a realistic two-line step —
the `ci-step-ran.sh` marker line, then the suite — merged into ONE segment whose head was
the marker script, and three suppression mutants passed while my single-line unit test
was green. It now works per logical line, and the regression test uses the two-line shape.
17 mutants, 0 missed, each caught by the intended assertion; baseline green. The
`docs/guard-inventory.md` residual list is rewritten as MEASURED reach — the previous one
was wrong rather than merely short, which cold review rightly called worse than silence.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
`docker/Dockerfile`'s web-build stage is gitless twice over — the build context is
`web/` + `design-system/` so there is no `.git`, and `node:22-bookworm-slim` ships no
git binary. Members of the SPA suite need one or the other, so running the suite there
required naming the ones that cannot run. That list was a population nothing derived:
#883 added a third member without updating the hand-written pair of `--exclude`s, and
because `Build & push image (amd64)` is `if: github.event_name != 'pull_request'` the
resulting red was unreachable on a PR. It landed on `main` and on the `v*` tag path
instead — every image build failed, `:latest` stopped being republished, and a release
cut would have failed at the image build.
Adding a third `--exclude` re-arms the trap, so the list is removed rather than
extended: the stage now lints, typechecks and BUILDS the SPA, and the suite runs once,
unfiltered, in `docker-build.yml`'s `test` job on a real checkout. `build` carries
`needs: [test, migrations, scan]`, so no image is published past a red suite.
`scripts/tests/test_image_build_delegates_the_spa_suite.py` holds both halves — the
negative one alone would be satisfied by deleting the `needs:` edge. Three populations,
all derived: tracked Dockerfiles and workflows from the git index, and which npm scripts
ARE the suite from `web/package.json` (so `test` is in and the Playwright `test:ui-e2e`
is out, with no exemption list). Publishing jobs come from the `docker/build-push-action`
step and the Dockerfile each builds from that step's own `file:` input, which is why
`ci-image.yml` is out of scope by derivation rather than by an entry that would outlive
its reason.
Four mutants witnessed red, each by the intended test: a filtered suite run put back
into the Dockerfile, the `needs:` edge deleted, and the gating run narrowed in both the
block and the single-line `run:` step forms.
Refs: #887
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
#859 was filed as a wrong STATED CAUSE. It was masking a live false-open in the merge gate.
Gitea reports a GLOB branch-protection rule with an EMPTY `branch_name` — the canonical
name lives only in `rule_name`. Measured 2026-08-30 on a scratch repo against 1.27.1.
jq's `//` fires on null and false but NOT on `""`, so `(.branch_name // .rule_name // "")`
resolved every glob rule to the empty string — a name with no metacharacters — and the
glob test, the entire basis of the classifier's undecidable-first ordering, never saw it.
Measured on the predecessor: glob `m*` (not requiring review-verdict/h10) beside plain
`main` (requiring it) resolved to `exact` on `main` and AUTO-GRANTED a scheduled merge,
while Gitea — ordering by Priority then plain-name-ness — may be applying `m*`. That is
#622's hole, reached through the ordering written to close it. Mirror case: a glob alone
resolved to `none` and DENIED about a rule that provably governs the base.
A name is now a non-empty string. Each field resolves to a NAME, a SKIP (absent/null/
empty — fall through), or POISON (present, wrong type — poisons whichever field carries
it). A rule with no usable name is a distinct `unreadable` verdict with its own operator
cause, instead of feeding `none`, whose whole authority is "the full rule list was read
and none matches". The short-circuit is STRUCTURAL: jq binds `as` eagerly, so the flat
form still evaluated `offs`/`nonascii` on the bad name and died before reaching the arm
meant to prevent that.
Also #859: `branch_protections` is fetched ONCE per run, not twice. The round trip is the
smaller half — it is mutable config, so two reads can disagree and the two arms then
decide about different repo states with neither able to notice.
#858: `verdict_script` resolves from `$repo_root`, not `$CLAUDE_PROJECT_DIR`. And the
finding that mattered more — `ETV_HOOK_FIRE_LIB` is `. `-SOURCED, so it is CODE running
before stdin is read and before `decide` exists. A first draft exempted it as "telemetry,
not a predicate"; cold review refuted that by execution: a decoy hook-fire-log.sh in an
env-var-named tree printing an allow and exiting 0 GRANTS THE MERGE, bypassing every
check. Classify a path by how it is CONSUMED, never by what it is called. This hook's copy
is self-located; the other twelve are #891 (high/security), which records the reachable
case — husky launches the prepush hooks by RELATIVE path, so the two roots diverge there.
check-required-contexts.sh gains an array-type gate (a JSON object previously printed
`nomatch`, a positive claim about server config from a body it cannot consume).
Verification: 1377 passed / 2 skipped; 11 declared mutants, 11 detected, disjoint
reddened sets; classifier executed across jq 1.8.2 and 1.6 with identical results; both
env-var tests ship a negative control, because the passing outcome is also what an inert
decoy produces.
Four cold review rounds plus a bounded prose check. Every round found defects the
previous round's fixes introduced — a type conflation that re-opened the auto-grant, a
comment asserting the opposite of the line its own commit changed, and a corrected
sentence whose identical twin survived in the same diff.
Docs: new record `process.hook-resolves-inputs-from-repo-root`; both inline sites cite it
rather than arguing it twice. docs/remote-state-inventory.md's row for the second read
updated. Follow-ups filed: #891 (the other 12 hooks), #895 ("all N tests green" claims).
fixes#858fixes#859
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
2026-08-30 11:33:50 +00:00
timothytimothyClaude Opus 5 (1M context) <noreply@anthropic.com>
`count_pr_mutations` treated an empty page past page 1 as proof it had reached the end of the PR
timeline. Gitea does not mean that: `ListIssueCommentsAndTimeline` applies the LIMIT/OFFSET in
`FindComments` at the DATABASE level and filters AFTERWARDS, dropping `CommentTypeCode` rows and
inaccessible cross-references into a nil slice that serializes as bare `null`. A page of 50 inline
review comments is byte-identical to a page past the end while later pages still hold events, and
rows are ASCENDING, so the events a fence looks for are the furthest from page 1. Fifty comments,
which a PR author can create on their own PR, truncated both walks at the same place: both counts
agreed, the sha comparison agreed, and an ABA force-push yielded an exemption `success` over a diff
no single head justified.
The walk no longer infers the end from an empty page BEFORE its cap. Such a page is skipped; the
loop reads every page to its 20-page cap and trusts the counts only when the LAST page came back
empty. An empty FIRST page and any unreadable shape still end the walk untrusted.
NARROWED, NOT CLOSED, and the docs say so in one unit: the page-20 terminator is still trusted for
the same unprovable reason, so the defeat now costs a timeline of over 1000 rows rather than ~100,
with the same 50-row filtered block pinned to offsets 950..999.
Measured at Gitea 1.27.1, ruling out the cheaper fixes: `X-Total-Count` on this endpoint is the
post-filter length of the PAGE, not a total (`?limit=1` returns 1 on a 14-row timeline), while
`/activities/feeds` returns a true total; `limit` clamps to 50; the only query params are `since`,
`before`, `page`, `limit`, so the paged and serialized sets cannot be made to agree.
Also: each page bounded `--connect-timeout 5 --max-time 15` and retried once, mirroring
`page_statuses`, because the walk went from ~2 requests to a fixed 20 and the third call site runs
after the exemption `success` is posted. Costs stated rather than hidden — worst case 40 requests
and 20 sleeps, wall-clock pessimum 620s per walk, and the suite roughly doubled (202s -> 474s).
Seven tests, each mutation-witnessed red; three reproduce the defeat against the shipped predecessor.
Two independent cold reviews plus a re-review of the fix: no Blocker or High in the code. Their real
finding was prose claiming the hole was closed, and cost arithmetic wrong twice. One reviewer claim
was refuted by execution.
Fixes#870
Refs: #803, #706, #664, #751, #893
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
The three recurrence arrays are read CONJUNCTIVELY by
AlternateScheduleSelector.GetScheduleForDate, so an empty set matches no date.
`?? []` on an omitted array therefore returned HTTP 200 while storing an
alternate-schedule or template item that could never apply, silently -- while the
read side (#823) already read a NULL column as the All*() sets.
Absent and explicitly-empty are two different requests and get two answers:
ABSENT (missing, or explicit null) normalizes to AlternateScheduleSelector.All*(),
the same symbols the read side substitutes; EXPLICIT [] is rejected with a 422
naming the consequence, via RecurrenceSetBounds called from both replace handlers.
The rejection lives in the handlers, not the controller, because
api.ffmpeg-profile-numeric-bounds' "accept an UNCHANGED bad value" rule binds
hardest here: both PUT paths are whole-list replaces, so rejecting a pre-existing
empty set would make every OTHER item in the list uneditable. That comparison
needs the stored row. The validated set is derived from `incoming`, so the
highest-Index catch-all -- whose recurrence the handler discards -- is excluded by
construction.
Verified: full ErsatzTV.Tests suite green; three mutation proofs with disjoint
reddened sets; live-E2E against a real instance confirmed an OMITTED property
round-trips as unrestricted (the Newtonsoft missing-property chain unit tests
cannot reach), an explicit [] returns the 422, and [] on the catch-all is accepted.
Cross-family cold review BLOCKED the first implementation with 3 findings, all real
and all fixed; re-review returned MERGEABLE.
Follow-up #894 filed: the SPA can still build the empty state the server rejects.
fixes#880
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The round-9 cross-family review found no Blockers and no Highs, and independently confirmed
the clause deletion it was asked to check. What it did find is that round 9 removed FIFTEEN
test definitions and added four — a net loss of eleven — where the commit message claimed
two. Verified against the parent: 227 definitions before, 216 after.
The cause is mechanical and worth naming, because it produces a green suite: the round-9
edits replaced whole source RANGES (`s[:start] + new + s[end:]`) whose end anchor was the
next test rather than the end of the one being rewritten, so everything in between went with
it. The suite then passed because the tests were GONE, not because the code was right — the
exact shape this issue exists to prevent, reproduced in its own test file.
Among the casualties were round 4's proofs for two earlier BLOCKERS:
- `test_a_generic_PENDING_with_no_mark_also_becomes_the_sentinel` and its mutation, which
pin the no-mark downgrade covering every re-derivable write rather than only `success`;
- `test_a_MALFORMED_creator_FIELD_...` and its mutation, which pin a wrong-typed field
taking the fault route rather than reading as absent and licensing a re-derive.
Also lost: both `$own`-exclusion proofs, the no-op-repair skip proof, the id-asymmetry pair
(the reviewer's named example), and two write-failure propagation proofs.
All 13 unintended deletions are restored verbatim from the parent commit and ALL PASS against
round 9's code, so nothing had regressed — the harm was the missing evidence, not the
behaviour. The two deletions that WERE intended stay deleted: a test superseded by
`..._still_refuses`, and the positive control round 9 inverted.
Prose: the comment above the unreadable-element guard still argued a malformed neighbour is
safe noise once the target row was found, eleven lines above code that now refuses
unconditionally — two adjacent blocks giving opposite accounts of one rule, and the stale one
licenses reinstating the Blocker.
Refs: #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The round-8 cross-family review found two more Blockers. Both are cases where a principle
this branch had already established was applied in one place and not the adjacent one.
## Sentinel text is not sentinel state
`ex_repair` and `ex_unverified` were set from the DESCRIPTION alone. A `success` carrying
`$REPAIR_DESC` verbatim — from a machine or an off-list account — therefore read as a
sentinel: the mid-run guard exited on it, and the mark's already-there test matched it and
returned without POSTing. A green stood on an unreviewed head, on a first-push event with no
successor guaranteed.
This is the same reasoning that removed the "this job's own output" exclusion one round
earlier: a description is not provenance. It is not state either. Both sentinels this job
writes are `pending` by construction, so requiring it costs nothing.
## An unreadable neighbour cannot be shown to be unrelated
Round 8 refused only when NO readable target row was found, reasoning that a malformed row
beside a good one is noise. An element whose `.context` cannot be read cannot be shown to be
a DIFFERENT context — so it may be a mangled rendering of this head's own rejection, and the
one-row-per-context invariant that would rule that out is exactly what a schema-corrupt
response has already broken. The branch's own POSITIVE CONTROL encoded the failing case: a
scalar beside an off-list `success`, which this branch re-derived and greened where
`origin/main` errored on the scalar and posted nothing. That test is inverted, not adjusted.
The cost is a stall on any head carrying a malformed element — the correct direction for a
required check, since it withholds a green rather than granting one.
## Two clauses deleted rather than proved
Chasing a proof for the mark's repair promotion showed its three clauses were MUTUALLY
REDUNDANT: each alone produces the outcome, so no single-clause mutation could show harm.
Tracing why revealed that two are unreachable as a sole cause — a repair sentinel at the
first read sets `ex_repair`, which forces `desc="$REPAIR_DESC"`, and one arriving mid-run is
caught by the sentinel guard unless this run is itself writing that string. So they are
redundant rather than unprovable, and they are gone. One clause, one mechanism, one proof.
refs #849
Refs: #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The rebase onto #889 resolved a CLAUDE.md hunk in favour of upstream, which kept #845's new
clause and discarded #849's — leaving the file asserting that a rejection landing inside a
run's own write window is "a separate and still-open route". Both edits belong: they touch
one sentence for different reasons.
This message also repairs the TRAILER BLOCK for the whole branch, which CI caught and local
runs did not. Every commit here ended:
refs #849
Decisions-Edit: yes
Co-Authored-By: ...
Git parses only the LAST paragraph as trailers, so the blank line put `Decisions-Edit: yes`
in the second-to-last one and it was never a trailer at all — `git log --format=%(trailers)`
showed only the Co-Authored-By pair. `refs #849` without a colon disqualifies that paragraph
independently. `decisions_validate.py` arms its rationale-prose exemption from ANY non-merge
commit in the range, so one correctly-formed block repairs all nine.
Refs: #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The first cross-family review in five rounds (Codex/GPT-5.6, once its quota reset). It found
a Blocker four same-family rounds had missed, and REVERSED two of round 7's fixes — which is
the more useful result, because both were made in response to a review and both overshot in
the direction the finding pointed.
## The Blocker: dropping unreadable elements became "no verdict exists"
Round 3 added `select(type == "object")` so a malformed NEIGHBOUR could not kill the step.
When it drops EVERY element, `first // {}` yields `{}`, all `ex_*` read empty, and the job
concludes no verdict exists — so a docs-only PR walks straight to the exemption. Measured:
`{"total_count":1,"statuses":[7]}` posts `Exempt: docs-only change` here and posted NOTHING
on `origin/main`, which raised jq error 5 and aborted under `set -e` before any write. An
input on which this branch greens a head that `main` fails closed on, and if that scalar is a
mangled rendering of the head's human `failure`, the rejection is what gets greened.
The asymmetry is now the rule: a malformed row BESIDE one we did read is noise; a malformed
row where we found NOTHING is the only evidence there was. The absence conclusion has to be
earned over a list with no unreadable elements in it.
## Two round-7 fixes that overshot
- **The arms judged both snapshots.** Round 6's review said they judged `$pre_*` while the
POST replaces `$ex_*`; I made both veto, which is the mirror defect — an opening row since
REPLACED by a machine `success` still vetoed, so the arm left that success gating the head.
They judge the current row alone now. The opening snapshot keeps exactly one job: it can
make the write STRONGER, never suppress it.
- **The "this job's own output" exclusion keyed on the DESCRIPTION.** A description is not
provenance. Any workflow with `code: write` can POST a `creator: null` row and any
repository writer can POST one with a creator, either wearing this job's text — so masking
a human `failure` with a lookalike `pending` bought an abstention, and the successor
re-derived it as ordinary machine output with the rejection below its own mark. Removed;
the attempt is recorded because it is the tempting one, and there is no issuer field that
could make it safe.
## A guard that could not be reached, folded into the one that can
The repair veto turned out unreachable: an `$ex_desc` of `$REPAIR_DESC` with a different
`$desc` is caught by the mid-run sentinel guard long before an arm runs, and when `$desc` IS
`$REPAIR_DESC` the promotion writes the same string. Rather than keep a guard no fixture can
reach — or delete it on the strength of a check three hundred lines away — the invariant is
enforced where it is local and provable: the mark carries the strongest description any
snapshot shows, then declines to write what is already there.
`ci.exemption-provenance` still called the post-final-count window a PERMANENT forged green
in its `rule:` frontmatter and body; the post-POST re-count made it transient two rounds ago.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
A fourth cold review of the tip. No Blockers, no High: it enumerated every POST site and
every exit and could not construct an input where this branch writes a `success` that
`origin/main` would not.
## The arms judged the wrong snapshot
`mark_declined_row_if_any`'s three refusals all read `$pre_*` — the FIRST read — while the
POST replaces whatever row is CURRENT. So a reviewer's verdict arriving between the two
reads slipped past every refusal written to protect it: the base mismatch clears
`ex_attributable` so the mid-run abstain declines, `pre_creator` is empty so the allow-list
loop declines, and the arm marks a row nobody evaluated. Executed trace, control and case.
Both snapshots are consulted now, and either one vetoes.
Recovery was not free, which is why it mattered: the next run's reconciliation counts that
`Review-verdict:` row as buried and upgrades to the human-only sentinel — exactly the cost
the refusal exists to avoid.
The arm also marked this job's OWN ordinary machine `pending`. Every PR past its first run
carries one, so "kept off the commonest path in this job" was true only of a head with no
status at all. Scoped on the DESCRIPTION rather than on `creator: null`, which would also
exclude a machine `success` from another workflow — the row this marking exists for.
## Two comments that invited a bug
- One still described the round-4 REGRESSION as the intended behaviour ("a malformed row
reads as no creator, hence re-derived"), two lines below the block recording that it was
fixed. Adjacent comments giving contradictory accounts of one line, and the stale one
licenses reinstating it.
- The fault token's justification said "no Gitea status field contains a NUL". The token is
SOH (0x01). That is not pedantry: `$'\000…'` is the EMPTY STRING in bash, so an editor
correcting the code to match the comment would make every legitimately-absent field
compare equal to the token and send every clean head down the fail-closed route — the gate
would stall every PR.
## Docs
The record quoted a predicate that no longer exists (`[ "$ex_desc" != "$pre_desc" ]`, now
`$row_replaced`); `docs/ci-cd.md` stated the reconciliation witness unconditionally when the
code degrades to a description match where the server omits `id`; one of the six unproven
clauses carried a wrong `because` (the conclusion holds via `(.id | numbers) // -1` over a
validated array, not via the schema-fault route, which governs a different endpoint's row);
and the record's own counts read as a contradiction cold — 20 surviving MUTANTS collapse
onto 6 distinct CLAUSES, several clauses admitting more than one disarming edit. The
run-by-run provenance moved to the issue, where `docs.no-session-narrative` says it belongs.
Two existing mutation proofs lost their binding to the reworded clauses and failed loudly
rather than measuring the unmutated body, which is what that count assertion is for. Rebound.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The `mechanics:` field ended with the same sentence twice — the enumeration of the
unreachable clauses was appended without removing the tail it replaced. Found by the
mutation-sweep agent while reading the record it was checking its own results against.
In its place, the number that makes the technique worth its cost: 60 mutants, 40 red, 20
survivors, on a tree that had already been through three per-finding review rounds by two
model families.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Codex was unavailable for this round (usage quota), so the cross-family reviewer was
replaced by a same-family agent doing one mechanical job: enumerate every security-bearing
clause the diff adds, disarm each, and run the WHOLE suite per mutant. 60 mutants, 40 red,
20 survivors — a yield no per-finding review in this series came close to, because a review
looks at what the diff says it does and a sweep looks at what the tests actually pin.
## Proved (nine)
- the description type test in the RECONCILIATION `buried` filter — exact twin of the
post-write one, which had a proof; without it a numeric description hard-errors
`startswith`, the count comes back unusable, and the genuine verdict on the next row is
lost with it;
- the `.status` / `.description` / `.id` type tests, parametrised over all four consumed
fields so a fifth cannot be added without a case (`.creator`'s was the only one proved);
- both retry loops — the combined read and `repair_status_to`'s second POST. Against a stub
that fails EVERY attempt a retrying reader and a one-shot reader are indistinguishable,
which is how a retry ships unexercised; the fixtures now fail only the first attempt;
- the mid-run guard's self-exemption, which is what stops a sentinel-writing run abstaining
on the row it was about to replace with an equivalent one;
- both repair-write failure paths (the repair and the post-POST replacement), reachable only
with a stub that lets the FIRST post through and fails the rest — with every post failing
the job dies on its own classification write and never reaches them;
- the two `state=pending` updates after a repair. The first is load-bearing beyond tidiness:
without it a repaired head re-enters the post-POST check and, on a retarget it then
observes, replaces `$REPAIR_DESC` with the weaker reconcilable sentinel — the same ordering
inversion the floor beside it exists to prevent, reached by another route.
## Declared unreachable (six), enumerated rather than counted
The path-predicate failure branch; the empty-`row` refusal; page 2's non-numeric length; the
`$witness` normalisation; and the two unusable-count arms. Each is defence in depth behind a
filter that makes its input well-formed for every case a fixture can pose — the same standing
exception the post-write unusable-count arm already carried.
That set has gone two -> five -> six across three rounds as the sweep widened. Naming them is
the point: an inventory that undercounts reads as a checked claim and talks the next reader
out of verifying, which is the same defect as inventing coverage — and this branch has
already had to correct that twice.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
A third cold review, which ran the mutants itself, found one measured direction regression
against `origin/main`, one ordering inversion, and four clauses this branch claims as
fixes that survived mutation of their own text.
## The regression
Round 3 type-tested the four consumed fields of the existing `h10` row and resolved a
failure to `""`. For `.creator` that means "no creator" — unattributable — which is a
LICENCE TO RE-DERIVE. Measured, same fixture, both bodies: a head carrying
`h10=failure` with `"creator": 7` posts `Exempt: docs-only change` here and posted NOTHING
on `main`, which died on `.creator.login` before any write. Fail-closed became fail-open.
The rationale that produced it came from #763, whose site is the POST-WRITE filter: there,
dying leaves a green already published, so dropping the row is the safe direction. Here the
alternative is dying BEFORE any write. The deferral rationale did not transfer — which is
the shape this repo has a record for.
A wrong TYPE is now distinguished from a legitimately ABSENT value: `null` is the machine
creator, an unset description and every field of the `{}` no-verdict row; anything else is
unknown state and takes the route an unreadable ELEMENT already took.
## The ordering inversion
`mark_declined_row_if_any` was scoped to "the head carries any row", so it fired on a head
carrying `$REPAIR_DESC` and replaced the human-only marker with the machine-clearable one —
inverting the ordering the SAME commit added a floor to protect at the repair site. One
mechanism, three writers, and only two had the rule.
It also buried a verdict an ALLOW-LISTED reviewer wrote for another base. "Declined" is
decided against this event's `$BASE_REF`, so such a row is still the right answer for the
base it names and the successor run for that base short-circuits on it; burying it costs a
manual re-post on an ordinary retarget-onto-the-reviewed-base flow. Membership is tested on
the raw creator, not on `ex_human`, which the base check has already cleared — the question
is who wrote the row, not whether it governs this diff.
## The unproven clauses
Four claims survived mutation, including the headline one. The witness fixture had been
designed AROUND its own discriminator — its comment said a seed with an unrelated id "would
make this run carry the sentinel forward … and the guard under test would never be reached",
which is a description of the test not reaching it. Eleven proofs added, covering the
witness-by-id, the head arm's own call site (two callers of one helper, one fixture), the
mark helper's result propagation, and the round-4 behaviour above.
`raced_why`'s human value is a named constant now: it is the one such value that is also a
PREDICATE, compared twice, and a drift in either copy silently downgrades the human
`::error::` — the only message that tells a reviewer their verdict was buried.
## Docs
The renamed sentinel literal in two places; three documents still asserting the fence
"writes NOTHING"; the record's `mechanics:` still describing round 2's witness; the
replacement-site list, which had grown by four; a residual pointing "below" at something
above it; and `CLAUDE.md`'s "closed", which is stronger than the record it points at — that
record lists six residuals including both endpoints failing at once. The proof inventory is
stated as an invariant (every clause with a predecessor is mutated back to it) rather than a
count that rots.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Two more cold reviews — cross-family (Codex/GPT-5.6) and a cold Claude reviewer that ran
the mutants itself — converged on two separate things: a remaining class of paths that
still left an unknown state standing, and, more importantly, that several clauses this
branch claimed as fixes SURVIVED mutation of the exact text they name.
## Behaviour
1. The reconciliation witness matches the CURRENT row's `id`, not merely a row with the
sentinel's description. Description alone is satisfied by an OLDER identical sentinel —
which is what a fixed point produces — so a read carrying only the earlier row cleared
the sentinel while the verdict buried under the current one ended up below the fresh
mark. Falls back to the description where the server omits `id`.
2. The two OBSERVED-mutation arms mark a head that carries a row this run declined, instead
of only abstaining. They are still right not to post their CLASSIFICATION — computed
against a base or head the PR may no longer have — but a declined row must not stay
authoritative for the whole window until a successor finishes, and for a PR's FIRST push
no successor is queued at all. Scoped to `pre_state` being non-empty, so the common path
stays quiet.
3. `replace_unknown_state` RETURNS a status. Its first version ended the failure arm with a
successful `echo`, so it reported 0 after both POSTs failed and the fence caller's
`exit 0` reported an abstention that had not happened.
4. An `id` difference counts only when BOTH reads supplied one. A response that omits `id`
beside one that includes it otherwise reads as a replacement, and this guard's reaction
is to abstain — over a row the classification had already declined.
5. Every element and every consumed field of the combined response is type-checked before
extraction, and a schema failure routes to the replacement. `.statuses` being an array
was checked; its ELEMENTS were not, so one scalar made `select(.context == $c)`
hard-error and `set -e` took the step down before any path could mark the head.
6. The path-predicate failure replaces rather than merely exiting, for the same reason.
7. `$UNVERIFIED_DESC` says "Status write", not "Exemption write". It is now written on paths
that grant no exemption at all, and it is the operator-facing text of a required check.
8. The no-op-repair skip keeps the human `::error::`. Skipping the WRITE is right — the head
already carries the strongest marker — but that message is the only place a reviewer is
told their verdict was buried. `raced_why` is a sentence now, not the token `human`.
## Proof
The cold reviewer measured three of the six round-2 claims surviving mutation of their own
clause, one against the verbatim predecessor from the previous commit. Nine proofs added:
the no-mark downgrade's SCOPE (not just the description it writes), the page-2 refusals, the
untrusted-fence write, the row-`id` comparison, the repair floor, the no-op skip, both `$own`
exclusions, the write-result return, and the both-ids-present rule.
Two of those needed the test double to grow: the combined-status stub emitted no `id` at
all, so the `ex_id` clause had never once run with a non-empty value; and POSTs always
succeeded, so both write helpers' failure arms were unreachable.
The `$own` exclusions and the no-op skip are OUTCOME-redundant — mutating either alone leaves
the post sequence unchanged, which is how duplicate guards hide each other. Their proofs
assert the LOG, because what the exclusions alone decide is whether the job reports a race
against its own row. One clause is left deliberately unproven and named as such in the record
and the guard inventory rather than counted: the path-predicate failure branch has no fixture
that can reach it.
## Also
Round 2 left two comment paragraphs duplicated verbatim and a block header narrower than its
block; both fixed. Stale prose corrected in the workflow ("dies WITHOUT posting", "post-write
verification never runs for it", "this block only runs after a `success`"), `docs/ci-cd.md`
("the fence never re-counts", "the history is read twice" — it is three now),
`ci.exemption-provenance` and `docs/guard-inventory.md`.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
Two independent cold reviews (Codex/GPT-5.6 cross-family, and a cold Claude reviewer in
its own worktree) converged on the same class: paths where "this job cannot establish
what is on the head" still resolved by leaving the head alone, which protects a real
verdict and leaves a forged one.
Behaviour:
1. The four page-2 completeness refusals now replace the unknown state too. They were
excluded on the reasoning that the probe fires when NO row for this context was on page
1, so there is no green of any provenance to leave standing — self-contradictory, since
the only reason page 2 is read is that the row may be beyond page 1, which the probe's
own message says. Accepted cost, stated in the record: a head with more CONTEXTS than
the 50-row cap stalls every run; measured 2026-08-29, this repo puts 8 on a `main` head,
and that case already stalled with an ABSENT check.
2. The no-mark downgrade covers every re-derivable write, not only `success`. Restricting
it analysed the wrong PR: the damaging case is one that IS exemptible and got the
generic `pending` only from a transient enumeration failure. That description carries no
marker, nothing verifies it without a mark, and the next run re-derives it into the
exemption with the human row below its own mark — route 2's damage through route 1's
condition. `$REPAIR_DESC` stays exempt, being stronger and not re-derivable.
3. The fence branch that cannot trust its retarget count while holding a derived `success`
writes the sentinel instead of abstaining. It is reached only after the classification
DECLINED to inherit the row the head carries, so posting nothing left that row current;
the message said the context "stays absent", true only of a head that had none.
4. Reconciliation needs a WITNESS: it may clear only over a complete history containing the
sentinel's own row. `ex_unverified` means the combined endpoint just returned that row
and `/statuses/{sha}` keeps one per POST, so a complete-but-empty history contradicts a
write that demonstrably happened — and `page_statuses` accepts an empty page 1 as
complete, which is what made it reachable. Both reviewers reproduced the clear-then-exempt
outcome. The shipped positive test used exactly that impossible fixture, so it was
pinning the defect; it now seeds the sentinel row, and an impossible-empty negative plus
a witness mutation proof were added.
5. The mid-run "did this row change" comparison now includes the row ID. The two sentinels
are byte-identical by design, so a mid-run replacement of one by another was invisible to
a state/creator/description triple. Measured 2026-08-29 (Gitea 1.27.1, head 736649b3):
the COMBINED endpoint carries `id` on every row, ids 14..30 ascending — the job had only
ever read ids from `/statuses/{sha}`. Where a server omits it both sides are empty and
the comparison degrades to the pre-existing text test.
6. The repair has a FLOOR — it may never write a description weaker than the one this run
decided — and is skipped when it would rewrite what is already there. Widening the gate
to every write meant a transient post-write read could rewrite a correct `$REPAIR_DESC`
carry-forward with the machine-clearable sentinel, reversing the ordering rule the
classification chain states.
Writing the sentinel and failing the job are separate decisions, which is why
`replace_unknown_state` and `replace_unknown_and_die` are two functions: the read refusals
were already non-zero exits on `main` and stay red; the fence branch exited 0 there and
still does, because an unreadable timeline is an ordinary hiccup and reddening every one is
noise this file elsewhere refuses to add.
Prose corrected where it now overclaimed: "the green never stands" after the post-POST
re-check is wrong — it is live between the POST and the repair, so the check makes a
permanent green TRANSIENT; "a later run reconciles this automatically" is wrong in the one
case where the replacement costs anything, since finding a masked verdict UPGRADES to the
human-only sentinel; and the mutation-proof framing claimed every mutant restores the exact
predecessor, when two do, one restores the shape #742 withdrew, and the rest disarm clauses
that have no predecessor. The quiet-timeline positive control now counts timeline walks,
because a single POST is also what a skipped re-check produces.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
The gate's post-write verification had five routes that all ended the same way — an
exemption `success`, or a generic `pending` a later run turns into one, standing over a
human `failure`.
Two of these were attempted inside #742 and withdrawn, and the withdrawal is what shaped
this change. That attempt withheld the exemption by writing a GENERIC `pending`, which is
exactly what a later run re-derives into `success` — it moved which run posted the forged
green rather than stopping it — and it had no retry path, because this workflow triggers
only on `pull_request_target` types, so a transient failure on a PR's last event stalled an
exempt PR until a human nudged it. The fix therefore needs two properties at once: sticky,
so a later run cannot re-derive it, and reconcilable, so a blip does not cost a head its
exemption permanently. Neither the repair sentinel nor a generic `pending` has both, which
is why there is now a second sentinel rather than a reuse of the first.
What changed:
1. No high-water mark => the exemption is WITHHELD before the POST and the head is marked
with the new `UNVERIFIED_DESC` sentinel. Withholding before the write rather than
posting and repairing matters because the defect is known in advance: publishing a green
to take it back opens a window branch protection, and an already-scheduled auto-merge,
can see.
2. Post-write verification runs after EVERY write, not only `success`. A generic `pending`
masks a rejection landing in its own write window just as well, and carries no marker,
so the next run re-derives it with the human's row now below THAT run's mark.
3. `.description` is type-tested before `startswith`. `(.description // "")` does not
replace a NUMBER, so `startswith` hard-errors on one, killing the whole count — the
genuine verdict beside the malformed row is lost with it.
4. The retarget count is re-taken AFTER the POST on the exemption path, closing the
PERMANENT forged green `ci.verdict-write-retarget-fence` listed as its residual 1. The
retarget axis only: a push after the POST moves the head, so the status no longer gates
that PR, while a retarget changes the effective diff with the sha unchanged.
5. An unreadable combined-status read retries once and then REPLACES the unknown state
instead of declining to write. Declining protects a real verdict and leaves a FORGED one
— an off-list `success` is the row #742 exists to revoke, revocation happens by
re-deriving it, and the job then went red on a status branch protection does not read.
One defect this introduced and fixed on the way: widening the post-write gate to every
write made the job match its OWN row, because the machine-sentinel arm selects on a null
creator. A run taking the carry-forward path POSTed `$REPAIR_DESC`, then found "a sentinel
above the mark", then repaired to the identical description. `--arg own "$desc"` excludes
it, by description rather than by id — the id of the row just written is not knowable
there.
Reconciliation is what bounds the stall: a later run pages `/statuses/{sha}` in full and
either finds a `Review-verdict:` row underneath the sentinel — an established fact, so it
upgrades to the repair sentinel, clearable only by a human — or finds none and clears it.
It is sound because the two endpoints disagree: a masked verdict is invisible on the
combined endpoint (latest row per context, which is the sentinel) and still present in the
per-POST history.
Tests: each fix is paired with a `test_MUTATION_…` proof that restores the exact
predecessor text through a new `_run_classify(mutate=…)` knob, whose count assertion is the
binding — a clause that has since moved substitutes zero times and fails loudly rather than
measuring the unmutated body. Two CHAINED tests feed run N's real output into run N+1,
because both sentinels are fixed points and a single hop cannot assert a fixed point: the
raced-`pending` repair must survive the run that would otherwise grant the exemption, and
the unverified sentinel must not decay while it cannot be reconciled.
Docs: new record `ci.verdict-unverified-write-sentinel`; the now-false guarantee prose in
`ci.verdict-write-retarget-fence` (its `rule:` frontmatter, the "resolves it" opener, "the
fence above closes", the truncating-block claim and residual 1), `ci.exemption-provenance`,
`docs/ci-cd.md`, `docs/remote-state-inventory.md` and `CLAUDE.md` corrected by concept
rather than by phrase, per the scope boundary recorded on the issue.
fixes#849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF
`review-verdict.yml` inherits an existing `review-verdict/h10=success` only from a status whose
`.creator.login` is on its `H10_REVIEWERS` allow-list (#742). `post-review-verdict.sh` wrote those
verdicts with whatever account owned the credential in the environment and never asked whose it was.
Two coupled values, nothing asserting the coupling, and the failure was the silent kind: the status
is written, the tool reports success, and the next `pull_request_target` event re-derives it and
posts over it. The PR stalls with no visible cause.
The writer now READS ITS OWN STATUS BACK, identifies that write by state and description, and
refuses — before the verdict comment, so the surviving half-state is the documented `ask` one —
unless the recorded creator is allow-listed. Measured after the write rather than probed before it:
that tests what Gitea recorded as the author, which is the value the gate reads, and needs no scope
beyond the repo access the POST already required.
Membership is required for a `success` ONLY, mirroring the gate's own asymmetry: a `failure` is
inherited from any attributable account, so requiring it there would refuse a verdict the gate
honours and leave an off-list reviewer no supported way to record a rejection.
The allow-list is DERIVED from the gate's own literal by the new `scripts/lib/h10-reviewers.sh` —
one declaration, not two plus a parity test. It is a parse rather than a shared declaration both
sides source because the gate runs against a checkout of the PR's BASE sha: a PR whose base predates
such a file would not have it, and a missing `source` under `set -euo pipefail` kills the job, which
posts no `review-verdict/h10` at all and blocks every merge including its own repair (#743).
`scripts/post-review-verdict.sh` moves BEHAVIOUR-ONLY -> MUTATION in the guard inventory, which the
manifest's own note called "the most valuable upgrade on this list". The declared clause lives in the
GATE: rewriting `H10_REVIEWERS` while the posting account stays fixed reddens the accept path only if
the writer reads the list live AND the comparison gates the outcome.
Two defects were caught by probing the live instance rather than re-reading the code. Reading `.state`
instead of `.status` per row would have refused EVERY verdict — a repo-wide deadlock, shipped green,
because the test shim replayed the POST payload as the read-back body and so agreed with the parser
by construction. Then a `(.status // .state)` fallback added as defensiveness recreated #845 exactly:
the writer would accept a shape the gate cannot read and report success.
Nine independent cold review rounds, all worktree-isolated, one cross-family (GPT-5.6 via Codex).
Round 8 caught the most important one: a `set -u` "correction" made mid-branch had inverted a TRUE
statement in live merge-gate code, because the probe used a plain `$UNSET` while the validator uses
`${#arr[@]}` — different shapes, different behaviour. Withdrawn wholesale; both libraries are
byte-identical to `main` again.
Verification: full `scripts/tests` suite green (1278 passed, 2 skipped); the declared mutation
executes every run and reddens its named proof with the manifest's `expect` string; every clause
disarmed individually and confirmed to redden its own named test; live probes against Gitea 1.27.1
for the row shape, the description round-trip, the paging order and the required-check list.
Docs: `ci.exemption-provenance` records the coupling as asserted rather than as a tracked residual,
plus `docs/ci-cd.md`, `CLAUDE.md`, `docs/guard-inventory.md`, `docs/remote-state-inventory.md`,
`ci.script-tests-job` and the `script-tests` population comment in `pr-checks.yml`.
Deferred: the refused-verdict residual (a non-inheritable status left standing with no comment) is
the `ask` half-state `release.verdict-writes-status-before-comment` designates as safe; a second
corrective write is the sticky-sentinel mechanism #849 is separately designing.
fixes#845
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Probed every reachable surface rather than stopping at the endpoint #853 already knew 404s: the dispatch API takes `ref` as a required free-form string with no allow-list; protected environments do not exist at 1.27.1 (0 of 308 documented paths mention "environment", secrets are org/repo/user-scoped only); the loaded `app.ini` sets two `[actions]` keys; and the CLI's sole Actions subcommand is `generate-runner-token`. So option 3 is unavailable.
Accepted on a different ground than the issue proposed. "Anyone with repository write can already do worse" is unfalsifiable and hides the cheaper route. The operative reason is that dispatch is not the cheapest path: `docker-build.yml`'s head-resolved `pull_request:` runs attacker-authored YAML, which reaches every secret in the store — six of its jobs hold `REGISTRY_PASSWORD` on that route and two are branch-protection required contexts. "Push a branch, open a PR" costs no act outside the ordinary contribution flow, where a dispatch costs one.
Corrections to #853's own table, verified against the tree: `dependency-scan.yml` references no secrets at all; the "four workflows" count is right.
Deliberately not applied: a `v*` tag protection (`tag_protections` is empty and 1.27.1 supports it) — protection-class config whose failure mode is a broken release cut, so it needs its own change and verification. Tracked with the `pull_request:` residual in #885.
The web UI was not swept, and the record says so explicitly rather than claiming exhaustiveness — a Gitea Actions control can exist with no API surface at all.
fixes#853
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
2026-08-30 01:18:03 +00:00
timothytimothyClaude Opus 5 (1M context) <noreply@anthropic.com>
`git fetch --depth=N` grafts a complete clone shallow. `scripts/ci-detect-docs-only.sh` applied a depth chosen for its three `fetch-depth: 2` consumers to `build`'s `fetch-depth: 0` checkout, so the `git describe --tags` in the next step found no reachable tag and a `|| echo v0.0.0` fallback turned that into a version: every `:latest` image shipped `InformationalVersion 0.0.0-<sha>` from 2026-07-17 (#416) until now.
Both fetch sites now go through `fetch_ref`, which passes `--depth` only when the checkout is already shallow. `Compute version and tags` fails the job instead of defaulting, so no `:latest` is published rather than a mislabelled one; releases are unaffected because the tag path never calls `describe`.
Ships a guard that drives the real script over real `file://` clones with a negative control, a declared clause mutation, and a decision record `ci.fetch-depth-never-grafts-a-complete-clone`.
fixes#836
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
`Complete<T>` (#807) makes SPA full-replace bodies fail typecheck when a builder omits a schema member. Nothing checked it was APPLIED: `completeRequest.guard.test.ts` proves the type's semantics and would stay green with every annotation deleted, and `test_optional_request_members.py`'s COVERED disposition — "the builder is annotated `Complete<T>`" — was a claim about another language's source that nothing verified.
Adds `completeAnnotationScan.ts` (compiler-API scanners) + `completeAnnotations.guard.test.ts`, with a synthetic-source fixture suite. Two derived populations: the `Complete<…>` annotations (SPA AST ∩ git index) and the droppable schemas (parsed from the generated `v1.d.ts`, a pass-through of the OpenAPI `required` array). It asserts a production annotation per schema dispositioned as needing one, NO annotation on the server-computed and load-bearing-omission schemas, that every `Complete<X>` resolves to a generated schema rather than a hand-written mirror, and set equality between droppable schemas and the reviewed dispositions. `test_complete_annotation_dispositions.py` cross-checks that table against the authoritative Python one and ships a declared, harness-executed mutation.
Found one live defect: `playouts.ts` declared two request types as hand-written mirrors SHADOWING generated schemas of the same name, so their `Complete<>` was checking a local copy rather than the contract — the #754 mechanism wearing the annotation meant to prevent it.
Eight review rounds, seven BLOCKED, two independent cold reviewers. A wrapper-signature scanner was built and REMOVED: every blocker traced to that one mechanism (obligation on the wrong population; reachability mistaken for protection, since `Complete<T>` is shallow; body discovery keyed on a parameter name, then parameter-vs-local; and finally `export function` → `export const` blinding the scanner and its cross-check together). Five defects from one mechanism, so the mechanism went rather than a sixth patch.
Residuals stated in §4b, the guard-inventory row and the record: per-SCHEMA not per-site or per-wrapper; token presence not liveness; the phantom direction unchecked (#777); a second `setupFiles` entry could discharge; and plugin-level population integrity borrowed from the sibling guard.
fixes#820
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Round-five review confirmed the decision record is coherent with no third
survivor of the empty reading, and returned three LOW findings. All are in
prose I wrote in the last two commits.
- The comment defending `(IsAbstract && !IsSealed)` cited
AlternateScheduleSelectorTests as an in-repo static-fixture witness. That
class IS static, but it merely NESTS its [TestFixture]es and declares no test
of its own, so it would fail the sibling "declares no runnable test" assertion
rather than demonstrating the point. The rule is right and the witness was
wrong, which is the worse of the two failures because a wrong example is what
a reader checks the rule against. No witness is cited now, and why is stated.
- A mis-bound `because` in `rule:`: "assigning a null and calling SaveChanges
SUCCEEDS ... because only the HTTP request records normalize with `?? []`".
The `?? []` clause explains how a null could REACH the entity; what makes the
save succeed is the column being nullable. A right observation with a wrong
cause attached. Split into the two claims.
- `signals:` carried the literal token `paths:` twice, an artifact of appending
the #823 path list to the existing one. It degrades the field the discovery
surface parses.
Also recorded from that review, and NOT changed: `MonthsOfYear ?? AllDaysOfMonth()`
survives the selector fixture and no date can kill it -- 1..31 contains every
valid month, so it is an EQUIVALENT mutant there rather than a coverage gap.
Its non-equivalent twin at the DTO boundary is pinned per-dimension by
RecurrenceLimitsMapperNullTests. Left alone deliberately: chasing an equivalent
mutant with a contrived date would buy nothing and cost the fixture's
readability.
Local gate: ErsatzTV.Tests 2091 passed / 6 skipped, Core.Tests 697/1 -- 0
failures. Format clean, no BOM. decisions_validate OK.
Refs #823
Refs #824
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zUmJZHhVP7kXg5DV237TW
Round-four review. One HIGH, again in the decision record, and the previous
commit message asserted this exact class was cleared. It was not.
THE HIGH, and the reason it recurred.
A second sentence still described the rejected reading: "The guard form is
`?? []` into a local rather than this record's Optional(x).Flatten(), a STATED
deviation". The shipped guard is `?? AllDaysOfWeek()`. That sentence is the one
that dictates guard FORM to the next implementer, so it would have taught the
`[]` reading the same record spends a paragraph calling data corruption -- and
it had already propagated into docs/decisions/README.md, the mandated entry
point, which carries `rule:` verbatim.
The mechanism, not the sentence, is the defect. I swept with a regex keyed on
"null" plus a reading word; this sentence talks about guard FORM and contains
neither, so it could not match. That is grepping the retracted WORDING instead
of sweeping the CONCEPT, which is exactly what this corpus warns about -- and
three rounds in a row have now found a defect introduced by the previous
round's targeted string edit. So the fix is not another targeted edit: the
whole `rule:` field was split into its 39 sentences and read back one by one
against the code. Everything below came out of that pass rather than a grep.
Its secondary damage is worth recording because it is the shape of a rationale
that outlives its claim: the deviation was justified by ".ToList() allocates
for nothing", which is now BOTH irrelevant to the choice AND false about the
shipped code, since AllDaysOfMonth()/AllMonthsOfYear() are themselves
Enumerable.Range(...).ToList() on exactly the null path it describes.
- The opening sentence of `rule:` prescribed Optional(x).Flatten() as THE
read-site form. It is the sentence most likely to be read in isolation, and
it is wrong for six of the eight columns. It now separates the universal half
(a LOCAL, never assigned back) from the half that is not (the substituted
value), and names where each applies.
- `signals:` had never been touched, so roughly 60% of `rule:` was unreachable
by the discovery surface built for it -- no AlternateScheduleSelector, no
mapper, no "unrestricted", and its paths: list named none of the files this
work touched. It also advertised "Optional Flatten hoisted local" as the
form, which is precisely what the six do NOT use.
- The body prose was still entirely about SongMetadata while `rule:` had grown
a whole second subject. Added the two results that contradicted the prior
reasoning, in prose, where a reader meets them.
A REAL BUG in my own guard, not just prose:
fixture.IsAbstract.ShouldBeFalse(...)
A C# `static class` compiles to `abstract sealed`, and NUnit runs tests
declared in one -- this repo already has such a fixture
(AlternateScheduleSelectorTests is `public static class`). So the check I added
one commit ago to reject an un-runnable fixture would have falsely reddened a
perfectly good static one. Now rejects an abstract BASE (abstract and NOT
sealed), which is the case NUnit actually cannot instantiate.
A SURVIVING MUTANT the added controls did not kill:
AnyDate was 2024-03-06. With a day <= 12 a CROSS-WIRED substitution survives
the whole fixture -- `DaysOfMonth ?? AllMonthsOfYear()` hands back 1..12, which
still contains day 6, so every assertion passes while the guard substitutes the
wrong set. Moved to 2024-03-20, still a Wednesday in March, outside 1..12.
Measured both ways rather than reasoned: the cross-wire mutant passes the old
fixture and FAILS 2 of 11 on the new one.
Also re-witnessed, because I had modified that file and never re-proved it:
restoring `??=` in LuceneSearchIndex reddens the LUCENE fixture (1 red, 1
green) -- the exact mirror of the Elastic mutation. Extracting
ThrowOnWarningLogger did not cost #701 its proof, and the two fixtures are
independently load-bearing in both directions.
The record is now 73 prose lines, over the 60-line WARNING ceiling. Stated
rather than trimmed: it is 42nd of 42 records over that line, and the added
content is distinct findings (a second subject, a migration analysis and three
residuals), not redundancy against a sibling.
Local gate: ErsatzTV.Tests 2091 passed / 6 skipped (the three fixtures' MySQL
halves), Core.Tests 697/1 -- 0 failures. Format clean, no BOM. decisions
validate OK.
Refs #823
Refs #824
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zUmJZHhVP7kXg5DV237TW
Round-three review findings. One HIGH, and it was in the durable artifact
rather than the code.
THE HIGH: the record stated the shipped reading and its inverse.
The semantic reversal (empty -> unrestricted) rewrote the residual and the
write-half of `media.nullable-primitive-collection-mutation` but left the
ORIGINAL reasoning standing two sentences earlier: "A null reads as EMPTY, so
the item matches nothing"; "the REJECTED alternative was the All*() set";
"SKIPPING the row is the conservative repair". The shipped code is
`?? AllDaysOfWeek()` -- precisely the alternative that passage calls rejected.
The previous commit then inserted residual (1), which reasons entirely FROM
the All*() reading, two sentences after the sentence denying it.
That is worse than a stale comment. A session resolving this key -- or reading
the MemPalace mirror, which carries `rule:` verbatim -- would have been told to
write the guard the other way, i.e. talked into the `[]` reading that the same
record elsewhere argues is data corruption one save later. Replaced the whole
passage, then swept the record for every other mention of the empty reading
rather than trusting the one replacement: the only survivor is the new sentence
that records EMPTY as the rejected alternative, which is the direction that
stops it being re-adopted.
THE MEDIUM: one arrangement did not close the hole it claimed to.
The discriminating control added last commit nulls DaysOfWeek against a
restrictive MonthsOfYear. It excludes "any NULL matches unconditionally" only
for that dimension. The review supplied the surviving mutant --
`if (item.MonthsOfYear is null) { return item; }` ahead of the checks -- and
traced it green through all nine tests. Verified by EXECUTION, not by reading:
applied to the previous fixture it passes; applied now it FAILS 1 of 11. Each
of the three dimensions is now nulled against a restriction on a different
dimension.
The rest, all from the same round:
- The coverage guard's test detection listed attribute TYPES, and each list
falsely reddened whatever it omitted: TestAttribute alone missed [TestCase],
and the three-type replacement missed [Theory]. Now decided by NUnit's own
ITestBuilder/ISimpleTestBuilder interfaces, which cannot fall behind the
vocabulary. It also dropped BindingFlags.Static (GetMethods() defaults to
including it), which would have falsely reddened a static test method.
- The same guard accepted an ABSTRACT fixture -- NUnit never instantiates one.
The indexer population already filtered IsAbstract; the fixture side now
mirrors it.
- The record's `mechanics:` still described the old `[Test]`-only clause, in
the same file the change edited.
- An <inheritdoc> made the ProgramScheduleAlternate empty-case test inherit a
docstring written from the PlayoutTemplate test's viewpoint.
Two more mutations executed:
- `if (item.MonthsOfYear is null) return item;` -> 1 red, 10 green. This is the
mutant that survived the previous head; it no longer does.
- an abstract type named in the covered set -> coverage guard red.
Local gate: ErsatzTV.Tests 2091 passed / 6 skipped (the three fixtures' MySQL
halves, skipping visibly without ETV_TEST_MYSQL_CONNECTION), Core.Tests 697/1
-- 0 failures. Format clean, no BOM on the touched set. decisions_validate OK.
Refs #823
Refs #824
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zUmJZHhVP7kXg5DV237TW
Follow-up commit (the branch is pushed, so not an amend). Two more cold
reviews landed on the previous head; both reported 0 Blocker and 0 High, and
these are their Mediums and Lows. Each fix carries its own witnessed mutation.
1. The selector fixture could not tell the fix from a much broader one.
Every null test set a NULL and expected the item SELECTED, so all of them
pass equally under "NULL means unrestricted" and under "any NULL makes this
item match unconditionally" -- a refactor short-circuiting the whole date
check on any null kept them green. Added the discriminating control: a NULL
DaysOfWeek paired with MonthsOfYear = [1] against a MARCH date must be None.
Only the narrow reading passes.
2. A_Null_Item_Does_Not_Disturb_Selection_Of_A_Later_Item never measured its
own docstring. The nulled item was unrestricted and at Index 0, so it always
won and the second item was never evaluated -- the stated invariant ("a null
on the first item must not decide the second") went unmeasured while the
test passed. Split into two: one where the nulled item genuinely does not
match, which measures that the loop CONTINUES; and one that pins the
index-order win separately.
3. The empty-preservation control existed for one of two identical mappers.
The anti-mutant test for "empty or null becomes All*" covered only
Playouts.Mapper; Scheduling.Mapper is a byte-identical triple in another
file and had none, so a defensive edit to it alone would have rewritten a
deliberately-empty user selection to 1..31 with the suite green. That is the
one-helper-two-callers shape this repo has been bitten by. Added the
matching test.
4. The coverage guard's [Test] clause did not check what its message claimed.
GetMethods() without BindingFlags returns INHERITED methods, so a fixture
that merely subclasses another satisfied it while driving the wrong indexer
-- and Values.Distinct() cannot catch that, since the two Types differ. It
also matched TestAttribute alone, so a future fixture written as [TestCase]
would have falsely reddened, and it accepted an [Explicit]/[Ignore]d fixture
that never runs, which is the "wired is not running" failure the guard
exists to prevent. Now DeclaredOnly, the full test-method vocabulary, and
Explicit/Ignore rejected at both method and fixture level.
5. Three residuals recorded on media.nullable-primitive-collection-mutation
that the previous head asserted nothing about:
- the LOUDNESS change, worst for an all-three-NULL ProgramScheduleAlternate,
which now matches unconditionally and shadows the default schedule where
it previously threw. Unreachable today, and a choice over an unreachable
state rather than a measured requirement -- said plainly.
- the normalization is ONE-WAY and WHOLE-LIST: both PUT paths are full
replaces, so editing any row persists All*() over EVERY NULL row in that
playout, and afterwards "the operator selected all 31" and "this is a
legacy row" are indistinguishable. An ordinary user action closes that
door.
- the WRITE side disagrees with the READ side about what ABSENCE means: an
omitted daysOfWeek normalizes to [] ("never applies") while a NULL column
reads as unrestricted, so an API client gets HTTP 200 and a row that
silently never fires. Filed as #880 rather than folded in here, because a
client omitting a field on a write is a different question from what a
legacy NULL meant.
Two more mutations executed, both witnessed:
- DaysOfWeek guard disarmed in Scheduling.Mapper -> 1 red, 3 green.
- A fixture with no DECLARED test named in the covered set -> coverage red.
Local gate (MySQL lane armed): ErsatzTV.Tests 2097 passed / 0 skipped,
Core.Tests 695/1 -- 0 failures. Format clean, no BOM on the touched set with
the population count asserted. decisions_validate OK.
Refs #823
Refs #824
Refs #880
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zUmJZHhVP7kXg5DV237TW
Both issues are #701 deferrals, and they land together because both rewrite
the same decision record.
#823 -- can a null reach one of the six collection-valued scalar columns?
MEASURED against a real TvContext on BOTH providers (SQLite, and MySQL 8.4
on an ephemeral server), because the reasoning available beforehand pointed
the wrong way. The two converters differ on their read side --
IntCollectionValueConverter maps null-or-blank to Array.Empty<int>(), while
EnumCollectionJsonValueConverter would dereference the result of
JsonConvert.DeserializeObject -- so the expectation was that a NULL row
behaves differently per column. NEITHER RUNS: EF does not invoke a value
converter for a NULL column at all. All six materialize as CLR null, the
int converter's null-to-empty branch is dead on this path, and unguarded
each .Contains in AlternateScheduleSelector throws NullReferenceException.
A NULL reads as UNRESTRICTED -- the All*() sets -- not as empty. This is
the whole semantic question and the first draft got it backwards. It is
decided by the one NULL reachable WITHOUT any code writing one: Sqlite's
20240113140741_Add_PlayoutTemplate_DaysOfMonth adds the column with
nullable:true and NO defaultValue, so a PlayoutTemplate row inserted before
it holds NULL and by construction had no day-of-month restriction. Reading
that as empty INVERTS the row's meaning and silently stops the template
applying at all. All*() preserves it, and is how "no restriction recorded"
is already represented (GetPlayoutAlternateSchedulesHandler,
PreviewBlockPlayoutHandler). What does NOT decide it, and was wrongly cited
in the first draft: the API request records normalize an omitted field with
`?? []`, but that is a client omitting a field on a WRITE and says nothing
about what a legacy database NULL meant.
Two read sites, not one. Guarding only the selector would have left the
entity->DTO mappers unguarded, and those feed the SPA: PlayoutScheduleEditors
spreads the collection (`[...template.daysOfMonth]` -> TypeError on a JSON
null) and playoutTemplateCalendar's appliesToDate -- an exact port of
GetScheduleForDate -- calls .includes on it. Both mappers now substitute the
SAME defaults, so the preview agrees with what is actually scheduled. Neither
guard is assigned back onto the entity, which is the
media.nullable-primitive-collection-mutation mechanism.
Reachability, stated precisely rather than overclaimed. All six are
nullable:true on both providers, but a nullable column does not produce a
NULL row: five of the six were present at CreateTable, so a NULL there still
needs code to write one, and on MySQL there is NO code-path-free NULL for any
of the six. The write path ACCEPTS a null (SaveChanges succeeds, stores SQL
NULL) but no caller supplies one today -- every production construction of the
two commands goes through the request records. That is a property of the code,
not a live caller; claiming otherwise would be the banned "it's AsNoTracking
today" argument pointed the other way.
#824 -- ElasticSearchIndex.UpdateSong had no regression test
Issue option 1 (a non-network transport) shipped, and needed no new package:
Elastic.Transport.InMemoryRequestInvoker is public in the pinned version and
ElasticsearchClientSettings(NodePool, IRequestInvoker) accepts it, injected
into the private _client the way #701 injects the Lucene IndexWriter.
UpdateItems never runs `_client ??= CreateClient()`, so the injected instance
is the one used.
Two traps there are load-bearing, both measured: the canned response must
carry an `X-Elastic-Product: Elasticsearch` header or the client's product
check throws UnsupportedProductException INTO UpdateSong's catch, and an empty
body fails to deserialize the same way. Either turns the fixture into a green
measurement of the error path -- which is how it first failed here, caught by
the ThrowOnWarningLogger. The document id is asserted as the LAST PATH SEGMENT,
not by substring: the index name carries digits, so ShouldContain would stop
discriminating for a song whose id collided with one.
Six mutations executed, each disarming ITS OWN clause alone:
- `??=` restored in ElasticSearchIndex only -> the Elastic fixture reddens on
"metadata.Artists should be null but was []" while the LUCENE fixture stays
GREEN. The #824 hole demonstrated, not described.
- DaysOfWeek guard disarmed in the selector -> 4 red, 3 green (DaysOfMonth and
MonthsOfYear unaffected). Each clause is independently load-bearing.
- DaysOfMonth guard disarmed in Playouts.Mapper -> 1 red, 2 green.
- Elastic dropped from the covered set / mapped to the SAME fixture as Lucene /
mapped to a class with no [Test] -> SearchIndexMutationCoverageTests reddens
on each.
That coverage guard is the boundary fix the issue asked for: the covered set is
compared against an ISearchIndex population DERIVED FROM THE ASSEMBLY. Its claim
stops where the check does -- no static check can establish that a named fixture
actually DRIVES its indexer, so it forces a human to look rather than proving
coverage. ThrowOnWarningLogger moved to ErsatzTV.Tests/Support so both fixtures
share it; the Lucene fixture's assertions are otherwise untouched, since it is a
witnessed proof artifact.
No production change in ElasticSearchIndex.cs -- #824 is coverage only.
Docs: testing.md gains a "Provider-parity fixtures" section naming all THREE
opt-in-MySQL fixtures and recording that CI runs none of them (#627);
docs/README.md gains the matching task signal; guard-inventory.md's
hand-written C# guard list goes from five files to six. Scheduling/Mapper.cs
loses the UTF-8 BOM it inherited, per #311 fix-as-you-touch.
Local gate (with the MySQL lane armed): ErsatzTV.Tests 2096 passed / 0 skipped,
Core.Tests 693/1, Infrastructure.Tests 114, Architecture.Tests 7, Scanner.Tests
1504 -- 0 failures in each. scripts/tests 1228 passed / 2 skipped. dotnet format
whitespace --verify-no-changes clean; BOM check over the touched set with the
population COUNT asserted, because a bare zsh loop silently checks one
concatenated filename. decisions_validate OK.
Fixes#823Fixes#824
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zUmJZHhVP7kXg5DV237TW
The guard asserted EXACT completeness over a population enumerated by a directory
walk, so an untracked .ts/.tsx under web/src/ entered it and failed as unregistered
on that developer's checkout while CI — which only ever checks out tracked files —
stayed green.
The glob still supplies file CONTENT; the POPULATION is now the git index, read by
web/vite-plugins/trackedSourceFiles.ts in Vite's own Node context and handed to the
app project as a virtual module. That reaches the index without admitting
@types/node to tsconfig.app.json, the obstacle that deferred this in #818.
Three mechanisms carry the proof, each added because the previous was measured
insufficient: a closed-form restatement of the shared scope predicate (sharing no
helper at any depth with what it checks); a second independent `ls-files --others`
query cross-checking the population; and real-git tests that execute the derivation
against a temp repository.
Six residuals are stated with their MEASURED fail-directions, and
testing.guard-derives-population-from-source gains a bounded exception plus the
closed-form criterion.
fixes#819
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
A force-push H1 -> H2 -> H1 spanning `pr-changed-files.sh`'s paging leaves its final
`.head.sha` comparison equal while the middle pages came from H2, so a mixed file list
could produce a docs-only exemption `success` no single head ever justified. The base
alias had been fenced since #706 by a monotonic `change_target_branch` count; the head
axis had nothing, and three contracts asserted otherwise.
`count_retargets` becomes `count_pr_mutations`: one timeline walk, two tallies, one shared
trust flag, a separate fence arm and diagnostic per axis. The advisory hook re-reads
`.head.sha` at the same hoist and off the same response as the base re-read. All three
overclaiming contracts are corrected, plus four paraphrases the first sweep missed.
Measured, not assumed: Gitea 1.27.1 still serves no `files` on `compare/{base}...{head}`;
every push is a `pull_push` event and its count cannot alias; PR #761 really went
`8798a1d -> 830a407 -> 8798a1d`; and Gitea creates the push comment BEFORE emitting the
synchronize notification, so a run cannot abstain on its own trigger.
Two pre-existing fail-opens in the shared walk were found by review and fixed: an empty
ARRAY first page was trusted on any page while the `null` arm required `page > 1`, and no
row was validated before `.type` was selected on.
NOT closed, and documented rather than overclaimed: the walk's `null` terminator is
defeatable, because Gitea pages before it filters (#870). The fence closes the ABA on a
timeline with no truncating block, not the ABA outright.
fixes#803fixes#664
Refs #870
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Closes#786 and #789, bundled because working either alone would build the artifact the other removes.
Every job in all six tracked workflows declares `env.CI_JOB_ROLE` (guard/report-only/none); the
`docker-build.yml` jobs also declare `env.CI_EXECUTION_CLASS` (toolchain/bare-runner). Both guard
populations derive from those markers; the `TOOLCHAIN_JOBS`/`BARE_RUNNER_JOBS` literals are deleted.
A missing or unrecognised marker is a hard failure in both checkers.
#789's literal had a real justification — set equality between two DERIVED sets is blind to a member
leaving both at once — so the marker is the anchor that replaces it, and the cost (proximity to the
`container:` block) is paid by a THIRD derivation from each job's own steps, which is also the only
check that sees the failure #789 filed: a .NET step moved into a bare-runner job, where no set
changes. The residual is disclosed: drop the block, flip the marker AND hide the tool behind a
script and all three go blind, bounded by the failure mode being a loud missing-binary crash.
#786's guard jobs join a machine-checked population: a new `test_workflow_job_guards.py` asserts set
equality both ways against a new "Workflow-job guards" table, and the four jobs with no dropped-step
guard each carry a recorded decision.
Two issue claims were refuted by measurement: #789's "editing docker-build.yml re-points the pin"
(the pathspec is `docker/ci` only) and #786's job count (17, not 15).
Four cold adversarial review rounds across two model families; rounds 1-3 BLOCKED, all findings
fixed and each fix demonstrated by reproducing the reviewer's own test. The recurring defect class
was prose drifting from code, including a mechanism claim in the decision record that execution
refuted. All five mutation proofs redden when their shipped detector is disarmed.
New decision record: `testing.workflow-declares-its-own-job-metadata`.
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Records `testing.verification-code-needs-its-own-proof`: the proof obligation follows the
VERDICT rather than the file, so it binds harnesses, wrappers, timeouts and checkers — not
only the files the guard population derives.
The issue asked for a stated position on whether non-guard checker scripts get mutation
proofs. The position as first written claimed `scripts/mcp_smoke.py` "cannot participate"
because driving it needs the gitignored `.mcp.json` and a cold-built language server. Cold
review refuted that by execution: it takes its config path and server name as positional
arguments. The record had failed its own headline rule on the one claim its decision rested
on, so this ships the proof instead of the exemption.
- `scripts/tests/test_mcp_smoke.py` — a hermetic stub JSON-RPC responder and six cases
pinning the defects the checker has already had, with the positive control as a fixture
the refusal tests depend on, so a node-id or `-k` selection cannot skip it.
- A declared clause in `mutation_manifest.py` targeting the unguessable request id, using
the `guard=test / target=script` shape that already exists for `mutation_harness_lib.py`.
Witnessed red: `id_init = 1` makes the pre-answer accepted at `initialize` (rc 9 -> 10),
and only that test moves.
`mcp_smoke.py` still gets no inventory row — one is rejected as a phantom (measured). The
row goes to the test file, which joins the derived population automatically.
Five cold-review rounds, four BLOCKED. Round 2 caught a `ruff format` red that would have
failed `script-tests`. Rounds 3-5 found only hand-maintained counts and uniqueness claims in
prose, three of them created by the previous round's fix; that class was deleted rather than
corrected again, per this record's own stop-and-subtract rule.
Docs updated in the same PR: `docs/README.md` task-signal map and `docs/guard-inventory.md`
(row, summary counts, scope-limit item 6).
fixes#796
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
CI's `Script lint and tests` job went red. Cause: I never ran ruff locally,
which this repo's Python convention requires after any .py change.
- E741 twice: `l` as a comprehension variable in the sort-order guard.
- `ruff format --check`: the file was correctly formatted on `main`; my edits
broke it. One of them left a docstring line at column 0, which `ruff format`
then "corrected" by over-indenting the rest of the paragraph — repaired at
the source rather than accepting that rewrite.
Verified the way CI does: local ruff is the pinned 0.12.11, and both
`ruff check` and `ruff format --check` run under bash over the full tracked
population (`git ls-files -z '*.py' '*.pyi' '*.ipynb'`, 46 files) are clean.
The population is counted, not assumed — an empty glob would pass vacuously,
which is the failure `scripts/tests` guards against elsewhere.
`scripts/tests` 1097 passed, 2 skipped after the reformat.
refs #763
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Population derived from `git ls-files`, not the issue's 9-key list (~21 claim sites).
Re-confirmed unchanged on 1.27.1: the distinct `skipped` commit-status state; `compare` serving
no `files`; no agent-side cancel route (REST route + swagger only); `branches: [main]` suppressing
the run off a non-main base.
Newly measured on four throwaway scratch bases, `main`'s rule never PATCHed: an absent required
context blocks an ORDINARY merge without needing `block_admin_merge_override` (that field governs
the FORCE path only), and `enable_bypass_allowlist` with an empty list is NOT a substitute for it.
Trap recorded: the PR API reports `mergeable: true` while such a merge is refused.
Left explicitly dated with reasons: push-supersession auto-cancel, `pull_request_target` overlap,
`--depth=1` no-merge-base, and the scope-enum/`reqRepoWriter`/403 items. Not a corpus sweep, and
`ci.actions-credential-scoping` now says so. `review-verdict.yml` untouched — #763 holds that file.
Five adversarial review rounds (21/12/9/6/2). Caveat: all same-model-family; Codex was rate-limited.
fixes#747
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
A sixth cold review found everything in round 8 clean except one line, and it
is the rule this branch keeps rediscovering: the test pinned the new
`raced_why` only by asserting the ABSENCE of the borrowed wording. Measured —
replacing the string with `zzz` left the suite green while an operator would
get `::error::… — zzz.` beside a sticky sentinel. The sibling test 330 lines
away states the rule and follows it; this one did not.
Now asserted positively, with the em-dash and full stop discriminating the
`::error::` reason from the `::warning::` text that continues ", which cannot
be true". The `zzz` mutation reddens it.
Three nits from the same review, all verified by execution rather than reading:
- the earlier fixture's row was excluded by the strict `> $since` because the
mark became its OWN id, not because it sat below the mark.
- the predecessor comment said `main` "warned only on `null`". True of the two
EMPTY shapes being contrasted; an empty body and a non-array object warned
as well. Scoped.
- `docs/ci-cd.md` and the record described the `::error::` as a two-way split
(found vs unverifiable). Round 8's whole argument is that a complete read
returning an IMPOSSIBLE answer is a third case, not a variety of the second
— which is the operator-facing point, since it decides whether to go looking
for an API failure that never happened. Both now say three.
The review re-verified, by comment-stripped diff, that round 8 changed no
executable line beyond the `raced_why` string and the if/elif restructure, and
independently reproduced both inertness measurements and the `origin/main`
predecessor behaviour.
Verification: `scripts/tests` 1097 passed, 2 skipped; decisions_validate and
build_decisions_catalog --check exit 0.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A fifth cold review confirmed the gate's behaviour is correct and proof-backed,
and blocked on three non-behavioural items. All three fixed; none touches the
shipped logic.
MEDIUM — the round-7 fixture narrated a raced human verdict it did not
construct. `null-page1-after-post` appended the row unconditionally, so it also
joined the PRE-write read and lifted the high-water mark above itself; removing
it changed nothing. The reviewer's suggested fix was to gate the append on the
post-write read. Measured after gating: still inert, because page 1 answers
`null` before any row reaches the wire.
So the row is gone rather than gated, and the prose now describes what the
fixture actually poses: a response asserting an empty history for a sha this job
wrote to must not be accepted as proof that nothing raced. Whether a verdict
really raced is not modelled and does not need to be — the response is not
evidence either way. A row the test cannot observe is decoration that reads as
coverage, which is the same class this branch has now been blocked on five
times.
LOW — the comment claimed the predecessor "at least produced a `::warning::`".
Half false, measured against `origin/main`: its `jq -e 'type == "array"'` gate
ACCEPTED `[]` silently and warned only on `null`. What is actually new is that
the paged walk reports such a read as a SUCCESS.
LOW — when the empty clause fired it set `ph_ok=no`, so the log said "could not
be read completely" beside a walk that completed on a validated terminator. The
answer was impossible, not unreadable, and an operator holding a sticky sentinel
needs to know which. It now carries its own `raced_why`, asserted by the test.
Both clauses mutation-proved: disarming the empty check, and reverting to the
borrowed wording, each redden the named test.
Verification: `scripts/tests` 1097 passed, 2 skipped; decisions_validate and
build_decisions_catalog --check exit 0.
refs #763
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A fourth cold review returned NOT-MERGEABLE on two Mediums. Both fixed, plus
its three Lows.
MEDIUM, and a defect this branch introduced. Tolerating a `null`/`[]` page 1 as
"complete, zero rows" is correct for the PRE-write caller — a head nothing has
posted to genuinely has no statuses — and impossible for the POST-write one,
which has just written a row to that sha. The body is well-formed, so nothing
retries it, and the walk reports success: `raced=0` concluded from a list that
cannot be real, on the one path whose failure direction is toward SUCCESS.
Worse than the code it replaced, which at least emitted a `::warning::` — a
logged fail-open had become an unlogged one. Reviewer measured both directions.
The post-write caller now rejects an empty result itself; the walk stays
caller-agnostic because the pre-write caller genuinely needs the empty answer.
This is NOT the withdrawn currency witness: that asked whether ANY row sat above
the mark, which an unrelated newer row satisfied while the rejection stayed
hidden, and it fired on schema-valid staleness. This asks only whether the list
is EMPTY — a state no unrelated row can produce and no ordering can disguise.
It carries neither defect. Proved by fixture; disarming it reddens the named
test, and the previously-uncovered `null`-at-page-1 clause is now covered too.
MEDIUM — the fourth overclaim of the same class, in the decision record body:
"Uncertainty must fail closed at both ends … Both repair now." The page-2 probe
was DELETED, not converted; it repairs nothing. It also contradicted the
record's own `rule:` ("the two directions are NOT symmetric") and the bullet
directly beneath it. Round 5 retracted this wording in `docs/ci-cd.md` only —
the sweep was by subject, not by the retracted words.
Also fixed:
- the record presented "an empty FIRST page is legitimate" as a property of
the walk; it is a property of the pre-write caller.
- `docs/ci-cd.md` called the numeric-only id comparisons a fix for mark
inflation; they are a TYPE guard, closing the string half. A corrupt but
genuinely numeric id still inflates the mark — not attacker-controllable,
since ids are server-assigned, and now stated rather than implied.
- `test_a_partial_mark_is_SAFE...`'s self-guard promised to detect that the
fallback ran; it keys on a warning emitted by a different condition, so
deleting the fallback left it green. Its sibling is what reddens; the
message now says what it actually pins.
- the order-faithful fixture appended the job's own POST after the reversal,
serving the NEWEST row on the OLDEST page — the opposite of DESC, in the one
fixture that exists to be ordering-faithful.
- "twice per walk" for the wasted sleep; it is once per walk, twice per run.
- a dead counter read in the DESC mode.
Rebased onto b16ec15d6 (the other session's #781/#799 docs work; no file
overlap, no conflicts).
Verification: `scripts/tests` 1097 passed, 2 skipped; fifteen executed mutations
across rounds 2-7; decisions_validate and build_decisions_catalog --check exit 0.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A fourth cold review (Opus, isolated worktree, tests/double/docs focus)
reported no correctness bugs in shipped behaviour but two coverage defects on
exactly the two things this change advertises. Both are closed.
The partial-mark fallback's safety is a claim ABOUT THE ORDERING — page 1 holds
the newest rows, so a walk that fails later still saw the true maximum. The
fixture pinning it served ASCENDING ids, i.e. the arrangement the design calls
unsafe, and passed anyway because the raced row's id sat above even the partial
mark. It could not distinguish safe from unsafe.
The stub now HONOURS the sort parameter: order-faithful modes serve DESC by
default and ASC when the request asks. The new fixture holds a PRE-EXISTING
base-mismatched verdict at id 7055 among 60 rows. Under DESC the salvaged mark
is 7059 and that row is below it — the exemption correctly stands. Under ASC
the mark would be 7049 and that untouched row tests as NEWER, a sticky repair
on a head nothing raced. So re-adding `sort=highestindex` now reddens by
BEHAVIOUR, not only by the structural assertion added in round 5. Measured:
re-adding it reds both tests.
Most modes stay ordering-blind on purpose and now say so: they test walk
COMPLETENESS, which is order-independent, and insertion order is what lets a
fixture place a row beyond page 1.
Also fixed:
- `null` is accepted as an empty page. An array-only gate is the exact shape
of #751 — `count_retargets` had one, the timeline really did return `null`
past the end, and the fence withheld EVERY exemption from the day it
shipped. The same narrowing here is worse, because this walk's failure is
the STICKY sentinel: every exempt PR would need a hand-posted verdict, per
head. Tolerating `null` cannot misread `[]`. Proved by fixture.
- the fail-closed comment said "past the 1000-row page cap"; the bound is 950,
as the walk's own comment and both docs already said.
- the docs claimed "only a read returning no rows at all abandons the mark".
False: a VALIDATED empty history yields a mark of 0 and is not abandoned —
that is the normal first run. What abandons it is a read that both FAILED
and returned nothing. Corrected in ci-cd.md and the record `rule:`.
- a comment pointed at the page-2 probe "a few lines further down"; it was
deleted, so the deixis pointed at nothing.
- the stub claimed its logical-read counter "is only reached on a SUCCESSFUL
page-1 serve" — measured false; it counts page-1 requests, retries included.
- five `(round N)` markers removed. A round number is session chronology and
does not parse for a reader who never saw it (`docs.no-session-narrative`);
an issue number does. The four that remain predate this change.
Verification: `scripts/tests` 1096 passed, 2 skipped. Thirteen executed
mutations across rounds 2-6. The reviewer independently re-ran the earlier
matrix and confirmed it, with one correction carried here: two of those
mutations redden MORE than their named test, so "each reddening exactly its
named test" was wrong — they redden at least it.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A third cold review (Opus, isolated worktree) returned NOT MERGEABLE with one
High and three Medium. All are addressed.
HIGH — the stated motivation was wrong, and self-contradictory once round 4
landed. Under the server default (`created_unix DESC`) page 1 holds the NEWEST
rows and ids are monotonic with `created_at`, so page 1 already carried the
true maximum id AND every row newer than the mark — the only rows the
post-write check selects on. A single-page read therefore missed a raced
verdict only if more than 50 rows were created INSIDE the write window, not
merely on "a head with more than 50 rows", which the issue, the comments and
the docs all asserted. Reviewer executed an order-faithful DESC stub: a
page-1-only reader repairs identically to the full walk.
What actually removed #761's stall is retiring #751's page-2 probe, not the
paging. The walk still earns its place, for a reason now stated instead of the
false one: it stops the gate's one fail-toward-SUCCESS path depending on an
undocumented ordering the server honours only coarsely (page 1 came back
`114,112,113,111,110`). That measurement was deleted in commit 1 and is
restored, since round 4's safety argument rests on exactly it.
MEDIUM/real defect — the string-id TWIN, live on `main` and one expression
away from the fix already made: `select((.id? // 0) > $since)`. jq orders
strings above every number, so a PRE-EXISTING row with `"id": "3"` reads as
newer than any mark, is counted as having raced the write, and gets the sticky
sentinel plus a false "was overwritten" on EVERY later run — a permanent
per-sha stall no re-trigger clears. Now numeric-only, with a test.
Also fixed: a non-empty history carrying no numeric id was collapsed to a mark
of 0 (making every pre-existing row look newer); it is now reported unusable
and the check is skipped. `sleep` no longer fires after the final attempt.
Three unpinned clauses now have tests, each proved by an executed mutation:
- the page cap is a refusal, not a terminator (1050-row fixture)
- the `::error::` found-vs-unverifiable distinction (forcing `raced_why=human`
reddened nothing before)
- the walk requests no sort order — a structural guard on round 4's
withdrawal, which nothing mechanical protected. It reads request LINES, not
comments, since the withdrawal note names the parameter to explain it.
Honest scoping, not new code: the test double is ordering-blind, so the paging
tests prove WALK COMPLETENESS, not that a real raced verdict would otherwise be
missed — under DESC it would not be. The stub comment and the docstrings now
say so rather than implying the stronger claim.
Docs: `ci-cd.md` and the record's `rule:` carry the corrected reachability, the
DESC dependency of the partial-mark fallback, and both rejected alternatives
stated as rejected alternatives rather than as draft chronology
(`docs.no-session-narrative`).
Verification: `scripts/tests` 1094 passed, 2 skipped; eleven executed
mutations across rounds 2-5, each reddening exactly its named test;
decisions_validate and build_decisions_catalog --check exit 0.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 3 added `sort=highestindex` to close a mid-walk-insert gap: under the
server default (`created_unix DESC`) a row inserted while the walk is running
lands at position 0, on a page already read, so the walk never sees it.
That fix and the round-2 partial-mark fallback are incompatible. ASC puts the
OLDEST rows on page 1, so an incomplete walk takes its high-water mark over
the oldest rows — leaving every pre-existing row above the mark and read as
"raced". That is a spurious STICKY repair on a head nothing raced, which is
precisely the #761 failure this whole issue exists to remove. Under the
default DESC the newest row is on page 1 by construction and ids are monotonic
with `created_at` (measured), so a partial mark is at or very near the true
maximum and "lower is safe" actually holds.
Two defects from one mechanism again, so the mechanism goes rather than
getting patched: the sort is withdrawn and the mid-walk-insert residual is
ACCEPTED and documented. It is bounded — a row arriving after this job's POST
is not one this job overwrote, and being newest it wins on the combined
endpoint branch protection reads.
Both the code comment and the docs record the withdrawal and the reason, so
the next reader does not re-adopt it.
Verification: `scripts/tests` 1090 passed, 2 skipped; the partial-mark mutation
still reddens `test_a_PRE_WRITE_paging_failure_still_yields_a_usable_high_water_mark`;
decisions_validate and build_decisions_catalog --check both exit 0.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two independent cold reviews (Codex GPT-5.6 cross-family, and an isolated
Opus agent) converged on the same blocker, which is fixed here along with
everything else they found.
BLOCKER — the mark walk turned a fail-closed case into a fail-open. The
high-water mark gates the post-write race check entirely: `max_id_before=-1`
skips it. Before paging, only a failure of the single page-1 request could
reach that. Requiring a COMPLETE walk newly routed a page-2 hiccup, an
over-cap history, or one malformed id on a later page into the same hole, so
a human rejection racing the write was left green where `main` repaired.
A partial list now still yields a mark: it can only be LOWER than the true
maximum, which makes the check more eager, never blinder. Only a read
returning no rows at all abandons it — the pre-existing #849 gap, unchanged
and now asserted by a test so it stays visible.
WITHDRAWN — the "currency witness". It produced two defects from one
mechanism, which is the signal to remove rather than patch twice: counting
ANY row above the mark does not witness this job's write, so a stale-but-valid
snapshot carrying an unrelated newer row passed while hiding a rejection; and
a schema-valid stale read is not retried, so one such response turned a
transient anomaly into a permanent sentinel. The hazard has no mechanism here
either — Gitea is a single instance with no read replicas. Removing it
restores the pre-change exposure on that path, a non-regression.
Also fixed, each a fail-open with a fixture and an executed mutation:
- `.creator` is type-tested before indexing. `.creator.login` on a non-object
exits jq 5 and `set -e` took the step down after the green was posted and
before the repair. Reproduced by both reviewers.
- the mark is the max over NUMERIC ids only. jq orders strings above every
number, so one `"id": "99999"` passed the numeric gate and inflated the
mark until nothing looked newer.
- an unusable `raced` count now repairs instead of "not acting on it".
- `sort=highestindex` (ASC, measured) so a row inserted mid-walk appends at
the end rather than at position 0 on a page already read. An unknown sort
value silently falls back to DESC, so this is insurance, not load-bearing,
and the comment says so.
- `ph_ok`/`ph_rows` renamed off `read_existing_verdict`'s `st_ok`. No live
bug, but a name collision in a 1400-line step.
Tests the reviews showed were missing, each proved by an executed mutation:
- verdict beyond a SHORT page (a deliberately unfaithful truncated response
— against a faithful double a short page is always the last, so the rule
"terminate only on an EMPTY page" was unobservable)
- pre-write paging failure still yields a usable mark
- pre-write read returning nothing abandons the mark and says so
- a TRANSIENT page failure is retried (the retry was unproven code: every
other error mode fails on every attempt, so disarming it reddened nothing)
- a string id cannot inflate the mark
- a malformed `creator` row does not kill the job
Stub corrections, both the same class as the earlier `[]`-vs-`null` gap: it
served one flat list (so paging was unobservable) and computed its own-post id
with `max()` over mixed str/int, which raised TypeError and made the string-id
test pass because the DOUBLE crashed rather than because the mark was right.
Mutation matrix, all executed, each reddening exactly its named test: retry
disarmed; numeric-max reverted; partial-mark fallback removed; short-page
terminates; page-1-only walk; post-write fail-closed flipped open; jq
type-guard reverted. The unusable-count arm is unreachable by any fixture and
is annotated as such rather than claimed as proved.
Verification: `scripts/tests` 1090 passed, 2 skipped; decisions_validate and
build_decisions_catalog --check both exit 0; terminator, clamp, sort order and
id monotonicity all re-measured live on Gitea 1.27.1.
refs #763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`review-verdict.yml` read the per-POST status history twice with a single
`?limit=100` request. `limit` clamps to the server-wide `MAX_RESPONSE_ITEMS`
(measured 50), so on a head carrying more rows than the clamp both reads saw a
partial list. The high-water mark was only page 1's maximum, and — the direction
that matters — a raced human verdict beyond page 1 was invisible to the
post-write race check, leaving a forged green over a rejection.
Both reads now walk to a validated empty page (`[]` on this endpoint, measured
2026-08-28 against PR #761's 114-row head: pages 1-2 return 50, page 3 returns
14, page 4 is `[]`), never terminating on a short page, under a 20-page cap and
retrying each page once. Correctness does not depend on the cap value.
This retires #751's page-2 "assume raced" probe, which repaired every head that
outgrew one page. It fired on Renovate PR #761: an `::error::` claimed a human
verdict had been overwritten on a head carrying none, and the sticky sentinel
then refused re-exemption on every later run.
Two properties replace it. Uncertainty now fails closed at both ends — the
unreadable-history branch warned and left the exemption green while the page-2
probe repaired on the same uncertainty, one check disagreeing with itself; this
is affordable only because paging removed the common trigger. And the post-write
read must witness the job's own write: reaching a validated empty page proves the
walk finished, not that it saw a current list, so at least one row above the
pre-write mark must exist because the job just posted one.
The `::error::` now distinguishes a verdict actually found from an unverifiable
read. The sentinel description stays generic — the classification recognises it
as a fixed point, so its wording is load-bearing.
The stub gained faithful paging (50-row slices, `[]` past the end, one snapshot
per logical read so a counter mode cannot describe two different histories across
pages) and, separately, modelling of the job's own POST appearing in the history
— which it had never done, so in its world every ordinary run looked like a head
nothing had been posted to. `own-write-invisible` withholds exactly that detail
as the negative control for the currency witness.
Mutation-proved by execution, one clause at a time:
- walk reads page 1 only -> RUNNING_PAST_PAGE_1_is_PAGED_and_the_exemption_
STANDS, raced_verdict_on_PAGE_2_is_detected_and_repaired and both UNREADABLE
history tests go red
- currency-witness zero branch deleted -> CANNOT_SEE_OUR_OWN_WRITE red
- fail-closed flipped to fail-open -> both UNREADABLE history tests red
Verification: `scripts/tests` 1085 passed, 2 skipped; decisions_validate and
build_decisions_catalog --check both exit 0.
fixes#763
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
§5.3's verdicts rested on a single surface, which manufactured four false zeros: codex is driven
through `codex exec` inside Bash, security-guidance and ralph-loop expose no tool at all and run as
hooks (1,086 executions each), and feature-dev is used through its agents. The audit also compared
current enablement against historical usage — six of the eight plugins it called "genuinely unused"
were disabled for 16 of the 30 corpus days.
The retirement half of #781 is answered *no* on evidence: the zeros split six ways and only one is
grounds for removal. Eight plugins are kept by operator decision.
#799's observation was correct and its cause is now established. serena was `false` in settings.json
until 2026-08-14T12:31Z, when a concurrent session enabled it; its tools appear in no transcript
before 12:42:54Z. #799's session started at 12:01Z and never reloaded, so its probe correctly found
nothing while the settings file already said `true`. serena is adopted and documented as the third
code-intelligence surface.
Four review rounds, two independent cold reviewers (one cross-family); rounds 1-3 BLOCKED.
fixes#781fixes#799
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Closes the push route into ci-image.yml (#744) and ships the persist-credentials guard that was waiting on it (#835).
ci-image.yml's push trigger had no branches: filter and was path-scoped to docker/ci/** AND to the workflow file itself. Gitea resolves a push workflow's definition from the pushed ref, so any branch push touching those paths ran that branch's own YAML on a docker-capable runner holding the credential that writes ersatztv:prod and the ersatztv-ci:<sha> five container: jobs execute.
Be precise about what the filter buys: it is loaded from the pushed ref like the rest of the file, so a branch that deletes it re-enables the route. This closes the DRIVE-BY case - publication as a side effect of an ordinary push - and is not a boundary against a writer who intends to run their own YAML. The wider class is #853.
The self-reference left both paths: and ci-image-pin's expected in the same change - a decided tradeoff with both prices stated, not a necessity. Branch publishing moves to workflow_dispatch, probed live: run 2340 on this branch published ersatztv-ci:43b1e45 and left :latest unchanged.
With both mechanical blockers gone, ci-image.yml's checkout takes persist-credentials: false (16 of 16) and scripts/tests/test_workflow_persist_credentials.py holds the convention with NO exemption list - git-index population, declared clause mutation re-run every suite, guard-inventory rows.
fixes#744fixes#835
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
`review-verdict.yml` decided whether an existing `review-verdict/h10` was worth INHERITING by
testing `.creator.login != null` — satisfied by any account's credential, including the `renovate`
bot's `RENOVATE_TOKEN`, a `write:repository` PAT that cannot be scoped down the way #697 scoped the
registry credential. The test is now membership in `H10_REVIEWERS="timothy"`, a literal in the
base-resolved definition.
The design that survived 11 cold review rounds:
* `read_existing_verdict` carries TWO flags. `ex_human` (attributable AND allow-listed) gates
INHERITANCE; `ex_attributable` gates the last-moment re-read, which asks the opposite question and
must stay broad. Narrowing both — the first draft — makes the job post its exemption over a
mid-run rejection, and the post-write repair does not cover that.
* The two calls no longer compute an identical predicate, so "changed" is made explicit: the
state/creator/description triple from the first read is snapshotted and compared.
* The allow-list governs an inherited `success` ONLY. An existing `failure` inherits on
attributability alone, because inheriting a rejection can only withhold an exemption while
re-deriving one can turn it green on an exempt PR. A symmetric rule was a measured fail-open.
* The post-write raced check stays broad — not because narrowing would let a rejection go green
(a real reviewer is on the list by construction), but for the misconfiguration case.
Two mechanisms were WITHDRAWN rather than patched a third time, and both withdrawals are recorded
in `ci.exemption-provenance` so they are not re-attempted: a `::warning::` annotation that produced
three defects in three rounds, and a post-write fix whose generic `pending` would have been
re-derived anyway and which had no retry trigger.
Verified: the inheritance predicate driven against the LIVE Gitea API on a probe-named context,
both allow-list directions; every clause mutation-proven against the shipped file; `scripts/tests`
1012 passed, 2 skipped.
Follow-ups filed: #845 (post-review-verdict.sh does not check its own account is allow-listed) and
#849 (post-write verification: three routes leaving an exemption `success` over a human `failure`,
plus the retarget fence's post-POST gap, plus the prose sweep that lands with the behaviour).
fixes#742
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
The verdict words lived in two hand-written shell copies — the `case` arms of
post-review-verdict.sh (write) and the POS_RE/NEG_RE regexes of
check-review-verdict.sh (read) — held together by nothing but a comment that had
already gone stale. scripts/lib/review-verdict-vocabulary.sh now declares them
once and both sides derive; neither script enumerates a verdict word any more.
Only the WORD SET moved. The grammar stays in check-review-verdict.sh, where
every #629 false-open actually lived.
No parity test: #774 shipped one and withdrew it after six rounds, because a
regex over shell source is not a shell parser. The proof is behavioural and
graded MUTATION — the harness restores the pre-#788 hardcoded POS_RE each run and
requires it to redden.
Enforcement is a DATA dependency, not a control-flow gate. Review round 1 found a
real fail-open in the first commit: `${#arr[@]}` is nounset-safe only for a
declared-empty array, and under `set -u` that error inside a function called as
`if ! validate` skips BOTH branches — so on the reader (deliberately no `set -e`)
an explicit BLOCKED @ head classified `positive`, exit 0. Validation now sets a
sentinel on its last line and the derived views refuse without it.
Six cold review rounds; rounds 2-6 found no fail-open across differential fuzzing
(4788 / 2612 / 7560 payloads, zero divergences from origin/main's grammar),
sentinel forgery, environment poisoning, declare -p evasion on bash 5.3 and 3.2,
path/symlink resolution and probe TOCTOU. Every malformation fails closed: reader
exit 2, writer exit 1 with nothing posted.
Also corrected: CLAUDE.md and release.review-verdict-gate both enumerated the
vocabulary without LGTM, a word the code has accepted since #629.
fixes#788
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Implements the three-level field-help pattern from #734 as a shared component: field name + optional one-sentence summary → a one-short-paragraph panel behind a consistent Info-icon trigger → a future external-docs deep link (`docsHref`, built and typed; no screen passes one yet).
Adopted on FFmpegProfilesScreen (9 fields), documented as docs/spa-conventions.md §15 with decision record `spa.field-progressive-disclosure`, and mirrored into the design-system prototype.
The panel is portalled to document.body: `.ctv-card` sets `overflow: hidden`, which clips a positioned descendant whatever its z-index, and one field's explainer rendered 12px of a 92px paragraph in every state of the Audio card.
A `::before` hover bridge was added and then WITHDRAWN — it held for a vertical descent onto the panel and failed for a diagonal one, leaving a safe sideways exit of 1.25px on an 18px icon. Hover reads the paragraph in place; the panel's interactive content is reached by pinning.
Four cold adversarial review rounds; the first three returned BLOCKED. They found five wrong copy claims across nine paragraphs and two vacuous tests in a row for the same mechanism.
Deferred with owners: #839 (placement verified by hand, not by a test) and #840 (the portal puts a docsHref link at the end of the tab order).
fixes#734
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
Both search indexers opened UpdateSong with
metadata.AlbumArtists ??= [];
metadata.Artists ??= [];
Artists/AlbumArtists hold the whole list in ONE COLUMN rather than
being navigations. So unlike the same `??= []` idiom on
Genres/Tags/Artwork all around them, the property IS the column value:
assigning it on a TRACKED entity flips the entry to Modified and the
next SaveChanges writes [] over a NULL column. This is the mechanism an
adversarial review demonstrated in #691, which is why that issue's
entity-level guard was reverted in favour of guarding at the read site.
Measured rather than reasoned about, per the issue's first done-when
box. Restoring ONLY the `??= []` clause (the real predecessor lines,
not a hand-written mutant) reddens the new fixture on
`metadata.Artists should be null but was []`; a probe variant with the
first two assertions replaced by prints reports STATE=Modified and the
raw column moving from NULL to "[]". Today's two feeds are both
AsNoTracking (SearchRepository.GetItemToIndex and GetAllSongs), so no
shipped caller loses data -- but that is a property of two callers, not
of the indexer, and #691 already recorded it as a loaded gun. The
fixture pins the indexer's own contract instead.
Removing the assignment is not sufficient alone: it was load-bearing
for the four reads below it, and deleting it by itself converts a
silent write into a live throw on every untagged song. Measured by
deleting only those two lines from the real predecessor file:
NullReferenceException, thrown at the foreach (cited by symbol: a line
number in a mutant that exists in no committed tree is unreproducible
by construction). The
exception type follows the read FORM, not the field -- foreach yields
NRE, string.Join/ToList yield ArgumentNullException -- and this PR
contains two of each, which is why no single exception-name grep
characterises the class. So each site moves together with its reads:
- LuceneSearchIndex.UpdateSong / ElasticSearchIndex.UpdateSong: hoist
Optional(...).Flatten().ToList() locals and read those.
- RefreshChannelDataHandler: the Scriban context took the raw nullable
lists (the issue's second item). The shipped _song.sbntxt only does
array.join, but a custom template is free to do anything.
The population was derived from the MODEL rather than from the issue's
file list, and the obvious derivation is wrong: "the IList<string>
properties under ErsatzTV.Core/Domain" returns two of eight. It misses
the six value-converted collections (ProgramScheduleAlternate and
PlayoutTemplate each carrying DaysOfMonth, MonthsOfYear, DaysOfWeek),
declared as plain ICollection<T> and made single columns only in
Data/Configurations -- and their storage differs (comma-separated text
for the int converter, JSON for the enum one), so the shared property
is "one scalar column", not the serialization. No site applies `??=`
to any of the six, so this defect has no instance there; whether a null
can REACH one at runtime is unverified and is filed as #823 rather than
asserted either way. Only the SongMetadata pair is left NULL in
practice, by FallbackMetadataProvider. Every site touching either field
was then swept; the remaining readers were already guarded by #691.
The fixture carries two anti-vacuity guards, both witnessed:
- A POSITIVE CONTROL (`writer.NumDocs.ShouldBe(1)`). Every other
assertion says something did NOT happen, so all of them hold
vacuously if UpdateSong never runs -- and it silently stops running
if a future refactor gates UpdateItems on `_initialized`, which this
fixture bypasses by injecting the writer. Verified BOTH directions:
with that gate added the control fails `NumDocs should be 1 but was
0`, and with the control removed the whole test PASSES while the code
under test is unreachable.
- A capturing logger, because UpdateSong wraps its body in a catch that
assigns metadata.Song = null -- severing a required relationship and
cascading the metadata to Deleted. Without it the probe silently
measures the error path; on the first run it did exactly that (a bare
ILanguageCodeService substitute NPEs inside AddLanguages). The raw
column helper also fails loudly on a missing row, since ExecuteScalar
returns CLR null for both "NULL column" and "no such row".
ElasticSearchIndex has no equivalent fixture -- it needs a stubbed
transport -- so its change is by inspection against the Lucene one, and
the gap is filed as #824 rather than covered by a source-text guard.
The whitespace-only churn in ElasticSearchIndex.cs is the #311
fix-as-you-touch format gate: it scopes to whole changed FILES.
`git diff -w` over that file shows only the two hunks above.
Local gate: ErsatzTV.Tests 2006 passed / 4 pre-existing skips,
Core.Tests 685/1 skip, Infrastructure.Tests 114, Architecture.Tests 7,
Scanner.Tests 1504 -- 0 failures in each. scripts/tests 874 passed / 2
skipped. dotnet format whitespace --verify-no-changes clean on the four
touched files, no BOM on any. decisions_validate OK.
Fixes#701
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Graded a nit and not blocked on, but it is a claim about a neighbouring subsystem that is
one notch too strong: the hook decides on the combined `.state` and its filters exclude
`skipped` (#593). Left as-is it would have taught the next reader that any non-success
context denies, which is what #593 exists to correct.
refs #772
Both remaining review findings were the same shape as the one before them, and it is the
shape this repo keeps recording: a claim corrected in one place, its copy left standing
somewhere else in the tree.
* `docs/ci-cd.md` said "gates nothing" in the small-lane paragraph while the section 1441
lines below said the opposite. A red preflight lands in the PR's combined status, which
the merge gate reads (#598) — what it does not do is SKIP the jobs it diagnoses, and
that is now the sentence in both places.
* Two docstrings in the preflight's test file still described the disarmed script as
warning and exiting 0. Built the mutant and ran it: it emits an error and exits 1. The
exit code separates nothing now that an unverifiable answer fails too — the DIAGNOSTIC
is what the mutation destroys, which is what `mutation_manifest.py` already said and
the prose next to it contradicted.
Nits from the same pass: the admin-cron URL is quoted (`?` globs in zsh, the operator's
shell); the retry assertion's message quoted a threshold it does not use; the arm table
omitted the malformed-credential shape the code and tests both have; `buildx inspect` no
longer `--bootstrap`s a builder just to read its name, and an empty capture no longer
produces a noisy `buildx use ""`.
Swept the tree for the shape rather than the two reported lines: the surviving "exits 0"
and "could-not-tell" hits are other subsystems, or the concept named as a concept.
refs #772
The advisory narrative check was right about both new passages: 'an earlier draft warned and
exited 0' and 'both cold reviews found it independently' only parse to someone who saw the
session. What a cold reader needs is that warning-and-exiting-0 is the natural way to write
this check and is wrong, and what it costs — which is now what the doc says.
refs #772
The re-review's one HIGH was mine and was the obvious one to miss: the previous commit changed
the preflight so an unverifiable answer FAILS, and left a `docs/ci-cd.md` paragraph two
screens away still saying "anything else is reported as could-not-tell". That paragraph is
the one an operator reads when the job goes red, and it would have talked them into
reinstating the defect. Replaced with the full arm table, including the two rows the first
draft got wrong and why.
* "gates nothing" was false in the way this repo has recorded before (#598): the
merge-consent hook reads the COMBINED status, so a red preflight blocks the merge like
any other red job. It does not SKIP the jobs it diagnoses; that is the accurate claim,
in ci-cd.md and in the remote-state row.
* The production retry defaults were evaluated by nothing — every test overrode both
knobs. A test now drops the overrides and measures three attempts and a real pause, so
editing the default to 1/0 (which would falsify the "a blip does not redden a PR"
argument) goes red.
* `journalctl -u gitea | grep ExecuteCleanupRules` is not a reproduction: that identifier
reaches the log only through slow-query warnings, so an empty grep on a healthy host
reads as "the rule never ran" — the inverse. Replaced with the admin cron API, which
answers deterministically.
* The recovery recipe's `docker buildx use default` needs the containerd image store to
`--push` (both named hosts have it, checked today) and mutated the operator's builder
selection without restoring it.
* The stub's comment claimed both halves of real curl's transport failure mattered; only
the exit status is observable, because `|| resp=""` discards what curl printed.
* The empty-half credential refusal echoed the username; it needs no value at all. The
401/403 arm aborts the remaining pins while 404 continues — deliberate, now stated.
* `curl -u "$VAR"` puts a credential in argv, and this job runs container-free on a shared
host. NOT fixed here: it is the shape all five `scripts/` callers already use, so fixing
one site leaves the class and splits the codebase. Filed as #821 and named at the site.
refs #772
refs #792
Two independent reviewers (one cross-family) converged on the same defect, and it was the
important one: the preflight WARNED and exited 0 on every answer that was not 200 or 404,
so a missing `curl`, a moved registry or a DNS change would have left it green forever —
"the check could not run" presenting as "the pin is fine", in a script whose own header
disclaimed exactly that. Unknown answers are now retried (3x, 5s) and then FAIL, with
wording kept distinct from the deleted case because the two send an operator to different
places.
Also from the reviews:
* An absent secret does not arrive as an unset variable. `${{ secrets.X }}:${{ secrets.Y }}`
interpolates to ":", a perfectly non-empty and perfectly useless credential, and the
tests covered only the unset shape. Both halves are now required, and the parametrised
test drives the production shape.
* HTTP 200 is not a manifest. A proxy or a login page answers 200 too, so the body is
fetched and matched for `schemaVersion` (a shell `case`, so no jq dependency and no
pipeline that can inject).
* The curl stub ignored `-u` and answered 200 regardless, so deleting the real `-u` would
have left the suite green while the live registry rejected every request. It now 401s an
unauthenticated read, as the registry does.
* The mutation's declared diagnostic changed with the script: now that unknown fails too,
the exit code no longer separates "deleted" from "could not check", so the proof turns on
the message and `expect` says so.
* docs/ci-cd.md: `scan` is no longer the only `docker-build.yml` job on the small lane, so
the tag-push exclusivity claim and the lane membership were both false. Fixed.
* "Immutable" was overstated: `ci-image.yml` tags `rev-parse --short HEAD`, so a dispatch or
a weekly no-cache run at the same HEAD republishes that tag from a rebuilt image. Stated,
along with what the rebuild recovery does NOT restore (mutable bases and apt, so equivalent
rather than bit-identical).
* The recovery recipe left you in a worktree checked out at the pin commit — where the
verify script does not exist, and where the workflow carries the pre-bump pin. It now
keeps `$repo`, returns, and removes the worktree. It also needed BuildKit's `http = true`
caveat: the container driver does not inherit the daemon's insecure-registries.
* The root cause carries its evidentiary limit and its reproduction commands, and says what
to conclude if a pin vanishes after server-management#842 lands (refuted, not re-applied).
* The `ci.required-job-step-execution-markers` carve-out named one container-free job; there
are two now, and the membership is what rots.
* The decision record's `''` YAML escapes leaked into rendered prose; "status, no comment ->
ask" is qualified (a prior positive verdict for the SAME head still satisfies condition
(c)); "exits 1" is "exits non-zero" (usage exits 2, jq its own status, signals 128+n).
refs #772
refs #792
Decisions-Edit: yes
#772 — the pinned CI toolchain image can be deleted out from under us, and when it was
(2026-08-11..13) all five `container:` jobs died at image pull, both required contexts
included, with the cause buried in each job's log. Root cause is registry-side and is now
established rather than guessed: an owner-level Gitea package cleanup rule (keep_count 15,
remove_days 1, remove_pattern `.*`, keep_pattern no 7-hex sha can match) deletes a sha tag
once 15 newer versions exist, and `ExecuteCleanupRules` ran nightly through the window. The
`ersatztv` package carries the same rule's fingerprint exactly — every sha tag older than
the 15-slot window is gone, every keep_pattern tag back to 26.3.1 survives. Version deletes
leave no audit row, so the specific run cannot be replayed; that limit is stated where the
claim is made. The durable fix belongs to the registry's repo: server-management#842.
What lands here is what a consumer of someone else's registry can do:
* `toolchain-preflight`, a container-free job (a job consuming the image could not run to
report it missing) resolving every pin against the registry and failing with a message
that names the tag and the recovery. Not a `needs:` of the jobs it diagnoses — gating
five jobs behind a checkout and one curl taxes every green run to speed up a rare red
one, and they already fail fast.
* Only HTTP 404 means gone. Everything else is could-not-tell, and rejected credentials
fail rather than pass as unknown — "the check could not run" must never present as
"the pin is fine".
* A recovery path that does not need CI: rebuild the SAME tag from the commit it names
and push it. The push half was verified against this registry on 2026-08-22 with a
throwaway package (created, resolved 200, deleted).
#792 — the reported defect was the exit code, and re-measuring says that premise is false:
every no-status path already exits 1, and eight refusal modes now assert it against the real
predecessor, where they pass. The observed 0 came from the invocation, not the script. What
WAS broken is the half-state the issue describes second: the comment was written before the
status, so every refusal left `Review-verdict: MERGEABLE @ <head>` on a PR with no gating
status behind it. The two writes are now ordered status-then-comment, which makes the only
reachable half-state the safe one — a status with no comment leaves the merge hook's
condition (c) with nothing to classify, which is an `ask`. The refusals themselves are
untouched. Ordering rather than compensating deletion: an orphaned-comment cleanup needs a
Gitea call, and these refusals are usually caused by Gitea being unreachable.
Proof for the ordering is the split against origin/main's script: the 8 orphan/ordering
tests go red there, the 8 exit-code tests stay green.
fixes#772fixes#792
Refs: server-management#842
Decisions-Edit: yes
`testing.guard-derives-population-from-source` (#774) was silent on the commonest
population in our own guards — files in a directory — and every one answered with a
filesystem walk. A walk is not authoritative: it reports build output, generated
shims and editor droppings, and differs per machine. #778 measured the cost by
getting the same population wrong three times in one PR.
CONVERTED (a completeness claim over tracked files): `test_guard_inventory.py`,
`test_hook_fire_log.py`, `test_ci_image_pin_population.py` (which also gained
`*.yaml`), `test_remote_state_inventory.py` (folded onto the shared derivation), and
`test_pr_changed_files.py` (not on the issue's list — found by sweeping the whole
repo).
ASSESSED AND RECORDED, not silently skipped: `_repo_copy` takes its file list from
the index for hermeticity though it makes no completeness claim;
`test_ci_dropped_step_guard.py` has no filesystem population at all; the decisions
corpus is recorded as unexamined rather than cleared; and the SPA page-size guard is
deferred to #819 with its obstacle documented. This is not "replace every glob".
`scripts/tests/tracked_files.py` is the single derivation.
`test_guard_populations_derive_from_git.py` proves it in two measured complements:
exhaustive removal catches a hardcoded `.exists()` admit and memoisation; the call
log catches an append-only source that yields nothing on this machine — #778's
shape — which removal cannot see because it has nothing to remove.
Twelve rounds of independent cold review, alternating model families in isolated
worktrees. The production derivations were confirmed sound every round; every
blocking finding after the first was in the proofs or in prose claims about them.
Counts over growing populations were removed rather than corrected, after three
drifted.
fixes#806
Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
2026-08-22 16:07:58 +00:00
333 changed files with 56203 additions and 2548 deletions
# --- Docs-only exemption: if every changed file is docs/process, skip the gate. ---
# The file list must be enumerated EXHAUSTIVELY, validated row by row, and bound to ONE head, or the
# exemption is unsafe. ALL of that now lives in scripts/pr-changed-files.sh — the single shared
# The file list must be enumerated EXHAUSTIVELY, validated row by row, and checked for head/base
# movement across the paging round trips, or the exemption is unsafe. (That check detects ONE-WAY
# movement only — this said "bound to ONE head" until 2026-08-28, ersatztv#803.) ALL of that now lives in scripts/pr-changed-files.sh — the single shared
# implementation, also called by .gitea/workflows/review-verdict.yml (ersatztv#649).
#
# Why it moved: this logic was written twice. This copy is ADVISORY (a failure produces a human
@@ -173,11 +186,11 @@ fi
# posted before ersatztv#632 and gets NO opinion, rather than denying every in-flight PR the day
# this lands. The window closes on its own — verdicts are per-head and short-lived, so every verdict
# posted after this carries the field.
# "Could not check" is a THIRD outcome, distinct from both "matches" and "no base recorded". Cold
# review found the first draft collapsing it into the latter: an unreadable status response yielded
# an empty `recorded_base`, which took the graceful-adoption path and skipped validation silently —
# "Could not check" is a THIRD outcome, distinct from both "matches" and "no base recorded".
# Collapsing it into the latter is a false-open: an unreadable status response yields
# an empty `recorded_base`, which takes the graceful-adoption path and skips validation silently —
# after which a later, successful status read could still auto-grant. A transient failure would then
# have produced a "merge gate: satisfied" message for a comparison that never happened. Every
# produce a "merge gate: satisfied" message for a comparison that never happened. Every
# unreadable input here therefore falls through to a human (`ask`), never to silence.
# RE-READ THE BASE HERE, ONCE, FOR EVERY PATH BELOW (ersatztv#778).
#
@@ -194,8 +207,8 @@ fi
# of `process.check-and-use-pins-a-version`, so the guard enforcing that rule must not break it.
#
# This re-read first landed inside the scheduled-auto-merge branch only, which fixed the branch-
# protection lookup and left the #632 retarget DETECTION below still reading the stale snapshot. Cold
# review demonstrated the consequence with this repo's own fixture: scheduled+retarget denied, while
# protection lookup and left the #632 retarget DETECTION below still reading the stale snapshot.
# Measured on this repo's own fixture: scheduled+retarget denied, while
# immediate+retarget auto-GRANTED. That is the twin-missed shape — a fix applied to the path where it
# was noticed — so the re-read is hoisted above every consumer rather than duplicated into each.
prjson_now=$(gq "repos/$owner/$repo/pulls/$pr")
@@ -212,6 +225,48 @@ fi
# From here on both names are the freshly-confirmed base; they are equal by the check above.
base_ref=$base_now
live_base=$base_now
# THE HEAD IS RE-READ AT THE SAME HOIST, FROM THE SAME RESPONSE (ersatztv#803).
#
# `$sha` comes from the PR snapshot at the top of this hook, and until 2026-08-28 every later check
# consumed that captured value: the CI combined status, the `review-verdict/h10` status, and the
# verdict-comment classification were all evaluated against `/commits/$sha/status` and `--head $sha`.
# A push landing in the gap — which includes the docs-only enumeration's up-to-forty round trips —
# was therefore checked against the commit it had just replaced, and the hook would report "a
# positive Review-verdict references the current head" about a head that was no longer current.
#
# This is the SAME defect the base had until #778 hoisted the re-read above, and it is fixed the same
# way rather than a different way. Reading `.head.sha` off `$prjson_now` — the response the base
# check already fetched — costs NO extra round trip, and it keeps the two axes on ONE snapshot, so
# they cannot disagree about which moment they describe. Two separate reads would answer about two
# different instants while reading as one check.
#
# DENY, not ask, and for the same reason the `stale` verdict class denies: a head that moved means
# the verdict this hook is about to accept covers an OLDER commit, which is a state we have
# positively established rather than failed to establish. An UNREADABLE `.head.sha` is the different
# case and asks.
#
# WHAT THIS DOES NOT CLOSE, said here rather than left to be inferred. A push landing after this
# check still passes, exactly as a retarget does — the file's rule against a second re-read applies
# unchanged (see the branch-protection block below), because two reads only move the window rather
# than closing it. That residual is bounded server-side and this hook is not what bounds it: the new
# head has no `review-verdict/h10` status, and that context is REQUIRED on `main`, so Gitea refuses
# the merge (#622). The hook's job here is to stop CLAIMING a head is reviewed when it can see that
# it is not — an advisory gate that states something false is worse than one that asks.
decide ask "H10 merge gate: PR #$pr reports no head commit (.head.sha) on re-read, so whether the review verdict still covers the current head could not be confirmed. Check the PR, then merge."
fi
if["$sha_now" !="$sha"];then
decide deny "H6/H10 merge gate: BLOCKED — PR #$pr's head moved from ${sha:0:7} to ${sha_now:0:7} while this gate was evaluating. Every check formed against ${sha:0:7} — the changed-file enumeration, the CI status and the review verdict — describes a commit that is no longer the one being merged (ersatztv#803). Re-review the current head and run: scripts/post-review-verdict.sh $pr MERGEABLE"
fi
# From here on `$sha` is the freshly-confirmed head; the two are equal by the check above. Mirrors
# `base_ref=$base_now` a few lines up, and is written for the same reason that one is: it makes the
# value every later check consumes the one that was just re-read, so a future edit moving a
# consumer above this point fails visibly rather than silently reading the stale capture.
sha=$sha_now
fi
if[ -n "$sha"];then
# This is the THIRD read of this endpoint in a worst-case hook run (the ordinary-CI branch and the
# scheduled-auto-merge branch each do their own). Sharing one snapshot would close a narrow
@@ -278,6 +333,50 @@ for n in $issues; do
fi
done
# ONE branch-protection READ per run (ersatztv#859). Two arms consume this endpoint — the scheduled
# path's `review-verdict/h10` required-check test, and the guard-scope freshness check at the bottom
# — and they used to issue independent GETs, so a scheduled auto-merge hit it twice (measured: the
# test stub recorded 2 URLs).
#
# THE ROUND TRIP IS THE SMALLER HALF. What matters is that branch protection is MUTABLE config: two
# reads can return two different answers, and the gap between them is a gap in which the two arms
# decide about different repo states — one concluding `review-verdict/h10` is required on the base
# while the other classifies a rule list that no longer says so. Neither arm can detect that; both
# would report confidently. Caching makes a single run internally consistent BY CONSTRUCTION, which
# is a property no retry or ordering change can supply.
#
# WHY #787 DID NOT ALREADY SHARE IT, since the obvious question is why two reads existed at all: the
# arms ask genuinely different QUESTIONS — one about `$base_ref` and its required contexts, one about
# `main` and snapshot freshness — so their classifications must stay separate. But they ask those
# questions of the same URL with the same credentials, so the RESPONSE is shareable even though the
# verdicts are not. Cache the bytes; never cache a verdict.
#
# This does NOT pin anything: protection can still change after the read, and the honest ceiling is
# unchanged (`process.check-and-use-pins-a-version`). It removes a second window, it does not remove
# the first.
bp_fetched=no
bp_cache=""
bp_cache_code=""
fetch_branch_protections(){
# Idempotent by design: every caller invokes it unconditionally and the FIRST one pays. A caller
# that had to know whether it was first would be a second place for the two arms to disagree.
if["$bp_fetched"= yes ];thenreturn 0;fi
bp_fetched=yes
local f
# A temp-file failure gets its own sentinel rather than an HTTP-shaped one, so each caller can
# keep the distinct message it had before this was shared. Reporting a mktemp failure as HTTP
# '000 — Gitea unreachable' would state a cause that did not happen, which is the defect class
# --- (a) CI combined status must be green (unless deferring to Gitea's own check-gate). ---
if["$mwcs" !="true"];then
[ -n "$sha"]|| decide ask "H6 merge gate: could not resolve PR #$pr head sha to check CI. Verify CI is green before merging."
@@ -402,7 +501,7 @@ else
# performs no matching and knows nothing about precedence, so a 200 from it means only "a rule
# with this NAME exists and lists this context", never "this context is required on this branch".
#
# It was used first, with the list consulted only on a 404, and cold review found what that left
# It was used first, with the list consulted only on a 404, and that design left a false-open
# behind: the precedence argument below guarded the 404 path while the 200 path — the one this
# repo actually takes — granted without it. Given a rule `main` requiring `review-verdict/h10` and
# a rule `m*` with better Priority that does not, Gitea applies `m*`, and the by-name hit on
@@ -410,13 +509,12 @@ else
# the twin rather than documenting it is the point: one fetch, one classifier, one argument, and
# no second path to keep in step. The ref no longer reaches a URL segment, so it needs no
# encoding either.
bp_file=$(mktemp)|| decide ask "H6/H10 merge gate: could not allocate a temp file to read branchprotection for '$base_ref'. Confirm the 'review-verdict/h10' required check manually before scheduling an auto-merge."
decide ask "H6/H10 merge gate: could not allocate a temp file to read branch protection for '$base_ref'. Confirm the 'review-verdict/h10' required check manually before scheduling an auto-merge."
decide ask "H6/H10 merge gate: the shared branch-protection rule classifier is missing or unreadable at $classifier, so which rule governs '$base_ref' — and therefore whether 'review-verdict/h10' is required on it — could not be derived (ersatztv#787). Restore the file, or confirm the required checks manually."
fi
bp_verdict=$(printf'%s'"$bp_list"| jq --arg b "$base_ref" -c -f "$classifier" 2>/dev/null ||true)
case$(printf'%s'"$bp_verdict"| jq -r '.verdict // ""' 2>/dev/null ||true) in
decide ask "H6/H10 merge gate: no branch-protection rule on this repo governs '$base_ref' decidably — a GLOB rule could govern it, or two rule names fold-equal, or a name is non-ASCII. This hook deliberately does not reimplement Gitea's glob matcher, so whether 'review-verdict/h10' is required on this base cannot be derived here (ersatztv#778). Confirm it in the repo's branch-protection settings, or merge immediately instead of scheduling.";;
undecidable)decide ask "H6/H10 merge gate: no branch-protection rule on this repo governs '$base_ref' decidably — a GLOB rule could govern it, or two rule names fold-equal, or a name is non-ASCII. This hook deliberately does not reimplement Gitea's glob matcher, so whether 'review-verdict/h10' is required on this base cannot be derived here (ersatztv#778). Confirm it in the repo's branch-protection settings, or merge immediately instead of scheduling.";;
none)bp_code=nomatch;bp="";;
# A DECLARED class of the classifier's contract (ersatztv#859), with its OWN sentinel — not
# merely its own arm. Giving it an arm that set `unreadable-rules`, the same value
# the catch-all sets, was measured to be a no-op: deleting that arm left the WHOLE suite
# green, because nothing downstream could tell the two apart. An arm no observation can
# distinguish is not a fix, it is a comment with syntax. (The invariant is "no test reddens",
# not a test count — a count goes stale the next time anyone adds one.)
#
# They are different findings and now say so. `unnamed-rule` means the list was READ and a rule
# in it carries no usable name; `unreadable-rules` means jq died or answered a word this hook
# does not know. Same decision (ask), different cause — and naming the cause accurately is the
# entire subject of this issue, so collapsing them here would have reproduced the defect being
# fixed, one arm over.
unreadable)bp_code=unnamed-rule;bp="";;
*)bp_code=unreadable-rules;bp="";;
esac
else
@@ -511,7 +612,6 @@ else
fi
bp=""
fi
rm -f "$bp_file"
# `nomatch` is the CLASSIFIER's verdict, deliberately not an HTTP code. Reusing 404 for it made
# this deny reachable from an HTTP 404 on the list read too — repo not found, or invisible to the
# credential, which Gitea also answers 404 — and then the reason claimed "the full rule list was
@@ -520,11 +620,26 @@ else
if["$bp_code"="nomatch"];then
decide deny "H6/H10 merge gate: BLOCKED — no branch-protection rule on this repo can govern '$base_ref' (the full rule list was read and none matches), so 'review-verdict/h10' is not a required check on it. A scheduled auto-merge is safe ONLY because that per-sha required check stops a commit pushed after scheduling from merging unreviewed (ersatztv#622). Restore branch protection on '$base_ref', or merge immediately (without merge_when_checks_succeed) once CI is green."
fi
# `unreadable-rules` is the CLASSIFIER failing on a 200 it could not parse — a numeric
# `branch_name` makes jq throw, and `//` does not catch it because it fires only on null/false.
# It gets its own sentinel for the same reason `nomatch` does: reporting "HTTP '000' — Gitea
# unreachable" about a successful 200 read states a cause that did not happen, which is the defect
# fixed one arm over for the deny.
# `unnamed-rule` is the classifier reporting a rule whose NAME it could not use. Two distinct
# shapes, and the reason string must cover both or it states a cause that did not happen: EITHER
# both fields supply no name (absent, null, or empty), OR one of them is present holding a
# non-string, which poisons the rule however good its sibling is. It is deliberately NOT reported as
# "no rule matches": a rule that cannot be read might be the rule Gitea is applying, so a list
# containing one supports no finding about which rule governs the base. That was the #859 defect —
# `""` is a valid name that matches nothing, so an unreadable rule DENIED with a stated cause that
# had not happened.
if["$bp_code"="unnamed-rule"];then
decide ask "H6/H10 merge gate: a branch-protection rule on this repo carries no name this hook can use — either both 'branch_name' and 'rule_name' are absent/null/empty, or one of them is present holding something that is not a string. Which rule governs '$base_ref', and whether 'review-verdict/h10' is required on it, therefore could not be derived. A rule that cannot be read might be the one Gitea applies, so this is deliberately NOT reported as 'no rule matches' (ersatztv#859). Inspect the branch-protection rules, or merge immediately instead of scheduling."
fi
# `unreadable-rules` is the CLASSIFIER failing on a 200 this hook could not turn into a verdict —
# jq died, or answered a word this contract does not define. It gets its own sentinel for the same
# reason `nomatch` does: reporting "HTTP '000' — Gitea unreachable" about a successful 200 read
# states a cause that did not happen, which is the defect fixed one arm over for the deny.
#
# A numeric `branch_name` was the worked example here until ersatztv#859 and no longer reaches this
# arm: it is not a usable NAME, so the classifier now classifies it rather than throwing on it, and
# it lands on `unnamed-rule` above with the cause that actually applies. The example is corrected
# rather than dropped, because it is the one shape a reader is likely to reach for when testing.
if["$bp_code"="unreadable-rules"];then
decide ask "H6/H10 merge gate: this repo's branch-protection rules came back in a shape this hook could not parse, so whether 'review-verdict/h10' is required on '$base_ref' is unknown. Check the rules manually, or merge immediately instead of scheduling."
fi
@@ -585,7 +700,19 @@ fi
# inside a fenced code block (documentation showing the convention counted as a real verdict), and a
# sha taken from the first `@<hex>` anywhere on the line (a markdown link could supply it). Every
# decision the classifier makes is documented there; this file only maps a class onto a hook decision.
decide ask "H10 merge gate: verdict classifier not found at $verdict_script, so the review state can't be derived. Confirm the review covered the latest commit before merging."
fi
@@ -614,6 +741,126 @@ case "$class" in
decide ask "H10 merge gate: unrecognized verdict classification '$class' for PR #$pr. Confirm the review covered the latest commit ($short) before merging.";;
esac
# --- (d) Guard-scope freshness (ersatztv#787): the committed mirror of `main`'s required status
# checks must still match the server. ------------------------------------------------------
# ORDERED LAST, and that is a severity argument rather than a stylistic one. Every check above
# can DENY; this one can only ever downgrade an otherwise-satisfied auto-grant to a prompt. Run
# earlier it would preempt those verdicts and report a stale guard scope at a reader whose merge
# is blocked for a completely different and more serious reason, and it would ask on payloads the
# checks above are about to reject anyway. Placed here it is also PAST the point where the two
# merge paths converge, so it covers both without duplicating anything.
# `scripts/tests/test_ci_dropped_step_guard.py` DERIVES which jobs must carry per-step execution
# markers from `.gitea/required-status-contexts.json`, because its CI job checks out with
# `persist-credentials: false` and cannot ask Gitea. That makes the snapshot the single
# hand-maintained input in the chain: a fourth required context added on the server leaves the
# snapshot — and therefore the guard's scope — silently behind, which is the whole of #787.
#
# THIS RUNS ON BOTH MERGE PATHS, deliberately, and it is placed here rather than beside the
# branch-protection read in the scheduled-auto-merge branch for that reason.
#
# WHAT IT DOES NOT COVER, said here rather than left to be discovered: a PR whose changed files are
# all docs/process — `.gitea/` included — exits at the docs-only passthrough far above, so this arm
# never runs for it. A PR that edits ONLY `.gitea/required-status-contexts.json` is docs-only BY
# CONSTRUCTION, and that is exactly the snapshot-NARROWING direction the decision record names as
# this design's residual. Excluding that path from the allow-list would not buy the protection it
# looks like it would: this arm compares the live server against the snapshot in the LOCAL CHECKOUT,
# not against the version the PR proposes, so it cannot see a narrowing that has not landed yet.
# What does hold is that the passthrough is a passthrough — a human prompt, never an auto-grant —
# which is the `.gitea/` treatment ersatztv#317 asked for. That read is inside
# `else` (mwcs = true) and never executes on an immediate merge, which is the common case; hanging
# the freshness check off it would fire it only when an auto-merge is armed. This file already
# records that exact defect one section up — the base re-read "first landed inside the
# scheduled-auto-merge branch only", with scheduled+retarget denied while
# immediate+retarget auto-GRANTED. Same shape, so it is not repeated here.
#
# It reads `main` (the branch the snapshot names), NOT `$base_ref`. That is a DIFFERENT question
# from the one the scheduled branch asks — "is review-verdict/h10 required on the base I am merging
# into" — so this is not a second copy of that classifier and the two cannot drift into disagreeing:
# they consume different fields of different rules for different decisions.
#
# ASK, NEVER DENY. Drift does not make THIS merge unsafe: Gitea enforces the live required set
# server-side, so a newly required context with no status blocks the merge on its own. What has gone
# stale is a guard's scope — a different artifact, on a different clock. Denying would state
# something false about the change in front of the reader. Every non-`match` class asks, so a
# comparison that could not be made is surfaced rather than skipped (`unknown` is not `fine`).
# ONE base for both the checker and the snapshot, and it is `$repo_root` — see
# `process.hook-resolves-inputs-from-repo-root` for why an env var may not select either
# (ersatztv#787, #858). The reason specific to THIS arm is that both halves of a comparison are
# resolved here: from two different roots the hook would classify one checkout's snapshot with
# another checkout's script — mismatched halves of a comparison whose entire job is to detect a
# mismatch — and answer `match` about a tree nobody asked about.
decide ask "H6 merge gate: $ctx_snapshot is missing, unreadable, or names no \`repo\`, so the dropped-step guard's scope could not be checked against branch protection — nor could it be established whether this snapshot even describes $owner/$repo (ersatztv#787). Restore the file, or check the required checks manually."
fi
# CASE-FOLDED, because Gitea resolves owner/repo case-insensitively: verified live, both
# `/repos/timothy/ersatztv` and `/repos/TIMOTHY/ErsatzTV` answer 200. A byte-exact compare would let
# any case variant sail through every other arm and SKIP this one, so drift would go unreported with
# no ask — the gate failing open on a spelling. The hook already treats case folding as
# decision-relevant one section up, where `MAIN` vs `main` makes the governing rule undecidable.
decide ask "H6 merge gate: the required-contexts checker is missing or not executable at $ctx_script, so whether the dropped-step guard's scope still matches branch protection on 'main' could not be derived (ersatztv#787). Check it manually, or restore the script."
fi
# THE SHARED READ (ersatztv#859). On a scheduled merge the arm above already fetched this; here that
# call is a cache hit, so the endpoint is read once per run instead of twice. On the IMMEDIATE path
# this is the only consumer and it performs the fetch itself, which is why the call sits AFTER the
# `[ ! -x "$ctx_script" ]` check above: a missing checker must ask without having touched the
# network, and a test pins exactly that by asserting no branch-protection URL was recorded.
fetch_branch_protections
if["$bp_cache_code"="mktemp-failed"];then
decide ask "H6 merge gate: could not allocate a temp file to read branch protection for the guard-scope freshness check (ersatztv#787)."
fi
ctx_code=$bp_cache_code
# ONE temp file, and it holds the checker's STDERR. Until ersatztv#859 this was `mktemp` for the
# payload plus an unmanaged `$bpf.err` beside it — a second path mktemp never created and therefore
# never made unpredictable. The payload now comes from the shared cache over a pipe, so the only
# thing still needing a file is the diagnostic, and it gets the mktemp'd one.
ctx_err=$(mktemp)|| decide ask "H6 merge gate: could not allocate a temp file for the guard-scope freshness check's diagnostics (ersatztv#787)."
if["$ctx_code"="200"];then
# stderr is KEPT, not sent to /dev/null. The checker exits 2 with a diagnostic on a usage error —
# an unreadable snapshot, a branch mismatch, a missing classifier — and discarding it made all of
# those arrive at the operator as the catch-all's "returned 'nothing'", which names no cause. That
# is the same states-a-cause-that-did-not-happen shape this arm was careful about elsewhere.
ctx_class=$(printf'%s'"$bp_cache"|"$ctx_script" --branch main --snapshot "$ctx_snapshot" 2>"$ctx_err"||true)
decide ask "H6 merge gate: the required status checks on 'main' no longer match .gitea/required-status-contexts.json (ersatztv#787). scripts/tests/test_ci_dropped_step_guard.py derives its marked-job scope from that snapshot, so until it is reconciled a required context may have NO dropped-step guard — a step the runner drops would conclude success and take that check green having done no work (ersatztv#756). Re-read the live list and update the snapshot in a PR (the guard will then demand markers for any newly required job, or an ACCOUNTED_ELSEWHERE entry naming what covers it). This does not make the merge in front of you unsafe — Gitea enforces the live required set server-side — so approve if you have judged it unrelated.";;
nomatch)
decide ask "H6 merge gate: no branch-protection rule governs 'main' at all, so the required status checks the dropped-step guard scopes itself to could not be confirmed (ersatztv#787). Branch protection on 'main' is what makes 'review-verdict/h10' load-bearing (ersatztv#743) — check it before merging.";;
undecidable)
decide ask "H6 merge gate: a glob branch-protection rule could govern 'main', so which rule's required contexts to compare against .gitea/required-status-contexts.json is not derivable without reimplementing Gitea's matcher (ersatztv#787). Confirm the required checks manually.";;
unreadable)
decide ask "H6 merge gate: branch protection for 'main', or .gitea/required-status-contexts.json itself, came back in a shape the required-contexts checker could not consume, so whether the dropped-step guard's scope is still current is unknown (ersatztv#787). Check the rules and the snapshot manually.";;
readfail)
decide ask "H6 merge gate: could not read branch protection for the guard-scope freshness check (HTTP '${ctx_code:-none}' — Gitea unreachable, or these credentials lack the repo-admin scope that endpoint needs), so whether .gitea/required-status-contexts.json is still current is unknown (ersatztv#787). Confirm the required checks on 'main' manually.";;
*)
decide ask "H6 merge gate: the required-contexts checker returned '${ctx_class:-nothing}', which is not a class this hook understands, so the dropped-step guard's scope could not be confirmed against branch protection (ersatztv#787).${ctx_diag:+ It said:${ctx_diag}}Check scripts/check-required-contexts.sh.";;
esac
fi# end of the guard-scope freshness arm (opened at `if [ "$ctx_repo_fold" = ... ]` above). The
# body is left unindented to match the rest of this file, which is flat throughout; the marker
# is here because the block is long enough that its extent is otherwise easy to misread.
if["$class"="positive"];then
# (a) CI + (b) all Done-when ticked + (c) positive verdict @ current head -> SATISFIED. Auto-grant.
# The reason string must not claim more than was actually checked: on the merge_when_checks_succeed
description:'Close one ersatztv issue or bundle in its own worktree via PR: claim, recon, implement, local gate, adversarial review before the push, fix loop, single push, PR, closing record',
constCOMMON=`Project: ersatztv, a fork of the ErsatzTV IPTV channel server (C#/.NET + a React SPA under web/). Shared checkout ${SHARED} is READ-ONLY for you: never commit there and never read its git log or HEAD as truth about main (process.shared-tree-readonly) — origin/main after a fetch is the only truth.
Issue(s) ${REF}: "${args.title}".
Issue body (condensed by a picker; read the real thing): ${args.body_summary}
DONE CONDITION: ${args.done_condition}
Read every issue in the bundle and all its comments yourself: curl -s -u "$ETV_GITEA_BASICAUTH" ${API}/issues/${ISSUE} and ${API}/issues/${ISSUE}/comments (the env var is set; never write the credential into a file or a commit).
Other slots of this session are working IN PARALLEL and will edit these files; do not touch them, and if your fix genuinely needs one of them, stop and report it instead of editing:
${JSON.stringify(args.avoid||[],null,1)}
Working rules, non-negotiable:
- Docs-first is a HARD RULE: read CLAUDE.md, then docs/README.md's task-signal map and ONLY the sections it points to for this task, then docs/contributing.md for the code you touch. Decisions resolve through docs/decisions/README.md by key, never by chasing a file path named in an old comment. Do not reverse-engineer conventions from source before reading these.
- Docs-update is part of done, same PR: an endpoint change updates docs/api-conventions.md's checklist and regenerates v1.json + endpoint-index.md via ./scripts/update-openapi.sh (build the app project first, then the script, then npm run generate:api under web/); a screen or route change updates docs/blazor-route-parity.md + docs/domain-model.md; a new or reversed convention gets a record under docs/decisions/records/<area>/ and a regenerated catalog (PYTHONPATH=. python3 scripts/build_decisions_catalog.py — the catalog docs/decisions/README.md is generated and shared with other slots: never hand-edit it, regenerate it, and resolve a rebase conflict in it by regenerating); a new or retitled doc updates docs/README.md.
- A TvContext model change needs a migration in BOTH providers: scripts/add-migration.sh <Name>.
- Tests are NUnit + Shouldly + NSubstitute in the existing *.Tests projects; vitest under web/. Pin the behaviour with a test that reddens when the fix alone is removed; never set ETV_UPDATE_GOLDENS or ETV_UPDATE_PLAYOUT_GOLDENS.
- Dependencies use Central Package Management: versions live only in Directory.Packages.props.
- Docs record the end state, never the investigation (docs.no-session-narrative): the path goes in the commit message and the issue comment. Date any measurement you write into a doc.
- Gitea labels take their own endpoint: POST ${API}/issues/{n}/labels {"labels":[100]} adds in-progress, DELETE ${API}/issues/{n}/labels/100 removes it; PATCH silently ignores labels.
- Never use bare git stash (the stash stack is shared across worktrees; commit WIP instead). Never push to main (it is refused server-side anyway). Never amend or force-push a pushed branch; a fix after the push is a new commit. Never cd out of your worktree except to read the shared checkout read-only.
- Kill only PIDs you started; never pkill by name — other sessions run dotnet and Playwright on this machine.`
constWORKTREE=`Worktree: ${WT} on branch ${BRANCH}. Check git -C ${SHARED} worktree list; if absent: git -C ${SHARED} fetch origin && git -C ${SHARED} worktree add ${WT} -b ${BRANCH} origin/main (absolute path, as written). Then give it its own web/node_modules: if cmp -s ${SHARED}/web/package-lock.json ${WT}/web/package-lock.json then cp -Rc ${SHARED}/web/node_modules ${WT}/web/node_modules, else (cd ${WT}/web && npm ci). Do ALL work inside ${WT}. If git commit is denied by the worktree-owner guard, the worktree belongs to ANOTHER session (orchestrated worktrees carry no marker): never overwrite the marker — STOP and report done=false with the guard's message. Commit as you go; every commit message ends with these trailer lines exactly:
${TRAILER}`
constCLAIM=`CLAIM FIRST, the four-way check from the kickoff (process.parallel-session-claim), for EVERY issue in the bundle: git -C ${SHARED} fetch origin; curl the open PRs (${API}/pulls?state=open&limit=50, page until empty) for a body saying fixes/refs ${REF}; ${issues.map(n=>`git -C ${SHARED} ls-remote --heads origin '*${n}*'`).join('; ')}; read each issue's comments for a claim that predates the label. If a PR, branch or comment shows another session already on ${REF} (other than this orchestrator's note, if any), STOP and report done=false with the evidence. Otherwise add the in-progress label and post a claiming comment naming branch ${BRANCH} and worktree ${WT}, on every issue in the bundle. If an issue body has no "## Done-when" section, append one (PATCH ${API}/issues/{n} with the full body): one unticked box per concrete completion criterion drawn from the issue, plus "- [ ] Adversarial review passed". The merge gate derives consent from those boxes; the orchestrator ticks them from your evidence, so write criteria that can be evidenced.`
constgateFor=(port,where)=>`LOCAL GATE (process.local-gate-before-push) — run it inside ${where} and read the real output; a skipped test is not a passing one:
- .NET: dotnet build the solution, then dotnet test on every test project that covers what you touched (ErsatzTV.Tests, ErsatzTV.Core.Tests, ErsatzTV.Scanner.Tests, ErsatzTV.FFmpeg.Tests, ErsatzTV.Architecture.Tests — all of them for anything under ErsatzTV.Core). Before any push touching .cs: BOM-check the touched set with od -A n -t x1 -N 3 <file> (efbbbf = BOM) and run bash -c 'dotnet format whitespace . --folder --verify-no-changes --include <files>' (process.bom-format-detection-recipe).
- SPA: cd web && npm run check:api && npm run lint && npm run typecheck && npm run build && npm test.
- scripts/, .claude/, .husky/, .gitea/: PYTHONPATH=. python3 -m pytest scripts/tests -q, plus ruff check and ruff format --check on any Python you touched. A new executable under scripts/ or .claude/hooks/ needs its row in docs/remote-state-inventory.md and, if it is a guard, in docs/guard-inventory.md — the suites say so.
- Docs: python3 scripts/check-doc-narrative.py --diff origin/main and answer what it flags (it is advisory, the rule is not).
- Live-E2E${args.needs_e2e?' IS REQUIRED for this change (write path or UI)':' only if you changed a write path or a screen'}: ETV_UI_PORT=${port} scripts/e2e-local.sh <fresh CONFIG_DIR> — port ${port} is yours; one run at a time in that worktree; curl the endpoints, never a browser tab; when done, kill the PID the launcher printed and nothing else. The launcher's pre-flight refuses a busy port and names the holder: report that, do not pick another port and never kill the holder.
- Builds on this Mac are capped at 3–4 concurrent and other slots are building too: run the .NET and web gates sequentially, not in parallel with each other.`
head_sha:{type:'string',description:'git rev-parse HEAD of YOUR WORKTREE after your last commit (not a PR head) — the finisher derives fix commits from these'},
patch_changed:{type:'boolean',description:'finisher only: true if the pre-push rebase changed the patch-id (a conflict resolved or an artifact regenerated)'},
facts:{type:'string',description:'what the docs the task-signal map names and the existing code say, with paths and decision keys'},
risks:{type:'string'},test_plan:{type:'string',description:'tests to add and the gate or E2E route that proves the done condition'},
},
}
letrecon=null
if(big){
phase('Recon')
recon=awaitagent(`${COMMON}
You are the recon agent. Read-only, in ${SHARED}. Read the docs the task-signal map names for this task, then find every fact an implementer needs to close ${REF} without re-deriving it: the exact handlers, components, signatures, call sites and guards, the existing tests, and which gate or E2E route proves the done condition. For a multi-site sweep use the csharp-lsp MCP tools, not the LSP tool (docs/local-lsp-tooling.md). Produce a concrete plan.`,
${recon?`Recon (verify what you rely on):\nPLAN: ${recon.plan}\nFACTS: ${recon.facts}\nRISKS: ${recon.risks}\nTEST PLAN: ${recon.test_plan}\n`:''}
You are the implementer. Close ${REF} completely: pin the behaviour with tests named for the branch they protect, update the docs the change obligates, commit. Then git fetch origin and rebase onto origin/main if it moved (never merge main in; regenerate generated artifacts), run the LOCAL GATE and STOP — do not push; reviewers read your worktree first, and a finisher pushes once after the review loop is clean. ${GATE}
Report done=true with the gate output when the worktree is ready for review, with pr_url empty and head_sha = git rev-parse HEAD of the worktree after your last commit.`,
${gateFor(e2ePort,'your own isolated worktree (never '+WT+')')}
Worktree ${WT}, branch ${BRANCH}, not yet pushed; diff: git -C ${WT} diff origin/main...HEAD. Read-only except scratch you create under /private/tmp; do not commit or push. NEVER run rm -rf, git worktree remove, git branch -D or any delete outside a directory you created under /private/tmp this session, and never build a path with .. segments. If you must build or run tests, do it in your own isolated worktree, never in ${WT}: git fetch ${WT}${BRANCH} && git checkout --detach FETCH_HEAD puts the unpushed branch there; run the .NET and web gates sequentially — other slots are building; E2E there on port ${e2ePort} (the GATE above is written for your worktree and that port).`
constLENSES=[
{key:'correctness',model:'opus',isolation:'worktree',prompt:'correctness against the done condition: run the gate and, for a write path or screen, the live-E2E route yourself, and read the output; try to break the change with the edge cases the issue and the docs name; check the pinning test actually reddens when the fix alone is reverted (mutate the clause, not the file).'},
{key:'conformance',model:'sonnet',prompt:'repo conformance: docs-update obligations met in this diff (endpoint → api-conventions + regenerated v1.json/endpoint-index; screen/route → blazor-route-parity + domain-model; convention → decision record + regenerated catalog; new doc → README index); no narrative in docs; every new script or hook has its inventory row; CPM respected; both-provider migration if the model changed; tests are NUnit/vitest in the existing projects; no BOM in touched .cs; no edit to a file another slot owns (listed above); commit trailers present; branch rebased on current origin/main; nothing pushed yet.'},
]
letxfamilyFailedRound=null
letxfamily=rubric?'codex':'not required (routine risk class under process.independent-review-rubric)'
You run the cross-family review — the diff touches a class where process.independent-review-rubric requires a reviewer from another model family, and you are only the runner. Write a prompt file under a directory you create in /private/tmp asking for an adversarial correctness and security review of the diff of branch ${BRANCH} against origin/main in ${WT} for issue(s) ${REF} with done condition "${args.done_condition}", listing findings as blocking / should-fix / nit with file and evidence, ending with a line VERDICT: merge or VERDICT: send-back. Run it EXACTLY like this, in the background, output to a file, stdin from /dev/null (it hangs otherwise): codex exec -C ${WT} -s read-only "$(cat <prompt>)" < /dev/null > <out> 2>&1 — then wait for the process to exit (poll pgrep on its PID with Monitor; measured 2026-07-28 in the #672 session, a real review took ~35 minutes for a 7-file diff) and read the file. Return its findings faithfully in the schema with ran=true; if the file has no VERDICT line the run failed (quota, tool error) — return ran=false, verdict merge, no findings, and put the file's tail in a single nit finding so the failure is visible; never invent a verdict.`,
xfamily=`codex could not run in round ${round} (${r?'no VERDICT line':'runner returned nothing'}); substituted a cold same-family review-only agent per process.independent-review-rubric — retry cross-family next window`
log(`${REF}: ${xfamily}`)
returnagent(`${reviewCommon(Number(args.port)+2)}
You are a COLD, review-only substitute for a cross-family reviewer that could not run. You have seen none of this branch before. Lens: adversarial correctness AND security of the diff against the done condition — the classes process.independent-review-rubric names (locks/concurrency, auth/security, API write paths, migrations, large C# diffs). Run the gate in your own worktree and read the output; report only what you verified, with evidence. blocking = done condition or a repo rule violated; should-fix = real defect; nit = style. Verdict send-back if any blocking.`,
Review round ${round} of the branch for ${REF}. Lens: ${l.prompt}
Be adversarial; report only what you verified, with evidence. blocking = done condition or a repo rule violated, or a test that passes for the wrong reason; should-fix = real defect; nit = style. Verdict send-back if any blocking.`,
You are the fixer. Reviewers found these problems in the unpushed branch; fix every blocking and should-fix one as new commits, or show with evidence why a finding is wrong:
if(blocking.length)return{issues,error:'blocking findings after two fix rounds; not pushed',blocking_remaining:blocking,history}
if(sendBack.length)return{issues,error:'should-fix findings still open after two fix rounds; not pushed — the orchestrator decides',should_fix_remaining:sendBack,history}
if(xfamilyFailedRound)return{issues,error:`the cross-family runner and its substitute both failed in round ${xfamilyFailedRound}; not pushed`,cross_family:xfamily,history}
constREVIEW_HISTORY=history.map(h=>`round ${h.round}: ${h.reviews.length} lens(es); ${countBy(h.reviews,'blocking')} blocking, ${countBy(h.reviews,'should-fix')} should-fix, ${countBy(h.reviews,'nit')} nit`+(h.fix?(h.fix_range?`; answered by the fix commit(s) in git log --oneline ${h.fix_range}`:'; answered without a new commit (findings refuted with evidence in the fixer report)'):'; clean — loop ended')).join('\n')
phase('Land')
constFINISH=`FINISH, in this order. Record the patch-id first: git diff $(git merge-base origin/main HEAD)..HEAD | git patch-id --stable. Then git fetch origin; if origin/main moved, rebase onto it (never merge main in; regenerate, never hand-resolve, generated artifacts — the decisions catalog by its generator), re-run the LOCAL GATE, and recompute the patch-id: report patch_changed=true if it differs. ${GATE}
Then ONE push: git push -u origin ${BRANCH}. Open the PR with the Gitea API (POST ${API}/pulls; head=${BRANCH}, base=main, title, body). The body must contain "fixes #N" for every issue in the bundle so the merge closes them, the root cause for a bug fix, the measured numbers, the review history VERBATIM as recorded by the workflow, one line per round, between the markers <<REVIEW HISTORY and REVIEW HISTORY>>:
<<REVIEW HISTORY
${REVIEW_HISTORY}
REVIEW HISTORY>>
${FIX_RANGES.length?`followed by what each fix commit changed, read from git show and not from memory, for exactly the commits git log --oneline lists in these ranges: ${FIX_RANGES.join('; ')}`:(history.some(h=>h.fix)?'and a sentence saying every finding was answered without a new commit, as the history block records':'and a sentence saying no fix commit exists because round one was clean')}, then the cross-family review status verbatim — "${xfamily}" — and every deliberately-left item with an issue number (file follow-up issues where needed). End the body with:
🤖 Generated with [Claude Code](https://claude.com/claude-code)
${SESSION_URL}
Arm the CI monitor: note the head sha and read ${API}/commits/<sha>/status once. Then the closing-an-issue skill (invoke it through the Skill tool if you have it, otherwise read .claude/skills/closing-an-issue/SKILL.md) with two modifications: do NOT close the issue — the merge closes it — and do NOT tick any "## Done-when" box; instead the "## Closing record" comment you post on each issue, linking the PR, ends with a "Done-when evidence" list giving, for every box, the command or artifact that evidences it — the orchestrator ticks from that. Remove nothing; the orchestrator removes the worktree after the merge. Report the PR URL, the head sha and patch_changed.`
constland=awaitagent(`${COMMON}
${WORKTREE}
You are the finisher. The branch has passed its review loop (${round} round(s)); nothing is pushed yet. ${FINISH}`,
if(!land.pr_url||!land.head_sha)return{issues,error:'finisher reported done without a PR URL or head sha — the branch may already be pushed; read its report before re-running',land,history}
log(`${REF} PR: ${land.pr_url||'none'} @ ${land.head_sha||'?'}${land.patch_changed?' (patch changed by the pre-push rebase)':''}`)
letpost_rebase_reviews=null
if(land.patch_changed){
log(`${REF}: patch changed on rebase — one more review round on the pushed head before any verdict`)
if(!post_rebase_reviews.length)return{issues,error:'the post-rebase review round produced no reviews (every lens failed); pushed, no verdict may be posted',pr_url:land.pr_url,head_sha:land.head_sha,cross_family:xfamily,history}
constlate=actionable(post_rebase_reviews)
if(late.length)return{issues,error:'blocking or should-fix findings on the pushed head after the pre-push rebase; no verdict may be posted',pr_url:land.pr_url,head_sha:land.head_sha,findings_remaining:late,cross_family:xfamily,history,post_rebase_reviews}
description:'Pick the next N ersatztv issues by the kickoff queue rules from scripts/select-queue.sh and live Gitea state, mutually non-colliding and avoiding what other slots hold, then adversarially verify the set',
phases:[{title:'Pick'},{title:'Refute'}],
}
// args: { taken: [{issues:[n], files:[...]}], closed: [n...], notes: 'free text', count: how many picks to return (default 3) }
consttaken=(args&&args.taken)||[]
constclosed=(args&&args.closed)||[]
constnotes=(args&&args.notes)||''
constcount=(args&&args.count)||3
constRULES=`Work read-only in /Users/timothy/ersatztv (the shared checkout; do not modify files, push, label or comment). Never read its git log or HEAD as truth about main: run git -C /Users/timothy/ersatztv fetch origin first, then read origin/main.
Read docs/handoffs/chicorytv-issue-queue.md fully — "Current phase", "Two concurrent tracks", the Selection and Bundles rules, and step 3's four-way claim check — and docs/handoffs/orchestration.md.
Ranking is NOT yours to derive: run ETV_GITEA_BASICAUTH="$ETV_GITEA_BASICAUTH" scripts/select-queue.sh 40 (the env var is already set) and take its order as given. It already excludes in-progress, parked, PRs, bot-authored issues and anything with an open blocker. Resolve only its CLAIM? and UMBRELLA? flags, by reading the flagged issue's body and comments.
Gitea REST: base http://192.168.1.95:3000/api/v1/repos/timothy/ersatztv, auth -u "$ETV_GITEA_BASICAUTH", curl only. Issue: GET /issues/{n}; comments: GET /issues/{n}/comments; open PRs: GET /pulls?state=open&limit=50 (page until a page comes back empty — the endpoint caps limit at 50). Remote branches naming an issue: git -C /Users/timothy/ersatztv ls-remote --heads origin '*<n>*'.
A pick is claimable only if the four-way check is clean: no open PR whose body says fixes/refs #n, no remote branch naming n, no claiming comment on the issue (a claim can precede the label), and the issue is still open after the fetch.
Bundles: after choosing an issue, scan its milestone, its cross-references and its labels for small independent siblings that are cheap to sweep in the same worktree; a bundle is one pick with several issue numbers. Never bundle issues that a taken slot already holds.
ALREADY TAKEN by this orchestrator (in flight, with the files each edits): ${JSON.stringify(taken)}
slug:{type:'string',description:'short kebab-case branch slug, e.g. null-font-family'},
rationale:{type:'string'},
body_summary:{type:'string',description:'body plus all comments, condensed but complete; include the Done-when section verbatim if the issue has one'},
risk:{type:'string',enum:['routine','rubric'],description:'rubric = touches locks/concurrency, auth/security, an API write-path handler, a DB migration, or will exceed ~150 changed C# lines (process.independent-review-rubric); needs a cross-family review'},
needs_e2e:{type:'boolean',description:'true for a write path or UI change (testing.live-e2e-prepush-timing)'},
skipped:{type:'string',description:'each higher-ranked issue skipped and the reason'}}}
constSCHEMA={type:'object',required:['picks'],properties:{picks:{type:'array',items:PICK,description:'in queue order; each later pick avoids the files of every earlier one'}}}
constVERDICT={type:'object',required:['refuted','reason'],properties:{refuted:{type:'boolean'},reason:{type:'string'},bad_picks:{type:'array',items:{type:'integer'},description:'issue numbers of the picks that fail, if not all'},better:{type:'array',items:{type:'integer'}}}}
phase('Pick')
constres=awaitagent(`${RULES}
Walk the selector's order and return up to ${count} issues or natural bundles, in that order, each of which (a) passes the four-way claim check, (b) edits no file a taken slot OR AN EARLIER PICK edits, (c) does not depend on another open issue (an earlier pick counts as open; a blocked-by dependency the selector already dropped), (d) is not a screen, handler or script an earlier pick is already on, (e) is not needs-hands or needs-the-user in disguise (a live-prod measurement nobody can take from here, a design question the body leaves open). Read each candidate's body and comments before accepting or rejecting it. Size is not a reason to skip: a large issue at the top of the queue is a pick, say size=large. Classify risk honestly — a write-path handler is rubric even when the diff is small. Stop early if the eligible queue runs out and say so in the last pick's skipped field; fewer than ${count} is fine, a colliding pair is not.`,{label:'picker',model:'sonnet',effort:'medium',schema:SCHEMA})
constpicks=(res&&res.picks)||[]
if(!picks.length)return{picks:[],refutations:[],note:'the picker returned no eligible pick',raw:res}
'ordering and claims: re-run scripts/select-queue.sh and the four-way claim check on every pick; refute if a higher-ranked eligible issue was skipped without a valid reason, the picks are out of selector order, or a pick is already claimed by a PR, branch or comment',
'collisions and classification: read the code each pick will touch; refute if any pick edits a file a taken slot or another pick edits, or the same docs section, or depends on an open issue; also refute a risk=routine pick that touches a lock, auth, an API write-path handler or a migration, and a needs_e2e=false pick that changes a write path or a screen',
].map((lens,i)=>()=>
agent(`${RULES}
Picks, in order:
${desc}
Lens: ${lens}. Try to refute; name the failing picks in bad_picks and a better ordering in better.`,{label:`refute:${i}`,model:'sonnet',effort:'medium',schema:VERDICT})))
description:'Resume a paused ersatztv branch: finish or fix, rebase onto origin/main, local gate, adversarial review, fix loop, push, PR body and closing record refreshed',
constCOMMON=`Project: ersatztv, a fork of the ErsatzTV IPTV channel server (C#/.NET + a React SPA under web/). Shared checkout ${SHARED} is READ-ONLY for you: never commit there and never read its git log or HEAD as truth about main (process.shared-tree-readonly) — origin/main after a fetch is the only truth.
Issue(s) ${REF}: "${args.title}".
YOUR BRIEF is the JSON file ${args.brief}: read it first with cat. It holds done_condition, context from the orchestrator, findings (the last review round) and recon where they apply.
Read every issue in the bundle and all its comments: curl -s -u "$ETV_GITEA_BASICAUTH" ${API}/issues/N and ${API}/issues/N/comments (the env var is set; never write the credential into a file or a commit).
Working rules, non-negotiable:
- Docs-first is a HARD RULE: read CLAUDE.md, then docs/README.md's task-signal map and ONLY the sections it points to for this task, then docs/contributing.md for the code you touch. Decisions resolve through docs/decisions/README.md by key.
- Docs-update is part of done, same PR: an endpoint change updates docs/api-conventions.md's checklist and regenerates v1.json + endpoint-index.md via ./scripts/update-openapi.sh (build the app project first, then the script, then npm run generate:api under web/); a screen or route change updates docs/blazor-route-parity.md + docs/domain-model.md; a convention gets a record under docs/decisions/records/<area>/ and a regenerated catalog (PYTHONPATH=. python3 scripts/build_decisions_catalog.py — the catalog is generated and shared with other slots: never hand-edit it, regenerate it, and resolve a rebase conflict in it by regenerating); a new doc updates docs/README.md. A TvContext change needs both providers' migrations via scripts/add-migration.sh.
- Tests are NUnit + Shouldly + NSubstitute; vitest under web/. Never set ETV_UPDATE_GOLDENS or ETV_UPDATE_PLAYOUT_GOLDENS. Dependencies only in Directory.Packages.props.
- Docs record the end state, never the investigation; the path goes in the commit message.
- Never use bare git stash. Never push to main. The ONLY sanctioned rewrite of a pushed branch is a rebase onto origin/main pushed with --force-with-lease (process.orchestrated-session); a fix is a new commit, never an amend. Never cd out of the worktree except to read the shared checkout read-only. Kill only PIDs you started.
Worktree: ${WT} on branch ${BRANCH}; it exists, do ALL work inside it. Give it its own web/node_modules if missing (cp -Rc from ${SHARED}/web when the lockfiles match, else npm ci). If git commit is denied by the worktree-owner guard, the worktree belongs to ANOTHER session (orchestrated worktrees carry no marker): never overwrite the marker — STOP and report done=false with the guard's message. Every commit message ends with these trailer lines exactly:
${TRAILER}`
constgateFor=(port,where)=>`LOCAL GATE (process.local-gate-before-push) — inside ${where}, real output, a skipped test is not a pass:
- .NET: dotnet build, then dotnet test on every test project covering what the branch touches (all of them for anything under ErsatzTV.Core); BOM-check touched .cs with od -A n -t x1 -N 3 and bash -c 'dotnet format whitespace . --folder --verify-no-changes --include <files>'.
- SPA: cd web && npm run check:api && npm run lint && npm run typecheck && npm run build && npm test.
- scripts/, .claude/, .husky/, .gitea/: PYTHONPATH=. python3 -m pytest scripts/tests -q, plus ruff on touched Python.
- Live-E2E${args.needs_e2e?' IS REQUIRED (write path or UI)':' only for a write path or screen change'}: ETV_UI_PORT=${port} scripts/e2e-local.sh <fresh CONFIG_DIR> — port ${port} is yours; one run at a time in that worktree; curl, never a browser tab; kill the PID the launcher printed when done and nothing else; a busy port is reported, never taken over.
- Run the .NET and web gates sequentially; other slots are building.`
constGATE=gateFor(args.port,WT)
constREBASE=`Rebase onto origin/main FIRST: git fetch origin; git rebase origin/main; resolve conflicts faithfully, keeping both sides' intent; regenerate generated artifacts rather than hand-resolving them. A commit titled "WIP: orchestrator checkpoint" holds uncommitted work from the paused session and must be folded into the commit it belongs to, never left in history — if it sits directly on that commit: git reset --soft HEAD~1 && git commit --amend --no-edit; otherwise: git commit --fixup=<target> is already its shape, so GIT_SEQUENCE_EDITOR=true git rebase --autosquash <target>~1 folds it non-interactively.`
verified:{type:'string',description:'exact gate commands run and their real output summary'},
left:{type:'string'},commits:{type:'string',description:'git log --oneline origin/main..HEAD'},pr_url:{type:'string'},head_sha:{type:'string',description:'git rev-parse HEAD of YOUR WORKTREE after your last commit (not a PR head) — the finisher derives fix commits from these'},
patch_changed:{type:'boolean',description:'finisher only: true if a second rebase before the push changed the patch-id'},
You are the implementer, continuing a paused session. Read git log and git show for the branch's commits first; a WIP checkpoint commit is the paused implementer's partial edit. ${REBASE} The brief's recon is a plan; verify what you rely on. Finish the done condition completely, with a regression test that reddens against the unfixed code. Run the LOCAL GATE and STOP without pushing; reviewers read the worktree first; report head_sha = git rev-parse HEAD of the worktree after your last commit. ${GATE}`,
You are the fixer, continuing a paused session. The PR is #${args.pr}. ${REBASE} Then the brief's findings are the last review round's: fix every blocking and should-fix one as new commits, or show with evidence why a finding is wrong. Run the LOCAL GATE and STOP without pushing; reviewers read the worktree first; report head_sha = git rev-parse HEAD of the worktree after your last commit. ${GATE}`,
${gateFor(e2ePort,'your own isolated worktree (never '+WT+')')}
Diff: git -C ${WT} diff origin/main...HEAD (rebased, not yet pushed). ${args.pr?`PR #${args.pr} exists: its pushed head, its body and any earlier closing record are INTENTIONALLY behind this worktree until the finisher pushes after this review loop and resyncs them — a stale PR head or body is not a finding, and neither is "not pushed".`:'No PR exists yet; the finisher opens it after this loop.'} Read-only except scratch you create under /private/tmp; do not commit or push. NEVER run rm -rf, git worktree remove, git branch -D or any delete outside a directory you created under /private/tmp this session, and never build a path with .. segments. If you must build or test, do it in your own isolated worktree, never in ${WT}: git fetch ${WT}${BRANCH} && git checkout --detach FETCH_HEAD puts the branch there; gates sequentially; E2E there on port ${e2ePort} (the GATE above is written for your worktree and that port).`
constLENSES=[
{key:'correctness',model:'opus',isolation:'worktree',prompt:'correctness against the done condition: run the gate and, for a write path or screen, the live-E2E route yourself, and read the output; try to break the change with the edge cases the issue and the docs name; check the pinning test reddens when the fix alone is reverted.'},
{key:'conformance',model:'sonnet',prompt:'repo conformance: docs-update obligations met; no narrative in docs; inventory rows for new scripts/hooks; CPM respected; both-provider migration if the model changed; no BOM in touched .cs; no WIP commit left in history; branch rebased on current origin/main; commit trailers present; PR body will carry fixes #N for each issue.'},
]
letxfamilyFailedRound=null
letxfamily=rubric?'codex':'not required (routine risk class under process.independent-review-rubric)'
You run the cross-family review required by process.independent-review-rubric; you are only the runner. Write a prompt file under a directory you create in /private/tmp asking for an adversarial correctness and security review of branch ${BRANCH} against origin/main in ${WT} for ${REF} with done condition from the brief, findings as blocking / should-fix / nit with file and evidence, ending with VERDICT: merge or VERDICT: send-back. Run EXACTLY: codex exec -C ${WT} -s read-only "$(cat <prompt>)" < /dev/null > <out> 2>&1 in the background, wait for the PID to exit (Monitor; measured 2026-07-28 in the #672 session, ~35 minutes for a 7-file diff), read the file, return its findings faithfully with ran=true; no VERDICT line means the run failed — return ran=false, verdict merge, no findings, and the file's tail in one nit finding; never invent a verdict.`,
xfamily=`codex could not run in round ${round} (${r?'no VERDICT line':'runner returned nothing'}); substituted a cold same-family review-only agent per process.independent-review-rubric — retry cross-family next window`
log(`${REF}: ${xfamily}`)
returnagent(`${reviewCommon(Number(args.port)+2)}
You are a COLD, review-only substitute for a cross-family reviewer that could not run. Lens: adversarial correctness AND security of the diff against the done condition in the brief. Run the gate in your own worktree; report only what you verified, with evidence. blocking = done condition or a repo rule violated; should-fix = real defect; nit = style. Verdict send-back if any blocking.`,
Review round ${round} of the branch for ${REF}. Lens: ${l.prompt}
Be adversarial; report only what you verified, with evidence. blocking = done condition or a repo rule violated, or a test that passes for the wrong reason; should-fix = real defect; nit = style. Verdict send-back if any blocking.`,
You are the fixer. Reviewers found these problems in the unpushed, rebased branch; fix every blocking and should-fix one as new commits, or show with evidence why a finding is wrong:
if(blocking.length)return{issues,error:'blocking findings after two fix rounds; not pushed',blocking_remaining:blocking,history}
if(sendBack.length)return{issues,error:'should-fix findings still open after two fix rounds; not pushed — the orchestrator decides',should_fix_remaining:sendBack,history}
if(xfamilyFailedRound)return{issues,error:`the cross-family runner and its substitute both failed in round ${xfamilyFailedRound}; not pushed`,cross_family:xfamily,history}
constREVIEW_HISTORY=history.map(h=>`round ${h.round}: ${h.reviews.length} lens(es); ${countBy(h.reviews,'blocking')} blocking, ${countBy(h.reviews,'should-fix')} should-fix, ${countBy(h.reviews,'nit')} nit`+(h.fix?(h.fix_range?`; answered by the fix commit(s) in git log --oneline ${h.fix_range}`:'; answered without a new commit (findings refuted with evidence in the fixer report)'):'; clean — loop ended')).join('\n')
phase('Land')
constFINISH=`FINISH: record the patch-id (git diff $(git merge-base origin/main HEAD)..HEAD | git patch-id --stable); git fetch origin; if origin/main moved again, rebase onto it, re-run the LOCAL GATE, and recompute the patch-id — report patch_changed=true if it differs. ${GATE}
Then git push --force-with-lease origin ${BRANCH}. ${args.pr?`Update PR #${args.pr}'s body (PATCH ${API}/pulls/${args.pr}) so it describes the branch as it now is`:`Open a PR (POST ${API}/pulls; head=${BRANCH}, base=main)`}: the body must contain "fixes #N" for every issue in the bundle, the root cause for a bug fix, the measured numbers, the review history VERBATIM as recorded by the workflow, one line per round, between the markers <<REVIEW HISTORY and REVIEW HISTORY>>:
<<REVIEW HISTORY
${REVIEW_HISTORY}
REVIEW HISTORY>>
${FIX_RANGES.length?`followed by what each fix commit changed, read from git show and not from memory, for exactly the commits git log --oneline lists in these ranges: ${FIX_RANGES.join('; ')}`:(history.some(h=>h.fix)?'and a sentence saying every finding was answered without a new commit, as the history block records':'and a sentence saying no fix commit exists because round one was clean')}, then the cross-family review status verbatim — "${xfamily}" — every deliberately-left item with an issue number, and end with:
🤖 Generated with [Claude Code](https://claude.com/claude-code)
${SESSION_URL}
Read ${API}/commits/<head-sha>/status once to arm the CI monitor. Post or update the "## Closing record" comment on each issue (the closing-an-issue skill's template) linking the PR, without closing the issue and without ticking any "## Done-when" box; end it with a "Done-when evidence" list naming, for every box, the command or artifact that evidences it — the orchestrator ticks from that. Report the PR URL, the head sha and patch_changed.`
constland=awaitagent(`${COMMON}
You are the finisher. The rebased branch has passed its review loop (${round} round(s)); nothing is pushed yet. ${FINISH}`,
if(!(land.pr_url||args.pr)||!land.head_sha)return{issues,error:'finisher reported done without a PR or head sha — the branch may already be pushed; read its report before re-running',land,history}
log(`${REF} PR: ${land.pr_url||args.pr} @ ${land.head_sha||'?'}${land.patch_changed?' (patch changed by the pre-push rebase)':''}`)
letpost_rebase_reviews=null
if(land.patch_changed){
log(`${REF}: patch changed on rebase — one more review round on the pushed head before any verdict`)
if(!post_rebase_reviews.length)return{issues,error:'the post-rebase review round produced no reviews (every lens failed); pushed, no verdict may be posted',pr_url:land.pr_url||args.pr,head_sha:land.head_sha,cross_family:xfamily,history}
constlate=actionable(post_rebase_reviews)
if(late.length)return{issues,error:'blocking or should-fix findings on the pushed head after the pre-push rebase; no verdict may be posted',pr_url:land.pr_url||args.pr,head_sha:land.head_sha,findings_remaining:late,cross_family:xfamily,history,post_rebase_reviews}
"source":"GET /repos/timothy/ersatztv/branch_protections -> the rule governing `main` -> status_check_contexts",
"why":"ersatztv#787. The committed mirror of the required status checks on `main`. It exists because the guards that make a required context trustworthy run in `pr-checks.yml::script-tests`, which checks out with persist-credentials:false and holds no Gitea credential, so it cannot ask the server. scripts/tests/test_ci_dropped_step_guard.py DERIVES its marked-job scope from `contexts` rather than repeating it as a literal, and scripts/check-required-contexts.sh compares this list against the live one wherever a credential does exist. Editing `contexts` by hand without re-reading the server is the one move that defeats both. The `repo` field exists because the merge-consent hook fires for whatever owner/repo the merge tool was called with: without it, merging a PR in another repo from an ersatztv session compares that repo's live contexts against THIS repo's mirror and reports a confident, flatly false finding about it.",
"contexts":[
"Build ErsatzTV Image / Build & test (.NET) (pull_request)",
echo "::error::git describe --tags --abbrev=0 failed, so this image would ship InformationalVersion 0.0.0-${SHORT} instead of a real version (ersatztv#836). The usual cause is a --depth fetch grafting this complete clone shallow; the two lines above say which."
echo "::error::git fetch of origin/${base_ref} failed, so this job cannot compute the changed-file set it derives its work from. That is a broken job, not an empty change set (ersatztv#746). Check the base branch still exists and that the runner can reach the repository."
exit 1
fi
if ! changed="$(git diff --name-only "origin/${base_ref}...HEAD")"; then
echo "::error::git diff against origin/${base_ref} failed, so the changed-file set could not be computed — do not read this as 'nothing changed' (ersatztv#746). If it reports no merge base, rebase this branch onto ${base_ref}."
exit 1
fi
echo "Changed files in this PR:"; printf '%s\n' "$changed"
if printf '%s\n' "$changed" | grep -Eq '^ErsatzTV/Controllers/Api/|^ErsatzTV\.Core/Api/'; then
echo "::error::git fetch of origin/${base_ref} failed, so this job cannot compute the changed-file set it derives its work from. That is a broken job, not an empty change set (ersatztv#746). Check the base branch still exists and that the runner can reach the repository."
exit 1
fi
if ! changed="$(git diff --name-only --diff-filter=ACM "origin/${base_ref}...HEAD" -- '*.cs')"; then
echo "::error::git diff against origin/${base_ref} failed, so the changed-file set could not be computed — do not read this as 'nothing changed' (ersatztv#746). If it reports no merge base, rebase this branch onto ${base_ref}."
exit 1
fi
echo "Changed .cs files in this PR:"; printf '%s\n' "$changed"
echo "Pins found in docker-build.yml: ${pins[*]} (${#pins[@]} distinct)"
@@ -107,10 +151,10 @@ jobs:
# in-repo remedy in that state: relax this length check in the same PR and say why. Note
# that ci-image.yml still tags with a plain `--short` (auto-scaled), so "always 7" is an
# empirical property of today's shallow clone, not an enforced invariant. Making the
# publisher emit `--short=7` is tracked as ersatztv#597. It is not blocked, just out of
# scope here: editing ci-image.yml re-points `expected` (above) at that commit, so it needs
# the branch's own publish-then-pin two-step (docs/ci-cd.md -> 'CI toolchain image') —
# ci-image.yml's push trigger has no branches: filter, so a feature branch does publish.
# publisher emit `--short=7` is tracked as ersatztv#597. That is no longer blocked by this
# job at all: since ersatztv#744, editing ci-image.yml does NOT re-point `expected`, so a
# `--short=7` change lands like any other PR. It does need a deliberate republish to take
# effect — see the note on `expected` above.
if [ "${#pins[0]}" -ne 7 ]; then
echo "::error::CI toolchain image pin ersatztv-ci:${pins[0]} is ${#pins[0]} chars, but ci-image.yml publishes 7-char tags (it tags with 'git rev-parse --short HEAD' from a fetch-depth:1 clone). A differently-sized abbreviation still resolves to the right commit, so this would pass every other check here — but NO such tag exists in the registry, and all five container: jobs would fail at image-pull time with 'manifest unknown'. Pin exactly: ersatztv-ci:${expected:0:7} (locally: git rev-parse --short=7 HEAD). See docs/ci-cd.md -> 'CI toolchain image'."
exit 1
@@ -121,7 +165,7 @@ jobs:
exit 1
fi
if [ "$pin_full" != "$expected" ]; then
echo "::error::CI toolchain image pin is stale: docker-build.yml pins ersatztv-ci:${pins[0]} ($pin_full), but docker/ci was last changed in $expected. Your jobs are testing an image that is NOT built from this PR's docker/ci. Let ci-image.yml publish the new :<sha>, then update the pin in ALL jobs to it (docs/ci-cd.md -> 'CI toolchain image')."
echo "::error::CI toolchain image pin is stale: docker-build.yml pins ersatztv-ci:${pins[0]} ($pin_full), but docker/ci was last changed in $expected. Your jobs are testing an image that is NOT built from this PR's docker/ci. Publish the new :<sha> — push this commit as branch HEAD and dispatch ci-image.yml on the branch (a branch PUSH no longer publishes, ersatztv#744) — then update the pin in ALL jobs to it (docs/ci-cd.md -> 'CI toolchain image')."
exit 1
fi
echo "Pin is current: ersatztv-ci:${pins[0]} resolves to $pin_full = docker/ci's last change."
@@ -134,16 +178,32 @@ jobs:
name:Docs update reminder
runs-on:small # seconds-long git diff; keep it off the build runners
if:github.event_name == 'pull_request'
env:
CI_JOB_ROLE:report-only
steps:
- name:Checkout
uses:actions/checkout@v4
with:
persist-credentials:false
fetch-depth:0
# `continue-on-error` for the same reason the two steps below carry it: this whole job
# is a non-blocking nudge, and an advisory red still joins the combined status the merge gate
# reads. Unmasking the fetch (ersatztv#746) makes a broken base LOUD in the log; it must not
# also make a warn-only job merge-blocking. The three jobs that genuinely gate on this diff —
# api-docs, format, decisions lifecycle — do redden on a failed fetch, which is where that
# belongs.
- name:Warn when a screen/route change skips the parity doc
echo "::error::git fetch of origin/${base_ref} failed, so this job cannot compute the changed-file set it derives its work from. That is a broken job, not an empty change set (ersatztv#746). Check the base branch still exists and that the runner can reach the repository."
exit 1
fi
if ! changed="$(git diff --name-only "origin/${base_ref}...HEAD")"; then
echo "::error::git diff against origin/${base_ref} failed, so the changed-file set could not be computed — do not read this as 'nothing changed' (ersatztv#746). If it reports no merge base, rebase this branch onto ${base_ref}."
exit 1
fi
echo "Changed files in this PR:"; printf '%s\n' "$changed"
screen_or_route=no
if printf '%s\n' "$changed" | grep -Eq '^web/src/screens/.+\.tsx$|^ErsatzTV/LegacyUiRedirects\.cs$'; then
@@ -169,7 +229,8 @@ jobs:
# python-using job on it declares this. Without it a missing interpreter is exit 127 — a RED
# advisory job joining the combined status, which is the one thing this step must never be.
#
# Both steps carry `continue-on-error` because the SCRIPT exiting 0 is not the whole invariant:
# Both steps OF THIS CHECK (setup-python + the narrative step; the parity nudge above has its
# own) carry `continue-on-error` because the SCRIPT exiting 0 is not the whole invariant:
# a setup-python download failure reddens the job just as effectively as a hit would, and an
# advisory red still joins the combined status the merge gate reads (ersatztv#598). Scope,
# stated rather than implied: this covers the two steps that exist to run the check. A failed
echo "::error::git fetch of origin/${base_ref} failed, so this job cannot compute the changed-file set it derives its work from. That is a broken job, not an empty change set (ersatztv#746). Check the base branch still exists and that the runner can reach the repository."
echo "::error::git fetch of origin/${base_ref} failed, so this job cannot compute the changed-file set it derives its work from. That is a broken job, not an empty change set (ersatztv#746). Check the base branch still exists and that the runner can reach the repository."
exit 1
fi
PYTHONPATH=. python3 scripts/decisions_validate.py --base "origin/${base_ref}" --head HEAD
@@ -83,7 +83,7 @@ main in) and re-run the local gate whenever the fetch shows movement.
Every task that closes a Gitea issue MUST complete ALL of these before it is considered done. Use `/done <issue>` to run through this automatically.
**Merge-consent is derived from state, not asserted (`## Done-when` convention — ersatztv#303 H6 + H10).** Any issue whose PR will merge to `main` should carry a `## Done-when` section in its **issue body** — a checklist of completion criteria (always include an "adversarial review passed" box; add per-issue criteria like tests-green, docs-updated, live-E2E). Two hooks derive merge-consent from it so a premature merge is blocked *by construction*, not by memory:
-`pretooluse-merge-consent.sh` (Claude PreToolUse on the Gitea merge tool) — **auto-grants** a merge (emits `permissionDecision: allow`, so **no** redundant mechanical prompt fires) only when the PR's CI is green **and** every `## Done-when` box on the linked issue (`fixes #N`) is ticked **and** a `Review-verdict:` comment references the PR's *current head sha* (**H10**); **denies** on an unticked box, red CI, or a stale/negative review verdict; **asks** (falls back to a human prompt) when it can't derive state (no linked issue, no `## Done-when` section, no `Review-verdict:` comment yet, no creds, Gitea down). On the auto-grant (satisfied) path the derived state **is** the consent — do not also ask conversationally to merge; a separate human confirmation is warranted only when the gate **asks** (ersatztv#314). **The H10 review-verdict convention**: after an adversarial/Codex review of a PR (or its latest fix commit), run **`scripts/post-review-verdict.sh <pr> <MERGEABLE|APPROVED|BLOCKED|NOT-MERGEABLE> [note]`** — it posts both the `Review-verdict: … @ <head-sha>` comment and the sha-bound `review-verdict/h10` commit status, proving the *latest* commit was reviewed rather than a stale earlier diff (ersatztv#242). Do not hand-write the comment: the **status** is the required check branch protection enforces, and a comment alone leaves it absent.
-`pretooluse-merge-consent.sh` (Claude PreToolUse on the Gitea merge tool) — **auto-grants** a merge (emits `permissionDecision: allow`, so **no** redundant mechanical prompt fires) only when the PR's CI is green **and** every `## Done-when` box on the linked issue (`fixes #N`) is ticked **and** a `Review-verdict:` comment references the PR's *current head sha* (**H10**); **denies** on an unticked box, red CI, or a stale/negative review verdict; **asks** (falls back to a human prompt) when it can't derive state (no linked issue, no `## Done-when` section, no `Review-verdict:` comment yet, no creds, Gitea down). On the auto-grant (satisfied) path the derived state **is** the consent — do not also ask conversationally to merge; a separate human confirmation is warranted only when the gate **asks** (ersatztv#314). **The H10 review-verdict convention**: after an adversarial/Codex review of a PR (or its latest fix commit), run **`scripts/post-review-verdict.sh <pr> <MERGEABLE|APPROVED|LGTM|BLOCKED|NOT-MERGEABLE> [note]`** — it posts both the `Review-verdict: … @ <head-sha>` comment and the sha-bound `review-verdict/h10` commit status, proving the *latest* commit was reviewed rather than a stale earlier diff (ersatztv#242). Do not hand-write the comment: the **status** is the required check branch protection enforces, and a comment alone leaves it absent.**The credential you post with must be an account on `H10_REVIEWERS` in `.gitea/workflows/review-verdict.yml`** (`timothy` today) — since ersatztv#742 the gate inherits an existing `success` only from an allow-listed creator (an existing `failure` is left alone on a weaker attributability test, so an attributable rejection VISIBLE AT THE FIRST READ is not re-derived into a green — a rejection landing later, inside a run's own write window, was a separate route and is NARROWED since ersatztv#849 — every path that cannot establish what the head carries now replaces that unknown state with a sticky sentinel instead of leaving it standing; see `ci.verdict-unverified-write-sentinel` for the residuals it names), and since ersatztv#845 the script ENFORCES that coupling rather than assuming it: it reads its own status back and refuses, before writing the verdict comment, unless the recorded `.creator.login` is on that allow-list — so a POSITIVE verdict posted with any other account fails loudly at your terminal instead of being reported as success. The gate still re-derives such a status on the next PR event — that part is unchanged; what the check removes is the tool telling you it worked. **The membership requirement is `success`-only**, mirroring the gate: a `BLOCKED` verdict is honoured from ANY attributable account, so an off-list reviewer can still record a rejection. **The status is still written** — the check runs after the POST, because it measures the creator Gitea recorded rather than what the credential claims — and what is withheld is the verdict COMMENT, which leaves the merge hook at condition (c) with nothing to classify, i.e. an `ask`. So a refused positive verdict leaves a green `review-verdict/h10` standing on that head that the gate itself will not inherit; branch protection binds the context NAME and not its issuer, so do not read that green as consent. The allow-list is derived from the workflow by `scripts/lib/h10-reviewers.sh`; it is never restated.
- **The gate is enforced server-side, per sha (ersatztv#622).** `review-verdict/h10` is a required status check on `main`. Because a commit status belongs to one sha, a commit pushed *after* an auto-merge is scheduled clears it and blocks the merge — closing the hole where `merge_when_checks_succeed` froze consent at scheduling time and Gitea later merged an unreviewed head. Renovate-authored and docs-only PRs are auto-passed by `.gitea/workflows/review-verdict.yml`, **except** when they touch `.claude/`, `.codex/`, `.gitea/`, `.husky/`, `scripts/` or `docker/ci/`. See `docs/ci-cd.md` → Review-verdict gate.
-`.husky/pre-push` → `prepush-donewhen.sh` — a fail-open backstop that blocks a direct `git push origin main` whose commits `fix #N` an issue with unticked boxes. **Since ersatztv#743 that push can no longer happen at all** (see below), so this hook is now belt-and-braces for a path the server refuses.
@@ -97,17 +97,30 @@ when finishing a task that closes an issue.
## Project Boundaries
**ersatztv OWNS**: ErsatzTV fork code (C#/.NET), channel/collection/schedule management, M3U/XMLTV generation, and the **`ersatztv` skill** — whose canonical copy is `.claude/skills/ersatztv/SKILL.md`**here**; `~/server-management/.claude/skills/ersatztv` is a symlink to it (ersatztv#617). Edit it in this repo; never fork a second copy.
**ersatztv OWNS** — *developing the fork*: the ErsatzTV fork code (C#/.NET), the `/api/v1` REST
surface, M3U/XMLTV generation, the `ErsatzTV.Mcp` server, CI and releases, and the **`ersatztv`
skill** — whose canonical copy is `.claude/skills/ersatztv/SKILL.md`**here**. Both
`~/server-management/.claude/skills/ersatztv` and `~/media-management/.claude/skills/ersatztv` are
symlinks to it (ersatztv#617, #755). Edit it in this repo; never fork a second copy.
**The split that is easy to get wrong** (ersatztv#755, `process.ersatztv-owns-code-not-operations`):
channel/collection/schedule *code* is owned here; **channel OPERATIONS against the running instance
are not**. Creating and editing channels, lineups, collections, schedules, playouts, logos and
overlays on the live ErsatzTV belong to `media-management`. Driving prod from here is in scope only
as *verification of a change this repo is shipping* (live-E2E, a release smoke test) — not as
day-to-day channel work.
**ersatztv does NOT own**:
- Channel/collection/schedule/playout **operations** against a live instance → media-management
- Jellyfin skill → server-management. `.claude/skills/jellyfin` here is a **relative symlink** to `~/server-management/.claude/skills/jellyfin` (ersatztv#617 — it had silently become a stale divergent copy). It therefore resolves only in a checkout at `~/ersatztv`, not inside a git worktree; that is inherent to the cross-repo symlink pattern server-management already uses (`beets`, `radarr`, `sonarr`, …).
**For infrastructure changes** (Docker, NFS, ports, Authelia): open an issue in `timothy/server-management`.
**For content/media sourcing questions** (what goes into channels, yt-dlp pipelines): open an issue in `timothy/media-management` once it exists; for now, `timothy/server-management`.
**For content/media sourcing questions and channel operations** (what goes into channels, yt-dlp
pipelines, editing a live channel): open an issue in `timothy/media-management`.
**For plan/audit reviews**: open `~/adversarial-reviewer` before significant architecture changes.
"description":"Committed stand-in for a user-authored scripted schedule: the ordered sequence of calls such a script makes against /api/v1/scripted/playout/build/{buildId}, replayed through the real ScriptedScheduleController by ScriptedScheduleControllerTests. Field names and casing match what the HTTP body binder accepts \u2014 the replay deserializes each body with ApiJsonSettings, the configuration Startup applies to AddNewtonsoftJson. Deliberately avoids wait_until / pad_until / pad_to_next (local day and time-of-day) and shuffle order, which would make the pinned snapshot machine-timezone- or seed-dependent. add_duration and pad_until_exact both end BETWEEN two content boundaries on purpose: an instruction that happens to end on one never reaches the engine's trim branch, and its trim flag is then witnessed by nothing.",
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.