Two independent cold reviews (Codex/GPT-5.6 cross-family, and a cold Claude reviewer in
its own worktree) converged on the same class: paths where "this job cannot establish
what is on the head" still resolved by leaving the head alone, which protects a real
verdict and leaves a forged one.
Behaviour:
1. The four page-2 completeness refusals now replace the unknown state too. They were
excluded on the reasoning that the probe fires when NO row for this context was on page
1, so there is no green of any provenance to leave standing — self-contradictory, since
the only reason page 2 is read is that the row may be beyond page 1, which the probe's
own message says. Accepted cost, stated in the record: a head with more CONTEXTS than
the 50-row cap stalls every run; measured 2026-08-29, this repo puts 8 on a `main` head,
and that case already stalled with an ABSENT check.
2. The no-mark downgrade covers every re-derivable write, not only `success`. Restricting
it analysed the wrong PR: the damaging case is one that IS exemptible and got the
generic `pending` only from a transient enumeration failure. That description carries no
marker, nothing verifies it without a mark, and the next run re-derives it into the
exemption with the human row below its own mark — route 2's damage through route 1's
condition. `$REPAIR_DESC` stays exempt, being stronger and not re-derivable.
3. The fence branch that cannot trust its retarget count while holding a derived `success`
writes the sentinel instead of abstaining. It is reached only after the classification
DECLINED to inherit the row the head carries, so posting nothing left that row current;
the message said the context "stays absent", true only of a head that had none.
4. Reconciliation needs a WITNESS: it may clear only over a complete history containing the
sentinel's own row. `ex_unverified` means the combined endpoint just returned that row
and `/statuses/{sha}` keeps one per POST, so a complete-but-empty history contradicts a
write that demonstrably happened — and `page_statuses` accepts an empty page 1 as
complete, which is what made it reachable. Both reviewers reproduced the clear-then-exempt
outcome. The shipped positive test used exactly that impossible fixture, so it was
pinning the defect; it now seeds the sentinel row, and an impossible-empty negative plus
a witness mutation proof were added.
5. The mid-run "did this row change" comparison now includes the row ID. The two sentinels
are byte-identical by design, so a mid-run replacement of one by another was invisible to
a state/creator/description triple. Measured 2026-08-29 (Gitea 1.27.1, head 736649b3):
the COMBINED endpoint carries `id` on every row, ids 14..30 ascending — the job had only
ever read ids from `/statuses/{sha}`. Where a server omits it both sides are empty and
the comparison degrades to the pre-existing text test.
6. The repair has a FLOOR — it may never write a description weaker than the one this run
decided — and is skipped when it would rewrite what is already there. Widening the gate
to every write meant a transient post-write read could rewrite a correct `$REPAIR_DESC`
carry-forward with the machine-clearable sentinel, reversing the ordering rule the
classification chain states.
Writing the sentinel and failing the job are separate decisions, which is why
`replace_unknown_state` and `replace_unknown_and_die` are two functions: the read refusals
were already non-zero exits on `main` and stay red; the fence branch exited 0 there and
still does, because an unreadable timeline is an ordinary hiccup and reddening every one is
noise this file elsewhere refuses to add.
Prose corrected where it now overclaimed: "the green never stands" after the post-POST
re-check is wrong — it is live between the POST and the repair, so the check makes a
permanent green TRANSIENT; "a later run reconciles this automatically" is wrong in the one
case where the replacement costs anything, since finding a masked verdict UPGRADES to the
human-only sentinel; and the mutation-proof framing claimed every mutant restores the exact
predecessor, when two do, one restores the shape #742 withdrew, and the rest disarm clauses
that have no predecessor. The quiet-timeline positive control now counts timeline walks,
because a single POST is also what a skipped re-check produces.
refs #849
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019T79beF1Ufid3dXju4yqkF