Files
ersatztv/docs/decisions/records/release/review-verdict-gate.md
T
timothyandtimothy 469d19852c
Build ErsatzTV Image / CI toolchain image resolves (push) Successful in 6s
Build ErsatzTV Image / Delimiter ban (release path) (push) Successful in 25s
Build ErsatzTV Image / Build & test (.NET) (push) Successful in 8m34s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (push) Successful in 6m17s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (push) Successful in 5m50s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (push) Skipped
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (push) Skipped
Build ErsatzTV Image / Build & push image (amd64) (push) Successful in 4m22s
fix(788): one declarative H10 verdict vocabulary, derived by both sides (#846)
The verdict words lived in two hand-written shell copies — the `case` arms of
post-review-verdict.sh (write) and the POS_RE/NEG_RE regexes of
check-review-verdict.sh (read) — held together by nothing but a comment that had
already gone stale. scripts/lib/review-verdict-vocabulary.sh now declares them
once and both sides derive; neither script enumerates a verdict word any more.

Only the WORD SET moved. The grammar stays in check-review-verdict.sh, where
every #629 false-open actually lived.

No parity test: #774 shipped one and withdrew it after six rounds, because a
regex over shell source is not a shell parser. The proof is behavioural and
graded MUTATION — the harness restores the pre-#788 hardcoded POS_RE each run and
requires it to redden.

Enforcement is a DATA dependency, not a control-flow gate. Review round 1 found a
real fail-open in the first commit: `${#arr[@]}` is nounset-safe only for a
declared-empty array, and under `set -u` that error inside a function called as
`if ! validate` skips BOTH branches — so on the reader (deliberately no `set -e`)
an explicit BLOCKED @ head classified `positive`, exit 0. Validation now sets a
sentinel on its last line and the derived views refuse without it.

Six cold review rounds; rounds 2-6 found no fail-open across differential fuzzing
(4788 / 2612 / 7560 payloads, zero divergences from origin/main's grammar),
sentinel forgery, environment poisoning, declare -p evasion on bash 5.3 and 3.2,
path/symlink resolution and probe TOCTOU. Every malformation fails closed: reader
exit 2, writer exit 1 with nothing posted.

Also corrected: CLAUDE.md and release.review-verdict-gate both enumerated the
vocabulary without LGTM, a word the code has accepted since #629.

fixes #788

Co-authored-by: Timothy <timothy@noreply.gitea.tblindustries.be>
2026-08-26 20:58:09 +00:00

11 KiB
Raw Blame History

key, title, status, since, supersedes, superseded-by, rule, signals, mechanics
key title status since supersedes superseded-by rule signals mechanics
release.review-verdict-gate 2026-07-12 — Review-verdict merge-gate: latest commit must be reviewed (#303 H10) active 2026-07-12 none none A PR may not merge until a `Review-verdict: <MERGEABLE|APPROVED|LGTM|BLOCKED|NOT-MERGEABLE> @ <head-sha>` comment references the PR's current head sha (short-sha prefix match against the verdict's OWN `@ <sha>` field, marker at COLUMN 0 (no indent, so indented code blocks cannot self-approve), whole-word verdict token, fenced code blocks stripped with markdown fence-length semantics, negative wins over positive on the same head); folds into the H6 merge-consent hook as condition (c). The grammar lives in ONE tested place, `scripts/check-review-verdict.sh` — #629 found three false-opens that survived because it was implemented inline and untested while this record described stricter behaviour than the code had. review-verdict, head-sha match, stale-review prevention, verdict false-open, MERGEABLE-LATER, fenced code block verdict, sha from a URL, unknown verdict token · paths: `scripts/check-review-verdict.sh`, `.claude/hooks/pretooluse-merge-consent.sh`, `.claude/settings.json` · issues: #303 (H10), #242, #629 `scripts/check-review-verdict.sh` (the grammar, + `scripts/tests/test_check_review_verdict.py`); `pretooluse-merge-consent.sh` (maps a class onto allow/deny/ask); `scripts/post-review-verdict.sh`; CLAUDE.md → Task Completion Protocol (H10 convention)

A PR may not merge until a Review-verdict: comment on it references the PR's CURRENT head sha — so the latest commit is proven-reviewed, not a stale earlier diff. This mechanizes the ersatztv#242 lesson ("re-review the fix commit, not just the initial PR diff": a review of an earlier revision does not license merging a head that carries un-reviewed follow-up commits). It folds into the existing H6 pretooluse-merge-consent.sh as condition (c), reusing its PR fetch, docs-only exemption, and Gitea-auth-from-env (no second hook → no detection drift, per the #303 methodology review).

Convention: after reviewing a PR (or its latest fix commit), post a PR comment (issue-style, not a Gitea formal-review body — the gate reads issues/{pr}/comments) whose line starts with the marker: Review-verdict: <MERGEABLE|APPROVED|LGTM|BLOCKED|NOT-MERGEABLE> @ <head-sha> (short ≥7-char or full sha). The gate counts a line as a verdict only when the marker is at line-start (after optional indent) — a comment that merely quotes the template mid-sentence (an instruction "please post: Review-verdict: MERGEABLE @ …", or the gate's own suggestion text echoed back) does not self-approve the merge (adversarial re-review false-open, folded pre-merge). It then classifies each verdict line by the sha in its @ <sha> field, matched to the head by git short-sha prefix semantics (head begins with the token, token ≥7 chars) — NOT a loose substring test, so an older sha that merely contains the head prefix does not count.

#629 — three of the protections described here were asserted but not implemented. The grammar lived inline in the hook with no tests, and each of these graded as a positive verdict until #629 (every one reproduced, then fixed, then mutation-verified):

  • the verdict token was prefix-matched, so MERGEABLE-LATER, APPROVED-PENDING-QA and LGTMish all read as positive. A token is now matched as a whole word, and one in neither vocabulary is classified unknown — never positive, and never guessed into a block either;
  • a verdict inside a fenced code block counted, because the line-start anchor is satisfied inside a fence. So documentation showing the convention was itself a verdict. Fenced blocks are now stripped, with fence state reset per comment body (blockquotes never needed handling — a > prefix already fails the anchor);
  • "the head prefix appearing in an unrelated URL on the line does not count" — this paragraph's own earlier claim — was false. The implementation took the first @<hex> anywhere on the line, so Review-verdict: MERGEABLE [x](https://e/@0123456) was graded against the link. The sha is now read from the verdict's own @ <sha> field, the one immediately following the token, which also makes multi-@ lines unambiguous.

A cross-family review of that first fix found three more, all reproduced before fixing — worth recording because each is the same fix done half-way:

  • fences were stripped for ``` only, but markdown also accepts ~~~;
  • the sha field matched {7,40} with no right boundary, so an over-long or malformed token was silently truncated into a passing one: @ <40-hex-head>f and @ <40-hex-head>ZZZ both graded as a verdict for head. The hex run is now matched whole, must end at a non-alphanumeric boundary, and its length is validated separately — an out-of-range token is rejected, never trimmed to fit;
  • fence stripping toggled on any line of three-or-more markers. Markdown closes a fence only with N-or-more of the same marker it was opened with, so a ```` block legitimately contains a
  • comment bodies were joined with a literal \x01BODY-BOUNDARY\x01 line so fence state could reset per comment. An in-band delimiter is forgeable by whoever writes the data — and here that is anyone who can comment on the PR: a body containing that line reset the fence mid-comment and exposed a verdict still inside an unclosed fence. Bodies are now carried out-of-band (one JSON-encoded string per line), which removes the class rather than escaping the sentinel.

A fourth round found the last one, and it is why the grammar was narrowed rather than patched again: markdown has a second code-block form — indented blocks (4 spaces or a tab) — which the fence stripper does not cover, so a pasted indented example still self-approved. Rather than add a second stripper (the same instance-not-class move that produced the previous three rounds), the marker must now sit at COLUMN 0. That removes every indentation-based ambiguity at once. The cost is that a verdict indented under a list item is ignored and classifies absent — which asks a human, the safe direction for a gate. Fence detection keeps its leading-whitespace tolerance, because stripping more is always safe.

Four review rounds, four sets of real findings, and rounds 24 each found a bug introduced while fixing the round before — every one the same fix applied half-way (one fence marker but not the other; a bounded left side but not the right; state reset per comment but via a delimiter the writer controls). For attacker-writable text, budget several rounds and prefer eliminating a class over enumerating its members.

A fifth round found raw HTML — <pre>, <code>, HTML comments — the third code-block form. It is stripped too, and hardening stopped there, deliberately.

Scope limit, stated so nobody re-derives it: the comment scanner is a best-effort heuristic, not a markdown parser. It handles the three code-block forms markdown actually has (fenced, indented via the column-0 rule, raw HTML) and errs toward stripping more, because every error in that direction can only withhold approval. It is not proof against every conceivable way to render text as non-prose, and chasing that was demonstrably not converging: five review rounds, five code-block forms, four of them introduced while fixing the previous one.

Stopping is safe because the comment is not the load-bearing gate. Since #622 the authoritative signal is the review-verdict/h10 commit status, written only by scripts/post-review-verdict.sh from explicit arguments — a comment cannot forge it, whatever it contains. The comment classifier is condition (c) of the PreToolUse hook, i.e. defense in depth on an agent's merge call. A residual false-open there means the hook does not object; it does not mean a merge happens.

The rule that generalises: when a heuristic keeps failing at the edges, check whether it is actually the thing enforcing the invariant before spending another round on it.

The lesson is about where a grammar lives, not about any one regex: this record described the intended behaviour accurately, the code did something looser, and nothing compared them. The grammar now lives in scripts/check-review-verdict.sh — one implementation, called by the hook and by post-review-verdict.sh's cross-check, covered by scripts/tests/test_check_review_verdict.py.

Two of those fixes also broke previously-green tests, which is the useful part: an over-long token became no-sha where a fixture expected stale (the fixture was 45 hex chars, so it had been testing the length guard while claiming to test the prefix rule), and jq -e exits 4 on no output, so an empty comment list started reading as a malformed payload — turning "no comments yet" into an input error on a gate whose callers fail closed.

The classifications:

  • a MERGEABLE/APPROVED/LGTM verdict whose @ <sha> is the current head → allow;
  • a negative verdict (BLOCKED/NOT-MERGEABLE) on the headdeny, and it wins over a positive one on the same head (a later BLOCKED retracts an earlier MERGEABLE; to retract, re-review head and post BLOCKED @ head). Staleness is symmetric on purpose: a negative for an older commit is stale exactly like a positive for an older commit, and does NOT override a fresh head-positive — otherwise a pre-fix BLOCKED @ oldsha would block forever even after the fix changes the sha and earns a fresh MERGEABLE @ head (the normal flow). So a genuine block must reference head, per the convention;
  • verdict comment(s) exist but reference only older commits → deny — the stale-review case #242 targets;
  • a verdict line whose token is in neither vocabulary (MERGEABLE-LATER, SHIP-IT, …) → ask (#629). Deliberately not read as approval, and deliberately not read as a block either — an unrecognized token means the reviewer's intent is unknown, so it goes to a human;
  • a Review-verdict: marker with no @ <sha> in its own field → ask (a lazy/quoted marker; not mislabelled as stale);
  • no Review-verdict: comment at all → ask (graceful adoption, mirrors H6's "no Done-when → ask": surface, don't hard-block a PR that hasn't adopted the convention yet);
  • comments unfetchable / head sha unresolvable → ask.

Scope: the Claude PreToolUse gate on the Gitea merge tool only. A direct git push origin main has no PR comments to check, so the .husky/pre-push backstop is not extended for H10 (the merge tool is the real merge path; docs-only PRs remain exempt via H6's file-set exemption). Rationale, as with the whole Wave-1/2/3 hook set: make the process rule a derivation/hook, not prose to remember (#303).