Files
ersatztv/scripts/tests/test_pr_changed_files.py
T
timothyandClaude Opus 5 fe00e0d71f
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 37s
PR Gates / Docs update reminder (pull_request) Successful in 40s
PR Gates / decisions lifecycle (pull_request) Successful in 45s
review-verdict/h10 Awaiting review verdict for fe00e0d
Review verdict / Set review-verdict status (pull_request_target) Successful in 14s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m23s
PR Gates / Script tests (pytest) (pull_request) Successful in 1m19s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m34s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 9m12s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 21m4s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 24m7s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
fix(698): compare the recorded base exactly, never parse it out
Review round 5 returned BLOCKED with one High, and it needed no forgery and no #697 —
just a branch name.

`main)evil` IS A VALID GIT BRANCH NAME (`git check-ref-format --branch 'main)evil'`
succeeds). A genuine human verdict earned while head H targeted it is written
`(base: main)evil)`. Truncating at the first `)` yields exactly `main`, which matches a
PR that has since been retargeted onto `main`, so the verdict is inherited over a
completely different diff.

I had asserted the opposite in a code comment one commit earlier — that a `)` in a
branch name "mismatches — safe direction". That was generalised from `feat/foo)bar`,
which does mismatch, and is false for EVERY branch whose name starts with the target
base. Two attempts at extracting this value have now been defeated (`##` last-marker by
an appended marker, `#` first-marker by this), so the lesson is the shape, not the
off-by-one: do not parse a value out of user- or attacker-influenced text when you can
compare against the exact expected literal instead.

The description must now END with the literal `(base: <this PR's base>)` AND contain
exactly ONE marker — the marker count kills the append trick without having to decide
which occurrence is authoritative. Pure shell (`${#}` arithmetic), no truncation to
abuse. Verified across all six shapes, including a PR that legitimately targets
`main)evil` (accepted) and `(base: )` (rejected). Absent markers remain accepted, since
verdicts predating #632 carry none.

Mutation-verified: restoring the truncating parse reddens only the new paren test, while
the appended-marker, matching-base and legacy tests stay green.

385 tests pass. Note for the record: pytest has never executed inside the review sandbox
in any of the five rounds, so the suite has only ever been run here.

Refs: #698
Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 20:51:57 +02:00

1545 lines
82 KiB
Python

"""Tests for `scripts/pr-changed-files.sh` — the SHARED PR file enumeration (ersatztv#649).
Why this file exists. The enumeration used to be written twice: once in
`.claude/hooks/pretooluse-merge-consent.sh` (ADVISORY — a failure produces a human prompt) and once
inline in `.gitea/workflows/review-verdict.yml` (ENFORCED — it writes the branch-protection-required
`review-verdict/h10` status). They drifted, and in the dangerous direction: four rounds of
ersatztv#643 hardening landed on the advisory copy and never reached the enforced one, so the copy
with real authority ended up strictly weaker than the copy without.
The specific thing this suite pins is the point of ersatztv#649's second Done-when box. A round-4
review traced that the enforced copy's fail-closed behaviour on a garbage response was INCIDENTAL,
not designed: `n` came back empty, `[ "$n" -lt 50 ]` errored to false, the loop ran to MAX_PAGES and
left complete=no. The right answer, reached through a bash arithmetic error that any refactor of the
loop could have silently flipped. Every failure-path test below therefore asserts a NON-ZERO exit
explicitly, so the behaviour is a contract rather than a coincidence.
Observable contract of the script:
exit 0 -> enumeration complete and bound to the expected head; stdout is the authoritative path set
exit 1 -> could not enumerate/verify; stdout meaningless, caller MUST withhold any exemption
exit 2 -> usage error
"""
from __future__ import annotations
import json
import os
import re
import subprocess
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parents[2]
SCRIPT = REPO_ROOT / "scripts" / "pr-changed-files.sh"
HOOK = REPO_ROOT / ".claude" / "hooks" / "pretooluse-merge-consent.sh"
WORKFLOW = REPO_ROOT / ".gitea" / "workflows" / "review-verdict.yml"
SHA = "a9e3e23abf337980ca4c05854f5b1e210099d08b"
OTHER_SHA = "b71c0d4e2f8a91b3c5d7e9f1a3b5c7d9e1f3a5b7"
# Serves paged `pulls/N/files`, plus the PR object the script re-reads to bind the enumeration.
CURL_SHIM = r'''#!/usr/bin/env python3
import json, os, sys, pathlib, urllib.parse
state = pathlib.Path(os.environ["STUB_DIR"])
args = sys.argv[1:]
url = [a for a in args if a.startswith("http")][-1]
if "/pulls/" in url and "/files" in url:
q = urllib.parse.parse_qs(urllib.parse.urlparse(url).query)
page = int(q.get("page", ["1"])[0])
pages = json.loads((state / "pages.json").read_text())
if page > len(pages):
print("[]"); sys.exit(0)
entry = pages[page - 1]
if entry == "ERROR": # transport failure on this page
sys.exit(22)
if entry == "GARBAGE": # 200 with a non-array body (proxy/error page)
print('{"message":"internal error"}'); sys.exit(0)
print(json.dumps(entry)); sys.exit(0)
if "/pulls/" in url:
# The PR object is read TWICE now: once before paging to bind the base, once after to bind both
# the head and the base (ersatztv#698). A counter distinguishes them, so a test can move either
# field at either end and the two guards can be isolated from each other.
counter = state / "pr_reads.txt"
n = int(counter.read_text()) if counter.exists() else 0
counter.write_text(str(n + 1))
sha = os.environ["STUB_SHA"]
alt = state / "pr_sha_after.txt"
if alt.exists() and n >= 1: # force-push landing between pagination round-trips
sha = alt.read_text().strip()
base = os.environ.get("STUB_BASE", "main")
base_before = state / "pr_base_before.txt"
base_after = state / "pr_base_after.txt"
if base_before.exists() and n == 0: # already retargeted when the job started
base = base_before.read_text().strip()
if base_after.exists() and n >= 1: # retargeted mid-enumeration
base = base_after.read_text().strip()
print(json.dumps({"head": {"sha": sha}, "base": {"ref": base}}))
sys.exit(0)
print("{}")
'''
def _rows(paths, status="modified"):
return [{"filename": p, "status": status} for p in paths]
@pytest.fixture
def enumerate_files(tmp_path):
bindir = tmp_path / "bin"; bindir.mkdir()
shim = bindir / "curl"; shim.write_text(CURL_SHIM); shim.chmod(0o755)
state = tmp_path / "state"; state.mkdir()
env = dict(os.environ)
env["PATH"] = f"{bindir}{os.pathsep}{env['PATH']}"
env["STUB_DIR"] = str(state)
env["STUB_SHA"] = SHA
env["ETV_GITEA_TOKEN"] = "stub"
env["ETV_GITEA_URL"] = "http://gitea.example"
env.pop("ETV_GITEA_BASICAUTH", None)
env.pop("GITEA_TOKEN", None)
class Handle:
def __init__(self, env_, state_):
self.env = env_
self.state = state_
def set_pages(self, *pages):
(state / "pages.json").write_text(json.dumps(list(pages)))
def head_moves_to(self, sha):
(state / "pr_sha_after.txt").write_text(sha)
def base_is_already(self, ref):
"""The PR targets `ref` before the first page is requested (retargeted pre-run)."""
(state / "pr_base_before.txt").write_text(ref)
def base_moves_to(self, ref):
"""The PR is retargeted to `ref` between the files pages and the binding re-read."""
(state / "pr_base_after.txt").write_text(ref)
def run(self, expected_sha=SHA, args=("timothy", "ersatztv", "42"),
expected_base="main"):
argv = ["bash", str(SCRIPT), *args, expected_sha]
if expected_base is not None:
argv.append(expected_base)
return subprocess.run(argv, env=self.env, capture_output=True, text=True)
def paths(self):
"""Assert success and return the enumerated path set."""
r = self.run()
assert r.returncode == 0, f"expected success, got {r.returncode}: {r.stderr}"
return [ln for ln in r.stdout.splitlines() if ln.strip()]
def fails_closed(self):
"""The whole point: a NON-ZERO exit, asserted, not inferred."""
r = self.run()
return r.returncode != 0
return Handle(env, state)
# --- The happy path, so the failure-path assertions below cannot pass vacuously ----------------
def test_complete_enumeration_returns_every_path(enumerate_files):
enumerate_files.set_pages(_rows(["docs/a.md", "ErsatzTV/Program.cs", "README.md"]))
assert enumerate_files.paths() == ["docs/a.md", "ErsatzTV/Program.cs", "README.md"]
def test_a_rename_contributes_BOTH_sides(enumerate_files):
"""One row, two paths — the `git mv` hole. Reading `.filename` alone hides the source."""
enumerate_files.set_pages([{"filename": "docs/innocuous-note.md", "status": "renamed",
"previous_filename": ".gitea/workflows/renovate.yml"}])
assert sorted(enumerate_files.paths()) == [".gitea/workflows/renovate.yml",
"docs/innocuous-note.md"]
def test_paths_on_a_later_page_are_included(enumerate_files):
enumerate_files.set_pages(_rows([f"docs/f{i}.md" for i in range(50)]),
_rows(["scripts/decisions_lib.py"]))
assert "scripts/decisions_lib.py" in enumerate_files.paths()
# --- Fail-closed contract: each of these MUST be non-zero, by design ---------------------------
def test_transport_failure_mid_pagination_fails_closed(enumerate_files):
"""The defect that started all of this: an errored page counted as zero rows and read as
'end of list', completing the enumeration over a PARTIAL list."""
enumerate_files.set_pages(_rows([f"docs/f{i}.md" for i in range(50)]), "ERROR",
_rows(["docs/tail.md"]))
assert enumerate_files.fails_closed()
def test_non_array_body_fails_closed(enumerate_files):
enumerate_files.set_pages("GARBAGE")
assert enumerate_files.fails_closed()
def test_an_OBJECT_OF_VALID_ROWS_isolates_the_top_level_array_check(enumerate_files):
"""Cold review found `test_non_array_body_fails_closed` passing for the wrong reason, and the
first attempt to fix it failed for a THIRD reason — worth recording, because both near-misses
look like coverage.
`jq`'s `all(.[]; …)` iterates an object's VALUES, so the top-level `type == "array"` check is
only load-bearing when those values would themselves validate:
* `{"message":"internal error"}` — values are strings, `.filename` on a string errors. Rejected
with the clause deleted, so it never isolated it.
* `{"filename":"docs/a.md","status":"modified"}` — a single row, but its values are still
strings. Same non-isolation, one level less obvious.
* `{"0": {"filename":"docs/a.md","status":"modified"}}` — values ARE valid rows. With the clause
deleted this validates, `length` is 1, the path is collected, and the next page ends the
enumeration cleanly: a non-array body enumerated as a complete docs-only list. That is the
fail-open, and only this shape exposes it.
"""
enumerate_files.set_pages({"0": {"filename": "docs/a.md", "status": "modified"}})
assert enumerate_files.fails_closed()
def test_rows_without_filename_fail_closed(enumerate_files):
"""`[{}]` is a well-formed array that yields no paths — a partial list wearing a valid shape."""
enumerate_files.set_pages([{}, {}])
assert enumerate_files.fails_closed()
def test_a_row_with_a_VALID_status_but_no_filename_isolates_the_filename_check(enumerate_files):
"""Same wrong-reason problem: `[{}]` is also rejected by the closed `.status` allow-list, since
an absent status is not in it. A row carrying a legitimate `status` and no `filename` removes
that second reason, leaving only the guard under test."""
enumerate_files.set_pages([{"status": "modified"}])
assert enumerate_files.fails_closed()
def test_array_of_scalars_fails_closed(enumerate_files):
enumerate_files.set_pages(["docs/a.md", "docs/b.md"])
assert enumerate_files.fails_closed()
@pytest.mark.parametrize("evil", ["safe.md\ndocs/Program.cs", "safe.md\rdocs/Program.cs"])
def test_CRLF_in_filename_fails_closed(enumerate_files, evil):
"""A newline splits one path into two lines that are each allow-list-matched separately, so
`safe.md\\ndocs/Program.cs` reads as two exempt paths while the real path ends in .cs."""
enumerate_files.set_pages(_rows([evil]))
assert enumerate_files.fails_closed()
def test_CRLF_in_previous_filename_on_a_NON_renamed_row_fails_closed(enumerate_files):
"""The hole one predicate wide: `previous_filename` is CONSUMED on every row, so it must be
VALIDATED on every row — not only where `.status == "renamed"` makes it semantically expected."""
enumerate_files.set_pages([{"filename": "docs/a.md", "status": "modified",
"previous_filename": "safe.md\ndocs/Program.cs"}])
assert enumerate_files.fails_closed()
def test_dotdot_path_component_fails_closed(enumerate_files):
"""The callers' allow-lists anchor `^docs/`, which `docs/../ErsatzTV/Program.cs` matches."""
enumerate_files.set_pages(_rows(["docs/../ErsatzTV/Program.cs"]))
assert enumerate_files.fails_closed()
@pytest.mark.parametrize("status", ["Renamed", "moved", ""])
def test_status_outside_the_closed_allow_list_fails_closed(enumerate_files, status):
"""Without a closed set, the `renamed => previous_filename REQUIRED` clause is dodgeable by any
other value, letting a `git mv` drop its source path and read as docs-only."""
enumerate_files.set_pages([{"filename": "docs/a.md", "status": status,
"previous_filename": "ErsatzTV/Program.cs"}])
assert enumerate_files.fails_closed()
def test_renamed_row_without_previous_filename_fails_closed(enumerate_files):
enumerate_files.set_pages([{"filename": "docs/a.md", "status": "renamed"}])
assert enumerate_files.fails_closed()
def test_head_differing_from_the_expected_sha_fails_closed(enumerate_files):
"""Narrowed to what this actually proves, per cold review.
The stub serves the alternate sha from the BINDING read (the one after paging; since
ersatztv#698 the script also reads the PR object BEFORE paging, to bind the base). So this
exercises "the final head does not equal the expected sha", not movement *during* enumeration.
The distinction matters: an A->B->A force-push round trip would restore the expected sha and
pass this check while the pages came from two different states. That race is inherent to
enumerating a mutable list over several round-trips against an API with no commit-pinned files
endpoint, and is tracked separately rather than papered over with a test name that implies it is
covered.
"""
enumerate_files.set_pages(_rows(["docs/a.md"]))
enumerate_files.head_moves_to(OTHER_SHA)
assert enumerate_files.fails_closed()
def test_missing_credentials_fails_closed(enumerate_files):
enumerate_files.set_pages(_rows(["docs/a.md"]))
enumerate_files.env.pop("ETV_GITEA_TOKEN", None)
assert enumerate_files.fails_closed()
@pytest.mark.parametrize("args", [("timothy", "ersatztv"), ("timothy", "ersatztv", "")])
def test_usage_errors_exit_2(enumerate_files, args):
enumerate_files.set_pages(_rows(["docs/a.md"]))
r = enumerate_files.run(args=args) if len(args) == 3 else subprocess.run(
["bash", str(SCRIPT), *args], env=enumerate_files.env, capture_output=True, text=True)
assert r.returncode == 2, r.stderr
# --- The base-ref binding (ersatztv#698 route 1) ------------------------------------------------
#
# `/pulls/{n}/files` diffs against the PR's LIVE base, so retargeting changes the answer without
# moving the head. Reproduced on the real instance as probe PR #703: opened into `main`, retargeted
# mid-run, enumerated docs-only, granted `review-verdict/h10=success` while its diff against `main`
# carried a C# file.
def test_a_FOUR_argument_call_is_a_usage_error_not_an_unbound_enumeration(enumerate_files):
"""The binding is REQUIRED, not optional.
This is the test that matters most for the shape of the fix. Had the base ref been added as an
OPTIONAL 5th argument, every existing caller would have kept compiling and kept running
unbound — an opt-out that is invisible at the call site, and the caller most likely to omit it
is the one that most needed it. Dropping the argument must be LOUD.
"""
enumerate_files.set_pages(_rows(["docs/a.md"]))
r = enumerate_files.run(expected_base=None)
assert r.returncode == 2, (
"a 4-argument call was accepted, so the base binding is effectively optional and any caller "
f"that forgets it silently enumerates against a mutable base (stdout={r.stdout!r})")
def test_an_EMPTY_base_ref_argument_fails_closed(enumerate_files):
"""The hook passes `.base.ref` straight from PR JSON; unparseable JSON yields an empty string."""
enumerate_files.set_pages(_rows(["docs/a.md"]))
r = enumerate_files.run(expected_base="")
assert r.returncode == 2, r.stderr
def test_a_PR_already_retargeted_before_the_enumeration_fails_closed(enumerate_files):
"""The pre-paging read. Without it, a PR retargeted before the job started would enumerate
against the scratch base with every page agreeing with every other page — internally consistent
and entirely wrong."""
enumerate_files.set_pages(_rows(["docs/a.md"]))
enumerate_files.base_is_already("probe/scratch-base")
assert enumerate_files.fails_closed()
def test_a_RETARGET_DURING_the_enumeration_fails_closed(enumerate_files):
"""The post-paging read. The head never moves here, which is the entire point: the existing
head-sha binding cannot see a retarget, so removing the base half of the binding leaves this
case exempted."""
enumerate_files.set_pages(_rows(["docs/a.md"]))
enumerate_files.base_moves_to("probe/scratch-base")
assert enumerate_files.fails_closed()
def test_positive_control_an_unmoved_base_still_enumerates(enumerate_files):
"""Without this, the three tests above could pass because the 5-argument call is broken outright
rather than because the binding works."""
enumerate_files.set_pages(_rows(["docs/a.md", "ErsatzTV/Program.cs"]))
assert enumerate_files.paths() == ["docs/a.md", "ErsatzTV/Program.cs"]
def test_the_base_comparator_is_the_BRANCH_NAME_never_the_TIP_SHA():
"""`.base.ref` is compared, never `.base.sha` — matching `post-review-verdict.sh` (ersatztv#632).
Comparing tips would fail every enumeration on every unrelated merge to `main`: a self-inflicted
deadlock dressed as a security control.
This is a STRUCTURAL assertion on purpose, and the previous version of this test is why. It was
written behaviourally as `head_moves_to(SHA)` — the sha that was ALREADY current — so it modelled
no movement at all and was simply a duplicate positive control. It would have passed just as
happily against a script comparing tip shas. The stub serves only branch names, so no behavioural
test in this harness can distinguish the two comparators; say so and assert the source instead.
"""
src = SCRIPT.read_text()
assert ".base.ref" in src, "the enumeration no longer reads .base.ref"
binding = [ln for ln in src.splitlines()
if "base_before=" in ln or "base_after=" in ln]
assert binding, "no base binding assignments found"
for ln in binding:
assert ".base.ref" in ln and ".base.sha" not in ln, (
f"the base binding compares a tip sha, which deadlocks on ordinary churn: {ln.strip()!r}")
def test_a_SHORT_page_does_not_end_the_enumeration(enumerate_files):
"""Gitea caps `limit` at the server-wide MAX_RESPONSE_ITEMS and may return fewer than asked.
'Fewer than 50 rows means last page' would complete over a partial list without any transport
error — so termination requires a validated EMPTY page. A 30-row page followed by code must be
seen."""
enumerate_files.set_pages(_rows([f"docs/f{i}.md" for i in range(30)]),
_rows(["ErsatzTV/Program.cs"]))
assert "ErsatzTV/Program.cs" in enumerate_files.paths()
# --- The same guard, under the runner's jq 1.6 ------------------------------------------------
#
# `test_transport_failure_mid_pagination_fails_closed` above does NOT isolate the explicit
# `if [ -z "${raw//[[:space:]]/}" ]` clause — cold review claimed this and mutation confirmed it:
# deleting that clause leaves the whole suite green on a developer Mac, because jq 1.8 rejects empty
# input on its own. jq 1.6 does not, and the runner ships 1.6 — so the one environment where the
# clause is load-bearing was the one environment with no coverage. That is the #643/#647 failure
# class exactly, reproduced in the test suite meant to prevent it.
#
# The shim is IMPORTED, not copied. A second copy of a version-quirk emulator is the same
# two-implementations-drift problem this whole issue is about, one level down.
from scripts.tests.test_merge_consent_exemption import _JQ16_SHIM # noqa: E402
@pytest.fixture
def enumerate_files_jq16(enumerate_files, tmp_path):
"""The same harness, plus a jq shim reproducing jq 1.6's empty-input `-e` exit status."""
jq = tmp_path / "bin" / "jq"
jq.write_text(_JQ16_SHIM)
jq.chmod(0o755)
return enumerate_files
def test_jq16_shim_actually_reproduces_the_quirk(enumerate_files_jq16, tmp_path):
"""Verify the verifier. A shim that failed to install would make the test below pass vacuously,
reporting the guard safe on jq 1.6 without ever exercising the quirk."""
assert enumerate_files_jq16 is not None # the fixture is what installs the shim
jq = str(tmp_path / "bin" / "jq")
empty = subprocess.run([jq, "-e", "."], input="", capture_output=True, text=True)
assert empty.returncode == 0, "the shim does not reproduce jq 1.6's empty-input exit 0"
real = subprocess.run([jq, "-e", ".a"], input='{"a":1}', capture_output=True, text=True)
assert real.returncode == 0 and real.stdout.strip() == "1", "the shim broke ordinary jq"
false = subprocess.run([jq, "-e", ".a"], input='{"a":false}', capture_output=True, text=True)
assert false.returncode == 1, "the shim broke jq's real -e semantics for a false result"
def test_transport_failure_mid_pagination_fails_closed_on_jq16(enumerate_files_jq16):
"""The property, asserted on the interpreter that actually runs it in CI."""
enumerate_files_jq16.set_pages(_rows([f"docs/f{i}.md" for i in range(50)]), "ERROR",
_rows(["docs/tail.md"]))
assert enumerate_files_jq16.fails_closed()
def test_a_docs_only_pr_still_enumerates_cleanly_on_jq16(enumerate_files_jq16):
"""Positive control: the shim must not make everything fail, or the test above proves nothing."""
enumerate_files_jq16.set_pages(_rows(["docs/a.md", "docs/b.md"]))
assert enumerate_files_jq16.paths() == ["docs/a.md", "docs/b.md"]
# --- The CALLER contract: a non-zero exit must withhold the exemption, on its own -------------
#
# This is the single line the whole extraction rests on, and it was the one guard nothing pinned. A
# review mutated the hook's `if files=$(...)` into `files=$(...) || true; files_complete=yes` — i.e.
# ignore the exit status entirely — and the ENTIRE suite still passed.
#
# It survives today only by REDUNDANCY: the script writes stdout once, immediately before `exit 0`,
# so every failure path also happens to yield empty stdout, and the hook's independent
# `[ -n "$files" ]` check catches it. That is exactly the shape this PR criticises elsewhere — safe
# by accident rather than by assertion. Any future change that streams pages, or prints a partial
# list before failing, turns it into a live false exemption.
#
# So the stub below FAILS *while emitting a perfectly docs-only list*, which is the one combination
# the redundancy cannot absorb. The positive control immediately after it proves the harness can
# actually observe the difference, rather than reporting "not exempt" for some unrelated reason.
HOOK_STUB_CURL = r'''#!/usr/bin/env python3
import json, sys
url = [a for a in sys.argv[1:] if a.startswith("http")][-1]
if "/pulls/" in url and "/files" not in url:
print(json.dumps({"head": {"sha": "%s"}, "body": "no linked issue"}))
else:
print("{}")
''' % SHA
def _mirror_tree(tmp_path, stub_body):
"""Lay out a minimal repo mirror so the hook resolves OUR stub as the shared script.
The hook finds the script via `${BASH_SOURCE[0]}/../..`, so the mirror must reproduce the real
`.claude/hooks/` + `scripts/` shape rather than just dropping the stub anywhere.
"""
hooks = tmp_path / ".claude" / "hooks"
hooks.mkdir(parents=True)
(hooks / "pretooluse-merge-consent.sh").write_text(HOOK.read_text())
scripts = tmp_path / "scripts"
scripts.mkdir()
stub = scripts / "pr-changed-files.sh"
stub.write_text(stub_body)
stub.chmod(0o755)
bindir = tmp_path / "bin"; bindir.mkdir()
curl = bindir / "curl"; curl.write_text(HOOK_STUB_CURL); curl.chmod(0o755)
env = dict(os.environ)
env["PATH"] = f"{bindir}{os.pathsep}{env['PATH']}"
env["ETV_GITEA_TOKEN"] = "stub"
env["ETV_GITEA_URL"] = "http://gitea.example"
env.pop("ETV_GITEA_BASICAUTH", None)
payload = {"tool_input": {"method": "merge", "owner": "timothy",
"repo": "ersatztv", "pull_number": 42}}
r = subprocess.run(["bash", str(hooks / "pretooluse-merge-consent.sh")],
input=json.dumps(payload), env=env, capture_output=True, text=True)
assert r.returncode == 0, r.stderr
# Exempt == passthrough == exit 0 with no decision JSON on stdout.
return r.stdout.strip() == ""
DOCS_ONLY_OUTPUT = 'printf "docs/a.md\\ndocs/b.md\\n"\n'
def test_a_FAILING_script_withholds_the_exemption_even_when_stdout_looks_docs_only(tmp_path):
exempt = _mirror_tree(tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY_OUTPUT + "exit 1\n")
assert exempt is False, (
"the hook granted a docs-only exemption from the stdout of a script that FAILED — the exit "
"status is not being checked")
def test_positive_control_the_same_output_with_exit_0_DOES_exempt(tmp_path):
"""Without this, the test above could pass for any unrelated reason and prove nothing."""
exempt = _mirror_tree(tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY_OUTPUT + "exit 0\n")
assert exempt is True, (
"the positive control failed, so the negative test above cannot be trusted to be measuring "
"the exit status at all")
def test_a_script_that_crashes_also_withholds_the_exemption(tmp_path):
"""Not every failure is a clean `exit 1` — a crash must not read as success either."""
exempt = _mirror_tree(
tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY_OUTPUT + "kill -TERM $$\n")
assert exempt is False
def test_a_MISSING_script_withholds_the_exemption(tmp_path):
"""The transition window: the shared checkout has no such script until this lands."""
hooks = tmp_path / ".claude" / "hooks"
hooks.mkdir(parents=True)
(hooks / "pretooluse-merge-consent.sh").write_text(HOOK.read_text())
(tmp_path / "scripts").mkdir()
bindir = tmp_path / "bin"; bindir.mkdir()
curl = bindir / "curl"; curl.write_text(HOOK_STUB_CURL); curl.chmod(0o755)
env = dict(os.environ)
env["PATH"] = f"{bindir}{os.pathsep}{env['PATH']}"
env["ETV_GITEA_TOKEN"] = "stub"
env["ETV_GITEA_URL"] = "http://gitea.example"
payload = {"tool_input": {"method": "merge", "owner": "timothy",
"repo": "ersatztv", "pull_number": 42}}
r = subprocess.run(["bash", str(hooks / "pretooluse-merge-consent.sh")],
input=json.dumps(payload), env=env, capture_output=True, text=True)
assert r.returncode == 0, r.stderr
assert r.stdout.strip() != "", "a missing shared script must not grant an exemption"
# --- Drift guard: the reason this file is worth having at all ----------------------------------
#
# Scope is BOTH callers. Until the enforced workflow was rewired (the follow-up half of ersatztv#649)
# this guard could only assert the hook, which left the copy with real authority — the one that
# writes the branch-protection-required `review-verdict/h10` status — unpinned. That asymmetry was
# the entire subject of the issue, so a guard that covered only the advisory side would have been
# the same mistake one level up.
CALLERS = pytest.mark.parametrize(
"caller", [HOOK, WORKFLOW], ids=["advisory-hook", "enforced-workflow"])
def _code_lines(path: Path) -> str:
"""Strip comment lines.
Both callers' comments legitimately discuss the `pulls/N/files` endpoint, and a future comment
writing `pulls/$pr/files?limit=100` as an example of what NOT to do would redden these tests —
which, since `script-tests` red blocks merges via the combined status, would block the repo over
a piece of prose. `#`-prefixed works for both files: YAML comments and the shell comments inside
the workflow's `run:` block share the marker.
"""
return "\n".join(
ln for ln in path.read_text().splitlines() if not ln.lstrip().startswith("#"))
@CALLERS
def test_the_caller_uses_the_shared_script(caller):
assert "scripts/pr-changed-files.sh" in _code_lines(caller), (
f"{caller.relative_to(REPO_ROOT)} no longer calls the shared enumeration")
@CALLERS
def test_the_caller_passes_the_BASE_REF_argument(caller):
"""Both callers must invoke the 5-argument form (ersatztv#698 route 1).
The runtime tests prove the SCRIPT rejects a 4-argument call. They cannot prove a caller still
makes a 5-argument one, and the failure is quiet in opposite directions at the two sites: the
workflow would die on `set -u` (fail-closed, but it takes every PR with it), while the hook would
pass an empty base and lose every docs-only exemption. Pin the call site itself.
Matching is deliberately narrow. A first draft keyed on "any line mentioning the script name" and
matched the enum_error MESSAGE string, failing for a reason that had nothing to do with the call.
The workflow also invokes through `"$ENUM"` rather than the literal path, so the alias is resolved
here and asserted to point at the shared script — otherwise this test could be satisfied while
`ENUM` pointed somewhere else entirely.
"""
code = _code_lines(caller)
if '"$ENUM"' in code:
assert re.search(r'^\s*ENUM=\S*scripts/pr-changed-files\.sh\s*$', code, re.M), (
f"{caller.relative_to(REPO_ROOT)} invokes \"$ENUM\" but ENUM is not assigned the shared "
"enumeration script")
invocations = [ln for ln in code.splitlines()
if re.search(r'\$\(\s*"(\$ENUM|[^"]*pr-changed-files\.sh)"', ln)]
assert invocations, (
f"{caller.relative_to(REPO_ROOT)} has no executable call to the shared enumeration")
for ln in invocations:
after = ln.split('"', 2)[2] if '"$ENUM"' in ln else ln.split("pr-changed-files.sh", 1)[1]
args = re.findall(r'"[^"]*\$[^"]*"', after)
assert len(args) >= 5, (
f"{caller.relative_to(REPO_ROOT)} calls the enumeration with {len(args)} quoted "
f"arguments, expected 5 including the expected base ref: {ln.strip()!r}")
# Counting five arguments is not enough — cold review caught that passing `"$SHA"` twice
# satisfied the count while stalling every real exemption. Name the fifth.
assert re.search(r'(?i)base', args[4]), (
f"{caller.relative_to(REPO_ROOT)} passes {args[4]} as the 5th argument; it must be the "
f"expected BASE ref: {ln.strip()!r}")
@CALLERS
def test_the_caller_does_not_reimplement_the_enumeration(caller):
"""Structural rather than behavioural on purpose. Behavioural equivalence tests would still pass
if someone pasted the loop back inline and kept it correct *that day* — which is exactly how the
drift happened the first time. What must be prevented is a SECOND implementation existing.
"""
# An inline `pulls/<n>/files?` fetch is the signature of a re-inlined copy. Match on the endpoint
# alone, NOT on `?limit=` — an earlier version anchored the query string, so a copy written as
# `files?page=1&limit=50` would have walked straight past a guard that exists to stop exactly
# that. Still evadable by a copy that builds the URL without a literal `?`, so this narrows the
# gap rather than closing it.
assert not re.search(r"pulls/\$?\{?\w+\}?/files\?", _code_lines(caller)), (
f"{caller.relative_to(REPO_ROOT)} appears to enumerate PR files inline again — "
"that is the duplication ersatztv#649 removed")
# --- The ENFORCED caller's own preconditions ---------------------------------------------------
def _workflow_steps():
import yaml
wf = yaml.safe_load(WORKFLOW.read_text())
return wf["jobs"]["set-verdict-status"]["steps"]
def _classify_step():
for step in _workflow_steps():
if "review-verdict/h10" in (step.get("run") or ""):
return step
raise AssertionError("no step in review-verdict.yml posts review-verdict/h10")
def _workflow_triggers():
"""The `on:` block, tolerating YAML 1.1's `on` -> True coercion.
`yaml.safe_load` parses the bare key `on` as the BOOLEAN True, not the string "on" — so the
obvious `wf["on"]` raises KeyError against a perfectly valid workflow. Looking up both is not
defensive padding: a test that died on a KeyError here would read as "the trigger assertion is
broken" rather than "the trigger changed", which is the wrong failure to hand a maintainer.
"""
import yaml
wf = yaml.safe_load(WORKFLOW.read_text())
on = wf.get("on", wf.get(True))
assert isinstance(on, dict), f"review-verdict.yml has no parseable `on:` mapping (got {on!r})"
return on
def test_the_workflow_trigger_is_pull_request_TARGET_scoped_to_main():
"""The definition-rewrite hole (ersatztv#672) — sibling of the base-ref checkout below.
That checkout binds the SCRIPTS this job runs to the base. It cannot bind the job DEFINITION:
Gitea resolves a `pull_request` workflow definition from the PR's own head, so a PR editing
review-verdict.yml ran its own rewritten copy and could post `review-verdict/h10=success` for
itself. Branch protection does not care who posted the context, and carries
`required_approvals: 0`.
Both halves are asserted because either alone closes nothing:
- `pull_request_target` resolves the definition from the base.
- `branches: [main]` keeps "the base" from being an attacker-pushed branch. Base resolution
without it merely moves the rewrite from the head to a scratch base — and since a commit
status is repo-global per sha (ersatztv#663), a success forged there is inherited by a later
real PR into `main` with the same head.
Parsed, not substring-matched, for the reason the checkout test gives — and here the text-level
version is actively broken rather than merely weak: `pull_request` is a PREFIX of
`pull_request_target`, so `"pull_request" in text` cannot tell the safe trigger from the
vulnerable one, and `"pull_request_target" in text` stays green when a plain `pull_request`
trigger is ADDED back alongside it.
"""
on = _workflow_triggers()
# EXACT SET, not "target present and plain absent". Cold review found the weaker pair of
# assertions green after ADDING `workflow_dispatch:` or `push:` alongside the safe trigger —
# both are ref-resolved and both get secrets, so either one restores an equivalent
# self-supplied-definition path while the test reports clean. Enumerating the two known-bad
# extra triggers would have the same hole one trigger later; pinning the whole set does not.
assert set(on) == {"pull_request_target"}, (
f"review-verdict.yml must trigger on `pull_request_target` and NOTHING else; got "
f"{sorted(map(str, on))}. Plain `pull_request` takes the workflow DEFINITION from the PR "
"head (ersatztv#672), and `push`/`workflow_dispatch` resolve it from an arbitrary ref — "
"any of them, ADDED ALONGSIDE rather than replacing, reopens the hole")
branches = (on["pull_request_target"] or {}).get("branches")
assert branches == ["main"], (
f"`pull_request_target.branches` is {branches!r}; it must be exactly ['main']. Base "
"resolution means the BASE branch supplies the gate, so an unfiltered trigger lets a PR "
"into an attacker-pushed base run that branch's rewritten copy (ersatztv#672)")
def test_no_OTHER_workflow_writes_the_review_verdict_status():
"""`review-verdict.yml` is the only workflow allowed to write `review-verdict/h10`.
The trigger test above hardens ONE file, and by construction cannot see the more general
problem: the gate is forgeable by *any* head-resolved workflow holding credentials that can POST
a commit status, not only by the gate's own definition. `docker-build.yml` is exactly that — it
triggers on `pull_request` (correctly; it builds the PR's code) and carries `ETV_STATUS_AUTH`,
which a probe confirmed can WRITE statuses, not merely read them. That residual is tracked in
ersatztv#697 and is NOT closed by this test.
What this test does close is the cheap regression: a second workflow quietly starting to write
the context. It is a guard against drift, not a security boundary — a workflow can still write
the status through an indirection this scan cannot see (a variable, a script it calls). Said
plainly so nobody reads a green here as "only one writer is possible".
"""
offenders = []
# `*.y*ml`, not `*.yml` — a workflow added as `.yaml` is just as executable and would otherwise
# be silently unscanned. Latent today (no `.yaml` exists), which is exactly when it is cheap.
for wf in sorted((REPO_ROOT / ".gitea" / "workflows").glob("*.y*ml")):
if wf.name == WORKFLOW.name:
continue
# CODE lines only — `docs`-style commentary in a workflow header legitimately discusses the
# context by name, and a raw-text scan would red on the prose explaining this very rule.
if "review-verdict/h10" in _code_lines(wf):
offenders.append(wf.name)
assert not offenders, (
f"{offenders} reference `review-verdict/h10` in executable lines. Only review-verdict.yml "
"may write the gate's own status; another workflow doing so is a forgery path (see "
"ersatztv#697) or, at best, a second implementation of the gate that will drift")
def test_the_workflow_checks_out_the_BASE_ref_never_the_head():
"""The self-rewrite hole. This workflow judges the PR, so the PR must not supply the code that
judges it: checking out `head.sha` would let a PR rewrite `scripts/pr-changed-files.sh` to
return an empty list and exempt itself.
Asserted through a YAML parse rather than a substring search, because the dangerous edit is
subtle — `head.sha` for `base.sha` is a nine-character diff in a `with:` block — and a
text-level `"base.sha" in text` check would still pass if a second checkout step took the head
afterwards and won.
"""
checkouts = [s for s in _workflow_steps() if "actions/checkout" in (s.get("uses") or "")]
assert len(checkouts) == 1, (
f"expected exactly one checkout step, found {len(checkouts)} — a second checkout can "
"silently overwrite the base ref with the PR head")
with_ = checkouts[0].get("with") or {}
ref = str(with_.get("ref", ""))
assert "pull_request.base.sha" in ref, (
f"the checkout ref is {ref!r}; it must be the PR's BASE sha, so a PR cannot rewrite the "
"gate that judges it")
assert "head" not in ref, f"the checkout ref {ref!r} references the PR head"
assert with_.get("persist-credentials") is False, (
"persist-credentials must be false — nothing here pushes, and a token left in .git/config "
"is handed to every script the job runs")
def test_the_workflow_runs_the_jq_preflight_in_FLOOR_mode_only():
"""`--expect` pins an exact jq version and fails when it drifts. That is right for the advisory
`script-tests` job and catastrophic here: this workflow writes `review-verdict/h10`, a REQUIRED
check on `main`, so a pin would turn any jq upgrade on the runner into a repo-wide merge
deadlock — a required gate failing because an upstream package manager did its job.
"""
# CODE only, for the reason `_code_lines` documents: the first draft of this assertion read the
# raw text and went red on the workflow's own comment explaining why `--expect` is banned here.
code = _code_lines(WORKFLOW)
# A bare `"jq-preflight.sh" in code` is NOT enough, and cold review was right to say so: the
# path also appears in the `if [ -x ./scripts/jq-preflight.sh ]` presence guard, so deleting the
# actual invocation would leave that substring behind and the assertion green. Require a line
# that INVOKES it.
steps = [s for s in _workflow_steps() if "jq-preflight.sh" in (s.get("run") or "")]
assert steps, (
"review-verdict.yml no longer runs the jq preflight, so the version its shell gates run "
"under is unobservable again (ersatztv#648)")
invocations = [ln.strip() for ln in steps[0]["run"].splitlines()
if re.match(r"^\s*(\./)?scripts/jq-preflight\.sh(\s|$)", ln)]
assert invocations, "the jq preflight is referenced but never actually invoked"
assert "--expect" not in code, (
"review-verdict.yml must run jq-preflight.sh in floor-only mode; --expect here deadlocks "
"every merge on `main` the day the runner's jq changes")
assert all("--expect" not in ln for ln in invocations)
# --- The ENFORCED caller's contract, EXECUTED ---------------------------------------------------
#
# The structural guards above prove the workflow *calls* the shared script. They cannot prove it
# reacts correctly when the script FAILS — and that is precisely the mutation that survived the last
# round on the hook side: making the caller ignore the exit status left the entire suite green,
# because every failure path also happened to produce empty stdout. So the stub below FAILS while
# emitting a perfectly docs-only list, the one combination that redundancy cannot absorb.
#
# The step's `run:` block is extracted from the YAML and executed directly. That is a real
# behavioural test of the shipped text — not a paraphrase of it — at the cost of not exercising the
# runner's step wiring, which no local test can reach anyway.
WORKFLOW_STUB_CURL = r'''#!/usr/bin/env python3
import json, os, pathlib, sys
args = sys.argv[1:]
url = [a for a in args if a.startswith("http")][-1]
out = pathlib.Path(os.environ["STUB_DIR"])
if "-X" in args and args[args.index("-X") + 1] == "POST":
payload = args[args.index("-d") + 1]
(out / "posted.json").write_text(payload)
(out / "posted_url.txt").write_text(url)
print("{}")
sys.exit(0)
if "/status" in url:
# Configurable. Hardcoding "no verdict yet" left the ersatztv#647 emptiness guard and the
# never-overwrite short-circuit unreachable: neither could be made to fire, so mutations
# deleting them survived the whole suite.
mode = os.environ.get("STUB_STATUS_MODE", "none")
if mode.startswith("appears-on-read:"):
# A human verdict that does NOT exist at the first read and DOES exist at the re-read made
# immediately before the POST (ersatztv#706). Models a reviewer posting BLOCKED while the job
# is still enumerating — the window the first read structurally cannot see.
nth = int(mode.split(":", 1)[1])
ctr = out / "status_reads.txt"
n = int(ctr.read_text()) if ctr.exists() else 0
ctr.write_text(str(n + 1))
if n + 1 < nth:
print(json.dumps({"statuses": []}))
sys.exit(0)
print(json.dumps({"statuses": [
{"context": "review-verdict/h10", "status": "failure",
"creator": {"login": "timothy"},
"description": "Review-verdict: BLOCKED @ a9e3e23 (base: main)"}]}))
sys.exit(0)
if mode == "transport-error":
# Real `gh()` is `curl -sf`: an HTTP error exits 22 with EMPTY stdout.
sys.exit(22)
if mode == "garbage":
print("<html>502 Bad Gateway</html>")
sys.exit(0)
if mode.startswith("existing:"):
# DECOY contexts either side, because real heads carry several (a live head showed 7 rows,
# the first an unrelated CI context). Their status is deliberately `pending`, NOT `success`:
# a `success` decoy triggers the very same short-circuit as a real verdict, so dropping
# `select(.context == $c)` produced an identical outcome and the mutation survived. With
# `pending` decoys, mis-selecting means no short-circuit — the job posts, and the test sees it.
#
# The row's PROVENANCE is configurable (ersatztv#698 route 3). A status POSTed by a user
# credential carries `.creator.login` and, for a real verdict, a `Review-verdict:`
# description; one POSTed by an Actions job carries `"creator": null`. Both were measured on
# the live combined endpoint. Defaults model a HUMAN verdict, so a test that does not opt out
# exercises the never-overwrite property.
creator = os.environ.get("STUB_STATUS_CREATOR", "timothy")
desc = os.environ.get("STUB_STATUS_DESC", "Review-verdict: MERGEABLE @ a9e3e23 (base: main)")
print(json.dumps({"statuses": [
{"context": "Build & test (.NET)", "status": "pending"},
{"context": "review-verdict/h10", "status": mode.split(":", 1)[1],
"creator": ({"login": creator} if creator else None), "description": desc},
{"context": "Functional E2E", "status": "pending"}]}))
sys.exit(0)
print(json.dumps({"statuses": []}))
sys.exit(0)
print("{}")
'''
def _run_classify(tmp_path, enum_stub: str | None, author: str = "timothy",
status_mode: str = "none", jq16: bool = False,
status_creator: str | None = "timothy",
status_desc: str = "Review-verdict: MERGEABLE @ a9e3e23 (base: main)"):
"""Execute the workflow's classify `run:` block with a stubbed enumeration script.
Returns the status payload the job POSTed, or None if it posted nothing.
"""
bindir = tmp_path / "bin"; bindir.mkdir()
curl = bindir / "curl"; curl.write_text(WORKFLOW_STUB_CURL); curl.chmod(0o755)
if jq16:
# Reproduce the runner's jq 1.6 (`-e` over EMPTY input exits 0, not 4). The shim is the one
# already used for pr-changed-files.sh in this file, so its fidelity is covered by that
# suite's own verify-the-verifier test.
jq = bindir / "jq"; jq.write_text(_JQ16_SHIM); jq.chmod(0o755)
scripts = tmp_path / "scripts"; scripts.mkdir()
if enum_stub is not None:
enum = scripts / "pr-changed-files.sh"
enum.write_text(enum_stub)
enum.chmod(0o755)
env = dict(os.environ)
env["PATH"] = f"{bindir}{os.pathsep}{env['PATH']}"
env["STUB_DIR"] = str(tmp_path)
env["STUB_STATUS_MODE"] = status_mode
env["STUB_STATUS_CREATOR"] = status_creator or ""
env["STUB_STATUS_DESC"] = status_desc
env.update({
"GITEA_TOKEN": "stub",
"BASE_URL": "http://gitea.example/api/v1",
"GITEA_BASE_URL": "http://gitea.example/api/v1",
"REPO": "timothy/ersatztv",
"PR": "42",
"SHA": SHA,
"BASE_SHA": OTHER_SHA,
# The base branch the event was raised for, threaded to the enumeration (ersatztv#698).
# Omitting it is not a soft failure: `set -u` kills the step, no status is posted, and the
# absent required check blocks the merge — fail-closed, but it would take every PR with it.
"BASE_REF": "main",
"AUTHOR": author,
"PR_URL": "http://gitea.example/timothy/ersatztv/pulls/42",
})
script = tmp_path / "step.sh"
script.write_text(_classify_step()["run"])
r = subprocess.run(["bash", str(script)], cwd=tmp_path, env=env,
capture_output=True, text=True)
# Assert the WIRING, not only the classification. The stub accepts every POST, so a status aimed
# at the wrong endpoint, sha, host or repo would otherwise leave these tests green while the real
# required check was never written.
#
# Two ways this check could disable itself, both found by cold review of an earlier draft:
# * it was guarded by `if url_file.exists()`, so deleting the recorder in the stub turned it
# into a no-op and every test stayed green — a verifier that silently opts out;
# * it compared only the URL SUFFIX, so a POST to the right path on the WRONG HOST OR REPO
# passed. Compare the whole URL against the env this job was given.
posted = tmp_path / "posted.json"
url_file = tmp_path / "posted_url.txt"
if posted.exists():
assert url_file.exists(), (
"a status was POSTed but its URL was not recorded — the wiring assertion below would "
"have silently skipped")
expected = f"{env['BASE_URL']}/repos/{env['REPO']}/statuses/{SHA}"
assert url_file.read_text().strip() == expected, (
f"the status was POSTed to {url_file.read_text().strip()!r}, expected {expected!r}")
return (json.loads(posted.read_text()) if posted.exists() else None), r
DOCS_ONLY = 'printf "docs/a.md\\ndocs/b.md\\n"\n'
def test_a_FAILING_enumeration_withholds_the_exemption_even_when_stdout_looks_docs_only(tmp_path):
"""The mutation that previously survived: ignore the exit status, trust stdout."""
posted, r = _run_classify(tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY + "exit 1\n")
assert posted is not None, f"the job posted no status at all: {r.stderr}"
assert posted["state"] == "pending", (
f"a docs-only exemption was granted from the stdout of a script that FAILED "
f"(state={posted['state']}, desc={posted['description']!r}) — the exit status is not being "
"checked")
def test_workflow_positive_control_the_same_output_with_exit_0_DOES_exempt(tmp_path):
"""Without this, the test above could pass for any unrelated reason and prove nothing."""
posted, r = _run_classify(tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY + "exit 0\n")
assert posted is not None, f"the job posted no status at all: {r.stderr}"
assert posted["state"] == "success", (
"the positive control failed, so the negative test above cannot be trusted to be measuring "
f"the exit status at all (state={posted['state']}, desc={posted['description']!r})")
def test_a_crashing_enumeration_also_withholds_the_exemption(tmp_path):
posted, _ = _run_classify(tmp_path, "#!/usr/bin/env bash\n" + DOCS_ONLY + "kill -TERM $$\n")
assert posted is not None
assert posted["state"] == "pending"
def test_a_MISSING_enumeration_script_withholds_the_exemption(tmp_path):
"""A PR whose BASE predates the script's introduction. It must post an actionable `pending`
rather than dying with no status — an absent required check blocks the merge either way, but a
stalled PR with no explanation is how a gate acquires a reputation for being flaky.
"""
posted, r = _run_classify(tmp_path, None)
assert posted is not None, f"the job posted no status at all: {r.stderr}"
assert posted["state"] == "pending"
def _emitting(*paths):
body = "".join(f'printf "%s\\n" "{p}"\n' for p in paths)
return "#!/usr/bin/env bash\n" + body + "exit 0\n"
def test_a_code_file_defeats_the_docs_only_exemption(tmp_path):
"""The workflow's own classification — deliberately NOT shared with the hook, whose allow-list
is wider because a match there falls through to a human prompt rather than a green status.
"""
posted, _ = _run_classify(tmp_path, _emitting("docs/a.md", "ErsatzTV/Program.cs"))
assert posted is not None
assert posted["state"] == "pending"
def test_a_BOT_pr_touching_a_protected_path_is_NOT_exempt(tmp_path):
"""`PROTECTED` guards the BOT exemption specifically, and mutation testing is how that got
stated correctly. The first version of this test used a docs-only+protected file list and
passed even with the `PROTECTED` clause deleted — `PROTECTED` (`.claude/ .gitea/ .husky/
scripts/ docker/ci/`) and `DOCS_ONLY` (`docs/`, root `*.md`) are DISJOINT, so on the docs-only
path that clause can never fire and the `DOCS_ONLY` check was doing all the work. The test
looked like it covered the self-exemption hole and covered nothing.
Renovate lands patch bumps unattended via Gitea's own auto-merge, so a bot PR that edits the
gate, CI, the hooks, or the scripts they call is the one path where an unreviewed change to the
merge gate could actually merge itself.
"""
posted, _ = _run_classify(
tmp_path, _emitting("ErsatzTV/Program.cs", "scripts/pr-changed-files.sh"),
author="renovate")
assert posted is not None
assert posted["state"] == "pending", (
"a bot PR editing the shared enumeration was auto-exempted — a PR that weakens the merge "
"gate must never be able to exempt itself from the merge gate")
def test_bot_positive_control_a_plain_bot_pr_IS_exempt(tmp_path):
"""Proves the test above measures `PROTECTED` and not merely 'bot PRs are never exempt'."""
posted, _ = _run_classify(
tmp_path, _emitting("Directory.Packages.props"), author="renovate")
assert posted is not None
assert posted["state"] == "success", (
"the bot exemption never fires at all, so the protected-path test above proves nothing")
# --- The STATUS READ: the ersatztv#647 guard and the never-overwrite short-circuit -------------
#
# These were unreachable until the stub's status response became configurable. Four mutations
# survived the full suite without them, including re-introducing the literal ersatztv#647 fail-open.
def test_a_transport_failure_on_the_STATUS_READ_posts_NOTHING(tmp_path):
"""`gh()` is `curl -sf`, so an HTTP error yields exit 22 and EMPTY stdout. Reading that as "no
verdict exists" would let the job post over a real human verdict. It must fail WITHOUT posting.
"""
posted, r = _run_classify(tmp_path, _emitting("docs/a.md"), status_mode="transport-error")
assert r.returncode != 0, "an unreadable status read must fail the job"
assert posted is None, "nothing may be posted when the existing verdict state is unknown"
def test_a_GARBAGE_status_response_posts_NOTHING(tmp_path):
"""A proxy error page is a 200 with a non-JSON body — not an absent verdict."""
posted, r = _run_classify(tmp_path, _emitting("docs/a.md"), status_mode="garbage")
assert r.returncode != 0
assert posted is None
# Two mutations to this read are NOT covered, and both are behaviourally equivalent rather than gaps:
# * `first` -> `last`: `select(.context == $c)` yields exactly ONE row, because the job reads the
# COMBINED status endpoint, which returns latest-per-context by contract. A fixture with two
# `review-verdict/h10` rows would be testing something the API does not produce.
# * `test_a_GARBAGE_status_response_posts_NOTHING` does not, on its own, defend the type guard:
# with the guard gone, the following `jq -r` fails on non-JSON and `set -e` kills the job anyway.
# That is fail-closed by REDUNDANCY. The guard's own behaviour is pinned by the jq-1.6 test below,
# which is where it actually matters.
@pytest.mark.parametrize("existing", ["success", "failure"])
@pytest.mark.parametrize("paths,author", [
(("ErsatzTV/Program.cs",), "timothy"), # non-exempt: only the short-circuit can stop it
(("docs/a.md",), "timothy"), # docs-only EXEMPT
(("Directory.Packages.props",), "renovate"), # bot EXEMPT
])
def test_an_existing_verdict_on_this_head_is_NEVER_overwritten(tmp_path, existing, paths, author):
"""A human verdict for this exact head may already exist — the reviewer ran
post-review-verdict.sh before this job finished, or the job re-ran. Re-posting would un-approve
a reviewed head, or (worse) approve one a human marked BLOCKED.
`failure` is the sharp case: that is a human saying NO, and an exemption posted over it would
turn a rejection into a merge.
"""
posted, r = _run_classify(tmp_path, _emitting(*paths), author=author,
status_mode=f"existing:{existing}")
assert r.returncode == 0, r.stderr
assert posted is None, (
f"overwrote an existing '{existing}' verdict on this head with {paths} as {author}")
# --- The `count -eq 0` guard, on the path where it is the ONLY guard ---------------------------
def test_an_EMPTY_enumeration_is_not_exempt_even_for_a_BOT(tmp_path):
"""The author must be a BOT for this to test anything.
With a non-bot author the blank line an empty list produces already fails `DOCS_ONLY`, so the
`[ "${count:-0}" -eq 0 ]` guard never decides the outcome — the same short-circuit that made an
earlier `PROTECTED` test vacuous. On the bot path that guard is the ONLY thing between an
unreadable-but-successful enumeration and an unattended `success`.
Verified by mutation: changing `grep -c .` to `grep -c ''` (counting the blank line, so
`count=1`) grants a bot PR `success` here while every other test stays green.
"""
posted, r = _run_classify(tmp_path, "#!/usr/bin/env bash\nexit 0\n", author="renovate")
assert posted is not None, f"the job posted no status at all: {r.stderr}"
assert posted["state"] == "pending", (
"an empty file list was treated as a bot exemption — the enumeration returning nothing is "
"not evidence that nothing was changed")
# --- Anchors in the classifier predicates ------------------------------------------------------
def test_the_BOT_match_is_whole_line_not_substring(tmp_path):
"""`grep -qxF` is anchored; plain `grep -qF` would exempt any author whose name CONTAINS a bot
name. `ova` is a substring of `renovate`."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"), author="ova")
assert posted is not None
assert posted["state"] == "pending", "a substring of a bot name was granted the bot exemption"
def test_DOCS_ONLY_anchors_the_markdown_extension(tmp_path):
r"""`[^/]*\.md$` must match only a top-level file ENDING in .md. Losing the `$` exempts
`evil.mdx`, and the docs-only exemption posts a green status with nobody in the loop."""
posted, _ = _run_classify(tmp_path, _emitting("evil.mdx"))
assert posted is not None
assert posted["state"] == "pending", "a .mdx file was accepted as docs-only"
def test_a_transport_failure_under_jq_1_6_STILL_posts_nothing(tmp_path):
"""The ersatztv#647 fail-open, tested where it actually lives: jq 1.6.
The shell emptiness check exists because `jq -e` over EMPTY input exits 4 on jq >= 1.7 but **0 on
jq 1.6, which is what the runner ships**. Remove that check and, on a dev Mac's 1.8, the guard
still fires and every test stays green — the bug is invisible locally.
An earlier version of this pinned the construct STRUCTURALLY instead, on the stated grounds that
"no behavioural test can catch this on a dev machine". That was wrong: this file already imports
`_JQ16_SHIM` for `pr-changed-files.sh`, so the runner's quirk is reproducible here. The structural
version was also weaker than it looked — it stripped only FULL-LINE comments, so leaving the
literal as a trailing comment on the surviving `if` satisfied it while the real guard was gone.
This test is strictly stronger: it catches that mutant, needs no comment-stripping, and fails for
the right reason. Verified by mutation under both jq versions.
"""
posted, r = _run_classify(tmp_path, _emitting("docs/a.md"),
status_mode="transport-error", jq16=True)
assert r.returncode != 0, (
"under jq 1.6 an unreadable status read must still fail the job — this is the exact "
"ersatztv#647 fail-open")
assert posted is None, "nothing may be posted when the existing verdict state is unknown"
def test_DOCS_ONLY_anchors_the_START_of_the_path_too(tmp_path):
r"""`^(docs/|...)` must match only at the start. Losing the `^` exempts `ErsatzTV/docs/Evil.cs`,
which is a C# file — fail-OPEN, and the sibling of the `$` case above."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/docs/Evil.cs"))
assert posted is not None
assert posted["state"] == "pending", "a path merely CONTAINING docs/ was accepted as docs-only"
# --- Route 2: a bot ACCOUNT does not attribute the CODE (ersatztv#698) --------------------------
#
# `AUTHOR` is `pull_request.user.login` — the PR's CREATOR, which is immutable. The head a PR points
# at is not. Force-push application code onto an open Renovate branch and the PR is still authored by
# `renovate`, still touches no protected path, and was exempted. Checking the PUSHER instead does not
# help: a git author/committer is self-asserted text. So the exemption is constrained by what a
# dependency bump can legitimately BE.
#
# The allow-list is measured, not guessed: across all 11 Renovate PRs this repo has ever had, the
# paths touched were `Directory.Packages.props` (10) and `.config/dotnet-tools.json` (1). The one
# historical outlier, PR #20, touched a `.csproj` AND two C# files — and received an unattended bot
# exemption for a source change.
def test_a_BOT_pr_carrying_a_CODE_file_is_NOT_exempt(tmp_path):
"""Route 2, stated as a behaviour: the hijacked-branch case."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"), author="renovate")
assert posted is not None
assert posted["state"] == "pending", (
"a PR authored by the renovate account was exempted while changing a C# file — the bot "
"identity is being read as attribution for code it did not write")
def test_a_BOT_pr_mixing_a_manifest_WITH_code_is_NOT_exempt(tmp_path):
"""The realistic shape of the attack: keep the manifest edit so the PR still looks like a bump,
and smuggle the code alongside it. A rule that asked 'does it touch a manifest' rather than 'is
EVERY path a manifest' would exempt this."""
posted, _ = _run_classify(
tmp_path, _emitting("Directory.Packages.props", "ErsatzTV/Program.cs"), author="renovate")
assert posted is not None
assert posted["state"] == "pending", (
"a manifest edit was enough to carry a C# file through the bot exemption — the allow-list "
"is being applied as 'any' rather than 'all'")
def test_bot_positive_control_a_REAL_dependency_bump_IS_still_exempt(tmp_path):
"""Renovate uses platformAutomerge, so breaking this deadlocks every dependency PR. The negative
tests above prove nothing if the exemption no longer works at all."""
posted, _ = _run_classify(tmp_path, _emitting("Directory.Packages.props"), author="renovate")
assert posted is not None
assert posted["state"] == "success", (
f"a plain dependency bump lost its exemption ({posted['description']!r}) — every Renovate PR "
"would now stall waiting on a human verdict")
@pytest.mark.parametrize("path", ["Directory.Packages.props", ".config/dotnet-tools.json"])
def test_every_manifest_in_the_allow_list_is_actually_exempt(tmp_path, path):
"""Pin the SET, not one example of it. A test that only ever exercises Directory.Packages.props
cannot see a typo in any of the other three alternations."""
posted, _ = _run_classify(tmp_path, _emitting(path), author="renovate")
assert posted is not None, path
assert posted["state"] == "success", f"{path} is in the allow-list but was not exempted"
def test_the_manifest_allow_list_is_ANCHORED(tmp_path):
"""Unanchored, `Directory.Packages.props` would match `evil/Directory.Packages.props.cs`. Same
class as the two DOCS_ONLY anchoring tests above, which is why it is tested the same way."""
posted, _ = _run_classify(tmp_path, _emitting("evil/Directory.Packages.props.cs"),
author="renovate")
assert posted is not None
assert posted["state"] == "pending", "the manifest allow-list matched mid-path"
def test_a_BOT_docs_only_pr_is_STILL_exempt_via_the_docs_rule(tmp_path):
"""The regression the restructure exists to prevent.
Written as an `elif` chain, a Renovate PR touching only `docs/` enters the bot branch, fails the
manifest test, and never reaches the docs-only branch — silently withdrawing an exemption the
docs-only rule grants on its own merits for ANY author. The predicates are therefore evaluated
independently and the decision made afterwards.
"""
posted, _ = _run_classify(tmp_path, _emitting("docs/note.md"), author="renovate")
assert posted is not None
assert posted["state"] == "success", (
"a docs-only PR lost its docs-only exemption merely because its author is a bot — the "
"exemptions are chained rather than composed")
def test_a_BOT_pr_touching_a_protected_path_is_still_NOT_exempt(tmp_path):
"""PROTECTED must keep outranking both exemptions, including the manifest one."""
posted, _ = _run_classify(tmp_path, _emitting("Directory.Packages.props", "scripts/evil.sh"),
author="renovate")
assert posted is not None
assert posted["state"] == "pending"
# --- Route 3: an inherited `success` is re-derived unless it is an attributable human verdict ----
#
# The short-circuit used to exit on ANY existing `success`, so an exemption this job wrote was
# indistinguishable from a verdict a human wrote. Obtained once, a forged success was thereafter
# accepted unchanged on every later run, because the guard exited before looking at the PR, the base,
# the author or the files.
#
# MEASURED on Gitea 1.25.4, on the COMBINED endpoint this job reads: a status POSTed with a user
# credential carries `.creator.login` (`timothy`), one POSTed by an Actions job carries
# `"creator": null`. Both halves are required, and the test is written in the POSITIVE direction —
# short-circuit only on something identified as a human verdict — so an unrecognised shape is
# re-derived rather than trusted.
def test_a_MACHINE_written_exemption_success_is_RE_DERIVED_not_inherited(tmp_path):
"""Route 3. The status looks exactly like one this job writes: creator null, `Exempt:` wording.
The PR now carries a C# file, so re-deriving must downgrade it to `pending`. Inheriting it would
leave a forged exemption standing forever.
"""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator=None,
status_desc="Exempt: docs-only change (no code, no protected path)")
assert posted is not None, (
"an existing machine-written success was inherited unchanged — this is route 3, and it is "
"how a forgery obtained once survives every subsequent run")
assert posted["state"] == "pending"
def test_a_success_with_NO_creator_but_a_VERDICT_LOOKING_description_is_re_derived(tmp_path):
"""Isolates the creator half. Both conditions are required; either alone is forgeable by the
other party. If a future Gitea populates `creator` for Actions, the description half still
fails — the guard degrades toward re-deriving, never toward trusting."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator=None,
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: main)")
assert posted is not None, "a creatorless status was trusted on the strength of its description"
assert posted["state"] == "pending"
def test_a_success_with_a_creator_but_an_EXEMPT_description_is_re_derived(tmp_path):
"""Isolates the description half, the mirror of the test above."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator="timothy",
status_desc="Exempt: docs-only change (no code, no protected path)")
assert posted is not None, "a status was trusted on the strength of its creator alone"
assert posted["state"] == "pending"
def test_an_UNRECOGNISED_success_shape_is_re_derived_rather_than_trusted(tmp_path):
"""The direction-of-test check. Written as 'skip if it looks machine-written', anything novel
would fall through to TRUSTED. Written as 'skip only if positively identified as human', novel
shapes are re-derived. This test fails under the first spelling and passes under the second."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator="", status_desc="")
assert posted is not None
assert posted["state"] == "pending"
@pytest.mark.parametrize("existing", ["success", "failure"])
def test_a_REAL_human_verdict_is_still_NEVER_overwritten(tmp_path, existing):
"""The property the short-circuit exists for, which the fix must not break.
`failure` is the sharp case: that is a human saying NO, and an exemption posted over it would
turn a rejection into a merge. Deliberately paired with a docs-only file list, so the job WOULD
have posted an exemption `success` had it not short-circuited — without that, a passing test
would prove only that nothing was posted for some unrelated reason.
"""
posted, r = _run_classify(tmp_path, _emitting("docs/a.md"),
status_mode=f"existing:{existing}",
status_creator="timothy",
status_desc=f"Review-verdict: MERGEABLE @ a9e3e23 (base: main)")
assert r.returncode == 0, r.stderr
assert posted is None, (
f"overwrote a human '{existing}' verdict written by timothy — the never-overwrite property "
"has been lost while fixing route 3")
def test_a_PENDING_status_from_a_previous_run_is_replaced_normally(tmp_path):
"""`pending` was never short-circuited and must stay that way, or a PR that becomes exempt after
an earlier pending run could never reach `success`."""
posted, _ = _run_classify(tmp_path, _emitting("docs/a.md"), status_mode="existing:pending",
status_creator=None, status_desc="Awaiting review verdict for a9e3e23")
assert posted is not None
assert posted["state"] == "success"
@pytest.mark.parametrize("path", ["web/package.json", "web/package-lock.json"])
def test_the_npm_manifests_are_NOT_exempt(tmp_path, path):
"""Cross-family review called an earlier draft's inclusion of these a Blocker, correctly.
`renovate.json` sets `enabledManagers: ["nuget", "github-actions", "dockerfile"]`, so Renovate does
not manage npm here at all — the entry bought nothing. Meanwhile `package.json` carries `scripts`
that CI EXECUTES (`npm ci`, `npm run build`), so exempting it lets a hijacked bot branch run
arbitrary shell in CI while every path still matches a "manifest" allow-list. Widening an exemption
to a code-execution vector for no operational benefit is strictly worse than the hole being closed.
"""
posted, _ = _run_classify(tmp_path, _emitting(path), author="renovate")
assert posted is not None
assert posted["state"] == "pending", (
f"{path} was exempted; package.json scripts execute in CI, and npm is not a managed ecosystem "
"in this repo")
# --- The `grep -q` + `pipefail` inversion (found by cross-family review of this PR) ---------------
#
# `grep -q` exits at its FIRST match, so the upstream writer takes SIGPIPE (141) once the path list
# exceeds the pipe buffer. Under `set -o pipefail` the pipeline is then a FAILURE even though grep
# MATCHED, inverting every guard built on `printf … | grep -q`. Reproduced at 171KB / 1901 paths,
# comfortably inside the enumerator's 2000-file cap.
#
# These are the regression guards, and they are the tests the original round did not have: every
# earlier case used a handful of short paths, far below the buffer, so the whole class was invisible.
# The construct PREDATES #698, so `main` carried this hole with no retarget or bot account required.
def _many_docs(n=1900):
return [f"docs/{'d' * 40}-{i:040d}.md" for i in range(n)]
def test_a_LARGE_pr_with_a_code_file_is_not_classified_docs_only(tmp_path):
"""The code file goes FIRST so `grep -qv` matches immediately and the writer is left with ~171KB
still to push — the exact shape that produced exit 141 and `docs_only=yes`."""
posted, r = _run_classify(tmp_path, _emitting("A.cs", *_many_docs()))
assert posted is not None, f"the job posted nothing: {r.stderr[-2000:]}"
assert posted["state"] == "pending", (
"a PR containing A.cs was exempted as docs-only because the classification pipeline was "
f"inverted by SIGPIPE on a large file list (desc={posted['description']!r})")
def test_a_LARGE_pr_touching_a_PROTECTED_path_still_voids_the_exemptions(tmp_path):
"""The worse direction: here the inversion makes the PROTECTED guard MISS, so a PR editing the
gate's own workflow falls through to the docs-only exemption and self-exempts."""
posted, r = _run_classify(tmp_path, _emitting(".gitea/workflows/review-verdict.yml", *_many_docs()))
assert posted is not None, f"the job posted nothing: {r.stderr[-2000:]}"
assert posted["state"] == "pending", (
"a PR editing .gitea/workflows/ was exempted — the protected-path guard was skipped on a "
f"large file list (desc={posted['description']!r})")
def test_a_LARGE_bot_pr_with_a_code_file_is_not_manifest_exempt(tmp_path):
"""The same inversion reached through the third guard added by this change."""
posted, r = _run_classify(tmp_path, _emitting("A.cs", *_many_docs()), author="renovate")
assert posted is not None, f"the job posted nothing: {r.stderr[-2000:]}"
assert posted["state"] == "pending"
def test_positive_control_a_LARGE_genuinely_docs_only_pr_IS_still_exempt(tmp_path):
"""Without this the three tests above could pass because large lists now fail outright, which
would be a merge deadlock rather than a fix."""
posted, r = _run_classify(tmp_path, _emitting(*_many_docs()))
assert posted is not None, f"the job posted nothing: {r.stderr[-2000:]}"
assert posted["state"] == "success", (
f"a large but genuinely docs-only PR lost its exemption (desc={posted['description']!r})")
def test_a_human_verdict_landing_MID_RUN_is_not_overwritten(tmp_path):
"""ersatztv#706, the half that IS mitigated here.
The first status read finds nothing, so the job proceeds to classify. A reviewer then posts a human
`failure` (BLOCKED). Without the re-read immediately before the POST, this docs-only PR would post
an exemption `success` OVER an explicit human rejection — turning a "no" into a merge, which is the
worst outcome in this class.
Note precisely what this does and does not prove: it pins the re-read, not atomicity. A verdict
landing between the re-read and the POST is still lost — there is no compare-and-set on Gitea's
status API. That remainder is #706, not a claim made here.
"""
posted, r = _run_classify(tmp_path, _emitting("docs/a.md"),
status_mode="appears-on-read:2")
assert r.returncode == 0, r.stderr
assert posted is None, (
"an exemption success was posted over a human BLOCKED verdict that landed while the job was "
"classifying — the pre-POST re-read is missing or ineffective")
def test_positive_control_no_late_verdict_still_posts_normally(tmp_path):
"""Without this, the test above would pass against a job that had simply stopped posting."""
posted, _ = _run_classify(tmp_path, _emitting("docs/a.md"), status_mode="none")
assert posted is not None and posted["state"] == "success"
# --- The PROTECTED branch must actually EXECUTE, not merely coincide with the right answer --------
#
# Round-3 review found `count_matching` being called before its definition, so it was
# `command not found` on every run and the PROTECTED branch never fired. Three "protected path" tests
# passed anyway, because a protected path is also not a manifest and not docs-only, so the job reached
# `pending` down a different route. Asserting the STATE could not see it; the guard was dead and the
# suite was green.
#
# The lesson generalises: when several branches produce the same outcome, asserting the outcome cannot
# tell you which branch ran. Assert the DISCRIMINATOR — here the reason string the branch writes.
def test_a_protected_path_is_rejected_BY_THE_PROTECTED_BRANCH(tmp_path):
posted, r = _run_classify(tmp_path, _emitting("scripts/evil.sh", "docs/a.md"))
assert posted is not None, f"the job posted nothing: {r.stderr[-2000:]}"
assert posted["state"] == "pending"
# The DISCRIMINATOR is the job's `Decision:` line, not the status description: for `pending` the
# description is always "Awaiting review verdict for <sha>", identical no matter which branch
# produced it. A first draft of this test asserted on the description and failed against a WORKING
# guard — the assertion has to be aimed at something that actually differs per branch.
assert "protected" in r.stdout.lower(), (
"the PR was not exempted, but NOT via the protected-path branch — it reached the same verdict "
f"by another route, so that guard may be dead. Decision log:\n{r.stdout[-800:]}")
def test_a_protected_path_defeats_the_BOT_exemption_by_the_protected_branch(tmp_path):
"""A bot PR whose every path IS a manifest, plus one protected path. Without a live PROTECTED
branch this still lands on `pending` (the protected file is not a manifest), so again only the
reason string distinguishes a working guard from a dead one."""
posted, r = _run_classify(
tmp_path, _emitting("Directory.Packages.props", ".gitea/workflows/renovate.yml"),
author="renovate")
assert posted is not None
assert posted["state"] == "pending"
assert "protected" in r.stdout.lower(), (
f"the bot exemption was refused, but not by the PROTECTED branch:\n{r.stdout[-800:]}")
@pytest.mark.parametrize("paths,author", [
(("docs/a.md",), "timothy"),
(("Directory.Packages.props",), "renovate"),
(("ErsatzTV/Program.cs",), "timothy"),
(("scripts/evil.sh",), "timothy"),
])
def test_the_classify_step_runs_without_SHELL_ERRORS(tmp_path, paths, author):
"""A cheap, general trap-catcher for the whole step.
`command not found`, `integer expression expected`, `unbound variable` — each of these silently
skips a branch inside an `if`/`elif` (the condition just evaluates false) while the job exits 0 and
posts a plausible status. That is precisely how the dead PROTECTED guard survived a green suite.
Scope, stated so this is not mistaken for more than it is: it catches guards that die NOISILY on
stderr. It is NOT a general liveness check — a clean mutation such as hardcoding `n_protected=0`
emits none of these and passes here. That case is covered by the branch-discriminator test above,
which is the actual liveness guard; this one is the cheap net for the whole error-emitting family.
"""
_, r = _run_classify(tmp_path, _emitting(*paths), author=author)
bad = [ln for ln in r.stderr.splitlines()
if "command not found" in ln
or "integer expression expected" in ln
or "unbound variable" in ln
or "syntax error" in ln]
assert not bad, f"the classification step emitted shell errors, so a guard is not running: {bad}"
def test_a_human_verdict_formed_against_ANOTHER_BASE_is_not_inherited(tmp_path):
"""Round-4 review: the sha-binding is escapable through the HUMAN verdict path.
Get a genuine `success` on head H while it targets scratch base S (benign diff there), then
retarget H onto `main`, where its diff contains unreviewed code. Creator is real, prefix is real,
so the short-circuit preserved it — a green required check over code nobody reviewed. The
merge-consent hook compares the base, but that is advisory; a merge through the Gitea UI or API
only sees the status.
`post-review-verdict.sh` already records the base it reviewed (`(base: …)`, ersatztv#632); this
asserts the gate actually READS it.
"""
posted, r = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: probe/scratch-base)")
assert posted is not None, (
"a human verdict formed against a DIFFERENT base was inherited unchanged — the reviewed diff "
f"is not this PR's diff. Log:\n{r.stdout[-600:]}")
assert posted["state"] == "pending"
def test_a_human_verdict_for_THIS_base_is_still_honoured(tmp_path):
"""The positive control the test above needs: matching bases must still short-circuit, or the
check has simply broken every verdict."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: main)")
assert posted is None, "a verdict formed against THIS base was re-derived; the base check is too strict"
def test_a_LEGACY_verdict_with_no_recorded_base_is_still_honoured(tmp_path):
"""Verdicts predating ersatztv#632 carry no `(base: …)`. Absent is deliberately not treated as a
mismatch: re-deriving over one would un-approve a genuinely reviewed head. Only a base that is
PRESENT and DIFFERENT is rejected."""
posted, _ = _run_classify(tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success",
status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23")
assert posted is None
def test_an_APPENDED_base_cannot_override_the_real_one(tmp_path):
"""The description is attacker-influencable by anyone who can POST a status (#697), so the parse
must not be trickable into reading a second, appended base.
This defeated the greedy `##` parse. The implementation no longer parses at all — it requires the
description to END with the exact literal marker AND to contain exactly one marker — so a second
appended marker makes the count 2 and is rejected outright.
"""
posted, r = _run_classify(
tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success", status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: probe/scratch) (base: main)")
assert posted is not None, (
"an appended '(base: main)' overrode the real recorded base, so a verdict formed elsewhere was "
f"inherited. Log:\n{r.stdout[-600:]}")
assert posted["state"] == "pending"
def test_an_EMPTY_recorded_base_is_treated_as_a_mismatch(tmp_path):
"""`(base: )` is not 'absent' — it is present and not equal to the PR's base, so it fails closed
rather than being waved through by the legacy-verdict allowance."""
posted, _ = _run_classify(
tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success", status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: )")
assert posted is not None and posted["state"] == "pending"
def test_a_branch_name_containing_a_PAREN_cannot_truncate_into_the_current_base(tmp_path):
"""Round-5 review, and the sharpest finding of the five: it needs no forgery and no #697.
`main)evil` is a VALID git branch name (`git check-ref-format --branch 'main)evil'` succeeds). A
genuine verdict earned while head H targeted it is written `(base: main)evil)`. Any implementation
that extracts the value by truncating at the first `)` gets exactly `main`, matches a PR that now
targets `main`, and inherits a verdict covering a completely different diff.
An earlier comment in the workflow asserted that a `)` in a branch name "mismatches — safe
direction". That was generalised from `feat/foo)bar` (which does mismatch) and is false for every
branch whose name STARTS with the target base. Hence the rule the code now follows: compare against
the exact expected literal, never parse a value out of attacker- or user-influenced text.
"""
posted, r = _run_classify(
tmp_path, _emitting("ErsatzTV/Program.cs"),
status_mode="existing:success", status_creator="timothy",
status_desc="Review-verdict: MERGEABLE @ a9e3e23 (base: main)evil)")
assert posted is not None, (
"a verdict recorded against branch 'main)evil' was inherited by a PR targeting 'main' — the "
f"base value is being truncated at the first ')'. Log:\n{r.stdout[-600:]}")
assert posted["state"] == "pending"