Files
ersatztv/scripts/tests/test_ci_toolchain_image_resolves.py
T
timothyandClaude Fable 5.1 a7d91bf15a
Build ErsatzTV Image / CI toolchain image resolves (pull_request) Successful in 35s
Build ErsatzTV Image / Delimiter ban (release path) (pull_request) Successful in 57s
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 37s
PR Gates / Docs update reminder (pull_request) Successful in 1m0s
PR Gates / decisions lifecycle (pull_request) Successful in 20s
PR Gates / Fix proofs (Proves trailers) (pull_request) Successful in 17s
review-verdict/h10 Review-verdict: MERGEABLE @ a7d91bf (base: main)
Review verdict / Set review-verdict status (pull_request_target) Successful in 45s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 9m25s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 6m17s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Skipped
PR Gates / Script lint and tests (ruff + pytest) (pull_request) Successful in 19m27s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 6m4s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 8s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 7s
fix(876): sweep session narrative out of hooks, workflows, scripts, tests and code comments; grow the detector to the process corpus
`docs.no-session-narrative` reaches every durable artifact, but its detector scanned only
`docs/**/*.md` and root markdown, and nothing had ever swept the rest. The issue named four sites
from one grep and called them a floor. Deriving the population instead — a whitespace-joined sweep
over every tracked file outside the detector, for the detector's own phrasings plus the attribution
and review-round class #812 found — gave 453 sites in 108 files at `fb5592971`, and a second pass
for phrasings the first list missed (hyphenated `round-N`, "an earlier version", "the reviewer
proved") added residuals in the same files. Every site was classified with #812's three
dispositions (CUT / SEVER / KEEP with its sub-kind) under the who-benefits test; the per-site
manifests are on the PR. The rejected designs, tested-and-rejected fixtures, measurements and
traps stay; the attribution of who found them and the round in which they were found go.

The detector's population grows to `.claude/`, `.gitea/`, `.husky/` and `scripts/` regardless
of extension, minus the detector and its own test (whose fixtures ARE the phrasings) and minus
`scripts/tests/fixtures/` (test data, including decision-record copies — the same reasoning as
the records' own exemption, and what keeps the record's depth measurement true), and `--all`
lists tracked REGULAR files only — a symlink's content is its target and a gitlink has none. The #812
argument for leaving `docs/superpowers/**` in the population runs the other way here: `--diff`
sees only ADDED lines, and 287 of the 453 sites were under 30 days old — this corpus is where
narrative is being added, so the advisory nudge has reach. Density agrees: 56 line-mode hits over
the 113 regular files the predicate admits, against 9 over 66 docs files before #812. `web/` and C# stay out on the same
measurement (3 of 74 PATTERNS-matching sites, ~4,600 files). The predicate did not grow: PATTERNS
matched 74 of 453 sites, and widening the word list to the attribution class is the treadmill
the withdrawn parity test ran on. The population oracle is restated over segments with the new
arms, the synthetic cross product gains the process heads and non-markdown extensions, a fixture
witnesses that a tracked symlink is neither scanned nor counted, a `.py.bak` axis separates a
by-name exemption from a `startswith` over the same tuple, and eight mutants (drop the process
arm, drop the by-name exemption, exempt by `startswith`, drop or add a prefix, drop the fixtures
exemption, list only markdown, drop the symlink filter, test the mode per row instead of per
path) each
redden it. A pre-existing silent drop in `--diff` goes with it: git tab-terminates a `+++`
filename that contains a space, and the kept tab made `is_scanned_path` refuse the file with no
notice — fixed, with a positive control and its own mutant.

Code is unchanged by construction, measured per file type against `origin/main`: Python modules
are AST-equal with docstrings stripped, except `#` lines inside the embedded fixture programs
(string literals) of three test modules; workflows differ only in `#` lines inside `run:` block
scalars; shell, C#, TypeScript and jq are equal with comment lines stripped. The stated
exceptions: the detector and its test, 26 vitest titles that carried review-round or severity
labels or a reviewer attribution (call sites whose title changed — every changed title line
walked back to its `it(` / `it.each(...)(` anchor, so a `' + '` concatenation counts once), two
registry note strings and the mutation manifest's prose fields. scripts/tests: 1565 passed.
Web: lint, typecheck, 1319 tests green. Closes #876.

Decisions-Edit: yes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PEcBoFw7ctrf3Nb7R7x7wk
2026-09-03 20:51:39 +02:00

304 lines
13 KiB
Python

"""Tests for `scripts/ci-toolchain-image-resolves.sh` (ersatztv#772).
The script answers one question — does the tag `docker-build.yml` pins still exist? — and the whole
value is in *which answers it refuses to round off*. A registry read has three outcomes, not two:
present, gone, and could-not-tell. Collapsing the third into either of the others is how a preflight
becomes decoration, so each is driven here through the real entry point with a stubbed `curl`.
`test_MUTATION_a_deleted_tag_is_reported_as_a_failure` is the load-bearing one and is declared in
`scripts/tests/mutation_manifest.py`. Note what it can and cannot turn on: since an unverifiable
answer fails the job too, disarming the `404` arm still exits non-zero, so the EXIT CODE separates
nothing. What the disarm destroys is the DIAGNOSTIC — the outage is reported as "could not verify",
which sends an operator to the registry's health instead of to the rebuild that fixes it.
"""
from __future__ import annotations
import os
import subprocess
import time
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parents[2]
SCRIPT = REPO_ROOT / "scripts" / "ci-toolchain-image-resolves.sh"
# Stands in for `curl -s -w '\n%{http_code}' -u <auth> -H Accept <url>`: prints a scripted body,
# a newline and the HTTP code, and logs the call. It VALIDATES `-u` rather than ignoring it — a stub
# that answers 200 whether or not the script authenticates would stay green if the real `-u` were
# deleted, which is the fidelity gap that lets a test double certify a script the live registry
# would reject on every request.
CURL_SHIM = r"""#!/usr/bin/env python3
import os, pathlib, sys
state = pathlib.Path(os.environ["STUB_DIR"])
args = sys.argv[1:]
url = [a for a in args if a.startswith("http")][-1]
tag = url.rsplit("/", 1)[-1]
auth = args[args.index("-u") + 1] if "-u" in args else ""
with (state / "calls").open("a") as fh:
fh.write(f"{url} auth={auth}\n")
# The live registry answers 401 to an anonymous read of ANY tag, present or deleted.
user, _, password = auth.partition(":")
if not user or not password:
print("{}\n401", end="")
sys.exit(0)
codes = dict(pair.split("=", 1) for pair in (state / "codes").read_text().split() if pair)
code = codes.get(tag, codes.get("*", "200"))
if code == "TRANSPORT":
# Only the EXIT STATUS is observable: the script's `|| resp=""` discards whatever curl printed,
# so what this reproduces is the non-zero exit, not the `\n000` real curl also emits.
sys.exit(7)
body = '{"schemaVersion": 2, "mediaType": "application/vnd.oci.image.manifest.v1+json"}'
if code == "200-NOT-A-MANIFEST":
code, body = "200", "<html><title>Sign in</title></html>"
print(f"{body}\n{code}", end="")
"""
WORKFLOW_TEMPLATE = """jobs:
test:
container:
image: 192.168.1.95:3000/timothy/ersatztv-ci:{pin}
"""
@pytest.fixture
def preflight(tmp_path):
bindir = tmp_path / "bin"
bindir.mkdir()
shim = bindir / "curl"
shim.write_text(CURL_SHIM)
shim.chmod(0o755)
state = tmp_path / "state"
state.mkdir()
(state / "codes").write_text("*=200")
workflow = tmp_path / "docker-build.yml"
workflow.write_text(WORKFLOW_TEMPLATE.format(pin="32747a0"))
env = dict(os.environ)
env["PATH"] = f"{bindir}{os.pathsep}{env['PATH']}"
env["STUB_DIR"] = str(state)
env["ETV_CI_WORKFLOW"] = str(workflow)
env["ETV_REGISTRY_AUTH"] = "stub-user:stub-pass"
# The retry PAUSE is what makes failing on an unknown affordable in CI and unaffordable in a
# test suite; the retry COUNT is behaviour, so it is kept and only the wait is removed.
env["ETV_CI_ATTEMPTS"] = "2"
env["ETV_CI_RETRY_SECONDS"] = "0"
class Handle:
def __init__(self):
self.env = env
self.state = state
self.workflow = workflow
self.script = SCRIPT
def set_codes(self, mapping: dict[str, str]):
(state / "codes").write_text(" ".join(f"{k}={v}" for k, v in mapping.items()))
def set_workflow_text(self, text: str):
workflow.write_text(text)
def calls(self):
log = state / "calls"
return log.read_text().splitlines() if log.exists() else []
def run(self, script: Path | None = None):
return subprocess.run(
["bash", str(script or SCRIPT)],
env=env,
capture_output=True,
text=True,
cwd=REPO_ROOT,
)
return Handle()
def test_a_pin_that_resolves_passes(preflight):
preflight.set_codes({"*": "200"})
result = preflight.run()
assert result.returncode == 0, result.stderr
assert "32747a0 resolves" in result.stdout
assert preflight.calls(), "the registry was never queried, so nothing was established"
def test_MUTATION_a_deleted_tag_is_reported_as_a_failure(preflight):
"""The outage of 2026-08-11..13, in one assertion.
Declared in `mutation_manifest.py`: replacing the `404` arm sends a deleted tag down the
could-not-verify path, which fails the job with the wrong story — a preflight that runs, reddens,
and still misses the only thing it was built to name.
"""
preflight.set_codes({"32747a0": "404"})
result = preflight.run()
assert result.returncode != 0, (
"a deleted tag did not fail the preflight — the 404 arm is not load-bearing:\n"
f"stdout={result.stdout}\nstderr={result.stderr}"
)
assert "IS GONE" in result.stderr, (
"a deleted tag was not reported as GONE — the 404 arm is not load-bearing. Since an "
"unverifiable answer now fails too, exiting non-zero no longer distinguishes 'the image is "
"deleted' from 'the check could not run', and only this message does:\n"
f"stderr={result.stderr}"
)
assert "32747a0" in result.stderr, "the message must name the tag the operator has to restore"
assert "server-management#842" in result.stderr, "and where the durable fix lives"
@pytest.mark.parametrize("code", ["TRANSPORT", "503"])
def test_an_unknown_answer_FAILS_and_is_not_reported_as_gone(preflight, code):
"""Warning and exiting 0 here is how a preflight becomes a no-op.
A missing `curl`, a moved registry or a DNS change all land in this arm, and each would have
been green forever. It fails — but with its own wording, because "could not verify" and "IS
GONE" send an operator to entirely different places.
"""
preflight.set_codes({"32747a0": code})
result = preflight.run()
assert result.returncode != 0, "an unestablished check must not report success"
assert "could NOT VERIFY" in result.stderr
assert "IS GONE" not in result.stderr, "could-not-tell must never be reported as gone"
def test_an_unknown_is_RETRIED_before_it_fails(preflight):
"""Retries are what make failing on unknown affordable rather than flaky."""
preflight.env["ETV_CI_ATTEMPTS"] = "3"
preflight.set_codes({"32747a0": "503"})
assert preflight.run().returncode != 0
assert len(preflight.calls()) == 3, f"expected 3 attempts, got {preflight.calls()}"
def test_an_ANSWER_is_not_retried(preflight):
"""404 and 200 are answers; retrying them would only slow the job down."""
preflight.set_codes({"32747a0": "404"})
assert preflight.run().returncode != 0
assert len(preflight.calls()) == 1, f"a 404 must not be retried, got {preflight.calls()}"
def test_HTTP_200_with_a_body_that_is_not_a_manifest_is_not_a_pass(preflight):
"""A proxy or a login page answers 200 too; the status line alone establishes nothing."""
preflight.set_codes({"32747a0": "200-NOT-A-MANIFEST"})
result = preflight.run()
assert result.returncode != 0
assert "not a manifest" in result.stderr
assert "IS GONE" not in result.stderr
@pytest.mark.parametrize("code", ["401", "403"])
def test_rejected_credentials_refuse_rather_than_pass(preflight, code):
"""The failure mode that would otherwise make this job green forever.
An anonymous read of this registry is 401 for a live tag and a deleted one alike, so treating
an auth failure as "could not tell, carry on" would turn a broken secret into a permanent,
silent pass.
"""
preflight.set_codes({"32747a0": code})
result = preflight.run()
assert result.returncode != 0
assert "rejected these credentials" in result.stderr
@pytest.mark.parametrize(
("value", "shape"),
[
(None, "unset"),
(":", "both secrets absent — WHAT THE WORKFLOW ACTUALLY PASSES"),
("user:", "password secret absent"),
(":pass", "user secret absent"),
("no-colon", "malformed"),
],
)
def test_unusable_credentials_refuse_BEFORE_querying_anything(preflight, value, shape):
"""The empty-halves cases are the ones that happen, and testing only `unset` misses them.
`ETV_REGISTRY_AUTH: ${{ secrets.REGISTRY_USER }}:${{ secrets.REGISTRY_PASSWORD }}` interpolates
a missing secret to the empty string, so a job with no secrets configured passes the non-empty
string ":" — which is a perfectly good non-empty value and a useless credential. The registry
answers 401 to it for a live tag and a deleted one alike.
"""
if value is None:
del preflight.env["ETV_REGISTRY_AUTH"]
else:
preflight.env["ETV_REGISTRY_AUTH"] = value
result = preflight.run()
assert result.returncode != 0, f"{shape}: reported success on a credential it cannot use"
assert "ETV_REGISTRY_AUTH" in result.stderr
assert preflight.calls() == [], "it must not query the registry it cannot authenticate to"
def test_the_credential_actually_REACHES_the_registry(preflight):
"""Anti-vacuity for every test above: the stub 401s an unauthenticated read, as the live
registry does, so a script that stopped passing `-u` would redden the whole file rather than
sailing through on a stub that answers 200 regardless."""
preflight.set_codes({"*": "200"})
assert preflight.run().returncode == 0
assert preflight.calls() == [
"http://192.168.1.95:3000/v2/timothy/ersatztv-ci/manifests/32747a0 auth=stub-user:stub-pass"
]
def test_a_workflow_with_no_pin_at_all_is_a_failure(preflight):
"""If the grep stops matching, the honest report is 'I found nothing', not 'all clear'."""
preflight.set_workflow_text("jobs:\n test:\n runs-on: ubuntu-latest\n")
result = preflight.run()
assert result.returncode != 0
assert "no ersatztv-ci pin found" in result.stderr
def test_every_distinct_pin_is_checked_and_one_gone_fails_the_job(preflight):
"""`ci-image-pin` bans a second pin; this must not silently check only the first one anyway."""
preflight.set_workflow_text(
WORKFLOW_TEMPLATE.format(pin="32747a0") + " image: 192.168.1.95:3000/timothy/ersatztv-ci:15d2439\n"
)
preflight.set_codes({"32747a0": "200", "15d2439": "404"})
result = preflight.run()
assert result.returncode != 0
assert "15d2439" in result.stderr
assert len(preflight.calls()) == 2, f"both pins must be queried, got {preflight.calls()}"
def test_the_grep_line_cannot_match_ITSELF(preflight):
"""The pin is found with the same expression `pr-checks.yml::ci-image-pin` uses.
That expression is written into this script's own source, so a careless pattern would find its
own text and 'check' a pin nobody wrote — and the same hazard sits in `pr-checks.yml`, whose
pin-count check greps the file this script's job now lives in. Feed the real script its own
source as the workflow file: the answer must be 'no pin found', not a query for `[0-9a-f]+`.
This also pins the second half of the property — the source carries no literal pin of its own,
so the file cannot go stale against a pin bump it does not participate in.
"""
preflight.set_workflow_text(SCRIPT.read_text())
result = preflight.run()
assert result.returncode != 0
assert "no ersatztv-ci pin found" in result.stderr
assert preflight.calls() == []
def test_the_PRODUCTION_retry_defaults_are_the_ones_that_run(preflight):
"""Every other test overrides the retry knobs, so nothing evaluated `${VAR:-default}` itself.
That matters because the defaults are the argument: "unknown fails" is only affordable if an
ordinary registry blip is absorbed first. Edited to 1 attempt / 0 seconds, this file would stay
green while a single transient 503 reddened every PR. So this one drops both overrides and
measures the real thing — three attempts, and a pause long enough to have actually happened.
"""
del preflight.env["ETV_CI_ATTEMPTS"]
del preflight.env["ETV_CI_RETRY_SECONDS"]
preflight.set_codes({"32747a0": "503"})
started = time.monotonic()
result = preflight.run()
elapsed = time.monotonic() - started
assert result.returncode != 0
assert len(preflight.calls()) == 3, f"the default attempt count is not 3 — got {len(preflight.calls())} call(s)"
assert elapsed >= 8, (
f"two pauses at the default 5s should clear the 8s floor; took {elapsed:.1f}s, so the pause "
"has been shortened out from under the 'a blip does not redden a PR' argument"
)