Files
ersatztv/scripts/tests/test_decisions_validate.py
T
timothyandClaude Opus 5 80818aa294
PR Gates / CI image pin matches docker/ci (pull_request) Successful in 33s
PR Gates / Docs update reminder (pull_request) Successful in 31s
PR Gates / decisions lifecycle (pull_request) Successful in 41s
Review verdict / Set review-verdict status (pull_request_target) Successful in 12s
Build ErsatzTV Image / Formatting (changed .cs conform to .editorconfig) (pull_request) Successful in 1m14s
Build ErsatzTV Image / API docs in sync (OpenAPI + endpoint index) (pull_request) Successful in 1m32s
PR Gates / Script tests (pytest) (pull_request) Successful in 2m8s
Build ErsatzTV Image / Build & test (.NET) (pull_request) Successful in 8m53s
Build ErsatzTV Image / Functional E2E (curl + UI contracts) (pull_request) Successful in 20m30s
Build ErsatzTV Image / EF migration integrity (SQLite + MySql) (pull_request) Successful in 23m19s
Build ErsatzTV Image / Build & push image (amd64) (pull_request) Has been skipped
fix(674,688): address independent review — restore the coarse bound's teeth
Two independent cold reviews (one cross-family) agreed on the top two findings.

1. The #688 fix was defeated by its own complement test. test_main_is_QUIET_about_
   drift asserted the drift notice was ABSENT while running main() over the LIVE
   corpus — whose failure condition is bit-for-bit v4's assertion, in the same
   blocking job, three functions down. p90 sat exactly on 60, so one over-ceiling
   record would have reddened it. Replaced with an IFF test that uses
   ceiling_calibration as its oracle, so it asserts the WIRING rather than the
   corpus's current state, plus a guard that at least one branch fires.

2. The coarse bound was nearly unfalsifiable. `0 < fraction_over < 1/3` accepted
   EVERY ceiling from 39 to 229 on the real corpus — including the ceiling of 200
   my own docstring offered as the case it catches, because one 230-line record
   keeps the count nonzero. That claim was simply false and is corrected. The floor
   is now a FRACTION (2%) and the cap 25%, which rejects 200/229/230 and 20, and
   accepts roughly 45..150. Headroom measured, not estimated: 37 consecutive
   over-ceiling additions, against ONE record to break v4.

3. yaml.safe_load raises a bare ValueError, not a YAMLError, on a well-shaped but
   impossible date (stale-after: 2026-06-31), which escaped as a traceback and
   killed the validator on any machine with PyYAML. The except is now deliberately
   broad, with a test.

4. PyYAML returns TYPED mapping keys, so a stray `1: x` made sorted(set|set) raise
   TypeError. Sorted with key=str.

5. The headroom prose was arithmetically wrong (~42/~40 where the real values are
   63/64; each addition moves numerator AND denominator) and the record counts were
   stale. Corrected against measurement.

6. test_adding_ordinary_records passed identically with its two additions removed.
   It now asserts the additions were counted, and that they break the v4 property
   while leaving v5 satisfied — guarded by `if`, never asserted, since whether v4
   currently holds is a fact about the live distribution and asserting it would
   rebuild the ratchet.

Also recorded honestly in docs.frontmatter-pyyaml-crosscheck: decisions-guard
installs no PyYAML, so in CI the cross-check always skips and script-tests already
caught both hazards — the CI delta is close to zero and the real fix is the local
loop plus the tool/suite agreement. And the new record was trimmed 64 -> 58 prose
lines: at 64 it moved p90 to 64 by itself, i.e. this PR would have reddened the old
blocking job. That is now cited in the test as the live demonstration.

Verification: 426 scripts/tests pass; ruff at baseline parity (47); validator exits
0; corpus back to p90=60, 18/183 over the ceiling, marks_tail and flags_minority
both true.

Decisions-Edit: yes
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:37:27 +02:00

1354 lines
61 KiB
Python

import subprocess
from pathlib import Path
from typing import Any
import pytest
import scripts.decisions_lib as dl
import scripts.decisions_validate as dv
def _rec(**kw: Any) -> dl.Record:
base: dict[str, Any] = dict(
heading="H",
source=Path("x"),
lineno=1,
status="active",
since="2026-01-01",
supersedes="none",
superseded_by="none",
signals="concept · paths: a/b.py · issues: #1",
)
base.update(kw)
return dl.Record(**base)
def _v(recs: list[dl.Record], **kw: Any) -> list[str]:
args: dict[str, Any] = dict(
archive_keys=set(),
catalog_ok=True,
removed=[],
rewritten=[],
archive_records=[],
demoted=[],
)
args.update(kw)
return dv.validate(recs, **args)
def test_two_active_same_key_fails():
assert any("more than one active" in e for e in _v([_rec(key="a.b"), _rec(key="a.b")]))
def test_bad_key_format_fails():
assert any("key format" in e for e in _v([_rec(key="BadKey")]))
def test_dangling_superseded_by_fails():
recs = [_rec(key="a.b", status="superseded", superseded_by="a.c@2026-07-01")]
assert any("superseded-by" in e and "a.c" in e for e in _v(recs))
def test_superseded_by_resolves_to_archive_key_passes():
# successor lives in archive → known via archive_keys, no dangling error
recs = [_rec(key="a.b", status="superseded", superseded_by="a.c@2026-07-01")]
assert not any("superseded-by" in e for e in _v(recs, archive_keys={"a.c"}))
def test_reciprocal_superseded_by_with_no_back_link_fails():
# B exists (has a record) but its `supersedes` does not point back to A → dangling back-link.
a = _rec(heading="A", key="a.b", status="superseded", superseded_by="a.c@2026-07-01")
b = _rec(heading="B", key="a.c", status="active", supersedes="none")
assert any("does not point back" in e and "a.c" in e for e in _v([a], archive_records=[b], archive_keys={"a.c"}))
def test_reciprocal_supersession_correct_pair_passes():
# active supersedes archived; archived is superseded-by active → reciprocal, no errors.
active = _rec(heading="Active", key="a.b", status="active", supersedes="a.c@2026-01-01")
archived = _rec(heading="Archived", key="a.c", status="superseded", superseded_by="a.b@2026-07-01")
errs = _v([active], archive_records=[archived], archive_keys={"a.c"})
assert errs == []
def test_archive_to_archive_reciprocal_pair_passes():
# A twice-reversed decision: both records live in the archive wing, pointing at each other.
a = _rec(heading="A", key="a.b", status="superseded", superseded_by="a.c@2026-07-01")
b = _rec(heading="B", key="a.c", status="retired", supersedes="a.b@2026-07-01")
errs = _v([], archive_records=[a, b], archive_keys={"a.b", "a.c"})
assert errs == []
def test_archive_to_archive_broken_back_link_fails():
# A.superseded-by=B but B.supersedes=none → the back-link is broken, and neither side is active
# (so this could previously only be caught by iterating archive_records, not decision_recs).
a = _rec(heading="A", key="a.b", status="superseded", superseded_by="a.c@2026-07-01")
b = _rec(heading="B", key="a.c", status="retired", supersedes="none")
errs = _v([], archive_records=[a, b], archive_keys={"a.b", "a.c"})
assert any("does not point back" in e and "a.c" in e for e in errs)
def test_active_record_with_superseded_by_fails():
recs = [_rec(key="a.b", status="active", superseded_by="a.c@2026-07-01")]
assert any("active record cannot already be superseded" in e for e in _v(recs, archive_keys={"a.c"}))
def test_removed_active_not_in_archive_fails():
assert any("removed from the active set" in e for e in _v([], removed=["2026-01-01 — Gone (#9)"]))
def test_rewritten_rationale_without_token_fails():
assert any(
"rationale" in e and "Decisions-Edit" in e
for e in _v([_rec(key="a.b")], rewritten=["2026-01-01 — Reworded (#9)"])
)
def test_clean_corpus_passes():
assert _v([_rec(key="a.b"), _rec(key="c.d")]) == []
def test_superseded_record_in_active_fails():
recs = [_rec(key="a.b", status="superseded", superseded_by="a.c@2026-01-01")]
assert any("relocate to docs/decisions/archive/" in e for e in _v(recs, archive_keys={"a.c"}))
def test_active_record_in_archive_fails():
assert any(
"must not live under docs/decisions/archive/" in e
for e in _v([], archive_records=[_rec(key="a.b", status="active")])
)
def test_missing_since_fails():
rec = dl.Record(
heading="H",
source=Path("x"),
lineno=1,
key="a.b",
status="active",
since=None,
supersedes="none",
superseded_by="none",
)
assert any("missing required metadata since" in e for e in _v([rec]))
def test_missing_signals_fails():
# `**Signals:**` is required (ersatztv#545): it is what MemPalace's keyword recall matches on, so
# a record without it ingests with weak metadata and under-surfaces ("no convention exists").
assert any("missing required metadata signals" in e for e in _v([_rec(key="a.b", signals=None)]))
def test_empty_signals_fails():
assert any("missing required metadata signals" in e for e in _v([_rec(key="a.b", signals="")]))
def test_signals_present_passes():
assert not any("signals" in e for e in _v([_rec(key="a.b", signals="concept · paths: x.py · issues: #1")]))
def test_demoted_heading_fails():
assert any(
"demoted to legacy-unmigrated" in e for e in _v([_rec(key="a.b")], demoted=["2026-01-01 — Some decision (#1)"])
)
def test_duplicate_metadata_block_fails():
# A migration bug left records with two stacked metadata blocks; the parser only reads the
# first, so only an explicit body scan (metadata_line_count) can catch the leftover second block.
body = (
"`key: a.b` · `status: active` · `since: 2026-01-01` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** first block rule.\n"
"\n"
"`key: a.b` · `status: active` · `since: 2026-01-01` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** second block rule.\n"
"\n"
"Rationale prose goes here."
)
rec = _rec(key="a.b", body=body)
assert dl.metadata_line_count(rec) == 2
assert any("duplicate metadata block" in e and "2 metadata blocks" in e for e in _v([rec]))
def test_single_metadata_block_passes():
body = (
"`key: a.b` · `status: active` · `since: 2026-01-01` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** the only rule.\n"
"\n"
"Rationale prose goes here."
)
rec = _rec(key="a.b", body=body)
assert dl.metadata_line_count(rec) == 1
assert not any("duplicate metadata block" in e for e in _v([rec]))
def test_contents_heading_is_skipped():
# docs/decisions/*.md topic files use "## Contents" as their index heading (the analog of
# decisions.md's "## Index") — it carries no metadata and must not be miscounted as a
# legacy-unmigrated record, same as "Index" already isn't.
assert "Contents" in dv.SKIP_HEADINGS
rec = dl.Record(
heading="Contents",
source=Path("x"),
lineno=1,
status="legacy-unmigrated",
since=None,
supersedes=None,
superseded_by=None,
)
# validate() itself excludes SKIP_HEADINGS from decision_recs, so a bare Contents record
# produces no errors (it isn't checked for required metadata, bad status, etc.)
assert _v([rec]) == []
# mirrors main()'s legacy-unmigrated count filter
unmigrated = [r for r in [rec] if r.status == "legacy-unmigrated" and r.heading not in dv.SKIP_HEADINGS]
assert unmigrated == []
def _git(cwd: Path, *args: str) -> None:
subprocess.run(["git", *args], cwd=cwd, check=True, capture_output=True, text=True)
def _write_decisions(path: Path, *, status: str, since: str, rationale: str) -> None:
path.write_text(
"## 2026-01-01 — Some decision (#1)\n"
f"`key: a.b` · `status: {status}` · `since: {since}` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** one-line current rule.\n"
"**Signals:** concept · paths: a/b.py · issues: #1\n"
"**Mechanics:** docs/foo.md\n"
"\n"
f"{rationale}\n",
encoding="utf-8",
)
def test_diff_engine_detects_rationale_rewrite_without_token(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
# (a) rewrite rationale prose only, no token
_write_decisions(
repo / "docs" / "decisions.md",
status="active",
since="2026-01-01",
rationale="**Rule:** this looks like metadata but is prose appended later.",
)
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "rewrite rationale")
removed, rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
assert removed == []
assert "2026-01-01 — Some decision (#1)" in rewritten
def _rewrite_repo(tmp_path: Path, monkeypatch, *commit_msg_paragraphs: str) -> tuple[list[str], list[str]]:
"""Seed a repo, rewrite the record's rationale under `commit_msg_paragraphs`, return (removed, rewritten).
Each paragraph is a separate `-m`, which is how git composes a message with a trailer block: the
LAST paragraph is the only one git parses for trailers (verified against git 2.55).
"""
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Reworded prose.")
_git(repo, "add", "-A")
_git(repo, "commit", *[a for p in commit_msg_paragraphs for a in ("-m", p)])
removed, rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
return removed, rewritten
_RECORD = "2026-01-01 — Some decision (#1)"
def test_diff_engine_allows_rewrite_with_trailer(tmp_path, monkeypatch):
removed, rewritten = _rewrite_repo(tmp_path, monkeypatch, "rewrite rationale", "Decisions-Edit: yes")
assert removed == []
assert rewritten == []
def test_diff_engine_trailer_is_case_insensitive(tmp_path, monkeypatch):
_removed, rewritten = _rewrite_repo(tmp_path, monkeypatch, "rewrite rationale", "Decisions-Edit: YES")
assert rewritten == []
def test_diff_engine_trailer_coexists_with_coauthor(tmp_path, monkeypatch):
# the repo's `commit-msg` hook mandates a Co-Authored-By trailer, so the real-world shape of an
# armed commit is a multi-trailer block — the marker must survive sharing it.
_removed, rewritten = _rewrite_repo(
tmp_path,
monkeypatch,
"rewrite rationale",
"Decisions-Edit: yes\nCo-Authored-By: Someone <s@example.com>",
)
assert rewritten == []
def test_diff_engine_prose_mention_of_retired_token_does_not_arm(tmp_path, monkeypatch):
"""ersatztv#609's live failure: a message *describing* the marker disarmed the whole body-diff.
This is the negative control the bare-substring matcher could not pass — under it, `token` went
True and all three rewrite comparisons were skipped while CI still reported green.
"""
_removed, rewritten = _rewrite_repo(
tmp_path,
monkeypatch,
"docs: explain the guard",
"No [decisions-edit] marker is needed here, because this commit only moves a record.",
)
assert _RECORD in rewritten
def test_diff_engine_quoted_trailer_mid_body_does_not_arm(tmp_path, monkeypatch):
"""The same class one level up: a commit that QUOTES an armed message as an example.
Git parses trailers only in the final paragraph, so a `Decisions-Edit: yes` line followed by more
prose is documentation, not a marker. Docs commits in this repo routinely carry such examples.
"""
_removed, rewritten = _rewrite_repo(
tmp_path,
monkeypatch,
"docs: document how to arm the guard",
"Contributors write:",
"Decisions-Edit: yes",
"...as the final paragraph of the message, alongside Co-Authored-By.",
)
assert _RECORD in rewritten
def test_diff_engine_negative_trailer_value_does_not_arm(tmp_path, monkeypatch):
# `Decisions-Edit: no` records a deliberate NON-edit; reading it as consent would recreate #609.
_removed, rewritten = _rewrite_repo(tmp_path, monkeypatch, "rewrite rationale", "Decisions-Edit: no")
assert _RECORD in rewritten
def test_diff_engine_folded_trailer_continuation_does_not_arm(tmp_path, monkeypatch):
"""A folded value must be judged whole: `no` + continuation ` yes` is the value "no yes".
Without `unfold`, git emits the continuation as its own line and ` yes`.strip() arms the
exemption on its own — inverting the author's explicit `no`.
"""
_removed, rewritten = _rewrite_repo(tmp_path, monkeypatch, "rewrite rationale", "Decisions-Edit: no\n yes")
assert _RECORD in rewritten
def test_diff_engine_merge_commit_body_does_not_arm(tmp_path, monkeypatch):
# On a pull_request event, actions/checkout lands on a synthetic merge commit whose body the
# forge composes (quoting the PR description) — a trailer parsed out of it is not an author's
# deliberate marker, so merge commits are excluded from the range.
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
base = subprocess.run(
["git", "rev-parse", "HEAD"], cwd=repo, check=True, capture_output=True, text=True
).stdout.strip()
_git(repo, "checkout", "-q", "-b", "feature")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Reworded prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "rewrite rationale") # the author never armed anything
# the forge's synthetic merge, whose body quotes a PR description ending in an example marker
_git(repo, "checkout", "-q", base)
_git(repo, "merge", "--no-ff", "-q", "feature", "-m", "Merge pull request #1", "-m", "Decisions-Edit: yes")
removed, rewritten, *_rest = dv._diff_findings(base, "HEAD")
assert removed == []
assert _RECORD in rewritten
def test_diff_engine_retired_bracket_token_no_longer_arms(tmp_path, monkeypatch):
# the pre-#609 habit (token appended to the subject) is retired, not silently honored.
_removed, rewritten = _rewrite_repo(tmp_path, monkeypatch, "rewrite rationale [decisions-edit]")
assert _RECORD in rewritten
def test_diff_engine_detects_appended_smuggled_rationale(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
# append-shape exploit: original prose line is PRESERVED, a smuggled **Rule:** paragraph is
# appended after a blank line, no edit marker on the commit. The old buggy _rationale() strips
# ANY `**Rule:**`/`**Signals:**`/`**Mechanics:**`/`` `key:`` line wherever it appears, so it
# silently strips the appended line too → base==head → bypass succeeds. The bounded strip only
# removes the contiguous top metadata block, so the appended paragraph survives as prose →
# base!=head → flagged.
_write_decisions(
repo / "docs" / "decisions.md",
status="active",
since="2026-01-01",
rationale="Original prose.\n\n**Rule:** smuggled rewrite that changes the actual meaning.",
)
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "append smuggled rewrite")
removed, rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
assert removed == []
assert "2026-01-01 — Some decision (#1)" in rewritten
def test_diff_engine_metadata_only_edit_is_free(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
_write_decisions(repo / "docs" / "decisions.md", status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
# metadata-only edit: status/since change, prose unchanged, no token
_write_decisions(
repo / "docs" / "decisions.md", status="superseded", since="2026-02-01", rationale="Original prose."
)
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "metadata edit")
removed, rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
assert removed == []
assert rewritten == []
def _write_archive(path: Path, *, rationale: str) -> None:
path.write_text(
"## 2026-01-01 — Some archived decision (#2)\n"
"`key: a.z` · `status: superseded` · `since: 2026-01-01` · `supersedes: none` · "
"`superseded-by: a.b@2026-07-01`\n"
"**Rule:** one-line historical rule.\n"
"**Signals:** concept · paths: a/z.py · issues: #2\n"
"**Mechanics:** docs/foo.md\n"
"\n"
f"{rationale}\n",
encoding="utf-8",
)
def test_diff_engine_detects_archive_rationale_rewrite_without_token(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs" / "decisions" / "archive").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
archive_path = repo / "docs" / "decisions" / "archive" / "x.md"
_write_archive(archive_path, rationale="Original archived prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
_write_archive(archive_path, rationale="Rewritten archived prose, no token.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "rewrite archived rationale")
_removed, rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
assert "2026-01-01 — Some archived decision (#2)" in rewritten
def test_diff_engine_detects_archive_record_removed(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs" / "decisions" / "archive").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
archive_path = repo / "docs" / "decisions" / "archive" / "x.md"
_write_archive(archive_path, rationale="Original archived prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
archive_path.unlink()
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "delete archived record")
removed, _rewritten, *_rest = dv._diff_findings("HEAD~1", "HEAD")
assert "2026-01-01 — Some archived decision (#2)" in removed
def test_diff_engine_demoted_migrated_to_legacy_unmigrated(tmp_path, monkeypatch):
repo = tmp_path / "repo"
(repo / "docs").mkdir(parents=True)
_git(tmp_path, "init", str(repo))
_git(repo, "config", "user.email", "t@example.com")
_git(repo, "config", "user.name", "T")
decisions_path = repo / "docs" / "decisions.md"
_write_decisions(decisions_path, status="active", since="2026-01-01", rationale="Original prose.")
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "base")
monkeypatch.chdir(repo)
# same heading, metadata block stripped: prose survives, migration silently reverted.
decisions_path.write_text(
"## 2026-01-01 — Some decision (#1)\n\nOriginal prose.\n",
encoding="utf-8",
)
_git(repo, "add", "-A")
_git(repo, "commit", "-m", "demote to legacy-unmigrated")
_removed, _rewritten, demoted = dv._diff_findings("HEAD~1", "HEAD")
assert "2026-01-01 — Some decision (#1)" in demoted
# ---- optional stale-after / Sources metadata (ersatztv#603) ----
def test_stale_after_absent_is_not_an_error():
assert _v([_rec(key="a.b")]) == []
def test_wellformed_stale_after_is_not_an_error():
assert _v([_rec(key="a.b", stale_after="2027-01-01")]) == []
def test_malformed_stale_after_fails():
for bad in ("2027-13-01", "01-01-2027", "soon", "2027-02-30"):
errs = _v([_rec(key="a.b", stale_after=bad)])
assert any("not a YYYY-MM-DD date" in e for e in errs), bad
def test_past_due_record_does_not_fail_validation():
"""A record going stale is the passage of time, not a defect in the commit under test."""
errs = _v([_rec(key="a.b", stale_after="2020-01-01")])
assert errs == []
def test_stale_records_reports_only_past_due_active():
today = dv.date(2026, 7, 25)
recs = [
_rec(key="past.due", heading="Past", stale_after="2026-07-24"),
_rec(key="due.today", heading="Today", stale_after="2026-07-25"), # today >= date → stale
_rec(key="not.yet", heading="Future", stale_after="2026-07-26"),
_rec(key="no.date", heading="Undated"),
_rec(key="arch.ived", heading="Archived", status="superseded", stale_after="2020-01-01"),
]
assert dv.stale_records(recs, today) == [("Past", "2026-07-24"), ("Today", "2026-07-25")]
def test_stale_records_ignores_malformed_date():
"""A malformed date is caught as a blocking format error; it must not also crash the notice path."""
assert dv.stale_records([_rec(key="a.b", stale_after="whenever")], dv.date(2026, 7, 25)) == []
def test_sources_is_optional_and_never_required():
assert _v([_rec(key="a.b")]) == []
assert _v([_rec(key="a.b", sources="ci run 1234 · docs/ci-cd.md")]) == []
def test_empty_stale_after_is_malformed_not_absent():
"""A blank/truncated value must not pass as 'absent' — it would silently never fire."""
errs = _v([_rec(key="a.b", stale_after="")])
assert any("not a YYYY-MM-DD date" in e for e in errs)
def test_noncanonical_iso_forms_are_rejected():
"""date.fromisoformat accepts these on py>=3.11; the field's contract is YYYY-MM-DD only.
Accepting them would also break the catalog's Review-due sort, which orders on the raw string.
"""
for bad in ("20270101", "2027-W01-1", "2027-01-01T00:00:00", " 2027-01-01 extra"):
errs = _v([_rec(key="a.b", stale_after=bad)])
assert any("not a YYYY-MM-DD date" in e for e in errs), bad
def test_malformed_stale_after_on_an_archive_record_is_caught():
arch = _rec(key="a.old", status="superseded", superseded_by="a.b", stale_after="soon")
errs = _v([_rec(key="a.b", supersedes="a.old")], archive_records=[arch], archive_keys={"a.old"})
assert any("not a YYYY-MM-DD date" in e for e in errs)
# ---- path <-> key correspondence and the filesystem one-active-per-key property (#610) ----
def test_path_key_mismatch_is_an_error(tmp_path, monkeypatch):
"""The filename IS the key. If they can drift, one-active-per-key stops being structural."""
records_dir = tmp_path / "docs" / "decisions" / "records"
monkeypatch.setattr(dv.dl, "RECORDS_DIR", records_dir)
monkeypatch.setattr(dv.dl, "ARCHIVE_DIR", tmp_path / "docs" / "decisions" / "archive")
wrong = _rec(key="ci.runner-placement", source=records_dir / "api" / "runner-placement.md")
assert any("does not match its path" in e for e in _v([wrong]))
right = _rec(key="ci.runner-placement", source=records_dir / "ci" / "runner-placement.md")
assert not [e for e in _v([right]) if "does not match its path" in e]
def test_legacy_multi_record_files_are_exempt_from_the_path_rule(tmp_path, monkeypatch):
"""A record still living in a multi-record file has no key-derived path to match."""
monkeypatch.setattr(dv.dl, "RECORDS_DIR", tmp_path / "records")
monkeypatch.setattr(dv.dl, "ARCHIVE_DIR", tmp_path / "archive")
legacy = _rec(key="ci.runner-placement", source=Path("docs/decisions.md"))
assert not [e for e in _v([legacy]) if "does not match its path" in e]
def test_one_active_per_key_is_unrepresentable_on_disk():
"""The old duplicate-key defect cannot be expressed in the split layout.
Two active records sharing a key would have to occupy the same path, so the filesystem refuses
it — this is the invariant moving from 'validator catches it' to 'cannot happen'. Asserted via
the derivation the migration uses, so a change to the path scheme breaks this test.
"""
import scripts.migrate_decisions_split as mig
a = _rec(key="ci.runner-placement", heading="One")
b = _rec(key="ci.runner-placement", heading="Two — a different title, same key")
assert mig.record_path(a) == mig.record_path(b), "same key must derive the same path"
def test_migration_aborts_on_a_destination_collision(tmp_path, monkeypatch):
"""Two records deriving the same path must abort, never silently overwrite one with the other."""
import scripts.migrate_decisions_split as mig
src = tmp_path / "docs" / "decisions.md"
src.parent.mkdir(parents=True)
src.write_text(
"# D\n\n"
"## 2026-01-01 — One (#1)\n"
"`key: ci.dup` · `status: active` · `since: 2026-01-01` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** r\n**Signals:** s\n\nbody one.\n\n"
"## 2026-01-02 — Two (#2)\n"
"`key: ci.dup` · `status: active` · `since: 2026-01-02` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** r\n**Signals:** s\n\nbody two.\n"
)
monkeypatch.setattr(dv.dl, "REPO_ROOT", tmp_path)
monkeypatch.setattr(dv.dl, "DECISIONS_MD", src)
monkeypatch.setattr(dv.dl, "TOPIC_DIR", tmp_path / "docs" / "decisions")
monkeypatch.setattr(dv.dl, "ARCHIVE_DIR", tmp_path / "docs" / "decisions" / "archive")
monkeypatch.setattr(dv.dl, "RECORDS_DIR", tmp_path / "docs" / "decisions" / "records")
with pytest.raises(SystemExit, match="collision"):
mig.main([])
def test_retitling_a_record_is_not_reported_as_a_removal(tmp_path):
"""The heading-rename trap (#610 box 2): renaming a heading used to read as an unlogged removal.
Matching on `key` makes a retitle a retitle. Driven through the real git-backed diff engine,
not the pure helpers, because the defect lived in how records were collected.
"""
import os
import subprocess
def g(*a, **kw):
return subprocess.run(["git", *a], cwd=tmp_path, capture_output=True, text=True, **kw)
rec = (
"# D\n\n## {title}\n"
"`key: ci.thing` · `status: active` · `since: 2026-01-01` · `supersedes: none` · `superseded-by: none`\n"
"**Rule:** r\n**Signals:** s\n\nUnchanged rationale prose.\n"
)
docs = tmp_path / "docs"
docs.mkdir()
g("init", "-q", ".")
g("config", "user.email", "t@e")
g("config", "user.name", "t")
(docs / "decisions.md").write_text(rec.format(title="2026-01-01 — Original title (#1)"))
g("add", "-A")
g("commit", "-qm", "base")
base = g("rev-parse", "HEAD").stdout.strip()
# retitle ONLY — same key, byte-identical rationale
(docs / "decisions.md").write_text(rec.format(title="2026-01-01 — Clearer, better title (#1)"))
g("add", "-A")
g("commit", "-qm", "docs: retitle the record")
head = g("rev-parse", "HEAD").stdout.strip()
assert head != base
cwd = os.getcwd()
os.chdir(tmp_path)
try:
removed, rewritten, demoted = dv._diff_findings(base, head)
finally:
os.chdir(cwd)
assert removed == [], f"a retitle was reported as a removal: {removed}"
assert rewritten == [], f"a retitle was reported as a prose rewrite: {rewritten}"
assert demoted == [], f"a retitle was reported as a demotion: {demoted}"
# --- #621: structural "one keyed record per file" over the record wings -----------------------
#
# The defect these cover: a file the dependency-free frontmatter reader cannot parse yields `[]`
# and vanishes from the corpus with NO error anywhere — validator OK, catalog "up to date", record
# absent. Only a NEWLY ADDED record is affected; an existing one disappearing is already caught by
# the no-vanish diff check, which structurally cannot see a record that never existed in the base.
def _wing(tmp_path) -> tuple[Path, Path]:
"""A tmp record-wing pair: (records_dir, archive_dir)."""
records = tmp_path / "docs" / "decisions" / "records"
archive = tmp_path / "docs" / "decisions" / "archive"
(records / "ci").mkdir(parents=True)
(archive / "ci").mkdir(parents=True)
return records, archive
_GOOD = (
"---\n"
"key: ci.good\n"
"title: '2026-01-01 — Good (#1)'\n"
"status: active\n"
"since: '2026-01-01'\n"
"supersedes: none\n"
"superseded-by: none\n"
"rule: 'a rule on one quoted line'\n"
"signals: 'concept · paths: a/b.py · issues: #1'\n"
"mechanics: 'x'\n"
"---\n\nRationale prose.\n"
)
def test_wing_faults_clean_corpus_is_silent(tmp_path):
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
assert dv.record_wing_faults(records, archive) == []
def test_wing_faults_block_scalar_record_fails_loudly(tmp_path):
"""The exact reproduction from #621: a YAML block scalar — the natural thing to reach for on
this corpus's very long `rule:` values — makes the whole record silently invisible."""
records, archive = _wing(tmp_path)
bad = records / "ci" / "blockscalar.md"
bad.write_text(_GOOD.replace("rule: 'a rule on one quoted line'\n", "rule: >-\n a long rule wrapped\n over two lines\n"))
# Precondition: this really is the silent-vanish case, not some other parse error.
assert dl.parse_file(bad) == [], "expected the reader to drop the record entirely"
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1, faults
assert "blockscalar.md" in faults[0]
assert "expected exactly 1" in faults[0]
def test_wing_faults_unterminated_frontmatter_fails_loudly(tmp_path):
records, archive = _wing(tmp_path)
bad = records / "ci" / "unterminated.md"
bad.write_text("---\nkey: ci.unterminated\nstatus: active\n\nNo closing delimiter.\n")
assert dl.parse_file(bad) == []
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1, faults
assert "unterminated.md" in faults[0]
def test_wing_faults_stray_note_file_fails_loudly(tmp_path):
"""A file under records/ that is not a record at all — no frontmatter whatsoever."""
records, archive = _wing(tmp_path)
(records / "ci" / "notes.md").write_text("# Scratch notes\n\nNot a record.\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1, faults
assert "notes.md" in faults[0]
def test_wing_faults_keyless_record_fails_loudly(tmp_path):
"""Parses to exactly one record, but carries no `key` — still a fault."""
records, archive = _wing(tmp_path)
(records / "ci" / "keyless.md").write_text(_GOOD.replace("key: ci.good\n", ""))
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1, faults
assert "no `key`" in faults[0]
def test_wing_faults_exempts_stripped_legacy_archive_files(tmp_path):
"""The #610 split left five KEYLESS stripped topic files at the archive TOP level (api.md,
scan.md, spa.md, startup.md, release-ci-governance.md) whose only content is a generated
"Records formerly in this file" index — that index is what keeps older date-based pointers
resolvable. They are indexes, not records, and must stay exempt. Their area-NESTED siblings
are real records and are checked."""
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
# a stripped legacy file at the archive top level: parses to a single keyless pseudo-record
(archive / "api.md").write_text("# api\n\n## Records formerly in this file\n\n- `api.thing`\n")
(archive / "README.md").write_text("# archive\n")
assert dv.record_wing_faults(records, archive) == []
# ...but a real archived record nested under an area IS checked
(archive / "ci" / "broken.md").write_text("# not a record\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "broken.md" in faults[0], faults
def test_validate_surfaces_wing_faults_as_errors(tmp_path):
"""Faults must arrive as validator ERRORS (exit 1), not warnings."""
errs = _v([_rec(key="ci.a", source=Path("docs/decisions/records/ci/a.md"), heading="A")],
wing_faults=["docs/decisions/records/ci/x.md: parsed to 0 records, expected exactly 1"])
assert any("x.md" in e for e in errs), errs
def test_real_repo_record_wings_are_all_parseable():
"""Positive control against the LIVE corpus: every one of the real record files parses to
exactly one keyed record. This is what makes the check's clean result meaningful rather than
vacuous — if the wings were empty or unreadable, the checks above would pass trivially."""
files = dv.record_wing_files()
assert len(files) > 100, f"record wings look empty ({len(files)} files) — check is vacuous"
assert dv.record_wing_faults() == []
def test_wing_faults_sees_a_DEEPER_nested_archive_record(tmp_path):
"""A one-level `archive/*/*.md` glob would exempt the top-level stripped files correctly but
silently skip anything nested deeper — a path escaping the check, which is the very failure
mode this guard exists to close. The exemption must be 'directly in archive/', not 'exactly
one level down'."""
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
(archive / "api.md").write_text("# api\n\n## Records formerly in this file\n") # still exempt
deep = archive / "ci" / "sub"
deep.mkdir(parents=True)
(deep / "broken.md").write_text("# not a record\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "broken.md" in faults[0], faults
# --- #621 cold-review findings ------------------------------------------------------------------
def test_main_actually_CALLS_the_wing_scan(monkeypatch, capsys):
"""The guard's only wiring was untested: deleting `wing_faults=record_wing_faults()` from
main() left the entire suite green, and a real block-scalar record vanished again with
`decisions-validate: OK`. Every other test either calls the collector directly or hands
validate() a hand-built list, so nothing pinned that main() invokes it at all — the #609
'marker that printed OK while doing nothing' defect, one level up."""
monkeypatch.setattr(dv, "record_wing_faults", lambda *a, **k: ["SENTINEL-WING-FAULT"])
rc = dv.main([])
assert rc == 1, "a wing fault must fail the validator"
assert "SENTINEL-WING-FAULT" in capsys.readouterr().err
def test_a_record_is_not_exempted_by_its_BASENAME(tmp_path):
"""`_NON_DECISION_FILES` is a set of TOPIC-dir names ({README, migration-map, retrieval-eval}).
Applying it to the wings meant a genuine record at `records/docs/retrieval-eval.md` was
silently skipped — and that path is forced, not hypothetical: the path<->key rule puts key
`docs.retrieval-eval` at exactly that filename. Worse, dl.active_files() applies the filter
only to the TOPIC_DIR glob, so such a file IS a corpus source while being exempt from the
guard."""
records, archive = _wing(tmp_path)
(records / "docs").mkdir(parents=True, exist_ok=True)
(records / "ci" / "good.md").write_text(_GOOD)
(records / "docs" / "retrieval-eval.md").write_text("# just a note, not a record\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "retrieval-eval.md" in faults[0], faults
def test_archive_toplevel_is_exempt_by_IDENTITY_not_by_location(tmp_path):
"""Exempting everything directly in `archive/` left that one directory unguarded: a new
unparseable `archive/foo.md` would vanish silently. The exemption is for the five #610 stripped
INDEX files, so it must test for that shape — one keyless record with a known generated
heading — not merely for sitting in that directory."""
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
(archive / "api.md").write_text("# api\n\n## Records formerly in this file\n\n- `api.thing`\n")
assert dv.record_wing_faults(records, archive) == [], "a real stripped index must stay exempt"
(archive / "foo.md").write_text("rule: >-\n wrapped\n value\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "foo.md" in faults[0], faults
def test_junk_frontmatter_key_from_a_split_value_is_faulted(tmp_path):
"""parse-to-WRONG, the case the structural check alone cannot see. A block scalar with an
UNINDENTED continuation containing a colon parses to ONE valid keyed record whose `rule` is
literally `>-`, plus a junk key from the continuation — silently truncating the real value.
PyYAML rejects this input, so the hand reader is more permissive than the writer."""
records, archive = _wing(tmp_path)
bad = records / "ci" / "corrupt.md"
bad.write_text(_GOOD.replace("rule: 'a rule on one quoted line'\n",
"rule: >-\nthe real rule: with a colon\n"))
recs = dl.parse_file(bad)
assert len(recs) == 1 and recs[0].key, "precondition: this parses to one KEYED record"
assert recs[0].rule == ">-", f"precondition: the real value was truncated, got {recs[0].rule!r}"
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "unrecognized key" in faults[0], faults
def test_an_empty_or_missing_record_wing_is_LOUD(tmp_path):
"""A guard that reports a clean corpus from a scan of nothing is the defect it exists to close,
applied to itself. Reachable via a partial checkout, a renamed directory, or a bad monkeypatch."""
missing = dv.record_wing_faults(tmp_path / "nope" / "records", tmp_path / "nope" / "archive")
assert missing and "missing or contains no" in missing[0], missing
records, archive = _wing(tmp_path) # exists but holds no *.md
empty = dv.record_wing_faults(records, archive)
assert empty and "missing or contains no" in empty[0], empty
def test_a_wing_root_README_is_exempt_by_PATH_not_by_basename(tmp_path):
"""`docs/decisions/archive/README.md` really exists (a hand-written directory README), so it
must be exempt — but by exact relative path, not by basename.
An earlier version excluded ANY wing-root `README.md` on the stated grounds that no such file
existed. That was false, and it would additionally have exempted a future
`records/README.md` — reintroducing the basename hole one directory over, in the wing that
matters most."""
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
(archive / "README.md").write_text("# archive\n\nHand-written directory README.\n")
assert dv.record_wing_faults(records, archive) == [], "archive/README.md must stay exempt"
(records / "README.md").write_text("# not a record\n")
faults = dv.record_wing_faults(records, archive)
assert len(faults) == 1 and "README.md" in faults[0], (
f"a README in the ACTIVE wing must NOT inherit the archive exemption: {faults}"
)
def test_an_empty_ARCHIVE_wing_is_legitimate(tmp_path):
"""Deliberate asymmetry with the active-wing check, pinned so nobody 'fixes' it into symmetry.
Zero active records means the scan measured nothing and any clean result is vacuous. Zero
ARCHIVED records just means nothing has been superseded yet — normal for a young repo and for
every fresh clone before the first supersession."""
records, archive = _wing(tmp_path)
(records / "ci" / "good.md").write_text(_GOOD)
assert dv.record_wing_faults(records, archive) == [], "an empty archive must not fault"
# --- #620: per-record ceiling replaces the aggregate budget --------------------------------------
_MINIMAL_RECORD = (
"---\n"
"key: ci.a\n"
"title: '2026-01-01 — A (#1)'\n"
"status: active\n"
"since: '2026-01-01'\n"
"supersedes: none\n"
"superseded-by: none\n"
"rule: 'r'\n"
"signals: 's'\n"
"---\n\nprose.\n"
)
def _rec_body(key: str, lines: int) -> dl.Record:
return _rec(key=key, heading=key, body="\n".join(f"prose line {i}" for i in range(lines)))
def test_oversized_records_flags_only_those_over_the_ceiling():
recs = [_rec_body("a.short", 10), _rec_body("b.exact", 60), _rec_body("c.long", 61)]
assert dv.oversized_records(recs, 60) == [("c.long", 61)], "the ceiling must be exclusive"
def test_oversized_records_sorts_longest_first():
recs = [_rec_body("a.mid", 80), _rec_body("b.big", 200), _rec_body("c.small", 70)]
assert [k for k, _ in dv.oversized_records(recs, 60)] == ["b.big", "a.mid", "c.small"]
def test_oversized_records_can_go_green():
"""The property the aggregate budget structurally lacked: a corpus that satisfies it.
An aggregate over a monotonically growing corpus can only ratchet — it goes red and stays red
until someone raises the number. A per-record ceiling is satisfiable and re-satisfiable, which
is what makes a red mean something."""
recs = [_rec_body("a.short", 10), _rec_body("b.short", 20)]
assert dv.oversized_records(recs, 60) == []
def test_budget_total_excludes_the_generated_catalog(tmp_path, monkeypatch):
"""#620: the catalog gains one row per record and cannot be consolidated away.
Counting it made the 'schedule a consolidation' number partly a record COUNT, pointing the
reader at work that cannot be done. Mutation-sensitive by construction: the catalog here is
large and distinctively sized, so re-adding it would change the total by exactly 500."""
topic = tmp_path / "docs" / "decisions"
records = topic / "records" / "ci"
records.mkdir(parents=True)
(records / "a.md").write_text(_MINIMAL_RECORD)
(topic / "README.md").write_text("\n".join(f"| row {i} |" for i in range(500)))
monkeypatch.setattr(dv.dl, "REPO_ROOT", tmp_path)
monkeypatch.setattr(dv.dl, "DECISIONS_MD", tmp_path / "docs" / "decisions.md")
monkeypatch.setattr(dv.dl, "TOPIC_DIR", topic)
monkeypatch.setattr(dv.dl, "ARCHIVE_DIR", topic / "archive")
monkeypatch.setattr(dv.dl, "RECORDS_DIR", records)
total = dv._budget_total()
assert total < 100, f"the 500-line generated catalog leaked into the total ({total})"
def test_real_corpus_ceiling_flags_a_nonempty_proper_minority():
"""Guards the calibration claim. This is the FIFTH version; the failures are the lesson.
v1 `max(under) <= 60 < min(over)` — true by construction of those two lists.
v2 a minimum gap WIDTH — but a ceiling of 200 also sits in a wide gap, so it passed.
v3 a 2-12% fraction band plus "clear air" measured against `min(over)` — the nearest
record ABOVE the ceiling. That made the test a hostage to an unrelated record: one
ordinary 62-line addition reddened it with the ceiling correctly placed, and the only
remedy the assertion admitted was to RAISE the ceiling.
v4 `p90 <= ceiling <= p95`. Scale-free and correct AS A DEFINITION, but an order statistic
over a SPARSE distribution is a STEP function. At #688 the lengths above the ceiling ran
60, 61, 62, 63 and then jumped to 81, so ONE new record moved p90 by twenty lines and
reddened the BLOCKING `script-tests` job for whoever happened to write it. It reproduced
twice live (#672, #706) and both times the only in-scope remedy was to trim the new
record to fit the constant — the ratchet pointed at record authors, which is precisely
what the v3 note says this whole design abolishes.
v5 SPLITS the claim by robustness instead of hunting for a better single assertion:
* the COARSE property — the ceiling flags a nonempty proper minority — is asserted HERE,
blocking. One record moves a fraction by at most 1/N, so routine growth cannot cross it.
* the FINE property — `p90 <= ceiling <= p95` — is now REPORTED by `main()` as a notice.
It is real signal about the CONSTANT drifting out of date, which is the passage of corpus
growth rather than a defect in the commit under test. That is the same reasoning
`stale_records` is built on, and it gets the same treatment.
Note what did NOT change: the ceiling is still 60, and the fine claim is still measured on
every run. v5 moves where each claim is enforced, it does not stop making them.
"""
recs = [r for r in dl.all_active_records() if r.key]
assert len(recs) > 100, f"corpus looks empty ({len(recs)}) — this check would be vacuous"
ceiling = dv.RECORD_CEILING_DEFAULT # the value the CLI actually uses; cannot drift from here
cal = dv.ceiling_calibration(recs, ceiling)
assert cal.flags_minority, (
f"the ceiling ({ceiling}) no longer flags a nonempty proper minority of records: "
f"{cal.n_over}/{cal.n} = {cal.fraction_over:.1%} are over it. At 0% it names nobody and "
f"signals nothing; above {dv.CEILING_MINORITY_MAX:.0%} it is cutting into the bulk of the "
f"corpus rather than marking its tail. Re-derive it from the distribution."
)
def test_ceiling_calibration_detects_drift_in_BOTH_directions():
"""The fine claim is asserted here, on a distribution the test OWNS.
This is the point of the v5 split: the property is still pinned, but against synthetic data
instead of the live corpus, so it cannot be reddened by someone else's record landing.
"""
# 100 records: 95 of 20 lines, 5 of 200. Index 90 lands in the short block and index 95 in the
# long one, so p90 == 20 and p95 == 200 — a wide, unambiguous tail boundary to aim at.
recs = [_rec_body(f"a.s{i}", 20) for i in range(95)] + [_rec_body(f"a.l{i}", 200) for i in range(5)]
assert [dv.ceiling_calibration(recs, 60).p90, dv.ceiling_calibration(recs, 60).p95] == [20, 200]
assert dv.ceiling_calibration(recs, 60).marks_tail, "60 sits between p90=20 and p95=200"
assert not dv.ceiling_calibration(recs, 10).marks_tail, "below p90 it cuts into the bulk"
assert not dv.ceiling_calibration(recs, 999).marks_tail, "above p95 it is parked among outliers"
# and the coarse property separates the same two failure modes
assert not dv.ceiling_calibration(recs, 999).flags_minority, "a ceiling nobody is over signals nothing"
assert not dv.ceiling_calibration(recs, 10).flags_minority, "95% over the ceiling is not a tail"
assert dv.ceiling_calibration(recs, 60).flags_minority
def test_the_coarse_bound_REJECTS_a_badly_placed_ceiling_on_the_REAL_corpus():
"""The blocking property must have teeth, demonstrated against the live distribution.
Review's strongest finding on the first draft: a floor of `fraction_over > 0` was nearly
unfalsifiable — it accepted every ceiling from 39 to 229, including the ceiling of 200 that the
docstring itself offered as the case it catches, because one 230-line record keeps the count
nonzero. A FRACTION floor is what restores the teeth, so pin it on the real corpus.
"""
recs = [r for r in dl.all_active_records() if r.key]
assert len(recs) > 100, "corpus looks empty — this check would be vacuous"
for bad in (200, 229, 230):
cal = dv.ceiling_calibration(recs, bad)
assert not cal.flags_minority, (
f"a ceiling of {bad} flags only {cal.n_over}/{cal.n} records and must be rejected, got {cal}"
)
assert not dv.ceiling_calibration(recs, 20).flags_minority, "a ceiling of 20 cuts into the bulk"
assert dv.ceiling_calibration(recs, dv.RECORD_CEILING_DEFAULT).flags_minority
def test_ceiling_calibration_is_empty_safe():
"""A vacuous corpus must report both claims FALSE, never a passing default."""
cal = dv.ceiling_calibration([], 60)
assert cal.n == 0 and not cal.marks_tail and not cal.flags_minority
def test_adding_ordinary_records_cannot_RED_the_blocking_property():
"""#688's Done-when, demonstrated rather than assumed.
Takes the REAL corpus and appends the two record shapes that actually reddened v4 in the wild:
#672's 62-line record and #687's 107-line one. Under v4 the 107-line addition moved p90 into
the sparse gap above the ceiling and failed the blocking job. Under v5 the coarse property
still holds, and it holds with room to spare.
"""
real = [r for r in dl.all_active_records() if r.key]
ceiling = dv.RECORD_CEILING_DEFAULT
grown = real + [_rec_body("new.typical", 62), _rec_body("new.substantial", 107)]
before = dv.ceiling_calibration(real, ceiling)
after = dv.ceiling_calibration(grown, ceiling)
# The additions must actually MOVE something, or this test proves nothing about them — review
# caught the first draft passing identically with the two records removed.
assert after.n == before.n + 2 and after.n_over == before.n_over + 2, (
f"the appended records were not counted: {before} -> {after}"
)
# The v4 property is the one that reddened CI, and these additions break it — that is the whole
# demonstration: the SAME corpus change that would have failed the old blocking assertion leaves
# the v5 one satisfied. Guarded by `if`, never asserted: whether v4 currently holds is a fact
# about the live distribution, and asserting it here would rebuild the very ratchet this test
# exists to document. It is not hypothetical: an earlier draft of the record this branch adds
# was 64 prose lines, which by itself moved p90 from 60 to 64 and broke v4 before either
# synthetic addition was appended — i.e. this PR would have reddened the old blocking job.
if before.marks_tail:
assert not after.marks_tail, (
f"these additions no longer break the old percentile property ({after}) — pick shapes "
f"that do, or this test has stopped demonstrating anything"
)
# ...and the v5 blocking property holds either way.
assert after.flags_minority, f"two ordinary additions reddened the blocking property: {after}"
# Headroom, stated as a number so a future reader can see how much slack there is: how many
# consecutive over-ceiling records it would take to breach the bound.
n, over = after.n, after.n_over
need = 0
while (over + need) / (n + need) <= dv.CEILING_MINORITY_MAX:
need += 1
assert need > 20, f"only {need} over-ceiling additions would breach the bound — too tight"
# --- #674: the validator cross-checks its own parse against PyYAML ------------------------------
_HAZARDS = {
# PyYAML REJECTS: the bare apostrophe closes the single-quoted scalar early.
"apostrophe": "rule: 'SQLite's LOWER() folds ASCII only'",
# PyYAML ACCEPTS but reads a DIFFERENT value: ` #` starts a comment, truncating the rule.
"unquoted-hash": "rule: use --flag #2 for this",
}
def _wing_with(tmp_path: Path, frontmatter_line: str) -> tuple[Path, Path]:
"""A record wing containing one file whose frontmatter carries `frontmatter_line`."""
records = tmp_path / "records" / "ci"
records.mkdir(parents=True)
(records / "a.md").write_text(
"---\n"
"key: ci.a\n"
"title: 'T'\n"
"status: active\n"
"since: '2026-01-01'\n"
"supersedes: none\n"
"superseded-by: none\n"
f"{frontmatter_line}\n"
"signals: 's'\n"
"---\n\nprose.\n"
)
archive = tmp_path / "archive"
archive.mkdir(parents=True)
return records, archive
@pytest.mark.parametrize("hazard", sorted(_HAZARDS))
def test_pyyaml_crosscheck_catches_frontmatter_the_hand_reader_accepts(tmp_path, hazard):
"""#674, both shapes. Hit twice in one session by two independent agents (#578, #651)."""
pytest.importorskip("yaml")
records, _ = _wing_with(tmp_path, _HAZARDS[hazard])
faults, ran = dv.pyyaml_frontmatter_faults(sorted(records.rglob("*.md")))
assert ran, "PyYAML is installed here, so the cross-check must have run"
assert faults, f"the {hazard} hazard slipped through the cross-check"
@pytest.mark.parametrize("hazard", sorted(_HAZARDS))
def test_the_hand_reader_really_IS_blind_to_these(tmp_path, hazard):
"""The positive control: pin the MECHANISM, so this suite cannot pass for the wrong reason.
If `record_wing_faults` ever started catching these on its own, the cross-check above could be
deleted and the tests would stay green while the guard vanished. Asserting that the pre-#674
machinery reports these files as CLEAN is what makes the cross-check's red meaningful — and it
is the exact state #674 was filed about: `decisions_validate.py` printed OK on input the
writer's own library rejects.
"""
records, archive = _wing_with(tmp_path, _HAZARDS[hazard])
assert dv.record_wing_faults(records, archive) == [], (
"the structural guard now catches this by itself — re-derive whether the PyYAML "
"cross-check is still the thing closing this gap"
)
def test_crosscheck_REPORTS_an_impossible_date_instead_of_crashing(tmp_path):
"""PyYAML raises a bare `ValueError`, not a `YAMLError`, for a well-shaped impossible date.
`stale-after: 2026-06-31` (June has 30 days) escaped an `except yaml.YAMLError` as a traceback,
killing the validator on any machine with PyYAML — including the Husky pre-commit hook. A check
documented as "strictly additive, must never be the reason the validator cannot run" must
REPORT this, so the except is deliberately broad.
"""
pytest.importorskip("yaml")
records, _ = _wing_with(tmp_path, "stale-after: 2026-06-31")
faults, ran = dv.pyyaml_frontmatter_faults(sorted(records.rglob("*.md")))
assert ran
assert faults and "REJECTS" in faults[0], faults
assert "ValueError" in faults[0], f"the fault must name the real exception: {faults[0]}"
def test_pyyaml_crosscheck_is_clean_on_the_REAL_corpus():
"""No false positives. A cross-check that flags correct records would be reverted within a day."""
pytest.importorskip("yaml")
files = dv.record_wing_files()
assert len(files) > 100, f"only {len(files)} wing files found — this check would be near-vacuous"
faults, ran = dv.pyyaml_frontmatter_faults(files)
assert ran
assert faults == [], "the cross-check disagrees with the live corpus:\n" + "\n".join(faults[:5])
def test_crosscheck_skips_cleanly_when_pyyaml_is_absent(tmp_path, monkeypatch):
"""The read path stays dependency-free (#674's second Done-when box).
`decisions-guard`, the Husky hooks and every contributor machine install nothing, so an absent
PyYAML must SKIP the cross-check rather than fault or crash — while every other check runs.
"""
records, _ = _wing_with(tmp_path, _HAZARDS["apostrophe"])
import builtins
real_import = builtins.__import__
def no_yaml(name, *a, **kw):
if name == "yaml":
raise ImportError("no yaml here")
return real_import(name, *a, **kw)
monkeypatch.setattr(builtins, "__import__", no_yaml)
faults, ran = dv.pyyaml_frontmatter_faults(sorted(records.rglob("*.md")))
assert ran is False, "an absent PyYAML must report that the check did not run"
assert faults == [], "a skipped check must not manufacture faults"
def test_main_ANNOUNCES_a_skipped_crosscheck(capsys, monkeypatch):
"""A skipped check that says nothing is the '#603 stale-after' defect: reports success, does
nothing. The skip is legitimate; staying quiet about it is not."""
monkeypatch.setattr(dv, "pyyaml_frontmatter_faults", lambda files: ([], False))
assert dv.main([]) == 0
err = capsys.readouterr().err
assert "cross-check" in err and "SKIPPED" in err, err
def test_main_FAILS_when_the_crosscheck_reports_a_fault(capsys, monkeypatch):
"""Wiring test: the faults must reach the exit code, not just be computed.
Without this, `wing_faults=record_wing_faults() + yaml_faults` could drop the second term and
every other test here would still pass.
"""
monkeypatch.setattr(dv, "pyyaml_frontmatter_faults", lambda files: (["x.md: bogus fault"], True))
assert dv.main([]) == 1
assert "bogus fault" in capsys.readouterr().err
def test_retired_budget_flag_says_it_is_ignored(capsys):
"""A retired flag must announce itself, not no-op silently.
Accepting `--budget 100` and enforcing nothing would be the same 'reports success while doing
nothing' defect this change exists to retire (#603 stale-after, #609 marker)."""
dv.main(["--budget", "100"])
err = capsys.readouterr().err
assert "--budget 100 is RETIRED and was IGNORED" in err, err
def test_no_budget_flag_means_no_retirement_warning(capsys):
dv.main([])
assert "RETIRED" not in capsys.readouterr().err
def test_main_reports_ceiling_drift_as_a_NOTICE_and_still_exits_0(capsys):
"""The fine claim's live wiring (#688). A ceiling of 999 is parked far above p95, so the drift
notice must fire — and must NOT turn the run red, which is the entire point of the v5 split."""
assert dv.main(["--record-ceiling", "999"]) == 0
err = capsys.readouterr().err
drift = [ln for ln in err.splitlines() if "drifted from the tail boundary" in ln]
assert len(drift) == 1, err
assert drift[0].startswith("::notice::"), f"drift must be a notice, not a warning: {drift[0]}"
def test_main_reports_drift_IFF_the_ceiling_is_off_the_tail_boundary(capsys):
"""The complement of the test above — asserting the WIRING, not the corpus's current state.
The obvious way to write this is `dv.main([]); assert "drifted" not in err`, and that is a trap
review caught: `main()` emits the notice exactly when `p90 <= 60 <= p95` is false over the LIVE
corpus, so such a test fails under precisely the condition #688 exists to stop failing — it
would move v4's assertion three functions down and leave it in the same blocking job. Today p90
sits exactly ON the ceiling, so ONE new over-ceiling record would have reddened it.
So the oracle is `ceiling_calibration` itself: whatever the corpus currently looks like, the
notice must be present iff the fine claim is false. The 999 case pins that at least one branch
is genuinely exercised, so this cannot pass by never firing.
"""
recs = [r for r in dl.all_active_records() if r.key]
seen_drift = False
for ceiling in (dv.RECORD_CEILING_DEFAULT, 999):
expected = not dv.ceiling_calibration(recs, ceiling).marks_tail
dv.main(["--record-ceiling", str(ceiling)])
err = capsys.readouterr().err
assert ("drifted from the tail boundary" in err) is expected, (
f"ceiling {ceiling}: expected drift notice={expected}, got the opposite"
)
seen_drift |= expected
assert seen_drift, "neither ceiling produced a drift notice — the test never exercised the notice"
def test_main_actually_REPORTS_the_ceiling_and_the_trend(capsys):
"""The new signal's live wiring was untested: `if oversized:` -> `if False:`, or bumping the
default ceiling to 999999, left every test green while main() reported nothing. Only the pure
function `oversized_records()` was covered — so the replacement signal could silently do
nothing, which is the exact defect this change exists to retire."""
dv.main([])
err = capsys.readouterr().err
assert "prose lines across" in err, "the aggregate trend notice must always print"
assert "exceed the" in err and "prose ceiling" in err, "the per-record ceiling warning must print"
over = [r.key for r in dl.all_active_records()
if r.key and dv.record_prose_lines(r) > dv.RECORD_CEILING_DEFAULT]
assert over, "precondition: the live corpus has at least one over-ceiling record"
assert any(k in err for k in over), "the warning must NAME the offending records"
def test_trend_notice_reports_record_prose_and_scaffolding_separately(capsys):
"""The notice used to print `_budget_total()` 'across N records', mixing two incompatible
definitions of prose: the total includes ~600 lines of headings and keyless scaffolding
belonging to no record, so dividing it by the record count and comparing that to the 60-line
ceiling compares incommensurable units — a small instance of the 'the metric is not what it
says it is' defect this change indicts."""
dv.main([])
err = capsys.readouterr().err
recs = [r for r in dl.all_active_records() if r.key]
record_prose = sum(dv.record_prose_lines(r) for r in recs)
assert f"{record_prose} prose lines across {len(recs)} records" in err, err
assert "non-record scaffolding" in err, err