Files
ersatztv/docs/handoffs/orchestration.md
T
timothyandClaude Fable 5.1 b1d5fbefcb fix(916): cross-family state is per round, and the orchestrator reads cross_family before a verdict
A substitute that failed in round one said nothing about the tree that lands
after round two, yet the flag was sticky and doomed the run; the xfamily
string was never reset either, so clearing the stickiness alone would have let
a stale "substitute ALSO failed" sentence into the PR body. Both reset at the
top of review(). The harness runner is round-aware (ran per round, its own
counter reset) and a two-round case pins the fix; restoring the sticky flag
reddens it in both scripts. Step 4 of the mechanics page tells the referee to
read cross_family, not only error, before posting on a rubric-class PR.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QqCpYFsKgnAnx6jVwrKiV
2026-09-05 13:48:15 +02:00

149 lines
12 KiB
Markdown

# Orchestrated sessions
How one session closes several issues at once: an orchestrator that never implements, subagents
that each own one issue in one worktree, and the repo's merge gate between every branch and `main`.
This page owns the mechanics (`process.orchestrated-session`). `orchestrator-prompt.md` is the
standing prompt that starts such a session; `chicorytv-issue-queue.md` still owns queue selection,
claiming and the HARD CONSTRAINTS, and nothing here relaxes them — an orchestrated session is
several single-issue sessions run by one referee, and each of them is bound by the whole kickoff.
## Roles
| Role | Model, effort | Does |
|---|---|---|
| Orchestrator | Fable or Opus, high | Reads verdicts and evidence, ticks the `## Done-when` boxes, posts the H10 verdict, merges through the consent hook, clears labels, removes worktrees, watches for collisions and hangs. Never picks, claims, codes or pushes. |
| Picker + 2 refuters | sonnet, medium | Apply `scripts/select-queue.sh` and the kickoff's claim and bundle rules to live Gitea state; return up to `count` mutually non-colliding picks. `.claude/workflows/ersatztv-pick-next.js`, args `{taken, closed, notes, count}`. |
| Recon (large issues) | Opus, high | Reads the docs the task-signal map names, finds the handlers, call sites, tests and guards, returns a plan. |
| Implementer | Opus high; sonnet medium for `size: small` | Claims, builds, rebases onto `origin/main`, runs the local gate, sits inside the review loop, supplies the evidence per Done-when box. |
| Reviewers | correctness: Opus high, in its own worktree; conformance: sonnet high; cross-family: Codex via `codex exec` for the rubric's risk classes, with a cold Opus review-only fallback when Codex cannot run | Run the gate themselves; `blocking` / `should-fix` / `nit`. |
| Fixer | same tier as the implementer | Answers findings; two rounds, then the orchestrator decides. |
| Finisher | sonnet, medium | The single push, the PR, the closing record. Ticks nothing. |
**Sizing is a rule** (`process.per-agent-model-routing`): every launch names its model and effort.
sonnet at medium for picking, refuting, small well-specified fixes and mechanical finishing; sonnet at
high for conformance review; Opus at high for recon, implementation, fixing and correctness review;
Fable for orchestrating and for the frontier escalations the kickoff lists, never for implementing.
Cross-family review is mandatory for locks/concurrency, auth/security, API write-path handlers, DB
migrations and diffs over ~150 changed C# lines (`process.independent-review-rubric`); the picker's
`risk` field decides. When Codex cannot run, the workflow substitutes a cold same-family review-only
agent and the PR body states the substitution and its reason, as that record requires.
`.claude/workflows/ersatztv-issue-build.js` runs claim → recon → implement → gate → review → fix →
land for one issue or bundle (args: `issues, slug, title, size, risk, area, needs_e2e, port,
body_summary, done_condition, files_likely, avoid, trailer`; `issues`, `trailer` and `port` are
required and the script refuses without them). `ersatztv-resume-branch.js` finishes or fixes a
paused branch from a JSON brief. Both are passed to the Workflow tool inline; the tool persists the
script and returns a path for later runs.
## Isolation
- One worktree per issue: `~/orca/workspaces/ersatztv/wt-<n>`, branch `<n>-<slug>`, cut from a
freshly fetched `origin/main` (`process.shared-tree-readonly`), by absolute path. Never under
`/tmp` — macOS purges it mid-session — and never the shared checkout.
- The commit guard (`pretooluse-worktree-guard.sh`) reads a `.claude-worktree-owner` marker that
`posttooluse-worktree-marker.sh` writes on `git worktree add`. Measured 2026-09-04: a workflow
agent's `git worktree add` writes no marker (project hooks do not run for a workflow's agents),
and a `~`-prefixed path from the session writes none either, so orchestrated worktrees carry no
marker and the guard fails open inside them — acceptable, because they all belong to one session.
A sibling *session's* worktree keeps its marker and its denial
(`process.foreign-worktree-plumbing-merge`), so a denial inside an orchestrated worktree means the
worktree is foreign: the scripts stop and report it, and never overwrite the marker.
- `web/node_modules` is per worktree: clone it from the shared checkout with `cp -Rc` when the
lockfiles match, otherwise `npm ci`. The shared copy is kept current by
`scripts/refresh-shared-checkout.sh` at session end.
- **One .NET-building slot at a time on this host; docs and Python slots may run beside it.** A
slot's correctness reviewer builds .NET in its own worktree, and on a rubric-class change whose
Codex runner cannot run the fallback is a second such reviewer, so a review round can be two
build-and-test pipelines on top of the implementer's lingering MSBuild node servers. Measured
2026-09-05 on the 16 GB Mac: three slots in review reached load 82 and swap exhaustion; one slot's
round alone took swap from 6.8 GB to 10.8 GB in three minutes with the two reviewers concurrent.
The scripts therefore start the fallback reviewer only after the lenses return (the Codex runner
builds nothing and runs beside them). The cap stays at one even so: the implementer's own build,
one reviewer pipeline and the MSBuild node servers that linger after each build already fill the
budget beside the host's other sessions; re-measure before raising it.
`process.build-concurrency-limits` is the standing rule (under 20% free RAM launch nothing that
builds, under 10% pause); the Agent hook's RAM gate does not see a workflow's agents, so the
orchestrator holds the count itself.
- **Live-E2E runs per worktree on the slot's own port.** `scripts/e2e-local.sh` refuses concurrent
runs *within one repo root* because each run re-copies that root's `wwwroot`; across worktrees
there is nothing shared but the ports, so every slot gets a distinct `port` argument, one run at a
time inside the slot, and the launcher's pre-flight refuses a busy port and names the holder. A
busy port is reported, never taken over; kill only the PID the launcher printed to you
(`testing.e2e-cleanup-scope-by-pid`).
- Reviewers create scratch only under `/private/tmp`, delete nothing outside it, and never build a
path with `..` segments. A reviewer that must build gets `isolation: worktree` and fetches the
unpushed branch into it (`git fetch <worktree> <branch> && git checkout --detach FETCH_HEAD`); it
never builds in the slot's worktree.
## Landing a branch
The order is the kickoff's (`process.pr-routine-sequence`, `process.local-gate-before-push`), with
the review loop *inside* the worktree, before the single push:
1. **Implementer**: rebase onto a fresh `origin/main`, then the local gate — `dotnet build` and
`dotnet test` on the touched test projects (the full solution for anything under
`ErsatzTV.Core`), `npm run check:api && npm run lint && npm run typecheck && npm run build &&
npm test` under `web/` for SPA changes, `PYTHONPATH=. python3 -m pytest scripts/tests -q` plus
`ruff` for anything under `scripts/`, `.claude/`, `.husky/` or `.gitea/`, the BOM check for
touched `.cs` (`process.bom-format-detection-recipe`), and live-E2E on the slot's port for a
write path or UI change (`testing.live-e2e-prepush-timing`).
2. **Reviewers** read the worktree diff (`git diff origin/main...HEAD`) and run the gate in their
own worktrees. A round carrying any `blocking` or `should-fix` finding goes back to the fixer
regardless of the lens's own verdict word (a `should-fix` is a real defect by definition); nits
alone end the loop. A round in which every lens failed is an error, never a clean round. After
two fix rounds an open finding stops the workflow before the push and the orchestrator decides.
The same filter applies to the post-rebase round on the pushed head. In a resume, the existing
PR's stale head and body are the expected state until the finisher pushes, never a finding. The
loop ends on a clean round, never on "round one's findings are fixed".
3. **Finisher**: fetch; if `origin/main` moved, rebase, re-run the gate, and compare the patch-id
before and after — a changed patch (a conflict resolved, an artifact regenerated) sends the
branch through one more review round before the verdict, because the reviewed tree must be the
pushed tree. Then **one push**, `git push -u origin <branch>`; open the PR with `fixes #N` for
every issue in the bundle, the review history the workflow recorded (one line per round, quoted
verbatim — the finisher never describes a fix it did not see), the cross-family review status,
and the session trailers; arm the CI
monitor on the head sha; post the `## Closing record` on each issue with the evidence for every
`## Done-when` box, ticking none.
4. **Orchestrator**: read the review evidence, not the summaries. Send the PR back for anything
that lets a route or test pass having done nothing. Tick each box whose evidence holds, then
re-read the head sha immediately before posting `scripts/post-review-verdict.sh <pr> MERGEABLE
<note>` naming the reviewers, the rounds and the workflow's returned `cross_family` status. Read
that field, not only `error`: a cross-family failure in the post-rebase round leaves the branch
pushed and the run successful, and only `cross_family` says the substitute also failed; a
rubric-class PR in that state gets no verdict until a cross-family or substitute review of the
pushed head has run. Tick the review box. Wait for CI — a `cancelled` job
reads as `failure` at the combined status endpoint, so resolve it via the run's jobs
(`ci.cancelled-is-not-a-verdict`).
5. **Merge through the Gitea merge tool with the full head sha**; the consent hook derives consent
from the ticked boxes, the sha-bound verdict and green CI and auto-grants
(`release.merge-consent-autogrant`). A merge that `asks` is a state the orchestrator could not
derive, so derive it — never answer the prompt from memory. Then `DELETE /issues/{n}/labels/100`
on each issue, `git worktree remove`, refill the slot.
A rebase changes the head sha and voids the verdict; a fix after the push is a new commit, never an
amend (`process.pr-routine-sequence`). A branch that conflicts after its push goes back to the fixer
for a rebase pushed with `--force-with-lease` — the one sanctioned rewrite, scoped in
`process.orchestrated-session` — another gate run and a fresh review of the rebased head. Push a
branch behind `origin/main` and the pre-push hook refuses it (`release.format-as-you-touch-rebase`).
## Resuming a paused branch
A branch with a PR whose verdict is stale (head moved), whose CI is red, or that Gitea reports
unmergeable is resumed by `ersatztv-resume-branch.js` with a brief holding the done condition, the
last findings and the orchestrator's context. Its fixer rebases first, re-runs the gate, sits in the
same review loop, and the finisher pushes with `--force-with-lease`; the orchestrator posts a fresh
verdict for the new head. A resumed session finishes every open branch before it picks anything new.
## Hangs and incidents
- A `dotnet test` or E2E run with no output for ten minutes is a hang, not a slow test. Kill it by
the PID you started, gently; re-run once before diagnosing.
- Under an API incident (`https://status.claude.com/api/v2/status.json`), mechanical finishers and
probe-driven reviews run on whichever model is healthy; large implementation waits. Back off
before retrying.
- A workflow that returns an empty or unexpected result is read from its `journal.jsonl` before it
is re-run; a resumed run replays the unchanged prefix from cache.
- Before any stop: the H12 audit (`process.issue-qualification-audit`),
`scripts/refresh-shared-checkout.sh` (`session.shared-checkout-refresh`), and a handoff memory
naming every open branch, its worktree, its PR and its next step.