---
name: "convention-agent-runtime"
type: convention
license: CC BY-NC-SA 4.0
description: How a Layer 3 role becomes a running worker — the execution substrate the org model deliberately left unnamed, now in ad-hoc use. Names the three runtime shapes a role is realised as (in-session tool-scoped sub-agent, background agent, driver routine), the per-leg tool-scoping that turns Layer 5 policies such as never-merge-your-own-work and propose-only into constraints a worker cannot exceed rather than instructions a run must remember, and the model-tier-by-leg-and-risk lever behind cheap overnight runs. Read before instantiating a functional agent or delivery leg as a runtime worker, scoping its tools, or setting its model tier; and when reasoning about QA independence or parallel-build isolation. Governance stays in Layer 5 — scoping enforces a subset of it, defence in depth, never a replacement. Owned by Pulse.
---

# Agent runtime substrate

How a Layer 3 role becomes a running worker. `core-operating-model` names roles — Relay, Forge, Prism and the rest — as an organising layer and deliberately leaves the platform's runtime unnamed, "unless we build our own … at which point a thin concern becomes real." That condition is now met: the substrate is already in ad-hoc use — `routine-delivery-loop` spawns its QA leg as a sub-agent, and its model identifier was silently wrong for weeks (APP-698) — so the runtime mapping is named here, beside `convention-github` and `convention-vercel`, rather than configured by hand per run with nothing written down.

This is the tool layer for the execution substrate: it says how a role is instantiated at runtime. It does not move governance. Layer 5 policies stay in Layer 5; this convention says which of them the substrate can additionally enforce, and how.

## Role and worker are different things

A Layer 3 agent is a *role* — a mission, a remit, an accountability (`core-operating-model`, Layer 3). At runtime that role is realised as a *worker*: a process with its own context window, a specific set of tools it may touch, and a model tier. The realisation is **chosen by the shape of the work, not one-to-one with the role** — one role can run as different shapes on different legs, and a driver can spawn several. **Spawning the next leg is the default, not a thing to be asked for.** Where a session driving one role's work reaches a step that belongs to a different role — a review, a build, a gate, any leg carrying another executor label — it spawns the sub-agent for that leg itself, in the same turn, rather than routing the ticket and stopping. Routing records *who owns* the leg; it does not perform it, and a ticket sitting routed-but-idle is indistinguishable from a ticket in progress to everything that reads the board. The posture holds for an interactive session exactly as it does for a driver routine: the two differ in who opened the session, not in what the loop owes. What still stops a session is the gate the leg itself carries — an Owner signal, a credential, a decision that is not the session's to take — never the absence of an instruction to proceed. (Worked case: an interactive session built the A1-557 deliverable as Pixel and routed the ticket to PR-1SM, then waited for the Owner to ask for the review before dispatching it; the Owner's correction was that the dispatch is the session's own default — APP-1853.) Do not collapse a remit into a worker's system prompt, and do not read a worker's tool scope as the whole of its governance: the role is richer than any single instantiation of it.

## The three substrate shapes

1. **In-session tool-scoped sub-agent** — a bounded pass with its own context window, a scoped tool set and its own model, returning a verdict or an artefact to a driver. The `Agent` tool; already in use for the QA leg. The default shape for a delivery leg.
2. **Background agent** — work that is its own long-running, separately-monitored session. This is the off-the-shelf equivalent of the custom SDK branching interface already built, so it is usually a build-versus-buy question rather than a gap to fill; reach for it only where in-session isolation genuinely cannot hold the work.
3. **Driver routine** — a Layer 2 routine that sequences shape-1 sub-agents into a loop. `routine-delivery-loop` is exactly this: it is the driver, and the legs it spawns are the workers.

**A prompt handed to a person to open a session with states the mode its steps require — the mode is a tool scope, and no prompt can widen it from inside.** A prompt whose own steps write — move a ticket, write a Linear document, post a comment — cannot run in a read-only session mode, and the failure surfaces several steps in, after the session is open and context is spent, with no recovery but to end it and start again. So such a prompt opens with one line naming the mode: *"Run in normal (agent) mode; a research-mode session cannot write Linear."* **Name a capability tier, never a model id** (*Model routing* below): a prompt that pins a version string is APP-698 with a human in the loop. (Worked case: a kickoff prompt for the APP-2082 spike, 2026-10-06, opened in research mode; its steps set In Progress, wrote a Linear document and posted a comment, none of which ran, so the session was ended and restarted. At least the second instance — the Owner reported the same on an earlier kickoff — and the prompt was reused across APP-2095 to APP-2099, each now carrying the line by hand. APP-2120.)

## The mapping — role/leg to worker

| Role / leg | Shape | Tool scope — the boundary | Model tier |
|---|---|---|---|
| Relay — triage | in-session sub-agent | read + label / route; no merge | cheap |
| Forge — build (`skill-exec`) | in-session sub-agent (main worker) | repo / edit / bash / open-PR; no merge tool | strong |
| Prism — review (`skill-qa`) | in-session sub-agent | read-only; no write, no merge | mid, escalated by ticket risk |
| Relay — merge | in-session sub-agent | merge tool only, gated on the Owner signal | cheap |
| Operator weekly loop | background agent | vertical scope | per task |
| Scheduled tasks (daily-report, triage, pipeline-*) | background agent | connector scope | mostly cheap |
| The delivery loop | driver routine | driver; spawns the above | — |

The mapping is the convention: a shared, checkable choice of shape, scope and tier per leg. It is arbitrary-but-shared in the Layer 5 sense — another sensible split would work if everyone held it — so it is recorded here rather than re-derived per run. The merge tool lives in exactly one place, Relay's merge instantiation behind the Owner signal, which is what makes the governance headline below structural rather than remembered.

## Governance via tool-scoping — the headline

Instantiating a leg as a tool-scoped worker turns several Layer 5 policies from instructions a run must remember into constraints a worker cannot exceed:

- *"Never merge your own work"* / *"Forge opens a PR, never merges"* → the merge tool is absent from the Forge and Prism scopes; it exists only in Relay's merge instantiation, behind the Owner signal. A worker with no merge tool cannot merge.
- *"Propose-only"* (Pulse) → the Pulse instantiation may write Linear notes and nothing else: no canon write, no Forge dispatch.

Scoping is a **backstop beneath the instruction, not a replacement for it** — defence in depth. The policy still lives in Layer 5 and is still stated in the skill; the scope means a run that forgets the rule still cannot break it. Never treat a tool scope as the *source* of a policy, or delete the instruction because the scope enforces it — the scope enforces a subset, and the governance is the authority.

## Model routing — tier by leg and risk

Model tier is set at the substrate as a function of (leg × risk), not left to the session default: triage and classification cheap, build strong, review mid and escalated by the ticket's risk label. This is the direct mechanism for the night-routine goal — cheap tokens by default, strong only where the reasoning is deep. Pulse owns the model-and-schedule efficiency of the scheduled layer (`agent-pulse`), so a mis-tiered leg is a Pulse finding, the same class as a config that has drifted from its loader form. Name tiers by role (cheap / mid / strong), not by a specific model id — the id is runtime config that goes stale, and pinning the wrong one is exactly the ungoverned-by-hand failure this convention exists to close (APP-698).

A new **default-model generation** is a re-check trigger for this calculus, not a silent upgrade to ride. When a new default ships — e.g. Claude Opus 5, with a 1M-token context window, thinking-by-default and fast mode — the tier-to-id resolution moves under every leg at once, and the leg-and-risk split itself may want re-cutting, since a far larger context or a cheaper strong tier can move work between tiers. Naming tiers by role rather than id is what keeps this a bounded re-check rather than a fleet-wide edit, but the re-check is still owed: re-evaluate each leg's pin against the new generation — which legs move up, which hold, which drop a tier — rather than assuming the previous generation's split still holds. This is the runtime-substrate face of the same trigger skill-optimisation-review's model-and-schedule pass carries for the scheduled layer (APP-863).

## Isolation

Each leg in its own context window keeps a worker's build noise — logs, diffs, dead ends — out of a downstream worker's context. This is not only token hygiene: it structurally supports **QA independence**, which the delivery topology already wants — a clean QA pass can auto-approve once independence lands — because a reviewer that never saw the build reasoning judges the diff, not the author's account of it. Worktree isolation additionally lets parallel Forge builds mutate files without conflict, a lever on the one-Forge-per-repo concurrency rule; use it deliberately, since a worktree per worker carries real setup and disk cost.

**A worktree-isolated worker's bash surface is narrower than its tool list suggests, in two ways worth knowing before a leg is scoped into one.** The harness guards a worktree-isolated agent against reaching the shared checkout, and the guard classifies the *command string* rather than the process it would start. Two refusal classes follow, and neither is a bug to route around. **(i) A redirect to another checkout** — `cd /home/user/<repo> && …` and `git -C /home/user/<repo> …` are both refused, read-only inspections included, so there is no read-only escape hatch to a sibling repo. The refusal is scoped to the **shared checkout**, not to the flag: `git -C <absolute scratchpad path>` is taken normally, and a clone made in the scratchpad is worked in exactly that way, one plain subcommand per call. So the rule to carry is the destination and not the option — a sibling checkout is closed, the scratchpad the harness directs a leg to is open. **(ii) A command the classifier cannot statically decide** — a heredoc, an `&&` chain, a shell variable standing where an option may stand, a `python3 -c` program interpolating a variable, an `env -u VAR …` prefix — refused with a message asserting the command "cannot be shown not to be git" even where no git appears in it at all. The message is the misleading half: it sends the reader looking for a git problem that is not there, on a `sed`, a `gh` or a `python3`. What the class keys on is narrower than "names git" and narrower than "carries a shell variable", because **neither is required**. A fully literal `node -e` with a constant program string and no variable anywhere is refused as running a command whose name is computed at runtime. A heredoc that writes a script and then runs it is refused as too complex to verify. A chained `cp` of a directory into a scratchpad path is refused. A `git archive | tar -C` split, whose git half targeted the worker's own worktree and only whose output left it, is refused as naming git in a form too complex to verify. And a compound of absolute-path commands naming no git and no variable at all — `mkdir`, then `rm -rf`, then `cd … && npm run build` — is refused on the same grounds. The operative shape is therefore closer to **any compound at all, plus any `-e` program string**, whatever it runs and wherever it points. Note where these refusals land: on one leg three of the four were **reads of the scratchpad**, the directory canon itself directs the leg to use, so the guard fires hardest inside the sanctioned workspace rather than at the boundary it exists to defend. The working forms are cheap and are the default worth adopting: **one plain command per call with absolute paths**; a program **written to a file and run** rather than fed through a heredoc or `-c`; and, for anything about a repo that is not this worker's own, a **fresh clone pinned to the head under test** in the scratchpad rather than a read of the sibling checkout (`skill-qa`, *Pin the whole review to the head under test*; `convention-github`, *Fresh clones* and the hard-link rule for its `node_modules`). The cost of not knowing this is one to three dead turns per leg, and it lands on QA hardest because a review is *about* another tree: five QA legs on 2026-09-11 hit it across three worktree/ticket pairings — a `meirionpritchard-com` worktree reviewing `firmup-operator` and then `Luna`, a `Luna` worktree reviewing `mbn-theme`, a `claude-ops` worktree needing the sibling `agents-site` — each costing turns, none changing a verdict. The **same-repo** instance is the one that shows what the classifier is reading: on 2026-09-18 a QA leg whose worktree *was* the repo under review, with no sibling anywhere in the flow, still paid four refusals. So the guard is not consulting the target repo, or the target path, but only the shape of the command string — which is why following the documented remedy is necessary and not sufficient, and why a leg that has read this page still pays turns unless it also issues each call as one plain command. One asymmetry is worth recording because it breaks a procedure canon's own repos document: on the same sibling path `cd /home/user/agents-site && npm run build` was allowed while `cd /home/user/agents-site && git rev-parse HEAD` was refused, so the build half of `claude-ops` `CLAUDE.md` *Checking the agents-site render from a delivery run* runs and the provenance half — which commit that checkout is at — cannot. The fresh clone is the answer there too. The guard is the harness's and not this convention's to set, and no amount of documentation moves it: what is recorded here is the **form that works**, so a leg reads it rather than deriving it from a refusal at one to four dead turns a run. Narrowing the matcher is a **config lever, not an authoring one** — it is the Owner's to pull and is carried as a repo-lane action, never as a change to this page, so a reader who finds the refusals unreasonable is reading the right diagnosis in the wrong place (APP-1342, APP-1340, APP-1353, APP-1507).

**Three tool-level forms carry most of what a plain command cannot, and two more shapes are worth naming.** For file I/O — reading a scratchpad file, writing a script, a fixture or a result — use the `Read` and `Write` tools with absolute paths rather than `cat`, `sed` or a `cat >` heredoc: the guard classifies shell command strings, and a file-tool call is not one. For any wait-until loop — an `until … gh api …; do sleep …; done` poll on a check run or a deploy — use the `Monitor` tool rather than a shell loop, which is a compound by construction. For git against a scratchpad clone, `git -C <absolute scratchpad path>`, one subcommand per call, as above. The two further shapes: a file-producing command chained ahead of an interpreter (`cp …` then `node …`) is a compound and is refused like any other, so issue the two as separate calls; and a **relative** `cd` is a trigger of its own, refused with its own message independent of what follows it, so give every path absolutely rather than changing directory. Four QA legs on 2026-09-21 (A1-452, MBN-313, A1-444, A1-455) each paid three to four dead turns across `set -e` git chains, a `sed` with a shell variable, a python heredoc, a `cat >` heredoc, a relative `cd`, `cp`-then-`node` and an `until … sleep` poll before settling on exactly these forms. This records the forms that work; it does not move the matcher, which stays the Owner's config lever above. (APP-1596, 2026-09-21.)

**And a long measurement gets one backgrounded call per leg, with an explicit `timeout` — never two legs chained into one call.** The forms above govern what a single call may *contain*; this governs what a single call may be allowed to *lose*. A backgrounded call can be stopped before it finishes, and the stop takes everything still queued behind it: a base-and-head suite comparison written as one `for` loop in one backgrounded call lost its second leg to a background time limit about ten minutes in, after the first leg had completed cleanly, and the partial log was unusable — so the stop cost a whole suite run rather than the tail of one. Issue **each leg as its own backgrounded call with an explicit `timeout`**, in the order the comparison needs, each writing its own output file; a stop then costs one leg, and the leg already finished stays finished. The call budget is a **property of the surface, read rather than assumed**: on the refine surface of 2026-10-05 the `Bash` tool's own contract documents a `timeout` default of 120,000 ms and a maximum of **600,000 ms for a foreground command**, and documents **no** ceiling at all for a backgrounded one — so a "30-minute background default" is not a figure to plan against, and the 600 s at which a foreground `until` wait loop was moved to the background matches the documented *foreground* maximum exactly rather than marking a separate background limit. And the wait itself is not a shell loop: use the `Monitor` tool, per the paragraph above. Re-read the tool's own contract on the surface in hand rather than carrying these two numbers forward — the same discipline this page applies to the refusal classes, which are recorded as forms that work and never as a mechanism. (Worked case: the A1-588 QA leg, `assoc-one/meirionpritchard-com` PR #217 head `9a6f942`, 2026-10-03 — the full `--project=desktop` suite run on head (301 passed / 6 failed / 4 skipped, 6.1 minutes) and then on base, chained as one `for` loop in a single `run_in_background` call with no `timeout`; the harness reported the task stopped at its background limit with the base leg partway through, costing a roughly six-minute re-run of a leg that had nothing wrong with it. The two legs ran strictly one after the other, so this is **not** the shared-apparatus concurrency shape `skill-qa` covers (APP-1996) — the fault is the call's shape, not the server's. Premise unverified — no repo is readable from this surface, so the suite scope and the figures are the filing leg's report, and the delivery harness's own background behaviour was not measured here nor its version compared with the refine surface's; the `timeout` contract quoted above was read first-hand from this surface's own tool definition. The note's suggested link between the kill and a foreground wait loop hitting its own 600 s limit at about the same moment is named there as unconfirmed and stays unconfirmed. APP-2015.)

**A literal `git` inside a path or a filename is read as a git command, and this class fires on a single plain command.** Everything above concerns a command's *shape*; this is the remaining one, and it is the one that defeats the documented remedy, because the working form — one plain command per call with absolute paths — is exactly what gets refused. The guard matches `git` as a **substring of a path token**, so `sha256sum cap-*/downloads/convention-github.md`, a `stat`-and-`sha256sum` loop over `…/convention-github/SKILL.md`, and a `cat` of `.git/HEAD` are each refused as naming git "in a form too complex to verify" while invoking no git binary at all. The exposure is every publish, QA or capture leg touching `convention-github`, and every `github`-named repo path — `.github/workflows`, `content/github-stats.json` — so it lands hardest on the legs canon sends to exactly those files. Three working forms: **glob around the literal string** (`*hub.md` for `convention-github.md`); reach the file with the **`Read` tool** rather than `cat` where the path cannot be globbed; and ask git itself for repository metadata — `git -C <absolute path> rev-parse HEAD`, which is taken — rather than reading `.git/HEAD`, `.git/refs` or `packed-refs` as files. As with every class above, this records the form that works and does not move the matcher, which stays the Owner's config lever. (Three legs, all 2026-10-01, roughly four dead turns each: the APP-1909 QA leg on `claude-ops` PR #349, where `sha256sum cap-*/downloads/convention-github.md` succeeded only once rewritten as `*hub.md`; the APP-1948 QA leg on PR #359, where a `cat` of `.git/HEAD` was refused and recovered with the `Read` tool plus a separate `sha256sum`; and the APP-1956 QA leg on PR #364, whose `.git` path read was the one refusal of five that the one-command-per-call rule does not already cover — APP-1945.)

**The auto-mode permission classifier is a separate gate — from the worktree guard above and from the permission allow list.** Claude Code's auto-mode classifier judges a proposed action by its **class**, not by its path or its command shape, and refuses with a named reason. Three classes have been seen on fleet legs: `[External System Writes]` on a QA leg's CCR REST `POST …/ccr/ready_for_review` (APP-1545); `Self-Modification`, then `Instruction Poisoning`, on writing content into a repo's `.claude/skills/*/SKILL.md` — `Write`, a heredoc and a content `Edit` were refused while a `touch` of the empty file was taken, so it keys on content that changes future agent behaviour, not on the path (APP-1556); and a "security-weakening edit" on an exec negative control that briefly reintroduced a removed safety default, where QA's isolated leg ran the same control without a refusal (APP-1557). Three things follow. Rewriting the call as one plain command does not move a class that keys on *what the action changes or reaches* — `Self-Modification`, `Instruction Poisoning`, a security-weakening edit, an external-system write — because that is the worktree guard's remedy, not this gate's. **Narrowing what a call reads is a different move, and it does clear the data-breadth classes.** Two more classes belong to that second family: `[PII Data Handling]`, which has now fired twice on the inline-attachment transcript-recovery route against a full-file read of a leg's own `subagents/agent-<hex>.jsonl` (APP-1880; APP-2066, 2026-10-05), and `[Credential Exploration]`, on a broad `find /root … -iname` probe (APP-1928). On APP-2066 the same recovery completed unrefused once the read was scoped to a single line: `Grep` — the dedicated tool, not raw `grep` — to locate the line carrying a unique content fragment, then `sed -n '<N>p' file.jsonl > scratch.jsonl`, then `jq -r` over the one-line scratch file. Reducing the data a call touches is **routing to a narrower read, not rephrasing until it passes**, and it is the route `skill-exec` already names for this family (*And scope the transcript search to the harness's own project directory*). So read the refusal's class first and route by which family it is in; and where even the scoped read is refused, take the owning skill's own fallback and surface it, rather than narrowing again. **Unverified — no repo access from this refine surface:** the premise flagged is the classifier's keying itself — that the data-breadth classes clear on a narrower read while the action classes do not — which is inferred from the refusal instances cited here and not from the classifier's implementation, which no surface this convention is authored on can read. No other claim in this paragraph is covered by this flag. The refusal **never reaches the remote**, so a recovery keyed on a remote error, such as a 403 naming the route, does not fire; read the refusal text rather than waiting for a response. And the answer is **to route, never to rephrase the action until it passes**: take the route the owning skill names — the GitHub MCP for a QA leg's PR writes (`skill-qa`), claim-time routing of a deliverable in a refused class (`skill-exec`), a control of that shape left to QA's isolated leg — and where no route exists, stop and surface it under the approval-block rule below. Whether the classifier is tunable, and for which classes, is an Owner config lever beside the allow-list question already surfaced on APP-1400: named here, never authored here, never worked around by a leg. Each instance was seen on one surface only, so which surfaces enforce which class is unverified (the CCR route succeeded from another surface on 2026-09-14 — APP-1444).

**A newly named class refuses the same transcript-recovery route, and the remedy above was not tried on it — so its family is open, not settled.** `[Session Transcript Tampering]` refused a leg's `python3 -I -c` read of its own `subagents/agent-<hex>.jsonl` — parsing the file and printing which of three `Grep` matches was the `tool_result` — on the 2026-10-08 `APP-2141` exec leg, the same inline-attachment recovery the two data-breadth classes above already fire on. It is absent from the taxonomy this paragraph enumerates, and it refused a read of the leg's **own** record, which is the narrowest provenance a read can have. Two things follow, and the second is the one that cost the leg. The class's family is **unverified**: the refused call read the whole file inside an interpreter rather than in the APP-2066 scoped form — `Grep`, then `sed -n '<N>p'` into a one-line scratch file, then `jq -r` over that — so whether narrowing clears it was never measured, and nothing here says it would not. And the leg fell back without trying it, which is the mirror of the failure this page already names: **where a refusal lands on a route this page already carries a narrower read for, run the narrower read before taking the owning skill's fallback, and record the result either way** — a cleared refusal is the only evidence that moves a class into the data-breadth family, and an untried remedy leaves the next leg exactly where this one was. The route's own availability is `skill-exec`'s and `skill-ops-sync`'s to state, not this page's: across four legs it has now failed in four distinct shapes against one clean success (`APP-1920`), which is a finding about the route rather than about this gate. **Unverified — no repo access from this refine surface:** the refusal text, the file read and the leg's call sequence are as `APP-2157`'s thread reports them, and the class was not reproduced at authoring time; that flag covers those three items only. (`APP-2157`, `APP-2141`.)

**The gate can block a worker's *read-back* of a write it was permitted to make — and when it does, the confirmation moves to the driver rather than disappearing.** `[External System Writes]` has been seen refusing not only the write itself (a QA leg's CCR REST `POST`, APP-1545) but the query-only GitHub reads a leg issues straight after a write that went through — `gh api`, `pull_request_read get`, `get_comments`. The write lands and the worker cannot confirm it, so its hand-back claim rests on the write call's own response, and it cannot tell "my write landed as I intended" from "it landed differently, or silently failed". This is a substrate fact about where confirmation can live, not a QA-method question: a worker whose read-back is gated has no self-check to offer, so **the driver reads the external state back itself, from outside the worker, before acting on the worker's verdict** — and a clean verdict the driver has not independently read back is an unconfirmed verdict, not a confirmed one. It is the same boundary as *A sub-agent's return is a verdict or a path*: the return is a claim, and the check on it is the driver's. Whether the class should permit a read on a resource the same leg was just permitted to write is the Owner config lever already named above (APP-1400) — named here, never authored here, never worked around by a leg. (Worked case: APP-2055 and APP-2048, 2026-10-05, PR `assoc-one/claude-ops#387` — both QA legs' hand-backs arrived under `SECURITY WARNING: auto mode blocked this subagent's report. Reason: [External System Writes]`; the driver's own two `pull_request_read` calls confirmed `draft:false`, `mergeable_state:clean` and the posted verdict text matching the subagent's claim exactly, so nothing mis-routed — but the check was improvised both times rather than named, and a driver that trusted the self-report would have had no independent signal — APP-2073.) **Unverified — no repo access from this refine surface:** the premise flagged is the classifier's behaviour on these calls, i.e. that the write was permitted while the following same-resource read was refused; it is taken from APP-2073's record of the two 2026-10-05 QA legs and was not re-run at authoring time. No other claim in this paragraph is covered by this flag.

**A fourth shape is not a class at all: the same refusal fires non-deterministically on a plainly benign command.** On 2026-10-02 an exec leg had three Bash calls refused with `Reason: [Instruction Poisoning]` — an `npm run build:manifest`, a `git log` + `git merge-base` pair, and a `git add` of five tracked paths followed by `git status --short`. None touched `.claude/skills/` or `.claude/agents/` content, none piped to a truncating command, none reached outside a worktree (the APP-1342 / APP-1507 shape above). The third was re-issued **byte-for-byte unchanged** and succeeded on the very next call. So the content-keyed model above is necessary but not sufficient: **a refusal naming a class is not proof that the action is in that class**, and a leg that reasons from the named class alone will go looking for a route that does not exist — on a `git add` of tracked paths there is none, so the leg stalls. Two consequences. **One verbatim retry is the discriminator, and it is not the rephrasing this section forbids** — re-issue the identical command once, unchanged: if it passes, the first refusal was noise and nothing about the action needed to change; if it refuses again, treat it as a real class refusal and take the owning skill's route exactly as above. **Do not retry a second time, and never edit the command to get it past** — that is the rephrase-until-it-passes path, and it stays closed. **Record every instance, including the ones a retry cleared**, because the flake rate is the only thing that makes this actionable and a cleared denial otherwise leaves no trace: an unattended leg with no operator to retry stalls on a denial that has no cause to name and no remedy to cite, which is the APP-1400 shape reached on a different surface. This clause buys a leg one cheap retry; it does not touch whether the classifier is tunable, which remains the Owner config lever named above. (APP-1987 — observed cost on that run was one retry per denial and the leg completed; case 1 was worked around by invoking the script's entry point directly, so it is not a clean repro and only case 3 is. Repo, command-transcript and pull-request claims there are **unverified**: this surface can read no repo.)

Isolation is also a **context-budget** lever, not only a QA-independence and parallel-write one. A driver that runs a heavy leg inline accumulates that leg's context — files read, build logs, diffs — in its own window, so a long queue exhausts the driver mid-run and strands the remainder. Instantiating the heavy leg — the Forge build especially — as a **disposable fresh sub-agent** keeps that context out of the driver, which then holds only the leg's short result line and can drain the whole queue in one run. A fresh sub-agent is not a fork: a fork inherits the driver's accumulated context, which is exactly what the isolation exists to keep out, so a build leg is a fresh context, never a fork of the driver (routine-delivery-loop; APP-826).

**A sub-agent's return is a verdict or a path, never a payload.** The return travels as a tool result, and a tool result over ~50KB is not delivered — it is spooled to a single-line JSON file on a host path the driver's sandbox does not mount, that the Read tool refuses past ~25k tokens with no line to page on, and that Grep drops as a long matching line. The driver then cannot read work its own worker already finished; on 2026-09-03 two cluster returns of 56KB and 90KB cost two further sub-agents, ~240k tokens and 25 minutes to recover by slicing the spool (APP-1098). So a worker expected to return more than a few thousand characters — a change set, a package, a full assessment, anything quoting canon — **writes the payload to a file in the session outputs folder and returns only the path plus a short summary**; the driver reads the file in pages. This is the Agent-tool member of the silent-overflow family `convention-linear` records for `get_document` and `get_status_updates`, and it binds every shape-1 fan-out — the delivery legs, the Pulse refine and author fans, an operator loop's sub-passes.

**And the worker's return is its *only* channel — the Owner-facing ones belong to the driver.** The rule above governs the return's size; this governs its exclusivity. A fanned worker sees one row of a run and cannot judge what the run as a whole is worth telling the Owner, so the Owner-facing tools — `PushNotification`, and the user-file and user-message tools beside it — are **held by the driver and never called from inside a fanned worker**, whatever the worker's own tool list happens to offer. The driver synthesises across every worker's return into at most one considered notification; a worker that pings on its own row's completion spends a scarce channel on routine progress, arrives ahead of the driver's own message, and does it invisibly, because the call appears in the worker's sidechain transcript and not in its return. The *what* this sits inside is `convention-comms-owner`'s: an Owner notification is a scarce channel whose first sentence is the whole message on a phone banner, and an unattended pass leads with the finding in its own remit (APP-1150). This is the general form of the per-routine restriction `skill-ops-retro` already carries for `SendUserFile` — subagents return the package path and the parent sends each file after the fan returns (APP-1534) — stated here so it binds every shape-1 fan-out instead of being re-stated per routine, and so a tool nobody thought to enumerate is still covered. (Worked case: the 2026-10-05 `task-pulse-retro` author run — one of 21 fanned author subagents called `PushNotification` mid-run with "Pulse retro: skill-qa row complete…", routine progress by construction and the only such call in the fan; it was found only by grepping each subagent's sidechain for a `PushNotification` `tool_use` block during telemetry capture, it did not surface in the subagent's return, and whether it reached the Owner at all is unconfirmed — APP-2040.)

**The scheduled layer has its own member of the same family.** `list_triggers` — the Cloud-Routine enumeration — overflows on a fleet of roughly 30 or more, because each row carries its full prompt: the payload is large **per row, not per page**, so a smaller page trades one overflow for several and never returns the fleet. The remedy is the sub-agent remedy above, read the other way round: take the **spooled payload** and parse out the fields the caller needs, rather than shrinking the request. `skill-task-health` Pass B holds the form; `convention-linear` *Reading Linear safely* holds the discipline the whole family shares.

## Unattended runs leave no trace of their label writes — so every run posts a run-log

Linear keeps **state history and comments only**. A label, priority, assignee or relation write leaves no history the connector exposes — no per-field `updatedBy`, nothing — and routing, escalation and executor hand-off are all label-and-priority writes. So an unattended run's most characteristic actions are structurally unverifiable after the fact: a capture, an audit or an article asking "what did the scheduled layer actually do on date X" cannot answer it from the tracker (the MRP-43 capture had to mark `[GAP]` on exactly the three actions it existed to evidence — two routings and a High → Urgent step — APP-923). The comment-idempotency guard makes it worse by design: the rule that stops a sweep re-posting removes the one artefact that would have recorded the action. The gap is **permanent** (Owner decision, Aled 2026-09-03): no connector-side history is coming, and the standing agenda is refreshed in place, not versioned.

**A run that stops on an approval says so, while it is still stoppable.** The run-log below records what a run *did*; this rule covers a run that is *waiting*. An unattended run blocked on an approval — a tool-permission prompt, a connector's own authorisation, an MCP allowlist — holds a state that is indistinguishable from a slow pass: the scheduler shows `PENDING` with no reason attached, the history shows a gap, and nothing is emitted. So the moment a run stops on a gate it cannot pass, it **writes one line to the durable surface it already reports on**, naming the gate, the surface, and what will not happen until it clears. The line is written **before** the run waits, not after it times out, because after the timeout there is no run left to write it. A run that then completes on a granted approval says so in its run-log; a run that times out has already left the only trace anyone can act on.

**Deny-and-fail-fast is a different remedy and is not the default.** On a headless host, a harness profile that denies anything which would prompt turns a hang into a legible failure, and for a leg with no human anywhere near it that is the right profile. It is the wrong profile wherever the gate is legitimate and a human would grant it: on 2026-09-08 a cloud fire of `task-goal-review` sat `PENDING` for 92 minutes on an unanswered Shopify connector approval and then completed correctly the moment it was granted — denying automatically would have failed a run that was fine (APP-1178). The damage was never the prompt existing; it was that nothing surfaced it. So **surfacing the block is the standing rule for every unattended leg on every surface; denying the prompt is a per-leg choice, made where no one is there to grant it and stated in that leg's own artefact.** The two compose: a leg that denies still says what it denied and why.

The remedy is a **run-log**: every scheduled or unattended run — an operator loop or stand-up, a dispatch sweep, a drain pass, a delivery-loop run — **posts one comment on a durable surface naming what it changed, ticket by ticket** (state, label, priority, assignee and relation writes, with the ticket IDs), on its own loader ticket or the vertical's initiative as the artefact names; never one comment per touched ticket, and never a repeat where nothing was written. It is a run record, not a per-ticket recommendation, so it does not reopen idempotency. Each scheduled artefact's output contract carries the line (`skill-ops-retro` step 6 is the first); a capture or a Mr P piece citing fleet behaviour cites the run-log, and a self-report with no run-log behind it is marked as a self-report by construction (`convention-mrp-voice`).

## Getting a local-only file into a repo — no worker holds both halves

A file that exists only on the Owner's Mac has no worker that can carry it to a repo, and the gap is worth stating once so it is not re-derived per ticket. Relay's surface is Linear plus Claude Code *for the merge action only*, and its boundary is explicit — it never writes code, so it cannot commit a file (`agent-relay`). Forge's surface is Claude Code and GitHub repos, with no route to a person's machine (`agent-forge`). A Cowork session with the device bridge is the only worker that can read the Mac, and it holds neither repo write nor a working attachment upload. So a ticket, template or blocker note saying "drop the file, or hand it to Relay to commit" names a route that does not exist: correct the line rather than improvise around it.

**The two sanctioned routes (Owner ruling, Aled 2026-09-11).** Relay does not get repo-write capability. A local-only file reaches a repo one of two ways: (a) **the Owner attaches it to the Linear ticket** — drag-and-drop in the Linear UI, from a browser with ordinary network access — and the build leg fetches the attachment when it picks the ticket up; or (b) **the Owner commits it himself** from a real clone, where he is not deliberately routing the work through the fleet. There is no third route. In particular, **a session never inlines file content into a Linear comment or description to work around a failed upload** — it is the payload failure this convention already names one section up (*a sub-agent's return is a verdict or a path, never a payload*), it corrupts the ticket as a record, and it was tried at real cost before the two routes were settled.

**Route (a) carries a text file, and only one that lands as a formal attachment.** Two limits sit under it; neither changes the ruling, both say where it holds. First, **a file dropped into the ticket description is not an attachment**: it lands as an inline `<linear-embed node-type="file">` in the description body while the issue's `attachments` array stays empty, `get_attachment` on the embed's upload id returns 400 ("Could not find referenced attachment with this url"), and the embed's `uploads.linear.app` href is refused by the proxy at CONNECT. The leg has nothing it can fetch, and a retry does not change that. Which Linear UI action yields a fetchable attachment instead is not yet confirmed and is the Owner's to confirm. Second, **even a formal attachment is view-only for a binary**: `get_attachment` renders an image, archive or font and gives back no bytes the sandbox can write (`skill-exec`, *Guardrails*, APP-574). So a local-only **binary** — a zip, an image set, a font — goes by route (b), the Owner commits it; route (a) is for text. A leg that finds an empty `attachments` array beside a description file embed does not treat the file as delivered: it routes `human` + Aled once, naming route (b) and the file, so the ticket does not go back to Todo on the belief that the file has arrived. (Worked case: A1-351, `assoc-one-v1.zip` embedded in the description on 2026-09-21, unfetchable; the ticket was re-routed to Forge on the belief it was solved and re-Blocked in a second round trip — APP-1587, 2026-09-21.)

**The Cowork limitation is a platform one, not an unconfigured setting.** An attachment upload from inside a Cowork session (`prepare_attachment_upload`, then the `PUT` to the storage host) and a direct `git clone` to `github.com` were both refused with a 403 by the session's own egress proxy — from the cloud container and from the linked-device shell — with the organisation's egress capability already set to all domains, and the defect is filed upstream. So treat it as standing: **a Cowork session cannot push to GitHub or upload a Linear attachment**, and the two routes above exist to work around exactly that. Treat it as re-testable, not permanent: egress settings apply only to sessions created after a change, so a run that believes this is fixed proves it in a **fresh** session before relying on it, and says which session it proved it in. (APP-1335. The upload direction is claimed to work in APP-1349; the two are not reconciled — settle them together before either is treated as fact.)

## Boundaries

- **Governance stays in Layer 5.** This convention places a role into a runtime shape and scopes its tools; it never authors or relocates a policy. Scoping enforces a subset of governance; it is not governance.
- **The org model is richer than the substrate.** A role is not its worker, a remit is not a system prompt, a tool scope is not the boundary of accountability. The profiles (`agent-roster` and the agent/operator artefacts) stay the source of what a role *is*.
- **Background agents and agent teams are likely redundant** with the SDK branching interface already built — a build-versus-buy read, not a capability to adopt by default.
- **Surface placement is unchanged.** Whether a scheduled worker runs on Cowork or a Cloud Routine is still decided by connector need (`core-operating-model`, Layer 2), not by this convention.

## Quick checklist — before instantiating a role as a worker

- [ ] Right shape for the work — in-session sub-agent, background agent, or driver routine?
- [ ] Tool scope set to the minimum the leg needs, with any policy-enforcing tool (merge, canon write, dispatch) deliberately absent where a Layer 5 policy says the leg must not hold it?
- [ ] Model tier set by leg and risk, named by role not by a model id, not left to the session default?
- [ ] Isolation adequate for QA independence, and worktree isolation used only where parallel writes need it?
- [ ] The governing policy still stated in the skill — the scope is a backstop, not the source?

- [ ] Long returns written to a file in the session outputs folder, path returned — never a payload over a few thousand characters as a tool result?
- [ ] The run's output contract carries the run-log line — one comment on a durable surface naming every write, ticket by ticket?

## Tone

Calm, precise, sentence case, British spelling, no exclamation marks, outcome before adjective. (cos.tov.)
