---
name: "convention-host-access"
type: convention
description: 'How an interactive Cowork or Claude session reaches — or structurally cannot reach — an always-on host, and what to do instead. A session cannot connect to a host over SSH (no ssh config, no key material on its filesystem, HTTPS-only egress that blocks TCP/22), so host work — deploys, restarts, log reads, credential placement — is routed rather than attempted: handed to the operator as a copy-pasteable command block he runs and pastes back, following convention-secrets for the never-print-a-secret discipline. Read before briefing or doing any host or SSH work from a session, so the structural dead end is not re-derived by failing each time. The host sibling of convention-github''s route-it-don''t-clone rule. Ruled: no host key ever lives in a cloud container''s setup script; where read-only host visibility is wanted, the host publishes outbound to a surface a session can read. Grown as host work recurs.'
license: CC BY-NC-SA 4.0
---

# Host access

How a session reaches — or structurally cannot reach — an always-on host, and what to do instead. bot.Trader now runs a live always-on host (Hetzner, APP-785) carrying the drawdown guard, so host work — deploys, restarts, log reads, credential placement — is a standing category, not a one-off. Nothing in canon said which surface can reach it, so each session re-derived the answer by failing. This writes it down.

## The rule — route host work, don't attempt an SSH it cannot reach

An interactive Cowork / Claude session **cannot reach a host over SSH**, and the correct move is to route the host work rather than attempt the connection first and route only after it fails. This is the host edge of the same rule `convention-github` holds for repos: the direct route is a structural dead end, so diagnose from the ticket and route, do not burn a step discovering the dead end.

Three independent reasons, each sufficient on its own, make the connection impossible from the session's sandbox:

1. **No SSH config** — there is no `~/.ssh/config`, so a host alias (`ssh bottrader`) resolves to nothing.
2. **No key material** — the private key lives on the operator's Mac, which is correct and must not change; nothing on the session filesystem can authenticate.
3. **No egress on TCP/22** — the environment's proxy carries HTTPS only, so outbound SSH times out even with a correct address and key.

**All three reasons are properties of the sandbox, so the rule binds the sandbox — not every Claude.** Each of the three is true of a cloud Cowork session and of any Claude Code session running in a hosted container. None of them is true of Claude Code running **locally on the Owner's own Mac**, where the SSH config and the key material already sit and the egress is his own network. State the distinction, because it is routinely collapsed: a session reading "a Claude session cannot reach the host" applies it to the Owner's own tooling and hands him a command block to paste for work that tooling could chain in one continuous run. The multi-round case is where the collapse costs most — an investigation whose next command depends on the last answer pays a full round-trip per step (worked case: the 2026-09-11 `work.Unblock` session on APP-929, APP-932 and APP-915, where establishing that PR #37's commit was an ancestor of neither the deployed checkout nor `origin/main` took four paste-back cycles). The judgement calls are unaffected either way and stay the Owner's: which alert receiver to use, whether to merge a PR that skipped review. **What is not settled here is whether the Owner's local Claude Code is an approved driver for host work, and under what bounds** — that is an authority grant, the same question the ruling below settles for containers, and it is his to make rather than a session's to assume. Until he rules it, the paste-back block remains the route from every session, and a run that notices the round-trip cost names it to him instead of reaching for a local runner. (APP-1332.)

Only one sandbox exists to spawn into, so a sibling probe session reproduces the same result — spawning does not escape it. (Worked case: a bot.Trader operator session was briefed to do host work over `ssh bottrader`; it installed an SSH client to rule out a missing binary, spawned a probe to test empirically rather than assert, and confirmed the connection was blocked on all three counts — APP-862.)

## What to do instead — the operator-run command block

Host work from an interactive session is done by **handing the operator a copy-pasteable command block to run on his Mac and paste the output back**, then reading that output. It completes real host work — deploys, restarts, log reads, credential placement — without the session ever touching the host.

The block follows `convention-secrets` (*command blocks that report state, never the secret*): it prints `SET`/`UNSET` and last-4 only, never a value, and states the expected-good result per command so the operator can tell it worked. Written for someone reading mid-task, it follows `convention-comms-owner`'s instruction register — one action per step, what he should see stated at each. (Worked case: the 2026-08-13 bot.Trader host deploy completed entirely this way.)

### The block must survive the paste

The block is tested by the paste, not by the author reading it back. The Owner's shell is interactive zsh on macOS; the author's is a Linux `bash -c` sandbox that tolerates what zsh does not, so the author's environment will not reproduce the failure — which is why this is canon rather than remembered. Three rules:

- **Explanation goes outside the fenced block.** Inside the fence, commands only. An inline `# comment` carrying `?`, `(`, `)`, `[`, `]` or `*` is globbed and executed under interactive zsh (`#` is not a comment there unless `interactive_comments` is set, which by default it is not), so the comment errors, the output interleaves with real results, and the Owner has to reconstruct which output belongs to which command. Put the annotation in prose or a numbered list above the block.
- **A command whose expected-good result is empty carries a distinguishable marker.** Terminate it with `&& echo NONE` or `|| echo NONE`, so ran-and-found-nothing and did-not-run are never the same transcript. This is the substantive rule; the comment syntax is only the trigger.
- **State the expected-good result per command**, as `convention-secrets` already requires, and read the pasted output against it.

(Worked case: the 2026-08-14 bot.Trader daily read issued a correct block whose six comment lines errored under zsh, and whose kill-switch check — expected result "empty" — produced no output; whether it had run at all was recoverable only because `halted: false` happened to be in the same block — APP-907.)

### Record how to reach the host

A host in standing use has its access details recorded in the repo's deploy doc (`deploy/README.md` for bot.Trader): the SSH alias, the provider and region, where the key lives (never the key), and the account or console needed to rebuild the box if it is lost. A host address is not a credential and `convention-secrets` does not bar it; a routed-host convention that never says where the host is describes a procedure whose first step is missing. The doc is a recovery artefact whose fidelity is tested only at the moment it is needed most, so it is read against the host at each deploy and corrected when it drifts — `skill-task-health`'s deployment-parity lane, not a new check. (Worked case: on 2026-08-15 an operator preparing a five-minute Owner-run block for APP-942 could not find the host's address anywhere in the repo, and the first substantive command then named `/etc/bot-trader`, which the box did not have — two round-trips on one task; APP-960, APP-929.)

## Measuring a live page from the sandbox — what works, and why the obvious routes fail

Built-surface measurement is the second half of every design-parity, design-QA and visual-regression ticket, and for a connector-surface agent it is the *only* route to the build: `agent-pixel` has no repo access, so the deployed page is what it measures. The sandbox does not make this easy, and each run rediscovers the same dead ends unless they are written down.

**Two sandboxes, opposite answers — read which one you are in first.** This section describes the interactive **workspace** image. The **delivery** sandbox the Cloud Routines run in is a different image and already carries a browser: Chromium under `/opt/pw-browsers` with `PLAYWRIGHT_BROWSERS_PATH` set, a global `playwright` under `/opt/node22/lib/node_modules`, and `chromedriver` on the path. A delivery run that reads the recipe below and concludes no browser exists has read a true statement about the wrong image, and defers a criterion it could measure in minutes (`skill-exec` step 3c, `skill-qa`). Versions and paths are properties of each image, not canon — read them, do not assume them. (APP-1186.) **And one CDP input route on that image is inert rather than absent:** `Input.synthesizeScrollGesture` with `gestureSourceType: "touch"` scrolls **0px** while `mouse` and `default` scroll normally, with no error raised — so a touch-modality check on this image uses an explicit `Input.dispatchTouchEvent` sequence (`touchStart`, a run of `touchMove`, `touchEnd`), which does scroll and does fire `pointerType: "touch"`. Re-measured 2026-10-05 on Chromium build 1194 / 141.0.7390.37 with Playwright 1.56.1. This is the same instrument-absence family as the socket-table read below, one level in: the call succeeds and carries no signal in either direction, so an inert gesture is recorded as **no signal** and never as evidence the page does not scroll. `skill-qa`'s obligation 5 carries the QA-side rules that hang off it. (APP-1975.)

**Why the obvious routes fail.** The `mcp__workspace__bash` image ships **no browser binary** and lacks `libxdamage1`. `playwright install --with-deps` **fails** — it shells out to `sudo`, and the sandbox sets no-new-privileges. The Chrome MCP needs the operator's Chrome connected, which an unattended run cannot assume. `curl`-and-read-the-CSS gets close but is **derivation, not measurement** — it cannot resolve cascade order, media-query precedence, flex resolution or negative-margin collapse, and it silently gets border-box arithmetic wrong.

**What works — observed on this image, 2026-08-22; pin versions and the missing shared object as facts of that image, not canon.**

1. `npx --yes playwright@1.49.0 install chromium` — without `--with-deps`.
2. `apt-get download libxdamage1 && dpkg -x libxdamage1_*.deb libs` — download-and-extract needs no root; `libxdamage1` was the single missing shared object.
3. `LD_LIBRARY_PATH=$PWD/libs/usr/lib/aarch64-linux-gnu node measure.js`, launching headless Chromium with `--no-sandbox --disable-dev-shm-usage`.

With that, `getBoundingClientRect()` and `getComputedStyle()` are available on every element at an exact viewport and DPR, and a screenshot can be measured pixel-by-pixel with PIL (already installed) to find where rules actually land — which is what caught the real cause on A1-295: Figma's centre-aligned strokes against CSS's inside borders, invisible to any CSS read.

Two gotchas. `waitUntil: 'networkidle'` never fires on a site whose animation or analytics keep sockets open — use `domcontentloaded` plus an explicit wait. Workspace files are visible to the Read tool only once copied into the mounted `outputs` folder.

**The stronger fix is configuration, not canon:** baking `libxdamage1` and a Chromium into the workspace image retires this recipe. That is the Owner's action and is named to him, not encoded here. (APP-1044.)

**A socket-table read is not an instrument in this sandbox, and an empty result is not a free port.** Observed 2026-09-26 in the delivery sandbox: with `next dev` started directly and `curl http://localhost:3000/` returning a real HTTP response, so the port was demonstrably bound, `lsof -i :3000` exited 1 with no output, `lsof -i tcp:3000` likewise, and `ss -ltnp` and `netstat -ltnp` both printed nothing. After the process was killed the same three tools reported nothing again — indistinguishable from the bound state — so on this image they carry no working signal in either direction, and only `curl` does. All three binaries are installed and all three commands run, which is the whole point: this is the instrument-absence family (`skill-exec` step 3c's `which <tool>` note) one level further in, where presence-testing the binary passes and the **socket-table read itself** returns nothing regardless of ground truth, plausibly a container or namespace restriction on reading other processes' socket tables. So a port-bound or port-free assertion from this surface is made by **attempting the connection**, and an empty `lsof`, `ss` or `netstat` result is recorded as **no signal**, never as evidence the port is free. *Flagged unverified, one premise named: whether a repo's own port check is affected turns on how that check is implemented, and this pass could not read one. `assertPortFree` in `firmup-operator` and in `Luna` may attempt a bind rather than read the socket table, which would sidestep this entirely; neither repo is readable from this cloud surface. The sandbox observation above does not depend on that mechanism and stands on its own; the effect on those scripts is not asserted here, and is worth re-testing from a surface that can read them, along with whether the restriction is general or specific to one run's container.* (APP-1797, APP-1584.)

## Ruled — host work is never un-routed; visibility comes outbound

Whether a session should *ever* reach a host directly turned on one security decision: should a private key live in the environment setup script or a cloud container? `core-operating-model` names the setup script as the durable home for harness config, so placing host access there was the obvious mechanical path. **The Owner ruled no (2026-09-03, APP-862): a host key, scoped or otherwise, never lives in a cloud container's environment setup script.** Host work stays routed — handed to the operator as a command block, never attempted from a session — and where read-only host visibility is wanted, **the host publishes outbound** to a surface a session can read (the health digest `task-bottrader-watch` and `task-bottrader-operator-loop` read, APP-914 / APP-1006); a session never reaches in.

The reasoning is the transferable part. Every other authority granted to an agent in this system is bounded structurally, not behaviourally — Aide writes only inside exact-named hold events, Forge opens PRs and cannot merge, Pulse proposes and cannot save. A shell has no equivalent bound: the grant would rest on the agent behaving, which is the class of guarantee this system avoids everywhere else. So the rule is not "hosts are dangerous" but **no unbounded grant**. Two specifics: opening TCP/22 is a change to the cloud environment's network policy, not a per-task grant, so it widens egress for every routine sharing that environment; and environment variables in a cloud environment are visible to anyone using it, which is a direct warning against parking a private key there. The honest counterweight is that routing is not the zero-risk option — its failure mode is manual credential placement, which has already cost a zero-order trading day; that exposure is now governed by `convention-secrets` (*One store is authoritative*, APP-896), not by loosening this rule.

## A claim quoted from a file is a claim about that file at a ref

A quotation from a repo file — "the file's own warning", "its `CLAUDE.md` says" — is an **assertion about that file at a named ref**, not a standing fact about the repo, so it is written with the ref and the short SHA it was read at. A claim that cannot be re-read from the current surface is written as **reported, never observed**. This bites hardest where the claim is inherited rather than made: a ticket authored from another note's quotation carries the quotation and not the check, and nothing downstream distinguishes the two, so the second ticket reads as well-grounded as the first while resting on nothing either run verified. `convention-comms-owner` (*Attribution — what the run saw against what it was told*) holds the family rule for every figure and claim a run repeats; this is that rule reaching file contents, and it is cited rather than restated here.

(Worked case: APP-1552 §2 stated that `meirionpritchard-com`'s `CLAUDE.md` warned explicitly about `pkill -f "next dev"`, and APP-1584 was authored from that filing note and framed its dev-server-stop criterion as replacing the existing warning's wording. `git log --all -p -- CLAUDE.md` against `origin/main` shows no such warning at any commit in the file's history, and `--all` covers every ref, so this is not a removed-without-trace case. The premise check held: the mismatch was reported on the ticket and the exec leg added the safe-form guidance fresh rather than fabricating a rewrite of content that was never there. Several sibling repos do carry `next`-process-kill guidance under other names, which is the likelier origin of the misattribution — APP-1797, APP-1584.)

## Relationship to the other conventions

`convention-github` holds the same rule for repos (*route it to Forge, don't try the clone*); this is its host sibling. `convention-secrets` holds the command-block discipline the routed work uses (report state, never the value) and the standing rule that a live credential never reaches a tracker or chat. `core-operating-model` holds the three-lane split this sits inside and names the setup script the ruling above keeps a host key out of.

Note (adjacent, ruled): bot.Trader holds the same credential in two stores — the host's `/etc/bottrader/env` and the GitHub Actions repo secrets — and they diverged once (the host on the new USD account, Actions still on the old GBP one, which cost a zero-order trading day). `convention-secrets` (*One store is authoritative — the multi-store case*) now names the host env file authoritative, Actions derived from it and never the reverse, with a startup assertion that logs the last-4 of the account actually loaded (Owner ruling 2026-09-03, APP-896). A host-touching session reads that rule before placing a credential in either store; the two stores are named here so the session knows both exist.

## Tone

Everything written follows `cos.tov`: calm, precise, sentence case, British spelling, no exclamation marks, outcome before adjective.

## Quick checklist before host work from a session

- [ ] Recognised that an interactive session cannot reach a host over SSH (no config, no key, no TCP/22), so routed rather than attempted?
- [ ] Host work handed to the operator as a copy-pasteable command block, output read back — not run from the session?
- [ ] Command block reports `SET`/`UNSET` and last-4 only, with the expected result per command (`convention-secrets`), one action per step (`convention-comms-owner`)?
- [ ] Commands only inside the fence, explanation outside; every empty-expected command ends `&& echo NONE`/`|| echo NONE`? (APP-907)
- [ ] The host's alias, provider, region, key location and rebuild console recorded in the repo's deploy doc, and the doc's paths match the box? (APP-960)
- [ ] No host key placed in any setup script or container environment — host work routed, and any host state a run needs taken from what the host publishes outbound, never reached in (Owner ruling 2026-09-03, APP-862)?
