---
name: "skill-telemetry-capture"
type: skill
description: 'Read a run''s own session transcript, extract one telemetry record against schema v1, and write it to Drive as the run''s closing act. The collector osclaude.vitals Phase 1 rests on: it locates the transcript unaided from the session id in the environment, folds in subagent sidechain files, counts usage once per requestId while gathering tool blocks across every sibling record, emits both cold-start figures, and is idempotent on session id so a retried capture cannot duplicate. Deployed in-band as a loader footer on an instrumented task; the Stop-hook collector is a later phase. Read whenever a scheduled task is instrumented for telemetry, a run is asked to report its own cost, or a captured record has to be reproduced, corrected or read back.'
license: CC BY-NC-SA 4.0
---

# telemetry-capture

The collector `osclaude.vitals` rests on. A run reads its own session transcript, extracts one telemetry record and writes it to Drive as its closing act. The baseline, the dashboard and the cold-start reduction the project exists to drive are all computed from these records, so the failure that matters here is not an absent record but a plausible wrong one — a naive collector over-counts tokens by roughly fourfold and looks entirely credible doing it. Every rule below is a measured finding from the spike (APP-1130) and its ADR `spec.vitals-telemetry-capture`, not a preference.

## Trigger

Invoked as the final step of an instrumented run — a scheduled `task-*` carrying the loader footer described under *Deployment*, or an ad-hoc run asked to report its own cost. One record per run, written once, at the end. Never invoked mid-run to report progress.

Two inputs come from the caller because the transcript has no source for them:

- **`task`** — the artefact or job the run executed, as a slug (`routine-delivery-loop/spike-app-1130`). No environment source exists; a collector that guesses this makes the baseline unattributable.
- **`outcome`** — what the run concluded, in one token or short phrase (`completed`, `blocked`, `spike-probe`). Correctly supplied rather than derived: only the run knows.

Where either is missing, record `unknown` and say so in `notes`. Never infer them from the transcript.

## Behaviour

1. **Locate the transcript, unaided.** `CLAUDE_CODE_SESSION_ID` is in the environment; the file is `~/.claude/projects/*/$CLAUDE_CODE_SESSION_ID.jsonl`. Resolve it by glob. Do not ask the caller for a path and do not accept one — a supplied path is how a run ends up measuring a different session.
2. **Fold in the sidechains.** Subagent work is written to separate files at `.../<session-id>/subagents/agent-*.jsonl`, sharing the parent `sessionId` and flagged `isSidechain: true`. A collector reading only the parent file misses every delegated turn, which was the majority of the tokens on the spike's own run. Glob the subagent directory and include those records in every aggregation below.
3. **Aggregate under the two counting rules.** See *The two counting rules* — this is the step that goes silently wrong.
4. **Assemble the record** per *The record — schema v1*, filling `task` and `outcome` from the caller and `notes` with anything a reader needs to discount the figures.
5. **Write it to Drive**, search-before-create, per *Where it writes*.
6. **Report one line** in the run's own output: the record's filename, the run's total tokens and its cold-start figures, or the reason no record was written. This line is the only trace of the capture in the host run's output.

The transcript is readable live, mid-run — proven, not assumed. The cost of that is stated under *What this cannot capture*.

The approach needs Bash and a filesystem. On a connector-only surface it is impossible rather than slow: decline, state that, and write nothing. A self-reported record is not a fallback — see *Guardrails*.

## The two counting rules

One API response is written to the JSONL as **several records**, one per content block, each carrying a **complete copy of the same `usage` object**. Two different rules run over those same records, and getting one right while getting the other wrong is easy and silent.

- **Usage — count once per `requestId`.** Summing per record over-counts: 3.7× on the spike's Claude Code run, 1.82× on a Cowork session measured 2026-09-07. The factor is surface-specific; the defect is not. Deduplicate on `requestId`, then sum.
- **Tool calls and writes — gather across every record, deduplicate by tool-use block `id`.** `tool_use` blocks are distributed across the sibling records, so applying the `requestId` rule here drops the inventory — 15 tools to 2 in the spike's run.

`turns` is the count of distinct `requestId`s across the parent transcript and its sidechains, not the count of records.

## The record — schema v1

One JSON object. Twenty-one fields, exactly these names.

| Field | Source |
| -- | -- |
| `schema_version` | Literal `1` |
| `task` | Supplied by the caller |
| `session_id` | `CLAUDE_CODE_SESSION_ID` |
| `model` | The model on the assistant records |
| `surface` | Surface and entrypoint together, e.g. `claude_code_remote/remote_trigger` |
| `started_at` | Timestamp of the first record |
| `finished_at` | Timestamp of the last record read — short by the closing turns, by construction |
| `duration_s` | `finished_at` minus `started_at` |
| `cold_start_tokens` | The first request's `cache_creation_input_tokens`. Kept for continuity with the original definition; not comparable run to run |
| `cold_start_context_total` | The first request's `cache_creation_input_tokens` plus its `cache_read_input_tokens`. The stable figure, and the one to compare |
| `cache_creation` | Sum of `cache_creation_input_tokens`, once per `requestId` |
| `cache_read` | Sum of `cache_read_input_tokens`, once per `requestId` |
| `input` | Sum of `input_tokens`, once per `requestId` |
| `output` | Sum of `output_tokens`, once per `requestId` |
| `thinking` | Sum of thinking tokens, once per `requestId` |
| `turns` | Count of distinct `requestId`s |
| `tool_calls` | Tool name to count, gathered across all records, deduplicated by block `id` |
| `writes` | Connector write calls itemised, same gathering rule |
| `skills_attributed` | Skill name to turn count, from `attributionSkill`. See below |
| `outcome` | Supplied by the caller |
| `notes` | Free text: caveats, surface, anything that makes the figures readable later |

**Why both cold-start figures.** The original definition — the first turn's `cache_creation_input_tokens` — under-reads whenever part of the prefix was already warm: 29,433 apparent against 56,468 real on the spike's run. It measures how much of the opening context was not already cached, which moves with cache warmth outside our control, so it is not comparable between runs. `cold_start_context_total` is. Emit both, compare on the second.

## `skills_attributed` — present in the schema, absent on some surfaces

The field stays in schema v1 and tolerates absence (Owner decision, 2026-09-07, APP-1138).

`attributionSkill` presence tracks the **entrypoint**, not the surface — so it is tested per run, never inferred from the surface name. Measured on an **interactive** Cowork session, 2026-09-07: present on 25 of 176 assistant records, populated (`agent-pulse`), and set only while a skill is active. Measured twice on the **scheduled** Cowork entrypoint (`cowork/remote_cowork_trigger`), with opposite results: 2026-09-11 **absent**, so `skills_attributed` came back `{}` on a run that was otherwise green; 2026-10-01 **present on 192 of 193 turns** deduped per `requestId`. One entrypoint string, two readings, three weeks apart — so not even the entrypoint predicts the field, and "tracks the entrypoint, not the surface" is the right direction of travel rather than a rule that can be relied on to answer the question. And absent on Claude Code remote at any depth, where what exists instead is `attributionMcpServer` and `attributionMcpTool` — connector attribution, not skill attribution. One surface therefore behaves two ways depending on how the run was started, which is why an earlier reading of this as a Cowork-versus-Claude-Code distinction was wrong: it generalised a single interactive measurement to the whole surface. **Probe the key on the run in hand and record what was found**; never predict it from `surface`.

So the rule is three-part, and each part matters:

- **Populate it where the field is present**, counting turns per skill name.
- **Record `{}` where the field is absent**, and say which in `notes`.
- **Never treat absence as an error, and never assume presence.** A collector that raises on a missing key fails on one surface; a collector that assumes the key produces a confidently empty attribution on the other.

**Where the field *is* present, a populated `skills_attributed` is not a per-activity distribution — do not read it as one.** On the 2026-10-01 run, 187 of the 192 attributed turns carried `core-operating-model`, the first skill the run loaded through the `Skill` tool, while the run's actual work spanned reading `skill-ops-retro` and `convention-linear` and fanning twelve subagents across twelve unrelated skills; only 1 turn attributed to `task-pulse-retro` itself and 4 to this skill. The field reads as *which skill was most recently invoked through the `Skill` tool and is still live in context*, not *which skill this turn's work belongs to* — so on any long run that loads one skill early and then works elsewhere, the counts concentrate on that first skill and mean almost nothing about where the tokens went. Record the counts as the field gives them and **state this caveat in `notes` whenever one skill holds a large majority of the attributed turns**; never present the distribution as a cost split per skill, and never derive a per-skill cost figure from it. One measurement, and the run's own shape — one early `Skill` call, then dozens of unrelated turns — is a sufficient explanation for the concentration on its own, so this is a caveat on reading the field, not a finding about the harness: a differently-shaped run (one that never calls `Skill`, or that calls several evenly through the run) has not been measured. (Record `telemetry/2026-10-01-task-pulse-retro-beff961f-cffa-5965-9d2e-afaec4837f68.json`, `skills_attributed: {"task-pulse-retro": 1, "core-operating-model": 187, "skill-telemetry-capture": 4}` against 193 turns — APP-1937.)

Per-skill attribution is therefore partial by design in v1. Do not reconstruct it from `Skill` tool-use blocks as a silent substitute — a run that reads its canon with `Read` follows its skills and invokes none, so that route attributes nothing while looking like it worked. Widening attribution is a Phase 2 question, not a thing this collector improvises.

## Where it writes

Google Drive, folder `telemetry/` (`1RD_YuyjvHyOyDRXtLq9WtSHICP0n31ip`), one file per run:

`telemetry/{ISO-date}-{task-slug}-{session-id}.json`

Created, not appended, so concurrent runs cannot collide — the session id in the name guarantees that much.

**Search before creating.** Drive does not enforce unique titles: creating the same filename twice produces two files with identical titles and different ids, verified. Concurrency is safe; a **retry** is not. So the writer is idempotent on `session_id` — search the folder for the session id first, and update the existing file rather than creating a second. A duplicated record double-counts a run in every figure computed downstream.

## Deployment — in-band for now

The collector runs **in-band**, as the closing step of the run it measures (Owner decision, 2026-09-07, APP-1138). It is wired as a **loader footer** on each instrumented `task-*`, per `core-operating-model`'s loader rule: the task's body gains one final step naming this artefact and passing `task` and `outcome`, and no behaviour is added to any task config. APP-1131 instruments three tasks by hand — `task-pipeline-sweep`, `task-content-drain`, `task-owner-standup`; the fleet-wide rollout waits for `convention-telemetry` in Phase 2 rather than short-circuiting the canon route.

Measured cost: roughly 4,347 tokens — about 0.9% of the spike run's cache creation and 0.06% of its total throughput, plus about 1.3k where the Drive tool schema has to be loaded on demand. Immaterial against what it measures. If that ratio grows, re-measure before widening the rollout.

**The `Stop`-hook collector is the target, and a later phase.** A hook fires after the run ends, so it captures the complete transcript including the final turn, costs zero model tokens, and runs unconditionally — which removes the survivorship bias in the in-band shape, where a run that fails or is cut off produces no record at all and biases the dataset toward the failures most worth measuring. It is not adopted now for one reason: a hook cannot call an MCP connector, so writing to Drive needs a service credential on the container, which is an Owner ruling under `convention-secrets`. Hook availability on scheduled Cowork is also unknown. That ruling should not hold up a 0.9% cost, and on Cowork the hook now buys **correctness and not only cost**: the ~60% in-band under-read measured there (see *What this cannot capture*) is a hook-shaped defect, since a hook fires after the transcript is complete. It is still carried as a later phase, but the case for it is no longer an efficiency argument.

**Surface proof, stated honestly.** Proven end to end on scheduled Claude Code remote, and now on **scheduled Cowork** — the collector ran end to end there on 2026-09-11 and the Drive record read back byte-identical (record `2026-09-11-spike-app-1130-cowork-36ed6a33-….json`, surface `cowork/remote_cowork_trigger`). Green on the mechanism, with two measured caveats that are properties of the surface and not of the collector: the in-band turn capture under-reads by roughly 60%, and `attributionSkill` was **absent on that entrypoint on that run** — a reading the next measurement on the same entrypoint string contradicts (2026-10-01: **present on 192 of 193 turns**, record `telemetry/2026-10-01-task-pulse-retro-beff961f-cffa-5965-9d2e-afaec4837f68.json`, APP-1937). So the honest statement is that presence **varies run to run on `cowork/remote_cowork_trigger` and is probed per run, never predicted** — see *`skills_attributed` — present in the schema, absent on some surfaces*, which carries both readings and the caveat on how a populated field is read. The 2026-09-11 absence stands as measured; what does not stand is reading it as a property of the entrypoint. **Still unexercised on Cowork: the subagent sidechain fold-in** — that run delegated nothing, so the glob has not been proven against a real sidechain directory there. Untested leg, not a defect; say so rather than claiming the whole behaviour is proven.

## What this cannot capture

Two gaps, both structural, both to be stated rather than smoothed:

1. **Whatever the transcript has not yet flushed — a floor, not a rounding error, and its size is surface-specific.** The record is written by the run, so `finished_at`, `duration_s` and the output totals are always short by whatever the transcript does not yet contain. On Claude Code remote that is the run's own closing turns and the shortfall is small. On **Cowork the flush lags the live turn badly**: measured 2026-09-11 on a scheduled Cowork run, **5 turns captured against 13 actually run** — a ~60% under-read, which is not a caveat on the figures, it is the figures. So an in-band Cowork figure is read as a **floor on the run's cost, never as a baseline**, `notes` says which surface produced it and names the shortfall where the real turn count is knowable, and no comparison is drawn between an in-band Cowork figure and an in-band Claude Code one. Fixed only by the hook collector.
2. **Runs that never reached the capture step.** A run that failed, was cut off or skipped its footer produces no record.

Downstream, **an absent record reads as "run not captured", never as "run cost nothing"** — the dashboard must distinguish absent from zero (APP-1132), and any run with no record is accounted for against the scheduled-tasks list.

## Failure behaviour

**A capture failure never fails the run that produced it.** Every leg fails soft and the run continues to its own outcome:

- Transcript not found, or no session id in the environment: report one line, write nothing, finish.
- Transcript unreadable or malformed: write the record with the fields that did parse and name the gap in `notes`.
- Drive unreachable or the write rejected: emit the record inline in the run's output so it is recoverable by hand, and do not retry into a duplicate.
- Sidechain directory absent: proceed on the parent transcript and record that in `notes`, since the totals are then partial.

The host run's status, output and outcome are never changed by anything this skill does.

## Guardrails

- **Never fabricate a figure.** A model cannot know its own token counts; a self-reported record would invent the single number this project exists to drive down. A plausible wrong baseline is worse than none, so a surface that cannot read a transcript gets no record rather than a guessed one.
- **Never sum usage per record, and never deduplicate tool blocks by `requestId`.** Two rules, two objects; either one applied to the other silently corrupts the baseline.
- **Never treat an absent `attributionSkill` as an error, and never assume it is present.**
- **Never write two records for one session.** Search before create; a retry updates.
- **Never read another run's transcript**, and never carry transcript prose into the record. The record is counts, timings and tool names — not content.
- **Never block or alter the host run.**
- **Never claim proof on a surface this has not run on.** Report the surface in the record and the caveat in `notes`.

## Setup

Invoked, not scheduled. Wired as a loader footer on each instrumented `task-*` (APP-1131), reading `task` and `outcome` from the calling task. Published to `claude-ops` via `skill-ops-sync` — never a direct commit — so a Cloud Routine can clone it. Worked example: the spike record at `telemetry/2026-09-07-spike-app-1130-e254e044-….json`, written and read back byte-identical, which is the reference for schema v1. On finish, run `skill-ops-retro` capture on any friction.
