You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Token report: include workflow subagents and add a per-agent wall-clock table #416
Rescoped at batch #417 sign-off (2026-10-01): the tokens-per-second column is dropped in favor of measured wall-clock per agent. The original body is in the issue edit history.
What to build
The token report posted at every /tm-advisor batch report and every /tm-kickoff wave end (.claude/skills/tm-kickoff/token-report.mjs) should show what a run really spent, and how long each agent took.
Completeness (bug). The report reads the lead transcript and <session>/subagents/agent-*.jsonl only. Agents spawned inside a Workflow keep their transcripts under subagents/workflows/<wf-id>/, which the report never reads, so a workflow's whole cost is left out without a warning. Found in #405 (docs/reviews/2026-09-28-ultracode-arm-379.md, sections 4 and 7): the merged script priced the arm at $4.28 against a measured $12.77, and the tm-review-changes run in that PR's kickoff session was missing from its report.
Per-agent wall-clock (new). For each agent (the lead, each dispatched seat, each workflow subagent), add a row with its wall-clock time. That shows which seats are slow, and whether a model or effort change sped them up.
Why wall-clock and not tokens per second, checked on 2026-10-01 against batch #403's transcripts:
Interactive transcripts carry only streaming placeholders for output_tokens (values 5, 8 and 16 across 32 messages in one subagent transcript, with no later line per message id holding a larger count). The only output figure is the chars/4 estimate, which leaves out thinking and ran 3.8-4.4x low in A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 (docs/reviews/2026-09-28-ab-native-vs-kickoff.md, "Cost-measurement gap"). Thinking is what an effort change moves, so a rate built on that estimate cannot answer the question.
The lead transcript's task notifications carry a measured <duration_ms> per dispatched agent, alongside <tool_uses> and <subagent_tokens>. subagent_tokens looks like the agent's final context size (319,875 for an agent with 5 tool uses over 60 s), not its output; do not present it as output.
Workflow subagents have no task notification of their own; the first-to-last timestamp span in each transcript is a candidate source. Check it against the workflow's own reported duration where one exists.
Acceptance criteria
The report includes workflow subagents under subagents/workflows/, grouped by role and by model like every other agent. A fixture test covers a session with a workflow.
The report names every subagent transcript directory it found but could not read, instead of leaving it out without a word.
A per-agent table lists role, model, calls, tool uses, estimated output tokens, and wall-clock time for each agent (the lead, each dispatched seat, each workflow subagent).
The report states where each wall-clock value comes from (for example, duration_ms in a task notification, or the first-to-last timestamp span in a transcript), and marks any estimated value.
Tests are written first, against fixture transcripts. They include an agent with no output and an agent with a zero or missing duration, which shows as "n/a" (no division by zero, no made-up value).
/tm-advisor section 5 and /tm-kickoff's wave-end report still post the report on every run, with the new table.
npm test is green, and .claude/.claude-plugin/plugin.json gets a version bump (the change touches .claude/skills/).
Non-goals
A tokens-per-second column.
Pricing headless claude -p child sessions started by a seat (for example A/B trial runs). Those stay reported by the package that ran them.
Real billing figures. The report stays at list prices.
Part of batch #417
Rescoped at batch #417 sign-off (2026-10-01): the tokens-per-second column is dropped in favor of measured wall-clock per agent. The original body is in the issue edit history.
What to build
The token report posted at every
/tm-advisorbatch report and every/tm-kickoffwave end (.claude/skills/tm-kickoff/token-report.mjs) should show what a run really spent, and how long each agent took.Completeness (bug). The report reads the lead transcript and
<session>/subagents/agent-*.jsonlonly. Agents spawned inside a Workflow keep their transcripts undersubagents/workflows/<wf-id>/, which the report never reads, so a workflow's whole cost is left out without a warning. Found in #405 (docs/reviews/2026-09-28-ultracode-arm-379.md, sections 4 and 7): the merged script priced the arm at $4.28 against a measured $12.77, and thetm-review-changesrun in that PR's kickoff session was missing from its report.Per-agent wall-clock (new). For each agent (the lead, each dispatched seat, each workflow subagent), add a row with its wall-clock time. That shows which seats are slow, and whether a model or effort change sped them up.
Why wall-clock and not tokens per second, checked on 2026-10-01 against batch #403's transcripts:
output_tokens(values 5, 8 and 16 across 32 messages in one subagent transcript, with no later line per message id holding a larger count). The only output figure is the chars/4 estimate, which leaves out thinking and ran 3.8-4.4x low in A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 (docs/reviews/2026-09-28-ab-native-vs-kickoff.md, "Cost-measurement gap"). Thinking is what an effort change moves, so a rate built on that estimate cannot answer the question.<duration_ms>per dispatched agent, alongside<tool_uses>and<subagent_tokens>.subagent_tokenslooks like the agent's final context size (319,875 for an agent with 5 tool uses over 60 s), not its output; do not present it as output.Acceptance criteria
subagents/workflows/, grouped by role and by model like every other agent. A fixture test covers a session with a workflow.duration_msin a task notification, or the first-to-last timestamp span in a transcript), and marks any estimated value./tm-advisorsection 5 and/tm-kickoff's wave-end report still post the report on every run, with the new table.npm testis green, and.claude/.claude-plugin/plugin.jsongets a version bump (the change touches.claude/skills/).Non-goals
claude -pchild sessions started by a seat (for example A/B trial runs). Those stay reported by the package that ran them.