Skip to content

Token report: include workflow subagents and add a per-agent wall-clock table #416

Description

@sv-tmueller

Part of batch #417

Rescoped at batch #417 sign-off (2026-10-01): the tokens-per-second column is dropped in favor of measured wall-clock per agent. The original body is in the issue edit history.

What to build

The token report posted at every /tm-advisor batch report and every /tm-kickoff wave end (.claude/skills/tm-kickoff/token-report.mjs) should show what a run really spent, and how long each agent took.

Completeness (bug). The report reads the lead transcript and <session>/subagents/agent-*.jsonl only. Agents spawned inside a Workflow keep their transcripts under subagents/workflows/<wf-id>/, which the report never reads, so a workflow's whole cost is left out without a warning. Found in #405 (docs/reviews/2026-09-28-ultracode-arm-379.md, sections 4 and 7): the merged script priced the arm at $4.28 against a measured $12.77, and the tm-review-changes run in that PR's kickoff session was missing from its report.

Per-agent wall-clock (new). For each agent (the lead, each dispatched seat, each workflow subagent), add a row with its wall-clock time. That shows which seats are slow, and whether a model or effort change sped them up.

Why wall-clock and not tokens per second, checked on 2026-10-01 against batch #403's transcripts:

  • Interactive transcripts carry only streaming placeholders for output_tokens (values 5, 8 and 16 across 32 messages in one subagent transcript, with no later line per message id holding a larger count). The only output figure is the chars/4 estimate, which leaves out thinking and ran 3.8-4.4x low in A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 (docs/reviews/2026-09-28-ab-native-vs-kickoff.md, "Cost-measurement gap"). Thinking is what an effort change moves, so a rate built on that estimate cannot answer the question.
  • The lead transcript's task notifications carry a measured <duration_ms> per dispatched agent, alongside <tool_uses> and <subagent_tokens>. subagent_tokens looks like the agent's final context size (319,875 for an agent with 5 tool uses over 60 s), not its output; do not present it as output.
  • Workflow subagents have no task notification of their own; the first-to-last timestamp span in each transcript is a candidate source. Check it against the workflow's own reported duration where one exists.

Acceptance criteria

  • The report includes workflow subagents under subagents/workflows/, grouped by role and by model like every other agent. A fixture test covers a session with a workflow.
  • The report names every subagent transcript directory it found but could not read, instead of leaving it out without a word.
  • A per-agent table lists role, model, calls, tool uses, estimated output tokens, and wall-clock time for each agent (the lead, each dispatched seat, each workflow subagent).
  • The report states where each wall-clock value comes from (for example, duration_ms in a task notification, or the first-to-last timestamp span in a transcript), and marks any estimated value.
  • Tests are written first, against fixture transcripts. They include an agent with no output and an agent with a zero or missing duration, which shows as "n/a" (no division by zero, no made-up value).
  • /tm-advisor section 5 and /tm-kickoff's wave-end report still post the report on every run, with the new table.
  • npm test is green, and .claude/.claude-plugin/plugin.json gets a version bump (the change touches .claude/skills/).

Non-goals

  • A tokens-per-second column.
  • Pricing headless claude -p child sessions started by a seat (for example A/B trial runs). Those stay reported by the package that ran them.
  • Real billing figures. The report stays at list prices.
  • Live or streaming throughput during a run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    size:M1 to 3 hours. Write a sub-plan first.type:feat

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions