Skip to content

feat: token report covers workflow subagents and adds a per-agent wall-clock table - #421

Merged
sv-tmueller merged 3 commits into
mainfrom
feat/416-token-report-workflows
Oct 1, 2026
Merged

sv-tmueller merged 3 commits into
mainfrom
feat/416-token-report-workflows

Conversation

@sv-tmueller

@sv-tmueller sv-tmueller commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Closes #416

Part of batch #417. Follows the architect's sub-plan on #416.

What changes

  • token-report.mjs reads workflow subagent transcripts under subagents/workflows/<runId>/. They count under the role workflow:<workflowName> (the Workflow launch's name, else workflow:<runId>), in the by-role and by-model tables. Model comes from each call, never from the sidecar alias.
  • New pure export perAgent() and a ## By agent table: Start (UTC), Role (with the agent() label for workflow agents), Model, Calls, Tool uses, Output (est.), Wall-clock, Source. No agent hex ids appear in the output.
  • Wall-clock sources: a dispatched seat uses the sum of duration_ms from task notifications when it has at least one and all are positive (copies in the three notification carriers are merged, tool-result text is never read). Otherwise it uses the sum of per-run transcript spans, marked (est.); a coordinator line starts a new run, so a resumed seat does not count the idle gap. Workflow subagents use their span (est.). The lead uses its first-to-last span in the window (est., includes idle time). No positive value gives n/a.
  • Paths under subagents/ that were found but could not be read (unknown directories, readdir or readFile errors) are named in Limitations and left out of every total.
  • render() prints the new table and lines only when there is something to show, so lead-dedupe.snapshot.md is unchanged. Neither SKILL.md needed an edit.
  • plugin.json 2.7.0 to 2.8.0. Other batch PRs also bump it; that conflict is for merge time.

Verification

  • npm test exits 0 (415 tests). Fixtures are synthetic only (fixtures/token-report/cli-config/projects/test-project/workflow-session*). Tests were written first and failed on the missing perAgent export before the implementation.
  • Read-only smoke run on a real kickoff session with a tm-review-changes run: 1 workflow:tm-review-changes role row, 8 workflow agent rows, 16 agent rows in total, nothing found-but-unread, no agent hex in the output. Nothing from the real transcripts is in the repo.

No new dependencies.

🤖 Generated with Claude Code

sv-tmueller and others added 2 commits October 1, 2026 10:24
…the token report

Red first: fixtures for a synthetic workflow session and failing tests for
perAgent, the By agent table and the unread-path line (issue #416).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…lock table

Workflow subagent transcripts under subagents/workflows/<run>/ were never
read, so a workflow's whole cost was missing from the report. They now count
under a workflow:<name> role. A new By agent table shows calls, tool uses and
wall-clock per agent, from measured task-notification durations where they
exist and from marked transcript spans otherwise. Paths that were found but
not read are named in Limitations. Bumps the plugin to 2.8.0.

Closes #416

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Tester report, fix round 1/3 (batch #417)

VERDICT: FAIL
COMMIT: f4cb270
FINDINGS:

  1. Agent hex id leaks into the posted report. When a subagent transcript cannot be read, the "Found but not read" Limitations line prints its relative path, e.g. subagents/agent-aaaa0001.jsonl and subagents/workflows/wf_fixture-001/agent-a0a00001.jsonl. The sub-plan and issue require no agent hex anywhere in the rendered output, since it is posted to public issues. The read-error test only checks that a run directory is named, so this is not pinned.
    Repro: S=<copy of fixtures cli-config>/projects/test-project/workflow-session/subagents; chmod 000 $S/agent-aaaa0001.jsonl; node .claude/skills/tm-kickoff/token-report.mjs --session workflow-session --config-dir <copy> (observed: hex in the Limitations line; expected: a hex-free path or placeholder).
  2. A workflow sidecar description containing a newline splits the table row, because only | is escaped. Low severity.
    Repro: write {"description":"a\nb|c"} into a wf_fixture-001 agent-*.meta.json in a fixture copy and run the CLI as above (observed: the row breaks across two lines; expected: one row).
    UNTESTED CLAIMS: No test names a transcript-read failure without exposing its hex (finding 1). The dispatch's "workflow-agent wall-clock sum close to 650,996 ms" does not hold as a literal sum on real data. The 8 rows sum to about 1,512 s because the agents run in parallel. Their union (first start to last end) is about 649 s, which matches the workflow's 650,996 ms and what the sub-plan says ("together cover"). Not a defect, but the check as worded in the dispatch is wrong.
    LESSONS: A "never print identifier X" rule needs a test per output path (unread paths, labels, errors), not only the main table.

Lead note: the dispatch wording in the untested-claims line was the lead's error (a sum of parallel spans, where the sub-plan meant their union). Not a finding against the PR.

An unreadable transcript is now named by its directory and a count,
and any agent-<id> in an unread entry is redacted, since the report is
posted to public issues. A newline in a workflow label no longer splits
its table row.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Tester report, re-test after tester fix round 1/3 (batch #417)

VERDICT: PASS
COMMIT: a8e2393
FINDINGS: none
UNTESTED CLAIMS: The sidecar description text itself is not redacted. A label containing the text agent-<hex> prints as-is in the By agent table (observed on a /tmp fixture copy). Labels are author-chosen agent() names, so I did not count this as a finding. No test covers it.

@sv-tmueller sv-tmueller mentioned this pull request Oct 1, 2026
4 of 19 tasks
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Reviewer report (batch #417)

VERDICT: APPROVE
STAGE: both clean
FINDINGS:

  1. .claude/skills/tm-kickoff/token-report.mjs:10, nit: the header comment still lists the pure exports as "(parse, aggregate, price, render)", so it leaves out the new perAgent export. Fix: change the line to // Exports are pure (parse, aggregate, perAgent, price, render); main() is the only.
    Evidence: line 10 against the new export function perAgent at about line 330.
  2. .claude/skills/tm-kickoff/token-report.mjs:299, nit: in transcriptTiming, start and first are both Math.min over the same timestamps, so first is redundant. Fix: delete let first = Infinity (line 299) and first = Math.min(first, t) (line 308), and change line 318 to return { start, spanMs: last - start, runSpanMs }.
    Evidence: lines 298-308, where both variables are updated the same way.
    CHECKS: n/a

@sv-tmueller
sv-tmueller marked this pull request as ready for review October 1, 2026 08:40
@sv-tmueller
sv-tmueller merged commit 74fc2b0 into main Oct 1, 2026
2 checks passed
@sv-tmueller
sv-tmueller deleted the feat/416-token-report-workflows branch October 1, 2026 08:54
sv-tmueller added a commit that referenced this pull request Oct 1, 2026
PR #421 merged first with 2.8.0, so this branch's 2.7.1 bump no longer
moves the version forward. Resolve to the next patch level.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sv-tmueller added a commit that referenced this pull request Oct 1, 2026
PRs #421 and #420 merged first and took 2.8.0 and 2.8.1. This branch adds
a workflow stage, so it takes the next minor version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token report: include workflow subagents and add a per-agent wall-clock table

1 participant