Skip to content

fix: token report prices Sonnet 5.5 and uses the lead's real output tokens - #427

Merged
sv-tmueller merged 3 commits into
mainfrom
fix/425-token-report-pricing
Oct 1, 2026
Merged

sv-tmueller merged 3 commits into
mainfrom
fix/425-token-report-pricing

Conversation

@sv-tmueller

@sv-tmueller sv-tmueller commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Closes #425

Part of batch #423.

The token report priced claude-sonnet-5-5 as unpriced and estimated the lead's output as visible chars / 4, though lead transcripts carry the real count.

  • Price table: add claude-sonnet-5-5 ($2 / $2.50 / $4 / $0.20 / $10 per MTok). Re-read the pricing page on 2026-10-01: the existing claude-opus-5-5 and claude-sonnet-5 rates are unchanged; retrieved is now 2026-10-01.
  • Parse: the lead's usage.output_tokens is read from the first line of each turn (every line repeats the final count). A lead turn with no count falls back to chars / 4 and is marked. Subagent output stays a chars / 4 estimate.
  • Render: header is "Output (est. where marked)"; estimated cells (and any by-model or total figure that includes an estimate) carry " (est.)". Limitations text updated.
  • Doc: one sentence in tm-ab-test/SKILL.md was stale and is corrected.
  • Version bump 2.9.0 to 2.9.1 (patch, a fix).
  • No new dependencies.

Checks: npm test passes (478 tests). Smoke run against a real recent lead session: lead row unmarked and in the thousands, claude-sonnet-5-5 priced, subagent rows marked (est.).

🤖 Generated with Claude Code

sv-tmueller and others added 2 commits October 1, 2026 13:01
The lead transcript carries real output_tokens and the price table lacks
claude-sonnet-5-5. Failing tests first.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e token report

Lead transcripts repeat the final output_tokens on every line of a turn, so the first line is safe to read. Subagent transcripts do not, so those stay a chars/4 estimate, now marked (est.) per cell. Adds claude-sonnet-5-5 to the price table (rates re-read from the pricing page 2026-10-01).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Tester report (round 1)

VERDICT: PASS
COMMIT: 46cc9ca
FINDINGS: none
UNTESTED CLAIMS:

  1. A bucket stays marked "(est.)" when a measured record arrives after an estimated one. Mutating bucket.outputEstimated ||= record.outputEstimated to = in addToBucket leaves all 478 tests green. The mixed by-model test only has the estimated record last. The shipped code is correct. Repro: change that one line in a copy of the repo and run npm test.
  2. Cosmetic: the edited tm-ab-test/SKILL.md line 242-244 runs past the wrapped width after "output, which it now reads from usage.output_tokens)." Only that sentence changed.
    Checked and clean:
  • npm test gave 478 pass, 0 fail.
  • node scripts/check-version-bump.mjs origin/main HEAD exited 0 (2.9.0 -> 2.9.1).
  • Sonnet 5.5 rates match the live pricing page: 2 input, 2.50 5m write, 4 1h write, 0.20 cache read, 10 output. The Opus 5.5 rates (4/5/8/0.2/20) and Sonnet 5 rates are unchanged.
  • The snapshot diff changes only Output headers and cells, the limitations notes, retrieved, and the lead-dedupe Cost cells. No Input, Cache or Wall-clock change, and no .jsonl changes.
  • parse() passes isLead explicitly. A lead line with no output_tokens falls back to the estimate and is marked "(est.)". Total and by-model rows are marked when any part is estimated.
  • The aggregate record() helpers now carry outputTokens and outputEstimated. 9 of 10 mutations were caught, including the per-role sum in both addToBucket and perAgent, the lead flag, the estimate fallback and the Total and by-model render. The one survivor is item 1.
  • My smoke run matches the developer's figures: lead row 160,256 unmarked, claude-sonnet-5-5 priced, subagent rows "(est.)", total $29.14.

@sv-tmueller

Copy link
Copy Markdown
Owner Author

Reviewer report (round 1)

VERDICT: APPROVE
STAGE: quality
FINDINGS:

  1. .claude/workflows/tests/token-report.test.mjs:234 (should-fix): no test catches it if bucket.outputEstimated ||= record.outputEstimated in addToBucket turns into =, so tester claim 1 holds. In every fixture, including both snapshots, an estimated record comes after the measured ones in its bucket, because parse() puts lead records first. A bucket that gets an estimated lead turn (no output_tokens) and then a measured one would lose its "(est.)" marker without any test failing. Fix: add an aggregate case where an estimated record comes before a measured record in the same bucket, and assert the bucket stays estimated.
    Evidence: token-report.mjs:215; in the records fixture (test.mjs:195-201), orphan-1, the estimated record, comes last in both the opus bucket and the totals.
  2. .claude/skills/tm-ab-test/SKILL.md:244 (nit): the edited line is 112 characters. The lines around it wrap at about 76, so tester claim 2 holds. Fix: rewrap lines 242-245 to the file's width.
    Evidence: line 244 is "output, which it now reads from usage.output_tokens). Measured output was 3.76x the estimate on A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 arm A".
    CHECKS: n/a

Resolve the plugin.json version to 2.10.1, above main's 2.10.0 from #428.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sv-tmueller
sv-tmueller merged commit cd630c1 into main Oct 1, 2026
2 checks passed
@sv-tmueller
sv-tmueller deleted the fix/425-token-report-pricing branch October 1, 2026 11:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token report: price Sonnet 5.5 and use the lead's real output tokens

1 participant