Skip to content

Batch: follow-ons-from-417 #423

Description

@sv-tmueller

Approved by the owner on 2026-10-01 (sign-off: dispatch). The proposal below is the contract, verbatim.

Packages


What I checked before slicing

  • The workflow loader bug is real on main. In all three tm- workflows (tm-review-changes.js, tm-review-codebase.js, tm-map-codebase.js), export const meta is built from SPEC and TIER_MODELS, and const SPEC comes before it. The Workflow runtime needs meta to be a plain literal and the first statement in the file. The installed 2.9.0 plugin has the same code, so /tm-review-changes probably won't load as shipped. A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 and feat: verify stage in tm-review-changes with spec-format item transforms #419 got around it with a modified copy. Nothing outside the runtime reads meta, so the fix stays contained.
  • The token report's output figures. In a real lead transcript, output_tokens holds values in the hundreds to thousands. In a subagent transcript the values are 2 to 16, so they are placeholders. That backs up Decision 3 from batch Batch: measurement-tools #417. The price table has claude-sonnet-5 but not claude-sonnet-5-5.

Proposed batch: follow-ons-from-417 (3 packages, one wave)

1. Make each tm- workflow's meta a literal first statement

Scope. In each of the three workflow files, meta becomes a plain literal and the first statement. It holds the same content as today: name, description, and each phase's title, detail and model. A new test reads each file and checks three things:

  • meta is the first statement.
  • meta evaluates in an empty node:vm context, which proves it has no outside references.
  • meta deep-equals what SPEC produces today, so editing SPEC without meta fails CI.

The architect picks whether the literal copies SPEC or SPEC reads its phases from meta. The sync test is required either way. After the tester passes, the lead runs the branch's tm-review-changes.js once, by path and unmodified, on the PR's own diff. The last such run cost about $2 and took about 3.5 minutes. Replying "dispatch" approves that cost.

Acceptance criteria

  • In all three tm- workflow files, export const meta is the first statement and a plain literal, with no identifiers, calls or interpolation.
  • A new test fails on today's main and passes after the fix. It checks position and evaluates the literal in an empty vm context.
  • The test checks that each meta deep-equals the value SPEC produces (name, description, and each phase's title, detail and model).
  • The lead's live run of the unmodified branch file loads. The PR body records that it loaded, plus cost and wall-clock time.
  • npm test passes, and the plugin.json version is bumped.

Size: S. Track: full, because it touches .claude/workflows and executable code.

Non-goals

  • Changing any stage, prompt or SPEC content.
  • Live runs of tm-review-codebase and tm-map-codebase. Each fans out over the whole repo and costs much more. The refusal happens when the file loads, so the static test covers both.
  • Porting the verify stage (review agent orchestration #344 P2).

2. Token report: price Sonnet 5.5 and use the lead's real output tokens

Scope. Sonnet 5.5 runs show up as "unpriced" and are left out of the total. #419 had to price them by hand. This package adds a claude-sonnet-5-5 entry to token-prices.json, read from the pricing page on the day of the change.

The report also estimates every agent's output as visible characters divided by 4. That leaves out thinking tokens and undercounts. The lead's transcript has the real usage.output_tokens, so the lead's rows use that number. Subagent transcripts only have placeholders, so their rows keep the estimate, and each row says whether its figure is measured or estimated.

Acceptance criteria

  • token-prices.json has a claude-sonnet-5-5 entry with all five rates, taken from the source URL, and retrieved is set to the check date.
  • A fixture test shows a Sonnet 5.5 record is priced and counted in the total.
  • Lead rows use the transcript's output_tokens, subagent rows keep the characters-divided-by-4 estimate, and a fixture test covers both.
  • The column header and notes say which figures are measured and which are estimated. The current line that says all output is estimated is replaced.
  • npm test passes, and the version is bumped.

Size: S. Track: full, because token-report.mjs is executable code under .claude/skills.

Non-goals

  • Re-pricing past batch reports.
  • Changing how wall-clock time is worked out.
  • Adding prices for models the policy doesn't pin.
  • Any other way of recovering subagent thinking tokens.

3. #407 (already filed): auto-must-fix severity floor in review prompts

Scope. The issue body stays as filed, with two changes:

Acceptance criteria: as filed, plus:

  • The critic or consolidate prompt says a floor finding can't be downgraded.

Size: S. Track: full, because it touches .claude/agents and .claude/workflows.

Non-goals: as filed, plus no change to the verify prompt.

Merge note. All three packages bump plugin.json from 2.9.0, so the second and third merges each need a one-line version rebase. Packages 1 and 3 both edit the two review workflow files, but in different places (the top-of-file meta versus PROMPTS). Suggested order: 1, then 3, then 2.

Deferred, and why

Things only you can do (not packages)


Decision log

Decisions made during the run are posted as comments on this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions