You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The workflow loader bug is real on main. In all three tm- workflows (tm-review-changes.js, tm-review-codebase.js, tm-map-codebase.js), export const meta is built from SPEC and TIER_MODELS, and const SPEC comes before it. The Workflow runtime needs meta to be a plain literal and the first statement in the file. The installed 2.9.0 plugin has the same code, so /tm-review-changes probably won't load as shipped. A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 and feat: verify stage in tm-review-changes with spec-format item transforms #419 got around it with a modified copy. Nothing outside the runtime reads meta, so the fix stays contained.
The token report's output figures. In a real lead transcript, output_tokens holds values in the hundreds to thousands. In a subagent transcript the values are 2 to 16, so they are placeholders. That backs up Decision 3 from batch Batch: measurement-tools #417. The price table has claude-sonnet-5 but not claude-sonnet-5-5.
Proposed batch: follow-ons-from-417 (3 packages, one wave)
1. Make each tm- workflow's meta a literal first statement
Scope. In each of the three workflow files, meta becomes a plain literal and the first statement. It holds the same content as today: name, description, and each phase's title, detail and model. A new test reads each file and checks three things:
meta is the first statement.
meta evaluates in an empty node:vm context, which proves it has no outside references.
meta deep-equals what SPEC produces today, so editing SPEC without meta fails CI.
The architect picks whether the literal copies SPEC or SPEC reads its phases from meta. The sync test is required either way. After the tester passes, the lead runs the branch's tm-review-changes.js once, by path and unmodified, on the PR's own diff. The last such run cost about $2 and took about 3.5 minutes. Replying "dispatch" approves that cost.
Acceptance criteria
In all three tm- workflow files, export const meta is the first statement and a plain literal, with no identifiers, calls or interpolation.
A new test fails on today's main and passes after the fix. It checks position and evaluates the literal in an empty vm context.
The test checks that each meta deep-equals the value SPEC produces (name, description, and each phase's title, detail and model).
The lead's live run of the unmodified branch file loads. The PR body records that it loaded, plus cost and wall-clock time.
npm test passes, and the plugin.json version is bumped.
Size: S. Track: full, because it touches .claude/workflows and executable code.
Non-goals
Changing any stage, prompt or SPEC content.
Live runs of tm-review-codebase and tm-map-codebase. Each fans out over the whole repo and costs much more. The refusal happens when the file loads, so the static test covers both.
2. Token report: price Sonnet 5.5 and use the lead's real output tokens
Scope. Sonnet 5.5 runs show up as "unpriced" and are left out of the total. #419 had to price them by hand. This package adds a claude-sonnet-5-5 entry to token-prices.json, read from the pricing page on the day of the change.
The report also estimates every agent's output as visible characters divided by 4. That leaves out thinking tokens and undercounts. The lead's transcript has the real usage.output_tokens, so the lead's rows use that number. Subagent transcripts only have placeholders, so their rows keep the estimate, and each row says whether its figure is measured or estimated.
Acceptance criteria
token-prices.json has a claude-sonnet-5-5 entry with all five rates, taken from the source URL, and retrieved is set to the check date.
A fixture test shows a Sonnet 5.5 record is priced and counted in the total.
Lead rows use the transcript's output_tokens, subagent rows keep the characters-divided-by-4 estimate, and a fixture test covers both.
The column header and notes say which figures are measured and which are estimated. The current line that says all output is estimated is replaced.
npm test passes, and the version is bumped.
Size: S. Track: full, because token-report.mjs is executable code under .claude/skills.
Non-goals
Re-pricing past batch reports.
Changing how wall-clock time is worked out.
Adding prices for models the policy doesn't pin.
Any other way of recovering subagent thinking tokens.
3. #407 (already filed): auto-must-fix severity floor in review prompts
Scope. The issue body stays as filed, with two changes:
I add one point to it: the floor sets severity, not whether a finding is true. So the critic prompt says it can't downgrade a floor finding, and the new verify stage can still reject one on the facts (for example, the "deleted test" was never deleted). The verify prompt doesn't change.
Acceptance criteria: as filed, plus:
The critic or consolidate prompt says a floor finding can't be downgraded.
Size: S. Track: full, because it touches .claude/agents and .claude/workflows.
Non-goals: as filed, plus no change to the verify prompt.
Merge note. All three packages bump plugin.json from 2.9.0, so the second and third merges each need a one-line version rebase. Packages 1 and 3 both edit the two review workflow files, but in different places (the top-of-file meta versus PROMPTS). Suggested order: 1, then 3, then 2.
The two trials (a second native-vs-kickoff pair, and /code-review versus the reviewer's quality pass):Record the headless trial traps from #400, #404 and #405 in tm-ab-test #418 set these up as the main work for this batch. I'm putting them after package 1 because the native-vs-kickoff judge is tm-review-changes, which currently loads only from a modified copy. That would add a known deviation to an expensive headless run. The /code-review trial also has no protocol yet. If you'd rather run them now with the workaround, say "revise".
Approved by the owner on 2026-10-01 (sign-off: dispatch). The proposal below is the contract, verbatim.
Packages
metaa literal first statement (PR fix: make each workflow meta a literal first statement #426 merged, 9c6669d)What I checked before slicing
tm-review-changes.js,tm-review-codebase.js,tm-map-codebase.js),export const metais built fromSPECandTIER_MODELS, andconst SPECcomes before it. The Workflow runtime needsmetato be a plain literal and the first statement in the file. The installed 2.9.0 plugin has the same code, so/tm-review-changesprobably won't load as shipped. A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 and feat: verify stage in tm-review-changes with spec-format item transforms #419 got around it with a modified copy. Nothing outside the runtime readsmeta, so the fix stays contained.output_tokensholds values in the hundreds to thousands. In a subagent transcript the values are 2 to 16, so they are placeholders. That backs up Decision 3 from batch Batch: measurement-tools #417. The price table hasclaude-sonnet-5but notclaude-sonnet-5-5.Proposed batch:
follow-ons-from-417(3 packages, one wave)1. Make each tm- workflow's
metaa literal first statementScope. In each of the three workflow files,
metabecomes a plain literal and the first statement. It holds the same content as today: name, description, and each phase's title, detail and model. A new test reads each file and checks three things:metais the first statement.metaevaluates in an emptynode:vmcontext, which proves it has no outside references.metadeep-equals whatSPECproduces today, so editingSPECwithoutmetafails CI.The architect picks whether the literal copies
SPECorSPECreads its phases frommeta. The sync test is required either way. After the tester passes, the lead runs the branch'stm-review-changes.jsonce, by path and unmodified, on the PR's own diff. The last such run cost about $2 and took about 3.5 minutes. Replying "dispatch" approves that cost.Acceptance criteria
export const metais the first statement and a plain literal, with no identifiers, calls or interpolation.vmcontext.metadeep-equals the valueSPECproduces (name, description, and each phase's title, detail and model).npm testpasses, and theplugin.jsonversion is bumped.Size: S. Track: full, because it touches
.claude/workflowsand executable code.Non-goals
SPECcontent.tm-review-codebaseandtm-map-codebase. Each fans out over the whole repo and costs much more. The refusal happens when the file loads, so the static test covers both.2. Token report: price Sonnet 5.5 and use the lead's real output tokens
Scope. Sonnet 5.5 runs show up as "unpriced" and are left out of the total. #419 had to price them by hand. This package adds a
claude-sonnet-5-5entry totoken-prices.json, read from the pricing page on the day of the change.The report also estimates every agent's output as visible characters divided by 4. That leaves out thinking tokens and undercounts. The lead's transcript has the real
usage.output_tokens, so the lead's rows use that number. Subagent transcripts only have placeholders, so their rows keep the estimate, and each row says whether its figure is measured or estimated.Acceptance criteria
token-prices.jsonhas aclaude-sonnet-5-5entry with all five rates, taken from the source URL, andretrievedis set to the check date.output_tokens, subagent rows keep the characters-divided-by-4 estimate, and a fixture test covers both.npm testpasses, and the version is bumped.Size: S. Track: full, because
token-report.mjsis executable code under.claude/skills.Non-goals
3. #407 (already filed): auto-must-fix severity floor in review prompts
Scope. The issue body stays as filed, with two changes:
Acceptance criteria: as filed, plus:
Size: S. Track: full, because it touches
.claude/agentsand.claude/workflows.Non-goals: as filed, plus no change to the verify prompt.
Merge note. All three packages bump
plugin.jsonfrom 2.9.0, so the second and third merges each need a one-line version rebase. Packages 1 and 3 both edit the two review workflow files, but in different places (the top-of-filemetaversusPROMPTS). Suggested order: 1, then 3, then 2.Deferred, and why
tm-review-codebase): next batch. Its new Verify phase goes into themetathat package 1 rewrites, and its live-run criterion needs/tm-review-codebaseto load. I'd also fold in the leftoverfinalizeReportguard-test nit from feat: verify stage in tm-review-changes with spec-format item transforms #419, since it's the same code path./code-reviewversus the reviewer's quality pass): Record the headless trial traps from #400, #404 and #405 in tm-ab-test #418 set these up as the main work for this batch. I'm putting them after package 1 because the native-vs-kickoff judge istm-review-changes, which currently loads only from a modified copy. That would add a known deviation to an expensive headless run. The/code-reviewtrial also has no protocol yet. If you'd rather run them now with the workaround, say "revise".Things only you can do (not packages)
Decision log
Decisions made during the run are posted as comments on this issue.