feat: verify stage in tm-review-changes with spec-format item transforms - #419
Conversation
…to renderers Both renderers looked up args["args.areas"], so the cap was always the default. Strip the args. prefix and coerce like MAX_AREAS. Add a closed set of item reducers (must_fix_deduped first) that run before the cap, an unknown name throws, and overflow is logged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A finding reached the critic from worker text alone. Verify each must-fix finding with an adversarial worker first. A dead verifier means unverified, not refuted, and the script owns the refuted and unverified lists so the report does not depend on model behavior. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Add end-to-end tests for refuted, over-cap and dead-verifier handling, and the render-path stub that makes the verify stage fire. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Tester report (initial test, batch #417)VERDICT: PASS
|
Reviewer report (batch #417)VERDICT: APPROVE
Lead note: nits 5 and 6 are PR-body text, not part of the diff; the lead applied both verbatim to the PR body. Findings 1-4 are left for the human review. |
The critic's verdict survived when finalizeReport removed all of its must-fix findings as refuted, so a report could read mustFix: [] with changes-requested. The verdict is what humans and the ab-test template read, so the script now sets approve in that case. Unverified findings are never removed, so they still block. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Reviewer fix round 1/3 (owner-requested before merge, batch #417): finding 1 (should-fix) applied in 69800c2 with the reviewer's exact change and two |
Tester report, re-test after review fix round 1/3 (batch #417)VERDICT: PASS
|
Reviewer report, re-review after fix round 1/3 (batch #417)VERDICT: APPROVE
Lead note: nit 2 (PR body) applied verbatim by the lead before merge, with (a) placed in the |
The Workflow runtime refuses to load a script whose meta is not a plain literal and its first statement. All three tm- workflows built meta from SPEC and TIER_MODELS below those constants, so the runtime refused them by name and by path (#405, #419). meta is now a literal at byte 0, copied from what the old expression produced. A new test fails if it moves, stops being a plain literal, or drifts from SPEC. Live run: the branch's tm-review-changes.js loaded unmodified by path on the first try (wf_170be40a-a33). Closes #424 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Closes #406
Part of batch #417 and of #344 (package P1).
What changed
items_transformfield naming a reducer from a closed set that each renderer implements. First entry:must_fix_deduped(flatten worker findings, keep must-fix, dedup on file + line + problem). The cap stays the one generic step and runs after the transform. An unknown name throws in both renderers, so a host without the reducer fails loudly instead of fanning out over unreduced items.items_capbug: both renderers looked upargs["args.areas"], so the cap was always the default. Fixed (strip theargs.prefix, coerce likeMAX_AREAS).tm-review-changes: newverifystage betweenreviewandconsolidate, worker tier, one adversarial verifier per must-fix finding (schema{ confirmed, note }), capped byargs.maxVerify(default 12). Consolidate gets confirmed, refuted and unverified findings as separate inputs.refutedandunverifiedarrays and strips refuted findings frommustFix(finalizeReport), and sets the verdict to approve when that strip emptiesmustFix, so this does not depend on model behavior.adapter-interface.md, README node label,tm-review-changesskill description, team-guide-rationale.md worked-example line.Report-level overflow AC (Claude Code runtime only)
The AC "findings past the cap appear in the report" is met by the Claude Code runtime (
tm-review-changes.js), the only renderer that builds consolidate inputs from runtime data. Hermes stubs those slots for every workflow, and the Codex renderer builds no consolidate prompt from stage results. In the Hermes and Codex renderers, overflow is a log line, like the existing stub-fallback log.Live run (lead, after tester PASS at 5741657)
tm-review-changesrun on this PR's own diff, 2026-10-01, run idwf_613b6f09-7e3,args: { base: "origin/main" }, with the lead checkout detached at57416579for the run and HEAD checked unchanged afterward (batch Batch: measurement-tools #417 Decision 5)..claude/workflows/tm-review-changes.jsby path ("export const meta = { name, description, phases }must be the FIRST statement in the script"), the same computed-metarefusal A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 recorded. The run used a scratch copy withmetaevaluated once and written as a literal first statement;diffagainst the branch file shows no other change (batch Batch: measurement-tools #417 Decision 2).verify:*dispatch, which is the D10 behavior for zero must-fix. The verify path is exercised by the tests on this PR, not by this live run.token-report.mjsfrom PR feat: token report covers workflow subagents and adds a per-agent wall-clock table #421's branch ata8e2393(unmerged, run read-only) against a scratch copy holding only this workflow's 8 transcripts. That script's price table has noclaude-sonnet-5-5entry, so the Sonnet figure is priced by hand at the published Sonnet 5.5 list rates (https://platform.claude.com/docs/en/about-claude/pricing, retrieved 2026-10-01: $2 input, $2.50 5m cache write, $4 1h cache write, $0.20 cache read, $10 output per MTok). Output tokens are that script's chars/4 estimate and exclude thinking, so the figure undercounts.duration_ms. On this diff that is well under the 600 s headless background-task ceiling; a run with must-fix findings adds the verify phase.finalizeReportremoves refuted findings fromreport.mustFixbut keeps the critic's verdict, so a report can readmustFix: []with verdictchanges-requested. The strip also matches exact file + line + problem, so a critic that rewords a refuted finding slips past it. The run's suggested fix: derive the verdict from the finalmustFixlength, with a test where the critic echoes only refuted findings. Fixed in 69800c2 (reviewer fix round 1/3): the verdict is recomputed to approve when only refuted findings were removed. The reworded-finding half stays out of scope (D2 specifies an exact key).REPORT_SCHEMAstill offersrefutedandunverified; the "Bounded by construction" comment reflow; the workflow's ownmaxVerifycoercion is tested only with 1;maxVerifyis not in the invoke comment or SKILL.md; the verify prompt does not mark the finding text as untrusted data;ITEM_TRANSFORMSis duplicated across the two renderers (guarded by the parametrized test).Verification
npm test(467 tests at 93bb0b1, 0 failures) andnode scripts/check-version-bump.mjs origin/main HEADboth pass locally.New tests:
item-transforms.test.mjs(both renderers, filter, dedup, cap, overflow log, unknown name throws),review-changes-verify.test.mjs(refuted, over-cap and dead-verifier handling end to end),hermes-verify-stage.test.mjs, plus additions tohermes-adapter,helpers,prompts-syncandrender-pathtests. Theitems_captest inhermes-adapter.test.mjsfailed on main (expected 1 dispatch, got 2).Notes for the reviewer
verifystage maps to thefact-checkerrole in both renderers (per the sub-plan). On Hermes that role's prompt prepends the fact-checker report format, while the stage schema is{ confirmed, note }. The task text andoutput_schemaask for the latter; worth a look in a live Hermes run.Verifyphase andverify:*dispatches lengthens the run; see the wall-clock figure above.🤖 Generated with Claude Code