Skip to content

Verify stage in tm-review-changes, with spec-format item transforms #406

Description

@sv-tmueller

Part of batch #417

Part of #344 (package P1 of the Claude-of-Tanks adoption plan, docs/plans/344-claude-of-tanks-adoption.md on the PR #345 branch).

What to build

Nothing in tm-review-changes re-checks a finding against the tree before it reaches a human. Workers go straight to the Opus critic, which verifies, consolidates and writes the report in one pass from worker text alone. The survey's source repo refuted 10 of 24 major findings once it added an adversarial check. That is one data point, from game visuals, not code review. This package adds the check and measures our own refuted rate.

It has two halves: a spec-format extension, then a stage that uses it.

Spec format. A verify stage needs its items to be "every must-fix finding, flattened across workers, deduped, capped". items_source is a dotted path and cannot express that. Add an items_transform field naming an entry in a small, closed set of reducers that each renderer (Hermes and Codex) implements. The first entry is must_fix_deduped: flatten the worker findings, keep severity === 'must-fix', dedup on file + line + problem, then cap.

  • Recommended option: a named transform. The architect confirms it or argues against it. Rejected: keeping the fan-out in JS only. The spec JSON is the portable contract (docs/architecture/adapter-interface.md), so a Hermes or Codex host would silently run a review with no verify stage. Rejected: an expression language, as over-engineering for one consumer.
  • The cap follows the MAX_AREAS precedent: an args field plus items_default_cap, starting at 12.

The items_cap bug (folded in, from batch #390's follow-ups). Both renderers read args[stage.items_cap] with the literal value "args.areas", so they look up args["args.areas"] and the cap is always the default. P1's cap goes through the same line, so fix it here.

Verify stage.

  • A new verify stage between review and consolidate. It runs on the worker tier, one dispatch per item, with schema { confirmed: boolean, note: string }.
  • The verifier prompt states the adversarial stance outright. It starts from the position that the finding is wrong or stale, and must reproduce it against the current tree or the finding does not survive. "Cannot reproduce", "already fixed" and "the evidence does not hold" all mean confirmed: false.
  • consolidate receives confirmed and refuted findings as separate inputs.
  • The report schema gains a refuted array. The report carries a refuted appendix, so a human can see what the stage removed and whether it removes too much.
  • Findings past the cap are reported, not dropped silently, the same way scoutDropped and ceilingReached work in tm-review-codebase.js.

These seven files must change together. Miss one and a test fails, costing a fix round.

Surface File Enforced by
Runtime SPEC .claude/workflows/tm-review-changes.js -
Portable spec .claude/workflows/specs/tm-review-changes.spec.json specs.test.mjs
Runtime prompts PROMPTS const in the workflow JS -
Portable prompts .claude/workflows/prompts/tm-review-changes.prompts.json prompts-sync.test.mjs
Second host .claude/adapters/hermes-renderer.mjs (collectSlotVals, inferRole) hermes-renderer-prompts.test.mjs
Tier pinning stage tier field effort-policy.test.mjs, render-path.test.mjs
Format contract docs/architecture/adapter-interface.md "Spec format" review

The Codex renderer needs the reducer too.

Live run. Once the tester passes, the lead (not the developer) runs the branch's tm-review-changes workflow once, on this PR's own diff, and records the results in the PR body. The owner pre-approved this run on 2026-09-28.

Acceptance criteria

  • npm test is green, including specs.test.mjs, prompts-sync.test.mjs, render-path.test.mjs, effort-policy.test.mjs and hermes-renderer-prompts.test.mjs.
  • A test shows must_fix_deduped filters, dedups and caps, in both the Hermes and Codex renderers.
  • A test that fails on today's code shows items_cap naming args.<field> resolves to the passed value.
  • A test shows a refuted finding is absent from mustFix and present in refuted.
  • Findings past the cap appear in the report, not silently dropped.
  • adapter-interface.md documents items_transform and the reducer set.
  • One live /tm-review-changes run on this PR's diff, with confirmed count, refuted count and cost recorded in the PR body.
  • The PR body records the live run's wall-clock time beside the confirmed count, refuted count, and cost (added by batch Batch: measurement-tools #417: tm-review-changes already ran past 600 s in A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405, and a verify phase makes it longer, which matters for headless judging).
  • version in .claude/.claude-plugin/plugin.json is bumped.

Non-goals

  • Porting the verify stage to tm-review-codebase (P2, next batch).
  • An expression language in items_source.
  • Codex live structured output (Codex spawn returns text).
  • Deleting the dry-run stub.
  • The auto-must-fix severity conditions (P3, its own issue).
  • An A/B rerun with a second model.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    size:M1 to 3 hours. Write a sub-plan first.type:feat

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions