Skip to content

Batch: measurement-tools #417

Description

@sv-tmueller

Approved contract (sign-off: dispatch, 2026-10-01). Raw need: "advise on the next issues to tackle. Consider the latest changes."

Refinement answers (owner, 2026-10-01)

Proposal: Batch: measurement-tools

Why this batch: every recent study had to work around broken measuring tools. This batch fixes the token report, the A/B harness and the review workflow, plus one small doc cleanup. The next batch can then run the two measurements that could change policy.

Four packages, all independent of each other. P1, P3 and P4 are existing issues. On dispatch their bodies are updated to the scope below and get the batch line. P2 is new.

P1 (#416, rescoped): Token report includes workflow subagents and adds a per-agent wall-clock table

Scope: token-report.mjs reads the lead transcript and subagents/agent-*.jsonl, but never subagents/workflows/<wf-id>/. So a workflow's whole cost goes missing without a warning ($4.28 reported vs $12.77 measured in #405). This package makes the report read those transcripts too, and say which transcript directories it found but couldn't read. It also adds a per-agent table so slow seats are visible. Speed is shown as measured wall-clock, not tokens per second. Interactive transcripts carry only streaming placeholders for output_tokens (checked 2026-10-01: values 5/8/16 across 32 messages), so a tokens-per-second figure would rest on an estimate that leaves out thinking.

Acceptance criteria:

  • The report includes workflow subagents under subagents/workflows/, grouped by role and by model like every other agent. A fixture test covers a session with a workflow.
  • The report names every subagent transcript directory it found but could not read.
  • A per-agent table lists role, model, calls, tool uses, estimated output tokens, and wall-clock time for each agent (the lead, each dispatched seat, each workflow subagent).
  • The report states where each wall-clock value comes from (for example, duration_ms in a task notification, or the first-to-last timestamp span in a transcript), and marks any estimated value.
  • Tests are written first, against fixture transcripts. They include an agent with no output and an agent with a zero or missing duration, which shows as "n/a" (no division by zero, no made-up value).
  • /tm-advisor section 5 and /tm-kickoff's wave-end report still post the report on every run, with the new table.
  • npm test is green, and plugin.json version is bumped.

Size: M. Track: full (executable code under .claude/skills/).
Non-goals: a tokens-per-second column; pricing headless claude -p child sessions; real billing figures; live or streaming throughput.

P2 (new): Record the headless trial traps from #400, #404 and #405 in tm-ab-test

Scope: The three trials of 2026-09-28 lost runs, money, or fix rounds to the same operational traps, and tm-ab-test mentions none of them. This package adds a checklist section to the skill. Each item gives the trap, its fix, and the report section it came from, so the next trial avoids it up front. These trials are the next batch's main work, so this lands first.

Acceptance criteria:

  • tm-ab-test's SKILL.md has a headless-run checklist. Each item states the trap, the fix, and its source report and section.
  • It covers at least:
    • every plugin copy disabled (including orchestrai@synced) and verified in the init event
    • dontAsk silently denying writes under .claude/
    • the 600 s background-task ceiling cutting off a workflow
    • total_cost_usd and modelUsage being cumulative across a --resume chain (take the last result event per session, don't sum)
    • modelUsage vs token-report.mjs's 3.8-4.4x-low estimate
    • a judge checkout that reproduces the runs' filesystem layout
    • the Glob allowlist gap
  • The 600 s item names the documented way to keep a headless session waiting for a background workflow, with a docs link. If none is documented, it says so with the check date.
  • npm test is green, and plugin.json version is bumped.

Size: S. Track: full (touches .claude/skills/).
Non-goals: a new script or code; re-running any trial; editing the three reports or any frozen protocol file.

P3 (#406, as filed plus one line): Verify stage in tm-review-changes, with spec-format item transforms

Scope: Unchanged from the issue. It adds an items_transform field with the must_fix_deduped reducer in both renderers, and fixes the items_cap lookup bug. It also adds an adversarial worker-tier verify stage between review and consolidate, a refuted array and appendix in the report, and over-cap findings reported rather than dropped. All seven surfaces change together. The one live /tm-review-changes run on the PR's own diff is included in this dispatch, as the owner pre-approved on 2026-09-28. The single addition: that run also records its wall-clock time. tm-review-changes already ran past 600 s in #405, and a verify phase makes it longer, which matters for headless judging.

Acceptance criteria: all eight from the issue, plus:

  • The PR body records the live run's wall-clock time beside the confirmed count, refuted count, and cost.

Size: M. Track: full.
Non-goals: as filed (no P2 port to tm-review-codebase, no expression language, no severity floor, no A/B rerun).

P4 (#415, two criteria settled): Tidy up after batch #403

Scope: Unchanged edits to the team-guide note and the lead-effort report nits. Two criteria change to match what is now known:

  • The owner step is done. settings.json shows modelSettings.claude-opus-5-5.effortLevel: xhigh as of 2026-10-01, which has the same effect as removing it.
  • The owner confirmed that typing /effort writes a per-model override into settings.json. So the team-guide note must say that a refine-phase /effort high persists across sessions until /effort xhigh is typed again.

Acceptance criteria:

  • The first eight criteria from the issue (team-guide "dispatch" or "file only", plus the seven report and data.json nits).
  • Owner step: already done. The developer confirms that a session log written after 2026-10-01 09:29 shows perTurnEffort xhigh, or says none exists yet.
  • The team-guide note says that /effort high writes a persistent per-model override (owner-confirmed 2026-10-01), and that typing /effort xhigh before "dispatch" or "file only" is what undoes it.

Size: S. Track: lean (only team-guide.md and docs/reviews/, no guarded paths, no code; the reviewer runs the check suite).
Non-goals: as filed (no frozen protocol edit, no tiers.lead, agent or test change, no new trials, no score changes).

How the run goes

  • Up to 3 packages run at once. P3 (the longest) starts first, alongside P1 and P2. P4 starts when a slot frees up.
  • P1, P2 and P3 all bump plugin.json. Whichever merges first, the other two each need a one-line rebase on the version.
  • Planned decision, logged here: the installed token-report.mjs can't see workflows, and P3's live run is a workflow. If P1's PR is ready by report time, the batch token report runs from P1's branch so that run is counted, and the comment says so.
  • Cost: roughly $30-50 at list price for the pipeline (for reference, batch Batch: lead-effort-and-ultracode #403's two packages cost $37.71), plus a few dollars for P3's live run.

Deferred to the next batch, and why

Left out, with a recommendation

Packages

Decision log

Decisions made during the run are posted as comments below.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions