Part of batch #417
What to build
The three trials of 2026-09-28 lost runs, money, or fix rounds to the same operational traps, and tm-ab-test mentions none of them. Add a headless-run checklist to the tm-ab-test skill. Each item gives the trap, its fix, and the report section it came from, so the next trial avoids it up front. The next batch's main work is two such trials (a second native-vs-kickoff pair, and a paired /code-review run), so this lands first.
Sources to extract from:
Acceptance criteria
Non-goals
- A new script or code.
- Re-running any trial.
- Editing the three reports or any frozen protocol file.
Part of batch #417
What to build
The three trials of 2026-09-28 lost runs, money, or fix rounds to the same operational traps, and
tm-ab-testmentions none of them. Add a headless-run checklist to thetm-ab-testskill. Each item gives the trap, its fix, and the report section it came from, so the next trial avoids it up front. The next batch's main work is two such trials (a second native-vs-kickoff pair, and a paired/code-reviewrun), so this lands first.Sources to extract from:
docs/reviews/2026-09-28-ab-native-vs-kickoff.md("The 11 deviations", "Cost-measurement gap") and its protocol file.docs/reviews/2026-09-28-lead-effort-comparison.md(sections 3, 4 and 8) and its protocol file.docs/reviews/2026-09-28-ultracode-arm-379.md(sections 3 and 5) and its protocol file ("Amendment 1", "Amendment 2").Acceptance criteria
tm-ab-test's SKILL.md has a headless-run checklist. Each item states the trap, the fix, and its source report and section.orchestrai@synced) and verified in the init eventdontAsksilently denying writes under.claude/total_cost_usdandmodelUsagebeing cumulative across a--resumechain (take the last result event per session, don't sum)modelUsagevstoken-report.mjs's 3.8-4.4x-low estimateGloballowlist gapnpm testis green, andversionin.claude/.claude-plugin/plugin.jsonis bumped (the change touches.claude/skills/).Non-goals