Skip to content

docs: add headless claude -p checklist to tm-ab-test - #420

Merged
sv-tmueller merged 3 commits into
mainfrom
feat/418-ab-test-headless-checklist
Oct 1, 2026
Merged

sv-tmueller merged 3 commits into
mainfrom
feat/418-ab-test-headless-checklist

Conversation

@sv-tmueller

Copy link
Copy Markdown
Owner

Closes #418. Part of batch #417.

Adds a ## Headless claude -p checklist section to .claude/skills/tm-ab-test/SKILL.md, with two one-line pointers from steps 3 and 4. Ten items, each with trap, fix (with its gate) and source report section: the 7 required by the issue plus reference-through-refs, spend-limit 429, and config-dir effort override.

The 600 s item is documented: CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS (checked 2026-10-01 at the env-vars and headless docs pages). Docs links are in the section.

No dependency added. Bumps plugin version 2.7.0 to 2.7.1 (touches .claude/skills/); other batch PRs also bump it, resolved at merge.

Verification: npm test 0, node scripts/check-version-bump.mjs origin/main HEAD 0, diff touches only the two files, no em dashes or banned phrases.

🤖 Generated with Claude Code

The 2026-09-28 trials (#400, #404, #405) lost runs, money or fix rounds to the same operational traps, and the skill named none of them. Each item gives the trap, the fix with its gate, and the source report section. Bumps the plugin to 2.7.1 because the change touches .claude/skills/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Tester report (initial test, batch #417)

VERDICT: PASS
COMMIT: a3e01d4
FINDINGS: none
UNTESTED CLAIMS: none for the acceptance criteria. Checks run:

  • npm test: 383 pass, 0 fail, exit 0. node scripts/check-version-bump.mjs origin/main HEAD: "2.7.0 -> 2.7.1", exit 0.
  • The diff touches only .claude/skills/tm-ab-test/SKILL.md and plugin.json; nothing under docs/reviews/.
  • The new text has 0 U+2014 and 0 U+2013, none of the banned phrases from process-core, and no local paths or provider hosts (only $CLAUDE_CONFIG_DIR and <slug> placeholders).
  • Concern 1: every cited report file and section exists, and each says what its item claims. This covers the A/B: native plan-and-execute vs the kickoff pipeline on real issues #400, A/B: lead effort, Opus 5.5 at high vs xhigh, on the advisor refine task #404 and A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 reports and protocols, including "The 11 deviations" items 1, 2, 5, 7 and 8, "Cost-measurement gap", "Limit event", "Cost cross-check", Amendments 1 and 2, "Known gap", and the 651 s review-workflow time (650,996 ms in the A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 report, section 7).
  • Concern 2: curl -sL returned 200 for env-vars, headless and permission-modes. The env-vars page says the variable is the ceiling in ms on idle waiting for background subagents and workflows after the final turn with -p. It also says idle waiting restarts each time Claude takes a turn to handle a background result. Default is 600000, 0 waits indefinitely, and it needs v2.1.182 or later. The #background-tasks-at-exit anchor exists on the headless page.
  • Concern 3: all 7 acceptance-criteria topics and the 3 planned extras are present, each with trap, fix and source. The 600 s item gives docs links and a check date of 2026-10-01.
    Unpinned wording, not a failure (the sources support the gist but not the exact framing):
  • $4.28 reported vs $12.77 measured is put down to missing subagents/workflows/. The A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 report says part of the gap is thinking tokens, which the estimate leaves out: the workflow-aware script gave $8.73.
  • "H2 and H3 ... leaving 0 of 3 runs valid" leaves out that H1 died of a 429. H3's Glob matched nothing; only H2 leaked.
  • "Judge passes are exposed too" is an inference: the 651 s tm-review-changes run was in-session, not a claude -p process.
    LESSONS: Check each cited number against the source's own definition, such as an estimate versus a measurement, before copying it into a checklist item.

@sv-tmueller

Copy link
Copy Markdown
Owner Author

Reviewer report, fix round 1/3 (batch #417)

VERDICT: CHANGES_REQUESTED
STAGE: spec
FINDINGS:

  1. .claude/skills/tm-ab-test/SKILL.md:130-131, must-fix. The sentence "Use placeholders in anything you write down: no local paths, no settings values, no provider hosts." is a new rule for trial records that nobody asked for. In the sub-plan, the placeholder line tells the developer how to write this section. It is not a rule for operators. The new rule also conflicts with how trials are recorded: the A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 protocol pins the exact --settings JSON (line 414) and the A/B: lead effort, Opus 5.5 at high vs xhigh, on the advisor refine task #404 report has a "9. Reproduction" section. Fix: delete that sentence, so the paragraph ends at "ultracode-arm-379 is A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405)."
    Evidence: sub-plan "Approach": "Use placeholders (...), not this machine's paths or settings values. No real provider host." The issue body and ACs have no such rule.
  2. .claude/skills/tm-ab-test/SKILL.md:184-185, should-fix (tester claim 2, upheld). The sentence blames Glob for all of round 1's 0 of 3. In fact H1 died of a 429, and H3's Glob matched nothing. Replace the sentence verbatim with: "In A/B: lead effort, Opus 5.5 at high vs xhigh, on the advisor refine task #404 round 1, H2 and H3 each ran Glob with a path one level above the trial root, which also covered a sibling trial root. H2's matches leaked two files from it; H3 matched nothing but was still invalid under the gate's literal text (H1 had died of a 429, so round 1 had 0 of 3 valid runs)."
    Evidence: docs/reviews/2026-09-28-lead-effort-comparison.md:80-108 and :129.
  3. .claude/skills/tm-ab-test/SKILL.md:241-243, should-fix (tester claim 1, upheld). The sentence blames the whole $4.28-vs-$12.77 gap on the missing workflow agents. The source credits about $6.43 of it to the 7 workflow subagents and the rest to the lead's output, which the estimate leaves out (thinking included). Replace verbatim with: "A script version that does not read subagents/workflows/ also misses workflow agents: in A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 it reported $4.28 against $12.77 measured, and about $6.43 of that gap was the 7 workflow subagents (the rest is lead output the estimate leaves out, thinking included)."
    Evidence: docs/reviews/2026-09-28-ultracode-arm-379.md:146-163 (section 4, "Cost cross-check").
  4. .claude/skills/tm-ab-test/SKILL.md:167-168, nit (tester claim 3, partly upheld). The inference holds, but the cited 651 s run was in-session, not under -p. Better evidence exists: A/B: native plan-and-execute vs the kickoff pipeline on real issues #400's judges ran this same workflow under claude -p. Replace verbatim with: "Judge passes are exposed too: A/B: native plan-and-execute vs the kickoff pipeline on real issues #400's judges ran tm-review-changes under claude -p (protocol section 8), and the same workflow took 651 s when A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 ran it in-session (report section 7)."
    Evidence: docs/reviews/2026-09-28-ab-native-vs-kickoff-protocol.md:97 and :420-427; docs/reviews/2026-09-28-ultracode-arm-379.md:262 and :297.
  5. .claude/skills/tm-ab-test/SKILL.md:103-104, nit. Earlier in the same paragraph, "The checklist" means templates/recording-checklist.md, so "the checklist's cost items below" can be read as pointing at the template. Replace verbatim with: "For a claude -p arm, record cost per "Recording cost" in the headless claude -p checklist below."
    Evidence: SKILL.md:100-102, "The checklist pins the exact commands for base commit, window, token usage...".
    CHECKS: n/a
    LESSONS: A sub-plan's authoring constraints (how to write the change) can leak into the deliverable as rules for its readers; check that instructions meant for the developer did not end up in the output.

@sv-tmueller sv-tmueller mentioned this pull request Oct 1, 2026
4 of 19 tasks
Drop the placeholder rule that conflicts with how trials are recorded,
correct the Glob and cost-gap attributions to match the trial reports,
and tighten two cross-references.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@sv-tmueller

Copy link
Copy Markdown
Owner Author

Tester report, re-test after review fix round 1/3 (batch #417)

VERDICT: PASS
COMMIT: 117bc1f
FINDINGS: none
UNTESTED CLAIMS: none new. The skill is prose, so no test pins its content; I checked it by hand against the issue and the cited reports. The 600 s docs URLs were not re-fetched this round (no WebFetch). They are unchanged from a3e01d4.
LESSONS: A re-test of a verbatim fix round is cheapest as a line-by-line diff against the reviewer's quoted replacement texts, then a spot-check of each cited source.

Checks run on 117bc1f:

  1. git diff a3e01d4 117bc1f touches only .claude/skills/tm-ab-test/SKILL.md (14 insertions, 9 deletions). All five hunks match the reviewer's replacement texts, re-wrapped. Finding 1 is the sentence deletion; findings 2 to 5 are verbatim. Nothing else changed.
  2. Each replacement matches its source.
  3. Original acceptance criteria still hold.
    • All 10 items (7 required, 3 extras) have Trap, Fix and Source.
    • The 600 s item has both docs links and "checked 2026-10-01".
    • Zero em dashes and no cliche phrases (grep).
    • No machine paths or real provider hosts (grep, only a prose "@" hit on plugin key names).
  4. npm test exits 0 (383 pass, 0 fail). node scripts/check-version-bump.mjs origin/main HEAD exits 0 (2.7.0 -> 2.7.1). The branch diff against main touches only SKILL.md and plugin.json.

@sv-tmueller

Copy link
Copy Markdown
Owner Author

Reviewer report, re-review after fix round 1/3 (batch #417)

VERDICT: APPROVE
STAGE: quality
FINDINGS:
Round-1 check: all five round-1 findings are fixed at 117bc1f. Finding 1's sentence is gone, and findings 2-5 now read word for word as the replacement texts. git diff a3e01d4 117bc1f touches only those five hunks. Pass 1 (spec) is clean on the whole diff. It covers the 7 AC items plus planned extras 8-10, each with trap, fix and source, plus both pointers. The 600 s item has docs links and the check date (I re-read docs/en/headless "Background tasks at exit" and env-vars on 2026-10-01; "drops its partial result" and the restart-on-turn wording match). The diff touches only SKILL.md and plugin.json (2.7.0 -> 2.7.1; origin/main is still 2.7.0, PR is MERGEABLE). Frontmatter and templates are unchanged, there are no em or en dashes and no banned phrases. Pass 2 found two nits.

  1. .claude/skills/tm-ab-test/SKILL.md:254-255, nit. The new claim "A/B: native plan-and-execute vs the kickoff pipeline on real issues #400's $35 stop rule used the estimate" (line 251-252) has no section in the Source line, though the AC asks for a source report and section per claim and the sub-plan names protocol section 10. Replace the Source bullet verbatim with: " - Source: A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 report "Cost-measurement gap" and "How adjusted R was computed"; A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 protocol section 10 ($35 stop rule); A/B arm: live ultracode on Opus 5.5 on a replayed merged issue #405 report section 4 "Cost cross-check"."
    Evidence: docs/reviews/2026-09-28-ab-native-vs-kickoff-protocol.md:553-564 ("list-price equivalent, per token-report.mjs").
  2. .claude/skills/tm-ab-test/SKILL.md:177-178, nit. "Set it inside the env -i wrapper" names a wrapper that the skill never introduces, so readers whose launch has no wrapper get an unclear instruction. Replace verbatim with: "If you launch through an env -i wrapper, set it inside the wrapper; anything set outside is dropped (the A/B: native plan-and-execute vs the kickoff pipeline on real issues #400 deviation 2 trap, hit with GH_CONFIG_DIR)."
    Evidence: SKILL.md has no other mention of env -i; the wrapper exists only in docs/reviews/2026-09-28-ab-native-vs-kickoff.md:240 and :272.
    CHECKS: n/a

PR #421 merged first with 2.8.0, so this branch's 2.7.1 bump no longer
moves the version forward. Resolve to the next patch level.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sv-tmueller
sv-tmueller merged commit 12e8c0f into main Oct 1, 2026
2 checks passed
@sv-tmueller
sv-tmueller deleted the feat/418-ab-test-headless-checklist branch October 1, 2026 08:56
sv-tmueller added a commit that referenced this pull request Oct 1, 2026
PRs #421 and #420 merged first and took 2.8.0 and 2.8.1. This branch adds
a workflow stage, so it takes the next minor version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Record the headless trial traps from #400, #404 and #405 in tm-ab-test

1 participant