Problem
Bunny's dark factory shells out to claude CLI for builder steps (specify, plan, implement) and gemini CLI for adversary steps (challenge, test-gen, verify). On a claude --remote cloud VM, the agent IS Claude — so shelling out to claude CLI is recursive and redundant.
How should bunny's architecture adapt to run natively inside a remote Claude Code session?
Key insight
On a remote VM, Claude is the host agent. It reads CLAUDE.md, understands the workflow, and can do the "claude steps" directly (writing specs, plans, implementations). Only the gemini steps need to shell out. This means the pipeline splits naturally:
| Step |
Local mode |
Remote mode |
| specify |
shell out to claude |
agent writes spec directly |
| challenge |
shell out to gemini |
shell out to gemini |
| plan |
shell out to claude |
agent writes plan directly |
| tasks |
shell out to claude |
agent writes tasks directly |
| test-gen |
shell out to gemini |
shell out to gemini |
| implement |
shell out to claude |
agent writes code directly |
| verify |
shell out to gemini |
shell out to gemini |
| retro |
shell out to claude |
agent writes retro directly |
Research questions
Architecture
- Should bunny detect it's running inside a remote session and switch behavior? Or should there be an explicit
--remote-mode flag?
- Can the structured output schemas (spec, plan, tasks, retro) be reused as direct prompts to the host agent instead of CLI calls?
- Should
bny remote build "desc" be a new command that composes the right prompt and calls claude --remote?
Auth & credentials
- Gemini API keys: how do they reach the cloud sandbox? Environment variables? Secrets manager? Network proxy?
- GitHub auth: the cloud VM uses a proxy — does this work with bunny's branch creation and PR workflow?
- Are there other credentials the pipeline needs (npm, bun, etc.)?
Parallel execution
- Can
bny next --parallel N launch N independent claude --remote sessions for different roadmap items?
- How do parallel builds coordinate? Separate branches? Lock files? Race conditions on
bny/state.md?
- What's the token budget impact of running multiple full pipelines simultaneously?
Adversary isolation
- Could builder (claude) and adversary (gemini) run on truly separate VMs?
- Communication only through git artifacts (specs, test files) — no shared process memory
- Does this improve adversarial quality or just add complexity?
Environment setup
- Cloud VMs have bun pre-installed — can we
bun run bin/bny.ts directly?
- Or should the setup script download the compiled binary from GitHub releases?
- What goes in the remote session's setup script? (
./dev/setup, gemini CLI install, env vars)
Prompt engineering
- What's the optimal prompt for
claude --remote to run a bny pipeline?
- Should it be a single big prompt or should bunny compose it from templates?
- How much of CLAUDE.md + roadmap + guardrails should be included vs. discovered by the agent?
Possible approaches
A. Thin wrapper — bny remote
New command that composes a prompt and calls claude --remote. The remote agent reads CLAUDE.md and runs steps manually. Bunny itself doesn't change — the remote agent just follows instructions.
B. Dual-mode pipeline
Bunny detects the execution environment. In remote mode, claude-steps become inline prompt execution (no shell-out). Gemini steps still shell out. Same binary, different code path.
C. Agent protocol
Bunny exposes its pipeline as a structured protocol (JSON task descriptions). Any agent (local claude, remote claude, gemini, human) can pick up tasks and execute them. Remote execution becomes just another consumer of the protocol.
Related
Problem
Bunny's dark factory shells out to
claudeCLI for builder steps (specify, plan, implement) andgeminiCLI for adversary steps (challenge, test-gen, verify). On aclaude --remotecloud VM, the agent IS Claude — so shelling out toclaudeCLI is recursive and redundant.How should bunny's architecture adapt to run natively inside a remote Claude Code session?
Key insight
On a remote VM, Claude is the host agent. It reads CLAUDE.md, understands the workflow, and can do the "claude steps" directly (writing specs, plans, implementations). Only the gemini steps need to shell out. This means the pipeline splits naturally:
claudegeminigeminiclaudeclaudegeminigeminiclaudegeminigeminiclaudeResearch questions
Architecture
--remote-modeflag?bny remote build "desc"be a new command that composes the right prompt and callsclaude --remote?Auth & credentials
Parallel execution
bny next --parallel Nlaunch N independentclaude --remotesessions for different roadmap items?bny/state.md?Adversary isolation
Environment setup
bun run bin/bny.tsdirectly?./dev/setup, gemini CLI install, env vars)Prompt engineering
claude --remoteto run a bny pipeline?Possible approaches
A. Thin wrapper —
bny remoteNew command that composes a prompt and calls
claude --remote. The remote agent reads CLAUDE.md and runs steps manually. Bunny itself doesn't change — the remote agent just follows instructions.B. Dual-mode pipeline
Bunny detects the execution environment. In remote mode, claude-steps become inline prompt execution (no shell-out). Gemini steps still shell out. Same binary, different code path.
C. Agent protocol
Bunny exposes its pipeline as a structured protocol (JSON task descriptions). Any agent (local claude, remote claude, gemini, human) can pick up tasks and execute them. Remote execution becomes just another consumer of the protocol.
Related