Capture LLM trajectories across OAuth and API-key runs - #1057
Capture LLM trajectories across OAuth and API-key runs#1057bingran-you wants to merge 74 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Devin Review found 3 potential issues.
2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 22861954d1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4c3ae11b1f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f905374331
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 744d43c58e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ffe3dcab87
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: effc13256a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: aa078e09c6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 200b1075a6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c3ac7f2dc4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b21b1aaeed
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5646fe8383
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bd768e05d1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep it up! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Outcome
BenchFlow now creates
trajectory/llm_trajectory.jsonlandtrajectory/llm_trajectory.manifest.jsonat rollout initialization for Claude Code and Codex, across API-key and native OAuth/subscription authentication. Setup failures, no-call runs, capture failures, mixed-role runs, and continuations all retain a truthful terminal artifact.Capture contract
benchflow-litellmprovider definition from allowlisted fields. Caller top-level fields, provider IDs, unselected providers, literal headers, and credential aliases are discarded; only the BenchFlow proxy URL,OPENAI_API_KEYbridge, selected model alias, Responses wire API, and disabled WebSocket flag survive.callback.jsonlandcapture_state.jsonartifacts are forced to 0700 and 0600 independent of inherited umask, and must be unreadable and unmodifiable by the agent; passwordless sudo/doas, root-group membership, effective capabilities, writable container-control sockets, probe anomalies, and integrity mismatches all fail closed to audit-only capture.pre_api_callbody—the payload handed to the provider—not proxy ingress. Missing post-transform bodies failrequest_completeclosed and downgrade the exchange to audit-only fidelity.agent_sessiondata because the files are agent-writable.agent_sessionfallbacks; ACP projection is the final audit-only fallback.replay_proxy/agent_sessionrows withrequest_complete=false; host ownership and sandbox shared-root custody remain separately recorded.payload_redacted=true; false or missing redaction custody fails closed before canonical results or trainer export.Lifecycle and integrity
.txtdata resource generated from that module at wheel-build time. Host and sandbox LiteLLM runtimes receive the exact resource beside the callback; no reflection or manually reconstructed dependency closure is used.Maintainability remediation
The exact-head thermo review findings are addressed:
LLMTrajectoryCaptureis reduced from 965 to 586 lines; typedClaudeOtelCollectorandNativeSessionCollectorown OTel lifecycle and native collection in a focused module;is_completed=trueonly when lifecycle provenance is internally consistent;Exact-head verification
Head:
272557493b0f056502b1a90bd8bbfeb749ba9316Thermo-nuclear remediation
Provider custody is now a dedicated orchestration module plus a packaged, shellcheck-able
provider_capture_custody.shresource; the deadsandbox_localargument and duplicated custody block are gone.Sandbox replay is a real stdlib-only Python module, packaged as an exact wheel resource and uploaded verbatim. Its quiesce/handler concurrency protocol is now linted, type-checked, and directly unit-tested.
payload_redacteddefaults false and is derived from the final serialized JSONL after the canonical structural redactor runs. A skipped or ineffective redactor fails training admission.Atomic writes, training-grade admission, private sandbox uploads, and sandbox-home resolution each have one canonical helper. The redaction compatibility re-export is removed.
Codex direct-provider configuration and BenchFlow proxy configuration are separate APIs; direct caller settings are preserved while proxy routes remain isolated.
Provider request paths now follow the actual LiteLLM call type (
/v1/messages,/v1/responses, or/v1/chat/completions).Healthy audit-only OAuth captures are reported as well-formed but non-training-grade, never malformed.
bench train validate --require-llm-trajectorynow rejects the non-training-grade category explicitly.A prepared OAuth role with no observed model call now terminates as a clean zero-exchange
no_model_callartifact that preserves canonical audit completion. Missing prepared roles still fail closed as soon as ACP reports a call or any captured row proves a call; regression coverage is tied to exact-head review findingr3888576789.Mixed API/OAuth usage is now reported as trusted
usage_source=mixed: provider and native token/cache totals are summed, source-specific detail and priced provider cost are preserved, and the aggregate cost remains unknown rather than pretending the unpriced OAuth leg was priced.Native ACP usage is accumulated only for roles actually using subscription auth, and role transitions reset cumulative checkpoints so provider-proxied ACP counters cannot be double-counted.
Primary setup consumes the first role's scoped environment, so Claude OAuth and Codex Azure credentials never need to coexist in the parent/global environment.
Every native-subscription reconnect refreshes credential files, registry auth uploads, native MCP settings, and web policy even when the role reuses the primary agent/model with an empty role environment. This closes exact-head review
r3888704109for OAuth → API → OAuth sequences after proxy cleanup removes the earlier login file.Proxy startup removes every registry-declared native-login route left by any prior role, not only the current harness's credential. Cleanup is no-follow, settings-aware, effective-home-aware, and fails closed when a runtime or safe path proof is unavailable.
Role teardown freezes and records the exact ACP wrapper process tree before termination, preventing its real CLI child from being orphaned under PID 1 while preserving unrelated task processes. Claude's OTel sink now runs as root-owned BenchFlow infrastructure with only the raw-body directory delegated to the agent, so it cannot invalidate the next API role's zero-process custody proof.
Training admission rejects contradictory
complete/provider_wiremanifests unless botherrorsandmissing_fieldsare present and empty; direct-predicate and canonical-results regressions cover both gap collections.OpenCode-family proxy mode replaces the full provider map, preventing image-baked literal keys or endpoints from bypassing the capture gateway while retaining unrelated settings.
OpenCode proxy launch now isolates every effective config authority from pinned OpenCode 1.18.11: alternate file/directory/inline env sources are unset; project discovery is disabled; HOME/XDG and test-home overrides are pinned; global,
~/.opencode, and redirected system-managed configs are removed; inline/file auth and remote well-known config are neutralized; and active-account state uses an empty in-memory DB. Regressionr3888738113seeds hostile values in every source.Continuation run-level model attribution now comes from finalized active role captures. A requested-but-unused live model cannot overwrite replay-only
config.json,result.json, or canonical/aggregate trainer rows; regressionr3888738115covers the zero-live-call path.Audit-only completion now cross-checks the strictly loaded JSONL rows against sidecar per-role cardinality and role/agent/model/auth/source/fidelity. Truncated or misattributed JSONL fails completion even when the sidecar remains valid (
r3888793594), and zero-call completion requires nonempty prepared-role evidence (r3888793598).Pi proxy mode proves the agent UID quiescent, no-follow removes
.pi/agent/models.jsonfrom every safe effective home, pins the canonical launch home, and fails capture trust closed on unsafe roots or cleanup errors before the unchanged manifest-owned launcher runs; direct-provider behavior remains unchanged.Provider usage and provider-API failure evidence are retained and aggregated across every retired role-scoped LiteLLM runtime plus the final runtime. Token/cache counters sum across rotations; cost sums only when every contributing runtime is priced, so partial pricing cannot masquerade as a complete total.
The real agent launch uses
setpriv --no-new-privs --reuid/--regid/--init-groupswhen supported. The custody probe is run through that same privilege boundary and requires Linux/proc/self/statusto reportNoNewPrivs: 1; the compatibilitysufallback therefore fails capture trust closed to audit-only instead of claiming provider-wire custody.MiMo proxy mode clears the alternate MIMOCODE_CONFIG override before its manifest-owned launcher writes both canonical configs; direct-provider mode remains unchanged.
Host and sandbox replay persist the recorded prefix actually consumed. Stitched artifacts include only that prefix, and an early stop fails closed with
recorded_replay_prefixmissing.Schema-v2 training admission cross-checks every row against the sidecar for fidelity, request/response completeness, redaction, capture source, auth mode, role, agent, model, and exact per-role cardinality; contradictory or empty role captures fail closed.
Every training-grade row must independently assert
role_attribution_complete=true; false or missing attribution evidence fails closed.Before API-proxy capture can be trusted, every agent—including agents without subscription-auth support—must prove a non-root sandbox UID has no live processes. The guard accepts only
pgrepexact no-match status 1; missing tooling, UID 0, a live process, or a probe error fails closed to audit-only, and no process is killed. For subscription-capable agents, registry-defined default/config-home/effective-home tokens are removed in the same root shell only after the guard passes and are verified absent through no-follow directory descriptors. Claude settings at every effective config home are atomically sanitized ofenv,apiKeyHelper,awsAuthRefresh, andawsCredentialExportwhile preserving unrelated policy, ownership, and mode; malformed JSON, symlinks, rewrite failures, and residual keys fail trust closed.HOMEandBENCHFLOW_AGENT_HOMEare pinned to the canonical sandbox home, unsafe roots and symlink traversal fail closed, and sandbox-home values are excluded from the host LiteLLM environment.Continuation live rows redact once at their durable boundary; crash-safety and authoritative post-cleanup writes are documented; replay state cannot fabricate a balanced journal.
LLMTrajectoryCaptureremains below 600 lines; the native capture suites are 967 and 715 lines; continuation trainer metadata was extracted sotest_orchestrator.pyis 897 lines; the branch has no under-1,000-to-over-1,000 crossing.Local verification
pytest tests/: 6,067 passed, 55 skipped, 7 deselected in 200.58s on exact head272557493b0f056502b1a90bd8bbfeb749ba9316.ruff format src tests tools --check: all 631 files formatted.ruff check .,ty check src/, andgit diff --check: pass.Real Docker authentication matrix
The authentication lanes used credential-isolated parent processes. Claude received only its OAuth subscription token; every Azure/OpenAI/Anthropic API-key variable was removed before launch. Codex received Azure credentials while every Claude credential was removed.
27255749): reward 1 with two redactedclaude_otel_raw_body/agent_session/oauth_subscriptionexchanges. Request/response completeness is true, the manifest has no errors or missing fields, canonical results reportis_completed=truebuttraining_ready=false, native ACP usage reports 43,562 tokens, and raw/base64/URL-encoded scans found zero credential occurrences.config.jsonandresult.jsoncontain no Azure marker. Both trainer exporters reject it withskipped_insufficient_capture_fidelity=1.27255749): the raw Azure Responses preflight returned HTTP 200 with a valid usage envelope, then the task scored reward 1. Canonical results aretraining_ready=trueandis_completed=true; the artifact contains two complete redactedlitellm_proxy/provider_wirerows, 28,797 positive-usage tokens, priced cost, no capture errors or missing fields, and zero raw/encoded credential occurrences. No Claude credential name or value was present.5646fe83): reward 1 with seven complete schema-v2 exchanges and no manifest errors or missing fields. Both Claude sessions aggregate into fourclaude_otel_raw_body/agent_session/oauth_subscriptionexchanges; Codex contributes threelitellm_proxy/provider_wire/api_keyexchanges. The return to the original OAuth role succeeds after the API role's all-agent credential cleanup. Mixed usage reconciles exactly: 87,104 native + 43,708 provider = 130,812 total tokens. Raw, base64, and URL-encoded scans found zero Claude/Azure credential occurrences.partial/mixedbecause it includes OAuth audit evidence. Codex rows retainprovider_wirefidelity, but the mixed rollout is intentionally non-exportable as a whole; Prime-SFT and TRL-SFT each write zero rows withskipped_insufficient_capture_fidelity=1.5646fe83): real Docker handoffs leave noagent-owned process from the outgoing role, keep the OTel sink in a root-owned BenchFlow infrastructure boundary, remove every registered native credential route before Codex proxy startup, and still restore OAuth configuration on the final Claude reconnect.Supplemental failure-mode runs were produced on earlier reviewed heads of this branch and remain covered by final-head regressions:
no_model_call/oauth_subscriptionmanifests.Trainer and package proof
--expected-rows 2 --source-jobs ... --require-llm-trajectorywithnon_training_grade_llm_trajectory=0.--require-tool-callsis intentionally not claimed for these converted rows: the source ACP trajectory records one custom shell action, but Codex's provider requests declare no trainer-format tool definitions. The validator correctly rejects that optional gate rather than fabricating tool schemas.benchflow-0.7.6.dev0wheel built from exact head27255749installed successfully in a clean Linux Python 3.12 container; packaged custody resources were present and thebench, evaluation, and trainer CLI surfaces passed smoke checks.Artifacts
/tmp/benchflow-pr1057-272557-live-claude-oauth.QnlVVk/2026-08-30__01-03-55/hello-world-task__862fa3cb/tmp/benchflow-pr1057-272557-live-codex-azure.P1aY13/2026-08-30__01-05-45/hello-world-task__b64cc03c/tmp/benchflow-pr1057-272557-claude-audit-export.KQsclo//tmp/benchflow-pr1057-272557-trainer.i7HLkW//tmp/benchflow-pr1057-272557-package.GdPwd1//tmp/benchflow-pr1057-final22-oauth-api-oauth._z4_k7ha/2026-08-30__00-04-19/hello-world-task__eb91faa1/tmp/benchflow-pr1057-final22-mixed-trainer.J04Xsx/Exact-head CI
test,pip-audit,parity,rollout-smoke,fixture-scenarios, and bothdetect-scopejobs are green on exact head272557493b0f056502b1a90bd8bbfeb749ba9316.Artifacts
/tmp/benchflow-pr1057-remediation-claude-oauth/2026-08-29__16-43-43/hello-world-task__f37883b1/tmp/benchflow-pr1057-remediation-codex-azure/2026-08-29__16-43-43/hello-world-task__fff6b39b/tmp/benchflow-pr1057-final15-codex-azure/2026-08-29__22-51-02/hello-world-task__f9043e20/tmp/benchflow-pr1057-final7-root-codex-azure/2026-08-29__20-35-47/hello-world-task__6a6a8648/tmp/benchflow-pr1057-remediation-claude-api/2026-08-29__16-45-34/hello-world-task__7b910a48/tmp/benchflow-pr1057-remediation-codex-oauth-json/2026-08-29__16-50-11//tmp/benchflow-pr1057-exact-claude-bedrock/2026-08-29__17-10-14//tmp/benchflow-pr1057-final15-trainer.zt29BT//tmp/benchflow-pr1057-final15-package.iJMIK4//tmp/benchflow-pr1057-final3-parity.cFCRow//var/folders/s8/dpnknm2j0kj9ww1r_nvtxdrr0000gq/T/benchflow-pr1057-abrupt-stop-in3z8w4d/