Skip to content

Add AgentRunProof to eval frameworks and harnesses - #81

Open
FU-max-boop wants to merge 1 commit into
benchflow-ai:mainfrom
FU-max-boop:add-agentrunproof
Open

Add AgentRunProof to eval frameworks and harnesses#81
FU-max-boop wants to merge 1 commit into
benchflow-ai:mainfrom
FU-max-boop:add-agentrunproof

Conversation

@FU-max-boop

@FU-max-boop FU-max-boop commented Aug 16, 2026

Copy link
Copy Markdown

What changed

Adds one entry to §5a · Eval frameworks & harnesses.

Why it belongs

AgentRunProof is a provider-free runtime-regression harness for the OpenAI Agents SDK. It drives the real Runner with deterministic scripts, compares run and run_streamed behavior, checks RunState resume invariants, and emits content-addressed records.

This is deliberately not presented as a model-quality evaluator. Its scope is SDK runtime correctness.

The public case study documents the evidence chain around two merged, member-authored upstream fixes:

The entry explicitly states that this is not OpenAI adoption or endorsement.

Evidence of outside use and review

Disclosure and maturity

I maintain AgentRunProof. It is still new, has no claimed default-branch package adoption, and remains unproven beyond the public evidence and external trial above. The inclusion case is reproducible runtime evidence, upstream diagnostic impact, and the disclosed outside review—not stars or self-reported benchmark results.

Verification

  • Checked README, MENTIONS, notes, and all-state issues/PRs for the project name and canonical URL; no duplicate found.
  • Verified the repository, case study, upstream PRs, LoopGain report, and third-party write-up links.
  • Rebased onto current main; git diff --check passes.
  • Changed only README.md with one annotated entry.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant