Skip to content

Add Dr. Bench deep-research benchmark - #75

Merged
xdotli merged 1 commit into
benchflow-ai:mainfrom
reacher-z:agent/add-dr-bench
Aug 20, 2026
Merged

Add Dr. Bench deep-research benchmark#75
xdotli merged 1 commit into
benchflow-ai:mainfrom
reacher-z:agent/add-dr-bench

Conversation

@reacher-z

@reacher-z reacher-z commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

What changed

Added Dr. Bench to the agent-specific evaluation section using the list's annotated benchmark format.

Why

Dr. Bench evaluates deep-research agents on 214 expert-curated tasks across 10 domains, scoring long reports for semantic quality, topical focus, and retrieval trustworthiness.

Checks

  • searched the README and all issue/PR states for the title, aliases, arXiv ID, and canonical repository
  • verified the canonical paper and repository metadata
  • ran git diff --check

Disclosure: submitted on behalf of Dr. Bench co-author Yuxuan Zhang.

@reacher-z
reacher-z marked this pull request as ready for review August 16, 2026 23:44
@xdotli
xdotli merged commit c912161 into benchflow-ai:main Aug 20, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants