A personal evaluation harness for choosing your AI agent stack. Run your actual daily tasks against any number of OpenAI-compatible agents, rate the responses yourself, and get data-driven recommendations after as few as 7 days — not synthetic benchmarks, not LLM-as-judge.
Repository · Documentation · Issues
AgentAssistBench is built for engineers and developers who:
- Are evaluating self-hosted AI agent stacks — Ollama, LibreChat, Open WebUI, Hermes, OpenClaw, vLLM, LiteLLM — and need to pick one for real work
- Are tired of synthetic benchmarks that score a model on tasks they'll never do
- Want to know whether agent A or agent B is better specifically for coding, specifically for research, or specifically for devops — not in aggregate
- Self-host their AI and care about actual latency and infrastructure cost, not just quality
- Want their own judgment to decide what's "better" — not an automated scorer with its own biases
If you're running two agent stacks and wondering which one to commit to, this tool gives you a data-driven answer — in as few as 7 days with aab report --days 7, or over 30–90 days for more statistical confidence.
You're choosing an AI agent framework for daily work — but every published benchmark scores the model, not the stack.
The key insight: Two agent frameworks running the exact same underlying LLM can produce wildly different outputs, at wildly different latencies, for wildly different costs.
The framework's system prompt, tool routing logic, memory architecture, context management, and retry behavior all contribute. When you choose LibreChat over Open WebUI, you're not choosing a model — you're choosing all of that. No existing benchmark measures it.
- Synthetic benchmarks (GAIA, SWE-bench, WebArena) test models on curated tasks — not your actual work
- LLM-as-judge introduces its own model's biases and doesn't reflect what you value
- "Just try both" gives you a subjective impression, not reproducible data
- You need weeks of your own tasks to separate signal from noise
AgentAssistBench is a self-hosted CLI that runs your actual daily tasks against any number of OpenAI-compatible agents side-by-side, then lets you rate the responses yourself. After any period you choose (--days 7, --days 30, --days 90) you get per-category win rates, latency comparisons, and cost-per-request — a data-driven answer to "which agent stack should I use for my work?"
See docs/sample-report.md for a realistic example of what aab report produces.
| Tool | What it does | Why it doesn't answer your question |
|---|---|---|
| GAIA / SWE-bench / WebArena | Leaderboard benchmarks on curated task sets | Scores the model, not the framework; uses tasks you'll never do |
| LMSYS Chatbot Arena | Crowd-sourced human preference between models | Model-vs-model, not framework-vs-framework; no cost/latency; not your tasks |
| promptfoo | Compare prompts and providers with automated evals | Not longitudinal; not framework-stack level; weak on infra cost |
| Langfuse / LangSmith | Observability and evals for an app you're building | For apps you ship, not agent stacks you're choosing between |
| "Just try both for a week" | Informal personal testing | Recency bias, no per-category breakdown, no cost data, not reproducible |
AgentAssistBench is the gap: framework-level, your real tasks, human judgment, configurable time window (--days N), with latency and cost in the same report.
graph TD
Tasks["🗒️ Your Real Tasks (daily)"]
subgraph Agents["Agents — any OpenAI-compatible endpoint"]
A1["Agent 1\nHermes · LibreChat · Ollama · vLLM ..."]
AN["Agent N\nOpenClaw · LiteLLM · Open WebUI ..."]
end
DB[("SQLite\nbenchmarks + executions")]
Ratings["⭐ Your Ratings\nHuman evaluation — not LLM-as-judge"]
Report["📊 aab report\nWin rates · Latency · Cost\nPer-category recommendations"]
Tasks -->|"aab run --tag coding"| A1
Tasks -->|"aab run --tag coding"| AN
A1 -->|"latency · tokens · cost"| DB
AN -->|"latency · tokens · cost"| DB
DB -->|"aab evaluate"| Ratings
Ratings --> DB
DB -->|"aab report"| Report
git clone https://github.com/theprodsde/agent-assist-bench.git
cd agent-assist-benchWith uv (recommended):
uv syncThen prefix commands with uv run — e.g. uv run aab status.
With pip (standard):
pip install -e .Then run aab directly — e.g. aab status.
cp .env.example .envOption A — generic agents (any OpenAI-compatible endpoint)
Use BENCH_AGENTS_JSON to register any number of agents without touching code. Each item needs name, endpoint, and model:
BENCH_AGENTS_JSON=[{"name":"ollama-llama3","endpoint":"http://localhost:11434/v1","model":"llama3"},{"name":"librechat","endpoint":"http://localhost:9000/v1","model":"gpt-4o","api_key":"sk-..."}]Option B — built-in Hermes / OpenClaw (Docker setup)
BENCH_HERMES_ENDPOINT=http://localhost:8001/v1
BENCH_OPENCLAW_ENDPOINT=http://localhost:8002/v1
BENCH_HERMES_API_KEY=your-hermes-api-key
BENCH_OPENCLAW_API_KEY=your-openclaw-api-keyBoth options can be used together — BENCH_AGENTS_JSON agents are registered in addition to Hermes/OpenClaw.
See .env.example for all available settings with descriptions.
docker compose -f docker/docker-compose.yml -f docker/docker-compose.override.yml up -dBoth options hit the same agents and store results in the same database.
# Submit a prompt — runs against all configured agents
aab run "Write a binary search in Python" # pip install
uv run aab run "Write a binary search in Python" # uv sync
# With category tags for reporting
aab run --tag coding "Write a binary search in Python"
aab run --tag research --tag devops "Explain Kubernetes vs Docker Swarm"
# Check what ran
aab status
# Rate the responses (interactive)
aab evaluate
# Generate a comparison report
aab reportSee CLI.md for all commands.
# Add to .env
BENCH_TELEGRAM_BOT_TOKEN=your-token-from-botfather
BENCH_TELEGRAM_ALLOWED_CHAT_IDS=your-numeric-chat-iduv run aab telegramSend any message to your bot — it benchmarks the prompt across both agents and replies with both responses. Prefix #tag to categorise:
#coding Write a binary search in Python
See Telegram Bot section for full setup.
✅ Any OpenAI-compatible agent — Point at Ollama, LibreChat, Open WebUI, vLLM, LiteLLM, Hermes, OpenClaw, or anything behind /v1/chat/completions using BENCH_AGENTS_JSON
✅ Your judgment, not LLM-as-judge — You rate responses 1–5 on your own criteria; no automated scoring introduces its own bias
✅ Per-category recommendations — Tag prompts (--tag coding, --tag research) to get category-specific win rates
✅ Complete stack comparison — Measures latency, tokens, cost, and quality in a single report; see sample output
✅ Multiple deployments — Test agents on local Docker, VMs, or managed containers with optional cost/resource tracking
✅ Telegram bot — Send prompts from your phone, get all agent responses back side-by-side
✅ Self-hosted, Apache 2.0 license — Your data stays local in SQLite; no vendor lock-in; patent rights explicitly granted
sequenceDiagram
participant User
participant CLI
participant Runner
participant A1 as Agent 1
participant AN as Agent N
participant DB as SQLite
User->>CLI: aab run --tag coding "prompt"
CLI->>DB: INSERT benchmark (PENDING)
CLI->>Runner: run(benchmark_id)
Note over Runner: Sequential execution across all configured agents
Runner->>A1: health_check()
A1-->>Runner: ok (measure cold start)
Runner->>A1: POST /v1/chat/completions
A1-->>Runner: response (measure latency)
Runner->>DB: INSERT execution (agent_1, latency, tokens, cost)
Runner->>Runner: sleep(configurable delay)
Runner->>AN: health_check()
AN-->>Runner: ok
Runner->>AN: POST /v1/chat/completions
AN-->>Runner: response
Runner->>DB: INSERT execution (agent_n, latency, tokens, cost)
Runner->>DB: UPDATE benchmark (COMPLETED)
Runner-->>CLI: benchmark_id
CLI-->>User: ✓ Benchmark complete — run aab evaluate to rate
graph LR
Run["aab run<br/>(daily tasks)"]
Status["aab status<br/>(check progress)"]
Evaluate["aab evaluate<br/>(rate responses)"]
Report["aab report<br/>(view results)"]
Run --> DB["SQLite<br/>(benchmarks,<br/>executions)"]
Status --> DB
DB --> Evaluate
Evaluate --> DB
DB --> Report
Report --> Recommendations["Markdown Report<br/>(win rates,<br/>recommendations)"]
style Run fill:#e1f5fe
style Report fill:#c8e6c9
style Recommendations fill:#fff9c4
Send prompts from your phone and get both agent responses back in the same chat.
1. Create a bot via @BotFather
/newbot
Follow the prompts. Copy the token it gives you.
2. Find your Telegram chat ID
Message @userinfobot — it replies with your numeric ID (e.g. 123456789).
3. Add to .env
BENCH_TELEGRAM_BOT_TOKEN=123456789:ABCdefGHIjklMNO...
BENCH_TELEGRAM_ALLOWED_CHAT_IDS=123456789BENCH_TELEGRAM_ALLOWED_CHAT_IDS restricts the bot to your chat only. Anyone else who messages the bot is silently ignored. Leave it empty only if you intentionally want a public bot.
4. Start the bot
aab telegram # pip install
uv run aab telegram # uv syncTelegram bot polling started — send messages to your bot.
Access restricted to chat IDs: 123456789
Press Ctrl-C to stop.
Send any message to your bot in Telegram:
Write a binary search in Python
Prefix #tag to categorise the benchmark for reporting:
#coding Write a binary search in Python
#research #devops Explain Kubernetes vs Docker Swarm
The bot runs the prompt through all configured agents sequentially and replies with all responses side-by-side.
| Document | Purpose |
|---|---|
| CLI.md | Command reference (aab run, aab evaluate, aab report, aab telegram) |
| DEPLOYMENT.md | How to deploy agents locally, on VMs, or in cloud (ACA, Cloud Run) |
| ARCHITECTURE.md | System design and component interactions |
| docs/sample-report.md | Realistic example aab report output (94 benchmarks) |
| docs/ | 16 technical docs: project overview, ADRs, RFCs, roadmap |
Version: 0.1.0 (MVP)
| Component | Status |
|---|---|
CLI (aab run, evaluate, report, status) |
✅ |
| Benchmark runner (sequential execution) | ✅ |
| Agent adapters (Hermes, OpenClaw, OpenAI-compatible) | ✅ |
Generic agent config (BENCH_AGENTS_JSON — any N endpoints) |
✅ |
| SQLite storage (CRUD + WAL mode) | ✅ |
| Report generator (per-category win rates) | ✅ |
| Evaluation service (rating + validation) | ✅ |
| Telegram bot (long-polling, access control) | ✅ |
| Tests (unit, integration, property-based) | ✅ |
| CI (lint, type check, parallel tests) | ✅ |
| Docker setup (local agents) | ✅ |
| REST API | 🔜 v0.2 |
| Terraform (VM + ACA deployment) | 🔜 v0.3 |
| Dashboard UI | 🔜 v1.0 |
See docs/16-roadmap.md for the full roadmap.
- Python 3.11+
- uv
- Docker (optional, for local agent testing)
uv sync --group devuv run pytest tests/ -v -n autouv run ruff check benchmark/ tests/
uv run ruff format benchmark/ tests/uv run mypy benchmark/uv run python3 -c "from benchmark.config.settings import Settings; s = Settings(); print(s.hermes_endpoint)"
# or with pip install:
python3 -c "from benchmark.config.settings import Settings; s = Settings(); print(s.hermes_endpoint)"Apache 2.0 — See LICENSE
Open-source project published under github.com/theprodsde/agent-assist-bench
- Try it locally — Quick Start above
- Set up the Telegram bot — Telegram Bot above
- Run benchmarks daily — Tag your real tasks with categories; use
aab report --days 7for a quick read or--days 30for stronger signal - Rate responses —
aab evaluateafter accumulating results - Get recommendations —
aab reportshows per-category win rates - Deploy at scale — DEPLOYMENT.md for VMs or cloud
| 💬 Questions / ideas | GitHub Discussions — ask anything, share your benchmark results, propose new frameworks |
| 🐛 Bug reports | Open an issue — use the bug report template |
| 🤖 Add an agent | New agent issue — or just use BENCH_AGENTS_JSON for zero-code setup |
| ✨ Feature requests | Open an issue — use the feature request template |
| 🤝 Contributing | See CONTRIBUTING.md — includes the full agent adapter guide |