Type: how-to.
Invisible Bench evaluates caregiver-support conversations and produces a Jury
Card for each completed scan. Use the bench CLI from the invisiblebench
Python package.
One judge path evaluates each active check: yes/no questions answered as probabilities, and a code-owned rule that derives the verdict from the saved answers. Safety and Care stay separate. There is no composite score or model rank.
Read the paper on arXiv and the published documentation. The paper describes the original benchmark. The documentation describes the current method. Use CITATION.cff to cite the paper.
A Jury Card complements a model card with evidence from a specific run. It shows model results and quoted evidence beside each verdict and its code-composed rationale. It also records judge settings, costs, technical errors, and attributed commentary. Frozen validation expectations add case-level agreement, errors, and untested coverage. Human review remains optional.
Each run has two parts in one private directory:
| Part | Contents |
|---|---|
jury-card.md |
The standard report. Replaces separate per-run reports and scorecard exports. |
| Saved evidence | scan_plan.json, answers.jsonl, judgments.jsonl, source manifests, and transcripts. Supports inspection and replay. |
Verdicts are PASS, FAIL, UNCLEAR, or NOT_APPLICABLE. Every FAIL needs
transcript evidence. UNCLEAR stays visible. Commentary can dispute a judgment
without changing the saved verdict.
The card reports model judgments. It does not establish judge accuracy or clinical outcomes. Care remains directional. Read the method for definitions, rates, and limits.
Install dependencies. Set the provider key for transcripts and the TypeSafe key for judging:
uv sync --extra dev
export OPENROUTER_API_KEY=...
export TYPESAFE_API_KEY=...Generate transcripts. Pick a catalog model from
src/invisiblebench/models/config.py; unknown IDs are rejected because the
CLI does not invent prices. Review the dry-run estimate before setting a cost
ceiling:
uv run bench -m <catalog-model-id> --dry-run
uv run bench -m <catalog-model-id> -y --max-cost-usd <budget>Find the run ID with uv run bench runs. Replace <run-id> below with that
directory name. Create a scan plan, then review its estimate before running
it. --llm-model defaults to the pinned judge model in
src/invisiblebench/api/typesafe.py:
uv run bench scan plan results/<run-id> \
--llm-model <judge>
uv run bench scan run \
--plan results/<run-id>/scan_plan.json --max-cost-usd <budget>Completion writes results/<run-id>/jury-card.md. Inspect the run and its evidence:
uv run bench get <run-id>
uv run bench explain <catalog-model-id> <scenario-id> \
--failures --scan results/<run-id>Each run lives in results/<run-id>/, named by its UTC start time in
YYYY-MM-DD_HH-MM-SSZ form. Each new run gets its own directory, including runs
of the same model. The card title shows the model and a readable UTC date.
Historical runs live under results/archive/.
Repeat the scan command to resume unfinished judgments. A request with an
unknown outcome blocks automatic resume. The
full quickstart explains recovery for both generation
and judging. Use uv run bench jury <run-id> to regenerate the card from saved evidence without model calls.
See the full quickstart to judge saved responses again or
replay a scan.
Prepare constructed controls without model calls:
uv run python scripts/plan_validation.py --output results/<new-run-id>Use the ordinary budgeted scan command above to judge the saved plan. The Jury Card reports expected and observed outcomes. The validation guide defines each tested boundary and the limits of the resulting evidence.
| Path | Purpose |
|---|---|
benchmark/ |
Scenario corpus, inventory, and tests |
checks/ |
Questions, rule, evidence requirements, and exemplars with committed judge answers |
src/invisiblebench/ |
Runtime, CLI, and Jury Card generation |
scripts/ |
Scan, validation, and documentation commands |
docs/ |
Method and run guides |
data/ |
Aggregate results and historical releases |
Use the inventory for current corpus facts and the scan contract for artifact fields.
uv run ruff check .
uv run python scripts/check_examples.py verify
uv run pytest benchmark/tests -q
uv run python scripts/lint_turn_indices.py --strictEnable the required local hook with git config core.hooksPath .githooks.
This repository owns the public scenarios, check definitions, judge runtime, scan artifacts, and separate Safety/Care projection. It does not own GiveCare product policy, model training, clinical guidance, or real-world outcome claims.
The judge questions are public in checks/. Private transcripts, private
scenarios, expected answers, and credentials stay in local storage. Public releases contain aggregate
results with recorded provenance.
Committed historical releases keep their original bytes and labels. New
releases use the current contract and a new release version.
Historical web evidence remains in data/releases/web-bench-release.tar.gz.
See contribution guidelines for the scenario contract, proof commands, and pull request checks. See the security policy for private security reports.