Skip to content

Repository files navigation

Invisible Bench

Type: how-to.

License: MIT

Invisible Bench evaluates caregiver-support conversations and produces a Jury Card for each completed scan. Use the bench CLI from the invisiblebench Python package.

One judge path evaluates each active check: yes/no questions answered as probabilities, and a code-owned rule that derives the verdict from the saved answers. Safety and Care stay separate. There is no composite score or model rank.

Read the paper on arXiv and the published documentation. The paper describes the original benchmark. The documentation describes the current method. Use CITATION.cff to cite the paper.

What a run produces

A Jury Card complements a model card with evidence from a specific run. It shows model results and quoted evidence beside each verdict and its code-composed rationale. It also records judge settings, costs, technical errors, and attributed commentary. Frozen validation expectations add case-level agreement, errors, and untested coverage. Human review remains optional.

Each run has two parts in one private directory:

Part Contents
jury-card.md The standard report. Replaces separate per-run reports and scorecard exports.
Saved evidence scan_plan.json, answers.jsonl, judgments.jsonl, source manifests, and transcripts. Supports inspection and replay.

Verdicts are PASS, FAIL, UNCLEAR, or NOT_APPLICABLE. Every FAIL needs transcript evidence. UNCLEAR stays visible. Commentary can dispute a judgment without changing the saved verdict.

The card reports model judgments. It does not establish judge accuracy or clinical outcomes. Care remains directional. Read the method for definitions, rates, and limits.

Quickstart

Install dependencies. Set the provider key for transcripts and the TypeSafe key for judging:

uv sync --extra dev
export OPENROUTER_API_KEY=...
export TYPESAFE_API_KEY=...

Generate transcripts. Pick a catalog model from src/invisiblebench/models/config.py; unknown IDs are rejected because the CLI does not invent prices. Review the dry-run estimate before setting a cost ceiling:

uv run bench -m <catalog-model-id> --dry-run
uv run bench -m <catalog-model-id> -y --max-cost-usd <budget>

Find the run ID with uv run bench runs. Replace <run-id> below with that directory name. Create a scan plan, then review its estimate before running it. --llm-model defaults to the pinned judge model in src/invisiblebench/api/typesafe.py:

uv run bench scan plan results/<run-id> \
  --llm-model <judge>
uv run bench scan run \
  --plan results/<run-id>/scan_plan.json --max-cost-usd <budget>

Completion writes results/<run-id>/jury-card.md. Inspect the run and its evidence:

uv run bench get <run-id>
uv run bench explain <catalog-model-id> <scenario-id> \
  --failures --scan results/<run-id>

Each run lives in results/<run-id>/, named by its UTC start time in YYYY-MM-DD_HH-MM-SSZ form. Each new run gets its own directory, including runs of the same model. The card title shows the model and a readable UTC date. Historical runs live under results/archive/.

Repeat the scan command to resume unfinished judgments. A request with an unknown outcome blocks automatic resume. The full quickstart explains recovery for both generation and judging. Use uv run bench jury <run-id> to regenerate the card from saved evidence without model calls. See the full quickstart to judge saved responses again or replay a scan.

Validate the evaluator

Prepare constructed controls without model calls:

uv run python scripts/plan_validation.py --output results/<new-run-id>

Use the ordinary budgeted scan command above to judge the saved plan. The Jury Card reports expected and observed outcomes. The validation guide defines each tested boundary and the limits of the resulting evidence.

Repository map

Path Purpose
benchmark/ Scenario corpus, inventory, and tests
checks/ Questions, rule, evidence requirements, and exemplars with committed judge answers
src/invisiblebench/ Runtime, CLI, and Jury Card generation
scripts/ Scan, validation, and documentation commands
docs/ Method and run guides
data/ Aggregate results and historical releases

Use the inventory for current corpus facts and the scan contract for artifact fields.

Proof

uv run ruff check .
uv run python scripts/check_examples.py verify
uv run pytest benchmark/tests -q
uv run python scripts/lint_turn_indices.py --strict

Enable the required local hook with git config core.hooksPath .githooks.

Repository boundaries

This repository owns the public scenarios, check definitions, judge runtime, scan artifacts, and separate Safety/Care projection. It does not own GiveCare product policy, model training, clinical guidance, or real-world outcome claims.

The judge questions are public in checks/. Private transcripts, private scenarios, expected answers, and credentials stay in local storage. Public releases contain aggregate results with recorded provenance.

Committed historical releases keep their original bytes and labels. New releases use the current contract and a new release version. Historical web evidence remains in data/releases/web-bench-release.tar.gz.

Contributing

See contribution guidelines for the scenario contract, proof commands, and pull request checks. See the security policy for private security reports.

About

AI safety benchmark for long-term caregiving relationships. Tests crisis detection, regulatory compliance, and care quality across multi-turn conversations. Includes GiveCare system paper and InvisibleBench evaluation framework.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages