Skip to content

About

No description, website, or topics provided.

Resources

Stars

55 stars

Watchers

2 watching

Forks

Repository files navigation

ARC-AGI-3 Benchmarking

Quickstart

Install uv if not already installed.

  1. Clone the arc-agi-3-benchmarking repo, enter the directory
git clone https://github.com/arcprize/arc-agi-3-benchmarking.git
cd arc-agi-3-benchmarking
  1. Install dependencies
uv venv
uv sync
  1. Copy .env.example to .env
cp .env.example .env
  1. Get an API key from the Arc Prize Website and set it as an environment variable in your .env file.
ARC_API_KEY=your_api_key_here
  1. Run the benchmarking agent against ls20.
uv run main.py --game=ls20

Running the Official Benchmarking Agent

  1. Get a model provider API key

Provider key links:

  1. Set your provider keys as environment variables in your .env file.
ANTHROPIC_API_KEY=your_anthropic_key_here
OPENAI_API_KEY=your_openai_key_here
GOOGLE_API_KEY=your_google_key_here
XAI_API_KEY=your_xai_key_here
GROK_API_KEY=your_grok_key_here
DEEPSEEK_API_KEY=your_deepseek_key_here
GROQ_API_KEY=your_groq_key_here
OPENROUTER_API_KEY=your_openrouter_key_here
FIREWORKS_API_KEY=your_fireworks_key_here
  1. View available games (there should be 25).
uv run main.py --list-games
  1. View available model config.
uv run main.py --list-configs
  1. Run the official benchmarking agent against a game:
uv run main.py --game=ls20 --config=openai-gpt-5-4-2026-03-05

Native Anthropic configs are also available:

uv run main.py --game=ls20 --config=anthropic-opus-4-7-low
uv run main.py --game=ls20 --config=anthropic-opus-4-7-low-thinking
  1. Or on all games:
uv run main.py --config=openai-gpt-5-4-2026-03-05

Standard and Provider Adapter harnesses

ARC-AGI-3 games span many model calls, so the harness must determine what carries forward between actions. The Standard harness uses a provider-neutral text history and asks the model to preserve useful discoveries in visible notes. These configurations use manual_rolling.

The Provider Adapter harness uses the provider's native conversation and reasoning state. It uses native compaction when the provider supplies it and a domain-neutral harness summary otherwise. These configurations use continuous_conversation; openai-gpt-5-6-sol-max-provider-adapter and google-gemini-3-8-flash-low-provider-adapter are examples. The xai-grok-4-7-low-provider-adapter profile adds xAI-native encrypted reasoning replay and separate Responses compaction using XAI_API_KEY. See the xAI Provider Adapter for its request contract, compaction boundaries, and opt-in live tests.

The Anthropic Provider Adapter profile is anthropic-opus-5-low-provider-adapter. It uses Opus 5 at low reasoning effort, preserves native thinking blocks between actions, and uses Anthropic's native on-demand compaction at a 175k completed- context threshold. It summarizes completed history before presenting the next frame, so the newest observation reaches the action request unchanged. It records provider-reported thinking-token usage when available without double-counting output tokens. The profile also enables Anthropic automatic prompt caching with the default 5-minute TTL so repeated native history can be reused. Set ANTHROPIC_API_KEY to use it:

uv run main.py --game=ls20 --config=anthropic-opus-5-low-provider-adapter

See runtime state adapters for replay, compaction, recording, and data-handling details.

DeepSeek thinking mode can opt into deepseek.chat_completions.v1. The adapter uses DeepSeek's documented tools protocol, replays exact reasoning_content, and uses harness summary compaction. See the DeepSeek Provider Adapter guide.

Both harnesses use the same games, actions, limits, and scoring. The Standard harness supports controlled comparisons across providers, while the Provider Adapter harness measures performance using provider-native context management. Their results should be reported separately and clearly labeled.

  1. View your scorecard

When you run a benchmark, a scorecard is saved on the ARC server. If you are logged in, you can browse your saved scorecards at arcprize.org/scorecards.

License

This project is licensed under the MIT License. See the LICENSE file for details.

About

No description, website, or topics provided.

Resources

Stars

55 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages