Skip to content

Frontier-evals sweep: Petri, Opus 4.6 interpretability audit, NLA, CoT-monitorability, Deployment Safety Hub, EVMbench - #85

Open
xdotli wants to merge 1 commit into
mainfrom
frontier-evals-sweep
Open

Frontier-evals sweep: Petri, Opus 4.6 interpretability audit, NLA, CoT-monitorability, Deployment Safety Hub, EVMbench#85
xdotli wants to merge 1 commit into
mainfrom
frontier-evals-sweep

Conversation

@xdotli

@xdotli xdotli commented Aug 20, 2026

Copy link
Copy Markdown
Member

Vetted via 4 parallel research subagents against the CONTRIBUTING.md bar, then spot-checked by fetching each live source myself. All additions are recent (Feb–Jul 2026) and I dropped several subagent recommendations that were either already listed, stale, or redundant with existing entries — see reasoning below.

§10 (Safety / adversarial evaluation)

  • Petri (meridianlabs-ai/inspect_petri) — confirmed canonical repo (Anthropic donated the project to Meridian Labs; the older safety-research/petri is now legacy v2). Live: 1.3k★. Verbatim-verified: 38 judging dimensions, 170+ seeds, used in the Claude 4 / Sonnet 4.5 System Cards and by UK AISI.
  • Claude Opus 4.6 System Card §6.5 — "Inhibiting internal representations of evaluation awareness": suppresses the model's internal eval-awareness representation via activation oracles/attribution graphs/SAE features and re-runs the alignment evals — a causal audit, not just a transcript read. Confirmed live in the system card TOC.
  • Natural Language Autoencoders (Anthropic) — decodes activations to natural language; verbatim numbers confirmed live: 26% eval-awareness on SWE-bench Verified, 16% on a destructive-coding test.
  • Evaluating chain-of-thought monitorability (OpenAI) — the primary methodology paper (not a specific system card), open-sourced at github.com/openai/monitorability-evals; g-mean² metric confirmed live on the GPT-5.4/5.6 system card pages.
  • Deployment Safety Hub — described structurally per CONTRIBUTING's rule that live-leaderboard numbers get pinned or dropped, not cited loose.

§9 (Agent-specific evaluation)

  • EVMbench — confirmed NOT already listed (checked against the existing SWE-Lancer/MLE-bench/PaperBench entries first). 117 vulnerabilities / 40 audits confirmed, independently audited by OpenZeppelin.

Declined / dropped (for your visibility, not included in this diff)

  • openai/frontier-evals repo as its own §5a entry — one subagent recommended it, but SWE-Lancer/MLE-bench/PaperBench are already individually listed (that subagent was wrong that they weren't); not worth a bullet just to cross-link.
  • GPT-5 "Production Benchmarks" saturation methodology — real but from the original Aug-2025 GPT-5 card, not recent enough on its own.
  • Claude Opus 4.5 sabotage evals, Gemini 3 Pro harmful-manipulation, Grok 4 RMF categories — model-card sweep correctly flagged these as better cited via their primary papers (SHADE-Arena, arXiv:2603.25326) or lacking methodological novelty.

…nguage Autoencoders, CoT-monitorability framework, Deployment Safety Hub, and EVMbench
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant