Measuring Autonomous AI Research
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
DataBench v1: 100 realistic analytical tasks across Q&A and open-ended prompts, run in a synthetic Hex workspace — built because existing analytics benchmarks test "overspecified pub trivia" rather than the vague, directional questions people actually ask.
Our most agentic Sonnet yet, with top-tier intelligence for coding and everyday professional work.
Pair Opus as an advisor with Sonnet or Haiku as an executor, and get Opus-level intelligence in your agents at a fraction of the cost.
Anthropic
Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.
Anthropic
Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks.
Anthropic
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.