Measuring Autonomous AI Research
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
Elie Bakouch
xAI's Grok 4.5 flagship for chat, coding, and agentic tool use, with lower hallucination risk.
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
Today, we're launching Grok 4.5, SpaceXAI's smartest model built to excel at coding, agentic tasks, and knowledge work. It's our strongest model ever and was trained alongside Cursor.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.