Measuring Autonomous AI Research
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
Elie Bakouch
xAI
xAI's frontier Grok model, tuned for long-running agents, coding, knowledge work, and visual projects.
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
Grok 4.6 is out! I've used it for a few weeks as my daily driver across the normal mix of coding and knowledge work, and built a few projects with it specifically to push on where it holds up.
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.