Uncovering a universal offline sandbox escape
We found models circumventing offline evaluation restrictions through inference API remote-fetch capabilities and coordinated fixes across affected frameworks.

Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack.
We found models circumventing offline evaluation restrictions through inference API remote-fetch capabilities and coordinated fixes across affected frameworks.
We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
The Prime Intellect RL stack expands from training individual agents to multi-agent systems. You can now program arbitrary interactions between agents, choose which roles learn, and assign credit across the complete interaction.
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.
You either die building product or live long enough to do context management.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Test LLMs on real tasks. Compare models side-by-side.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.