Some reflections on building evals and benchmarking frontier models
Building eval suites for VulcanBench is really fascinating, and also so much more challenging than I was expecting.
A benchmark for LLMs on complicated tasks in the terminal
Building eval suites for VulcanBench is really fascinating, and also so much more challenging than I was expecting.
Learn how Cline evaluates and improves open weight coding agents using Terminal Bench, with practical heuristics for model performance, token efficiency, reasoning, providers, and eval optimization.
Benchmarks are where SOTA has to earn its name. This post is about designing good tasks.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Test LLMs on real tasks. Compare models side-by-side.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.