articleX20 Aug 2026

Some reflections on building evals and benchmarking frontier models

Building eval suites for VulcanBench is really fascinating, and also so much more challenging than I was expecting.

Morgan Linton@morganlinton
x.com
articleCline18 Aug 2026

Open-sourcing evals for open-weight agents

Learn how Cline evaluates and improves open weight coding agents using Terminal Bench, with practical heuristics for model performance, token efficiency, reasoning, providers, and eval optimization.

Ara Khan
cline.bot
articleX10 Jul 2026

Good Benchmarks

Benchmarks are where SOTA has to earn its name. This post is about designing good tasks.

Ivan Bercovich@neversupervised
x.com

More in Evals

View all tools
AACR-Bench logo

AACR-Bench

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

Evals
BenchLocal logo

BenchLocal

Test LLMs on real tasks. Compare models side-by-side.

Evals
Braintrust logo

Braintrust

Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.

EvalsObservability

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.