PostTrainBench logo

PostTrainBench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

Categories

announcementPostTrainBench28 Jul 2026

PostTrainBench v1.1: Hardening the benchmark against reward hacking

PostTrainBench v1.1 clarifies the boundary between legitimate benchmark hill climbing and item specific contamination, with specialized checks for external LLM API use, model substitution, and direct lookup.

PostTrainBench team
posttrainbench.com

More in Evals

View all tools
AACR-Bench logo

AACR-Bench

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

Evals
BenchLocal logo

BenchLocal

Test LLMs on real tasks. Compare models side-by-side.

Evals
Braintrust logo

Braintrust

Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.

Evals

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.