Evals

21 tools in Evals.

AACR-Bench logo

AACR-Bench

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

Evals
aimock logo

aimock

Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies.

Evals
BenchLocal logo

BenchLocal

Test LLMs on real tasks. Compare models side-by-side.

Evals
Braintrust logo

Braintrust

Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.

Evals
Dynobox logo

Dynobox

Cross-harness testing for multi-step agent flows

Evals
Evalite logo

Evalite

Evalite makes evals simple. Test your AI-powered apps with a local dev server.

Evals
ExtractBench logo

ExtractBench

A benchmark for schema-guided extraction from real enterprise documents. 370 documents, 4,869 pages, 67 document types, each with its own JSON Schema. Scored on value accuracy, completeness, and evidence.

EvalsExtraction
GEPA logo

GEPA

Optimize any text — prompts, code, agent architectures, configurations — using LLM-based reflection and Pareto-efficient evolutionary search. If you can measure it, you can optimize it.

Skills & PromptsEvals
Harbor logo

Harbor

Framework for evaluating and improving agents.

Evals
M

MemConflict

Benchmark and evaluation toolkit for long-term memory systems under memory conflicts.

Evals
Ori Eval logo

Ori Eval

Find the best model for your project. Ori Eval runs your agent and model on your prompts, asserts on the tools it called, and grades the answers with an LLM judge, so you catch regressions and pick the best model before you ship.

Evals
Phoenix logo

Phoenix

AI Observability & Evaluation

Evals
PostTrainBench logo

PostTrainBench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

Evals
Prime Intellect logo

Prime Intellect

Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack.

EvalsSandboxesCompute
R

r0b0bench

Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.

Evals
Raindrop logo

Raindrop

Monitor your AI Agent the right way. Get alerted when your agent fails in production, trace exactly what went wrong, and prove your fix worked.

Evals
SlopCodeBench logo

SlopCodeBench

SlopCodeBench: Measuring Code Erosion Under Iterative Specification Refinement

Evals
Supabase Evals logo

Supabase Evals

This repo answers how well can agents use Supabase across various tasks.

Evals
Terminal-Bench logo

Terminal-Bench

A benchmark for LLMs on complicated tasks in the terminal

Evals
vitest-evals logo

vitest-evals

A vitest extension for running evals.

Evals
V

VulcanBench

Open source, clear, transparent, real world llm benchmarks

Evals

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.