Evals

23 tools in Evals.

AACR-Bench logo

AACR-Bench

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

Evals
aimock logo

aimock

Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies.

Evals
BenchLocal logo

BenchLocal

Test LLMs on real tasks. Compare models side-by-side.

Evals
Braintrust logo

Braintrust

Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.

EvalsObservability
DataBench logo

DataBench

DataBench scores frontier AI on the analytics work that matters — reasoning, reporting, and investigation on realistic, messy warehouse data.

Evals
Evalite logo

Evalite

Evalite makes evals simple. Test your AI-powered apps with a local dev server.

Evals
ExtractBench logo

ExtractBench

A benchmark for schema-guided extraction from real enterprise documents. 370 documents, 4,869 pages, 67 document types, each with its own JSON Schema. Scored on value accuracy, completeness, and evidence.

EvalsExtraction
GEPA logo

GEPA

Optimize any text — prompts, code, agent architectures, configurations — using LLM-based reflection and Pareto-efficient evolutionary search. If you can measure it, you can optimize it.

Skills & PromptsEvals
Harbor logo

Harbor

Framework for evaluating and improving agents.

Evals
LLM-as-a-Verifier logo

LLM-as-a-Verifier

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and medical agentic benchmarks.

Evals
LocalMaxxing CLI logo

LocalMaxxing CLI

The official CLI for localmaxxing.com — run local LLM inference speed tests and quality benchmarks, then submit results from your terminal.

Evals
M

MemConflict

Benchmark and evaluation toolkit for long-term memory systems under memory conflicts.

Evals
Ori Eval logo

Ori Eval

Find the best model for your project. Ori Eval runs your agent and model on your prompts, asserts on the tools it called, and grades the answers with an LLM judge, so you catch regressions and pick the best model before you ship.

Evals
Phoenix logo

Phoenix

AI Observability & Evaluation

EvalsObservability
Pipette logo

Pipette

Compare foundation models on real devices across accuracy, throughput, latency, memory, quantization, runtime, and hardware.

Evals
PostTrainBench logo

PostTrainBench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

Evals
Prime Intellect logo

Prime Intellect

Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack.

EvalsSandboxesCompute
R

r0b0bench

Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.

Evals
SlopCodeBench logo

SlopCodeBench

SlopCodeBench: Measuring Code Erosion Under Iterative Specification Refinement

Evals
Supabase Evals logo

Supabase Evals

This repo answers how well can agents use Supabase across various tasks.

Evals
Terminal-Bench logo

Terminal-Bench

A benchmark for LLMs on complicated tasks in the terminal

Evals
vitest-evals logo

vitest-evals

A vitest extension for running evals.

Evals
V

VulcanBench

Open source, clear, transparent, real world llm benchmarks

Evals

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.