AACR-Bench
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
21 tools in Evals.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies.
Test LLMs on real tasks. Compare models side-by-side.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
Cross-harness testing for multi-step agent flows
Evalite makes evals simple. Test your AI-powered apps with a local dev server.
A benchmark for schema-guided extraction from real enterprise documents. 370 documents, 4,869 pages, 67 document types, each with its own JSON Schema. Scored on value accuracy, completeness, and evidence.
Optimize any text — prompts, code, agent architectures, configurations — using LLM-based reflection and Pareto-efficient evolutionary search. If you can measure it, you can optimize it.
Framework for evaluating and improving agents.
Benchmark and evaluation toolkit for long-term memory systems under memory conflicts.
Find the best model for your project. Ori Eval runs your agent and model on your prompts, asserts on the tools it called, and grades the answers with an LLM judge, so you catch regressions and pick the best model before you ship.
AI Observability & Evaluation
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack.
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
Monitor your AI Agent the right way. Get alerted when your agent fails in production, trace exactly what went wrong, and prove your fix worked.
SlopCodeBench: Measuring Code Erosion Under Iterative Specification Refinement
This repo answers how well can agents use Supabase across various tasks.
A benchmark for LLMs on complicated tasks in the terminal
A vitest extension for running evals.
Open source, clear, transparent, real world llm benchmarks
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.