Ori Eval: Find the Best Model for What You're Building
Ori Eval runs your agent on your own prompts, asserts on the tools it called, and grades answers with an LLM judge. Find the best model for what you're building.

Find the best model for your project. Ori Eval runs your agent and model on your prompts, asserts on the tools it called, and grades the answers with an LLM judge, so you catch regressions and pick the best model before you ship.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Test LLMs on real tasks. Compare models side-by-side.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.