AACR-Bench
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
DataBench scores frontier AI on the analytics work that matters — reasoning, reporting, and investigation on realistic, messy warehouse data.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.