articleSupabase31 Jul 2026
Introducing Supabase Evals
Our open-source benchmark for how well AI coding agents build with Supabase.
Matt Rossman
supabase.com
This repo answers how well can agents use Supabase across various tasks.
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
Test LLMs on real tasks. Compare models side-by-side.
Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.