BenchLocal
Test LLMs on real tasks. Compare models side-by-side.

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.