DataBench logo

DataBench

Hex

DataBench scores frontier AI on the analytics work that matters — reasoning, reporting, and investigation on realistic, messy warehouse data.

Categories

articleX13 Aug 2026

Introducing DataBench

DataBench v1: 100 realistic analytical tasks across Q&A and open-ended prompts, run in a synthetic Hex workspace — built because existing analytics benchmarks test "overspecified pub trivia" rather than the vague, directional questions people actually ask.

x.com

More in Evals

View all tools
AACR-Bench logo

AACR-Bench

An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.

Evals
BenchLocal logo

BenchLocal

Test LLMs on real tasks. Compare models side-by-side.

Evals
Braintrust logo

Braintrust

Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.

EvalsObservability

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.