DocETL
Open-source toolkit, built by the EPIC Data Lab at UC Berkeley, for creating LLM-powered pipelines that extract, transform, and link knowledge from unstructured documents.
Repository for building knowledge graphs from specific datasets using generative language model through ollama
Open-source toolkit, built by the EPIC Data Lab at UC Berkeley, for creating LLM-powered pipelines that extract, transform, and link knowledge from unstructured documents.
A benchmark for schema-guided extraction from real enterprise documents. 370 documents, 4,869 pages, 67 document types, each with its own JSON Schema. Scored on value accuracy, completeness, and evidence.
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.