anydoc
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
9 tools in Extraction.
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
Open-source toolkit, built by the EPIC Data Lab at UC Berkeley, for creating LLM-powered pipelines that extract, transform, and link knowledge from unstructured documents.
A benchmark for schema-guided extraction from real enterprise documents. 370 documents, 4,869 pages, 67 document types, each with its own JSON Schema. Scored on value accuracy, completeness, and evidence.
Repository for building knowledge graphs from specific datasets using generative language model through ollama
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
Port from Google's LangExtract to Typescript
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Industrial-strength Natural Language Processing (NLP) in Python
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.