Open-source platform for extracting structured data from documents using AI.
-
Updated
May 15, 2025 - JavaScript
Open-source platform for extracting structured data from documents using AI.
Extract and download key-value pairs, tables, and paragraphs from your scanned pdf, jpg, and png documents as CSV files.
AI-powered complaint management for pharma QA (API/FDF manufacturing). LangGraph agents + Groq LLMs extract complaint details from documents into a React form, with an audit trail and QMS-aware chat assistant. React, FastAPI, MySQL, Docker Compose.
Claude Code plugin: auto-extracts PDF, DOCX, XLSX into Markdown with embedded images. Pure JavaScript, no native modules
Document schema extraction framework for regulated industries. Parse complex documents into versioned, comparable structured data with citation tracking and audit trails.
Documentation for the DocumentPro MCP server — extract and classify structured data from invoices, POs, receipts, and tax forms with any MCP-compatible AI agent
Automated invoice processing pipeline built on n8n — extracts structured data from PDFs using DeepSeek, validates and stores in PostgreSQL, alerts via Discord
A Node.js-based analysis tool that extracts text from .pdf, .docx, and .md files to generate AST (Abstract Syntax Tree) outputs using remark and markdown-it.
Add a description, image, and links to the document-extraction topic page so that developers can more easily learn about it.
To associate your repository with the document-extraction topic, visit your repo's landing page and select "manage topics."