Free open-source web software for signing PDF (alone or with others) and also organize pages, edit metadata and compress pdf
-
Updated
Jul 22, 2026 - JavaScript
Free open-source web software for signing PDF (alone or with others) and also organize pages, edit metadata and compress pdf
🧾Self-hosted AI invoice processing pipeline — extracts PDF data with Gemini AI, matches against Purchase Orders using 12+ business rules, and auto-routes to Payment, Review, or Rejection. Built with n8n, NocoDB, Node.js & Docker.
Browser web extension to extract key CPAP report data from website or PDFs and generate a structured clinical summary locally.
A responsible document revision tool for source-aware, personal, and transparent writing.
Turns redlined construction drawings into structured checklists; hybrid PDF-parse / vision extraction with a deterministic eval harness. Node + React.
Scalable GenAI-powered system to extract structured invoice data from PDFs & images, detect errors, and process large batches efficiently.
A privacy-focused tool to extract structured content (tables, text) from PDFs into editable Word & CSV formats. Features offline translation and runs locally. Powered by Docling & Flask.
PDF MCP server — give your AI agent PDF parsing, extraction & accessible-HTML tools over the Model Context Protocol. Connect Claude, Cursor, or ChatGPT to okraPDF.
An NLP & ontology-based pipeline that makes large collections of financial documents queryable through semantic search — transforming unstructured PDFs into a structured, searchable knowledge base.
open source PDF processing toolkit
Local-first desktop palette tool for scientific figures, journal graphics, PDF color extraction, and publication-ready exports for macOS and Windows.
A code-first engineering reference for automated customs brokerage & HS code classification pipelines — HTS mapping, duty calculation, document ingestion, and CBP ACE transmission.
Supplier PDF-to-Excel/CSV workflow with structured extraction, confidence scoring, validation flags, and human-review cues.
Maxcavator 2.0 is an intelligent, AI-native PDF Data Extraction and Retrieval-Augmented Generation (RAG) system. It fundamentally changes how you understand and interact with your PDF documents by instantly extracting complex structures (sections, tables, images), generating robust RAG indices.
Add a description, image, and links to the pdf-extraction topic page so that developers can more easily learn about it.
To associate your repository with the pdf-extraction topic, visit your repo's landing page and select "manage topics."