A package for parsing PDFs and analyzing their content using LLMs.
-
Updated
Aug 6, 2024 - Python
A package for parsing PDFs and analyzing their content using LLMs.
LyraPDF: convert a PDF to JSON or MarkDown
Python library to extract transaction data from Indian Mutual Fund CAS (Consolidated Account Statement) PDFs — supports CAMS and KFintech — into CSV, DataFrame, JSON, or dict
This repository will assist you in scrapping data from multiple websites. It will identify, download and classify the latest pdf files published on a website as per the users requirement. This can be used for automating various operations involved in market research.
LLM for generating synthetic data from published papers
Inspired by PEStudio and Didier Stevens' tools
yet another pdf texts and tables extractor
Extract text, images, tables, and metadata from PDF files using Python. Built with PyPDF2, PyMuPDF, pdfplumber, and pdfminer. Helpful for practicing document parsing and data extraction tasks.
Production-grade Multilingual RAG API powered by FastAPI, BAAI/bge-m3 embeddings (100+ languages), ChromaDB, and Groq (Llama 3.3 70B). Features multi-format ingestion (PDF, DOCX, PPTX, XLSX, TXT) and grounded answers with deterministic source citations.
ATS Resume & Job Description Analyzer powered by deterministic NLP entity extraction, symbol-preserving tokenization, and calibrated TF-IDF machine learning scoring.
This is an AI-powered chatbot that automatically engages with users who sign up on your website. It can answer questions about your company based on PDF documents and send automated emails.
To associate your repository with the pdfparser topic, visit your repo's landing page and select "manage topics."