PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
-
Updated
Jul 23, 2026 - Java
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Free open-source web software for signing PDF (alone or with others) and also organize pages, edit metadata and compress pdf
Use TradeRepublic in terminal and mass download all documents
JavaScript bindings for MuPDF
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.
Convert your PDFs and EPUBs into audiobooks effortlessly. Features intelligent text extraction, customizable text-to-speech settings, and efficient processing for low-resource systems.
Turn PDFs into clean, structured Markdown
Claude Code and Codex SKILLs for PDF, Excel, Word, and PowerPoint manipulation — extraction, forms, formulas, tracked changes, adapted from Anthropic skills.
An MCP server that lets Claude Code and other AI agents work through large PDFs without overflowing their context — search by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts.
Self-healing PDF extraction that flags what it can't read instead of dropping it — and now certifies any extractor's output, catching silently-dropped pages. #2 of all tools, #1 free on opendataloader-bench (0.903). 7-tool MCP. MIT, free.
A professinal CLI workflow for PhD students to extract, analyze, and visualize academic papers into structured Markdown and Obsidian Canvas.
Translate many large PDF Reports for free using Python.
Want to search arXiv papers, fetch metadata, and extract full-text PDFs without leaving your editor? This MCP server connects any MCP-compatible client (Claude Code, etc.) directly to arXiv.
X-ray for documents: lossless PDF & PPTX extraction to JSON with bounding boxes, fonts, and colors — CLI, HTTP API, and an in-browser playground. Rust + PDFium.
Multimodal web content extraction engine backed by Qt WebEngine.
Fast pure-Rust PDF extraction library and CLI by Clark Labs Inc. — 10–50x faster than pdfplumber for text, word, table, layout, image, and metadata extraction.
Add a description, image, and links to the pdf-extraction topic page so that developers can more easily learn about it.
To associate your repository with the pdf-extraction topic, visit your repo's landing page and select "manage topics."