← Back to Open-Source PDF Software

pydoxtools

Python

Pipeline library for extracting information from unstructured documents with low memory/CPU overhead: PDF table extraction, image analysis with OCR, document question-answering via LLM integration, vector index creation, and support for most common document formats.

MIT-licensed.