⬢github Python · 116 ★ +116 since we first saw it · pushed 41 min ago · MIT
beatrizalmeidaf/papero-pdf-text-extractor
Fast, open-source PDF text extraction API. Files never stored.
papero is a CPU-only, no-ML Python library/API that extracts structured content from PDFs: reading order for multi-column layouts, tables, LaTeX formulas, figures with captions, and bounding boxes for every block. Output goes to Markdown, JSON, Word, or Excel, with a browser demo and OCR for scanned pages; files aren't stored.
Why now: Featured on Hacker News as a lightweight PDF parser handling layout, tables, formulas and bounding boxes — appealing amid heavy ML-based document parsers.
Who it is for: Developers building RAG pipelines, search indexing, or document-conversion tools who want fast structure-aware PDF extraction without GPU/ML dependencies.
pdfpdf-extractionpdf-processingpdf-to-textpdf-toolstext-extraction
Stars over our 22 snapshots: 0 to 116, since 5 h ago.
Where people talked about it
API: https://socialmediatrends-api.osmike.com/v1/repos/beatrizalmeidaf/papero-pdf-text-extractor