MikeTrendsTrends right now

⬢github Python · 116 ★ +116 since we first saw it · pushed 41 min ago · MIT

beatrizalmeidaf/papero-pdf-text-extractor

Fast, open-source PDF text extraction API. Files never stored.

papero is a CPU-only, no-ML Python library/API that extracts structured content from PDFs: reading order for multi-column layouts, tables, LaTeX formulas, figures with captions, and bounding boxes for every block. Output goes to Markdown, JSON, Word, or Excel, with a browser demo and OCR for scanned pages; files aren't stored.

Why now: Featured on Hacker News as a lightweight PDF parser handling layout, tables, formulas and bounding boxes — appealing amid heavy ML-based document parsers.

Who it is for: Developers building RAG pipelines, search indexing, or document-conversion tools who want fast structure-aware PDF extraction without GPU/ML dependencies.

pdftext-extractionpythonocrragapi

pdfpdf-extractionpdf-processingpdf-to-textpdf-toolstext-extraction

Open on GitHub →

Stars over our 22 snapshots: 0 to 116, since 5 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/beatrizalmeidaf/papero-pdf-text-extractor