Lucas Vitti

Prensa

site: prensa.lucas.mat.br updated

Prensa is a small service that takes a PDF — or Word, PowerPoint, Excel, HTML, EPUB, image — and gives back clean Markdown. Headings, lists and tables the model understands natively, instead of the same pages as images stuffing the context with tokens for nothing. It came from a personal itch: feeding scanned PDFs to Claude cost too much, answered poorly, and nobody wanted to build yet another pipeline for it. A press flattens the document; text is what's left.

Three engines behind one API: markitdown (Microsoft) for everything that isn't a PDF — DOCX, PPTX, XLSX, HTML, EPUB, ZIP; PyMuPDF for text-layer PDFs with a simple layout; and Docling (IBM) for scanned PDFs, with Portuguese + English OCR and complex tables. engine=auto picks the cheapest that fits — markitdown for non-PDF, PyMuPDF for text PDFs, Docling only when the first pages have no text layer or contain tables. Results are cached: same document, same engine, same options → one conversion, even when three different projects request it.

Three access tiers: public at / (no sign-up, 10 MB, fast engines only, nothing stored); signed-in at /app/ (Authentik SSO, 50 MB, Docling unlocked, history + API keys); and admin (everything, server-level keys). Under the hood, FastAPI + SQLite + two Docker containers (API + Docling worker, both read-only and unprivileged; the worker has no network at all). It also exposes an /mcp endpoint behind Bearer auth — plug straight into Claude Code without writing a shim.