Prensa
Prensa is a small service that takes a PDF — or Word, PowerPoint, Excel, HTML, EPUB, image — and gives back clean Markdown. Headings, lists and tables the model understands natively, instead of the same pages as images stuffing the context with tokens for nothing. It came from a personal itch: feeding scanned PDFs to Claude cost too much, answered poorly, and nobody wanted to build yet another pipeline for it. A press flattens the document; text is what's left.
Three engines behind one API:
markitdown (Microsoft) for everything that isn't
a PDF — DOCX, PPTX, XLSX, HTML, EPUB, ZIP;
PyMuPDF for text-layer PDFs with a simple layout;
and Docling (IBM) for scanned PDFs, with
Portuguese + English OCR and complex tables.
engine=auto picks the cheapest that fits — markitdown
for non-PDF, PyMuPDF for text PDFs, Docling only when the first
pages have no text layer or contain tables. Results are cached:
same document, same engine, same options → one conversion, even
when three different projects request it.
Three access tiers: public at /
(no sign-up, 10 MB, fast engines only, nothing stored);
signed-in at /app/ (Authentik SSO,
50 MB, Docling unlocked, history + API keys); and
admin (everything, server-level keys). Under the
hood, FastAPI + SQLite + two Docker containers (API + Docling
worker, both read-only and unprivileged; the worker has no
network at all). It also exposes an /mcp endpoint
behind Bearer auth — plug straight into Claude Code without
writing a shim.