dolico — document extraction
Upload a PDF, scan, DOCX, XLSX, PPTX, ODF, RTF, EPUB, CSV, Markdown or plain text
document and get its text, tables and structure back as JSON or Markdown.
How it works
Every page is routed independently to the cheapest engine that can read it: direct
text extraction where the text is already in the file, OCR where it is a scan, and a
vision model for scans the OCR tiers read badly. You see which engine read each page
and what it scored — before paying.
Pricing
Per page, by the engine that read it, with a minimum charge per document. The first
page is shown in full for free, so you can check the extraction before you buy the
rest. One payment, no account, no subscription.
Privacy: no third-party AI sees your document
Every engine runs on our own hardware. Text extraction is a local library; scanned
pages go to a local PaddleOCR service, and the hardest of them to a local vision
model. There is no call to OpenAI, Anthropic, AWS Textract or Google Document AI
anywhere in the pipeline. Uploads that are never paid for are deleted within a day;
purchases are deleted after their retention window. See
privacy and terms.
Self-hosting
The same system runs as containers inside your own infrastructure, so nothing
crosses your network boundary at all — no licence check, no telemetry, no
dependency on this site. That is the point for anyone handling privileged, medical
or otherwise confidential documents that a cloud OCR API cannot legally touch. See
self-hosting.
This page needs JavaScript to upload a document.