Document handling

Which inputs Ovrin reads — PDFs with and without a text layer, scans, images and office formats — and how it decides which reading each page gets.

Ovrin is designed to handle a wide range of document inputs, including PDFs, images, scanned files, and office documents.

Supported inputs

The upstream project describes support for:

  • PDFs with a text layer
  • PNG, JPEG, TIFF
  • scanned PDFs
  • DOCX, XLSX, CSV

Handling model

The product does not hardcode a single document type. It treats receipts, invoices, bank statements, transcripts, and contracts as similar extraction problems expressed as typed Go schemas.