Image and office inputs

Extracting from images and office documents: PNG, JPEG, TIFF and WebP, plus DOCX, XLSX and CSV, and what geometry each format can actually report.

Ovrin is not limited to PDFs. It also works with image inputs and office-style formats such as DOCX, XLSX, and CSV, while using OCR or other reading strategies only where needed.

Why this is useful

Many documents flow through mixed source types. A document pipeline should accept those inputs without requiring a separate implementation for each format.