Embedded text or scanned pages: two PDF routes, one Skavio Flow review workflow

Skavio Flow brings PDFs with embedded text and scanned documents into the same workflow for reviewing fields and exporting structured data. What changes is how the text is obtained: embedded PDF text is read locally, while scans use OCR. Neither route removes the need to check extracted dates, totals and line items against the source.

1. Inspect the PDF—not just its appearance

An invoice generated by accounting software and a scan of a paper invoice may look similar on screen, but their text can be stored differently. Try selecting and copying a representative heading, date and amount. This is a useful clue, not proof of how the PDF was created: a scanned document can also contain selectable text added by an earlier OCR process.

2. Define the fields before processing

Choose the information you need for the handoff, then use custom extraction fields or a saved company workflow in Flow. For an invoice, an illustrative field set might include the supplier name, invoice number, invoice date, currency, total and line-item details. These are examples, not mandatory or default fields. Reusing a saved workflow lets subsequent documents follow the same extraction requirements, rather than redefining the task for each supplier PDF.

3. Separate text acquisition from field extraction

Flow reads embedded PDF text locally and uses Worker Fabric OCR for scans. Structured extraction then uses the configured OpenAI model to turn the acquired text into the requested fields. These are distinct stages: obtaining readable words is not the same as identifying an invoice total or reconstructing its line items. That distinction matters even for PDFs with embedded text. A table may be stored as separately positioned text rather than as ready-made rows and columns. The presence of a text layer therefore does not guarantee a clean business record. Local PDF text reading also does not mean the entire Flow job runs offline or on your device.

4. Review evidence and check corrections

Both input routes lead to Flow’s downstream review-and-export workflow, with source evidence and review warnings. Use these to guide your checks, and compare extracted values with the original document before accepting them. For blurry or low-resolution scans, obtaining a clearer source is sensible preparation—not a guarantee that OCR errors will disappear. A Skavio acceptance check on October 1, 2026 used two synthetic PDF invoices and verified extraction and export. It also verified that a scalar correction—a correction to a single field value—regenerated results without a second job. That is limited functional evidence, not an accuracy benchmark or proof of every possible line-item editing operation.

5. Export the reviewed fields—not merely the recognized text

After review, choose the output that fits the next step. Flow provides XLSX with document and line-item worksheets, alongside CSV and JSON. The useful handoff is the reviewed structure: OCR supplies text from a scan, while field extraction and review prepare that text for finance or operations. Flow extraction costs 25 shared credits per page, with a minimum of 25 credits per document. Web processing tools and the Processing API use the same Pricing V2 wallet. Company signup receives 1,000 trial credits, subject to the current signup policy.

Bring an invoice or scanned order to Skavio Flow and review the extracted fields before exporting.