SKAVIO / JOURNAL
SKAVIO JOURNAL · PROCUREMENT WORKFLOWS
Microsoft Support — Keeping leading zeros and large numbers
Skavio Flow turns orders, invoices, scans and text into JSON, CSV and XLSX using custom fields and saved company workflows. For procurement teams, the useful shift is not simply making a document readable. It is defining the information the business needs once, checking extracted values against the source, and producing reviewed structured data for downstream reuse.
Define the fields before processing the document
Consider an illustrative procurement team receiving purchase orders and supplier documents in different layouts. Its first task is to decide what a usable record should contain—not to reproduce every element on the page. Flow supports custom extraction fields and saved company workflows, giving the team a reusable starting point. Separate information about the whole document from information about individual order lines. OpenPeppol’s November 2024 Order syntax offers a useful reference for that distinction, covering identifiers, dates, trading parties, delivery information and order lines. It is inspiration for field design, not evidence of Skavio standards compliance.
- Document-level information to consider: purchase-order number, issue date, buyer, supplier and currency.
- Delivery information to consider: delivery address and requested delivery date, where present.
- Line-level information to review: product code, description, quantity, unit, unit price and line amount.
Extract from representative files—not an assumed perfect layout
Flow reads embedded PDF text locally, uses OCR for scans, and uses the configured OpenAI model for structured extraction. This is not an entirely local processing pipeline. Begin with representative purchase orders and supplier documents from the intended workflow. Test the formats and scan conditions the team actually encounters. Support for orders and scans does not establish reliable extraction across every supplier document category or layout; a useful pilot checks whether the chosen fields can be extracted and reviewed in those specific files.
Make source review the gate before export
Flow includes source evidence and review warnings alongside extracted results. Use them to check values against the original document before treating the output as a business record. Warnings support that review; they should not be assumed to flag every possible error. Pay particular attention to identifiers, dates, currencies, quantities, units, prices, totals and line items. An order number that looks plausible can still be wrong, and an extracted total is not proof that the document’s arithmetic has been validated. Human review remains the step that determines whether the data is ready for downstream use.
- Check order references and product codes character by character where ambiguity matters.
- Verify date interpretation, currency and quantity units against the source.
- Review line items and totals separately rather than assuming one confirms the other.
Choose an export that preserves the meaning of the data
Flow offers JSON, CSV and XLSX, including XLSX document and line-item worksheets. Choose the format around the next task: spreadsheet review, tabular import or application processing. For downstream JSON design, a useful proposed model separates document-level fields from an array of line items. JSON supports objects and arrays, but this is a design recommendation—not Flow’s verified export contract. Inspect the actual output before building mappings or assuming a particular nesting structure. Spreadsheets need another safeguard: identifier columns should remain text. Microsoft documents that Excel can remove leading zeros and lose precision in numeric values beyond 15 significant digits. When importing CSV, set purchase-order numbers and product codes to text so spreadsheet conversion does not alter them.
Reuse the workflow, then validate the next supplier
Once the field selection and exported results meet the team’s needs, reuse the saved company workflow for subsequent documents. Keep the review step in place, especially when a new supplier or document layout enters the process. Reusing field definitions makes the process repeatable; it does not make every extraction automatically correct. Teams developing a downstream integration can access Flow extraction through the Skavio Processing API. That still requires application-specific mapping and validation; it is not a built-in ERP connector or automatic posting feature. Under canonical Pricing V2, Flow costs 25 shared credits per page, with a minimum of 25 credits per document. Web processing and API jobs use the same wallet. Company signup receives 1,000 trial credits, subject to the current signup policy.
The reusable asset is the reviewed field structure—not just the text recovered from a document.
Sources
- OpenPeppol — Order syntax, November 2024 release — Separates order identifiers, dates, trading parties, delivery information and order lines; used only as a field-design reference, not a Skavio compliance claim.
- IETF RFC 8259 — JSON Data Interchange Format — Defines JSON objects, arrays and value types; supports the proposed downstream design discussion, not a claim about Flow’s exact export schema.
- Microsoft Support — Keeping leading zeros and large numbers — Documents Excel’s handling of leading zeros and numeric precision beyond 15 significant digits, and importing identifier columns as text.