Intelligent Document Processing
Invoices, remittances, purchase orders and supporting schedules still arrive as PDFs, scans and photographs. Somebody reads each one and types it into a system that will later reconcile it against a document it cannot see. Document processing removes the typing, matches automatically, and routes only the genuine exceptions to a person.
1
Sample
Real documents, not ideal ones
2
Extract
Field level capture
3
Match
Against your records
4
Route
Exceptions to people
5
Monitor
Accuracy over time
How We Build It
01
Document Sampling and Analysis
We work from a representative sample of what actually arrives, including the poor scans and the supplier whose layout changes quarterly. Building against clean examples produces a pipeline that fails on contact with reality.
- Representative sample across suppliers and formats
- Volume and arrival pattern by document type
- Quality range including scans and photographs
- Fields required per document type
02
Extraction Build
We build extraction for the fields that matter rather than attempting the whole document. Confidence scoring is configured per field so low confidence values are flagged instead of silently accepted.
- Field level extraction with confidence scoring
- Handling of multiple layouts per document type
- Line item extraction where required
- Thresholds set per field by consequence of error
03
Matching and Validation
Extracted data is matched against your records: invoice to purchase order to receipt, remittance to open items. Validation catches what extraction cannot, such as an arithmetically impossible total.
- Matching against orders, receipts and open items
- Tolerance rules agreed with finance
- Arithmetic and business rule validation
- Duplicate detection across periods
04
Exception Routing
Anything failing extraction, matching or validation goes to a person with the document and the reason displayed together. The objective is that a human sees the exceptions and nothing else.
- Exceptions routed with document and reason side by side
- Routing rules by exception type and value
- Correction captured to improve future extraction
- Full audit trail of every automated and manual action
05
Monitoring and Improvement
Accuracy is measured continuously rather than assumed from an initial test. Supplier layout changes are the most common cause of degradation and they arrive without warning.
- Extraction accuracy tracked per field and supplier
- Alerting when accuracy falls below threshold
- Periodic retraining from captured corrections
- Reporting on volume, straight through rate and savings
What You Receive
- Extraction pipeline configured for your document types
- Matching and validation rules agreed with finance
- Exception routing with document and reason displayed together
- Full audit trail of automated and manual actions
- Accuracy monitoring with threshold alerting
- Reporting on volume, straight through rate and time saved
Indicative Timeline
A first document type usually takes six to ten weeks. Additional types are faster because the matching and routing framework already exists and only extraction differs.
- Sampling and analysis: one to two weeks
- Extraction build and tuning: three to four weeks
- Matching, validation and routing: two weeks
- Parallel running and handover: one to two weeks
What We Process
Automation is applied to documents arriving in volume with a stable set of fields.
Supplier Invoices
Header and line item extraction, matched to purchase orders and goods received.
Remittance Advice
Payment allocation against open items, including partial and consolidated payments.
Purchase Orders
Capture of orders arriving as documents rather than through a system integration.
Bank Statements
Statement capture where electronic feeds are unavailable from the institution.
Delivery Notes
Proof of delivery matched to orders and invoices to complete the three way match.
Claim Documents
Supporting documentation captured, validated and routed for assessment.
Frequently Asked Questions
How accurate is extraction?
On clean, consistent documents it is high. On poor scans and variable layouts it is lower, which is exactly why confidence thresholds and exception routing exist. The design assumes imperfection rather than promising perfection.
What happens when it gets something wrong?
Low confidence values are flagged before they enter your system. Where a value is confidently wrong, validation and matching normally catch it, and the correction feeds back to improve future extraction.
Do we still need people in accounts payable?
Yes, doing different work. They stop keying and start handling genuine exceptions and supplier queries. Volume per person rises substantially; the role changes rather than disappearing.
Can it handle handwritten documents?
Partially and unreliably. Printed text extracts well, handwriting much less so. Where handwritten documents form a meaningful share of volume we say so during sampling rather than discovering it in build.
What if a supplier changes their invoice layout?
It happens regularly and accuracy monitoring is designed to catch it. Extraction degrades for that supplier, the threshold alert fires, and it is retuned. Without monitoring, the first sign is usually a reconciliation problem weeks later.
Is our data sent to a third party?
That depends on the design and it is a decision you make. Where documents contain personal information, processing can be kept inside your environment, which costs more but keeps the POPIA position simpler.
Related Services
This sits inside our AI Services practice. Related work: Workflow Automation for the approval routing that follows capture, and Financial and Management Accounting if you would rather outsource the function than automate it.
