orwelllab
Automation / System guide

Intelligent document processing: test it on your own documents first

A vendor-neutral guide to pulling data out of invoices, forms and other business documents, with an eight-step test plan to run on your own documents before anything goes live.

In brief
  • The UK government's Extract tool cut the time to create a planning record by around a half (40 to 55%), with every record reviewed (i.AI).
  • Reading the page image scored 92.71% on scanned invoices, against 64.03% for converting it to text first (Berghaus et al., 2025).
  • On fixed layouts, table-based OCR reached an F1 of 1.0 in 0.97 seconds, where direct LLM extraction took 13.4 to 13.6 seconds (Wang and Shen, 2025).
  • A clean run on about 300 documents shows, with 19-in-20 confidence, an error rate under one in 100.

Vendor demos use clean documents. Yours aren't. Intelligent document processing reads incoming invoices, delivery notes, forms and contracts, pulls out the fields you need and sends them to your systems, with a person checking anything the system isn't sure of. It replaces manual keying: at 1,000 documents a month, our illustrative example releases about 49 hours a month.

So test it on your own post first.

Intelligent document processing at a glance

AI document processing is a pipeline of steps around a model, and a person stays on the ones that need judgement.

Intelligent document processing at a glance
AreaWhat it covers
Tasks it takes overCollecting documents from inboxes, portals, scanners and phone photos; sorting them by type; pulling out the fields you need, including tables and line items; checking them against rules and your records; filling the target system; filing the original with its extracted data
InputsPDFs (text or scanned), images and phone photos, Word and Excel files, email bodies. Typical SME sets include supplier invoices, delivery notes, purchase orders, application forms, ID documents, contracts, certificates and statements
Systems it connects toEmail and shared inboxes, SharePoint or Google Drive, your accounting package, CRM, ERP or practice system, and a spreadsheet where nothing else exists
Where a person stays in the loopAny field below its confidence threshold; any document type the system has not met before; a regular sample of straight-through documents; changes to the rules or thresholds
What triggers an exceptionA required field missing or below threshold; a value that fails a rule (a date out of range, totals that do not add up, an unknown supplier or customer); an unreadable page; a document it cannot classify; a duplicate

What IDP does, in four steps

PDF data extraction is one of four steps:

  1. Classify. What kind of document is this?
  2. Extract. Which fields does it hold, including tables and line items?
  3. Validate. Do the values make sense against your rules and records?
  4. Route. Into the target system, or to a person.

A text PDF reads as it is; scans and photos need optical character recognition (OCR) or a vision model. Tables are the hard part.

Paper hasn't gone. A 2025 AIIM survey of enterprises in the US and German-speaking Europe, sponsored by IDP vendors, found 61% of IDP processes still include paper.

Extraction underpins most systems in our Insights guides. For invoice VAT fields, posting and duplicates, see automated invoice processing.

LLM extraction vs template OCR: where each one wins

AI data extraction with a large language model (LLM) isn't always the right tool. Independent tests split on how messy the documents are.

Which approach for which documents
Document situationBetter approachEvidence
Scanned or photographed documents, varied layoutsA vision LLM reading the page imageBerghaus et al.: 92.71% on scanned invoices, against 64.03% for text-first
Many suppliers, each with its own layoutAn LLM, with no template per supplierSame study; template tools need one template per layout
Thousands of documents in one fixed layout (your own forms, one portal's export)Template or table-based OCR, which gives the same answer every timeWang and Shen: F1 of 1.0 in 0.97 seconds
Excel, CSV, Word or text PDFParse the file directly; no OCR neededWang and Shen's structured result came from Word and Excel files
Fields where a wrong value is costly (amounts, dates, bank details)Either, plus a rule check and human review below thresholdPatel et al.: 686 false dates against 304 correct ones
Arabic or bilingual documentsA vision LLM, tested on your own samplesKITAB-Bench: only 65% accuracy on Arabic PDF conversion

LLMs win on messy input. On scanned receipts, Berghaus and colleagues found a vision model got 87.46% of fields on scanned receipts, against no more than 47.00% when the page was converted to text first.

Templates win on volume. Scored on F1, where 1.0 is perfect, Wang and Shen's table-based pipeline reached an F1 of 1.0 in 0.97 seconds on structured documents, where direct LLM extraction needed 13.4 to 13.6 seconds each.

Hand-coded patterns are a poor fallback. On synthetic cheques, Patel and colleagues found an OCR plus regular-expression baseline scored an F1 score of 0.395. The same study showed how LLMs fail: by over-extracting. One model returned 686 false dates against 304 correct ones. A confidence score wouldn't catch that. A rule check would.

So most SME set-ups are hybrids: parse what's structured, template the high-volume fixed layouts, send the rest to an LLM. It's the case for mixing rules and agents.

How to test extraction accuracy on your own documents

No sales demo covers this. Run these eight steps, in order, on any tool you're weighing up.

  1. Build a test set from real documents. Use recent documents in the mix you actually receive: every type, the worst scans, phone photos, the odd supplier. Don't let the vendor choose.
  2. Size it with the rule of three. If a system makes no errors on n documents, you can say with 19-in-20 confidence that its true error rate is at most 3 in n (Hanley and Lippman-Hand, 1983). So a clean run on about 300 documents shows an error rate under one in 100.
  3. Label the right answers first. Before the system runs, a person keys the correct value for every field. That's your ground truth.
  4. Measure at field level, not document level. Score every field (supplier, date, total, VAT number, each table row) as right, wrong or missing. A strong document-level score can hide a field that's wrong one time in three.
  5. Set a confidence threshold per field. Above it, the field goes straight through; below it, a person checks. At each threshold, record the straight-through rate and the number of errors that slipped through above it. The second number is the one that matters.
  6. Add rule checks. Totals add up, dates fall in a sensible range, the supplier or customer exists, the VAT number has a valid format. Rules catch the confident mistakes.
  7. Run shadow mode on live documents. Give it two to four weeks alongside your team, releasing nothing. Our shadow mode note shows how.
  8. Agree go-live criteria before you start. For example: accuracy at or above your team's on the same documents, no errors above threshold on critical fields, and a straight-through rate that pays its way.

The UK government worked this way. i.AI tested Extract on over 400 existing digital records from over 30 planning authorities, compared its output directly with records made by hand and set accuracy thresholds in advance. Outputs needed only minor edits in around two-thirds (59 to 72%) of cases.

Test scorecard (fill in from your own test set)
FieldDocuments testedCorrectWrongMissingThresholdStraight throughErrors above threshold
Supplier or customer name
Document date
Total
VAT or TRN number
Line items (per row)

Handling mixed document types

One inbox can carry invoices, credit notes, statements, delivery notes and remittances. Test classification on its own: a misfiled document gets the wrong fields.

Give each document family its own model or prompt and threshold. Split multi-document PDFs, such as five invoices in one scan, before extraction, and test the split too. Statements are a document type of their own, covered in supplier statement reconciliation. Extracted delivery notes feed three-way matching.

Keying time and cost at 1,000 documents a month

Data entry automation is where the hours show. Take a business keying 1,000 mixed documents a month (illustrative, not a client result). The only measured public benchmark covers planning records. So every input except pay is our assumption, and pay is the median UK full-time hourly pay plus employer National Insurance and pension.

Inputs for 1,000 documents a month
InputValueSource
Documents a month1,000Example volume
Minutes to key one document by hand4 minutesOrwell assumption
Documents passing every check (straight through)Half of themOrwell assumption, to be measured in shadow mode
Minutes to check and correct a flagged document2 minutesOrwell assumption
Spot checks of straight-through documents1 in 20, 1.5 minutes eachOrwell assumption
Hourly costabout £22.75 an hourONS ASHE 2025 median of £19.67 an hour, plus employer National Insurance at 15% on earnings above £5,000 a year and pension at 3% of qualifying earnings between £6,240 and £50,270
Outputs: keying time before and after
OutputValue
Keying time todayabout 67 hours a month
Time still needing a personabout 17 hours a month
Hours releasedabout 49 hours a month, worth about £1,123 a month
Released if nothing goes straight through (every document checked)about 33 hours a month, in line with i.AI's measured saving of around a half (40 to 55%)
Released if only three in ten go straight throughabout 43 hours a month
Capacity with today's hoursabout 3,850 documents a month

One cross-check. i.AI measured review and correction at 5 to 20 minutes per planning record, against our 2 minutes per document, because a typical business document has far fewer fields. Shadow mode will give you your own figure.

Released hours aren't cash unless someone uses them. Data-entry staff are likely paid below the median too, so treat the value as an upper estimate.

Priced in dirhams, those hours are worth about AED 2,831 a month at our UAE rate of about AED 57 an hour. No official UAE pay data by occupation exists, so the rate rests on an assumed package of AED 8,000 a month and gratuity at 21 days' basic pay for each of the first five years. For Emirati staff the private-sector floor is AED 6,000 a month.

Document extraction benchmarks

Independent evidence on document extraction
MeasureFigurePublisher, dateNotes
Time to create a record by handbetween 15 minutes and 2 hoursi.AI (UK government), Extract evaluation, June 2026Planning documents, many scanned or handwritten
Time for AI to draft the recordtwo minutes on averagei.AISame evaluation
Human review and correction5 to 20 minutesi.AIEvery record reviewed
Total time savedaround a half (40 to 55%)i.AIPlanning records from councils
Outputs needing only minor editsaround two-thirds (59 to 72%) of casesi.AISame evaluation
Field accuracy, scanned receipts, image against text-first87.46% of fields on scanned receipts, against no more than 47.00% when the page was converted to text firstBerghaus et al., arXiv, August 2025Exact match, open datasets
Field accuracy, scanned invoices, image against text-first92.71% on scanned invoices, against 64.03%Berghaus et al.Same method
Table-based OCR on fixed layoutsF1 of 1.0 in 0.97 secondsWang and Shen, arXiv, October 2025Synthetic, highly repetitive documents
Direct LLM extraction time13.4 to 13.6 secondsWang and ShenSame test set
OCR plus regex baselinean F1 score of 0.395Patel et al., arXiv, September 2026Synthetic cheques
Arabic: vision models against traditional OCRBetter by an average of 60% in character error rateKITAB-Bench (MBZUAI), arXiv, February 2025Accepted at ACL 2025

No vendor figures here. Academic results use public or synthetic datasets, not yours, and arXiv papers are preprints.

Choosing IDP software: build, buy or configure

  • Buy a packaged tool when your documents are common types (invoices, receipts) and your target system has an integration.
  • Configure a general model with your own fields and rules for many document types at moderate volumes.
  • Build when documents are unusual, rules are yours alone, or the output lands in a bespoke system.

The usual build-or-buy question applies: bespoke or off the shelf. Whichever route you take, run every candidate through the eight-step test plan on the same test set.

How Orwell builds document processing

Our business automation work on documents runs in six stages, test plan first.

  1. Map. We map the documents and where each field goes.
  2. Test. With your team, we build the labelled test set and run the eight steps.
  3. Shadow. It runs on live documents, releasing nothing, until it meets the go-live criteria.
  4. Gate. Low-confidence or rule-breaking fields stop at an approval gate.
  5. Go live by type. Highest-volume document type first. We measure cost per document before and after.
  6. Monitor. A weekly sample of straight-through documents is checked by a person; drift alerts fire when a layout changes.

What document processing won't do

  • It won't be perfect, and doesn't have to be. What counts is catching and correcting errors.
  • It won't read what isn't legible. Bad scans and difficult handwriting go to a person.
  • It won't catch a confident wrong answer on its own. That takes rule checks.
  • It won't stay accurate unwatched. Layouts change; accuracy drifts.
  • It won't make decisions about people. Anything with a significant effect on a person keeps a human on it.
  • It won't handle Arabic PDFs reliably yet. In KITAB-Bench the best model managed only 65% accuracy on Arabic PDF-to-Markdown conversion.

Gartner's forecast is that over 40% of agentic AI projects will be cancelled by the end of 2027. Test before you commit.

UK and UAE rules for extracted data

IDP itself isn't regulated; the rules attach to the data and what it's used for.

UK

  • Accuracy principle. UK GDPR requires personal data to be accurate, but the ICO says an AI system needn't be 100% statistically accurate to comply, as long as errors are caught and corrected. Field-level testing and a human check below threshold are how you show it.
  • Automated decisions. The Data (Use and Access) Act 2025 rules on automated decision-making (UK GDPR Articles 22A to 22D) came into force on 5 February 2026. Extraction alone isn't a decision, but if its output feeds one with significant effects on a person (an application, an ID check, a claim), safeguards apply, including human review on request.
  • Making Tax Digital. If extracted data feeds VAT records, copy and paste is not a digital link. Extracted values must flow into your accounting software through an API or a file import.
  • Retention. Keep the original and the extracted record with your VAT records for at least 6 years.

UAE

  • Arabic documents. Records and documents given to the Federal Tax Authority must be in Arabic, although the FTA may accept another language and ask for a translation (Federal Decree-Law No. 28 of 2022 on Tax Procedures, Article 5). Bilingual extraction should keep the Arabic original alongside any English reading.
  • E-invoicing. Businesses with revenue of AED 50 million or more go first. The Ministry of Finance says PDFs, Word documents, images, scanned copies and emails are not e-invoices, so UAE e-invoicing replaces most domestic PDF invoices with structured data. Delivery notes, contracts, IDs and foreign invoices will still arrive as documents.
  • Data protection. The executive regulations for the federal Personal Data Protection Law (Federal Decree-Law No. 45 of 2021) had still not been issued in March 2026. DIFC and ADGM have their own laws.

Common questions

What is intelligent document processing?

Intelligent document processing (IDP) is software that reads documents such as invoices, forms, delivery notes and contracts, pulls out the fields you need and sends them to your systems. It classifies each document, extracts the data, checks it against rules and records, and routes anything uncertain to a person. It replaces manual keying, not the person who checks the exceptions.

How is IDP different from OCR?

OCR turns an image of text into text. IDP goes further: it works out what kind of document it is, which values matter, whether they make sense and where they should go. Newer IDP tools can use a vision language model instead of, or alongside, OCR; reading the image directly scored 92.71% on scanned invoices, against 64.03% for converting to text first.

How accurate is AI data extraction?

It depends on your documents, which is why you test on them. One independent study found a vision model read 87.46% of fields on scanned receipts, against no more than 47.00% when the page was converted to text first, and fixed layouts can reach near-perfect scores. The UK government's Extract tool needed only minor edits in around two-thirds (59 to 72%) of cases.

How many documents should we test before going live?

Enough to make the result mean something. By the rule of three, if a system makes no errors on n documents you can be 19-in-20 confident its error rate is at most 3 in n, so about 300 documents without an error shows a rate under one in 100. Use real documents in the mix you actually receive.

Is IDP worth it for a small business?

At low volumes, often not on time alone. In our worked example, 1,000 documents a month take about 67 hours a month to key, and automation releases about 49 hours a month. Halve the volume and you halve the gain, so start with the document type you receive most.

Can IDP read Arabic documents?

Yes, with care. Modern vision models beat traditional OCR on Arabic by an average of 60% in character error rate, but Arabic PDF conversion still reached only 65% accuracy for the best model tested. Test bilingual and Arabic documents as their own group.

Tell us which documents your team keys in today, where the data ends up and roughly how many arrive each month, and bring a few real examples to the call, especially the awkward ones. After an informal scoping chat we'll send a price range for automating the extraction. Document mixes, volumes and target systems differ too much for a published price to be honest. Select "Business automation" on the form.

Sources

Every figure in this article links back to the source below it was checked against.

  1. i.AI (UK government Incubator for AI): Extract evaluation Official source · checked 1 October 2026
  2. Berghaus et al.: Multi-modal vision vs text-based parsing for invoice processing (arXiv, August 2025) Research · checked 1 October 2026
  3. Wang and Shen: Hybrid OCR-LLM framework for copy-heavy document extraction (arXiv, October 2025) Research · checked 1 October 2026
  4. Orwell Lab calculation from Hanley and Lippman-Hand (1983) Our calculation from official rates · checked 1 October 2026
  5. AIIM and Deep Analysis: IDP Survey 2025 (vendor-sponsored, August 2025) Research · checked 1 October 2026
  6. Patel et al.: Robustness, cost and governance trade-offs for VLMs in templated document extraction (arXiv, September 2026) Research · checked 1 October 2026
  7. Heakl et al.: KITAB-Bench, Arabic OCR and document understanding benchmark (arXiv, February 2025) Research · checked 1 October 2026
  8. Orwell Lab calculation from ONS ASHE 2025 and HMRC 2026 to 2027 employer rates Our calculation from official rates · checked 1 October 2026
  9. GOV.UK (HMRC): Rates and thresholds for employers 2026 to 2027 Official source · checked 30 September 2026
  10. GOV.UK: Workplace pensions, what you, your employer and the government pay Official source · checked 30 September 2026
  11. Orwell Lab calculation from stated UAE pay assumptions and Federal Decree-Law No. 33 of 2021 (inputs: UAE Legislation, download) Our calculation from official rates · checked 1 October 2026
  12. MoHRE: minimum wage for Emiratis in the private sector raised to AED 6,000 a month (December 2025) Official source · checked 1 October 2026
  13. Gartner: over 40% of agentic AI projects will be cancelled by end of 2027 (June 2025) Research · checked 30 September 2026
  14. ICO: Guidance on AI and data protection, accuracy and statistical accuracy Official source · checked 1 October 2026
  15. legislation.gov.uk: Data (Use and Access) Act 2025 (Commencement No. 6) Regulations 2026 Official source · checked 1 October 2026
  16. GOV.UK (HMRC): VAT Notice 700/22, Making Tax Digital for VAT Official source · checked 1 October 2026
  17. GOV.UK (HMRC): Record keeping for VAT, VAT Notice 700/21 Official source · checked 1 October 2026
  18. UAE: Federal Decree-Law No. 28 of 2022 on Tax Procedures, Article 5 Official source · checked 1 October 2026
  19. UAE Ministry of Finance: UAE Electronic Invoicing Guidelines v1.1 (1 June 2026) Official source · checked 1 October 2026
  20. UAE Ministry of Finance: UAE eInvoicing initiative page Official source · checked 1 October 2026
  21. Chambers and Partners: Data Protection and Privacy 2026, UAE trends and developments (March 2026) Research · checked 1 October 2026
About this guide

Part of Orwell Lab’s system guides: what one AI system takes off a team, with the time, cost and capacity worked out from the pay figures, benchmarks and labelled assumptions shown above. Each figure was checked on the date shown beside its source.

Discuss your own process ↗

A useful next step.

Tell us what you’re building, what could work better and where you want to take the business.

Book a free discovery call