How to Run an Accurate Proof of Concept for Receipt OCR

Most OCR vendors quote a headline accuracy number. Here is the methodology to stress-test that number properly, against your own receipts, on your own terms. It works for evaluating any vendor, including us.
- Headline accuracy figures from vendors are almost always measured on their own curated test sets. A proper POC measures accuracy on your receipts.
- The gold standard approach uses a blind held-back test set that the vendor never trains on, giving you an unbiased accuracy reading.
- Ground truth is harder to establish than most teams expect. A trustworthy vendor will check your ground truth data and flag discrepancies before reporting results.
- If you lack enough data for a blind test, a single-batch pragmatic POC is still better than taking a vendor's quoted figure at face value.
Every receipt OCR vendor has a headline accuracy number, and almost every one of those numbers is real in some narrow sense and misleading in every practical sense. It was measured on the vendor's own test set, on receipt formats the model was built for, under controlled conditions that have nothing to do with your data. A proper proof of concept replaces that number with a measurement that actually matters: accuracy on your receipts, on your formats, with your edge cases. This post lays out the methodology we use at Tabscanner and that any serious team should apply when evaluating any receipt OCR vendor, including us.
Why headline accuracy figures fail in practice
An accuracy figure only means something relative to the test set it was measured on. A model that scores 98% on a benchmark of US grocery receipts may perform significantly worse on European fuel receipts, hotel folios, or any format it has not been tuned for. The gap between benchmark accuracy and production accuracy is where most OCR evaluations go wrong.
Receipts are genuinely diverse. A single retailer may run different point-of-sale systems in different countries, producing layouts that look nothing alike. The same chain in France and Portugal may generate receipts with different field positions, different date formats, different tax line structures. A model that handles one will not automatically handle the other.
The only honest measure is performance on a sample drawn from the actual population you will process in production. Everything else is marketing.
Before you start: what to count as a format
Before collecting samples, align on what a format means. For POC purposes, a format is a distinct receipt layout produced by a specific retailer's POS system in a specific region. Two receipts from the same brand can be different formats if one comes from a French store and the other from a Portuguese store, or if the retailer changed its POS vendor between those two territories.
This distinction matters because training and configuration happen at the format level. If you lump different layouts together in your sample batch, the results become hard to interpret and the model harder to tune accurately.
A practical rule: if two receipts from the same brand look visually distinct in their field positions, date formatting, or tax presentation, treat them as separate formats in your sample collection.
Path 1: the gold standard POC
Use this path whenever you have enough samples to set aside a held-back test set. It is more work upfront, but it produces an accuracy figure you can actually trust.
The ground truth problem is more common than you expect
Establishing accurate ground truth is genuinely difficult. Manual transcription introduces errors. Teams use different conventions for handling ambiguous fields, a truncated product name, a partially visible price. If two people transcribed different halves of the dataset without agreeing on a standard, you will have inconsistencies built into your benchmark before the vendor processes a single receipt.
This is not a criticism of the teams doing the work. It is a structural problem with manual data labelling at scale. The relevant point for a vendor evaluation is this: a vendor who reviews your ground truth and flags potential errors before reporting final results is demonstrating something important about how they think about accuracy. They are not just optimising their own score. They are trying to give you a true reading.
When Tabscanner runs a POC and the results come back, we ask to see the client's ground truth. We review it, flag discrepancies and present a reconciled accuracy figure alongside the raw one. That reconciled figure is the one that reflects real model performance.
Path 2: the pragmatic POC
Not every team has the volume or the time to run a full blind test. If your receipt corpus is small, or if you are doing an initial feasibility check before committing to a full evaluation, a single-batch POC is still a meaningful step forward.
What to measure and how to report it
Field-level accuracy is more useful than document-level accuracy. A document-level score tells you what percentage of receipts had every field correct. A field-level score tells you how each individual field performs, which is the number you actually need if you are building, for example, an expense tool where the total and date are critical but a missing middle line item is tolerable.
Measure accuracy separately for each field type (merchant, date, total, tax, line item description, line item price) and separately for each receipt format. An aggregate accuracy figure across all fields and all formats is too coarse to act on. A vendor who reports only the aggregate is hiding the distribution.

The real cost of a one percent error rate versus ten percent
Error rates compound. A one percent field error rate on a total field means one in every hundred expense claims carries an incorrect amount. At ten thousand claims a month, that is a hundred errors entering your finance system every month, each one requiring investigation, correction and potential reprocessing.
At industry-typical error rates of ten to fifteen percent for general-purpose OCR applied to receipts without specialist training, the same volume produces between one thousand and fifteen hundred errors per month. The operational cost to catch and correct those errors, staff time, delayed reimbursements, audit exposure, is rarely included in the headline price comparison between OCR vendors.
A POC that measures real field-level error rates on your actual receipts lets you project these costs accurately before you commit to a production integration.
What a trustworthy vendor does differently during a POC
Beyond the mechanics of the test, the way a vendor behaves during a POC tells you a great deal about how they will behave in production. A few things to observe.
Do they tell you which formats are under-sampled and ask for more data before they start? A vendor who trains on three receipts per format and does not flag that as thin is prioritising a fast turnaround over a useful result. Do they report accuracy by field and by format, or only as a single aggregate? The aggregate hides problems. Do they ask to see your ground truth and cross-check it, or do they report results against whatever you gave them without question? And when accuracy on a particular format is lower than expected, do they explain why and propose a path to improvement, or do they just move on?
None of these are difficult tests. But vendors who do all of them consistently are treating accuracy as an engineering problem rather than a sales exercise.
Scale the POC to the data you actually have
A rigorous methodology does not require thousands of receipts. The key is proportionality. If you have ten receipt formats and thirty samples per format, you can run a meaningful blind test with a twenty percent hold-out. If you have two formats and fifteen samples total, you cannot run a clean blind test, and Path 2 is the honest choice for now.
Be clear with your vendor about what you have. A vendor who pushes you toward Path 1 even when your sample count does not support it is setting up a test that will produce noisy, hard-to-interpret results. Scale the rigour to your data, not to the vendor's preference.
As your production volume grows, revisit the evaluation. A POC run at launch with fifty receipts tells you something useful. A re-evaluation run six months later with five hundred receipts drawn from production tells you something much more precise.
Path 1 uses a blind held-back test set for an unbiased accuracy reading. Path 2 is a reasonable fallback when sample volume is limited.
A vendor who checks your ground truth and flags errors before reporting results is not trying to inflate their score. They are trying to give you a true reading. That is the difference between a sales exercise and an engineering one.
Tabscanner processes receipts from across the globe, from grocery and fuel to hospitality and retail, and we apply this methodology to every vendor evaluation we run. Get in touch to talk through your receipt mix, your sample volume and which path makes sense for your evaluation.
Get in touch to see how we can boost your accuracy
Tell us what you are building and we will show you how Tabscanner's receipt OCR handles your receipts and your accuracy targets.
Get in touch →
