Custom Fields in Receipt OCR: Why Generic Extraction Fails Loyalty Campaigns

Generic OCR vendors can read a receipt. Whether they return the right fields, structured the right way, for a specific campaign is a different question entirely. Here is what that gap costs and how custom field training closes it.
- Generic OCR extracts what is on a receipt but does not guarantee the granularity, consistency or structure a specific campaign requires.
- Incorrect item-quantity extraction means customers are wrongly rewarded or miss out entirely, breaking campaign integrity at scale.
- Discount codes appear differently across retailers and POS systems; only custom-trained fields can harmonise them reliably across formats.
- Tabscanner's Pro service adds custom field training and cross-format mapping so the API returns exactly what your campaign logic needs.
Receipt OCR is not a binary capability. Every serious vendor will tell you they can extract data from a receipt, and technically most of them can. The gap that matters is not whether a field is found but whether it is found reliably, classified correctly, and returned in a structure that your campaign logic can act on without manual correction. For a loyalty program or a promotional campaign running across thousands of submissions a day, that gap has a direct cost.
What generic OCR actually delivers
A general-purpose receipt OCR model is trained to find and return the fields that appear most commonly across receipts: merchant name, date, total, tax, and a list of line items. For a broad expense-management use case where the goal is to digitise the document, that coverage is often sufficient. The output does not need to align with a campaign schema; it just needs to be readable and reasonably accurate.
Campaigns impose a different requirement. A loyalty program rewarding the purchase of a specific product needs to know not just that the product appeared on the receipt but that it appeared in the right quantity, at the right price, on the right date, and that the extraction can be trusted enough to trigger a reward without a human checking every claim. A promotional campaign built around discount codes needs that code extracted consistently whether the receipt comes from a large supermarket chain, an independent retailer, or a convenience format with a different point-of-sale system.
Generic models are not trained for that level of specificity. They were built for breadth, not depth. When a campaign requires a field that does not conform to the model's training distribution, or when a field that exists is formatted differently across the retailer estate the campaign covers, accuracy degrades and campaign logic starts making wrong decisions.
Example one: item quantity and campaign reward integrity
Consider a CPG brand running a promotion that rewards customers for buying three or more units of a specific product in a single transaction. The campaign mechanics are simple. The proof of purchase is a receipt photo. The extraction has to confirm the right product and the right quantity.
A generic OCR model will typically return a line-items array. Whether the quantity field for each item is reliably populated depends entirely on how consistently receipts in that category format their quantity column, and how well the model was trained on that specific product name and its common abbreviations across different retail formats. In practice, quantity is one of the fields most prone to misreads: it can be confused with pack size, price-per-unit, or simply missed when the receipt uses an unconventional layout. A model not trained on the specific product or the specific retailer estate will have an elevated miss rate on precisely the field that determines whether the reward fires.
The cost of getting this wrong compounds quickly. Customers who bought three units and are not rewarded contact support, lose trust in the program, and do not re-engage. Customers who bought one unit and are incorrectly rewarded represent a direct financial leak and a campaign integrity failure that is difficult to audit after the fact. At any meaningful campaign scale, a few percentage points of error on the quantity field translates to a real budget exposure and a measurable increase in support overhead.

Example two: discount codes across a mixed retailer estate
Discount codes present a different structural problem. A promotional campaign that distributes a unique code redeemable at multiple retailers needs to verify that the code appeared on the receipt. The challenge is that different retailers and different POS systems label, position and format that field in different ways. One chain prints it beneath the subtotal line labelled PROMO. Another prints it in the header block next to a barcode label. A third formats it as a separate line item in the body of the receipt. A fourth abbreviates it differently again.
A generic OCR model has no concept of campaign-level field harmonisation. It will return what it finds in a consistent schema based on its training, which means the discount code may be captured as part of the line items array, as a standalone field, as a footer element, or not at all depending on which retailer's format it happens to handle well. Mapping that inconsistency into a single structured field the campaign logic can validate against requires work, and without custom training it is work that falls on the engineering team processing the raw output, not on the OCR layer where it should be solved.
When that mapping is incomplete or inconsistent, valid discount codes are rejected and the customer experience breaks. Invalid or duplicate codes may pass validation if the extraction missed or misread the field. Neither outcome is acceptable for a campaign that is live and generating submissions at volume.
What custom fields and custom training actually mean
A custom field is a named output field in the API response that is defined for a specific campaign rather than derived from the generic model's default schema. Instead of receiving whatever the model decides to put in line_items, a campaign using a custom field receives a dedicated, typed field, for example item_quantity_product_x or promo_code, that is trained to find and extract that specific value reliably.
Custom training means the underlying model is adapted using examples drawn from the actual receipt formats the campaign will encounter. Rather than relying on a general model's handling of a field category, the model sees real receipts from the relevant retailers, labelled with the correct extraction targets, and learns the specific layout cues, label variants and position patterns that signal that field in that retailer's format. The result is a model that performs on the campaign's actual input distribution, not on a hypothetical average receipt.
At a practical level, this means that when a new campaign is set up with Tabscanner Pro, the integration is not just a matter of pointing an API key at an endpoint. The data team works through the receipt formats the campaign covers, defines the fields the campaign logic needs, and trains the model on representative samples. The output schema the campaign receives is designed around what the campaign needs to know, not around what a generic model happens to return.
Where generic OCR accuracy figures mislead
Receipt OCR vendors typically publish accuracy figures as a single headline number. That number is usually computed against a broad benchmark dataset covering many receipt types and a standard set of fields. It is a reasonable measure of general-purpose performance. It is not a measure of performance on the specific field a campaign depends on.
A model can achieve high accuracy overall while performing considerably worse on a specific minority field, particularly one that varies by retailer format or that appears infrequently in the training distribution. The quantity field for a niche product category, or a discount code printed in an unusual position, may fall well outside the distribution the headline figure was measured on. Campaign operators who select a vendor based on headline accuracy without testing their specific extraction requirements are effectively accepting unknown error rates on the fields that actually matter to their campaign.
The practical consequence is that errors compound. If a campaign processes a large volume of submissions and the extraction error rate on a critical field is even a few percent, the absolute number of incorrect reward decisions is significant. At a high volume of submissions, a modest error rate produces a substantial number of wrong outcomes, split roughly between false positives (fraud exposure) and false negatives (customer experience damage). Custom training is not a premium feature; it is the mechanism by which the accuracy figure becomes relevant to the campaign.
How Tabscanner Pro addresses this
Tabscanner is a receipt OCR API that turns photos and scans of receipts into structured, line-item data for expense management, loyalty programs and market research. The standard API handles the general case well: a wide range of global receipt formats, standard fields returned as typed structured data, confidence scores on each field so consuming applications can route uncertain results to review rather than acting on them automatically.
Tabscanner Pro extends that foundation with custom field definition and custom model training for specific campaigns. A campaign that needs a reliable quantity field for a named product gets a model trained on that product's appearance across the retailer formats in scope. A campaign that needs a harmonised discount code field across a mixed retailer estate gets a custom mapping layer trained on the actual label variants and positions those retailers use. The output schema reflects the campaign's data requirements, not a generic default.
The confidence score on each custom field is especially useful for campaign operations. Rather than processing every submission the same way, campaigns can apply a confidence threshold: high-confidence extractions flow directly to reward logic, lower-confidence results are held for a lightweight review step. That pattern keeps fully automated throughput high while protecting campaign integrity at the margins where the extraction is less certain.
Evaluating receipt OCR for a campaign: the right questions
When a loyalty platform, CPG brand or campaign manager is evaluating receipt OCR vendors, the headline accuracy figure is a starting point, not a conclusion. The questions that determine whether a vendor will perform on a real campaign are more specific: can the vendor define custom output fields aligned to the campaign schema, can the model be trained on the actual retailer formats the campaign covers, and what does the confidence score distribution look like on the specific fields the campaign logic depends on.
Testing with representative samples from the actual retailer estate before committing to a vendor is the most direct way to answer those questions. A vendor that performs well on a generic benchmark but has not been tested on the campaign's actual input distribution is an unknown risk in production.
Generic OCR vs Custom-Trained Fields
Generic OCR
line_items[ ]
returned, variable structure
quantity
unreliable across formats
promo_code
inconsistent, position-dependent
Custom-Trained Fields
line_items[ ]
structured to campaign schema
item_qty_[product]
trained on target product formats
promo_code
harmonised across all retailers
Bar width represents extraction reliability on campaign-critical fields across a mixed retailer estate.
The fields a generic model returns versus what a custom-trained model returns for the same campaign-critical data points.
A model can achieve high accuracy overall while performing considerably worse on the specific field a campaign depends on. The headline figure is a starting point, not a conclusion.
If you are building or evaluating a receipt-based loyalty program, promotional campaign or rewards platform, get in touch to discuss how custom field training can be applied to your specific retailer formats and campaign schema.
Get in touch to see how we can boost your accuracy
Tell us what you are building and we will show you how Tabscanner's receipt OCR handles your receipts and your accuracy targets.
Get in touch →
