Invoice & receipt OCR

Invoice and receipt OCR that ends in a record, not a text dump.

Accounts payable intake is high volume and low reasoning — exactly the shape where a Flash-class model beats a flagship on cost without giving anything up.

What comes out

Not a wall of extracted text. A record: vendor, invoice number, issue date, line items with quantities and amounts, tax, and total — in the shape your ledger workflow already expects. You define the fields; the model fills them.

Why invoice OCR usually has three parts

Conventional invoice OCR runs in three parts: it reads the text, a parser pulls out the fields, and a model handles the parts the parser could not. Three components, three bills, three things to monitor.

The first two exist because the model could not see the page. Flash can. Send the image and it returns the fields directly — and it keeps the layout information a text pipeline discards, which is exactly the information that tells you which number is the total and which is a subtotal.

Where the volume is

  • Supplier invoices, digital or scanned
  • Employee expense receipts, including phone photos
  • Credit notes and statements
  • Non-English supplier documents, extracted and translated together

What we will not claim

We are not going to put an accuracy percentage on this page. Accuracy on invoice extraction depends on your suppliers, your scan quality and your schema — three things we have not seen yet.

What we will do instead is run a sample of your own invoices through it before you commit, and show you what came back. That number is real and it is yours.

  • Sample test first
  • Your documents
  • No invented figures
Get a quote

What it means for the team

Finance

Fewer documents stuck in a manual queue, and a cost per document that behaves predictably as volume grows.

  • Predictable per-document cost

Engineering

One component instead of three. No OCR service to run, no parser to maintain when a supplier changes their template.

  • One call
  • No parser to maintain

Legal

A negotiable agreement and a data processing agreement, so invoice data does not become the reason the project stalls.

  • DPA available
  • Negotiable terms

Questions about invoice extraction

Can it handle scanned invoices, not just digital PDFs?

Yes — scanned and photographed documents are the main use case. Because the model reads the page rather than a text layer, a scan is not a second-class input.

What about multi-page invoices with attachments?

They fit. The context window is 1M tokens, so a long invoice with supporting pages can go in as one job rather than being split.

Is this invoice OCR software, or an API?

An API, not a boxed product. There is no installer, no desktop app and no per-seat license — you send the document to an endpoint and the record comes back. If what your team actually needs is invoice OCR software a clerk can operate by hand, that is a different category of tool and we would rather say so than sell around it.

What does invoice OCR processing cost at volume?

It is a per-token rate rather than a per-page license, so cost tracks how much text each document actually contains — a two-line credit note and a 40-page supplier statement are not the same unit. Send us a sample and a monthly volume and we will work out what that means for you before you commit.

I only need to convert an invoice to OCR text. Is that enough?

If plain text is genuinely all your downstream process needs, a cheaper OCR-only tool may well be enough and we will say so. Most teams arrive looking to put an invoice to OCR and then find that the text is the easy half: the hard half is getting the fields into a shape your ledger can ingest, which is what happens here in the same call.

We scan receipts — does OCR for receipts work the same way?

Yes. OCR receipt scanning is the same call: the image goes in, the fields come out. Receipts are harder than invoices in one respect — they are photographed, folded, sometimes faded — which is exactly the case a model that reads the page handles better than a positional template.

Can we define our own output schema?

Yes. You specify the fields and the structure, and the model fills them. That is the difference between a record your ledger can ingest and a text blob someone has to re-key.

How do you handle suppliers who change their invoice template?

A model reading the page is far more tolerant of template changes than a positional parser — it is reading for meaning rather than matching coordinates. That is one of the practical reasons teams move to this approach.

Does it work for receipts as well as invoices?

Yes. Phone photos of receipts work, which is why expense capture is a common second workload once invoice intake is running.

How does this compare to Amazon Textract or Google Cloud Vision?

They are purpose-built document services; this is a general model that can see and reason. If you need fixed document primitives, they fit well. If you need the content understood and structured to your schema, a model does. Textract comparison · Cloud Vision comparison

INVOICE PARSING

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Input

Scanned or photographed invoice

PDF · JPEG · PNG

DeepSeek V4.1 Flash

Output

Vendor, line items, tax, totals

one JSON record

  • Reads scans and photos directly
  • Whole document in one call
  • No separate OCR stack
Scanned or digital, one call in, one record out.Diagram of the workload shape. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.