Image input
PNG, JPEG, and the pages of a PDF. Photographed documents work — they do not need to be clean scans.
Flash takes image and video input directly. That removes an entire component from your pipeline — and the bill that came with it.
The conventional document pipeline has three parts: an OCR service that turns pixels into text, a parser that turns text into fields, and a model that interprets the fields. The first two exist because the model could not see.
Flash can see. Send it the image and it returns the fields — including the layout cues that a text-only pipeline throws away. Headers, stamps, table structure and handwriting all survive the trip, because nothing was flattened into text first.
PNG, JPEG, and the pages of a PDF. Photographed documents work — they do not need to be clean scans.
Video goes in as well as images. We tested this directly against the API rather than taking the model's word for it.
You define the fields and the shape. The model fills them, so downstream code gets a record rather than a wall of text.
Non-English documents read as well as English ones, and you can extract and translate in the same call.
When we tested image and video input against the API directly, the model told us in conversation that it only accepted text. It was wrong — the images and video went through and were read correctly. We mention it because it is a good reason to test rather than ask.
We are not going to quote you an accuracy figure we cannot stand behind. What we will do is run your own sample documents through it before you commit — that is the only number that matters, and it is specific to your documents.
Yes, including the row and column structure. Because the model sees the page rather than a text dump, table structure survives.
It handles it, with the caveat that any vision model's accuracy on handwriting depends heavily on the handwriting. Send us a sample and we will tell you honestly what we see.
Usually not. Photos taken on a phone, skew, shadows and moderate blur are within range. Extreme cases — a crumpled receipt photographed in the dark — may still need attention, and you will find that out in testing.
They are purpose-built document and vision services; this is a general model that can see. If you need detection primitives, they are a better fit. If you need the content understood and structured, a model is. We wrote up the Textract comparison and the Cloud Vision comparison.
SCANNED INPUT
read natively
Sample layout. No real invoice data.
STRUCTURED OUTPUT
image in, fields out
{
"vendor": "…",
"invoice_number": "…",
"issue_date": "…",
"line_items": [
{ "description": "…", "qty": … },
{ … }
],
"tax": …,
"total": …
}
Field names shown; values omitted.
Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.