DeepSeek OCR

OCR without a separate OCR service.

Flash takes image and video input directly. That removes an entire component from your pipeline — and the bill that came with it.

What this replaces

The conventional document pipeline has three parts: an OCR service that turns pixels into text, a parser that turns text into fields, and a model that interprets the fields. The first two exist because the model could not see.

Flash can see. Send it the image and it returns the fields — including the layout cues that a text-only pipeline throws away. Headers, stamps, table structure and handwriting all survive the trip, because nothing was flattened into text first.

Image input

PNG, JPEG, and the pages of a PDF. Photographed documents work — they do not need to be clean scans.

  • Native vision
  • Photos and scans

Video input

Video goes in as well as images. We tested this directly against the API rather than taking the model's word for it.

  • Verified by test
  • 2026-09-26

Output in your schema

You define the fields and the shape. The model fills them, so downstream code gets a record rather than a wall of text.

  • Your schema
  • Structured output

Multilingual sources

Non-English documents read as well as English ones, and you can extract and translate in the same call.

  • Any language
  • Extract + translate

A detail worth knowing

When we tested image and video input against the API directly, the model told us in conversation that it only accepted text. It was wrong — the images and video went through and were read correctly. We mention it because it is a good reason to test rather than ask.

Questions about OCR

How accurate is it compared to a dedicated OCR service?

We are not going to quote you an accuracy figure we cannot stand behind. What we will do is run your own sample documents through it before you commit — that is the only number that matters, and it is specific to your documents.

Does it handle tables?

Yes, including the row and column structure. Because the model sees the page rather than a text dump, table structure survives.

What about handwriting?

It handles it, with the caveat that any vision model's accuracy on handwriting depends heavily on the handwriting. Send us a sample and we will tell you honestly what we see.

Do we still need a pre-processing step?

Usually not. Photos taken on a phone, skew, shadows and moderate blur are within range. Extreme cases — a crumpled receipt photographed in the dark — may still need attention, and you will find that out in testing.

How does this compare to Amazon Textract or Google Cloud Vision?

They are purpose-built document and vision services; this is a general model that can see. If you need detection primitives, they are a better fit. If you need the content understood and structured, a model is. We wrote up the Textract comparison and the Cloud Vision comparison.

Diagram of the flow: two stages side by side.

SCANNED INPUT

read natively

Sample layout. No real invoice data.

STRUCTURED OUTPUT

image in, fields out

{
  "vendor": "…",
  "invoice_number": "…",
  "issue_date": "…",
  "line_items": [
    { "description": "…", "qty": … },
    { … }
  ],
  "tax": …,
  "total": …
}

Field names shown; values omitted.

Scanned input on the left, structured record on the right. Field names are shown; values depend on your schema.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.