OCR API for developers

An OCR API that speaks the SDK you already import.

You do not need another client library. The endpoint follows the OpenAI chat-completions schema, so the code you have works with a different base URL.

A Python OCR API call, start to finish

An image goes in as content, the fields come back as the model's response. No upload endpoint, no job polling, and no separate OCR extraction service to authenticate against.

That is illustrative, not a runnable sample — your schema, error handling and retry policy are yours to set. Ask us for a key and we will help you get the first call working.

What you get

Long documents in one call

1M tokens of context and up to 384K of output. Multi-page documents do not need a chunking pipeline in front of them.

  • 1M context
  • 384K output

Images and video

Image input is native, and video input works too — we tested both against the API directly rather than trusting the documentation.

  • Verified by test
  • 2026-09-26

Concurrency up to 2,500

High enough that your throughput is more likely to be limited by your own infrastructure than by the API.

  • 2,500 concurrent

Caching and off-peak

Cached input is billed far below fresh input, and off-peak calls are billed at half the peak rate. Both matter at volume.

  • Caching
  • Half rate off-peak

Questions from developers

Do I need a new SDK?

No. Use the openai package or anything compatible with it, and change the base URL and key. That is the whole migration for most codebases.

Can I get a key to test with before we buy?

Yes. Ask for a limited trial key in the form and run it against your own documents. We would rather you tested on real data than on our sample.

What image formats are supported?

Common raster formats including PNG and JPEG, sent inline as base64 or by URL. PDF pages go in as images. Tell us your specific format if it is unusual.

Is there a batch endpoint?

Batch work pairs well with off-peak billing, where the rate is half the peak rate. Whether that fits your latency budget depends on your workload — tell us and we will work it out with you.

What happens when the model gets a document wrong?

Design for it, as you would with any extraction pipeline: validate the record against your own rules, and route anything that fails to a human queue. We will help you work out where the confidence thresholds should sit.

OCR API

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Input

Image bytes over the API

PNG · JPEG · PDF pages

DeepSeek V4.1 Flash

Output

JSON back in your schema

same client you already use

  • Native image input, no OCR step
  • OpenAI-compatible request shape
  • Structured output you define
Image bytes in, your schema out, through the client you already have.Diagram of the workload shape. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.