Long documents in one call
1M tokens of context and up to 384K of output. Multi-page documents do not need a chunking pipeline in front of them.
You do not need another client library. The endpoint follows the OpenAI chat-completions schema, so the code you have works with a different base URL.
An image goes in as content, the fields come back as the model's response. No upload endpoint, no job polling, and no separate OCR extraction service to authenticate against.
from openai import OpenAIimport os, base64, json client = OpenAI( base_url="https://api.workhorseapi.com/v1", api_key=os.environ["WORKHORSE_API_KEY"],) img = base64.b64encode(open("invoice.png", "rb").read()).decode() r = client.chat.completions.create( model="deepseek-flash", messages=[{"role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img}"}}, {"type": "text", "text": "Return the invoice fields as JSON."}, ]}],) record = json.loads(r.choices[0].message.content)
That is illustrative, not a runnable sample — your schema, error handling and retry policy are yours to set. Ask us for a key and we will help you get the first call working.
1M tokens of context and up to 384K of output. Multi-page documents do not need a chunking pipeline in front of them.
Image input is native, and video input works too — we tested both against the API directly rather than trusting the documentation.
High enough that your throughput is more likely to be limited by your own infrastructure than by the API.
Cached input is billed far below fresh input, and off-peak calls are billed at half the peak rate. Both matter at volume.
No. Use the openai package or anything compatible with it,
and change the base URL and key. That is the whole migration for most codebases.
Yes. Ask for a limited trial key in the form and run it against your own documents. We would rather you tested on real data than on our sample.
Common raster formats including PNG and JPEG, sent inline as base64 or by URL. PDF pages go in as images. Tell us your specific format if it is unusual.
Batch work pairs well with off-peak billing, where the rate is half the peak rate. Whether that fits your latency budget depends on your workload — tell us and we will work it out with you.
Design for it, as you would with any extraction pipeline: validate the record against your own rules, and route anything that fails to a human queue. We will help you work out where the confidence thresholds should sit.
OCR API
Input
Image bytes over the API
PNG · JPEG · PDF pages
DeepSeek V4.1 Flash
Output
JSON back in your schema
same client you already use
Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.