DeepSeek V4

Two V4 tiers. One endpoint. Change the model name.

Most workloads do not need a reasoning model — they need a cheap one that is good enough. Knowing which tier you actually need is most of the saving.

The two tiers

From DeepSeek's published model documentation, verified 2026-09-27
V4.1 FlashV4.1 Pro
Model name to senddeepseek-flash deepseek-v4-pro
Context window1M tokens1M tokens
Maximum output384K tokens384K tokens
Image inputYesNo
Video inputYesNo
ConcurrencyUp to 2,500Up to 500
Built forVolume work: extraction, classification, summarization Harder reasoning where accuracy matters more than cost

Notice what is not different: context and output limits are identical. The tiers differ in reasoning depth and in whether the model can see. That makes the choice simpler than it looks.

Choosing between them

Flash is right when…

The job is high volume and low reasoning: pulling fields out of documents, sorting items into categories, summarizing, translating, reading images.

  • High volume
  • Low reasoning
  • Needs vision

Pro is right when…

The answer has to be right and the reasoning is genuinely multi-step — analysis, complex code work, anything where a wrong answer costs more than the tokens.

  • Multi-step reasoning
  • Accuracy-critical
  • No vision

The mistake worth avoiding

The most common pattern we see is a workload that started on the strongest available model and never got revisited. It worked, nobody had a reason to touch it, and the workload grew tenfold underneath it.

Run a sample of your real traffic through Flash and compare the output against what you get today. If Flash is good enough — and for extraction and classification it usually is — you have found your saving without touching the architecture.

Questions about the models

Are DeepSeek V4 Pro and DeepSeek V4 Flash the same as V4.1?

V4.1 is the current point release of the V4 line, so if you are searching for DeepSeek V4 Pro or DeepSeek V4 Flash, these are the two tiers that answer to those names: V4.1 Pro and V4.1 Flash. When a newer point release ships we will say so on this page rather than quietly renaming things.

Can we use both tiers?

Yes, and most teams should. Route the volume work to Flash and escalate the hard cases to Pro. It is the same endpoint and the same client — only the model name changes.

Why does Pro not support images?

That is how DeepSeek built the tiers. Practically it means a vision-dependent workload runs on Flash, which is also the cheaper tier — so the constraint rarely bites.

Is 1M context real, or marketing?

It is the published figure for both tiers, and it is why you can send a whole contract or claims file in one call instead of building a chunking pipeline. Test it on your longest document before you rely on it.

Do you serve any other models?

No. We serve the DeepSeek V4 family. If you need a broad catalog behind one key, a router like OpenRouter is the better fit — we wrote up that comparison too.

TWO TIERS, ONE ENDPOINT

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Input

Your workload, by shape

high volume or hard reasoning

DeepSeek V4.1 Flash / V4.1 Pro

Output

Flash or Pro, same endpoint

swap the model name only

  • Flash: image and video input
  • Pro: reasoning, no vision
  • Both: 1M context, 384K output
Same endpoint either way. The tier is a parameter, not a migration.Diagram of the workload shape. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.