Pricing

35% below DeepSeek's published list price.

That is the anchor for every quote. Your own figure depends on volume, cache hit ratio and how much of the load can move off-peak — tell us those and we will send it, usually within one business day.

Where the number comes from

Our rate is 35% below DeepSeek's published list price. Their price list is public — open it and check the arithmetic yourself. We would rather you did. What follows is DeepSeek API pricing as it applies to your workload: the published list is where DeepSeek pricing starts, not where your own number lands.

What we cannot put on a page is your figure. Off-peak scheduling bills batch work at half the peak rate, and prompt caching cuts the cost of repeated input sharply. How much of your workload can use either one is what moves the number — and that differs by customer. So we ask two questions: how many tokens a month, and what are you running today.

What we can tell you up front

These are properties of the model's published pricing, not of our quote — they apply whatever rate we agree.

  • Off-peak calls are billed at half the peak rate
  • Cached input tokens are billed far below fresh ones
  • You pay per token — no seat licenses, no platform fee
  • One monthly invoice, payable by bank transfer
  • Prices exclude applicable taxes

What changes your rate

  • Monthly volume
  • Workload mix
  • Cache hit ratio
  • Peak / off-peak split
  • Contract length

The two biggest levers are almost always prompt caching and moving batch work off-peak. We will look at both with you before quoting.

Get a quote

How the quote works

STEP 01

You send two numbers

Monthly token volume, and what you pay today. If you do not know the token count yet, send the workload description and we will estimate it.

STEP 02

We send a rate

Usually within one business day, from a named person. If we cannot beat your current cost, we say so instead of quoting anyway.

STEP 03

You decide, then we contract

Nothing is charged before a signed agreement. Terms are negotiable and a data processing agreement is available.

ONE LINE CHANGES

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Your application

The OpenAI SDK you already ship.

api.workhorseapi.com

Same request shape. Same streaming.

WORKHORSE ENDPOINT

DeepSeek V4.1 Flash

1M context · 384K output

response returns the same way

WHAT WE HANDLE AROUND IT

  • Contract you can negotiate
  • One monthly invoice
  • A named account contact
  • Data processing agreement
Billing sits on our side of the line. You see one endpoint and one invoice.Diagram of the request path. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.