Request shape
Chat-completions with tool calling and streaming. If your harness speaks that, the change is a base URL.
If you are wiring up an agent harness, the endpoint is a configuration value. For most harnesses that is one line to change, and we handle everything commercial around it.
A harness is the scaffolding around a model: the loop that calls tools, keeps state, retries, and decides when it is done. It is real engineering, and it is the part you should not have to rewrite. The model endpoint underneath it is a configuration value.
Harnesses built on the OpenAI-compatible request shape take a base URL and a key. Point both at us and the harness keeps working — same tools, same prompts, same loop. If your harness is built around a different provider's native SDK, tell us and we will be honest about whether this is a five-minute change or a real project.
Chat-completions with tool calling and streaming. If your harness speaks that, the change is a base URL.
1M tokens of context and up to 384K of output. Agent loops accumulate context fast — this is usually more headroom than you need.
Agent loops re-send a growing prefix on every step. Prompt caching is what keeps that affordable, and it is usually the first thing we look at with you.
Ask us for a limited key and run your real harness against it. Testing on your own workload beats any benchmark we could show you.
Yes — tool calling and streaming are part of the chat-completions shape, and the official SDKs work unchanged. If your harness has a provider abstraction, add us as another provider.
There are open-source harnesses on GitHub; we are not one of them. We are the endpoint a harness points at, not the harness itself — so if you found yours in a repository, the change is in its config rather than its code: base URL and API key, then run it against your own workload.
Concurrency on Flash goes up to 2,500, and context is 1M tokens. Very long autonomous runs are usually limited by your own timeout budget rather than the API. Tell us your peak concurrency and we will confirm what we can commit to in writing.
Yes, and it is a sensible way to start. Keep your direct account configured, send a share of traffic to us, and compare on your own numbers.
The harness is yours and the model is DeepSeek's — we provide the endpoint and the commercial relationship. If something breaks, you have a named contact rather than a ticket queue, and we will help you work out which layer is at fault.
We serve DeepSeek V4.1 Flash and V4.1 Pro. If your harness needs a different family entirely, we are not the right provider and we will say so.
WE HANDLEYOU DO
1
Tell us your volume
Monthly token usage, and what you are running today. If we cannot beat your current cost, we say so.
2
We quote and contract
A named account contact, a negotiable agreement, an invoice.
- base_url="…"
+ base_url="…"
# nothing else changes3
You change one line
Point your existing client at our endpoint. Keep a fallback.
Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.