Blog

Notes

Three things we end up explaining on most calls. Written for the engineer who has to make the decision, not for a search engine.

What's in here

Three notes. Read the one you need — they do not depend on each other.

Migration

Tokenization, streaming edge cases, and the system prompts that quietly stop being portable.

  • 2 October 2026
  • Migration
What actually changes

Long context

What a very large window removes from your roadmap, and what it does not remove.

  • 2 October 2026
  • Engineering
What a 1M window changes

Procurement

The eight questions we would ask if we were the buyer — with our own answers linked.

  • 2 October 2026
  • Procurement
What to ask before you sign

What actually changes when you move to the DeepSeek API

  • 2 October 2026
  • Migration

The short version: less than you expect, and not in the places you expect. The request shape is compatible with the OpenAI SDK, so for most teams the first change is a base URL and a key, and the first request succeeds. That success is what makes the next part annoying — it looks finished before it is.

Tokenization differs, so your limits move. Token counts are not portable between model families. A prompt that sits comfortably inside a context window on your current provider can land somewhere else entirely here, and the same is true of a completion budget. Before you cut over, count tokens for a representative sample of your real traffic with the tokenizer that actually applies, not with a character estimate.

Streaming edge cases are where integrations break. Chunk boundaries, how finish reasons arrive, and what a stream looks like when it ends early are all places where two implementations that both call themselves compatible behave differently. If your client has retry logic, test it against a deliberately interrupted stream rather than assuming the happy path generalizes.

System prompts are not a portable asset. A prompt tuned against one model carries the habits of that model. Instructions that were doing real work may now be redundant, and instructions that were redundant may now be load-bearing. Re-run your evaluation set rather than your intuition.

The practical method: run both paths side by side for a week on real traffic, log the outputs, and compare them on the cases your users complain about. That is a duller plan than a cutover weekend, and it is the one that does not produce a rollback.

Long context in practice: what a 1M token window changes

  • 2 October 2026
  • Engineering

A very large context window removes a whole category of engineering work. It does not remove the need to think about what you put in it. Teams that treat "it fits" as "it is understood" tend to discover the difference in production.

Retrieval quality is not uniform across the window. Information placed deep inside a very long input is not attended to as reliably as information near the edges. If your use case depends on a specific clause in a long document, test that clause at different positions rather than assuming a single successful run generalizes.

Latency grows with input, and it is not linear in your users' patience. A request that carries an entire document takes meaningfully longer than one carrying a page. If your interface is interactive, measure the long-input path end to end before you promise a response time — the model is only one component of it.

Structure beats volume. A long input with clear delimiters, a stated task and an explicit output format outperforms the same content pasted as a wall of text. If you are assembling the input programmatically, spend the effort on the scaffolding rather than on squeezing in one more document.

The pattern that works: give the model the task at both ends of a long input, keep the middle for evidence, and validate on your own hardest cases rather than on a demo document.

What to ask a provider before you sign

  • 2 October 2026
  • Procurement

These are the questions we get asked by good procurement teams, and the ones we would ask if we were buying. They apply to us as much as to anyone else, so the answers we would give are linked where we have written them down.

  • Who holds the upstream account? If the answer is "we do", ask what happens to your service if that account is suspended. The answer should be a process, not a reassurance.
  • What happens when a model is deprecated? Model lifecycles end. Ask for the notice period you will receive and who is responsible for telling you.
  • Who do we call, and when? A named person with a defined response time is worth more than a support portal with a ticket queue.
  • What is in writing? Ask which of the provider's promises appear in the contract, and which exist only on the marketing site. The gap between those two lists is the real product.
  • What is the exit path? How do you leave, what happens to your data, and how long does it take. Ask before you need it.
  • What do you do with our data? "We do not sell it" and "we do not train on it" are different statements. Get the one you actually need.
  • Which certifications do you hold? Ask for the list, and treat a vague answer as a no.
  • What does the invoice look like? Entity name, currency, payment terms, and whether tax is included. Finance will ask eventually, and it is easier to answer now than during onboarding.

A provider who answers all eight directly is telling you something. So is one who answers three and changes the subject.

Written by the team behind this site. These notes describe our own experience and general engineering practice; they are not legal or procurement advice, and specific figures about the model service are quoted from the provider's public documentation.

BLOG · WHAT WE WRITE ABOUT

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

WHAT WE PUBLISH

Migration, OCR, volume, billing

the parts that decide a rollout

HOW WE WRITE

From work we actually did

numbers we can point at

WHAT WE DO NOT DO

No listicles, no filler

nothing written to fill a calendar

  • Three articles to start with
  • Every claim has a source
  • Written for a technical reader
Three notes: migration, long context, and what to ask before signing.Diagram of the process. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.