Migration
Tokenization, streaming edge cases, and the system prompts that quietly stop being portable.
What actually changesThree things we end up explaining on most calls. Written for the engineer who has to make the decision, not for a search engine.
Three notes. Read the one you need — they do not depend on each other.
Tokenization, streaming edge cases, and the system prompts that quietly stop being portable.
What actually changesWhat a very large window removes from your roadmap, and what it does not remove.
What a 1M window changesThe eight questions we would ask if we were the buyer — with our own answers linked.
What to ask before you signThe short version: less than you expect, and not in the places you expect. The request shape is compatible with the OpenAI SDK, so for most teams the first change is a base URL and a key, and the first request succeeds. That success is what makes the next part annoying — it looks finished before it is.
Tokenization differs, so your limits move. Token counts are not portable between model families. A prompt that sits comfortably inside a context window on your current provider can land somewhere else entirely here, and the same is true of a completion budget. Before you cut over, count tokens for a representative sample of your real traffic with the tokenizer that actually applies, not with a character estimate.
Streaming edge cases are where integrations break. Chunk boundaries, how finish reasons arrive, and what a stream looks like when it ends early are all places where two implementations that both call themselves compatible behave differently. If your client has retry logic, test it against a deliberately interrupted stream rather than assuming the happy path generalizes.
System prompts are not a portable asset. A prompt tuned against one model carries the habits of that model. Instructions that were doing real work may now be redundant, and instructions that were redundant may now be load-bearing. Re-run your evaluation set rather than your intuition.
The practical method: run both paths side by side for a week on real traffic, log the outputs, and compare them on the cases your users complain about. That is a duller plan than a cutover weekend, and it is the one that does not produce a rollback.
A very large context window removes a whole category of engineering work. It does not remove the need to think about what you put in it. Teams that treat "it fits" as "it is understood" tend to discover the difference in production.
Retrieval quality is not uniform across the window. Information placed deep inside a very long input is not attended to as reliably as information near the edges. If your use case depends on a specific clause in a long document, test that clause at different positions rather than assuming a single successful run generalizes.
Latency grows with input, and it is not linear in your users' patience. A request that carries an entire document takes meaningfully longer than one carrying a page. If your interface is interactive, measure the long-input path end to end before you promise a response time — the model is only one component of it.
Structure beats volume. A long input with clear delimiters, a stated task and an explicit output format outperforms the same content pasted as a wall of text. If you are assembling the input programmatically, spend the effort on the scaffolding rather than on squeezing in one more document.
The pattern that works: give the model the task at both ends of a long input, keep the middle for evidence, and validate on your own hardest cases rather than on a demo document.
These are the questions we get asked by good procurement teams, and the ones we would ask if we were buying. They apply to us as much as to anyone else, so the answers we would give are linked where we have written them down.
A provider who answers all eight directly is telling you something. So is one who answers three and changes the subject.
BLOG · WHAT WE WRITE ABOUT
WHAT WE PUBLISH
Migration, OCR, volume, billing
the parts that decide a rollout
HOW WE WRITE
From work we actually did
numbers we can point at
WHAT WE DO NOT DO
No listicles, no filler
nothing written to fill a calendar
Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.