Synthetic data

Synthetic data built from your own seeds, not from nothing.

A model inventing data from a bare prompt produces something that looks plausible and tests nothing. Generated from your own examples and rules, it can be genuinely useful.

The distinction that matters

What is synthetic data, in one line? Records a model produced, rather than records you collected. That definition covers both the version that works and the version that does not — which is why the distinction below is the part that actually matters.

Synthetic data has a bad reputation for a good reason. Asked to invent examples from nothing, a model produces text that reads well and contains no signal — it inherits the model's own assumptions rather than the structure of your problem.

The version that works starts from real material: a set of documents you already have, and the rules that govern what a valid record looks like. The model then varies within that structure rather than inventing structure. That is a meaningfully different product.

Seed material

Your own documents and examples provide the structure. The more representative your seeds, the more useful the output.

  • Your documents
  • Your examples

Rules and schema

You state the constraints — field ranges, relationships, formats. The generated data has to satisfy them, not just look right.

  • Your constraints
  • Validated output

Large batches

Up to 384K tokens of output in a single response, which is enough to generate a dataset in far fewer calls than a small-output model would need.

  • 384K output
  • Fewer calls

Off-peak generation

Generation is batch work with no latency requirement, which makes it a natural fit for off-peak billing — half the peak rate.

  • Half rate off-peak
  • Batch friendly

Be careful with this one

If you are generating data to train or evaluate a model, remember that the generator's own biases come along with it. Synthetic data is a supplement to real data, not a replacement — and evaluating a model on data generated by a model tells you less than it appears to. We will say so if your use case looks like it is heading that way.

Questions about synthetic data

Can it generate data that follows our schema exactly?

That is the point of supplying rules alongside seeds. You state the constraints and the output has to satisfy them. Validate the result against your own checks before using it — we will help you set those up.

How much seed data do we need?

Enough to show the model the structure you want, which is usually fewer examples than people expect. Send us what you have and we will tell you honestly whether it is enough.

Is this suitable for training data?

It can augment a training set. It should not be your only data, and evaluating a model against synthetic data alone will flatter it. We will say so if that is where you are heading.

Can it handle structured records as well as text?

Yes — JSON records, tabular rows, and structured documents are all reasonable targets. Tell us the shape you need.

SYNTHETIC DATA

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Input

Seed documents and a schema

your own examples and rules

DeepSeek V4.1 Flash

Output

Generated records at volume

shaped by your schema

  • Generate from your own seeds
  • Long output for large batches
  • Off-peak halves the rate
Seeds and rules in, validated records out.Diagram of the workload shape. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.