PDF to structured data

PDF in. Structured records out. No chunking pipeline.

The hard part of document extraction from PDFs is rarely the model — it is the scaffolding you build around it to handle documents that do not fit. Here, they fit.

What usually makes this hard

A PDF is a description of where ink goes, not a description of what the document means. Table rows and columns are just glyphs at coordinates. Most AI data extraction pipelines rebuild the meaning with heuristics — and the heuristics break on the documents that matter most.

Then there is the size problem. Long documents exceed a model's context, so you chunk them, which means you now need to decide what a chunk boundary does to a table that spans it. That decision is where most of the engineering time goes.

What changes here

  • A 1M token context window — long documents go in whole
  • Up to 384K tokens of output, for large extraction batches
  • Scanned pages read as images, so a text layer is optional
  • Output shaped to your schema, not to a generic table format

Where it fits best

  • Financial statements
  • Contracts and schedules
  • Claims files
  • Forms with mixed layouts
  • Scanned archive material

If your documents are uniform and machine-generated, a deterministic parser may well be cheaper. The model earns its place on the messy, mixed and scanned material.

Questions about PDF extraction

Do we still need to chunk documents?

Usually not. With a 1M token context window, a long document goes in as one job, which removes the boundary problem entirely. Very large archives are still batched at the document level.

Can it preserve table structure?

Yes. Because the model can see the page, row and column structure survives — which is precisely the information a text-layer extraction tends to lose.

What if the PDF is a scan with no text layer?

It still works. Pages are read as images, so a missing text layer is not a blocker.

Can we get one record per document, or per row?

Whichever your schema needs. You define the output shape, including whether line items are a nested array or separate records.

How do we handle documents in other languages?

Non-English documents extract fine, and you can translate in the same call if you need the output in English.

DOCUMENT EXTRACTION

Diagram of the flow: three stages connected by arrows, with the middle stage highlighted.

Input

Contract · claims file · form

up to 1M tokens of context

DeepSeek V4.1 Flash

Output

Structured record, field by field

schema you define

  • Long documents go in whole
  • No chunking pipeline to build
  • Output shape is yours to set
Whole document in, schema-shaped record out.Diagram of the workload shape. Not a screenshot.

Tell us what you're running.

Send us your monthly volume and what you use today. We usually reply within one business day with a price and a named contact.