Automation MCP Server Features Blog Pricing Contact
Deterministic first · AI when needed · One JSON shape

Extract invoice data from any PDF or e-invoice

Invoices arrive as pristine Factur-X, as plain UBL, and as a photo of a crumpled page. One extraction API reads them all into the same JSON: the embedded XML when there is one, so the result is exact, and AI with honest confidence scores when there is not. No SDK to install; one REST call from any language.

POST /v1/extract/json · no AI involved
curl -X POST https://api.invoicexml.com/v1/extract/json \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]"
200 OK · read from the embedded XML
{
  "invoice": {
    "invoiceNumber": "RE-2026-0142",
    "seller": { ... },
    "totals": { ... },
    "lines": [ ... ]
  }
}
Three routes, chosen by input

The document decides the endpoint

Extraction is not one problem. An e-invoice already contains its data; a scanned PDF only pictures it; and some invoices carry whole documents inside them. Each case has its own endpoint, and all three return within the same API, key, and error model.

If your input is

E-invoices and hybrid PDFs

POST /v1/extract/json

Factur-X, ZUGFeRD, UBL, CII, XRechnung: the XML is parsed directly, so every field is exact. Deterministic, no AI involved, no confidence caveats.

  • Accepts a PDF or a UBL/CII XML file
  • Values come from the XML, never from OCR
  • 4006 NoEmbeddedXml when the PDF has no data layer
If your input is

Ordinary and scanned PDFs

POST /v1/parse/json

AI reads the page when there is no XML to read: typed, scanned, or photographed, up to 10 pages, with per-area confidence scores in every response.

  • Confidence for seller, buyer, tax, and line items
  • Treat scores under 0.70 as needs-review
  • 4008 NotAnInvoice, 4009 MultipleInvoices
If your input is

Embedded supporting documents

POST /v1/extract/attachments

Timesheets, delivery notes, and contracts embedded in an e-invoice (BG-24), returned together as a ZIP, copied byte-for-byte.

  • Accepts a hybrid PDF or a UBL/CII XML file
  • Original filenames and types preserved
  • 4013 NoAttachments when there is nothing inside
The fallback pattern

Deterministic first, AI second

When you cannot know in advance what a mailbox will contain, wire the two endpoints in sequence. Try /v1/extract/json first: if the document carries XML you get an exact read. When it answers 4006 NoEmbeddedXml, the PDF has no data layer, so send the same bytes to /v1/parse/json and check the confidence that comes back.

1 · extractHybrid and native e-invoices resolve here, exactly, with no AI in the path.
2 · on 4006Only plain PDFs fall through. The AI reads them and scores its own certainty.
3 · routePost confident reads onward; send anything under 0.70 to a human first.
400 · /v1/extract/json on a plain PDF
{
  "title": "No embedded XML found",
  "errorCode": 4006,
  "valid": false
}
200 · same file, /v1/parse/json
{
  "invoice": { ... },
  "confidence": {
    "overall": 0.94,
    "areas": {
      "sellerIdentification": 0.97,
      "buyerIdentification": 0.95,
      "taxCalculation": 0.92,
      "lineItems": 0.91
    }
  }
}
One JSON shape

What comes out can go straight back in

Both extraction endpoints return the same BT-mapped invoice model that the create endpoints accept. Read a received document, adjust what you need, and post the result to /v1/create to issue a compliant UBL, CII, XRechnung, Factur-X, or ZUGFeRD document. Received-to-reissued is a pipeline, not a mapping project.

  • Field names follow the EN 16931 business terms, the same across all three routes
  • Deterministic results carry no confidence block, so its presence tells you AI was involved
  • Uploads up to 20 MB, processed in memory and deleted on response
extraction output → create input
// 1. read whatever arrived
invoice = extract_or_parse("received.pdf")

// 2. issue it as a compliant e-invoice
curl -X POST /v1/create/xrechnung \
  -d '{ "invoice": '$invoice' }'
Compliance

Secure by Architecture

Your invoices are processed in memory and returned in the same response. Zero data retention is not a policy we enforce, it is an architecture we built.

Zero data retention

Processed in volatile memory only. Never written to disk, never queued, never backed up.

EU-only processing

Servers in Frankfurt, Germany. No transfers outside the European Economic Area.

No use of your data

Never used for analytics, never to train AI models, never shared with third parties.

Certified infrastructure

SOC 2 Type II, ISO/IEC 27001 and PCI-DSS at the platform layer, held by our infrastructure provider.

Read the full security overview

GDPR compliant by design · nothing stored on our servers

Extraction questions, answered

Which extraction endpoint should I call first?

Start deterministic. If the document might be an e-invoice or a hybrid PDF, call /v1/extract/json: it reads the actual XML, so the result is exact and no AI is involved. Only when it answers 4006 NoEmbeddedXml, meaning the PDF has no machine-readable layer, send the same file to /v1/parse/json and let the AI read the page.

Does invoice extraction use AI?

Only where nothing else works. /v1/extract/json and /v1/extract/attachments are fully deterministic: they parse the embedded or supplied XML and never guess. /v1/parse/json is the AI endpoint for PDFs with no structured layer, and it says so honestly by scoring its own confidence in every answer.

Which file types are accepted?

/v1/extract/json and /v1/extract/attachments take a PDF or a UBL/CII XML file. /v1/parse/json takes a PDF only, typed, scanned, or photographed, up to 10 pages. Uploads are capped at 20 MB on every extraction endpoint.

How should I use the confidence scores?

Each AI-parsed invoice returns an overall score and four area scores (seller, buyer, tax calculation, line items), from 0.0 to 1.0. At 0.90 and above the read was unambiguous; between 0.70 and 0.89 spot-check high-value documents; below 0.70 route the invoice to a human before it reaches your books. The comparison is yours to make in code, which keeps the policy in your hands.

Can it read receipts?

No, by design. The parser extracts invoices: documents with line items, parties, and a tax breakdown. A receipt without invoice line items is rejected with 4008 NotAnInvoice instead of being force-fitted into an invoice shape, and a file containing several invoices returns 4009 MultipleInvoices so you can split it first.

What can I do with the extracted JSON?

The envelope matches the /v1/create request body. That makes round trips one-liners: read a received document with extraction, adjust fields, and post the result straight to invoice creation to issue a compliant UBL, CII, XRechnung, Factur-X, or ZUGFeRD document.

What about the other documents embedded in an e-invoice?

E-invoices can carry supporting documents (timesheets, delivery notes, contracts) as embedded attachments (BG-24). /v1/extract/attachments returns them all as a ZIP, copied byte-for-byte. Documents referenced only by URL are listed in the invoice but not fetched, and an invoice with no attachments answers 4013 NoAttachments.

Start free today

Ready to automate your invoices?

Validate, convert and embed compliant e-invoices through one API. Start your 30-day free trial. No credit card required.

GDPR Compliant No credit card required Setup in minutes
Peppol UBL
Factur-X
EN 16931
142 / 142 passed
Compliant
PDF/A-3 embedded