Invoices arrive as pristine Factur-X, as plain UBL, and as a photo of a crumpled page. One extraction API reads them all into the same JSON: the embedded XML when there is one, so the result is exact, and AI with honest confidence scores when there is not. No SDK to install; one REST call from any language.
curl -X POST https://api.invoicexml.com/v1/extract/json \ -H "Authorization: Bearer YOUR_API_KEY" \ -F "[email protected]"
{
"invoice": {
"invoiceNumber": "RE-2026-0142",
"seller": { ... },
"totals": { ... },
"lines": [ ... ]
}
}
Extraction is not one problem. An e-invoice already contains its data; a scanned PDF only pictures it; and some invoices carry whole documents inside them. Each case has its own endpoint, and all three return within the same API, key, and error model.
Factur-X, ZUGFeRD, UBL, CII, XRechnung: the XML is parsed directly, so every field is exact. Deterministic, no AI involved, no confidence caveats.
AI reads the page when there is no XML to read: typed, scanned, or photographed, up to 10 pages, with per-area confidence scores in every response.
Timesheets, delivery notes, and contracts embedded in an e-invoice (BG-24), returned together as a ZIP, copied byte-for-byte.
When you cannot know in advance what a mailbox will contain, wire the two endpoints in sequence. Try /v1/extract/json first: if the document carries XML you get an exact read. When it answers 4006 NoEmbeddedXml, the PDF has no data layer, so send the same bytes to /v1/parse/json and check the confidence that comes back.
{
"title": "No embedded XML found",
"errorCode": 4006,
"valid": false
}
{
"invoice": { ... },
"confidence": {
"overall": 0.94,
"areas": {
"sellerIdentification": 0.97,
"buyerIdentification": 0.95,
"taxCalculation": 0.92,
"lineItems": 0.91
}
}
}
Both extraction endpoints return the same BT-mapped invoice model that the create endpoints accept. Read a received document, adjust what you need, and post the result to /v1/create to issue a compliant UBL, CII, XRechnung, Factur-X, or ZUGFeRD document. Received-to-reissued is a pipeline, not a mapping project.
// 1. read whatever arrived invoice = extract_or_parse("received.pdf") // 2. issue it as a compliant e-invoice curl -X POST /v1/create/xrechnung \ -d '{ "invoice": '$invoice' }'
Your invoices are processed in memory and returned in the same response. Zero data retention is not a policy we enforce, it is an architecture we built.
Processed in volatile memory only. Never written to disk, never queued, never backed up.
Servers in Frankfurt, Germany. No transfers outside the European Economic Area.
Never used for analytics, never to train AI models, never shared with third parties.
SOC 2 Type II, ISO/IEC 27001 and PCI-DSS at the platform layer, held by our infrastructure provider.
GDPR compliant by design · nothing stored on our servers
Start deterministic. If the document might be an e-invoice or a hybrid PDF, call /v1/extract/json: it reads the actual XML, so the result is exact and no AI is involved. Only when it answers 4006 NoEmbeddedXml, meaning the PDF has no machine-readable layer, send the same file to /v1/parse/json and let the AI read the page.
Only where nothing else works. /v1/extract/json and /v1/extract/attachments are fully deterministic: they parse the embedded or supplied XML and never guess. /v1/parse/json is the AI endpoint for PDFs with no structured layer, and it says so honestly by scoring its own confidence in every answer.
/v1/extract/json and /v1/extract/attachments take a PDF or a UBL/CII XML file. /v1/parse/json takes a PDF only, typed, scanned, or photographed, up to 10 pages. Uploads are capped at 20 MB on every extraction endpoint.
Each AI-parsed invoice returns an overall score and four area scores (seller, buyer, tax calculation, line items), from 0.0 to 1.0. At 0.90 and above the read was unambiguous; between 0.70 and 0.89 spot-check high-value documents; below 0.70 route the invoice to a human before it reaches your books. The comparison is yours to make in code, which keeps the policy in your hands.
No, by design. The parser extracts invoices: documents with line items, parties, and a tax breakdown. A receipt without invoice line items is rejected with 4008 NotAnInvoice instead of being force-fitted into an invoice shape, and a file containing several invoices returns 4009 MultipleInvoices so you can split it first.
The envelope matches the /v1/create request body. That makes round trips one-liners: read a received document with extraction, adjust fields, and post the result straight to invoice creation to issue a compliant UBL, CII, XRechnung, Factur-X, or ZUGFeRD document.
E-invoices can carry supporting documents (timesheets, delivery notes, contracts) as embedded attachments (BG-24). /v1/extract/attachments returns them all as a ZIP, copied byte-for-byte. Documents referenced only by URL are listed in the invoice but not fetched, and an invoice with no attachments answers 4013 NoAttachments.
Validate, convert and embed compliant e-invoices through one API. Start your 30-day free trial. No credit card required.