You are not looking for an SDK
An invoice-parsing SDK has to bundle the hard parts: OCR binaries or a cloud dependency anyway, layout models that age, template libraries someone must maintain, and a retraining story for every supplier whose invoice looks different. The result is a heavyweight dependency that still phones home for the difficult documents.
The REST shape dissolves all of it. Your integration is the HTTP client your runtime already ships, the request is one multipart POST, and the models, OCR, and format knowledge live server-side where they are updated without a release on your side. Language coverage stops being a feature matrix: C#, Node.js, Python, PHP, Java, Go, and anything else with an HTTP client are all first-class, as the three implementations below show.
Two kinds of PDF, two endpoints
Every invoice PDF that reaches your intake belongs to one of two populations, and the single most valuable design decision in a reception pipeline is telling them apart:
Hybrids and e-invoices. A ZUGFeRD or Factur-X PDF carries the complete invoice as embedded XML; a pure CII or UBL file is that XML on its own. For these, parsing the pixels would be malpractice: the exact data is present, and POST /v1/extract/json reads it deterministically. No AI, no scores, no review queue.
Everything else. A PDF generated from a template with no structured layer, a scan of a paper invoice, a photographed one. Here there is nothing to extract, only pixels and text to understand, and that is POST /v1/parse/json: the AI route, with confidence scores that say how well it went.
The two endpoints return the same JSON envelope, which is what makes the fallback pattern later in this guide two lines instead of an adapter layer. The envelope also matches the request body of the create endpoints, so a parsed supplier invoice can be re-issued or converted onward without remapping.
AI parsing: /v1/parse/json
POST /v1/parse/json accepts a PDF (and only a PDF), typed or scanned, up to 10 pages, as multipart/form-data. It runs OCR and semantic field extraction in a single call and returns the invoice mapped to EN 16931 business terms:
curl -X POST https://api.invoicexml.com/v1/parse/json \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "[email protected]"
The response carries the structured document under invoice and the scoring under confidence:
{
"invoice": {
"invoiceNumber": "R-84112",
"issueDate": "2026-09-22",
"currency": "EUR",
"seller": { "name": "Papier Krause GmbH", "vatIdentifier": "DE811230999", "...": "..." },
"buyer": { "name": "Ihre Firma GmbH", "...": "..." },
"lines": [ { "quantity": 12, "item": { "name": "Kopierpapier A4 80g" }, "...": "..." } ],
"totals": { "payableAmount": 214.20, "...": "..." }
},
"confidence": {
"overall": 0.91,
"areas": {
"sellerIdentification": 0.97,
"buyerIdentification": 0.95,
"taxCalculation": 0.88,
"lineItems": 0.84
}
}
}
Two behaviours are worth knowing because they are deliberate. First, any embedded XML in the uploaded PDF is ignored on this endpoint; if you want the embedded XML read instead, that is /v1/extract/json. Second, an issue date the AI could not actually read stays null rather than being defaulted, because a stamped-in "today" looks like data and is not on the document.
Working with confidence scores
The confidence object is the difference between OCR output you paste into a review screen and a read you can automate on. It carries an overall score and four area scores, each between 0.0 and 1.0: sellerIdentification, buyerIdentification, taxCalculation, and lineItems. The overall value is the average of the areas the AI reported.
A policy that has held up well in production pipelines:
| Signal | Policy |
overall at or above 0.9 | Eligible for automatic booking, subject to your own business checks (known supplier, plausible totals, no duplicate invoice number). |
| Any score below 0.7 | Needs review. Route the document and the parsed JSON to a human queue; the low area tells the reviewer where to look first. |
taxCalculation below your booking bar | Never auto-book, whatever the overall score says. Amount errors are the expensive kind, and this area scoring low is the model telling you the arithmetic did not reconcile cleanly. |
An area score of null | The AI reported no score for that area. Treat it as unscored rather than as zero, and let it trigger review if the area matters to your flow. |
The point of the between-space (0.7 to 0.9) is that it is yours to tune: start conservative, measure how often reviewers change anything, and move the bar with evidence. The scores make that measurable, which no unscored parser output can offer.
POST /v1/extract/json is the other half of intake: it accepts a hybrid PDF (ZUGFeRD, Factur-X) or a standalone e-invoice XML file (CII or UBL), reads the structured data directly, and returns the same JSON envelope, with exact values and no confidence object, because nothing was estimated:
curl -X POST https://api.invoicexml.com/v1/extract/json \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "[email protected]"
When the uploaded PDF carries no embedded XML, the endpoint does not quietly switch to AI. It returns HTTP 400 with errorCode: 4006 (NoEmbeddedXml), and that explicitness is the feature: your pipeline decides whether the AI route runs, your logs show which route produced every booking, and deterministic data is never silently mixed with estimated data.
The fallback pattern in three languages
Put together, intake is: try /v1/extract/json first, and on error code 4006 fall back to /v1/parse/json, carrying the confidence policy from above. Here is the complete pattern.
C#, with HttpClient and System.Text.Json, no packages:
using System.Net.Http.Headers;
using System.Text.Json;
async Task<(JsonDocument Invoice, bool NeedsReview)> ReadInvoiceAsync(
HttpClient http, string path)
{
var extract = await PostPdfAsync(http, "/v1/extract/json", path);
if (extract.StatusCode == System.Net.HttpStatusCode.OK)
{
// Deterministic read: exact values, nothing to review.
var body = JsonDocument.Parse(await extract.Content.ReadAsStringAsync());
return (body, NeedsReview: false);
}
var problem = JsonDocument.Parse(await extract.Content.ReadAsStringAsync());
var errorCode = problem.RootElement.GetProperty("errorCode").GetInt32();
if (errorCode != 4006) throw new InvalidOperationException($"Extract failed: {errorCode}");
// 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
var parse = await PostPdfAsync(http, "/v1/parse/json", path);
parse.EnsureSuccessStatusCode();
var parsed = JsonDocument.Parse(await parse.Content.ReadAsStringAsync());
var confidence = parsed.RootElement.GetProperty("confidence");
var overall = confidence.GetProperty("overall").GetDecimal();
var tax = confidence.GetProperty("areas").GetProperty("taxCalculation");
var needsReview = overall < 0.7m
|| tax.ValueKind == JsonValueKind.Null
|| tax.GetDecimal() < 0.7m;
return (parsed, needsReview);
}
static async Task<HttpResponseMessage> PostPdfAsync(
HttpClient http, string endpoint, string path)
{
using var form = new MultipartFormDataContent();
var pdf = new ByteArrayContent(await File.ReadAllBytesAsync(path));
pdf.Headers.ContentType = new MediaTypeHeaderValue("application/pdf");
form.Add(pdf, "file", Path.GetFileName(path));
return await http.PostAsync(endpoint, form);
}
Node.js, with built-in fetch and FormData, Node 18 or later:
import { readFile } from "node:fs/promises";
const BASE = "https://api.invoicexml.com";
const headers = { Authorization: `Bearer ${process.env.INVOICEXML_API_KEY}` };
async function postPdf(endpoint, path) {
const form = new FormData();
form.append("file", new Blob([await readFile(path)], { type: "application/pdf" }), path);
return fetch(`${BASE}${endpoint}`, { method: "POST", headers, body: form });
}
async function readInvoice(path) {
const extract = await postPdf("/v1/extract/json", path);
if (extract.ok) {
// Deterministic read: exact values, nothing to review.
return { ...(await extract.json()), needsReview: false };
}
const problem = await extract.json();
if (problem.errorCode !== 4006) {
throw new Error(`Extract failed: ${problem.errorCode} ${problem.title}`);
}
// 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
const parse = await postPdf("/v1/parse/json", path);
if (!parse.ok) throw new Error(`Parse failed: ${(await parse.json()).errorCode}`);
const result = await parse.json();
const { overall, areas } = result.confidence;
const needsReview =
overall < 0.7 || areas.taxCalculation == null || areas.taxCalculation < 0.7;
return { ...result, needsReview };
}
Python, with requests:
import os
import requests
BASE = "https://api.invoicexml.com"
HEADERS = {"Authorization": f"Bearer {os.environ['INVOICEXML_API_KEY']}"}
def post_pdf(endpoint: str, path: str) -> requests.Response:
with open(path, "rb") as f:
return requests.post(
f"{BASE}{endpoint}",
headers=HEADERS,
files={"file": (os.path.basename(path), f, "application/pdf")},
)
def read_invoice(path: str) -> dict:
extract = post_pdf("/v1/extract/json", path)
if extract.status_code == 200:
# Deterministic read: exact values, nothing to review.
return {**extract.json(), "needs_review": False}
problem = extract.json()
if problem.get("errorCode") != 4006:
raise RuntimeError(f"Extract failed: {problem.get('errorCode')}")
# 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
parse = post_pdf("/v1/parse/json", path)
parse.raise_for_status()
result = parse.json()
confidence = result["confidence"]
tax = confidence["areas"]["taxCalculation"]
needs_review = (
confidence["overall"] < 0.7 or tax is None or tax < 0.7
)
return {**result, "needs_review": needs_review}
Three languages, one shape, zero dependencies beyond an HTTP client. That is the whole "SDK".
The error codes that shape the pipeline
Errors arrive as RFC 7807 problem responses with a stable numeric errorCode extension, so the branches above never parse message text. Four codes do most of the work in a reading pipeline:
| Code | Name | Meaning and the sane reaction |
4006 | NoEmbeddedXml | The PDF sent to /v1/extract/json has no embedded invoice XML. Not a failure of the document, just the wrong route: fall back to /v1/parse/json. |
4008 | NotAnInvoice | The AI looked and the document is not an invoice: a manual, a contract, a letter, a form, or a receipt without line items. Do not retry; route the file out of the invoice flow. This rejection is what keeps hallucinated bookings out of your ledger. |
4009 | MultipleInvoices | The file contains more than one distinct invoice. The response describes one invoice or none, never a merge; split the file upstream and resubmit page ranges. |
4007 | PdfError | The file could not be processed as a PDF (corrupt, encrypted, or mislabeled). Surface it to whoever supplied the file. |
The 10-page limit on /v1/parse/json is enforced for the same reason 4008 exists: a typical invoice runs one to five pages, and a 60-page PDF hitting the parser is nearly always a statement, a contract, or a bundle that should have been split. The full list of codes lives in the error handling documentation.
Embedded attachments as a ZIP
E-invoices can carry more than the invoice: EN 16931's BG-24 lets a document embed supporting files, such as timesheets, delivery notes, or the original order, base64-encoded inside the XML. POST /v1/extract/attachments takes the same inputs as /v1/extract/json (a hybrid PDF or an e-invoice XML) and returns every embedded attachment as a ZIP download:
curl -X POST https://api.invoicexml.com/v1/extract/attachments \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "[email protected]" \
--output attachments.zip
The payloads are decoded and streamed into the archive verbatim, never parsed or interpreted server-side, and externally referenced documents (BT-124 URLs) are deliberately not fetched: you get exactly what the sender embedded, nothing pulled from the network. If the document carries no attachments, the response says so with errorCode 4013 rather than returning an empty archive.
Get started
The fastest evaluation is your own worst PDF: the scan, the photographed one, the supplier whose layout defeats every template. POST it to /v1/parse/json and read the scores. Create a free InvoiceXML account → and get 100 credits for free, no credit card required.
Processing is stateless throughout: documents are handled in memory, purged when the response ships, and never used to train any model.
Related resources:
InvoiceXML is a REST API for European e-invoice compliance covering ZUGFeRD, Factur-X, XRechnung, Peppol UBL, and CII. Stateless processing, GDPR compliant by architecture, and callable from any stack: C#, Node.js, Python, PHP, Java, Go, or anything else with an HTTP client.