Automation MCP Server Features Blog Pricing Contact

Reading PDF Invoices in Code: Scanned PDF to JSON with Confidence Scores

Most searches for an invoice parser SDK end at a REST call: one endpoint reads typed, scanned, and photographed invoice PDFs into EN 16931-shaped JSON with per-area confidence scores, and a second reads e-invoices and hybrid PDFs deterministically, with no AI involved. This guide covers both routes, the fallback pattern that connects them, the confidence thresholds worth automating on, and the error codes that keep non-invoices out of your books, with code in C#, Node.js, and Python.

The search that leads here usually says "SDK": an SDK to read PDF invoices, an invoice parser SDK, a library to get supplier PDFs into the ERP. The honest answer is that the SDK era of this problem is over. Reading an invoice PDF well takes OCR, layout understanding, and semantic mapping onto invoice fields, none of which belongs in a client library that you version, ship, and retrain. What replaced it is one multipart REST call that accepts the PDF and returns the invoice as structured JSON, scored for how much you should trust it.

This guide covers the reading side of the InvoiceXML API end to end: the AI parser for ordinary and scanned PDFs, the deterministic extractor for e-invoices that carry their data as embedded XML, the fallback pattern that chains them, the confidence thresholds worth automating on, and the error codes that keep the pipeline honest. Working code in C#, Node.js, and Python throughout.

One scope note up front: the parser reads invoices. Receipts without line items, along with manuals, contracts, and letters, are deliberately rejected rather than force-fitted into invoice fields, and this guide will show exactly how that rejection surfaces in code.

You are not looking for an SDK

An invoice-parsing SDK has to bundle the hard parts: OCR binaries or a cloud dependency anyway, layout models that age, template libraries someone must maintain, and a retraining story for every supplier whose invoice looks different. The result is a heavyweight dependency that still phones home for the difficult documents.

The REST shape dissolves all of it. Your integration is the HTTP client your runtime already ships, the request is one multipart POST, and the models, OCR, and format knowledge live server-side where they are updated without a release on your side. Language coverage stops being a feature matrix: C#, Node.js, Python, PHP, Java, Go, and anything else with an HTTP client are all first-class, as the three implementations below show.


Two kinds of PDF, two endpoints

Every invoice PDF that reaches your intake belongs to one of two populations, and the single most valuable design decision in a reception pipeline is telling them apart:

Hybrids and e-invoices. A ZUGFeRD or Factur-X PDF carries the complete invoice as embedded XML; a pure CII or UBL file is that XML on its own. For these, parsing the pixels would be malpractice: the exact data is present, and POST /v1/extract/json reads it deterministically. No AI, no scores, no review queue.

Everything else. A PDF generated from a template with no structured layer, a scan of a paper invoice, a photographed one. Here there is nothing to extract, only pixels and text to understand, and that is POST /v1/parse/json: the AI route, with confidence scores that say how well it went.

The two endpoints return the same JSON envelope, which is what makes the fallback pattern later in this guide two lines instead of an adapter layer. The envelope also matches the request body of the create endpoints, so a parsed supplier invoice can be re-issued or converted onward without remapping.


AI parsing: /v1/parse/json

POST /v1/parse/json accepts a PDF (and only a PDF), typed or scanned, up to 10 pages, as multipart/form-data. It runs OCR and semantic field extraction in a single call and returns the invoice mapped to EN 16931 business terms:

curl -X POST https://api.invoicexml.com/v1/parse/json \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]"

The response carries the structured document under invoice and the scoring under confidence:

{
  "invoice": {
    "invoiceNumber": "R-84112",
    "issueDate": "2026-09-22",
    "currency": "EUR",
    "seller": { "name": "Papier Krause GmbH", "vatIdentifier": "DE811230999", "...": "..." },
    "buyer": { "name": "Ihre Firma GmbH", "...": "..." },
    "lines": [ { "quantity": 12, "item": { "name": "Kopierpapier A4 80g" }, "...": "..." } ],
    "totals": { "payableAmount": 214.20, "...": "..." }
  },
  "confidence": {
    "overall": 0.91,
    "areas": {
      "sellerIdentification": 0.97,
      "buyerIdentification": 0.95,
      "taxCalculation": 0.88,
      "lineItems": 0.84
    }
  }
}

Two behaviours are worth knowing because they are deliberate. First, any embedded XML in the uploaded PDF is ignored on this endpoint; if you want the embedded XML read instead, that is /v1/extract/json. Second, an issue date the AI could not actually read stays null rather than being defaulted, because a stamped-in "today" looks like data and is not on the document.


Working with confidence scores

The confidence object is the difference between OCR output you paste into a review screen and a read you can automate on. It carries an overall score and four area scores, each between 0.0 and 1.0: sellerIdentification, buyerIdentification, taxCalculation, and lineItems. The overall value is the average of the areas the AI reported.

A policy that has held up well in production pipelines:

SignalPolicy
overall at or above 0.9Eligible for automatic booking, subject to your own business checks (known supplier, plausible totals, no duplicate invoice number).
Any score below 0.7Needs review. Route the document and the parsed JSON to a human queue; the low area tells the reviewer where to look first.
taxCalculation below your booking barNever auto-book, whatever the overall score says. Amount errors are the expensive kind, and this area scoring low is the model telling you the arithmetic did not reconcile cleanly.
An area score of nullThe AI reported no score for that area. Treat it as unscored rather than as zero, and let it trigger review if the area matters to your flow.

The point of the between-space (0.7 to 0.9) is that it is yours to tune: start conservative, measure how often reviewers change anything, and move the bar with evidence. The scores make that measurable, which no unscored parser output can offer.


Deterministic extraction: /v1/extract/json

POST /v1/extract/json is the other half of intake: it accepts a hybrid PDF (ZUGFeRD, Factur-X) or a standalone e-invoice XML file (CII or UBL), reads the structured data directly, and returns the same JSON envelope, with exact values and no confidence object, because nothing was estimated:

curl -X POST https://api.invoicexml.com/v1/extract/json \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]"

When the uploaded PDF carries no embedded XML, the endpoint does not quietly switch to AI. It returns HTTP 400 with errorCode: 4006 (NoEmbeddedXml), and that explicitness is the feature: your pipeline decides whether the AI route runs, your logs show which route produced every booking, and deterministic data is never silently mixed with estimated data.


The fallback pattern in three languages

Put together, intake is: try /v1/extract/json first, and on error code 4006 fall back to /v1/parse/json, carrying the confidence policy from above. Here is the complete pattern.

C#, with HttpClient and System.Text.Json, no packages:

using System.Net.Http.Headers;
using System.Text.Json;

async Task<(JsonDocument Invoice, bool NeedsReview)> ReadInvoiceAsync(
    HttpClient http, string path)
{
    var extract = await PostPdfAsync(http, "/v1/extract/json", path);
    if (extract.StatusCode == System.Net.HttpStatusCode.OK)
    {
        // Deterministic read: exact values, nothing to review.
        var body = JsonDocument.Parse(await extract.Content.ReadAsStringAsync());
        return (body, NeedsReview: false);
    }

    var problem = JsonDocument.Parse(await extract.Content.ReadAsStringAsync());
    var errorCode = problem.RootElement.GetProperty("errorCode").GetInt32();
    if (errorCode != 4006) throw new InvalidOperationException($"Extract failed: {errorCode}");

    // 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
    var parse = await PostPdfAsync(http, "/v1/parse/json", path);
    parse.EnsureSuccessStatusCode();
    var parsed = JsonDocument.Parse(await parse.Content.ReadAsStringAsync());

    var confidence = parsed.RootElement.GetProperty("confidence");
    var overall = confidence.GetProperty("overall").GetDecimal();
    var tax = confidence.GetProperty("areas").GetProperty("taxCalculation");
    var needsReview = overall < 0.7m
        || tax.ValueKind == JsonValueKind.Null
        || tax.GetDecimal() < 0.7m;

    return (parsed, needsReview);
}

static async Task<HttpResponseMessage> PostPdfAsync(
    HttpClient http, string endpoint, string path)
{
    using var form = new MultipartFormDataContent();
    var pdf = new ByteArrayContent(await File.ReadAllBytesAsync(path));
    pdf.Headers.ContentType = new MediaTypeHeaderValue("application/pdf");
    form.Add(pdf, "file", Path.GetFileName(path));
    return await http.PostAsync(endpoint, form);
}

Node.js, with built-in fetch and FormData, Node 18 or later:

import { readFile } from "node:fs/promises";

const BASE = "https://api.invoicexml.com";
const headers = { Authorization: `Bearer ${process.env.INVOICEXML_API_KEY}` };

async function postPdf(endpoint, path) {
  const form = new FormData();
  form.append("file", new Blob([await readFile(path)], { type: "application/pdf" }), path);
  return fetch(`${BASE}${endpoint}`, { method: "POST", headers, body: form });
}

async function readInvoice(path) {
  const extract = await postPdf("/v1/extract/json", path);
  if (extract.ok) {
    // Deterministic read: exact values, nothing to review.
    return { ...(await extract.json()), needsReview: false };
  }

  const problem = await extract.json();
  if (problem.errorCode !== 4006) {
    throw new Error(`Extract failed: ${problem.errorCode} ${problem.title}`);
  }

  // 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
  const parse = await postPdf("/v1/parse/json", path);
  if (!parse.ok) throw new Error(`Parse failed: ${(await parse.json()).errorCode}`);

  const result = await parse.json();
  const { overall, areas } = result.confidence;
  const needsReview =
    overall < 0.7 || areas.taxCalculation == null || areas.taxCalculation < 0.7;

  return { ...result, needsReview };
}

Python, with requests:

import os
import requests

BASE = "https://api.invoicexml.com"
HEADERS = {"Authorization": f"Bearer {os.environ['INVOICEXML_API_KEY']}"}


def post_pdf(endpoint: str, path: str) -> requests.Response:
    with open(path, "rb") as f:
        return requests.post(
            f"{BASE}{endpoint}",
            headers=HEADERS,
            files={"file": (os.path.basename(path), f, "application/pdf")},
        )


def read_invoice(path: str) -> dict:
    extract = post_pdf("/v1/extract/json", path)
    if extract.status_code == 200:
        # Deterministic read: exact values, nothing to review.
        return {**extract.json(), "needs_review": False}

    problem = extract.json()
    if problem.get("errorCode") != 4006:
        raise RuntimeError(f"Extract failed: {problem.get('errorCode')}")

    # 4006 NoEmbeddedXml: an ordinary PDF, so take the AI route.
    parse = post_pdf("/v1/parse/json", path)
    parse.raise_for_status()

    result = parse.json()
    confidence = result["confidence"]
    tax = confidence["areas"]["taxCalculation"]
    needs_review = (
        confidence["overall"] < 0.7 or tax is None or tax < 0.7
    )
    return {**result, "needs_review": needs_review}

Three languages, one shape, zero dependencies beyond an HTTP client. That is the whole "SDK".


The error codes that shape the pipeline

Errors arrive as RFC 7807 problem responses with a stable numeric errorCode extension, so the branches above never parse message text. Four codes do most of the work in a reading pipeline:

CodeNameMeaning and the sane reaction
4006NoEmbeddedXmlThe PDF sent to /v1/extract/json has no embedded invoice XML. Not a failure of the document, just the wrong route: fall back to /v1/parse/json.
4008NotAnInvoiceThe AI looked and the document is not an invoice: a manual, a contract, a letter, a form, or a receipt without line items. Do not retry; route the file out of the invoice flow. This rejection is what keeps hallucinated bookings out of your ledger.
4009MultipleInvoicesThe file contains more than one distinct invoice. The response describes one invoice or none, never a merge; split the file upstream and resubmit page ranges.
4007PdfErrorThe file could not be processed as a PDF (corrupt, encrypted, or mislabeled). Surface it to whoever supplied the file.

The 10-page limit on /v1/parse/json is enforced for the same reason 4008 exists: a typical invoice runs one to five pages, and a 60-page PDF hitting the parser is nearly always a statement, a contract, or a bundle that should have been split. The full list of codes lives in the error handling documentation.


Embedded attachments as a ZIP

E-invoices can carry more than the invoice: EN 16931's BG-24 lets a document embed supporting files, such as timesheets, delivery notes, or the original order, base64-encoded inside the XML. POST /v1/extract/attachments takes the same inputs as /v1/extract/json (a hybrid PDF or an e-invoice XML) and returns every embedded attachment as a ZIP download:

curl -X POST https://api.invoicexml.com/v1/extract/attachments \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]" \
  --output attachments.zip

The payloads are decoded and streamed into the archive verbatim, never parsed or interpreted server-side, and externally referenced documents (BT-124 URLs) are deliberately not fetched: you get exactly what the sender embedded, nothing pulled from the network. If the document carries no attachments, the response says so with errorCode 4013 rather than returning an empty archive.


Get started

The fastest evaluation is your own worst PDF: the scan, the photographed one, the supplier whose layout defeats every template. POST it to /v1/parse/json and read the scores. Create a free InvoiceXML account → and get 100 credits for free, no credit card required.

Processing is stateless throughout: documents are handled in memory, purged when the response ships, and never used to train any model.

Related resources:


InvoiceXML is a REST API for European e-invoice compliance covering ZUGFeRD, Factur-X, XRechnung, Peppol UBL, and CII. Stateless processing, GDPR compliant by architecture, and callable from any stack: C#, Node.js, Python, PHP, Java, Go, or anything else with an HTTP client.

Start free today

Ready to automate your invoices?

Validate, convert and embed compliant e-invoices through one API. Start your 30-day free trial. No credit card required.

GDPR Compliant No credit card required Setup in minutes
Peppol UBL
Factur-X
EN 16931
142 / 142 passed
Compliant
PDF/A-3 embedded