Benchmark · delivery notes · version 1, measurements July to September 2026
What the engine really does, measured on 195,000 real documents.
Not a brochure percentage: the share of documents validated, with and without human touch, field-level accuracy against human-verified ground truth, behavior on issuers never seen before, throughput — with the full method, replayable on your documents.
The corpus
Real production, not a demo set.
195,514
real documents processed in production at one customer (January–June 2026, four markets) — a flow of 8,000 a day, 2.8 million a year
2,365
distinct layouts observed; 1,714 issuers read; 10,531 shippers in the customer’s reference list
36
transport document types recognized, in 14 families; delivery notes make up 97 % of the flow
6 fields
date · issuer · consignee · document type · number · line items
The documents are scans from real supplier and carrier flows, in all their variety: degraded or skewed scans, handwritten annotations (3,899 documents flagged), international consignment notes in foreign languages (1,475), stamps and carbon copies. They have been processed on the server deployed on the customer’s premises since September 1, 2026. The visuals published on this site are synthetic samples: no real document is shown.
What operations actually see
Documents processed with no human touch.
99.3 %
of documents validated — 97.9 % with no human touch, 1.4 % validated in bulk after a handful of human corrections taught the system their layout
195,514 documents, production database, September 12, 2026
0.7 %
awaiting human review at the time of the snapshot
same database
This rate rises with use. Every human validation or correction is a time-stamped verdict that teaches the system the layout: one verdict immediately re-adjudicates the documents of the same layout still pending — a median of 197 documents per verdict; on September 1, 11 verdicts validated 1,022 documents in one go. The 1.4 % “validated after correction” are exactly that: a few dozen human corrections, 2,819 documents learned. For audit, 1.5 % of automatically validated documents are re-served to a person.
Field-level accuracy
Against human-verified ground truth, field by field.
Counting rule: a field is correct if its value, once normalized (case, accents, whitespace, ISO date), is identical to the one established by a human annotator. Three levels are published, from the most demanding to the floor.
1 · Documents validated automatically, acceptance set with human-verified ground truth (September 2026)
| Field | Accuracy | Documents |
|---|---|---|
| Issuer | 99.4 % | 336 documents |
| Date | 98.9 % | 285 documents |
| Document number | 94.4 % | 285 documents |
| Full core (number + date + issuer all correct together) | 92.9 % | 336 documents |
Document type is measured separately: the annotator’s taxonomy is finer than the engine’s, so the gap is one of vocabulary, not of reading. Simulated end-to-end accuracy across the whole flow, strict policy: 95.0 % (± 0.8) on 700 documents.
2 · Issuers the engine had never seen (1,000 documents, 127 issuers, July 2026)
| Field | Accuracy |
|---|---|
| Document number | 96.4 % |
| Date | 97.2 % |
| Issuer | 87.7 % on the raw read · 97.9 % after the resolution layer |
3 · Floor: the reading service alone, with no layout memory and no resolution (63,838 documents, July 2026)
| Field | Accuracy |
|---|---|
| Document number | 97.7 % |
| Date | 97.5 % |
| Document type | 95.3 % |
| Issuer | 94.3 % |
| Consignee | 94.2 % |
On the clean subset of this sweep (38,849 documents), the four fields number, type, date and issuer are all correct together on 90.6 % of documents.
Throughput
On the baseline deployment, at the customer’s site.
10,216 / hr
documents per hour, average over a full campaign (190,643 documents, September 1, 2026)
7,788 / hr
sustained in steady state, full pipeline
2.31 s
per document, isolated request
Baseline configuration: one Ubuntu server, 12 cores, 64 GB of memory, one server-class GPU, database on the same server, no inbound connection. An outage of the reading service loses nothing: documents wait and are caught up on restart.
Replay the test on your documents
The only number that matters: the one on your documents.
- 01
A representative sample
A few hundred real documents, in the variety of your flows: your big suppliers and your small ones, the clean ones and the damaged ones.
- 02
The same method
Same fields, same counting rule, same confidence threshold. Field-level results, with the list of cases sent for review.
- 03
You decide on numbers
The pilot sets the measured target for go-live. If your documents are not a fit, we tell you before, not after.
Frequently asked questions
About the measurement.
Why publish all this rather than a single percentage?
Because a percentage on its own says nothing about which fields, which documents, or how many human validations. Two engines both advertised “at 99 %” can behave very differently on unknown issuers or damaged scans. The method makes the figure comparable and verifiable — and it is the one we replay on your documents.
What is the difference between “documents validated” and “accuracy”?
The first measures what operations actually see: the share of documents that go through the pipeline — straight through, or validated in bulk after a person corrected a layout once. The second measures, field by field, whether the extracted value is identical to the ground truth established by a human annotator. Both are published; neither is enough on its own.
Are the measurements taken on documents the engine had never seen?
Yes, on two sets: 1,000 documents from 127 issuers absent from the engine’s reference data, and an acceptance set of 697 documents with human-verified ground truth. The sweep of 63,838 documents measures the floor of the reading service alone across the whole flow.
What happens on uncertain fields?
Every field carries a confidence score; below the threshold, the document is validated with a reservation or sent to a person. The principle is never to guess silently: a flagged error costs one review, a silent error costs a dispute. Every human correction immediately re-adjudicates the documents of the same layout still in the queue.
Can I run the test on my own delivery notes?
Yes. Send us a representative sample — a few hundred documents is enough — and we return field-level results, with the cases sent for review. It is the starting point of every project.
Do you publish throughput?
Yes: 10,216 documents per hour on average over a full campaign, 7,788 sustained in steady state, on the baseline deployment installed at the customer’s site (one server, one server-class GPU).
Send us 200 delivery notes. We send back the numbers.
30 minutes to scope the sample and the flow, then field-level results.