Transport document AI

Your delivery notes, CMRs and transport orders become TMS data — no more rekeying.

A document engine specialized in transport, in production at one customer on a flow of 8,000 documents a day — 2.8 million a year: 99.3 % of documents validated — and the system learns from every correction —, 100 % handled, up to 10,000 documents an hour on the baseline deployment. It reads, understands, checks and writes into the TMS you already have. On-premise, flat annual license, no per-page pricing.

99.3 % of documents validated195,514 real documents in production2,365 layouts · 10,531 shippers36 document types, 14 familiesUp to 10,000 documents / hourOn-premiseKeep your TMS

The proof

Measured in real production, not on a demo.

99.3 %

of documents validated — 97.9 % with no human touch, 1.4 % taught to the system by a handful of human corrections. 195,514 real documents, production database.

99.4 %

issuer accuracy on documents validated automatically, against human-verified ground truth; date 98.9 %, document number 94.4 %.

195,514

real documents processed in six months at one customer — a flow of 8,000 a day, 2.8 million a year.

2,365

distinct layouts read, 1,714 issuers observed, 10,531 shippers in the reference list.

36

transport document types recognized, in 14 families — type, number, date and issuer extracted on each.

10,216 / hr

documents per hour, campaign average on the baseline deployment — 100 % of documents handled.

A figure is only worth its method: which fields, which documents, how much human-in-the-loop validation. We publish all of it, and we replay the same method on your own documents during the pilot, before any commitment.

The measurement method, field by field →

Fields extracted

Not raw text: data.

Classic OCR returns characters. The engine returns named fields, normalized and matched against your reference data, ready to enter the TMS.

On every recognized document (36 types, 14 families)

Document type · number · date · issuer — out of the box, with no setup.

On the delivery note, and on your fine-tuned types

Full extraction, in production: the consignee, matched to your sites, and the line items — references, quantities, units, weights, batches, packages.

Document date

Whatever the written format, normalized.

Issuer

Supplier, carrier or shipper — identified even when the letterhead changes.

Consignee

Site, warehouse or customer delivered to.

Document type

Delivery note, CMR, order, POD, invoice: classified automatically within the flow.

Document number

Plus the related numbers: purchase order, transport order, customer reference.

Line items

References, descriptions, quantities, units, weights, batches, packages — as a structured table.

{
  "type": "bon_de_livraison",
  "numero": "BL-2026-04871",
  "date": "2026-09-09",
  "emetteur": { "nom": "…", "siren": "…" },
  "destinataire": { "site": "Entrepôt Perpignan" },
  "commande_ref": "CMD-77120",
  "lignes": [
    { "ref": "MV-310", "designation": "Carottes sable — sac 10 kg", "qte": 60, "unite": "sacs", "poids_kg": 600.0 }
  ],
  "confiance": { "date": 0.99, "emetteur": 0.98, "destinataire": 0.97, "numero": 0.96 },
  "etat": "AUTO",
  "a_valider": []
}

Field names are those of the production API schema (French keys): type, number, date, issuer, consignee, order reference, line items, per-field confidence, status, fields to review.

Examples

Document → data, real extraction.

Four synthetic delivery notes — invented companies, addresses and references, no customer data — run through the production engine on September 12, 2026. The document on the left, the output on the right, with the fields sent for human review highlighted.

Synthetic sample — Clean PDF: the document on the left, the JSON extraction on the right
Clean PDF. Full read, automatic validation.
Synthetic sample — Noisy office scan, slight skew: the document on the left, the JSON extraction on the right
Noisy office scan, slight skew. Number and date read but not legibly found on the page: sent for review.
Synthetic sample — Handwritten annotations: the document on the left, the JSON extraction on the right
Handwritten annotations. Hand-corrected quantity and driver’s reservation detected: line sent for review.
Synthetic sample — Document in Italian: the document on the left, the JSON extraction on the right
Document in Italian. Same read, different language.

Synthetic samples; recognition and field reading by the production services. No real document is published.

How it works

OCR to read, reasoning to understand.

Two stages: an OCR robust to scans and photos, then a specialized model that interprets the document the way a dispatcher would — which issuer, which order, which lines, which reservation. That second stage, built on millions of transport documents, is what separates it from template-based OCR.

  1. 01

    Capture

    PDF dropped via API, into a watched folder, a SharePoint library or by email: the document comes in through the channel you already use.

  2. 02

    Read

    The document is deskewed and OCR reads the text, scans included; the specialized model understands what each zone means — not just the characters.

  3. 03

    Structure

    Fields are extracted, normalized and matched against your reference data: suppliers, items, open orders.

  4. 04

    Check

    Every field carries a confidence score. What is certain goes through; what is not is highlighted for a person.

  5. 05

    Write

    Validated data is written into your TMS or ERP, or exported (JSON, CSV, API, webhook). The document is archived with its extraction.

Non-negotiable principle: the AI never guesses silently. Every field carries a confidence score; below the threshold, a person decides, and their correction is used to fine-tune the model on your layouts.

A system that learns

Every correction your team makes is worth two hundred.

The engine does more than read: it learns from the dispatcher. When a person validates or corrects a field, that verdict is memorized for the document’s layout, and every document of the same layout still pending is re-adjudicated on the spot — a median of 197 documents per verdict, measured in production. That is how the share of validated documents went from 63 % cold to 99.3 % over 195,514 documents: a few dozen human corrections, thousands of documents learned.

One verdict, one layout learned

A dispatcher’s correction is authoritative: it is memorized, time-stamped, and a refuted value never comes back through automatic validation.

Immediate re-adjudication

Documents from the same issuer still in the queue are revalidated in one go. On September 1, 2026, eleven verdicts validated 1,022 documents.

Controlled, not blind

Automatic cross-checks run continuously, and 1.5 % of automatically validated documents are re-served to a person as a random audit.

The figures and the method →

Your TMS

Keep your TMS. We automate what goes into it.

Switching TMS to get automated document capture means running a migration project to solve a data-entry problem. The GTEK engine is an intelligence layer added on top of whatever system you run today: REST API, webhooks, SFTP drop or shared folder, watched mailbox, CSV and JSON exports. Mapping to your fields is set during implementation, and nothing is written into the TMS without going through the check.

Automated order entry in your TMS: the details →

On-premise & license

Your volume doubles? Your bill does not.

The document-capture market bills per page, per credit or by meter: cost tracks your activity, and the meter resets every year. The GTEK engine is deployed on-premise and licensed annually. The millionth document costs the same as the first: nothing.

Flat annual license
No per-document fees
No per-page fees
No credits to buy
Unlimited volume*
Data processed on-premise

* within your server’s capacity — the only limit is physical, not contractual.

Your documents stay in-house

The engine runs on a server inside your premises. Delivery notes, CMRs, negotiated rates, customers: nothing is sent to a remote AI service. It even works without an Internet connection.

No per-page pricing

One flat annual license, unlimited volume within your server’s capacity. At several thousand documents a day, the cost is known in advance — it does not climb with your activity.

The model fine-tunes on your layouts

Your validations feed the model deployed on your premises. It gets better on your suppliers, your sites, your habits — and that progress stays with you.

Comparison

Transport engine, generic OCR or a TMS switch.

GTEK transport engine Generic online OCR AI bundled with a new TMS
Built for transport documents Yes — measured on 195,514 real delivery notes, 2,365 layouts No, generic models Sometimes, on the vendor’s layouts
Document types 36 types in 14 families; full extraction on delivery notes, fine-tuning on your other types A few predefined families The vendor’s
Adapted to your documents Fine-tuned on your layouts during the pilot Templates to configure Depends on the vendor
New supplier, unknown layout Template-free: read with no setup; sent for review if in doubt A new template to build Depends on the vendor
Where your documents go On your server, on-premise The vendor’s cloud The vendor’s cloud
Your TMS You keep it You keep it You switch TMS
Cost at high volume Flat annual license, unlimited volume within your server’s capacity Per-page or credit-based — climbs with volume Bundled into a TMS contract
Accuracy 99.3 % of documents validated; field-level accuracy published Claimed, rarely detailed per field Claimed
Throughput Up to 10,000 documents / hour (baseline deployment, measured) Depends on the plan and API quotas Depends on the vendor

And versus a public chatbot on your documents? Sovereignty & GDPR →

What it costs you today

How much does transport document entry cost you today?

People, hours per week, profile: in 30 minutes we work out the hours consumed each month and the capacity they represent — on your numbers, not ours.

Work it out with us

Frequently asked questions

Logistics document OCR: your questions.

What exactly are your figures worth?

They come from real production at one customer: 195,514 delivery notes processed, 99.3 % validated — 97.9 % with no human touch, 1.4 % validated in bulk after a handful of human corrections taught the system their layout — and that rate rises with use. Against human-verified ground truth, on documents validated automatically, the issuer is correct 99.4 % of the time, the date 98.9 %, the document number 94.4 %. On issuers never seen before, the number is read at 96.4 % and the date at 97.2 % before any resolution. A damaged document is still processed — 100 % of documents are handled — and its uncertain fields are flagged, never invented. The full method is on the benchmark page.

How fast is it?

10,216 documents per hour on average over a full campaign, 7,788 sustained, measured on the baseline deployment installed at a customer’s site — one server, one server-class GPU. Actual throughput on your documents and your server is measured during the pilot.

Do you read handwritten, damaged or foreign-language documents?

Yes, and they are part of the real flow that was measured: degraded or skewed scans, handwritten annotations (3,899 documents flagged by the detector), international consignment notes in foreign languages (1,475). All are processed — no document is rejected. In every case, when the read is uncertain, the field is highlighted and a person reviews it: the AI never guesses silently.

Which document types do you recognize?

Thirty-six transport document types, grouped into fourteen families, identified on more than 3,500 real pages: delivery note, domestic and international consignment note (CMR), transport note, transport confirmation, transport order, purchase order, order confirmation, picking list, dispatch note, pallet sheet, loading list, collection note, pick-up note, invoice, return note, and their Spanish and Italian variants. On each of them the engine extracts the type, number, date and issuer out of the box. Full extraction — consignee, line items — is trained on the delivery note and extends to your other types through fine-tuning.

Can you handle a document specific to my company?

Yes. The model is fine-tuned on any document type from your own samples: we define the fields to extract together, measure on a sample during the pilot, then go live with human-in-the-loop validation of the uncertain cases.

What if a supplier sends a layout we have never seen?

The engine is template-free. On 1,000 documents from 127 issuers it had never seen, it reads the number at 96.4 % and the date at 97.2 % on the raw read alone. The first documents from an unknown issuer are validated with a reservation or by your team, and each validation memorizes the layout for the next ones.

Do I have to switch TMS or ERP?

No. The engine is added to what you already have: API, webhooks, SFTP drop or folder, watched mailbox, CSV/JSON exports. Mapping to your TMS fields is done during implementation.

Do my documents leave the company?

No. The engine is deployed on a server inside your premises. Documents, extractions and the model stay with you; nothing is sent to a remote AI service.

How is it priced?

A flat annual license with unlimited volume within your server’s capacity — no per-page, per-document or credit fees. At several thousand documents a day, that is what keeps the cost predictable. The amount is quoted after a free 30-minute first call, based on scope and server.

What exactly does the license cover?

One legal entity, one production environment, for your internal needs — with an unlimited number of documents and users within your server’s capacity. An additional site or subsidiary, a second production server, an additional document family or a higher service level is an extension, quoted separately.

How long does it take to get started?

A few weeks for a first flow: server installation, connection of the intake channel and the TMS, a measured pilot on your documents, then go-live with human-in-the-loop validation of the uncertain cases.

Can we test on our own documents before deciding?

Yes — that is the rule: we measure accuracy on a sample of your real documents during the pilot, field by field, before any commitment.

Send us your documents. We measure in front of you.

30 minutes, remote, free: we look at where your documents come from, where they need to go, and show you the extraction on a case close to yours.