All models

Document Intelligence

Read anything, understand everything, execute the workflow

Extraction is the first step and the least interesting one. This model works out what a document is, what it means once the amendments and the context are applied, whether it agrees with your system of record, and then does the thing the document was always going to trigger — the posting, the write-back, the filing — routing to a person only what genuinely needs one.

Classify and linkUnderstand in contextExecute downstream

The problem

Text on a page is not the same as knowing what to do about it

Most document tooling stops at characters. It returns fields with a confidence score and leaves the actual work untouched: nobody has decided which document this is, which version of it is in force, whether the numbers agree with the ledger, or what should now happen. So the extracted values land in a queue and a person opens the source document anyway — to check the amendment, to find the matching record, to key the result into the system that was the point of the exercise. The reading was never the bottleneck. The judgement and the downstream action were.

Capability

From page to meaning

What a document is, what it says, and what it means once everything else on file is taken into account.

Classification and linking

Each document is identified by what it is rather than by which folder it arrived in, then linked to the entity, counterparty, agreement or account it belongs to. An amendment attaches to its master agreement, a receipt to its claim, a statement to its account — so the document arrives already in its context.

Extraction with per-field confidence

Native digital files, scanned pages and photographs taken on a phone are read the same way, including rotated, creased and partially legible pages. Every field carries its own confidence and a pointer to the region it came from, so a low-confidence value is a specific question rather than a reason to distrust the whole document.

Context resolution

The terms in force are the terms after the last amendment. Renewals, side letters, addenda, revised schedules and superseding versions are read together with the original, and the model works from the resulting position rather than from whichever page it happened to open.

Rules applied to values, not to templates

Policy, contractual terms, entitlements, limits and tax treatment are evaluated against the extracted values themselves. A new supplier, a reformatted statement or an unfamiliar layout does not need a new template, because the rules were never bound to a layout.

Capability

From meaning to completed work

Agreement with the system of record, a route for what cannot be resolved, and the downstream action carried out.

Reconciled against the system of record

Values are checked against what your systems already hold — the purchase order, the agreement, the account, the prior period, the ledger balance. A document that agrees is evidence; a document that does not is a named difference with both sides shown.

Exception routing with the page open beside the field

What cannot be resolved is routed to the person who can resolve it, with the source page rendered next to the field in question, the rule that fired, and the record it disagreed with. The reviewer confirms or corrects in one place, and the correction is recorded against their name.

Workflow execution

The write-back to the source system, the journal, the schedule update, the filing, the notification — the action the document was always going to trigger is completed, not queued. Straight-through where everything agrees, held where it does not.

Evidence kept with the outcome

The page, the extracted values, the version of the rules applied, the reconciliation result and the resulting action are stored together. When somebody asks in nine months why a figure is what it is, the answer opens rather than gets reconstructed.

Scope

The documents it already reads

Leases and contracts, including their amendments and schedules. KYC and onboarding forms, with identity documents and supporting evidence. Loan and credit files, from application packs to sanction letters and repayment schedules. Invoices and receipts, native or photographed. Bank statements, account statements and transaction files, whatever shape the exporting system produced. The list is not a template library — an unfamiliar document type is a configuration conversation, not a rebuild.

How it works

From arrival to completed action

  1. 01

    Arrive

    Documents come in from email, upload, shared drives, portals, scanners and phone cameras, and are held with their original file so nothing downstream depends on a copy.

  2. 02

    Classify and link

    The document type is identified from the content, and the document is attached to the counterparty, agreement, account or case it belongs to.

  3. 03

    Extract

    Fields, tables and clauses are read from native, scanned or photographed pages, each with its own confidence and a pointer back to the region it came from.

  4. 04

    Resolve context and apply rules

    Amendments and prior versions are folded in to establish the terms in force, and policy, contractual and tax rules are evaluated against the resulting values.

  5. 05

    Reconcile

    The result is checked against the system of record, and any difference is classified by cause rather than reported as a single mismatch.

  6. 06

    Execute or route

    Where everything agrees, the downstream action is carried out. Where it does not, the item is routed with the page, the rule and the conflicting record attached, and completes once resolved.

Questions

Frequently asked

How is this different from OCR?
OCR turns an image into text. That is the first step here and the least interesting one. This model decides what the document is, links it to the agreement or account it belongs to, works out which terms are actually in force after amendments, applies your rules to the values, reconciles against your system of record, and then completes the downstream action. If all you need is text off a page, an OCR engine is cheaper and you should use one.
Do we have to build a template for every document type or supplier?
No. Rules are applied to the extracted values rather than to positions on a page, so a new supplier's layout or a reformatted statement does not require a new template. Genuinely new document types are a configuration conversation about the fields, the rules and the downstream action — not a rebuild.
What happens when the document is a bad photograph?
It is read, and the fields that could not be read confidently are flagged individually rather than the whole document being rejected. The reviewer sees the page region beside the field, so a two-second confirmation replaces a re-scan request in most cases.
How do you handle amendments and superseded versions?
An amendment is linked to its master agreement and read together with it, so the position the model works from is the position after the last change. Where documents conflict and the precedence is not clear, that is an exception with both versions shown rather than a silent choice.
Can it write back into our systems, or does it just hand us a file?
It writes back. Executing the workflow — the posting, the schedule update, the status change, the filing — is the point of the model, and integration with the systems that hold the record is part of the deployment rather than a later phase.
What accuracy do you quote?
We do not publish a headline extraction accuracy, because it is meaningless without the document mix it was measured on. We would rather run your own documents through in a sandbox and show you the field-level result, the exceptions and the straight-through rate on your material.

Send us the documents you think are unreadable

A folder of leases with their amendments, a batch of photographed receipts, a set of statements from four different banks. We will show you what is read, what is reconciled, what is routed and what completes on its own.

Last reviewed

Essential cookies are required for the site to function and cannot be switched off. Everything else is off until you switch it on, and you can change or withdraw your choice at any time from the Cookie settings link in the footer. The Cookie Policy lists the cookies we set and how long each one lasts.

No choice recorded yet