Dijit.app Dijit.app Documentation

Documentation/ Checking the document header/ How a document is read

How Dijit reads a document

Knowing what the system does with a document explains why some fields always come out right and others want a look. There are four stages and it takes under three seconds a page.

The four stages

StageWhat happens
1. RecognitionOCR turns the image of the document into text you can find things in.
2. ExtractionThe AI models work out what each value is: which of those bits of text is the issuer, which the date, which the net amount.
3. CheckingThe document is tested against itself: that net plus taxes is the total, that the lines add up to the net, that the tax breakdown agrees, and that this is not a document you already have.
4. CodingNominal accounts are filled in from your company's history and the rules you have set.

What comes out of stage three is what decides the status the document shows in the list.

OCR and artificial intelligence

They are two different things and you need both:

Which is why the system does not depend on a template per supplier: it does not need to know in advance where each one prints what.

Accuracy, and what it turns on

Average accuracy on reading and coding is 99 % on an original PDF. On photographed documents it follows the quality of the image.

Where the document came fromHow it behaves
A PDF made by the sender's systemThe text is inside the file. The best case there is.
A scanned PDFNeeds recognising. Good, with a clean scan.
A photographSensitive to framing, lighting and shadows.
A thermal receipt or a worn photocopyThe worst case: the original has already lost information.

The practical upshot: when a document is always read badly, look at how it arrives before you correct it again and again. Improving the source settles more than correcting the result.

What it learns from your history

The system does not treat every document as if it were the first:

Which is why there is more checking in the first month than in the sixth: every correction takes work off the future.

What it will not do

It does not fill in data that is not on the document. If a value is not there, or cannot be read, the field is left empty and the document is flagged for a look. An empty field is not a reading failure: it is the record that there was nothing legible there.

Nor does it decide for you: it checks and it flags, but approving is always down to a person or to the route you have set up.

Where it all happens

The models run inside Dijit.app's own Azure estate, on servers in the European Union. Customer documents are never used to train models.

↑ Back to top

Last reviewed: 11 September 2026 · The Dijit.app team