Why scans are different
A born-digital PDF contains its text as characters with positions. A scanned PDF contains an image of a page. Nothing can compare the words on two scans until something has read them, which is what optical character recognition does: it adds an invisible text layer over the image.
What works
If your scanner, copier or document management system applies OCR on ingest, the PDF has a text layer and the comparison works on it. The engine detects that the text came from OCR and treats formatting differences with suspicion, because OCR guesses bold and italic from glyph shapes and gets it wrong often. Substantive text changes are still reported normally.
What does not work
- A scan with no text layer at all. The comparison cannot see any text and reports that the documents have no comparable content.
- Very poor scans, where OCR misreads characters. Those misreads show up as changes that are not really there. Rescan at 300 dpi or better.
- Handwritten additions. OCR does not read handwriting reliably, so a handwritten change to a printed contract will not appear in the change list. Check signed copies visually as well.
Recommended workflow for signed copies
Compare the OCR-layered scan of the signed copy against the born-digital version that was sent for signature. Expect a small amount of OCR noise. Anything substantive in the change list is a real edit and should be investigated. Enterprise deployments can include an OCR step for documents that arrive without a text layer.
Try it on your own documents. Upload two versions on the homepage. Documents up to 30 pages need no account, and nothing is stored.