Three checks by hand
Try to select a line of text. On a plain scan nothing is selected, or the whole page is selected as one picture. Search for a word you can see: a scan without a text layer finds nothing. Zoom in until the letters are large: printed text stays sharp, and scanned text breaks into pixels and shows the paper’s grain.
The first two checks are fooled by a scan that carries a hidden OCR layer. Its text can be selected and searched, but it is only as good as the program that added it, and that program may not have read your language at all.
What the analyser measures
For each page it asks how much of the page its largest image covers. Measured, scanned pages are covered completely, 1.00, and the pages of a book with figures on them 0.007, so the bar at 0.8 sits in an empty gap between the two. A page with under 20 characters of text has no text layer worth the name, only a page number or a stamp.
The resolution is taken from the largest image on the page, not the first. A scanned letter can carry a 248 DPI letterhead strip over body text scanned at 355 to 397 DPI, and the body is what OCR reads. Below 150 DPI the analyser warns that the scan is too low for reliable OCR, and between 150 and 299 DPI that accuracy will be reduced.
A scan with a little text on it is still a scan
A 20-page Hindi book carried a website address on every page as a watermark, 87 Latin characters in its first pages. That was enough for its text layer to look like English, so it was read as English and came back with no Devanagari at all. A full-page picture with fewer than 200 characters of text over it is now treated as a scan whatever those characters say, and the same book reads as Hindi.
A hidden OCR layer can be in the wrong language
A UPSC Telugu exam paper had been through Acrobat’s Paper Capture, which has no Telugu. Its hidden layer reads ~Lese:i <t9~~so ~~Las where the page has Telugu, and it was long enough to pass for real text. Measured across every scanned page on hand with a hidden layer, a correct English layer has 1 to 8% of its words carrying a symbol no word carries, and another program’s OCR of an Indian script it did not know has 21 to 94%, over 73 pages of three documents. Pages like that are read again from their images, and the paper went from 0 Telugu characters to 6,197.
Why it matters before you convert
OCR is the slow part. Measured on production, a scanned page takes a median of 17 to 21 seconds in Gujarati, Malayalam and Hindi, and a dense Telugu page over a minute. Analyse estimates how many pages of a PDF will be read by OCR and how long that will take. A 158-page Malayalam PDF whose text layer does not read is reported as about an hour of OCR, too long for one conversion, with a suggestion to split it into parts of up to 23 pages.