How to check if a PDF is scanned

A scanned PDF holds a picture of each page. Some carry a hidden layer of text added by an OCR program, which makes them look like ordinary PDFs until the text is used. Knowing which pages are scans, how well they were scanned and whether their text can be trusted decides which tool reads them.

Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

Three checks by hand

Try to select a line of text. On a plain scan nothing is selected, or the whole page is selected as one picture. Search for a word you can see: a scan without a text layer finds nothing. Zoom in until the letters are large: printed text stays sharp, and scanned text breaks into pixels and shows the paper’s grain.

The first two checks are fooled by a scan that carries a hidden OCR layer. Its text can be selected and searched, but it is only as good as the program that added it, and that program may not have read your language at all.

What the analyser measures

For each page it asks how much of the page its largest image covers. Measured, scanned pages are covered completely, 1.00, and the pages of a book with figures on them 0.007, so the bar at 0.8 sits in an empty gap between the two. A page with under 20 characters of text has no text layer worth the name, only a page number or a stamp.

The resolution is taken from the largest image on the page, not the first. A scanned letter can carry a 248 DPI letterhead strip over body text scanned at 355 to 397 DPI, and the body is what OCR reads. Below 150 DPI the analyser warns that the scan is too low for reliable OCR, and between 150 and 299 DPI that accuracy will be reduced.

A scan with a little text on it is still a scan

A 20-page Hindi book carried a website address on every page as a watermark, 87 Latin characters in its first pages. That was enough for its text layer to look like English, so it was read as English and came back with no Devanagari at all. A full-page picture with fewer than 200 characters of text over it is now treated as a scan whatever those characters say, and the same book reads as Hindi.

A hidden OCR layer can be in the wrong language

A UPSC Telugu exam paper had been through Acrobat’s Paper Capture, which has no Telugu. Its hidden layer reads ~Lese:i <t9~~so ~~Las where the page has Telugu, and it was long enough to pass for real text. Measured across every scanned page on hand with a hidden layer, a correct English layer has 1 to 8% of its words carrying a symbol no word carries, and another program’s OCR of an Indian script it did not know has 21 to 94%, over 73 pages of three documents. Pages like that are read again from their images, and the paper went from 0 Telugu characters to 6,197.

Why it matters before you convert

OCR is the slow part. Measured on production, a scanned page takes a median of 17 to 21 seconds in Gujarati, Malayalam and Hindi, and a dense Telugu page over a minute. Analyse estimates how many pages of a PDF will be read by OCR and how long that will take. A 158-page Malayalam PDF whose text layer does not read is reported as about an hour of OCR, too long for one conversion, with a suggestion to split it into parts of up to 23 pages.

What the analyser tells you

Scanned, no text layer
Every page is a picture. Use OCR.
No text layer on pages …
Those pages are pictures; the rest have text. PDF to DOCX reads each page the way it needs.
… scans with a text layer over the image
A hidden OCR layer. How accurate it is was not checked; a fresh OCR reading is the way to compare.
… whose text layer, added by another program’s OCR, does not read as text
The hidden layer is not the page’s language. PDF to DOCX reads those pages from the images.
Scanned below 150 DPI
Too low for reliable OCR. Rescan at 300 DPI if you can.

Frequently asked questions

Can a scanned PDF be made searchable?

Yes. OCR Correction reads the scan, lets you check the words it is least sure of, and exports a PDF in which the page stays the scan, with the recognised text laid invisibly over each word, so it can be searched, selected and copied.

What resolution should I scan at?

300 DPI, which is also the resolution pages are rendered at for OCR. Below 150 DPI the analyser warns that OCR will not be reliable.

Why can I select text on my scanned PDF?

Because an OCR program has already added a hidden text layer. It may be accurate, or it may be another language’s guess at yours; the analyser says which.

How long will OCR take?

The median page took 17 to 21 seconds on production in Gujarati, Malayalam and Hindi, and dense pages take longer. Analyse gives an estimate for your file before you convert it.

My PDF is part scan and part text. Is that a problem?

No. PDF to DOCX reads the text pages from their text and the scanned pages by OCR, page by page.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems