Hindi text not extracting from a PDF: the four causes

When Hindi will not come out of a PDF, it fails in one of four ways, and each has a different fix. What you get when you paste, whether Latin letters, boxes or nothing, is the first clue. The analyser above tells you which one you have.

On the page

भारत में हिन्दी भाषा बोली जाती है

Copied and pasted

Hkkjr esa fgUnh Hkk"kk cksyh tkrh gS

A line set in Kruti Dev 010 and saved as a PDF, then copied out of it. The pasted letters are the keys a typist presses on the Kruti Dev keyboard.
Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

Cause one: a legacy font such as Kruti Dev

Kruti Dev, DevLys and Chanakya draw Devanagari in the places an English font keeps its letters, so a PDF made from a Kruti Dev document stores those letters. Pasted, they are what the typist pressed: भारत is Hkkjr. The file’s own map, when it has one, points back to the same Latin letters, so it agrees with the garbage.

The analyser names the font, “Legacy font detected: Kruti Dev 010” for the line above, and Legacy PDF to DOCX converts it to Unicode Hindi. DevLys is converted the same way, as a build of Kruti Dev, and Chanakya, the newspaper layout, has a table of its own.

Cause two: the pages are scans

A scanned PDF holds a picture of each page and no text at all, so selecting does nothing or selects the whole page as one image. Nothing is broken and nothing can be extracted: the text has to be read from the picture. For a two-page scanned government agenda the analyser reports “2 pages · scanned, no text layer” and that the pages were scanned at 150–299 DPI, where OCR accuracy is reduced. Hindi OCR reads such pages, and 300 DPI scans read best.

Some scans carry a hidden text layer added by an earlier OCR pass, which is why a scan can sometimes be selected. The analyser reports those as scans with a text layer, and says that nobody has checked how accurate that layer is.

Cause three: a Unicode font with no map back to text

A PDF can use a proper Unicode Hindi font and still paste as garbage. When a font is embedded by glyph number (an Identity encoding) and without its ToUnicode map, the page draws correctly, because drawing needs only glyph numbers, but extraction has nothing to turn them back into letters. It returns boxes, (cid:12) markers or unrelated characters: in a test file built this way the word Identity copied out as *EFOUJUZ.

The analyser lists such a font with Unicode mapping “None” and names it in the extraction warning. OCR is the dependable fix, because it reads the page as it is drawn.

Cause four: the PDF forbids copying

Whoever made a PDF can set permissions that ask readers not to allow copying, printing or editing. Many readers obey, so copy is greyed out or pastes nothing while the text inside is perfectly good. The analyser reports “This PDF blocks text extraction” and shows which permissions are set. If a password is needed just to open the file, nothing inside it can be inspected until it is opened with that password.

What the analyser tells you

Legacy font detected: Kruti Dev 010
Cause one. Convert with Legacy PDF to DOCX.
Scanned, no text layer
Cause two. The pages are pictures; use Hindi OCR.
Unicode mapping: None
Cause three. The named font has no map back to characters; OCR reads the page instead.
This PDF blocks text extraction
Cause four. The permissions forbid copying.
Text extraction is reliable
None of the four. The file’s text should copy correctly; if one PDF reader still mangles it, try another.

Frequently asked questions

Why does my Hindi paste as Hkkjr instead of भारत?

The text was set in Kruti Dev, which draws भारत from the keys H, k, k, j and r. The PDF stores the keys, so that is what copies out. Converting it through a Kruti Dev table gives Unicode Hindi back.

Do Marathi and Sanskrit PDFs have the same problems?

Yes. They are written in Devanagari and set in the same fonts, Kruti Dev included, so all four causes apply exactly as they do to Hindi.

My PDF has Hindi and English, and only the Hindi breaks. Why?

The English is set in an ordinary font whose codes are the letters it draws. Causes one and three affect only the text set in the problem font, so a mixed page breaks line by line.

Should I use Hindi PDF to Word or Legacy PDF to DOCX?

For a scan or a PDF set in a Unicode font, Hindi PDF to Word. For a PDF set in Kruti Dev, Legacy PDF to DOCX, which reads the text through the font’s table.

Can I copy Kruti Dev text and fix it later?

Only by converting it. Pasted Kruti Dev text stays a string of Latin letters that no search, spell check or screen reader understands until it is converted to Unicode.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems