Cause one: a legacy font such as Kruti Dev
Kruti Dev, DevLys and Chanakya draw Devanagari in the places an English font keeps its letters, so a PDF made from a Kruti Dev document stores those letters. Pasted, they are what the typist pressed: भारत is Hkkjr. The file’s own map, when it has one, points back to the same Latin letters, so it agrees with the garbage.
The analyser names the font, “Legacy font detected: Kruti Dev 010” for the line above, and Legacy PDF to DOCX converts it to Unicode Hindi. DevLys is converted the same way, as a build of Kruti Dev, and Chanakya, the newspaper layout, has a table of its own.
Cause two: the pages are scans
A scanned PDF holds a picture of each page and no text at all, so selecting does nothing or selects the whole page as one image. Nothing is broken and nothing can be extracted: the text has to be read from the picture. For a two-page scanned government agenda the analyser reports “2 pages · scanned, no text layer” and that the pages were scanned at 150–299 DPI, where OCR accuracy is reduced. Hindi OCR reads such pages, and 300 DPI scans read best.
Some scans carry a hidden text layer added by an earlier OCR pass, which is why a scan can sometimes be selected. The analyser reports those as scans with a text layer, and says that nobody has checked how accurate that layer is.
Cause three: a Unicode font with no map back to text
A PDF can use a proper Unicode Hindi font and still paste as garbage. When a font is embedded by glyph number (an Identity encoding) and without its ToUnicode map, the page draws correctly, because drawing needs only glyph numbers, but extraction has nothing to turn them back into letters. It returns boxes, (cid:12) markers or unrelated characters: in a test file built this way the word Identity copied out as *EFOUJUZ.
The analyser lists such a font with Unicode mapping “None” and names it in the extraction warning. OCR is the dependable fix, because it reads the page as it is drawn.
Cause four: the PDF forbids copying
Whoever made a PDF can set permissions that ask readers not to allow copying, printing or editing. Many readers obey, so copy is greyed out or pastes nothing while the text inside is perfectly good. The analyser reports “This PDF blocks text extraction” and shows which permissions are set. If a password is needed just to open the file, nothing inside it can be inspected until it is opened with that password.