Why copying a Kruti Dev PDF gives Latin letters
Kruti Dev is built on the slots of an English font. Where an English font keeps H it draws half a भ, and where it keeps k it draws the stem that completes it, so भ is typed Hk. The PDF stores H and k, and its own map from codes to letters says H and k, because that is what those slots are in any other font. Every copy, every search and every ordinary PDF-to-Word converter reads that map, and gets Hkkjr for भारत.
The page looks right only because the PDF carries the Kruti Dev font that draws those codes. The converter does not use the map. It reads the codes themselves through a table built for Kruti Dev 010 and checked in the real font, then puts every sign where Unicode wants it: the short i typed before its consonant goes after it, and the reph typed after its letter goes before it.
What comes back in the Word file
Unicode Devanagari set in Noto Sans Devanagari, so it opens on any computer without Kruti Dev. The document is rebuilt page by page: one Word section for each printed page, with the page’s size and margins, and paragraphs with their alignment, indents, type sizes and bold. Ruled tables become Word tables where a grid can hold them, and a dense form comes out longer than the printed one. Pictures that are figures on a page are placed on it, and a running head or page number goes into the Word header. English and numbers the PDF sets in an ordinary font such as Times or Calibri come across as they are.
A word processor still breaks lines its own way, so a page can end a line or two from where the PDF’s did. The Word file is for editing and reusing the text, and it opens with a note that the mapping is unverified, because it is.
Measured
Fourteen real PDFs typed in Kruti Dev: government tenders and orders, a pension booklet, and a 103-page gazette notification. Converted, they hold 19,885 words, and 88.0% of them are in Tesseract’s Hindi or Devanagari word lists, between 80% and 97% document by document. That is not an accuracy figure: the words outside the lists include names, place names and abbreviations that are right. One printed line of every document was checked against the conversion by eye, and all fourteen matched.
Measuring found faults, and each is fixed: a reph after a letter the stem completes was dropped, so वार्मिंग read वाम्िांग; an anusvara typed before its vowel sign stayed there; and a time typed 03%00 read ०३ः००. What is left is mostly in the files: of the 40 sign sequences no Hindi word holds, 29 are one tender where the typist pressed the nukta key after a place name down a table column, and the rest are doubled keys and stray signs.
Where it still goes wrong
Dozens of codes in Kruti Dev draw glyphs nobody has identified, mostly rare conjuncts; they come through as characters that are plainly not Devanagari, so they are easy to spot. A dense application form comes out longer than the printed one, because a word processor wraps the text in its boxes its own way. One booklet sets 1,651 characters of Kruti Dev in a font with no name at all, and those come through as the typed keys. A letter lost before the file reached you cannot be recovered.
A scanned PDF has no codes to read, only pictures of letters; it needs OCR instead. And some PDFs do not include the Kruti Dev font on every page, so a viewer shows those pages in Latin letters. They still convert, because the codes are there.