How Kruti Dev stores Hindi
Kruti Dev follows the Remington typewriter layout. Most consonants are typed as a half form and completed by the vertical stem on k, so भ is Hk. The short i is typed before the consonant it follows, which is how कविता becomes dfork, and the reph is typed after the letter it sits over, so कर्म is deZ. The converter reads the keys and puts each sign back where Unicode wants it.
The layout leaves no room for punctuation. The comma key draws ए and the full-stop key draws ण्, so a typist who wants a comma switches to another font for it, as the typist of the publication above did, and the danda is on A. Converted text follows the font, not the key: a comma typed inside a Kruti Dev run is ए.
A letter one position out
Drawn in the real font, 32 of Kruti Dev’s 33 consonants matched their Unicode shapes and झ did not. The key > already draws a whole झ, with its stem, and the table had read it as half a letter waiting for one. So a typist’s समझ came out as समझ् and समझा lost its ा. Written through the table and drawn in the font, six test words holding झ went from one right to six. It hid for a long time because the words it broke look like typing mistakes rather than a mapping fault.
A report that श and ष were swapped was checked the same way and was false: the font draws श and ष exactly where the table reads them. OCR confuses those two letters, which is why a check made by reading output back suggested a swap that is not there.
Measured
Fourteen real PDFs typed in Kruti Dev, from government tenders to a 103-page gazette notification, hold 19,885 words once converted, and 88.0% of them are in Tesseract’s Hindi or Devanagari word lists; fourteen DevLys PDFs hold 17,025 words at 95.2%. The words outside the lists include names, places and abbreviations that are right, so this is not an accuracy figure. One printed line of every document was checked against the conversion by eye, and all matched.
A Hindi PageMaker publication reads 122 of its 126 words exactly as its Unicode source. Eleven of its letters had been turned into curly quotes by Word or PageMaker, because श् and ष् sit on the quote keys; the file still says which key was pressed, so they are read back as the letters they were. The four words left are damage no key explains: three nuktas stored as ऽ, and quotation marks saved as question marks.
Where it still goes wrong
Kruti Dev 010 has a table, and Kruti Dev 714 is read as a variant of it: drawn side by side the two fonts put the same letters on the same keys except the digits, which 714 draws in Devanagari and 010 in Western form, and the codes for श् and ष् that Word’s curly quotes produce, which the two draw the other way round. A file is read as the face it names draws, and where a document was typed in a third build of 010 that swaps those two letters, the reading holding more real words is chosen. DevLys 010 is that swapped build under another name, and is converted; Chanakya has a table of its own, for PDFs. Dozens of codes in the font draw glyphs nobody has identified yet, mostly rare conjuncts: they come through as characters that are plainly not Devanagari, not as plausible wrong letters, so they are easy to spot. And a letter lost before the file reached you, such as a nukta saved as ऽ, cannot be recovered; the curly quotes above can be, and are.