Why a Kruti Dev Word file shows Latin letters
Kruti Dev is built on the slots of an English font. Where an English font keeps H it draws half a भ, and where it keeps k it draws the stem that completes it, so भ is typed Hk. Word stores H and k. On a computer with Kruti Dev the page looks right; anywhere else, in a search, in a mail preview or on a phone, it reads Hkkjr for भारत. Changing the font to a Unicode one does not help, because the letters stored are H and k.
The converter goes through the file run by run and asks which face Word draws each run in: the run’s own font, then its character style, the paragraph style and the document defaults, the way Word decides it. A run drawn in Kruti Dev or DevLys has its keys read through the table and every sign put where Unicode wants it: the short i typed before its consonant goes after it, and the reph typed after its letter goes before it. A run in Times, Arial or Calibri is English and is left alone. A word split between two runs, because half of it is bold, is read whole.
What comes back
The same Word file, as .docx. Text in Kruti Dev becomes Unicode Devanagari in Noto Sans Devanagari, set in all four of the run’s font slots, with the language and complex-script properties Word uses for Hindi, so bold and size stay on the Hindi. Everything else is untouched: tables, headers, footers, footnotes, alignment, indents, spacing, pictures, and every run in an English font.
Kruti Dev draws small for its point size, so text kept at the same size would grow: its letters are about 1.22 times as tall in Noto Sans Devanagari. Converted text is scaled to keep its height, so 14 pt Kruti Dev becomes 11.5 pt. Without that, a 21-page form rendered 36 pages; scaled, it renders 23. The note that the mapping is unverified goes into the file’s Comments property, not into your document’s text.
Measured
Ten real government Word files typed in Kruti Dev and DevLys, from nine offices: circulars, orders, forms, letters and a tender, eight of them in the older .doc format. Converted, they hold 21,229 words of five letters or more, and 89.9% of them are in Tesseract’s Hindi or Devanagari word lists, between 86% and 94% file by file. That is not an accuracy figure: names, place names and abbreviations outside the lists can be right. No code was left unread. Pages rendered from the converted files read line for line as the originals. Seven of the ten end within a page of the original length, and three come out shorter (106 pages of 119, 13 of 17, 41 of 53), because letters matched in height take less width in Noto.
Measuring found faults, and each is fixed. Two files were typed wholly on the plain keys, 3,429 and 45,255 characters with nothing above basic ASCII, and were refused before; the converter now goes by the font each run is set in. Four files set headings in Calibri while typing them on Kruti Dev keys, and those are now read when they read as Hindi words, while English labels such as Name of or Loan Agreement are left alone. A colon typed on the visarga key after a digit or an anusvara is read as the colon it is. And one tender set 11 KV Jaw across runs so that KV J stood alone and read as Hindi; a word shared between runs is now judged whole.
Where it still goes wrong
Only Kruti Dev and DevLys files were measured. Other legacy fonts the converter has a table for are converted the same way, unmeasured. Hindi typed on Kruti Dev keys under an English font name is converted only when its reading gives Hindi words, so a very short line in such a font can stay as keys. Of the 79 sign sequences no Hindi word holds, most are in the files themselves: keys pressed twice, and ँ typed as ू followed by the candra key. They are left as typed.
Noto Sans Devanagari is wider or narrower than Kruti Dev letter by letter, so lines break in other places and a page can end a line or two from where it did. A .doc is opened with LibreOffice first and comes back as .docx. Codes in Kruti Dev that nobody has identified come through as characters that are plainly not Devanagari, so they are easy to spot. A scanned document has no keys to read; it needs OCR.