What converting a legacy font means
A legacy font does not change what a file stores, only what it draws. Kruti Dev draws half a भ where an English font draws H and completes it with the stem on k, so a Hindi document typed in it stores H, k and the other keys the typist pressed. Converting it means reading each code, or run of codes, through a table that says which letter the font draws there, and then putting the letters in Unicode order: several of these fonts store a vowel sign before the consonant it follows, or a reph after the syllable it sits on.
The table has to be right for the font, not just for the script. Kruti Dev 010 and Kruti Dev 714 put different letters on the same keys, and three Gujarati books set in Krishna read at 9.8% words through the LMG Arun table and 95.1% through their own.
Why the tables were checked in the real fonts
A table can agree with itself perfectly and still be wrong. Reading back what the same table wrote proves nothing, because a wrong entry is wrong in both directions. So each table was checked against the font it is for: real words were written through it, drawn in the real legacy font, and read back by OCR beside the same words drawn in a Unicode font. That found faults nothing else could. In Kruti Dev, झ sat one position out, so समझ and समझा were read wrong. In Anu, a two-byte entry swallowed the first byte of the next letter, so పదివరాలు read పదివలు.
OCR misreads correct words too, so the numbers below are floors. And every table is still marked unverified: nobody who reads each script has checked every entry.
Fonts recognised but not yet converted
Shivaji and Walkman-Chanakya for Devanagari, Shree-Lipi’s Gujarati and other non-Devanagari faces (its Telugu SHREE-TEL7 is read, so far from one document), and Akruti fonts other than the Marathi and Oriya ones are named by the analyser when a PDF uses them, and there is no table to convert them yet. One Telugu face is not read either: CVTEMeghna. Assamese legacy fonts, Geetanjali and Ramdhenu, have no table. A page with no legacy text at all, a scan, needs OCR instead.
What the result keeps, and what it does not
From a PDF, the Word file carries the recovered text in a Unicode font for its script, with the sizes and bold the PDF states, the book’s figures, and its running head moved into the page header. English and numbers the PDF sets in an ordinary font such as Times or Calibri come across beside it. It is rebuilt page by page, one Word section for each printed page with that page’s size and margins, and ruled tables become Word tables where a grid can hold them. A word processor still breaks lines its own way, so a page can end a line or two from where the PDF’s did. And the document opens with a note that the mapping is unverified, because it is.