Why a Chanakya PDF copies as symbols
Chanakya, like Kruti Dev, draws Devanagari in the slots of an 8-bit Latin font, but its layout is its own: a full letter is usually a half form plus a stem, the short i is typed before its consonant, and the reph after its syllable. The PDF stores those codes, and the map it carries for copying turns them into symbol characters, so every copy, search and ordinary converter gets ∑§Ë for की.
Some newspapers’ text layers are worse: the letters the font keeps at codes 0x80 to 0x9F are missing from the copied text altogether, so मुख्यमंत्री loses its ख and गिरफ्तार its फ. The converter reads the codes themselves through a Chanakya table, drawn and checked in the real font, and puts every sign where Unicode wants it.
One name, three builds
Documents set in Chanakya do not all put the same thing on every key. Some newspapers put the digits 1 to 9 on keys that other builds use for half letters, and one exam paper put क् and ख् on keys of their own. The converter reads each document every way these builds allow and keeps the reading that gives words, so a page of prices does not come out as strings of half letters. Nothing has to be chosen.
Walkman-Chanakya, a different font with a similar name, has a layout of its own and is not read yet. Headlines a newspaper sets in some other face are not read either: only the text set in Chanakya is converted.
Measured
Twenty-six real PDFs set in Chanakya, from at least seven publishers: newspaper pages, a party document and a spiritual monthly typeset in PageMaker, two school-board exam papers made in Word, and a Gita commentary. Converted, they hold 101,746 words, and 94.0% of them are in Tesseract’s Hindi or Devanagari word lists: between 93% and 98% document by document, and 79.5% for the Gita commentary, whose Sanskrit verse and compounds the lists do not hold. That is not an accuracy figure; names and rare words outside the lists can be right. No code was left unread, and one printed line of each document was checked against the conversion by eye.
Measuring found faults, and each is fixed. A glyph read as the ra-foot is the subscript na, so अग्नि had come out अग्रि and प्रश्न प्रश्र; the ya-sign used after letters with no stem was not read, so पाठ्यक्रम and बुद्ध्या lost it; a candrabindu typed before ू stood before it; and on pages printing times and prices, digits had been read as half letters.
Where it still goes wrong
Of the 101,746 words, 17 hold a sign sequence no Hindi word has, mostly keys typed twice. A few codes in the font draw glyphs nobody has identified; they come through as characters that are plainly not Devanagari, so they are easy to spot. A newspaper page of boxed stories is the hardest layout there is, so check the order of the stories in the Word file.
Text in Walkman-Chanakya or in a headline face other than Chanakya stays as it was. A scanned page has no codes to read and needs OCR instead.