How Anu stores Telugu
An Anu font builds a letter from parts. A consonant is a body glyph followed by its own talakattu, the tick on top, and a vowel sign is typed in place of the talakattu, not after it. Subscript consonants are separate glyphs typed after the syllable, and the ra-foot is typed before the letter it sits under. So one Telugu syllable is two to five codes, in the order the typist pressed them, and none of those codes is a Telugu character.
A PDF made from such a file usually carries a map from codes to characters, and the map is wrong: it names the Latin and symbol characters those codes would be in an English font. That is why the sample above copies out the way it does, and why an ordinary PDF to Word conversion returns the same Latin letters and calls it done. The converter reads the codes themselves, joins the parts into syllables and puts them in Unicode order.
One layout for the whole family, and a second one
Twelve test words were drawn in each of the 86 Telugu faces Anu Script Manager 7 installs, beside a Unicode font. Every face draws the same letters from the same codes, so one table reads Priyaanka, Gowthami, Anupama, Pallavi, Brahma, Kranthi and the rest, whatever the weight.
Some books do not follow it. One book embeds a Priyaanka with the same glyphs at different codes: read with the usual table, 1.7% of its words were real Telugu. Its own font outlines were drawn and matched against the installed face, glyph by glyph, which gave a second layout that reads the same book at 73.8%. The converter does not trust the font’s name for this. It reads a face’s text through both layouts and keeps the one that produces fewer words Telugu cannot spell.
Faults found on real books
Most wrong words came from one kind of fault: a table entry two codes long that swallowed the first code of the next letter. పదివరాలు read పదివలు, because the entry for వ took the body of ర. అబ్బాయ్ read అబ్బాయు, because the pollu before య’s tail was read as nothing. నర్సింహం read నర్సింహాం, because an entry hid the tail that completes హ. Each was settled by cropping the printed word from the page, or by drawing the codes in the real Priyaanka font, and each fix was measured over every Anu book on hand before it was kept.
The other large fault was not in the table. A PDF gives a page’s text as one long run, and the converter was joining lines with no space, so the last word of each line was welded to the first of the next: 54 welded words on ten pages of one novel. Lines are now cut where the print cuts them.
Measured
A Telugu reader compared printed lines with the conversion, word by word: 288 lines, 2,475 words, drawn at random from twelve Anu-set books and documents, in four rounds. The first three rounds found 2, 7 and 2 wrong words, and every one was traced to a table entry and fixed. The fourth round, 610 words from lines not seen before, found none. The reader marked six words in it, and each turned out to be printed that way: the book’s own misspelling, which the conversion keeps.
Those lines are set mostly in three faces, Priyaanka (1,428 of the words), GowthamiMedium (543) and AnupamaMedium (428). Faces the sample barely reached, such as PriyaankaBold with 62 words, are read by the same table but have had far less checking.
Over whole books, 116,750 converted words, 75.8% are in Tesseract’s 221,189-word Telugu list. The words outside it are mostly names, loanwords and joined forms the list does not hold, so this number shows movement between versions, not an error rate.
Where it still goes wrong
The reader’s check sampled lines, it did not read every word, and a wrong table entry is wrong everywhere it occurs. So read the result before relying on it, most of all for names and numbers. One Telugu face is not read at all, CVTEMeghna, and a book set in it is refused with a message saying so. A few subset fonts keep no usable codes, and a document set mostly in one is refused too.
The Word file is rebuilt page by page, with the print’s paragraphs, sizes, bold, pictures and running heads. Unicode Telugu fonts set wider than Anu, so the text is set at the size where its letters stand as tall as the print’s, and a long book can still come out with more pages than the print has. A scanned book has no codes to read and needs Telugu OCR.