Anu to Unicode: convert Telugu set in Anu fonts

Anu Script Manager is what Telugu books, newspapers and DTP shops have been set in for decades: Priyaanka, Gowthami, Anupama, Pallavi and their relatives. A file set in them stores the font’s own codes, so its text copies out as accented Latin letters and no search finds a Telugu word in it. Upload the PDF or the PageMaker publication and get the whole document back in Unicode Telugu, read through a table built for the Anu layout and checked line by line against the printed page by a Telugu reader.

Takes PDF, PageMaker .pmd files up to 25 MB · Files deleted within 4 hours · No sign-up

ఈ పేజీని తెలుగులో చదవండి

The PDF’s text, as PyMuPDF reads it

á düe÷#ês¡+q≈£î dü+ã+~Û+∫ s¡TdüTeTT @yÓTÆHê e]Ô+∫q#√

Read as Unicode

ఈ సమాచారంనకు సంబంధించి రుసుము ఏమైనా వర్తించినచో

A line of a Telugu petition set in Priyaanka. The left side is the text the PDF itself carries; the right is the Anu table’s reading, which is what the printed page shows.
What do you have?
Back to Tools
🔤

Legacy PDF to DOCX

A PDF whose text was set in a legacy Indic font, recovered as editable Unicode Word — read from the font’s own table rather than by OCR.

How to use?

Leave it on detect if you are not sure. The ones marked “not yet” will tell you so rather than converting badly.

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

How Anu stores Telugu

An Anu font builds a letter from parts. A consonant is a body glyph followed by its own talakattu, the tick on top, and a vowel sign is typed in place of the talakattu, not after it. Subscript consonants are separate glyphs typed after the syllable, and the ra-foot is typed before the letter it sits under. So one Telugu syllable is two to five codes, in the order the typist pressed them, and none of those codes is a Telugu character.

A PDF made from such a file usually carries a map from codes to characters, and the map is wrong: it names the Latin and symbol characters those codes would be in an English font. That is why the sample above copies out the way it does, and why an ordinary PDF to Word conversion returns the same Latin letters and calls it done. The converter reads the codes themselves, joins the parts into syllables and puts them in Unicode order.

One layout for the whole family, and a second one

Twelve test words were drawn in each of the 86 Telugu faces Anu Script Manager 7 installs, beside a Unicode font. Every face draws the same letters from the same codes, so one table reads Priyaanka, Gowthami, Anupama, Pallavi, Brahma, Kranthi and the rest, whatever the weight.

Some books do not follow it. One book embeds a Priyaanka with the same glyphs at different codes: read with the usual table, 1.7% of its words were real Telugu. Its own font outlines were drawn and matched against the installed face, glyph by glyph, which gave a second layout that reads the same book at 73.8%. The converter does not trust the font’s name for this. It reads a face’s text through both layouts and keeps the one that produces fewer words Telugu cannot spell.

Faults found on real books

Most wrong words came from one kind of fault: a table entry two codes long that swallowed the first code of the next letter. పదివరాలు read పదివలు, because the entry for వ took the body of ర. అబ్బాయ్ read అబ్బాయు, because the pollu before య’s tail was read as nothing. నర్సింహం read నర్సింహాం, because an entry hid the tail that completes హ. Each was settled by cropping the printed word from the page, or by drawing the codes in the real Priyaanka font, and each fix was measured over every Anu book on hand before it was kept.

The other large fault was not in the table. A PDF gives a page’s text as one long run, and the converter was joining lines with no space, so the last word of each line was welded to the first of the next: 54 welded words on ten pages of one novel. Lines are now cut where the print cuts them.

Measured

A Telugu reader compared printed lines with the conversion, word by word: 288 lines, 2,475 words, drawn at random from twelve Anu-set books and documents, in four rounds. The first three rounds found 2, 7 and 2 wrong words, and every one was traced to a table entry and fixed. The fourth round, 610 words from lines not seen before, found none. The reader marked six words in it, and each turned out to be printed that way: the book’s own misspelling, which the conversion keeps.

Those lines are set mostly in three faces, Priyaanka (1,428 of the words), GowthamiMedium (543) and AnupamaMedium (428). Faces the sample barely reached, such as PriyaankaBold with 62 words, are read by the same table but have had far less checking.

Over whole books, 116,750 converted words, 75.8% are in Tesseract’s 221,189-word Telugu list. The words outside it are mostly names, loanwords and joined forms the list does not hold, so this number shows movement between versions, not an error rate.

Where it still goes wrong

The reader’s check sampled lines, it did not read every word, and a wrong table entry is wrong everywhere it occurs. So read the result before relying on it, most of all for names and numbers. One Telugu face is not read at all, CVTEMeghna, and a book set in it is refused with a message saying so. A few subset fonts keep no usable codes, and a document set mostly in one is refused too.

The Word file is rebuilt page by page, with the print’s paragraphs, sizes, bold, pictures and running heads. Unicode Telugu fonts set wider than Anu, so the text is set at the size where its letters stand as tall as the print’s, and a long book can still come out with more pages than the print has. A scanned book has no codes to read and needs Telugu OCR.

Frequently asked questions

Which Anu fonts does it read?

The Telugu faces of Anu Script Manager: Priyaanka, Priyaanka Uni, the Gowthami, Anupama, Pallavi and Pragathi weights, Brahma, Kranthi and the others, which all share one layout. A Telugu reader has checked lines set in Priyaanka, GowthamiMedium and AnupamaMedium; the rest are read by the same table with less checking. CVTEMeghna is not read.

Anu font ni Unicode loki ela marchali?

Upload the PDF or the PageMaker file above and download the Word file it gives back. The whole document is converted at once, and the Telugu comes back in Unicode.

Do I need Anu Script Manager or the font installed?

No. The converter reads the codes the file stores, and the Word file is in Unicode, set in Noto Sans Telugu. Neither you nor the person you send it to needs an Anu font.

Can it read a PageMaker file?

Yes. A .pmd publication is read from PageMaker’s own records, so paragraphs set in English faces come across beside the Telugu, and the Word file keeps the publication’s sizes, bold, alignment and indents. A typist’s 30-page publication of petitions is kept as a fixed test for this.

What about a Word file typed in Anu?

The Legacy tool reads a Word file whose runs name an Anu face. No real typist’s Anu Word file has been measured the way the PDFs were, so this page offers PDFs and PageMaker files, and a Word result needs checking.

Can I turn Unicode Telugu back into Anu for PageMaker?

Yes. DOCX to PageMaker-ready writes the font’s own codes into a file PageMaker places, set in Priyaanka Uni, with English kept in an ordinary face. Drawn in the real font, 803 of 871 test words read back as themselves. ృ is drawn with the nearest glyph PageMaker keeps, and ఋ is written the way Telugu typists write it in this font.

The book has a spelling mistake. Does the converter fix it?

No. It reproduces what is printed, mistakes included. In the reader’s last check, all six words marked were the books’ own misspellings.

What happens to my file?

It is stored encrypted while it converts and removed by the cleanup that clears every conversion, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Related tools

Other legacy font pages