Convert legacy Indian-language fonts to Unicode

Anu, Kruti Dev, Bijoy, Nudi and the other DTP fonts of the last thirty years put an Indian script where an English font keeps its letters. A document set in one reads only while that exact font draws it. IndicPDF reads the codes through a table built for each font family and writes real Unicode, from a PDF, a Word file or a PageMaker publication.

Takes PDF, Word .docx or .doc, PageMaker .pmd files up to 25 MB · Files deleted within 4 hours · No sign-up

What do you have?
Back to Tools
🔤

Legacy PDF to DOCX

A PDF whose text was set in a legacy Indic font, recovered as editable Unicode Word — read from the font’s own table rather than by OCR.

How to use?

Leave it on detect if you are not sure. The ones marked “not yet” will tell you so rather than converting badly.

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

What converting a legacy font means

A legacy font does not change what a file stores, only what it draws. Kruti Dev draws half a भ where an English font draws H and completes it with the stem on k, so a Hindi document typed in it stores H, k and the other keys the typist pressed. Converting it means reading each code, or run of codes, through a table that says which letter the font draws there, and then putting the letters in Unicode order: several of these fonts store a vowel sign before the consonant it follows, or a reph after the syllable it sits on.

The table has to be right for the font, not just for the script. Kruti Dev 010 and Kruti Dev 714 put different letters on the same keys, and three Gujarati books set in Krishna read at 9.8% words through the LMG Arun table and 95.1% through their own.

Why the tables were checked in the real fonts

A table can agree with itself perfectly and still be wrong. Reading back what the same table wrote proves nothing, because a wrong entry is wrong in both directions. So each table was checked against the font it is for: real words were written through it, drawn in the real legacy font, and read back by OCR beside the same words drawn in a Unicode font. That found faults nothing else could. In Kruti Dev, झ sat one position out, so समझ and समझा were read wrong. In Anu, a two-byte entry swallowed the first byte of the next letter, so పదివరాలు read పదివలు.

OCR misreads correct words too, so the numbers below are floors. And every table is still marked unverified: nobody who reads each script has checked every entry.

Fonts recognised but not yet converted

Shivaji and Walkman-Chanakya for Devanagari, Shree-Lipi’s Gujarati and other non-Devanagari faces (its Telugu SHREE-TEL7 is read, so far from one document), and Akruti fonts other than the Marathi and Oriya ones are named by the analyser when a PDF uses them, and there is no table to convert them yet. One Telugu face is not read either: CVTEMeghna. Assamese legacy fonts, Geetanjali and Ramdhenu, have no table. A page with no legacy text at all, a scan, needs OCR instead.

What the result keeps, and what it does not

From a PDF, the Word file carries the recovered text in a Unicode font for its script, with the sizes and bold the PDF states, the book’s figures, and its running head moved into the page header. English and numbers the PDF sets in an ordinary font such as Times or Calibri come across beside it. It is rebuilt page by page, one Word section for each printed page with that page’s size and margins, and ruled tables become Word tables where a grid can hold them. A word processor still breaks lines its own way, so a page can end a line or two from where the PDF’s did. And the document opens with a note that the mapping is unverified, because it is.

The fonts there is a table for

“Read back” means real words written through the table, drawn in the real legacy font and read back by OCR, counted only where the same words drawn in a Unicode font read back too.

ScriptFontsWhat was measured
TeluguAnu family: Priyaanka, Priyaanka Uni, Kranthi, Brahma, and the Pallavi facesA Telugu reader compared 2,475 words of printed lines from twelve books with the conversion, in four rounds: the first three found 11 wrong words, all fixed, and the fourth found none in 610. Scored over 116,750 words of real Telugu books, 75.8% are in Tesseract’s 221,189-word Telugu list; the rest include names, loanwords and joined forms the list does not hold, as well as real misreadings.
Hindi, Marathi, SanskritKruti Dev 010, with Kruti Dev 714 and DevLys 010 as variantsFourteen real Kruti Dev PDFs, government tenders and orders, a pension booklet and a gazette notification: 19,885 words, 88.0% in Tesseract’s Hindi or Devanagari word lists. Fourteen DevLys PDFs of recruitment notices: 17,025 words, 95.2%. One printed line of each was checked against the conversion by eye. A Hindi PageMaker publication reads 122 of its 126 words as its Unicode source.
HindiChanakya (newspapers)Twenty-six PDFs from at least seven publishers, newspaper pages, a party document, a monthly, two exam papers and a Gita commentary: 101,746 words, 94.0% in Tesseract’s Hindi or Devanagari word lists, one printed line of each checked against the conversion. PDFs only.
MarathiAkruti Marathi: AkrutiMar_BYogini, AkrutiMar_BAditi, AkrutiMar_BVijay94.8% of 4,930 words of a nine-page book are in Tesseract’s Marathi list. PDFs only, and one book.
Marathi, HindiShree-Lipi: SHREE-DEV7 and SHREE-DEV-E (8-bit), SHREE-DEVX (read by glyph number)Twelve issues of two Marathi monthly magazines, 636 pages: 90.7% of 160,505 long words are in Tesseract’s Marathi list, one printed line of each checked against the page. One publisher. SHREE-DEVX is measured on one Hindi report, 76% of its long words in the Devanagari list. PDFs only.
BengaliBijoy: SutonnyMJ291 of 322 words of two real documents read back. The misses that were checked by eye were the OCR misreading correct words.
TamilTAB-ELCOT-Kovai67 of the 73 building blocks it spells with are confirmed two independent ways, in the real font. No real TAB-ELCOT-Kovai PDF has been converted yet, so reading one is untested.
TamilTSCII 1.7: TSCInaimathi, TSCMylai, TSCKanna and the TSCu facesBuilt from the published TSCII 1.7 standard and checked on five e-texts of a public Tamil library that publishes each work in TSCII and in Unicode: 99.72% to 99.97% of 478,079 characters match the Unicode edition, and every remaining difference was traced to a line-end word break, the edition’s own losses or the typist’s spelling. PDFs only.
TamilBamini (Bamini Plain, and a misspelt Baamini)Read off the font glyph by glyph. On three real PDFs, a 57-page handbook, a 24-page prospectus and a journal’s pages, 40% to 43% of long words are in Tesseract’s Tamil list, where a correct Unicode edition of Tamil prose scores 27%. Bamini typists spell ர் and ரி with the kaal and ூ as ு plus the kaal, and those are read as meant. Writing for PageMaker: of 745 words drawn in the real font where a Noto control reads back, 702 read back as themselves, and the misses checked by eye were the OCR misreading correct words. Placed in PageMaker 7.0, a test file and a four-page journal arrived byte for byte, 4,396 and 7,942 Bamini bytes, and drew correctly.
KannadaNudi 01 e (and 01 k, which differs only in its digits)217 of 244 words of a test document read back.
GujaratiLMG Arun342 of 373 words of two real documents read back.
GujaratiKrishna, KrishnaBold, KrishnaItalic, Girdhar or Giridhar, ChitraRead off the fonts three Dada Bhagwan books embed: 95.1% of the body’s word occurrences are in Tesseract’s Gujarati list. PDFs only.
MalayalamML-TT Karthika165 of 183 words of a test document read back.
OdiaAkruti Oriya (AkrutiOri -99 faces)290 of 304 words read back.
PunjabiAsees158 of 180 words from Tesseract’s own Punjabi vocabulary read back.

Frequently asked questions

Which legacy fonts can IndicPDF convert?

Anu (Priyaanka, Kranthi, Brahma and the Pallavi faces) for Telugu, Kruti Dev 010, Kruti Dev 714 and DevLys 010 for Hindi, Marathi and Sanskrit, Chanakya for Hindi newspapers, Akruti Marathi, Shree-Lipi (SHREE-DEV7 and SHREE-DEVX) for Marathi and Hindi and SHREE-TEL7 for Telugu, Bijoy (SutonnyMJ) for Bengali, TAB-ELCOT-Kovai, TSCII and Bamini for Tamil, Nudi 01 e for Kannada, LMG Arun and Krishna for Gujarati, ML-TT Karthika for Malayalam, Akruti Oriya, and Asees for Punjabi.

Does the font need to be installed on my computer?

No. The conversion reads the codes in the file, not the drawn page, so it needs neither the font nor print quality. The font matters only in the other direction, when Unicode text is turned back into a legacy font to place in PageMaker: then the operator’s machine must have it.

Can I turn Unicode text back into a legacy font?

Yes, for every table except Akruti Marathi and Krishna, which read only. DOCX to PageMaker-ready writes the font’s own bytes into a file PageMaker places, with English kept in an ordinary face.

How accurate is the conversion?

The table above gives what was measured for each font, and on how much text. None is perfect, and a wrong entry is wrong everywhere it occurs, so read the converted text before relying on it.

My font is not in the list. What now?

Analyse a PDF first: it names the fonts in it and says whether there is a table. If the font is not listed, tell us which one it is. If the pages are scans, OCR reads them whatever font was used to print them.

Why did an ordinary PDF to Word converter give me Latin letters?

A general converter trusts the codes and the PDF’s own map, which for a legacy font point at Latin letters. It writes those into Word and reports success. Only a table for the font turns them into the script.

What happens to my file?

It is stored encrypted while it converts and removed by the cleanup that clears every conversion, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Related tools

Other legacy font pages