How to check if a PDF uses a legacy font

A PDF set in a legacy Indian-language font looks normal and behaves wrongly: its text copies, searches and converts as nonsense. Checking first tells you which tool will work, and saves converting it with one that will not.

Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

What makes a font a legacy font

Unicode gives every letter of every Indian script its own number, so ఆ or अ is the same character whatever font draws it. A legacy font does not use those numbers. It borrows the numbers of English letters and punctuation and draws its own shapes there, so a document set in it is only readable while that exact font is drawing it.

These fonts were the standard for Indian-language DTP for a long time, so newspapers, government forms, books and court papers are full of them. The ones met most often are Anu, Priyaanka and Kranthi for Telugu; Kruti Dev, DevLys and Chanakya for Hindi and Marathi; SutonnyMJ, the Bijoy font, for Bengali; TAB-ELCOT-Kovai and Bamini for Tamil; Nudi for Kannada; LMG Arun for Gujarati; and Shree-Lipi across several scripts.

Three ways to check by hand

Copy a line and paste it into Notepad. Latin letters and symbols where the page shows an Indian script is the clearest sign there is. Search the PDF for a word you can see on the page: a legacy PDF will not find it, because the word is not in the file as text. Or open the font list, which Adobe Acrobat Reader shows under File › Properties › Fonts and the pdffonts command prints on the command line, and look for a name such as Kruti Dev 010 or Priyaanka.

The font list is the weakest of the three. Fonts are often embedded under names that mean nothing: one 154-page Telugu novel carries Anu fonts named TT1D00O00 and TTE2371118O00.

What the analyser checks instead

It reads the codes each font draws against the conversion tables IndicPDF has, for Anu, Kruti Dev, Akruti Marathi, Shree-Lipi, Bijoy, TAB-ELCOT-Kovai, TSCII, Bamini, Nudi, LMG Arun, ML-TT Karthika, Akruti Oriya and Asees, and names the family when the codes read as that script. Where the codes cannot be read but the font name matches a table, it accepts the name only if the font’s own text is not already Unicode Indian script, so a Unicode document is not flagged because of what its font is called. DevLys and Chanakya are read too, DevLys as a build of Kruti Dev. It also recognises Shivaji, Akruti (other than Akruti Oriya and Akruti Marathi) and Shree-Lipi by name, and says plainly that there is no conversion table for them yet.

What it cannot see

A legacy font none of the tables has met, embedded under a meaningless name, can pass without being named. If the analyser names no legacy font but its scripts panel says the text is almost all Latin while you can see Telugu, Hindi or Tamil on the page, trust your eyes: that combination is a legacy encoding. Nor can a font say which language a Devanagari document is in, because Kruti Dev is used for Hindi, Marathi and Sanskrit alike.

Why check before converting

The wrong tool fails quietly. A general-purpose PDF to Word converter trusts the file’s map, writes the same symbols into Word and reports success; nobody notices until someone tries to edit or search the result. A legacy PDF needs Legacy PDF to DOCX, a scan needs OCR, and a locked PDF needs its owner’s permission, and the analyser says which of those you have in one pass.

What the analyser tells you

Legacy font detected: Kruti Dev 010
There is a table for this font. Legacy PDF to DOCX converts the text to Unicode.
… and IndicPDF has no table for this font yet
The font was recognised by name, but converting it would be guesswork. Tell us which font it is.
Unicode mapping: Legacy encoding
Shown for each font in the fonts table, with the family it belongs to.
Extracted text reads as Latin
The file’s text is Latin letters while its fonts draw an Indian script.
Unicode mapping: None
Not a legacy font, but its text still cannot be copied reliably: the font has no map back to characters.

Frequently asked questions

Which legacy fonts can IndicPDF convert?

Anu, Kruti Dev, Akruti Marathi, Shree-Lipi, Bijoy, TAB-ELCOT-Kovai, TSCII, Bamini, Nudi, LMG Arun, ML-TT Karthika, Akruti Oriya and Asees. Each table was checked by drawing text in the real font, but rare characters are not all covered, so read the converted text before relying on it.

Is a font with Uni or Unicode in its name always a Unicode font?

No. Priyaanka Uni is one of the Anu legacy fonts. Only the analyser’s reading of the codes, or a copy-and-paste test, settles it.

Can one PDF use legacy and Unicode fonts together?

Yes, and it is common. A Telugu application may set its Telugu in Anu fonts and its English address in Arial and Times. Only the text in the legacy font is affected, and the analyser lists each font separately.

Can I check a Word file instead of a PDF?

The analyser reads PDFs. Legacy DOCX to PDF reads a Word file typed in a legacy font and rebuilds it as a Unicode PDF.

Does a legacy font make OCR necessary?

No. The codes are in the file, so reading them through the font’s table is more accurate than reading the printed page. OCR is for pages that have no text at all.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems