Why your Telugu PDF copies out as garbage

The page shows Telugu. Select it, copy, paste, and you get a row of accented Latin letters and symbols. The PDF is not damaged and your computer is not missing a font: the text was set in a legacy font, and the file only knows that font’s codes.

On the page

సమాచార హక్కు చట్టం

Copied and pasted

düe÷#ês¡ Vü≤≈£îÿ #·≥º+

The heading of a Right to Information application set in Kranthi, one of the Anu family of fonts, and the same heading as text extraction reads it out of the file.
Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

What is actually in the file

For years Telugu DTP was done in fonts such as Anu, Priyaanka, Kranthi and Brahma, which put Telugu shapes in the places an English font keeps its letters. The document stores those codes, and the font turns each code into a Telugu shape on the page. A PDF made from it carries the codes and the font together, so it prints and displays perfectly.

What it does not carry is which Telugu letter each code stands for. Copying hands over the codes, and the editor you paste into draws them in an ordinary font, where they are ü, ÷ and ≤. The PDF usually has a map that is meant to turn codes back into characters, called ToUnicode, and the Anu fonts in the application above have one. It maps every code to a Latin character, because that is what the font says it is. A check that only asks whether the map exists passes exactly the files that fail.

Why installing a Telugu font does not help

The damage is in the copied text, not in how it is displayed. Pasting into a document set in Gautami, Nirmala UI or Noto Sans Telugu changes nothing, because the pasted characters really are Latin letters and symbols, not Telugu letters shown in the wrong font. Changing the font of the pasted text does not bring the Telugu back either. The text has to be converted, code by code, into Unicode Telugu.

How to tell for sure

The analyser above lists every font in the PDF and reads the codes each one draws. For the application in the sample it reports “6 pages · Telugu · legacy font detected”, names Brahma, Kranthi, Priyaanka and PriyaankaBold as Anu Telugu, and says text extraction fails for 100% of the text. The English address on the same page is set in Arial and Times and copies out correctly, which is why only some lines break.

It identifies the family from the codes rather than trusting the font name, because names are unreliable. A 154-page Telugu novel names several of its fonts TT1D00O00, TT1D39O00 and TTE2371118O00, and the analyser identifies all three as Anu fonts from what they draw.

Getting the Telugu back

Legacy PDF to DOCX reads each code through a table built for the Anu family and writes real Unicode Telugu into a Word document: the heading above comes back as సమాచార హక్కు చట్టం. It reads the file’s codes rather than the printed page, so print quality does not matter, but the table is not perfect and the result is worth reading before it is relied on. If the analyser says some pages have no text layer, those pages are pictures of text, and Telugu OCR is the tool for them.

What the analyser tells you

Legacy font detected: Anu Telugu
The Telugu was set in an Anu font. Copied text is the font’s codes. Convert it with Legacy PDF to DOCX.
Text extraction fails for most of the text
Names the fonts responsible and the share of the document’s text set in them.
Extracted text reads as Latin
The same finding from the other side: the page shows Telugu and the file’s text is Latin letters.
No text layer on pages …
Those pages are images. Nothing can be copied from them; they need OCR.
This PDF blocks text extraction
The file’s permissions ask readers not to allow copying. The text may be fine; the lock is the problem.

Frequently asked questions

Is my Telugu PDF corrupted?

No. A PDF set in a legacy font is working exactly as it was made to: it draws the page correctly and stores the font’s codes. It was simply never made to give its text back as Unicode.

Why do the English lines copy correctly but the Telugu does not?

The English is set in an ordinary font such as Arial or Times, whose codes are the letters they draw. Only the text set in the legacy Telugu font is affected, which is why a mixed page breaks line by line.

Will an ordinary PDF to Word converter fix it?

Usually not. A general-purpose converter trusts the file’s own map, so it writes the same Latin symbols into Word and reports success. The text has to be read through a table for the legacy font.

Is Priyaanka Uni a Unicode font?

No. Despite the name, Priyaanka Uni is one of the Anu legacy fonts. Gautami, Nirmala UI and Noto Sans Telugu are Unicode fonts, and text set in them copies out as real Telugu.

I still have the PageMaker file. Is that better than the PDF?

Often, yes. A .pmd stores the text as the codes the operator typed, and PageMaker to PDF reads it directly without PageMaker installed.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems