Why your Gujarati PDF copies out as garbage

The page shows Gujarati. The pasted text is symbols and accented letters, or Gujarati with its conjuncts missing. Either the text was set in a legacy font, or the font’s map back to letters has gaps in it. The analyser above tells you which, and the fix differs.

On the page

સ્વાર્થ છે અને સ્વાર્થ રહ્યો ત્યાં પ્રેમ હોય નહીં.

The PDF’s text

V‰Î◊˝ »ı ±fiı V‰Î◊˝ flè΢ I›Î_ ≠ı‹ ˢ› fiËŸ.

A line of Dada Bhagwan’s પ્રેમ, set in the legacy Krishna font in PageMaker 7.0: the PDF’s text as PyMuPDF and Poppler both read it, and the same line read through the Krishna table.
Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

Cause one: a legacy font such as Krishna or LMG Arun

Gujarati DTP used fonts that put Gujarati shapes where an English font keeps its letters, so the file stores those codes and pastes them as symbols. LMG Arun is the most downloaded of them on the one font site that counts every face, and Legacy PDF to DOCX reads it through a table built from the font: 342 of 373 words from two real documents, written through the table and drawn in the real font, read back correctly.

Three Dada Bhagwan books, પ્રેમ, અહિંસા and વાણી વ્યવહારમાં, are set in Krishna, a face no table covered and no font file on hand could explain. Their table was read off the fonts the PDFs themselves embed: every code drawn from its own outline, and settled against every word of the three books that holds it. Through it, 95.1% of the body’s words are in Tesseract’s Gujarati word list; through the LMG Arun table, the same bytes give 9.8%. One code shows why this had to be done by drawing: 0xAD is પ્ર, and it is also the soft hyphen, which a text renderer hides, so it first looked blank.

Cause two: a Unicode font whose map leaves out the conjuncts

A 727-page ebook, જંગલી પશ્ચિમમાં સાહસો, is set in Arial Unicode MS, a Unicode font, and in its first hundred pages the PDF’s map leaves out 84 of the conjuncts and half forms it draws, 8,958 times. Extracted, ક્રિયા came out as ἵṀયા in one reader and યા in another. PDF to DOCX reads those glyphs from the font’s own shaping rules. On pages 11 to 60 the share of words in Tesseract’s Gujarati list went from 78.6% to 95.2%, and 4,769 (cid markers went to none. What is left is the book’s own: it is a machine translation that leaves English inside some words.

Cause three: a scan

A scanned page holds a picture of the text and nothing to copy, or a hidden layer added by an OCR program. Gujarati OCR reads the picture. One thing to check afterwards is dates and numbers: ર and the digit ૨ are drawn almost alike, and the Gujarati OCR page explains how that is handled.

Shruti is not a legacy font

Shruti, the Gujarati font Windows ships, is a Unicode font, and text set in it copies out as Gujarati. If a Shruti PDF pastes wrongly, the cause is a gap in its map or a scan, not the font. The analyser checks what the text reads as, not the font’s name.

What the analyser tells you

Legacy font (Krishna)
Cause one. Convert with Legacy PDF to DOCX.
Legacy font (LMG-Arun)
Cause one. Convert with Legacy PDF to DOCX.
Legacy font, no table yet
A legacy face such as a Shree-Lipi Gujarati face, recognised but not convertible yet.
Conjuncts are missing from the text layer
Cause two. PDF to DOCX rebuilds them from the font.
Scanned, no text layer
Cause three. Use Gujarati OCR.

Frequently asked questions

Is Krishna the same font as Harikrishna or GIST-TLOTKrishna?

No. C-DAC’s GIST-TLOTKrishna is a different font. The Krishna table reads faces named exactly Krishna, KrishnaBold, KrishnaItalic, Girdhar or Giridhar and Chitra, and it reads PDFs only.

Which other Gujarati legacy fonts are covered?

LMG Arun and Krishna. Terafont, Shree-Lipi’s Gujarati faces, Gopika, Saral and the rest have no table yet; the analyser names Shree-Lipi when it sees it.

Can I turn Unicode Gujarati into LMG Arun for PageMaker?

Yes. DOCX to PageMaker-ready writes LMG Arun bytes, and a test placed in PageMaker 7.0 arrived byte for byte when the steps in the instructions that come with the file were followed, starting with unticking Convert quotes in the Place dialog.

Why do the dates in my scan come out with ર?

Because ર and the digit ૨ look almost the same. Gujarati OCR repairs the case it can decide safely, a token made of that letter followed by Gujarati digits, and leaves ordinary words alone.

Is my Gujarati PDF corrupted?

No. It displays and prints correctly. It only lacks a correct way back from the page to Unicode text, which the conversion supplies.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems