Why your Marathi PDF copies out wrong

The page shows Marathi, and the copied text has holes in it: conjuncts gone or replaced by stray symbols, vowel signs on the wrong letter, or whole lines of Latin letters. There are four causes, and each needs a different fix. The analyser above tells you which one your PDF has.

On the page

मला जे म्हणायचं होतं ते ललितला बरोबर समजलं.

The PDF’s text, as PyMuPDF reads it

मला जे ᭥हणायचं होतं ते लिलतला बरोबर समजलं.

A line of व. पु. काळे’s झोपाळा, an ebook set in Arial Unicode MS. म्ह comes out as a stray symbol and the i-sign of ललित lands on the wrong letter. Poppler, another common reader, drops the conjunct altogether: हणायचं.
Back to Tools
🔍

Analyse

What have you got, and what will break? Upload a batch of PDFs, DOCX files and scans: each is checked for legacy fonts, a text layer, scan resolution, signatures, protection and invalid sign sequences, and pointed at the tool it needs.

How to use?

Used when a scan is opened in Correct.

Drop files here or click to choose them

PDF, DOCX, PNG, JPEG or TIFF · up to 500 files, 25 MB each

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

Cause one: the font’s map leaves out the conjuncts

Some Marathi ebooks, made with calibre and set in Arial Unicode MS, carry a map from glyphs back to letters that names only the 61 letters and signs the font has codes for. The half forms, conjuncts and rephs the page is full of have no entry, so text extraction returns nothing, a (cid:7013) marker or a stray symbol for each of them. The text also comes out in the order it was drawn, so the i-sign stands before its consonant, which turns आणि into आिण.

The font itself knows what each of those glyphs is, because its shaping rules built them from letters. PDF to DOCX reads those rules for this font and puts the letters back in order. On two such books the share of words found in Tesseract’s Marathi word list went from 67.4% to 91.9% and from 67.2% to 87.1%, and the 9,893 and 8,887 (cid markers went to none. Most of what is left out of the list is names, dialect such as म्हनं and घिऊन, and loanwords.

Cause two: Kruti Dev

Marathi was typed in Kruti Dev for years, as Hindi was. A PDF of such a document stores the keys the typist pressed, so it pastes as Latin letters. Legacy PDF to DOCX reads those keys through a Kruti Dev 010 table, the same one it uses for Hindi, and Marathi’s ळ, which Kruti Dev keeps on the G key, comes through with the rest.

Cause three: Akruti

Books set with Akruti’s Marathi fonts store the typed keys too, and their embedded fonts renumber them, so an ordinary converter produces Ùðó ÐððòçÃð¨î where the page says मी नास्तिक का आहे. Legacy PDF to DOCX reads the AkrutiMar_B faces through their own table.

Cause four: a scan, or another program’s OCR over it

A scanned page is a picture, and needs OCR. The trap is a scan that already carries another program’s OCR: one Marathi book, scanned on its side, came with another program’s OCR as a hidden text layer, in Latin letters. It selects and copies as text, and none of it is Marathi. Such pages are read again from the page images: on twelve pages of that book, 0 Devanagari characters became 28,265.

A sideways page is also where script detection used to go wrong. Lying on their side, the letters were taken for Bengali, Gujarati or Punjabi on 7 of that book’s 11 body pages and for Hindi on others, and not one page was detected as Marathi. The page is now turned upright before its script is judged, and 10 of the 11 are detected as Marathi.

What the analyser tells you

Conjuncts are missing from the text layer
Cause one. The font’s map leaves out its conjuncts; PDF to DOCX rebuilds them from the font.
Legacy font (Kruti Dev 010)
Cause two. Convert with Legacy PDF to DOCX.
Legacy font (Akruti Marathi)
Cause three. Convert with Legacy PDF to DOCX.
Text layer does not read: read by OCR
Cause four, when another program’s OCR is over the scan. PDF to DOCX reads those pages from the images.
Scanned, no text layer
Cause four, a plain scan. Use OCR.

Frequently asked questions

Why does आणि paste as आिण?

The PDF stores its text in the order the glyphs are drawn, and the i-sign is drawn to the left of its consonant. A reader that trusts that order hands it over as written. Converting puts the sign back after the letter it belongs to.

Will installing a Marathi font fix the pasted text?

No. What is pasted is missing letters or holds the wrong characters, whatever font displays it. The text has to be read again, from the font’s rules, its table or the page image.

Is my PDF damaged?

No. Every one of these PDFs prints and displays correctly. What each lacks is a correct way back from the drawn page to Unicode text, and that is what the conversion supplies.

Which tool do I need?

For a map with missing conjuncts, or another program’s OCR over a scan, PDF to DOCX. For Kruti Dev or Akruti, Legacy PDF to DOCX. For a plain scan, OCR. The analyser above says which applies.

Does the same thing happen to Hindi PDFs?

Yes. Hindi is written in the same script and set in the same fonts, so the same four causes apply.

Does the analyser keep my PDF?

Not for long. The PDF is stored encrypted while the report is built, and it is kept afterwards only so you can open it in Correct from the report without uploading it again. The same cleanup that clears every conversion removes it, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Tools for what it finds

Other PDF problems