Cause one: the font’s map leaves out the conjuncts
Some Marathi ebooks, made with calibre and set in Arial Unicode MS, carry a map from glyphs back to letters that names only the 61 letters and signs the font has codes for. The half forms, conjuncts and rephs the page is full of have no entry, so text extraction returns nothing, a (cid:7013) marker or a stray symbol for each of them. The text also comes out in the order it was drawn, so the i-sign stands before its consonant, which turns आणि into आिण.
The font itself knows what each of those glyphs is, because its shaping rules built them from letters. PDF to DOCX reads those rules for this font and puts the letters back in order. On two such books the share of words found in Tesseract’s Marathi word list went from 67.4% to 91.9% and from 67.2% to 87.1%, and the 9,893 and 8,887 (cid markers went to none. Most of what is left out of the list is names, dialect such as म्हनं and घिऊन, and loanwords.
Cause two: Kruti Dev
Marathi was typed in Kruti Dev for years, as Hindi was. A PDF of such a document stores the keys the typist pressed, so it pastes as Latin letters. Legacy PDF to DOCX reads those keys through a Kruti Dev 010 table, the same one it uses for Hindi, and Marathi’s ळ, which Kruti Dev keeps on the G key, comes through with the rest.
Cause three: Akruti
Books set with Akruti’s Marathi fonts store the typed keys too, and their embedded fonts renumber them, so an ordinary converter produces Ùðó ÐððòçÃð¨î where the page says मी नास्तिक का आहे. Legacy PDF to DOCX reads the AkrutiMar_B faces through their own table.
Cause four: a scan, or another program’s OCR over it
A scanned page is a picture, and needs OCR. The trap is a scan that already carries another program’s OCR: one Marathi book, scanned on its side, came with another program’s OCR as a hidden text layer, in Latin letters. It selects and copies as text, and none of it is Marathi. Such pages are read again from the page images: on twelve pages of that book, 0 Devanagari characters became 28,265.
A sideways page is also where script detection used to go wrong. Lying on their side, the letters were taken for Bengali, Gujarati or Punjabi on 7 of that book’s 11 body pages and for Hindi on others, and not one page was detected as Marathi. The page is now turned upright before its script is judged, and 10 of the 11 are detected as Marathi.