Marathi or Hindi
Marathi and Hindi share the Devanagari script, so the page itself does not say which one it is, and each has its own recognition model. The pipeline decides from the letters: it reads a sample and counts ळ, a letter Marathi uses constantly and Hindi does not. With enough of them the page is read with the Marathi model; with too little text to judge, it stays with Hindi, the long-standing default, because reading Hindi through the Marathi model measured worse (1.5% of characters wrong against 3.6%). Choosing Marathi yourself skips the guess.
Books scanned on their side, and another program’s text layer
Two things that are common in Marathi book scans used to defeat the tool, and both are fixed. A book scanned sideways was detected as Bengali, Gujarati or Punjabi, because the check that verifies the script read its trial samples with the letters lying on their side; the page is now turned upright first, and a sideways Marathi novel detects as Marathi on 10 of its 11 body pages. And a scan put through another program’s OCR carries an invisible Latin text layer that does not read; the tool now recognises that layer as junk and reads those pages from the pixels. Twelve pages of that novel went from no Devanagari at all to 28,265 characters.
Measured accuracy
On our ground-truth corpus the pipeline read 100% of the characters of its one Marathi test page correctly, 1,453 characters of prose. One clean digital page is far too small a sample to call that a rate, so here is a harder one. We read six pages each of two scanned Marathi history books, which have no reference text, and counted the words of five letters or more that are in Tesseract’s own Marathi word list: 73.7% of 1,573 words and 75.3% of 1,231. That is not an accuracy figure either. Most of the words outside the list are names and places, such as हेस्टिंग्ज, नंदकुमार and महावंस, which are right.
Where it still goes wrong
The letter most often misread is ळ itself, the one that tells Marathi from Hindi: in those scans त्यामुळे came back as त्यामुक्ठे and त्यावेळी as त्यावेढी, and ट was read as ठ in दृष्टीने. Worn or faint old print loses thin strokes and matras. Handwriting is not supported; this reads printed Marathi.