Marathi OCR — Extract Marathi Text from Scans and Images

Upload a scanned Marathi PDF or an image of a Marathi page and get editable Unicode Marathi back. Free, ad-free, no sign-up.

मराठी · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

100% character accuracy on our published test corpus.

Marathi or Hindi

Marathi and Hindi share the Devanagari script, so the page itself does not say which one it is, and each has its own recognition model. The pipeline decides from the letters: it reads a sample and counts ळ, a letter Marathi uses constantly and Hindi does not. With enough of them the page is read with the Marathi model; with too little text to judge, it stays with Hindi, the long-standing default, because reading Hindi through the Marathi model measured worse (1.5% of characters wrong against 3.6%). Choosing Marathi yourself skips the guess.

Books scanned on their side, and another program’s text layer

Two things that are common in Marathi book scans used to defeat the tool, and both are fixed. A book scanned sideways was detected as Bengali, Gujarati or Punjabi, because the check that verifies the script read its trial samples with the letters lying on their side; the page is now turned upright first, and a sideways Marathi novel detects as Marathi on 10 of its 11 body pages. And a scan put through another program’s OCR carries an invisible Latin text layer that does not read; the tool now recognises that layer as junk and reads those pages from the pixels. Twelve pages of that novel went from no Devanagari at all to 28,265 characters.

Measured accuracy

On our ground-truth corpus the pipeline read 100% of the characters of its one Marathi test page correctly, 1,453 characters of prose. One clean digital page is far too small a sample to call that a rate, so here is a harder one. We read six pages each of two scanned Marathi history books, which have no reference text, and counted the words of five letters or more that are in Tesseract’s own Marathi word list: 73.7% of 1,573 words and 75.3% of 1,231. That is not an accuracy figure either. Most of the words outside the list are names and places, such as हेस्टिंग्ज, नंदकुमार and महावंस, which are right.

Where it still goes wrong

The letter most often misread is ळ itself, the one that tells Marathi from Hindi: in those scans त्यामुळे came back as त्यामुक्ठे and त्यावेळी as त्यावेढी, and ट was read as ठ in दृष्टीने. Worn or faint old print loses thin strokes and matras. Handwriting is not supported; this reads printed Marathi.

Frequently asked questions

Marathi PDF madhla text kasa copy karaycha?

If the PDF is a scan, upload it above and the pages are recognised as Marathi text you can copy and edit. If it is a digital PDF whose text copies as nonsense, it was probably set in a legacy font; the Marathi PDF to Word page explains that case.

स्कॅन केलेल्या मराठी PDF मधून मजकूर कसा काढायचा?

वर PDF अपलोड करा. पाने मराठी मजकूर म्हणून ओळखली जातात आणि तो युनिकोडमध्ये परत मिळतो, जो तुम्ही कॉपी आणि संपादित करू शकता.

Will it mistake my Marathi page for Hindi?

Not if you choose Marathi. On Auto, a page with enough ळ is read as Marathi; a page with very little text may be read as Hindi, which reads most Marathi words correctly but misses some. Picking the language removes the guess.

Can it read old Marathi books?

Yes, if the print is legible. Scans of old books score lower than clean prints, and a page scanned on its side is turned upright before it is read.

What file types can I use for Marathi OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages