Malayalam OCR — Extract Malayalam Text from Scans and Images

Upload a scanned Malayalam PDF or an image of a Malayalam page and get editable Unicode Malayalam back. Free, ad-free, no sign-up.

മലയാളം · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

94.0% character accuracy on our published test corpus.

Two ways to spell one letter

Malayalam writes six chillu letters, the consonants with no vowel of their own, such as ൽ and ൻ, and Unicode has two spellings for each: one character, or the consonant plus chandrakkala plus an invisible joiner. They look the same, but a search for one does not find the other. The recognition model writes the old three-character spelling, so ജനങ്ങൾ came back as ജനങ്ങള് followed by a joiner. Every result here is rewritten to the single character. The model also puts an invisible non-joiner after every chandrakkala that ends a word; it changes nothing on screen, drawn in three Malayalam fonts, and it breaks searches, so that one is dropped. Inside a word the non-joiner is kept, because there it chooses the visible chandrakkala over a conjunct.

Conjuncts on their own

Our Malayalam test file ends with a list of conjuncts set one after another, out of any word, and that section reads far worse than the prose: about 23% of its characters come back wrong, against 6% in running text. Recognition leans on the letters around a cluster, and a list gives it none. It is kept out of the figure below, because nobody writes Malayalam that way.

Measured accuracy

On our ground-truth corpus, 2 Malayalam documents with 3,834 characters of prose, this pipeline reads 94.0% of characters correctly (5.97% character error rate). They are clean digital renders of one text, one of them in a rounded handwriting-style font. On a real novel whose PDF text layer is broken, six pages read by OCR put 72.7% of their words of five letters or more in Tesseract’s own Malayalam word list; the book is written in a regional dialect the list does not hold, so that is a floor, not an error rate.

Where it still goes wrong

On the test pages the same few confusions recur: പ്പ read as ട്ട (പുഷ്പങ്ങൾ as പുഷ്ടങ്ങൾ), ന read as പ in a conjunct (സ്വപ്നം as സ്വപ്പം), and the ya sign ്യ read as ൃ. The au sign may come back as either of its two correct spellings. Handwriting is not supported; this reads printed Malayalam.

Frequently asked questions

Malayalam PDF il ninnu text engane copy cheyyam?

If the PDF is a scan, upload it above and the pages are recognised as Malayalam text you can copy and edit. If a digital PDF copies as symbols and Latin letters, the Malayalam PDF to Word page explains why and converts it.

സ്കാൻ ചെയ്ത മലയാളം PDF-ൽ നിന്ന് ടെക്സ്റ്റ് എങ്ങനെ എടുക്കാം?

മുകളിൽ PDF അപ്‌ലോഡ് ചെയ്യുക. പേജുകൾ മലയാളം ടെക്സ്റ്റായി തിരിച്ചറിഞ്ഞ്, പകർത്താനും തിരുത്താനും കഴിയുന്ന യൂണികോഡിൽ തിരികെ ലഭിക്കും.

Will ൽ and ൻ search correctly in the result?

Yes. Chillu letters are written as the single Unicode character, which is what current keyboards type and what search expects.

A word I can see on the page does not match when I search. Why?

Malayalam can spell the same letters with invisible joiners, and the two spellings look identical. The result is written the way current keyboards type, with single-character chillus and no joiner at the end of a word, so a search typed on a keyboard finds it.

What file types can I use for Malayalam OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages