Extraction or recognition
A Malayalam PDF saved from a Unicode document holds its text, so converting it is extraction and comes out close to exact. A scanned Malayalam PDF holds pictures of pages, so it is recognised first and inherits recognition’s mistakes, described on the Malayalam OCR page. You upload both the same way.
A text layer that sends conjuncts to Latin letters
Several Malayalam PDFs we tested look right and copy as a mix of Malayalam and Latin-1 characters, like സി·ുവുംകാർ³ിയും: the PDF’s map from glyphs to letters sends its conjuncts to the wrong characters. The converter spots that damage page by page and reads those pages by OCR instead of trusting the map, and the result is written with the same repairs the OCR tool makes, so its chillu letters are the single Unicode characters. Recognition takes about 23 seconds a page, so a long book is refused at upload with an instruction to split it, rather than failing after an hour.
Legacy Malayalam fonts
Older Malayalam DTP was set in legacy fonts that store keys rather than letters. Three layouts have tables here, each built from the font’s own glyphs rather than taken from another converter: ML-TT Karthika, the family of ML-Revathi, ML-Leela, ML-Anjali and ML-Thunchan, and Akhila. A PDF set in one of them is read automatically. None has yet been checked by a reader who reads Malayalam, and the converted document says so.
What is kept
Text, paragraph breaks and reading order are kept. Fonts and exact layout are approximated, since the aim is a document to edit, and Malayalam is set in a Unicode Malayalam font so it opens anywhere.