Two ways to spell one letter
Malayalam writes six chillu letters, the consonants with no vowel of their own, such as ൽ and ൻ, and Unicode has two spellings for each: one character, or the consonant plus chandrakkala plus an invisible joiner. They look the same, but a search for one does not find the other. The recognition model writes the old three-character spelling, so ജനങ്ങൾ came back as ജനങ്ങള് followed by a joiner. Every result here is rewritten to the single character. The model also puts an invisible non-joiner after every chandrakkala that ends a word; it changes nothing on screen, drawn in three Malayalam fonts, and it breaks searches, so that one is dropped. Inside a word the non-joiner is kept, because there it chooses the visible chandrakkala over a conjunct.
Conjuncts on their own
Our Malayalam test file ends with a list of conjuncts set one after another, out of any word, and that section reads far worse than the prose: about 23% of its characters come back wrong, against 6% in running text. Recognition leans on the letters around a cluster, and a list gives it none. It is kept out of the figure below, because nobody writes Malayalam that way.
Measured accuracy
On our ground-truth corpus, 2 Malayalam documents with 3,834 characters of prose, this pipeline reads 94.0% of characters correctly (5.97% character error rate). They are clean digital renders of one text, one of them in a rounded handwriting-style font. On a real novel whose PDF text layer is broken, six pages read by OCR put 72.7% of their words of five letters or more in Tesseract’s own Malayalam word list; the book is written in a regional dialect the list does not hold, so that is a floor, not an error rate.
Where it still goes wrong
On the test pages the same few confusions recur: പ്പ read as ട്ട (പുഷ്പങ്ങൾ as പുഷ്ടങ്ങൾ), ന read as പ in a conjunct (സ്വപ്നം as സ്വപ്പം), and the ya sign ്യ read as ൃ. The au sign may come back as either of its two correct spellings. Handwriting is not supported; this reads printed Malayalam.