Why Telugu is harder to recognise than Latin text
Telugu is an abugida: each consonant carries an inherent vowel, and other vowels are written as signs (matras) attached around the letter. Consonant clusters are written by stacking a subscript form (ottu) beneath the base letter, so a single visual unit can be two or three letters tall. That vertical stacking is where recognition usually fails — the subscript sits below the baseline, is smaller than the base letter, and is the first thing lost when a scan is low-resolution or slightly blurred. A dropped ottu does not produce obvious garbage; it produces a different, real Telugu word.
What this tool does about it
Telugu pages get a second recognition pass at the word level: words the first pass read with low confidence are cropped and re-read on their own, which recovers clusters that were lost in the full-page read. Character repairs are deliberately contextual rather than dictionary-based — a correction only fires on a specific confusion in a specific position. There is no fuzzy match against a word list, because snapping an uncertain word to its nearest dictionary neighbour silently turns correct text into different correct-looking text, and nothing downstream can detect that.
Measured accuracy
On our ground-truth corpus — 4 Telugu documents, 7,604 characters of prose — this pipeline reads 97.5% of characters correctly (2.51% character error rate). Those documents are clean, single-column, digitally rendered pages. A photograph taken at an angle, a faded photocopy or a page with columns and images will score worse. We publish the method alongside the number because an accuracy claim without one is not checkable.
Measured on a scan of our Telugu stress test
We also turned our 4-page Telugu stress test file into an image-only PDF, at 300 and at 200 dpi, and read it with this tool as a scan. The file holds every vowel, consonant, vowel sign and conjunct, with English, punctuation and a table. Its running prose came back with 0.2% and 0.0% of characters wrong: one word of 66 at 300 dpi, none at 200. The whole file came back with 16.4% and 15.4% of its 4,511 characters wrong, because most of it is charts. About 39% of the conjunct chart and 35% to 40% of the lone vowels were misread or missed. Letters set a space apart often come back run together. The Kannada and Hindi words in the file's table are not read at all, because a page is read in one Indic script. These are clean images, not photographs of paper, so they are the best case for a scan.
Measured on real books
We took 20 pages from 10 books on Telugu Wikisource, pages that volunteers had proofread against the scan and a second reader had checked, and read the scans with this tool. On clean modern print (6 pages from 3 books) 2.8% of characters came back wrong. On old letterpress, scanned for the Digital Library of India and the Million Book Project (14 pages from 7 books), 16.2% came back wrong, mostly where the type is worn, broken or faint. Old books also use two letters the Telugu recognition model cannot write at all, ఁ (arasunna) and ఱ: all 94 of them on those pages came back as something else, ఁ most often as ం and ఱ as whichever letter it looked most like. Running heads and page numbers are not counted.
What still gives it trouble
Old letterpress with worn or broken type, the letters ఁ and ఱ, letters and conjuncts standing alone rather than in words, dense conjunct stacks at small point sizes, decorative or display fonts that depart from standard letterforms, faint photocopies where the ottu has faded, and pages where Telugu and English alternate mid-sentence. If a result looks wrong, a sharper, flatter, evenly lit scan usually helps more than any setting. On our stress test, 200 and 300 dpi read equally well.