Telugu OCR — Extract Telugu Text from Scans and Images

Upload a scanned Telugu PDF or a photograph of a Telugu page and get back editable Unicode text. Free, ad-free, and no sign-up.

తెలుగు · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

97.5% character accuracy on our published test corpus.

Why Telugu is harder to recognise than Latin text

Telugu is an abugida: each consonant carries an inherent vowel, and other vowels are written as signs (matras) attached around the letter. Consonant clusters are written by stacking a subscript form (ottu) beneath the base letter, so a single visual unit can be two or three letters tall. That vertical stacking is where recognition usually fails — the subscript sits below the baseline, is smaller than the base letter, and is the first thing lost when a scan is low-resolution or slightly blurred. A dropped ottu does not produce obvious garbage; it produces a different, real Telugu word.

What this tool does about it

Telugu pages get a second recognition pass at the word level: words the first pass read with low confidence are cropped and re-read on their own, which recovers clusters that were lost in the full-page read. Character repairs are deliberately contextual rather than dictionary-based — a correction only fires on a specific confusion in a specific position. There is no fuzzy match against a word list, because snapping an uncertain word to its nearest dictionary neighbour silently turns correct text into different correct-looking text, and nothing downstream can detect that.

Measured accuracy

On our ground-truth corpus — 4 Telugu documents, 7,604 characters of prose — this pipeline reads 97.5% of characters correctly (2.51% character error rate). Those documents are clean, single-column, digitally rendered pages. A photograph taken at an angle, a faded photocopy or a page with columns and images will score worse. We publish the method alongside the number because an accuracy claim without one is not checkable.

Measured on a scan of our Telugu stress test

We also turned our 4-page Telugu stress test file into an image-only PDF, at 300 and at 200 dpi, and read it with this tool as a scan. The file holds every vowel, consonant, vowel sign and conjunct, with English, punctuation and a table. Its running prose came back with 0.2% and 0.0% of characters wrong: one word of 66 at 300 dpi, none at 200. The whole file came back with 16.4% and 15.4% of its 4,511 characters wrong, because most of it is charts. About 39% of the conjunct chart and 35% to 40% of the lone vowels were misread or missed. Letters set a space apart often come back run together. The Kannada and Hindi words in the file's table are not read at all, because a page is read in one Indic script. These are clean images, not photographs of paper, so they are the best case for a scan.

Measured on real books

We took 20 pages from 10 books on Telugu Wikisource, pages that volunteers had proofread against the scan and a second reader had checked, and read the scans with this tool. On clean modern print (6 pages from 3 books) 2.8% of characters came back wrong. On old letterpress, scanned for the Digital Library of India and the Million Book Project (14 pages from 7 books), 16.2% came back wrong, mostly where the type is worn, broken or faint. Old books also use two letters the Telugu recognition model cannot write at all, ఁ (arasunna) and ఱ: all 94 of them on those pages came back as something else, ఁ most often as ం and ఱ as whichever letter it looked most like. Running heads and page numbers are not counted.

What still gives it trouble

Old letterpress with worn or broken type, the letters ఁ and ఱ, letters and conjuncts standing alone rather than in words, dense conjunct stacks at small point sizes, decorative or display fonts that depart from standard letterforms, faint photocopies where the ottu has faded, and pages where Telugu and English alternate mid-sentence. If a result looks wrong, a sharper, flatter, evenly lit scan usually helps more than any setting. On our stress test, 200 and 300 dpi read equally well.

Frequently asked questions

Can it read scanned Telugu books and old documents?

Yes, if the print is legible. Recognition quality tracks scan quality closely: 300 dpi or higher, flat pages and even lighting make a much larger difference than any option in the interface. Very old print with faded or broken letters will produce errors.

Are Telugu conjuncts and matras preserved correctly?

Conjuncts are written as base plus virama plus subscript in Unicode, and matras attach to their consonant, so correctly recognised text renders identically to the original in any Unicode-aware application. Recognition errors on stacked conjuncts are the most common failure mode and are described above.

Does it handle Telugu and English on the same page?

Yes. Telugu recognition runs alongside English, so mixed pages — Telugu prose with English technical terms in brackets, for example — come back with both scripts intact.

What file types can I use for Telugu OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages