Tamil OCR — Extract Tamil Text from Scans and Images

Upload a scanned Tamil PDF or a photo of a Tamil page and get editable Unicode text back. Free, ad-free, no sign-up.

தமிழ் · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

99.1% character accuracy on our published test corpus.

Why Tamil recognises comparatively well

Tamil has a smaller consonant inventory than most Indic scripts and, crucially, far less vertical stacking. Where Telugu and Kannada build conjuncts by placing a subscript form beneath the base letter, Tamil generally writes consonant clusters in sequence and marks a bare consonant with a dot above it, the pulli (்). Letters therefore stay on the line and stay separable, which is exactly what a recogniser needs. That is reflected in the measurements below.

Where it still goes wrong

Several Tamil letters differ only in a small stroke or loop, and those pairs are the usual source of substitutions at low resolution. Vowel signs that attach to both sides of a consonant (as in ெ ா combinations) can be split across a line break by a naive reader. Grantha letters used for Sanskrit-derived sounds — ஜ, ஷ, ஸ, ஹ — appear less often in training material and are correspondingly less reliable.

Measured accuracy

On our ground-truth corpus — 2 Tamil documents, 3,286 characters of prose — this pipeline reads 99.1% of characters correctly (0.88% character error rate), the strongest result of any script we measure. The corpus is clean digital renders; photographs and photocopies score lower. Two documents is a small sample, and we would rather say so than imply more coverage than we have.

Frequently asked questions

Does it handle old Tamil print and palm-leaf transcriptions?

Standard printed Tamil works well. Very old typefaces with letterforms that differ from modern print will produce more errors, and handwritten or palm-leaf material is not supported — this recognises printed text.

Is Tamil output compatible with Unicode Tamil fonts?

Yes. Output is standard Unicode Tamil in NFC form, so it works with any Unicode font and any Tamil input method.

What about Tamil mixed with English?

Supported. Tamil recognition runs alongside English, so pages that mix the two come back with both intact.

What file types can I use for Tamil OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages