Kannada OCR — Extract Kannada Text from Scans and Images

Upload a scanned Kannada PDF or a photo of a Kannada page and get editable Unicode text back. Free, ad-free, no sign-up.

ಕನ್ನಡ · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

95.1% character accuracy on our published test corpus.

Why Kannada is tall

Kannada is a sister script of Telugu and builds clusters the same way: the second consonant is written as an ottakshara, a small subscript form beneath the base letter. Vowel signs sit beside or above the consonant, and some, such as ೊ and ೋ, are drawn in several parts. A single syllable can therefore be two or three layers tall. The subscript is the smallest and lowest part of it, and the first to be lost on a blurred or low-resolution scan. Losing it rarely produces obvious garbage. It usually produces a different, real Kannada word.

The anusvara that reads as a zero

Kannada draws the anusvara ಂ and the digit zero ೦ almost identically. On the live service, all three anusvaras on one Kannada page came back as zeros, ಬೆಂಗಳೂರು as ಬೆ೦ಗಳೂರು, at confidences of 95, 93 and 88, so confidence gave no warning at all. The repair is a rule of spelling rather than a list of words: Kannada does not put a digit inside a word, so a zero attached to a Kannada letter is turned back into the anusvara. A number standing on its own is left alone.

Measured accuracy

On our ground-truth corpus — 1 Kannada document, 2,176 characters of prose — this pipeline reads 95.1% of characters correctly (4.87% character error rate). That is one clean digital render, which is a small sample, and printed or photocopied pages will do worse. The test file also lists ottaksharas one after another out of context; that part scores far lower and is reported separately, because nobody writes Kannada that way.

What still gives it trouble

Deep ottakshara stacks at small point sizes, faded photocopies where the subscripts have broken up, display typefaces, and lines where Kannada and English alternate word by word. Handwriting is not supported; this reads printed Kannada. A rescan at 300 dpi or higher helps more than anything else.

Frequently asked questions

Can it read Kannada government orders and circulars?

A printed or scanned order, yes. Many older Karnataka government PDFs were not scanned but typed in Nudi, and still contain their text in a legacy encoding. For those, the legacy converter reads the text exactly, where OCR would only approximate it.

Are ottaksharas kept in the output?

Yes. An ottakshara is stored in Unicode as base consonant, virama and consonant, and any Kannada font draws it stacked. Recognition errors on dense stacks are the most common failure, as described above.

Will Kannada numerals survive the anusvara repair?

Yes. The repair only changes a zero that is attached to a Kannada letter inside a word. Dates, page numbers and figures written in Kannada digits are left as digits.

What file types can I use for Kannada OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages