Bengali OCR — Extract Bengali Text from Scans and Images

Upload a scanned Bengali PDF or a photograph of a Bangla page and get editable Unicode text back. Free, ad-free, no sign-up.

বাংলা · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Optical Character Recognition (OCR). Online & Free

Convert Scanned Documents and Images into Text

Drop PDF or image here. PDF, JPG, PNG, TIFF supported.

How to recognize text from image?

1

Upload your file

Click "Choose Files" or drag and drop your scanned PDF or image onto the upload area.

2

Select language

Choose the language of the text in your document from the dropdown for best OCR accuracy.

3

Download text

Click "Recognize" and download the extracted text file once processing is complete.

Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

94.8% character accuracy on our published test corpus.

What Bengali script asks of a recogniser

Bengali hangs its letters from a headstroke, the matra, that runs along the top of a word and joins the letters into one band, much as Devanagari does. Conjuncts (juktakkhor) are often fused ligatures whose shape does not show the letters inside them: ক্ষ and জ্ঞ look like single letters, not like two. The vowel signs ি, ে and ৈ are drawn to the left of their consonant but stored after it, and ো and ৌ wrap around both sides. So correct Unicode output does not follow the left-to-right order of the marks on the page, and a reader that copies the visual order gets the spelling wrong.

Bengali or Assamese?

Assamese is written in the same script, so identifying the script does not identify the language. The two differ in a few letters: Assamese writes ৰ and ৱ where Bengali writes র and ব. When a page is Bengali script, the pipeline reads a sample and counts those two letters. If they make up at least 2% of it, the page is read as Assamese, otherwise as Bengali. When the sample cannot be read, it falls back to Bengali, which is by far the more common of the two. If you know the language, choosing it removes the check entirely.

Measured accuracy

On our ground-truth corpus — 1 Bengali document, 3,006 characters of prose — this pipeline reads 94.8% of characters correctly (5.16% character error rate). One document is a small sample, and it is a clean digital render; printed books, newspapers and photocopies will score lower. The same test file also lists conjuncts out of context, one after another, and that section reads far worse. It is reported separately rather than averaged in, because it is not prose.

How the page is prepared, and why only once

A scan is prepared several ways before recognition, as plain grayscale and with two kinds of threshold, and the best-scoring version is kept. On Bengali that choice turned out to be dangerous when made page by page: recognition confidence barely moved between versions while the error rate moved enormously, and one threshold that looked just as confident produced more than twice the errors. Choosing afresh for every page kept landing on it. The version is now chosen once, for the whole document. Handwriting is not supported; this reads printed Bengali.

Frequently asked questions

Can it read Bengali books and newspapers?

Printed Bengali in standard typefaces, yes. Books scan better than newspapers, which have tighter columns, thinner paper and lower print quality. Scan at 300 dpi or more with the page flat; that does more for the result than any setting.

My document is in Assamese. Will it be read as Bengali?

Not if it uses the Assamese letters ৰ and ৱ in ordinary quantity; the pipeline checks for them before choosing. A very short page gives that check little to go on, so choose Assamese yourself when you know it.

Are Bengali conjuncts kept correctly in the output?

A conjunct is stored in Unicode as consonant, hasanta (্) and consonant, and any Bengali font draws it as the joined form. When a conjunct is recognised correctly it looks the same as the original. Dense conjuncts at small sizes are the most common place for errors.

What file types can I use for Bengali OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages