Bengali PDF to Word — Convert Bengali PDFs to Editable DOCX

Convert a Bengali PDF into an editable Word document, with the text as real Unicode Bangla and the paragraphs in order.

বাংলা · Accepts PDF, JPG, PNG, TIFF up to 25 MB · Files deleted within 4 hours

Back to Tools
📋

PDF to DOCX

Extract text from PDFs and PageMaker documents and convert them into editable Word files, recovering legacy Indic fonts as Unicode.

How to use?

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted
Unicode output

NFC-normalised text that pastes into any editor and stays searchable.

No ads, no sign-up

Free to use. Uploads removed within 4 hours.

Measured, not claimed

94.8% character accuracy on our published test corpus.

Three kinds of Bengali PDF

A Bengali PDF made from a Unicode document carries its text inside, and converting it is extraction. A scanned Bengali PDF is pictures of pages and has to be recognised first, with the accuracy described on our Bengali OCR page. A third kind is common in Bangladesh and West Bengal: text typed in Bijoy. It looks like Bengali on the page, but the file holds something else, and neither extraction nor OCR is the right tool for it.

Bijoy and SutonnyMJ

Bijoy was released by Ananda Computers in 1988 and set Bangladesh’s newspapers, government print and textbooks for decades. Text typed in Bijoy stores ordinary Latin letters that the SutonnyMJ font draws as Bengali: বাংলা is stored as evsjv. Copy from such a PDF and those Latin letters are what you get. Our legacy converter reads the stored letters through a Bijoy table built by drawing every code in SutonnyMJ itself. Choose Bengali — Bijoy (SutonnyMJ) in its script list rather than relying on detection. The table has not yet been checked by a Bengali reader, and the converted document says so.

What comes through

Text, paragraph breaks and reading order come through. Fonts and exact layout are approximated, because the aim is a document you can edit rather than a copy of the page. Conjuncts are written with the hasanta, so any Unicode Bengali font on the machine opening the file draws them joined.

Frequently asked questions

Text copied from my Bengali PDF comes out as letters like evsjv. What is that?

The document was typed in Bijoy, where Bengali is stored as Latin letters that only SutonnyMJ draws correctly. Use the legacy converter with Bengali — Bijoy selected. OCR would only approximate text the file already holds exactly.

Will the Bengali text be editable and searchable in Word?

Yes. The DOCX holds real Unicode Bengali, so search, copy and spell-check all work. Word needs a Bengali font to display it; Windows and macOS both ship one.

Can I convert a scanned Bengali PDF to Word?

Yes. The pages are recognised first and then written into the DOCX, so expect the accuracy of recognition rather than the near-exact result of a digital PDF.

What file types can I use for Bengali OCR?

PDF, JPG, PNG and TIFF, up to 25 MB per file. Scanned PDFs are rasterised page by page before recognition, so a PDF with no text layer works the same as a photograph.

Is the output real Unicode text?

Yes. Output is standard Unicode, normalised to NFC, so it copies into Word, Google Docs or any editor and stays searchable. It is not an image of text and not a legacy font encoding.

Do I need an account?

No. The tool is free and requires no sign-up, no email and no payment. There are no advertisements on any page.

What happens to my files?

Uploads and results are stored encrypted and removed by a cleanup that runs every 15 minutes, so nothing is kept longer than 4 hours after upload. Files are processed to produce your result and are not used for anything else.

Should I use Auto Detect or pick the language myself?

Pick the language when you know it. Auto detection reads the script from the page and then verifies its guess, but a page with only a few lines, heavy noise, or mixed English gives it less to work with. Manual selection removes that uncertainty entirely.

Related tools

Other Indian languages