Chanakya PDF to Unicode: newspaper Hindi you can edit

Upload a PDF set in Chanakya, the font Hindi newspapers and many presses typeset in, and get it back as a Word file in Unicode Hindi. No copy-paste, no retyping, no paragraph limits. Copy a line out of such a PDF and you get strings like ∑§„Ê Á∑§ ‚⁄U∑§Ê⁄U instead of कहा कि सरकार, because the PDF’s own map from codes to letters points at symbols. The converter ignores that map and reads the codes the font draws.

Takes PDF files up to 25 MB · Files deleted within 4 hours · No sign-up

यह पेज हिन्दी में पढ़ें

The PDF’s text, as PyMuPDF reads it

∑§„Ê Á∑§ ‚⁄U∑§Ê⁄U ‚÷Ë ’Ê…∏ ¬˝÷ÊÁflÃÙ¥ ∑‘§ ‚ÊÕ π«∏Ë „Ò–

Read through the Chanakya table

कहा कि सरकार सभी बाढ़ प्रभावितों के साथ खड़ी है।

A line of a Hindi newspaper page set in Chanakya. Left: the text the PDF’s own map gives, which is what copying and ordinary converters get. Right: the same line read from the font’s codes; the printed page shows the right-hand line.
Back to Tools
🔤

Legacy PDF to DOCX

A PDF whose text was set in a legacy Indic font, recovered as editable Unicode Word — read from the font’s own table rather than by OCR.

How to use?

Leave it on detect if you are not sure. The ones marked “not yet” will tell you so rather than converting badly.

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

Why a Chanakya PDF copies as symbols

Chanakya, like Kruti Dev, draws Devanagari in the slots of an 8-bit Latin font, but its layout is its own: a full letter is usually a half form plus a stem, the short i is typed before its consonant, and the reph after its syllable. The PDF stores those codes, and the map it carries for copying turns them into symbol characters, so every copy, search and ordinary converter gets ∑§Ë for की.

Some newspapers’ text layers are worse: the letters the font keeps at codes 0x80 to 0x9F are missing from the copied text altogether, so मुख्यमंत्री loses its ख and गिरफ्तार its फ. The converter reads the codes themselves through a Chanakya table, drawn and checked in the real font, and puts every sign where Unicode wants it.

One name, three builds

Documents set in Chanakya do not all put the same thing on every key. Some newspapers put the digits 1 to 9 on keys that other builds use for half letters, and one exam paper put क् and ख् on keys of their own. The converter reads each document every way these builds allow and keeps the reading that gives words, so a page of prices does not come out as strings of half letters. Nothing has to be chosen.

Walkman-Chanakya, a different font with a similar name, has a layout of its own and is not read yet. Headlines a newspaper sets in some other face are not read either: only the text set in Chanakya is converted.

Measured

Twenty-six real PDFs set in Chanakya, from at least seven publishers: newspaper pages, a party document and a spiritual monthly typeset in PageMaker, two school-board exam papers made in Word, and a Gita commentary. Converted, they hold 101,746 words, and 94.0% of them are in Tesseract’s Hindi or Devanagari word lists: between 93% and 98% document by document, and 79.5% for the Gita commentary, whose Sanskrit verse and compounds the lists do not hold. That is not an accuracy figure; names and rare words outside the lists can be right. No code was left unread, and one printed line of each document was checked against the conversion by eye.

Measuring found faults, and each is fixed. A glyph read as the ra-foot is the subscript na, so अग्नि had come out अग्रि and प्रश्न प्रश्र; the ya-sign used after letters with no stem was not read, so पाठ्यक्रम and बुद्ध्या lost it; a candrabindu typed before ू stood before it; and on pages printing times and prices, digits had been read as half letters.

Where it still goes wrong

Of the 101,746 words, 17 hold a sign sequence no Hindi word has, mostly keys typed twice. A few codes in the font draw glyphs nobody has identified; they come through as characters that are plainly not Devanagari, so they are easy to spot. A newspaper page of boxed stories is the hardest layout there is, so check the order of the stories in the Word file.

Text in Walkman-Chanakya or in a headline face other than Chanakya stays as it was. A scanned page has no codes to read and needs OCR instead.

Frequently asked questions

Chanakya font PDF ko Unicode me kaise convert kare?

Upload the PDF above and download the Word file it gives back. The whole file is read at once, and the Hindi comes back in Unicode.

चाणक्य फ़ॉन्ट की PDF को यूनिकोड में कैसे बदलें?

ऊपर PDF अपलोड करें और जो Word फ़ाइल मिले उसे डाउनलोड करें। पूरी फ़ाइल एक साथ पढ़ी जाती है, और हिन्दी यूनिकोड में आती है।

Why does text copied from a Hindi newspaper PDF come out as symbols?

Because the PDF is set in a legacy font such as Chanakya, and the map it carries for copying points at symbol characters. The page looks right only because the font draws those codes as Hindi letters.

Is Chanakya the same as Kruti Dev?

No. Both put Devanagari on the keys of an English font, but on different keys, so a Kruti Dev converter turns Chanakya text into other letters. Each is read through its own table here.

What about Walkman-Chanakya?

It is a different font with its own layout, and it is not read yet. Text in it stays as it was.

My newspaper PDF is scanned. Will this work?

No. A scan holds pictures of letters and no codes to read. Hindi OCR reads the pictures instead.

What happens to my file?

It is stored encrypted while it converts and removed by the cleanup that clears every conversion, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Related tools

Other legacy font pages