Akruti Marathi to Unicode: convert AkrutiMar PDFs

Akruti’s Marathi fonts store the keys the typist pressed, not Marathi letters. An ordinary PDF-to-Word conversion turns a PDF set in them into lines of accented Latin letters, which is what IndicPDF’s own did until this table existed. This one reads those keys through an Akruti Marathi table and gives back Unicode Marathi in a Word file.

Takes PDF files up to 25 MB · Files deleted within 4 hours · No sign-up

The PDF’s text

Ùðó ÐððòçÃð¨î ¨îð ¡ðè÷

Read as Unicode

मी नास्तिक का आहे

The running title of a Marathi booklet, set in AkrutiMar_B faces, as PyMuPDF extracts it from the PDF and as the Akruti table reads it.
Back to Tools
🔤

Legacy PDF to DOCX

A PDF whose text was set in a legacy Indic font, recovered as editable Unicode Word — read from the font’s own table rather than by OCR.

How to use?

Leave it on detect if you are not sure. The ones marked “not yet” will tell you so rather than converting badly.

Files are stored encrypted
Deleted within 4 hours
SSL Encrypted

How Akruti stores Marathi

Most consonants are stored as a half form and completed by a bar, so म is Ù followed by ð. The short i is typed before its consonant and the reph after its syllable. क, क्र and फ are two glyphs, a body and a hook, and a vowel sign may be typed between them: ¨÷î is के. The converter reads each piece and puts the letters in Unicode order.

Why other tools see nothing here

The Akruti fonts that book embeds renumber their codes 1 to 91. Read by those codes, the text is control characters; read the way PyMuPDF and IndicPDF’s ordinary route read it, it is the accented letters in the sample. Neither is Marathi. The typed keys survive only as the names the PDF gives each glyph, and that is where this converter reads them from. Until it did, the ordinary route wrote the Latin letters into Word, and the legacy route did not notice a legacy font at all.

The colon key is a second trap. Inside a word it is the visarga: all 41 colons followed by a letter in the book are स्वतः, दुःख, निःसंशय and their like. At the end of a word it is usually a label, as in प्रकाशक: or मूल्य:, so there a Devanagari word list decides: स्वतः becomes a word with a visarga, and संकेत: keeps its colon.

Measured

On मी नास्तिक का आहे, nine pages, the Word file holds 4,930 words, 94.8% of them in Tesseract’s Marathi word list, with no sign sequence Marathi cannot spell. It agrees with a reference converter on 4,977 of 4,979 words, but that converter came with the mapping, so the agreement is not independent evidence. The two words that differ are ऐ, which the book types as ए and a sign: this reads याऐवजी where the other reads the impossible एे. A rendered page of the result reads word for word as printed. The sample is one book.

Where it still goes wrong

The table is named for the three faces that book embeds: AkrutiMar_BYogini, AkrutiMar_BAditi and AkrutiMar_BVijay. Another Akruti release may place a letter differently, and nobody has checked one. It reads PDFs only, not Word files typed in Akruti, and it cannot write Unicode back into Akruti. Text the same pages set in other fonts does not come across: the book’s years and prices in Times and its web and email addresses are left out, and so is a back-cover blurb set in an Akruti Unicode font whose map is broken.

Frequently asked questions

Which Akruti fonts does it read?

AkrutiMar_BYogini, AkrutiMar_BAditi and AkrutiMar_BVijay, measured on a real book. Other AkrutiMar faces may share their layout, but that has not been checked.

Can I convert a Word file typed in Akruti?

Not yet. Only PDFs have been measured, so this page offers PDFs only.

Does it work for Hindi set in Akruti?

Akruti’s Devanagari fonts for Hindi have not been measured, and the analyser names them without converting them. Hindi typed in Kruti Dev 010 is covered.

Is Akruti Oriya supported too?

Yes, through a separate table for the AkrutiOri -99 faces, in both directions.

Why does a colon sometimes become a visarga?

Because the Akruti colon key is the visarga inside a word, as in स्वतः and दुःख. At the end of a word it is kept as a colon unless a Devanagari word list knows the word only with a visarga.

What happens to my file?

It is stored encrypted while it converts and removed by the cleanup that clears every conversion, so nothing is kept longer than 4 hours after upload. Files up to 25 MB, no account needed.

Related tools

Other legacy font pages