Extraction or recognition
A Marathi PDF saved from a Unicode document holds its text, so converting it is extraction and comes out close to exact. A scanned Marathi PDF holds pictures of pages, so it is recognised first and inherits recognition’s mistakes, described on the Marathi OCR page. You upload both the same way.
A text layer with no conjuncts in it
Some Marathi e-books were made with calibre and set in Arial Unicode MS, a Unicode font, and still copy badly: the PDF’s map from glyphs to letters names only the plain letters, so every half form, conjunct and reph is lost or comes out as a stray symbol, and म्हणायचं copies as हणायचं. The converter reads what each glyph is from the font’s own shaping rules and puts the signs back in order. On two such books the share of words in Tesseract’s Marathi list went from 67.4% to 91.9% and from 67.2% to 87.1%. A third book, set in Noto Sans Devanagari UI with the same kind of gap, went from 51.2% to 77.4%.
Legacy Marathi fonts
Older Marathi publishing was typeset in legacy fonts that store keys rather than letters, and copying from them gives Latin letters or symbols. Two families have tables here and pages of their own. Shree-Lipi, measured on twelve issues of a Marathi magazine (636 pages, 90.7% of long words in the Marathi list), and Akruti Marathi, measured on one book (94.8%). Either is read automatically when you upload the PDF.
What is kept, and what is not
Text, paragraph breaks and reading order are kept; fonts and exact layout are approximated, since the aim is a document to edit. One spelling is left as the source typed it: अ with the candra sign (अॅ), which Marathi writes for the English a in बॅंक and which some typists key as two characters.