Cause one: a legacy font such as Krishna or LMG Arun
Gujarati DTP used fonts that put Gujarati shapes where an English font keeps its letters, so the file stores those codes and pastes them as symbols. LMG Arun is the most downloaded of them on the one font site that counts every face, and Legacy PDF to DOCX reads it through a table built from the font: 342 of 373 words from two real documents, written through the table and drawn in the real font, read back correctly.
Three Dada Bhagwan books, પ્રેમ, અહિંસા and વાણી વ્યવહારમાં, are set in Krishna, a face no table covered and no font file on hand could explain. Their table was read off the fonts the PDFs themselves embed: every code drawn from its own outline, and settled against every word of the three books that holds it. Through it, 95.1% of the body’s words are in Tesseract’s Gujarati word list; through the LMG Arun table, the same bytes give 9.8%. One code shows why this had to be done by drawing: 0xAD is પ્ર, and it is also the soft hyphen, which a text renderer hides, so it first looked blank.
Cause two: a Unicode font whose map leaves out the conjuncts
A 727-page ebook, જંગલી પશ્ચિમમાં સાહસો, is set in Arial Unicode MS, a Unicode font, and in its first hundred pages the PDF’s map leaves out 84 of the conjuncts and half forms it draws, 8,958 times. Extracted, ક્રિયા came out as ἵṀયા in one reader and યા in another. PDF to DOCX reads those glyphs from the font’s own shaping rules. On pages 11 to 60 the share of words in Tesseract’s Gujarati list went from 78.6% to 95.2%, and 4,769 (cid markers went to none. What is left is the book’s own: it is a machine translation that leaves English inside some words.
Cause three: a scan
A scanned page holds a picture of the text and nothing to copy, or a hidden layer added by an OCR program. Gujarati OCR reads the picture. One thing to check afterwards is dates and numbers: ર and the digit ૨ are drawn almost alike, and the Gujarati OCR page explains how that is handled.
Shruti is not a legacy font
Shruti, the Gujarati font Windows ships, is a Unicode font, and text set in it copies out as Gujarati. If a Shruti PDF pastes wrongly, the cause is a gap in its map or a scan, not the font. The analyser checks what the text reads as, not the font’s name.