What a paste box does well
For a line typed in plain Kruti Dev 010, the paste converter measured here and the whole-file conversion give the same Hindi. Its code reads the keys, moves the short i after its consonant and the reph before its letter, and it even reads the curly quotes Word puts on the keys that type श् and ष्, and the digits a PDF viewer copies as ƒ „ …. If all you have is a paragraph of Kruti Dev text, a paste box is fine.
What a document loses in a paste box
Copying flattens the page first. What you paste is lines of text: the page breaks, the table grid, the running head and the pictures stay behind. Then the paste box reads every character as a Kruti Dev key, including the English, the web addresses and the file numbers the PDF set in an ordinary font. A file reference number on one notice came back as Devanagari letters and signs, with only its digits left as they were. On a government notice that is most of the numbers that matter, so you copy around every English run or retype it.
And a converter built for Kruti Dev 010 reads the other builds the 010 way. DevLys 010 and the Kruti Dev build of a gazette notification put श् and ष् on the two curly quotes the other way round, so प्रश्न comes out प्रष्न and आवश्यक comes out आवष्यक.
What the whole file gives back
The conversion reads the codes the PDF stores, not the text a viewer copies, and rebuilds the document around them: one Word section for each printed page, with its size and margins; paragraphs with their alignment, indents and type sizes; bold; the running head and page numbers in the Word header; pictures on their pages; and English left in its own font. Ruled tables become Word tables where a grid can hold them — 219 tables in 29 of the 32 documents below. Which build of the font a document uses is decided from the document itself, by which reading of its curly quotes gives words.
Measured
Thirty-two real public PDFs typed in Kruti Dev and DevLys — government tenders and orders, recruitment notices and a gazette notification — were copied the way a PDF viewer copies them, and the copied text was run through the published code of a widely used online Kruti Dev to Unicode converter. The whole files went through this converter.
The whole-file conversions kept 15,806 English words, in 27 of the 32 documents. The paste converter’s output held no Latin word at all: every one became Devanagari letters. In 31 of the 32 documents, a larger share of the long words were words in Tesseract’s Devanagari word list after the whole-file conversion. On the eleven DevLys documents, of 64 places where the pages print प्रश्न, आवश्यक or उद्देश्य, the paste converter wrote 18 with ष्.
One paste converter was measured, by running its own published code; others differ, and some also take a .docx file, which was not measured.
Where it still goes wrong
The whole-file conversion has limits of its own. A dense application form comes out longer than the printed one, because a word processor wraps the text in its boxes its own way: twelve printed pages of one application form take eighteen. The Gazette’s pages of mixed English and Hindi run over too, fourteen pages into twenty-two. A scanned PDF has no codes to read and needs OCR. And codes nobody has identified yet come through as characters that are plainly not Devanagari, so they are easy to spot.