[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: An improved method for scanning documents



Shawn B. wrote:
For a while I was assisting Mike Harvey to OCR source code from his Nibble scans. Anything less than 600 DPI was a major headache. For example, I must have easily spent over 20 hours of time converting just one program (in failed attempts) scanned at 300 DPI. When we tried at 600 DPI, the quality of the OCR (using OmniPage Pro 14/15 Acrobat 7) was very good, but the shear number of fixups... when started over and tried at 600 DPI took only about 6 hours or so to get 200 lines of AppleSoftt BASIC tested and running.

For text and manuscript, there'd be very little effort for the human. But for source code (HEX especially and ASM doubly and BASIC last) it is a huge pain to verify the sometimes confused "B" for "8" (vise versa) and "0" (zero) for "O" (capital O) and "l" (small L) for "1" (number 1). Sometimes the $ would get messed up and the paranetheses are always a pain. In all, you have to hand very each and every line of code and without special help (a special mono-spaced font) that more clearly defines the differences (underlining numbers for example) you're in for a treat.

It would take a lot less than 6 hours to type in a 200 line Applesoft
program, so I'd say that OCR is worse than useless in this case.

-michael

Music synthesis for 8-bit Apple II's!
Home page:  http://members.aol.com/MJMahon/

"The wastebasket is our most important design
tool--and it is seriously underused."