[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: Apple II Reference Manual Addendum PDF now online



On Tue, 28 Jun 2005 13:44:01 GMT, pausch@saaf.se (Paul Schlyter)
wrote:

>In article <Rc-dnV4HaNNOel3fRVn-vQ@comcast.com>,
>Michael J. Mahon <mjmahon@aol.com> wrote:
>............. 

--<snip>--
>
>Otoh OCR'ing is done at the character level, not the word level.
>And if the OCR program is good, and encounters something it cannot
>interprete, it will display a graphical representation of that character
>and then ask you what ASCII character you should assign to it.
>In that way, the OCR program will "learn" and become better and
>better.
>

--<snip>--

A problem that I ran into when deciding whether to attempt to OCR the
Computists that I am scanning, wasn't when the program ran into
characters it didn't understand, but when it came across characters
that it mistakenly thought it knew and substituted incorrect ones.  

The other problem was that, yes, the actual OCR is done a character at
a time, however the software that did the proofing worked like a
spell-checker, so it would often choke on, for example, "$FE02",
either making a best guess or giving up entirely.  This resulted in a
tedious proofing process - approving every single line and character
in an assembly dump, for example - and I ended up having to go through
manually again and correct all the errors made by the recognition
stage.

Font changes and paragraph realignments in the middle of a scanned
section often throw the OCR software for a loop, as well.

These were common problems that cropped up in the most recent versions
OmniPage, ABBYY, and ScanSoft Paperport.

- Mike
maginnis@tarnover.org

The Computist Project
http://www.computist-project.net