[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: Scanning?



Mike Maginnis wrote:

On Fri, 15 Apr 2005 12:19:20 -0700, "Michael J. Mahon"
<mjmahon@aol.com> wrote:


Mike Maginnis wrote:


--<snip>--


A searchable PDF sounds like Nirvana--but experience indicates that
OCR'd documents inevitably have errors that make it through the
proofing process--and the process is much more difficult as well
as errorprone.

Scanning is a fine solution.

You didn't mention what tools you have available for compression.
I find that Photoshop 7's "Save for web.." option is very versatile
and effective.

For text, pre-processing to increase contrast and drop out the
background "white" noise, followed by .gif compression with 4
levels (black, 2 grays, and white) to be excellent.

If you can still see the paper grain in the white background,
then you cannot achieve good compression--noise doesn't compress!

-michael

8-voice music synthesizer using NadaNet networking!
Home page:  http://members.aol.com/MJMahon/


I've got Photoshop CS, but I've not played with image manipulation
levels or compression, so I'm in over my head with that.  I made the
unfortunate decision to purchase a SnapScan One-Touch scanner a while
back, so I'm limited in my image acquiring methods; the One-Touch
won't talk to anything but its own rather rudimentary software.

Good TWAIN scanners are down in the $50 range, so that shouldn't
be too much of a problem.  Or if your scanner can save in .TIF or
any other "standard" format, that could work, too.

Just load the image into Photoshop, adjust levels (primarily by
dropping the white level below the "grayest" parts of the paper,
and raising the black level above the "grayest" parts of the black
ink).  Leave the mid-tones, since they provide useful information
on letter shape.

Then in the file menu, select "Save for web..." and select GIF format
with, say 4 grayscale values (so a couple of mid-tones are preserved).

You will see a preview of the compressed picture with its compressed
size, and you can play with the parameters to see what works best.

It is a joy to use compared to "cut and try" methods.

Photoshop is one of the finest justifications for owning a computer,
so it's worthwhile to "break the ice".  ;-)

Most of the errors in an OCR documents could probably be caught with a
spellchecker - the hard part comes with character-by-character
proofing the sections of programming code.

You might think so, but most computer documentation is loaded with
peculiar abbreviations and words set off by special fonts.  OCR is
a real bear.  It's enough to give you a new appreciation of the
human eye-brain combination.  ;-)

-michael

8-voice music synthesizer using NadaNet networking!
Home page:  http://members.aol.com/MJMahon/