[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]
Re: Syndicomm Scam
<gids.rs@sasktel.net> wrote:
>The OCR isn't perfect
OCR packages that I have collected since starting with the stuff in the 80's
when you had to replace every unwanted O with a 0 and every unwanted l with
an L have improved immensely. However they are not all equal and some are
even not all that good even today.
I have bought so many scanners over the years that these days I just use the
best ones in my collection like Abby Fine Reader Sprint Plus 6.0, already 9
years old. A 4-in one MX (MX-430, MX-450, etc) scanner from Canon can be
bought with ink cartridges for around $50, and I still "swear by" Abby when
I OCR with my Canon. I've had the Scanmen, HP's, Lexmarks, and their
bundles over the years and have done integration for data aquisition using
Textbridge, Caire and others... programmed sampling applications too in the
early days of scanners.
Anyway, for computer generated text or text with serifs, OCR packages all
work pretty well even with crappy images.
I like the OCR packages that output rtf (which I used-to handbomb in my text
editor to create Windows Help files a couple of decades ago). Rtf just goes
right into my MSWORD where I resave as a word document and do a spell check
and reformat if necessary. Then I use OpenOffice Star Writer to open my
finished Word Files and create PDF files. It's all really simple.
Now as far as digital Cameras go, I've had a bunch over the years and still
do, Sony, Canon, and even the old Epson Cameras, JVC movie cameras, the
whole bit... I don't care how many megapixels you have on a digital camera,
the modern scanner will always produce a consistently high quality scan
suitable for OCR and the camera will not necessarily. However if you have a
good OCR package and someone with a scanner does not you will of course
win... but that's not my point.
For copying bound books you may be right in your case, but its faster and
better for me, with what little experience I have, to simulate a "fax-style
aquisition" on the image itself after scanning to a non-lossy format like
bmp and using mid-pass filters to get rid-of the noise, then re-saving in a
mono format of around 300 or 600 dpi before OCR'ing.
Did I mention that I prescan and then scan only the area without the shadow
on a bound book? It takes only seconds on todays fast and wireless machines
and peripherals.
Anyway if a scanner is working correctly, the sampling is perfect, and
better than a jittery hand in questionable light, and can't be seriously
promoted as superior to scanning on a modern flatbed scanner, even with a
good OCR package. I won't say it's silly to suggest... each to his own:)
Back in the day (80's, early 90's) when I used a graphics arts camera with a
40 foot long bed to take photos of wallpaper and flooring for 3d rendering
when I wrote kiosk software simulations for many of the major paint and
decorating companies using "Targa-Style" and later "True-Color" VGA boards,
incorrectly shaded areas would produce invalid light source data from
sampling, even with the studio lighting I used which would produce a
pillowing effect (dark on the edges of course since color temperature cools
away from the viewpoint) which needed to be adjusted and evened out
algorithmically before applying grayscale shadow masks to add light source
information for accurate 3-d modeling of the colors onto a real digitized
photo so ladies would buy paint and stuff:)
That tells me something about camera sampling. Edge-detection algorithms too
worked better with scanned rather than even the highest quality photos. If
you are familiar with how some of the convolution agorithms work (high-pass,
low-pass) you will understand also how model look-up and trainable ocr has
evolved... and ultrasound, wav to midi, and all the rest of the sample
image processing that too has evolved.
Anyway that's my 2 cents worth, FWIW. You wondered why anyone would bother,
now you know why I would bother... I think I've still got a number of
scanner interface dev-kits somewhere that enable a guy to sample right off
the head itself, and I am sure I have some model lookup algorithms for text
recognition hidden away somewhere else here.
But I do have something else handy anyone who is interested can take a look
at... available from one of my websites:
ReadMe: http://www.teacherschoice.ca/wsquash.txt
DownLoad: http://www.teacherschoice.ca/wsquash.zip
Talks a little about how these filters for sampling work.
Good Day Eh!
Bill