[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: Syndicomm Scam



<gids.rs@sasktel.net> wrote:

>The OCR isn't perfect

OCR packages that I have collected since starting with the stuff in the 80's 
when you had to replace every unwanted O with a 0 and every unwanted l with 
an L have improved immensely. However they are not all equal and some are 
even not all that good even today.

I have bought so many scanners over the years that these days I just use the 
best ones in my collection like Abby Fine Reader Sprint Plus 6.0, already 9 
years old. A 4-in one MX (MX-430, MX-450, etc) scanner from Canon can be 
bought with ink cartridges for around $50, and I still "swear by" Abby when 
I OCR with my Canon.  I've had the Scanmen, HP's, Lexmarks, and their 
bundles over the years and have done integration for data aquisition using 
Textbridge, Caire and others... programmed sampling applications too in the 
early days of scanners.

Anyway, for computer generated text or text with serifs, OCR packages all 
work pretty well even with crappy images.

I like the OCR packages that output rtf (which I used-to handbomb in my text 
editor to create Windows Help files a couple of decades ago). Rtf just goes 
right into my MSWORD where I resave as a word document and do a spell check 
and reformat if necessary. Then I use OpenOffice Star Writer to open my 
finished Word Files and create PDF files. It's all really simple.

Now as far as digital Cameras go, I've had a bunch over the years and still 
do, Sony, Canon, and even the old Epson Cameras, JVC movie cameras, the 
whole bit... I don't care how many megapixels you have on a digital camera, 
the modern scanner will always produce a consistently high quality scan 
suitable for OCR and the camera will not necessarily. However if you have a 
good OCR package and someone with a scanner does not you will of course 
win... but that's not my point.

For copying bound books you may be right in your case, but its faster and 
better for me, with what little experience I have, to simulate a "fax-style 
aquisition" on the image itself after scanning to a non-lossy format like 
bmp and using mid-pass filters to get rid-of the noise, then re-saving in a 
mono format of around 300 or 600 dpi before OCR'ing.

Did I mention that I prescan and then scan only the area without the shadow 
on a bound book? It takes only seconds on todays fast and wireless machines 
and peripherals.

Anyway if a scanner is working correctly, the sampling is perfect, and 
better than a jittery hand in questionable light, and can't be seriously 
promoted as superior to scanning on a modern flatbed scanner, even with a 
good OCR package. I won't say it's silly to suggest... each to his own:)

Back in the day (80's, early 90's) when I used a graphics arts camera with a 
40 foot long bed to take photos of wallpaper and flooring for 3d rendering 
when I wrote kiosk software simulations for many of the major paint and 
decorating companies using "Targa-Style" and later "True-Color" VGA boards, 
incorrectly shaded areas would produce invalid light source data from 
sampling, even with the studio lighting I used which would produce a 
pillowing effect (dark on the edges of course since color temperature cools 
away from the viewpoint) which needed to be adjusted and evened out 
algorithmically before applying grayscale shadow masks to add light source 
information for accurate 3-d modeling of the colors onto a real digitized 
photo so ladies would buy paint and stuff:)

That tells me something about camera sampling. Edge-detection algorithms too 
worked better with scanned rather than even the highest quality photos. If 
you are familiar with how some of the convolution agorithms work (high-pass, 
low-pass) you will understand also how model look-up and trainable ocr has 
evolved... and ultrasound, wav to midi,  and all the rest of the sample 
image processing that too has evolved.

Anyway that's my 2 cents worth, FWIW. You wondered why anyone would bother, 
now you know why I would bother... I think I've still got a number of 
scanner interface dev-kits somewhere that enable a guy to sample right off 
the head itself, and I am sure I have some model lookup algorithms for text 
recognition hidden away somewhere else here.

But I do have something else handy anyone who is interested can take a look 
at... available from one of my websites:

ReadMe: http://www.teacherschoice.ca/wsquash.txt
DownLoad: http://www.teacherschoice.ca/wsquash.zip

Talks a little about how these filters for sampling work.

Good Day Eh!

Bill