[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

SPEECH SYNTHESIS teaching lab exercise



               TEACHING LAB EXERCISE ON SPEECH SYNTHESIS

I have just written a lab exercise for my Sensation/Perception (psychology)
class.  The exercise is based on the Echo II speech synthesizer card.  This
afternoon, I tried it with college students and it worked GREAT.  One of the
best labs this semester.  Results turned out just as expected.

Why mention this to the A2 forum?  First, to affirm that original materials
are still being produced.  Second, those who have the Echo II speech interface
card may wish to get the packet from me.  For $1.00 I will send the following:
(a) Paper copy of the seven page lab handout (with text file on flip side of
the disk),  (b) a 5.25" self-booting disk with the Echo II software and also
the stimulus files for the lab exercise,  (c) photocopy of the two-part
article on speech synthesis and speech recognition (Stephen Rosenthal, A+
Magazine, May 1985, 3(5), 34-44), diskette mailer and postage.

If anyone wishes to see the complete lab handout, it is a seven page single-
spaced text file.  I will be happy to transmit by e-mail to requesting
individuals.

What follows is the introduction section of the paper as a sampler.
Transmit to me directly via e-mail if you want the $1.00 mailed packet
or the transmission of the full length text-file.

NOTE: If anyone out there has an Echo II speech card that they wish to
donate to an educational institution (this college!), please contact me
via e-mail!  I will guarantee that it will be used productively. sebugg

00000000000000000000000000000000000000000000000000000000000000000000000000
00000000000000000000000000000000000000000000000000000000000000000000000000

Perception 407                                                    Dec 1993
Presbyterian College                                           Page 1 of 7


       SPEECH SYNTHESIS AND TOP-DOWN PROCESSING IN SPEECH RECOGNITION


                               Stephen Buggie


       How intelligent is artificial intelligence (AI)?  One aspect of AI is
computerized speech.  Speech synthesis, analogous to human expressive/
productive speech, is a familiar part of modern daily experience.   We are
acquainted with talking clocks, calculators, vending machines, control
panels, teaching toys, or other products using microprocessor speech
technologies.

       Today's lab exercise will introduce speech synthesis.  An Apple II
computer equipped with text-to-speech hardware and software will generate
understandable speech.  Unlike the limited vocabulary of other electronic
talkers (watches, calculators, etc.), the Apple's Echo II interface hardware
can read an unlimited variety of text files,  but with limited accuracy!
The Echo II and related speech synthesizers reproduce the phonemes of
English within the constraints of a few hundred pronunciation rules.  But
the exceptions are spoken with occasional errors, resulting in a distinctly
audible "accent."

       The other aspect of speech is speech recognition, known also as
receptive or comprehensive speech.  The technology of computerized speech
recognition is less advanced than is the case for speech synthesis.
Automated speech recognition is limited in several ways: (a) Recognition
vocabulary size is limited to no more than several hundred words;  (b) The
computer must be "enrolled;" that is, it must be given in advance samples of
the particular speaker's voice;  (c) The computer requires distinct pauses
or breaks between words to be recognized, a condition that is unnecessary
for fluent human listeners.

       Computerized speech recognition will not be demonstrated in this lab
exercise, although details of this technology are available (Rosenthal,
1985, Part One).  Instead, natural (human) recognition of synthetic speech
will be demonstrated and evaluated.

       Human speech comprehension is superior to computerized speech
recognition in the inference of meaning from incomplete speech stimuli.
Computerized speech recognition is based on auditory template matching, a
primitive cognitive process analogous to trying pieces in a jigsaw puzzle
one at a time until a match is achieved.  The template of each vocabulary
word is tried, and the stimulus is identified from the best matching
template.  But template matching has disadvantages: it is slow when the
vocabulary is large, and it especially is prone to failure when the stimulus
quality is poor or degraded by background.  For example, automated speech
recognition is greatly handicapped if the stimulus is masked by other
sounds, especially if the background sounds are voices.  A computer would
have a rough time listening to a conversation in a crowded party situation!

       Human listeners' superiority over computers in speech recognition
skills is attributed to the fact that humans effectively apply top-down
processing to extract meaning from partially heard speech messages.
Top-down processing includes the listener's expectations, attitudes, past
experiences, or knowledge of the speech message's context.   Listeners'
perceive better if the incoming message is anticipated to be speech, if the
conversational topic is known, or if the speaker's lip movements are
visible.  Phoneme recognition is aided if the phonemes appear in the context
of familiar words, and words are better perceived when they appear in
sentence contexts.  Even if a phoneme is missing or masked by other sounds,
the listener may use the word or sentence cont