[Martin Taylor 970306 14:40]
Tracy Harms, apparently Wed, 5 Mar 1997 15:26:50 -0700
I'm curious to see more examples of how experienced PCT psych-sorts
critique specific non-PCT theories. ... I nominate the outline of a
psychology problem, and prominent contenders in theory, which I found at
this URL:
http://www.psy.utexas.edu/psy/diehllab/Den_Speech.htm
I'm hoping to provoke something more detailed than "A load of hogwash!" or
its near equivalent, if anybody cares to volunteer the level of attention
that would involve.
I've looked at the page in question. It lists three classes of theory about
speech recognition, and proposed what is claimed to be an advance on one
of them--specifically that analyzing the auditory waveform in the way
the auditory system does, but noting that the information about any one
phoneme is smeared throughout at least a whole word (he omits "at least"),
would lead to improved speech recognition technology.
In my view, he has been ill-taught. What he proposes has been talked about
for many years. Some people still buy it, some reject it on the grounds
that there's no evidence that human-like preprocessing is better than
preprocessing based on the characteristics of the signal. And human-
like processing is computationally expensive. But it isn't
a new idea. As for the "smearing" of the information, just about any
reasonably successful recognition system has to take account of it,
as has been well known for many years.
Where his taxonomy of theories falls down, I think, is that he omits a
whole class of theories, a class in which I think any successful speech
understanding system must fall. That is the class of theories in which
the intention of the talker is a datum along with the acoustic waveform.
There's a hint of this in the "Motor Theory" class, but it's not explicit
there. All of the theories he mentions use a one-way flow of data, from
the speech waveform to the recognized word(s).
A long time ago ARPA (which became DARPA and then ARPA) supported research
in speech recognition that considered the higher levels of "understanding"
but used very crude waveform analysis. To some degree the more successful
programs under this heading might be classed in the group that uses the
talker's intentions (I may be stretching the point a bit; I'm thinking
of syntactic context--if the talker is enquiring about airline schedules
and has said "a flight to" then the next word is likely to be a city
name).
It seems probable to me that a successful speech recognition program will
never be made that uses only acoustic waveform data. My own prejudice is
that a hierarchic control "tracker" of some kind is what will do it. A
control hierarchy will, at many levels of abstraction, control perceptions
that track the values coming from the acoustic waveform analysis. But
the tracker will be working from references based on perceptions of the
talker's intentions. In a way, it's a bit like the Motor Theory, but
in another way it is profoundly different, since what it would track is
not the talker's articulator functions, but the talker's goals, all the
way from, say, whether the talker wanted to get a window opened, down to
whether the talker wanted to purse the lips for 10 msec.
That last paragraph is my prejudice and speculation. And it's not worked out
in detail. It is what we were trying to begin studying in the "Little
Baby" project a couple of years ago.
I don't think the cited URL has much to say about speech recognition,
either in the classical context, or in the context of PCT.
Martin