articulation and sound

[From Bill Powers (2003.08.05.0557 MDT)]

This is mostly for Bruce Nevin (and any other linguists), who I hope will
check my conclusions here and maybe even agree.

Practicing making the sounds of Chinese has confused me because it's not
clear which parts of the task involve control of sound and which involve
control of articulation. Obviously -- it seems -- the articulations are
what produce the sounds, so the sounds should be hierarchically above the
articulations. This morning I realized that this is not true.

I now think that the muscles produce both articulations and sounds, in
parallel and simultaneously. Muscles squeeze the diaphragm to create air
pressure and flow; they tighten the vocal cords to produce sound if the
diaphragm is producing air flow; they move the tongue, jaw, lips, and small
flaps here and there to create sensations of pressure and to vary the
sounds made by air moving through cavities and constrictions. This all
happens, I now believe, no higher than the configuration order in our
perceptions. Of course changes in these configurations then produce
transitions and transitions are chained to produce events that we call
words, but when sound is first produced and shaped, the perception appears
now to be that of a configuration of simultaneous sensations of different
kinds: articulatory and auditory.

I was led to this in part by seeing the author of an articulation diagram
apparently making the same mistake I made about 55 years ago when it cost
me a job. When I was a freshman at the University of Colorado (before
transferring to Northwestern the next year), I auditioned for a part-time
job as an announcer at a local radio station. I was rejected because I
could not say "ladies" correctly when reading a commercial message. I tried
again and again, but every time, the "s" hissed loudly into the microphone
and the person interviewing me said, "No, no, it's _ladies_" with the S
sounding like a Z. I tried everything, but couldn't get the tip of my
tongue to vibrate and make the Z sound at the end of the word.

Reading the articulation diagram I mentioned yesterday, I noticed that the
author described the sound written "r" in Pinyin as being made by holding
the tip of the tongue close to the roof of the mouth. He said, "The tip of
the tongue vibrates." I tried it, and was reminded immediately of the
interview at the radio station when I tried, in vain, to make the tip of my
tongue "vibrate" to make a Z sound. What this author and I both failed to
understand was that the felt vibration was not a cause but an effect. It
was the effect of the fundamental frequency of the buzz generated by the
vocal cords. If I had only known, I could have let my voice continue
through the "S" of "ladies", and I might have got the job. I could have
been a very old radio announcer by now, darn it.

What this boils down to is that in order to make the Z sound, whether
forward in the mouth or further inside, the muscles that must be used are
not in the tip of the tongue but in the throat where they tighten the vocal
cords. I didn't use the right muscles, so I couldn't make the sound/feel of
a Z at the end of a word.

Maybe the last piece of the puzzle fell into place when I thought that the
goal of my little project was to look at Pinyin letters and hear them make
the correct Chinese sounds, the way we seem to hear printed words in
imagination. Or _see_ a mouth formed a certain way and _imagine_ the sound,
or _hear_ the sound and _imagine_ how your mouth feels when you make it,
and perhaps _imagine_ the printed word at the same time. This showed me
that all these elements of speech are at the same level. We can perceive
some of them while imagining the rest, or produce sound and articulation at
the same time with or without reading the actual words.

The physical fact is (to sum up) that motor actions of speech make both
perceived articulations and perceived sounds through physical interactions
of flowing air and cavity shapes. Neither sound nor articulation is prior
to the other. They are both at the same level of perception.

Bruce N, what do you say?

Best,

Bill P.

[From Rick Marken (2003.08.05.1105)]

Bill Powers (2003.08.05.0557 MDT)--

I now think that the muscles produce both articulations and sounds, in
parallel and simultaneously.

What this author and I both failed to
understand was that the felt vibration was not a cause but an effect.

I guess it depends on what is meant by "articulation". If it refers to the shape
of the vocal tract then, in the case of the felt vibration of the tongue, it's not
really the articulation that is perceived but, as in the case of sound, a
consequence of the articulation. As you say, the felt vibration was "...the effect
of the fundamental frequency of the buzz generated by the vocal cords". Voicing is
considered an articulation, I think (voiced/unvoiced), so the relevant perception
controlled when making the "s" sound is a consequence of the articulation. I think
you are right that non-sound perceptual consequences of speech muscle actions
(like the tongue vibration) are part of the speech perception that is controlled
when we speak. One evidence for this, it seems to me, is the difficulty one has
saying certain things after a Novocain injection, which just takes out sensory
perception from some part of the mouth. The motor functions are all intact as are
the auditory sensory functions. So the difficulty must result from lack of
ability to produce some of the the non-sound perceptual consequences of speaking
actions.

Interesting!

Best

Rick

···

---
Richard S. Marken, Ph.D.
Senior Behavioral Scientist
The RAND Corporation
PO Box 2138
1700 Main Street
Santa Monica, CA 90407-2138
Tel: 310-393-0411 x7971
Fax: 310-451-7018
E-mail: rmarken@rand.org

[From Rick Marken (2003.08.06.1120)]

Bruce Nevin (2003.08.06 13:08 PDT)
Bill Powers (2003.08.05.0557 MDT)–

Neither sound nor articulation is prior to the
other. They are both at the same level of perception.
Yes. But references for articulations are calibrated according to their
affect on sounds, not vice versa.
I think this is right: articulations are controlled as the means of controlling
sound. So I disagree with Bill: I don’t think sound and articulation are
at the same level of control. I think the sounds of speech are at
a higher level of control than articulation, with articulations controlled
as the lower level means of controlling sound.
I have been doing some speech experiments on myself since Bill posted
his observations about learning Chinese and my subjective impression is
that I don’t control for articulatory patterns at all when I talk. I can
barely tell what the hell my mouth (oral cavity, tongue, palates, etc)
is doing when I speak. I think I just do whatever I have leaned to do with
my vocal apparatus in order to make sounds like “pot 'o coffee” and “Paraguay”
(great examples, by the way, Bruce).

The most convincing demonstration (to me) of articulation as a means
of producing intended sound patterns comes from pressing my tongue against
the bottom of my mouth (way below my front teeth; it feels like it must
be pressing against my soft palate) and then trying to say words like “Paraguay”.
When I’m affecting a Spanish accent I roll the r in “Paraguay”. But
I can’t do that with my tongue pressed against the bottom of my mouth.
But I can get something like the sound of the rolled r by making a guttural
roll in my throat. I came up with this articulatory compensation for the
lack of tongue rolling “automatically” in an effort to hear “Paraguay”
in something like the way I usually hear it. And, indeed, I am able
to hear something remarkably close to the “rolled r” version of “Paraguay”
using the guttural substitute.

What I see here is a completely different articulation (guttural roll)
being used to produce an intended sound: “Paraguay” with rolled r. This
makes me think that perceptions of articulation are not parallel controlled
components of the words we speak. Articulations are, I believe, varied
as necessary to produce the sounds we intend to produce. Since there articulations
are typically observed when no disturbances (like food in the mouth) are
present, it looks like particular articulations are used to produce particular
sounds. But I think this apparent one-to-one mapping of action (articulation)
to result (sound) is an artifact of observing speech in conditions where
the need for variation in action is reduced or eliminated.

Best

Rick

···

Richard S. Marken, Ph.D.

Senior Behavioral Scientist

The RAND Corporation

PO Box 2138

1700 Main Street

Santa Monica, CA 90407-2138

Tel: 310-393-0411 x7971

Fax: 310-451-7018

E-mail: rmarken@rand.org

[From Bruce Nevin (2003.08.06 13:08 PDT)]

Bill Powers (2003.08.05.0557 MDT)–

This is mostly for Bruce Nevin (and any other
linguists), who I hope will

check my conclusions here and maybe even agree.

First, about z vs. s, and both vs. r. The difference between z and s is
the same as the difference between v and f, between p and b, and so
on.
voiceless p t k f
s sh …
voiced b
d g v z zh (as in azure) …
The conclusion that you drew, those many years ago, that the difference
was due to the tongue vibrating, was incorrect. Whatever sensation of
vibration that you may be aware of at the tongue tip is merely
transmitted from the pressure wave in the air stream, due to vibration of
the vocal folds.
There is less tension in the occlusion with the tongue for z than there
is for s. There are two reasons. First, the tension is a function of air
pressure, being the critical amount to cause sufficient turbulence
without stopping the air flow, and there is less air pressure when the
vocal folds interrupt the air flow. This is generally true of voiced
consonants vs. voiceless. Secondly, it takes less constriction behind the
gingival ridge to produce audible turbulence when the airstream is
already articulated with the pressure wave of voicing.
In American English, at the end of a word like ladies, z tends to
be devoiced, but is still distinct from s because (as just noted) it is
more lax. Say lates’ answer (“latest”, without the final
t), then say ladies and gentlemen. You have to kind of sneak up on
yourself to pronounce these naturally while listening acutely. For
broadcasters, particularly with the microphones of that time, there must
be a deliberate and artificial-feeling effort to fully voice the z so
that it doesn’t sound too sibilant. Say ladies again, normally.
Say a zzz sound as in imitating the buzzing of bees. Then say
ladies again, only end with a brief segment of that zzzsound. ladiezz. It feels and sounds funny. Almost like three
syllables.

The sound written [r], on the other hand, is a trill, produced indeed
with the tip of the tongue vibrating. The effect is primarily
aerodynamic, with the tongue in appropriate configuration and tension
(just enough flexibility & slenderness at apex of tongue) and with
sufficient airflow to maintain a trill. However, electromyographic
studies indicate there may be more muscle activity than is warranted just
to maintain configuration and tension. It’s pretty hard to tease these
apart. Maintaining the tongue configuration against turbulence requires
efforts very similar to causing brief (partial) interruptions to airflow.
The process of learning to produce a trilled r seems like creating pulses
with the tongue. Start with a flap, e.g. the medial t in

    pot 'o

coffee

    Paraquay

    Butter

    barato

(Spanish: “cheap”)

These are examples of a flapped r in English that is indistinguishable
from a flapped r in Spanish. A flap is one tap of the tongue tip.
Consider also the indistinguishability of ladder and latter in
unemphasized speech. For a trill as in e.g. Spanish perro “dog”
burro, barra there’s a series of taps of the tongue tip on the gingival
ridge interrupting the air flow. You have to learn to thin your tongue
tip and hold it appropriately so that it flaps in the air stream.
Remember learning to whistle?

Although this is not what is going on with z vs. s, they do have acoustic
and articulatory similarities such that in the history of languages and
in dialect differentiation there are frequent changes from r to z or s
and vice versa, so there is a solid basis, and precedent as it were, for
your difficulties as a job applicant.

Practicing making the sounds of Chinese has
confused me because it’s not

clear which parts of the task involve control of sound and which
involve

control of articulation. Obviously – it seems – the articulations
are

what produce the sounds, so the sounds should be hierarchically above
the

articulations. This morning I realized that this is not
true.

Yes, control of acoustic inputs is not hierarchically above control of
articulation, it is in parallel; but nonetheless, references for
articulations are calibrated according to their affect on sounds, and not
vice versa. (This point comes up again below, so if you have thoughts
about it that might be a better place to put them.)

I now think that the muscles produce both
articulations and sounds, in

parallel and simultaneously.

Yes, clearly. But when you hear the sound of t it is too late to amend it
to p, you repronounce the syllable or the word or even the entire phrase.

Muscles squeeze the diaphragm to create
air

pressure and flow; they tighten the vocal cords to produce sound if
the

diaphragm is producing air flow; they move the tongue, jaw, lips, and
small

flaps here and there to create sensations of pressure and to vary
the

sounds made by air moving through cavities and constrictions. This
all

happens, I now believe, no higher than the configuration order in
our

perceptions.

Yes.

Of course changes in these configurations then
produce

transitions and transitions are chained to produce events that we
call

words,

Also syllables, and in parallel morphemes that are combined (like dis-
and place and -ing), and in parallel words (like place and displacing).

but when sound is first produced and shaped,
the perception appears

now to be that of a configuration of simultaneous sensations of
different

kinds: articulatory and auditory.

[… example discussed above.]

Maybe the last piece of the puzzle fell into
place when I thought that the

goal of my little project was to look at Pinyin letters and hear them
make

the correct Chinese sounds, the way we seem to hear printed words in

imagination. Or see a mouth formed a certain way and imagine the
sound,

or hear the sound and imagine how your mouth feels when you make
it,

and perhaps imagine the printed word at the same time. This showed
me

that all these elements of speech are at the same level. We can
perceive

some of them while imagining the rest, or produce sound and articulation
at

the same time with or without reading the actual words.

This sounds like your conception of category recognition. It also sounds
like associative memory.

The physical fact is (to sum up) that motor
actions of speech make both

perceived articulations

Well, mostly sensations of touch and pressure, maybe configuration for
jaw angle, but tongue configuration seems very tough to sense - all those
muscles in there!.

and perceived sounds through physical
interactions

of flowing air and cavity shapes. Neither sound nor articulation is
prior

to the other.

If you mean hierarchically prior, yes (see above). If you mean temporally
prior, it depends. Closure for e.g. /t/ is felt before release of that
closure, which (word-initially) is the first point in time that any sound
is heard.

They are both at the same level of
perception.

Yes. But references for articulations are calibrated according to their
affect on sounds, not vice versa. As a learning process, the calibrating
is done over series of utterances, like setting a sight.

There’s more, of course, concerning dialects & ‘accents’, unconscious
mimickry, acting and ‘passing’ and otherwise taking on a persona of one
kind of person or another, retention of a ‘foreign’ accent after years,
etc. These matters involve higher levels of control (probably related to
self image) parallel to higher levels of control for the informational
aspects of understanding and formulating speech or writing. But for now
that invites confusion.

    /Bruce

Nevin

···

At 09:09 AM 8/5/2003, Bill Powers wrote:

[From Bill Powers (2003.08.06.1132 MDT)]

Bruce Nevin (2003.08.06 13:08 PDT)]--

Thanks for the long, and as usual lucid, reply,

First, about z vs. s, and both vs. r. The difference between z and s is
the same as the difference between v and f, between p and b, and so on.

voiceless p t k f s sh ...
voiced b d g v z zh (as in azure) ...

Yes, once I learned about voicing I figured all this out. Weird that the
answer didn't occur to me for so long.

There is less tension in the occlusion with the tongue for z than there is
for s. There are two reasons.

Makes sense, but I'm mostly concerned here, for obvious reasons, with
subjectively observable facts.

The sound written [r], on the other hand, is a trill, produced indeed with
the tip of the tongue vibrating.

I don't think this applies to the letter written [r] in Pinyin. It's a
strange sound, because it comes from much farther back in the mouth than
the gingival (thanks for the word) ridge. According to the articulation
diagram, the tip of the tongue points straight up or even a little toward
the back, near the roof of the mouth. There's nothing to flap against, as
with the forward [r]. The result is to make the usual hisses and ch sounds
more hollow than when you say them in front.Just compare a gingival sh with
the same sound made at the roof of the mouth. There are lots more high
frequencies in the frontal version. The Pinyin [r] is the voiced version of
the ch made in the same (back) position, as near as I can figure. You can
indeed feel a tongue vibration from your voice, but it's not a rolled r.
You, Bruce, could probably get a lot closer to a correct description by
reading technical descriptions that I wouldn't understand. How about that
book on how to pronounce the world's languages that you brought to the
meeting? What was the title?

Yes, control of acoustic inputs is not hierarchically above control of
articulation, it is in parallel; but nonetheless, references for
articulations are calibrated according to their affect on sounds, and not
vice versa. (This point comes up again below, so if you have thoughts
about it that might be a better place to put them.)

...

If you mean hierarchically prior, yes (see above). If you mean temporally
prior, it depends. Closure for e.g. /t/ is felt before release of that
closure, which (word-initially) is the first point in time that any sound
is heard.

They are both at the same level of perception.

Yes. But references for articulations are calibrated according to their
affect on sounds, not vice versa. As a learning process, the calibrating
is done over series of utterances, like setting a sight

I think the distinction we're both trying to get at here is whether the
priority is determined perceptually or physically. There are certain moves
that have to be made before some sounds can occur; that's a matter of
physics. But there's no need for the perceptual systems to know the
physics. If you try to say /t/ without holding the closure briefly, you'll
just get a different sound. Learning of this sort isn't logical. You just
keep changing things until the result is what you intended. Physical
explanations tend to take us to too high an order of perception. I don't
think they describe what we learn at the lower levels.

I wrote to Isaac that I think the relationship between kinesthetic
articulation and auditory speech is much like the relationship between
kinesthetic and visual space in motor behavior. The perceptual systems
adapt -- "calibrate" -- to each other, slowly. The physical facts get into
the act, but only indirectly, by making some perceptions depend on others.

I think that one could actually start making sense of all this, in time.

By the way, I still recall your presentation at the meeting with pleasure.
What an organized mind you have, and how well you grasp your subject! Your
talk was as much art as science.

Best,

Bill P.

[From Bill Powers (2003.08.06.1208 MDT)]

Bruce Nevin (2003.08.06 13:08 PDT) --

Typo. I meant that the Pinyan [r] is the voiced version of the _sh_ sound
made at the roof of the mouth -- not ch.

Besg,

Bill P.

[From Bruce Nevin (2003.08.06 16:00 PDT)]

Bill Powers (2003.08.06.1132 MDT)–

Bruce Nevin (2003.08.06 13:08 PDT)–

The sound written [r], on the other hand, is a
trill, produced indeed with

the tip of the tongue vibrating.

I don’t think this applies to the letter written [r] in Pinyin. It’s
a

strange sound, because it comes from much farther back in the mouth
than

the gingival (thanks for the word) ridge. According to the
articulation

diagram, the tip of the tongue points straight up or even a little
toward

the back, near the roof of the mouth. There’s nothing to flap against,
as

with the forward [r]. The result is to make the usual hisses and ch
sounds

more hollow than when you say them in front. Just compare a gingival sh

s is gingival (you don’t speak a dialect where s, z, t, d are fronted to
postdental) and sh is behind that, palatal.

with the same sound made at the roof of the
mouth. There are lots more high

frequencies in the frontal version.

Right, lower as the point of occlusion moves back. Lip rounding may
contribute if you are not careful. Smile and move from s through sh to
Pinyin r (below) in a continuum. Then produce any of these sounds while
broadening and rounding your lips in a continuum. You’ll hear the
perceived pitch of the turbulence move up and down.

The Pinyin [r] is the voiced version of

the [sh] made in the same (back) position, as near as I can figure. You
can

indeed feel a tongue vibration from your voice, but it’s not a rolled
r.
[Your correction inserted.]
Ah yes, my mistaken assumption. You’re right. Pinyin r represents what is
called a retroflex articulation, with the tongue tip curled up toward the
roof of the mouth. In English, we have it in a vowel sound, as in the
final syllable of butter (unless you pronounce this as r-less
‘buttah’), and it functions as a consonant if it is a ‘glide’ transition
between syllables, e.g. around. (Some speakers of English
pronounce r with the dorsum [“back”, main body] of the tongue
humped toward the roof of the mouth, rather than the tip retroflexed) and
they get the acoustic effect of lowering the ‘pitch’ with lip rounding as
for w.)
In Mandarin Chinese, it’s a fricative in the same family as z. In
the phonetic transcription alphabet of the International Phonetic
Association (IPA) it’s in the column of retroflex consonants
<http://www.arts.gla.ac.uk/IPA/pulmonic.html>,
the z with a little tail below to the right. One approach to producing it
is to pronounce English retroflex r and then close the occlusion until
you get the sound of frication. At
<http://www.wikipedia.org/wiki/Pinyin#Pronunciation>
they say that it is “similar to the English r in ‘rank’ with a bit
of the initial sound in French ‘journal’ in it”.

But it’s also a vowel quality in Mandarin. The same source says that
after a preceding vowel (ar, er, ) it’s similar to English vowel plus
retroflex r – obviously for those dialects of English that pronounce r
after vowels – except that the retroflexion extends through the vowel
sound: “ar: like a, but pronounced with the tongue curled up against
the palate; like rhotic are in North American English. … er: like e,
but pronounced with the tongue curled up against the palate; similar to
the vowel in rhotic her in English.” Other vowels are not mentioned,
so maybe its distribution is restricted.

The difference is probably in whether it’s syllable-initial (consonant)
or syllable-final (quality of the preceding vowel).

[…]

that

book on how to pronounce the world’s languages that you brought to
the

meeting? What was the title?

Ladefoged, Peter, & Ian
Maddieson, The sounds of the world’s languages, Blackwell,
1996.

[…]

I think the distinction we’re both trying to
get at here is whether the

priority is determined perceptually or physically. There are certain
moves

that have to be made before some sounds can occur; that’s a matter
of

physics. But there’s no need for the perceptual systems to know the

physics. If you try to say /t/ without holding the closure briefly,
you’ll

just get a different sound. Learning of this sort isn’t logical. You
just

keep changing things until the result is what you intended. Physical

explanations tend to take us to too high an order of perception. I
don’t

think they describe what we learn at the lower levels.

Right. We have two distinct questions: what perceptions are controlled,
and how is control learned.

I wrote to Isaac that I think the relationship
between kinesthetic

articulation and auditory speech is much like the relationship
between

kinesthetic and visual space in motor behavior. The perceptual
systems

adapt – “calibrate” – to each other, slowly. The physical
facts get into

the act, but only indirectly, by making some perceptions depend on
others.

That’s interesting, and yes does seem apt. Is that mutual calibration in
each perceptual modality according to the other demonstrated and
documented in Isaac’s literature?

    /Bruce

Nevin

···

At 02:03 PM 8/6/2003, Bill Powers wrote:

[From Bill Powers (2003.08.07.0650 MDT)] --

I have been doing some speech experiments on myself since Bill posted his
observations about learning Chinese and my subjective impression is that I
don't control for articulatory patterns at all when I talk. I can barely
tell what the hell my mouth (oral cavity, tongue, palates, etc) is doing
when I speak. I think I just do whatever I have leaned to do with my vocal
apparatus in order to make sounds like "pot 'o coffee" and "Paraguay"
(great examples, by the way, Bruce).

Before you come to any final conclusions, I recommend doing a lot more
experimenting. What I have found is that much of what I have called speech
"sounds" consists of kinesthetic sensations, not auditory ones. The buzzing
sound of Z, for example, is felt in the tip of the tongue! Also, say "the"
a few times and examine the "th" sound. It sounds partly like your tongue
between your teeth. That's not a sound.

I think what we need are some other examples of "instrumental" controlled
variables; that is, variables that are controlled in order to control other
variables at the same level. Example: I curl my fingers around a glass in
order to hold it in position for drinking.The behavior of the fingers and
the glass seems related in the same way the behavior of the tongue is
related to saying "th" This perhaps ought to be cast in terms of several
levels of perception in one modality being related to control of several
levels of perception in another modality, because there are sensations,
configurations, and transitions (etc) in both modalities at once, but the
first set is controlled to have controlled effects on the second set.

Another: I set the temperature of my shower by pulling the faucet knob out
and turning it to a specific angle. I know that in a minute or so, the
temperature will be close to right. Then, of course, I turn the handle
while feeling the water, because it's never _exactly_ right. In a similar
vein, turning a knob on a gas stove to a specific number or mark and only
then looking at the flame to finish adjusting it. One configuration being
controlled because of its effect on another controlled configuration.

This is how I'm thinking about speech: a kinesthetic modality being
controlled because of its effects on the auditory modality _at the same
level_. Isaac Kurtzer sent me an abstract of a paper about an experiment in
which articulatory variables were disturbed in a way that had no measurable
or sensible effects on sounds. The paper reported that the disturbance was
resisted. Of course in a hierarchical model that doesn't prove that
articulation is the primary controlled variable (as the authors thought).
But it does suggest doing the opposite experiment, varying the sound
without changing the articulation (altering phonemes before they are heard,
and so forth). I've tried to get somewhere with that but have had
programming difficulties.

I think there's no doubt that in speech the highest priority variable is
auditory. But this does not make it a higher level variable, unless we can
show that configurations in the sounds of speech are functions of
articulatory sensations, which I'm pretty sure is not the case.

There are some articulations, as you say, that change when disturbed so as
to preserve the sound being made. But with this new consideration of
_instrumentality_, this does not have to mean that the auditory variables
are of higher level. If your chin is held down against a rest, your head
goes up and down to say "mama". You can also say "mama" with your lips
remaining slightly opened and using a crosswise finger to block and open
the gap. There are physical interactions amoung different effects of
controlled variables, and we use them to produce the experiences we want.
If one aspect of the situation changes we can usually find an alternate
means of creating the same effect. This doesn't mean there is a
hierarchical relationship among the various means at the same level.

I don't mean to be dogmatic here. I'm just exploring an idea.

Best,

Bill P.

[From Bill Powersa (2003.08.07.0734 MDT)]

Bruce Nevin (2003.08.06 16:00 PDT)--

In Mandarin Chinese, [r is] a fricative in the same family as z. In the
phonetic transcription alphabet of the International Phonetic Association
(IPA) it's in the column of retroflex consonants
<http://www.arts.gla.ac.uk/IPA/pulmonic.html&gt;, the z with a little tail
below to the right. One approach to producing it is to pronounce English
retroflex r and then close the occlusion until you get the sound of
frication. At <Pinyin - Wikipedia; they
say that it is "similar to the English r in 'rank' with a bit of the
initial sound in French 'journal' in it".

Thanks, those were useful references. especially the second one.

But it's also a vowel quality in Mandarin.

Haven't run across that yet. The source did list ar in the vowel column. I
think the classifcation is sort of optional -- another source says that
only the consonants n, ian, and r can occur at the end of Chinese
syllables, making r into a consonant. Of course we pronounce the sound, not
the classification.

I have to remind myself not to spend all my time talking about Mandarim in
English as opposed to practicing it.

Best,

Bill P

[From Rick Marken (2003.08.07.1020)]

Bill Powers (2003.08.07.0650 MDT) --

>I have been doing some speech experiments on myself since Bill posted his
>observations about learning Chinese and my subjective impression is that I
>don't control for articulatory patterns at all when I talk.

Before you come to any final conclusions, I recommend doing a lot more
experimenting.

Ok. Will do. The research equipment (my mouth) is really quite inexpensive, even
amortizing the dental costs (mainly cleaning, luckily) :wink:

What I have found is that much of what I have called speech
"sounds" consists of kinesthetic sensations, not auditory ones. The buzzing
sound of Z, for example, is felt in the tip of the tongue! Also, say "the"
a few times and examine the "th" sound. It sounds partly like your tongue
between your teeth. That's not a sound.

It doesn't "sound" like that to me. It just sounds like "th". I can even make
that sound with my tongue starting behind my top front teeth; my tongue never goes
between my teeth and I still hear "th".

I think what we need are some other examples of "instrumental" controlled
variables; that is, variables that are controlled in order to control other
variables at the same level. Example: I curl my fingers around a glass in
order to hold it in position for drinking.The behavior of the fingers and
the glass seems related in the same way the behavior of the tongue is
related to saying "th" This perhaps ought to be cast in terms of several
levels of perception in one modality being related to control of several
levels of perception in another modality, because there are sensations,
configurations, and transitions (etc) in both modalities at once, but the
first set is controlled to have controlled effects on the second set.

Yes.

Another: I set the temperature of my shower by pulling the faucet knob out
and turning it to a specific angle. I know that in a minute or so, the
temperature will be close to right. Then, of course, I turn the handle
while feeling the water, because it's never _exactly_ right. In a similar
vein, turning a knob on a gas stove to a specific number or mark and only
then looking at the flame to finish adjusting it. One configuration being
controlled because of its effect on another controlled configuration.

This is how I'm thinking about speech: a kinesthetic modality being
controlled because of its effects on the auditory modality _at the same
level_.

Interesting possibility. But I think it's more important to get the hierarchical
relationships right before worrying about whether the variables involved are at
the same level perceptually, according to the PCT hierarchy. I think what we have
to do is build a model of speech (sound) production and then test it. If it turns
out that the model works best when kinesthetic perceptions are controlled at the
same level as sound perceptions then that is evidence that, indeed, speech
involves the simultaneous control of sound and kinesthetic sensations. The tests
of such a model would, of course, use PCT methodology (test for the controlled
variable). But I think a model is needed before we can get anywhere on this.

Isaac Kurtzer sent me an abstract of a paper about an experiment in
which articulatory variables were disturbed in a way that had no measurable
or sensible effects on sounds. The paper reported that the disturbance was
resisted. Of course in a hierarchical model that doesn't prove that
articulation is the primary controlled variable (as the authors thought).

Of course not. You'll resist disturbances to your running pattern when you run to
catch a ball but that doesn't mean that the running pattern is the primary
controlled variable when catching a ball.

But it does suggest doing the opposite experiment, varying the sound
without changing the articulation (altering phonemes before they are heard,
and so forth). I've tried to get somewhere with that but have had
programming difficulties.

Right! This is the test we need. Someone must have done this. I would be surprised
if people did not change their articulation to compensate for changes in the
perceptual result of those articulations.

I think there's no doubt that in speech the highest priority variable is
auditory. But this does not make it a higher level variable, unless we can
show that configurations in the sounds of speech are functions of
articulatory sensations, which I'm pretty sure is not the case.

I don't understand this. Isn't it likely to be necessary to vary articulatory
sensations (doing which has particular physical consequences) as the means of
producing various sound sensations (which are a physical consequence of the
physical correlates of the articulatory sensations)?

There are some articulations, as you say, that change when disturbed so as
to preserve the sound being made. But with this new consideration of
_instrumentality_, this does not have to mean that the auditory variables
are of higher level. If your chin is held down against a rest, your head
goes up and down to say "mama". You can also say "mama" with your lips
remaining slightly opened and using a crosswise finger to block and open
the gap. There are physical interactions amoung different effects of
controlled variables, and we use them to produce the experiences we want.
If one aspect of the situation changes we can usually find an alternate
means of creating the same effect. This doesn't mean there is a
hierarchical relationship among the various means at the same level.

I think I need to see a model in order to understand this.

I don't mean to be dogmatic here. I'm just exploring an idea.

Feel free;-)

Best

Rick

···

--
Richard S. Marken, Ph.D.
Senior Behavioral Scientist
The RAND Corporation
PO Box 2138
1700 Main Street
Santa Monica, CA 90407-2138
Tel: 310-393-0411 x7971
Fax: 310-451-7018
E-mail: rmarken@rand.org

[From Bill Powers (2003.08.07.1157 MDT()]

Rick Marken (2003.08.07.1020)--

I can even make
that sound with my tongue starting behind my top front teeth; my tongue
never goes
between my teeth and I still hear "th".

That's an interesting point. I wonder if there are any sounds that can be
understood only if made in one particular physical way. It may be that
while you can hear one sound in isolation created in different ways, when
you put it together with other sounds the options become fewer. This could
relate to what Bruce Nevin has been saying about "contrasts." If you simply
make the sounds "f", "s", "th" one after the other in various orders, can
someone else (or you listening to a recording) tell the difference if you
make the th in a nonstandard way? And how about a word like "puffs" -- can
you make the ff sound in any way but with lip to teeth, and be understood,
while still ending with a sound that sounds like the standard s?

Also, I think we sometimes tend to hear our own reference signals instead
of our own words when we speak, and of course we do not hear our own speech
exactly as others hear it. We may think we're saying something in the
standard way when we're not. I suggest trying the "th" sound by recording
it and listening to the playback. I'm going to try that with the Chinese
sounds I think I'm making, now that you've brought this up. Ouch, I'll bet
I'm not doing as well as I thought I was.

Interesting possibility. But I think it's more important to get the
hierarchical
relationships right before worrying about whether the variables involved
are at
the same level perceptually, according to the PCT hierarchy.

That's a somewhat confusing proposition, since the variables, according to
the PCT hierarchy, _are_ the perceptions. What's coming out of this for me
is the question of whether a higher-order perception can be composed of
elements of lower-order perceptions of more than one modality. I think it can.

I think what we have
to do is build a model of speech (sound) production and then test it. If
it turns
out that the model works best when kinesthetic perceptions are controlled
at the
same level as sound perceptions then that is evidence that, indeed, speech
involves the simultaneous control of sound and kinesthetic sensations.

Here's a thought about that. We say that a higher-order system works by
varying the reference signals of lower-order systems. So immediately we
think of a sound configuration being controlled by varying reference
signals for sound sensations, and the sound sensations being controlled by
varying reference signals for sound intensities. But how do we vary sound
intensities? There isn't any lower level than intensity control. Sound
intensities have to be controlled by tightening and loosening muscles in
the output functions of the intensity-control systems.

OK, now how do we "mouth" a word without making any sounds? Try "feel,"
which involves some visible and feelable configurations of the mouth. What
we're doing now is controlling the mouth configurations with which we would
normally make the spoken word "feel" which is represented here by marks
instead of sounds but is, somehow, the "same word." If you see me mouth
that word silently, you can probably guess what word I'm saying,
particularly if I hold my lips open so you can see my tongue. Now I'm
controlling a kinesthetic configuration by varying reference signals for
kinesthetic sensation control systems, which in turn are controlling
kinesthetic sensations by varying reference signals for kinesthetic
intensity control systems. And the kinesthetic intensity control systems
are controlling by ... varying the tensions in muscles, the SAME muscles
that I use when I say I'm controlling sound intensities.

What we _don't_ have in the above descriptions is a sound configuration
control system sending reference signals to a kinesthetic sensation control
system, or a kinesthetic configuration control system sending reference
signals to a sound sensation control system.

Now I think that in reality we have a word (or I suppose the word is
"morpheme") configuration control system which is controlling a
configuration made up of both auditory and kinesthetic sensations. And the
auditory and kinesthetic sensations are being controlled by sending
reference signals to _effort_ intensity control systems, so either form of
control uses the same or overlapping sets of muscles. To SCREAM a sound
LOUDLY you also have to SQUEEZE with your diaphram HARD, squeezing air out
FAST. You can hardly do either one without doing the other. At bottom only
a single set of muscles is used, and control is happening simultaneously in
the kinesthetic and auditory hierarchies. The branches merge at the level
where we experience words, which is why we can experience a somewhat
subdued version of a spoken word on the basis of either modality alone.
Feeling and hearing our own speech sensations at the same time provides the
canonical experience of speech.

Note that a person who loses hearing can continue speaking quite
intelligibly for a long time, but gradually loses intelligibility because
the sound sensations no longer can be checked to see if they match the
intended sounds. If the higher level consisted only of sound control, I
would expect a much more rapid and thorough loss of intelligibility. I
suppose there must be a comparable condition in which a person loses the
ability to say intelligible words because of some distortion in the
kinesthetic sensations (I think you mentioned dental use of Novocain). In
the latter case function can be somewhat restored as long as it is at all
possible to make the sounds. In the former case, intelligibility can be
maintained, maybe, through other means of getting feedback about the sounds
one is making.

One last example occurs to me. Have you ever watched a movie in which the
sound was half a second or so out of synch with the picture? The result, in
closeups where the mouth is prominent, is very strange for me. I keep
wanting the visual configurations to "match" the sound configurations. Of
course that's impossible since these are two different modalities. Clearly
the experience of hearing the person talk is partly disrupted if the visual
configurations are not the ones that usually go with the sound
configurations. I'm used to experiencing them together as part of the same
perception. Of course I can see the mouth movements with no sound at all,
or hear the sounds without being able to see the mouth movements, with no
problem.

As I get slowly more deaf, I find myself moving so I can see the speaker's
mouth when I am having trouble hearing. That seems to make the words
clearer. I don't experience the non-auditory aspect of the speech I am
hearing as a separate thing, or haven't until now, but clearly it's there.

I'm also finding that as I learn how the mouth is supposed to feel in
speaking Chinese, the feel seems to suggest the sound more and more, and
when I see the letters, my mouth moves toward the appropriate configuration
and I imagine the feel of chuffs of air or buzzes. While trying not to
bother Mary, I say the words with no sound at all, going strictly by the feel.

The tests
of such a model would, of course, use PCT methodology (test for the controlled
variable). But I think a model is needed before we can get anywhere on this.

I think we have as much of a model as we're likely to construct for a
while, but that we need to consider it in more detail, as above.

> I think there's no doubt that in speech the highest priority variable is
> auditory. But this does not make it a higher level variable, unless we can
> show that configurations in the sounds of speech are functions of
> articulatory sensations, which I'm pretty sure is not the case.

I don't understand this. Isn't it likely to be necessary to vary articulatory
sensations (doing which has particular physical consequences) as the means of
producing various sound sensations (which are a physical consequence of the
physical correlates of the articulatory sensations)?

I know, I know, this makes perfect sense, but I don't think it holds true
here. Controlling some perceptions involves affecting others, but the
relationship is not always hierarchical. It's like saying that controlling
the position of your hand while reaching out entails affecting the position
of your elbow, and vice versa, or drawing on a blackboard entails creating
a screeching sound. The two kinds of experience grow out of the same
underlying causes, so neither one can be said to cause the other or depend
on the other.

That seems to make sense today. Tomorrow, we'll see.

Best,

Bill P.

[From Bruce Nevin (2003.08.07 17:08 EDT)]

Bill Powers (2003.08.07.0650 MDT)–

much of what I have called speech

“sounds” consists of kinesthetic sensations, not auditory ones.
The buzzing

sound of Z, for example, is felt in the tip of the tongue! Also, say
“the”

a few times and examine the “th” sound. It sounds partly like
your tongue

between your teeth. That’s not a sound.

Well, yes, but there is also sound of course, or Mary would be asking why
your sudden interest in thistles when you say something like this’ll
do
.

It’s an interesting exercise, and one now of particular interest to you
perhaps for learning Mandarin, to train yourself to be more aware of
what’s going on during speech. As Rick says

Rick Marken (2003.08.06.1120)–

my subjective impression is that I don’t
control for articulatory patterns at all when I talk. I can barely tell
what the hell my mouth (oral cavity, tongue, palates, etc) is doing when
I speak. I think I just do whatever I have learned to do with my vocal
apparatus

You can’t control what you can’t perceive, yes, but maybe it’s also true
that its harder to learn to control what you’re not aware
of perceiving. And for reasons that I’ve suggested at other times, for
cultural conventions to work it may be necessary that they be controlled
out of awareness. We have good reason to mistrust those who we know have
a facility at manipulating the variables from which we deduce their
persona. But that’s another topic, and too far afield for now. More to
the point here, you can train your awareness of what you are doing as you
speak, and learn finer control of your ability to make different speech
sounds independently of “whatever you have learned to do with your
vocal apparatus” in order to speak English.
If you produce the initial fricative of this continuously without
moving your tongue back to form a vowel you certainly do feel the
vibration of the pressure wave produced by your vocal folds in the air
stream. You can turn your voice off and on while continuously forcing air
through the aperture between your tongue and your teeth, alternating
between the voiced dh of this (written with a lowercase
delta) and the voiceless th of thing (written with a
lowercase theta). Don’t say either this or thing, but just
alternate dhthdhthdhthdhth by producing a voice or not. (Drop into
saying this or thing from time to time until you are sure
you have it right.) While you’re alternating the voice on and off,
everything else stays the same. Because everything else is constant, you
can more clearly notice the sensations of the edges of your tongue
touching your teeth, but dropping away from contact just at the tip,
leaving a narrow, flat tunnel for the air to pass through. You can feel
the increase of air pressure when voicing stops, with its hissing
sound.

I think what we need are some other examples
of “instrumental” controlled

variables; that is, variables that are controlled in order to control
other

variables at the same level.

Nice term. And yes, I think you’re on the right track here. But instead
of instrumental controlled variables I’d call them instrumental
dependencies between controlled variables.

Rick Marken (2003.08.06.1120)–

Bruce Nevin
(2003.08.06 13:08 PDT)

Bill Powers (2003.08.05.0557 MDT)–

Neither sound nor articulation is prior to the
other. They are both at the same level of perception.
Yes.
But references for articulations are calibrated according to their affect
on sounds, not vice versa.
I think this is right:
articulations are controlled as the means of controlling sound.

This is confusing the source of the dependency between auditory input and
tactile input. What is going on is that perceptual inputs in one modality
are linked through the environment to perceptual inputs in the
other.

We control variables by means of dependencies in the environment all the
time. I’m making letters appear on a cathode ray tube by pressing my
fingers down on an array of little cuboid pieces of plastic.

The dependencies between the tactile sensations of speaking and the
sounds of speech are not in the hierarchy, they are in the environment,
as in other examples of ‘instrumental’ dependencies between controlled
variables such as those Bill identified:

Bill Powers (2003.08.07.0650 MDT)–

Example: I curl my fingers around a glass
in

order to hold it in position for drinking. The behavior of the fingers
and

the glass seems related in the same way the behavior of the tongue
is

related to saying “th”

As you said earlier, the dependency is a matter of physics - physiology
and aerodynamics - not a function of how the nervous system is hooked
up.

Here, I think you’re talking about the correlation of the visual
perception of the configuration of your fingers curling around the glass
with the kinesthetic perception of it. To pick up the glass you control
both at once, and you set the references for the kinesthetic perceptions
(configuration of hand and configuration of arm) according to the visual
perceptions (configuration of hand and location of hand) and the tactile
perception (touch of fingers on glass).

This is how I’m thinking about speech: a
kinesthetic modality being

controlled because of its effects on the auditory modality _at the
same

level_. Isaac Kurtzer sent me an abstract of a paper about an experiment
in

which articulatory variables were disturbed in a way that had no
measurable

or sensible effects on sounds. The paper reported that the disturbance
was

resisted. Of course in a hierarchical model that doesn’t prove that

articulation is the primary controlled variable (as the authors
thought).

I’m very interested in this! Can you send me the reference?

But it does suggest doing the opposite
experiment, varying the sound

without changing the articulation (altering phonemes before they are
heard,

and so forth). I’ve tried to get somewhere with that but have had

programming difficulties.

Last I knew, this had to be done in hardware (Remember Houde et al? Never
replied). Are processors fast enough now?

I think there’s no doubt that in speech the
highest priority variable is

auditory. But this does not make it a higher level variable, unless we
can

show that configurations in the sounds of speech are functions of

articulatory sensations, which I’m pretty sure is not the
case.

Just as in picking up the glass the highest priority variable is the
tactile sensation of fingers gripping glass, and that is obviously at a
lower level than either the visual or the kinesthetic configuration
perception!

There are some articulations, as you say, that
change when disturbed so as

to preserve the sound being made. But with this new consideration of

instrumentality, this does not have to mean that the auditory
variables

are of higher level. If your chin is held down against a rest, your
head

goes up and down to say “mama”. You can also say
“mama” with your lips

remaining slightly opened and using a crosswise finger to block and
open

the gap. There are physical interactions among different effects of

controlled variables, and we use them to produce the experiences we
want.

If one aspect of the situation changes we can usually find an
alternate

means of creating the same effect. This doesn’t mean there is a

hierarchical relationship among the various means at the same
level.

This perhaps ought to be cast in terms of
several

levels of perception in one modality being related to control of
several

levels of perception in another modality, because there are
sensations,

configurations, and transitions (etc) in both modalities at once, but
the

first set is controlled to have controlled effects on the second
set.

Yes indeed! This is perhaps most elaborate in the correlations between
our control of language perceptions all the way up and down the hierarchy
and our control of nonverbal perceptions all the way up and down the
hierarchy.

Rick Marken (2003.08.06.1120)–

The most convincing demonstration (to me) of
articulation as a means of producing intended sound patterns comes from
pressing my tongue against the bottom of my mouth … and then trying to
say words like “Paraguay”. When I’m affecting a Spanish
accent I roll the r in “Paraguay”. But I can’t do that
with my tongue pressed against the bottom of my mouth. But I can get
something like the sound of the rolled r by making a guttural roll in my
throat.

This is simple physics. Pulling your tongue tip down forces the dorsum of
the tongue up toward the velum and (dangling from it) the uvula. The flow
of air between the back of the tongue and the uvula makes the uvula flap.
Voilá! You are now producing a (Parisian) French r, a uvular trill. Also
heard in many dialects of German. And you have a demonstration of how the
‘same’ phoneme can be pronounced as an apical trill (tongue tip) in some
dialects and a uvular trill in others, and how the ‘same’ phoneme can
change its phonetic implementation among s, z, trilled r, uvular r, etc.
over time.

I came up with this articulatory compensation
for the lack of tongue rolling “automatically” in an effort to
hear “Paraguay” in something like the way I usually hear
it. And, indeed, I am able to hear something remarkably close to
the “rolled r” version of “Paraguay” using the
guttural substitute.

What I see here is a completely different articulation (guttural roll)
being used to produce an intended sound: “Paraguay” with rolled
r. This makes me think that perceptions of articulation are not
parallel controlled components of the words we speak. Articulations
are, I believe, varied as necessary to produce the sounds we intend to
produce.

Surely. But not in real time, but rather over a series trials until the
reference value is reset, as in setting a sight. A ventriloquist
practices a lot before being able to produce normal-sounding
speech without moving the lips. Having done all that practicing, she has
a set of kinesthetic references for the ‘ventriloquism dialect’.

Since there articulations are typically
observed when no disturbances (like food in the mouth) are present, it
looks like particular articulations are used to produce particular
sounds. But I think this apparent one-to-one mapping of action
(articulation) to result (sound) is an artifact of observing speech in
conditions where the need for variation in action is reduced or
eliminated.

Without practice, impediments to normal articulation result in failure to
produce the desired sounds. The young woman with the hole in her cheek
(reported in the body-piercing paper that I mentioned in my talk and in
the Festschrift) was unable to produce consonants like p and f until she
learned to hold her tongue against the inside of her cheek to plug the
hole and maintain sufficient air pressure. This was not an instant
variation of means to control the auditory variable, such that there was
no error in the acoustic output when the disturbance was introduced.

    /Bruce

Nevin

···

At 09:34 AM 8/7/2003, Bill Powers wrote:
At 11:20 AM 8/6/2003, Richard Marken wrote:
At 11:20 AM 8/6/2003, Richard Marken wrote:
At 11:20 AM 8/6/2003, Richard Marken wrote:

[From Rick Marken (2003.08.08.0800)]

Bruce Nevin (2003.08.07 17:08 EDT)–
Rick Marken (2003.08.06.1120)–

I think this is right: articulations are controlled
as the means of controlling sound.
The dependencies between the tactile sensations of speaking and the sounds
of speech are not in the hierarchy, they are in the environment, as in
other examples of ‘instrumental’ dependencies between controlled variables
such as those Bill identified:
I agree. That’s all I meant, really: sound is an environmental function
of articulation. Whether there is a hierarchical dependency between any
control loops involved is something to be determined by research.
Rick Marken (2003.08.06.1120)–
The most convincing demonstration (to me) of
articulation as a means of producing intended sound patterns comes from
pressing my tongue against the bottom of my mouth … and then trying to
say words like “Paraguay”. When I’m affecting a Spanish accent I
roll the r in “Paraguay”. But I can’t do that with my tongue pressed
against the bottom of my mouth. But I can get something like the sound
of the rolled r by making a guttural roll in my throat.

This is simple physics.

Well, it’s not that simple. But it is physics.

I came up with this articulatory compensation
for the lack of tongue rolling “automatically” in an effort to hear “Paraguay”
in something like the way I usually hear it. And, indeed, I am able
to hear something remarkably close to the “rolled r” version of “Paraguay”
using the guttural substitute.
What I see here is a completely different articulation (guttural roll)
being used to produce an intended sound: “Paraguay” with rolled r. This
makes me think that perceptions of articulation are not parallel controlled
components of the words we speak. Articulations are, I believe, varied
as necessary to produce the sounds we intend to produce.

Surely. But not in real time

Actually, it was in real time. the first time I pressed my tongue to my
lower palate and said “Paraguay” I made the “r” with the guttural roll.

but rather over a series trials until the reference
value is reset, as in setting a sight.
The reference value for the lower order articulation? If so, then you are
proposing a hierarchical relationship between articulatory and sound control,
which is OK with me but I thought you and Bill were arguing against that
model.
Without practice, impediments to normal articulation
result in failure to produce the desired sounds.
It depends on how large the impediment (disturbance). But of course
that’s true; we learn to articulate sounds in the context of the ambient
disturbances to controlled articulatory and sound variables. I think I
was able to do the guttural compensation so well because I spent a lot
of time in grad school singing and playing guitar while disturbing by articulatory
apparatus with various beverages.
Best regards

Rick

···

Richard S. Marken, Ph.D.

Senior Behavioral Scientist

The RAND Corporation

PO Box 2138

1700 Main Street

Santa Monica, CA 90407-2138

Tel: 310-393-0411 x7971

Fax: 310-451-7018

E-mail: rmarken@rand.org

[From Bill Powers (2003.08.08.0745 MDt)]

Bruce Nevin (2003.08.07 17:08 EDT)--

I'm approaching saturation on this subject for now, but will probably take
it up again before long when I have some new problems.

Isaac Kurtzer sent me an abstract of a paper about an experiment in
which articulatory variables were disturbed in a way that had no measurable
or sensible effects on sounds. The paper reported that the disturbance was
resisted. Of course in a hierarchical model that doesn't prove that
articulation is the primary controlled variable (as the authors thought).

I'm very interested in this! Can you send me the reference?

Here's the body of Isaac's post:

···

=========================================================================
In this vein is an intresting article published recently in Nature,
"Somatosensory basis of speech production" Tremblay et al, Vol 423, pp 866-
869. It should be noticed the adaptation occurs for "silent" speech as well
as vocalized speech, but not just opening and closing the mouth.

Here's the abstract:

The hypothesis that speech goals are defined acoustically and maintained by
auditory feedback is a central idea in speech production research. An
alternative proposal is that speech production is organized in terms of
control signals that subserve movements and associated vocal-tract
configurations. Indeed, the capacity for intelligible speech by deaf speakers
suggests that somatosensory inputs related to movement play a role in speech
production but studies that might have documented a somatosensory component
have been equivocal. For example, mechanical perturbations that have altered
somatosensory feedback have simultaneously altered acoustics. Hence, any
adaptation observed under these conditions may have been a consequence of
acoustic change. Here we show that somatosensory information on its own is
fundamental to the achievement of speech movements. This demonstration
involves a dissociation of somatosensory and auditory feedback during speech
production. Over time, subjects correct for the effects of a complex
mechanical load that alters jaw movements (and hence somatosensory feedback),
but which has no measurable or perceptible effect on acoustic output. The
findings indicate that the positions of speech articulators and associated
somatosensory inputs constitute a goal of speech movements that is wholly
separate from the sounds produced.

Later,

Bill P.

[From Bruce Nevin (2003.08.10 07:28 EDT)]

Rick Marken (2003.08.08.0800)--

Bruce Nevin (2003.08.07 17:08 EDT)--

Rick Marken (2003.08.06.1120)--

The most convincing demonstration (to me) of articulation as a means of producing intended sound patterns comes from pressing my tongue against the bottom of my mouth ... and then trying to say words like "Paraguay". When I'm affecting a Spanish accent I roll the r in "Paraguay". But I can't do that with my tongue pressed against the bottom of my mouth. But I can get something like the sound of the rolled r by making a guttural roll in my throat.

This is simple physics.

Well, it's not _that_ simple. But it is physics.

I came up with this articulatory compensation for the lack of tongue rolling "automatically" in an effort to hear "Paraguay" in something like the way I usually hear it. [...]. Articulations are, I believe, varied as necessary to produce the sounds we intend to produce.

Surely. But not in real time

Actually, it was in real time. the first time I pressed my tongue to my lower palate and said "Paraguay" I made the "r" with the guttural roll.

OK, we're talking about two things here. First, the 'physics' (simple or not) was what "came up with" this new articulation. You didn't come up with it by reorganization or by otherwise resetting any references, it just happened because pushing the tip of your tongue down bunched the body of your tongue up against the uvula. Then you noticed the sound. If I had tried to describe to you how to produce a French/German uvular trill so that you could establish reference values for it, it probably would have been difficult. (I say that based on experience with other people.)

The second thing is that I am talking about how references are set, for example when you imitate someone's accent. You do what it feels to you should produce that way of talking, you hear what it sounds like, you try it a bit differently, reorganizing until it sounds right to you.

but rather over a series trials until the reference value is reset, as in setting a sight.

The reference value for the lower order articulation? If so, then you are proposing a hierarchical relationship between articulatory and sound control, which is OK with me but I thought you and Bill were arguing against that model.

At a higher level where syllables are recognized, and (in parallel) where morphemes are recognized, you have inputs of diverse modalities. That's Bill's proposal. Rather like the way category recognizers are supposed to work. Either the acoustic input or the tactile input or both. Now suppose you have the tactile input without the acoustic input. The syllable-recognizer is the source of the reference signal for the acoustic input, so in imagination you hear what you must be saying even though that noise is masking the sound of your own voice. And suppose you have the acoustic input but not the tactile input. The syllable-recognizer is the source of the reference signal for the tactile input, so you feel, in imagination, what it must feel like to produce that sound. From there it is not a great step to see how acoustic feedback, although too delayed for real-time control, enables syllable-control systems to adjust reference values for tactile control so that, over a series of trials, the syllables come to sound like what you are hearing from others.

With this caveat):

Bill Powers (2003.08.07.1157 MDT)--

I think we sometimes tend to hear our own reference signals instead
of our own words when we speak, and of course we do not hear our own speech

If you prefer your remembered references for syllables, stress patterns, and intonation patterns over the way these people you have to live with talk, you retain your foreign accent even after years. Clara Bow couldn't change her Brooklyn slum accent, even though she had lived for years in California (said to be the reason her career died).

         /Bruce Nevin

···

At 08:06 AM 8/8/2003, Richard Marken wrote:

[From Bill Powers (2003.08.10,0856 MDT)]

Bruce Nevin (2003.08.10 07:28 EDT)--

At a higher level where syllables are recognized, and (in parallel) where
morphemes are recognized, you have inputs of diverse modalities. That's
Bill's proposal. Rather like the way category recognizers are supposed to
work.

If you tell me something often enough, it will eventually sink in. Of
course: words are categories, at the category level. Below that level we
have the things that are eventually brought together to make words:
intensities, configurations, transitions, and relationships. You've been
muttering "categories" for several posts now, and I guess this time I heard it.

Notice how this also makes words into functions of _visual_ as well as
kinesthetic variables. At the category level, the origins of the input
signals no longer matter; the result is a single perceptual signal that we
call a word. That signal tells us a specific word is occurring or has just
occurred. The actual "look and feel" (and sound) of the word is made up of
all the lower-level perceptions occurring at that time. However, neither
the look, the feel, or the sound of the word is the word. The word is the
sense of the class of experience that is being evoked, a perception that of
course can also be aroused by the experience that also belongs to that class.

That's only a trial balloon. Or a pretty something to be held up and
admired for a while before we get back to work.

Bill Powers (2003.08.07.1157 MDT)--

I think we sometimes tend to hear our own reference signals instead
of our own words when we speak, and of course we do not hear our own speech

If you prefer your remembered references for syllables, stress patterns,
and intonation patterns over the way these people you have to live with
talk, you retain your foreign accent even after years. Clara Bow couldn't
change her Brooklyn slum accent, even though she had lived for years in
California (said to be the reason her career died).

That's true, but I meant something a little different. I meant that
sometimes we hear what we're intending to say instead of what we're
actually saying, and what other people hear us saying. At the level of
meaning this certainly happens: "I threw the horse over the fence some
hay." But I think it can happens with ordinary words, too. You say
something and only hear what you said a second later, and correct yourself:
"Did I say I want the ham? I meant
the chicken." And sometimes, I have noticed, people will say the wrong word
and NOT realize they said it. How could they not recognize saying a whole
unintended word? I can only guess that they're imagining the meaning they
want to convey and miss the fact that the word actually said doesn't have
that meaning.

And why not at a lower level still? I might imagine that I am making a
Third Tone, while actually producing, in the listener's ears, a meaningless
gargle.

Best,

Bill P.

···

        /Bruce Nevin

[From Rick Marken (2003.08.11.1830)]

Bruce Nevin (2003.08.10 07:28 EDT) --

Rick Marken (2003.08.08.0800)--

>the first time I pressed my tongue to my
>lower palate and said "Paraguay" I made the "r" with the guttural roll.

OK, we're talking about two things here. First, the 'physics' (simple or
not) was what "came up with" this new articulation.

Not really. I came up with the articulation, which was already part of my
articulatory repertory, and the result (via the physics) was a sound that was a
reasonable approximation to the rolled r.

You didn't come up with
it by reorganization or by otherwise resetting any references

I think so. I'm pretty sure that the necessary reorganization happened back in
those grad school singing days to which I referred. I believe that my sound
control systems -- the ones controlling for the perception of rolled r -- were
already organized to use the guttural roll articulation as a means of producing
the desired sound when my tongue was "out of the loop" (being pressed against
my soft palate).

it just
happened because pushing the tip of your tongue down bunched the body of
your tongue up against the uvula. Then you noticed the sound.

It seemed like an active process to me. I had to do stuff with my throat and
other stuff back there that I don't do when I normally roll my r. I just
seemed to "know" (unconsciously) how to get that rolled r sounds using this
other vocal means.

I agree that perceptions of articulation can be part of what we perceive when
we perceive speech at the category level. That would certainly explain why
people can perceive words when they read lips our simply mouth words
themselves. But I don't think perceptions of articulation are essential for
recognizing spoken language. Otherwise, how could we understand what people say
when our mouths are full.

Best

Rick

···

--
Richard S. Marken
MindReadings.com
marken@mindreadings.com
310 474-0313

[From Bill Powers (2003.08.11.0815 MDT)]

Rick Marken (2003.08.11.1830)--

Thanks for forwarding the message about the tracking program. I picked up
the wrong message to use for a reply. Hope everyone enjoys the program.

Just to update those interested:

A month or two ago, John Flach of Wright State College wrote a joint paper
with Rick and me on a PCT model of pursuit tracking. The conclusion of the
paper was that long-established ideas about adaptation of human control
systems were upset by finding that a control system could control under
highly variable conditions -- without any change in its internal
characteristics. The appearance of adaptation under different load
conditions proved to be an illusion. This isn't to say that adaptation
never occurs, but it shows that control systems can operate over such a
broad range of environments that adaptation is necessary only under the
most extreme changes. John felt this was a highly significant finding. Rick
and I both admired him for coming to this view, as he was the co-author of
a book, just published, in which adaptation was featured in at least one
major chapter (chapter 14, for those who have the book, "Control theory for
humans" by Jagacinski and Flach)..

After the CSG meeting, John reported that his colleagues (and his
boss/co-author) were reluctant to accept this finding, and that they were
saying it would not hold true for compensatory tracking (stationary target,
force disturbance applied to hand or joystick). A coworker came up with an
analytical solution of the equations for the compensatory tracking case
which seemed to show that indeed, there was no adaptation when the control
system retained constant charcteristics. Flach went through the derivation
and had to concur.

However, I found that this derivation contained a mistake, which evidently
both Flach and his coworker had made, When the mistake was removed, the
equations came to look much more like those for the pursuit tracking case.
At the moment, I have left it to Flach to complete the reanalysis, which
uses mathematical methods (complex variables) I have allowed to become
extremely rusty (they were none too shiny to begin with).

In the meantime, I rewrote my simulation of the pursuit-tracking case so it
could be switched to do the compensatory tracking case as well, and found
that the apparent but illusory adaptation occurred just as strongly as in
the pursuit case. That is the program sent yesterday. Rick redid his own
spreadsheet model and found the same thing. We are awaiting Flach's
response (I sent my report on the mistake Friday night and we haven't heard
back yet). By the way, I can ber sure that no adaptation occurs in the
simulation, because I didn't include any way for the control system to
alter its own parameters.

"Flach" is pronounced "flak", which John says he's gotten plenty of in his
life.

Best,

Bill P.

[Bruce Nevin (2003.08.12 07:28 EDT]

Bill Powers (2003.08.10,0856 MDT)–

words are categories, at the category
level.

We have several domains of perception and control in parallel. Language
is one, distinct from its referents. A word is a perception in the
language domain. The category perception that you’re looking for has a
word (or it could be a set of words and phrases) among its inputs.

It’s fairly easy to identify words by a stochastic process on phonemes,
and there is good evidence that people do this (statistical learning
theory). The language-domain perception is then an input to
your category recognizer along with perceptual constructs in other
domains.

Notice how this also makes words [categories,
rather] into functions of visual as well as

kinesthetic variables. At the category level, the origins of the
input

signals no longer matter; the result is a single perceptual signal that
we

call [rather, that we name by or with or by means of] a word. That signal
tells us a specific word is occurring or has just

occurred.

No, the signal from a word recognizer is one of the possible inputs to
the category recognizer.

The actual “look and feel” (and
sound) of the word is made up of

all the lower-level perceptions occurring at that time. However,
neither

the look, the feel, or the sound of the word is the word. The word is
the

sense of the class of experience that is being evoked, a perception that
of

course can also be aroused by the experience that also belongs to that
class.

A word is distinct from its meaning. (Reminds me of the maxim about maps
and territories, but not the same.)

Bill Powers
(2003.08.07.1157 MDT)–

I think we sometimes tend to hear our own
reference signals instead

of our own words when we speak, and of course we do not hear our own
speech

If you prefer your remembered references for syllables, stress
patterns,

and intonation patterns over the way these people you have to live
with

talk, you retain your foreign accent even after years. Clara Bow
couldn’t

change her Brooklyn slum accent, even though she had lived for years
in

California (said to be the reason her career died).

That’s true, but I meant something a little different. I meant that

sometimes we hear what we’re intending to say instead of what we’re

actually saying, and what other people hear us saying.

We’re in somewhat jostling agreement. I was starting from there and
jumping on to a comment based on why we hear what we intend to say
rather than what we actually say. Why do we sometimes prefer a replay in
imagination to actual perceptual input? Why do we sometimes delay
updating our self-image to accord with what we see in the mirror?

    /Bruce

Nevin

···

At 11:16 AM 8/10/2003, Bill Powers wrote:

[From Bill Powers (2003.08.12.1920 MDT)]

Bruce Nevin (2003.08.12 07:28 EDT]--

words are categories, at the category level.

We have several domains of perception and control in parallel. Language is
one, distinct from its referents. A word is a perception in the language
domain. The category perception that you're looking for has a word (or it
could be a set of words and phrases) among its inputs.

I think I agree. What I'm looking for is the way language is connected to
the nonverbal experiences that are its meanings. "Association" doesn't tell
us much -- it just describes the phenomenon, without telling us what's
actually happening. Maybe that's as close as we can get.

It's fairly easy to identify words by a stochastic process on phonemes,
and there is good evidence that people do this (statistical learning
theory). The language-domain perception <word> is then an input to your
category recognizer along with perceptual constructs in other domains.

Yes to the last, but I doubt the first. You can get effects resembling the
things we do with statistics using circuits that do no statistical
computations. The "probability" that something has a certain form can be
mimicked by a circuit that simply averages the output of some filter or
summator. I would not want to say that a brain process was statistical
unless it performed explicitly statistical calculations.

I'm trying to name something that is not really a word -- that is, it is
neither a visual nor an auditory configuration, transition, event, or
relationship. "Word" is probably not the word I'm looking for. It's a
linguistic entity that is a symbol for some nonverbal experience. Any part
of it can evoke the rest of it as the same perceptual signal which says
that this entity and not some other is present.

A word is distinct from its meaning. (Reminds me of the maxim about maps
and territories, but not the same.)

Yes, that's one of the conditions I'm accepting. Somehow a word acts as a
pointer to some other experience.

I am getting a strong sense that I don't know enough to pursue this idea
much further right now.

We're in somewhat jostling agreement. I was starting from there and
jumping on to a comment based on why we hear what we intend to say rather
than what we actually say. Why do we sometimes prefer a replay in
imagination to actual perceptual input? Why do we sometimes delay updating
our self-image to accord with what we see in the mirror?

To say that we "prefer" the reference signal implies that we have observed
both the reference signal and the actual perceptual signal and decided to
confine our attention to the reference signal. While that can happen, it's
not the sort of thing I wanted to describe. What I had in mind would be
more like trying to do something hard, and not being very good at it, and
concentrating very hard on what it is one _wants_ to do. If too much
imagination gets into the process, it can seem that one _is_ doing the
right thing, while attention is not on the actual perception which is quite
different.

This happens all the time on CSGnet. People write things that, in
retrospect, it's pretty clear that they didn't mean to write -- writing
wasn't for didn't, or Bruce for Lloyd, errors that aren't just slips of the
finger. _And then posting them_. This is the kind of error Rick calls --
heck, I forgotten what he calls it. But it means that you do or say
something which is not at all what you meant, but you fail to realize that
it wasn't what you meant. I can't think of any way for this to happen
except for attention to be on the imagined reference signal to the extent
of missing the fact that the perception totally failed to match it.

Oh, well. that's probably not what I meant to say anyway.

Best,

Bill P.