social environment

[From: Bruce Nevin (Wed 930519 12:26:11)]

Rick Marken (930518.2200) --

···

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
If, by "reference perceptions" you mean "reference signal" settings then
HPCT already has an account of how they are arrived at; they are varied
by higher level systems as the means by which they control their
perceptions. If by "reference perceptions" you mean the state of the
perceptual variable that corresponds to the reference signal setting
currently in effect, then PCT already provides an account of that --
negative feedback control of perceptual input. If by "reference settings"
you mean the fixed settings of the intrinsic reference signals, then
these are presumably determined by the genes.
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

The forming of arcs and circles is universal. It can hardly be spoken
of as normative precisely because it follows from first principles that
apply to all humans regardless of their culture or history.

In the case of the gather demo, we might ask why it is that all the
individual control systems have the same setting for a reference
perception for proximity to other individual control systems. What would
happen if an "Arab" or "Latin American" control system were among
the "American" ones, with the proximity setting much closer? How would
the others respond if that one pressed closer to the speaker, in such a
way as appearing to stand between other individuals and the speaker?
Could notions of politeness and offense arise out of such interactions?

The use of "s" vs. "th" to distinguish words (as in English sin vs. thin)
or of [y] vs. [u] to distinguish words (as in French tu vs. tout) is
normative and language-specific (culture-specific, group-specific),
not universal. It may follow from first principles of HPCT that
apply to all humans (I assume that it does), but if so it does so
indirectly.

The relevant perception appears to be one of maintaining mutual distinctness
(contrast) among the words used in communicating with others on whom one
relies for cooperative attainment of goals. The shapes of the words
must be set as reference perceptions equivalently in the several co-members
of the cooperating group. "Equivalently" means in terms of the contrasts
or distinctions that must be maintained between words (so that the
words are recognizably distinct from one another), and not in terms
of the detailed perceptual image of each word (in acoustic, kinesthetic,
tactile, and other perceptions). The vocabulary and the contrasts are
necessarily pre-set in speaker and hearer. That is, the setting of the
reference perceptions for words and for the perceptual means for
maintaining contrasts among the words must be established first for
linguistic communication to be possible.

I am leaving out of consideration here the correlation of word-
perceptions with non-word perceptions (differently in different
languages), the perceptions of dependency relations among words, the
correlation of word-construction perceptions with non-word perceptions
(differently in different languages), the association of gestural
perceptions such as expressive intonation and "body language" with
these linguistic elements, having affective perceptual associations,
and so on. As a simple example of the first of these, in two languages
that are after all pretty closely related:

        English Italian

        floor piazza (I think) Under your feet now
        floor plano ... of a building
        plane plano geometric surface

Let us suppose that the hypothesis is true, that all these language-
specific, culture-specific, in some cases family-specific or couple-
specific agreements about settings of reference perceptions arise
from first principles of perceptual control theory within environmental
constraints. We still have to account for how it is that the settings
that are normal and agreed to by this pair of people over here are
so vastly different from the settings that are normal and agreed to by
this other pair of people over there. You don't have that problem
for arm movements or for the formation of arcs and circles and other
such non-normative phenomena. We (we) do have that problem to
account for your current process of reading and understanding these
words and sentences, ma dhen katalavenis katholu tora, except insofar
as you can project expectations from context.

I have granted the hypothesis, subject of course to demonstration.
I am posing additional questions that must be answered before
we can say very much that is useful about category perceptions
and higher-level perceptions dependent upon categorization.

(I believe you are confusing "there is more" with disagreement,
and that you are presuming that disagreement betrays failure to
understand the theory. I'm not concerned with competing for who
understands HPCT best (though I do appreciate Wayne's kind words).
But when you say or suggest that I don't understand the theory I
would prefer that you point out something that I indeed don't
understand.)

Think of a child learning to throw and catch a ball. Reference
perceptions are set based upon experience of trial and error in
throwing and catching a ball. As we understand matters, the
ball's trajectories on different trials have their basis in
properties of the environment, expressable mostly in terms of
physical laws (observed regularities), with adjustments for such
things as wind. Someone worried about the veracity of the theory
from a Realist stance might say that the child calibrates reference
perceptions for conformity to properties of the environment. PCT
itself does not require this.

The same child (at a younger age) learning the language spoken by
her parents sets reference perceptions based upon trial and error
in understanding words and in having her words understood. The
word shapes that she produces, and the contrasts among the words
that she has at any given time learned to produce, have their basis
in properties of the environment, namely, the word shapes and
distinctions that her parents and others produce and recognize.
But these properties of the environment are in the environment
as a byproduct of perceptual control by other people--specifically,
as a byproduct of their control of perceptions that the child is
learning to control in a way that they can recognize. In this
respect, the properties of the child's environment that give rise
to words and phonemic distinctions (in that order, and then to
more words once the system of distinctions is controlled) is unlike
the physical properties of the environment that imputedly give rise
to such things as parabolic trajectories. We really have no choice
about attributing these "social" properties to the (social)
environment, since we must give an HPCT account of how all those
people in the child's environment produce and recognize the words
that the child is learning to produce and recognize, and how they
keep them distinct from one another by the language-specific
distinctions that the child is learning to control.

In doing this, we must carefully distinguish our account of how
the adults do it (once they have learned the distinctions and many
of the words, etc.) from our account of how the child learns to do it.
The principles demonstrated in the Gather demo presumedly apply
to the child's learning. The child does whatever results in
cooperative attainment of her various and changing goals.
Cooperation itself, relationships upon which she may rely for
cooperation, and indeed the learning of the language itself, seem
to become goals, which the child abandons in favor of crying if
intrinsic variables as for hunger are not satisfied. So if
intrinsic variables are seen as "driving" the process, it is only
in an indirect, ultimate sense.

The word shapes and distinctions in currency in the child's
environment are givens for the child, just as surely as are the
regularities noticed by physics. The physical properties
of the environment (including the acoustic properties of the
vocal tract, such as those that make p t ch k easier to keep distinct
from one another than substituting e.g. for ch some other sound
intermediate between ch and k) are fixed and invariable, so far as
we can determine. They are not historically contingent, although
physical theories are. In contrast with this, the social properties of
the environment are variable. This is because they are variable
behavioral byproducts of other people's control of perceptions.
And in consequence, these given properties of the social environment
are historically contingent. They do not cause behavior. They do
"cause" the particular values to which the child sets her reference
perceptions, by control of which she participates in social
transactions, in the somewhat forced sense in which one might say
that physical laws "cause" her settings of reference perceptions
for throwing and catching a ball. The patterns constituting a
given language at one stage do not cause the different patterns
constituting that language at a later stage of its history. But
it is an observed fact that language change (barring competition
between alternative languages) gives the appearance of being
gradual and continuous, with the differences between one time-slice
of observation and the next being explicable in terms of similarity
of changed settings for an articulatory reference perception or
acoustic similarity of articulatory reference perception A at setting x
substituted for reference perception B at setting y, and the like.
This is largely because the child's choices for producing recognizable
renditions of adult speech are constrained by physical properties
of the environment.

In the very ingenious cross-purposed control of three lines, the
process of learning reference perceptions is not itself modelled.
Suppose other control systems in the environment were successfully
co-controlling three lines (with each other plus an invisible
disturbance). Suppose a control system lacking the needed settings
of reference perceptions "wanted" to do likewise with an offering
partner. (I think "wanting to" means controlling the perception
in imagination, but lacking the perceptual control to provide
actual inputs in place of the imagination loop.)

Modelling the learning process by simple reorganization gives us the
hard case, the one requiring a *very* patient and persistent partner.
The answers you provided (quoted at the beginning of this message) get
us only this far.

Now suppose we have some pre-established set of arbitrary reference
perceptions associated 1-1 with other perceptions such as line, left,
right, middle (or this, that), you, me, etc., such that if control
system A produces "line" then control system B imagines a line and
then (if possible) attends to an actual rather than imagined
perception of a line, and vice versa. In other words, the perceptions
used for communicating between A and B must be pre-set in A and in B
such that each may produce repetitions recognizable by the other and
such that each associates imagined perceptions with them for which
they can substitute actual perceptions that they can agree are
"the same thing" in their shared environment, etc. In other words,
give them a shared language. Can we now model the other cases of
communicating and acquiring the settings of reference perceptions
that are needed for the odd control system to turn that imagined
perception of "doing what those others are doing" into actual
perceptual control?

Can you now see why I believe this question of how we acquire
reference perceptions is important? Can you see that HPCT does
not give a coherent account of the skillful use of language
with reference to perceptions, the unskillful use of language
with reference to actions, the pointing and gestures and
jumping up and down, and the other means of accomplishing this end
that the human participants demonstrated?

        Bruce
        bn@bbn.com

From Tom Bourbon (930520.1410)

Bruce,

I have been working on a reply for some time now. When I logged on a few
minutes ago, I saw there were more posts on this topic. I will ignore
them for now, in order to finish this round of thoughts on the subject. If
I repeat anyting that is already out there, you will know the reason.

First, I have a few short comments on part of your reply to Rick, concerning
the GATHER program. Next, I discuss the question of what people are "really
doing" when we observe their behavior. Finally, I explore some implications
of the fact that, without modification, the PCT model predicts interactive
tracking data data where the interactions ocur in what are typically treated
as unrelated domains -- social; bi-manual, or inter-hemispheric; 2-D
movement around one joint, or intra-hemispheric.

[From: Bruce Nevin (Wed 930519 12:26:11)]

Rick Marken (930518.2200) --

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
If, by "reference perceptions" you mean "reference signal" settings then
HPCT already has an account of how they are arrived at; they are varied
by higher level systems as the means by which they control their
perceptions. If by "reference perceptions" you mean the state of the
perceptual variable that corresponds to the reference signal setting
currently in effect, then PCT already provides an account of that --
negative feedback control of perceptual input. If by "reference settings"
you mean the fixed settings of the intrinsic reference signals, then
these are presumably determined by the genes.
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

The forming of arcs and circles is universal. It can hardly be spoken
of as normative precisely because it follows from first principles that
apply to all humans regardless of their culture or history. ...

First principles? As in: organisms control many of their own
perceptions, by way of uncontrolled, necessarily variable, actions?
I believe you answered this affirmatively later in your post.

In the case of the gather demo, we might ask why it is that all the
individual control systems have the same setting for a reference
perception for proximity to other individual control systems. ...

The answer would be: because the person who selected the reference
signals and gain factors for that particular run made them equal.

... What would
happen if an "Arab" or "Latin American" control system were among
the "American" ones, with the proximity setting much closer? How would
the others respond if that one pressed closer to the speaker, in such a
way as appearing to stand between other individuals and the speaker? ...

Fortunately, Bill's programming of GATHER allows novices at
modeling to run simulations that answer those questions. Give one
of the characters a reference to be closer than the others, or a
higher gain for proximity than those of the others, and watch the
results. The results look sufficiently familiar that first- !time
viewers of the program characterize them as: "Like a (child-puppy)
trying to be closer to (adults-people) than the (adults-people)
want them to be." Or, "Like a 'pushy' person in a crowd." It
looks as though first principles, not acquired norms, might produce
these recognizable social events as readily as arcs and rings.

Were you asking why, on average (normative statistics), more people
from one culture are better modeled with reference signals and
gains of particular magnitudes, while those from another culture
require different magnitudes, but in both cultures people control
perceived proximity? Could it be that control of perceived
proximity is as much a given as control of perceived core body
temperature, but that unlike the control loop for body temperature,
the one(s) for proximity come equipped with a few more-readily-
adjustable parameters? (While thinking about your question, I
wonder if Clark McPhail and his colleagues have any field data on
the average radii of arcs and rings formed in different societies.)

Could notions of politeness and offense arise out of such interactions?
.....

Politeness and offense from whose viewpoint? Certainly not from
that of a character in GATHER, but certainly so from that of people
who observe the simulations. Many observers spontaneously describe
various characters in terms such as "pushy, aggressive, polite,
nice, frenetic, sluggish, patient." Of course, no matter the
impressions held by observers, the characters have no such
"traits." They do not even have reference signals for perceptions
like those, nor do they even have the perceptions. None of this is meant to
imply that people never think of their own actions as "polite," "boorish,"
and the like; only that an agent need not think in those terms for there to
occur social interactions we characterize in those terms. And all from first
principles.

Bruce, the remainder of your long post dealt more with culture-
specific, contingent features of language, as you contrast them
with arm movement and the actions of the GATHER characters. One of the
examples you used was of people communicating while one of them tries to
learn the cooperative tracking task. You acknowledged that whether tracking
is social or not, it entails a person acting through an environment that can
disturb the person's intended perceptions, but you seemed to hold to the
idea that the social environment differs qualitatively from the non-social
one.

For speech, just as for arm waving, once a person intends to experience a
particular result, the person is obligated to act through the environment,
CANCELLING the effects of independent influences, when they lead to
deviations from the intended experience; NOT TINKERING WITH their effects
when they produce a match with intended experiences. At this level, the
origins of effects from the environment do not matter. The person acts to
cancel the net disturbance on the intended experience -- however complex it
might appear on careful analysis, the environment reduces to a net
disturbance. From the perspective of the person, all that matters is the
state of the intended perceptions.

While tracking effectively, the person is NOT controlling arm movements,
or wrist positions, or grip force, or any of the infinity of variables that
we can identify from the outside. The reference signals for the features
of action that affect those variables -- vary, continuously, reflecting
error from the next higher level in the system. Error at the highest
level(s) depends on how well I am doing what I am really DOING -- I am
controlling my perceptions of position-of-cursor-relative-to-target. I am
not "doing" my actions. I am doing my perceptions of something entirely
different from my actions. My actions are the uncontrolled means by which
I achieve the intended perceptions.

At any moment, I might decide to control my perceptions of one or more of
those incidentally affected variables. Depending on which one(s) I select,
I might or might not be able to continue tracking. Whichever perceptions I
select, my actions will be controlled by the net disturbance from the
environment.

I think the situation is exactly the same when I talk. While I am talking,
I am not controlling speech sounds. Exactly as was the case for tracking,
the reference signals for my actions vary, depending on error from above,
which depends on what I am really doing. I am not doing phonemes, words,
syntax and the like. I may be "talking a couple through the cooperative
task," which means I now rely on the magic of words to control the same
perceptions of the screen that I earlier controlled through the magic of my
arm. Once again, I am doing my perceptions of something other than my
actions. My actions -- in this case my speech behavior -- are the
uncontrolled means to that end. In this case, it certainly matters whether
the other two people and I share a language, but if we don't, I will try
other ways. I might "jump, point and grunt them through the task," rather
than talk them through, in much the way that traders, tavellers and lovers,
across the world and the centuries, have tried other ways when language no
longer served as a means for them to achieve their intended experiences.
They were doing their intended experiences, not the actions of speech.

The magic of words spoken to other people works through links in the
environment that are less-tight than those when a person grabs a control
stick. Perhaps that difference in tightness has something to do with the
seemingly more constrained, but slowly-drifting, specificity of the sounds,
hence the actions, that work in a particular language community.

One more topic: whether, and when, control phenomena are sufficiently
different that we must invoke different explanations. The "cooperative
tracking task" can be performed equally well by -- here comes my list --
one person using two hands, two people each using one hand, four people each
pulling a string attached to a handle, two PCT models, or a PCT model and a
person, or two with strings, or .... The models reproduce and predict
the results, or participate in them, equally well, no matter the mix of
agents. Admittedly, it is "just tracking," but the simplest PCT model
works across what are usually thought of as independent theoretical domains:
social interaction, bi-manual (hence, interhemispheric) coordination,
coordination around a joint (hence, intrahemispheric). How much farther, in
either direction, can we push the model before it no longer works? I don't
know the answer, but I will enjoy trying to learn it.

In the cognitive and neurosciences, it has become popular to speak of a
"society of mind," or a "modular mind," or a "social mind," in the sense that
mind (you can substitute brain, or mind-brain) can be construed as a
"community" of interacting, but otherwise independent, "processes,"
"modules," or "functions." (This all seems to be a reworking of the older
"faculty psychology.") In those cases, the bridges between levels of
investigation and discourse are metaphorical. And of course, none of the
major players speaks of the phenomenon of control, or of PCT.

In contrast to those metaphorical links, the PCT model (even the simple
one-level model I am pushing in search of its limits; I will use no HPCT
model before its time) .... Let me start again. Already, control phenomena
from several domains, at different levels of observation and analysis, can
be duplicated by interactions between two of the simplest possible PCT
models. I believe (but cannot yet prove) that we can push the model farther
than that, in one direction as an explanation for control at the level of
physiology; in another, as an explanation for control of more complex
perceptions (than tracking) during social interactions, including those
during conversations. As Bruce says, we have yet to produce the working
models for all of those phenomena.

[From Bruce Nevin (Mon 930524 14:14:52)]

( Rick Marken (930520.1100) ) --

V ===========================================================
... a prevailing mis-
understanding about how hierarchical control works. It does
not work because the system has found a way to set the
"right" references; it works becuase the system has found a
way to vary references so that they end up producing the
"right" perception -- the perception defined as "right" by
other references.

even these [highest-level]
reference signals are variables; they can be varied by reorganiztion
or by higher level systems that might develop as the result of
reorganization. The only constant reference signals in the PCT model
are the intrinsic references -- which are almost certainly not neural
signals but, more likely, structural properties of cells or cell components
that specify what must be true (physiologically) if the system is to
survive.
^ ===========================================================

Let's take the reference signal for the perception of saying "no".
This is a single event perception. Also a syllable perception
in parallel, as may be seen if "no" is part of the event perception
of saying "nobody". (The same syllable perception in "notice" without
the event perception for the word "no".) Concurrently a category
perception for the phoneme /n/ and another for the phoneme /o/.
(We'll pretend the diphthong is really just one vowel, for simplicity.)
The phoneme-category perceptions are inputs to the syllable detector.
The phoneme-category perceptions together with the syllable perception
are inputs to the "no" word-detector. Other lower-level perceptions
of nasalization (lowered velum), apical articulation (tongue tip against
upper alveolar ridge), tongue backing and raising with lip rounding
(for the /o/ vowel) are inputs to the phoneme detectors. Acoustic
and visual cueues are also inputs to the phoneme-category detectors.
Now, pick any one of these and there's not a whole lot of variation
that is possible. Less nasalization and you get "doe". Less tongue
raising and you get "gnaw". And so on.

So the word and syllable detectors specify the reference signals for the
phoneme detectors. The reference signal for producing /n/ is probably
the same for "no", "on", "paining", and "painting". That the behavioral
outputs differ in these four kinds of cases is probably a matter of the
physics of moving the tongue around rather than variation of the reference
signal.

The error signal from the word-detector for "no" goes to lower levels as
reference signals for producing the phonemes /n/ and then /o/ in that
order in some cases ("I said `no!'") but just /n/ in others ("didn't I?!"),
and there are other forms such as non- in nonsmoking, a- in aseptic. So
OK, there is variation there (as for all of the reduction system of
language), but the variation is not in the reference input to a single
ECS but in the selection of which lower-level ECSs receive the signal.
This sounds to me like Bill's (Marc Olsen's) "place-coded perceptions"
(Bill Powers (930520.0830 MDT)).

I just don't see this freeforall of variation of reference signals going
on here. Give me some help connecting your principles to my examples.

[From Rick Marken (930520.1300)]

I was going to look in BCP this weekend for talk of the reference signal
being played back from memory. Unfortunately, I didn't get enough time
to track down the discussion of the reference signal. It may be that
I remember some discussion on the net too.

(Rick)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
I said:

What you
describe is exactly what happens in Tom's cooperation experiment

You say:

I don't understand how it is the same.

What I meant was that a conversation is similar to what is going on
in Tom's experiment. In order to control a higher order variable (say,
"being polite") both people in the conversation have to be controlling
for a similar variable (like both subjects in Tom's experiment had to control
for "three lines in a row").
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

You're jumping too far ahead again. In effect, you're changing the subject.
I'm not talking about cooperative control of conversational turns or "being
polite". What I said:

me (Wed 930519 12:26:11):

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
The use of "s" vs. "th" to distinguish words (as in English sin vs. thin)
or of [y] vs. [u] to distinguish words (as in French tu vs. tout) is
normative and language-specific (culture-specific, group-specific),
not universal. It may follow from first principles of HPCT that
apply to all humans (I assume that it does), but if so it does so
indirectly.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

So far as I can see, there is no parallel between Tom's experiment and
control of perceptions of contrast between words (or between parts of words).
You might get an analogy of some kind when B didn't understand and A
subsequently pronounces words more carefully, but that's not the normal
case that I wanted you to consider.

You were the only one to take a stab at the results of the speech experiments.
Bravo! And thanks. I'll give the results in a moment. I'm intrigued by
your comment on getting a "different vowel" when you mixed vowels with
noise "(harmonically spaced equal amplitude sine waves)". I'm interested
because if this works as I think you are saying, it could provide a way
of disturbing acoustic feedback of one's speech in real time, to determine
the relative importance of acoustic and other perceptions in the control
of word perceptions. What was the intended vowel and what was the
heard vowel? Was it the same change for both low-pass and high-pass
filtering of the noise, that is, the same perceived vowel?

Your guesses were pretty good, but there are still some surprises here,
which I think reveal some interesting things about the control of speech.

(Rick)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Experiment 1: In one ear (say, the right ear) play the acoustic signal
for syllables, only with gaps (silence) in place of the transitions.
In the other ear (the left) play the glissandi for various syllables,
appropriately timed to match the gaps.

In experiment 1, what do you hear in the right ear
and what do you hear in the left?

My guess is that the right ear heard, if not the "correct" syllables,
at least some syllables; the configuration systems that hear syllables
were probabaly able to piece the left ear stuff in to produce the
syllable perception. The left ear probably sounded like clicks --
sensations that were just happening in parallel to the syllable
configuations.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Let's call the acoustic signal with the gap the "base", and the other
signal the "chirps". You hear the original "correct" syllables as
though they were input to the ear that actually receives the base
signal. In the ear receiving the chirps, you hear only the chirps.

If you turn off the chirps, you hear syllables in the ear listening to
the base, but they are ambiguous as to what stop consonant is present
(e.g. ambiguous as to whether it is /b/ or /d/ or /g/).

(Rick)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Are there changes as amplitude is
decreased equally in both channels, and if so, what?

As amplitude is decreased equally in both channels? If so, my guess
is that new syllables were heard as the amplitude decreased;
'baa' became 'daa' perhaps. But there would still be the separate
clicks and syllables in each channel.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

I wasn't clear enough here. Decrease the amplitude in both channels
equally until the chirps are no longer audible. You still hear the
original "correct" syllables in the "base" ear, nothing in the "chirp"
ear. Detectors for speech sounds appear to work below the threshold
required for awareness of constituent elements of speech sounds.

(Rick)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Experiment 2: play the syllable /ba/ three times while displaying
a face pronouncing /bE vE dhE/ perfectly synchronized with the
acoustic signal. (/E/ is the vowel in "bed", /dh/ is the consonant
in "the".). [I corrected /th/ to intended /dh/ here & below--BN]
what do you hear in the headphones?

I think they would hear what they see -- they hear /bE vE dhE/. I think
this is because language perception has a visual component and
the perceptual functions that produce these syllables probably
have inputs from the visual system. If the sound input to these
functions is not radically different than what could go with the
visual input, then the visual input could be the main determiner
of the magnitude of the perceptual signal generated by each
syllable configuration detector.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Not bad. But what you hear is /ba va dha/. The visual cue is preferred for
consonant articulation, the acoustic cue is preferred for vowel quality.
/ba ba ba/ is heard as /ba va dha/ when seen as /bE vE dhE/. I wonder
if the visual cue for lip rounding with a vowel /o/ or /u/ would take
precedence, or if on the other hand it would visually obscure the
tongue position for /dh/.

In any case, this "multimodal" integration is very like Bill's notion
of how category perception might work. Lip readers learn to do without
acoustic cues. Input from recognizers for words and for word dependencies
seems to enable anyone to perceive intended pronunciations when control
of neighboring sounds/articulations is disturbed because of mutual
interference (coarticulation), or when gain is low (unstressed words,
predictable words). The reduction system of language spins out of this.

This experiment was summarized in Liberman and Mattingly (1983), along
with much else of interest. Based on experimental work at Haskins
Laboratories and elsewhere, they advance the view that speech perception
is the perception of the intended speech gestures of the speaker.
They draw analogies to e.g. auditory localization, an equally
"compelling" perception of location where the evidently "constituent"
perception of time delay between the two ears is not separable.
They are stuck in a command-control notion of neural processes,
and (partly in consequence of this) they assume that speech is
made possible by biologically innate neural mechanisms in a speech
module. It is not too hard to set these disturbances aside.
The reference:

Liberman, Alvin M., and Ignatius G. Mattingly. 1983.
        The motor theory of speech perception revised.
        _Cognition_ 21 (1983) 1-36.

There is more recent work out of Haskins that I have not yet been able
to connect with this.

(Tom Bourbon (930520.1410) ) --

Thanks for your response and very interesting and useful comments. I am
sure that I am underestimating what can be accomplished with PCT at just
one level. I admire and greatly respect the work that you are doing.
I hope it is not seen as wrong to try to apply PCT "prematurely" to
my field.

We are agreed that when members of a group of people share a setting
for a controlled perception, it is not because some superordinate
social-control system imposed a given setting upon them. Yet in
the system comprising the gather demo and its configurer/user, that
is precisely what is being modelled. The configurer/user imposes
a setting for the proximity perception on all the individuals in
the model, from the outside. How might this coming to agreement
be put into the model?

Again, you describe how individuals in the gather model can be given
different settings from the outside, and how the

(Tom)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
results look sufficiently familiar that first- !time
viewers of the program characterize them as: "Like a (child-puppy)
trying to be closer to (adults-people) than the (adults-people)
want them to be." Or, "Like a 'pushy' person in a crowd."
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Again, the fact that these perceptions are not in the individuals in
the model means that relevant perceptions controlled by humans
are not being modelled. Rightly or wrongly, people do have these
perceptions and do attempt to control them. Perhaps they are wrong
and doomed to failure when they attempt to control what after all are
judgements and other stretched analogies, and they would do better to
control perceptions at a lower level (as the Buddhists have been telling
us). Nonetheless, that is what people do, and we aim to model real people,
I think. Is there a perception of being different from others that
people control? "Be different, but don't let anyone know it" is what
a gay friend told me 30 years ago in Florida--for good reason, in his
experience! Perhaps it is a perception of being perceived as being
not a member, not one with whom cooperative action is reliably expected?
There is something primitive and fundamental about the in-group/out-group
thing, something close to intrinsic error, I think. Capture that,
and a great deal would emerge as byproducts of low-level perceptual
control.

("Like a (child-puppy) trying to be closer to (adults-people) than the
(adults-people) want them to be." Or, "Like a 'pushy' person in a crowd."
Polite, offensive, pushy aggressive, nice, frenetic, sluggish, patient, etc.
As you say, the observed agents do not have to have these perceptions for
human observers to perceive them as having them and even as controlling
them. The agents in gather are incapable of having them. If they were
capable of having such perceptions, would they not probably perceive
themselves as well as others in such terms?

Let's look at speech vs. arm waving. I have the intention of communicating
to you that you are doing something that is not what I wanted you to do.
I have a range of means for doing this, including arm-waving, words such
as "no", and things in between like the vocal gesture "uh-uh". If from
this range I control a perception of saying "no" I don't have a lot of
freedom how to say it. You have to be able to recognize, from whatever
I utter, that I intend to say the word "no". Saying "oxi" won't cut it
for you (unless you know Greek), and anyway that would be a different
word-perception. What's involved in your recognizing that I intend
to say the word "no"? What are the settings that we must have in
common. (Remember, I have the intention of saying the word "no". The
intention to communicate denial is at a higher level, expressible by
various words and gestures.)

(Tom)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
I think the situation is exactly the same when I talk. While I am talking,
I am not controlling speech sounds. Exactly as was the case for tracking,
the reference signals for my actions vary, depending on error from above,
which depends on what I am really doing. I am not doing phonemes, words,
syntax and the like.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

I think you are confusing awareness with intention, i.e. control of
perceptions. My attention may be on talking a couple through the
cooperative task. Beneath that focus of attention I at a particular
point have the intention of uttering the word "no" in a way that the
people I am talking to recognize my intention of uttering that word,
and behind that (as well as behind concurrent gestures, etc.) to
recognize my intention of denying something indicated by other words
and/or gestures. Beneath the intention to utter "no" I have the
intention to control the speech sounds symbolized here /n/ and /o/.
I am aware of the communicative task. Beneath that awareness and
in service of it I am indeed doing phonemes, words, syntax, and all
the rest.

(Tom)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
I may be "talking a couple through the cooperative
task," which means I now rely on the magic of words to control the same
perceptions of the screen that I earlier controlled through the magic of my
arm.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Different magic. No person has to recognize the intention behind a given
behavioral output with my arm. I don't have to control those behavioral
outputs for recognizability. I could control the lines on the screen
with visual perception of my arm movements cut off, and I could perform
the teaching task (perhaps more effectively) if the couple learning
the cooperative task could not see my arm. (They could be introduced
to the effects of the mouse or cursor in a separate step of the presentation.)

(Tom)

V++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
The magic of words spoken to other people works through links in the
environment that are less tight than those when a person grabs a control
stick. Perhaps that difference in tightness has something to do with the
seemingly more constrained, but slowly-drifting, specificity of the sounds,
hence the actions, that work in a particular language community.
^++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Arm movements work to control the lines in any social environment.
A given sequence of utterances (in the manner of Chuck's note cards, say)
work in one social environment but not in another. Why is that?

I am glad to see you grappling with these questions. The questions have
to be asked and seriously considered. I do not believe that PCT
provides quick and easy answers--yet.

        Bruce
        bn@bbn.com

From Tom Bourbon (930528.0810)

This is a long reply to a post from Bruce Nevin, concerning language and
social interactions. I composed it during two late-night sessions, for
posting this morning. I did not have access to the mail yesterday; if there
were additional posts on this theme, I have not seen them.

[From Bruce Nevin (Mon 930524 14:14:52)]

( Rick Marken (930520.1100) ) --

V ========================== .... a
prevailing mis-
understanding about how hierarchical control works. It does not work
because the system has found a way to set the "right" references; it
works becuase the system has found a way to vary references so that
they end up producing the "right" perception -- the perception defined as
"right" by other references.

even these [highest-level]
reference signals are variables; they can be varied by reorganiztion or by
higher level systems that might develop as the result of reorganization.
The only constant reference signals in the PCT model are the intrinsic
references -- which are almost certainly not neural signals but, more
likely, structural properties of cells or cell components that specify what
must be true (physiologically) if the system is to survive.

···

=================================

BN:
Let's take the reference signal for the perception of saying "no". This is
a single event perception. Also a syllable perception in parallel, as may
be seen if "no" is part of the event perception of saying "nobody". (The
same syllable perception in "notice" without the event perception for the
word "no".) Concurrently a category perception for the phoneme /n/ and
another for the phoneme /o/. (We'll pretend the diphthong is really just
one vowel, for simplicity.) The phoneme-category perceptions are inputs
to the syllable detector. The phoneme-category perceptions together with
the syllable perception are inputs to the "no" word-detector. Other lower-
level perceptions of nasalization (lowered velum), apical articulation
(tongue tip against upper alveolar ridge), tongue backing and raising with
lip rounding (for the /o/ vowel) are inputs to the phoneme detectors.
Acoustic and visual cueues are also inputs to the phoneme-category
detectors. Now, pick any one of these and there's not a whole lot of
variation that is possible. Less nasalization and you get "doe". Less
tongue raising and you get "gnaw". And so on.

++++++++++++++++++++++++++++++++

TB:
Examples like these occur reliably in discussions of differences between
speaking and tracking. (No putdown intended, only an observation.)
Each time they are raised, I try to understand the implication (sometimes
the claim) of great significance in the fact that changes in configurations
of the speech apparatus lead to changes in the sounds that are produced
and perceived. Every time, I come back to the thought that those
configurations vary any way necessary to produce perceptions (requested
by a higher level) of consequences other than the configurations and
sounds themselves. I still cannot see why the facts of configuration-
sound associations are different from, for example, the facts of various
configuration-movement associations we would observe were we to as
carefully monitor the activities of various motor units and joints in hands
and arms during arm waving, stick wiggling, gesturing and typing.

The constraint of "any way necessary" applies in either case; in neither
case does it imply "more than necessary," that is to say, "more than is
effective," with effective defined as production of the perception
requested in a higher order reference signal. Arm waving for its own
sake, as a perceived consequence of muscle actions, can be achieved by
many configurations of muscles and joints; waving an arm in the patterns
recognized by naval pilots as the signal to abort a landing on an aircraft
carrier is another matter. And it is another matter again for the stylized
hand movements of a Balinese dancer, or a Hindu practicing a mudra, or
a gang member in Los Angeles flashing a sequence of hand-signs. In
every case, a little more tension here or there, just a bit more pronation,
or extension, or rotation, and a whole other configuration occurs.
Whether the resulting configuration is significant to the actor or to any
particular observer, or class of observers, is yet another matter.

A personal example came to mind as I paused after typing those lines and
drummed my fingers on the computer table to accompany the Bach
"Brandenburg Concerti" playing on the radio. Those drummings,
contrasted with my now-resumed typing, contrasted with the movements
were I to type, not on my large keyboard, but on the tiny keys of the
Psion palmtop sitting on the shelf. Drumming can occur in many ways;
my movements when I use the big keyboard to type this message are
more constrained -- I can;t get too exc ited ort it fallls apart -- and my
readers will not understand. And if readers are to recognize what I type
on the Psion, I am even more severely constrained -- a tiny bit too much
extension of the right forefinger, not quite the porper arching, and I make
"u" not "j", and so it goes.

The same muscles and joints produce each of those results. In each
case, muscle activity varies, to produce a different intended result -- the
thing the actor is doing, not the actions seen by an observer.

+++++++++++++++++++++++++++++++++

BN:
So the word and syllable detectors specify the reference signals for the
phoneme detectors. The reference signal for producing /n/ is probably the
same for "no", "on", "paining", and "painting". That the behavioral
outputs differ in these four kinds of cases is probably a matter of the
physics of moving the tongue around rather than variation of the
reference signal.

++++++++++++++++++++++

TB:
Bill Powers already commented on this section.

[From Bill Powers (930525.1930)]

My point is that the reference signal that calls for producing the
perception of "no" is not itself the word "no." It is just a signal, which
specifies saying no not because of anything special about the reference
signal, but because it enters a certain comparator. Only its destination
makes it a "no" reference signal. And all that makes the corresponding
perceptual signal represent the word "no" is the way it is derived from
perceptual signals indicating that certain phonemes have been perceived.
The only way to get the "no" signal to appear is to supply signals from
phoneme detectors within a certain range of temporal and amplitude
limits. It's this concept that made Mark Olsen's remarks on "place"
coding ring a bell with me -- assuming this is anything like what he
meant. As you noticed.

The big objection to this whole picture is that this doesn't seem at all like
the way we experience "no" or phonemes or anything else. Some quality
seems to be missing. But what is missing, I think, is the context of all the
other perceptual signals that are in awareness at the same time.

..

Higher systems DO control their own perceptions by VARYING lower-level
perceptions (through manipulation of reference signals) -- but don't forget
that the higher level systems are perceiving different aspects of the
world. A system that tells a lower system to produce "no" does not
perceive the word "no". It does not know it is telling the lower system to
produce a word. Only the lower system can do that or know that; there
are no words at the higher levels. The higher level system perceives
something that is derived from the presence of the "no" signal and other
signals, and in this derived perception there is no longer any trace of
"no". To awareness, which can span several levels at a given time,
everything from the intensities and sensations on up coexists in the
experienced world, but a single level in the hierarchy always perceives
only its own specialized aspect of the situation.

+++++++++++++++++++++++++++++++

TB:
I second Bill's remarks, but he did not directly address the last part of
your paragraph, on the physics of moving the tongue. I am not sure I will
address it either, not in the way you might have intended, but it prompted
me to spend some time working my way through some of my
disorganized thoughts on that topic. Tongue wagging (talking) and eye
rolling (saccades, versions, etc) are often presented as special cases that
defy, or pose serious challenges to, the PCT model. (I am not saying that
*you* intended either of those, Bruce.) It has always seemed to me that
tongues and eyeballs DO enjoy rather special status: each is equipped
with an abundance of motor units and operates pretty much within the
protective confines of a body cavity or chamber (protruding eyeballs and
tongues are not all that common), relatively free from direct disturbances
(gravity excepted -- and generally overlooked by theorists). Those
circumstances, more than any other, probably contribute to easily-formed
impressions that feedforward, not feedback, controls the movements of
eyeballs, and that fixed commands (or reference signals) control the
movements of the speech apparatus. Those two systems operate in what
is probably as nearly an undisturbed environment as a human will ever
encounter -- precisely the kinds of conditions that allow a command-
driven model, like the one in "Models and Their Worlds," to seem
legitimate; and those are the conditions that lead Rick's "Blind Men" to
conclude that a control system is a command-driven (translate that as
'fixed-reference-signal') system. A few well-placed disturbances quickly
reveal that something else is going on. (Over a year ago, Gary Cziko
posted an ingenious demonstration of controlling speech for meaning, not
production, in spite of disturbances. What was that example?) People vary
their articulatory actions any wey necessary, the first time, when the
sensors for that system are blocked by local anesthesia administered by a
dentist, and they often produce intelligible speech (i.e., they create
a sufficient number of recognizible perceptions) even if they are chewing,
smoking, or (if they are PCT people who experiment on themselves) if they
hold their tongue against the roof or floor or side of the mouth, or against
the teeth, or between their teeth, when they hang upside down. And when the
disturbance is too much to work around, they no longer produce intelligible
speech.

Under typical conditions, the spectrograms of repeated utterances of a
particular speech sound are closely associated with a repeated productions
of similar articulatory movements, but it does not follow necessarily that
the movements result from a particular pattern of commands (as is sugested
in plan- or command-driven theories) or from invariant reference signals
(as in some interpretations of control theory).
+++++++++++++++++++

Bruce quotes Rick again:

(Rick)

++++++++++++++++++++++++++++++++
I (Rick) said:

What you (Gary)
describe is exactly what happens in Tom's cooperation experiment

You (Gary) say:

I don't understand how it is the same.

What I (Rick) meant was that a conversation is similar to what is going
on in Tom's experiment. In order to control a higher order variable (say,
"being polite") both people in the conversation have to be controlling for
a similar variable (like both subjects in Tom's experiment had to control
for "three lines in a row").
++++++++++++++++++++++++++++++++

BN:
You're jumping too far ahead again. In effect, you're changing the
subject. I'm not talking about cooperative control of conversational turns
or "being polite". What I (Bruce) said:

me (Wed 930519 12:26:11):

++++++++++++++++++++++++++++++++
The use of "s" vs. "th" to distinguish words (as in English sin vs. thin) or
of [y] vs. [u] to distinguish words (as in French tu vs. tout) is normative
and language-specific (culture-specific, group-specific), not universal. It
may follow from first principles of HPCT that apply to all humans (I
assume that it does), but if so it does so indirectly.
++++++++++++++++++++++++++++++++

So far as I (Bruce) can see, there is no parallel between Tom's experiment
and control of perceptions of contrast between words (or between parts
of words). You might get an analogy of some kind when B didn't
understand and A subsequently pronounces words more carefully, but
that's not the normal case that I wanted you to consider.

+++++++++++++++++++++++++++++++++
+
Rick replied:
[From Rick Marken (930524.1500)]

Whatever you say. I thought you were describing a situation where lower
level outputs had to be produced in a cooperative or coordinated manner;
Tom's experiment demonstrates one way of implementing cooperation
between control systems -- when these systems are housed in different
bodies. If you are not dealing with an example of cooperative control in
the control of contrasts between words then, fine, maybe Tom's
experiment is not a good parallel. But I'm pretty sure that it is.

++++++++++++++++++++++

TB (now):
(Whether the systems are housed in different bodies, or in one, makes no
difference.) I have the impression Bruce and Rick are talking about the
cooperation experiment at two very different levels. Not surprisingly,
Bruce seems to be keying in on my observations about the many ways
people varied their actions to communicate, while Rick seems to be
focusing on the modeling of interaction between two control systems.
Both levels are there to see. It is true, as Bruce says, that I did not (I
CANNOT) model the conversations, gestures and other means of
communicating. As things stand, the experiment only provides a setting
rich in actions by which participants intend to communicate.

I think Rick was pursuing the idea that cooperative interactions between
the models might offer a parallel to some aspects of speech production.
If he was, I agree. The two models produce the perceptual signals that
satisfy the (in this case, implicit) higher level reference of "three lines in
a specific configuration." I described "three lines in a row," but any
combination of two controllers could produce any of many possible
configurations. A change in the reference signal in one model, or in both,
leads to a new configuration. In a two-level system, a change in the
higher-level reference for a particular pattern would, if not matched by a
perceptual signal of that pattern, create error that acts as reference
signals for the two systems I modeled, and each would independently
produce one of the two two-line relationships that together result in
"three lines in the new configuration." (Hard to say; easy to do!) In that
case, the cooperative modeling illustrates how the same two independent
lower-level control loops can cooperatively produce any one of many
perceptual signals, corresponding to many different states of the
environment and to many differeny refeence signals from a higher level.
In none of those cases would the lower systems "know what they were
doing," in the sense of creating the three-line configuration. Each would
perceive its own little piece of a bigger pattern about which it was as
blissfully ignorant as Fowler and Bandura are about PCT. Were the model
fleshed out a bit, with a level of control loops below the ones I modeled,
the new loops could produce movements, with no knowledge of lines, or
of relationships between two lines, much less between three.

I imagine Rick was also thinking of how neither system controls its own
actions, which are controlled by the disturbance acting on the middle line
and -- in a way that still fascinates me -- by the actions of the other
system. Given a system with a reference signal for a particular
perception, anything that disturbs that perception "controls" the actions
of the system. I think those features of interactions between systems
might have implications for more elaborate HPCT models of speech
production.
++++++++++++++++++++++

Bruce commented on a post from me (TB):
(Tom Bourbon (930520.1410) ) --

BN said:
We are agreed that when members of a group of people share a setting
for a controlled perception, it is not because some superordinate social-
control system imposed a given setting upon them. Yet in the system
comprising the gather demo and its configurer/user, that is precisely what
is being modelled. The configurer/user imposes a setting for the proximity
perception on all the individuals in the model, from the outside. How
might this coming to agreement be put into the model?
+++++++++++++++++++

TB (now):
In the GATHER demo, the configurer/user "inserts" the reference signals,
and the gain factors as well. But when they do, I hope they do not
assume they can do the same things to living control systems. That is
the advantage of being a modeler/configurer; you have a chance to play
the part of a god. (I assume that only a god-like entity could so easily
tweak the reference signals of living systems.)

How might a coming to agreement be put into the model? If you mean
the appearances of coming to agreement, the characters in GATHER
sometimes seem to do that already; if you mean two or more systems
each acting to produce intended perceptions of coming to agreement,
they would need hierarchies deep enough to accommodate the setting of
reference signals for that level of perception. That is long way from the
ingenious programming of the multiple loops, all in parallel at the same
level, that Bill placed into the GATHERERS.
++++++++++++++++++

BN:
Again, you (Tom) describe how individuals in the gather model can be
given different settings from the outside, and how the

(Tom)

++++++++++++++++++++++++++++++++
results look sufficiently familiar that first- !time viewers of the program
characterize them as: "Like a (child-puppy) trying to be closer to (adults-
people) than the (adults-people) want them to be." Or, "Like a 'pushy'
person in a crowd."
++++++++++++++++++++++++++++++++

Again, the fact that these perceptions are not in the individuals in the
model means that relevant perceptions controlled by humans are not
being modelled.
++++++++++++++++

TB (now):

Of course that is the case. And it will always be the case, in modeling at
any conceivable level of complexity and detail. (That's why it is called
"modeling.")

In fact, a GATHERER controls only four or five modeled perceptions that
might be relevant to humans (or to birds, fish, dogs, apes, ants, robots
or any other social entities I can imagine). Amazing, isn't it, that an
observer could attribute so much knowledge and personality to a few
simple entities controlling so few relevant perceptions?
++++++++++++++++++++++++

BN:
Rightly or wrongly, people do have these perceptions and do attempt to
control them. Perhaps they are wrong and doomed to failure when they
attempt to control what after all are judgements and other stretched
analogies, and they would do better to control perceptions at a lower
level (as the Buddhists have been telling us). Nonetheless, that is what
people do, and we aim to model real people, I think.
++++++++++++++++++++

TB (now):

Of course people can have those perceptions. And, yes, sometimes they
are wrong and doomed to failure in their attempts to control them. Can
you suggest how we might model control of those perceptions? I don't
know how to do it. But I believe the primitive PCT modeling we *can*
do will be the substrate on which the really interesting models, like the
ones you request, will grow. Contemporary PCT models have scarcely
begun to swim in the primordial ooze.
++++++++++++++++++++

BN quotes and comments:
("Like a (child-puppy) trying to be closer to (adults-people) than the
(adults-people) want them to be." Or, "Like a 'pushy' person in a
crowd." Polite, offensive, pushy aggressive, nice, frenetic, sluggish,
patient, etc. As you (Tom) say, the observed agents do not have to have
these perceptions for human observers to perceive them as having them
and even as controlling them. The agents in gather are incapable of
having them. If they were capable of having such perceptions, would
they not probably perceive themselves as well as others in such terms?
+++++++++++++++++++

TB(now):
I think you are agreeing with what I said. Observers often see the
characters as "having" certain traits, qualities, or motives. I used that as
an example of how an observer can be wrong about what an actor-agent
is doing. But I also said that, given an HPCT control system of sufficient
complexity, those perceptions might occur within the system and could
become perceptions requested in reference signals.
++++++++++++++++++++++++++++

BN:
Let's look at speech vs. arm waving. I have the intention of
communicating to you that you are doing something that is not what I
wanted you to do. I have a range of means for doing this, including arm-
waving, words such as "no", and things in between like the vocal gesture
"uh-uh". If from this range I control a perception of saying "no" I don't
have a lot of freedom how to say it. You have to be able to recognize,
from whatever I utter, that I intend to say the word "no". Saying "oxi"
won't cut it for you (unless you know Greek), and anyway that would be
a different word-perception. What's involved in your recognizing that I
intend to say the word "no"? What are the settings that we must have
in common. (Remember, I have the intention of saying the word "no".
The intention to communicate denial is at a higher level, expressible by
various words and gestures.)
+++++++++++++++++++++

TB (now):
Agreed, Bruce. If you want to hear yourself say "no" you do not have
much freedom how you say it. In fact, you have NO freedom -- you must
do whatever is necessary to hear "no." So, too, in tracking, where it may
look to an observer as though I am free to move my arm and hand any
way I wish, but as soon as I decide to track, I am no longer free: I must
act to cancel the net disturbance acting on my controlled variable. And
once I grasp the handle in a particular way, or plant my elbow on the
table, or press my arm against my side, or hold a coffee cup in the same
hand as the handle ... . In every one of those cases, I give up even more
"freedom;" I put myself more certainly and rigidly (and vulnerably) under
the control of disturbances to my intended perception. I give up degrees
of freedom in my actions and am forced to use a smaller number of free
control loops, acting in more seemingly stereotypic ways. I have chosen
my perceptions, and I don't have a lot of freedom how to control them.

I do not mean to be flip or glib when I say this: The fact that a person's
choice of words is constrained does not seem to me to be different from
the constraints on every other kind of control behavior. And the fact that
some of the constraint seems to depend on the cultural or community
affiliations of the speaker confirms, for me, the idea that people (in the
universal sense) vary their actions (utterances) any way necessary to
achieve the control they are really doing -- which is hardly ever the
utterances we call speech. Most of the time. speech is like hand wiggling
and jabbing and turning; both are unintended means to intended ends.
(Babies, when they babble, seem to be among the very few who utter
speech sounds for their own sake.)
++++++++++++++++++++++

Bruce quotes:

(Tom)

++++++++++++++++++++++++++++++++ I
think the situation is exactly the same when I talk. While I am talking, I
am not controlling speech sounds. Exactly as was the case for tracking,
the reference signals for my actions vary, depending on error from above,
which depends on what I am really doing. I am not doing phonemes,
words, syntax and the like.
++++++++++++++++++++++++++++++++

I think you are confusing awareness with intention, i.e. control of
perceptions. My attention may be on talking a couple through the
cooperative task. Beneath that focus of attention I at a particular point
have the intention of uttering the word "no" in a way that the people I am
talking to recognize my intention of uttering that word, and behind that
(as well as behind concurrent gestures, etc.) to recognize my intention of
denying something indicated by other words and/or gestures. Beneath
the intention to utter "no" I have the intention to control the speech
sounds symbolized here /n/ and /o/. I am aware of the communicative
task. Beneath that awareness and in service of it I am indeed doing
phonemes, words, syntax, and all the rest.
++++++++++++++++++++++++

TB (now):
I don't think I was confusing awareness and intention. When I track, I
control (specify and hold invariant) my perceptions of the relationship I
intend to see. To control those perceptions of relationship, I cannot also
control (specify and hold invariant) all of the perceptions in control loops
below the level where the reference signal for the relationship is set.
Sometimes I can become aware of perceptions at those lower levels, but
they "just happen" while I control the perceived relationship. The
reference signals for the perceptual signals in lower levels cascade down
as errors from higher levels, but I do not intend them, nor am I aware of
most of them. I" do not intend all of the perceptual signals that occur
every moment, nor do (does?) "I" control the actions. I-the-intender can
remain serenely aloof while I-the-control-hierarchy control perceptions at
levels rarely know by the intender. Each of us inside here is ignorant of
the others and of what they are doing.
+++++++++++++++++++++++

BN quotes:

(Tom)

++++++++++++++++++++++++++++++++
The magic of words spoken to other people works through links in the
environment that are less tight than those when a person grabs a control
stick. Perhaps that difference in tightness has something to do with the
seemingly more constrained, but slowly-drifting, specificity of the sounds,
hence the actions, that work in a particular language community.
++++++++++++++++++++++++++++++++

BN:
Arm movements work to control the lines in any social environment. A
given sequence of utterances (in the manner of Chuck's note cards, say)
work in one social environment but not in another. Why is that?
+++++++++++++++++++++

TB (now):

Different intended perceptions.
+++++++++++++++++++++++++++

Until later,
Tom Bourbon